# Decision Models Index

Canonical page: https://amplifying.ai/decision-models

Sources checked October 2, 2026. This is a selected list, so some projects will be missing. The order is not a ranking.

Give a decision model some content and a few questions. It returns choices, scores, or probabilities your code can use. Each entry identifies whether it is a model, an inference method, a platform, or an integration. Test how well the probabilities hold up on your own tasks.

## Start here

These editorial picks highlight API support, deployment choices, and published code or training resources. Use the benchmarks below to compare measured performance.

### TypeSafe Jev

The reference API · Hosted API

TypeSafe’s Choice, Score, and Noul format is the reference for many entries here. Start with its docs to understand how these models work.

[Details and sources](https://amplifying.ai/decision-models#jev)

### Cloudflare Clef

Cloud deployment and open weights · Hosted API · Open weights

Cloudflare offers Clef and the smaller Clef-flash through Workers AI and as downloadable weights. Both accept text, images, and video.

[Details and sources](https://amplifying.ai/decision-models#clef)

### Perplexity Decisions

A hosted API you can also self-host · Hosted API · Open weights

Perplexity publishes the 27B model behind its Decisions API, along with inference code. You can try the service before setting up your own GPU.

[Details and sources](https://amplifying.ai/decision-models#pplx-decider)

### Kev

Run and train your own model · Open weights

Jared Palmer’s model family comes with training recipes and frozen evaluation sets. It’s a useful starting point if you want to inspect or adapt a decision model.

[Details and sources](https://amplifying.ai/decision-models#kev)

### Laya

Smaller models for local use · Open weights

Laya’s 322M and 421M encoders offer a smaller local option. The family includes multilingual checkpoints and community runtimes for Apple Silicon and browsers.

[Details and sources](https://amplifying.ai/decision-models#laya)

### Databricks ai_decide

Decisions inside your data platform · Beta · Managed platform

For teams already using Databricks, ai_decide brings typed questions into SQL and REST workflows over governed data. Databricks manages the underlying model.

[Details and sources](https://amplifying.ai/decision-models#databricks-ai-decide)

## What is in the catalog

Featured entries appear first, followed by the remaining entries grouped by type and name.

Open weights includes complete checkpoints and adapters that need a separate base model. Some downloads require accepting access terms. Open code means the implementation is public; each entry lists the license and deployment requirements.

- Models: Models you can call through an API or download and run yourself.
- Inference methods: Ways to get decisions from an existing language model. The base model affects the results.
- Platforms: Services and local tools for running decision models. Check which model each one uses.
- Integrations: Examples built with a decision model, such as Neon’s redaction app using Jev.

| Entry | Publisher | Kind | Access | Size |
| --- | --- | --- | --- | --- |
| Jev | TypeSafe AI | Model | Hosted API | Not disclosed |
| Clef | Cloudflare | Model | Hosted API, Open weights | 27B |
| pplx-decider-v1-27b | Perplexity | Model | Hosted API, Open weights | 27B |
| Kev | Jared Palmer | Model | Open weights | 0.8B / 4B / 9B / 27B |
| Laya | Convai Innovations | Model | Open weights | 322M / 421M |
| Databricks ai_decide | Databricks | Platform | Managed platform | Underlying model can change |
| AgentJev | malevrigns | Model | Open weights | 0.6B |
| AutoTrust JEV | AutoTrust AI | Model | Open weights | 9B / 27B |
| basal-1.0 | Remek / rkinas | Model | Open weights | 1.5B / 4.5B |
| bekko-system-one | hotchpotch | Model | Open weights | 17M / 68M / 400M |
| Bespoke Nimble | Bespoke Labs | Model | Open weights | 9B |
| Clef-flash | Cloudflare | Model | Hosted API, Open weights | 9B |
| CLM v0.1-8B | Contrastive-LM research team | Model | Open weights | 8B backbone + projection heads |
| Cua-S1 | Cua | Model | Open weights | Nano / 4B variants |
| d1 | Liquid AI | Model | Hosted API | Not disclosed |
| Decider | Mapika | Model | Open weights | 0.8B / 2B / 4B / 12B / 35B-A3B |
| Decision Machine 1 | milliseconds.ai | Model | Hosted API | Not disclosed |
| Drex 1.5 | Nace AI | Model | Hosted API | Under 10B (publisher figure) |
| GLiDE | Fastino Labs | Model | Hosted API | Not disclosed |
| GLiNER2.5-Decide | Fastino Labs | Model | Open weights | 340M (publisher figure) |
| Imajev | mohit67890 | Model | Open weights | 2B / 4B / 9B |
| Intern-Decision | InternLM | Model | Open weights | Checkpoint-dependent; 4B documented |
| Jeeves | PostHog | Model | Open weights | 9B |
| Julia-1 | Supersonic Labs | Model | Open weights | 144.3M |
| lev | Interfaze AI | Model | Open weights | 4B backbone |
| Mercury Decide | Inception | Model | Hosted API | Not disclosed |
| NanoJev | Tianyu Chen | Model | Open weights | 0.6B |
| OneJev | OmniJev | Model | Open weights | 0.8B / 4B / 9B / 27B |
| OpenJev | OpenJev project | Model | Open weights | 27B |
| OpenThai-SystemOne | iApp Technology | Model | Open weights | 0.8B |
| Solar Decide | Upstage | Model | Hosted API | Solar Mini 4 backbone |
| Span-01 / Span-01 Lite | Respan | Model | Hosted API | Not disclosed |
| Standard One | Standard Thinking | Model | Open weights | 3B / 8B |
| Strands Decider | Strands Labs | Model | Open weights | 1.9B |
| Surogate Rune | Invergent | Model | Hosted API, Open weights | 26.5B total / ~4B active |
| Tev1 | Together AI | Model | Hosted API, Open weights | 4B / 0.8B experimental |
| U2-Decision | Unisound | Model | Open weights | 4B backbone |
| Von | wfzyx | Model | Open weights | 395M |
| XOR | Xyne / Juspay | Model | Open weights | 35B total / ~3B active |
| AnyJev | Nokia Applied Research | Inference method | Open code | Depends on base model |
| djev | Davipar | Inference method | Open code | DiffusionGemma 26B-A4B backbone |
| LLM2Jev | Yinsongxu and contributors | Inference method | Open code | Depends on base model |
| Reflex | kshetrajna12 | Inference method | Open code | Base-model dependent; 4B default |
| SemIf (formerly OpenJev) | TheoLeeCJ and contributors | Inference method | Open code | Depends on base model |
| Decisions API | OpenAI | Platform | Hosted API | Powered by GPT-6 Luna |
| Instinct | ZooWork | Platform | Hosted API | 4B / 27B configurations |
| llama.cpp decision models | ggml-org | Platform | Open code | Depends on selected model |
| Ollama decision models | Ollama | Platform | Open code | Depends on selected model |
| Ollaya | ollaya-dev | Platform | Open code | Depends on selected model |
| Jev Ultrafast | Browser Use | Integration | Open code | Uses TypeSafe Jev + a text model |
| Neon pg_redact example | Neon / Rishi Raj Jain | Integration | Open code, Hosted API | Uses TypeSafe Jev |

## Jev

Model · TypeSafe AI

Ask several questions about the same input and get back choices, rubric scores, or yes/no probabilities. The stable jev-latest alias currently points to jev-1.13.0.

Inputs: Text / JSON. Size: Not disclosed.

Access: Hosted API.

License: Proprietary service.

Deployment: TypeSafe API and supported gateways; official Python and TypeScript SDKs.

Limits: Use a specific version for comparisons because aliases can change. TypeSafe does not document downloadable weights.

- [Documentation](https://docs.typesafe.ai/introduction)

- [Versions](https://docs.typesafe.ai/models)

## Clef

Model · Cloudflare

Cloudflare’s larger decision model, built on Qwen3.8-27B. A joint decision head scores the allowed answers to your questions. It accepts the Jev/SystemOne request format.

Inputs: Text / JSON / images / video. Size: 27B.

Access: Hosted API, Open weights.

License: Apache-2.0.

Deployment: Workers AI or self-hosted weights. Cloudflare announced the family on October 1, 2026.

Limits: Cloudflare ran the published benchmarks. To compare them with other results, check that the tasks and serving setup match.

- [Model card and weights](https://huggingface.co/Cloudflare/clef)

- [Launch](https://blog.cloudflare.com/clef-decision-models/)

## pplx-decider-v1-27b

Model · Perplexity

Powers Perplexity’s Decisions API. It returns probabilities for each allowed answer or rubric level, and a probability for yes/no questions.

Inputs: Text / JSON / images. Size: 27B.

Access: Hosted API, Open weights.

License: Apache-2.0 weights; hosted service terms.

Deployment: POST /v1/decisions at $0.04 per million input tokens, with output tokens free, or self-host the published weights and inference code.

Limits: The open release includes an AutoJev inference implementation. Compare hosted and self-hosted runs using the same checkpoint and settings.

- [Model, API, and pricing](https://docs.perplexity.ai/docs/decisions/quickstart)

- [Weights and inference code](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b)

## Kev

Model · Jared Palmer

A family built on Qwen, with downloadable weights, training recipes, and frozen evaluation sets. You can mix Choice, Score, and Noul questions in one request.

Inputs: Text / JSON. Size: 0.8B / 4B / 9B / 27B.

Access: Open weights.

License: Apache-2.0.

Deployment: Local CUDA or Apple Silicon via MLX; a SystemOne-compatible server. Kev-4B also runs natively in llama.cpp through ggml-org/Kev-4B-GGUF.

Limits: Each model card lists its memory needs and tested context length. These vary by size.

- [Repository and model cards](https://github.com/jaredpalmer/kev)

- [llama.cpp weights](https://huggingface.co/ggml-org/Kev-4B-GGUF)

## Laya

Model · Convai Innovations

An encoder family with English, multilingual, and typed-workflow versions. Its router chooses a checkpoint for each request.

Inputs: Text / JSON. Size: 322M / 421M.

Access: Open weights.

License: Apache-2.0.

Deployment: Local Python runtime or native llama.cpp serving with ggml-org/Laya-GGUF; separate community ports support MLX, ONNX, and browser use.

Limits: Check the language coverage and tested context length for the version you choose. Runtime ports use the same model family.

- [Repository and variants](https://github.com/NandhaKishorM/laya)

- [Model card](https://huggingface.co/convaiinnovations/laya)

- [Community MLX runtime](https://github.com/mizorewww/laya-mlx)

- [llama.cpp weights](https://huggingface.co/ggml-org/Laya-GGUF)

## Databricks ai_decide

Platform · Databricks

Ask Noul, Choice, and Score questions about your governed data from SQL or REST. The function is in beta and supports several questions per call.

Inputs: Text / structured data. Size: Underlying model can change.

Access: Managed platform.

License: Service terms; model licenses apply.

Deployment: Supported Databricks regions and compute; Runtime 15.4 LTS or later, with 18.2 recommended. Not SQL Classic.

Limits: Databricks manages the function and may change the model behind it. It has not announced this as a separate model family.

- [Function and requirements](https://docs.databricks.com/aws/en/sql/language-manual/functions/ai_decide)

- [Announcement](https://www.databricks.com/blog/introducing-aidecide-make-fast-decisions-your-governed-data)

## AgentJev

Model · malevrigns

A Qwen3-0.6B model with a candidate-scoring head for questions about logs, diffs, and workflow state. It scores several questions without generating text.

Inputs: Text / agent state. Size: 0.6B.

Access: Open weights.

License: Apache-2.0.

Deployment: Self-host with the repository’s custom model loader and server. Its quickstart pins a specific checkpoint revision.

Limits: The reported Typed Decisions score measures agreement with teacher labels on 400 cases. It does not measure coding-agent task completion.

- [Repository and evaluation scope](https://github.com/malevrigns/agent-jev)

- [Weights](https://huggingface.co/aimeigaoshou/agent-jev)

## AutoTrust JEV

Model · AutoTrust AI

Combines a fast decision path with a reasoning path. The JEV-27B-VL version retains vision for questions about images and text.

Inputs: Text / images in VL version. Size: 9B / 27B.

Access: Open weights.

License: Apache-2.0.

Deployment: Downloadable weights and decision-serving code; the VL quickstart uses vLLM on a GPU with at least 80 GB of memory.

Limits: Separate from TypeSafe’s Jev despite the name. AutoTrust’s image recommendation report is its own experiment on one dataset, not a general visual benchmark result.

- [Model family](https://huggingface.co/autotrust)

- [VL weights and serving code](https://huggingface.co/autotrust/JEV-27B-VL)

- [VL release report](https://huggingface.co/blog/autotrust/autotrustjev-27b-vl-a-decision-model-that-learned)

## basal-1.0

Model · Remek / rkinas

Models trained for Polish and English typed decisions, with particular attention to Polish tasks. The server averages two option orders and applies per-question-type calibration.

Inputs: Polish / English text and JSON. Size: 1.5B / 4.5B.

Access: Open weights.

License: Apache-2.0.

Deployment: Local inference engine and SystemOne-compatible HTTP service; published BF16, FP8, and NVFP4 variants.

Limits: A Polish specialist. Its general English benchmark results are weaker, and the repository documents known notice-period label corrections. At most ten options per question.

- [Models, results, and limitations](https://github.com/rkinas/basal)

## bekko-system-one

Model · hotchpotch

Small models that score candidate answers from instructions and context. The release includes training data, training recipes, and evaluation code.

Inputs: Text. Size: 17M / 68M / 400M.

Access: Open weights.

License: Code MIT; check each checkpoint license.

Deployment: Python and ONNX releases, with examples for local and browser use.

Limits: Experimental. The author also runs S1MB, and some training datasets supply evaluation tasks. Use the separate generalization tests when judging unfamiliar tasks.

- [Models, training, and evaluations](https://github.com/hotchpotch/bekko-system-one)

## Bespoke Nimble

Model · Bespoke Labs

A Qwen3.5-9B LoRA fine-tune that scores allowed answers directly. Bespoke publishes its training recipe and explains how it chose contrasting training examples.

Inputs: Text. Size: 9B.

Access: Open weights.

License: See checkpoint license.

Deployment: Apple Silicon or NVIDIA GPU; also available through Ollama’s decision-model support.

Limits: Accepts text only, even though the base model has a vision component. The flat schema allows up to 255 options per field.

- [Repository and constraints](https://github.com/bespokelabsai/nimble)

- [Weights](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B)

## Clef-flash

Model · Cloudflare

The smaller Clef model, built on Qwen3.5-9B. It uses the same question format and joint decision head as the 27B version.

Inputs: Text / JSON / images / video. Size: 9B.

Access: Hosted API, Open weights.

License: Apache-2.0.

Deployment: Workers AI or self-hosted weights; Jev/SystemOne-compatible API.

Limits: The smaller model may suit a different deployment budget. Speed and accuracy still depend on the task.

- [Model card and weights](https://huggingface.co/Cloudflare/clef-flash)

- [Launch](https://blog.cloudflare.com/clef-decision-models/)

## CLM v0.1-8B

Model · Contrastive-LM research team

Encodes the state and candidate actions separately, then scores how well each pair matches. Repeated candidate actions can reuse their embeddings.

Inputs: Text / state-action pairs. Size: 8B backbone + projection heads.

Access: Open weights.

License: Apache-2.0.

Deployment: The contrastive-lm package serves trained projection heads with a frozen Qwen3-8B encoder, using vLLM for embeddings.

Limits: The heads require the encoder and pooling setup they were trained with. Published coding-agent verifier results use additional fine-tuning, not this checkpoint as released.

- [Checkpoint and limitations](https://huggingface.co/Contrastive-LM/CLM-v0.1-8B)

- [Code and training recipe](https://github.com/Contrastive-LM/CLM)

## Cua-S1

Model · Cua

Chooses among supplied computer-use actions. The 4B 0.2 release has separate text and multimodal adapters, trained with rewards from completing tasks in live GUI environments.

Inputs: Text / screenshots. Size: Nano / 4B variants.

Access: Open weights.

License: Adapters Apache-2.0; code MIT; base models have separate terms.

Deployment: Download the Cua-S1 adapters and base model; run them with the Cua repository. Nano and earlier 4B research checkpoints remain available.

Limits: These are computer-use research models. Offline action-choice accuracy and completion of a multi-step GUI task measure different things.

- [4B 0.2 adapters](https://huggingface.co/cua-ai/cua-s1-4b-0.2)

- [Nano checkpoint](https://huggingface.co/cua-ai/cua-s1-nano-0.1)

- [Source and recipes](https://github.com/trycua/cua/tree/main/libs/cua-s1)

## d1

Model · Liquid AI

Liquid’s model answers Choice, Score, and Noul questions in the same request, without generating response tokens.

Inputs: Text / JSON. Size: Not disclosed.

Access: Hosted API.

License: Hosted service terms.

Deployment: Liquid Decisions API; documentation demonstrates d1:free with TypeSafe SDKs.

Limits: The d1 docs do not list downloadable weights. Check the free route’s capacity limits before relying on it in production.

- [Decision-model documentation](https://docs.liquid.ai/lfm/models/decision-models)

## Decider

Model · Mapika

A family of trained decision models with public training recipes and serving code. Small GGUF releases run locally; a separate 2B vision variant accepts images.

Inputs: Text / images in vision variant. Size: 0.8B / 2B / 4B / 12B / 35B-A3B.

Access: Open weights.

License: Apache-2.0; check each base model’s terms.

Deployment: The decider-ai package supports PyTorch and GGUF runtimes, with CUDA and selected Apple Silicon configurations.

Limits: The project also publishes decider-chat configurations using unchanged base weights. Those are inference configurations, distinct from its trained checkpoints.

- [Checkpoints, runtimes, and evaluations](https://github.com/Mapika/decider)

## Decision Machine 1

Model · milliseconds.ai

Handles yes/no checks, classification, ratings, quoted answers, extraction, entities, and verification. Its API uses a different request format from Jev.

Inputs: Text. Size: Not disclosed.

Access: Hosted API.

License: Hosted service terms.

Deployment: Hosted API at api.milliseconds.ai/v1; choose the capability that matches the output you need.

Limits: Capability limits differ. For example, classification accepts 2–64 labels; extraction and verification have their own schemas.

- [Model launch](https://milliseconds.ai/blog/why-we-built-decision-machine-1/)

- [Capabilities](https://docs.milliseconds.ai/concepts/choosing-a-capability)

## Drex 1.5

Model · Nace AI

Returns choices, yes/no probabilities, and scores over a 128K context window. Nace also offers tuning on a company’s own decision data.

Inputs: Text / JSON. Size: Under 10B (publisher figure).

Access: Hosted API.

License: Proprietary service and deployment terms.

Deployment: Hosted API, with managed, private, and hybrid deployments offered by Nace.

Limits: Nace discloses training on the training splits of Decision Index datasets. Private deployment is offered commercially; it does not establish a public open-weights release.

- [Model and evaluation details](https://www.nace.ai/drex)

- [1.5 announcement](https://x.com/NaceAI/status/2105319241436328256)

## GLiDE

Model · Fastino Labs

Makes an initial decision, then spends more time reasoning when it is uncertain. Announced September 30, 2026.

Inputs: Text / JSON. Size: Not disclosed.

Access: Hosted API.

License: Hosted service terms.

Deployment: Fastino API, model fastino/glide, through /v1/systemone. The 40K context budget includes thinking tokens.

Limits: Fastino ran the published Decision Index comparison itself. Its launch score was not a public leaderboard submission.

- [Launch and API example](https://fastino.ai/blog/introducing-glide-the-first-thinking-decision-model)

## GLiNER2.5-Decide

Model · Fastino Labs

A small encoder that makes decisions, extracts spans and relationships, and enforces constraints between outputs. You define the schema. Released September 24, 2026.

Inputs: Text. Size: 340M (publisher figure).

Access: Open weights.

License: Apache-2.0.

Deployment: The gliner2 Python library on CPU or GPU; supports local and air-gapped deployment.

Limits: It has its own schema and constraint interface. Check compatibility before using it with a Jev client.

- [Launch and specifications](https://fastino.ai/blog/gliner-2-5-decide-open-weight-decision-model)

- [Weights](https://huggingface.co/fastino/GLiNER2.5-Decide)

## Imajev

Model · mohit67890

Answers typed questions about photos and application state, including an explicit unknown answer when evidence is missing. The project publishes weights, training details, and visual tests.

Inputs: Text / JSON / images. Size: 2B / 4B / 9B.

Access: Open weights.

License: See release and base-model licenses.

Deployment: Local serving with calibration files; includes an MLX backend and playground.

Limits: English only, with at most two images and eight questions per request. Some public test results influenced checkpoint selection, and an unknown answer is not guaranteed.

- [Releases, evaluation, and limits](https://github.com/mohit67890/imajev)

## Intern-Decision

Model · InternLM

Answers questions about text and images using Qwen3.5 models and parallel answer markers. The project shares training and inference code, calibration tools, and links to weights.

Inputs: Text / images. Size: Checkpoint-dependent; 4B documented.

Access: Open weights.

License: See checkpoint and dependency terms.

Deployment: Native Hugging Face or XTuner inference; HTTP service and browser demo.

Limits: Download the weights and training data separately from the code. The accuracy and calibration reports cover different evaluation sets.

- [Repository and model collection](https://github.com/InternLM/Intern-Decision)

## Jeeves

Model · PostHog

PostHog’s Jev-style classifier can think before answering. It combines Qwen3.5-9B with a pointer head and a diffusion drafter.

Inputs: Text. Size: 9B.

Access: Open weights.

License: Code MIT; check weight terms.

Deployment: CUDA or Apple Silicon MPS; Jev-compatible API; published weights and training code.

Limits: Thinking takes extra time. Compare its reasoning mode and one-pass mode separately.

- [Repository and modes](https://github.com/PostHog/jeeves)

- [Weights](https://huggingface.co/PostHog/jeeves)

## Julia-1

Model · Supersonic Labs

A small mmBERT-based encoder for classification, routing, scores, and yes/no questions. Its model card reports scenario-classification tests across 52 locales.

Inputs: Text / JSON; multilingual. Size: 144.3M.

Access: Open weights.

License: Apache-2.0.

Deployment: Python on CPU or CUDA, a separate ONNX/WebGPU release, or ggml-org’s GGUF through llama.cpp.

Limits: The native runtime accepts 2–20 options per question. Its 8K context limit has a smoke test, but the published accuracy tests used 1K inputs. Training code is not part of the release.

- [Model card and evaluations](https://huggingface.co/SupersonicLabs/Julia-1)

- [llama.cpp weights](https://huggingface.co/ggml-org/Julia-1-GGUF)

## lev

Model · Interfaze AI

A Qwen3.5-4B LoRA adapter that answers typed questions about shared context. It supports TypeSafe’s request format and reads probabilities for the supplied options.

Inputs: English text / JSON. Size: 4B backbone.

Access: Open weights.

License: Apache-2.0.

Deployment: The original adapter requires the Qwen base model and its serving setup. ggml-org also publishes a GGUF for native llama.cpp serving.

Limits: Some reported evaluation numbers predate a serving update that changed how large choice sets are scored. Match the serving version when comparing results.

- [Adapter and evaluation notes](https://huggingface.co/interfaze-ai/lev)

- [llama.cpp weights](https://huggingface.co/ggml-org/lev-GGUF)

## Mercury Decide

Model · Inception

Inception’s decision endpoint uses the Jev schema for choices, scores, and yes/no questions. OpenRouter lists its release as September 30, 2026.

Inputs: Text. Size: Not disclosed.

Access: Hosted API.

License: Hosted service terms.

Deployment: OpenRouter, including a free route with a 32,768-token context window.

Limits: The free route has rate limits. Check the hosted route’s current availability and limits.

- [Hosted model and limits](https://openrouter.ai/inception/mercury-decide:free)

- [Company launch](https://x.com/_inception_ai/status/2105375133003325768)

## NanoJev

Model · Tianyu Chen

A small model that makes decisions in parallel, with one checkpoint for Maze, Snake, and ViZDoom. The project includes the full training pipeline.

Inputs: Structured game states. Size: 0.6B.

Access: Open weights.

License: See repository and checkpoint terms.

Deployment: Downloadable checkpoint, dataset, and local game demonstrations.

Limits: The published tests focus on games. Performance on business classification tasks has not been established here.

- [Repository and task scope](https://github.com/TianyuCodings/NanoJev)

- [Weights](https://huggingface.co/C-Tianyu/NanoJev)

## OneJev

Model · OmniJev

A multimodal family that scores allowed answers to typed questions in one forward pass. The release includes training data and local serving examples.

Inputs: Text / screenshots / photos / video. Size: 0.8B / 4B / 9B / 27B.

Access: Open weights.

License: Apache-2.0.

Deployment: Published Hugging Face checkpoints with PyTorch and llama.cpp serving instructions.

Limits: Check each checkpoint’s evaluation and runtime requirements. A video-capable interface alone does not establish accuracy on your video tasks.

- [Repository and model collection](https://github.com/OmniJev/OneJev)

## OpenJev

Model · OpenJev project

Answers typed questions about text, web pages, and screenshots. The project publishes multilingual evaluations and several quantized releases.

Inputs: Text / JSON / screenshots. Size: 27B.

Access: Open weights.

License: Weights CC BY-NC 4.0; helper code Apache-2.0.

Deployment: vLLM with the project’s decision helper, text-only MLX variants, or ggml-org’s GGUF with image support in llama.cpp.

Limits: Weights permit non-commercial use; commercial use needs a separate license. This is a distinct project from SemIf, which previously used the OpenJev name. Check the exact conversion because image support differs across formats.

- [Model card and license](https://huggingface.co/openjev/openjev)

- [ggml-org conversion](https://huggingface.co/ggml-org/OpenJev-GGUF)

## OpenThai-SystemOne

Model · iApp Technology

A Qwen3.5-0.8B text model continued-pretrained on Thai, then trained for Choice, Score, and Noul questions with a 256-way decision head.

Inputs: Thai / English text and JSON. Size: 0.8B.

Access: Open weights.

License: Apache-2.0.

Deployment: Python package and SystemOne-compatible server on CUDA, MPS, or CPU. Training and evaluation scripts are public.

Limits: The vision encoder is removed. Test Thai and English separately; multilingual support does not imply equal performance in both languages.

- [Repository and training pipeline](https://github.com/iapp-technology/openthai-systemone)

- [Weights and evaluations](https://huggingface.co/iapp/OpenThai-SystemOne)

## Solar Decide

Model · Upstage

Built on Solar Mini 4, with a 512K context window and the Jev/SystemOne request format. OpenRouter lists its release as September 28, 2026.

Inputs: Text. Size: Solar Mini 4 backbone.

Access: Hosted API.

License: Hosted service terms.

Deployment: OpenRouter; $0.05 per million input tokens and no output-token charge in the checked listing.

Limits: Test your longer documents: a 512K context limit says how much input fits, not how well the model uses all of it.

- [Hosted model and pricing](https://openrouter.ai/upstage/solar-decide)

## Span-01 / Span-01 Lite

Model · Respan

Checks agent traces for behaviors you describe in plain language. For each behavior, it returns probabilities for present, absent, and not observable.

Inputs: Text / agent traces. Size: Not disclosed.

Access: Hosted API.

License: Hosted service terms.

Deployment: Respan’s hosted classifier and Lite tier; follow the launch page’s quickstart.

Limits: Designed for behavior detection. Its answer categories differ from the Choice/Score/Noul format.

- [Model and quickstart](https://www.respan.ai/blog/introducing-span-1)

## Standard One

Model · Standard Thinking

Ministral-based models with merged weights, adapters, and a SystemOne server. The vision component is retained, so requests can include image data URLs.

Inputs: Text / images. Size: 3B / 8B.

Access: Open weights.

License: Apache-2.0.

Deployment: Self-hosted serving, including SGLang and published quantized variants.

Limits: The model card does not report a separate image-decision benchmark. The 8B v2 weights were updated September 26, 2026.

- [Model card, versions, and server](https://huggingface.co/StandardThinking/StandardOne-8B)

## Strands Decider

Model · Strands Labs

Built on Qwen3.5-2B, with a pointer head and LoRA adaptation. It answers choice, score, and yes/no questions for agent workflows.

Inputs: Text. Size: 1.9B.

Access: Open weights.

License: Apache-2.0.

Deployment: The strands-decider CLI and Python package; published weights and training resources.

Limits: The published calibration results apply to the tested tasks. Set confidence thresholds using your own labeled examples.

- [Repository, weights, and evaluations](https://github.com/strands-labs/strands-decider)

## Surogate Rune

Model · Invergent

A Gemma-based model for choices, yes/no checks, and scores. It can reason before answering uncertain questions when thinking mode is enabled.

Inputs: Text / JSON / images. Size: 26.5B total / ~4B active.

Access: Hosted API, Open weights.

License: Apache-2.0; download gate applies.

Deployment: Hosted API or the Surogate serving engine. Current v3 files are BF16 safetensors despite the repository’s GGUF name.

Limits: Downloading requires accepting the repository’s conditions. Thinking changes latency and is not available with images in the documented release.

- [Model card and weights](https://huggingface.co/surogate/rune-26b-a4b-GGUF)

- [API and deployment](https://surogate.ai/labs/rune/)

## Tev1

Model · Together AI

Together AI’s experimental Qwen3.5 fine-tunes. The repository includes the data-building and supervised training recipe; its example limits answers to the supplied option letters.

Inputs: Text. Size: 4B / 0.8B experimental.

Access: Hosted API, Open weights.

License: Code MIT; weights have separate terms.

Deployment: Published checkpoints, Together fine-tuning and serving, or supported local runtimes.

Limits: Together describes the token log probabilities as model preferences, not calibrated confidence. The development benchmarks were used more than once.

- [Recipe and limitations](https://github.com/togethercomputer/tev1)

- [4B checkpoint](https://huggingface.co/togethercomputer/Tev1-4B-experimental)

- [Serverless announcement](https://x.com/togethercompute/status/2102882216950763814)

## U2-Decision

Model · Unisound

A Qwen3.5-4B LoRA adapter for safety checks, model routing, and general typed decisions. The public server uses the name u2_decision_preview.

Inputs: Text / JSON. Size: 4B backbone.

Access: Open weights.

License: Adapter Apache-2.0; base model has separate terms.

Deployment: Download the adapter from GitHub with Git LFS and obtain Qwen3.5-4B separately. Includes a /v1/systemone server and evaluation scripts.

Limits: The release contains adapter weights, not a complete base model. Its routing evaluation has only 12 cases; broader routing claims need more testing.

- [Adapter, API, and evaluations](https://github.com/Unisound-LLM/u2_decision_4b)

- [Product page](https://maas.unisound.com/models/u2-decision)

## Von

Model · wfzyx

A ModernBERT encoder that returns Choice, Noul, and Score answers. It comes with Python and TypeScript clients and a SystemOne-compatible server.

Inputs: Text / JSON. Size: 395M.

Access: Open weights.

License: Apache-2.0.

Deployment: CPU with OpenVINO, CUDA, ROCm, or Apple MPS.

Limits: English only. The author reports that confidence can be misleading on unfamiliar tasks.

- [Repository and limitations](https://github.com/wfzyx/von)

## XOR

Model · Xyne / Juspay

A Qwen3.6-35B-A3B fine-tune for typed questions about text and visual content. The release includes merged weights and a SystemOne-compatible serving bundle.

Inputs: Text / images / video. Size: 35B total / ~3B active.

Access: Open weights.

License: Apache-2.0.

Deployment: Self-host with the pinned, patched SGLang runtime. Xor 1.2 is available on the release/xor-1.2 branch.

Limits: Pin the release revision and use its serving bundle. The model card’s JevBench numbers are self-run public-subset results; send images or video separately.

- [Weights and release instructions](https://huggingface.co/juspay/xor)

- [Company announcement](https://x.com/juspay/status/2103047740276240418)

## AnyJev

Inference method · Nokia Applied Research

A toolkit for getting structured decisions from existing language models. It tests how answer order affects results and compares ways to speed up inference.

Inputs: Depends on base model. Size: Depends on base model.

Access: Open code.

License: Apache-2.0 code; upstream model licenses.

Deployment: Local models and supported serving engines, including vLLM.

Limits: Results depend on which base model and scoring method you use.

- [Repository and evidence](https://github.com/nokia-applied-research/AnyJev)

## djev

Inference method · Davipar

Adds typed decision outputs to DiffusionGemma through vLLM, with multimodal support and a local playground.

Inputs: Text / images. Size: DiffusionGemma 26B-A4B backbone.

Access: Open code.

License: Code Apache-2.0; base model terms apply.

Deployment: Self-host the upstream model with the repository’s runtime changes and decision layer.

Limits: The project does not release newly trained weights. Its performance report calls for matched baselines and further held-out quality evaluations.

- [Code, scope, and performance notes](https://github.com/Davipar/djev-dev)

## LLM2Jev

Inference method · Yinsongxu and contributors

Scores decisions during a local model’s input-processing pass. Includes text and image examples, plus research on optional fine-tuning.

Inputs: Text / images, model-dependent. Size: Depends on base model.

Access: Open code.

License: See repository and base-model terms.

Deployment: Repository inference and evaluation tools with compatible local checkpoints.

Limits: Choose a compatible base model and serving setup separately; the project provides neither a single fixed model nor a service guarantee.

- [Repository](https://github.com/Yinsongxu/LLM2Jev)

- [Research paper](https://arxiv.org/abs/2610.02076)

## Reflex

Inference method · kshetrajna12

Reads answer probabilities from a frozen model and averages two option orders. The current stable setup uses Qwen3.5-4B without training new weights.

Inputs: Text / images. Size: Base-model dependent; 4B default.

Access: Open code.

License: Code MIT; base model terms apply.

Deployment: Local SystemOne endpoint on CUDA. A smaller WebGPU version demonstrates the same approach in a browser.

Limits: Older benchmark entries used a LoRA configuration. Do not treat those scores as results for the current frozen-model setup.

- [Current status and evaluation history](https://github.com/kshetrajna12/reflex)

## SemIf (formerly OpenJev)

Inference method · TheoLeeCJ and contributors

Reads answer scores from existing open models without retraining them. Includes calibration and scoring for questions about a shared input. It follows the Jev interface, without claiming to reproduce Jev’s architecture.

Inputs: Text / JSON. Size: Depends on base model.

Access: Open code.

License: MIT code; upstream model licenses.

Deployment: CUDA, MLX, MPS, llama.cpp, and a browser demo; Databricks publishes a deployment notebook.

Limits: SemIf runs on other models and is independent of TypeSafe. It does not name a single proprietary checkpoint.

- [Repository](https://github.com/TheoLeeCJ/SemIf-OpenJev)

- [Databricks deployment](https://www.databricks.com/blog/running-open-jev-sql-databricks)

## Decisions API

Platform · OpenAI

OpenAI announced a limited preview on September 29, 2026. Define questions and allowed answers to classify content, route requests, or choose an agent’s next action.

Inputs: Content + questions and allowed answers. Size: Powered by GPT-6 Luna.

Access: Hosted API.

License: Limited-preview service terms.

Deployment: Limited preview. The announcement identifies GPT-6 Luna as the underlying model.

Limits: The launch post does not establish general availability, public pricing, or compatibility with the Jev request format.

- [Official developer announcement](https://x.com/OpenAIDevs/status/2105003318917697873)

## Instinct

Platform · ZooWork

ZooWork’s API answers questions with choices, scores, and yes/no probabilities. It offers three Qwen configurations: unchanged 27B, unchanged 4B that scores options in two orders, and a 4B fine-tuned for decisions.

Inputs: Text / JSON. Size: 4B / 27B configurations.

Access: Hosted API.

License: Hosted service terms; check individual weight releases.

Deployment: Paid API through ZooData, or a separate free evaluation preview with its own account and key.

Limits: The published comparison uses 231 public questions. Two configurations use unchanged Qwen weights; only instinct-tuned-4b adds a decision fine-tune.

- [Models and access](https://instinct.zoowork.ai/)

- [API reference](https://instinct.zoowork.ai/docs/)

## llama.cpp decision models

Platform · ggml-org

Run decision models locally through the native /v1/systemone endpoint. The October 2 release supports Julia-1, Laya, Kev-4B, lev, and OpenJev, with Choice, Score, and Noul questions.

Inputs: Text / JSON / images with supported models. Size: Depends on selected model.

Access: Open code.

License: Runtime MIT; model licenses differ.

Deployment: Update llama.cpp, then start a supported GGUF, for example: llama serve -hf ggml-org/Kev-4B-GGUF. Router mode loads models on demand.

Limits: Image support depends on the model. OpenJev’s weights are non-commercial. The launch lists Cloudflare Clef as coming next; its reported timings use one RTX PRO 6000, not a typical laptop.

- [Launch and examples](https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp)

- [Merged implementation](https://github.com/ggml-org/llama.cpp/pull/29818)

## Ollama decision models

Platform · Ollama

Run Jev-style requests locally with models such as Nimble and Tev1. Ollama added this support on September 29, 2026.

Inputs: Depends on selected model. Size: Depends on selected model.

Access: Open code.

License: Runtime and model licenses differ.

Deployment: Ollama’s local /v1/systemone endpoint with a supported decision checkpoint.

Limits: Ollama runs the model you download. That model keeps its own license and capabilities.

- [Release and usage](https://ollama.com/blog/ollama-now-supports-jev-style-decision-models)

## Ollaya

Platform · ollaya-dev

Downloads and serves local decision models through a TypeSafe-compatible API. Supports several encoder and decoder backends.

Inputs: Depends on selected model. Size: Depends on selected model.

Access: Open code.

License: Runtime and model licenses differ.

Deployment: Local CLI and server; use each checkpoint’s supported backend and calibration.

Limits: Runs models such as Laya and Kev; it does not add another model family.

- [Repository and supported models](https://github.com/ollaya-dev/ollaya)

## Jev Ultrafast

Integration · Browser Use

Builds a list of browser elements and actions, then asks Jev which operation and target to choose. A separate language model writes text when an action needs it.

Inputs: Browser state. Size: Uses TypeSafe Jev + a text model.

Access: Open code.

License: Check repository and API terms.

Deployment: Run the public Python agent locally with Browser Harness and your API keys. The cloud offering has a waitlist.

Limits: This is a browser-agent integration. Its recorded demo timings apply to those tasks and include more than the decision-model call.

- [Official repository and measurements](https://github.com/browser-use/jev-ultrafast)

## Neon pg_redact example

Integration · Neon / Rishi Raj Jain

A Postgres example that uses Jev to identify personal information in messages. It stores the labels as JSONB and uses SQL functions to hide text according to the reader’s role.

Inputs: Support-message text. Size: Uses TypeSafe Jev.

Access: Open code, Hosted API.

License: Example code; Neon and TypeSafe service terms apply.

Deployment: Run the example app with Neon Postgres and a TypeSafe API key. Neon published the walkthrough on October 1, 2026.

Limits: The model is Jev; Neon provides the database integration. The example can miss sensitive text, so it does not establish complete PII detection.

- [Neon walkthrough](https://neon.com/blog/pg-redact-content-aware-pii-redaction-with-jev-enforced-in-postgres)

- [Example source](https://github.com/rishi-raj-jain/pg-redact)

## Existing benchmarks

Benchmark sources checked October 2, 2026. These groups ran the evaluations; we have not reproduced them. Compare scores within the same test suite, version, and serving setup.

### Decision Index

Community benchmark maintained by multimodalart and contributors

Compares decision models across knowledge, language, retrieval, tools, and human preferences. Version 0.2.1 combines 38 benchmarks into a score adjusted for chance performance.

How to read it: The index changes its tasks and scoring between versions. A provider’s own run of the suite is separate from a result accepted by the public leaderboard. Check the submission notes and training-data disclosures.

- [Leaderboard and methodology](https://huggingface.co/spaces/multimodalart/jev-decision-index)

### JevBench / Benchmark Heaven

Run by Benchmark Heaven

Ranks models using decision quality, calibration, speed, and cost. The repository includes the scoring code, results, and public test questions. Some evaluation questions stay private.

How to read it: Check the methodology version before comparing scores. The ranking depends on how the four measures are weighted, and some cost and latency figures use estimates or adjustments. Results for private questions are reported in aggregate.

- [Method and code](https://github.com/fstandhartinger/jevbench)

- [Leaderboard](https://benchmarkheaven.com/jev-models)

### System One Mosaic Benchmark (S1MB)

Run by hotchpotch, who also makes a model

Tested 23 models across 137 English Noul, Choice, and Score tasks in its September 30, 2026 snapshot: 14,009 cases in total. You can download the code, data, and results.

How to read it: Borda Score tells you how models rank against each other. It is not an accuracy percentage. The author’s model trained on datasets that also supply evaluation tasks; a separate synthetic test checks how well it handles unfamiliar tasks.

- [Report and methodology](https://huggingface.co/blog/hotchpotch/system-one-mosaic-benchmark)

- [Evaluation code](https://github.com/hotchpotch/S1MB)

### sysone-bench

Run by instax-dutta

Gives Jev 1.13.0, Laya 0.3.11, and Qwen2.5-1.5B exactly the same inputs from a sealed manifest. Qwen uses constrained decoding. The evaluation split contains 952 cases and 1,240 scored decisions.

How to read it: The results apply to these three configurations. Keep the local CPU setup and hosted API separate when comparing speed.

- [Protocol, results, and code](https://github.com/instax-dutta/sysone-bench)

### Clef evaluation

Run by Cloudflare using Decision Index tasks

Cloudflare ran Clef, Clef-flash, Jev, Kev, Laya, and a DiffusionGemma implementation on Decision Index tasks. The report includes timings and a separate run of TypeSafe’s workflow tests.

How to read it: Cloudflare makes Clef and Clef-flash. These are its own runs of a community test suite; other groups’ runs may use different setups.

- [Decision Index suite](https://huggingface.co/spaces/multimodalart/jev-decision-index)

- [Cloudflare report](https://blog.cloudflare.com/clef-decision-models/)

- [Cloudflare leaderboard](https://clef-evals.workers-ai-mle.workers.dev/)

### Cua-Bench-S1

Run by Cua, which also makes Cua-S1

Tests choices among fixed computer-use actions from screenshots or accessibility trees. Publishes text and multimodal results, dataset hashes, and separate live-environment task-completion results.

How to read it: Read the split and modality for each row. Training exposure differs across models, some task families are very small, and the report documents labeling artifacts. Offline action choices do not measure a complete agent workflow.

- [Tasks, results, and limitations](https://github.com/trycua/cua/tree/main/libs/cua-bench-s1)

### U2-Decision evaluation sets

Run by Unisound

The public release includes 500 safety cases, 212 general decision cases, and 12 routing cases, with scripts for reproducing each evaluation.

How to read it: These are provider-specific tests. The routing set is especially small, so its score offers limited evidence about broader routing performance.

- [Data and evaluation scripts](https://github.com/Unisound-LLM/u2_decision_4b)

### Bespoke Nimble public benchmarks

Run by Bespoke Labs

Tests Nimble and Jev on 3,880 records from 13 subsets labeled by people. Tasks include routing, moderation, retrieval, entailment, and rubric scoring. Fixed record lists let you rebuild the same subsets.

How to read it: Scores measure agreement with the people who labeled the data. Task sizes and label distributions vary. These tests are separate from Nimble’s synthetic training and holdout sets.

- [Datasets, results, and reproduction](https://github.com/bespokelabsai/nimble/blob/main/docs/PUBLIC_BENCHMARKS.md)

### Kev frozen evaluation suites

Run by Kev’s developer

Separately tests new examples from familiar training sources and examples from new sources. Reports accuracy, Brier score, calibration, and how many decisions can be automated at a fixed error rate.

How to read it: Use the same suite version, checkpoint, and data split. Jev’s development-set runs and Kev’s test-set scores answer different questions. The published evaluations use fp32, which also differs from default serving precision.

- [Evaluation protocol](https://github.com/jaredpalmer/kev#benchmarks)

- [Frozen suites](https://huggingface.co/datasets/jaredpalmer/kev-suites)

## Compare on your own task

Give each model the same inputs and questions. Keep labeled cases aside for testing, then measure accuracy, confidence, the share of decisions you can automate, latency, and cost. Record the model versions and hardware or endpoints. We have not independently reproduced the providers’ results.

We found projects through GitHub and X, then checked their docs, model cards, repositories, and serving-platform listings. We include documented company releases and open projects with published weights, working inference code, or a distinct language or task focus. Related sizes share an entry; ports of the same weights do not count as new model families. Selected platforms and integrations have their own categories.

[Coding Models Index](https://amplifying.ai/models) · [All agent-readable resources](https://amplifying.ai/llms.txt)