Skip to content
New research:What Dots Actually Choose

Decision Models Index

Give a decision model a message and a few questions. It returns answers your code can use, like which team should handle a ticket or how urgent it is.

Start here

Six places to start, depending on how you want to use a decision model.

Run and train your own model

Kev

Open weights

Jared Palmer’s model family comes with training recipes and frozen evaluation sets. It’s a useful starting point if you want to inspect or adapt a decision model.

Details and sources for Kev

Smaller models for local use

Laya

Open weights

Laya’s 322M and 421M encoders offer a smaller local option. The family includes multilingual checkpoints and community runtimes for Apple Silicon and browsers.

Details and sources for Laya

These editorial picks highlight API support, deployment choices, and published code or training resources. Use the benchmarks below to compare measured performance.

Browse the catalog

Browse models you can call or download, methods that use existing models, and tools for running them. Use the category filters to narrow the list.

Open weights includes complete checkpoints and adapters that need a separate base model. Some downloads require accepting access terms. Open code means the implementation is public; each entry lists the license and deployment requirements.

For measured results, see the benchmarks below.

Loading the decision-model catalog…

Recent developments

Fastino introduces GLiDE

GLiDE spends extra time reasoning when its initial decision is uncertain. Fastino offers it through a hosted API.

View entry: Fastino introduces GLiDE

Databricks adds ai_decide

The beta function answers questions about governed data through SQL or REST. Databricks can change the model behind it.

View entry: Databricks adds ai_decide

OpenAI previews Decisions API

The API uses GPT-6 Luna to classify content, route requests, and choose an agent’s next action. OpenAI announced limited-preview access.

View entry: OpenAI previews Decisions API

What decision models return

Send a message, document, or other input, along with the questions you want answered. APIs often call this input the state. You could ask these three questions about the same support ticket:

Choice · pick an option

Which team owns this issue?

Billing · Support · Engineering

Score · apply a rubric

How soon does this need attention?

Can wait · This week · Today

Noul · check a statement

Does the message describe an outage?

Probability of yes, from 0 to 1

These examples show TypeSafe’s three question types. They are not live model responses. Other models may use different request formats or score ranges. TypeSafe’s docs explain the format.

Existing benchmarks

Several groups have tested these models. Some publish a leaderboard; others share a test suite you can run yourself. We checked the sources below on October 2, 2026, but haven’t reproduced their runs.

Community benchmark maintained by multimodalart and contributors

Decision Index

Compares decision models across knowledge, language, retrieval, tools, and human preferences. Version 0.2.1 combines 38 benchmarks into a score adjusted for chance performance.

The index changes its tasks and scoring between versions. A provider’s own run of the suite is separate from a result accepted by the public leaderboard. Check the submission notes and training-data disclosures.

Run by Benchmark Heaven

JevBench / Benchmark Heaven

Ranks models using decision quality, calibration, speed, and cost. The repository includes the scoring code, results, and public test questions. Some evaluation questions stay private.

Check the methodology version before comparing scores. The ranking depends on how the four measures are weighted, and some cost and latency figures use estimates or adjustments. Results for private questions are reported in aggregate.

Run by hotchpotch, who also makes a model

System One Mosaic Benchmark (S1MB)

Tested 23 models across 137 English Noul, Choice, and Score tasks in its September 30, 2026 snapshot: 14,009 cases in total. You can download the code, data, and results.

Borda Score tells you how models rank against each other. It is not an accuracy percentage. The author’s model trained on datasets that also supply evaluation tasks; a separate synthetic test checks how well it handles unfamiliar tasks.

Run by instax-dutta

sysone-bench

Gives Jev 1.13.0, Laya 0.3.11, and Qwen2.5-1.5B exactly the same inputs from a sealed manifest. Qwen uses constrained decoding. The evaluation split contains 952 cases and 1,240 scored decisions.

The results apply to these three configurations. Keep the local CPU setup and hosted API separate when comparing speed.

Run by Cua, which also makes Cua-S1

Cua-Bench-S1

Tests choices among fixed computer-use actions from screenshots or accessibility trees. Publishes text and multimodal results, dataset hashes, and separate live-environment task-completion results.

Read the split and modality for each row. Training exposure differs across models, some task families are very small, and the report documents labeling artifacts. Offline action choices do not measure a complete agent workflow.

Run by Unisound

U2-Decision evaluation sets

The public release includes 500 safety cases, 212 general decision cases, and 12 routing cases, with scripts for reproducing each evaluation.

These are provider-specific tests. The routing set is especially small, so its score offers limited evidence about broader routing performance.

Run by Bespoke Labs

Bespoke Nimble public benchmarks

Tests Nimble and Jev on 3,880 records from 13 subsets labeled by people. Tasks include routing, moderation, retrieval, entailment, and rubric scoring. Fixed record lists let you rebuild the same subsets.

Scores measure agreement with the people who labeled the data. Task sizes and label distributions vary. These tests are separate from Nimble’s synthetic training and holdout sets.

Run by Kev’s developer

Kev frozen evaluation suites

Separately tests new examples from familiar training sources and examples from new sources. Reports accuracy, Brier score, calibration, and how many decisions can be automated at a fixed error rate.

Use the same suite version, checkpoint, and data split. Jev’s development-set runs and Kev’s test-set scores answer different questions. The published evaluations use fp32, which also differs from default serving precision.

A leaderboard can help you pick models to try. Compare scores from the same test version, and check whether a model trained on the test’s source data. Labels, confidence measurements, and network time can all change how you read the results.

How to compare

Give each model the same inputs, questions, and answer choices. Keep some labeled examples aside before you tune anything. Include awkward cases where none of the choices fits, and count the errors that would cause problems in your application.

Check whether a model’s confidence matches how often it gets things right. Then measure how many decisions you can automate at an acceptable error rate, and how many still need review.

Time the whole request, including the network if you use an API. Record the model version, hardware, input length, questions per call, and number of simultaneous requests. Report the median and 95th-percentile response time alongside cost per decision.

We found projects through GitHub and X, then checked their docs, repositories, model cards, and serving-platform listings. Each entry links to those sources. We haven’t run our own comparison of these models or combined their benchmark scores.

We include documented company releases and open projects with published weights, working inference code, or a distinct language or task focus. This list will miss some projects. Related sizes share an entry; ports of the same weights do not count as new model families. Selected platforms and integrations have their own categories.

Compare coding modelsExplore coding agents