The reference API
TypeSafe Jev
Hosted API
TypeSafe’s Choice, Score, and Noul format is the reference for many entries here. Start with its docs to understand how these models work.
Details and sources for TypeSafe JevGive a decision model a message and a few questions. It returns answers your code can use, like which team should handle a ticket or how urgent it is.
Six places to start, depending on how you want to use a decision model.
The reference API
Hosted API
TypeSafe’s Choice, Score, and Noul format is the reference for many entries here. Start with its docs to understand how these models work.
Details and sources for TypeSafe JevCloud deployment and open weights
Hosted API · Open weights
Cloudflare offers Clef and the smaller Clef-flash through Workers AI and as downloadable weights. Both accept text, images, and video.
Details and sources for Cloudflare ClefA hosted API you can also self-host
Hosted API · Open weights
Perplexity publishes the 27B model behind its Decisions API, along with inference code. You can try the service before setting up your own GPU.
Details and sources for Perplexity DecisionsRun and train your own model
Open weights
Jared Palmer’s model family comes with training recipes and frozen evaluation sets. It’s a useful starting point if you want to inspect or adapt a decision model.
Details and sources for KevSmaller models for local use
Open weights
Laya’s 322M and 421M encoders offer a smaller local option. The family includes multilingual checkpoints and community runtimes for Apple Silicon and browsers.
Details and sources for LayaDecisions inside your data platform
Beta · Managed platform
For teams already using Databricks, ai_decide brings typed questions into SQL and REST workflows over governed data. Databricks manages the underlying model.
Details and sources for Databricks ai_decideThese editorial picks highlight API support, deployment choices, and published code or training resources. Use the benchmarks below to compare measured performance.
Browse models you can call or download, methods that use existing models, and tools for running them. Use the category filters to narrow the list.
Open weights includes complete checkpoints and adapters that need a separate base model. Some downloads require accepting access terms. Open code means the implementation is public; each entry lists the license and deployment requirements.
For measured results, see the benchmarks below.
Loading the decision-model catalog…
Serve Julia-1, Laya, Kev-4B, lev, or OpenJev through /v1/systemone. Image support depends on the model; Clef is still listed as coming next.
View entry: llama.cpp adds native decision-model supportClef and Clef-flash accept text, images, and video. You can call them on Workers AI or download the Apache-2.0 weights.
View entry: Cloudflare adds two open modelspplx-decider-v1-27b is available through the Decisions API and as Apache-2.0 weights with inference code.
View entry: Perplexity releases its decision modelGLiDE spends extra time reasoning when its initial decision is uncertain. Fastino offers it through a hosted API.
View entry: Fastino introduces GLiDEThe beta function answers questions about governed data through SQL or REST. Databricks can change the model behind it.
View entry: Databricks adds ai_decideThe API uses GPT-6 Luna to classify content, route requests, and choose an agent’s next action. OpenAI announced limited-preview access.
View entry: OpenAI previews Decisions APISend a message, document, or other input, along with the questions you want answered. APIs often call this input the state. You could ask these three questions about the same support ticket:
Which team owns this issue?
Billing · Support · Engineering
How soon does this need attention?
Can wait · This week · Today
Does the message describe an outage?
Probability of yes, from 0 to 1
These examples show TypeSafe’s three question types. They are not live model responses. Other models may use different request formats or score ranges. TypeSafe’s docs explain the format.
Several groups have tested these models. Some publish a leaderboard; others share a test suite you can run yourself. We checked the sources below on October 2, 2026, but haven’t reproduced their runs.
Community benchmark maintained by multimodalart and contributors
Compares decision models across knowledge, language, retrieval, tools, and human preferences. Version 0.2.1 combines 38 benchmarks into a score adjusted for chance performance.
The index changes its tasks and scoring between versions. A provider’s own run of the suite is separate from a result accepted by the public leaderboard. Check the submission notes and training-data disclosures.
Run by Benchmark Heaven
Ranks models using decision quality, calibration, speed, and cost. The repository includes the scoring code, results, and public test questions. Some evaluation questions stay private.
Check the methodology version before comparing scores. The ranking depends on how the four measures are weighted, and some cost and latency figures use estimates or adjustments. Results for private questions are reported in aggregate.
Run by hotchpotch, who also makes a model
Tested 23 models across 137 English Noul, Choice, and Score tasks in its September 30, 2026 snapshot: 14,009 cases in total. You can download the code, data, and results.
Borda Score tells you how models rank against each other. It is not an accuracy percentage. The author’s model trained on datasets that also supply evaluation tasks; a separate synthetic test checks how well it handles unfamiliar tasks.
Run by instax-dutta
Gives Jev 1.13.0, Laya 0.3.11, and Qwen2.5-1.5B exactly the same inputs from a sealed manifest. Qwen uses constrained decoding. The evaluation split contains 952 cases and 1,240 scored decisions.
The results apply to these three configurations. Keep the local CPU setup and hosted API separate when comparing speed.
Run by Cloudflare using Decision Index tasks
Cloudflare ran Clef, Clef-flash, Jev, Kev, Laya, and a DiffusionGemma implementation on Decision Index tasks. The report includes timings and a separate run of TypeSafe’s workflow tests.
Cloudflare makes Clef and Clef-flash. These are its own runs of a community test suite; other groups’ runs may use different setups.
Run by Cua, which also makes Cua-S1
Tests choices among fixed computer-use actions from screenshots or accessibility trees. Publishes text and multimodal results, dataset hashes, and separate live-environment task-completion results.
Read the split and modality for each row. Training exposure differs across models, some task families are very small, and the report documents labeling artifacts. Offline action choices do not measure a complete agent workflow.
Run by Unisound
The public release includes 500 safety cases, 212 general decision cases, and 12 routing cases, with scripts for reproducing each evaluation.
These are provider-specific tests. The routing set is especially small, so its score offers limited evidence about broader routing performance.
Run by Bespoke Labs
Tests Nimble and Jev on 3,880 records from 13 subsets labeled by people. Tasks include routing, moderation, retrieval, entailment, and rubric scoring. Fixed record lists let you rebuild the same subsets.
Scores measure agreement with the people who labeled the data. Task sizes and label distributions vary. These tests are separate from Nimble’s synthetic training and holdout sets.
Run by Kev’s developer
Separately tests new examples from familiar training sources and examples from new sources. Reports accuracy, Brier score, calibration, and how many decisions can be automated at a fixed error rate.
Use the same suite version, checkpoint, and data split. Jev’s development-set runs and Kev’s test-set scores answer different questions. The published evaluations use fp32, which also differs from default serving precision.
A leaderboard can help you pick models to try. Compare scores from the same test version, and check whether a model trained on the test’s source data. Labels, confidence measurements, and network time can all change how you read the results.
Give each model the same inputs, questions, and answer choices. Keep some labeled examples aside before you tune anything. Include awkward cases where none of the choices fits, and count the errors that would cause problems in your application.
Check whether a model’s confidence matches how often it gets things right. Then measure how many decisions you can automate at an acceptable error rate, and how many still need review.
Time the whole request, including the network if you use an API. Record the model version, hardware, input length, questions per call, and number of simultaneous requests. Report the median and 95th-percentile response time alongside cost per decision.
We found projects through GitHub and X, then checked their docs, repositories, model cards, and serving-platform listings. Each entry links to those sources. We haven’t run our own comparison of these models or combined their benchmark scores.
We include documented company releases and open projects with published weights, working inference code, or a distinct language or task focus. This list will miss some projects. Related sizes share an entry; ports of the same weights do not count as new model families. Selected platforms and integrations have their own categories.