A support message arrives with two requests: cancel a subscription and explain a charge. Your application needs to choose a queue, decide whether a person should respond, and assign a priority. A paragraph of generated text would give the application more work to do. A decision model can return the values the workflow needs.
Jev from TypeSafe takes information such as a customer message, along with questions that specify the allowed answer format. It returns choices, yes/no probabilities, or scores. Since our look at what developers are building with Jev, the available implementations have expanded to managed APIs, downloadable checkpoints, and smaller classifiers for more specific jobs.
These options move different responsibilities into your application. A managed API leaves inference infrastructure with the provider. A local model puts serving and hardware in your hands. A classifier may cover routing and tagging while leaving scores or other decisions to separate code. The first question is which of those jobs you need to replace.
DataRoot Labs is an AI R&D company that develops and evaluates AI systems for clients. This review draws on official documentation, model cards, and published source code checked on October 11, 2026. We have not yet completed hands-on testing of every system covered here. The tables describe published capabilities, API contracts, and deployment requirements; they do not rank accuracy.
Hosted API comparison
Scroll horizontally to explore all columns →
| Product / route | Input content | Decision output | Integration contract |
|---|---|---|---|
TypeSafe Jev Baseline | Text | Yes/no, choice, score | State + named questions |
OpenAI Decisions Public beta | Text, images | Yes/no, choice, score | Input + question array |
Text, images | Yes/no, choice, score | State + named questions | |
Text, images | Yes/no, choice, score | State + named questions | |
Liquid d1 Via OpenRouter | Text | Yes/no, choice, score | State + named questions |
Microsoft- Public preview | Text | Yes/no, choice, score | Named questions + Foundry deployment |
Text, images | Yes/no, choice, score | State + named questions | |
Text | Yes/no, choice, score | State + named questions | |
Text | Yes/no, choice, score | State + named questions | |
Text | Yes/no, choice, score | State + named questions | |
Text | Labels, multi-label | Text + classification tasks | |
Together Tev1 Experimental | Text | One of 2–24 choices | Chat messages + labelled options |
With named questions, the application assigns a key to each judgment. OpenAI instead accepts an array, while GLiNER exposes classification tasks and Tev1 uses Chat Completions. A matching request shape still leaves response fields, refusals, and score semantics to check before switching.
The Liquid row describes the text route through OpenRouter. Liquid’s direct API also supports images. Hosted d1 and the downloadable d1 checkpoints are separate deployment options.
Local deployment comparison
Scroll horizontally to explore all columns →
| Model / size | Device / runtime | Memory or hardware evidence | Weights license |
|---|---|---|---|
CUDA · custom loader | Tested on H200; no stated minimum | Apache 2.0 | |
CUDA · custom loader | Tested on H200; no stated minimum | Apache 2.0 | |
Liquid d1- 3.12B | CPU / CUDA / Apple MPS | Minimum not specified | LFM Open License · revenue condition |
CPU / CUDA / Apple MPS | Minimum not specified | LFM Open License · revenue condition | |
CUDA · separate decision head | ~49 GiB weights, plus working memory | Apache 2.0 | |
GLiNER2.5- 340M | CPU / GPU · gliner2 | Minimum not specified | Apache 2.0 |
H2O- 4B class | CUDA · vLLM + shim | 9.1 GB weights; tested on 32 GB GPU | Apache 2.0 |
Kev 0.8B / 4B / 9B / 27B | CUDA / Apple MLX | 4B guidance: 24 GB GPU or 32 GB Mac; 27B needs substantially more | Apache 2.0 |
Laya 421M; multilingual 322M | CPU / CUDA / MPS / XPU | Minimum not specified | Apache 2.0 |
Strands Decider, Qwen variant 2B-class base + adapter | CPU / CUDA / MPS | Publisher checked L4 and CPU; minimum not specified | Apache 2.0 |
Tev1- 4B class | Local setup needs validation, per card | Requirement not established | Weight license still being finalized; code is MIT |
Bespoke Nimble 9B v2 9B base + ~173 MB adapter | CUDA / Apple MLX after merge | ~18 GB base weights, plus runtime and merge memory | Apache 2.0 adapter and base |
Local input and output contracts. Hardware fit is only one part of the comparison:
- Images and video frames: local Clef and Clef-flash accept text, JSON, images, and video frame arrays, with yes/no, choice, and ordinal-score outputs. Their hosted API has separate input limits.
- Images, with a separate audio option: d1-3B, pplx-decider, and H2O-Lightning document image inputs and the same three decision types. Strands’ Qwen 2B release needs its optional vision mode. d1-omni also accepts speech; a request may contain images or audio, but not both.
- Text-based typed decisions: Kev and Laya accept text or structured state and return yes/no, choice, and ordinal scores. An ordinal score can be the probability-weighted average of the levels, rather than a selected integer. Laya’s separate schema wrapper projects these decisions into boolean, enum, or integer fields.
- Narrower or differently packaged outputs: GLiNER2.5-Decide classifies text into one or several labels; numeric labels remain classes. Tev1 selects one of 2–24 letters. Nimble returns boolean and enum fields, with an explicit scorer option for expected ordinal values.
The linked model cards and repositories supply the configurations above. Weights-only memory, tested hardware, and a minimum requirement are different measurements. We have not measured local memory use ourselves; context length, batch size, precision, and temporary allocations during loading affect the total. For example, Kev's 4B guidance does not apply to its 27B model, whose repository describes about 51 GB of weights and roughly 66 GB with serving buffers.
Apache-licensed weights do not carry a per-call model fee, but running them still consumes hardware, electricity, or cloud capacity. Liquid's license guide describes no-charge commercial use subject to a $10 million annual-revenue threshold and separate terms beyond that threshold. The license sets the eligibility conditions. Tev1's code license does not resolve its fine-tuned weight terms. Nimble's 173 MB download is an adapter, not a complete 9B model.
Hosted alternatives to Jev
OpenAI Decisions API
OpenAI Decisions is a direct functional alternative for bounded judgments. Its public beta uses gpt-6-luna and supports predicate, choice, and score questions over text and images. It exposes classification and routing separately from text generation within the OpenAI API.
There is an integration change to account for. The Decisions API contract uses input and an array of questions, rather than Jev’s state and map of named questions. Responses also use a different structure, and the application must handle refusals.
Our OpenAI Decisions use-case review examines how developers use these judgments within a larger application.
Cloudflare Clef and Clef-flash
Cloudflare’s Clef family offers two routes: call a model through Workers AI or run the released weights yourself. That makes it relevant to teams already using Cloudflare and to teams evaluating a move from a managed service to their own infrastructure.
The Clef API exposes typed decisions over shared state and supports embedded images. Its hosted image contract has specific limits, including four images per request; an image-capable model is not automatically a replacement for every existing visual workflow. Clef-flash provides a smaller variant with its own context limit.
The published Clef weights use Apache 2.0. Its local loader includes the decision-specific head and preprocessing; loading it as an ordinary chatbot would omit that interface.
Liquid AI d1
Liquid’s decision-model family covers a hosted d1 API and two public checkpoints. d1-3B accepts text and images; the smaller d1-omni-600M supports text paired with images or audio. Its card describes audio training on English speaker–assistant requests and a 30-second clip limit.
The public checkpoints support applications that make decisions on a local device. A visual inspection tool, for example, could ask whether a photo meets a defined condition and return a score without requesting a written explanation. That is a proposed application, not a result from our own test.
The license deserves attention before deployment. The d1-3B release uses the LFM Open License v1.0. Liquid’s license guide describes a commercial revenue threshold of $10 million and separate licensing above it. Hosted d1 and the public checkpoints are separate deployment options, so family membership alone does not establish identical results.
Microsoft-Decision-1
Microsoft-Decision-1 is available in Microsoft Foundry public preview. Its deployment documentation describes typed judgments over text or JSON: yes/no probabilities, categorical choices, and ordered scores.
The integration uses Foundry deployment and authentication. The request identifies the deployed model, so adopting it involves more than changing the hostname in a Jev client.
Microsoft describes internal applications involving feedback, quality checks, and incident-related workflows. These are vendor-reported uses, rather than independent demonstrations of accuracy on a reader’s workload.
Perplexity pplx-decider
Perplexity’s Decisions API accepts text, JSON, and images in a shared state, with named noul (yes/no), choice, and score questions. It provides an API route for classification, visual judgments, and ordered ratings without a generated explanation.
Perplexity also publishes pplx-decider-v1.1-27b under Apache 2.0. This is a substantial local deployment: the model card lists roughly 49 GiB for the backbone weights in BF16, a 16-bit number format, before runtime memory. The release includes a separate decision head, the output component that turns the model’s representation into typed answers.
Inception Mercury Decide
Mercury Decide has a documented API for classification, routing, and scoring. Requests go to /v1/decisions with the model mercury-decide, shared state, and a map of named questions. The response preserves those names and returns typed answers with probabilities.
Its interface follows the shared-state, named-question pattern used by System One APIs. For example, one request can ask whether a message concerns billing, which team should receive it, and how urgent it is. The questions are evaluated independently against the supplied state.
Mercury Decide does not support chat completions, streaming, or tool calling through this endpoint. Applications still own the action that follows a decision. That makes it a candidate for an existing decision step, rather than a replacement for an agent’s entire conversation loop.
Upstage Solar Decide
Solar Decide is Upstage’s beta decision endpoint on Solar Mini 4. It exposes choice, yes/no, and score outputs and documents a 512K context window.
The System One API reference accepts a string or JSON state and questions keyed by name. A long policy document can therefore form the shared context for several decisions. The advertised window is an input allowance; it does not, by itself, demonstrate that a decision remains accurate when the relevant paragraph is buried near the end.
Upstage marks the endpoint as beta and warns that the schema may change before general availability.
Fastino: GLiDE and GLiNER2.5-Decide serve different needs
GLiDE is Fastino’s hosted decision model. It supports named yes/no, choice, and score questions, with additional reasoning for uncertain decisions. Its documentation says input usage sums internal passes across questions. That matters when estimating the cost of a large state with many questions.
GLiNER2.5-Decide is a compact English text classifier with Apache 2.0 weights. Its classification contract supports caller-defined labels, multiple classification tasks, and multi-label results. This makes it relevant to routing, tagging, or moderation pipelines that need local inference.
The distinction affects implementation. GLiNER’s classification labels do not automatically become an ordered score rubric, and its classic classification interface does not enforce relationships between tasks. GLiDE exposes a different decision contract.
Local models and training-oriented alternatives
H2O-Lightning-4B
H2O-Lightning-4B is an Apache-licensed model built on Qwen3.5-4B, with text and image decision support. Its release includes weights, serving configuration, examples, and a shim, or translation layer, in front of the vLLM serving engine.
The published shim matters as much as the checkpoint. It requests candidate-token probabilities, applies the configured temperature, and constructs typed answers. It sends one completion per question, so “one token” should not be read as one identical unit of computation for an entire multi-question request.
Kev
Kev is a family of local decision models with a System One-style API, pretrained releases, and a fine-tuning workflow. The project provides 0.8B, 4B, 9B, and 27B variants, along with serving support for CUDA and Apple Silicon through MLX.
That range is useful when the deployment constraint is concrete: a small local service, a Mac-based application, or a larger GPU installation. The Kev-4B model card distinguishes contexts the runtime accepts from contexts for which the authors validated accuracy. Those are different selection criteria.
Kev is an independently trained implementation; API resemblance does not make it the original Jev model.
Laya
Laya takes a smaller text-model approach, with English, multilingual, and typed-decision checkpoints and a router that selects between them. It publishes code, weights, benchmarks, examples, and fine-tuning material under an Apache-licensed project.
Its English and multilingual checkpoints cover intake classification and local text processing. The project’s own results also make domain adaptation an important part of the story: a checkpoint trained for typed decisions and a general English checkpoint should not be treated as interchangeable.
The multilingual quickstart documents a default limit of 1,024 tokens. For longer documents, it specifies model="multilingual" and max_len=8192. That raises the input limit to 8,192 tokens; accuracy on long documents still needs checking against the intended workload.
Strands Decider
Strands Decider provides code for training, evaluating, and serving small models, plus public checkpoints such as the Qwen3.5-based 2B release. Its examples cover choice, yes/no, and score questions, with a local HTTP server as well as a command-line interface.
This is useful when a team wants to inspect and adapt the decision-model pipeline, not just consume an endpoint. The repository includes tests and multiple runtime paths and belongs to Strands Labs, the community’s experimental arm. That experimental status is part of its deployment context.
Check the exact code revision and checkpoint for the intended device. At the review date, the README distinguished released installation paths from an MLX extra that required installing from the repository. The Qwen 2B release documents image inputs through --vision and publishes image evaluations for that checkpoint. Its card also reports confident answers when images are missing, a relevant failure case for a visual workflow.
Together Tev1
Together’s Tev1 experiment is particularly useful for understanding the training route. The public repository includes data preparation, training and evaluation material for a Qwen3.5-4B-based decision model.
Its output contract is narrower than Jev’s. The Tev1 model card describes a state, a question, and 2–24 labelled options; the model returns a letter that application code maps back to an answer. It retains a standard next-token language-model head and uses the Chat Completions interface. The repository explicitly distinguishes log probabilities from calibrated confidence.
Treat it as an experimental choice model and a training reference. The code’s MIT license does not settle the terms of every artifact: at the review date, the card said the fine-tuned weights’ release license was still being finalized. That should be resolved before a commercial deployment.
Bespoke Nimble
Bespoke Nimble combines local typed decisions with a published data-curation and evaluation approach. The reference interface handles flat boolean and enum schemas over text, with scoring code that converts allowed-answer logits into probabilities. Its score_fields option adds expected ordinal values for designated enum fields.
Its 9B v2 release is a LoRA adapter, so deploying it requires the specified base model as well as the adapter. The release documents its prompt contract and probability temperature. Loading the adapter alone does not reproduce all of that behavior.
Nimble exposes the training and scoring components of a decision system. The repository distinguishes how its Mac and CUDA scorers handle shared context, and the v2 card explains that its temperature default was transferred from an earlier checkpoint. Those details are more useful for planning an experiment than treating “local Jev” as a complete technical specification.
A narrower alternative for agent traces
Respan span-01
If the job is to evaluate agent behavior, Respan’s span-01 has a dedicated trace-scoring interface. It reads a trace and caller-defined behaviors, returning probabilities for present, absent, and not observable. The third outcome is useful: a transcript that stops after an agent offers a refund cannot establish whether the customer accepted it.
The span-01 quickstart uses a dedicated scoring endpoint. This is a task-specific alternative for trace evaluation, rather than a general replacement for every Jev choice or score request. Its value should be judged on the behaviors an application needs to detect.
What to compare before switching
Start with the decisions, not the model names. Take a labelled sample from one workflow, including ambiguous inputs and cases with missing evidence. Use development cases to choose prompts, checkpoints, and thresholds. Reserve a separate holdout for the final error-and-coverage comparison. This is our suggested evaluation design, rather than a benchmark result.
Measure four things together:
- Wrong automatic decisions. Count errors among the cases the system actually acts on. A high overall accuracy can conceal a costly failure category.
- Coverage at the chosen threshold. Record how much work the model completes and how much it sends to a person or another model. Compare candidates at a similar error tolerance.
- End-to-end latency. Include networking, image preparation, queueing, retries, and fallback calls. Report tail latency as well as the median.
- Cost per completed workflow. Include repeated context, multiple questions, internal reasoning where applicable, and the cost of running local hardware.
Then check the contract. What happens if none of the options fits? Are questions independent? Does an ordered score return a level, an expected value, or both? Does an oversized input fail or get truncated? The answers can change application behavior even when two providers accept similarly named fields.
Questions about switching
Can I switch by changing the API URL?
Do not assume so. OpenAI uses input and an array of questions; Jev uses state and a map of named questions. Tev1 uses Chat Completions with labelled choices. Even similar request fields can produce different response shapes or confidence values, so the adapter is part of the migration.
Which alternatives can run locally?
Clef, Liquid’s public d1 checkpoints, pplx-decider, GLiNER2.5-Decide, H2O-Lightning, Kev, Laya, Strands Decider, and Nimble provide public model artifacts and local inference instructions. Their hardware needs and licenses differ considerably. Start with the specific checkpoint and runtime, rather than a family name.
Is a structured-output LLM also an alternative?
For some workflows, yes. A generative model with Structured Outputs can be a better fit when the same call must extract a more complex object or produce an explanation. Schema conformance still does not establish that the values are correct. Keep the existing implementation in the evaluation as a baseline.
Should confidence thresholds carry over from Jev?
Re-evaluate them. For example, GLiDE documents choice confidence as the gap between the two leading probabilities. An application that interprets every field named confidence as the probability of being correct would already be making the wrong assumption. Select thresholds against labelled cases and the consequences of a wrong decision.
Keep the current workflow as the baseline and compare it with one candidate that meets the actual requirements. Add a local model if data control, offline operation, or available hardware makes local serving relevant. Staying with the current implementation is also a valid result.

