TypeSafe released Jev on 15 September 2026. It evaluates developer-defined questions and returns choices, scores or probabilities for application code to use. Two weeks later, public projects are using it to select browser actions, classify documents, find code and manage agent context.
How the API works
A request supplies context and one or more questions: choose a category, assign a score on a defined scale, or estimate whether a statement is true. The developer specifies the categories or scoring criteria.
This support-routing request asks Jev to choose a team for a customer who reports being charged twice:
curl --fail-with-body https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "jev-1.13.0",
"state": "The same invoice was charged twice. Please check the second payment.",
"questions": {
"department": {
"type": "choice",
"instructions": "Choose the team that should review this request.",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Software fault or broken feature",
"sales": "Inquiry about a new purchase",
"other": "None of these teams clearly fits"
}
}
}
}'
Illustrative output, not a live API result. The invented probabilities below demonstrate selected fields from answers.department:
{
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.94,
"technical": 0.02,
"sales": 0.01,
"other": 0.03
}
}
Application code uses the selected category, billing, to route the message.
Choice answers also include a confidence field, omitted from this excerpt, that summarises the probability distribution. A routing application can use it to send uncertain cases for human review. The cutoff needs testing on the application’s own data.
Pydantic CTO David Montague describes the reason for such compact outputs in his article on Jev evaluators: “Writing an explanation can be useful, but you don't always need one for every production response.”
Ten early uses
The following examples draw on project documentation and authors’ reports available on 29 September 2026.
1. Browser actions — Browser Use and Stagehand
Browser Use’s prototype asks Jev to select an operation and a page element, then uses a language model when text is required. Browserbase’s Stagehand implementation selects among candidate actions with an LLM fallback. Browserbase’s Kyle Jeong distinguishes these hybrid systems from standalone Jev agents, writing of the latter that “even the best demos aren’t ready to be deployed to production.”
2. Document classification and splitting — DocJev
Jerry Liu’s DocJev extracts page text, then uses Jev to classify pages and locate document boundaries within a packet. Its evaluation report measures text extraction separately from Jev calls and notes that the labels used to judge accuracy did not receive human review.
3. Agent evaluation — LangSmith
LangSmith’s Jev integration evaluates agent responses against a rubric. LangChain’s experiment used five fixed weather-agent responses, each scored 100 times for quality and pass/fail against a human reviewer’s reference labels. The repeated runs measure consistency; the five distinct responses define the scope of the accuracy comparison.
4. Semantic PostgreSQL queries — pg-jev
pg-jev adds natural-language conditions to SQL—for example, filtering support tickets by whether a customer wants to cancel. It sends row data to Jev over HTTPS, batches requests and caches answers within the database session.
5. Agent memory selection — fast-jev-compaction
This Claude Code plugin frees context space by retaining, truncating or removing earlier tool calls and results. Retained text stays verbatim. To make those decisions, Jev receives conversation text and tool arguments, with tool-result contents replaced by status and length notes; additional input is shortened when needed to fit the request.
6. Code retrieval — jevgrep
jevgrep accepts a question about code behaviour and uses Jev to judge the relevance of folders, files and declarations. It returns verbatim source excerpts with line references for a coding agent to inspect. The tool targets searches where the agent knows what the code should do but has not found the responsible files.
7. Semantic HTTP routing — hono-jev-router
In this experimental Hono router, developers describe requests in natural language and attach a handler to each description. Jev evaluates whether an incoming request matches each description; the router dispatches to the first one above the configured probability threshold. The author explicitly excludes authentication and authorisation from its intended use.
8. Support routing — Entagl
Entagl replayed 402 routing decisions from two client workspaces. Against its reviewed answer key, Jev performed better at selecting an action and worse at choosing a specialist than the original systems. The authors associate many specialist errors with missing context: Jev saw recent messages but was not told which specialist already owned the conversation. One reviewer judged disagreements, ambiguous cases were excluded and rare outcomes oversampled.
9. Archive screening — hn-oracle
Anthony Maio’s hn-oracle uses Jev to select Hacker News comments that contain predictions for subsequent analysis. The repository reports a pilot and projects the cost of processing the full archive.
10. Writing feedback — Undertone
Undertone’s README describes a writing interface that sends eight questions about a draft in one Jev call, covering qualities including tone, urgency and formality. Application code converts the answers into colour, labels and phrase-level tips displayed beside the text.
What TypeSafe wants to build
TypeSafe’s manifesto proposes software built from callable AI decisions, with code handling exact computation. Its Bitterest Lesson essay argues that choosing the right task to train for matters more than improving algorithms or adding compute. In his Latent Space interview, CEO Diogo Almeida identifies intelligence per dollar as a priority.
Choosing a first trial
DRL would start with one existing classification or routing step whose correct outcomes can be labelled. Replay the same inputs through Jev and the current system, supplying the same context to both. Keep arithmetic and date comparisons in code: TypeSafe lists these among Jev’s known weaknesses.
Measure cost and latency for the complete task, including fallback calls. Record incorrect automatic decisions separately from cases sent for review. That distinction shows whether a proposed confidence threshold reduces errors or simply shifts more work to the fallback.

