Table of contents
Key takeaways
- Jev is a System One model from TypeSafe AI that returns typed decisions, not generated text.
- Three primitives you can request: Choice picks an option, Score rates a rubric, Noul answers yes or no.
- One request carries many questions, evaluated in parallel, so response time barely moves.
- TypeSafe trains Jev with RLCD, so higher confidence should mean higher accuracy.
- A decision model works as one layer, with a generative model still handling the rest.
What is Jev?
Jev is a non-generative AI model that takes an input state plus one or more typed questions and returns structured answers with per-option probabilities and confidence scores. The state can be a string, structured data such as a JSON object, or an array of messages. The questions are declared explicitly, with the answer space defined up front. The model evaluates that state against each question and returns a typed value your code consumes without any parsing step.
Traditional LLMs work differently. They return a string, and turning that string into a decision means prompting for a specific format, parsing the result, and handling the cases where the format drifts. Jev skips that path altogether. TypeSafe frames the problem as a mismatch: language models are built to produce text for humans, so using one to make a judgment that software will act on introduces a translation layer where type errors live.
Who makes Jev?
Jev is made by TypeSafe AI, a San Francisco company founded in 2024, led by cofounder and CEO Diogo Almeida, previously a researcher at OpenAI and a listed author on the InstructGPT paper, alongside cofounders Erik Gafni and Sasha Sheng. Public reporting puts the company's seed round at $40 million led by DCVC, at a reported valuation of around $200 million.
Jev came out of stealth in September 2026. After an initial early access period, TypeSafe removed its waitlist, and the current production version is Jev-1.13.0, with Jev-latest available as a rolling alias.
What is a System One model?
A System One model is an AI model that evaluates an input state and returns a typed decision with a probability in a single pass, without producing a reasoning trace. TypeSafe named this model class after System 1, the fast, intuitive mode of thinking Daniel Kahneman describes in Thinking, Fast and Slow.
- System 1 is fast, automatic, and intuitive. It recognizes a face, detects urgency in a message, or spots that something looks wrong, and it returns a judgment without an explanation.
- System 2 is slow and deliberate. It works through a proof or weighs a decision step by step, and it handles problems where the answer must be constructed rather than recognized.
Jev is TypeSafe's System One model, built for the high-volume operational decisions that production software makes continuously: routing a ticket to a queue, scoring a request for risk, classifying urgency, and selecting which model handles a task. These decisions require a typed output the next line of code can consume directly, not a chain of thought. Reasoning models suit System 2 tasks, but on routine operational decisions they add latency and token cost for reasoning the task does not require.
How does Jev work?
The Jev request model is deliberately narrow. You post a state and a map of typed questions to a single decision endpoint, naming each question yourself, and Jev returns an answer object keyed the same way, plus token usage. Before looking at how a single call behaves, it helps to see what you can actually ask.
The three primitives you can request
TypeSafe exposes three AI primitives. Each asks a different type of question and returns a different type of answer, and all three are composable in the same call.
- Choice: select one option from a defined set. You supply an option-to-description map, and Jev picks one option and returns it with per-option probabilities and a confidence score.
- Score: rate the state against ordered levels, such as a two-to-ten-point rubric. It returns a numeric score, the underlying distribution, and confidence.
- Noul: evaluate a yes or no question. It returns a single probability between 0 and 1, with no separate confidence field.
Question type | Goal | Returns |
|---|---|---|
Choice | Choose one option from a list | choice, probabilities, confidence |
Score | Score the state on a rubric | score, probabilities, confidence |
Noul | Is this statement true? | noul (0-1) |
None of this makes one model better than the other. It makes them suited to different points in
Each question carries its own instructions and criteria, so you define the task in natural language while the answer stays bounded by the type. The full schema, field-level details, and code examples are in the TypeSafe Jev docs.
Parallel multi-question evaluation
Every question in a request is evaluated in parallel and in isolation against the same state, in one pass. According to TypeSafe, adding questions barely changes the response time, and each extra question only adds its own cheap token cost. Input is priced at $0.042 per million input tokens, with output tokens free, subject to the rate limits and context window published on TypeSafe's pricing page.
Isolation matters as much as parallelism. Because each question sees the state independently, a long list of questions doesn't degrade the way a single crowded prompt does. There's no context rot from stacking twelve instructions into one request. It also means questions can't influence each other, which is what you want when each one drives a different branch in your code.
The practical pattern is this: instead of one prompt asking a model to return a JSON object with eight fields, you send one request with eight typed questions and get eight independently evaluated answers back.
Calibration and confidence scores
Calibration is what makes a probability worth acting on: outcomes assigned 0.2 should come true roughly 20% of the time, so higher confidence means higher accuracy in aggregate. TypeSafe calls its training method reinforcement learning for calibrated decisions, or RLCD, alongside RLHF, which produced chat assistants, and reinforcement learning with verifiable rewards, which produced reasoning models. No paper on RLCD has been published, so treat the mechanism as vendor-stated.
Confidence is a separate statistic derived from that distribution: flat yields low confidence, peaked yields high. Probability tells you how likely an answer is; confidence tells you how decisive the distribution behind it was. The usual pattern is three-tier gating: act above a high threshold, confirm in the middle band, route low-confidence cases to a person.
How to use Jev in a multi-model setup
In a multi-model setup, Jev handles fast structured decisions while a frontier LLM handles reasoning, drafting, and code generation. Connecting each model directly requires separate provider accounts, billing relationships, API credentials, and failure handling, and every additional model adds another set of each.
nexos.ai, as an all-in-one AI platform, replaces separate provider accounts, billing relationships, and API credentials with a single layer between your application and every model provider. Jev and frontier LLMs run through one interface, sharing the same API keys, access policies, and request logs. That shared layer gives teams three practical advantages:
- Task-fit routing. nexos.ai sends each request to the model best suited to the task. Classification, urgency flagging, and pre-execution checks are System 1 tasks, and routing them to Jev reduces latency and per-token cost compared with a frontier LLM. Code refactoring and multi-step debugging stay on the frontier LLM.
- Swappable models without code changes. Model rankings for specific tasks shift with each provider release. A model-agnostic AI setup keeps application code stable while the underlying model changes, so adopting a new model version is a configuration change rather than a refactor. The nexos.ai AI Gateway exposes one endpoint across all supported providers and models, which makes that configuration change possible.
- Cost attribution and centralized governance. Per-token costs compound as request volume grows. nexos.ai breaks down usage by team, project, and API key, which shows which workloads drive spend. nexos.ai applies access rules, model permissions, and budget limits once at the gateway instead of in each application.
The nexos.ai System one guide documents a complete Jev decision call, including the format of typed questions and returned probabilities when Jev runs through the nexos.ai AI Gateway alongside other models. Teams evaluating Jev can use the guide to test the integration against their existing stack before deployment.
Keeping spend visible across teams
Cheap models create a specific blind spot. When a call costs a fraction of a cent, nobody instruments it, and then an agent loop fires it fifty thousand times a day and the line item stops being small. The risk isn't the unit price. It's volume nobody is watching.
Useful visibility means spend broken down by model, team, project, application, and individual developer, in something closer to real time than a monthly statement. That's the difference between finding out at the invoice and catching it the same afternoon. LLM monitoring at the gateway layer captures request-level detail across every model in the mix, so a decision model and a frontier model show up in the same view rather than in two separate dashboards.
Swapping and falling back between models
Any external model is a dependency that can rate-limit you, slow down, or go offline. If a decision model sits in the critical path of a routing step, its availability becomes your availability.
Fallback chains handle this at the infrastructure layer instead of in application code: define a primary model and an ordered list of alternates, and failed or timed-out requests reroute automatically. Failures become transparent reroutes rather than incidents.
The same mechanism covers planned changes, such as pinning a specific version like Jev-1.13.0 while you validate a newer one, then promoting it once your thresholds still hold. Version aliases are convenient, but pin the exact version when your code depends on tuned confidence thresholds, because a silent upgrade can shift the distribution underneath them. More on this in AI for developers.

Run every AI model in one Gateway
One secure, lightweight layer to access, route, and manage 200+ models.

Monitor AI usage

Control AI spend

Optimize AI costs
What can you build with Jev?
The realistic use cases for a decision model are the small judgment calls buried inside larger workflows, the points where a system needs one decision before it can continue. Here are three patterns that show the shape clearly.
Support ticket triage
A ticket arrives as unstructured text. In one request you can ask: which department should handle this (Choice), how urgent is it on a five-level rubric (Score), how frustrated does the customer sound (Score), and is this a billing dispute (Noul). Four typed answers come back, each with probabilities, and ordinary code routes the ticket from there. No prompt engineering for output format, no parsing, no retry when the model returns prose instead of JSON.
Model routing inside an agent loop
A decision model can sit in front of a generative model and assess the incoming request: is this boilerplate or a hard reasoning task, does it need tool access, how much context does it require. The answer selects the model, and the probabilities and confidence stay available downstream so the loop can escalate a borderline case to a stronger model rather than guessing. This is a common pattern when comparing where each model earns its cost.
Tool-call and action checks
Before an agent executes a proposed action, a lightweight classification step can evaluate it: does this action match what the user asked for, is it reversible, does it fall inside the permitted scope. The result is a typed answer with a confidence value, which the surrounding code uses to proceed, ask for confirmation, or hand off to a person.
Treat this as one input to a decision rather than a guarantee. The model reads the state as data, and a check like this belongs alongside conventional permission logic rather than in place of it. For more patterns, see AI agent use cases.
Jev vs a general LLM for classification
Both Jev and general LLM can classify text, and read natural-language instructions to do it. The difference is what comes back. A general LLM returns text that you prompt into a format and then parse, while Jev returns a typed decision with probabilities and a confidence value already attached. That changes the cost, the latency, and the amount of code around the call. And that trade sharpens across five dimensions.
Dimension | Jev (System One) | General LLM |
|---|---|---|
Output | Typed decision + probabilities | Free text / structured text |
Primary use | Fast classification | Reasoning, generation |
Latency | Very low (vendor-claimed) | Higher |
Cost per call | Lower (vendor-claimed) | Higher |
Confidence scores | Native | Not native |
None of this makes one model better than the other. It makes them suited to different points in the same workflow, which is the argument for running both and letting each take the work it fits. Testing that on your own tasks is the only way to see where the line falls for you, and it's worth putting the two side by side to compare AI models on the requests your system actually handles.
PRO TIP: Reach for Jev-style structured decisions when you need a fast yes/no, a label, or a score you can act on in code, and keep a general LLM for the reasoning and generation steps.
When to use Jev and when not to
A decision model is excellent inside its range and poor outside it, and the boundary is sharp enough to be worth stating plainly. Start with the work it's built for.
Tasks that fit Jev
- Routing decisions: picking a queue, a department, a downstream model, or a branch.
- Urgency and priority flags: ordinal judgments about time sensitivity.
- Content classification: labeling text against a defined set of categories.
- Ordinal scoring: rating a state against a rubric with fixed levels.
- Pre-execution action checks: evaluating a proposed action before it runs.
The pattern across all five is a bounded answer space, fuzzy criteria that resist hard-coded rules, and a result that code acts on immediately. But some tasks still need a general model.
Tasks that need a general AI model
- Writing code: Jev isn't trained to generate, and chaining choices to fake it works poorly.
- Drafting content: any open-ended text output.
- Multi-step reasoning: chains of inference and levels of indirection.
- Open-ended generation: anything without a declared answer space.
Two more limits are worth knowing, because TypeSafe documents them directly. Jev is not a calculator: it doesn't count reliably, and it reads dates as text rather than as ordered quantities, so arithmetic, counting, and date comparison belong in your own code. And accuracy drops as irrelevant content grows in the state, so keep the input tight.
Not every request needs a frontier model, and not every request suits a structured-decision model. Matching the task to the class of model that fits it is the actual skill.
Run every AI model in one Gateway
One secure, lightweight layer to access, route, and manage 200+ models.
Run every AI model in one Gateway
One secure, lightweight layer to access, route, and manage 200+ models.