What Is Jev? TypeSafe's Decision Model, Explained
Jev is TypeSafe AI's first System One model: it answers typed questions about your data with probabilities instead of writing text. Input costs $0.042 per million tokens, output is free, and replies arrive in roughly 70 to 500 milliseconds.
Most AI models are built to write. Jev is built to decide.
Released in September 2026 by TypeSafe AI, a San Francisco lab, Jev answers typed questions about data you send it and returns a choice, a score or a probability, with no sentences attached. It is not a chatbot and it will not write your email. It exists for the moment in software where a program needs a judgment call and currently has to ask a language model for one, then parse the reply and hope.
One disambiguation first, because search results mix them up: this article is about the AI model. JEV is also the abbreviation for Japanese encephalitis virus, which is a mosquito-borne disease and has nothing to do with any of this.
The problem Jev is built for
TypeSafe's documentation states the case plainly: "Large language models (LLMs) are designed to produce text for humans to read. When you need a model to make a judgment that your code will consume, that creates a mismatch."
Anyone who has shipped an AI feature knows the shape of that mismatch. You want to know which team a support ticket belongs to. You ask an LLM, instruct it firmly to reply with only one word, and it returns billing nine times out of ten and Based on the content, this appears to be a billing issue. on the tenth. So you write a parser, then a retry, then a fallback, and the feature spends most of its code defending against the model's desire to talk.
Jev removes the text entirely. You define the answer space; it picks from it.
What it actually returns
TypeSafe calls them primitives, and there are three. Each asks a different kind of question:
| Question type | What it asks | What comes back |
|---|---|---|
| Choice | Pick one option from a list you define | The choice, a probability for every option, and a confidence value |
| Score | Rate the input against ordered levels you describe | A score, probabilities across the levels, and confidence |
| Noul | Is this statement true? | A single number from 0 to 1 |
A real response from the docs, for a support message about a failing Stripe integration, looks like this:
{
"department": { "choice": "technical", "confidence": 0.78,
"probabilities": { "technical": 0.85, "billing": 0.15, "sales": 0.0 } },
"frustration": { "score": 1.0, "confidence": 1.0 },
"is_urgent": { "noul": 1.0 }
}
That is the whole output. No prose, nothing to parse, nothing that might arrive in a different shape tomorrow.
Note the asymmetry: Choice and Score carry a confidence value, Noul returns only its probability. It matters when you design around the response.
Probabilities are the point, not a bonus
The number beside each answer is what separates this from a classifier you could have trained yourself.
TypeSafe says System One models "are trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty." Calibrated means that when the model says 80%, it should be right about 80% of the time across many such predictions.
The docs are also honest about the limit of that, in a line more vendors should copy: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct."
So a probability is a statement about the long run, not a guarantee about the ticket in front of you. Used properly, it lets your code do something an LLM's confident paragraph never could: act automatically above a threshold, send the middle band to a human, and escalate the rest.
Confidence is a second, different signal
Confidence is not the same as probability, and conflating them is the easiest mistake to make here.
Probability says how likely each option is. Confidence describes the shape of that distribution: whether the model has settled on one answer or is split. A Choice spread evenly across three options has a clear winner by a nose and almost no confidence; one option at 0.85 has both.
For Score questions it is subtler, because the levels are ordered. Probability split between "frustrated" and "very frustrated" is a narrow disagreement; split between "calm" and "very angry" is the model telling you it has no idea. Confidence captures that distance.
The practical pattern TypeSafe documents is confidence-gated routing: the answer tells you what, and confidence tells you whether to act.
Why it is quick and cheap
The numbers are the reason people are paying attention.
Input costs $0.042 per million tokens, and output is free, per TypeSafe's models page. End-to-end latency runs roughly 70 to 500 milliseconds, with independent user reports putting the median near 76ms.
TypeSafe's own benchmark claims Jev is 193.6 times faster and 444.6 times cheaper than GPT-5.6 Terra on one workflow. Treat that with the scepticism any vendor benchmark deserves: it is the company's own evaluation, on tasks it chose, and it represents the favourable end. The honest summary is that for narrow classification work, Jev is in a different cost and latency class from a frontier LLM, which is unsurprising, because it is doing a far smaller job.
The architecture helps too. Every question in a request is evaluated in parallel and in isolation against the same state, so asking twelve questions costs barely more time than asking one. TypeSafe's own cookbook on batching reports that putting 13 questions into a single call was 12.2 times cheaper and 10 times faster than asking separately, with no change in the answers.
What Jev cannot do
This list matters more than the feature list, because the failure mode here is buying the wrong kind of tool.
- It writes nothing. No replies, no summaries, no code, no explanation of its reasoning.
- Text only. The docs say it evaluates strings, JSON objects and arrays of text. No images, audio or video.
- The answer space is yours to define. It cannot discover a category you did not list. If your options are billing, technical and sales, every ticket becomes one of those, including the one about a partnership.
- Limited context. TypeSafe lists 64k tokens per request, with 32k for the state plus the longest question. Cloudflare's listing shows a 32,000-token context window. Either way, this is not a model for reasoning across a 300-page contract.
- No reasoning chain. If you need the model to work through several dependent steps, decompose the problem into separate questions and combine them in your code, which is TypeSafe's stated recommendation anyway.
Where the name comes from
"System One" is borrowed from Daniel Kahneman's Thinking, Fast and Slow, where System 1 is fast, intuitive judgment and System 2 is slow, deliberate reasoning. TypeSafe's framing is that most AI products currently use a System 2 tool for System 1 work: a large reasoning model asked to decide whether an email is spam.
It is a useful mental model even if you never use the product. Look at the AI calls in your own application and ask how many are judgments rather than compositions. For most teams it is the majority.
Who is behind it
TypeSafe AI is a San Francisco lab founded by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida is a former OpenAI researcher. Reports put the company's funding at around $40 million.
Launch was busy: Jev reached general availability in late September 2026, and demand was high enough that signups were briefly paused before reopening. TypeSafe's homepage now says there is no waitlist.
Jev is not the only one doing this
A second product appeared in the same category at almost the same moment. Perplexity's Decisions API answers yes/no, multiple-choice and rubric questions with probabilities rather than prose, using a model called pplx-decider-v1-27b, priced at $0.04 per million input tokens with free output.
Two independent launches at nearly identical prices, weeks apart, is the clearest signal that "decision model" is becoming its own product category rather than a single company's idea. For buyers that is good news, because it means a second source and price pressure.
Should you care?
You should look at it if your application already calls an LLM to classify, route, score, moderate, rank or extract, and you are paying frontier prices and waiting seconds for a one-word answer.
You can ignore it if you use AI to write, summarise, chat or generate. Jev does none of that, and it is not competing for that work. It sits alongside those tools, often in front of them, deciding which request is worth sending to the expensive model at all.
You should be cautious if your decisions are high-stakes and you were hoping probabilities would remove the need for human review. Calibration is a property of a batch, not an individual answer, and TypeSafe says so itself.
If you have decided it fits, our guide to how to use Jev covers the API key, the first request and the gateways it is available through. Our AI coding tools ranking covers the assistants developers use day to day, and the AI chatbots list covers the general-purpose models Jev is designed to work with rather than replace.
FAQ
What is Jev in simple terms?
Jev is an AI model that answers questions about your data with a choice, a score or a probability instead of writing text. Software sends it a piece of state and some typed questions, and gets back values the code can branch on directly.
Is Jev a large language model?
No. TypeSafe describes it as transformer-based but not an LLM, because it does not generate text. It returns typed decisions with probabilities, and it cannot write replies, code or explanations.
How much does Jev cost?
Input costs $0.042 per million tokens and output tokens are free, according to TypeSafe's models page. That is roughly $42 per billion input tokens.
How fast is Jev?
TypeSafe reports end-to-end latency of about 70 to 500 milliseconds, and independent user reports put the median near 76 milliseconds. Asking several questions in one request barely increases the time, because they are evaluated in parallel.
Is JEV the same as Japanese encephalitis virus?
No. JEV is also the standard abbreviation for Japanese encephalitis virus, a mosquito-borne disease. Jev the AI model from TypeSafe is unrelated, and most search results for the bare acronym refer to the virus.
Does Jev replace ChatGPT or Claude?
No. It does a different job. Chat assistants write and reason; Jev decides. A common pattern is to use Jev to triage or filter, then send only the cases that need it to a larger model.
Related tools
ChatGPT
The most widely used AI chatbot for writing, coding, and research.
Claude
Anthropic's AI assistant known for careful reasoning and long context.
Gemini
Google's multimodal AI assistant integrated across Search, Workspace, and Android.
Perplexity
AI search engine that answers questions with cited sources.