On September 15, 2026, TypeSafe AI came out of stealth with forty million dollars led by DCVC and a model called Jev. The founder is Diogo Almeida, who co-authored the InstructGPT paper. TypeSafe's own materials describe him as a co-inventor of the training method behind ChatGPT, which is the company's framing; reinforcement learning from human feedback predates his involvement.
Jev doesn't write. You hand it some messy state, an email, a support ticket, a blob of JSON, along with questions you've defined in advance. It returns typed answers. Every answer carries a calibrated confidence number. There's no prose to parse and no way for it to answer outside the options you declared. The state is read once and every question runs against it at the same time, so ten questions about one ticket cost roughly what one question costs.
A choice
One option from a set you supplied. Nothing outside the list can come back.
A score
A position on a rubric you wrote. Your scale, your anchors.
A yes-no
A binary answer with a probability attached, so you know how sure it is.
TypeSafe calls the category System One models, after the fast intuitive half of Kahneman's split. Sébastien Dubois has the better description: a semantic if-statement you can afford to call thousands of times.
It costs four point two cents per million input tokens. Output tokens are free.
The name is the tell. Jev is short for William Stanley Jevons, the economist behind the paradox where more efficient steam engines caused England to burn more coal rather than less. Make a thing cheap enough and people do vastly more of it.
They were right about that. They were also standing inside their own paradox, and the clock started faster than anyone expected.
Seventy-two hours
On September 18, Convai Innovations published Laya on Hugging Face under Apache 2.0. Three checkpoints, 421 million parameters for the English model, weights and training method and fine-tuning notebook all public. Same shape of interface: state in, typed questions, calibrated probabilities, no output tokens.
It was also faster. Laya reports median latency around 33 milliseconds on a Tesla T4 against Jev's 236 to 276. An independent head-to-head run on byte-identical inputs put Jev ahead on triage, guardrails, moderation and multilingual intent, with Laya ahead on a couple of classification suites. Quality is genuinely split. Speed isn't.
Jev
236 to 276 ms median latency. Closed weights, gated API, forty million in funding.
Laya
Around 33 ms on a Tesla T4. Apache 2.0, 421M parameters, published three days later.
Quality
Split. Jev leads on triage, guardrails, moderation and intent. Laya leads on some classification.
Seventy-two hours isn't enough time to train a model from a blog post. Convai's Nandakishor M has said publicly that he'd already built along these lines, watched the idea arrive as a well-funded breakthrough, and chose to rebuild it properly rather than argue about it. That's his account rather than something I can verify independently, but it's the only thing that explains the timeline.
Then came everyone else. The curated awesome-jev list now carries fourteen engines and five runtimes; the GitHub topic shows around thirty repositories. SemIf takes an open-weights model you already have and reads the raw scores assigned to each permitted answer, training nothing. Kev ships fine-tuned open models behind a Jev-compatible API. Jevlike is small enough to train on a CPU. LangChain shipped an integration. Zapier added a block.
Why it was so easy to copy
Sean Goedecke published the technical case on September 16, the day after launch, and it's held up.
His argument: most of the speed comes from inference strategy rather than a new architecture. Prefill the response so the only thing left to produce is a single token, constrain that token to the options the user supplied, and read the probabilities off the logits. One forward pass, a probability per option, nothing to parse. He tested the approach on Qwen2.5-1.5B-Instruct and got a two to three times speedup over ordinary structured output. His read is that there may be no substantial technical moat here, and that a model with no test-time compute will cap out around the strength of non-reasoning LLMs.
He also took apart the hallucination claim. A model that can only emit legal options still picks wrong options. Constraining the shape of an answer says nothing about its content, which is the same thing OpenAI's structured-output documentation has always said about schema adherence.
Constraining the shape of an answer says nothing about its content.
The underlying fact is plainer still. Every language model, on every token, computes a score for every possible next word before picking one. Those scores have always been there. Reading them off a constrained set is a decision to stop throwing information away.
TypeSafe built more than a weekend hack. Their model is trained so the confidence numbers are honest, which is harder than it sounds and is the part nobody reproduced in three days. But the category they created, the thing their launch taught thousands of engineers to want, turned out to be reachable from weights people already had.
Eleven days: the student catches the teacher
On September 26, AutoTrust AI published JEV-27B under Apache 2.0. Weights, training code, serving code, evaluation reports, all public. A 108.9-million-parameter decision block sitting on a frozen Qwen3.8-27B backbone, about 0.4 percent of the model, trained in roughly 9.2 B200-hours.
In AutoTrust's own runs of both models across six public decision benchmarks, JEV-27B averages 84.07 against Jev's 83.85. Those are self-reported numbers from the party with an interest, and the gap is 0.22 points, which is noise. Treat it as parity rather than victory.
Parity is the point. What makes it worth the section is how they got there.
JEV-27B was distilled from Jev itself. About two thirds of its training corpus, roughly 498,000 of 741,000 rows, carries Jev 1.13's own output distributions, collected through a hosted gateway and published as an Apache-2.0 dataset by a third party. The rest comes from open sources. The student learned to imitate the teacher's probabilities, and on 25,376 held-out questions labeled the same way, the average divergence between their distributions is about 0.017, where zero means identical. It copies the teacher's errors too. On a poker spot where a solver always checks, Jev shoves at 0.62 and JEV-27B shoves at 0.63.
Lineage
Distilled mostly from Jev 1.13's own published output distributions, via a public third-party corpus. Shares no weights or code with TypeSafe.
Fidelity
Divergence from Jev's distributions of about 0.017, where zero means identical. It reproduces the teacher's mistakes along with its judgments.
Parity
84.07 against Jev's 83.85 across six benchmarks, in AutoTrust's own runs of both. A 0.22-point gap is noise. Call it a tie.
So the sequence runs: ship a closed model, publish calibrated probabilities as the product, and watch those probabilities become the training signal for an open model that matches you in under two weeks.
The output was the API. The API was the leak.
The thing you can't keep
Framing doesn't stay yours.
The open-source wave didn't steal TypeSafe's weights. It couldn't; there's no paper and no public model. What it took was the description. Once somebody says out loud "typed decisions with calibrated probabilities, in one pass," the people with the GPUs know exactly what to build. TypeSafe published that for free in a launch post, because they had to, because that's what a launch is.
And the description was the valuable part. The techniques were available to anyone. What didn't exist before September 15 was a name for the shape of problem they solve, a story about why you'd want one, and three primitives clean enough that a developer could see their own backlog inside them. Nate B. Jones sorted the early builds into four patterns: adding interpretation to an existing pipeline, asking new questions of old data, deciding where deeper reasoning is worth paying for, and giving small units of content their own judgment. None of those required Jev. All of them required somebody to point at the category first.
TypeSafe's terms still bar you from using Jev's output to build a competing product, which tells you exactly where they think the risk lives, and JEV-27B is what that clause was written against. AutoTrust states plainly that it shares no weights or code with TypeSafe and isn't affiliated with them; the corpus it trained on was assembled and published by someone else. Whether a term like that survives contact with a third-party dataset is a question for lawyers, and the model is already downloadable either way. The separate clause at launch that forbade publishing benchmarks is gone, removed in the September 23 revision, after it didn't go over well on Hacker News.
What the model asks of you
The demo everyone shares is a hot dog.
Someone asks Jev whether a hot dog counts as a sandwich. The interesting part is what they had to write to get an answer: a description of what qualifies, with a supporting description of what counts as food at all. Change the written criteria and the probability moves. Jev came back at 73 percent.
That's a rubric, and the person had to know what they meant before the model would run.
Most coverage calls this a schema, which is correct and insufficient. A schema is a data shape. What Jev demands is that you pull your criteria out of your head, write them down, and commit to them before anything executes. The model can't hedge. It can't produce a paragraph of qualified prose that lets you feel answered without being answered. It returns one of your options and a number saying how sure it is.
If your criteria were vague, you find out immediately. The confidence comes back at 0.51 and there's nowhere to hide.
TypeSafe's documentation pushes the same discipline: decompose the questions, ask atomic ones, and let your own code combine the answers. Anyone who has built a real decision framework will recognize the instruction, because building one has always required it.
The numbers, and what they aren't
TypeSafe's homepage claims 193.6 times faster and 444.6 times cheaper, and the company itself notes those figures sit on the higher end of real-world gains.
Independent reports are smaller and still large. Matthew Berman categorized 724 ads in forty seconds for nine cents. One developer sorted about twenty thousand items for a dollar forty-five. Another measured roughly 34 times cheaper and 6 times faster on a tax-document classifier, against their own previous setup rather than against a frontier model in general. Alex Volkov compacted a Claude Code session from about a million tokens to 86,000 in roughly a second by scoring each tool call instead of writing a summary. One team cleared 9,081 product-matching pairs for 32 cents in thirteen minutes, on a stage they'd designed in June and shelved because it cost too much.
The speed is real. The quality figure needs more care.
TypeSafe skipped the public leaderboards and scored Jev against workflows they wrote themselves. The reference answers weren't verified ground truth; they were the averaged predictions of two frontier models. TypeSafe discloses this and notes the bias. Jev scores about 67.8 percent where the best comparator scores 74.1. On invoice processing it drops to 61.8 against 79.1.
Agreement with the frontier rather than correctness. That's a real thing to measure, and a different thing from what most readers will assume.
Three limits worth writing on the wall.
Numbers, dates and adversarial content
TypeSafe's own documentation admits trouble with all three. Use it where the inputs are prose, not arithmetic.
A missing answer still gets answered
If the right answer is missing from your option list, the model will still confidently pick one of the wrong ones. The option list is your responsibility.
A number, never a reason
Simon Willison, after testing it on ranking Bay Area cities, came away with results that raised obvious bias questions and advice worth repeating: a model like this needs more evaluation than an LLM, not less.
The transfer
You don't need to care about decision models for this part.
Run the test on whatever you've built. If somebody described your thing clearly to a room of competent people, how long before one of them has it? For TypeSafe the answer was three days, and the person who got there first had done the work already and was waiting for someone to name it.
That's the uncomfortable shape of it. The naming created the market, and the naming is the one asset that can't be held. Most people building things have this backwards: they guard the artifact and give away the framing, when the framing was what moved and the artifact was reachable all along.
The naming created the market, and the naming is the one asset that can't be held.
If you want to use Jev for real work, access isn't the obstacle any more. TypeSafe dropped the waitlist on September 20, paused signups within two days when the demand arrived, and reopened on the 27th. Anyone can create an account, though the five dollars of free credit that came with the first opening isn't offered to new signups for now. Vercel, OpenRouter, Cloudflare and Netlify all serve it too. Or skip the hosted question entirely and run Laya or JEV-27B on a machine you own, where your data never leaves the building.
Which, given what the last three weeks demonstrated, may be the more durable choice anyway.
Questions people ask
What is a System One AI model?
It's TypeSafe AI's name for a model that doesn't write prose. You give it some state and questions you defined in advance, and it returns typed answers, a choice, a score, or a yes-no, each with a calibrated confidence number. The name borrows the fast, intuitive half of Kahneman's System One and System Two split.
What is Laya, and how does it compare to Jev?
Laya is an open Apache-2.0 decision model from Convai Innovations, published on Hugging Face on September 18, 2026, three days after Jev launched. It reports median latency around 33 milliseconds on a Tesla T4 against Jev's 236 to 276. An independent head-to-head found quality split: Jev ahead on triage, guardrails, moderation and multilingual intent, Laya ahead on some classification suites.
Does constraining the output stop hallucination?
No. A model that can only emit legal options still picks wrong options. Constraining the shape of an answer says nothing about its content, and if the right answer is missing from your option list the model will still confidently choose one of the wrong ones.
Can an open model copy a closed one just from its answers?
That's what happened here. JEV-27B was trained largely on about half a million rows labeled with Jev's own published probability distributions, gathered into a public dataset by a third party, and it ended up close enough to match Jev's six-benchmark average within noise and to reproduce its mistakes. When calibrated probabilities are the product you sell, every answer you return is a training example you've handed over.
Are Jev's benchmark numbers a measure of accuracy?
No. TypeSafe scored it against reference answers averaged from two frontier models, not verified ground truth, and disclosed that. The figures measure agreement with frontier models. The speed and cost advantages are independently reported and large.