Review Machine · Master Review · AI decision model
Jev, by TypeSafe AI
The company says 193x. The tests that were not the company's say 5x to 25x.
Across twelve sources read on 22 September 2026, Jev is TypeSafe AI's early-access decision model whose 193x-faster and 445x-cheaper headline is self-measured against two other models' answers, and whose three independent tests found 5x to 25x faster and 1.6x to over 200x cheaper on narrow classification tasks, with accuracy that held up in both that scored it.
Maker TypeSafe AI, San Francisco, founded 2024Launched 15 September 2026, early accessPrice $0.042 per million input tokens, output freeFunding $40 million seed led by DCVC; no weights or paper publishedRead 22 September 2026Skip to the verdictRead as text

Verdict age
ProvisionalTests few · press quickJev launched on 15 September 2026 and everything here is from the seven days after it. The independent tests are from 15 and 16 September, the editorial pieces from 16 to 21 September, and the company's pages were read on 22 September. It is in early access and the current build is jev-1.13.0; anything measured this week may not hold next month.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 12 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
12 sources, 15 Sep – 22 Sep 2026. Editorial and forum tiers thin.
By tier, heaviest first
The launch post of 15 September 2026 and the homepage, read 22 September.
The Register of 16 September, Tom's Hardware and The New Stack of 21 September. All report the claims; The New Stack adds one developer's out-of-scope test.
One analysis, 16 September 2026, that traced each multiple to its source and gathered the independent figures.
None. There is no usage figure beyond the thread's points and a Vercel uptake claim reported second-hand.
One Hacker News launch thread, 1,942 points and 510 comments, read 22 September 2026.
Two independent tests with published methods, 15 and 16 September; two hands-on guides, 16 and 17 September; one reference page for the funding.
Date window
15 Sep 2026 to 22 Sep 2026
Editorial reviews 16 Sep – 21 Sep 2026 · News 16 Sep 2026Everything here is from the seven days after launch: the company's post of 15 September, the independent tests of 15 and 16 September, editorial pieces from 16 to 21 September, and the company's homepage and the thread read on 22 September 2026.
Couldn’t read, so not used
- Every.to, Mini-Vibe Check, 15 September 2026body not readable; its figures are cited as reported by The Cherry Creek News
- AIwire, TypeSafe AI emerges from stealth with $40M, 16 September 2026redirect loop; funding taken from Wikipedia citing Forbes and TechCrunch
One independent test is cited second-hand because its body was not readable, and it is labelled that way wherever it appears.
Read, then rejected
- KuCoin news flash, 16 September 2026 · a one-paragraph restatement of the company's claim
- NextBigFuture, 16 September 2026 · repeats the launch post without a figure of its own
- Karmactive explainer · no byline, no date on the page, nothing beyond the company's numbers
- MindStudio, Beam, eesel, explainx and similar vendor blogs · explainers by companies selling adjacent products; none ran a test
- Ground.news aggregation · an aggregator page, not a source
What they measured
Two independent tests with numbers of their own, and the company's demo, which works out to 75x rather than 193x.
Near Here, 50 event listings, 16 September 2026
Near Here · 16 SepNear Here, response time
0.59sJev
Near Here · 16 Sep3.40sGemini 3.5 Flash-Lite
Near Here · 16 Sep2.90sMistral Small 4
Near Here · 16 SepNear Here, cost per 1,000 decisions
$0.043Jevagainst $2.496 for Gemini and $0.370 for Mistral
Near Here · 16 SepGood Start Labs, 6,003 rubric checks, 15 September 2026
91.5% agreementwith DeepSeek V4.1 Flash; agreement, not accuracy
Good Start Labs · 15 Sep$160per M gradedagainst $33,000 with Fable 5.1
Good Start Labs · 15 SepThe company's own demo
0.114sJev, against 8.566 s for GPT-5.6 Terra; 75x, not 193x
TypeSafe AI · read 22 SepThe homepage's latency demo works out to 75x. The 193.6x figure comes from the workflow evals, scored against two other models' averaged answers.
Where they agree
Typed decisions in one pass, sub-second, $0.042 a million tokens in, no text out. Every source, the company's included, describes the same object.
One tick per editorial review, in order of publication:TR The RegisterTH Tom's HardwareTN The New Stack
It returns decisions, not text
3/3Choice, score or yes-no with a probability, in one parallel pass; it cannot write a sentence. T1Official T2Editorial Other
Sub-second in every measurement
0/30.114 s in the company's demo, 0.59 s at Near Here, about 0.5 s at Good Start Labs, 0.35 s at Every.to as reported. T1Official T2News Other
Input priced at $0.042 a million tokens, output free
1/3The company's pages, The Register and the practical guide give the same figure. T1Official T2Editorial Other
Accurate on the narrow tasks tried
0/348 of 50 at Near Here against 43 and 42 for two small LLMs; 91.5% agreement with a DeepSeek judge at Good Start Labs. Other
The output cannot be malformed
1/3Valid values are fixed by the schema, so there is no type error; every source that repeats the zero-hallucination claim qualifies it the same way. T1Official T2Editorial Forum Other
The launch thread
1,942 points and 510 comments on Hacker News, read 22 September 2026. Forum
1942points510comments
Points are attention, not a measurement.
What keeps coming up against it
The multiple, in the company's own inconsistency and in three independent tests. No reasons with the answer, in two guides. Nothing published to check, in the thread and the coverage.
The headline multiple is the company's own
Self-measured · three independent tests, all narrowerthe homepage says 193.6x faster and 444.6x cheaper; the launch post says 40x to 200x and calls the workflow figures the high end of real-world gains, from four workflows the team built, scored against the average of GPT-6 Astra and Fable 5.1.
TypeSafe AI · read 22 SepTypeSafe AI · 15 Sep
traces the spread from 193.6x on the homepage to 20x to 200x in the founder's thread, works the homepage latency demo out to 75x, and collects the independent figures.
The Cherry Creek News · 16 Sep
the independent tests measured smaller gains on narrow tasks.
- Near Here 5.8x faster and 58x cheaper than Gemini 3.5 Flash-Lite; 4.9x faster and 8.6x cheaper than Mistral Small 4, on 50 event listings
- Good Start Labs, The Cherry Creek News about $160 against $33,000 per million graded answers with Fable 5.1; 1.6x cheaper than DeepSeek V4.1 Flash
- The Cherry Creek News 0.35 s against 8.83 s for Fable 5.1, catching fewer defects, at Every.to as reported
Near Here · 16 SepGood Start Labs · 15 SepThe Cherry Creek News · 16 Sep
3/3The Register, Tom's Hardware and The New Stack report the company's figures; none ran a test of its own. Repeats
A number, no reason
1/3Probabilities without a rationale; the practical guide and DataCamp both name audits and regulated decisions as where that bites. T2Editorial Other
It reads literally and dislikes noise
1/3Answers exactly the question asked, weak on counting and dates, and accuracy falls as the state fills with material the question does not need. T2Editorial Other
Not the same job as a chat model
1/3A model that only emits listed values is not competing with a text model on the same task, so 'hallucination-free' is not a like-for-like win. T2Editorial Forum Other
Nothing published to check
1/3No weights, no paper, three component names; one commenter says they rebuilt something similar in two hours on open weights. T2Editorial Forum Other
The launch thread's doubts
510comments
The thread's main objections: the original 'frontier model' framing, the hallucination claim, the Doom demo running on game state rather than pixels, and replicability on open weights. Forum Opinion, not measurement
One thread, six days old. Loud is not the same as right.
Where they split
Four places the evidence pulls two ways, starting with whether 193x is a claim or a ceiling.
Whether the headline is a claim or a ceiling
calls 193.6x the high end of real-world gains, in the same post that publishes it.
Against that: reads the same sentence as marketing ahead of evidence, and finds no independent test within a factor of seven on speed.
Both quote the companythe disagreement is about tone; on the numbers, the company and its critic agree that the independent figures are smaller.
What Jev is cheaper than
about $160 against $33,000 per million graded answers, with Fable 5.1 on the other side.
Against that: reports the same lab at 1.6x against DeepSeek V4.1 Flash.
Both truethe saving is real and it is almost entirely a function of what you would otherwise have called.
Speed against catching things
faster and more accurate: 48 of 50 against 43 and 42, with no valid event wrongly rejected.
Against that: 25x faster than Fable 5.1 and caught fewer defects.
Two tasksone binary listing check, one defect hunt over prose; neither result says what yours would be.
Attention against evidence
1,942 points, 510 comments, and Vercel reporting fast uptake per The New Stack.
Against that: two independent tests, two narrow tasks, no weights and no paper.
Soa week in, the interest is measured and the model mostly is not.
Who it's for, who should pass
For bounded questions at volume, and as a cheap reader in front of an expensive one. Pass if you need text, a reason, or someone else's measurement first.
It suits you if
- You have a bounded question and a lot of it.Routing, triage, eligibility and rubric checks are the fit every write-up names, and both independent tests were exactly that shape.
- You want a cheap second reader before an expensive call.The practical guide and Good Start Labs both land on Jev as a filter in front of a frontier model.
- You can act on a probability without a reason.The output is a value and a confidence; if your code path needs no more, the price is hard to argue with.
Pass if
- You need text out.It cannot write a sentence, and the company says so first.
- You need the decision explained.There is no rationale, only a number; DataCamp and the practical guide both name audits as where that fails.
- Your input is noisy or numeric.The practical guide reports accuracy falling with irrelevant state; The New Stack reports weaker handling of numbers than text.
- You need someone else's measurement before you commit.A week old, early access, and the three independent tests are three narrow tasks.
This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.
The verdict
Fast and cheap in every test, by a smaller multiple than the homepage, on tasks with a fixed set of answers.
Jev is TypeSafe AI's early-access model that returns typed decisions instead of text, whose 193x-faster and 445x-cheaper headline is the company's own measurement scored against two other models' answers, and whose three independent tests found 5x to 25x faster and 1.6x to over 200x cheaper on narrow classification tasks, with accuracy that held up in both that scored it.
Confidence, by tier
The company's post and homepage agree on what it does, what it costs and what it cannot do; they disagree with each other on the multiple.
Two independent tests with published methods and numbers, both on narrow binary tasks, and two hands-on guides.
Three established publications within a week of launch, all reporting the claims and one adding a hands-on caveat; none ran its own test.
One analysis that traced every multiple to its source and collected the independent figures.
One launch thread, 510 comments, six days old.
- Editorial evidence
- under a month old
- Newest report
- 16 Sep 2026
- Read
- 22 Sep · month 0
It is faster than they say it is slow. It is not 193 times faster than anything anyone else has measured.
Rests onThree independent tests read on 22 September 2026 put Jev at 5x to 25x faster than the models it was set against. The 193.6x figure is TypeSafe's own, from workflows its team built, scored against the average of GPT-6 Astra and Fable 5.1, which the company's own post calls the high end of real-world gains.
Sources
12 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialIntroducing System One Models & Jev15 Sep
- T1OfficialHomepage, headline claims and pricingread 22 Sep
- T2EditorialTypeSafe AI debuts model for machines that plays Doom, by Thomas Claburn16 Sep
- T2EditorialTypeSafe AI's Jev offers an alternative to LLMs that claims to be 193x faster and 445x cheaper, by Bruno Ferreira21 Sep
- T2EditorialTypeSafe launched Jev because sequential LLMs are 'totally useless for computers', by Adrian Bridgwater21 Sep
- T2NewsTypeSafe's Jev claims 193x faster and 444x cheaper. Its own eval scores against two other models' answers, by Sofia Reyes16 Sep
- ForumIntroducing System One Models and Jev, 1,942 points, 510 comments16 Sep
- OtherVerification is the bottleneck, by Alex Duffy15 Sep
- OtherTesting TypeSafe Jev, Mistral and Gemini for local event validation, by Jon Reed16 Sep
- OtherJev: TypeSafe's System One model that never hallucinates, by Matt Crabtree16 Sep
- OtherHow to use Jev: a practical guide, by Prosper Otemuyiwa for Valyu17 Sep
- OtherJev (AI model), citing Forbes 15 September and TechCrunch 18 September 2026 for fundingread 22 Sep
MethodOn 22 September 2026 we read two of TypeSafe AI's own pages, three pieces from established publications, one news analysis of the claims, two independent tests with their own figures, two hands-on write-ups, one Hacker News launch thread and one reference page, by web search and plain page fetch, and synthesised them with AI; we ran nothing and tested nothing.
Dates without a year are 2026.
The review as text
Jev is the first model from TypeSafe AI, a San Francisco company that came out of two years in stealth on 15 September 2026 with $40 million of seed funding and a founder, Diogo Almeida, who worked on RLHF, InstructGPT and ChatGPT at OpenAI. It does not write text. You send it a state and typed questions, and it returns a choice, a score or a yes-no answer with a probability attached, in what the company measures at 70 to 500 milliseconds. We read twelve sources about it on 22 September 2026: two of the company's own pages, three pieces from established publications, one news analysis of its claims, three hands-on write-ups of which two are independent tests with their own numbers, one launch thread with 510 comments, and one reference page for the funding.
The headline is the argument. TypeSafe's homepage says 193.6x faster and 444.6x cheaper than LLMs; its own launch post says 40x to 200x faster and calls the workflow figures the high end of real-world gains; the three independent tests we could read measured Jev at 5x to 25x faster than the models it was set against, and cheaper by anything from 1.6x to over 200x depending on which model sat on the other side. Every one of those tests was a narrow classification task, and in the two that scored accuracy, Jev did well.
Consensus
Everyone we read agrees on what Jev is and what it is not. It takes unstructured input and returns typed decisions in one parallel pass instead of generating tokens one at a time, which is where the speed comes from, according to the company's post, The Register, Tom's Hardware, The New Stack and the practical guide. It cannot generate text at all, which the company says plainly and every write-up repeats. Input costs $0.042 per million tokens and output is free, on the company's pages and in The Register and the guide.
The reading is strong on shape and thin on independent measurement. The company's headline multiples come from four workflows its own team built, scored against the average answer of GPT-6 Astra and Claude Fable 5.1, which The Cherry Creek News points out measures agreement with two other models rather than correctness, a limit TypeSafe's own post acknowledges. No architecture, weights or technical paper has been published, per Wikipedia's summary of the coverage and the launch thread, where one commenter said they had rebuilt something similar in two hours on open weights.
Recurring strengths
It is fast, in every measurement anyone has made. The company's demo shows 0.114 seconds against 8.566 for GPT-5.6 Terra. Near Here measured 0.59 seconds against 3.40 for Gemini 3.5 Flash-Lite and 2.90 for Mistral Small 4. Good Start Labs reports roughly half a second a call. The Cherry Creek News cites Every.to at 0.35 seconds against 8.83 for Fable 5.1.
The output cannot be malformed. Because valid answers are defined in the schema in advance, Jev cannot return a value outside them, which the company calls zero hallucinations and DataCamp, The Register and the launch thread all qualify the same way: a valid value can still be the wrong one.
It was accurate on the narrow tasks that were tried. Near Here ran 50 event listings through it and got 48 right, against 43 for Gemini 3.5 Flash-Lite and 42 for Mistral Small 4, with no valid event wrongly rejected. Good Start Labs found it agreed with DeepSeek V4.1 Flash on 91.5% of 6,003 rubric checks, and says in the same paragraph that agreement is not accuracy.
The price makes redundancy affordable. Near Here paid $0.043 per thousand decisions against $2.496 for Gemini and $0.370 for Mistral. Good Start Labs puts a million graded answers at about $160 against $33,000 with Fable 5.1, and argues for running several cheap graders rather than one expensive one.
Recurring complaints
The multiples are the company's own. The Cherry Creek News lays out the spread: 193.6x and 444.6x on the homepage, 40x to 200x in the launch post, 20x to 200x and 40x to 400x in the founder's thread, and a homepage latency demo that works out to 75x. The independent figures it collects are narrower: about 5x speed and 8.6x cost against Mistral Small 4 at Near Here, 1.6x cost against DeepSeek V4.1 Flash at Good Start Labs, and at Every.to a 25x speed gain alongside fewer defects caught than Fable.
It reads literally and it has no reasons. The practical guide reports that Jev answers exactly the question asked, struggles with counting and date arithmetic, and loses accuracy as the state fills with material the question does not need. DataCamp notes it returns probabilities without a rationale, which matters wherever a decision has to be explained. Tom's Hardware adds that it can misclassify and can be attacked adversarially.
The comparison is not like for like. The Register's objection to "hallucination-free" is that a model which only emits values from a list is not doing the job a chat model does, so the two are not competing on the same task. The launch thread makes the same point about the Doom demo, which ran on structured game state rather than pixels, and The New Stack reports one developer who pushed encoded image data through it and got 35% accuracy.
Nothing is published to check. No weights, no paper, no architecture beyond the names of three components: a new model architecture, a custom sampler, and a training method the company calls Reinforcement Learning for Calibrated Decisions.
Where reviewers split
Whether the headline number is a claim or a ceiling. TypeSafe's own post calls 193.6x the high end of real-world results and The Cherry Creek News reads the same sentence as the marketing running ahead of the evidence. Both are quoting the company. What neither disputes is that no independent test has come within a factor of seven of it on speed.
What Jev is cheaper than. Good Start Labs' 200x figure is against Fable 5.1 grading the same answers; The Cherry Creek News' 1.6x figure from the same lab is against DeepSeek V4.1 Flash. The cost advantage is real and it is almost entirely a function of what you would otherwise have used.
Speed against catching things. Every.to's test, as reported by The Cherry Creek News, is the one result where Jev was faster and caught less; Near Here's is the one where it was faster and caught more.
Who it suits
You have a bounded question and a lot of it. Routing, triage, eligibility, a rubric check: every write-up describes the same fit, and both independent tests were exactly that shape.
You want a cheap second reader before an expensive call. The practical guide and Good Start Labs both land on Jev as a filter in front of a frontier model.
You can act on a probability without a reason. The output is a value and a confidence, and the price is hard to argue with.
Who should pass
You need text out. It cannot write a sentence, and the company says so first.
You need the decision explained. There is no rationale, only a number, and DataCamp's point about audits stands.
Your input is noisy or numeric. The practical guide reports accuracy falling with irrelevant state, and The New Stack reports weaker handling of numbers than of text.
You need someone else's measurement before you commit. It is a week old, in early access, and the three independent tests are three narrow tasks.
This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.
One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Get the weekly verdict.
Sources
Official
- TypeSafe AI, Introducing System One Models & Jev, 15 September 2026
- TypeSafe AI, homepage, read 22 September 2026
Editorial and news
- The Register, TypeSafe AI debuts model for machines that plays Doom, 16 September 2026
- Tom's Hardware, TypeSafe AI's Jev offers an alternative to LLMs, 21 September 2026
- The New Stack, TypeSafe launched Jev because sequential LLMs are "totally useless for computers", 21 September 2026
- The Cherry Creek News, TypeSafe's Jev claims 193x faster and 444x cheaper, 16 September 2026
Independent tests and write-ups
- Near Here, Testing TypeSafe Jev, Mistral and Gemini for local event validation, 16 September 2026
- Good Start Labs, Verification is the bottleneck, 15 September 2026
- DEV Community, How to use Jev: a practical guide, 17 September 2026
- DataCamp, Jev: TypeSafe's System One model that never hallucinates, 16 September 2026
Developer thread
Reference
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.