Artificially generatedJEV Hype vs. Reality Check
Jack Rudenko watched a week of launch demos. The model is fast and cheap. The comparisons are still the wrong ones.
JEV, the decision model from TypeSafe AI, left stealth on 15 September 2026. Jack Rudenko published his read on 22 September. A call takes 70 to 500 milliseconds. Input costs $0.042 per million tokens. Output costs nothing. He does not argue with the speed. He argues with the opponent.
Almost nobody in the thread wanted this sentence.
Jack Rudenko watched a week of launch demos and wrote the thing almost nobody in the thread wanted to say: the new model is fast and cheap, and most of the comparisons are still the wrong ones.
His essay is Jev: Solving yesterday's problems, faking wins with overengineered demos, and never facing the real competition, published 22 September 2026. The feed he describes had been sorting inboxes, tagging mail, detecting language, scoring sentiment, and playing Pong. Every clip ended on the same punchline: fast and cheap, and a chat model cannot touch this.
Both claims can be true. The opponent in the clip is still a giant generalist. That is the whole complaint, and it is worth slowing down, because the product underneath the clips is a real change in shape.
State in. A typed decision out. One call.
JEV is TypeSafe AI's System One model. You give it some state and a list of typed questions. It gives you decisions: a choice from a list, a score on a rubric you set, or a yes or no with a probability between 0 and 1. TypeSafe calls that last one Noul. All of them run in one call. The output drops into code. There is no paragraph to scrape for the word you wanted.
The mechanism, the three question types, and the vendor's own speed claim are already written up on this site, in German: Jev: kein Chatbot. Ein Entscheidungsmodell. This piece does not rebuild that page. It stays with Rudenko, and with the one outside test that agrees with him.
Asking a text generator to return the word "billing" is the wrong sport.
The slide races a novelist. The job is a label.
Every public clip he watched puts JEV next to a GPT-class model. TypeSafe's own headline, as he reports it: 193 times faster and 444 times cheaper than the average of GPT-6 Astra and Fable 5.1, on four internal workflows. Ably's Pong demo: 227 milliseconds a move, against 2.5 to 3.5 seconds for Haiku, Gemini Flash, and GPT-5.6 Sol.
Against a text generator, the gap is real. A big model doing ticket routing takes seconds and bills every word it emits. JEV bills the input at $0.042 per million tokens and bills the output at zero. Calls land between 70 and 500 milliseconds. Arize describes the same price and the same latency band. The miracle, on those terms, is not invented.
Rudenko thinks the miracle is aimed at the wrong opponent. Classification already had answers. A fine-tuned BERT has been the standard since 2018, when Jacob Devlin and his co-authors published the paper: small, about 110 million parameters, happy on a laptop CPU. Before that, logistic regression on word counts. Before that, regular expressions, still running more production routers than the demos admit.
Count the network, and a boring BERT is already in the room.
He wants the column the clips leave off. On a single CPU request, he puts an unoptimized BERT-base at about 100 milliseconds, a distilled one at about 60, an optimized one under 10. JEV's 70 to 500 milliseconds include the trip to someone else's machine. Same band, sometimes faster, sometimes not. Batch the local model on a GPU and he says the item time falls under a millisecond. That is where a hundred-fold gap opens, and it opens toward the box you already own.
Cost follows the same split. $0.042 per million input tokens looks small beside Opus. Beside hardware you already pay for, the local model often wins after setup. No per-token markup. No round trip. These CPU figures are his, not a bench we re-ran. Treat them as his estimate. The direction is the point he is forcing: classifier versus classifier.
One outside team published the comparison he was asking for. Parallel, in a post by Vlad Shulman on 18 September 2026, ran JEV on work they say they do billions of times a day. Zero-shot reranking matched one of their own rerankers at NDCG@10 of 0.7. Topic classification and query freshness: their models won. Cost per document was higher, and they say why: they own the machines their classifiers run on. Their summary is his summary. The headline speed and cost are against autoregressive chat models. Specialized classifiers still often win on both.
No labels. The question changes next week.
A fine-tuned BERT cannot take a brand-new question on Thursday without a training loop: examples, a run, an eval, a deploy, and the same loop again when the labels move. JEV takes the question as a string at request time.
That covers a real set of jobs. A router while you still do not know the categories. Research while the taxonomy is moving. People who will never train a model. Agent guardrails where the question is different for every tool. Parallel's own close is the same rule: if you need a classifier and you have not trained one, start here, then check the quality on your examples.
He stops at the next morning. The day the categories freeze and the volume goes up, the labels JEV just produced are a training set. Take them and fit a small model you own. Use the API to bootstrap. The production path is the model on your machine.
Returning a score is a format. Being right is a product.
The demos mix up two sentences. JEV returns a score. The score is good. Only the first sentence follows from the product. The second is a separate question, and on that question he calls it mid-tier, like the other generalists.
TypeSafe's own dashboard, by his reading: 67.8 percent agreement with the reference across their workflows. Opus 5 in workflow mode lands at 74.1 percent. On invoices the gap widens, 61.8 percent against 79.1 percent. He likes that they published it, and that they say the multipliers sit at the high end of real gains, on evals their own team wrote.
A new wrapper on the same family of training data does not grow a new sense of taste. At MadAppGang he fed deliberately sloppy text to JEV, Opus 5, and GPT-5.6 Sol and watched every model score it high. JEV did not stand apart. So "great at scoring" means great at emitting a number.
One date in that paragraph does not fit the rest of his essay. He places the writing test in May 2024. He also places the launch on 15 September 2026. A test in May 2024 cannot have included a model that left stealth in September 2026. Read the scores as his claim. Read the month as a slip.
The job he wants is not the inbox.
Put a spec next to a diff and ask how far apart they are. A fine-tuned BERT has no dataset called "your spec versus your implementation." A chat model can do it in a paragraph you then have to parse, slowly, at frontier prices. JEV can return a typed score and a probability in under half a second, for a fraction of a cent.
That is the shape he wants in an eval harness. Underneath, checks with no model in them: object diffs, screenshots, pass or fail by a rule. Above that, a judge that asks a few typed questions. Did the change cover the spec? How far did it drift? Did it ship anything the spec never asked for? The last one is a hallucination check with a probability attached. It is the question he says he would pay for.
He names the limits himself. JEV reads text and JSON. A UI check needs the DOM, or a vision model that describes both screens before JEV scores the descriptions. State plus the longest question has to fit in 32,000 tokens, so the unit is a component, not a repository. He has not published the head-to-head against Opus 5 on implementation quality. He says that comparison is still open, and that he will publish it either way. Until that page exists, the second job is a proposal with a shape, not a result.
Keep the clip. Add DistilBERT. Or add a regex.
His ask to the people posting the demos is small. Keep the clips. The latency is real, and the clips are how a new shape gets seen. When the benchmark goes up, add one column: a fine-tuned DistilBERT, or a regular expression if the job is that small. If JEV still wins, you found a use. If it loses, you found out before a waitlist became your router.
A model that returns a decision instead of prose is the right idea. Measuring it against a novelist, then putting a cloud call where three lines of regex would do, is the mistake. Classification already had a local answer. JEV solved the zero-shot version, and opened a second job: score a result against a reference when nobody has built a training set. He thinks that second job is the one worth testing.
The rest of the feed is still playing Pong.
| Claim | Basket | Where it comes from |
|---|---|---|
| JEV left stealth on 15 September 2026. TypeSafe AI. Diogo Almeida, who worked on InstructGPT, founded it. | Fact | Jack Rudenko, 22 September 2026. Parallel's product description points at the same TypeSafe docs. |
| Calls take 70 to 500 milliseconds. Input is $0.042 per million tokens. Output is not billed. | Fact, vendor figures | Rudenko. The same band and the same price are in Arize's write-up of 18 September 2026. |
| 193 times faster and 444 times cheaper than the average of GPT-6 Astra and Fable 5.1, on four internal workflows. | Fact as a vendor headline | Rudenko, reporting TypeSafe. Not an outside replication. |
| Pong at 227 milliseconds a move, against 2.5 to 3.5 seconds for Haiku, Gemini Flash, and GPT-5.6 Sol. | Fact as a described demo | Rudenko, on Ably's demo. We did not retime the video. |
| BERT-base about 100 ms on one CPU, distilled about 60, optimized under 10. A GPU batch under 1 ms an item. | His estimate | Rudenko. No independent timing in this piece. Direction: local classifier versus API, once the network is counted. |
| Zero-shot reranking at NDCG@10 of 0.7, comparable to one internal reranker. Topic and freshness: internal models won. Cost per document higher. | Fact | Vlad Shulman, Parallel, 18 September 2026. They run the classifiers. |
| Dashboard agreement: JEV 67.8 percent on their workflows, Opus 5 at 74.1 percent. Invoices 61.8 against 79.1. | Reported | Rudenko's reading of TypeSafe's dashboard. We did not open a second copy of that board. |
| A May 2024 writing test included JEV. | Conflict | The essay dates the MadAppGang scores to May 2024 and the launch to 15 September 2026. Both sentences are in the piece. They cannot both describe one test. |
| Text and JSON only. 32,000 tokens. No published head-to-head on implementation quality. | Fact, as his limit | Rudenko. He says the Opus 5 comparison on spec-versus-diff is still open. |
- Jack Rudenko, the essay22 September 2026. The week of demos, the dashboard reading, the Pong timing, the May 2024 line, the spec-versus-diff proposal.https://www.linkedin.com/pulse/jev-solving-yesterdays-problems-faking-wins-demos-never-jack-rudenko-0ss8c/
- Vlad Shulman, Parallel18 September 2026. NDCG@10 of 0.7, topic and freshness, cost per document, and the line that headline speed is against chat models.https://parallel.ai/blog/testing-jev
- Arize, on JEV as a judge18 September 2026. Same latency band, same input price, output not billed.https://arize.com/blog/typesafe-jev-llm-judge/
- TypeSafe, introductionThe vendor's own page for the decision model. Parallel cites it.https://docs.typesafe.ai/introduction
- Jev: kein Chatbot. Ein Entscheidungsmodell.Our explainer, 22 September 2026, in German. Choice, Score, Noul, and the vendor speed claim.https://mcgrinsey.com/magazin/jev-entscheidungsmodell-typesafe/
- BERT, Devlin and others, 2018The paper behind the "since 2018" line. Bidirectional transformers for language understanding.https://arxiv.org/abs/1810.04805
Cut-off: 24 September 2026. The dashboard percentages are Rudenko's reading. The Parallel numbers are Parallel's. The May 2024 date stays in the fact table because the essay prints it.


