PostHog's Jeeves adds autoregressive reasoning to Jev-style decision models, trading speed for accuracy
- PostHog published jeeves, an open source decision model that runs an autoregressive reasoning pass before emitting a calibrated probability, extending the Jev family of single-forward-pass decision models.
- It ships as a 9B model, and the reasoning pass is expensive: one commenter measured a p90 of about 17 seconds, and another spent over 30 minutes on an M5 Pro with 48GB to decide just 100 German soccer tweets for irony.
- On that same irony benchmark the tester got 68 correct against Jev's 79, so jeeves lost to Jev on accuracy while beating other open decision models he had tried.
- The repo coins noul for a yes/no output: a Bernoulli trial that returns the calibrated probability the statement is true, from 0.0 to 1.0, instead of picking one side, and includes a diffusion drafter plus calibration, checkpoint and distributed training modules.
- The repository is 20 commits old with 153 stars and 9 forks, and commenters argue the whole approach is just test-time compute scaling with a typed output interface.
Hacker News opinions
Where does the term noul even come from for yes/no answers? Digging in, it returns the calibrated probability that a statement is true instead of picking a side. I hate it.
It's literally a Bernoulli trial. If you want to be pedantic, every digital boolean in a transistor is one of these too.
The whole no-hallucinations pitch rests on that. Sure, it doesn't hallucinate, but only because it forces you to make the final call yourself.
I like the name. It's short and distinct, which fits a primitive, and it draws a useful connotation with boolean. In Bayesian statistics this is just called credence, weird that they invented a new word.
If the model does autoregressive reasoning before deciding, doesn't that throw away most of what a Jev-style model buys you, the single forward pass and cheap calibrated probabilities? Or is the point to keep the typed output and probability interface while getting better accuracy on the hard cases?
LLM reasoning, and chain of thought before it, is basically test-time compute scaling. None of this should surprise anyone.
I think the hybrid route wins: part of your pipeline stays Jev-like and part goes to a traditional LLM depending on the nature of the work.
The 9B size looks off to me. Too big for the small stuff like Laya, too small for real LLM-level decisions. I'd want benchmarks against both smaller and bigger models.
What's the point if it's p90 17 seconds? At that rate I'll just use an LLM. The beauty of Jev is that it's dirt cheap and insanely fast.
Hold your horses on dirt cheap. For spam detection Luna came out 20 percent cheaper than Jev in my cases because of prompt caching, though not as fast.
Jev gives calibrated decision probabilities at low accuracy, so why not show both numbers in the same table? That's the comparison everyone wants.
I ran it on my own irony-detection benchmark over German soccer tweets. On an M5 Pro with 48GB it took over 30 minutes for 100 tweets, the thinking really is slow. It scored below Jev but above the other open decision models I've tested, 68 correct against Jev's 79.
You don't need to post-train anything for this. Just have an LLM think, then force it to output specific JSON with a prefill after the think block, and put good conditioning examples in the prompt.
That reply misses the point. A classifier is a subset of generative text, and Jev is cheap and fast enough to sprinkle across an app where an LLM would be laggy and unnecessarily expensive. The value is a new approach that unlocks new use cases.
Glad someone adapted the diffusion drafter, that part is nice. Still, jev is getting its lunch eaten apparently in under two weeks.
One week of AI hype is enough to close a billion-dollar term sheet with VCs, so it doesn't matter much.
In my experience Jev is only faster because it's a small, shitty model. Anyone have numbers against an ordinary 9B LLM with structured outputs, on both accuracy and speed?
How general are these Jev-type models really? Has anyone run a broad cross-domain eval on them? And is there anything like this that takes image input?
Jev models are about to have their lunch eaten in under two weeks.