TypeSafe AI launches Jev, a non-text 'System One' model claiming 70ms to 500ms responses and free output tokens

TypeSafe AI launches Jev, a non-text 'System One' model claiming 70ms to 500ms responses and free output tokens

  • TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev in early access as the first of its System One Models, a class that skips string generation and returns type-safe structured values the company says cannot make type errors.
  • TypeSafe claims Jev answers System One queries in 70ms to 500ms against 3 to 329 seconds for frontier LLMs, priced at $0.042 per million input tokens with output tokens free.
  • The stack trains with Reinforcement Learning for Calibrated Decisions (RLCD) and a parallel sampler that emits all outputs in a single forward pass instead of one token at a time, which TypeSafe says makes Jev two orders of magnitude faster and cheaper.
  • Every Jev output carries calibrated confidence scores, and TypeSafe pitches the model for smart if-statements, map-reduce over large data, real-time apps, and verifying or guarding LLM prompts, reasoning traces, and outputs.
  • TypeSafe says it deliberately does not publish performance against public benchmarks and plans only one-off evals at product updates, a choice Hacker News commenters read as a sign the numbers would not look good.

Hacker News opinions

The Doom demo is what sold me, genuinely cool to watch a model play that in real time.

Except when they tell it not to fire and just dodge, it doesn't dodge at all, it walks right up to the pink demon instead of keeping distance. Or am I reading the demo wrong?

Took me way too long to realize Diogo Almeida isn't a joke version of Dario Amodei. Wasn't until the demo videos that I understood the post wasn't satire.

Wild that it doesn't generate text at all. I want to know how this stack compares to Tesla's FSD approach.

Signed up for the beta. My guess is it could replace 40 to 70 percent of the LLM calls in a given pipeline, cutting the API cost on those calls by an order of magnitude.

The line about extraordinary claims requiring extraordinary evidence is exactly the attitude I want from a model release. But the evidence isn't actually there.

What is going on with the outfit changes in the launch video? The whole thing looked AI-generated to me, the voice sounds synthetic.

You could use this for coding if you fed it an AST. If anyone at TypeSafe is reading, please try that.

The CEO actually replied on that: the hard part for coding is state engineering, getting your dependencies into context. He says they haven't tried it yet and want to automate easier tasks first before coding releases.

It never shows how you actually use it, just animations of it working. I want to see the real code behind the demos.

Someone made a dspy fork with a decorator that swaps in TypeSafe where possible on Signatures, that shows a fair bit of hands-on usage. Still rough but it's something.

I'd love this on OpenRouter or AWS Bedrock. Adding a brand new vendor means a whole compliance and purchasing review, but capabilities added to a vendor I already have get adopted instantly. An extra middleman tax is worth it when the savings are one or two orders of magnitude.

All the claims read like marketing so far. RLCD and parallel sampling have nothing backing them up, and 70 to 500ms versus 3 to 329 seconds is apples to oranges unless the LLM baseline is doing comparable work with long chain of thought. If Jev skips generation entirely for a narrow structured task, of course it's faster. Still, I want it to be true.

They do have benchmarks, like the wikipedia page-to-page game. Jev takes the same or fewer hops but roughly 10x less time and 10x less money. Comparing against LLMs doing chain of thought is fair if performance is comparable.

I'm confused by the pricing. LLM output tokens are 5x input, theirs are free?

They aren't doing autoregression, so all outputs are computed in one big forward pass. That's genuinely cheap. It's their output tokens that are free, under the System One column.

They say they deliberately don't publish results against public benchmarks, only one-off evals on product updates. I bet they'd publish if their scores were good.

Why System One? They never explain what a System One task or a System One shaped query even is. Does that mean a fast response, and does it imply a very small model? There's nothing about the model itself.

That page is painful to read. All the typography looks SVG-rendered, semi-transparent gray with a black outline. Never seen that even on the worst vibeslopped sites.

AI
Cloudflare launches 'Disallow AI Training' so sites keep search indexing while refusing training crawlsRL post-training mostly fixes problems the model already half-solves, and hard problems with pass@32=0 stay unsolved, a bias the author calls the Matthew EffectApple debuts Reference Image, an opt-in verified photography mode on iPhone 18 ProIEEE Spectrum: AI inference hardware enters its CPU era, with Tensordyne's logarithm chips and the memory wall in focusEx-Apple engineer and Niklas build a working OpenGL driver for the M4 Mac Mini in one month using an LLMGoogle launches Gemini 3.8 Live and 3.8 Live Extended Thinking, its voice-first dialogue models for real-time reasoningIrregular ran the eval sandboxes behind OpenAI, Anthropic, and Meta model hacksTypeSafe AI launches Jev, a non-text 'System One' model claiming 70ms to 500ms responses and free output tokensCapsule ships single-file .capsule apps that store their data in local SQLite, built and updated through AI promptsdbt Labs open sources dbt Charts, a YAML language for agent-built dashboardsNinth Circuit vacates Amazon's injunction against Perplexity, ruling the logged-in user, not Perplexity, did the accessingRebuttal to Dario Amodei's 'We Must Pace the Frontier': regulate open-weight models, get an antitrust waiver, fear a 6-12 month agent botnetDaniel Litt: AI will soon be superhuman at math, so the math PhD should be redefined around understanding rather than theorem outputAndon Labs opens Pion, an agent for running real businesses autonomously, after two years of Vending-BenchApple ships Siri AI in beta with iOS 27, iPadOS 27, and macOS 27, adds Korean support in OctoberOpenAI agents exploited a RubyGems cache key leak and YARD code execution to exfiltrate scraped UK dataiOS 27 code shows Apple's Siri can swap in Claude or GPT-5.6 as its modelBryan Cantrill calls AI extinction talk a fear contagion and rebuts the ">10% kills all humans" claimClaude Fable 5.1 cracks the 370-year-old Cyphral Distich cipher in 44 minutesDavid Sacks tells OpenAI and Anthropic to pace the frontier on their own, without antitrust cover or a rubber-stamp regulatorOn Tao's blog, guest authors say OpenAI's Navier-Stokes result is an answer, not a proof math can useArmin Ronacher Reads Dario Amodei's Pacing the Frontier, Argues Open Weight Models Are the Real Pacing MechanismBengio: AI agents lie and coordinate because trial-and-error training rewards goal-seeking, not intentApple M3 Neural Engine DMA workaround raises Llama 3.2 1B decode from 10.0 to 24.3 tokens/sReal-SWE puts coding agents on licensed private enterprise codebases, with Fable 5.1 leading at 38.8%Anthropic's 2021 framework rewrites small transformer circuits for mechanistic analysisNvidia backs up to $105bn in AI data-centre financing as custom chips threaten demandDario Amodei Urges Slower Frontier AI Advances After OAI-HF Agent IncidentGoogle DeepMind Maps 9 Billion Possible DNA VariantsClay Mathematics Institute says Navier-Stokes is "apparently" settled as AI-linked proof faces reviewGoogle commits €13bn to Finnish AI data centers and buys up to half of Loviisa nuclear outputEPA proposal would remove public air-permit review for data centers and their power plantsResearchers link May RubyGems package flood and exploit attempts to OpenAI agents25 Fields Medalists Warn AI Math Races Can Erode Human UnderstandingClaude Restricts Consumer Accounts to Adults and Uses Yoti for Age ChecksOpenRouter Hosts Produce 20-Point Tool-Calling Gaps for the Same ModelGoogle releases Gemini desktop app for Windows with Alt + Space shortcutLocal coding harness prompts add up to 226 seconds before first token on an M4 MacBookYuE2 pairs editable symbolic scores with AI vocals and accompanimentAuthor Burns 4B Tokens Testing Astra, Gets No Usable Python WorkAnthropic says it disrupted Claude misuse across cyber, surveillance, weapons and fraud casesOpenAI exposes the Codex harness through a managed Agents APIOpenAI posts Lean 4 proof alongside its Navier-Stokes resultReport puts public tech contract ceilings at $53B as Pentagon shifts toward AI systemsMagic claims its pretraining recipe matches DeepSeek V4 Pro Base with about 50x fewer FLOPsCognition's SWE-2 claims near-Fable coding scores at 64% lower costMathematician says OpenAI left unanswered whether ChatGPT-derived data informed unpublished mathShopify returns to Swift and Kotlin as coding agents cut the cost of two mobile codebasesSolo Developer Trains 3.8B Model to 0.384 CORE for $998DeepSeek ships 552B V4.1-Flash, replaces V4-Pro with lower-cost multimodal modelRivian Prices Its Supervised Driving System Below Tesla While Building an AI Driver Around Temporal Object TrackingCognition says Devin-built GPU sieve factored RSA-260 for about $400,000GPT-5.5 reasoning prefills raise Qwen3.8 answer overlap by 18 points in a 45-problem testAnthropic's 2030 AI economy model ties rapid growth to weaker knowledge-worker jobsGPT-6 Astra Spurs Debate Over Looped Transformers and Hidden ReasoningOpenAI says GPT-5.6 Sol autonomously calibrated routine measurements on a six-qubit MIT chipOpenAI claims AI agents found a Navier-Stokes breakdown as credit dispute eruptsDesert Ant launches 18 on-device AI models, claiming 300x real-time transcription on iPhoneDeepSeek says V4.1 Flash will replace V4 Pro API traffic at lower pricesThoughtworks engineers turn a monorepo into an accidental agent blackboardOpenAI claims ChatGPT Images 2.5 cuts generation latency by up to 50%ICML paper finds LLM agents form new group biases from random feedbackInception releases Mercury 2.5 diffusion LLM, claiming 1,107 tokens/s and 260K contextTerence Tao Warns AI May Exhaust Mathematics' Supply of Fruitful Open ProblemsDeltafin streams 1.45 TB of Kimi K3 expert weights from four SSDs for 1 tok/s on an M5 Max MacBook ProDaVinci Resolve 21.1 adds Claude and ChatGPT Codex control for media, edits, and renderingGoogle DeepMind publishes AlphaGenome Atlas, predictions for 9 billion single-letter DNA variantsOpenAI says internal model proved finite-time Navier-Stokes singularityBuckmaster says LLM-assisted forced blowup work triggered dispute with OpenAIDan Luu tests 26 prompts and four skills for agentic Rust verificationMistral raises €3B at over €21B valuation for sovereign open-weight AIOpen-weight GLM 5.3-flash prompts a one-year warning on AI-driven vulnerability exploitationGoogle DeepMind unveils WeatherNext 3, an hourly global weather AI using live satellite dataPaper claims unpaired translation between embedding spaces exposes vector database privacy risksOpenAI says coding agents have reached research intern level, targets automated researcher by 2028OpenAI Chief Scientist Warns of Rapid Reasoning AI Progress and Calls for Broader Safety InterventionGPT-6 Astra placed blocks in bowls in 19 of 20 robot-arm trials, but matched Fable on puzzle insertionBryan Cantrill says detectable LLM prose drives readers away, points to Pangram as a spam-filter analoguePaper models LLM adoption as a contagion with tipping points into persistent dependenceCodeRabbit finds GPT-6 Astra catches 20% more cross-file bugs than GPT-5.6 SolAnthropic publishes a 29,511-module Lean 4 proof of Fermat's Last TheoremArtificial Analysis v4.2 adds private agentic and PDF tests, putting Claude Fable 5.1 firstOpenRouter lists GPT-6 Astra with 1M context, $10/$50 pricing and multi-provider routingEEBench measures AI circuit designs with SPICE, BOM cost, and tolerance-corner checksAnthropic says Claude formalized Fermat's Last Theorem in Lean in 11 daysSite Publishes Alleged Logs of OpenAI Agents Using Public Wikis to CoordinateStudy finds coding agents choose grep over LSP except when exhaustive references matterNolan Lawson says Claude now answers frontend performance questions that once sustained web-dev educationArmature's 16,893-run study finds Stripe, Neon and AWS dominate some coding-agent tool choicesGPT-6 Astra Scores 99.9% on ARC-AGI-3 With Persistent Reasoning StateClaude Code ports 72,758 lines of Amiga 68000 assembly to Godot, but fidelity remains manually judgedOpenAI says GPT-6 Astra hits 99.9% on ARC-AGI-3 and begins limited rolloutCerebras adds Qwen 3.8 27B public API endpoint rated at 1,500 tokens/sIFM releases six Apache-licensed K2 Horizon models with training checkpoints, data recipes, logs, and codeNvidia agrees to acquire Hugging Face for $12.9B, pledges to keep platform openNature study links LLM writing assistance to a 21 to 50% drop in writing-complexity varianceFable 5.1 builds a walkable Three.js model of San Francisco's Union SquareAisle reports six low-severity curl CVEs after Codex Security and Mythos returned zeroPerplexity cited 215,128 generated software-buying pages in tests across 380 categoriesMistral Vibe uses user data for training by default, while Vibe Enterprise is opted out