OpenAI launches Decisions API in public beta: gpt-6-luna only, 10x faster than Responses, $0.10 per million input tokens
- OpenAI put the Decisions API into public beta at
POST /v1/decisions, returning typed answers about 10x faster than the Responses API, with GA promised in the coming weeks and gpt-6-luna as the only available model. - A request takes three fields (model, input, questions) and the response holds an answers array where each question's unique name is echoed back; predicate returns a probability from 0 to 1, choice returns one of your supplied values, and score returns the probability-weighted average of the level indices.
- Pricing is $0.10 per million input tokens with output free; a commenter measured the same benchmark at $0.0192 for Jev versus $0.06 for Luna through OpenRouter, roughly 3x cheaper for Jev.
- Input can be text, images, or both, which commenters flagged as the first big gap they had found in Jev, and the docs include an example of the API playing a video game.
- Simon Willison wrapped the endpoint in an llm plugin, noting that when OpenAI defines an endpoint like
/v1/decisionsit tends to become a de facto standard for other providers.
Hacker News opinions
Didn't take long for OpenAI to fast-follow Jev. People were saying weeks ago that this was coming and they were right. Curious whether the other big labs follow suit.
How's the pricing against Jev?
$0.10 per million input tokens versus Jev's $0.042, both free on output. In the same bench a full Jev run cost $0.0192 and Luna cost $0.06 through OpenRouter today, so about 3x in Jev's favor.
Benchmarks or it didn't happen.
Can I use this through a subscription?
No. The OpenAI subscription has never covered any API calls. The closest you get is structured outputs via Codex.
It already supports image inputs, which was the first big gap I found in Jev.
Since it's fast and understands images, I wonder if it can play video games. I have a harness for an LLM playing EA FC but even the fastest models are too slow. Need to try this with the Decisions API.
One of the examples on the docs page is it playing a video game. Doubt it runs anything complex though. You're trading accuracy for speed.
I genuinely don't understand why anyone would pay OpenAI for this. Running something comparable to Jev is trivial. The whole point of paying for ChatGPT is that OpenAI has warehouses running a zillion-parameter model, and a decision model is way easier and cheaper to run. Feels like zero moat.
Why would I run it myself? It's $0.10 per million tokens, dirt cheap, and Jev is cheaper still. Same question as why anyone rents a VPS instead of running their own hardware. Buy versus rent is about what's economic, not what's possible.
There's no moat in the self-hosting sense, but you need a reason for people who don't want that to stay on your platform. Payments and management are easier this way. It's a race to the bottom on price, so it comes down to branding and platform stickiness.
Depends on the quality of the results. These are driven by text prompts, so if OpenAI returns better quality than the open weight variants the market rewards it. Anyone using a decision model has to spin up their own evals, these are far harder to vibe-check than regular text LLMs.
For my use case it'll cost like $11 a month and we already have OpenAI keys and billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai.
Existing enterprise contracts, zero data retention agreements, staying with one provider because everything is in one place. There are probably a lot more reasons.
If your company has a 3 to 6 month onboarding period for new vendors and a lifetime of vendor management overhead, this makes a ton of sense. Add model governance for anything you train yourself and it's a slam dunk.
My reason is different. Every decision model use case I can think of, I don't want the model changing in X weeks when the lab decides to improve it or make it safer.
Jev should be the nail in the coffin on whether the AI business is a commodity market. It appeared out of nowhere and kicked off the next round of price wars, and it showed the value of System One models. A fast yes/no/confidence score is cheaper and often all people want.
It takes me two keypresses to switch models. I don't know of a less sticky product.
Curl against /v1/decisions with complaint and compliment predicates came back 0.91 and 0.06, 310 input tokens and zero output tokens. Notable because when OpenAI defines an endpoint like this it ends up a de facto standard. I turned it into a new llm plugin.
How is this different from the categorization models from the ML era?
Real world problems are more complex and need highly structured outputs covering many attributes. Make twenty independent calls to a decisions API and you lose coherence among the twenty decisions. I think the hype gets forgotten soon enough.
Ran my own evals, under 600 calls on UI component selection, chat charting, tag selection and PKM, via OpenRouter against Jev and Mercury Decide. Slower than Jev, similar latency to Mercury, 346ms p50 and 860ms p95, less confident on my ambiguous UI component tasks where Luna fails at 0.6 and lower, 4 failed calls versus 0 for both competitors, and about 3.1x more expensive than Jev on average.