OpenAI Decisions API Returns Predicates, Choices, and Scores About 10x Faster Than Responses at $0.10 per Million Input Tokens
OpenAI opened the Decisions API to public beta on Oct 6, 2026: typed predicates, choices, and scores on gpt-6-luna at $0.10 per million input tokens with free outputs, about 10x faster than the Responses API.

OpenAI opened the Decisions API to all developers in public beta on October 6, 2026. The dedicated POST /v1/decisions endpoint evaluates text, images, or both and returns typed answers your code can branch on, without asking a chat model to write prose you then parse.
In the official developer docs and the OpenAI Developer Community announcement, OpenAI says Decisions runs about 10x faster than the same gpt-6-luna model through the Responses API. Input costs $0.10 USD per million tokens. Output tokens, cache reads, and cache writes are free for this endpoint. General availability is expected in the coming weeks. gpt-6-luna is the only supported model today.
The launch lands in the same product lane as Microsoft’s Decision-1 scorer, Cloudflare’s Clef models, and TypeSafe’s Jev. The common bet is simple: agent stacks need cheap, structured gates more often than they need another long essay.
What OpenAI confirmed
A Decisions request has three parts: model, shared input, and a questions array. Each question gets a unique name so the response can echo answers by name. OpenAI documents three question types.
| Type | What it returns | Typical use |
|---|---|---|
predicate |
Probability from 0 to 1 that a condition is true | Damage checks, policy flags, relevance gates |
choice |
One of your supplied values, plus per-option probabilities and confidence | Department routing, content categories |
score |
Probability-weighted average of ordered level indices, plus confidence | Severity rubrics, quality grades |
OpenAI draws a hard product line. Use Decisions when you need a probability, a fixed choice, or a rubric score. Use Structured Outputs on the Responses API when you need an object that matches your own JSON schema or a written explanation. Use function calling when the model must request a tool with arguments.
| Item | OpenAI claim | Status |
|---|---|---|
| Public beta start | October 6, 2026 | Confirmed (community announcement) |
| Endpoint | POST /v1/decisions |
Confirmed (docs) |
| Model | gpt-6-luna only |
Confirmed |
| Speed vs Responses | About 10x faster | Company-stated; not independently audited in launch materials |
| Price | $0.10 / 1M input tokens; output free | Confirmed |
| Caching | No cache-read or cache-write charges; caching unavailable in beta | Confirmed in docs and staff replies |
| Image input | Inline base64 data URLs only; no hosted HTTP image URLs | Confirmed |
| Compliance | ZDR and HIPAA for eligible customers; US and EU (EEA + Switzerland) residency | Confirmed |
| Accuracy / calibration | Not published in the Decisions guide | Open: docs tell you to set thresholds on your own labeled examples |
How the three answer types behave
The docs walk through concrete examples. A predicate can inspect a product photo for a crack, tear, or dent and return something like a 0.92 probability of visible damage. A choice can route “I was charged twice for my order” to billing with per-option probabilities. A score can grade issue severity across ordered levels such as Cosmetic, Workaround available, and Fully blocked.
Score math is explicit. Level indices start at 0. If the model assigns probabilities 0.1, 0.7, and 0.2 across three levels, the returned score is 1.1, which sits between the second and third labels. That is useful for queues that need a continuous signal, not only a hard label.
You can pack independent questions into one request against the same input. Dependent decisions still need separate calls. OpenAI also warns builders to write observable criteria, keep choices distinct, and define adjacent score levels with clear boundaries.
Price, latency, and the competitive lane
At $0.10 per million input tokens with free outputs, Decisions matches Luna’s ordinary input list price on Responses while removing output charges for this path. Regional processing premiums and long-context multipliers still apply. Staff confirmed in the community thread that caching is not available yet in the beta.
That pricing sits above Microsoft’s Microsoft-Decision-1 launch at $0.042 per million input tokens with free outputs, and above Cloudflare’s open-weight Clef and Clef-flash Workers AI models. OpenAI’s counter is distribution, image support, ZDR and HIPAA eligibility, and a Playground that is already live for experimentation.
OpenAI has not published a public accuracy leaderboard for Decisions. Community early testers on the announcement thread reported mixed results against TypeSafe Jev, including cases where predicate questions tracked expected frequencies better than choice questions. Treat those as early developer reports, not a formal bakeoff. The docs themselves tell you to calibrate thresholds on labeled samples from your own app.
For context on the broader GPT-6 rollout that Decisions sits on top of, see our coverage of GPT-6 and Intelligent UI reaching free ChatGPT users. Decisions is the API-side routing product, not the chat UI.
Where it fits in an agent stack
The useful mental model is a gate, not a writer. Before an agent opens a browser, files a ticket, or spends money, a Decisions call can answer yes or no, pick a department, or score severity. That is the same layer Microsoft is selling with Decision-1 and that Anthropic’s own agent eval problems keep highlighting: lots of small control points beat one slow frontier call at every fork.
Voice is part of the pitch too. OpenAI says you can combine Decisions with the Live API so spoken requests select actions and report results back. SDK support already lists minimum versions for Python, JavaScript, Go, Ruby, and Java.
What Decisions does not replace is generation. If your workflow needs explanations, extracted fields, or tool argument construction, you still call Responses or tools. Decisions becomes an extra hop that should earn its keep in latency and token spend.
What it means for Indian developers
USD pricing comes first. At $0.10 per million input tokens with free outputs, Decisions is cheaper than paying Luna for generated JSON when your only need is a label or score, and more expensive than Microsoft-Decision-1’s $0.042 list on Foundry and OpenRouter. Indian teams already budgeting OpenAI keys can try it without a new vendor review, which matters when procurement is slow.
Practical first tests look familiar: support ticket routing, RAG answer quality gates, image moderation for marketplace listings, and allow or deny checks before any tool that writes to the outside world. Because images must be inline base64, watch request size and mobile upload costs. Because caching is off in beta, repeated classification of the same long policy text will bill every time.
Compliance is a real differentiator for banks, hospitals, and SaaS exporters selling into the US and EU. If you already cleared OpenAI for ZDR or HIPAA, Decisions rides that path. Compare that with Anthropic’s recent Claude Haiku 5.5 pricing cut: Haiku still generates text; Decisions is sold as a typed decision surface.
What remains open
Three caveats stay live. First, the 10x claim is OpenAI’s comparison to Responses on Luna, not a third-party latency study across regions. Second, accuracy and calibration data are not in the launch guide, so production thresholds are your job. Third, beta limits are real: one model, no caching, base64-only images, and a GA date described only as “coming weeks.”
Even with those gaps, the product direction is clear. As agents multiply tool calls, the stack that wins is often a frontier model for hard work plus a fast typed decision layer for every boring fork. OpenAI just put that layer on the public API with a clean USD price and a Playground you can poke today.
Frequently Asked Questions
What is the OpenAI Decisions API?
It is a public-beta endpoint that evaluates text and images and returns typed predicate probabilities, fixed choices, or rubric scores, instead of free-form chat text.
How much does the Decisions API cost?
With gpt-6-luna, input costs $0.10 USD per million tokens. OpenAI says there are no output, cache-read, or cache-write charges on this endpoint. Regional and long-context multipliers can still apply.
How fast is it compared with the Responses API?
OpenAI states Decisions is about 10x faster than using GPT-6 Luna through the Responses API. Independent, multi-region latency studies are not part of the launch docs.
Can it replace Structured Outputs?
No. OpenAI says use Decisions for probabilities, choices, and scores. Use Structured Outputs when you need a custom JSON schema or a written explanation.
When will it leave beta?
OpenAI says it expects general availability in the coming weeks. No exact GA date is published in the Decisions guide as of October 10, 2026.