Skip to content

OpenAI Decisions API opens to all developers in beta

OpenAI's Decisions API is now in public beta. It returns a probability, a pick or a score instead of text, runs on GPT-6 Luna and costs $0.10 per 1M input tokens.

By Tech AI Wire Team

3 min read

XLinkedIn
Screenshot of OpenAI's Decisions API documentation, with a notice that the API is in public beta and runs on gpt-6-luna.

By the numbers

per 1M input tokens, with no output charge
$0.10
per request in OpenAI's demo, vs 1.6 s via the Responses API
~150 ms
answer types: predicate, choice and score
3
steps right in Every's preview test, vs 73 for Jev
76 of 78

OpenAI opened its Decisions API to all developers as a public beta on October 6, 2026, according to AlphaSignal. The API does not write text. It answers a fixed question with a probability, one pick from a list, or a score. OpenAI says that makes it about 10 times faster than asking the same question through its Responses API. For apps that make thousands of small choices, such as routing support tickets, speed and price are the whole point.

What the Decisions API does

Most AI apps ask a large language model, or LLM, to make small calls all day. Is this message spam? Which team should get this ticket? The model writes its answer as text, word by word, and the app then reads it back.

The Decisions API skips the writing. You send it evidence, which can be text, images or both, plus a question. It returns a typed answer instead of prose. OpenAI's documentation lists three answer types:

Answer typeWhat you get back
PredicateA probability between 0 and 1, for yes-or-no questions
ChoiceOne of the options you supplied, with probabilities and a confidence value
ScoreA probability-weighted score across ordered levels, with confidence

Calls go to a new endpoint, POST /v1/decisions. The only model it runs today is GPT-6 Luna, the documentation says.

Price, limits and availability

The documentation lists a price of $0.10 per 1 million input tokens. Tokens are the small chunks of text a model reads. There is no charge for output, cache reads or cache writes. Regional processing and very long inputs cost extra, though the page does not give those rates. AlphaSignal reported on launch day that no Decisions-specific pricing had been published yet, so the figure may have been added since.

Other details from the documentation:

  • Status: public beta, with general availability expected "in the coming weeks"
  • Images: must be sent inline as base64 data URLs; links to hosted images and file IDs do not work
  • Data: US and Europe data residency, plus Zero Data Retention and HIPAA support for eligible customers
  • SDKs: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0 and Java 4.78.0

Rate limits, the maximum input size and the maximum number of options are not stated.

How fast is it?

The speed claim comes from OpenAI's own demo. It routed 10,000 customer requests between Billing, Technical and Sales teams. Each decision took about 150 milliseconds, against 1.6 seconds through the Responses API, AlphaSignal reports. OpenAI has published no method for that test.

The publication Every ran early checks during the preview. In an offline replay of a computer-use task, the Decisions API got 76 of 78 steps right, AlphaSignal reports. TypeSafe's Jev, a rival decision service, got 73. On sorting message threads, the two tied for accuracy and Jev was faster. Both samples were small.

A crowded new category

OpenAI first showed the API as a limited preview on September 29, XenoSpectrum reports. It arrives in a busy field. TypeSafe's Jev launched in September and claims 70 to 500 milliseconds per answer. Cloudflare followed with Clef, open-weight decision models, and AWS released Strands Decider 2B.

XenoSpectrum argues the idea is less new than it looks. Structured Outputs can already force a model to pick from fixed options. What changes is a separate API built for many fast, repeated calls.

What this means for developers

  • Move your fixed-choice calls first. Routing, spam checks, guardrails and tool picks fit the three answer types. Leave writing and reasoning to a normal model.
  • Use the probability as a gate. Act on high-confidence answers and send low-confidence ones to a bigger model or a person.
  • Run your own latency test. The 150 ms figure is OpenAI's, with no method published. Time it on your real traffic before you redesign anything around it.
  • Budget on input only. At $0.10 per million input tokens with no output charge, cost depends on how much evidence you send. Trim long prompts and resize images.
  • Encode images first. Hosted image URLs are rejected, so convert them to base64 before the call.
  • Treat it as beta. Pin your SDK version and keep a fallback path until general availability.

Start by replaying a day of logged decisions through the API. Then compare its answers with what your current model chose, and check how often they agree.

Sources

  1. Decisions API guide - OpenAI
  2. OpenAI's Decisions API Promises Routing Answers 10x Faster for Developers - AlphaSignal
  3. OpenAI Decisions API preview on GPT-6 Luna - XenoSpectrum

Related articles

The weekly digest

One email every Friday with the week's top stories from all six desks: AI, dev tools, coding, the tech industry, startups and what's next. Free, no spam.

Unsubscribe anytime with one click.