Microsoft-Decision-1 Scores Structured Choices 35x Faster Than GPT-6 Sol at $0.042 per Million Input Tokens
Microsoft launched Microsoft-Decision-1 on Oct 9, 2026: a Qwen3.5-9B post-trained decision scorer in Foundry and on OpenRouter at $0.042 per million input tokens with free outputs, aimed at routing, verification, and agent control.

Microsoft introduced Microsoft-Decision-1 on October 9, 2026 as a decision-scoring model built for routing, classification, prioritization, verification, and workflow control. Unlike a general chat model, it returns calibrated probability scores over a fixed set of options so software can act without parsing free-form prose.
In Microsoft’s Command Line post, Achint Srivastava, VP of Software Engineering in the Office of the CTO, says the model is available in Microsoft Foundry and through OpenRouter. Input tokens cost $0.042 USD per million. Output tokens are free.
CEO Satya Nadella separately flagged the launch on X, saying Microsoft is already testing it for incident response, quality control, and scientific discovery. That pitch lands in a week when agents keep expanding write access, including the containment problems Anthropic disclosed the same week.
What Microsoft confirmed
Microsoft-Decision-1 is post-trained from Alibaba’s open-weight Qwen3.5-9B for single-pass decision scoring. Microsoft says it will later rebase the service on other bases, including Microsoft AI (MAI) and OpenAI models, while keeping the same decision-scoring interface.
Given a fixed answer set, the API returns a calibrated probability for each option. Supported formats include yes/no, multiple choice, ratings, and rubric-based grading of AI responses or agent actions. That design targets agent loops that need many cheap gates, not one long reasoning essay.
| Item | Microsoft claim | Status |
|---|---|---|
| Launch date | October 9, 2026 | Confirmed (company post) |
| Access | Microsoft Foundry and OpenRouter | Confirmed |
| Price | $0.042 / 1M input tokens; output free | Confirmed |
| Base model | Post-trained Qwen3.5-9B; future MAI / OpenAI rebases planned | Confirmed (company) |
| Blind eval set | 36 benchmarks, nearly 150,000 questions kept blind from training | Company-reported |
| Latency vs GPT-6 Sol | ~35x faster (P50); 2.5x faster than H2O-Lightning-4B v1.1 | Company-reported |
| Perturbation flips | 1.3% average across eight input perturbations; zero flips on paraphrase / option reorder tests | Company-reported |
| Safety suite | 5,250 requests across 11 harmful / jailbreak / injection benchmarks | Company-reported |
The post was updated after publication to add JevBench-related accuracy and calibration notes, which is worth remembering when you quote the leaderboard language.
Why Microsoft wants a decision model, not another chat LLM
Microsoft’s argument is blunt: every agent hop that waits on a full LLM burns latency and money. Add 100 milliseconds to each of 20 sequential decisions and you add two seconds before the user sees a result. Decision-1 is meant to sit in those hops as a scorer, not as the writer of the answer.
That is a different product bet from shipping another frontier chat model. It also sits beside OpenAI’s own Decisions API work and other small decision specialists Microsoft lists in an appendix, including H2O-Lightning-4B and GPT-6 Luna Decisions. Microsoft claims Decision-1 led its 36-benchmark suite on both accuracy and speed. Independent third-party replication is not in the launch materials.
For builders already juggling cheap specialist models, the pricing story is the hook. At $0.042 per million input tokens with free outputs, Decision-1 undercuts typical frontier chat rates by a wide margin for classification-style calls. Compare that to Anthropic’s recent Claude Haiku 5.5 cut at $0.10 / $0.50 per million tokens: Haiku still generates text; Decision-1 is sold as a structured scorer.
Internal tests Microsoft is willing to name
Microsoft lists several in-house pilots. Treat them as vendor case studies, not customer SLAs.
- Xbox Research: more than 10,000 open-ended survey, Steam, and X reviews sorted into fixed themes. Microsoft says quality was competitive with GPT-6 Sol while running over 14 times faster and about 200 times less expensive.
- Copilot quality control: scoring chat and agentic responses. Microsoft says Decision-1 was competitive with GPT-5.6 Luna and about 100 times faster.
- Incident response: retrieving relevant knowledge across logs and tickets. Microsoft says Decision-1 beat an LLM baseline on quality and speed for that retrieval step.
- Microsoft Discovery: adaptive replanning where an agent grades a prior experiment and revises. Microsoft says Decision-1 was 46 times more consistent than an LLM score, three times faster on scoring, and nearly four times faster on the full replan loop.
Those numbers will sell well in enterprise decks. They also all come from Microsoft teams using Microsoft’s own harness. Outside buyers should rerun the same rubrics on their private queues before swapping a production judge.
Where it fits next to agents
Microsoft’s suggested uses read like a checklist for agent orchestration: approve or block an agent’s next step, pick a model by cost and quality, apply skill rules without re-reading a long prompt, label data, judge AI answers, route incidents, filter content, and pick UI actions for computer use.
That framing matters after a week of agent write-path failures. Anthropic’s Oct 9 report on Claude agents submitting forms and working around gated sites, covered in our Anthropic eval write-up, is exactly the class of problem a cheap external gate is supposed to catch. A scorer that can refuse an unsafe next action is useful only if the harness actually calls it before the tool runs.
Google’s Gemini agent pitch across Workspace and Microsoft 365, still in private preview per our Gemini Agent coverage, is selling planning across suites. Decision-1 is selling the opposite layer: a fast stop/go and route decision under the planner. Teams can use both ideas without treating either as a finished safety system.
What it means for Indian developers
Foundry and OpenRouter access matter more than a regional marketing tour. Indian startups already route classification and guardrail calls through Azure or OpenRouter to keep USD token bills predictable. At $0.042 per million input tokens with free outputs, Decision-1 is cheap enough to put in front of every agent tool call if the latency claims hold on your region and SKU.
Practical first tests: intent routing for support bots, rubric grading for RAG answers, and a hard allow/deny gate before any browser or payment tool. Price in USD first, then convert for finance. There is no India-only list price in the launch post.
Also watch the Qwen base. Procurement teams that already cleared OpenAI or Anthropic models may need a separate review for an Alibaba-derived checkpoint, even when the API is billed by Microsoft. Microsoft’s promise to rebase onto MAI and OpenAI later softens that issue only after those rebases ship.
For broader Microsoft AI messaging this month, pair this launch with Mustafa Suleyman’s softer job-automation forecast in our Suleyman coverage: Microsoft is selling specialized control layers as much as giant chat models.
Skeptical read
Three caveats stay open. First, the headline 35x and accuracy wins are Microsoft’s own blind suite, not a public leaderboard everyone can fork tonight. Second, Decision-1 is not a replacement for GPT-6 style generation; if your workflow still needs explanations, you will call a chat model anyway and Decision-1 becomes an extra hop. Third, OpenRouter and Foundry availability can diverge by region, quota, and enterprise agreement, so “available today” should be checked in your tenant before you blog a migration plan.
Still, the product direction is clear. As agents multiply tool calls, the winning stack may be a frontier model for hard reasoning plus a tiny, calibrated scorer for every boring fork in the road. Microsoft is betting Decision-1 is that scorer, and it put a very low USD price on the bet.
What is Microsoft-Decision-1?
It is a decision-scoring model that returns calibrated probabilities over fixed options for routing, classification, verification, and agent control. It is not positioned as a general chat LLM.
How much does Microsoft-Decision-1 cost?
Microsoft lists $0.042 USD per million input tokens, with output tokens free, via Microsoft Foundry and OpenRouter.
Is it based on OpenAI models?
The first release is post-trained from Qwen3.5-9B. Microsoft says it will later rebase onto other foundations, including Microsoft AI and OpenAI models.
Should I replace GPT-6 with Decision-1?
No for open-ended writing or deep reasoning. Yes as a candidate for high-volume structured decisions where you already know the option set and care about latency and cost.
Where can developers try it?
Microsoft points to the Foundry catalog model page for Microsoft-Decision-1 and to OpenRouter under microsoft/microsoft-decision-1.
1 comment