Urgent.News

What's breaking now, across thousands of outlets.

AI

Microsoft-Decision-1 Scores the Choices an Agent Has to Make

TL;DR Microsoft-Decision-1 is a small language model that returns a probability score for each option in a list you hand it, so a program can route a request, grade a response or decide what an agent does next. Microsoft published the model on 9 October 2026 in Microsoft Foundry, with access through OpenRouter planned. The model was built by post-training Qwen3.5-9B for single-pass decision…

Microsoft-Decision-1 is a language model designed to provide probability scores for a fixed set of answer options, enabling software to make decisions based on the returned scores. Developed by post-training Qwen3.5-9B, Microsoft released the model on October 9, 2026, with access through OpenRouter planned. The model scored highest in a benchmark comparison covering nearly 150,000 questions, and boasts a P50 latency that is about 35 times faster than GPT-6 Sol.

Microsoft identified five engineering challenges during development: speed, quality that generalizes, robustness across paraphrased inputs, probability calibration, and safety. To test the model, Microsoft internally used Xbox Research to sort over 10,000 pieces of feedback, the Copilot team to measure response quality, and Microsoft Discovery to grade experiments.

When integrating Microsoft-Decision-1 into a system, users must provide their own labeled set of tickets, a baseline to compare against, and an understanding of which errors carry financial consequences. Unlike LLMs, decision models are designed to deliver structured outputs that software can act on immediately, making them suitable for tasks such as routing support tickets, grading AI responses, or checking tool calls against a rubric.

Microsoft-Decision-1 is priced at $0.042 per million input tokens, with output tokens being free, making it more affordable for high-volume routing and grading. The model's architecture allows it to be easily integrated into existing AI agent systems, providing a cheap judge alongside an expensive writer. By removing generated text from steps where the output is not read, the model improves efficiency and reduces costs.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

NatureQuest — Anti-Retention Outdoor Companion Powered by Local Qwen2.5

This is an official submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass . What I Built NatureQuest is a local-first outdoor adventure companion that turns a few free minutes…

  • NatureQuest promotes anti-retention through outdoor adventures.
  • Powered by local Qwen2.5 model for personalized micro-adventures.
  • Users input time and surroundings to generate safety-conscious tasks.

I built an outdoor game that uses AI to get you off your phone

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What if an app's most important feature was knowing when to get out of your way?

  • GrassQuest is an outdoor exploration game using AI to generate environment-specific quests.
  • The app aims to encourage players to put phones away and spend time exploring outdoors.

More from Sunday 11 October →