60 بالعربي

Microsoft Launches Decision Model Based on Chinese Qwen to Compete with Jev

Microsoft has launched the 'Microsoft-Decision-1' model built on Alibaba Cloud's Qwen3.5-9B model, claiming it is 2.8 times faster than Jev with 83.5% accuracy across 36 benchmarks.

October 11, 2026
Microsoft Launches Decision Model Based on Chinese Qwen to Compete with Jev

Microsoft has launched a new model named "Microsoft-Decision-1" built on the Qwen3.5-9B model from the cloud computing unit of China's Alibaba, according to a report by The Register. The company is offering the model through its "Microsoft Foundry" platform, with plans to make it available soon on "OpenRouter," a platform that aggregates multiple models into a single interface.

Microsoft stated that the model is 2.5 times faster than H2O-Lightning-4B and 2.8 times faster than Jev in latency testing, achieving an accuracy of 83.5% across 36 benchmarks and ranking second in confidence score at 92.2%, behind Quyet-1.0-Large. The company added that it is more than 20 times cheaper than OpenAI's GPT-6 Sol model in text classification tasks, according to the report.

"Decision-1" is categorized as a decision model, meaning a model designed primarily to produce structured outputs that software can act on directly, as explained by Achint Srivastava, Vice President of Software Engineering in the CTO's office at the company. The company set the usage price at $0.042 per million input tokens, while offering output tokens free of charge.

Microsoft stated that it will soon rebuild "Decision-1" on its proprietary models and models from OpenAI, without citing a reason for the move. The model reflects a growing trend among major companies to build on open-weight models from Chinese labs to reduce costs and accelerate launches, amid fierce competition with models like Jev in the enterprise applications market.

What Do These Terms Mean?

Open-weight model: An artificial intelligence model whose weights are published, allowing it to be run or modified locally, unlike closed models whose inner workings are not accessible to the public.

Latency: The time it takes for a model to respond to a request; a lower latency means a faster user experience, especially in real-time applications.

Token: The smallest unit of text processed by a model, which may be a word or a part of a word; usage is billed per million tokens.

Share
Keywords