Microsoft launches MAI-Code-1.1-Flash: 25% more efficient, quarter the cost
Microsoft has released MAI-Code-1.1-Flash, an upgrade to its coding model that delivers higher quality code with 25% greater token efficiency and at a quarter of the cost of the version launched in June at Microsoft Build. Now integrated into GitHub Copilot, the model shows concrete gains: a 22% improvement on Terminal-Bench 2.1 in Copilot CLI, a 15% improvement on .NET tasks, and a 4% rise in code survival. Token streaming is 25% faster, and the model uses 25% fewer tokens per task. This efficiency is achieved through optimization across hundreds of thousands of reinforcement-learning environments in Copilot. The move reflects a hill-climbing approach to model refinement and signals intensifying competition in coding AI, yet the metrics are self-reported, so independent verification is still needed.
Microsoft launches MAI-Code-1.1-Flash: 25% more efficient, quarter the cost
MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model we launched in June at Microsoft Build. This small, efficient, coding workhorse is now in production in GitHub Copilot.
Key takeaway
Microsoft's MAI-Code-1.1-Flash delivers 25% fewer tokens and quarter cost, making high-quality AI code generation dramatically more affordable and accessible for developers.
What happened
Microsoft announced the release of MAI-Code-1.1-Flash, an updated version of its coding model, now live in GitHub Copilot. The company claims the new model produces higher quality code with 25% greater token efficiency and at a quarter of the cost compared to the June Build version.
Specific improvements include a 22% gain on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks. Microsoft reports code survival rose 4% and return visits increased 9%, with tokens streaming 25% faster and using 25% fewer tokens per task.
Evidence
MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model launched in June.
Microsoft · attributed
MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model we launched in June at Microsoft Build.
The model achieved a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks.
Microsoft · attributed
The result: a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks.
Code survival rose 4% and return visits increased 9%.
Microsoft · attributed
Most importantly, code survival rose 4% and return visits increased 9%.
Tokens stream 25% faster and the model uses 25% fewer tokens to complete a task.
Microsoft · attributed
In GitHub Copilot tokens stream 25% faster and the model uses 25% fewer tokens to complete a task.
Why it matters
This price-performance leap intensifies competition in the AI coding assistant market, putting pressure on rivals like OpenAI and Anthropic to match efficiency gains.
Limits and uncertainties
All metrics are self-reported by Microsoft and lack independent verification.
The announcement provides no direct comparison with competing models beyond internal benchmarks.
Practical implications
For developers using GitHub Copilot, expect faster response times and lower token usage, which can reduce cloud costs.
Enterprises may see lower AI coding expenses, enabling broader deployment of AI-assisted development across teams.
What to watch
Monitor independent benchmark results and third-party evaluations of MAI-Code-1.1-Flash.
Watch for competitive responses from other AI labs, including pricing adjustments or efficiency claims.