Qwen 3.8 27B scores 52 on Artificial Analysis Intelligence Index
Simon Willison reports that Qwen 3.8 27B has posted a score of 52 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Luna (max) and trailing GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) by one point on the same metric. That places a 27-billion-parameter open-weight Qwen release at levels associated with far larger systems, including a 753B GLM model cited in related coverage. Early operator accounts describe Qwen3.8:27b running as a top orchestrator on a 24GB M4 Mini within Hermes Agent, suggesting some deployments can move off cloud APIs and high-end data-center GPUs. The reported benchmark gap is narrow, but the score reflects one composite index rather than verified performance across every real workload, so builders should confirm fit before swapping production stacks.
Qwen 3.8 27B scores 52 on Artificial Analysis Intelligence Index
Simon Willison reports Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index. That matches GPT-5.6 Luna (max) and sits one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) on the same benchmark.
Key takeaway
Qwen 3.8 27B reaches a 52 Artificial Analysis Intelligence Index score, tying GPT-5.6 Luna (max) despite using far fewer parameters than nearby leaders.
What happened
Simon Willison reports that Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Luna (max) and sitting one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) on the same benchmark.
Related coverage frames the result against much larger models, citing GLM-5.2 at 753B parameters, while a separate operator post describes Qwen3.8:27b serving as the top orchestrator in Hermes Agent on a 24GB M4 Mini.
Evidence
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index.
Simon Willison · attributed
Simon Willison reports Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index.
The 52 score matches GPT-5.6 Luna (max) on the same benchmark.
Simon Willison · attributed
That matches GPT-5.6 Luna (max) and sits one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) on the same benchmark.
Related coverage cites GLM-5.2 at 753B parameters in comparison to Qwen 3.8 27B.
Simon Willison · attributed
Qwen 3.8 27B has scored 52 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Luna and trailing only slightly behind much larger models like GLM-5.2 (753B) and DeepSeek V4 Pro (1.6B).
An operator reports Qwen3.8:27b running as a top orchestrator on a 24GB M4 Mini in Hermes Agent.
Teknium (X) · attributed
Qwen3.8:27b has claimed its seat in my @NousResearch Hermes Agent as top orchestrator in my M4 mini 24GB ma
r/LocalLLaMA discussion links Artificial Analysis Qwen3.8-27B benchmarks to DeepSeek V4 and GPT-5.6 Luna Max.
r/LocalLLaMA Top · attributed
Artificial Analysis benchmarks indicate that Qwen3.8-27B achieves performance levels comparable to top-tier models like DeepSeek V4 and GPT-5.6 Luna Max.
Why it matters
Operators weighing local or lower-cost inference stacks now have a concrete benchmark anchor suggesting a 27B open-weight model can sit near models associated with much larger parameter counts.
Limits and uncertainties
The packet reports a single Artificial Analysis Intelligence Index score, not independent verification across diverse production workloads.
The r/LocalLLaMA item excerpt contains only submission metadata, limiting direct community evidence beyond headline-level discussion.
Comparative parameter counts for DeepSeek V4 Pro appear only in related-article analysis rather than in Simon Willison's primary report.
Practical implications
Teams running agent orchestration on consumer Apple Silicon should benchmark Qwen3.8:27b against their current cloud or local model before committing.
Cost planners can use the 52-point index result as a starting point when comparing a 27B open-weight stack to GPT-5.6 Luna (max) and larger GLM or DeepSeek variants.
What to watch
Whether Artificial Analysis updates the Intelligence Index or publishes task-level breakdowns that confirm or narrow the one-point gap to GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max).
Additional operator reports on real-world agent orchestration quality for Qwen3.8:27b on 24GB-class local hardware beyond the Hermes Agent post.