Upstage drops 250B open-weight MoE with 15B active parameters

Korean AI company Upstage released a new open-weight mixture-of-experts model with 250B total parameters and 15B active per token. Early reporting puts its performance in the same band as DeepSeek-V4-Flash, with notably strong Korean-language capability.
Key takeaway
A Korea-rooted lab is shipping a large sparse open-weight model aimed at competitive quality at lower active compute, expanding sovereign open-model supply beyond the usual US-China shortlist.
Context
The reported architecture is a classic efficiency bet: keep a very large parameter pool while activating only 15B weights per token, which matters more for inference cost and latency than raw parameter count alone. Positioning against DeepSeek-V4-Flash frames the release as a frontier-adjacent coding and general model, not a niche research demo.
For builders, open weights plus a favorable active-parameter budget can matter as much as the download itself, especially where Korean-language quality and local deployment constraints dominate. The claim still rests on secondary reporting rather than a fully audited third-party leaderboard, so treat parity language as directional until primary evals land.
Numbers to know
- 250BTotal parameters reported for the Upstage MoE
- 15BActive parameters per token reported for the Upstage MoE