LLMgram · AI News · 2026-08-08

DeepGrove open-sources Maple-Preview 20B ternary reasoning LLM

DeepGrove open-sources Maple-Preview 20B ternary reasoning LLM

DeepGrove has released Maple-Preview as an open-source reasoning model built on 20-billion-parameter ternary-weight architecture, positioning it as state-of-the-art within its weight class. According to the company's launch announcement on X, the 20B-A1B model targets competitive math reasoning at speeds exceeding 200 tokens per second on a Mac Mini M4, with reported throughput five to sixteen times faster than Gemma 4, Qwen3.5, and gpt-oss baselines. The release extends the growing line of locally runnable reasoning LLMs aimed at developers who want strong problem-solving without cloud dependence. Operators evaluating on-device inference stacks should treat performance figures as vendor-reported until independent benchmarks confirm claims, but the open-source availability lowers the barrier to direct testing on Apple Silicon hardware.

Sources

DeepGrove open-sources Maple-Preview 20B ternary reasoning LLM

DeepGrove open-sources Maple-Preview 20B ternary reasoning LLM

DeepGrove introduces Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM that the team calls SOTA in its weight class. The launch post claims IMO-level problem solving at 200+ tokens/s on a Mac Mini M4, 5–16× faster than Gemma 4, Qwen3.5, and gpt-oss.

Key takeaway

DeepGrove's open Maple-Preview pairs ternary 20B-A1B weights with claimed IMO-tier reasoning and very high Mac Mini M4 throughput versus major open models.

What happened

DeepGrove announced Maple-Preview, an open-source 20B-A1B ternary-weight reasoning large language model. According to the launch post on X, the team describes the model as state-of-the-art within its weight class.

The announcement claims the model achieves IMO-level problem solving at more than 200 tokens per second on a Mac Mini M4. It also states Maple-Preview runs five to sixteen times faster than Gemma 4, Qwen3.5, and gpt-oss.

Evidence

  • DeepGrove open-sourced Maple-Preview as a 20B ternary reasoning LLM

    X · attributed

    DeepGrove open-sources Maple-Preview 20B ternary reasoning LLM

  • The team calls Maple-Preview SOTA in its weight class

    X · attributed

    an open-source 20B-A1B ternary-weight reasoning LLM that the team calls SOTA in its weight class

  • The launch post claims IMO-level problem solving at 200+ tokens/s on a Mac Mini M4

    X · attributed

    The launch post claims IMO-level problem solving at 200+ tokens/s on a Mac Mini M4

  • Maple-Preview is reported as 5–16× faster than Gemma 4, Qwen3.5, and gpt-oss

    X · attributed

    5–16× faster than Gemma 4, Qwen3.5, and gpt-oss

Why it matters

If throughput claims hold under replication, Maple-Preview could shift the benchmark for local reasoning LLMs on Apple Silicon without cloud APIs.

Limits and uncertainties

Performance and reasoning claims come from DeepGrove's own launch post on X, not independently verified benchmarks.

The evidence packet includes only a single primary source: the X announcement.

Practical implications

Teams building on-device reasoning stacks on Apple Silicon can download and benchmark Maple-Preview directly against Gemma 4, Qwen3.5, and gpt-oss baselines cited in the launch post.

What to watch

Independent replication of the 200+ tokens/s and IMO-level reasoning benchmarks on Mac Mini M4 and comparable hardware.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: X