LLMgram · AI News · 2026-08-09

Pokee AI launches Isaac 28B with 10M-token context on one GPU

Pokee AI launches Isaac 28B with 10M-token context on one GPU

Pokee AI has unveiled Pokee-Isaac 28B, a 28-billion-parameter agentic model positioned as the first frontier-class system with a 10-million-token context window deployable on a single GPU starting from hardware comparable to an NVIDIA RTX 4090. The company cites proprietary non-decoder-only architecture, 93.3% RULER benchmark performance at 10M tokens, prefill throughput up to 137,000 tokens per second, and weight fine-tuning from Qwen3.6-27B. For teams building long-context agents without multi-GPU infrastructure, those claims describe a sharply different cost and operations envelope if validated outside vendor testing. Every metric and architectural detail currently comes from a single Pokee AI post on X, so independent replication and third-party benchmarking are still required before treating the release as settled production guidance.

Sources

Pokee AI launches Isaac 28B with 10M-token context on one GPU

Pokee AI launches Isaac 28B with 10M-token context on one GPU

Pokee AI released Pokee-Isaac 28B, a 28B-parameter agentic model it bills as the first real 10M-token context frontier-class model deployable on a single GPU starting from RTX 4090 or equivalent. It claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill on a proprietary non-decoder-only architecture, with some weights fine-tuned from Qwen3.6-27B.

Key takeaway

If substantiated, Pokee-Isaac 28B could let single-GPU setups run 10M-token agent workloads that today typically demand much heavier hardware.

What happened

Pokee AI announced Pokee-Isaac 28B on X, describing a 28-billion-parameter agentic model designed for extreme long-context use.

The post claims 10M-token context on one GPU from RTX 4090-class hardware, 93.3% RULER at that scale, 137K tokens/s prefill, a proprietary non-decoder-only design, and fine-tuning from Qwen3.6-27B weights.

Evidence

  • Pokee AI released Pokee-Isaac 28B as a 28B-parameter agentic model.

    @Pokee_AI on X · attributed

    Pokee AI released Pokee-Isaac 28B, a 28B-parameter agentic model

  • The model is billed as deployable on a single GPU starting from RTX 4090 or equivalent with a 10M-token context window.

    @Pokee_AI on X · attributed

    the first real 10M-token context frontier-class model deployable on a single GPU starting from RTX 4090 or equivalent

  • Pokee AI claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill.

    @Pokee_AI on X · attributed

    It claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill

  • The architecture is described as proprietary non-decoder-only with some weights fine-tuned from Qwen3.6-27B.

    @Pokee_AI on X · attributed

    on a proprietary non-decoder-only architecture, with some weights fine-tuned from Qwen3.6-27B

Why it matters

Verified long-context performance on consumer-grade single-GPU hardware would reshape how teams scope retrieval, memory, and deployment costs for agent systems.

Limits and uncertainties

All performance and architecture details come from a single Pokee AI post on X with no independent verification cited in the packet.

Benchmark scores, throughput figures, and the claim to be the first real 10M-token frontier-class model are self-reported by the vendor.

Practical implications

Teams planning long-context agent stacks should treat single-GPU 10M-token deployment as unconfirmed until independent RULER and throughput results appear.

Operators on RTX 4090-class hardware may want to track weight release and deployment guidance before reallocating workloads away from multi-GPU setups.

What to watch

Independent RULER benchmark replication at the full 10M-token scale.

Publication of model weights, inference tooling, and reproducible prefill throughput measurements on RTX 4090-class hardware.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: X