Pokee AI launches Isaac 28B with 10M-token context on one GPU
Pokee AI has unveiled Pokee-Isaac 28B, a 28-billion-parameter agentic model positioned as the first frontier-class system with a 10-million-token context window deployable on a single GPU starting from hardware comparable to an NVIDIA RTX 4090. The company cites proprietary non-decoder-only architecture, 93.3% RULER benchmark performance at 10M tokens, prefill throughput up to 137,000 tokens per second, and weight fine-tuning from Qwen3.6-27B. For teams building long-context agents without multi-GPU infrastructure, those claims describe a sharply different cost and operations envelope if validated outside vendor testing. Every metric and architectural detail currently comes from a single Pokee AI post on X, so independent replication and third-party benchmarking are still required before treating the release as settled production guidance.
Pokee AI launches Isaac 28B with 10M-token context on one GPU
Pokee AI released Pokee-Isaac 28B, a 28B-parameter agentic model it bills as the first real 10M-token context frontier-class model deployable on a single GPU starting from RTX 4090 or equivalent. It claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill on a proprietary non-decoder-only architecture, with some weights fine-tuned from Qwen3.6-27B.
Key takeaway
If substantiated, Pokee-Isaac 28B could let single-GPU setups run 10M-token agent workloads that today typically demand much heavier hardware.
What happened
Pokee AI announced Pokee-Isaac 28B on X, describing a 28-billion-parameter agentic model designed for extreme long-context use.
The post claims 10M-token context on one GPU from RTX 4090-class hardware, 93.3% RULER at that scale, 137K tokens/s prefill, a proprietary non-decoder-only design, and fine-tuning from Qwen3.6-27B weights.
Evidence
Pokee AI released Pokee-Isaac 28B as a 28B-parameter agentic model.
@Pokee_AI on X · attributed
Pokee AI released Pokee-Isaac 28B, a 28B-parameter agentic model
The model is billed as deployable on a single GPU starting from RTX 4090 or equivalent with a 10M-token context window.
@Pokee_AI on X · attributed
the first real 10M-token context frontier-class model deployable on a single GPU starting from RTX 4090 or equivalent
Pokee AI claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill.
@Pokee_AI on X · attributed
It claims 93.3% RULER at 10M tokens and up to 137K tokens/s prefill
The architecture is described as proprietary non-decoder-only with some weights fine-tuned from Qwen3.6-27B.
@Pokee_AI on X · attributed
on a proprietary non-decoder-only architecture, with some weights fine-tuned from Qwen3.6-27B
Why it matters
Verified long-context performance on consumer-grade single-GPU hardware would reshape how teams scope retrieval, memory, and deployment costs for agent systems.
Limits and uncertainties
All performance and architecture details come from a single Pokee AI post on X with no independent verification cited in the packet.
Benchmark scores, throughput figures, and the claim to be the first real 10M-token frontier-class model are self-reported by the vendor.
Practical implications
Teams planning long-context agent stacks should treat single-GPU 10M-token deployment as unconfirmed until independent RULER and throughput results appear.
Operators on RTX 4090-class hardware may want to track weight release and deployment guidance before reallocating workloads away from multi-GPU setups.
What to watch
Independent RULER benchmark replication at the full 10M-token scale.
Publication of model weights, inference tooling, and reproducible prefill throughput measurements on RTX 4090-class hardware.