Kog claims 30x faster LLM inference on existing GPUs with 3,000 TPS demo
French startup Kog claims that software optimization can extract 30x faster LLM inference from conventional GPUs, presenting a demo of 3,000 tokens per second on the open-sourced Laneformer 2B model. The company now targets large model acceleration via memory bandwidth optimization, a shift from fine-tuning smaller models. With 200 tangible business leads and plans to demonstrate a 10x speedup on a major model by September, Kog is betting on deep GPU engineering. The catch: the demo used a purpose-built 2B parameter model, so scaling to large LLMs remains unproven. If successful, this could lower inference costs and improve latency for AI providers without requiring hardware upgrades, though the promise hinges on bridging the gap to large model sizes.
Kog claims 30x faster LLM inference on existing GPUs with 3,000 TPS demo
French startup Kog claims software optimization can unlock 30x faster LLM inference on existing GPUs. Its demo hit 3,000 tokens per second on the open-sourced Laneformer 2B model, but the promise for large LLMs still faces a gap.
Key takeaway
Software-only optimization could redefine GPU economics for inference, but Kog must first demonstrate its 30x claim on large models rather than a 2B parameter demo.
What happened
According to TechCrunch, French startup Kog claims that software optimization can unlock 30x faster LLM inference on existing GPUs, with a demo achieving 3,000 tokens per second on the open-sourced Laneformer 2B model. The company, led by solo founder Gaël Delalleau, argues that GPUs are misunderstood and that newer hardware's memory bandwidth can be exploited for faster decoding.
Kog has shifted focus to accelerating larger models after learning customers aren't prepared to fine-tune small ones, and expects to demonstrate a 10x speedup on a major model by September. The startup reports 200 tangible business leads, with software engineering as the likely first use case, and plans to raise a Series A after proving customer traction.
Evidence
French startup Kog claims software optimization can unlock 30x faster LLM inference on existing GPUs.
TechCrunch AI · attributed
French startup Kog claims software optimization can unlock 30x faster LLM inference on existing GPUs.
Kog's demo achieved 3,000 tokens per second on the open-sourced Laneformer 2B model.
TechCrunch AI · attributed
Its demo hit 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open sourced Laneformer 2B.
Kog has 200 tangible business leads from its tech preview.
TechCrunch AI · attributed
"We had 200 tangible business leads," CEO Gaël Delalleau told TechCrunch.
Why it matters
If Kog scales its approach, AI providers could cut inference costs and latency without new hardware, potentially reshaping infrastructure spending and competitive dynamics in the AI stack.
Limits and uncertainties
The demo used a purpose-built 2B parameter model, and the promise for large LLMs still faces a gap.
The approach is hands-on and time-consuming, requiring weeks or months per GPU, limiting chip support.
Practical implications
AI providers may consider evaluating Kog's inference engine for latency-sensitive applications, but should validate performance on their specific large models.
Software engineering workflows, where response times are critical, could be early adoption targets.
What to watch
Kog's planned September demonstration of a 10x speedup on a major model.
Whether the 30x claim extends to large models beyond the 2B demo.
Adoption by enterprises with existing GPU infrastructure.