Skip to content
Breaking:

French Startup Kog Targets GPU Optimization to Boost LLM Inference Speeds

Led by a former cybersecurity researcher, the 11-person team is optimizing low-level hardware routines on standard Nvidia and AMD GPUs to challenge custom silicon.

By The Company Wire4 min read
Share
Kog — French Startup Kog Targets GPU Optimization to Boost LLM Inference Speeds
Kog — French Startup Kog Targets GPU Optimization to Boost LLM Inference Speeds. Photo: TechCrunch AI.

As custom silicon vendors like Cerebras attract public market interest following a May debut, French startup Kog is taking an alternative path by attempting to extract dramatically higher performance out of conventional datacenter hardware. The Paris-based company is building software designed to accelerate artificial intelligence inference directly on standard processors already widely deployed in corporate data centers, such as Nvidia's H200 and AMD's MI300X.

The startup gained initial industry visibility in May after publishing a technical preview that demonstrated 3,000 per-request tokens per second on its open-sourced 2-billion parameter model, Laneformer 2B. Chief Executive Officer Gaël Delalleau told TechCrunch AI, which first reported the company's developments, that the demonstration yielded roughly 200 commercial leads from organizations eager to bypass current throughput limits in AI workloads.

Initial prospective client feedback indicated that enterprise users are reluctant to fine-tune smaller, lightweight models. Consequently, Kog shifted its primary technical focus toward accelerating larger frontier models. The company is initially targeting software engineering environments and prompt-based application generation platforms, where latency directly impedes developer workflows and limits revenue opportunities for platform operators.

Delalleau, a solo founder who previously co-founded Stribe and competed four times as a finalist in DEFCON’s Capture the Flag competition, brings a background in solid-state physics from École Polytechnique and offensive cybersecurity. He relies on low-level binary analysis and assembly language reverse-engineering techniques to squeeze extra bandwidth out of standard GPU architectures, countering industry arguments that GPUs are inherently inefficient at token decoding.

Kog’s deep hardware-level software optimization places it alongside entities such as Stanford University’s Hazy Research, distinguishing it from higher-level hardware-agnostic software frameworks like France's ZML. However, this intensive engineering approach requires significant manual labor. With an 11-person team spending weeks or months dissecting individual chip architectures, the startup currently faces limits on how many hardware models it can support concurrently.

To support broader chip compatibility over time, Kog plans to eventually integrate its optimization workflows into automated agent pipelines. Co-led in its seed round by Varsity VC—a firm co-founded by Delalleau’s former Stribe partner Kamel Zeroual—Kog is backed by Bpifrance, Scaleway, and the French Tech 2030 program. The company aims to demonstrate a tenfold speed improvement on a major language model by September, which it expects will pave the way for a Series A funding round.

Sources

  1. TechCrunch AI

Company: Kog

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.