News

Kog develops faster GPU inference for agentic workflows

What’s the deal? French startup KogDealroom has a profile for this one. Try Dealroom → is developing a hardware-software inference stack designed to make long, iterative AI workloads run substantially faster. The company argues that GPUs may be better suited to agentic workflows than commonly assumed, provided the surrounding system reduces the communication and synchronisation overhead that can interrupt generation.

How it works: Kog says its inference engine is built to deliver up to 3,000 tokens per second per request, compared with approximately 100 tokens per second for ChatGPT. Its system combines a low-latency engine with a parallel architecture, including the Kog Communication Library, which is intended to improve tensor-parallel scaling across high-end GPUs, and LaneFormer, which delays inter-device communication by one layer so computation can continue without a synchronisation pause. The company describes the approach as a hardware-software co-design that keeps GPU compute running continuously.

Why it matters: Agentic coding and reasoning systems often need to draft, test, lint and refine outputs repeatedly. Kog says that reducing each iteration from minutes to seconds could make deeper reasoning practical at product speed, with up to 30 refinement loops in the time a standard stack completes one. Its website positions the technology for AI coding agents and other real-time agentic workflows, while TechCrunch reports that the company is challenging the idea that GPUs are inherently poorly suited to these use cases.

Read more: Kog · TechCrunch on X

More top stories