Kog is going deeper to squeeze more inference out of GPUs
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
In the competitive world of AI inference, companies Cerebras and Kog are striving to maximize the potential of GPUs. Cerebras gained attention in its IPO debut, while Kog, a French startup, is betting on tapping more power from conventional GPUs. Kog's demonstration in May revealed that "extremely fast single-request decoding" is achievable on datacenter GPUs, such as AMD MI300X and NVIDIA H200.
While some might have been skeptical about this approach extending to laptop GPUs, interest in Kog's promise grew, with 200 business leads reported. CEO Gaël Delalleau anticipates software engineering as the initial use case, as users often face delays when relying on AI workflows. Kog's software, the Kog Inference Engine (KIE), aims to accelerate larger models to meet market demand, aiming for "30x faster LLM inference."
Recent demos showed a remarkable 3,000 tokens per-request per second with a small 2 billion-parameter model called Laneformer 2B. Kog's CEO sees newer GPUs as having higher memory bandwidth that can be unlocked through software optimization. Kog's approach is more comparable to Stanford University lab Hazy Research in focusing on GPU acceleration at a deeper level.
The startup's unique background stems from its CEO's solid-state physics background and cybersecurity expertise, which have shaped his mindset for maximizing GPU capabilities. With a team of 11 people, Kog's focus on individual GPUs limits its scope but could lead to broader support for more chips and models in the future. While proving the method's effectiveness on LLMs is crucial for Kog's growth, the startup has already secured support from Scaleway and France's Bpifrance and French Tech 2030's program.
Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.