Close Menu
TechurzTechurz
    What's Hot

    Kog is going deeper to squeeze more inference out of GPUs

    August 14, 2026

    Investors sue Selena Gomez alleging fraud tied to her mental health startup

    August 13, 2026

    Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.

    August 13, 2026
    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    Tech Pulse
    • Kog is going deeper to squeeze more inference out of GPUs
    • Investors sue Selena Gomez alleging fraud tied to her mental health startup
    • Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.
    • Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard
    • How a $250 million acquisition collapsed into allegations of fraud and forged signatures
    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    TechurzTechurz
    • Home
    • Tech Pulse
    • Future Tech
    • AI Systems
    • Cyber Reality
    • Disruption Lab
    • Signals
    TechurzTechurz
    Home - Opinion - Kog is going deeper to squeeze more inference out of GPUs
    Opinion

    Kog is going deeper to squeeze more inference out of GPUs

    TechurzBy TechurzAugust 14, 2026No Comments5 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Kog CEO Gaël Dellalleau at AMD AI Dev Day
    Share
    Facebook Twitter LinkedIn Pinterest Email

    The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.

    The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and NVIDIA H200 GPUs it used for its demo.

    Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch. 

    Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.

    Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.

    The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”

    This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open-sourced Laneformer 2B. 

    Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.

    Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper-level focus on GPU acceleration.

    Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, has nothing to do with his new one — other than his former cofounder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.

    Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.”

    As for hacking, the four-time finalist at at DEF CON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.”

    The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future.

    In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program.

    For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said. 

    When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

    deeper GPUs inference Kog squeeze
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleInvestors sue Selena Gomez alleging fraud tied to her mental health startup
    Techurz
    • Website

    Related Posts

    Opinion

    Investors sue Selena Gomez alleging fraud tied to her mental health startup

    August 13, 2026
    Opinion

    Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.

    August 13, 2026
    Opinion

    Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard

    August 12, 2026
    Add A Comment
    Latest Tech Pulse

    College social app Fizz expands into grocery delivery

    September 3, 20252,290

    12 Father’s Day E-Card Sites That Are Actually Good

    June 4, 202523

    SolarSquare in talks to raise up to $60M as India’s rooftop solar market draws major VC interest

    May 23, 202622
    Stay In Touch
    • YouTube
    • WhatsApp
    • Twitter
    • Pinterest
    • LinkedIn

    Techurz helps readers stay ahead of digital change with clear, practical, future focused technology intelligence written today,searched tomorrow.

    X (Twitter) Pinterest YouTube LinkedIn WhatsApp
    Company
    • About Us
    • Contact Us
    • Our Authors / Editorial Team
    • Write For Us
    • Advertise
    Policy
    • Editorial Policy
    • Privacy Policy
    • Terms and Conditions
    • Affiliate Disclosure
    • Cookie Policy
    • Disclaimer
    • DMCA
    Explore
    • AI Systems
    • Cyber Reality
    • Future Tech
    • Disruption Lab
    • Signals
    • Tech Pulse
    • Sitemap

    Join the Techurz Brief

    The future does not arrive suddenly.
    Stay ahead with fast, sharp tech signals.

    Type above and press Enter to search. Press Esc to cancel.