Nvidia’s $20 Billion Groq Bet Is Going Live Before the End of 2026. Here’s What It Means for Investors.
Nvidia says its Groq 3 LPX inference chip is now in production, with Nebius lined up as the launch customer before the year's end. The announcement caps an e...
Nvidia Corporation (NASDAQ:NVDA) said on Aug. 24 that Groq 3 LPX, its new low-latency artificial intelligence (AI) inference system, is in full production and will come online before the end of 2026. Nebius will be the first AI cloud to adopt it through the Nebius Token Factory platform.
The production announcement came exactly eight months after Nvidia signed a $20 billion licensing agreement with the chip designer, Groq Inc, on Dec. 24, 2025.
Missed Nvidia in 2009? This Rare Signal Is Flashing Again. In 2009, a "Double Down" signal flashed for a little-known chipmaker called Nvidia. For the first time in years, that same "Total Conviction" signal is flashing for a company 1/100th the size of Nvidia. Continue »
Nvidia's $20 billion Groq deal closed on Dec. 24, 2025
At the end of last year, Groq, a neocloud and semiconductor start-up focused on creating low-latency chips for AI inference, signed a non-exclusive licensing agreement with Nvidia.
At the same time, Nvidia hired Groq founder and CEO Jonathan Ross, president Sunny Madra, and much of the engineering staff. It held off on actually acquiring Groq as a company, however, and Groq remains an independent entity.
Nvidia reportedly paid $20 billion in cash for the assets.
Groq 3 LPX debuted at GTC with 256 LPUs and 35x efficiency gains
According to the company, the liquid-cooled Groq 3 rack contains 256 language processing units (LPUs) and is part of Nvidia's Vera Rubin platform. Nvidia's GPUs handle AI training and large-context "prefill" work, while LPX accelerates inference where speed is king.
Groq's design leans on on-chip memory and other design trade-offs that make it especially suited to inference, but bring limitations that make it less so for other tasks like training. Even with inference, however, they work best in conjunction with Nvidia's premier chips working alongside them, which is why the company is positioning LPX beside GPU racks rather than as a stand-alone replacement for them.
LPX can be paired with Nvidia's new Vera Rubin chips without customers changing their CUDA workflows -- the software that dominates as the standard base layer across the AI industry. CUDA is a critical reason -- maybe the critical reason -- Nvidia has dominated for as long as it has.
The company says it estimates that a system that pairs the LPX and Vera Rubin would process inference workloads with as much as 35 times the throughput for each megawatt of power used. Given that access to electricity is one of the most pressing constraints in the industry at the moment -- and likely will be for some time -- efficiency is paramount.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.