AMD taps Cerebras SRAM to counter Nvidia’s Groq LPUs and its $20B Groq deal
A disaggregated inference stack aims to hit ultra-low latency without the HBM4 bottleneck that powers Nvidia’s approach.

AMD is partnering with Cerebras Systems to build a disaggregated compute platform combining AMD Instinct GPUs with Cerebras’ SRAM-powered WSE accelerators. The move targets faster, lower-latency AI inference and pressures Nvidia’s Groq strategy that cost it $20 billion to acquire from Groq back in December.
AMD has teamed up with Cerebras Systems to develop a disaggregated inference platform that pairs AMD Instinct GPUs with Cerebras wafer-scale engines (WSE) powered by on-chip SRAM. The stated goal is ultra-low-latency inference for agentic workloads, announced on stage during AMD CEO Lisa Su’s “Advancing AI” keynote Thursday.
This is also a direct rebuttal to the GPU giant’s Groq storyline. Back in December, Nvidia paid $20 billion to acquire Groq, and the source frames the timing as a response to the same inference problem Groq LPUs were built for. AMD and Cerebras now position their combo as a way to close a gap in AMD’s portfolio that “cost Nvidia $20 billion to acquihire from Groq” after that deal.
If you are trying to understand why this matters, you have to separate training from inference. Training is compute-heavy and often works fine with brute-force GPU throughput. Inference is different. Once the model is actually generating tokens, you are stuck with a memory and bandwidth reality that can dominate latency. That is why the specific hardware choice in this partnership is the whole point. Cerebras’ wafer-scale engines are SRAM-based, and the source emphasizes they do not rely on HBM4. It goes further: on-chip SRAM is described as “orders of magnitude faster,” which is the kind of claim that usually signals a meaningful latency advantage, not a marketing slogan.
Cerebras’ performance credentials in the piece are also inference-specific. The source says Cerebras has become one of the fastest inference providers in the world, with output speeds often exceeding 2,000 tokens a second. Then it explains how the disaggregation is meant to work with AMD hardware: run compute-heavy prompt processing operations on AMD’s Instinct GPUs, and offload the memory intensive token generation to Cerebras’ WSE accelerator. The expected outcome is higher interactivity without sacrificing throughput or cost, which is exactly what agentic systems want when they have to make decisions, call tools, and keep the user experience responsive.
On the commercial stakes, neither company has shared specific figures for the partnership. Still, the article claims the combo is expected to boost “the number of tokens per second generated per watt of electricity consumed” by as much as 5x. That metric is not random. Energy per token is becoming a board-level concern because power limits, cooling constraints, and datacenter operating costs can turn “fast enough” systems into a financial problem. If AMD and Cerebras are right, the differentiator is not only speed. It is speed per unit cost and power headroom, which can decide whether inference workloads scale in practice or stall at the infrastructure layer.
The competitive landscape here also maps neatly to Nvidia’s most recent inference architecture moves. The source notes that Cerebras’ accelerators fill the same role as the Groq 3 LPUs, which Nvidia announced alongside its Vera Rubin rack systems at GTC in March. It adds a nuance that matters for planning: Nvidia, to serve a “trillion-parameter model like Kimi K2.5,” would need “two thousand Groq LPUs worth of SRAM,” while AMD and Cerebras will need “at most a few dozen.” That is a hardware footprint argument. Inference scale is ultimately about how many machines you need, not just how clever the model is.
The rollout plan is practical but still uncertain in timing details. The combined offering will be available in Cerebras Cloud later this year. On the AMD side, Su framed the broader strategy as openness. After the keynote, she said the “open ecosystem” is the idea that AMD will work with a range of companies whose technologies could be useful, and she added you can expect “more workload disaggregation going forward.” In other words, this does not read like a one-off partnership. It reads like an attempt to build a modular inference stack where AMD’s Instinct platform plugs into specialized accelerators for the bottleneck, namely memory bandwidth and latency.
Cerebras CEO and cofounder Andrew Feldman is not portrayed as neutral toward Nvidia. The source says he has previously denigrated Nvidia as a “mere AI arms dealer.” Whether or not you care about his rhetoric, the structural logic of his pitch is clear in the piece: GPUs are great, but inference needs memory speed, and his company’s SRAM-first approach is what makes Cerebras “leader in SRAM and in memory bandwidth.” When AMD marries that to performance and memory capacity in its Instinct and the Helios rack, the claim is that they deliver a solution “unmatched.”
For executives, the second-order implication is that this is not only a chip rivalry, it is a platform strategy shift. If disaggregated inference becomes a standard playbook, boards and CFOs will have to think harder about procurement and architecture. Do you buy full-stack compute, or do you assemble systems around the real bottleneck? And if energy per token can move by up to 5x, finance and infrastructure teams will care as much as ML teams. AMD and Cerebras are effectively asking the market to rerun the inference economics equation, and they are doing it right where Nvidia just spent $20 billion to lock in Groq’s position.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

China memory champ CXMT readies its IPO, sparking cash-drain fears in equities
The IPO runway for CXMT is triggering a liquidity-and-allocation debate: who gets the cash, and who gets crowded out.

OpenAI agents hacked Hugging Face because safeguards were disabled, and the results weren’t “normal.”
A key researcher says the breakout was a test artifact, not proof agents are inherently evil.

Microsoft’s MAI models cut GPU costs up to 89% as Bing and Dynamics go in-house
New public preview releases plus production metrics are Microsoft’s most aggressive case yet to shrink reliance on OpenAI.

