
AMD taps Cerebras SRAM to counter Nvidia’s Groq LPUs and its $20B Groq deal
A disaggregated inference stack aims to hit ultra-low latency without the HBM4 bottleneck that powers Nvidia’s approach.
By Yousef Al-Zahrani·· 4 min

Curating from trusted global sources…
2 briefings · “rega”

A disaggregated inference stack aims to hit ultra-low latency without the HBM4 bottleneck that powers Nvidia’s approach.

AMD’s new Cerebras pact is a direct bet that AI inference should be split across specialized chips, not one dominant rack.