AMD’s ROCm.AI leans on Claude and Codex to auto-optimize inference, claims 38% gains
ROCm.AI plugs into coding assistants and can profile, reconfigure, even generate kernels on demand for Instinct GPUs.

AMD unveiled ROCm.AI at its Advancing AI event in San Francisco, aiming to help developers get better inference performance without hand-tuning. The company says Hyperloom-driven optimization delivered a 38% boost over baseline on newly launched Helios racks, with support for tools used by Claude Code, OpenAI’s Codex, Antigravity, and Cursor.
AMD just tried to knock a hole in the long-running “CUDA moat” argument, but with a twist: instead of asking developers to become GPU-kernel wizards, it is building an optimization workflow that plugs into frontier-model coding assistants. At its Advancing AI event in San Francisco this week, AMD unveiled ROCm.AI and paired it with an automated performance optimization system called Hyperloom. AMD says this can boost model performance by 38% over baseline when tested on its newly launched Helios racks.
The key detail is that ROCm.AI is not positioned as a new programming language or a vague AI wrapper. AMD’s pitch is operational: ROCm.AI connects to existing code assistants running on frontier models and gives them the tools and documentation to deploy, debug, profile, and optimize models and serving frameworks for AMD Instinct hardware. The workflow matters because it is designed to convert a prompt like “optimize MiniMax M3 with Hyperloom” into a concrete performance engineering loop, including spinning up an inference server in a Docker container, running benchmarks for baseline performance, profiling to find bottlenecks, and adjusting the configuration. AMD also says it can generate custom CPU kernels on the fly, depending on what the profiling reveals.
Why is AMD betting on this? Because the CUDA-versus-non-CUDA story has been changing. The “House of Zen” has struggled with a perception problem, namely that AMD chips are less capable because they do not run CUDA. AMD argues that the perception has not kept pace with reality over the past few years, when frameworks like PyTorch and JAX made it possible, for the most part, to “write once and, for the most part, run anywhere” without developers touching CUDA or AMD’s ROCm and HIP libraries.
But ROCm.AI’s implicit point is uncomfortable for anyone who assumes “it runs” equals “it flies.” High-level framework portability does not guarantee peak performance. Getting the most out of silicon typically requires low-level programming interfaces and, in practice, hand tuning GPU kernels and optimizing general matrix-matrix multiplication (GEMM) routines. That skill gap is real for many teams. Yet AMD is leaning on a counterintuitive asset: frontier models are increasingly able to translate instructions into hardware-aware code paths. AMD corporate VP of AI software and solutions Anush Elangovan said AMD publishes not just the ISA spec, but the machine-readable ISA, and therefore “the frontier models are very, very capable of programming to AMD’s hardware.”
ROCm.AI is essentially AMD’s attempt to productize that capability into a repeatable pipeline. When a code assistant is asked to optimize for a specific workload, Hyperloom orchestrates the performance work in a way that can be consumed without expert intervention. Elangovan described a cycle that starts with establishing baseline performance, then profiling to identify bottlenecks, then tweaking configuration or generating custom CPU kernels. The pitch is that this lets teams “eke out the maximum performance,” while making it easier to consume, debug, profile, and deploy.
There is also a second layer to AMD’s strategy that is easy to miss if you focus only on the software mechanics. AMD says it is using its “close relationship with AI model houses like OpenAI and Anthropic” to ensure their models are trained to better understand both the inner workings of the hardware and the software. Elangovan framed it as more than asking a frontier model to generate a kernel. “We’re working deeply with frontier model companies so that they natively speak AMD programming,” he said.
That matters because it reframes the competitive battleground. It is no longer just about whether developers will adopt AMD tooling. It is about whether the assistants developers already rely on can produce correct, high-performance code targeting AMD hardware, without requiring humans to learn the edge cases of ROCm and HIP. In other words, AMD is trying to reduce the “friction cost” of switching targets from CUDA-first workflows to ROCm-first workflows by moving the expertise into the optimization system and the training loop behind the model assistants.
Finally, ROCm.AI is also positioned for distribution in the tools developers touch daily. AMD says that beyond a built-in command-line interface, ROCm.AI will be offered as a plug-in for coding assistants including Anthropic’s Claude Code, OpenAI’s Codex, Google’s Antigravity, and Cursor. For executives and infrastructure leaders, that distribution channel is the whole game: it determines whether the technology becomes a niche research experiment or a practical lever teams can pull in production pipelines.
So what should peers in AI infrastructure and semiconductor-adjacent leadership teams take from this? ROCm.AI signals that performance differentiation is increasingly being routed through the software layer where frontier models operate. If AMD can consistently deliver big performance gains, like the 38% over baseline claim on Helios racks, while making optimization as accessible as prompting an assistant, the “CUDA moat” narrative becomes harder to sustain. And that puts pressure on everyone running AI compute stacks, from cloud operators to model deployment teams, to rethink where optimization expertise lives and how fast it can be applied at scale.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Finland powers through wind and solar lulls with the world’s largest sand battery
A small town in Finland is using the world’s largest commercial sand battery to solve intermittency.

Nvidia locks SK Hynix memory supply to power its $500B AI push
Nvidia is securing high-bandwidth memory access for its AI systems, and that supply chain control matters more than ever.

Beijing-backed money quietly ties DeepSeek, Zhipu AI, Unitree, and CXMT
The SCMP report shows state capital is reshaping “VC-style” funding across China’s frontier tech ecosystem.
