← All company news

ADVANCED MICRO DEVICES INC · News & developments

AMD Milks Up to 38% Extra Inference Throughput from Existing Instinct Silicon via ROCm Software Updates

Source: AMD Technical BlogsBullish

AMD published its MLPerf Inference 6.1 results across six model families, reporting throughput and latency improvements on unchanged Instinct MI355X hardware alongside submissions for its MI350X and new MI350P PCIe card. Continued ROCm software optimization lifted eight-GPU GPT-OSS-120B throughput by 28% in Offline and 38% in Server scenarios, while reducing Wan-2.2 SingleStream latency by 70%. Separately, Crusoe delivered record aggregate token throughput at 512 GPUs, reaching 5.75 million Offline tokens per second on GPT-OSS-120B and 2.90 million on DeepSeek-R1.

Why this matters

The commercial bottleneck in AI inference is rarely raw peak silicon math; it is software runtime maturity and cluster communication overhead. When AMD can squeeze 38% more throughput out of identical MI355X silicon simply through ROCm driver and attention kernel refinements, it changes customer total cost of ownership without consuming incremental wafer capacity at TSMC or requiring new tape-outs. Delivering 95% scaling efficiency from eight to 72 GPUs and record throughput on Crusoe's 512-GPU deployment chips away at Nvidia's primary competitive moat: the perception that deploying non-CUDA clusters incurs a software management tax. If these synthetic gains transfer directly into enterprise serving frameworks, AMD gains pricing power to defend Data Center margins ahead of its planned MI400 HBM4 generation.

Read the original source — AMD Technical BlogsExplore ADVANCED MICRO DEVICES INC research →