💻 Computing

AMD Says Its New GPU Is 34× Faster. Read Footnote 7.

AMD launched the Instinct MI455X at Advancing AI 2026, claiming 34 times the token throughput of its predecessor. The footnote says the benchmark ran FP4 precision on DeepSeek V4 Flash. Decompose the 34× and the real architectural improvement is closer to 6–8×. The number that actually matters for customers is buried in a different footnote: 30% more tokens per dollar than Nvidia’s Vera Rubin.

Thirty-four times. That was the number Lisa Su put on screen at AMD’s Advancing AI 2026 keynote on July 23 in San Francisco, describing the token throughput leap from the Instinct MI355X to the new Instinct MI455X. Applause rippled through the audience. Press releases went out, and headlines across the tech press echoed the claim uncritically, each outlet amplifying a number that sounds too clean to be a straight comparison. What almost nobody did was read Footnote 7. Almost nobody does. Ever.

Footnote 7 is twelve lines of dense methodology text at the bottom of AMD’s press release. It states, in full: “Based on measurements and calculations by AMD Performance Labs in July 2026, for the AMD Instinct MI455X GPU to determine measured token throughput at high, medium and low interactivity points run on Deepseek V4 Flash with FP4 serving compared to AMD Instinct MI355X GPU.”

Two phrases do all the work: the benchmark model is DeepSeek V4 Flash, a Mixture-of-Experts architecture where only a fraction of the model’s parameters activate on each token. And the serving precision is FP4, a 4-bit floating-point format that the MI455X supports natively in hardware and the MI355X does not. Together they produce the largest possible multiplier from a generational comparison, and AMD chose them deliberately. How much of 34× is real architecture? How much is measurement theater?

Decomposing the 34×

Start with precision. FP4 stores each model weight in four bits. Its predecessor ran FP8. Eight bits per weight. In the decode phase of large language model inference, where the GPU generates one token at a time and the bottleneck is loading model weights from high-bandwidth memory, halving the bit width of every weight approximately doubles the number of weights the memory system can deliver per clock cycle. Not architecture but arithmetic: a new generation adopts a new number format and the throughput benchmark immediately reflects it.

Contribution to the 34×: roughly 2×, with no silicon improvement required to capture it.

Next: architecture. CDNA 5, the MI455X’s underlying compute architecture, brings redesigned compute units, higher clock frequencies, improved instruction scheduling, and wider internal datapaths compared to CDNA 4 in the MI355X. Generational GPU architecture improvements in the data center segment have historically delivered between 1.5 and 2.5 times the raw compute throughput, a range that encompasses everything from AMD’s own CDNA 3-to-4 transition to Nvidia’s Ampere-to-Hopper leap, and that range has held remarkably steady across process node shrinks, packaging innovations, and memory interface changes for the past decade of accelerator generations, which makes any claim of 34× in a single generation a number that demands explanation rather than celebration. Nvidia’s own Hopper-to-Blackwell transition delivered approximately 2.2× on comparable dense workloads. AMD’s CDNA 4-to-5 jump likely falls in a similar band, with the exact figure depending on workload shape, memory pressure, and compiler maturity.

Architecture contributes somewhere between 2 and 2.5× to the total, depending on how much of the compute pipeline the workload actually saturates.

Memory bandwidth comes next: its predecessor shipped with eight stacks of HBM3E delivering approximately 6.4 TB/s of bandwidth. AMD has not disclosed the MI455X’s memory specifications, but the combination of CDNA 5 packaging and what is almost certainly HBM4 or an expanded HBM3E configuration suggests a bandwidth number in the 10 to 12 TB/s range. For inference decode, where memory bandwidth determines throughput more directly than compute, a 1.6 to 1.9× bandwidth increase translates almost linearly into faster token generation.

Memory bandwidth contributes an estimated 1.6 to 1.9× to the total, and on MoE models where weight loading dominates, this factor carries disproportionate weight.

Model selection is the final lever, and it may be the most consequential: DeepSeek V4 Flash is a Mixture-of-Experts architecture. In an MoE architecture, each token activates only a small subset of the model’s total parameters. Active parameter counts might be 30 to 50 billion even though total model size exceeds a trillion parameters, which means the ratio of memory traffic to compute is extremely high: the GPU spends most of its time loading weights, not multiplying them. Any improvement to memory bandwidth or weight compression (like switching from FP8 to FP4) produces a disproportionately large throughput gain on MoE models compared to dense models like Llama 3 70B, where compute and memory demands are more balanced. Running the same benchmark on a dense 70-billion-parameter model at FP8 would produce a substantially smaller multiplier.

Multiply the components together: 2× (FP4) × 2.2× (architecture) × 1.75× (memory bandwidth) = approximately 7.7×. Where does the remaining gap to 34× come from? Likely a combination of compiler and software stack improvements in ROCm (AMD’s MI355X software was notoriously less optimized at launch than its successor), FP4 quantization producing superlinear throughput gains on sparse MoE attention patterns, and the possibility that the MI355X baseline measurement used an earlier, less-optimized serving stack. AMD’s footnote does not specify the MI355X’s software version, serving framework, or whether FP8 quantization was even applied to the MI355X measurement.

The Number That Actually Matters

Buried in Footnote 1 is the competitive metric: AMD Helios delivers up to 30% more inference tokens per dollar than a leading competitive rack-scale solution, identified in the fine print as the Nvidia Vera Rubin NVL72. The benchmark ran Kimi K2 Thinking, a thinking model from Moonshot AI with a 32K input and 8K output window. AMD calculated results across low, medium, and high interactivity operating points, then computed tokens per dollar using “hourly pricing projection of system GPUs based on market conditions.”

This comparison is honest because the variables are controlled: same rack form factor, same workload class, same pricing methodology. And the result, 30%, is a fraction of 34×. It reflects the real economic advantage a customer can expect when choosing AMD Helios over Nvidia Vera Rubin for inference at rack scale. Thirty percent is meaningful at gigawatt deployment volumes, where even single-digit cost advantages compound into hundreds of millions of dollars in annual savings, a figure that dwarfs the engineering cost of qualifying a second GPU vendor and rewriting the inference serving stack. It is not the kind of number that makes keynote audiences gasp, and that tension between the number that sells GPUs and the number that actually predicts economic outcome is the entire story of semiconductor marketing in the AI era.

Anthropic’s 2-Gigawatt Order

A partnership announcement that landed the same week tells a larger story. Anthropic, the $965 billion AI company behind Claude, committed to deploying up to 2 gigawatts of AMD Instinct MI455X GPUs in Helios rack-scale solutions. Two gigawatts is an extraordinary number, and putting it in concrete terms reveals the scale of what Anthropic is building:

Each AMD Helios rack contains 72 MI455X GPUs and 18 sixth-generation EPYC Venice CPUs. Industry-standard power density for a GPU-dense inference rack runs between 100 and 130 kilowatts, accounting for GPUs, CPUs, networking, and cooling overhead. Nvidia’s comparable GB200 NVL72 rack draws approximately 120 kilowatts. Using that figure as a proxy for Helios:

MetricEstimate
Total power commitment2,000 MW
Estimated rack count~16,700 Helios racks
Total MI455X GPUs~1.2 million
GPU cost at $25K each~$30 billion
Total system cost (GPU + CPU + networking)~$45–55 billion

Those numbers carry significant uncertainty. AMD has not disclosed MI455X pricing. The $25,000-per-GPU estimate is based on the MI300X’s list price of $10,000 to $15,000 adjusted upward for the die size increase, HBM4 memory, and the generational architecture premium that Nvidia has trained the market to accept. If MI455X pricing lands closer to $35,000, the GPU cost alone rises to $42 billion. That phrase “up to” preserves optionality on both sides: Anthropic is not obligated to deploy the full 2 gigawatts, and AMD is not obligated to deliver it on any specific timeline. Still: 2 gigawatts.

For context, Anthropic closed a $65 billion Series H round in May 2026 at a $965 billion post-money valuation. Annualized revenue exceeds $47 billion. SK Group Chairman Chey Tae-won, speaking at the South Korea AI summit on July 25, noted that Anthropic was “regularly checking on future chip availability as it moves beyond developing AI models to building computing capacity.” Anthropic is transitioning from renting compute to owning infrastructure, and AMD just became its preferred silicon vendor for inference.

The Competitive Landscape at Rack Scale

AMD’s AAI 2026 roster of partners tells a story about where Nvidia’s pricing power is weakest. OpenAI expects to bring Helios racks online beginning in Q4 2026, with deployments accelerating through 2027. Meta is validating sixth-generation EPYC platforms in its labs and has begun testing Helios workloads. Cerebras announced a collaboration to combine its ultra-low-latency compute with AMD Helios’s high-throughput rack-scale infrastructure for inference serving, an architecture that pairs Cerebras’s wafer-scale chip for the fastest possible individual responses with AMD’s rack-scale throughput for the highest possible aggregate token volume, attacking both dimensions of inference economics simultaneously.

Every one of these partnerships targets inference, not training. For training, Nvidia’s NVLink and InfiniBand interconnect ecosystem remains effectively unmatched, and switching costs for multi-thousand-GPU training clusters are prohibitive. Inference is different. Fundamentally. Inference is embarrassingly parallel at the request level, far less dependent on tight GPU-to-GPU communication, and running 24 hours a day at volumes where a 30% cost-per-token reduction compounds fast. Nvidia’s approximately 74% share of the AI chip market, estimated by The Information, is concentrated in training. Inference is where the market fractures.

AMD’s roadmap extends through 2030: MI500 Series GPUs in 2027, MI600 in 2028, Zen 7 CPUs in 2028, and Zen 8 in 2030. Next-generation Helios 500 and Helios 600 rack-scale solutions will integrate these future chips with Pensando networking. For hyperscale buyers, the message is clear: AMD is not offering a one-generation alternative to Nvidia. It is building a multi-year parallel infrastructure track, and it is pricing that track below Nvidia’s.

The Footnote 8 Problem

AMD’s third headline metric introduces a different measurement issue. The MI350P, positioned as a drop-in inference accelerator for existing infrastructure, reportedly delivers 4.2 times more tokens per second per dollar than “the competition.” Footnote 8 identifies the competitive reference: the Nvidia RTX PRO 6000, a workstation-class GPU built on the Blackwell architecture.

The MI350P server was priced at $327,238 while the RTX PRO 6000 server was priced at $265,928 based on an OEM list price captured on July 16, 2026, and the benchmark ran Llama 3.3 70B Instruct at FP8 across concurrency levels from 1 to 512 with the 4.2× figure representing the peak per-concurrency ratio rather than the median or average across all tested configurations.

Categorical mismatch: the RTX PRO 6000 is a workstation GPU designed for professional visualization, local AI inference, and creative workflows. It is not Nvidia’s data center inference accelerator. Comparing the MI350P to an H200, an L40S, or even a B200 would produce a different and likely smaller multiplier. AMD chose the comparison that generated the largest headline number, which is exactly what every semiconductor company does, and exactly what the footnotes exist to reveal.

Limitations

This analysis relies on estimation where AMD has not disclosed specifications. The MI455X’s memory bandwidth, die size, compute unit count, and TDP have not been publicly confirmed. Our decomposition of the 34× claim into component factors uses historical generational improvement ranges and architectural reasoning, not measured hardware data. Our 2-gigawatt deployment cost estimates assume a per-GPU price that AMD has not announced and a rack power draw extrapolated from Nvidia’s comparable system. If AMD’s actual rack power is significantly different, the rack count and total GPU figures shift accordingly.

We also do not know whether the MI355X baseline in the 34× benchmark was measured with a fully optimized ROCm stack or an earlier, less-mature software version. AMD’s software ecosystem has historically lagged Nvidia’s CUDA by six to twelve months at new product launch, meaning the MI355X’s measured throughput at the time of the MI455X comparison may understate what the MI355X can achieve with current software. If the MI355X baseline was artificially low, the 34× figure is inflated beyond what even the precision and architecture changes explain, and AMD’s decision not to disclose the baseline software configuration leaves that question permanently unresolvable from public data alone.

The Strongest Counterargument

AMD would argue that the 34× figure is exactly what customers experience when upgrading from an MI355X rack to an MI455X rack running production inference. The customer does not care whether the throughput gain comes from FP4, architecture, memory bandwidth, or software maturity. They care about tokens per second on the workloads they actually run, and DeepSeek V4 Flash is representative of the MoE models that dominate frontier inference deployment in 2026. By this reasoning, decomposing the 34× is an academic exercise that misses the practical point: if you buy the new GPU and run your production model, you get 34 times the throughput, no decomposition needed, no footnotes to read.

That argument has real force, but the counterargument to the counterargument is that the 34× does not hold across model types, precision formats, or workload shapes. A customer running Llama 3 405B at FP8 on the MI455X will not see a 34× improvement over the MI355X. The 30% tokens-per-dollar advantage over Nvidia, measured on a different model at a different precision, is far more representative of what a diverse production fleet will experience. And the market, which sent AMD’s stock down 2.29% on announcement day despite the 34× headline and the Anthropic partnership and the OpenAI commitment and the Meta validation and the six-year roadmap, appears to have decided that 30% more tokens per dollar is the number it should price, not 34×.

What You Can Do

If you are evaluating GPU procurement for inference at scale, focus on the tokens-per-dollar metric, not the generational throughput multiplier. AMD’s 30% advantage over Nvidia Vera Rubin on Kimi K2 Thinking (Footnote 1) is the number closest to what you will experience in a mixed production environment. Request benchmark results on your specific model architecture, at the precision you intend to deploy, before committing procurement dollars.

If you hold AMD stock, understand that the market priced in a significant product launch weeks before AAI 2026 and reacted negatively to the details. The 2-gigawatt Anthropic partnership and the OpenAI and Meta collaborations represent real revenue pipeline, but the “up to” qualifier means deployment timing and volume are uncertain. Watch Q3 and Q4 2026 data center revenue for evidence that these partnerships are converting to purchase orders.

If you are a journalist covering semiconductor launches, read the footnotes. Every 34× has a Footnote 7. The methodology disclosures exist precisely because the headline numbers are engineered to be as large as possible within the bounds of technical accuracy. Competitive comparisons like the 30% in Footnote 1 are always more informative than generational multipliers like the 34× in Footnote 7. And the same principle applies to every GPU launch from every vendor, including Nvidia. applies to every GPU launch from every vendor, including Nvidia.

The Bottom Line

AMD’s Advancing AI 2026 event delivered a genuine product milestone. The MI455X is a major architectural step forward. The Helios rack-scale system represents AMD’s most credible challenge to Nvidia’s data center dominance, and the Anthropic, OpenAI, and Meta partnerships signal that hyperscale buyers are diversifying their inference supply chains. None of that is in question, and the partnerships alone would make this the most significant AMD data center product launch since the original EPYC disrupted Intel’s server monopoly in 2017. What is in question is the 34× number, which conflates a precision format transition (FP4 replacing FP8), architectural improvement, memory bandwidth gains, a software maturity gap between product generations, and the selection of a maximally memory-bandwidth-sensitive MoE model into a single headline multiplier. Multiply out the real architectural and memory gains and you get roughly 6 to 8×, with the remainder attributable to the FP4 precision shift and workload selection. Thirty percent more tokens per dollar versus Nvidia’s best rack is the honest competitive metric, and it is the number that will determine whether AMD’s inference business scales. It is the better number. But it does not make the keynote.