Home » Blog » AMD and Nvidia Built Very Different CPUs for Agentic AI AMD and Nvidia Built Very Different CPUs for Agentic AI

AMD and Nvidia Built Very Different CPUs for Agentic AI

AMD spent two days in San Francisco arguing that the server CPU has stopped being a supporting part. The number carrying that argument: a data center CPU market worth roughly $220 billion by 2030, inside a total compute opportunity near $2 trillion.

Worth being precise about what that figure is. AMD’s slide states a CAGR above 50%. Assuming that rate covers the five years from 2025 through 2030, working backward from $220 billion implies a starting market of roughly $28 to $29 billion. It is a vendor forecast, not an independent consensus. At least one investment research view going into the event framed the opportunity nearer $120 billion.

The forecast was the headline. The footnotes were more interesting.

The unit of comparison quietly changed

Some of AMD’s 6th Gen EPYC headline claims are not stated in cores or SPECrate. They are stated in agents: agents per rack, agents per watt, agents per CPU dollar. One footnote compares “the most agents per rack at a 100kW power envelope.”

The same footnote explains how the count is produced. “Agent counts are estimates derived from available CPU thread resources used as a proxy under a consistent theoretical workload.”

So the agent numbers are thread counts in a better suit.

That is not an accusation. AMD published the methodology, which is more than most vendors bother with, and it publishes conventional and modeled workload results alongside. The two should not be conflated. The narrower point: using hardware threads as the numerator structurally favors high-thread-density designs. Final ranking still depends on the processor price, power limit, socket count and rack assumptions chosen for the model, all of which AMD also selects.

Nvidia frames the same product category around a different workload set and a different success metric.

Two designs, one thesis, 168 cores of disagreement

EPYC “Venice” leads with thread density and rack-scale capacity. AMD’s own footnotes put the top part at 256 cores and 512 threads with SMT enabled, 16 channels of MRDIMM at 12.8 GT/s, PCIe Gen 6, and boost frequencies to 5 GHz on select SKUs, built on TSMC’s 2nm node across multiple compute dies and two I/O dies. In AMD’s modeled comparison, selected EPYC configurations deliver up to 2.8 times the estimated agents per watt of selected Arm-based alternatives.

Vera leads with loaded per-core performance, memory latency and predictable data movement, and Nvidia’s technical write-up is unusually explicit about the reasoning. Eighty-eight custom Olympus cores sit on a monolithic mesh, with up to 50% higher IPC than Grace, a 10-wide decode unit, a neural branch predictor, and a graph prefetcher aimed at the indirect memory access patterns in agent memory traversal.

The specific claims are narrower than the marketing shorthand suggests, and the differences matter:

  • More than 1.8x higher sandbox performance under full load versus x86, measured across code compilation, code analysis and Python sandbox workloads. Not a universal per-core figure.
  • 40% lower peak memory latency than x86 CPUs, with over 90% of peak bandwidth sustained under load from 1.2 TB/s of LPDDR5X.
  • 50% faster core-to-core data movement than CPUs that fragment compute across dies, which is a direct shot at chiplet designs.
  • Memory subsystem power under 30W, against well over 100W for DDR5 configurations. This is the source of Nvidia’s “half the power” language. It describes the memory subsystem, not the package. Vera itself is a configurable 250W to 450W part.

Both can be right, which is the actual problem

Neither vendor has shown a universal advantage across representative production agent traces, and both can be correct under their chosen definitions. “Agent” is doing an enormous amount of unexamined work in each pitch.

If an agent is a cheap, mostly-idle sandbox waiting on a tool call, then a model response, then a database, it is a concurrency problem. A high-thread-density design has the advantage when sandboxes are lightly active and memory capacity, rack power and concurrency dominate. It holds that advantage only until something else saturates: memory bandwidth, last-level cache, NUMA (non-uniform memory access) links, scheduler overhead, storage or network I/O.

If an agent is a long serial chain where each step blocks the next, sold against a completion-time budget, a faster-core design becomes more attractive. That is not automatic either. More cores can serve more requests concurrently while still holding tail latency, given enough memory bandwidth.

Plausible production fleets contain both patterns. A coding agent is latency-critical while a person waits, and embarrassingly parallel when it fires off a thousand test suites.

A skeptical operator will reach for a third reading first: that “agentic CPU” is branding over ordinary server consolidation, and the rational buy is whatever delivers the most threads per dollar per watt. That view resembles conventional consolidation procurement, where utilization, compatibility and cost per unit of capacity outweigh workload branding. It stops holding only if latency-bound agent serving grows into a revenue line large enough to justify its own SKU class.

The measurements that would actually decide it

Workload What to measure
Mostly-idle agent sandboxes Active environments per rack, memory per environment, threads per watt
Interactive coding agents p95/p99 CPU-step latency, compile time, tool execution latency
GPU host nodes PCIe and interconnect bandwidth, memory bandwidth, GPU utilization
Batch agent evaluation Completed workflows per hour and per kWh under concurrency
Mixed enterprise fleet Compatibility, licensing, RAS (reliability, availability, serviceability), total cost

The prediction that follows is unglamorous. Sandbox orchestration tends toward higher thread density where memory capacity and rack power permit. GPU-host duty is decided more by accelerator connectivity, memory bandwidth, topology and the ability to keep expensive accelerators busy. Latency-bound tool execution favors stronger loaded per-thread performance. Large operators will likely run more than one CPU profile.

The market as it stands in mid-2026

AMD’s own slide puts it near 46% of x86 server CPU revenue, against roughly 33% of x86 unit shipments. A genuine record. Read the denominator carefully: x86 server CPU revenue, not all data center CPUs and not all servers.

A second figure is circulating with the denominator dropped. IDC’s tracker put non-x86 systems at 47.9% of worldwide server revenue in Q1 2026, $58.7 billion of a $122.6 billion market, up 107% year over year. That is system revenue, heavily weighted by accelerated platforms where GPUs, memory and networking carry most of the price. It is not an Arm CPU share number. What it shows is how fast the revenue denominator around conventional x86 servers is being reshaped.

Nvidia’s finance chief said the company has visibility to nearly $20 billion in total CPU revenue this fiscal year across Grace and Vera. That covers CPUs sold inside Superchip and rack-scale systems as well as standalone, so the accounting basis is not directly comparable with merchant server processor revenue. It is still a large number for a company whose data center CPUs use Arm rather than x86.

Intel is not absent, and the shape of its answer is telling. Clearwater Forest reaches 288 E-cores per socket on 18A, the highest core count of the three, though built from density-optimized E-cores rather than AMD’s mix of high-density and high-frequency parts. The portfolios are not equivalent. But Intel and AMD have converged on the same thesis from different process positions, leaving Nvidia alone on the contrarian side.

What this means for server refresh and asset recovery

The architectural debate reaches asset disposition through one mechanism: rack density changes how many older systems are needed to deliver a fixed amount of work.

When comparisons are drawn at a fixed 100kW per rack, the question stops being which CPU is faster and becomes how much old equipment has to leave to make room. A dual-socket 2019-era server with 32 or 64 total cores does not lose a benchmark so much as a floor-space and power-budget contest. Each round of that arithmetic pushes another wave of gear out of production racks, along with the DDR4 RDIMMs, older Xeon Scalable and early EPYC processors, and boot drives inside them.

Two practical notes, both tendencies rather than rules. Components and whole units typically move on different clocks: a retired server generally sells for what a refurbisher will pay for a complete configuration, while processors and memory can behave differently, particularly under the DRAM pricing pressure visible through the second and third quarters of 2026. Pulling and grading parts costs labor, and whether it pays depends on the SKU mix.

Timing usually matters more than decommission plans assume. All else equal, recovery value on a CPU generation tends to be highest before its replacement ships in volume, though supply shortages, extended OEM support, compatibility-driven demand and export restrictions have all interrupted that curve before.

Venice enters cloud and OEM deployment through the back half of this year. AMD describes an annual cadence across its portfolio, but the server CPU architecture steps land roughly every two years: Zen 7 “Florence,” “Ferrara” and “Fidenza” in 2028, Zen 8 “Ravenna” in 2030. A published roadmap through 2030 is also a published depreciation schedule for everything already installed. Trays of 4th and 5th Gen EPYC or Sapphire and Emerald Rapids parts waiting in a storage cage on a decision are usually losing value while they wait, which is the argument for getting used server processors valued early in a refresh rather than at the end of it. A similar calculation runs on the other side of the rack, where the economics of owning accelerators versus renting them have moved sharply in two years.

What would settle it

Part of the missing standard already exists. MLCommons announced Agentic Inference for MLPerf Inference in July, replaying 990 multi-turn trajectories and more than 30,000 client-issued turns across coding and enterprise workflow domains, on thinking-capable models at up to 262K tokens of context. It measures the tradeoff between aggregate serving throughput and per-agent progress, which is a real advance over vendor slides.

It also runs inside the MLPerf Endpoints framework, which means it exercises the model-serving path: context growth, KV-cache behavior, throughput under concurrency. That is the GPU side of the loop. It will not settle the CPU argument unless submissions also expose what sandbox execution, compilation, orchestration and tool runtimes cost outside the model endpoint.

Which is the thing to watch. Whether MLCommons extends the suite far enough into CPU-side tool execution and sandbox orchestration to distinguish a high-thread-density processor from a high-loaded-performance one. And what large operators actually order for GPU head nodes versus standalone agent hosts. Those purchase orders will describe the workload more honestly than either keynote managed.