
Intel unveils a 256-core Diamond Rapids Xeon
TL;DR
- Diamond Rapids tops out at 256 P-cores, double the 128 in today’s Xeon 6980P. Intel detailed it at Hot Chips on 24 August; it ships in the second half of 2027.
- Hyperthreading is gone, so 256 cores means 256 threads — the same count the 128-core 6980P already delivers. SMT returns in the next generation, Coral Rapids.
- The package is 22 pieces of silicon: 16 core chiplets, 4 base tiles carrying 1.28 GB of cache, and 2 fabric hubs that now hold all the memory controllers and I/O.
- 16 memory channels at 12,800 MT/s with MRDIMMs, around 1.6 TB/s peak. There is no 8-channel part, so filling every channel takes sixteen DIMMs per socket instead of eight.
Intel laid out Diamond Rapids in detail at Hot Chips 2026 on Monday 24 August. It is the Xeon Intel plans to sell in 2027, and most coverage is already calling it Xeon 7.
The headline is 256 P-cores in one socket. That is double the 128 in today’s Xeon 6980P. Intel already ships a denser part, the 288-core Xeon 6+ built on E-cores. But on the big-core line this is the first doubling since Granite Rapids.
The thread count did not move. The 6980P runs 128 cores and 256 threads. Diamond Rapids runs 256 cores and 256 threads, because Intel dropped hyperthreading for this generation.
What Intel put on the slides
Intel’s own summary is short. Up to 256 cores with 1.28 GB of last-level cache, against 504 MB in the 6980P. Then 16 memory channels at 12,800 MT/s, the transfer rate a DIMM runs at. Then 128 lanes of PCIe Gen 6 and CXL 3.0, the coherent interconnect that ties processors to accelerators and memory devices.
The instruction set moves too. APX is new, and it doubles the general-purpose registers a core can hold values in, from 16 to 32. That can cut the spills and reloads compiled code makes to memory. AMX, the matrix engine Intel uses for on-CPU inference, gets extra instructions rather than a rewrite, and AVX10.2 comes with it.
The cores are Panther Cove, though Intel did not detail the microarchitecture. The chip is built on Intel 18A-P, a higher-performance version of the process already behind the Xeon 6+ parts launched in June.
Memory is where the generational jump is largest: twelve channels become sixteen. Ordinary DDR5 still works, up to 8,000 MT/s. With MRDIMMs the ceiling is 12,800. An MRDIMM runs two ranks at once and multiplexes their output through a buffer on the module. That is how it clears rates a conventional RDIMM cannot.
Sixteen channels at 12,800 MT/s is about 1.6 TB/s, which is the figure Intel’s slides carried. That is peak theoretical bandwidth, though, not a measured one.
Twenty-two pieces of silicon in one package
The packaging is the biggest structural change, though, and Intel calls it a Fan-out Fabric. A Diamond Rapids package is 22 separate pieces of silicon bonded together. Sixteen of them are core chiplets carrying 16 cores each. Four are base tiles, the layer underneath that holds the shared cache. Then two more are fabric hub tiles that handle everything else.
The split is the interesting part. On Granite Rapids the memory controllers sit on the compute dies. On Diamond Rapids they move off into the two fabric hubs, along with the PCIe and CXL ports and the accelerator blocks for crypto, data movement and compression.
So compute and I/O scale more separately than they did. Core count and the memory path are no longer tied to the same die.
The 1.28 GB of shared cache lives on the four base tiles, 320 MB apiece. The core chiplets stack directly on top using Foveros Direct, a copper-to-copper bond that welds two dies face to face with no solder bumps.
Traffic between the base tiles and the fabric hubs runs over copper on the substrate. It uses UCIe-S, an industry-standard die-to-die interface, rather than the bridge technology Intel developed in-house. That removes one interface barrier to mixing in third-party chiplets later, though this package is all Intel silicon.
The thread arithmetic
Intel dropped SMT from this generation. SMT is simultaneous multithreading, the feature that lets one physical core present two logical CPUs to the operating system. Intel first signaled the change at Computex in June. Monday added the core number that turns it into a thread-count story.
Why Intel dropped it
Intel has never given a Diamond-Rapids-specific reason. The one explanation on record comes from the client side. Lead x86 architect Stephen Robinson, discussing the same removal on Panther Lake laptop chips, said SMT “isn’t necessarily as valuable” once hybrid cores are in the mix, leaving a core that is “maybe a little bit easier and less expensive and maybe can go a little bit faster.” That reasoning is contingent on hybrid. Panther Lake hands throughput work to E-cores; Diamond Rapids has none to hand it to.
Intel’s own CEO has said the removal cost the company. On the data-center line, Lip-Bu Tan wrote in his Q2 2025 letter to stakeholders: “Moving away from SMT put us at a competitive disadvantage. Bringing it back will help us close performance gaps.” Diamond Rapids ships without it anyway.
SMT returns in Coral Rapids. CEO Lip-Bu Tan said so on the July earnings call: “Multithreading will be coming in the Coral Rapids.”
What it costs
One thread per core is not new to Xeon: the 288-core Xeon 6+ also runs without SMT. But those are E-cores built for scale-out throughput, where nobody expected it. Diamond Rapids is the big-core line, where per-core-licensed workloads live.
For latency-sensitive work, no SMT is arguably a feature: fewer noisy neighbors, and one less thing to disable in BIOS.
For per-core-licensed consolidation, it raises the bar. Licensing models vary, but where software is counted by physical core, threads are the part that arrives free. VMware counts “the total number of physical CPU cores across all ESXi hosts”, with a 16-core minimum per socket, and says nothing about threads at all. So removing SMT does not reduce the licence count. Whether it reduces what that licence buys comes down to per-core throughput.
How it lines up
Set the three parts side by side:
| Chip | Cores / threads | Availability |
|---|---|---|
| Xeon 6980P | 128 / 256 | shipping now |
| Diamond Rapids | up to 256 / 256 (no SMT) | 2H 2027 |
| EPYC 9996 (Venice) | up to 256 / 512 | in production, platforms Q4 2026 |
Per licensed physical core, Venice exposes two hardware threads and Diamond Rapids exposes one. That is not twice the throughput, and on some latency-sensitive workloads SMT is worth nothing or worse than nothing. What it is, is an execution resource AMD has on this generation and Intel does not. The number that settles it is measured throughput per licensed core, and nobody has that yet.
AMD has noticed. Its compute and enterprise AI VP, Madhu Rangarajan, said in April that Intel’s “interesting choices on multithreading” would help AMD win enterprise share, citing “implications to your licensing.”
Consolidation does not help either. Under VMware’s rule, one 256-core socket and two 128-core sockets both count as 256 licensed cores. Collapsing two boxes into one saves power, rack space and a chassis, but not a single licensed core.
On memory the two are unusually close. Both run 16 channels and MRDIMM-12800. AMD publishes 1,638 GB/s of peak bandwidth for the 9996, and sixteen channels at 12,800 MT/s gives Intel the same arithmetic. Neither publishes a sustained figure.
I/O is less settled than it looks. Intel disclosed 128 lanes of PCIe Gen 6. AMD’s spec page for the 9996 lists PCIe 6.0 x96, while press briefed at launch reported 128 lanes. Until AMD reconciles those, it is not a comparison to lean on. Everything else that decides a purchase order, Intel has not published at all.
What Intel did not say
A lot, because the session stayed architectural. No clock speeds, no power limits, no SKUs, no pricing and no release date came out of it.
Power is also the column missing from the table above. Published: the Xeon 6980P is a 500W part. Not published: any figure at all for Diamond Rapids. What circulates instead is a leak: the 256-core part at 650W, standard models up to 192 cores at 500W. A leak is not a specification. Anyone sizing power and cooling for a 2027 refresh still has nothing official to work from.
The 2027 date comes from Computex in June, not from this week. Press in the room read it as the second half of 2027.
There is also no 8-channel part. Intel cancelled the 8-channel Diamond Rapids last November. It said it was “simplifying the Diamond Rapids platform with a focus on 16-channel processors and extending its benefits down the stack to support a range of unique customers and their use cases.”
That has a consequence nobody is putting next to the core count. The 6700P and 6500P class is Intel’s mainstream server platform, and eight memory channels are part of its design. If the successor is 16-channel only, filling every channel takes sixteen modules per socket instead of eight, and filling every channel is what exposes the platform’s full bandwidth. A buyer can under-populate and give up bandwidth, or buy the DIMM count and hold capacity flat with smaller modules. Either way the memory decision changes shape before the CPU price is quoted.
What it means for the hardware already in the racks
Nothing until 2027, which is the honest answer, because none of these ships before then. The reason it matters is timing rather than technology.
AMD announced the 9996 with full specifications in July and says the silicon is in production, with partner platforms due in the fourth quarter. Intel’s answer is not expected until the second half of 2027. So, the decision is between what can be bought now and holding off. What can be bought now is Turin or Xeon 6, at whatever pricing the ramp allows. Venice platforms follow in Q4. What is being held off for is a part whose clocks, power and price nobody has seen.
The Granite Rapids and Sapphire Rapids boxes those replace come out of racks on lease and depreciation schedules, not on Intel’s launch calendar.
But when they do come out, the parts outlast the chassis. Xeon 6 and Xeon 5 processors, the DDR5 RDIMMs beside them and the accelerators in the PCIe slots are worth counting and pricing separately, rather than going out with the metal. What any one-part fetches depend on the SKU, the capacity, the condition and who is buying that week. That is the market BuySellRam works in, on both sides of it: sell RAM for the DDR5 modules, sell CPU for the processors.
The number to watch from here is not 256. Intel has confirmed that ceiling. At Computex in June it talked about 192, so how far down the stack 256 reaches is a real question. It is the smaller one.
The bigger one is what a Diamond Rapids core delivers at whatever clock and power it ships at, measured per physical core. That is the unit enterprise licences count. Until those numbers exist, 256 is an architectural fact rather than a consolidation result.