d-Matrix Claims 20x Bandwidth Density Over NVIDIA Rubin In New Chip
by
Zak Killian
—
Monday, August 24, 2026, 03:40 PM EDT
A photo of d-Matrix's Pavehawk 3D DRAM test chip. - Image: d-Matrix
In AI inference, you have two phases of the workload: prefill and decode. In most cases, decode is overwhelmingly the more time-consuming portion of the workload because it's strictly memory-bandwidth bound. You can have all the compute in the world, and that's great for prefill, but decode doesn't care. There have been various strategies to attack this problem, including startups like Groq (whose tech was recently purchased by NVIDIA) and d-Matrix offering chips with massive SRAM to deliver enormous memory bandwidth. d-Matrix now has a new product it's showing off, and it attacks the bandwidth problem directly from a different angle: below.
Raptor is the chip, and 3D DRAM is the tech. - Image: d-Matrix via ServeTheHome (click for big)
You know how AMD's second-generation 3D V-Cache, used on the Ryzen 9000 processors, places the SRAM cache die underneath the CPU cores to improve thermals? That's basically what's going on in d-Matrix's "Raptor" accelerator chip, except the Silicon Valley startup is using up to four layers of DRAM instead of SRAM for improved capacity. The density is still significantly lower than what's possible with HBM, but d-Matrix says it can achieve some 20 times the bandwidth per area and 13.5 times reduced power per gigabyte transferred.
A slide comparing SRAM, HBM, and 3D DRAM. - Image: d-Matrix via ServeTheHome (click for big)
Both of these metrics target major concerns for AI processors as the limits of HBM become apparent. Because HBM stacks sit adjacent to the compute die on a silicon interposer, bandwidth is fundamentally choked by the fact that you can only squeeze so many I/O pins along the chip's physical edges. Pushing data horizontally across that interposer also requires dedicated interface drivers (the PHY) that consume massive amounts of power, often accounting for hundreds of watts in high-end packages. By using 36-micron face-to-face bonding to fuse a TSMC N4 logic die directly on top of the DRAM, Raptor eliminates the PHY entirely. Electrons travel vertically over mere micrometers across the entire surface footprint of the die, cutting I/O energy down to roughly 0.37 pJ/bit while unlocking SRAM-class bandwidth.
d-Matrix's 3D DRAM trades density for speed and efficiency. - Image: d-Matrix via ServeTheHome (click for big)
The reduced capacity versus HBM isn't really a huge concern either given the inference focus of Raptor. These chips aren't suited for training enormous frontier models, but then, that's not what they're for. At 32GB per card, Raptor lands in a pragmatic middle ground between tiny, ultra-expensive SRAM blocks and massive HBM stacks. To host giant LLMs, d-Matrix relies on scaling these chips out horizontally into rack-scale clusters, like a 72-card setup, all tied together over high-speed fabrics. In a cluster configuration, the aggregate memory bandwidth scales linearly into petabytes per second, allowing the system to easily distribute parameters across nodes and handle multi-million-token context windows without choking during the decode phase.
d-Matrix's claims are extremely ambitious. - Image: d-Matrix via ServeTheHome (click for big)
There's surprisingly little information about the capabilities or performance of the Raptor product itself besides the slide above, where d-Matrix claims it can provide around 1,000 tokens per second, per user, while serving a "3T-class model at 1M context." That's presented alongside benchmarks that show Raptor (presumably a whole rack) producing some 2,121 tokens/second/user for 32 users with GLM 5.2, and 785 tokens/second/user with 32 concurrent users while running the new Kimi K3 model.
The Hot Chips presentation from d-Matrix is mostly about how the company was able to overcome the challenges of stacking DRAM under logic, which isn't a novel concept in itself. Of course, specifications and projections are only that, and it doesn't seem like there has been a live demo of the product yet. We'd love to see if d-Matrix can deliver on these promises. The company does have products shipped, so this isn't some pie in the sky startup promising the world and delivering nothing. Hopefully we get more information on Raptor soon, including third-party benchmarks!
A 30-year PC building veteran, Zak is a modern-day Renaissance man who may not be an expert on anything, but knows just a little about nearly everything.
TOP STORIES
Which New GPU Is For You?
KEEP INFORMED
Stay updated with the latest news and updates. Subscribe to our newsletter!