NVIDIA Unveils NVHBM Memory To Turbocharge AI Chip Speeds By 30%

hero nvidia nvhbm processor render
NVIDIA NVHBM Example Render - Image: NVIDIA
On Monday, we reported on SiFive's move to bring the open-source RISC-V CPU architecture into the datacenter with the BigSky server platform, and as part of that reporting, we mentioned that SiFive had also partnered with NVIDIA to integrated NVLink Fusion into its future products. We didn't really explain what NVLink Fusion is, though. That understanding is critical to understand today's announcement regarding its new NVHBM tech, which is a novel high-bandwidth memory technology destined for future processors, be they GPUs, CPUs, AI accelerators, or whatever else.

NVLink Fusion is an architecture initiative and ecosystem program that opens up NVIDIA's proprietary NVLink high-speed interconnect fabric to third-party chipmakers. Basically, it allows the company's partners to license NVLink IP and implement it into their own chips so that they can mix and match their own silicon (or even another firm's) alongside NVIDIA processors in a single rack. Companies like Amazon, Marvell, MediaTek, Fujitsu, Arm, Cadence, and Samsung (not exhaustive) have signed on to support NVLink fabric.

NVHBM, then, is the next step in this program. NVHBM is, as NVIDIA describes it, "a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs." XPUs in this case meaning "processors", as the "X" in this case is used to mean "whatever kind of" Processing Unit.

nvlink fusion nvhbm example
NVIDIA NVHBM Illustration - Image: NVIDIA

NVHBM isn't just a licensable IP block for an HBM controller, though; in fact, NVHBM is a different structure altogether for integrating HBM into processors. Normally, in most processors, the memory controller is integrated into the processor die, which means it's consuming area there. What NVIDIA is doing is what AMD already did with the Radeon RX 7900 XTX: moving the memory controller onto a separate die. However, where RDNA 3 used Memory and Cache Dies (MCDs) that connected to external GDDR memory packages, NVHBM stacks the HBM itself on top of the MCU.

NVIDIA claims that by integrating the memory controller into the 3D HBM stack, it can offer "up to 30% greater memory bandwidth and 15% lower HBM power consumption." It also reportedly frees "up to 25% more area" on the processor's die versus regular HBM4E. On its developer blog, NVIDIA says "these co-designed architectural improvements translate into a significant 30% overall end-to-end performance increase per XPU." Impressive stuff, assuming it pans out that way.

Interestingly, NVIDIA isn't exactly competing with the memory vendors here; the company says it is "establishing a standard NVHBM implementation" that will be "available from multiple memory providers." The company already has one partner on board; Amazon's Annapurna Labs, a division of AWS, has committed to using NVLink Fusion on its Trainium4 processors, and it is apparently working with NVIDIA on NVHBM in some capacity too, although the firm hasn't committed to actually doing anything specific with NVHBM yet.
Zak Killian

Zak Killian

A 30-year PC building veteran, Zak is a modern-day Renaissance man who may not be an expert on anything, but knows just a little about nearly everything.