Toward efficient device HBM

Yu Zhao
yuzhao@google.com

Agenda

1 Intro
~5 mins
2 Case Study: zHBM for NV GPUs
~15 mins
3 Roundtable
~25 mins

Aggregate Hyperscaler Spend (2025–2026 Annual Run-Rate)

Globally, the total DRAM market reached ~$115B–$130B+ per year in 2025–2026. Data centers now account for ~70% of global DRAM revenue, with the major hyperscalers consuming ~65–75% of all Server DRAM and ~85–90% of all HBM.

Metric (Annualized 2025–2026) CPU Memory (Server DDR5 / LPDDR5X / CXL) GPU / ASIC Memory (HBM3 / HBM3e / HBM4)
Global Market TAM (Memory Vendor Revenue) ~$40B – $52B / yr ~$35B – $46B / yr (Up from ~$4B in 2023 and ~$16B in 2024)
Hyperscaler Share of Global Volume ~65% – 75% ~85% – 90%
Direct Memory BOM Spend (Paid to SK Hynix, Samsung, Micron) ~$26B – $38B / yr ~$30B – $40B / yr
Effective CapEx Spend (Incl. GPU vendor margin markup) ~$26B – $38B / yr (Bought direct / ODM pass-through) ~$70B – $95B+ / yr (~3–4x markup on merchant GPUs; 1x on custom ASICs)

CPU Memory Tiering & Page Lifecycle

Click any flow filter or diagram block to inspect mechanisms & latency tiers

CPU Cores & MMU • L1/L2/LLC + TLB & Page Tables • Sets PTE Accessed / Dirty bits • Traps Page & NUMA Hint Faults • PMU Sampling (PEBS / IBS) Byte-Addressable Load/Store Hot/Cold Page Tracking • MGLRU / DAMON • PTE A-bit scans • NUMA Hint Faults (PROT_NONE) • CXL CMS / Hotness Counters Classifies Hot vs. Cold Pages Tier 0: Local DDR DRAM CPU-Attached NUMA Node (~80ns) Hot Pages • Active workingset • PTE Mapped (A=1) • Promoted or newly faulted Aging (A=0) Cold Pages • Unreferenced for K intervals • Reclaim / Demotion candidates • Selected by kswapd / MGLRU • Unmapped on demotion/swap Tier 1A: CXL.mem Expander CPU-less NUMA Node (~170–250ns) • Byte-addressable via CXL.mem • Direct CPU load/store (no fault) • Warm/Cold pages via migration • NUMA hint fault detects hotness Hardware Cache-Coherent HDM Tier 1B: Compressed RAM zswap / zram Pool (~1–5µs fault) • Page-granular (Swap PTE entry) • LZ4 / ZSTD / IAA compression • ~2–3x effective RAM density • Decompressed on do_swap_page Software/Accel Swap Cache Tier 2: Disk Swap NVMe / SSD (~15–100µs) • Block I/O swapfile • Coldest anonymous pages & zswap LRU • Major page fault reads back to DRAM Highest Capacity ~80ns Direct CXL.mem Load/Store (Cacheline Granular, ~200ns) PTE A-bit / Faults Demote Promote Compress Decompress Swap Out Major Page Fault (do_swap_page Read-In to Tier 0 DRAM)
Architecture Overview: Pages age in Tier 0 DRAM via MGLRU/PTE tracking. Cold pages are demoted via migrate_pages() to byte-addressable CXL.mem or compressed into zswap/zram before overflowing to NVMe Disk Swap.
3 Tiers • Interactive

Agenda

1 Intro
~5 mins
2 Case Study: zHBM for NV GPUs
~15 mins
3 Roundtable
~25 mins

NVIDIA GPU Architecture Terms (via CPU & OS Analogies)

A mapping of core NVIDIA GPU compute, memory management, and hardware offload engines to their closest CPU and OS/MM equivalents.

NVIDIA GPU Term CPU / OS Analogy Role & Architectural Mechanism
Streaming Multiprocessor (SM) Compute Core Block CPU Physical Core (with wide SIMD + SMT threads) The fundamental compute building block of an NVIDIA GPU (e.g., 132 SMs on H100; 148–160 SMs on B200). Each SM contains its own instruction schedulers, a massive 256 KB Register File, scalar/vector ALUs (CUDA Cores), matrix units (Tensor Cores), and on-chip L1/Shared Memory.
GPU MMU (GMMU) Address Translation & Faults CPU MMU / IOMMU (Page Tables, TLBs, PTE A-bits & NUMA Faults) Manages virtual-to-physical translation across local HBM, peer GPU HBM, and host CPU DRAM (via Unified Virtual Memory / HMM / ATS). Supports replayable page faults (stalling thousands of threads while the OS/driver migrates missing pages) and Hardware Access Counters that track remote page hotness to trigger automatic page promotion.
HBM & Unified L2 Cache Device Memory Hierarchy Local DDR5 DRAM + Shared LLC (L3 Cache) HBM (High Bandwidth Memory) is 3D-stacked DRAM on the GPU package interposer delivering 3.35–8.0 TB/s (~10x CPU DDR5 bandwidth) but with constrained capacity (80–288 GB). Fronted by a massive GPU-wide Unified L2 Cache (50–126 MB) shared across all SMs, acting like a CPU socket's Last-Level Cache (L3).
Copy Engine (CE) Async Data Movement System DMA Engine (e.g., Intel DSA / CBDMA) Dedicated hardware DMA engines independent of the SMs. CEs execute asynchronous memory copies (cudaMemcpyAsync), page zeroing, and bulk page migration between CPU DRAM, peer GPUs (over NVLink), and local HBM with zero SM compute overhead, enabling compute and memory tiering to overlap.
Decompression Engine (DE) Hardware Offload Block Hardware Compression Accelerator (e.g., Intel IAA / QAT) Fixed-function hardware engine in Hopper & Blackwell (up to 800 GB/s throughput supporting LZ4, Snappy, Deflate/Gzip, ANS) that decompresses pages/streams directly into HBM at line rate—completely bypassing both CPU cores and GPU SMs. Ideal for compressed memory tiering (zswap-like GPU paging) and fast weight/KV-cache loading.
NVLink & NVLink-C2C Coherent Interconnect Socket Interconnect (UPI / Infinity Fabric) + CXL.mem High-bandwidth interconnect (900–1,800 GB/s per GPU) linking GPU-to-GPU (NVLink) and CPU-to-GPU (NVLink-C2C on Grace-Hopper/Blackwell). Enables cache-coherent, byte-addressable load/store and atomic access across host CPU DRAM and GPU HBM as a single unified NUMA-like memory fabric.

LLM KV Cache & How It Causes GPU OOM

To avoid recomputing attention over all previous tokens at every generation step (O(N²) compute), LLMs cache each token's Key (K) and Value (V) vectors in GPU HBM. In OS terms, if Model Weights are the fixed .rodata binary image, the KV Cache is the per-session dynamic heap—growing linearly with sequence length and concurrency until HBM is exhausted.

Typical GPU HBM Footprint During Long-Context / High-Concurrency Inference KV Size = 2 (K,V) × Layers × KV_Heads × Head_Dim × Bytes × Seq_Len × Batch
Model Weights (~Fixed)
Activations
Active & Idle KV Cache (Grows w/ Tokens)
KV Overflow → GPU OOM!
0 GB Example: Llama-3 70B (~0.32 MB/token) • 32 concurrent 128K-context requests = ~1.3 TB of KV Cache (>16× single 80GB HBM) HBM Limit (80–192 GB)
KV Cache / OOM Factor CPU / OS Analogy Why It Triggers GPU Out-of-Memory (OOM)
Unbounded Per-Token Growth O(Batch × Seq_Len) Scaling Unbounded Per-Connection Heap (Unknown allocation size at start) Every generated token appends new K/V vectors across all transformer layers (e.g., ~0.3–1.0 MB per token). Because an LLM server cannot predict how many tokens a prompt will generate, concurrent requests simultaneously expanding their contexts can suddenly exceed free HBM mid-generation.
Idle Multi-Turn & Prefix Caches Hot vs. Cold Workingset Cold Anonymous / Page-Cache Pages (Sitting in Tier 0 RAM between requests) In multi-turn chat, code assistants, and agentic workflows (waiting on tool calls or user think-time), huge KV caches sit completely idle ("cold") in expensive HBM for seconds or minutes. Retaining them in HBM starves active ("hot") requests of memory and causes OOM under traffic spikes.
Hard HBM Wall (No Default Tiering) Fail-Fast vs. Graceful Paging Running mlock() System-Wide with Swap Disabled On a CPU, memory pressure triggers background reclaim (MGLRU → zswap / CXL / SSD). On a GPU, hitting 100% HBM triggers a hard cudaErrorOutOfMemory crash or forces the serving engine to evict & recompute KV caches from scratch—unless cold KV pages are tiered out to CPU DRAM or compressed memory.

Compressibility of LLM KV Cache

To halve memory footprint and double matrix-engine throughput without sacrificing model quality, modern LLMs transitioned from FP32 (32-bit) to 16-bit floating-point formats—Heavily favoring BF16 (bfloat16) over FP16, because dynamic range matters far more than decimal precision:

FP32 (Single Precision) 32 bits (4B)
1s
8b Exp
23b Mantissa

Wide dynamic range (±3.4×1038), but 2× memory & bandwidth cost.

FP16 (IEEE Half) 16 bits (2B)
1s
5b Exp
10b Mantissa

Narrow 5-bit exponent (max 65,504) suffers from overflow/underflow on LLM outlier activations.

BF16 (Brain Float) 16 bits (2B)
1s
8b Exponent
7b Mant

Same 8-bit exponent range as FP32 with fewer noisy mantissa bits—preventing overflow & concentrates entropy.

Why Standard Byte Compressors Fail on Floating-Point KV Caches: While the exponent bits across KV cache tensors have very low entropy (values cluster tightly in magnitude), the pseudo-random mantissa bits interleaved into every 2-byte element break byte-level string repetition—rendering general-purpose LZ77 algorithms (LZ4, Snappy) completely ineffective (~1.000×), whereas float-aware compression (kvcomp) achieves >1.5× lossless compression even at 4KiB page size.

Average Compression Ratios on FP16 KV Cache Corpora
Compressors / Block Sizes 4KiB 64KiB 2MiB
kvcomp (custom) 1.515× 1.525× 1.527×
Bitcomp (nvCOMP, USHORT, algo=0) 1.200× 1.237× 1.134×
Snappy (zippy) 0.999× 1.000× 1.000×
LZ4 0.996× 0.996× 0.996×

zHBM: In-HBM Compression for NVIDIA GPUs

zHBM runs on NVIDIA GPUs just like zram runs on CPUs—while zram compresses cold pages in CPU DRAM, zHBM transparently compresses cold pages in GPU HBM to expand effective HBM capacity and prevent KV cache OOMs without leaving the GPU package.

1. Hot / Cold Tracking (GMMU) No Hardware Accessed (A) Bit in PTE • 1a. PF-based Access Detection Unmaps PTE → GPU access faults = Hot • 1b. Periodic Scanning Scans & ages un-faulted pages → Cold 2. Compressor: kvcomp on SMs Runs on Streaming Multiprocessors • Custom lossless BF16 KV compressor • Exploits low-entropy 8-bit exponents • ~1.52× ratio at 4KiB–2MiB blocks GPU HBM: Uncompressed Pages Hot BF16 Pages (Active KV) • Mapped in GMMU PTEs • Direct SM load / store access • Kept hot via 1a. access faults 1b. Scan 1a. Fault Cold BF16 Pages (Idle KV) • Unreferenced across scan windows • Selected for zHBM compression • Unmapped & freed after packing 3. Compressed Data Storage (in HBM) Segregated-Fit Allocation Size Class A (e.g., 2.5 KiB slots) KV Page #4 KV Page #9 KV Page #12 Size Class B (e.g., 2.7 KiB slots) KV Page #2 KV Page #7 Free Size Class C (e.g., 3.0 KiB slots) KV Page #1 KV Page #5 Free Zero External Fragmentation • ~1.52× Density PTE Scan Cold Page Compress via kvcomp on SMs → Store in Size-Segregated HBM Class Decompress (on Fault)
zHBM Component CPU / Linux MM Analogy Design & Architectural Mechanism
1a. PF-based Access Detection Hardware A-bit Because the GPU MMU (GMMU) currently does not support a hardware Accessed (A) bit in PTEs, zHBM detects page references via page-fault based access detection—unmapping candidate PTEs so any GPU access traps into a replayable GMMU fault and marks the page hot.
1b. Periodic Scanning Pressure Driven (MGLRU / kswapd) Periodically scans HBM address ranges to age pages across intervals, separating actively referenced ("hot") working-set pages from unreferenced ("cold") KV cache pages eligible for compression.
2. Custom Compressor for BF16 (kvcomp) Crypto API (e.g., lz4) Uses kvcomp, a custom lossless KV cache compressor designed specifically for BF16 (bfloat16) tensors and executed directly on GPU Streaming Multiprocessors (SMs) at high throughput (~1.52× compression where LZ4/Snappy yield 1.00×).
3. Segregated-Fit Allocator zsmalloc Pool Allocator Manages compressed pages in HBM using segregated storage (segregated-fit allocation)—grouping compressed blocks into discrete size classes to eliminate fragmentation and enable fast, deterministic allocation and decompression on page faults.

KV Cache Latency & Throughput Comparison

Comparison of effective throughput and end-to-end latency to access or restore 1 GB of KV Cache per GPU (e.g., ~32K tokens on Llama-3 70B with TP=8) across memory tiers and recomputation.

Source / Method Hardware Path & Bottleneck Effective Throughput (per GPU) Latency for 1 GB KV (~32K tokens @ TP=8) Compute / System Impact
Local HBM (Uncompressed Hit) Direct SM load/store via HBM3/HBM3e (3.35–8.0 TB/s) 3,350 – 8,000 GB/s ~0.13 – 0.30 ms Zero overhead; baseline decoding path.
Uncompressing from zHBM (In-HBM kvcomp) Read compressed HBM → SM kvcomp → Write HBM (~1.66× HBM traffic) ~800 – 2,000+ GB/s (SM-bandwidth bound) ~0.5 – 1.5 ms Stays 100% inside HBM; ~4×–10× faster than GB200 LPDDR5X and ~20×–50× faster than x86 DDR5/PCIe.
Host DRAM (CPU Memory) PCIe Gen5 x16 (64 GB/s raw) via Copy Engine (CE) (or NVLink-C2C: ~400 GB/s) + Host DDR5 (~200–280 GB/s per 4-GPU socket) / LPDDR5X (~400 GB/s per 2-GPU Grace)
~35 – 55 GB/s (x86 DDR5 / PCIe5)
~180 – 200 GB/s (GB200 LPDDR5X / C2C)
~20 – 30 ms (x86)
~5.0 – 5.5 ms (GB200 -- less common)
Bottlenecked by shared host DDR5/LPDDR5X controllers across GPUs.
Distributed KV Storage (SSD / Remote DRAM) NVMe-oF SSD (3–10 GB/s) or RDMA (100–400G RoCE/IB) (12–45 GB/s)
~3 – 10 GB/s (SSD)
~12 – 45 GB/s (RDMA)
~100 – 300 ms (SSD)
~25 – 85 ms (RDMA -- less common)
Competes with East-West inference/training network fabric; high tail latency.
Recomputing the KV Cache (Prefill Pass) Full Transformer prefill (O(S²) attention + FFN across all layers) ~0.4 – 2.0 GB/s (effective KV gen rate)
~800 – 1,400 ms (at 32K)
~6,000+ ms (at 128K)
Scales quadratically O(S²); hijacks 100% of Tensor Cores and stalls active batch.

Agenda

1 Intro
~5 mins
2 Case Study: zHBM for NV GPUs
~15 mins
3 Roundtable
~25 mins