Globally, the total DRAM market reached ~$115B–$130B+ per year in 2025–2026. Data centers now account for ~70% of global DRAM revenue, with the major hyperscalers consuming ~65–75% of all Server DRAM and ~85–90% of all HBM.
| Metric (Annualized 2025–2026) | CPU Memory (Server DDR5 / LPDDR5X / CXL) | GPU / ASIC Memory (HBM3 / HBM3e / HBM4) |
|---|---|---|
| Global Market TAM (Memory Vendor Revenue) | ~$40B – $52B / yr | ~$35B – $46B / yr (Up from ~$4B in 2023 and ~$16B in 2024) |
| Hyperscaler Share of Global Volume | ~65% – 75% | ~85% – 90% |
| Direct Memory BOM Spend (Paid to SK Hynix, Samsung, Micron) | ~$26B – $38B / yr | ~$30B – $40B / yr |
| Effective CapEx Spend (Incl. GPU vendor margin markup) | ~$26B – $38B / yr (Bought direct / ODM pass-through) | ~$70B – $95B+ / yr (~3–4x markup on merchant GPUs; 1x on custom ASICs) |
Click any flow filter or diagram block to inspect mechanisms & latency tiers
migrate_pages() to byte-addressable CXL.mem or compressed into zswap/zram before overflowing to NVMe Disk Swap.
A mapping of core NVIDIA GPU compute, memory management, and hardware offload engines to their closest CPU and OS/MM equivalents.
| NVIDIA GPU Term | CPU / OS Analogy | Role & Architectural Mechanism |
|---|---|---|
| Streaming Multiprocessor (SM) Compute Core Block | CPU Physical Core (with wide SIMD + SMT threads) | The fundamental compute building block of an NVIDIA GPU (e.g., 132 SMs on H100; 148–160 SMs on B200). Each SM contains its own instruction schedulers, a massive 256 KB Register File, scalar/vector ALUs (CUDA Cores), matrix units (Tensor Cores), and on-chip L1/Shared Memory. |
| GPU MMU (GMMU) Address Translation & Faults | CPU MMU / IOMMU (Page Tables, TLBs, PTE A-bits & NUMA Faults) | Manages virtual-to-physical translation across local HBM, peer GPU HBM, and host CPU DRAM (via Unified Virtual Memory / HMM / ATS). Supports replayable page faults (stalling thousands of threads while the OS/driver migrates missing pages) and Hardware Access Counters that track remote page hotness to trigger automatic page promotion. |
| HBM & Unified L2 Cache Device Memory Hierarchy | Local DDR5 DRAM + Shared LLC (L3 Cache) | HBM (High Bandwidth Memory) is 3D-stacked DRAM on the GPU package interposer delivering 3.35–8.0 TB/s (~10x CPU DDR5 bandwidth) but with constrained capacity (80–288 GB). Fronted by a massive GPU-wide Unified L2 Cache (50–126 MB) shared across all SMs, acting like a CPU socket's Last-Level Cache (L3). |
| Copy Engine (CE) Async Data Movement | System DMA Engine (e.g., Intel DSA / CBDMA) |
Dedicated hardware DMA engines independent of the SMs. CEs execute asynchronous memory copies (cudaMemcpyAsync), page zeroing, and bulk page migration between CPU DRAM, peer GPUs (over NVLink), and local HBM with zero SM compute overhead, enabling compute and memory tiering to overlap.
|
| Decompression Engine (DE) Hardware Offload Block | Hardware Compression Accelerator (e.g., Intel IAA / QAT) |
Fixed-function hardware engine in Hopper & Blackwell (up to 800 GB/s throughput supporting LZ4, Snappy, Deflate/Gzip, ANS) that decompresses pages/streams directly into HBM at line rate—completely bypassing both CPU cores and GPU SMs. Ideal for compressed memory tiering (zswap-like GPU paging) and fast weight/KV-cache loading.
|
| NVLink & NVLink-C2C Coherent Interconnect | Socket Interconnect (UPI / Infinity Fabric) + CXL.mem |
High-bandwidth interconnect (900–1,800 GB/s per GPU) linking GPU-to-GPU (NVLink) and CPU-to-GPU (NVLink-C2C on Grace-Hopper/Blackwell). Enables cache-coherent, byte-addressable load/store and atomic access across host CPU DRAM and GPU HBM as a single unified NUMA-like memory fabric.
|
To avoid recomputing attention over all previous tokens at every generation step (O(N²) compute), LLMs cache each token's Key (K) and Value (V) vectors in GPU HBM. In OS terms, if Model Weights are the fixed .rodata binary image, the KV Cache is the per-session dynamic heap—growing linearly with sequence length and concurrency until HBM is exhausted.
| KV Cache / OOM Factor | CPU / OS Analogy | Why It Triggers GPU Out-of-Memory (OOM) |
|---|---|---|
| Unbounded Per-Token Growth O(Batch × Seq_Len) Scaling | Unbounded Per-Connection Heap (Unknown allocation size at start) | Every generated token appends new K/V vectors across all transformer layers (e.g., ~0.3–1.0 MB per token). Because an LLM server cannot predict how many tokens a prompt will generate, concurrent requests simultaneously expanding their contexts can suddenly exceed free HBM mid-generation. |
| Idle Multi-Turn & Prefix Caches Hot vs. Cold Workingset | Cold Anonymous / Page-Cache Pages (Sitting in Tier 0 RAM between requests) | In multi-turn chat, code assistants, and agentic workflows (waiting on tool calls or user think-time), huge KV caches sit completely idle ("cold") in expensive HBM for seconds or minutes. Retaining them in HBM starves active ("hot") requests of memory and causes OOM under traffic spikes. |
| Hard HBM Wall (No Default Tiering) Fail-Fast vs. Graceful Paging |
Running mlock() System-Wide with Swap Disabled
|
On a CPU, memory pressure triggers background reclaim (MGLRU → zswap / CXL / SSD). On a GPU, hitting 100% HBM triggers a hard cudaErrorOutOfMemory crash or forces the serving engine to evict & recompute KV caches from scratch—unless cold KV pages are tiered out to CPU DRAM or compressed memory.
|
To halve memory footprint and double matrix-engine throughput without sacrificing model quality, modern LLMs transitioned from FP32 (32-bit) to 16-bit floating-point formats—Heavily favoring BF16 (bfloat16) over FP16, because dynamic range matters far more than decimal precision:
Wide dynamic range (±3.4×1038), but 2× memory & bandwidth cost.
Narrow 5-bit exponent (max 65,504) suffers from overflow/underflow on LLM outlier activations.
Same 8-bit exponent range as FP32 with fewer noisy mantissa bits—preventing overflow & concentrates entropy.
Why Standard Byte Compressors Fail on Floating-Point KV Caches:
While the exponent bits across KV cache tensors have very low entropy (values cluster tightly in magnitude), the pseudo-random mantissa bits interleaved into every 2-byte element break byte-level string repetition—rendering general-purpose LZ77 algorithms (LZ4, Snappy) completely ineffective (~1.000×), whereas float-aware compression (kvcomp) achieves >1.5× lossless compression even at 4KiB page size.
| Compressors / Block Sizes | 4KiB | 64KiB | 2MiB |
|---|---|---|---|
| kvcomp (custom) | 1.515× | 1.525× | 1.527× |
| Bitcomp (nvCOMP, USHORT, algo=0) | 1.200× | 1.237× | 1.134× |
| Snappy (zippy) | 0.999× | 1.000× | 1.000× |
| LZ4 | 0.996× | 0.996× | 0.996× |
zHBM runs on NVIDIA GPUs just like zram runs on CPUs—while zram compresses cold pages in CPU DRAM, zHBM transparently compresses cold pages in GPU HBM to expand effective HBM capacity and prevent KV cache OOMs without leaving the GPU package.
| zHBM Component | CPU / Linux MM Analogy | Design & Architectural Mechanism |
|---|---|---|
| 1a. PF-based Access Detection | Hardware A-bit |
Because the GPU MMU (GMMU) currently does not support a hardware Accessed (A) bit in PTEs, zHBM detects page references via page-fault based access detection—unmapping candidate PTEs so any GPU access traps into a replayable GMMU fault and marks the page hot.
|
| 1b. Periodic Scanning |
Pressure Driven (MGLRU / kswapd)
|
Periodically scans HBM address ranges to age pages across intervals, separating actively referenced ("hot") working-set pages from unreferenced ("cold") KV cache pages eligible for compression. |
2. Custom Compressor for BF16 (kvcomp)
|
Crypto API (e.g., lz4)
|
Uses kvcomp, a custom lossless KV cache compressor designed specifically for BF16 (bfloat16) tensors and executed directly on GPU Streaming Multiprocessors (SMs) at high throughput (~1.52× compression where LZ4/Snappy yield 1.00×).
|
| 3. Segregated-Fit Allocator |
zsmalloc Pool Allocator
|
Manages compressed pages in HBM using segregated storage (segregated-fit allocation)—grouping compressed blocks into discrete size classes to eliminate fragmentation and enable fast, deterministic allocation and decompression on page faults. |
Comparison of effective throughput and end-to-end latency to access or restore 1 GB of KV Cache per GPU (e.g., ~32K tokens on Llama-3 70B with TP=8) across memory tiers and recomputation.
| Source / Method | Hardware Path & Bottleneck | Effective Throughput (per GPU) |
Latency for 1 GB KV
(~32K tokens @ TP=8)
|
Compute / System Impact |
|---|---|---|---|---|
| Local HBM (Uncompressed Hit) |
Direct SM load/store via HBM3/HBM3e (3.35–8.0 TB/s)
|
3,350 – 8,000 GB/s | ~0.13 – 0.30 ms | Zero overhead; baseline decoding path. |
Uncompressing from zHBM
(In-HBM kvcomp)
|
Read compressed HBM → SM kvcomp → Write HBM (~1.66× HBM traffic)
|
~800 – 2,000+ GB/s (SM-bandwidth bound) | ~0.5 – 1.5 ms | Stays 100% inside HBM; ~4×–10× faster than GB200 LPDDR5X and ~20×–50× faster than x86 DDR5/PCIe. |
| Host DRAM (CPU Memory) |
PCIe Gen5 x16 (64 GB/s raw) via Copy Engine (CE) (or NVLink-C2C: ~400 GB/s) + Host DDR5 (~200–280 GB/s per 4-GPU socket) / LPDDR5X (~400 GB/s per 2-GPU Grace)
|
~35 – 55 GB/s
(x86 DDR5 / PCIe5)
~180 – 200 GB/s
(GB200 LPDDR5X / C2C)
|
~20 – 30 ms
(x86)
~5.0 – 5.5 ms
(GB200 -- less common)
|
Bottlenecked by shared host DDR5/LPDDR5X controllers across GPUs. |
| Distributed KV Storage (SSD / Remote DRAM) |
NVMe-oF SSD (3–10 GB/s) or RDMA (100–400G RoCE/IB) (12–45 GB/s)
|
~3 – 10 GB/s
(SSD)
~12 – 45 GB/s
(RDMA)
|
~100 – 300 ms
(SSD)
~25 – 85 ms
(RDMA -- less common)
|
Competes with East-West inference/training network fabric; high tail latency. |
| Recomputing the KV Cache (Prefill Pass) | Full Transformer prefill (O(S²) attention + FFN across all layers) | ~0.4 – 2.0 GB/s (effective KV gen rate) |
~800 – 1,400 ms
(at 32K)
~6,000+ ms
(at 128K)
|
Scales quadratically O(S²); hijacks 100% of Tensor Cores and stalls active batch. |