Speaker
Description
As modern foundation models and their deployments demand exponentially larger memory capacity, high-bandwidth memory (HBM) has become an unprecedented driver of datacenter capital expenditures (CapEx). While the majority of discussions surrounding memory tiering have focused on leveraging disk swap, in-memory compression and CXL.mem to expand system capacity and reduce the total cost of ownership (TCO) of host DRAM, HBM on AI accelerators presents a distinct yet critical frontier for TCO optimization.
This proposal explores the feasibility and architectural requirements for extending the traditional Linux memory management (MM) concepts, such as hot/cold tracking, page migration and data compression, to accelerator-attached HBM. We evaluate existing kernel mechanisms to discuss how much infrastructure can be reused and brainstorm with the audience to seek new opportunities.