Speakers
Description
As the memory hierarchy deepens at both ends, with HBM adding a faster tier on top and CXL adding cheaper, slower capacity below, keeping hot data in the fast tier and shedding cold data downward makes page migration central to NUMA, tiered-memory and coherent CPU-GPU systems (where device memory is exposed as NUMA nodes).
Profiling move_pages(2) on EPYC Zen 6 shows that the folio copy dominates migration (~96% of time for a 2MB THP), making it the primary scaling bottleneck. Yet this path is largely sequential: folios are copied one at a time by a single CPU, while DMA engines, idle cores and memory bandwidth go unused.
To tap that idle hardware, this work separates the folio content copy from the rest of migration and hands it to a pluggable migrator. The batch path:
- unmaps a batch of folios and flushes the TLB once,
- asks a migrator to copy the eligible folios,
- marks copied folios so the move phase skips the per-folio copy,
- completes the move through the existing flow [1].
A faster copy is not the whole story: lower rmap overheads, a simpler migration architecture and a composable core matter just as much. The work therefore proceeds in three directions: accelerating the dominant copy phase, batching the rmap walks and restructuring the migration core so that these and other independent optimizations compose instead of becoming special cases.
Accelerating the copy
Three complementary, measured directions:
1. dcbm - a DMA-based migrator using dmaengine devices, for bulk copy across multiple channels.
2. mtcopy - multi-threaded CPU copy on idle cores, a software fallback where no DMA/offload engine is available [9].
3. Bulk folio_copy() for large folios: one copy over the contiguous range, independent of the batch and offload framework [2].
Measured with move_pages() on 1GB anonymous memory, DRAM -> DRAM, batch copy offload makes 2MB THP migration several times faster: ~6x with PTDMA (Zen 3, 16 channels), ~3.5X with SDXI (Zen 6, 1 channel) and ~3.8x with multi-threaded CPU copy (mtcopy, 8 threads).
Separately, a related demotion effort uses non-temporal stores to reduce cache pollution and CXL-side read traffic [3].
Beyond the copy: the rmap phase
Once the copy is offloaded, rmap work dominates for PTE-mapped large folios (mTHP) because the unmap and restore walks handle each PTE separately. Batching those walks is the next lever, and its payoff grows with folio size: on Zen 3, a 1MB folio reaches ~5.6x over vanilla with DMA offload and rmap batching, compared to ~2x with offload alone [7].
Reworking the core so these compose
Several efforts optimize different migration phases, and each one adds flags or special cases to the current monolithic path [1][3][4]. They also run into the same limitation: ->migrate_folio() receives only enum migrate_mode, which describes the blocking semantics but not the migration intent (for example, promotion versus demotion). So, new behavior ends up either extending migrate_mode or going through a side channel. For example, non-temporal demotion added MIGRATE_ASYNC_NON_TEMPORAL_STORES, while the DMA offload records FOLIO_CONTENT_COPIED in migrate_info.
Two changes are posted as RFC[8]:
-
Pass a small context struct (migrate_control) through migrate_pages() into ->migrate_folio(), carrying mode, reason, and room for future attributes. This separates blocking behavior from intent, so the copy phase knows why the folio is migrating and copy policy (cached, non-temporal, offload, and so on) can follow.
-
Split the single list in migrate_pages() into separate classes, each with its own list and retries: hugetlb folios, movable_ops pages and LRU folios. This removes the type checks scattered through the shared batch path and makes the flow easier to follow.
Discussion
I'd like to use the session to validate the direction of the migration core work and how ongoing efforts should fit together.
Key questions:
- Interface: should migrate_pages() and ->migrate_folio() take a context struct? It touches every caller and callback. Is that one-time churn acceptable, and is this the right shape?
- Policy: Should each policy, such as cache intent, be stated explicitly by the caller, or should the core derive it from the migration reason?
- Core structure: is the per-class migration model right shape for maintainability? movable_ops pages still use page->lru, new_folio_t, put_folio_t. What should a non-folio migration interface look like?
- Synchronous fallback: after the async pass, should the remaining folios be retried through a direct single-folio path, or should they keep reusing migrate_pages_batch() and avoid a second single-folio path?
- Copy helpers: plain memcpy() regresses on CPUs without FSRM. Should bulk folio_copy() use a new copy_pages() helper, following the clear_pages() pattern, or should memcpy() be improved for this case?
- Engine selection: who should select the copy engine: the caller, the provider (as in the current RFC), or the core, based on caller constraints and provider capability? Should folio size, batch size or topology drive that choice?
- Batching trade-offs: What is an acceptable trade-off between throughput and first-folio latency? Should the batch limit be a tunable per platform?
Roadmap:
Migration core redesign and context passing; migration-side rmap batching, bulk copy within large folios, batch copy and DMA offload, Scatter-gather copy (DMA_MEMCPY_SG); SDXI migrator [6] and DCBM driver; topology-aware thread placement for multi-threaded copy and CPU accounting for multi-threaded copy; hot page promotion (pghot [5]) reusing the offload path.
References
[1] Migration batch/offload, RFC V6: https://lore.kernel.org/all/20260630-shivank-batch-migrate-offload-v6-0-da95d7e8b8a2@amd.com
[2] Bulk folio_copy optimization, RFC v1: https://lore.kernel.org/all/20260427142036.111940-4-shivankg@amd.com
[3] Non-temporal demotion, RFC V2: https://lore.kernel.org/linux-mm/20260730-rfc-nt-demote-v2-0-452dbe3b5073@zptcorp.com/#t
[4] migrate_mode/migrate_folio() ABI extensibility (Memory Hotness and Promotion call):
https://lore.kernel.org/linux-mm/37010218-ecad-4f75-961d-686d51b6677d@amd.com
[5] pghot, V8: https://lore.kernel.org/all/20260728054356.291998-1-bharata@amd.com
[6] SDXI, V3: https://lore.kernel.org/all/20260605-sdxi-base-v3-0-4d38ca2bdffe@amd.com
[7] Migration-side rmap batching, RFC V2: https://lore.kernel.org/linux-mm/20260813-migrate-rmap-batch-v2-0-3c5424c555c7@amd.com
[8] Migration redesign and policy, RFC V1: https://lore.kernel.org/linux-mm/20260902-migrate-refactor-shivank-v1-0-9dcca87669c4@amd.com
[9] mtcopy RFC v3: https://lore.kernel.org/all/20250923174752.35701-7-shivankg@amd.com