Linux Plumbers Conference 2026
507 October, Prague, Czechia
The Linux Plumbers Conference is the premier event for developers working at all levels of the plumbing layer and beyond.
-
-
10:00
→
13:33
Android MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128The Android Micro Conference brings the upstream community and Android systems developers together to discuss issues and changes to the Android platform and their dependencies and interactions with the Linux kernel, allowing for collaboration on solutions for upstream.
Some highlights of progress made since last year’s MC:
- On 16k kernels, a set of recommendations were put together about how to reduce the memory footprint on 16kb kernels
Also related to 16k kernels, work on writing a memory driver that will be used during debugging to allocate any type of memory on the kernel (UNMOVABLE, MOVABLE, RECLAIMABLE, CMA, etc), which was proposed at LPC Tokyo 2025. - Discussions with attendees who had solved similar 4k / 16k compatibility issues pointed toward using a minimal VM. This feedback shifted the approach away from dynamic linker workarounds and led to investigations around AArch64 per-process page sizes to provide a more robust compatibility mode.
- HW/SW Design Recommendations for 16kB Devices also delivered actionable design recommendations to our industry partners, helping ensure their future hardware is natively 16kB compatible.
- The talk on Pixel upstreaming talk helped improve visibility of the project. There was a great conversation with Mark Brown during the talk about regulators which helped nail down the solution and move things along on the list. Connecting with maintainers and developers during the conference helped increase the project's credibility.
- At LPC, we got a chance to communicate detailed plans to transition from ashmem to memfd, and as there were no objections, that work is progressing as outlined and is expected to release in an upcoming version of Android.
Potential discussion topics for this year include:
- How Android is dealing with the memory crunch (likely multiple talks/discussions)
- Android USB Stack Updates: Userspace AOA and Multiport Device Mode
- Cuttlefish in Debian
- Wattson for power analysis
- Increasing Rust kernel driver usage with Android
- Debugging GBL with EFI debug support
- Pixel upstreaming updates
- …and more!
Key Attendees:
- Suren Baghdasaryan
- Kalesh Singh
- T.J. Mercier
- Juan Yescas
- William McVicker
- Alice Ryhl
- Matthew Maurer
- Yifan Hong
- Neill Kapron
- Paul Liu
- Peter Griffin
MC leads:
- Suren Baghdasaryan surenb@google.com
- Amit Pundir amit.pundir@linaro.org
- Mostafa Saleh smostafa@google.com
- Sumit Semwal sumit.semwal@linaro.org
- John Stultz jstultz@google.com
- Karim Yaghmour karim.yaghmour@opersys.com
-
10:00
Android MC Intro 1m
-
10:01
Vendor hooks in Android kernels? Why? What do vendors do with them? 14m
Android’s Generic Kernel Image (GKI) enforces a single binary kernel, restricting Android ecosystem partners from modifying core kernel code. To allow partner-specific optimizations, Google introduced "vendor hooks", based on Linux tracepoints, to act as in-kernel callback registration points. Partners use loadable kernel modules to register handlers for these hooks to implement partner-specific features.
But what are they actually used for?
In this session we'll discuss the motivation behind providing vendor hooks and some of the common optimization patterns implemented by partners across various kernel subsystems including scheduler and mm. The discussion will be based on the actual usage of these hooks across major Android partners.
Speaker: Todd Kjos (Google) -
10:15
Kernel Lock Contention Hotspots Causing Frame Drops on Android Mainline 15m
Problem Statement
Android UI rendering is latency-sensitive — at 90Hz, each frame has only ~11ms time budget. When a UI-critical thread (e.g., RenderThread) is blocked in the kernel, the frame is likely to be dropped.
We profiled 40 popular Android applications on a Pixel 6 (kernel 6.18.0-mainline, 90Hz refresh rate), using ftrace lock_contention events correlated with Perfetto FrameTimeline data, to confirm causal links between lock contention and dropped frames.
The results show that approximately 14% of all dropped frames are causally attributable to kernel lock contention. The contention concentrates in 5 subsystems that account for over 90% of all lock-caused frame drops.
We would like to present these findings and discuss optimization directions with the community.
Top 5 Problems
1. Mali GPU Driver
The Mali GPU kernel driver uses multiple locks in its internal paths such as job submission and scheduling, and the RenderThread is observed to be occasionally blocked on these locks, causing frame drops.
To discuss: What optimization approaches are available for reducing lock contention in the Mali GPU driver? Is this a known issue?
2. mmap_lock
Several mm paths contend on the same per-process mmap_lock on both write side (e.g., mmap, mprotect) and read side (page faults).
RFC PATCH: https://lore.kernel.org/lkml/20260804095135.45897-1-zhanghongru@xiaomi.com/
To discuss: We have frame-drop causation data showing mmap_lock contention still impacts Android UI performance on 6.18. How effective are current mitigations in practice? What remains to be addressed?
3. ZRAM Compression
We observed priority inversion in the ZRAM swap-in path with durations exceeding 20ms. Threads such as RenderThread, Bluetooth, and SystemUI are observed to be blocked on this lock.
RFC PATCH: https://lore.kernel.org/all/20260805005545.66112-1-baohua@kernel.org/
To discuss: Possible approaches to avoid priority inversion in the ZRAM swap-in path
4. Binder IPC Allocator
We observed lock contention in the Binder buffer allocation path with durations exceeding 8ms. Threads such as RenderThread, SystemUI, and app main threads are observed to be blocked on this lock.
RFC PATCH: https://lore.kernel.org/all/20260805152752.1924434-1-zhangbo56@xiaomi.com/
To discuss: Is this a known issue? What approaches can reduce lock contention in the Binder allocation path?
5. cgroup threadgroup rwsem
We observed lock contention in the process fork and exit paths (cgroup_threadgroup_rwsem) with durations exceeding 11ms. The RenderThread is observed to be blocked on this lock.
To discuss: Is this a known issue? What approaches can reduce lock contention on cgroup_threadgroup_rwsem?
Speakers: Mr Barry Song (Xiaomi Corporation), Mr Bo Zhang (Xiaomi Corporation), Mr Hongru Zhang (Xiaomi Corporation) -
10:30
RT tasks in sched-ext support 15m
The upstream Linux kernel explicitly excludes real-time (RT) tasks from the sched-ext extensible scheduler framework, restricting sched-ext BPF schedulers to only manage SCHED_NORMAL/SCHED_BATCH/SCHED_IDLE tasks. However, Android's production workloads present a fundamentally different reality: many performance-critical scenarios — including audio pipelines, camera capture, display composition, and game rendering — depend on tight co-scheduling of RT and non-RT tasks to meet latency and throughput targets simultaneously. Optimizing only the CFS portion of the scheduler in isolation leaves significant headroom on the table.
A concrete example is pipeline scheduling identifying related task chains and placing critical tasks onto specific pipeline CPUs to improve locality and predictability. The problem is that RT tasks remain outside sched-ext control — even when sched-ext builds and protects such a pipeline, an uncoordinated RT task can land on a pipeline CPU, break the intended task-to-core layout, and degrade performance. The same blind spot appears in load tracking, where per-CPU utilization and frequency selection must account for RT load that a BPF scheduler managing only fair-class
tasks cannot see.This talk explores the gap between upstream's conservative stance and Android's practical needs — including why upstream currently draws this line — surveys the combined RT+sched-ext scheduling scenarios that matter most on Android, and discusses what vendor hook infrastructure is currently required from the Android Common Kernel (ACK) to bridge this gap safely and
maintainably in the short term. Beyond that immediate step, the session aims to align customers, Google, and Qualcomm, and to open a discussion with the upstream community on a possible long-term direction in mainline.Speakers: Aiqun Yu (Maria), Tengfei Fan -
10:45
BPF FD Loader: Integrity-Verified BPF on Android 15m
On modern Android devices, eBPF loading is strictly confined to early boot by a privileged loader alongside SELinux lockdown of BPF_PROG_LOAD. While this protects the kernel attack surface, it leaves the platform fundamentally rigid: platform services and hardware partners cannot deploy or activate BPF programs on demand. Furthermore, the approach of individual cryptographic signature verification fails in a decentralised client ecosystem comprising several SoC vendors, OEMs, and partners, where managing runtime key-rings in the Generic Kernel Image (GKI) imposes an unsustainable operational burden.
To address this, we introduce the BPF FD Loader driver in Android. Inspired by finit_module(), it allows the kernel to load BPF programs directly from verified file descriptors rather than untrusted userspace memory.
Because BPF lacks an in-kernel linker, our solution leverages upstream BPF Light Skeletons. At runtime, an authorised userspace daemon simply passes an open file descriptor of this ELF to the Android driver via ioctl(). The kernel driver reads the backing file directly before executing the loader, verifying that it originates from an authenticated, read-only partition sealed by Android Verified Boot (AVB) and dm-verity.
In this talk, we present the end-to-end architecture and discuss our roadmap towards standardising file-based BPF loading in the Android ecosystem.
Speaker: Siddharth Nayyar (Google) -
11:00
CPU Power Regression Testing with Wattson 15m
Wattson is a trace based power estimation tool designed around perfetto to estimate CPU power consumption on ARM64 devices using a statistical per-SoC model. With support of both the Pixel 6 (gs101) and SM8750 upstream, you can now use the Wattson tool to detect CPU power regressions in the Linux kernel. This talk dives into how Google is using Wattson to catch CPU power regressions including:
- What is needed to enable testing.
- Walk-through of the metrics collected and how to use Wattson.
- Examples of regressions Wattson has caught.
- Expanding into other subsystems like GPU and TPU/NPU power estimates.Ultimately, we’d like to connect with the Linux community to see who wants to collaborate on establishing CPU power regression testing upstream.
Speaker: William McVicker -
11:15
Modernizing Android USB: Userspace AOA, Multi-UDC, Type-C, and Framework APIs 15m
The Android USB stack was originally architected around the constraints of early smartphones. Initially designed primarily for phones with a single USB Micro-B device port, the stack lacks the modern API surface required for advanced USB applications. While incremental changes have been implemented out of necessity, a larger refactoring is required to better support modern USB features like Type-C and multiple-UDC topologies.
This presentation will highlight some of the limitations of the existing frameworks, provide an update on current projects, and solicit feedback on the new architecture plans and API surfaces.
Topics for Discussion:
- Userspace AOA via function_fs: Deprecating the out-of-tree kernel AOA driver and moving protocol handling entirely to userspace.
- Upstream Type-C Integration: Hooking the framework directly into the Type-C connector class, allowing applications to query cable capabilities and provide safe role-swap hints to the TCPC.
- Multi-UDC & Per-Port Configuration: Refactoring the framework to break the single-UDC assumption, allowing for independent per-port configurations on laptop-class devices.
- Modular Function Provider Patterns: Replacing legacy broadcast mechanisms with strict provider patterns and robust APIs with proper error handling.
Speakers: George Chan (Google - Android Security), Neill Kapron (Google) -
11:30
Coffee Break 30m
-
12:00
Space sharing between super and data partitions on Android system 15m
The Android Super partition relies on static reserved space to guarantee OTA upgrade capability. However, this fixed allocation method possesses an inherent defect: if the reserved space is too large, it permanently occupies flash memory and cannot be utilized by the data partition; if it is too small, OTA upgrades will fail due to insufficient space. To address the industry pain points of static partition isolation and low flash utilization, we propose a space-sharing mechanism between the super and data partitions. Distinct from traditional reserved space schemes, this technology enables space sharing between the two partitions. The core idea is that in daily usage scenarios, idle physical blocks within the Super partition are shared with the data partition to expand the storage space available to users; when the super partition requires expansion for OTA upgrades or system maintenance, it can dynamically requisition idle physical blocks from the data partition without causing any data corruption throughout the process.
Speaker: Yangtao Li (vivo) -
12:15
Android GBL + Dynamic Partitions 2.0 15m
Hi I'm here to discuss the developments in android GBL + CF support, and talk about android dynamic partitions 2.0 feature (resizable super partition) to accommodate large OS updates.
- Quick GBL + Cuttlefish recap
- Cuttlefish default booting off GBL
- New android_esp partition
- Fastbootd deprecation
- u-boot implementation
Android Dynamic Partitions 2.0 allow for dynamic resizing of super partition vs. userdata during OTA to support larger os updates. We will cover the different solutions proposed and in plans of getting adopted.
- Samsung F2FS patches to support dynamic data partition
- bootloader support
- liblp support in GBL
Speaker: Daniel Zheng -
12:30
Generic Boot Loader on Android platforms 15m
Bootloaders play a vital role in the Android boot process, but the ecosystem has long been fragmented across silicon vendors and OEM-specific implementations. To address this, Google introduced the Generic Boot Loader (GBL), a Rust-based EFI application in AOSP, with the goal of standardizing Android boot logic.
This session focuses on how Qualcomm is adapting its boot chain to integrate GBL while preserving existing firmware and platform capabilities. It describes the Qualcomm-specific changes required to expose core UEFI services such as block I/O, memory allocation, RNG, etc. to GBL, along with the firmware adjustments needed to present these services in a clean and consistent way. The session also covers Qualcomm’s implementation of device tree selection and fixups through GBL-defined interfaces, as well as the handling of Android-critical boot data including kernel command line and bootconfig handoff.
In addition, the talk discusses Qualcomm’s support for GBL-specific UEFI protocols that enable Android requirements such as Verified Boot and Fastboot, and how these fit into the existing Qualcomm boot flow. Overall, the session provides a practical look at the platform changes needed to bring GBL into the Qualcomm boot chain and move Android boot architecture closer to a more maintainable, upstream-aligned model.Speaker: Naina Mehta -
12:45
Actionable Strategies to mitigate memory footprint on 16kb kernels 15m
With the release of Android 17, partners are required to support the 16 KB developer option on devices meeting specific hardware requirements, such as CPU compatibility and memory configurations exceeding 8 GiB. While this transition significantly improves system performance and unlocks hardware efficiencies, Android partners frequently observe a noticeable memory footprint increase, and they ask "How can we reduce this memory footprint?". This regression is primarily driven by internal fragmentation, mapping more memory than what is needed and page-granularity alignment constraints across both userspace and kernel allocations.
This presentation provides a concrete, engineering-focused recipes for Partners to reduce the memory footprint in:
Userspace
- Fileback mappings
- Anonymous Memory
- Shmem memory
Kerne memory
- Kernel modules
- CMA memory usage
- Buffer allocations in drivers
Speakers: Juan Yescas (Google), Kalesh Singh (Google) -
13:00
16KB page cache and mTHP on Android 15m
This topic proposes exploring the possibility of using 16 KB for both
page cache and mTHP on Android.Collaborating with Kalesh Singh, Ryan Roberts, David Hildenbrand and
others, Xiaomi is exploring the use of 16 KB large folios for both page
cache and anonymous memory.We have posted an RFC patchset for large folio support in F2FS, which
has been working well on Android devices so far:https://lore.kernel.org/lkml/20260622160830.324455-1-zhaonanzhe@xiaomi.com/
In this discussion, we would like to present the data we have collected
using 16 KB page cache and mTHP, and discuss the following topics:-
Pros and cons of using 16 KB large folios with 16 KB base pages.
-
Memory fragmentation and direct reclaim issues observed with 16 KB
large folios, along with potential compaction optimizations for a
lightweight mechanism targeting 16 KB folios only. -
Potential Scudo and other component optimizations to reduce memory
footprint while using 16 KB as the management granularity. These
optimizations could also be beneficial when using 16 KB as the base
page size.
Speakers: Barry Song, Bo Zhang (Xiaomi Corporation), nanzhe zhao (Xiaomi Corporation) -
-
13:15
Compatibility for 4KB Applications on Large Page-Sized Systems 15m
As the ecosystem shifts toward larger page sizes for enhanced performance, preserving functionality for legacy applications remains a significant challenge. This talk will introduce a per-process 4KB compatibility mode on native 16KB kernel. Primarily for legacy Android applications, this approach intentionally prioritizes functional correctness over memory efficiency, allowing legacy software to run seamlessly while ensuring native 16KB applications experience zero compatibility overhead.
We plan to discuss to the design, current validation status, and the subsystem integration work required to bridge the gap between 4KB legacy user spaces and 16KB native environments; and outstanding challenges of ensuring consistency and correctness in a mixed-mode e4KB PTEs --> 16KB Folio environment
Speakers: Frederick Mayle (Google), Kalesh Singh (Google)
- On 16k kernels, a set of recommendations were put together about how to reduce the memory footprint on 16kb kernels
-
10:00
→
18:30
Birds of a Feather (BoF): Birds of a Feather (BoF) No A/V "Club D" (Prague Congress Centre)
"Club D"
Prague Congress Centre
53-
10:00
AF_XDP BoF 45m
AF_XDP has grown considerably in functionality and hardware support, but recent netdev discussions have exposed gaps in semantics, queue lifecycle, offloads, and cross-driver testing. This BoF is intended to align on the main pain points and identify concrete next steps.
Topics can, e.g., be:
- AF_XDP semantics: copy vs zero-copy, multi-buffer, UMEM constraints
- TX/RX metadata and offloads
- Queue ownership, RSS contexts, netkit queue leasing
- Driver capability discovery and consistency
- xskxceiver gaps and AF_XDP hardware compliance testing
- Running AF_XDP tests in drivers/net/hw / NIPA CI
- Open issues and next steps
Speaker: Mr Björn Töpel -
10:45
Attested TLS for Confidential AI Agents: Insights and Lessons Learnt from High- and Critical-severity CVEs on Early Attestation 45m
We aim to bootstrap some discussions on the intersections of the following topics:
- Agentic AI
- Confidential Computing (MC on this already exists)
- Attestation
- Attested TLS
To initiate discussion, we will present a few slides to introduce the above, then pose some open questions. We very much welcome presentations by attendees on related topics. If you are interested in presenting, please email me with the title and expected time.
We will share insights from the several CVEs and GitHub Security Advisories (GHSAs) on Early Attestation. In particular, we will share our discovered critical-severity CVEs, such as CVE-2026-92701, CVE-2026-92702, CVE-2026-100835, each of CVSS 9.1, and other critical- and high-severity security concerns against early attestation, and discuss with the community whether early attestation is really necessary.
Speaker: Muhammad Usama Sardar (TU Dresden) -
11:30
Coffee break 30m
-
12:00
Live Update BoF 45m
The primary topic of discussion will be LUO/KHO ABI and what type of backwards and forwards compatibility should be guarantee for the ABI. There are different opinions from various parts of the community and this would be a chance to bring out all views and come to an agreement. More context can be found at these mailing list threads: [0] [1].
Other list of potential topics, time permitting:
- TMPFS preservation [2] [3].
- Getting rid of scratch/bootmem areas.
- KHO enablement on different architectures like Loongarch or PowerPC.
- Making KHO work with crash reservations and KASAN.
This list is open to suggestions.
Speaker: Pratyush Yadav -
13:30
Lunch 1h 30m
-
16:30
Coffee break 30m
-
10:00
-
10:00
→
13:35
Device and Specific Purpose Memory MC "Small Theatre" (Prague Congress Centre)
"Small Theatre"
Prague Congress Centre
105The Device and Specific Purpose Memory Microconference is proposed as a space to discuss topics that cross MM, Virtualization, and Memory device-driver boundaries. Beyond CXL this includes software methods for device-coherent memory via ZONE_DEVICE, physical memory pooling / sharing, and specific purpose memory application ABIs like device-dax, hugetlbfs, and guest_memfd. Some suggested topic areas include, but not limited to:
NUMA vs Specific Purpose Memory challenges
Core-MM services vs page allocator isolation
CXL use case challenges
Hotness Tracking and Migration Offloads
ZONE_DEVICE future for Accelerator Memory
ZONE_DEVICE future for CXL Memory Expansion
PMEM, NVDIMM, and DAX "legacy" challenges
Memory hotplug vs Device Memory
Memory RAS and repair gaps and challenges
Dynamic Capacity Device ABI (sparse memfd?)
Confidential Memory challenges
DMABUF beyond DRM use cases
virtiomem and virtiofs vs DAX and CXL challenges
Peer-to-peer DMA challenges
CXL Memory Pool Management
Device Memory testingWhy not the MM uConf for these topics? One of the observations from MM track at LSF/MM/BPF is that there is consistently an overflow of Device Memory topics that are of key interest to Memory device-driver developers, but lower priority to core MM developers.
Key Attendees:
John Groves
Jason Gunthorpe
David Hildenbrand
John Hubbard
Alistair Popple
Gregory Price
Jonathan Cameron
Dave Jiang
Davidlohr BuesoProgress made on topics discussed at 2025 Plumbers:
Patches available: To online or not online CXL memory?: https://lore.kernel.org/all/20260321150404.3288786-1-gourry@gourry.net/
Patches available: CXL HDM-DB support for Linux: https://lore.kernel.org/all/20260315202741.3264295-1-dave@stgolabs.net/
Patches available: Unifying sources of page hotness information: https://lore.kernel.org/all/20260323095104.238982-1-bharata@amd.com/
Patches available: Protected DMAbufs and its dynamic memory assignment woes: https://lore.kernel.org/all/20250911135007.1275833-1-jens.wiklander@linaro.org/
Patches available: DAMON-based Pages Migration for {C,G,X}PU [un]attached NUMA nodes: https://lore.kernel.org/all/20251208062943.68824-1-sj@kernel.org/
Partially merged: FAMFS Update: Status, DAX Challenges & Use Cases: https://lore.kernel.org/all/69e7d1949ebcc_7d12a10098@iweiny -mobl.notmuch/"Device Memory" Background:
"Device Memory" is a catch-all term for the collection of platform
technologies that add memory to a system outside of the typical "System RAM" default pool. Compute Express Link (CXL), a coherent interconnect that allows memory and caching-agent expansion over PCIe phys, is one such technology. GPU/AI accelerators with hardware coherent memory, or software coherent memory (ZONE_DEVICE::DEVICE_PRIVATE), are another example technology.The problem is how to keep Device / Specific Purpose memory contained to its specific consumers while also offering typical core-mm services. Solutions to that problem potentially intersect mechanisms like numactl, hugetlbfs, memfd, and guest_memfd. For example, guest_memfd is a kind of specific-purpose memory allocator.
-
10:00
Intro/Welcome 5m
-
10:05
A Compressed RAM Service 30m
Compressed RAM (where hardware offloads compression) presents a particularly novel problem for the kernel: the device fundamentally lies about its true capacity - while the kernel is written to assume any
struct pageit can get will always be backed by real capacity.Unlike zswap/zram - these devices provide cacheline/byte access to compressed memory, their memory can remain page-table mapped and page-cache present without generate faults on read-access (no COW, no software decompression step etc).
Lets discuss what it would take to formalize support for a compressed RAM service that otherwise "looks like" normal memory (migratable, mappable, page-cache-able - but maybe not directly writable).
All but one feature required to achieve support already exists in the kernel (or has existed at one time).
- Anon Memory support: Page Table Write Protection (COW, KSM)
- Page Cache support: Clean Cache (clean page eligible only)
- Reclaim and Demotion
- Memory Ballooning (dynamic sizing)
- Free Page Reporting (stale data trimming)
- Memory Tiering (Hotness / NUMA Balancing)
- Strictly controlled NUMA memory allocation (private nodes)
I will present a tested compressed ram service (mm/cram.c) that achieves near-native performance to DRAM under TAOBench and FIO benchmarks, and has been tested under most in-tree filesystems for pagecache correctness.
Speaker: Gregory Price (Meta) -
10:35
Unordered I/O (UIO) support and P2P paths 30m
The PCIe Unordered I/O (UIO) feature (introduced in v6.1) relaxes the strict ordering rules of the PCIe fabric, providing benefits such as the avoidance of head-of-line (HOL) blocking. CXL v3.2 specification incorporates P2P UIO access into the HDM space, enabling the peer access from non-CXL capable accelerators (e.g., GPUs) over the PCIe bus.
Enabling UIO in the Linux kernel involves certain considerations:
- Coherency management while allowing CXL HDM space for both normal and P2P accesses (Ex: coupling with back-invalidation).
- Managing a UIO capable P2P route between a provider (e.g., CXL HDM) and requester (e.g, PCIe GPU).This session will discuss:
- Enumeration of UIO within the PCI and CXL layers.
- CXL-side aspects of a P2P UIO capable memory region (including HDM-DB).
- PCI driver design for UIO and associated P2P-DMA considerations.
- Potential race conditions during enumeration and the mapping P2P routes.Speaker: Arun George (Samsung Semiconductor) -
11:05
CXL Performance Monitoring - Status and Outlook 20m
The CXL specification defines a Component Performance Monitoring Unit (CPMU) register interface for performance monitoring of CXL devices. The Linux kernel includes a CPMU driver that exposes an interface to collect hardware events from CXL memory devices through the perf subsystem.
The current driver supports poll-based event counting using perf stat , providing events such as clock ticks, DDR CAS read/write counts, and CXL M2S/S2M protocol events. However, interrupt-based sampling is not implemented — the CPMU capability register defines interrupt support in the CXL spec, but the kernel driver does not utilize it. As a result, perf top and other sampling-based tools are not available.
This talk gives an overview of the current CPMU driver implementation, its usage with the perf tool, and the limitations encountered. We also look at how CPMUs relate to standard DRAM profiling with uncore PMUs and EDAC. We discuss adding interrupt support and extending CPMU discovery beyond CXL endpoint devices to root ports and switch ports.
Speaker: Robert Richter (Advanced Micro Devices) -
11:30
Coffee Break 30m
-
12:00
Using Platform-provided Hints for Hot Page Detection and Promotion 25m
Hardware platforms continue to expose useful and actionable memory access information to the OS in various ways. Sources of such information include CPU-level instruction/op sampling mechanisms (like AMD IBS and ARM SPE), PMU-based precise sampling (like Intel PEBS), and device-side facilities such as the CXL Hotness Monitoring Unit (HMU). These platform-provided hints can be used by sub-systems like NUMA balancing, reclaim and hot page promotion in tiered memory systems to make better placement decisions, especially on systems with CXL-attached or accelerator memory where the cost of a bad placement is high.
The IBS Memory Profiler [1] on future AMD processors is one concrete example of such a hint source. It is a second, lightweight IBS instance on the chip, dedicated to memory access profiling, and provides per-access information (virtual/physical address, source/target NUMA node) at low overhead. This talk intends to share the experience of working with the IBS Memory Profiler and to look at the ways and challenges in making such a hardware source usable by the kernel in general.
For any such platform-provided hint to be useful, a sub-system that can collate and act on hotness information from multiple sources in a producer-agnostic manner is needed. The pghot subsystem [2] is being developed for this purpose. It moves the hot page promotion engine out of NUMA balancing into a sub-system of its own and exposes a uniform interface so that multiple producers (NUMA hint faults, AMD IBS, CXL HMU, etc.) can feed into a single promotion policy. pghot has been discussed at LPC 2023 [3], LSF/MM/BPF 2025 [4] and LPC 2025, and is a recurring topic on the biweekly Linux Memory Hotness and Promotion community call. In this talk, pghot is treated as the enabling base infrastructure and not as the main subject.
The talk will also touch upon the producer-side ABI that pghot should commit to, where filtering and aging of the access information should reside, the coexistence with NUMA balancing-driven sampling on capable hardware, and a set of workloads and metrics for fairly comparing different hotness sources.
[1] AMD IBS Memory Profiler documentation: https://docs.amd.com/v/u/en-US/69205_1.00_AMD64_IBS_PUB
[2] pghot v7 patchset: https://lore.kernel.org/linux-mm/20260504060924.344313-1-bharata@amd.com/
[3] Using hardware hints for optimal page placement, LPC 2023: https://lpc.events/event/17/contributions/1513/
[4] Unifying sources of page hotness information, LSF/MM/BPF 2025: https://lore.kernel.org/linux-mm/20250319124552.0000344a@huawei.com/Speaker: Bharata Bhasker Rao (AMD) -
12:25
DAMON (Data Attributes Monitoring/Operations Engine)-based {C,G,X}PU [un]attached NUMA Pages Migration 25m
In the last LPC, we introduced a plan to extend DAMON (Data Access MONitor) for migrating pages around NUMA nodes based on their access pattern. Based on on/offline feedback, we continued discussions and development in the upstream community.
As a result of the collaborations, we made a concrete plan and a roadmap for the goal, including support of extensions for h/w features such as AMD IBS, Intel PEBS and ARM SPE. Major stakeholders agreed on the roadmap that currently planned to deliver the first working version by mid 2027. By the time, it will no longer be called DAMON (Data Access MONitor) but DAMON (Data Attributes Monitoring/Operation eNgine).
In this session, we will introduce the evolved design of the in-kernel framework and user interface. We will share the entire roadmap with the expected timeline and the status as of the time of the session. We will discuss pros and cons of the designs and possible adjustments on roadmap, timeline, and priorities of short term future works.
Questions we will discuss include but not limited to below:
- What is the plan and the current status of DAMON for NUMA system memory optimizations?
- Does the plan makes sense?
- Any of planned work items need [de]prioritizations?
- Monitoring extension for AMD IBS, Intel PEBS, ARM SPE, and/or page faults.
- DAMOS integration with new monitoring framework.
- Can we make it just works without any user inputs?
- Is there a plan to use it in real products?
- Will the classical page table accessed bit based access monitoring be deprecated? What is the plan and the status?
- What is the current and future test plan?
- Myths and truths of DAMON for NUMA and general use cases.
Speaker: SJ Park -
12:50
XMFS: A Cross-node Express Memory File System 20m
Cross-node Memory as a Linux Storage Tier:Exploring POSIX-based Shared Memory
Background & Motivation
Memory-semantic interconnects such as CXL 3.0 and Huawei United Bus make remote memory directly addressable. This raises a question for Linux: should cross-node memory become another storage tier that can be exposed through existing POSIX filesystem interfaces?Our Exploration: XMFS
To explore this question, we built XMFS, an in-kernel prototype that exports cross-node shared memory through a standard POSIX filesystem interface. Initial experiments with metadata-intensive workloads and container image sharing suggest the approach is practical.Engineering Challenges Encountered
During development we repeatedly encountered four areas where existing kernel abstractions are either missing or insufficient:
• Metadata synchronization: Current VFS and filesystem metadata management assume a single kernel instance coordinating namespace and inode updates. What generic kernel mechanisms are needed to efficiently synchronize metadata across multiple kernels without introducing centralized bottlenecks?
• Software cache coherence: The Linux page cache and MM subsystem rely on hardware cache coherence within a machine. When memory is shared across nodes without hardware coherence, should software cache-coherence policies live inside individual filesystems, the MM subsystem, or new generic kernel infrastructure?
• Integrating hybrid transports: Linux currently exposes memory-semantic interconnects and RDMA through different subsystems and programming models. How should these be integrated so that filesystems can transparently use both without exposing transport-specific behavior to applications?
• Capacity management & Tiering: Linux already provides memory tiering, page migration, and NUMA-aware memory management. Should cross-node shared memory participate in these mechanisms, and what interfaces are needed between filesystems and MM to support it?Topics for Discussion
Rather than presenting XMFS as a finished filesystem, we hope to discuss whether Linux should treat cross-node memory as a first-class storage tier, which abstractions belong in VFS/MM rather than individual filesystems, and what a practical upstream path could look like. We believe these questions will become increasingly relevant as memory-semantic interconnects become more widely available.Speaker: Yifan Qiao -
13:10
DAXFS: What Does Shared-Memory Storage Need from DAX, VFS, and CXL Memory Management? 20m
When multiple kernels or CXL-connected hosts share byte-addressable memory, every existing filesystem option pays a copy per participant: tmpfs replicates content N times, erofs and fscache keep a private page cache per kernel. To ground the discussion we bring DAXFS, a prototype filesystem that runs directly on DAX memory with no block layer: one shared namespace, a cooperative page cache in the shared region, and mounts backed by a raw physical range or a dma-buf, so GPUs reach file data zero-copy through the backing buffer. Unlike famfs, where a userspace master preallocates files and distributes metadata to read-only clients out of band, DAXFS is masterless: metadata lives in the shared region itself and any participant can create files and COW-write pages concurrently through lock-free CAS updates, no metadata server, no out-of-band log.
The killer use case we are building toward is checkpoint and restore of AI agent sandboxes. Thousands of short-lived agent instances share one mapping of the base rootfs and model weights, with writable state as COW overlay pages in shared memory. Today the image carries a single shared overlay; giving each sandbox its own sealable overlay would make checkpoint a seal operation and restore (or forking N agents from one checkpoint) just a map: no gigabytes serialized through a block device, cold start drops from minutes to seconds.
Building it surfaced gaps bigger than one filesystem, and we want the MC to weigh in:
-
DAXFS bypasses fs/dax entirely (memremap plus raw PFN mappings) because fs/dax assumes a filesystem privately owns its dax_device. How do shared multi-writer DAX mappings become supported, not a special case?
-
Is a filesystem the right ABI for pooled memory carved into named shareable objects, or should this be device-dax, a sparse memfd (per the DCD discussions), or guest_memfd-shaped?
-
DAXFS mounts from an imported dma-buf today; exporting per-file extents as dma-bufs is the natural next step. What are the lifetime and coherence rules when the exporter is a filesystem over shared device memory, with P2P DMA targeting it?
-
Who owns allocation when N hosts share a Dynamic Capacity Device, and how does DCD extent release interact with live filesystem data?
-
Where is the line between memory core MM manages and memory a cross-kernel service manages, and how do hotness tracking and migration see pages no single kernel owns?
Speaker: Mr Cong Wang (Multikernel Technologies) -
-
10:00
-
10:00
→
18:30
LPC Refereed Track "Small Hall" (Prague Congress Centre)
"Small Hall"
Prague Congress Centre
215-
10:00
Untangling convoluted performance regression 45m
In this talk I will speak about a performance regression in DB2 backup speed reported by one of SUSE's customer last year. Due to various reasons the analysis was rather convoluted so I will go through the dead ends we have explored as well as leads which eventually allowed us to track down and fix the problem. Overall we demonstrate on a practical example how various tools for analyzing IO related performance regressions can be used.
Speaker: Jan Kara -
10:45
The slab allocator sheaves post-mortem 45m
Sheaves are a new percpu caching layer for the Linux kernel's slab allocator (specifically, its only remaining implementation, SLUB). To some extent it's a return to the former SLAB implementation's percpu arrays (callled magazines in the original Bonwick's paper), but avoiding the pitfalls that the SLAB implementation had, thus attempting to get the best of both SLAB and SLUB approaches.
In 6.18 sheaves were merged and enabled for maple node and VMA caches. Later in 7.0 they were enabled for all caches and the original cpu slabs and cpu partial slabs caching layer was removed. This talk will discuss the new implementation, explain the tradeoffs involved, the challenges and performance regression reports encountered on the way. We'll also look at the lessons learned, and ongoing/future work that the sheaves caching has enabled.
Speaker: Vlastimil Babka (SUSE Labs) -
11:30
Coffee Break 30m
-
12:00
Modern Developments with NFS in Linux 45m
The Linux Kernel's NFS server and client has been undergoing a lot of changes recently. This talk will cover some of the latest developments in the Linux NFS Client and Server in the last few years. Including:
- Dynamic threading
- New iomodes (buffered, direct and dontcache)
- Directory delegations
- POSIX ACLs for NFSv4
- Signed filehandles
- Delegated timestamps
- Multigrain timestamps
- Netlink upcalls
Speaker: Jeffrey Layton -
12:45
Modernizing Kernel Boot Options: Resolving cmdline and Bootconfig Discrepancies 45m
The current Linux kernel command-line subsystem is very simple and easy to define, but it seems to have several problems, such as inconsistent API naming, drivers and the kernel sharing the same command-line options, and discrepancies between documentation and command-line option definitions. Furthermore, some options require special handling and cannot be supported by Bootconfig, an extension of kernel command-line options, but there is no way to indicate which options these are.
This session will present ideas for solving these problems, or discuss whether to continue as is.
Speaker: Mr Masami Hiramatsu (Google) -
13:30
Lunch Break 1h 30m
-
15:00
Nova — Building an NVIDIA GPU Driver in Rust Upstream 45m
Nova is an open-source NVIDIA GPU driver being developed entirely upstream from day one — and in Rust. This talk presents the current status and roadmap of the project, describes the upstream development process, and dives into the driver's architecture and how Rust shapes it.
On the technical side, we present Nova's architecture: the split into nova-core, nova-drm, vGPU support, and fwctl, the rationale behind this decomposition, and how the components interact across bare-metal and virtualized environments. We follow up with implementation details showing how Rust's type system and ownership model help enforce the boundaries between these components at compile time in the context of the driver model and the DRM subsystem — and show why the DRM subsystem provides a challenge in this regard.
Finally, we examine the Hardware Abstraction Layer (HAL) architecture common in GPU drivers, where multiple generations of hardware must be supported through composable abstraction layers. We discuss how Rust's trait system and generics provide stronger compositional guarantees than C when building and maintaining these layered abstractions.
Speakers: Danilo Krummrich, John Hubbard (NVIDIA) -
15:45
Long-term latency monitoring of real-time Linux systems 45m
With the PREEMPT_RT configuration being merged for the 6.12 release a milestone of a twenty years lasting journey was reached: Linux officially became an RTOS! During that time many technical issues have been resolved and many features have been added that made Linux even better even for non real-time users. A specific challenge which had to be tackled was testing and proving the real-time behavior. While for classical RTOSes traditionally a path analysis was carried out, this is close to impossible for a modern operating system such as Linux - not just because of the complexity of the software: Modern processors do come with a lot of performance with the price of being non-deterministic due to several levels of caches, speculation engines and other techniques. As a result the real-time behavior of modern systems has to be evaluated. This is why the OSADL QA Farm was born, doing comparable measurements on a huge variety of systems collecting long-term data to prove stability in the field. But even after 20 years of operation work is not done yet. Linux is evolving rapidly and the test scenarios (also for real-time) have to adopt. Apart from that Open Source RTOSes are also approaching small processors. This is why the OSADL QA Farm was recently extended with Zephyr tests.
This presentation gives an overview on best practices for evaluating the real-time behavior of Linux (and other) systems, sharing the experience from the OSADL QA Farm. It also wants to serve as a basis for discussion on how measurements shall be carried out in future and how data can be efficiently shared.Speaker: Jan Altenberg -
16:30
Coffee Break 30m
-
17:00
Firmware-mediated accelerators for edge AI 45m
A well-represented class of accelerators for AI in edge deployments aren't programmed directly by the Linux kernel in the host CPU, but by firmware running on a companion core.
During the first half of 2026 alone, we have seen three different drivers submitted to the mailing list for this specific type of hardware, by their respective vendors or on their behalf (TI C7x, NXP Neutron and Qualcomm QDA).
This talk will describe in detail a proposal that aligns with the upstreaming requirements in the drm/accel subsystem and streamlines the functionality that is needed inside the Linux kernel and the UAPI. The architecture will be described in detail: firmware, kernel and userspace.
An important benefit of the approach is that users will be able to accelerate their workloads without having to integrate any vendor BSPs or vendor-specific software.
Speaker: Tomeu Vizoso (NPU drivers - Independent contractor) -
17:45
20 years of pahole : alive and kicking! 45m
The first cset for pahole is from October 24, 2006, a long time ago, from the first goals of helping reorganize the Linux kernel networking data structures, that was done beyond my expectations, helping countless open source projects to view its data structures with great precision and flexibility, to becoming a swiss army knife tool to convert type information from DWARF to CTF and, crucially, to BTF, becoming part of the kernel build process to enable BPF CO-RE, it has kept earning its keep.
Recent advances in DWARF tag and language support, plus features that exist but aren't well publicized, are subjects I want to communicate.
Coverage analysis, a quickly growing set of regression tests, support for DWARF tags for C++ and Rust concepts, support for DWZ, partial units, supporting modern DWARF present in distro userlands are topics of recent improvement that, by the 20th anniversary, will surely be ready to talk about.
Integration with perf, namely in having pahole and perf work together in areas such as data-type profiling has been a perennial source of requests that, by now, with some help from new and controversial friends, should finally become a reality.
Speakers: Alan Maguire (Oracle), Arnaldo Carvalho de Melo (Red Hat Inc.)
-
10:00
-
10:00
→
13:30
Linux System Monitoring and Observability MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128The Linux System Monitoring and Observability MC brings together developers, maintainers, system engineers, and researchers to tackle unsolved problems in understanding, monitoring, and maintaining the health of Linux systems at scale.
Engineers managing millions of Linux servers face monitoring and observability challenges that no single team can solve alone. This track provides a forum to surface those challenges, share partial approaches, and leave with concrete next steps.
The goal is to have these engineers together to discuss the direction and strategy for better monitoring of Linux systems.
Track Objectives
- Surface the most pressing unsolved problems in Linux monitoring and observability
- Identify gaps in existing kernel interfaces, tooling, and infrastructure
- Build consensus on priorities and approaches for the upstream community
- This track invites participants to bring their hardest open questions, pain points, and gaps in current tooling, so the community can collaboratively work toward solutions.
Target Audience
- Hyperscaler Engineers: System reliability engineers who encounter monitoring gaps at scale that others may share
- Kernel Developers: Contributors working on tracing, performance counters, and diagnostic interfaces who want to understand real-world pain points
- Monitoring Tool Developers: Creators of observability platforms who have hit kernel or infrastructure limitations
- System Administrators: Operations teams who can articulate what breaks, what's missing, and what's too hard
- Performance Engineers: Specialists who can identify where current observability falls short for optimization work
Problems involving any of the following (but not limited to) are in scope:
- eBPF/BPF: Tracing and monitoring programs: limitations, missing features, safety constraints
- ftrace/perf: Kernel tracing infrastructure gaps
- Runtime Sanitizers: KFENCE, KASAN: coverage gaps, performance trade-offs, production usability
- Hardware Interfaces: EDAC, MCE, ACPI error reporting, missing integrations, inadequate interfaces
- bpftrace, systemd, netconsole: usability and scalability issues
- kdump/crash/drgn: crash analysis workflow pain points
- perf, memory profilers, below, strobelight, OpenTelemetry: analysis gaps and scaling challenges
Things that made progress given and were discussed in the Micro conference:
-
Lack of NMI on some architecture and how to collect information about it. (Breno)
https://lore.kernel.org/all/rs4igmsjrm6r2aio4nbe5jos3vcqk2u4bjhltjwtj2pn3cquip@kv3grgec7qrb/
https://lkml.org/lkml/2026/3/30/1280 -
Improvements in the LAVD monitoring system (Gavin Guo)
-
Page owner tracking (Mauricio)
https://lore.kernel.org/all/20251205231721.104505-1-mfo@igalia.com/ -
Kmemleak detection in the fleet (Breno)
https://lore.kernel.org/all/20260323-kmemleak_report-v1-1-ba2cdd9c11b9@debian.org/ -
Memory failure and clean crashes
https://lore.kernel.org/all/20260413-ecc_panic-v3-0-1dcbb2f12bc4@debian.org/ -
Track kernels doing kexec
https://lore.kernel.org/all/20260309-kho-v8-0-c3abcf4ac750@debian.org/ -
TCP Reset Observibility (Jason)
https://mailarchive.ietf.org/arch/msg/tcpm/d27ntz9UM4tb4-cxJNxfYw8zCSE/ -
Always-on 7x24 network latency monitor (Jason)
-
Who is planning to submit topics being discussed in the MC to SOSP 2026
-
Relay monitoring (Jason)
-
Diagnostic check and api in relay and future upstreaming discussion
-
Improving memcg statistics collection (JP)
https://lore.kernel.org/all/20260401203752.643259-1-jp.kobryn@linux.dev/ -
General improvements for DAMON (SJ)
-
10:00
Performance Tools and Prompts 20m "Club H"
"Club H"
Prague Congress Centre
128A new era is upon performance engineering where we ask AI to do much of our analysis and thinking: we run "performance prompts" (and skills and agents) instead of "performance tools." These can analyze all sources, including flame graphs, bcc and bpftrace tools (eBPF), Ftrace, PMCs, and MSRs; they can also propose new metrics and tools. How well does it currently work, and what does this mean for the future of system monitoring and observability, including for flame graphs and eBPF tools. There are many questions to ponder (more questions than answers), including: Should one day we include prompts, skills, and agents in the kernel tree, just as we do performance tools? This talk discusses questions and challenges that we will all likely be facing, and new opportunities, including example performance prompts.
Speaker: Brendan Gregg (OpenAI) -
10:20
Unified collection of kernel bug reports 15m "Club H"
"Club H"
Prague Congress Centre
128Debugging end user kernel issues is a problem
Reporting bugs is hard - important information is scattered across many log files, crash files, sysfs entries, etc. Most end users don't know where or how to file a bug. They might not even know that a crash has happened (just a glitch on the screen or a log entry somewhere).
When a bug is reported, triage can be difficult. Likely only one report but is it really a one-off issue, never to be seen again? Was all the important info included? Were the logs collected but two hours after the bug occurred and have wrapped? ...
Bugs can cascade - a non-fatal issue can sometimes lead to a fatal issue later in time. A timeline of reports from a single system/boot is needed to identify cause/effect relationships. Bugs can also cross subsystem boundaries and a bug logged against driver X might really have been caused by module Y but team Y might never get to know about it because team X didn't pass it on properly (or at all).
Solution: Remove the user
Automatically generate a dump file - a daemon can ensure it captures all important information (e.g. devcoredump and other sysfs entries plus dmesg, syslog, etc. plus timeline info such as boot time, hashed processor id, etc.) and immediately at the time of the bug occurrance. Daemon then sends the bug report a server.
First level server can be local, but ultimately bugs are sent to a global server. Server does first step triage - build timeline of bug reports with the same session id, check for known instability issues earlier in the timeline, match new bug to existing entries, etc. It would have plugins to help analyse driver specific data blocks (e.g. devcoredump files) for customised triage. Sends notification to driver/module owner when a new bug is identified and provides database for investigating bugs - developers can see how many times a given bug has really occurred, what platforms/hardware/kernels are affected, etc.
Speaker: John Harrison (Igalia) -
10:35
What Bugs Matter? Continuous Linux Kernel Bug Monitoring 20m "Club H"
"Club H"
Prague Congress Centre
128Linux kernel bug reports vary widely in quality: some are duplicates or false positives, while others describe exploitable vulnerabilities or issues with little practical impact. Maintainers still need to determine what is real, what matters, and what deserves attention first.
At the same time, bug findings are increasingly fragmented across different tools and workflows, including fuzzers, static analyzers, human analysis, AI-assisted tools, and patch-review systems. For example, Sashiko may uncover pre-existing bugs while reviewing patches, but those findings can remain scattered across individual review reports rather than being collected and tracked in one place.
We are developing a continuous Linux kernel bug monitoring and veirfy platform. The platform aggregates findings from multiple sources, maps them to kernel subsystems, and gives maintainers a view of currently observed bugs affecting the code they maintain. For each finding, it attempts to generate and execute a reproducer on a KASAN-enabled kernel to help validate the report. It also assesses potential impact. For example, whether a bug may enable privilege escalation or has limited practical impact. Confirmed bugs can also include AI-generated analysis and draft patches to help reduce the effort required to understand and fix the issue.
Building on the mailing-list discussion[1], we would like to explore what actionable kernel bug observability should look like. How should heterogeneous bug signals be validated, prioritized, and presented to subsystem maintainers? How can automated analysis help distinguish important bugs from noise without creating yet another stream of low-quality reports?
[1] https://lore.kernel.org/all/20260830115546.3942129-1-yuantan098@gmail.com/
Speaker: Yuan Tan -
10:55
Detecting Unhealthy Kernels at Scale 20m "Club H"
"Club H"
Prague Congress Centre
128This talk explores Meta's approaches to identifying production issues rooted in the Linux kernel. Our investigation process, spanning detection, correlation, and deep triage, relies on a unified framework of telemetry and monitoring tools. A common challenge arising in hyperscale kernel releases is that of robustly comparing kernel performance across lifecycle phases, despite vast variations in deployment size. Our findings highlight the efficacy of combining flexible data aggregation with statistically grounded analysis, to streamline and automate the release process.
Speaker: Sam Crossley (Meta) -
11:15
We don't know what we're looking for until we find what we're looking for 15m "Club H"
"Club H"
Prague Congress Centre
128What started out as an investigation for potential isolation issues in the fleet led us down into a few different rabbit holes. We started with trying to find patterns where the kernel lock holders are unable to run, causing other tasks to wait on said locks. We'll discuss ways to monitor said patterns across the fleet at a broader scale and then zoom into some interesting case studies to see what we found including:
- OOM killer prints to the console while holding a lock
- Act of monitoring causes isolation issues in the fleet for sensitive services
- Hard partitioning from CPUSETS causes lock holder unable to get CPU time
- Chain of lock waiters -> Lock holder sleeping on synchronize_rcu -> CPU holding up grace period
Also we'll discuss how to expand on monitoring for scheduling issues and/or issues in the fleet that causes stalls that prevent system wide progress
Speakers: David Dai, Shakeel Butt -
11:30
Coffee Break 30m "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128 -
12:00
Following the bread crumbs: Tracing packets through the kernel with per-skb BPF metadata 20m "Club H"
"Club H"
Prague Congress Centre
128Operating a network at scale, we routinely need to answer a deceptively simple question: what path did this packet take through the kernel, and where did it get dropped? On our edge, a single packet can cross several network namespaces, get encapsulated in GRE or IPIP, be encrypted with IPsec, and pass through multiple nftables chains before it leaves the box. Tagging the payload the way we tag L7 requests is not an option here: it's opaque, encrypted, and rewritten along the way.
We'll walk through our attempt to build an
mtr-like tool that, instead of a list of hosts, shows the kernel waypoints a packet hits — interfaces, netns boundaries, tunnel encap/decap, IPsec encrypt/decrypt, nftables verdicts, and drops — each with a timestamp.We started with Retis, Red Hat's eBPF-based packet tracer, which already does most of the heavy lifting. We'll cover what worked out of the box and where we hit its limits:
- It correlates events by hashing the skb data area (
skb->head); we found
tracking byskbaddress to be both cheaper (~15% fewer trigger events on a busy box) and a better fit. - The "tcpdump model" mismatch: Retis filters at every probe, capturing many
flows. For a traceroute-like UX we instead want to lock on to a packet once,
then have downstream probes fire purely on identity.
This is where the observability problem becomes a kernel interface problem. To follow a packet across transformations, and to let BPF programs annotate a packet at one hook and read it back at another, we need per-packet state that lives and dies with the
sk_buff. The BPF-map "side-stash" works today but is costly: to avoid leaking entries it has to hook theconsume_skbrelease path, which fires for every one of the hundreds of thousands of packets passing through the system each second, even when we only care about a handful of them. So we started an upstream effort to add a better building block: a BPF metadata buffer embedded in an skb extension chunk.We'll update the audience on where that work has moved since we last presented it at Netdev 0x1A in July, and on the challenges we faced wiring Retis up to use it.
We'd like to hear whether others have felt the same gap in production, and if not, how they observe a packet's journey through their systems today.
Speaker: Jakub Sitnicki (Cloudflare) - It correlates events by hashing the skb data area (
-
12:20
Crash Dumps at a Bare-Metal Cloud Provider - Capture, Dedup, and AI-Assisted Analysis. 20m "Club H"
"Club H"
Prague Congress Centre
128CoreWeave is a bare-metal cloud provider. A typical machine in our fleet has at least 2TB of RAM, and customers tend to run close to that limit. That has two immediate consequences: the crash-kernel memory reservation matters a lot (every GB we reserve is a GB the customer doesn't get), and when a machine panics the resulting crash dump is huge. Because it's so big, the downtime during a crash, and how soon we can hand the machine back to the customer, ends up gated almost entirely on makedumpfile. At the fleet level it gets even more challenging as bugs triggered by a bad kernel upgrade or some fancy customer workload lead to getting tons of crash dumps at once and all of them are due to the same bug. Finally there’s us, a lean team of 2 kernel engineers at CoreWeave that have to look at all the incoming crash dumps. We lean on AI a lot, at least for triaging failures and knocking out the easy RCAs. That field has progressed and gotten more accurate over time but we still get the occasional confident hallucination.
Our ideas and projects so far:
-
(Project in-progress at the time of submission) Fingerprint the failure before we even try to capture a dump. Compute a cheap signature of the crash from within the crash kernel, compare it against what we've already seen, and if it's a known one, skip collection altogether. (This probably needs a network connection out of the crash kernel to check against a central store, and we're not sure yet how painful that is to pull off)
-
(Idea) Decouple capturing the crashed memory from turning it into a proper dump. Instead of running makedumpfile to completion before we reboot, grab the crashed memory (maybe with some very light filtering), reboot early to give the machine back, and do the heavy transform/filter later. (There may be Kexec HandOver-related discussion here).
-
(Idea) An alternative on the above that we're weighing is to skip the general-purpose dump entirely and use drgn+sdb to take ultra-filtered dumps on the spot, driven by a preconfigured list of data structures we care about.
-
(Project in-progress at the time of submission) Use out-of-band kernel introspection from the DPU to grab ultra-filtered dumps the moment we know a panic happened, without even waiting for the crash kernel to come up (we got drgn+sdb on the DPU reading host memory over RDMA).
-
(Implemented) For the AI triaging/RCA side, hand the agent a Docker container preloaded with the crash dump plus matching debug info, the kernel source for that exact version, and some hints on where to look for existing issues (bug trackers, the mailing list, etc.). We also drop in a small agents.md whose main job is to keep the agent from spinning in endless loops.
Topics we'd like to discuss:
-
How do folks capture dumps on multi-TB hosts at fleet scale? Full vs. aggressively filtered, local NVMe vs. a network target, and what capture times are people living with? And how do you settle on the crash-kernel memory reservation?
-
Is anyone else deduplicating crash dumps? If so, what crash-signature/dedup schemes are you using, and is there any appetite for converging on a shared, drgn-based triage library instead of everyone rolling their own?
-
How are people using AI tools for crash dump analysis, and which setups have worked best? How do you keep the confident hallucinations in check?
-
Is anyone else doing out-of-band introspection of production hosts (DPU, BMC, PCIe P2P, RDMA)? How do you handle it, and do you use drgn?
-
How do you keep historical crash dump data around? Jira and/or database plus Grafana? Something else?
Speaker: Serapheim Dimitropoulos (CoreWeave) -
-
12:40
Breakpoints in drgn 20m "Club H"
"Club H"
Prague Congress Centre
128Breakpoint support for drgn is currently underway (https://github.com/osandov/drgn/issues/626). Support is planned for:
- KGDB
- Linux kernel QEMU guests
- Live kernels in production (a sequel to my 2023 LPC session).
- Userspace processes (ptrace)
Each of these targets has its challenges and deficiencies. I will discuss these problems as well as next steps.
Speaker: Omar Sandoval
-
10:00
→
13:30
Nova GPU & DRM Rust Workshop "Club B" (Prague Congress Centre)
"Club B"
Prague Congress Centre
53This workshop will center on Nova, the upstream Rust-based kernel driver for NVIDIA GPUs, and on Rust in the DRM subsystem in general.
On the Nova side, discussion topics will include the design and evolution of the firmware APIs exposed by the GPU System Processor (GSP), in particular the new GMC APIs, as well as user-space submission interfaces, compute APIs, and interactions with the core kernel (device / driver APIs; locking and lifetimes; memory management APIs). Beyond Nova, the workshop will cover the shared Rust DRM infrastructure — device / driver core and initialization, TTM/GEM, GPUVM, job submission and scheduling — and how to properly tie these components into driver lifecycle design.
Potential key participants are members of the Nova team at NVIDIA and Red Hat, contributors from the DRM and Rust-for-Linux communities, and developers of parallel Rust driver efforts such as Tyr (Arm, Collabora, Google), the Asahi AGX driver, and rvkms.
The workshop aims to keep Nova and the Rust DRM infrastructure closely tied to the needs of the graphics / compute stack in Linux, and to foster collaboration around shared challenges in GPU driver design.
-
10:00
→
13:30
Safe Systems with Linux MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53Description/Motivation
As Linux continues to be deployed in systems with varying criticality constraints, the need for consistent linkage between requirements, code, and tests becomes increasingly important at the higher assurance levels. Establishing such traceability can improve development and testing efficiency, supports necessary analysis, and reduces long‑term maintenance risks.
This MC addresses key challenges in expectation management, documentation, testing, and artifact sharing within the Linux kernel ecosystem. While tests are commonly contributed alongside code, the underlying requirements they validate are typically not documented in a structured manner. This creates significant “tribal knowledge” within subsystems, leading to technical debt when maintainers stop working on subsystems or subsystem expertise is lost in other ways.
Given the feedback from last year's "Safe Systems with Linux" miniconference[1], we are pivoting away from the original guidance in 2024 of publicly documenting the kernel design in the code, and focusing on expressing the requirements and traceability as side car data structures. This makes the information machine‑readable, maintainable, and scalable without inhibiting kernel development velocity.
Building on the 2025 discussions and the progress made over the past year, the goal of this MC is to gather wider input from maintainers and developers across different subsystems on the proposed approach and its practical adoption in upstream workflows.
Potential Topics
-
Technical Debt Reduction
How capturing expected behavior and design intent as structured requirements enables maintainers to validate functionality during refactoring (e.g., language transitions such as C→Rust) and supports onboarding of new contributors. -
Requirements-Driven Testing
How linking requirements to specific tests and code paths can increase test efficiency, improve coverage understanding, and allow automated validation of expected behavior. -
Semantic Aspects of Kernel Requirements
How to document expected kernel behavior while accounting for design constraints, architectural dependencies, and interactions between subsystems. -
Progress on Linux Kernel Requirements Framework
How the SPDX‑based template for low-level requirements is evolving, what has been learned from early pilots, and how broader adoption as sidecar metadata could be enabled. -
Practical Implementation Challenges
How to balance detailed requirements documentation with the realities of fast‑paced kernel development, and what workflows or structures can minimize friction. -
Required tools for automation
How tooling can generate, validate, and track requirements, tests, and other work products, increasing dependability and reducing manual effort throughout kernel development. -
Connecting with Other Kernel Quality Initiatives
How the requirements approach can integrate with existing kernel quality, testing, and sustainability initiatives, and where collaboration can reduce duplication and improve adoption. -
Industry Adoption
How safety-critical industries are beginning to leverage these developments for certification and compliance purposes.
How their safety engineers can participate in contributing formalized requirements to the kernel and providing linkage. -
Requirements as an Education Tool
How linux kernel documentation can mine the requirements, and help new contributors understand kernel functionality and design intent and attract new upstream developers
Objective
The MC aims to bring together kernel maintainers, developers, safety architects, and industry stakeholders to advance the adoption of structured requirements and traceability practices to complement the Linux kernel existing development workflows. It will focus on aligning documentation, testing, and tooling in a coherent workflow, and addressing remaining technical and organizational challenges in building dependable and safety‑relevant systems with Linux.
Potential Participants
- Gabrielle Paoloni
- Chuck Wolber
- Luigi Pellecchia
- Alessandro Carminati
- Paul McKenney
- Julia Lawall
- Sasha Levin
- Steve Rostedt
- Thomas Gleixner
- Shuah Khan
- Gustavo Padovan
- Wolfram Sang (Renesas BSP)
- Kate Stewart
- Philipp Ahmann
- Nicole Pappler
References
[1] LPC 2025 Safe Systems with Linux MC: https://lpc.events/event/19/sessions/221/#20251212
-
10:00
Aspects of Dependable Linux Systems 10m
In regulated industries, Linux is widely used due to its strong software capabilities in areas such as dependability, reliability, and robustness. These industries follow best practices in terms of processes for requirements, design, verification, and change management. These processes are defined in standards that are typically not accessible to the open source kernel community.
However, since these standards represent best practices, they can be incorporated into structured development environments like the Linux kernel even without the knowledge of such standards. The kernel development process is trusted in critical infrastructure systems as it already covers many process elements directly or indirectly.
The purpose of this session is to initiate a discussion on what is currently available and what may be missing in order to enhance the dependability and robustness of Linux kernel-based systems. How can the artifacts be connected? Where are the appropriate places to maintain them? And who is the best responsible for each element of the development lifecycle?
Speakers: Kate Stewart (Linux Foundation), Philipp Ahmann (Etas GmbH (BOSCH)) -
10:15
Towards Program Verification of the Linux Kernel Library XArray 25m
XArray is a data structure used in many Linux kernel components, most notably the page cache. Its API contracts are not precisely documented, making it challenging to understand caller obligations and which invariants any changes to the implementation must maintain over the structure. Since its integration into the Linux source tree in 2019, errors in the use of the library have caused bugs such as memory leaks [1, 2], and the library code itself has suffered from bugs like race conditions and null pointer dereferences [3, 4]. Recent case studies [5, 6] suggest that applying program verification to Linux kernel source code is a promising means to make requirements structured, explicit, and machine-checkable against the implementation.
This talk presents progress verifying the core load and store APIs of the XArray library. So far, we have verified the xa_load() function and its callees, and are working on the verification of xa_store(). We will discuss how the structured specification of the original library can be used to build confidence in the Rust re-implementation that is under active development [7]. By formalizing and checking the API’s precise contracts, we can validate the parity of the Rust port’s behavior in critical kernel clients.[1] Matthew Wilcox. “mm/huge_memory: Fix xarray node memory leak.” url: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?h=v6.19&id=69a37a8ba1b4
[2] Yang Yang. “swap_state: update shadow_nodes for anonymous page.” url: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?h=v6.19&id=5649d113ffce
[3] Matthew Wilcox. “XArray: Disallow sibling entries of nodes.” url: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?h=v6.19&id=63b1898fffcd
[4] Matthew Wilcox. XArray: Fix xas_create_range() when multi-order entry present. url: https://github.com/torvalds/linux/commit/3e3c658055c002900982513e289398a1aad4a488
[5] Julia Lawall, Keisuke Nishimura, and Jean-Pierre Lozi. 2024. Should We Balance? Towards Formal Verification of the Linux Kernel Scheduler. SAS 2024.
[6] Julia Lawall, Keisuke Nishimura, and Jean-Pierre Lozi. 2025. Understanding Linux Kernel Code through Formal Verification: A Case Study of the Task-Scheduler Function select_idle_core. OLIVIERFEST '25.
[7] Daniel Gomez, "rxarray," commit 0ae6dd57e31d, Linux kernel development tree. url:
https://git.kernel.org/pub/scm/linux/kernel/git/da.gomez/linux.git/commit/?id=0ae6dd57e31d9e436ee935288767f53e118db8c0Speaker: Corinn Tiffany (Inria) -
10:45
Common pitfalls and errors, when using Linux in Safety Applications 45m
Safety-oriented use of the Linux Kernel presents a class of obstacles that are normally absent from other types of use, and therefore can be easily overlooked. Sometimes what would normally constitute a strength can actually turn into a weakness. Or aspects that are peculiar of creation of physical products (e.g. cars, robots) can present unexpected challenges.
This talk wants to raise awareness about these uncommon aspects.Speaker: Igor Stoppa (nvidia) -
11:30
Coffee Break 30m
-
12:00
Scaling the Design Definition 25m
The fundamental unit of design is a requirement (i.e. "testable expectation"). All forms of design can be expressed as a (directed acyclic) knowledge graph of requirements. A software project that documents its design in this manner can derive both source code and automated test from the same idea. Without this, there is no guarantee that test and implementation reliably reflects the same idea.
In virtually all open-source projects, documenting the design to this level of detail is rarely (if ever) considered. While the benefits are easy to understand, retrofitting a design on an existing project and then keeping it up to date is a barrier to entry that few are willing to consider.
In this troubleshooting session, Chuck will describe the problem in more detail, answer relevant background questions, and solicit ideas for potential approaches to retrofitting designs and keeping them updated.
Speaker: Chuck Wolber -
12:30
Defining Linux Kernel Requirements and Test Specifications out of the Kernel Tree 25m
Following the objections in having requirements and even code documentation traceable to Kernel testing as optional part of the Linux Kernel development process, the ELISA project decided to start defining requirements and test specifications in a separate repository.
This session will present the current status for the requirements and test specifications framework and traceability to upstream code and test case implementations.Speakers: Gabriele Paoloni (Red Hat), Kate Stewart (Linux Foundation) -
13:00
From Manual Argument to Machine-Readable Graph: Toward a Safety SBOM for Linux 25m
A safety case is the structured argument that a system is acceptably safe: claims, the evidence supporting them, and the assumptions under which the argument holds. Building and maintaining one is largely manual, and keeping it consistent as a design evolves is harder still.
SPDX 3.1's Functional Safety profile changes the starting point by giving the safety case a machine-readable form. It models requirements and their refinement, verification activities, pass/fail evaluations, evidence, and assumptions of use—the same structure safety practitioners already reason about, now expressed as a graph that tools can produce, exchange, and check.
This talk introduces the profile and its data model: what each element means, and how claims link to the evidence that supports them and the assumptions they depend on. We show how this representation can be used to generate a safety SBOM for Linux—a safety case generated and recomputed from project content rather than assembled by hand—building on the work of the ELISA Architecture working group and its current efforts on requirements. We argue this graph-based approach is the foundation for an automated handshake between upstream evidence and downstream safety arguments, enabling the kind of contract-style consumption of Linux components that safety-critical projects need.Speaker: Nicole Pappler -
13:25
Wrap up 5mSpeakers: Kate Stewart (Linux Foundation), Philipp Ahmann (Etas GmbH (BOSCH))
-
-
10:00
→
18:30
Toolchains Track "South Hall 1 B" (Prague Congress Centre)
"South Hall 1 B"
Prague Congress Centre
158-
10:00
Security Features status update 30m
Another year of work is behind us, with lots of progress across GCC, Clang, and Rust to provide the Linux kernel with a variety of security features. Let's review and discuss where we are with parity between toolchains, approaches to solving open problems, and exploring new features.
Parity reached since last year:
- Various little behavioral corner-case bug fixesIn progress:
- Overflow Behavior Types (needed in GCC)
- forward-edge CFI (GCC KCFI at v14)
- coverage-sanitizer stack-depth tracking (needed in GCC)Stalled / needs attention:
- -fbounds-safety language extension (slow in Clang)
- __strong typedef (needs design finalized)
- Link Time Optimization for GCC kernel support
- backward-edge CFI (x86 CET shadow stack, kernel mode)Speakers: Justin Stitt (Google), Kees Cook (Google) -
10:30
Compiler-Based Context Analysis for Compile-Time Lock Safety 30m
Compiler-Based Context Analysis is a language extension which enables statically checking that required contexts are active (or inactive) by acquiring and releasing user-definable “context locks”. An obvious application of this feature is lock-safety checking for the kernel's various synchronization primitives, verifying at compile time that locking rules are not violated. This session will begin with a brief overview of this new infrastructure that relies on Clang's Thread Safety Analysis (-Wthread-safety), but the focus will be on how developers can use it to annotate locking requirements to catch concurrency bugs before they ever make it into a kernel binary.
Currently, the analysis is opt-in by default and requires declaring which modules and subsystems should be analyzed, as enabling it tree-wide currently results in numerous false positive warnings. In the remainder of the talk, we will discuss strategies and best practices on broadening Context Analysis coverage across the entire kernel tree.
Reference: https://docs.kernel.org/dev-tools/context-analysis.html
Speaker: Marco Elver (Google) -
11:00
BPF support in the GNU Toolchain 30m
In this activity we will first provide a very brief update on the status of the port of GNU binutils and GCC to the BPF target, with emphasis on the level of support for extant BPF programs and the kernel BPF selftests. Then we will address a set of particular issues for which we need feedback and/or consensus from the BPF kernel and clang/LLVM hackers.
Speakers: Cupertino Miranda, David Faust, Jose E. Marchesi (GNU Project, Oracle Inc.), Vineet Gupta -
11:30
Coffee Break 30m
-
12:00
Are your types the same as mine? 30m
This session is aimed at discussing the benefits (and requirements) of having a unified type representation that can be used in tracing tools to facilitate type management, type compatibility checks, etc. With type info from the kernel, type info from userspace, and types that may be defined in tracing scripts, the hoops that one may need to jump through to make it all work as a single environment where cross-stack tracing can be done are a bit excessive.
Let's discuss ongoing efforts to work on this, and look at the needs, wants, etc to help drive this in the right direction where a single type management system can satisfy all requirements, and avoid duplication of effort, and painful conversions between systems.
Speaker: Kris Van Hees (Oracle USA) -
12:30
Towards real-time ABI compatibility assurance and smarter module versioning with libctf 30m
The CTFv4 extension of the CTF file format into a superset of BTF is nearly complete. Once it's working, what can we do with systemwide type information for all C programs and the kernel?
One possibility we explore in this talk is to use some simple linker extensions to provide real-time, linear-time ABI checking in ld.so to determine for every running program whether any of the C libraries it uses has changed incompatibly since the program was linked in a way that breaks the program, without linking ld.so with libctf or requiring it to do anything more difficult than a few equality comparisons and strcmps.
The end goal in userspace is to be able to see something like this
# emacs warning: libxml.so.2.13.9: incompatible symbol: linked against xmlDocPtr xmlReadMemory (const char *, int, const char *, const char *, int) but running against xmlDocPtr xmlReadMemory (const char *, int, const char *, const char *)(a contrived example: xmlReadMemory has not actually broken ABI. ld.so invokes a helper program that uses libctf to print that output, but does not itself need to use libctf.)
It also seems likely to me that the same infrastructure that implements the above can be reused to implement a better modversions without the manifold pain points of the current implementation.
This is only partly written so far: we explore the design of this scheme in part so that the audience can spot any serious errors before it causes widespread breakage, and to see if people can suggest improvements to the general design. (I'm certain that modversions will add exciting extra wrinkles to this.)
Speaker: Nick Alcock -
13:00
Adding C library wrappers for all Linux syscalls - The Return 30m
At Linux Plumbers Conference 2025 in Tokyo I gave part 1 of this talk, this year I want to revisit progress made against the use cases presented last year including progress made on raw futex.
In review we'll look at more applications and if we we are able to solve their most thorny calls to syscall with actual libc wrappers.
I'll include a more formal review of what is currently missing in glibc, collecting the list of syscalls for x86_64 and comparing to the kernel.
I will close again with a call to action for always having syscalls available as C library calls even if they could be used behind the back of the implementation.
I will propose a formal mentoring program. If you want to get a syscall wrapper into glibc I'll work with you to make it happen.
Speakers: Mr Carlos O'Donell (Red Hat), Florian Weimer (Red Hat) -
13:30
Lunch Break 1h 30m
-
15:00
Advancing PGO, AutoFDO, and Propeller in the Linux Kernel 30m
AutoFDO and Propeller are upstream and PGO is available out of tree, but none of them supports kernel modules, which run 4–5% of kernel cycles on Google's servers and carry performance-sensitive vendor drivers on Android. Part 1 presents per-module profiles for all three behind one Kbuild interface, and a path to upstream PGO by reusing compiler-rt's profile code. Part 2 shows why function-level coverage misleads once functions are split, and introduces line-level attribution and Cycles per Byte (CPB) to measure how well a kernel's hot code is laid out.
Speaker: Rong Xu (Google) -
15:30
Upgrading Runtime Leaks to Compile-Time Errors: Adopting Clang’s require_explicit_initialization 30m
The Linux kernel makes heavy use of aggregate struct initializers to configure subsystem interfaces, driver registries, and object lifecycles (such as struct kobj_type and struct device_type). However, omitting mandatory struct members—such as a kobject's release callback—may result in silent memory leaks or deferred runtime panics.
Clang offers a static alternative attribute ‘require_explicit_initialization’, which triggers the -Wuninitialized-explicit-init warning if a designated field is left uninitialized during aggregate construction . This talk is proposed as a collaborative discussion to brainstorm how we can adopt this safety mechanism in kernel headers (e.g., via a macro like __required_init) . We will discuss the candidate structures that would benefit most, evaluate the engineering friction of enforcing this across cross-tree refactors, and address compiler parity issues—specifically how to handle GCC fallback and whether a similar diagnostic can be proposed for GCC.
Speaker: Ian Rogers (Google) -
16:00
Kage: device driver isolation in the Linux kernel with LFI 30m
Drivers in Linux run with full kernel privileges and account for a significant share of exploitable vulnerabilities. Kage is an experimental kernel subsystem that runs kernel modules inside of in-kernel sandboxes using LLVM's Lightweight Fault Isolation (LFI) feature. Kaged drivers still run in the same hardware privilege level and address space as the rest of the kernel, but are compiled such that control-flow and memory accesses are confined to a subset of virtual memory. This keeps transition costs cheap and reduces the engineering effort needed to retrofit isolation onto existing drivers. LFI is designed to provide secure isolation even for arbitrarily malicious code through the use of a machine code verifier that keeps the compiler out of the trusted code base.
This talk will walk through a prototype version of Kage, which can load LFI-built kernel modules into an isolated portion of the address space, and demonstrate how drivers can be ported to run in sandboxes. Kage uses BTF type information to automatically check call signatures and replace kernel pointers with opaque handles on calls between the kernel and the module. The current prototype can sandbox simple computational kernel modules, and is expected to grow in capabilities in order to handle MMIO, buffer sharing, DMA (via the IOMMU), and interrupt routing, with the end goal of being able to isolate real, complex drivers, such as USB, network, or possibly even GPU drivers.
Speaker: Zachary Yedidia (Stanford University and Google) -
16:30
Coffee Break 30m
-
17:00
GCC Rust support for Linux 30m
A conversation about the Linux kernel requirements from GCC Rust and how to promote GCC Rust as another compiler for the Rust components of the Linux kernel. The technical and community issue.
Speakers: David Edelsohn (NVIDIA), Philip Herron (Embecosm) -
17:30
Crossing borders between user space and kernel space: GDB/Valgrind/drgn close cooperation 15m
While is GDB is primarily a user space tool, drgn is a kernel space one. Using them together may improve user experience as drgn can help GDB to get access to various kernel structures. Right now, GDB can show what a thread is doing in user space, drgn can show what it is doing in the kernel — but the two tools do not talk to each other. I would like to discuss what it would mean to cross that border.
The crossing is technically straightforward: drgn is a Python library, and GDB already has a Python API, making it natural to call drgn from GDB scripts. As a concrete demonstration, I wrote a small GDB Python command, drgn_why_sleeping, which retrieves the kernel stack of a sleeping thread via drgn and displays it alongside GDB's user-space backtrace — showing the full kernel path where bt alone shows only a generic syscall frame.
This raises further questions: could drgn's per-frame register and local variable access allow GDB users to inspect kernel frames interactively? Is a remote variant worth pursuing for kgdb setups? Could drgn help trace where in the kernel an error originates when stepping over a failing syscall? Valgrind could also benefit — since drgn does not use ptrace, it could attach to a Valgrind-instrumented process simultaneously and resolve invalid memory addresses to named kernel structures. And in the other direction: could user-space tools like GDB or Valgrind offer anything useful back to drgn?
Speaker: Alexandra Petlanova Hajkova -
17:45
Beyond Coverage: Per-Task Function Boundary Extraction for Kernel Contract Verification 15m
KCOV's trace-pc provides edge coverage feedback for kernel fuzzers, but captures no data-flow context. Two syscalls hitting identical basic blocks with different argument values (e.g., vfs_open with O_RDONLY vs O_WRONLY|O_TRUNC) are indistinguishable to the fuzzer. This "semantic gap" causes coverage saturation on value-dependent state transitions in complex subsystems (binder, io_uring, ksmbd).
ftrace function graph tracer support arg/ret values tracing during runtime. But their is a couple of challenge(FTRACE_REGS_MAX_ARGS 6) regarding architecture calling convention from register, and currently can't support structure and their members on the function.
On the other hands, Rust For Linux kernel modules can't supported by current tracing framework.
We need consensus on extending LLVM's SanitizerCoverage to capture function arguments and return values, including automatic struct field decomposition from DWARF metadata, with a kernel-resident per-task ring buffer.
What I Implemented
Two new SanitizerCoverage modes in LLVM:
Emits LLVM RFC submitted
-
__sanitizer_cov_trace_args(arg_idx, arg_size, val, offsets, num_fields)at function entry for each parameter. -
__sanitizer_cov_trace_ret(ret_size, val, offsets, num_fields)before each return.
The pass uses DICompositeType metadata to extract struct field offsets/sizes at compile time and emits them as global constant arrays. At runtime, the kernel backend reads fields via
copy_from_kernel_nofault()into a per-task lock-free mmap'd buffer at/sys/kernel/debug/kcov_dataflow.A native rustc path (rustc built against the custom LLVM) enables Rust kernel module instrumentation without a post-compilation pipeline. The only known method for capturing Rust function arguments at runtime given -O2 DWARF elision.
Discussion Topics Needing Agreement
-
Callback ABI stability
Should__sanitizer_cov_trace_args/retbe considered stable SanitizerCoverage API, or experimental? What's the path from-fsanitize-coverage=trace-argsto an accepted upstream flag? -
DWARF metadata dependency
The pass requires-gfor struct expansion. Currently it gracefully degrades: without-g,F.getSubprogram()returns null and the function is silently skipped. With-gbut missing type info for a param, it records as scalar (offsets=NULL, num_fields=0). Should this behavior be documented as the contract? -
Interaction with optimizations
At -O2, scalar arguments are spilled toallocaso the kernel callback receives a uniform pointer. Should the pass mark these as side-effecting to prevent DSE, or accept the current behavior (compiler preserves them because the call itself is a side effect)? -
Rust integration path
Currently requires building rustc against the custom LLVM (native path) or using a wrapper (rustc emit-llvm-ir then opt). Is there appetite for native-Zsanitizer-coverage=trace-argssupport in upstream rustc? The kernel's Kbuild does not expose intermediate IR for external instrumentation, making the native path mandatory for Rust kernel modules. -
Kernel-side callback design
The kernel backend usescopy_from_kernel_nofault()for safe pointer reads and a bit-31 recursion guard (needed becausecopy_from_kernel_nofaultitself is instrumented under INSTRUMENT_ALL). Should the LLVM pass remain fully target-agnostic, or emit hints for the runtime? -
Scope control
Per-module opt-in (KCOV_DATAFLOW_file.o := y) vs. global (CONFIG_KCOV_DATAFLOW_INSTRUMENT_ALL). The pass currently instruments all functions with DISubprogram. Should it support function-level filtering via attributes (e.g.,__attribute__((no_sanitize("dataflow"))))? -
Performance implications
With INSTRUMENT_ALL: +9.5% .text, +133% syscall latency. Per-module: +8.3% on instrumented paths. The per-callback cost is ~27ns dominated bycopy_from_kernel_nofault. Is this acceptable for the SanitizerCoverage framework, or should it be a separate pass?
Why Rust Kernel Modules Need This
Existing observability tools completely fail on Rust kernel code:
- ftrace: Cannot hook Rust functions (no
__fentry__prologue — rustc doesn't emit-mfentry) - kprobes/fprobe: Rust symbol mangling makes targeting impractical; even with correct names, only captures raw register values (no struct expansion)
- drgn/vmcore:
rustc -O2elidesDW_AT_locationfor all parameters —frame.locals()returns{} - eBPF: Requires manual struct layout specification; Rust
#[repr(Rust)]types have no stable ABI
Our LLVM pass operates on Rust-generated IR before codegen — the only point where both debug metadata and argument values coexist.
Verified: Rust eight_struct_args_rust selftest produces correct output at -O2 on CI result: https://github.com/yskzalloc/kcov-dataflow/actions/runs/29659128011/job/88119068007
do_el0_svc({0x4, 0x4, 0xffffffff, 0x0, 0x0, 0x0}) 0x0 = B... sb_start_write() file_start_write() 0x0 = _RNvCsdfZGIOKgjaD_22eight_struct_args_rust13write_handler() [eight_struct_args_rust] 0x11 = rsf_1(0x11) [eight_struct_args_rust] 0x33 = rsf_2(0x11, {0x11, 0x22}) [eight_struct_args_rust] ... 0x30c = rnsf_8({0xa8, 0x11, 0x11, 0x11, 0x11, 0x11, 0x11, 0x11, 0x11}) [eight_struct_args_rust] 0xaa = rsf_fwd_inner(0x11, {0x11, 0x22}, {0x11, 0x22, 0x33}, {0x11, 0x22, 0x33, 0x44}) [eight_struct_args_rust] 0xaa = rsf_fwd(0x11, {0x11, 0x22}, {0x11, 0x22, 0x33}, {0x11, 0x22, 0x33, 0x44}) [eight_struct_args_rust] rsf_ret_struct(0x0, {0x11, 0x11}, 0x11) [eight_struct_args_rust] 0xfffffdffc0356e80 = rust_helper_krealloc_node_align(0x0, 0x8, 0x8, 0xcc0, 0xffffffff) 0x6f00000009 = rust_helper_krealloc_node_align(0x0, 0x10, 0x8, 0xcc0, 0xffffffff) 0xf200006e00000011 = rust_helper_krealloc_node_align(0x0, 0x18, 0x8, 0xcc0, 0xffffffff) 0x2500006e00000011 = rust_helper_krealloc_node_align(0x0, 0x20, 0x8, 0xcc0, 0xffffffff) 0xaa = rsf_heap(0x11, {0x11, 0x22}, {0x11, 0x22, 0x33}, {0x11, 0x22, 0x33, 0x44}) [eight_struct_args_rust] 0x0 = rust_helper_krealloc_node_align(0x11, 0x0, 0x1, 0x0, 0xffffffff) 0x0 = rust_helper_krealloc_node_align(0x11, 0x0, 0x1, 0x0, 0xffffffff) 0x0 = rust_helper_krealloc_node_align(0x11, 0x0, 0x1, 0x0, 0xffffffff) 0x0 = rust_helper_krealloc_node_align(0x11, 0x0, 0x1, 0x0, 0xffffffff) 0x1 = _RNvCsdfZGIOKgjaD_22eight_struct_args_rust13write_handler(0xffffa64d1231, 0x1, 0x0) [eight_struct_args_rust] ...Implementation vs. ftrace funcgraph-args
ftrace's funcgraph-args (available since 6.x) captures 6 raw register values but:
- Requires-mfentry(no Rust support)
- No struct field expansion (just pointer addresses)
- Shared per-CPU ring buffer with preempt_disable per event
- Under load: events silently overwritten/dropped
- Stack-passed args (> 6) not capturedkcov-dataflow:
- Per-task private mmap buffer (no contention, no preempt_disable)
- Automatic struct field expansion from DWARF
- All args including stack-passed (alloca'd at compile time)
- Works on both C and Rust identically
- Zero-copy consumer via mmapCurrent Status
- Repository (kernel + LLVM + rustc + selftests + CI):
https://github.com/yskzalloc/kcov-dataflow - CI (x86_64 + arm64, 6 selftests):
https://github.com/yskzalloc/kcov-dataflow/actions - Tested on linux-next 7.2.0-rc3 with custom clang/LLVM 23 and rustc 1.99-nightly
- All selftests pass on both x86_64 (KVM) and arm64 (native) Github CI
Why This Needs Face-to-Face Discussion
- This work crosses 3 communities (LLVM, kernel, Rust-for-Linux) that rarely meet simultaneously.
- The callback ABI, optimization interaction, and Rust pipeline decisions require input from all 3.
- LPC Toolchains Track is the only venue where clang/LLVM developers, kernel maintainers, and Rust-for-Linux contributors are in the same room.
References
- LLVM RFC: https://discourse.llvm.org/t/rfc-sanitizercoverage-add-fsanitize-coverage-trace-args-trace-ret/91026
- LLVM PR: https://github.com/llvm/llvm-project/pull/201410
- Kernel patch v3: https://lore.kernel.org/all/20260902-b4-kcov-dataflow-rfc-v3-v3-0-824609f279f5@est.tech/
- arXiv paper: https://arxiv.org/pdf/2606.00455
- LWN: https://lwn.net/Articles/1092429/
Speaker: Mr Yunseong Kim (Ericsson Software Technology) -
-
10:00
-
10:00
→
18:30
eBPF Track "South Hall 1 A" (Prague Congress Centre)
"South Hall 1 A"
Prague Congress Centre
158The eBPF Track is going to bring together developers, maintainers, and other contributors from all around the globe to discuss improvements to the Linux kernel’s eBPF subsystem and its surrounding user space ecosystem such as libraries, loaders, compiler backends, related system tooling as well as eBPF use cases.
The gathering is designed to foster collaboration and face to face discussion of ongoing development topics as well as to encourage bringing new ideas into the development community for the advancement of the eBPF subsystem.
The track will be composed of talks, 30 minutes in length (including Q&A discussion).
eBPF Track's technical committee: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko
-
10:00
Rust-BPF updates 30m
Support for all of Rust required rethinking of how the verifier processes instructions. Stack liveness, SCEV, indirect calls are known building blocks while more fundamental rewrite is still necessary. The talk will cover completed and upcoming work areas in the verifier, LLVM, GCC, rustc, libbpf.
Speaker: Alexei Starovoitov -
10:30
Support large-size arguments and Rust based exception handling 30m
Large-size Arguments
Current bpf verifier will reject a function if one of its arguments
is a union/struct or more than one register size (except __int128 type).
Such limitation forces users to work around codes to fit bpf prog
requirement, but such limitation does not exists for other languages,
like normal C, rust, etc. Without such limitation, users will be able
to write more elegant codes.Rust Based Exception Handling
Current C based bpf prog do not have language level exception
handling. The bpf echosystem added bpf_throw() kfunc to allow
exception which does architecture unwinding and allows some
code at main prog level.Rust bpf supports more flexible exception handling at language
level. It allows bpf callbacks at individual function exception
level, e.g. to release certain references etc. At each function
level, it is possible there are multiple possible exception
handling, e.g., A -> B -> C. If C tiggered exception, bpf
callbacks can be called for C, then B, then A.In any case, Rust can provide lots of flexibility when something
wrong during verification.Speaker: Yonghong Song -
11:00
Extending the BPF Type Format to support more languages 30m
Now that kernel modules can be built in Rust [1],[2], it seems timely to look at programming language constructs outside of C to assess their applicability to BTF representation. Suggestions include
- slices/"fat pointers" with type and size
- tagged unions
- generics
- interfaces
- lifetime tags
Rough suggestions will be made for how to represent these as BTF kinds to trigger discussion.
[1] https://docs.kernel.org/rust/index.html
[2] https://lwn.net/Articles/991719/Speaker: Alan Maguire (Oracle) -
11:30
Coffee Break 30m
-
12:00
Slaying the BPF Verifier: Safe & Static Kernel Extensions Using Rust 30m
While the BPF verifier ensures kernel safety, its rigid static rules heavily bottleneck the development of complex kernel extensions. Last year’s LPC introduced Rex (Rust Extensions) to address these gaps. By leveraging Rust, Rex aims to guarantee memory and type safety at compile time rather than relying on the restrictive verifier. While this initial iteration utilized runtime protection to manage panics and address the halting problem, it resulted in considerable architectural and performance trade-offs.
This talk introduces the revamped, statically compiled panic-free Rex. By fully leaning into Rust’s static typing, borrow checker, and a dedicated compiler pass, this new architecture closely mirrors the verifier’s safety profile but at compile time. It empowers developers to build large-scale kernel programs using dynamic loops, complex data structures, and natural programming patterns.
To prove this model's viability, we will present a successful Proof of Concept integrating the new Rex with the ghOSt scheduler. We will walk through the implementation details, demonstrating how Rex handles complex scheduling logic safely and efficiently entirely in Rust.
Ultimately, we seek to engage the community in a dialogue regarding the robustness of this architectural framework and the strategic roadmap for upstream integration. We are eager to assess the broader interest in the evolved Rex and solicit feedback on potential refinements to further optimize our model.
Speaker: Aniket Gattani (Google Inc.) -
12:30
blitmus: Litmus-testing the eBPF memory model on real hardware 30m
The eBPF instruction set has quietly grown a concurrency surface: load_acquire/store_release instructions, arena atomics now JIT-compiled across x86, arm64, riscv and powerpc, resilient queued spinlocks, and lock-free ring buffers. eBPF programs routinely run concurrently across CPUs and increasingly coordinate through shared memory. Yet unlike the Linux kernel — which has a formal, herd7-executable memory model (LKMM) and the klitmus7 tool to exercise it — eBPF has no formal memory model and no way to check that the ordering a program requests actually survives the verifier and each architecture's JIT.
This talk presents blitmus (https://github.com/puranjaymohan/blitmus), a tool that converts standard LKMM litmus tests into eBPF programs and runs them on real hardware. Each process of a litmus test becomes a BPF program pinned to a specific CPU; a userspace harness races them across millions of iterations, histograms the outcomes, and compares them against the test's exists clause and expected result — the same methodology klitmus7 and herd7 use, but exercising the BPF runtime end to end. Because the same litmus test can be run unchanged on every architecture, blitmus can surface JIT ordering bugs (e.g. a missing barrier for cmpxchg) that are otherwise silent and catastrophic for concurrent BPF code.
I will demo the tool converting and running the LKMM litmus suite through BPF, walk through the results on aarch64 and x86-64 (including weak behaviors observed for store-buffering and a Result: Never test that blitmus flags as violated), and use that case to discuss tool fidelity: how faithfully a batched, test-run-based harness models true multi-CPU races, and where the current barrier/interleaving strategy needs to improve.
I would like to use the session to discuss next steps with the community:
(1) growing the harness to cover BPF-specific primitives — arena spinlocks, rqspinlock, ring buffer reserve/commit ordering, per-CPU maps;
(2) a proper litmus parser and cross-architecture CI (qemu on riscv/ppc/s390/loongarch) wired into selftests/bpf as a memory-model conformance suite;
Speaker: Puranjay Mohan (Meta) -
13:00
Adding KASAN support for JIT-compiled BPF Programs 30m
KASAN (Kernel Address Sanitizer) is a powerful developer tool for detecting
use-after-free and out-of-bounds memory accesses in the kernel code.
However, not all memory accesses performed by the kernel are covered by
KASAN monitoring. BPF programs are a major example: when they are
translated by the in-kernel JIT compiler, the kernel directly emits new
native instructions that then escape KASAN instrumentation. Closing this
gap would shed some light on potential bugs in the eBPF verifier or JIT
compilers that would be hard to investigate otherwise.This talk will present the current effort funded by the eBPF Foundation
aiming to introduce KASAN support for eBPF programs: we will discuss the
main architectural points, going over how the JIT compiler emits calls to
__asan_*()functions before each memory access, discussing the register
save/restore strategy, and the interactions with the BPF verifier.As a first draft of this work has been introduced at the LSFMMBPF
conference 2026 (Zagreb, Croatia), and as many revisions have been sent and
discussed on the BPF mailing list since then, this talk will also act as an
update, highlighting the main changes since the first draft, as well as the
remaining difficulties and issues to solve.Speaker: Alexis Lothoré (Bootlin) -
13:30
Lunch Break 1h 30m
-
15:00
BPF Kernel Verifier meets GCC: A saga of arranged marriage 30m
BPF Verifier has troubles grok'ing GCC generated code (having evolved with LLVM over time). This talk is about my recent work on BPF GCC and the kernel verifier and efforts to make them like each other more.
Speaker: Vineet Gupta -
15:30
Evolution of Value Tracking in the BPF Verifier 30m
While safety and performance get most of the credit for BPF's success, a safe and fast program that couldn't do anything interesting would be rather useless. Much of the flexibility that makes BPF programs interesting today (e.g. indexing into a map value at a computed offset, walking packet data, reading a stack array at a variable index) was actually not there at the start.
In this talk, we'll give a thousand-foot overview that goes from the initial tracking of constant/unknown all the way to cnum, covering topics like scalar IDs, sub-register bounds, bounds syncing, as well as value-tracking bugs along the way.
Speaker: Shung-Hsi Yu (SUSE) -
16:00
Proof-Carrying Verification for eBPF 30m
Abstract
eBPF provides a safe way to extend the kernel functionality. To ensure safety, the kernel verifies the memory safety and termination of eBPF programs and is thus the security boundary for eBPF. It must accept untrusted programs from userspace and determine, under strict time and memory limits, whether they are safe to execute.
Today, the verifier symbolically executes program paths, tracks abstract register and memory states, and uses state pruning and repeated loop exploration to establish safety. Keeping both proof discovery and proof checking in the kernel has the important advantage that the kernel does not need to trust the compiler or any userspace component.
The cost of this design is that the kernel has to employ a sophisticated analysis capable of reconstructing invariants from eBPF bytecode. Compilers may already discover source-level relationships and loop invariants, but these facts are not normally conveyed to the kernel. When the verifier cannot recover invariants, it must conservatively reject the program. At the same time, the complexity of the analysis makes it difficult to audit and maintain this security boundary, and verifier bugs can result in malicious programs being accepted. Verification failures can also be difficult to predict and relate back to the source program.
We are exploring a proof-carrying verification architecture that separates proof generation from proof checking. An untrusted userspace proof generator analyzes the program, discovers invariants, and uses an SMT solver to discharge safety obligations. If successful, it emits a certificate containing the invariants and proof steps needed to validate the program. If an obligation cannot be proved, the tool can use compiler metadata and solver output to distinguish concrete counterexamples from unsupported reasoning or missing invariants, and produce source-oriented diagnostics.
Inside the kernel, a domain-specific checker validates the program and certificate together. The certificate supplies facts such as basic-block and loop invariants, pointer bounds, and relationships between program values. The checker does not search for invariants or invoke an SMT solver. It validates each claimed state transition and entailment using a restricted set of eBPF-specific reasoning rules.
For example, a loop certificate may provide an invariant relating an induction variable, a packet pointer, and data_end. The checker verifies that the invariant holds on entry, is preserved across the back edge, and is strong enough to justify each memory access. It does not need to infer the invariant or repeatedly explore the loop.
We have built a preliminary end-to-end prototype that generates and checks certificates for a subset of eBPF programs, including examples with control flow, loops, scalar relationships, and memory-safety obligations. Our early experiments indicate that certificate checking can perform particularly well on larger programs for which abstract interpretation requires extensive path exploration, state merging, or repeated loop analysis. In these cases, the checker benefits from being given the required invariants directly and can avoid much of the search performed by the existing verifier. These results are still preliminary and currently apply only to the supported subset. At LPC, we will present the supported feature set, representative examples, and an initial comparison of verification behavior and runtime, together with the cases where the approach does not yet apply.
A central design problem is defining the safety policy shared by the proof generator and checker. This policy must cover instructions, helper calls, kfuncs, program contexts, pointer accesses, and kernel-managed resources, while evolving with eBPF and its kernel interfaces. Its format, maintenance model, and relationship to existing verifier semantics are part of the research question.
The talk will present the architecture, certificate language, prototype, and initial results. We will also discuss certificate size, malformed or adversarial certificates, unsupported features, policy evolution, verifier transformations, and the relationship between checked bytecode and JIT semantics.
We seek feedback from the eBPF community on whether this is a useful and technically plausible way to structure safety checking, which parts of current verifier semantics are hardest to express as explicit proof obligations, and whether a maintainable shared safety policy is possible.
Relation to Prior Work
Proof-carrying verification for eBPF was previously explored by Exoverifier, which separates proof generation from proof checking and emits general proof objects based on a Lean formalization of eBPF semantics. More recently, VEP proposed an annotation-guided toolchain with userspace verification, a specialized compiler, and a lightweight bytecode-level proof checker. BCF takes a hybrid approach: it retains the kernel verifier’s abstract interpretation, while using machine-checkable proofs generated in userspace to guide abstraction refinement and improve precision.
Our work explores another point in this design space. Rather than relying on general theorem-prover objects, source annotations, or refinement of the existing verifier analysis, we are designing a restricted certificate language around eBPF-specific invariants and state transitions. The checker accepts only a fixed set of reasoning steps, with explicit limits on certificate structure and checking work. Our preliminary prototype suggests that this model is sufficient for a useful subset of programs while remaining closely aligned with current eBPF concepts. The talk will compare these approaches and discuss which parts of eBPF verification fit restricted certificate checking and which may require richer reasoning.
Speaker: Martin Fink (Technical University of Munich) -
16:30
Coffee Break 30m
-
17:00
Safely Optimizing eBPF for Hardware and Workloads 30m
eBPF users pay for verified safety with performance. Across 27 microbenchmarks, we find that eBPF code runs 1.7–2.1x slower than the same code compiled natively. Most of the time is lost in the JIT, which translates one instruction at a time and leaves many hardware features unused. The JIT also does not adapt to the workload, such as stable map entries or branches. This is by design: the eBPF instruction set is designed for verification, not execution, and the kernel must trust the JIT and keeps it simple to keep the trusted computing base (TCB) small, so optimizing it is harder than optimizing a userspace JIT. We ask: how can we get the speed back without growing the TCB much?
We start with the hardware. We propose verified inline kfuncs (kinsn), which are kfuncs that BPF programs use like instructions, and can be verified and inlined by the JIT. Each carries a short eBPF body that the verifier checks at every call, while the JIT inlines native code in its place. We can also formally verify the eBPF and the native code is same. Seven such operations speed up microbenchmarks by up to 24% and Cilium and Katran by up to 12%.
kinsn fixes one operation at a time, but registers, branches and calls still slow down the whole program. What if the kernel just accepted native code? Then we propose kprog: it runs the whole program as native code, and the verifier checks an eBPF simulator of the native code, instead of JIT-compiling the eBPF into native code. You can then use any language or any compiler to compile your kernel extensions. A simulator is far easier and simpler to prove than an optimizing JIT, and applications run up to 1.83x faster.
Many workload patterns appear only at run time. To use them, we also propose Speculative ReJIT, which rewrites an unmodified application's bytecode in userspace and resubmits it to the existing verifier, so correctness stays in userspace and safety stays in the kernel.
We would like feedback on the kinsn interface, on when an operation should be applied at all, and on whether kprog and Speculative ReJIT are worth pursuing.
Speakers: YUSHENG ZHENG, Hao Sun (ETH Zurich) -
17:30
BPF, sysctls and programmable kernel decision points 30m
Traditionally the kernel community have introduced new sysctl tunables in places in the code where there is no "one-size-fits-all" answer. With BPF we have the opportunity to make such decision points programmable. A simple example of this is the kernel socket acceptq length, set via a combination of the listen() backlog and somaxconn sysctl. While the traditional advice has been to set this to a very high value to allow for connection bursts, doing so can lead to unacceptable connection latencies for applications, particularly when the application is doing CPU-intensive processing that limits connection acceptance rate. We only drop SYNs when the acceptq is full, but in such cases it would be beneficial to drop earlier to allow the service to recover. Having a BPF program attached could detect such conditions and handle this case by dropping SYNs. Having a dynamic response - potentially informed by wider system state - would be helpful here and in other cases. Wider discussion of decision points like these for BPF programmability is explored.
Speaker: Alan Maguire (Oracle) -
18:00
eBPF Security: The Uneven Landscape 30m
Although eBPF is widely used to extend the Linux kernel, running programs in kernel space introduces critical security risks. Existing technical work often overlooks eBPF's entire security lifecycle. This talk aims to address this gap by systematically analyzing eBPF vulnerabilities, mitigations, and architectural limits. By reviewing research papers and CVEs, we will map real-world exploits to specific components, revealing that current defenses target isolated attacks rather than systemic architectural risks. Finally, we will present key takeaways, outline open directions toward lifecycle-aware, composition-safe security models, and highlight understudied non-verifier components.
Speakers: Gürkan Gür (Zurich University of Applied Sciences ZHAW), Louie Wolf (Zurich University of Applied Sciences ZHAW)
-
10:00
-
15:00
→
18:30
Birds of a Feather (BoF): No A/V "Club B" (Prague Congress Centre)
"Club B"
Prague Congress Centre
53-
15:00
Debugging CPU isolation / nohz_full 45m
Installing a CPU isolated workload for extreme low-latency expectations can be challenging. Requirements and settings have been recently documented upstream but practice is another story. Let's discuss that around a live example.
Speaker: Frederic Weisbecker (Suse) -
15:45
Rust for Linux Office Hours 45m
A Rust-related BoF to work together on several topics, to answer questions or resolve pain points from attendees, to get to know people interested in Rust and Rust for Linux and generally to have some more time for discussion on top of the Rust MC.
The list of topics will be developed closer to the conference (and more may be added during the conference too), but some examples of potential topics would be:
- Review and discussion of particular patch series.
- Prototyping of a small project, e.g. implementing a kernel module.
- Providing assistance with
pin-initand Klint.
Please feel free to join!
Speakers: Alice Ryhl (Google), Mr Andreas Hindborg (Samsung), Benno Lossin, Boqun Feng, Danilo Krummrich, Dr Gary Guo, Miguel Ojeda -
16:30
Coffee Break 30m
-
17:45
Updates and Future Directions for Resctrl RDT/MPAM/PQOS/CBQRI 45m
Significant progress has been made on the resctrl subsystem recently, driven by architectural evolutions across multiple hardware vendors. Key developments include Intel's separate control and monitor domains, ARM's IOMMU, NVIDIA's CPU-less MPAM support and MBW MAX hard limits, AMD's monitor counter assignment, and initial RISC-V inclusion.
Following last year's highly productive BoF, which successfully guided upstream development over the past several months, this session will bring the broader resctrl ecosystem together again—including maintainers and developers from Intel, ARM, NVIDIA, AMD, RISC-V, Google, Fujitsu, and Alibaba. While we will briefly review recent milestones, the primary focus of this session will be face-to-face discussion to address current architectural bottlenecks, maintain a clean cross-architecture abstraction layer, and reach consensus on future upstream design paths.
Speakers: Fenghua Yu (NVIDIA), Ben Horgan, Reinette Chatre
-
15:00
-
15:00
→
18:30
Gaming on Linux MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128The Gaming on Linux Microconference welcomes the community to discuss a broad range of topics around improvements for Gaming devices running Linux. Gaming on Linux has pushed the kernel to improve in several areas and has helped create new features for Linux, such as the futex_waitv() syscall, the Unicode subsystem, HDR support, sched_ext, and much more. Although some of these were initially created for gaming use cases, they now have more generalized use.
The potential topics for this year are around a lot of subsystems in the kernel, including:
- Tracking workchain-granularity deadline detection in sched_ext for improved latency and frame rates
- The current state of upstreaming a GPU cgroup controller
- Using the cgroup v2 MM controller to manage GPU-mapped DRAM
- Autonomous kernel-level performance monitoring and improvement for gaming on Linux workloads
- Resource management on gaming devices using Orchestrator
- Improvements in debug data collection
- Current challenges in standardized benchmarking of gaming on Linux
- Distro support for Gaming on Linux
Last year was the first edition of the Gaming on Linux MC, and the focus was fairly broad across various parties and subsystems in the kernel. Since last year, Linux has continued to increase its footprint in the global market share of gaming platforms. Thanks to advancements in the kernel and the larger ecosystem in e.g. Valve Proton, the vast majority of games now run on Linux at equal parity to Windows, or better.
As this MC continues to mature and find its rhythm, we hope to emphasize this year's discussions on concrete problems in the kernel and surrounding ecosystems that need to be addressed to enable Linux to become the preeminent gaming platform.
-
15:00
Optimizing Gaming Performance and Power Efficiency with Intel LPMD 30m
Gaming performance remains critical for user experience, but power efficiency has become equally important, particularly for battery-powered devices. Also improves thermal management and reduce thermal throttling. LPMD offers a solution for enhancing energy efficiency by dynamically selecting optimal processor configurations and power slider settings based on CPU and GPU utilization.
This session demonstrates how Intel LPMD can be configured to improve energy efficiency during gaming workloads while maintaining acceptable performance levels. Will present some performance metrics showing improvements in power consumption and thermal behavior. The discussion will cover LPMD configuration, implementation considerations, and get some suggestions from gaming community.Speaker: Srinivas Pandruvada -
15:30
GPU Job Scheduling: Current Development 30m
The Linux kernel is responsible for managing the dependency graph of userspace's compute and graphics shaders. Moreover, it handles the GPU load balancing and tries to guarantee forward progress and deadlock resistance.
Historically, most graphics drivers have used drm_sched, a problematic legacy code base, for these tasks. More recently, Rust drivers are ramping up their own infrastructure to achieve the same, while trying to learn from the mistakes of the past and to leverage the features of the Rust programming language for robustness and reliability. Notably, engineers at Collabora and Red Hat collaboratively develop a successor in Rust for drm_sched, drm::JobQueue.
This talk briefly informs about the aforementioned developments and then opens room for discussions which may help the developers (who are – in part – concerned with the compute aspect of GPUs) understand the wider community's pains and requirements, such as from the graphics and video game parties. Topics that could be discussed are, for example:
- Future of GPU scheduling: more control is handed over to the GPU's firmware, making it difficult or impossible for the kernel to prioritize real time applications such as games.
- What are the major issues in the Linux kernel regarding stability, e.g., crashing video games?
- One reason why drm_sched is in bad shape is that super-priority was given to performance, for example by omitting locks. Fixing some of these missing locks has been rejected, due to reported massive performance regressions. Would video game developers argue that ~100% reliability is worth significant sacrifices with frames per second?
- Used hardware: Which GPUs and drivers are most important for gaming on Linux, and which would the community desire to be supported better the most?
Speakers: Daniel Almeida (Collabora), Philipp Stanner -
16:00
EPP-boost: Per-core util-based performance boosting in AMD p-state driver 30m
I've observed that on the Steam Deck, changes to the AMD p-state driver can have a significant impact on tail latencies and stale frame numbers. I've submitted an RFC upstream which proposes a per-CPU EPP boost heuristic: https://lore.kernel.org/all/20260728073150.54964-1-void@manifault.com/.
We should discuss gaps in the current cpufreq / AMD p-state driver, and how to best address frequency scaling for gaming workloads in general.
Speaker: David Vernet (Meta) -
16:30
Coffee Break 30m
-
17:00
Emulating x86 atomics on arm64 30m
Atomics can be challenging enough by themselves, and emulating them can be even worse. The lack of correctness will crash the application sooner or later, and the lack of performance will be very noticeable.
One key difference between x86 and arm64 is how they handle atomic operations on unaligned addresses. While on x86 such operations are supported transparently, on arm64 they raise a SIGBUS and the application is terminated. This creates a problem when emulating x86 games on arm64, since we can’t modify their source code to avoid unaligned atomic operations. So, we need a mechanism to handle such operations on top of arm64.
This talk will introduce the topic and describe an implementation design, benchmark results and how other OSes approached this issue.
LKML link: https://lore.kernel.org/lkml/20251117160841.334224-1-andrealmeid@igalia.com/
Speaker: André Almeida (Igalia)
-
15:00
→
18:30
Kernel Memory Management MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53Some people say that 2026 is the year of Linux Memory Management. Others weirdly disagree.
In any case, there is plenty to discuss, as MM is as busy as ever.
We are looking for topics that would be of interest to the kernel memory-management community.
In particular, we are also interested in topic suggestions from outside the core kernel community, including userspace, drivers, architectures, and other areas that affect memory management in the kernel.
Example topics that might be worth discussing this year include:
- Making (m)THP/large folios first-class citizens
- Supporting gigabyte THPs: allocators, compaction, policies
- Better policies: applying eBPF and friends sensibly in MM
- Polishing memory reclaim: making MGLRU less special
- Ongoing challenges with memdescs conversion
- Can we make device memory less special?
- Letting the kernel manage special-purpose memory
- Improving page promotion/demotion for memory tiering
- Challenges with hypervisor live-update integration
- Towards deprecating hugetlb: mshare, memory reserves
- The future of swap: missing pieces, cleanups, and do we still need zram?
- The future of memcg: new resources, optimizations, and cleanups
- Doing more with less memory (RAM is getting expensive ...)
A microconference talk should provide enough context to enable an open discussion about the topic being presented. In particular, presentation-focused talks with little room for discussion are not what we are looking for.
-
15:00
Reducing kernel stack memory (on arm64) 30m
There have been a bunch of attempts [1] to reduce the memory consumed by kernel stacks for x86 and arm64, largely based around the idea of dynamically growing the kernel stack allocation based on page faults. This poses what appear to be insurmountable challenges, as it introduces complexity into the architecture exception entry code (which needs to be able to transition cleanly to a new stack) but also means that the kernel must be able to allocate memory from any context at all.
I would like to discuss and explore alternatives to dynamic stack allocation based on some initial rework I have started of the arm64 exception entry code[2]. Specifically:
- Revisiting the decision to move from 8k to 16k stacks
- Changing the kernel stack size at boot time
- Changing the kernel stack size per task
- Controlling the kernel stack size from userspace
I think it would also be useful to talk briefly about the sticking points with dynamic kernel stacks and whether there is anything that can be carried forward from that work.
My motivation is based on Android vendors having various out-of-tree implementations of the dynamic kernel stack series and I would like to see if there are alternatives that we can provide upstream so that we can all stop carrying the broken stuff around in our pockets.
[1] https://lore.kernel.org/all/20260424191456.2679717-1-stevensd@google.com/
[2] https://git.kernel.org/pub/scm/linux/kernel/git/will/linux.git/log/?h=overflow-stackSpeaker: Will Deacon -
15:30
Better Anon (Swap) Readahead 20m
When handling swap page faults, the kernel often don't reads just a single page. For legacy HDDs, an entire cluster is read in. For SSDs, we check the VMA and page tables to see if nearby pages are likely to be needed soon, using a basic statistic-driven heuristic. For compressed RAM (like zram, and potentially soon zswap), traditional readahead is skipped entirely. Meanwhile, enabling (m)THP swap-in currently acts as kind of a speculative fetching.
The current implementation is not really ideal. For example, why can't we extend (m)THP support to SSDs by merging readahead detection with (m)THP lookaround? Why can't we enable a lightweight readahead for compressed RAM as well? We also need a smarter heuristic, where workingset shadow entries could provide a much better hint than basic per-VMA statistics?
With so much active development in the swap subsystem, the goal of this session is to consolidate these scattered paths, and explore whether the swap readahead mechanism can share common, useful infrastructure.
Speaker: Kairui Song (Tencent) -
15:50
DMA-buf cgroup accounting for proxy allocators: attributing buffers to the right consumer 20m
On embedded and automotive Linux systems, a single daemon often allocates DMA-buf memory on behalf of other clients. Today, most of that memory is invisible to cgroup accounting, only system-heap buffers can carry a charge via __GFP_ACCOUNT, and even then it lands on the allocator’s cgroup rather than the app that requested the buffer.
The misattribution has real consequences on any system that uses cgroup memory limits to drive reclaim or enforce resource budgets, as untracked buffers silently skew pressure signals. For example, on Android, where app memory limits depend on per-app accounting; or on automotive systems, where strict Freedom From Interference (FFI) is a key requirement and a central allocator absorbing the charge of its clients across domains breaks any meaningful resource isolation argument.
In this session we will discuss the design space for fixing the proxy-allocator problem, presenting the context from our proposal in upstream. Should the charge be attributed at allocation time, or before export time? Should the allocator identify the target via a pidfd (natural for approaches like Binder, where the allocator already knows the sender_pid) or a cgroupfd? What happens with buffers that may belong to different cgroups across their lifetime?
Speakers: Albert Esteve (Red Hat), T.J. Mercier (Google) -
16:10
Page fault locking 20m
In 2023 we introduced per-VMA locks to solve contention and priority inversion on the mmap_lock for multithreaded programs. While successful, the initial implementation did not solve every case of contention and some workloads can still demonstrate problems. This session will discuss ways of fixing the remaining problems.
Speaker: Matthew Wilcox (Oracle) -
16:30
Coffee Break 30m
-
17:00
Direct map fragmentation 30m
On supported configurations, the direct map is built using large
PMD/PUD-level block mappings. This is expected to bring many of the
benefits of hugepages: improved TLB hit rate, less page table walking,
smaller page table memory footprint.This optimisation is most effective when the direct map gives access to
the entire physical memory with uniform permissions. Unfortunately, a
growing number of features need to remove pages from the direct map or
modify permissions/attributes at page granularity, causing PMD/PUD
blocks to be split down to PTEs.According to results presented by Mike Rapoport at LSF/MM 2023 [1], the
performance impact of such fragmentation is negligible for data
accesses. However, recent measurements on x86 and arm64 suggest a more
nuanced picture, and other factors such as power consumption should also
be considered.This session will briefly summarise key results, and then discuss
possible ways to reduce fragmentation.Questions for discussion:
- Should direct map permissions be modelled as an allocation property
in the buddy allocator, or managed by higher-level interfaces such as
execmem or secretmem? - Can x86 and arm64 use common logic for splitting and coalescing pages?
A few interesting sources of direct map fragmentation:
- execmem (may use PMD-sized pools; direct map PMDs will still be split
if writing code) - secretmem
- Confidential VMs including guest_memfd (encrypted/unmapped pages)
- pkeys-based page table protection
Proposals to reduce fragmentation:
- 2021/01 - PMD-sized pools for secretmem [2]
- 2021/04 - Grouped vmalloc [3]
- 2023/03 -
__GFP_UNMAPPEDbased on dedicated cache [4] - 2026/02 - pkeys-protected page table pools [5]
- 2026/03 -
__GFP_UNMAPPEDbased on migratype/pageblock [6] - 2026/06 - Collapsing blocks in secretmem [7]
- 2026/06 - EXECMEM_ROX_CACHE for arm64 with direct map PMD coalescing [8]
[1] https://lwn.net/Articles/931406/
[2] https://lore.kernel.org/lkml/20210121122723.3446-8-rppt@kernel.org/
[3] https://lore.kernel.org/lkml/20210405203711.1095940-1-rick.p.edgecombe@intel.com/
[4] https://lore.kernel.org/all/20230308094106.227365-1-rppt@kernel.org/
[5] https://lore.kernel.org/linux-hardening/20260227175518.3728055-1-kevin.brodsky@arm.com/
[6] https://lore.kernel.org/lkml/20260320-page_alloc-unmapped-v2-0-28bf1bd54f41@google.com/
[7] https://lore.kernel.org/all/20260603104624.36390-1-lance.yang@linux.dev/
[8] https://lore.kernel.org/all/20260611130144.1385343-1-abarnas@google.com/Speaker: Kevin Brodsky (Arm) - Should direct map permissions be modelled as an allocation property
-
17:30
Accelerating Page Migration and Making the Migration Core Composable 30m
As the memory hierarchy deepens at both ends, with HBM adding a faster tier on top and CXL adding cheaper, slower capacity below, keeping hot data in the fast tier and shedding cold data downward makes page migration central to NUMA, tiered-memory and coherent CPU-GPU systems (where device memory is exposed as NUMA nodes).
Profiling move_pages(2) on EPYC Zen 6 shows that the folio copy dominates migration (~96% of time for a 2MB THP), making it the primary scaling bottleneck. Yet this path is largely sequential: folios are copied one at a time by a single CPU, while DMA engines, idle cores and memory bandwidth go unused.
To tap that idle hardware, this work separates the folio content copy from the rest of migration and hands it to a pluggable migrator. The batch path:
- unmaps a batch of folios and flushes the TLB once,
- asks a migrator to copy the eligible folios,
- marks copied folios so the move phase skips the per-folio copy,
- completes the move through the existing flow [1].A faster copy is not the whole story: lower rmap overheads, a simpler migration architecture and a composable core matter just as much. The work therefore proceeds in three directions: accelerating the dominant copy phase, batching the rmap walks and restructuring the migration core so that these and other independent optimizations compose instead of becoming special cases.
Accelerating the copy
Three complementary, measured directions:
1. dcbm - a DMA-based migrator using dmaengine devices, for bulk copy across multiple channels.
2. mtcopy - multi-threaded CPU copy on idle cores, a software fallback where no DMA/offload engine is available [9].
3. Bulk folio_copy() for large folios: one copy over the contiguous range, independent of the batch and offload framework [2].Measured with move_pages() on 1GB anonymous memory, DRAM -> DRAM, batch copy offload makes 2MB THP migration several times faster: ~6x with PTDMA (Zen 3, 16 channels), ~3.5X with SDXI (Zen 6, 1 channel) and ~3.8x with multi-threaded CPU copy (Zen 3, mtcopy, 8 threads).
Separately, a related demotion effort uses non-temporal stores to reduce cache pollution and CXL-side read traffic [3].
Beyond the copy: the rmap phase
Once the copy is offloaded, rmap work dominates for PTE-mapped large folios (mTHP) because the unmap and restore walks handle each PTE separately. Batching those walks is the next lever, and its payoff grows with folio size: on Zen 3, a 1MB folio reaches ~5.6x over vanilla with DMA offload and rmap batching, compared to ~2x with offload alone [7].
Reworking the core so these compose
Several efforts optimize different migration phases, and each one adds flags or special cases to the current monolithic path [1][3][4]. They also run into the same limitation: ->migrate_folio() receives only enum migrate_mode, which describes the blocking semantics but not the migration intent (for example, promotion versus demotion). So, new behavior ends up either extending migrate_mode or going through a side channel. For example, non-temporal demotion added MIGRATE_ASYNC_NON_TEMPORAL_STORES, while the DMA offload records FOLIO_CONTENT_COPIED in migrate_info.
Two changes are posted as RFC[8]:
-
Pass a small context struct (migrate_control) through migrate_pages() into ->migrate_folio(), carrying mode, reason, and room for future attributes. This separates blocking behavior from intent, so the copy phase knows why the folio is migrating and copy policy (cached, non-temporal, offload, and so on) can follow.
-
Split the single list in migrate_pages() into separate classes, each with its own list and retries: hugetlb folios, movable_ops pages and LRU folios. This removes the type checks scattered through the shared batch path and makes the flow easier to follow.
Discussion
I'd like to use the session to validate the direction of the migration core work and how ongoing efforts should fit together.
Key questions:
- Interface: should migrate_pages() and ->migrate_folio() take a context struct? It touches every caller and callback. Is that one-time churn acceptable, and is this the right shape?
- Policy: Should each policy, such as cache intent, be stated explicitly by the caller, or should the core derive it from the migration reason?
- Core structure: is the per-class migration model right shape for maintainability? movable_ops pages still use page->lru, new_folio_t, put_folio_t. What should a non-folio migration interface look like?
- Synchronous fallback: after the async pass, should the remaining folios be retried through a direct single-folio path, or should they keep reusing migrate_pages_batch() and avoid a second single-folio path?
- Copy helpers: there is no range version of copy_page() for bulk folio copy. memcpy() works, but without FSRM on x86 it falls back to an unrolled movq loop instead of copy_page()'s rep movsq. Should we add copy_pages(), like clear_pages()?
- Engine selection: who should select the copy engine: the caller, the provider (as in the current RFC), or the core, based on caller constraints and provider capability? Should folio size, batch size or topology drive that choice?
- Batching trade-offs: What is an acceptable trade-off between throughput and first-folio latency? Should the batch limit be a tunable per platform?
Roadmap:
Migration core redesign and context passing; migration-side rmap batching, bulk copy within large folios, batch copy and DMA offload, Scatter-gather copy (DMA_MEMCPY_SG); SDXI migrator [6] and DCBM driver; topology-aware thread placement for multi-threaded copy and CPU accounting for multi-threaded copy; hot page promotion (pghot [5]) reusing the offload path.
References
[1] Migration batch/offload, RFC V6: https://lore.kernel.org/all/20260630-shivank-batch-migrate-offload-v6-0-da95d7e8b8a2@amd.com
[2] Bulk folio_copy optimization, RFC v1: https://lore.kernel.org/all/20260427142036.111940-4-shivankg@amd.com
[3] Non-temporal demotion, RFC V2: https://lore.kernel.org/linux-mm/20260730-rfc-nt-demote-v2-0-452dbe3b5073@zptcorp.com/#t
[4] migrate_mode/migrate_folio() ABI extensibility (Memory Hotness and Promotion call):
https://lore.kernel.org/linux-mm/37010218-ecad-4f75-961d-686d51b6677d@amd.com
[5] pghot, V8: https://lore.kernel.org/all/20260728054356.291998-1-bharata@amd.com
[6] SDXI, V3: https://lore.kernel.org/all/20260605-sdxi-base-v3-0-4d38ca2bdffe@amd.com
[7] Migration-side rmap batching, RFC V2: https://lore.kernel.org/linux-mm/20260813-migrate-rmap-batch-v2-0-3c5424c555c7@amd.com
[8] Migration redesign and policy, RFC V1: https://lore.kernel.org/linux-mm/20260902-migrate-refactor-shivank-v1-0-9dcca87669c4@amd.com
[9] mtcopy RFC v3: https://lore.kernel.org/all/20250923174752.35701-7-shivankg@amd.comSpeakers: Shivank Garg (AMD), Zi Yan (NVIDIA) -
-
18:00
Tackling Isolation Issues Due to Memory Pressure 30m
On a densely shared host, a routine event breaks tenant isolation: a task holding a shared kernel lock (filesystem metadata, a control-plane mutex/rwsem) allocates inside the critical section, the allocation falls into memory reclaim, and until reclaim finishes every other task waiting on that lock — from any cgroup — pays the stall. Shared kernel locks aren't mediated by cgroups or memcg limits, so this is cross-tenant priority inversion, and overcommit makes it worse.
In this talk we will present the data we have gathered across the Meta fleet on this problem, and share how we think the solution should take shape. Rather than a finished design or patch series, the goal is to lay out the problem space, the constraints any fix must satisfy, our current direction, and the open questions and to get early feedback from the upstream community.
Speaker: Shakeel Butt
-
15:00
→
18:30
System Boot and Security MC "Small Theatre" (Prague Congress Centre)
"Small Theatre"
Prague Congress Centre
105The System Boot and Security Microconference remains a key venue for engineers and researchers focused on the intersection of firmware, bootloaders, and kernel security. For 2026, we continue our mission to address the persistent friction encountered when upstreaming security-focused boot improvements. Our goal is to bridge the gap between low-level hardware initialization and the Linux kernel requirements.
Experience shows that integrating complex security technologies - like TrenchBoot or advanced attestation frameworks - often stalls due to architectural disagreements or a lack of cross-project coordination. We want to bring together developers from different layers of the stack (firmware, bootloaders, and kernel) to identify these bottlenecks.
We are interested in deep technical deep-dives, but also in the legal, licensing, and organizational aspects that often dictate how-and if-security standards can be integrated into the open-source ecosystem.We invite proposals covering technical updates, work-in-progress, and strategic discussions on:
- TrenchBoot, tboot,
- TPMs, HSMs, secure elements,
- Roots of Trust: SRTM and DRTM,
- Intel TXT, SGX, TDX,
- AMD SKINIT, SEV,
- ARM DRTM,
- Growing Attestation ecosystem,
- IMA,
- TianoCore EDK II (UEFI), SeaBIOS, coreboot, U-Boot, LinuxBoot, hostboot,
- Measured Boot, Verified Boot, UEFI Secure Boot, UEFI Secure Boot Advanced Targeting (SBAT),
- shim,
- boot loaders: GRUB, systemd-boot/sd-boot, network boot, PXE, iPXE,
- UKI,
- u-root,
- OpenBMC, u-bmc,
- legal, organizational, and other similar issues relevant to people interested in system boot and security.
Progress in various areas after LPC 2025:
CVE (CVE-2026-33697) of score 7.8/10 discovered and responsibly disclosed
Security advisoryPresentations, Slides, Videos etc. after LPC'25
-
Who Authenticates Linux? Rethinking PAM & NSS in the Age of Cloud Identity
-
TrenchBoot Linux Kernel Developments
- Android Boot, DRTM, UKIs As an alternative solution to DRTM for ARM Mobile systems, Android's Virtualization team has delivered an EL2 stub
-
15:00
Linux Crypto API is dead, long live Linux Crypto API v2? 20m
Linux Crypto API has recently come under fire due to a series of discovered vulnerabilities, like Copy Fail and friends. This resulted in a drastic response from some maintainers: deprecation of the Linux Crypto API userspace interface. However, it feels like cracking a nut with a sledgehammer, as some vulnerabilities, like Dirty Frag, do not even use the userspace interface directly for the exploit.
Doing crypto in the kernel on behalf of userspace code may bring some security and efficiency advantages:
- the key material is not exposed to the process address space anymore, which can save the key from very common memory bugs (remember Heartbleed?)
- abstracting the key material as a kernel object (e.g., a file descriptor) allows sharing the key usage across unprivileged applications without giving those applications access to the key material (software HSM/KMS use case, cloud workloads, K8s)
- the kernel is best positioned to abstract modern cryptographic hardware and allow applications to utilize it and reap the performance benefits
- highly-constrained embedded systems may have fewer dependencies as they don't have to pull extra crypto libraries such as OpenSSL or gcrypt into their firmware
One of the reasons quoted for the deprecation is that the userspace Linux Crypto API interface is not efficient anyway and there are very few users. But maybe there are few users, because it is not very efficient and apparently insecure? What if instead of disabling it we improve it or even redesign? For example:
- replace the splice()-based zero-copy interface (which was targeted by the vulnerabilities in the first place and is not truly zero-copy anyway) with a modern primitive such as io_uring
- fix the limitation of authenticated encryption being able to process only ~64k of data
- replace the confusing network-based AF_ALG protocol with something simpler, such as plain device files/sysfs
- use eBPF to expose Linux Crypto API to userspace in a more constrained and controllable manner
- limit the number of algorithms exposable to userspace – only general-purpose algorithms should be exposed
The goal of this session is to reach a shared view on whether the userspace Linux Crypto API is worth saving and, if so, what a v2 should look like.
Speaker: Ignat Korchagin (Citadel) -
15:25
Early Attestation is Very Harmful. Why is it Required? 20m
Abstract:
This talk is a follow-up of LPC'25, where we introduced the security problem to the community for a suitable approach of attested TLS protocols. We would like to present updates to the work since LPC'25 and discuss open problems with the community so that the feedback can be accommodated in the standardization in the IETF. In particular, we share our discovered critical-severity CVEs, such as CVE-2026-92701, CVE-2026-92702, CVE-2026-100835, each of CVSS 9.1, and other critical- and high-severity security concerns against early attestation, and ask the community whether early attestation is really necessary.Discussion:
We would like to discuss:Background:
Transport Layer Security (TLS) is a widely used protocol for secure channel establishment. However, it lacks an inherent mechanism for validating the security state of the workload and its platform. To address this, remote attestation can be integrated into TLS, which is named attested TLS protocol.As presented in LPC'25, there are two key approaches for this integration, namely early (intra-handshake) and standard (post-handshake) attestation. We also present insights from formal verification using the state-of-the-art symbolic security analysis tool ProVerif that led to the discovery of the critical- and high-severity CVEs.
Current project partners include CoreWeave, University of Namur, Stevens Institute of Technology, Shanghai Guan An Information Technology, Switch, Oxford University, Sureshot Labs, PHI-OMEGA, HashCloak Inc, Stoffel Labs Inc, EMILIA Protocol, ESILV, Birmingham City University, Verily, Arm, Linaro, Siemens, Huawei, Intuit, Axis, Bonn-Rhein-Sieg University of Applied Sciences, and Barkhausen Institut. By this talk, we hope to inspire more open-source contributors to this project.
The attendees will gain technical insights into the latest developments of standardization of attested TLS protocols in the IETF, and will be able to provide feedback on the standardization effort of attested TLS.
Benefits to the ecosystem
Our thorough formal security analysis shows that early attestation is potentially vulnerable to replay, relay, diversion, and misbinding attacks. Our analysis shows that:- Early attestation adds unnecessary complexity in the handshake.
- Early attestation results in several security and privacy problems, such as those captured in published CVEs.
- All use cases of early attestation can be handled with standard attestation while offering better security properties.
Because of critical-severity CVEs and fundamental architecture problems in early attestation, we believe early attestation is very harmful to the TEE attestation ecosystem. We would like to discuss with the community whether it is necessary.
Speaker: Muhammad Usama Sardar (TU Dresden) -
15:45
Extending BPF Signing into an IMA Integrity Lifecycle 20m
eBPF programs enter the kernel as privileged extensions and, once admitted, become part of the system’s trusted computing base. Prior LPC discussions established BPF signing as a way to track the provenance of these programs, while deliberately separating signature verification from the policy that consumes it [1,2]. IMA was identified as one possible consumer through LSM integration, but its enforcement semantics were left open. Follow-up discussion highlighted unresolved issues involving dynamically generated programs, trusted loaders, and applications such as Cilium and bpftrace.
We developed IMA-BPF as a concrete continuation of this work. IMA-BPF connects BPF signatures to IMA policy, measurement, appraisal, and TPM-backed attestation. It preserves signer provenance for programs created by signed loaders, allowing resident programs to be re-evaluated after the original loader disappears. We also introduce reappraisal and purge mechanisms, allowing programs that no longer satisfy current policy or trust state to be detached and removed through BPF’s link and reference based lifetime mechanisms.
This talk will explain how IMA-BPF extends BPF signing into an integrity lifecycle, evaluated through Cilium vulnerability remediation. We hope to discuss three areas: the practicality of adopting signing across real BPF workloads; how the BPF ecosystem should standardize provenance for dynamically generated programs, such as providing common reference tooling in different languages; and how our initial reappraisal and revocation policies and mechanisms should be refined. Although our prototype demonstrates this lifecycle end to end, we hope to use the discussion to improve the design and identify a practical path toward upstream adoption.
References
- KP Singh, BPF Signing and IMA Integration, https://lpc.events/event/16/contributions/1357/
- KP Singh, Discussion: What’s Next for BPF Signing, https://lpc.events/event/19/contributions/2167/
Speaker: Hongming Qiu (University of Illinois Urbana-Champaign) -
16:10
Attestation-Based Storage Provisioning for Linux: Towards a Standard Model 20m
Linux has no standard model for provisioning and unlocking encrypted persistent storage at boot in environments where the underlying platform cannot be fully trusted. Today, every deployment stack solves this independently: custom boot agents, post-provisioning scripts, and platform-specific unlock logic. The mechanism exists — LUKS2 tokens and the cryptsetup plugin interface — but the workflow around it does not. Every team that wants attestation-backed storage unlocking ends up writing the same three things from scratch: a script to register token metadata at provisioning time, a plugin to perform attestation and key retrieval at boot, and custom logic to handle failure cases. None of this is reusable across platforms or deployable through standard Linux tooling.
This session proposes that the LUKS2 token mechanism and the cryptsetup plugin interface should become the standard interoperability contract between Linux storage tooling and external attestation-aware key services. In this model, a Key Broker Service holds encryption keys and releases them only after the system proves its identity through remote attestation. The LUKS2 header carries the token metadata that tells systemd-cryptsetup which plugin to invoke at boot, making the storage layer independent of any specific attestation technology, hardware platform, or key service implementation. The provisioning side is handled by systemd-repart, which registers the token metadata in the LUKS2 header at first-boot — no separate import step, no custom scripts.
For this model to work reliably end-to-end, three specific gaps need to close. First, the cryptsetup plugin error contract is undefined: today systemd-cryptsetup silently falls through to passphrase unlock when a plugin fails, with no distinction between permanent attestation failure and transient network unavailability. A KBS plugin that fails because attestation was rejected should behave very differently from one that timed out waiting for the network — but there is currently no standard way to express that. Second, the LUKS2 token JSON schema beyond type and keyslots is unspecified: KBS-specific metadata such as service endpoint URLs is stored in private fields with no agreed naming, making provisioning tools and unlock plugins non-interoperable without bilateral coordination. Third, there are no packaging or initrd inclusion guidelines for token plugins, leaving distributions without a clear answer on how to ship and discover them at boot.
The session will walk through the current state of the code in systemd-cryptsetup and systemd-repart, identify precisely where these gaps exist today, and open a discussion on what standardisation is needed. The goal is agreement on reusable, upstream building blocks for attestation-driven storage provisioning — so that every platform and deployment stack building on Linux does not have to solve this problem independently.
Speaker: Nandakumar Raghavan -
16:30
Coffee Break 30m
-
17:00
Secure Boot: Getting your SHIM signed - The process and Its challenges 20m
This talk walks through the end-to-end process for Linux distributions to have their SHIM submitted for review and eventual signing by Microsoft, surfacing the challenges faced by both distribution teams and the SHIM review team.
We examine this process in depth from dual vantage points: distribution maintainers navigating submission requirements -- including SBAT entries, CVE patching, DBX and revocation policies -- and the SHIM review team evaluating submissions.
We also address the Partner centre for Windows Hardware and what Linux distribution teams need to eventually get the SHIM signed by Microsoft.
Speaker: Sherif Nagy -
17:25
Generic Boot Flow for Type 1 Hypervisors: UKI with systemd-stub 20m
ARM SystemReady defines a standard firmware-to-OS boot path via UEFI, but provides no standard mechanism for booting Type 1 hypervisors such as Gunyah or Xen. Platforms relying on Type 1 hypervisors to isolate security- or latency-sensitive workloads, lack a standardized boot flow, which creates fragmentation across mobile, IoT, and compute ecosystems.
We have implemented a solution for this across three areas. First, we propose extensions to the UKI format to bundle a Type 1 hypervisor and its static guest VM images in a single UEFI secure-boot authenticated image. Second, we discuss the required changes to systemd-stub to parse the extended UKI to extract and pass configuration data to the hypervisor at boot. Third, we have prototyped support to the systemd-boot tooling to control and maintain this configuration.
By aligning with both the Xen and Gunyah communities, the goal is to enable a vendor-neutral, standards-compliant boot flow for any Type 1 hypervisor, using UKI.
Topics to be discussed:
1. UKI extensions for Type 1 hypervisors
2. Interface between systemd-stub and hypervisor
3. Configuration data passed to the hypervisorSpeaker: Anushka Nabar (Qualcomm) -
17:45
Booting Android with Linux EFI Boot Stub 20m
Android is evaluating the Linux EFI boot stub to standardize OS handoff and reduce bootloader feature duplication.
We present blockers encountered with embedded hardware constraints, and aim to discuss potential solutions.Topics for Discussion:
-
EL2 Elevation and KVM within an EFI context
- The Problem: KVM requires booting the kernel at EL2. Since Android firmware operates at EL1, it must elevate to EL2 before kernel execution.
- Discussion: What are the options to elevate to EL2 before jumping to the kernel?
- Override UEFI
ExitBootServicesto hijack the boot flow? - Install a custom UEFI protocol callback that executes EL switch and takes over kernel jump?
- Override UEFI
-
Physical Image Placement and PE/COFF Relocation
- The Problem: The UEFI PE/COFF loader copies and relocates the kernel image to a dynamically allocated buffer address.
- The Constraint:
- Boot Latency: Copying large kernel payloads increases boot latency.
- Defeat Architectural Expectations: Hardware platforms might require the kernel image loaded at specific address ranges. The PE loader moving the kernel image around defeats this.
- Discussion: How do we eliminate unnecessary allocations and respect physical image placement done by the firmware/bootloader.
Contact us: android-gbl@google.com
Speaker: Yi-Yo Chiang (Google) -
-
18:10
Arm DRTM beyond the reference model: dynamic launch on real platforms and Linux 20m
Arm's DRTM architecture (DEN 0113) defines a dynamic root of trust for measurement: a DCE-Preamble triggers a dynamic launch event, EL3 firmware (TF-A) measures a protected payload into a fresh chain of trust, and the dynamically launched measured environment (DLME) continues under a smaller TCB. Yet a working implementation exists only against the Base AEM FVP TF-A's plat_drtm hooks (SMMU-based DMA protection, the DRTM address map, and measurement) live solely under the fvp board, and the Neoverse reference-design platforms that model real server silicon ship no DRTM support at all.
This session looks at what it takes to move Arm DRTM off the reference model and make it useful to Linux: porting the plat_drtm platform layer onto the Neoverse reference-design ports; delivering the DCE-Preamble as a UEFI/EDK2 application rather than a bootloader patch; measuring into PCR 17/18 with a DRTM event log; and the kernel-side plumbing to consume and attest the launch. Drawing on 3mdeb's TrenchBoot work and years in the Arm DRTM ecosystem, we want to map the upstreaming path across TF-A, EDK2, and the kernel, and name the gaps and their owners.
Desired outcome: consensus on a minimal, reproducible Arm DRTM reference that the community can build on FVP-first, then real silicon, and a shared list of the upstream changes needed to get there.
Speaker: Piotr Król (3mdeb)
-
15:00
→
18:30
VFIO/IOMMU/PCI MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128The PCI interconnect specification, the devices that implement it, and the system IOMMUs that provide memory and access control to them have become the de-facto standard for connecting high-speed components. The specification continues to expand with features spanning address translation (ATS/PRI), I/O virtualisation (SR-IOV/PASID/SVA), high-performance data movement (RDMA, P2PDMA, CXL/DOE), and hardware security (CMA, IDE, SPDM), targeting everything from server and desktop computing to embedded SoC platforms, virtualisation, and IoT devices.
The kernel code that enables these features requires tight coordination between PCI devices, the IOMMUs they sit behind, and the VFIO layer that exposes them to userspace for direct access and device passthrough. Kernel interfaces and userspace APIs across all three subsystems must evolve together to keep the overall design coherent.
Following the success of the VFIO/IOMMU/PCI MC at LPC 2017, 2019, 2020, 2021, 2022, 2023, 2024 and 2025, the 2026 edition will continue to bring together PCI, IOMMU, and VFIO developers to resolve cross-subsystem design questions that cannot be settled on individual mailing lists alone.
Video recordings from 2025 are available: LPC 2025 - VFIO/IOMMU/PCI MC. Older recordings can be found on the official Linux Plumbers Conference YouTube channel and the archived LPC 2017 VFIO/IOMMU/PCI MC web page at Linux Plumbers Conference 2017.
The tentative schedule will provide an update on the current state of the VFIO/IOMMU/PCI kernel subsystems, followed by focused discussion of current issues related to the proposed topics.
The following were outcomes of last year's successful Linux Plumbers MC:
-
The TEE I/O initiative for confidential compute advanced through a phased plan spanning PCI, VFIO, IOMMUFD, and KVM. Phase 1 (PCI core support for authenticated and encrypted links via the TEE Security Manager) is largely complete, and most of the dma_buf, VFIO, and IOMMUFD infrastructure for Phase 3 (host-side private MMIO and DMA) has been merged. The critical remaining issue is KVM integration: an RFC for mapping from KVM's guest_memfd within IOMMUFD is under active review. The community also reached consensus on using Sub-Stream IDs (PASID) for hardware-level isolation when transitioning devices between shared and secure states.
-
A new generic IOMMU page table infrastructure was merged in 6.19 for AMD and Intel VT-d, replacing six architecture-specific implementations with a unified codebase. Since the merge, RISC-V Svpbmt support, small VA support for AMD, and preserve/unpreserve/restore callbacks have been posted, with fixes flowing through the IOMMU tree for v7.1.
-
Building on the generic page tables, design consensus was reached on the Hyper-V pvIOMMU and an RFC v1 has been posted. Key decisions: simple gIOVA without PASID is acceptable, the HV IOMMU driver will not cover ARM SMMU, per-device PASID space is required, and hypercall-based enumeration without ACPI tables is acceptable. Follow-up work is addressing IOTLB tagging with domain ID differences on Intel VT-d and AMD IOMMU.
-
SDXI (Smart Data Acceleration Interface) and Smart Data Cache Injection (SDCI) via PCIe TLP Processing Hints (TPH) were presented, demonstrating up to 70% memory bandwidth reduction in 100G NIC benchmarks. Since the conference, a SDXI driver series has been posted and VFIO PCIe TPH support patches are under review.
-
The PCI power control (pwrctrl) framework rework reached design consensus and has since been merged, resolving timing flaws where host controller drivers deasserted PERST# signal before endpoints were powered. The new API restores PCIe spec compliance and enables proper D3cold suspend flows, with controller drivers already adopting the new interface.
Tentative topics that are under consideration for this year include (but are not limited to):
PCI: CCIX/CXL expansion memory and accelerator management; DOE; IDE; CMA; SPDM; IOASID allocation; INTX/MSI IRQ domain consolidation; Gen-Z interconnect fabric; error handling and management (AER, DPC, APEI, EDR); power management and ASPM; P2PDMA; resource claiming/assignment consolidation; DMA ownership models; Thunderbolt, DMA, RDMA, and USB4 security.
VFIO: I/O Page Fault (IOPF) for passthrough devices; SVA interface; SR-IOV/PASID integration; PASID in SR-IOV virtual functions.
IOMMU: /dev/iommufd development; IOMMU virtualisation; IOMMU driver SVA interface; DMA-API layer interactions and the move towards generic dma-ops for IOMMU drivers; IOMMU core changes (e.g., tighter integration with the device-driver core).
If you are interested in participating in this MC and have topics to propose, please use the Call for Proposals (CfP) process. Additional topics may be added based on CfP submissions.
Otherwise, join us in discussing how to help Linux keep up with the new features added to the PCI interconnect specification. We hope to see you there!
Key Attendees:
Alex Williamson, Benjamin Herrenschmidt, Bjorn Helgaas, Dan Williams, Ilpo Järvinen, Jacob Pan, James Gowans, Jason Gunthorpe, Jonathan Cameron, Jörg Rödel, Kevin Tian, Krzysztof Wilczyński, Lorenzo Pieralisi, Lu Baolu, Manivannan Sadhasivam.
Contacts:
- Alex Williamson (alwilliamson@nvidia.com)
- Bjorn Helgaas (helgaas@kernel.org)
- Jörg Rödel (joro@8bytes.org)
- Lorenzo Pieralisi (lpieralisi@kernel.org)
- Krzysztof Wilczyński (kwilczynski@kernel.org)
-
15:00
Secure vIOMMU architecture and implementation challenges 25m
Modern cloud computing increasingly relies on confidential virtual machines (VMs) - isolated environments
where sensitive workloads run securely, even from the underlying hypervisor. A key challenge in this space
is enabling these confidential VMs to safely communicate with hardware devices without exposing their data
to the host system or hypervisor.Virtual IOMMU (vIOMMU) is an extended feature of AMD's IOMMU that allows IOMMU hardware to directly access
and process virtual machine data structures. Building on this foundation, AMD's Secure vIOMMU extends
vIOMMU capabilities to support SEV-TIO (Secure Encrypted Virtualization - Trusted IO) device assignment to
confidential virtual machines. Specifically, it enables secure communication channels between confidential
virtual machine and IOMMU hardware enabling reliable management of device DMAs including guest TLBs without
hypervisor intervention.This presentation introduces the Secure vIOMMU architecture and explains how its components work together
-- including the IOMMU, AMD's Security Processor, the hypervisor, Virtual Machine Manager (VMM) and
confidential VMs. We will also cover supported operational modes (vTOM and Guest Page Table) for secure
data transmission between Trusted I/O devices and confidential VM.Finally, we will discuss the interface challenges that arise across various kernel modules including IOMMU,
ccp, KVM, SEV-TIO, etc, and explore potential approaches to address them.SEV-TIO specification: https://docs.amd.com/v/u/en-US/58271_0.91
Secure vIOMMU initial code: https://github.com/AMDESE/linux-iommu/tree/sviommu/tsm0407_v619_v0Speakers: Mr Suravee Suthikulpanit, Vasant Hegde -
15:25
IOMMU Page Table Observability and Reclamation 20m
The current Linux memory management framework lacks the mechanism for tracking IOMMU page table (IOPT) consumption across sparse virtualization workloads. By utilizing the IOMMUFD and VFIO frameworks, hypervisors dynamically map and unmap massive IOVA regions, often leaving intermediate page table directories stranded and invisible to the core kernel. This session talks about the current state of IOPT observability and attempts to propose new features and enhancements for IOPT memory reclamation.
The session proposes a design for an IOMMU Page Table Shrinker that allows lockless and lazy reclamation of stranded IOPT memory, specifically by the intermediate tables while employing the generic_pt to keep the shrinker arch-agnostic. We invite a discussion to explore strategies for lockless tracking, page table branch severance, and lifecycle management using RCU state machines to ensure that the core kernel can harvest stranded memory under pressure without penalizing the IOMMU mapping paths.Speakers: Logan Odell (Google), Pranjal Shrivastava -
15:45
IOMMUFD Support for Microsoft Hypervisor Hosts 25m
Device assignment in Microsoft Hypervisor (/dev/mshv) host environments has requirements beyond existing KVM-oriented workflows, including direct HWPT attachment for assigning a device to a guest while preserving existing Stage-2 page tables, and VMM-assigned virtual device information to facilitate hypercall-based device operations.
We plan to publish an RFC before LPC describing the proposed iommufd/VFIO plumbing for MSHV device assignment. At LPC, we would like to discuss the remaining design questions and potential follow-on work, including alignment with Xen and how the same model could support future use cases such as TDISP or page table sharing in general.
Speaker: Jacob Pan -
16:10
arm-smmu contig hint support and performance impact 20m
This talk covers ARM SMMU contiguous-hint support in Linux stage-1 page tables and the ARM-SMMU PMU support used to measure its impact. The work enables larger effective mappings, including 64K regions built from 16 adjacent 4K entries, and extends IOMMU map/unmap support for mixed and larger mapping sizes. In parallel, ARM-SMMU PMU integration exposes TBU and TCU counters through perf, including access, TLB allocate/read/write, cache-hit, and page-table-walk events, with StreamID-based filtering for workload attribution.
We present page-table behavior observed in tracing alongside translation-side measurements from PMU counters. Early measurements show measurable gains from mixed 4K/64K mappings with contiguous hint enabled: lower TLB allocation pressure and lower elapsed time versus a 4K-only baseline.
The talk also highlights different pmu counters used, invalidation effects, and how we can evaluate real SMMU translation performance.
Speakers: Prakash Gupta (Qualcomm), Mr Vijayanand Jitta (Qualcomm) -
16:30
Coffee Break 30m
-
17:00
VFIO PCIe TPH Userspace API & Progressive Security Policy 20m
PCIe TPH enables cache steering for high-performance P2PDMA, RDMA and SDXI workloads. Today Linux only supports TPH operations within host kernel space; userspace/VFIO passthrough environments lack a standard interface to resolve and program steering tags.
I’ve posted a complete patch series adding native VFIO TPH support, aligned with community incremental security policy design. The series provides unified handling for CPU memory (via root port ACPI _DSM), dma-buf, zero-tag and literal raw ST sources, supporting cross-device P2P DMA scenarios. Three new VFIO_DEVICE_FEATURE ioctls cover feature opt-in, multi-source tag resolution and batch ST table programming.
Current upstream challenges include:
Automated selftest coverage without physical TPH hardware;
Lock granularity performance tradeoffs for batch ST programming;
Security boundary definition for cross-device dma-buf TPH metadata access.
The core uAPI and security logic are stable and targeted for v7.3 merge window in late August. Post-merge follow-up plans include extending QEMU to emulate TPH capabilities and implementing full guest passthrough support with VMM-managed _DSM emulation.
This session will briefly introduce the design, then focus on resolving the outstanding upstream challenges and aligning on VM passthrough roadmap.Speaker: Mr Chengwen Feng (Huawei) -
17:20
Two Years of Making PCIe Hotplug Reliable: What Landed, What Didn't, and Where Do We Go 25m
In 2024, hotplug on our NVMe storage fleet — NTB behind Microchip/Switchtec switches — was
not stable on mainline, so we ran Keith Busch's out-of-tree bus-locking series on a private
branch. It held. Then a new hardware generation arrived, we needed a kernel we could
validate against, and that meant mainline: 6.8, now 6.17. Going upstream meant giving up the
structural work that had been keeping us alive, and we started hitting problems again. This
is the story of that round trip, and where I think it should go next.The good news is real, so I lead with it: since Keith Busch's ALPSS 2024 talk "The Problem
with PCI Subsystems," the subsystem has closed most of what bit us. The child use-after-free
atpci_dev_waitis fixed (11a1f4bc4736, v6.11). pciehp suppresses reset- and DPC-induced
link events (v6.16). DPC recovery holds a port reference (a1ed752bc7cb, v7.1).
reset_subordinateis hotplug-safe (8238cb69c01f, v7.1). DPC is decoupled from per-device
AER binding for v7.3 (97ca178c899d), which resolves — properly — a layering hack I had
been carrying in production. The point fixes landed.The structural work did not come with us, and its absence is where the new problems live.
Hierarchy mutation still sits behind one globalpci_rescan_remove_lock; Keith's per-bus
locking and subordinate-bus refcounting are still unmerged, two years on. Where exclusion
became coordination, the budgets do not match:pci_dpc_recovered()waits 4 s for a
recovery bounded at roughly 62 s. And a failure nobody was watching turned up — after a
switch-subtree reset, 3 of 8 endpoints cannot place a 1 GiB BAR that boot places without
complaint.I am not here to complain; my crash was fixed and I will say so. I bring crash dumps, source
analysis, a reproducible switch topology, and fleet-scale test capacity — the validation
bandwidth the subsystem said it was short on — and three questions about the way forward. I
want to leave the room agreeing on which one is worth fixing next, and who validates it.Speaker: Maciej Grochowski -
17:45
Describing P2PDMA capabilities in ACPI 25m
The kernel uses a whitelist do determine whether peer-to-peer DMA is supported between devices. But constantly amending the whitelist is a maintenance burden.
We are proposing to extend the ACPI HMAT table with bandwidth and latency characteristics for P2PDMA traffic between PCI host bridges (and between devices below the same host bridge).
This will remove the need for a whitelist.
It also closes a gap in virtualization: VMs have no visibility of the physical PCI topology and thus cannot determine P2PDMA reachability. The proposed HMAT extension solves this: VMMs such as qemu may present a synthesized HMAT table to VMs which reflects the physical capabilities. This replaces prior proposals such as Nvidia's Virtual Peer-to-Peer Approval Capability.
Last not least, the proposed HMAT extension allows users to populate PCI slots for optimal P2PDMA performance.
We will briefly present the draft of this ACPI Code First ECN and explain how it fits into the Linux P2PDMA model and path-selection algorithm. Afterwards we hope for a lively discussion on its merits.
Speakers: Leon Romanovsky, Lukas Wunner -
18:10
Testing the Untestable: Using QEMU for Testing PCI Endpoint Framework 20m
The Linux PCI Endpoint Framework provides the software infrastructure for a Linux system to present itself as a PCIe Endpoint device to an external Host. It consists of an Endpoint Controller layer that abstracts the hardware, an Endpoint Function layer that implements device behavior, and a ConfigFS interface for userspace to bind them together. For testing, the upstream pci_endpoint kselftests are used. But running this kselftest has always required two physical machines connected by a physical PCIe link. This makes automated testing impractical and limits coverage to whoever has the right hardware.
This talk presents a software only test setup that removes the hardware requirement entirely. Two QEMU instances communicate over a Unix socket, one acting as the Endpoint Controller and the other as the Root Complex. A new generic EPC driver on the Endpoint side talks to a virtual Endpoint Controller device emulated by QEMU, which forwards commands over the socket to the RC side where QEMU emulates a PCIe Endpoint Function that the Host kernel enumerates and tests against. The existing kselftests run unmodified against this setup.
This talk will walk through how the two VM design works, what trade-offs were made in the socket protocol and what remains to be done for full CI integration.
Speaker: Manivannan Sadhasivam
-
-
19:30
→
22:30
Evening Event 3h
-
10:00
→
13:33
-
-
10:00
→
13:30
AI-Assisted Open Source Development MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128Overview
AI coding tools (LLMs, code assistants, AI agents) are rapidly becoming part of the developer workflow across the software industry. Open source communities are beginning to grapple with how these tools intersect with their development processes — from code generation and review assistance to documentation, debugging, and large-scale refactoring. This microconference will bring together maintainers, developers, and tooling experts to discuss the practical realities, policies, risks, and opportunities of AI-assisted development in the open source ecosystem.
The goal is not to debate whether AI tools will be used — developers are already using them — but to align on how communities should adapt their processes, what guardrails are needed, and where these tools can deliver the most value with the least risk.
Example Subtopics
- AI-assisted code generation and review — practical experiences, failure modes, and disclosure norms
- Large-scale refactoring with AI assistance, such as C-to-Rust conversions
- AI for debugging, crash analysis, and root cause identification
- Policy and process implications: attribution, copyright, licensing, and trust in AI-generated contributions
- AI-powered test case generation and fuzzing guidance
- Building project-aware AI tooling with domain-specific context and integration with existing development infrastructure
Key People Who Should Attend
- Major subsystem and project maintainers who are receiving AI-assisted contributions and need to make policy decisions
- Developers actively using AI tools in their open source workflows who can share real-world experiences
- AI tooling developers building tools targeted at open source and systems-level development
- Linux Foundation / legal experts for the policy, licensing, and attribution discussion
Previous Related Sessions
While there has not been a dedicated AI microconference at LPC before, related discussions have touched on adjacent topics in microconferences such as Kernel Testing & Dependability, Rust for Linux, and Toolchain. This would be the first session to bring together the AI-specific cross-cutting concerns that span all of these areas.
Expected Outcomes
- Community alignment on disclosure and attribution requirements for AI-assisted contributions
- Identification of high-value, low-risk use cases where AI tools should be encouraged
- Concrete next steps for building project-aware AI tooling
- Framework for evaluating AI-generated code in the review process
-
10:00
Sashiko current status 45m
Current status of Sashiko. Also in this block: #140 "Citation Needed: Treating the kernel's AI review context like code" (Fuad Tabba), and Chris Mason and others on review prompts.
Speaker: Roman Gushchin (Google) -
10:45
Sashiko feedback and feature requests 30m
Open discussion with everyone: feedback on Sashiko and feature requests.
Related talks: #140 Fuad Tabba (Google); #166 Andrea Righi (NVIDIA); #298 Konstantin Sinyuk (Intel)
-
11:15
Automating Kernel Bug Fixes: Syzbot's AI-Assisted Patching Workflow 15m
Syzbot reports around 2,000 new fuzzer-detected findings in the Linux kernel each year. Even for straightforward bugs where the root cause is obvious, the manual effort of drafting, testing, and sending patches constitutes a significant effort on the kernel developers side.
To facilitate this process, we have launched an AI-assisted fix generation workflow in syzbot built around human-in-the-loop review:
* Two-Stage Pipeline: Candidate patches are sent to a moderation list where reviewers reply with natural-language feedback or commands (#syz reject,#syz upstream). The AI agent interprets comments, answers questions, and iterates on new versions while preserving tags like Reviewed-by or Tested-by.
* Human Accountability: When approved via the#syz upstreamcommand, the human reviewer's Signed-off-by: tag is attached in compliance with Linux Kernel AI Coding Assistants guidelines before forwarding the patch to public mailing lists.Status & Discussion: The system is in a POC stage and undergoing active development. As of September 2026, 30 of the syzbot AI-assisted patches have been merged upstream, more are under review/discussion.
In this session, we will focus on:
* System Overview: More detailed workflow description and how the system works under the hood.
* Lessons Learned: Based on the current experience, what went well and what needs more attention.
* Community Discussion: How to make the tool more useful and increase review efficiency.Speaker: Aleksandr Nogikh (Google) -
11:30
Coffee Break 30m
-
12:00
Boro: Do We Need More AI-assisted Kernel Tooling? 15m
AI-assisted tooling is increasingly becoming part of Linux kernel development workflows, particularly for patch review and feedback generation. Sashiko has become the de facto standard for reviewing publicly posted patch series on mailing lists.
While Sashiko can also be used beyond the mailing-list workflow, its design is primarily optimized around upstream review interactions. This can result in a less natural fit when it's applied to more iterative or downstream-focused development workflows. In particular, it's not specifically designed for reviewing backported patch series locally or tracking upstream follow-ups related to those changes (e.g., distro kernel maintenance, applying security fixes, etc.).
These gaps motivated the development of Boro (https://github.com/NVIDIA/boro), an AI-assisted kernel development tool designed for local-first, interactive usage, focused on improving review and testing iterations before public submission, and supporting downstream workflows such as backporting and distribution kernel development.
This talk will present Boro's design and open a discussion on whether some of its functionality could be integrated into Sashiko.
Speaker: Andrea Righi (NVIDIA) -
12:15
Automated kernel CVE backports 15m
The Problem: CVEs
- CVEs: Common Vulnerabilities and Exposures
- February 2024, Linux kernel project becomes the Linux kernel CVE Numbering Authority (CNA)
- The number of Linux kernel CVEs skyrockets
- Red Hat customers expect CVE fixes/mitigations, with some having Service Level Agreements for delivery within X number of days, depending on severity
- Red Hat did not get an increase in kernel developers commensurate with the increase in kernel CVEsThe Solution: Automation
- Don’t manually do things that can be done by robots! (*)
- Assignment of the work to the right team
- Identification of the patch fixing the CVE
- Backport of the fix (possibly with AI coding assistance)
- Submission of a merge request containing the fix
- Validation of the merge request
- Testing of the fix(*) This does not necessarily mean AI, but it also does now
Speaker: Jarod Wilson (Red Hat) -
12:30
From Findings to Fixes: AI-Assisted Linux Kernel Bug Triage and Remediation 15m
AI-assisted tools are producing an increasing number of Linux kernel bug findings. While these tools can uncover genuine bugs, many findings still require validation before they can be treated as actionable issues. Reports may be false positives, lack a working reproducer, duplicate existing issues, or fail to trigger under the reported conditions.
We would like to share our experience triaging over 1,000 Linux kernel bugs, with a focus on validating findings, understanding the privilege required to trigger them, and assessing their security impact. These steps help distinguish false positives and low-impact corner cases from bugs that deserve higher remediation priority.
To support this process, we are developing a continuous Linux kernel bug collection and verify platform[1] inspired by syzbot. It aggregates findings from multiple sources, like sashiko, and attempts to generate and execute reproducers for them. For bugs with runtime evidence, the platform further analyzes trigger privilege, impact, severity and generating patches.
[1] https://lore.kernel.org/all/20260830115546.3942129-1-yuantan098@gmail.com/
Speaker: Yuan Tan (NebuSec) -
12:45
Open discussion: using fix generation effectively 15m
Open discussion led by Josef with the audience on using AI fix generation effectively.
Related talks: #447 Aleksandr Nogikh; #482 Yuan Tan; #315 Jarod Wilson
Speaker: Josef Bacik (Facebook) -
13:00
Talk to Your Crash Dumps: Debugging a Real-Time Kernel Stall with drgn and AI 15m
Analyzing kernel crash dumps with drgn is powerful but demands fluency in drgn's Python API, kernel data structure layouts, and the right helper functions for each subsystem. drgn-mcp is an MCP (Model Context Protocol) server that exposes drgn's debugging capabilities as structured tools that AI assistants can call, enabling kernel developers to investigate crash dumps through natural language instead of scripting.
This talk is a live demo of a real-world investigation. I will use drgn-mcp to analyze a vmcore from a production RT kernel hang where rtmutex contention leads to a system-wide stall. This is a real customer bug affecting RHEL RT kernels. Starting from the vmcore with no prior analysis, I will walk through the entire investigation interactively, driven by natural language, with the LLM choosing which drgn tools to invoke to uncover the root cause live on stage.
Speaker: Mr Wander Costa (Red Hat) -
13:15
Wrap-up 15m
Session wrap-up and next steps.
Speaker: Josef Bacik (Facebook)
-
10:00
→
18:30
Birds of a Feather (BoF) "Club D" (Prague Congress Centre)
"Club D"
Prague Congress Centre
53-
13:30
Lunch 1h 30m
-
15:00
Linux CVEs, vulnerability management and agentic security workflows 45m
This year has been marked by the rise of cyber-capable AI models, flooding maintainers with increasingly better vulnerability reports. More vulnerabilities have been reaching the headlines, resulting in emergency mitigations across different deployments. In response, many organizations have adopted agentic solutions to scan, triage, reproduce, and patch vulnerabilities.
During this year's BoF, we plan to not only update the community on our ongoing efforts with CVE triaging (https://github.com/cloud-lts/linux-cve-analysis), but also to discuss these common agentic efforts. Our objective is to boost cross-organizational collaboration and support the security community in navigating these new challenges.
Speaker: Damiano Melotti (Google) -
15:45
Safety with Linux 45m
Opportunity to discuss and share experiences about ongoing attempts to use Linux in safety applications e.g. automotive.
Problems identified, technical obstacles, tentative solutions.
Starting the discussion, NVIDIA with ongoing work, in the space of adapting existing features to safety, but open to anyone who wants to share their experience and perspective.
Speaker: Igor Stoppa (nvidia) -
16:30
Tea Break 30m
-
13:30
-
10:00
→
18:30
Birds of a Feather (BoF): No A/V "Club B" (Prague Congress Centre)
"Club B"
Prague Congress Centre
53-
10:00
Kernel Sanitizers Office Hours 45m
The Linux kernel has numerous tools to detect bugs, among them a family of dynamic program analysis called "sanitizers": Kernel Address Sanitizer (KASAN), Kernel Memory Sanitizer (KMSAN), Kernel Concurrency Sanitizer (KCSAN), and the Undefined Behaviour Sanitizer (UBSAN). Although not a bug-detecting sanitizer itself, we will also cover Kernel Coverage (KCOV), which utilizes similar compiler instrumentation to collect and expose execution feedback.
Knowing when to apply which sanitizer in the kernel development process may not always be obvious: each sanitizer is dedicated to finding a different class of bugs, and each introduces some amount of performance and/or memory overhead. Not only that, each sanitizer also provides a range of options to tweak their abilities.
This session is dedicated to briefly introducing each kernel sanitizer, the bug classes they help detect, and important gotchas when using them.
The rest of the session is dedicated to answering questions around each of the sanitizers, KASAN, KMSAN, KCSAN, and UBSAN. Feel free to also share success stories that may give other attendees only starting out with some of the sanitizers ideas how to best apply them.
Speakers: Aleksandr Nogikh (Google), Alexander Potapenko (Google), Dmitry Vyukov (Google), Justin Stitt (Google), Kees Cook (Google), Marco Elver (Google), Pimyn Girgis (Google) -
11:30
Coffee break 30m
-
12:00
x86 BoF 45m
This will be an x86-focused session. Broadly speaking, anything that might affect arch/x86 is on topic, except where there may be a more focused discussion occurring, like around Confidential Computing or KVM.
We'll talk about new ISA, hardware security processes and structural issues affecting arch/x86.
Speaker: Dave Hansen -
12:45
DRM Fabric: Vendor-Neutral Topology Infrastructure for Scale-Up Accelerator Interconnects 45m
Modern GPUs and dedicated AI accelerators are increasingly connected through scale-up interconnect fabrics such as AMD xGMI today and UALink in the near future. Yet Linux lacks common topology infrastructure for reporting, vendor-neutrally, which accelerators are directly connected, through which ports, and in what state. Vendors expose fragments privately — e.g.
amdgpu's xGMI sysfs — and prior per-driver proposals, such as XeLink, never became shared Linux infrastructure. As a result, topology semantics are being defined by vendor-private uAPIs rather than by a shared Linux control-plane contract for monitoring and fabric management.Building on the LPC 2025 "Toward Mainline Linux Support for UALink" BoF, we propose DRM Fabric: a vendor-neutral, protocol-agnostic DRM topology infrastructure for scale-up interconnects, using the model:
fabric → endpoint → port → peerVendor drivers populate it through a thin provider API; userspace queries it through
drm-fabric, a YAML-defined generic-netlink family consumed throughynltooling, following the netlink uAPI patterndrm_rasintroduced to DRM. No data path: load/store traffic and memory semantics remain in the vendor driver, so the infrastructure remains valid across vendors and fabrics.The implementation is split into two RFC series. Phase 1 targets provider-owned fabrics, such as an xGMI-like implementation, where the driver discovers topology and reports it read-only to userspace, including notifications, optional port statistics, and consistent multipart enumeration. Phase 2 adds privileged provisioning for software-defined, UALink-style fabrics: creating fabric containers, attaching orphan endpoints, setting administrative state, and provisioning peers on userspace-managed ports, while providers continue to report actual hardware state. The RFCs include the core infrastructure, an in-tree synthetic provider (
fabricsim, modeled onnetdevsim), and KUnit pluspyynl/kselftests exercising the infrastructure and uAPI end-to-end without hardware. The next step is a hardware-backed provider for an already-shipping scale-up fabric.We want LPC feedback on the DRM representation, the boundary between provider-owned topology and privileged userspace provisioning, whether the proposed provisioning primitives are sufficient, and how AMD, Intel, UALink Consortium members, and other accelerator vendors should co-develop the common infrastructure and provider interfaces.
Speaker: Mr Konstantin Sinyuk (Intel) -
13:30
Lunch 1h 30m
-
15:00
Bridging the Linux Tux and the FreeBSD Beastie – Linuxulator, LinuxKPI, and LLM-Assisted Porting Challenges on License Issue 45m
The boundary between the Linux and FreeBSD ecosystems is often loose and porous,while developers continuously building bridges to maximize hardware support and software availability across both platforms.
Two mechanisms carry that weight on shoulders: the Linuxulator, a native
implementation of the Linux syscall ABI that runs unmodified Linux binaries,
and LinuxKPI, a shim layer that lets FreeBSD build largely unmodified Linux
drivers -- the DRM drivers via drm-kmod, mac80211-based wireless drivers, and a growing set of others -- against a subset of the Linux in-kernel API.Depending directly on the ripples from Linux's advancements, Linuxulator tracks syscall, vDSO, /proc, netlink and seccomp behaviors and LinuxKPI tracks an in-kernel API that is explicitly not stable, which makes every DRM or wireless driver during release cycle into an aerobics workout.
Hence, this Birds of a Feather (BoF) session is then hereby proposed to serve as an information exchange hub for developers working at the intersection of Linux and FreeBSD kernel mechanisms.
Furthermore, FreeBSD is not the only Frankenstein here: Fuchsia's starnix and illumos LX zones reimplement parts of the exposed surface of Linux userspace, and several BSDs have maintained downstream of DRM shims of their own. So if you are also an amphibian developer other operating systems. Don't hesitate to join us :-)
During this session, we may discuss topics including but not limited to :
- Linuxulator (Linux Emulation Layer): The current state, limitations, and future enhancements of running unmodified Linux binaries on FreeBSD. E.g. RISC-V architecture.
- LinuxKPI: The compatibility framework designed to ease the porting of complex Linux kernel drivers (such as DRM and wireless drivers) to FreeBSD. We will discuss the friction points driver maintainers face and how the Linux community's API changes impact downstream compatibility.
- The LLM Porting and License Issues: With the rise of Large Language Models (LLMs), developers are actively experimenting with AI to assist in translating or porting drivers from Linux to FreeBSD. However, this introduces complex license compatibility questions i.e. GPLv2 vs. BSD. And thus shall deserve a chat with beers in my humble opinion. And here's the living witness of such kind of work : aqtion-freebsd-aq2 work.
Related discussions could be found in AsiasBSDCon 2026
Speakers: "Ruinland" ChuanTzu Tsai (Andes Technology), Ed Maste (FreeBSD Foundation), Mr Li-Wen Hsu (FreeBSD Foundation) -
15:45
F2FS update 45m
Touch F2FS topics:
- Large folio
- Sub-page block support (16KB page on 4KB block)
- Device Aliasing for AndroidSpeaker: Jaegeuk Kim (Google) -
16:30
Coffee break 30m
-
17:00
Improving THP placement policies 45m
Transparent Huge Page placement has historically been (and is) quite simplistic and overeager. This normally results in users turning off THP completely, or for selected workloads using any number of available toggles.
I'd like to discuss realistic ways in which we could meaningfully improve this, including past ideas such as thp=auto.
Several relevant aspects would include:
- adapting THP placement based on filesystem/storage needs
- adapting THP placement based on userspace memory accesses
- adjusting and improving expectations on MADV_COLLAPSE
- getting rid of hard-to-tune heuristics like max_ptes_none & othersSpeaker: Pedro Falcato (SUSE Labs) -
17:45
DAMON Nano Conference 45m
DAMON is a kernel subsystem for efficient data access/attributes monitoring and memory operations. It received many updates last year. Many more updates are in progress. Multiple ideas for future works are flooding. In this session, multiple people who work on DAMON will present and discuss what changes have recently been made, what changes are ongoing, and what should be the next work.
Schedule
Note that the schedule could be updated until the last minute.
17:45-17:48 Opening, SJ Park. Slides
17:48-17:55 Breaking through Accessed Bit Limits of DAMON, SJ Park, Ravi Jonnalagadda, Akinobu Mita. Slides
17:55-18:05 Beyond Weighted Interleaving: Bandwidth-Driven Memory Tiering with DAMON, Ravi Jonnalagadda. Slides
18:05-18:15 DAMON for Huge Pages
- Guiding THP Decisions with DAMON Memory Monitoring, Asier Gutierrez. Slides
- Host-side DAMON Hotness under KVM/THP: Granularity Gaps and Evidence for Demotion, Lian Wang, Kunwu Chan. Slides
18:15-18:25 DAMON in the AI Cloud: Monitoring Real-World GPU Workloads, Krishna Iyer. Slides
18:25-18:30 Joint QnA and ClosingAuthors (Alphabetically sorted)
Akinobu Mita
Software engineer at Fixstars Corporation.Asier Gutierrez
Senior Software Engineer at Huawei. Experienced Senior Software Developer specializing in Linux development, kernel-level debugging, and high-load, high-availability systems. I currently work for Huawei. I have worked for Intel, IBM and Yandex in the past.Krishna Iyer
Krishna Iyer is a Software Engineer at Crusoe specializing in low-level systems, driver development, and Linux build systems. Previously at Cisco-Meraki, he now focuses on researching memory access patterns using DAMON for AI workloads on GPU infrastructure.Kunwu Chan
Kunwu Chan is a Linux kernel developer with experience in kernel performance analysis and tuning, subsystem design, and memory safety. He works on DAMON, memory management, swap, MGLRU, and RCU.Lian Wang
Lian Wang is a Linux kernel developer working on memory management, with a focus on DAMON, huge pages, swap, and virtualized and large-memory systems.Ravi Jonnalagadda
Ravi Jonnalagadda is a Sr Manager at Micron working on Linux kernel memory management for CXL and tiered-memory systems. He co-developed the weighted interleaving memory policy and added the node_eligible_mem_bp DAMOS quota goal. His current work lets DAMON take access samples from hardware sampling units.SJ Park
SJ is a kernel programmer with a strong focus on memory management. He develops and maintains DAMON, a Linux kernel subsystem designed for efficient data access monitoring and access-aware system operations.Speaker: SJ Park
-
10:00
-
10:00
→
13:30
Devicetree MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53The Devicetree Microconference focuses on discussing and solving problems present in the systems using Devicetree as firmware representation. This notably is Linux kernel and U-Boot, which share the Devicetree bindings and sources, but also can cover topics relevant to Zephyr or System Devicetrees. Systems using Devicetree are majority of embedded boards, mobile devices and ARM64 laptops.
Ongoing problems, being discussed last year in LPC 2025 or previous years:
-
Status of DTS validation against DT schema among SoC platforms: are we getting to error-free dtbs_check anywhere? What are the blockers in achieving compliance, what is the progress.
-
Hot-pluggable hardware with Devicetree overlays (addons) - on-going efforts, discussed also on LPC 2024. Current work includes changing the DTB format (RFC patches posted).
-
Sharing DT bindings and DTS sources with U-Boot (aka OF_UPSTREAM): progress and what are the obstacles?
-
Shall we migrate all of_property_read_xxx() calls in Linux drivers to device_property_read_xxx() to handle also ACPI?
-
Fixing common pattern of unconditional device_init_wakeup() in drivers which makes it impossible to disable it via Devicetree, since wakeup-source is bool.
-
Style-checker (aka checkpatch) for DTS - tool automating all style related reviews. Discussed in 2025, but no tool got wide acceptance.
-
Power sequencing for enumerable busses - is it done yet? Discussion in 2025 suggested that at least MDIO is suffering from lack of generic solution for power sequencing.
-
How to choose and apply overlays, when vendor wants to ship many of them with single image. Many Android builds follow such approach. No generic properties/bindings were accepted so far. Discussed also in 2025.
-
DTB selection on EFI systems like arm64 laptops or embedded boards: How to store, update and choose the DTB to pass to the Linux kernel? The problem might be solved by Ubuntu Stubble, so is it considered a community consensus? What is still missing?
Key attendees:
AngeloGioacchino Del Regno, Arnd Bergmann, Bartosz Golaszewski, Bjorn Andersson, Chen-Yu Tsai, Conor Dooley, Douglas Anderson, Geert Uytterhoeven, Hervé Codina, Konrad Dybcio, Luca Ceresoli, Michal Simek, Nishanth Menon, Rob Herring, Saravana Kannan, Thierry Reding, Wolfram SangExpected attendees (very likely to come): Arnd Bergmann, Bartosz Golaszewski, Bjorn Andersson, Chen-Yu Tsai, Conor Dooley, Geert Uytterhoeven, Hervé Codina, Konrad Dybcio, Luca Ceresoli, Michal Simek, Wolfram Sang
-
10:00
DT bindings and sources shared with U-Boot: current status and ongoing work 20m
Since early 2024, U-Boot has been synchronizing the DTS and bindings of each Linux release into the U-Boot source tree using git subtree functionality. Since this transition, some platforms like Amlogic or Qualcomm have immediately switched, and some older platforms have been actively migrating to upstream Linux DT sources and bindings while maintaining minimal local U-Boot DTS modifications.
Neil will summarize the current state of U-Boot's integration of these upstream DTS, provide an update on ongoing migrations, and outline the remaining work required to fully eliminate local U-Boot DTS modifications.
Speaker: Neil Armstrong (Linaro) -
10:20
Status of the DTS Validation in the Linux Kernel 20m
The great benefit of Devicetree bindings in the current DT schema format is the ability to validate the correctness of DTS (Devicetree sources) against those bindings. However, once validation was introduced, we discovered that many in-kernel DTS files simply did not pass.
Continuing such summary from 2025, what is the status of dtbs_check now? Which platforms have the most warnings and which are warning-free? What improved over last year?
This is a similar talk to one in 2025 LPC showing current stage of dtbs_check.
Speaker: Krzysztof Kozlowski (Qualcomm) -
10:45
SDCA and Devicetrees 20m
Last year's LPC discussion explored whether ACPI-defined hardware descriptions could be reused on systems booting with Devicetree, and where the boundary should exist between the two firmware ecosystems.
Since then, several related efforts have continued across the kernel community, having intermediate representation for MIPI DISCO tables for SDCA, emerging ACPI-DT hybrid approaches. At the same time, new platform requirements—particularly on ARM64 laptops and embedded systems—continue to challenge the traditional "ACPI or DT" model.
This session will review the current state of these efforts, How should systems without ACPI tables should work with drivers like SDCA, discuss practical use cases where combining ACPI and Devicetree provides value, examine ongoing hybrid firmware work, and identify remaining technical ad architectural challenges. The goal is to gather feedback on whether the community should continue evolving hybrid solutions, standardize common mechanisms, or pursue alternative directions given that we have more data points.
Key Highligths:
- Ongoing efforts.
- Current Status of work on SDCA.
- ACPI-DT hybrid mode and future.
- Systems without ACPI tables(mobile devices).
Hans is covering ACPI-DT Hybrid mode in detail https://lpc.events/event/20/contributions/2363/Speaker: Srini Kandagatla -
11:10
Streamlining DT Bindings to Protect Upstreaming Momentum 20m
Corporate adoption of an "upstream-first" strategy is gaining traction, with companies increasingly willing to allocate engineering bandwidth to contribute hardware support directly to the Linux kernel. However, integrating these contributions smoothly remains a practical challenge. The review process frequently experiences friction during Device Tree (DT) binding discussions, which can extend timelines, consume allocated engineering bandwidth, and occasionally discourage newcomers to the open-source ecosystem.
This discussion explores ways to accelerate DT binding adoption so developers can maintain their focus on driver implementation. We will discuss current friction points and propose solutions to streamline the process. Specifically, we will evaluate the concept of "experimental bindings"—exploring whether provisional bindings could be used under certain circumstances to unblock driver development and maintain contributor momentum without compromising the long-term quality of hardware descriptions.
Speaker: Amit Kucheria -
11:30
Coffee Break 30m
-
12:00
Device Tree Addons: Describe hot-pluggable extension boards 25m
Device Tree has been extremely successful at describing non-discoverable hardware in Linux and other embedded systems. However, it remains fundamentally monolithic: each Device Tree describes a complete system, and reuse across independently developed hardware components is limited.
This becomes problematic for modular platforms where expansion boards, mezzanines, daughter cards, or hot-pluggable hardware are expected to work across multiple base boards. While Device Tree Overlays provide a mechanism to modify an existing hardware description, they require detailed knowledge of the target system. They do not offer a robust way to describe addon boards independently of the base platform nor a way to use the same addon board description to describe the same board connected on multiple connectors available on a single base board.
The recently proposed Device Tree Addon format [1] aims to address those limitations. It allows a single addon DTB to describe an extension board that can be used on different connectors of a given base board, or on different base boards, without requiring board-specific customization.
The design goal of Device Tree Addon is illustrated in slides used for a talk given at ELCE 2025 [2]. The relevant slides sections are the following:
- Use case and goals
- The connector abstraction.The "Connector: current status" section available in [2] is now obsolete. The Device Tree Addon concept emerged from discussions [3] that followed the ELCE 2025 conference.
This session will briefly present Device Tree Addon concepts recently
introduced and available in the RFC implementation [1]. Beyond introducing the format, the focus will be on the remaining challenges that must be addressed before broader adoption.Topics for discussion include remaining design questions, integration with
Device Tree bindings and validation rules, standardization efforts in Devicetree Specification (DTSpec). We will also examine the work still required in dtc, libfdt, bootloaders, and the Linux kernel.The goal of the session is to gather feedback from Device Tree maintainers,
kernel developers, bootloader developers, and users of modular hardware platforms to help shape the future direction of Device Tree Addons.A Bootlin blog post about Device Tree addons (based on its current state) is available [4].
[1] https://lore.kernel.org/all/20260826094950.1088288-1-herve.codina@bootlin.com/
[2] https://bootlin.com/pub/conferences/2025/elce/ceresoli-hotplug-status.pdf
[3] https://lore.kernel.org/all/20250902105710.00512c6d@booty/
[4] https://bootlin.com/blog/device-tree-addons-describing-non-discoverable-addon-boards-and-connectors/Speaker: Hervé Codina -
12:30
Using DT overlays to manage a collection of related boards 25m
It is commonplace that many boards in the real world have many sibling or cousin boards that are 95-99% the same as each other. Some examples:
- During development, most boards go through several revisions. A board might have proto0, proto1, evt0, evt1, evt1.1, dvt1, pvt, and mp revisions. These revisions almost the same with just small changes. While only "mp" (mass production) devices should end up in the public's hands, old prototype devices continue to be used for many years within the companies supporting the hardware.
- Many boards have several small variants that are nearly the same. As an example, the Pixel 10, Pixel 10 Pro, and Pixel 10 Pro XL are based on the same design and are nearly the same.
- Many boards are based on a reference design and only make small changes from that reference design. Chromebooks are one example of this. "sc7180-trogdor" was a reference design and a pile of different ODMs took this design and made small changes. Outside of Chromebooks, many phones are heavily based on the reference design of the SoC vendor.
Products that are highly similar to each other usually end up sharing a single binary image that is shipped out. That single image contains all needed device trees. Firmware on the product can figure out which specific device it's running on and can pick the correct device tree.
The common practice in the upstream device tree repository today is that each of these very similar boards needs its own complete device tree. The device tree is often shared using #includes of .dtsi files, but the final result is a pile of .dtb files that are mostly the same.
There are a few problems we have today. Two of them that are perhaps the most important:
- All of those device trees take up a lot of space. It would be nice if we didn't need a whole pile of .dtb files that are nearly the same.
- While firmware has all the details about the exact device it's running on, there is no standard way to mark device tree files so firmware can find the right one.
The device-tree overlay concept should help with point #1 in reducing the amount of space things take up. Trying to get agreement about how this should look has been difficult, though. One question that came up is if a "base" device tree was required to be a board or not. For supporting Pixel 10 hardware, it was most convenient for the base device tree to represent the SoC itself since some Pixel 10 devices had a different revision of the SoC. A second question that came up is how to deal with the top-level device-tree "compatible" in this case. It was unclear if we needed some way to "merge" the compatible based on all the overlays included which then led into questions about what needed to be in this top-level "compatible" to begin with.
Trying to solve point #2 has come up again and again, but somehow we never end up with any resolution. The typical refrain is that we need a grand unified design that all device-tree loaders will agree to use, but this always fails to materialize.
Both problems have downstream solutions. As an example, Pixel phones ship with devicetree overlays to avoid some of the duplication. Device trees and their overlays on Pixel phones have special properties in them that are used to help firmware pick the right one at boot time. Trying to upstream these solutions has met with difficulty though.
Let's have a discussion where we look at prior attempts and see if we can make some more forward progress.
Some threads I've been involved in (this is by no means a complete list of threads about the topic):
Speaker: Doug Anderson -
13:00
OS managed Devicetree overlays through UKI add-ons 25m
SBCs like the Arduino UNO Q can have a number of extension boards, like the UNO Media Carrier or an Arduino shield connected. On top of this extra hardware may be connected through e.g. CSI camera connectors and DSI display connectors.
Each of these possible hardware addons needs a DT overlay to work. This requires some method for the user to select which DT overlays to use.
There have been several attempts a decade ago to allow loading DT overlays from Linux userspace for this, but non has ever been accepted upstream.
UKI addons can provide an updated DTB, the purpose of this session is to present and discuss a proposal to use this to manage DT overlays from userspace through UKI addons:
Phase 1: UKI addons with a replacement DTB file embedded
-
For systems booting with secureboot, create DTBs with overlays merged in for popular hardware addon combinations and build a signed UKI addon for each such DTB on the distros build-system.
-
For systems without secureboot any overlay combination can be supported by building a merged DTB locally and creating an UKI addon from that DTB.
Phase 2: UKI addons with a DTBO file embedded
- For this phase the UKI stub needs to be extended with the capability to merge DTBOs from addons into an earlier loaded DTB, allowing any combination of overlays by the distro providing a signed UKI addon for each available overlay.
Speakers: Agathe Porte (Qualcomm), Hans de Goede (Qualcomm) -
-
-
10:00
→
13:30
KVM MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128KVM (Kernel-based Virtual Machine) enables the use of hardware features to
improve the efficiency, performance, and security of virtual machines
created and managed by userspace. KVM was originally developed to host
and accelerate "full" virtual machines running a traditional kernel and
operating system, but has long since expanded to cover a wide array of use
cases, e.g. hosting real time workloads, sandboxing untrusted workloads,
deprivileging third party code, reducing the trusted computed base of
security sensitive workloads, etc. As KVM's use cases have grown, so too
have the requirements placed on KVM and the interactions between it and
other kernel subsystems.The KVM Microconference will focus on how to evolve KVM and adjacent
subsystems, with a strong emphasis on all things guest_memfd.Potential Topics:
- 1GiB hugepage support for guest_memfd[1]
- KVM Userfault, or: demand paging support for guest_memfd[2]
- Removing guest memory from the host kernel's direct map[3]
- Eliminating "struct page" for guest_memfd
- Paravirtual scheduling
- Nested virtualizaton optimizations, e.g. PV APIs for "nested" VMsSuccesses from LPC 2024:
- KVM x86's mediated virtual PMU support landed in 7.0[4]
- pKVM support for protected anonymous memory landed in 7.1[5]
- In-place private<=>shared conversion for guest_memfd[6] is nearing
inclusion (likely 7.2 or 7.3)[1] https://lore.kernel.org/all/cover.1747264138.git.ackerleytng@google.com
[2] https://lore.kernel.org/all/20250618042424.330664-1-jthoughton@google.com
[3] https://lore.kernel.org/all/20260410151746.61150-1-kalyazin@amazon.com
[4] https://lore.kernel.org/all/20251206001720.468579-1-seanjc@google.com
[5] https://lore.kernel.org/all/177505732748.363663.16964917665296494635.b4-ty@kernel.org
[6] https://lore.kernel.org/all/20260507-gmem-inplace-conversion-v6-0-91ab5a8b19a4@google.com-
10:00
mmu_notifier: can we use it better? 30m
KVM uses the mmu_notifier infrastructure to keep the guest mapping up to date. After having completely rewritten the memory management of KVM/s390 to use mmu_notifiers, I have noticed some of the shortcomings in how KVM uses the notifiers.
Some of areas where KVM's usage of the notifiers, or the mmu_notifier infrastructure itself can be improved:
- passing the reason code to the KVM callback when unmapping
- more fine-grained reason codes
- passing back flags from the notifier callback
I will show how and why each of the above points would be useful for KVM with some concrete examples.
Speaker: Claudio Imbrenda (IBM) -
10:30
Exploring KVM and SMMUv3 Stage-2 Page Table Sharing for arm64 pKVM 30m
As Protected KVM (pKVM) progresses on arm64, achieving full isolation requires hypervisor-controlled IOMMUs to prevent DMA attacks. While there has been progress on SMMUv3 support for pKVM via trap-and-emulate on the mailing list [1], a major architectural design question remains unresolved: managing the stage-2 translation tables.
Currently, the series relies on maintaining a shadow stage-2 page tables for the SMMUv3.
Architecturally, it is possible to share these page tables between the CPU and the SMMUv3 in a similar way the kernel does SVA for stage-1, which would reduce memory overhead. However, KVM has different design constraints than user-space processes, and we are expected to run on a wide range of hardware without introducing impractical limitations or assumptions brought by SVA.Some of the main points:
- Dealing with page faults (Pinning, BBM…)
- IO TLB invalidation and DVM
- Cache coherency of the IOMMU and page tablesThe goal of this session is to discuss the viability of shared stage-2 page tables and reach a path forward for DMA isolation on pKVM.
[1] https://lore.kernel.org/all/20260501111928.259252-1-smostafa@google.com/
Speaker: Mostafa Saleh (Google) -
11:00
How to support multiple backends for guest_memfd? 30m
To improve performance of CoCo VMs, there's active work on huge pages for guest_memfd. The first "backend" in the works for providing huge pages is HugeTLB.
Frank summarized interest in having devdax/ZONE_DEVICE memory as another backend [1] for guest_memfd.
If guest_memfd is to be the guest memory provider of KVM, it has to support (almost) any memory backend that can be configured in a memslot today through VMAs, going beyond memory backends to other things can be mmap()-ed, like perhaps MMIO though VFIO.
I would like to get ideas and feedback on how guest_memfd should be extended to support different backends.
Are flags what we really want to go with? Will we end up supporting many flags? Using HugeTLB as an example, flags/fields can encode use of HugeTLB, HugeTLB page sizes, but how would the HugeTLB mount be encoded? Should there be some way to encode size limits of the mount? Is that a required feature?
Could we perhaps reuse existing interfaces to communicate the memory config to guest_memfd? What if we hand fds (opened on specific mounts) where guest_memfd can grab config from? Should guest_memfd perhaps use pre-allocated memory from the fd?
[1] https://lore.kernel.org/all/CAPTztWajm_JLpp9BjRcX=h72r25ELrXeGkOXVachybBxLJGS=g@mail.gmail.com/
Speaker: Ackerley Tng -
11:30
Coffee Break 30m
-
12:00
In-Kernel Caretaker for Orphaned VMs 30m
As cloud infrastructure continues to push toward zero-downtime host maintenance, extending the capabilities of kexec-based Live Update to minimize guest disruption is becoming increasingly critical. This proposal introduces the architectural concept of an "Orphaned VM"—a virtual machine that actively executes guest instructions on isolated physical hardware while completely decoupled from a host operating system or a VMM during a host reboot. This execution asymmetry is achieved by introducing the "Caretaker," a specialized bare-metal primitive interpose layer that runs within the hypervisor's privileged hardware state to trap and resolve some VM exits locally while the host kernel is offline. This session will explore the mechanisms required to orchestrate this transition using the LUO, focusing on how vcpufd structures are preserved across the reboot gap, how physical CPUs are shielded from boot-time reset signals, and how asynchronous hardware challenges like timekeeping drift, stray interrupts, and guest-to-guest IPI routing are mitigated during the management gap.
RFC Reference & Discussion Thread: For full architectural background and ongoing community feedback, see the original mailing list discussion at https://lore.kernel.org/all/afEwWZksU0Fw61oT@plex
Speaker: Pasha Tatashin -
12:30
Orphaned VMs: Which VM Exits can we handle? 30m
Orphaned VMs [0] is proposed to be the next evolution of hypervisor live update. It relies on a specialized micro-hypervisor, called the Caretaker, which handles VM Exits during the time between old kernel shutting down and new kernel booting up. The Caretaker reduces the downtime observed by the VM during live update by servicing VM Exits during this transition period.
The Caretaker is not a full hypervisor or VMM, so it cannot service all the VM Exits. If it cannot handle a VM Exit, it must pause the VM and wait for the next kernel to take over. The success of the Caretaker in reducing observed VM downtime relies on which VM Exits it can and can't service, and how frequently they occur.
In this talk, we will present the types and frequency of VM Exits observed under some sample workloads. We will also categorize the VM Exits into ones which the Caretaker can handle and the ones which it can't.
By combining the frequency of VM Exits and whether they can be handled, we can estimate observed VM downtime, and decide how effective Orhpaned VMs can be.
[0] https://lore.kernel.org/all/afEwWZksU0Fw61oT@plex/
Speakers: Pratyush Yadav, Tarun Sahu (Google) -
13:00
KVM and Live Update 30m
During this session, the maintainers of KVM welcome discussion around other topics related to live update and the Caretaker that did not fit in the previous hour, including:
* guest_memfd live update
* stable ABI for live update
* reusing the Caretaker concept for pKVMThe above list is just a suggestion, as we will take other topics from the audience if there are any.
Speakers: Paolo Bonzini (Red Hat, Inc.), Sean Christopherson (Google)
-
10:00
-
10:00
→
18:30
Kernel Summit Track "South Hall 1 B" (Prague Congress Centre)
"South Hall 1 B"
Prague Congress Centre
158-
10:00
Kernel CVEs at AWS Scale: Two Years of Empirical Findings 45m
Since the Linux kernel project became a CVE Numbering Authority (CNA) in February 2024, organizations maintaining custom kernels have faced a flood of CVE disclosures.
We present empirical findings, lessons learned, and the key challenges that the kernel team responsible for Amazon Web Services fleet infrastructure encountered over the following two years while handling the steady flow of kernel CVEs.
We summarize our approach to assessing kernel CVEs in a warehouse-scale environment, pruning our code base to reduce the attack surface, along with quantitative results from field data. Specifically, we examine
- how NVD/CVSS scores and third-party assessments correlate with our final in-house ratings;
- how community guidance, including stable tree tracking and domain-specific risk assessment, measurably improved our development velocity, kernel maintenance, and kernel adoption strategy; and
- how backporting practices introduce regressions and compound risk.
We then present data on patching velocity and CVE-related trends across LTS versions. Although our data comes from a single environment, we believe these findings generalize to any organization that maintains in-house or custom kernels.
We will close with an open discussion on emerging developments (including where AI-assisted tooling fits into the CVE triage and patching pipeline) and invite attendees to challenge, extend, or contradict our findings with their own data.
Speakers: Mr Dylan Johnson (Amazon Web Services), Mr Justinien Bouron (Amazon Web Services) -
10:45
Regressions & tracking them: current state, plans, and what do you want? 45m
Provide a quick "state of the union" about the Linux kernel regression ecosystem in the first part of the session before spending the second discussing what improvement the members of the audience wish for in this area.
The first part is meant to take less then half of the allotted time and will cover things like:
- KernelCI and regzbot joined forces -- what this means for the future.
- Tracking regressions with regzbot, the regression tracking bot: status and future plans.
- Linux development workflow patters that lead to regressions or delay resolving them.
The second half will be audience driven and might discuss details in the areas raised earlier, things like the following, or whatever the audience is interested it:
- Regression tracking with regzbot still useful in the age of AI?
- How to improve interaction with parties interested in the space of CI, regressions, and tracking them – like kernel maintainers of distros that regularly update their kernels to latest stable or longterm series.
- Is there interest in trees with "pending" and "wip" regression fixes?
In case I'm invited to the kernel maintainers summit I'll also collect topics with regards to regressions the audience wants me to bring up at kernel maintainer summit a few days later.
Speaker: Thorsten Leemhuis -
11:30
Coffee Break 30m
-
12:00
Challenges in multi-tenant GPU sharing 45m
We would like to share some of the challenges we have encountered in a multi-tenant GPU environment, discuss the solutions we have explored, and gather feedback from the community.
To provide context for the audience, we will start by giving an overview of our multi-tenant architecture. We then plan to discuss several areas where we have encountered challenges, including:
- GPU reset and debugging
- Identifying the problematic tenant session among many concurrent sessions
- Missing resource management capabilities, especially around memory cgroups
- Lock contention in drivers
- Power distribution across different system components
We would like to close the session with an open discussion on the issues and solutions presented. If the right audience is present, we would also like to discuss what a reasonable roadmap could look like for improving Linux support for multi-tenant GPU platforms.
Speakers: Boqun Feng, Gregoire Pean, Jatin Kataria -
12:45
Maintainer-oriented features of b4 45m
B4 is already a well-known tool for retrieving and applying series from lore.kernel.org. Recently, it gained several new features that can make the life of a maintainer a bit easier:
- b4 review - helps with code review and full series lifecycle management
- b4 bugs - integrates distributed bug tracking into your workflow
This session will go over the new features and how they can assist overloaded maintainers in keeping a handle on the stream of incoming changes.
Speaker: Konstantin Ryabitsev (The Linux Foundation) -
13:30
Lunch Break 1h 30m
-
15:00
Rust for Linux 45m
Rust for Linux is the project adding support for the Rust language to the Linux kernel. This talk will give a high-level overview of the status and the latest news around Rust in the kernel since LPC 2025.
Speaker: Miguel Ojeda -
15:45
Leveraging Rust's Field Projections in the Kernel and Beyond 45m
Several Rust contributors and I have been collaborating on a novel language feature called "Field Projections". The feature is still being designed and implemented, so now is a good opportunity to experiment with it to see what new APIs it unlocks for the Kernel and other projects. We also want to investigate any gaps in our current design that prevent important use-cases from being supported.
Rust makes heavy use of custom pointers filling the gap (& going beyond) between references (
&Tand&mut T) and raw pointers (*const Tand*mut T). Currently, they have two major issues: ergonomics and feature parity with builtin references andBox<T>. Raw pointers and their extensions (e.g.NonNull<T>) have especially bad ergonomics. Our Field Projection language feature aims to remedy these shortcomings of custom (dumb & smart) pointers; our current approach is an ambitious generalization ofDerefthat supports a plethora of custom pointers fromMyBox<T>(that has all the properties ofBox<T>) toVolatilePtr<T>(which is only using{read,write}_volatileto access the pointee). Our proposal integrates tightly with existing features such as place expressions, operations on places, the borrow-checker, and autoref.Speaker: Benno Lossin -
16:30
Coffee Break 30m
-
17:00
DRM: handling runtime requirements for device components. 45m
The DRM susbsytem grew up from the desktop GPUs, where the device is a single unit, powered on and off only at the important runtime points. For the embedded display controllers it's no longer true. The display pipeline can consist of several different devices, each having its own runtime power up and down code points. Handling device power on and off in the existing atomic callbacks makes the code fragile: it's too easy to create a disbalance of calls or to miss a register access from one of the points.
In this talk I would like to point out these issues and trigger a discussion about possible ways to solve the issue.
Speaker: Mr Dmitry Baryshkov (Qualcomm) -
17:45
Bringing Linux DRM Display Panel support in the modern age 45m
Since the introduction of the first Samsung DSI panel, the Linux DRM panel API has been a crucial piece of software for enabling displays across diverse architectures, but it has not evolved alongside modern graphics stacks. Currently, the API lacks atomic DRM API support and the ability to adapt power setups during mode changes. Furthermore, it fails to support advanced Display Driver IC (DDIC) features that modern hardware heavily relies on, including:
- Standby and advanced power states
- Advanced color management
- Dynamic rate switching
- Command mode self-refreshThis lack of evolution has led to severe fragmentation between upstream and vendor downstream trees for advanced devices support, creating a heavy maintenance burden and making native hardware support incredibly difficult.
The goal would be to outline these architectural limitations and trigger a discussion on how to collaboratively modernize the panel API. By standardizing advanced DDIC capabilities and fully embracing the atomic DRM API, we hope to establish a unified path forward for the entire Linux community.
Speaker: Neil Armstrong (Linaro)
-
10:00
-
10:00
→
19:45
LPC Refereed Track "Small Hall" (Prague Congress Centre)
"Small Hall"
Prague Congress Centre
215-
10:00
RCU Pseudo-Transactions: Bridging the gap between RCU and STM 45m
RCU data structures are notoriously complex to design mainly due to the
need to carefully manage how mutations are made observable to concurrent
readers.As a general solution to this problem, I am proposing a novel
transaction-based synchronisation mechanism: "RCU Pseudo-Transactions"
(rcu_txn).It applies both to userspace and kernel. It allows publishing complex
data structure mutations atomically to RCU readers, and synchronizing
updates from multiple writers, with minimal overhead on the read-side
(low-bit pointer tag check, predicted branch on rcu_dereference), and no
size overhead on the data structure nodes. It is composable: an object
can belong to multiple transaction-aware data structures, and
transactions allow it to become visible (or hidden) atomically.Speaker: Mathieu Desnoyers (EfficiOS Inc.) -
10:45
Checkpoint/Restore for RDMA: Saving and Restoring Live Queue Pairs with CRIU 45m
RDMA delivers high-throughput, low-latency networking by bypassing the kernel and letting applications communicate directly with the hardware. CRIU, by contrast, works by freezing running processes and serializing their state so they can be restored later. Bringing the two together is difficult precisely because of what makes RDMA fast: RDMA bypasses the kernel abstractions CRIU would normally use to checkpoint. CRIU already migrates live TCP connections but relies on filesystem attributes and a hook in the socket interface. RDMA is conceptually similar but the interfaces required to save/restore look very different. Our work aims to close this gap with an approach we demonstrate on NVIDIA ConnectX and BlueField devices using existing SRIOV VF migration support, as well as on RXE/Soft-RoCE. In both cases the RDMA connection survives checkpoint/restore intact: when peers are checkpointed together, the connection is never torn down and neither side sees QP errors or forced reconnects.
This matters most for machine learning, where RDMA carries communication for large distributed training and inference jobs that are expensive to start and stop. The ability to checkpoint and restore these jobs enables defragmenting a cluster to improve utilization, recovering seamlessly from hardware failures, time-sharing expensive resources between seasonal workloads (for example, inference by day and training by night), and migrating jobs to cheaper resources as availability changes. It also standardizes the save/restore workflow across frameworks, simplifying resource management for infra owners. Crucially, these jobs typically run on bare metal, so virtual machine live migration—the main existing alternative—will not be adopted by many would-be users.
To get there, the talk will first propose the concrete kernel interfaces required to support checkpoint/restore for RDMA, and explain how our implementation in CRIU uses them to checkpoint and restore a connection. Then we will then turn to mlx5, which has supported live migrating RDMA connections inside of QEMU VMs for some time. By reusing the same device capabilities and firmware APIs that already power SRIOV VM live migration, we show how realistic machine learning workloads can be checkpointed and restored on Linux using networking hardware available today.
Speaker: Raphael Norwitz (nvidia) -
11:30
Coffee Break 30m
-
12:00
Improving kernel test coverage using stress-ng 45m
The Linux kernel is constantly growing and evolving; unfortunately, corner-case regressions can creep into code in every release. Gcov test coverage can find infrequently used code paths that may contain issues. This presentation discusses how such techniques are used to improve kernel testing with stress-ng and the challenges in reaching full test coverage.
Speaker: Colin King (stress-ng) -
12:45
Rust SPDM in the Kernel 45m
Security Protocols and Data Models (SPDM) is used for authentication, attestation and key exchange. SPDM is generally used over a range of transports, such as PCIe, MCTP/SMBus/I3C, ATA, SCSI, NVMe or TCP.
From the kernels perspective SPDM is used to authenticate and attest devices. In this threat model a device is considered untrusted until it can be verified by the kernel and userspace using SPDM. As such SPDM data is untrusted data that is possibly from a mallicious device. The SPDM specification is also complex, with the 1.2.1 spec being almost 200 pages and the 1.3.0 spec being almost 250 pages long.
As such we have the kernel parsing untrusted responses from a complex specification, which sounds like a possible exploit vector. This is the type of place where Rust excels!
Over the last few years there has been gradual momentum building for SPDM support in the kernel and an implementation written in Rust. This implementation is in charge of authenticating and attesting untrusted and potentially malicious devices in the kernel using Rust code. The kernel also needs to allow userspace to apply security policies and allow remote verifiers to verify the running system, even with a possible malicious kernel.
This talk is going to cover the current status of the SPDM Rust implementation, how and why we got here and then discuss next steps for getting it merged into mainline.
It's not even over once SPDM is supported in the kernel though, as there are a range of more complex features that need to be supported. We can also talk about future features and what they might look like, ensuring we don't step on any PCI TSM feet.
Speaker: Alistair Francis -
13:30
Lunch Break 1h 30m
-
15:00
Virtio-GPU for Automotive: Implementing Libkrun + Vhost-User. 45m
Automotive hardware architectures are consolidating standalone Electronic Control Units (ECUs) into centralized compute platforms. A major challenge in this architecture is safely and efficiently sharing a single GPU across multiple isolated virtual machines. For example, systems must run critical instrument clusters, infotainment setups, and ADAS pipelines simultaneously without risking cross-domain interference.
This presentation tackles this challenge by introducing an architecture that combines libkrun, a process-based KVM virtualization library, with the vhost-user protocol to split device emulation into separate processes. By executing the virtio-gpu backend independently via virglrenderer, this approach achieves fault isolation, zero-copy transfers, and a reduced attack surface compared with traditional Type-1 hypervisors. We have implemented headless GPU compute acceleration, verified using AMD Radeon graphics via virgl, allowing offscreen rendering and ADAS sensor preprocessing. We have also implemented software scanout display output using the gfxstream backend, with DMABUF zero-copy scanout for virglrenderer in progress.
However, productizing this architecture has revealed concrete specification gaps, where the virtio-gpu specification and Linux kernel implementation diverge. During our ongoing implementation of display output paths, such as UPDATE and cursor paths, we encountered some blockers where the written specification and the kernel driver handle headless state and display configurations differently.
This talk focuses on two concrete topics from our implementation experience:
- Spec vs. Kernel Reality on Headless Operation: The Virtio-GPU Spec (v1.4 §5.7.4) requires a minimum of 1 scanout, yet the Linux kernel (virtgpu_kms.c) gracefully accepts and handles 0. We will discuss how to reconcile the specification to natively support headless, compute-only automotive workloads without forcing VMMs to waste resources on dummy display allocations.
- Display Output & Device Infrastructure in libkrun: We will present the two GPU display scanout paths we are implementing: Software scanout via gfxstream (pixel copy in message payload) and DMABUF zero-copy scanout via virglrenderer and also covering their tradeoffs in latency, memory usage, and backend compatibility. We will discuss how implementing GPU display support required adding generic SHMEM region mapping and BACKEND_REQ protocol features to libkrun's vhost-user framework, and how this infrastructure then enabled support for other vhost-user devices like virtio-media (camera/decoder passthrough) with minimal additional effort.
Eventually, we want to engage kernel maintainers, virtio specification editors, and VMM developers to discuss how the specification and its implementation can be improved to build a way forward for automotive virtualized graphics.
Session Timeline & Core Discussion Points (45 Minutes)
- Architecture & Status (10 mins): High-level overview of the libkrun + vhost-user-gpu stack, where it stands relative to Type-1 hypervisors, and a status update on PR #717 (working headless compute vs. pending display paths).
- Spec vs. Kernel: num_scanouts Divergence (20 mins): Open discussion on the conflict where the kernel accepts num_scanouts == 0 while the spec forbids it. Questions: Should the spec be amended to support headless operation? Should the kernel enforce spec compliance? When spec and implementation conflict, which is authoritative? What are the implications for VMM developers and automotive use cases?
- Display Output & Device Infrastructure in libkrun (10 mins): The two GPU scanout paths: Software scanout (gfxstream, full-frame pixel copy) vs DMABUF zero-copy (virglrenderer, FD passing via SCM_RIGHTS), their tradeoffs and current status. How implementing GPU display required adding generic SHMEM region and BACKEND_REQ protocol support to libkrun, which then enabled vhost-user support for virtio-media with minimal additional work.
- Q&A and Upstream Planning (5 mins)
Speaker: Dorinda Bassey (Red Hat) -
15:45
The State of the Kconfig Ecosystem 45m
Part 1: Tooling
Researchers continue to be fascinated by the Linux kernel’s usage of Kconfig, and academic papers have been regularly published on it for almost 20 years now. And with these papers often comes tools. Some examples include detecting dead configuration options, unmet dependency bugs, and generating config files that compile affected lines of C code from patches, among many others. We take a look at which tools are still being maintained post-publication, and discuss interesting tools that have since been abandoned. Can they be picked up by the open source community? And what kind of tooling is still missing?Part 2: The Many Implementations of Kconfig
Kconfig, a language originally introduced for use in the Linux build system, has grown to over 200,000 lines of usage in Linux itself, and has been adopted by many other open source projects, like coreboot, BusyBox, Zephyr, and more. However, Kconfig does not have a specification like other languages, and is implemented differently in each of these projects that uses it. We take a look at how these implementations differ, and discuss the feasibility of a unified spec and implementation.Speaker: Julian Braha -
16:30
Coffee Break 30m
-
17:00
Devicetree-ACPI hybrid mode 45m
Currently when booting in Devicetree mode the kernel will fully disable the ACPI subsystem. On WoA Snapdragon laptops where the factory Windows OS actually boots using the ACPI tables this is not necessarily desirable.
The purpose of this session is to present and discuss a proposal for a new DT-ACPI hybrid mode, in which while booting with Devicetree:
-
The ACPI tables are still parsed and ACPI fwnodes are made available for device-drivers to use for (extra) information.
-
Some devices may even be fully enumerated through ACPI e.g. enumerate I2C clients through ACPI for an I2C controller which itself is described in DT.
-
Going futher: use ACPI GPIO-IRQ event handlers + I2C opregion support to let ACPI handle a laptops embedded controller connected over I2C and using the ACPI battery device (backed by the EC) to expose battery state information in a laptop-model agnostic way like how laptop batteries are handled on x86 laptops.
Note on current laptops Linux cannot boot using ACPI due to some information missing from the ACPI tables. People are working on changing this so that for future WoA Snapdragon laptops Linux can boot using ACPI only without requiring Devicetree.
An early RFC patch-series implementing 1. + 2. has been posted upstream.
Speaker: Hans de Goede (Qualcomm) -
-
18:45
Celebration of Life: Dan Williams 1h
-
10:00
-
10:00
→
13:30
Tracing MC "Small Theatre" (Prague Congress Centre)
"Small Theatre"
Prague Congress Centre
105Description:
Visibility into the Linux kernel has always been critical for debugging and validating the execution of the code. The never ending challenge is to be able to trace the code without causing extra overhead, as tracing is most useful in a production environment.Possible topics for this year include:
- Updating the deferred stack tracer for sframes.
- A light weight lock stat tracer
- More read1ng of user space from syscall tracepoints
- Additions to the persistent ring buffer
- Adding error injection via trace points and kprobes
- Doing more with synthetic events
- Rewriting the histogram/trigger/synthetic event code
And much more
What has been done before
Here's the enhancements that were added to Linux tracing that were derived from the previous Tracing MC session:- libside has been released for better user space tracepoint hooking
- Faultable system call tracepoints
- We have a new Runtime Verification maintainer!
Key Attendees:
- Steven Rostedt
- Masami Hiramatsu
- Mathieu Desnoyers
- Ian Rogers
- Gabriele Monaco
- Namhyung Kim
- Arnaldo Carvalho de Melo
- Tomas Glozar
- Peter Zijlstra
- Jens Remus
-
10:00
Probe events with typecast using BTF 20m
BTF fetcharg has been introduced since v6.5 for fprobe and kprobe events for fetching function parameters by name, and now we intrdouced typecast feature for BTF. This typecast feature is not only casting type, but also, the series supports nested typecasts, container_of(), this_cpu_ptr(), and "current" task structure access. With these features, we can access more context local data from dynamic trace events.
I would like to show how you can use these features and discuss what features we can provide with BTF, what limitations we have now etc.Speaker: Mr Masami Hiramatsu (Google) -
10:22
Build ID + offset sampling in perf events 20m
Performance event sampling traditionally captures virtual addresses, resolving them to file offsets and symbols post-facto via mmap metadata. Conversely, BPF enables a direct approach: recording the file's build ID and offset within the sample itself. While virtual addresses are smaller inline, the required mmap event stream may lead to much larger raw data files.
In this talk, we examine this new sampling method and its implementation trade-offs. The goal is to gather feedback to land the pending patchset and explore hybrid sampling schemes that selectively combine both approaches to achieve the smallest possible footprint.
Speaker: Ian Rogers (Google) -
10:44
Developing Next Generation Perf Tools With Python 20m
Traditionally, the Linux perf tool has relied on a text user interface (TUI) based on libslang and embedded Python/Perl interpreters for scripting support. However, developing robust user interfaces in C is tedious and error-prone, and using perf itself as the interpreter integrates poorly with broader programming language ecosystems.
In this talk, we will describe how we refactored perf into a set of libraries, enabling it to be imported and used directly as a native Python module. Building on this foundation, we will showcase new tools featuring rich, interactive console UIs, culminating in a console-based flame graph analysis tool. This modular library approach dramatically accelerates tool development, paving the way for a more agile and extensible Linux performance tooling ecosystem.
Speakers: Ms Alice Rogers, Ian Rogers (Google) -
11:06
Bridging perf and ftrace: Recording perf in ftrace buffers and perf data to trace.dat conversion 20m
The tracefs system allows for fast flexible tracing of events. Not only is it designed for speed (around 100 nanoseconds per event), it also has a fully functional filtering system and a way to build on events via probes and synthetic events.
Allowing perf events to be injected into the tracefs ring buffer with these filters can allow for things like seeing how many cache misses happen between schedule switch events. Perf data could be recorded at any event and even filtered based on the value of the event fields.
There is a Proof of Concept patch set that does this but the interface is very poor. This session will be about the best way to create the interface to allow perf events to be injected into tracefs ring buffers.
The perf tool is best for statistical profiling and performance analysis, and ftrace/trace-cmd for event timeline visualization. While perf captures comprehensive performance data with minimal overhead, understanding temporal relationships and event sequences remains challenging through its native interfaces. In contrast, KernelShark provides powerful graphical insights but traditionally requires separate trace-cmd recordings. Since perf record already captures tracepoint events alongside other performance metrics, running multiple commands to obtain both statistical and timeline data is redundant in production or embedded environments.
This discussion will be about both moving perf events into trace.dat files as well as being able to inject perf events directly into the tracefs ring buffer.
Speakers: Madhavan Srinivasan, Steven Rostedt, Tanushree Shah -
11:30
Coffee Break 30m
-
12:00
Sharing trace infrastructure between in-tree and OOT tracers 20m
The current situation regarding LTTng vs upstream Linux:
1) There are maintainers who push for everything to be in tree
2) There are maintainers who are proponents for no-GPL-export when there are no in-tree users
3) Most of the tracer common facilities are used by tracers which do not compile as modules (only builtin)
4) Linus Torvalds stated that LTTng will stay out of treeAs a consequence, LTTng has no way to use common tracer facilities
without kernel patches.This context is not favorable for collaboration of LTTng developers with upstream. It is hard to justify spending time on kernel infrastructure collaboration when the resulting APIs cannot be used by out-of-tree tracers.
I am open to suggestions to improve this situation.
Speaker: Mathieu Desnoyers (EfficiOS Inc.) -
12:22
Lightweight Preempt and IRQ Tracepoints for Production RT Kernels 20m
Debugging latency issues in production RT kernels often requires
understanding where and why preemption and interrupts are being
disabled. Today, enabling thepreempt_disable/enableand
irq_disable/enabletracepoints requires pulling in heavyweight
infrastructure — either the preemptoff/irqsoff latency tracers or the
full lockdep IRQ tracking — that carries too much overhead for
production deployments. This forces engineers to reproduce customer
issues in lab environments, which is often impractical for intermittent
latency problems.This talk presents a patch series that introduces
CONFIG_TRACE_PREEMPT_TOGGLEandCONFIG_TRACE_IRQFLAGS_TOGGLE, two
new user-selectable kernel options that enable these tracepoints
independently, with minimal overhead. The tracepoints are gated behind
static keys.Speaker: Mr Wander Costa (Red Hat) -
12:44
Handling stack traces in tracing 20m
Stack tracing of events can be very useful, for both kernel stack tracing as well as user space stack tracing. Stack traces can fill the buffer quickly with many duplicate stacks. Having a way to consolidate them would make it possible to store even more data. There's been efforts to do this but how to implement it and the interface is still an ongoing subject.
On top of that, tracing user space can have the same issue. With the deferred stack trace, it is now possible to do more when taking the stack trace (like reading the vma and finding what files are associated with the stack as well as reporting the actual file offset instead of the virtual memory of the task). Work for this has been done as well but it also has issues. Figuring out how to implement this in a way that everyone is satisfied would be a goal of this topic.
Speaker: Steven Rostedt -
13:06
DEPT (DEPendency Tracker): A Path to Mainline and Solving the False Positive Challenge 20m
DEPT (DEPendency Tracker) is a runtime dependency tracking framework
that detects potential deadlocks by tracking wait/event relationships
rather than lock acquisition order. Unlike lockdep, which is limited to
typical locking primitives, DEPT can detect deadlocks involving general
synchronization mechanisms such as folio locks, completions, DMA fences,
and other wait/event-based patterns.This talk presents the current state of DEPT and addresses the critical
challenge blocking its mainline inclusion: false positive handling
through subsystem annotations.Any runtime dependency tracking tool faces the inherent challenge of
distinguishing real deadlocks from intentionally ordered synchronization
patterns. Lockdep also encountered this problem at its inception and
solved it through extensive subsystem-specific annotations (lock classes,
subclasses, etc.). DEPT currently lacks sufficient annotation for other
than typical locking primitives because it just started.This talk proposes a collaborative path forward: (1) core framework
stabilization to finalize annotation APIs, (2) subsystem pilot programs
with maintainers if any (3) gradual mainline inclusion enabled
incrementally as annotations mature.I seek community discussion on: annotation design that minimizes
maintainer burden, default behavior when annotations are missing,
testing infrastructure, and coexistence strategy with lockdep. Like
lockdep before it, achieving practical utility requires extensive
subsystem-specific annotations. This talk charts the collaborative path
to get there.Speakers: Byungchul Park, Yunseong Kim (Ericsson Software Technology)
-
10:00
→
18:30
eBPF Track "South Hall 1 A" (Prague Congress Centre)
"South Hall 1 A"
Prague Congress Centre
158The eBPF Track is going to bring together developers, maintainers, and other contributors from all around the globe to discuss improvements to the Linux kernel’s eBPF subsystem and its surrounding user space ecosystem such as libraries, loaders, compiler backends, related system tooling as well as eBPF use cases.
The gathering is designed to foster collaboration and face to face discussion of ongoing development topics as well as to encourage bringing new ideas into the development community for the advancement of the eBPF subsystem.
The track will be composed of talks, 30 minutes in length (including Q&A discussion).
eBPF Track's technical committee: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko
-
10:00
eBPF Research: What's Going On In Academia? 30m
The body of academic work on eBPF is growing so large and scattered that it’s hard to see the forest for the trees. Papers span all kinds of topics and conferences, vary wildly in quality, and are often dense and hard to parse.
This talk will present the dominant trends of research. To that end, we will first explain some heuristics to identify "high-quality papers"—it is not the number of citations!—and what it means for a paper to be considered high-quality. We’ll then map these standout papers to some of the teams behind them.
Finally, we will see that only specific research topics are leading to contributions upstream and explain why that might be. We will discuss how the kernel community could encourage research on specific problems, should it choose to do so.
Speaker: Paul Chaignon (Isovalent) -
10:30
bpf_fault: Custom Page Fault Handling with eBPF 30m
Page faults, which occur when a program accesses a virtual memory page that is not mapped to physical memory, are traditionally handled by the operating system. However, many applications benefit from running custom page fault handling logic. For example, some applications may seek to prefill newly-faulted pages with content, or intercept writes in order to make a copy of the original contents. Linux’s userfaultfd interface enables some of these use cases by offloading fault handling to userspace. Unfortunately, it suffers from significant limitations, namely high overhead, poor scalability, and a design that precludes its use in libraries.
We present bpf_fault, an experimental framework that allows applications to run page fault handlers directly within the Linux kernel using eBPF. By eliminating the overheads and complexity associated with userfaultfd, bpf_fault reduces fault latency by 2.8-6.1x and eliminates userfaultfd’s scalability bottleneck. We integrate bpf_fault with several applications, including VM live snapshots in Firecracker and QEMU (eliminating tail latency spikes caused by snapshots), JVM garbage collection, and more. We also design a novel lazy dynamic linking mechanism for Linux that defers relocations to fault time using bpf_fault, a use case impossible with userfaultfd, which reduces dirty memory usage of widely-used applications like Chrome and Clang by up to 50%. Together, these results demonstrate that eBPF-based fault handling can improve performance and reduce resource usage in widely-deployed applications.
Speaker: Tal Zussman (Columbia University) -
11:00
Programming Across the VM Boundary with eBPF: Lessons from Revisiting Phantom Tracker 30m
eBPF programs can exchange data efficiently with other programs and userspace through maps, ring buffers, and kfuncs, but these mechanisms stop at the boundary of a kernel instance. The Linux kernel has no generic, safe, low-latency mechanism for eBPF programs in a guest and host to communicate through shared memory across the VM boundary. Roy’s Google Summer of Code (GSoC) project [1] reimplements Himadri’s in-kernel thesis [2] prototype of Phantom Tracker using eBPF and, in turn, highlights the need for paravirtualized shared-memory drivers that expose reliable communication APIs to eBPF programs across the VM boundary.
The thesis models vCPU lifecycle states using the distinction between phantom and viable vCPUs. A phantom vCPU satisfies two conditions: (1) it is runnable but waiting in the run queue of a pCPU on the host, and (2) a worker thread of the guest parallel application that had been running on that vCPU is now stalled because the vCPU is not executing. Conversely, any vCPU that does not satisfy both conditions is considered viable. Across the VM boundary, Phantom Tracker correlates guest-side information about which vCPUs are running worker threads of the parallel application with host-side scheduler wake-up and context-switch events involving those vCPUs. The GSoC project uses a configurable eBPF timer that aggregates these observations and computes a per-VM metric called the phantom average. Guest userspace parallel runtime libraries, such as libgomp, can use this metric to adapt the application’s degree of parallelism at runtime and minimize the number of phantom vCPUs.
Communication between the host and guest is implemented using QEMU’s Inter-VM Shared Memory (IVSHMEM) device [3]. While the thesis prototype relied on custom IVSHMEM drivers tied to custom kernels, the eBPF implementation resulted in the development of new eBPF-compatible drivers [4]. By presenting the lessons learned while working on this GSoC project, this talk discusses the broader applicability of our eBPF-compatible IVSHMEM drivers within the QEMU/KVM virtualization stack and invites discussion on the scope for upstreaming them.
[1] https://summerofcode.withgoogle.com/programs/2026/projects/iXbp42du
[2] https://inria.hal.science/tel-05438117v3
[3] https://www.qemu.org/docs/master/system/devices/ivshmem.html
[4] https://github.com/himadrics/phantom-tracker/tree/main/pvsched-shmemSpeaker: Roy Nchang (National Advanced School of Engineering (ENSPY), Yaoundé) -
11:30
Coffee Break 30m
-
12:00
shirudo: BPF live-patching infra for the agentic era 30m
In times of Fable/Mythos or equivalent LLMs, security fixes and attack-surface hardening increasingly needs to land on production systems now, but data-center fleets, Kubernetes nodes, or embedded/air-gapped devices all typically share long patch-and-reboot cycles.
BPF is the natural vehicle for on-the-fly live mitigations and runtime visibility - to the kernel itself as well as to userspace apps - and in the age of AI-assisted engineering the natural author of those patches is an agent working in a close loop with the operator. We'll present shirudo, which is an agentless BPF-based security platform tailored for exactly this: There is deliberately no config DSL, because agents work far better with code directly. shirudo also fully embraces xattrs, signed BPF and seals all its assets via BPF LSM in order to defend against untrusted root tampering with bpf. In this talk we walk through the operator/target node workflow, architecture internals, demo its capabilities, and discuss gaps and next steps on shirudo, BPF kernel and libbpf loader side.
Speakers: Daniel Borkmann (Isovalent), John Fastabend (Isovalent) -
12:30
One Layer of the Onion: A Daemonless eBPF LSM for Confidential VMs in a WhatsApp TEE 30m
Host-side hardening for a confidential VM is not one control, it's an onion. AMD SEV protects guest memory, MetalOS and a measured/verified boot chain establish the platform, signing and provisioning gate what lands on disk, and process isolation constrains the runtime. This talk is about one specific layer of that onion, the eBPF LSM that enforces binary identity and process protection at runtime. WhatsApp runs user workloads inside AMD SEV-backed TEEs where the guest is a QEMU process; eBPF LSM is how we add a defense-in-depth layer for that host without becoming a new single point of trust.
The core of the talk is the BPF and the goal is to give attendees a full picture of how we use BPF at Meta to supplement confidential workloads in WhatsApp. Specifically, we attach a set of LSM programs bprm_creds_from_file for exec-time identity, ptrace_access_check/ptrace_traceme for anti-trace, and task_kill for anti-signal and drive them entirely from BPF maps keyed by role. The design choice we want to dig into is persistence through pinning that has been discussed in other talks from Tetragon discussed in this lwn article and fully explore how we pin, configure our maps, set up keychains used for binary identification and how we track processes. We will also talk about the need for options for logging as is done in the initial article, but it will be a small portion to ask for opinions and discuss some other options that we are exploring, as opposed to standing exclusively behind UDP packet sending that has been discussed previously. We will mention potential malware opportunities from this approach. Our method will be contrasted with pros and cons of a typical resident daemon. In the use case at Meta, a run-to-completion init binary loads the programs, pins every program, link, and map into bpffs, and exits. Because the LSM links survive pivot root, enforcement is live before the confidential workload starts. We'll walk through the map layout, how policy is expressed per-role, and how pinned state lets the enforcement layer survive independently of any userspace processes. We will also discuss some of the strategies we use to disallow malicious userspace processes from removing these programs once they are in place and how we manage the potential fallout from this.
On identity, we'll show the in-kernel verification path in detail. At exec, the bprm_creds_from_file hook reads an extended attribute on the binary user.bpfj.policy.exec naming its role, which forces calls to bpf_get_fsverity_digest and bpf_verify_pkcs7_signature to check the binary's fs-verity digest and detached PKCS7 signature against a kernel keyring seeded from that role's certificates enrolling the process into a role only if its on-disk contents are signed for that role. Here the BPF layer leans on fs-verity, the keyring subsystem, and our custom signing pipeline do the heavy cryptographic lifting. Our programs make the runtime authorization decision from that verified identity. We'll cover the kfuncs we depend on, the sleepable-LSM constraints, and the pitfalls we eliminated by moving to signature-gated, exec-time enrollment.
Finally, we show what verified identity buys at runtime with QEMU pinned to a verified role, the ptrace and task_kill hooks enforce per-role allow-lists so nothing on the host can attach a debugger to, or signal, the confidential VM closing common paths for extracting or faulting guest state. We'll be explicit about the layer's limits what it does not defend against and which sibling controls cover those gaps — so the audience sees where a focused eBPF LSM fits in a real confidential-computing threat model.Speakers: Mr Joshua Lilly (Meta), Liam Wisehart -
13:00
Scaling BPF LSM hooks and error injection across the kernel 30m
In the modern era, Linux kernel CVEs might accumulate faster than fleets can reboot into patched kernels. In some cases, this takes not even days. BPF-based mitigations can block vulnerable code paths at runtime, no reboot needed.
However, at the moment, BPF is far from being a golden bullet. Two mechanisms on how BPF can alter an execution path, LSM Hooks and error injection, are naturally limited: by the set of existing LSM hooks and by the [short] whitelist of ALLOW_ERROR_INJECTION functions. Many subsystems, such as different parts of net/, device drivers, etc., have no BPF security coverage.
In the first part of the talk, we investigate how subsystems currently lacking BPF hook coverage can be equipped with it and present tooling and guidelines to support adding that coverage more systematically.
In the second part, we discuss the error injection topic. One recent radical attempt, killswitch, allows any function to be altered. While this ultimately solves the problem, this is not really a solution which can be kept under control. Thus we discuss what might be done to substantially extend the set of functions eligible for error injection, while keeping the mechanism firmly under control.
Speaker: Anton Protopopov (Isovalent at Cisco) -
13:30
Lunch Break 1h 30m
-
15:00
Application-Layer Parsing in eBPF 30m
Recent advances in the Linux kernel have enabled increasingly complex kernel offloads with eBPF. Despite this, parsing application-layer protocols, e.g. HTTP, remains a challenge. The reason for this is the self-describing structure of such protocols, which typically requires more state and more complex control flow than a transport-layer protocol. This is unfortunate because supporting L7 protocols in eBPF opens up new opportunities to optimize many applications that were previously deemed too complex, e.g. web servers, web application firewalls, or L7 service proxies.
This talk introduces Beeper, a novel architecture to parse HTTP directly in eBPF, without the need for kernel modules. To circumvent eBPF’s stringent limitations, it constructs Aho-Corasick-like DFAs in user space, and leverages them in kernel space to identify and extract relevant headers from the message buffer. We will discuss the challenges Beeper must address to parse HTTP/1.1 and HTTP/2, and show its advantage by serving HTTP requests directly from eBPF.
Speaker: Laurin Brandner (ETH Zürich) -
15:30
Inline DDoS protection for cloud-native game servers at PlayStation with eBPF/XDP 30m
At PlayStation, we see DDoS attacks of multiple terabits per second targeting game servers. Traditional DDoS mitigation systems can be costly, slow to react, and difficult to place close enough to the ingress points of the network.
This talk presents a token-based eBPF/XDP architecture for inline DDoS protection. We show how distributing short lived tokens to legitimate clients allows us to make the most of XDP’s position early in the stack for whitelisting and minimise wasted cycles on unwanted traffic. We also complement the overall architecture with an eBPF based agent on the K8s workers that i) removes the need to expose backend clusters directly to the Internet ii) forms a transparent overlay between the public-facing DDoS protection layer and private game-server clouds and iii) solves challenges in integrating with Kubernetes CNIs and cloud environments. In order to tackle the unique requirements of frequently changing routing state that needs to be globally distributed, we also propose a control plane architecture built on CNCF xDS that allows us to propagate tokens at high rates close to our network ingress points, constantly updating the eBPF maps that drive our routing decisions.
Finally, we share lessons from building and operating an eBPF based DDoS protection architecture at global scale. We discuss the challenges of integration with cloud providers and their CNIs, as well as return path optimisations that make the most of Playstation’s backbone network.
Speakers: Babis Stylianopoulos (Sony Interactive Entertainment), Jeffrey Barendse (Sony Interactive Entertainment) -
16:00
BPF ksock: bringing network sockets to BPF 30m
BPF-based agents aim to be as transparent as possible while minimizing CPU and memory overhead. Real-world experience from projects such as Cilium’s Tetragon has shown that moving more functionality directly into the kernel is an effective strategy. However, one remaining limitation for observability and logging is the lack of an API for sending data over the network directly from BPF programs, keeping user space in the critical path.
Efforts to provide BPF programs with networking capabilities through new kfuncs have been discussed at the past two LSF/MM/BPF summits in Montreal and Zagreb. An initial approach based on the netpoll infrastructure was proposed but ultimately rejected.
This talk will recap the motivation behind the current patch sets, summarize the discussions so far, and introduce the current design. We will then demonstrate several ways BPF programs can use the new API, ranging from basic examples to practical and more unexpected use cases. Finally, we will discuss possible future extensions to the API, such as support for additional BPF program types and TCP sockets.
Speakers: Kornilios Kourtis (Isovalent), Mahé Tardy (Isovalent) -
16:30
Coffee Break 30m
-
17:00
BPF-RBACd: Delegating BPF permissions with high granularity 30m
BPF usage has traditionally required system-wide capabilities, and giving an application access to using BPF is an all-or-nothing proposition. With BPF tokens, we gained the ability to delegate BPF capabilities to user namespaces with more granularity, and with an LSM we can increase granularity further.
Both BPF token usage, and an LSM, require a userspace implementation of the policy enforcement mechanism.
bpf-rbacd(the "BPF Role-Based Access Control daemon") is such an implementation, which runs as a system service and supports granting permissions to applications or containers on the system using either BPF token delegation or syscall proxying. A policy language restricts which subset of BPF an application is allowed to use, with high granularity, enforced through an LSM written in BPF.We are working on making
bpf-rbacda core part of the Fedora and RHEL distributions. In this talk we'll present the architecture ofbpf-rbacdand solicit feedback from the community on the design of the system, in the hope that this can prove useful to other distributions and operators as well.Speaker: Toke Høiland-Jørgensen (Red Hat) -
17:30
eBPF on Wheels - Automotive and Industrial Use Cases 30m
The industry increasingly adopts Linux for automotive and industrial use cases, since companies like Red Hat or Canonical engaged in development of Linux platforms in safety-critical domains. Using container technologies or just a bootable images on restricted silicon, now software operates cyber-physical processes. Those platforms are often cloud-connected because of maintainability, attracting new threat actors to that domain. Defenses in this area are limited, because of blind spots in endpoint protection solutions regarding electronics.
The missing integrations can be found in non-IP bus systems, such as the CAN bus as well as on-board interfaces to flash chips, FPGA or the protocols themselves like SPI or I2C. While less common in the IT world those missing capabilities are missing out in comprehensive observability resulting in unnoticed cyber attacks against IOT platforms.
This talk showcases the use of eBPF for protocols like CAN, DMA and SPI to create observability as well as defenses against cyber attacks in automotive or industrial contexts, such as secure updates, injection attacks and non-IP filters.
Speaker: Mr Reinhard Kugler (SBA Research) -
18:00
Pluggable Runtime Verification (RV) monitors with BPF 30m
RV is a lightweight method for verifying system behavior at runtime using, for instance, deterministic automata. Currently, RV monitors must be implemented in-kernel, meaning any new monitor requires going through the upstream kernel development process.
We can replicate the existing monitor infrastructure in BPF mapping kernel primitives to BPF equivalents such as maps and ring buffers, while reusing common logic where possible.
This allows to develop, test, and deploy domain-specific monitors entirely from userspace, with all the perks of the BPF tracing infrastructure.In the talk we will cover an implementation using BPF struct_ops for mostly seamless integration with the in-kernel RV framework and tools, BPF monitor lifecycle control via the rv command line tool (i.e. registration, activation, and tracing), and various tradeoffs to keep a similar experience between different monitor implementations.
Speaker: Gabriele Monaco (Red Hat Inc.)
-
10:00
-
15:00
→
18:30
Kernel Testing & Dependability MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128The Kernel Testing & Dependability Micro-Conference (a.k.a. Testing MC) focuses on advancing the current state of testing of the Linux kernel and its related infrastructure.
Building upon the momentum from previous years, the Testing MC's main purpose is to promote collaboration between all communities and individuals involved with kernel testing and dependability. We aim to create connections between people working on related projects across the wider ecosystem and foster their development. This should serve applications and products that require predictability and trust in the kernel.
We ask that all discussions focus on identified issues, aiming to find potential solutions, alternatives, and concrete next steps. The Testing MC is open to all topics related to testing and dependability on Linux, not necessarily limited to the kernel itself.
In particular, topics of interest for Linux Plumbers Conference 2026 include:
- KernelCI and related infrastructure: Maestro, kci-dev, dashboard and API improvements, KCIDB-ng, pull-mode lab support, and integration with Tuxmake, TuxRun, and related tooling
- Expanding production use of testing infrastructure and improving how developers consume, triage, and act on test results
- Improving interoperability between KUnit and kselftest, including UAPI testing, running kernelspace tests from userspace, and unified reporting workflows
- Continued evolution of KUnit itself, including better support for parameterized tests, improved tooling, and broader adoption throughout the kernel
- Improving kselftest and related frameworks, including output consistency, KTAP compliance, parser and tooling improvements, and better handling of skips, nesting, and other real-world test results
- Building, running, and testing in-kernel Rust code, including Rust doctests and other Rust-oriented test workflows
- Improving sanitizers and dynamic analysis tools, including KFENCE, KCSAN, KASAN, UBSAN, and related debugging infrastructure
- Using Clang and compiler-assisted features to improve test coverage, diagnostics, and reproducibility
- Consolidating toolchains, build environments, and reference setups to improve reproducibility, consistency, and quality control
- Targeted fuzzing of internal kernel functions and other techniques to extend coverage beyond traditional syscall fuzzing
- Patch-series fuzzing, regression detection during review, and other ways to shift testing and fuzzing earlier into the development cycle
- Kernel benchmarking, performance evaluation, and shared infrastructure for tracking and bisecting performance regressions
- Determining which test coverage infrastructures are most effective for kernel quality assurance, and how coverage should be measured
- Improving traceability between requirements, code, tests, results, and hardware or lab metadata
- Regression testing for safety and dependability, including prioritization of critical configurations, platforms, and test suites
- Identifying missing features needed to support assurance in safety-critical systems
- Moving toward more test-driven kernel release practices for both mainline and stable trees
- Exploring how SBOMs and related metadata contribute to kernel dependability and assurance
- Better ways to share, normalize, store, and analyze test results across projects, labs, and communities
- AI-assisted testing and review workflows, including patch triage, regression-risk estimation, test selection, test generation, bug localization, and evidence-based validation of LLM-assisted results
Things accomplished since LPC 2025:
- KernelCI reached the last KCIDB-ng milestone, moving KCIDB submission ingestion closer to the Django backend and decoupling the KCIDB schema from the dashboard database
- KernelCI improved project health with API and system-resource monitoring, and increased backend test coverage from about 40% to nearly 70%, including benchmark tests
- Progress was made on pull-mode lab support in Maestro, enabling labs behind firewalls or with different internal setups to participate more easily in KernelCI
- Tuxmake, TuxRun, and TuxSuite/LAVA-related tooling continued to be integrated more closely with KernelCI, with tuxmake, tuxrun, and tuxlava moved under the KernelCI GitHub namespace
- kci-dev continued to mature as a developer-facing CLI, with v0.1.9 and v0.1.10 adding packaging improvements, better regression comparison and issue triage workflows, improved validation and reporting ergonomics, and general workflow fixes
- Continued work on KUnit and kselftest integration was discussed on the linux-kselftest mailing list, including the KUnit UAPI testing framework series for running UAPI-oriented tests under KUnit
- KUnit tooling and parser follow-up also continued on the linux-kselftest mailing list, including fixes for nested test result handling and better parsing of skipped tests from kselftest output
- KTAP standardization work continued, with ongoing discussion around aligning kselftest and KUnit output formats and the KTAP v1 format now documented in the official kernel documentation
-
15:00
Testing MC Intro 5m
Welcome to the Kernel Testing & Dependability MC
Quick overview of the 2026 edition and usual Plumbers briefing
Speakers: Arisu Tachibana, Guillaume Tucker, Sasha Levin, Shuah Khan (The Linux Foundation) -
15:05
kci-dev: What Changed, What Works, and What Kernel Developers Still Need 20m
kci-dev was created as a standalone command-line tool that allows kernel developers and maintainers to interact directly with KernelCI. Since its initial releases, the project has grown beyond its original role as a thin client for triggering jobs and retrieving results.
The latest development cycle, including the v0.1.11 release, introduced a reusable Python library interface, direct submission of external build results to KCIDB, storage-related commands, broader result filtering, improved Maestro and Dashboard validation, machine-readable validation reports, distribution packaging workflows, and an experimental Model Context Protocol server for automation and AI-assisted tooling. Significant reliability work also addressed automated bisection, job watching, network timeouts, malformed output, configuration handling, and inconsistent Maestro behaviour.
These changes are moving kci-dev from a collection of CLI commands toward a common integration layer for kernel testing workflows. However, several important pieces are still missing before it can become a practical everyday tool for a broader group of kernel developers and maintainers.
Key gaps include first-class testing of email patch series through b4 and Patchwork, reproducible pre-submit test plans, reliable comparison of results between revisions, classification of regressions versus flaky or infrastructure failures, consistent interfaces across Maestro, KCIDB and the KernelCI Dashboard, richer lab and hardware metadata, caching and multi-tree reporting, and stable extension points for external tools.
This session will present the current state of kci-dev, distinguish mature functionality from experimental work, and discuss which missing workflows should be prioritised. The goal is to agree on interfaces, responsibilities and concrete contributions needed to make kci-dev a dependable bridge between kernel development workflows and KernelCI.
Speaker: Arisu Tachibana -
15:25
Kernel Builds Deserve an Identity: Portable Kernel Artifacts and Traceable Test Results 20m
Our CI systems build thousands of kernels a day, yet the kernel build itself has no canonical identity and no portable distribution format. Reproducing the exact binary behind a regression report is guesswork, and correlating KCIDB results back to a build is convention, not verification.
KBI (Kernel Bundle Image, Apache-2.0) is a concrete starting point for fixing this. It packages vmlinuz, modules, initrd, BTF, firmware, and config as a standard OCI image and derives a deterministic, independently recomputable build identity from the artifacts. Existing registries handle distribution and signing, which maps naturally onto pull-mode labs. Module and eBPF add-ons bind to a specific build identity, turning silent runtime mismatches in labs into early, explainable rejections.
Proposed discussion points:
-
What should define kernel build identity, and how does it interact with reproducible builds?
-
Could Tuxmake and KernelCI Maestro emit OCI kernel bundles, and could KCIDB adopt a verifiable build identity field?
-
How far should artifact metadata go toward a kernel SBOM for safety and traceability use cases?
-
Could it solve bare metal kernel CI testing with multikernel?
Speaker: Mr Cong Wang (Multikernel Technologies) -
-
15:45
Rethinking KUnit Configuration 25m
The
kunit.pytool currently spreads its configuration across three places:kunitconfigfiles, which contain Kconfig entries for the kernel being tested;qemu_configpython scripts, which configure architecture- and emulator-specific options; and command-line arguments, which specify what is being done (building, testing, parsing, etc.), and any options specific to the run (filters, build options, output configuration, etc.)This means that, in order to completely configure a test run, three different configuration sources need to be managed, all of which are in different formats, and one of which isn’t even a file.
If we can implement a combined configuration format, this will not only simplify configuring KUnit tests, but also allow enhancements to
kunit.pywhich manage configurations more explicitly. With first-class configurations,kunit.pycould — for example — support running multiple different configurations as part of a single execution.Would this be useful, and if so, how should we do it? In particular:
- Should we extend the Kconfig-based ‘kunitconfig’ files with additional configuration lines?
- Would a python-based configuration (more like the qemu_config files) be sufficiently simpler?
- How important is maintaining compatibility with existing configs (it shouldn’t be hard to do so)?
- Is doing multiple test runs from within kunit.py useful? (And, if not, are people doing this using other scripts/CI systems?)
- If so, how should results be handled? We can merge KTAP results (but it’s trickier to do so in a way which preserves the streaming nature of results).
- What else can’t be expressed in existing KUnit configurations which would be useful?
Speaker: David Gow -
16:10
Optimizing local kernel build testing 20m
For the past 15 years, Arnd has been using a custom kernel build setup for regression testing linux-next kernels. Initially set up purely for the maintenance of the SoC tree, this has also helped in a number of other initiatives for cross-tree code changes and produced thousands of bugfixes across every major subsystems of the kernel.
In this session, Arnd will discuss some of the lessons learned from this work, and how some of the same ideas can be applied to other test setups and CI systems.
Speaker: Arnd Bergmann (Linaro) -
16:30
Coffee Break 30m
-
17:00
Kselftest coverage with kcov 15m
We can use KCOV configuration options and a simple tool to measure coverage to find gaps in individual selftests. Let's discuss how and if it is helpful in improving and enhancing the existing tests strictly to identify the gaps in critical coverage.
Speaker: Shuah Khan (The Linux Foundation) -
17:15
KFuzzTest: Fuzzing Internal Kernel Functions, One Year On 15m
At LPC 2025 we introduced KFuzzTest, a framework for exposing stateless and low-state internal kernel functions, such as complex data parsers and the like, directly to a userspace fuzzer, reaching code that system-call fuzzers struggle to exercise. Developers define targets alongside their functions using a simple macro-based API, with constraints and type annotations compiled into dedicated ELF sections for automatic discovery. This follow-up reports on a year of progress and charts the path ahead.
Since last year, the patch series has matured through upstream review, with a follow-up revision that simplified the design and removed the dependency on syzkaller, lowering the barrier to adoption. With the framework stabilizing, we turn to a broader question: KFuzzTest's defining feature is a uniform, low-boilerplate interface for invoking internal functions, and that interface is useful well beyond a single fuzzing engine.
We will explore two directions. First, integration with KUnit: a fuzzing harness is naturally expressed as a specialized unit test, suggesting a guiding principle that if a function can be unit-tested, it can be fuzzed, and reusing existing test infrastructure rather than standing up new machinery.
Second, automated harness generation: writing fuzz harnesses is routine, mechanical work, and KFuzzTest's minimal interface is well suited to being driven by LLM-based agents, both to author targets and to drive fuzzing loops. We see offloading this boilerplate as a concrete value proposition.
This presentation will cover what changed over the past year, lessons from the review process, and these avenues for extending KFuzzTest's reach. We hope to use the session to gather community feedback on the most promising directions for upstreaming.
Speaker: Ethan Graham (Student at ETH Zurich) -
17:30
LLM-Assisted Fuzzing with Syzkaller 15m
Coverage-guided fuzzing has proven to be highly effective at discovering Linux kernel vulnerabilities. Since its appearance in 2016, syzkaller - especially through the automated syzbot platform - has reported over 14,000 findings to the public kernel mailing lists.
Despite the success, traditional fuzzing methods struggle to reach deep code paths within complex kernel subsystems, even when provided with exact descriptions of the Linux kernel interfaces and fine-grained coverage feedback. Generating programs to reach deeply nested Linux kernel locations often requires a nuanced semantic understanding of target source code - a bottleneck that historically demanded substantial manual engineering. Now, the ever increasing capabilities of LLMs offer us the opportunity to address this problem in a fully automated manner.
We will present an intelligent, context-aware bug discovery workflow that integrates LLM-driven seed program generation with coverage-guided fuzzing using Syzkaller. By leveraging coverage feedback from the fuzzing engine, our approach selectively targets weakly covered Linux kernel regions and automatically synthesizes context-relevant seed programs to reach them. These seeds are subsequently fed back into the fuzzer, enabling coverage expansion into previously unexplored code regions and facilitating bug discovery.
During the experiments, we have successfully increased the testing coverage of the Linux kernel from 30% to 35% on the syzbot platform, exercising over 70,000 previously uncovered basic blocks. In total, by injecting the discovered seeds into the syzkaller pipeline, we've already discovered 450+ new, reproducible kernel bugs, including numerous UAF, OOB accesses and memory leaks.
Benchmark testing of the seed program generation workflow has demonstrated that it can reach roughly 16% of previously never reached locations (a random sample of N=500). We are confident that further improvements and future LLM models will let us achieve even higher effectiveness.
Speakers: Aleksandr Nogikh (Google), Chuyang Wang -
17:50
Testing the kernel where the fleet runs: KernelCI in the cloud 20m
KernelCI provides continuous testing for the Linux kernel. Over the past year we extended testing into the cloud. AWS lab is now integrated upstream and visible on the KernelCI dashboard.
We will describe the architecture and how it maps onto KernelCI's pull-mode lab model so other cloud vendors can attach their own infrastructure. We will share the gotchas of testing on remote virtualized machines.
The cloud unlocks special test usecases from tiny, memory-pressured instances to machines with hundreds of cores and terabytes of RAM.
Additionally, EC2 exposes nested KVM on c8i instances, and we run kvm-unit-tests inside these guests to exercise the kernel's own virtualization stack.KernelCI already runs tests on AWS infrastructure. We will discuss next steps and improvements like performance testing.
Speakers: Mr Norbert Manthey (AWS), Stanislav Uschakow (AWS) -
18:10
Security Lessons from Auditing the Linux accel/ethosu Driver 20m
The Linux accel/ subsystem is still relatively young, yet it already exposes
complex userspace interfaces for machine learning accelerators. These drivers
process user-controlled command streams, DMA descriptors, tensors, and region
metadata, making input validation critical.This talk presents the results of a security audit of the accel/ethosu driver
that led to seven upstream fixes merged in 2026, with the fixes subsequently
backported to supported stable kernel series where appropriate. The issues
included out-of-bounds accesses, arithmetic overflows, incorrect index
validation, sentinel misuse, and insufficient validation of DMA-related
metadata.Rather than focusing on individual bug fixes, the presentation identifies the
recurring implementation patterns that produced these vulnerabilities and
discusses practical techniques for preventing similar issues in future
accelerator drivers. Topics include overflow-safe arithmetic, structured input
validation, region and descriptor verification, and opportunities for improving
fuzz testing of accelerator interfaces.The session concludes with discussion of possible common validation helpers,
security review guidelines for new accelerator drivers, and ways to improve the
long-term robustness of the Linux accel/ subsystem as ML hardware becomes
increasingly common.Speaker: Muhammad Bilal (Individual)
-
15:00
→
18:30
Live Update MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128Proposal
Live Update is a specialized reboot process where selected devices are kept operational and kernel state is preserved and recreated across a kexec. For devices, DMA and interrupts may continue during the reboot.
The primary use-case of Live Update is to enable hypervisor updates in cloud environments with minimal disruption to running virtual machines. During a Live Update, a VM can pause and its state is stored to memory while the hypervisor reboots. PCIe devices attached to those VMs (such as GPUs, NICs, and SSDs), are kept running during the Live Update. After the reboot, VMs are recreated and restored from memory, reattached to devices, and resumed. The disruption is limited to the time it takes to complete this entire process.
With Live Update infrastructure in place, other use-cases may emerge, like for example preserving the state of GPU doing LLM, freezing running containers with CRIU, and preserving large in-memory databases.
The Live Update and state persistence functionality touch on different parts of the kernel and this microconference aims to bring together people from different subsystems. Upstream support for Live Updates is still in its infancy and there are a lot of unsolved aspects that will benefit from direct communication.
Key problems that will be discussed:
Support for memfd/guest_memfd/hugetlb/tmpfs Preserving the state of VFIO, IOMMUFD, and IOMMU drivers. Preserving vCPUs and Orphaned Virtual Machines LUO systemd integration Integration of Live Update with PCI and Device Model Leveraging suspend/resume functionality for device state preservation Optimizing kernel shutdown and boot times.
Last year achievements:
Following “Memory persistence over kexec” BoF at LPC 2024 we we landed support for Kernel KHO, LUO, and memfd preservation.
Expanding the BoF to a full blown MC helped defining key data structures required for Live Update stability and isolating them into a dedicated kho/abi/ directory under include/linux.
Duing 2025 edition of Live Update MC we finalized the objectives and design for making KHO stateless and it’s now transitioned to a radix tree for memory preservation.
Key attendees:
- Alex Graf
- Alex Williamson
- Ben Herrenschmidt
- Bjorn Helgaas
- David Matlack
- David Rientjes
- David Woodhouse
- Evangelos Petrongonas
- Jason Gunthorpe
- Josh Hilke
- Luca Boccassi
- Michał Cłapiński
- Mike Rapoport
- Pasha Tatashin
- Pratyush Yadav
- Samiullah Khawaja
- Vipin Sharma
-
15:00
LUO support in systemd 15m
systemd v261 added native support for LUO. System services, user services and nspawn containers can preserve data across kexec in a simple and transparent manner, using the existing File Descriptor Store API.
This talk will explore what is implemented and how to use it, and what are the next steps.
Speaker: Mr Luca Boccassi (Microsoft) -
15:15
Resolving Complex FD Dependencies for Live Update 15m
Preserving subsystem state across a live update introduces a complex challenge: managing the dependencies of File Descriptors. Interconnected subsystems—such as VFIO and IOMMUFD with a lot of shared state, IOMMUFD's dependency on memory providers memfd/guest_memfd or KVMfd/guest_memfd dependency requires preservation ordering to guarantee state immutability and performance.
Currently, the burden of orchestrating this dependency falls on userspace (VMMs), which must navigate complex API contracts, which will become much more complex with circular dependencies between subsystems down the road.
This session talks about FD dependencies with examples and it proposes a unified, kernel-driven dependency resolution architecture that is handled by Live Update Orchestrator by introducing some extensions.
Looking forward to your feedback on this.
Speaker: Samiullah Khawaja -
15:30
Extending Memory State Preservation Across Warm Reboots Using KHO/LUO/memfd 20m
Motivation:
Kexec Hand-Over (KHO) helps minimize downtime during kernel updates by preserving state across the kexec reboot. However, there are some scenarios which need a deeper form of reboot to be effective -- for example, performing a firmware update or a device-tree refresh, or even to avoid known hardware/platform bugs that make the kexec path itself unreliable. In such cases, the kernel's "warm reboot" facility could be leveraged, which performs a full reboot (Firmware + Bootloader + OS/kernel init), but with the guarantee that this entire flow leaves DRAM contents intact (distinct from a "cold reboot", where DRAM is also cleared). We would like to extend KHO (and its memory state preservation feature via LUO & memfd, in particular) to also work across warm reboots, so as to minimize downtime during update scenarios that necessitate a warm reboot.
Key Challenges:
When trying to extend KHO/LUO/memfd to warm reboots, we ran into 2 key technical challenges, described below. We were able to solve the first one (supported by a working prototype), but we need the community's insights & feedback to help solve the second challenge (our proposed solution approaches are outlined below).
Challenge 1: Discovering preserved KHO state in the next kernel (without passing an FDT)
During a kexec, the handover from the first kernel to the second one is explicit -- by passing an FDT with info embedded as to where to find the preserved KHO state in memory. However, in the case of a warm reboot, there is no opportunity for the first kernel to do a similar hand-off, and yet, the second kernel should be able to discover where the KHO state was preserved, in order to restore it.
We solved this problem by using a classic Computer Science technique -- i.e., by adding a level of indirection :-). While the KHO state can be allocated & preserved anywhere in memory, the KHO metadata info is kept in a small header and it is passed to the next kernel. On UEFI platforms an EFI variable can be used for this purpose, whereas on DT based platforms, the header can be kept in a pre-determined location, and a DT node added to inform the kernel about it. This retains the KHO allocation flexibility. We also added checksums to verify the integrity of the preserved metadata to make this mechanism robust against stale/circumstantial data left in that fixed location.
Challenge 2: Preventing firmware from clobbering preserved KHO state across a warm reboot
The warm reboot flow includes all the usual init sequence across components like firmware and bootloader, leading up to launching the Linux kernel. The memory used by these components before Linux is launched can be termed as "FW-Transient", as they are only needed during that transition period. One common optimization technique used on memory-constrained systems is to allow firmware to temporarily re-use the same memory region that Linux has access to, but only when the kernel isn't running. Thus, FW-Transient memory during a warm reboot transition intentionally overlaps with Linux's usable memory, to avoid wasted permanent memory carveouts for firmware's temporary use. However, this introduces a problem for KHO, where the firmware can trample on memory regions that the kernel expects to pass unmodified to the next kernel across a warm reboot.
Delving deeper into the platform/firmware constraints around the use of FW-Transient memory, we found that while the memory maps vary widely between platforms, they do remain fixed (pre-determined at design time) for a given platform. This implies that while we can't move FW-Transient memory to cooperate with KHO, KHO can instead discover the boundaries of this region to avoid saving state-to-be-preserved in that region. But the exact mechanism for how to demarcate and export that FW-Transient region info from the firmware to the kernel across a variety of platforms (UEFI vs DT) remains an open question.
Using this insight, we have 2 solution approaches that we would like to propose to solve this challenge. The first one is to influence the sizing and placement of the kho-scratch region such that it fully overlaps with FW-Transient memory. If feasible, this would be an elegant solution because KHO won't have to do anything else to prevent firmware from clobbering its preserved state, as the kho-scratch region is meant for non-preserved allocations -- by definition. However, if there are any memory-management constraints that prevent kho-scratch placement in this way, we could also explore migrating preserved KHO state out of the FW-Transient region just before the first kernel's shutdown. This latter approach has its own potential pitfalls, such as difficulties in maintaining KHO state integrity post page migration and the cost of migrating pages in a performance sensitive path (i.e., the need for fast reboot to minimize downtime).
We would like to present these challenges and solution approaches and brainstorm with the Linux kernel, Firmware, and Bootloader communities at LPC on the best approach to solve this problem to enable KHO across warm reboots.
Speakers: Mr Prasanna Kumar T S M (Microsoft), Srivatsa Bhat (Microsoft) -
15:50
Live Update Compatibility 20m
Data that is serialized and preserved through a live update may have formatting differences between two kernels. This can affect the compatibility between two kernels when trying to perform a Live Update. Current efforts have been to use compatibility strings with implicit version numbers, like "memfd-v2" to help provide compatibility data. Changes to the formatting of the preserved data may not affect the compatibility if the new kernel understands the format of the old version. For example, an optional feature may be available on a new kernel that is not strictly needed. If the old kernel does not support that feature and the new kernel does, we still want to allow a Live Update between these versions.
This discussion is for a proposed method of serializing all supported versions of a component and having the new kernel negotiate the version to deserialize. The current proposal uses file handlers as the first user of this scheme, but this will apply to any component that needs to preserve and restore serialized data.
Speaker: Logan Odell (Google) -
16:10
Guest_Memfd Preservation For Live Update 20m
Guest_memfd is a specialized, guest-first memory subsystem within the Linux kernel, specifically designed for KVM. It provides an isolated file-descriptor-based approach to managing Virtual Machine memory.
For this MC, I want to propose the topic on Preservation of guest_memfd with LUO.
I will discuss the current development on guest_memfd preservation during kernel Liveupdate and (If time permits) future topics like supporting preservation for guest_memfd with private memory.
Link to patches: HereSpeaker: Tarun Sahu (Google) -
16:30
Coffee Break 30m
-
17:00
Performance improvements for live-update reboots 15m
Many use-cases of live-update require it to be fast. For cloud providers, executing host kernel updates without impacting guest workloads demands low-downtime live-update reboots.
We will discuss what improvements were already done to the Linux kernel, the ongoing work and future plans. This talk will also explain to developers how to configure their systems to achieve the best reboot performance.
Speaker: Michał Cłapiński (Google) -
17:15
Preserving tmpfs Files Across Live Update 15m
The Live Update Orchestrator (LUO) can preserve memfds across kexec. But a memfd has no path, so it can't replace a file on tmpfs, such as VMM binaries on hosts without local storage. Preserving a tmpfs file's contents is easy. Retrieving it is hard: LUO always returns a new anonymous file, with no way to give it a mount, path or mount options.
This talk compares three proposals:
- RETRIEVE_INTO_FD: userspace recreates the mount and directories, and the kernel fills an empty file.
- tmpfs mount and file handlers: the kernel preserves and recreates the mount and the files in its root directory.
- Retrieve as memfd, then move: retrieve a memfd and move its folios into tmpfs with an extended splice(SPLICE_F_MOVE).
We weigh the new uAPI, ABI and kernel code each one needs against what it supports. The goal is to agree on a direction for upstream.
Speaker: David Matlack (Google) -
17:30
Live Update IOMMU (vIOMMU with Arm SMMUv3) 20m
The current IOMMU Live Update framework establishes the mechanism for preserving IOMMU hardware state and translation units (HWPTs) across a live update. By utilizing the IOMMUFD and Kernel Handover (KHO) framework. This session talks about the current state of Liveupdate IOMMU and the upcoming features, challenges and enhancements.
The session dives deeper into the design challenges for enabling vIOMMU preservation to allow preserving the nested translation state, specifically guest-owned Stage 1 tables and metadata, with a focus on the ARM SMMUv3 nested translation architecture. We explore strategies for preservation, restoration, and lifecycle management of nested translation structures and associated metadata to ensure that the incoming kernel can resume Stage 1 + Stage 2 translation without guest-side disruption.
Looking forward to your feedback and participation.Speakers: Pranjal Shrivastava, Samiullah Khawaja -
17:50
VFIO Live Update Road to Continuous DMA 20m
Live Update enables host kernel updates with minimal guest downtime by preserving VM state across
kexec. While memory preservation (via KHO/LUO) covers VM RAM, pass-through PCIe devices assigned to VMs (e.g., GPUs, NICs, NVMe, accelerators) require kernel support to maintain device state across the reboot boundary.This talk presents the design of VFIO Live Update (Phase 1, recently posted in v5 [1]) and our roadmap toward continuous DMA [2]. Phase 1 establishes base support for preserving VFIO
cdevdevice file descriptors acrosskexecusing the Live Update Orchestrator (LUO) and Kexec Handover (KHO).We will discuss:
- Current State: What is the current state of VFIO patches?.
- Upstream Path: We will discuss few of the concerns regarding VFIO patches.
- Data storage and cleanup: Clarification on what to save and what should be cleanup during error.
- Future series: Future series toward making VFIO LU feature complete like:noiommu -> iommufd -> SR-IOV VFs).
Looking forward to get community feedback on our approach and align on the future direction of VFIO Live Update.
References
Speaker: Vipin Sharma (Google) -
18:10
Challenges in Preserving Virtual Networking State and dmabuf across Kexec Live Updates 20m
As cloud infrastructure scales and Confidential Computing adoption increases, host-level Live Update via kexec has become vital for maintaining zero-downtime operations. While progress has been made in preserving guest memory (guest_memfd, hugetlb) and hardware device states (VFIO), a critical gap remains in the network stack: preserving virtual networking topologies and stateful software-defined networking (SDN) backends, as well as high-performance memory buffers like dmabuf and devmem.
When a host undergoes a kexec warm reboot, virtual devices (such as netkit, veth, netkit, bonding) and software transport layers (RXE for Soft-RoCE, SIW for Soft-iWARP) lose their underlying kernel configurations, ring buffer indices, and state machines. Concurrently, zero-copy performance mechanisms like dmabuf and TCP devmem face the challenge of memory lifecycle decoupling—how to safely hand over page references and page tables across the kexec boundary without freeing the buffers or causing in-flight DMA corruption.
This interactive session aims to pinpoint the architectural constraints of virtual network state preservation and explore how to extend the Kexec HandOver (KHO) framework or complementary Live Update Orchestrators (LUO) to solve these stateful handover and high-performance memory recycling dilemmas.
Key Questions / Discussion Topics:
1. Generic State Management
Can we provide a generic state-management framework for Live Update, so individual drivers do not need to implement the full LUO/KHO callback lifecycle independently?2. Preservation Ordering & Quiescence
How should Live Update coordinate the preservation of DMA-BUF with PCI, VFIO, and IOMMU, especially when fence cleanup and DMA quiescence require these components to remain operational?3. State Synchronization
How can we synchronize restored DMA-BUF kernel state with hardware state, including IOMMU mappings, device contexts, rings, interrupts, and doorbells, without causing data corruption or IOMMU faults?4. Dependency & Recovery Ordering
How can Live Update enforce dependencies and restore ordering across components, such as PCI → VFIO → IOMMU and NIC → Netkit → RXE?5. Failure Handling & Dependency Status
How should DMA-BUF Live Update determine whether its dependencies are successfully preservable, and what should happen if PCI, VFIO, or IOMMU Live Update cannot be preserved?Speaker: yanjun zhu
-
15:00
→
18:40
Power Management and Thermal Control MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53The Power Management and Thermal Control micro-conference is about all things related to saving energy and managing heat. Among other things, we care about CPU, platform and device power-management mechanisms, thermal control support, and power capping. In particular, we are interested in improving and extending thermal control support in the Linux kernel and utilizing energy-saving features of modern hardware.
The general goal is to facilitate cross-framework and cross-platform discussions in order to improve energy-awareness and thermal control in Linux.
Since the previous iteration of this micro-conference, several topics covered by it have been addressed or work is in progress to address them, including:
- Thermal zones suspend and resume relocation closer to device suspend and resume, respectively: https://lore.kernel.org/linux-pm/12871778.O9o76ZdvQC@rafael.j.wysocki/
- Step-wise thermal governor improvements: https://lore.kernel.org/linux-pm/12745610.O9o76ZdvQC@rafael.j.wysocki/
- Support for latency limits in system-wide power management idle states: https://lore.kernel.org/linux-pm/20260205-topic-lpm-pmdomain-device-constraints-v2-0-61f7be7d35ac@baylibre.com/
- Support for fine-grained sync_state in generic PM domains: https://lore.kernel.org/linux-pm/20260410104058.83748-1-ulf.hansson@linaro.org/
- Support for remote processor cooling: https://lore.kernel.org/linux-pm/20260127155722.2797783-1-gaurav.kohli@oss.qualcomm.com/ and https://lore.kernel.org/linux-pm/20260419182203.4083985-1-daniel.lezcano@oss.qualcomm.com/
The topics that we would like to cover this year include, but are not limited to:
- Support for sync_state in more subsystems beyond generic PM domains
- Remaining rough edges in kernel thermal control support
- Cooling devices and thermal zones with parents
- thermald improvements, enhancements and development process
- Latency-focused QoS for kernel devices - implementation and experiments
- A lacking policy for power/perf-management of NVMe/UFS/eMMC/SD storage
The key people we would like to participate in the session are Rafael Wysocki, Ulf Hansson, Daniel Lezcano, Lukasz Luba, Srinivas Pandruvada, and Viresh Kumar.
-
15:00
Thermal Framework Corner Cases: Bugs or Undefined Behavior? 20m
The Linux thermal framework has proven to be flexible and robust over the years. However, some aspects of its event handling have become increasingly difficult to reason about, resulting in inconsistent behaviors and corner cases.
Examples include trip point updates while the temperature has already fallen within the hysteresis range after the notification was generated, inconsistent handling of thermal zones across system suspend and resume, and other situations where the expected behavior is not clearly defined or documented.
Some of these issues have been reported over the years, while others have surfaced as new users and platforms have adopted the framework. Addressing them individually without a common understanding of the expected semantics risks introducing further inconsistencies.
The goal of this discussion is to identify the current limitations of the thermal framework, establish which behaviors are considered bugs versus intentional design choices, prioritize the issues that should be addressed, and agree on a roadmap for improving the framework while preserving compatibility with existing users.
Speaker: Dr Daniel Lezcano (Qualcomm) -
15:20
Power Cooling Devices: Bridging Thermal and Powercap 20m
The Linux thermal management and power capping frameworks currently evolve independently, despite relying on the same Energy Model (EM) to describe the power-performance characteristics of devices.
The thermal framework regulates temperature by estimating a sustainable power budget using the power_allocator governor. A PID control loop computes the power reduction required to maintain a target temperature, and the Energy Model translates this budget into Operating Performance Point (OPP) constraints applied through thermal cooling devices.
The powercap framework, and in particular the Dynamic Thermal Power Management (DTPM) controller, also relies on the Energy Model to enforce power budgets. However, instead of manipulating OPPs directly, it distributes power limits across a hierarchy of power domains.
Although both frameworks address closely related problems and use the same underlying model, there is currently no mechanism to connect them. This raises the question of whether a common abstraction could allow the thermal framework to express power constraints through the powercap hierarchy rather than through device-specific OPP constraints.
One possible direction is the introduction of a power cooling device, acting as a bridge between the thermal and powercap frameworks. Such an abstraction could enable thermal governors to allocate power budgets to powercap domains while leaving power distribution and enforcement to the powercap subsystem. This approach may also open the door to hierarchical power budgeting and more coordinated thermal management across heterogeneous devices.
The goal of this discussion is to explore whether this direction makes sense from an architectural perspective, identify the challenges involved, and agree on a set of incremental milestones that would allow such an integration to be developed and upstreamed progressively.
Speaker: Dr Daniel Lezcano (Qualcomm) -
15:40
Introducing Hardware Thermal Governor 20m
Modern System-on-Chip designs increasingly incorporate built-in hardware thermal controllers to manage thermal conditions efficiently. Intel platforms, beginning with the Lunar Lake generation, feature integrated platform temperature controllers capable of autonomous thermal management. However, not all platform designs directly interface temperature readings and thermal thresholds with these hardware governors, limiting their effectiveness.
This presentation introduces a kernel-level approach that leverages existing thermal zone infrastructure to provide temperature data and threshold information directly through a new hardware thermal governor, bypassing user space interactions.
A proof-of-concept implementation will be demonstrated, showcasing the integration of thermal zones with hardware-based thermal control mechanisms.
Speaker: Srinivas Pandruvada -
16:00
Devfreq for uncore DVFS 20m
SoC uncore components, such as interconnects, system-level caches, and memory controllers, can account for a substantial share of package power, and their operating frequency bounds the achievable bandwidth and latency. Core DVFS is well supported by cpufreq, but there is no generic upstream mechanism for uncore DVFS - a gap felt most on server platforms. We propose building uncore DVFS on top of devfreq, the DVFS framework for non-CPU devices. The device model fits, but driving it from real utilization data exposes challenges worth discussing.
This topic intends to cover:
1. Motivation - why we need kernel-side uncore DVFS.
2. Existing upstream techniques - devfreq and vendor drivers.
3. Issues encountered - no way for a driver to obtain uncore PMU IDs and create perf events; no suitable events or scaling policy for uncore.
4. Proposals and analysis - a perf core helper, an event monitoring and scaling policy, and prototype measurements.Speaker: Jie Zhan (HiSilicon) -
16:20
Coffee Break 25m
-
16:45
Autonomous CPU performance selection: a new cpufreq governor? 20m
ACPI CPPC "autonomous selection" lets the platform pick the CPU performance level itself, within OS-provided min/max bounds and guided by an Energy Performance Preference (EPP) hint. Today cppc_cpufreq exposes this as a per-policy auto_select sysfs toggle layered on top of whatever scaling governor is attached. The result is confusing: once autonomous mode is on, the attached governor (schedutil, ondemand, ...) keeps issuing frequency requests that the hardware ignores, and scaling_governor no longer reflects what actually drives frequency.
cpufreq provides two ways to drive frequency, and CPPC autonomous mode fits neither. Scaling governors use the driver's ->target() callback to apply OS-selected operating points. setpolicy drivers (intel_pstate and amd-pstate "active"/EPP modes) instead let the hardware choose, but CPPC has requirements it does not cover. With setpolicy, the mode is selected globally for all CPUs at the driver level, whereas CPPC's registers are per-CPU and can be controlled per policy. setpolicy also exposes only the performance and powersave policies and provides no ->target() path, while the CPPC spec allows an OS desired_perf hint even when autonomous selection is enabled. So CPPC maps cleanly onto neither model.
This session proposes a third option: modelling autonomous, EPP-biased operation as a per-policy cpufreq governor. Selecting the governor enables autonomous mode and programs the EPP and min/max bounds; switching away restores OS frequency control. EPP is exposed as a per-policy governor tunable, and the autonomous state is reported read-only.
The broader question for the micro-conference is how cpufreq should represent hardware-autonomous selection in general:
- Is a new cpufreq governor the right way to expose autonomous selection at all?
- If a governor is the right model, should it be CPPC-specific or generalized so other drivers can share it?
- How should it relate to the existing setpolicy model used by intel_pstate and amd-pstate?
- Should such a governor also pass utilization-based desired_perf hints to the hardware?
- How should the EPP control be exposed to userspace (naming, accepted values, and consistency with the existing energy_performance_preference interface)?
- What is the migration path away from the current sysfs toggle without breaking existing users?
Speaker: Sumit Gupta -
17:05
Amortizing CPU wakeup costs 20m
Transitioning a CPU into and out of idle has a non-negligible energy overhead. During low-utilization states, frequent background wakeups repeatedly pay this "wakeup tax." Because the Energy Model only prices active OPP power and not idle-state transition costs, EAS defaults to spreading tasks across a cluster rather than packing them.
Prototyping with sched_ext on a Pixel 9 running Linux 7.1 as a proof-of-concept, we measure the power and latency tradeoffs of batching non-urgent wakeups onto already-active CPUs. We show that reducing idle exits measurably lowers CPU power, and that the tail-latency penalty of packing comes from committing tasks to a specific CPU at wakeup time. Deferring CPU selection until dispatch time achieves the same power savings at a fraction of the latency cost.
Discussion points:
- Energy Aware Scheduling: Should the Energy Model expose an idle-exit cost term, or rely on a low-utilization packing heuristic?
- Cluster-scoped deferral: Can we support cluster-scoped task push-pull (similar to RT pushable tasks) instead of committing at select_task_rq()? Would this make it easier to batch?
- Classifying latency-tolerance: What task QoS hints and system-state signals should govern wakeup deferral?
Speaker: Samuel Wu (Google) -
17:25
Current developments in the Common Clk Framework 20m
At last year's Linux Plumbers Conference, we had some great discussions about how to fix clock tree propagation in the Common Clk Framework (https://lpc.events/event/19/contributions/2152/). Taking that feedback into account, a v8 patch set has been posted that solves the problem in a simple manner.
https://lore.kernel.org/linux-clk/20260327-clk-scaling-v8-0-86cd0aba3c5f@redhat.com/
The KUnit tests demonstrate the problem where a clock can unknowingly change the rate of its parent and siblings to suboptimal frequencies. This can have adverse effects on subsystems that need precise rates, such as DRM and sound. The patch set needs more eyes from the community to get across the finish line. We’d like to use part of this session to go over the approach and figure out how to get it merged.
We can use the remainder of the time to discuss suggestions for ways to replace the clk subsystem's global prepare lock with a more fine-grained locking mechanism.
Speaker: Mr Brian Masney -
17:45
Short Break 5m
-
17:50
Hardware Pressure on x86: Does the Linux Scheduler Need to Know When You're Throttled? 20m
The scheduler's task wake-up logic and capacity-aware load balancing rely on arch_scale_cpu_capacity() to determine how much work a CPU can absorb. It also has logic to account for the effects of transient thermal- and power-driven capacity loss. This mechanism, known as hardware pressure, is currently not used on x86: arch_scale_cpu_capacity() returns a fixed constant on non-hybrid systems, and even on hybrid systems it is a self-normalizing ratio that hides uniform, proportional capacity loss across all cores, since the reference CPU used for normalization throttles along with the rest. No plumbing exists to feed x86 thermal/power status signals (PROCHOT, RAPL power-limitation bits, HWP capability changes) into capacity accounting.
This gap has consequences. Task placement at wake-up and EAS's
overutilized-detection gate can silently keep placing or retaining tasks on a throttled CPU, because the capacity term used for comparison never reflects the throttling; the load balancer may likewise
move tasks onto throttled CPUs. This is not a purely theoretical
concern: single-core thermal throttling is common even on non-hybrid parts, and Intel Speed Select Technology’s Core Power makes persistent, policy-driven asymmetric power allocation across cores.This talk proposes a design for exposing hardware pressure on x86 and demonstrates it against concrete use cases on real hardware. The goal is not to present a finished solution, but to solicit feedback on whether this problem is worth solving, and if so, whether the proposed solution makes sense.
Speaker: Ricardo Neri (Intel Corporation) -
18:10
Dynamic Runtime Prediction of CPU Idle-State Exit Latency 20m
Benefits of Accurate Exit Latency can have:
More Accurate hrtimer Expiration
Better CPU Idle level Selection
Improved Support for Latency-Sensitive SystemsCpu different low power state can have different exit latency. And the exit latency may be affected by:
- current cpu frequency
- Different firmware version
- Different hardware difference and etc.
So compile time static idle exit latency is not sufficient. Hence propose to have dynamic Runtime Prediction of CPU Idle-State Exit Latency
More Accurate hrtimer Expiration
Due to exit latency, the actual execution time of a timer interrupt is often later than the programmed expiration time. With an accurate estimate of exit latency, the timer expiration can be advanced accordingly, improving timer firing precision.QQ Music Version Max Delay (ns) Average Delay (ns) Original 2,446,254 618,558 Original 2,859,657 664,312 Original 1,916,403 466,087 Optimized 629,050 86,448 Optimized 674,887 107,095 Optimized 868,674 102,682 Honor of Kings (30 Hz) Version Max Delay (ns) Average Delay (ns) Original 2,309,583 14,470 Optimized 360,256 8,234Note that data is collected from an old kernel version and legacy qcom platform.
Better CPU Idle-State Selection
The current TOE cpuidle governor estimates idle duration using Exit Latency / 2. A more accurate exit latency estimation allows the governor to derive an idle duration closer to the actual value, leading to more appropriate idle-state selection.Improved Support for Latency-Sensitive Systems
Accurate exit latency prediction is particularly beneficial for latency-sensitive workloads, such as:
• PREEMPT_RT systems
• Real-time applications
• Interactive workloads like audio scenarioObtaining Accurate Exit Latency at Runtime (Monitor)
• Exit latency is fundamentally defined as:
Exit Latency = System Resume Timestamp − Wakeup Event Timestamp
• The current cpuidle governor already compares predicted idle duration against actual idle duration. However, the existing measurement only covers the interval between the last instruction before entering idle and the first instruction after wakeup. Exit latency itself is typically estimated using a static value defined in the device tree.
• In practice, exit latency varies dynamically depending on runtime conditions and system state.
• There are many possible wakeup sources. To accurately determine the wakeup-event timestamp, measurements are restricted to timer-based wakeups because the timer expiration time is known.
• To ensure measurement accuracy, only idle exits triggered exclusively by timer interrupts should be considered, i.e., the timer interrupt is the only pending interrupt when the CPU wakes up.
• CPU idle states are also dynamic. As additional CPUs enter idle, the cluster-level idle state may change. Therefore, the actual idle state must be determined dynamically both when entering and exiting idle, ensuring that collected historical data is correctly associated with the corresponding idle level.Exit Latency Prediction (Predict)
• Exit latency is influenced by multiple factors and continuously changes during runtime.
• A sliding-window-based approach can be used to collect multiple exit-latency samples.
• Outliers are filtered out, and the average of the remaining samples is used as the predicted exit latency.Utilizing Exit Latency (Control)
More Accurate Idle-Time Estimation
The cpuidle governor can use the predicted exit latency to improve idle-duration estimation accuracy, resulting in better idle-state selection.Improved Timer Accuracy on Broadcast-Timer Platforms
On platforms supporting a broadcast timer, multiple CPUs may need to wake up simultaneously. Since a broadcast timer interrupt can only be delivered to a single CPU initially, additional wakeup delay is introduced for the remaining CPUs.
To compensate for this effect, the exit latencies of multiple CPUs can be accumulated and applied to the first awakened CPU, producing a more accurate timer expiration schedule.Summary of Benefits
Accurate exit latency prediction provides the following advantages:
• More accurate idle-time estimation for the cpuidle governor.
• Improved CPU idle-state selection.
• Higher timer firing accuracy through compensation for exit latency.
• Reduced wakeup latency on broadcast-timer platforms.
• Better performance for latency-sensitive systems such as PREEMPT_RT.Open Discussions
• During evaluation, a certain percentage of outlier samples is still observed within the sampling window. Currently, these samples are removed directly to reduce their impact on prediction accuracy.
• On platforms running multiple virtual machines, the timer virtualization mechanism may affect the effectiveness and accuracy of exit-latency prediction and compensation. Further investigation is required.Speakers: Aiqun Yu (Qualcomm), Mr Cong Zhang (Qualcomm)
-
15:00
→
18:30
RISC-V MC "Small Theatre" (Prague Congress Centre)
"Small Theatre"
Prague Congress Centre
105LPC 2026: RISC-V Microconference
The RISC-V ecosystem continues to expand rapidly, with new silicon like the RVA23-compatible SpacemiT K3, a steady cadence of ratified and vendor-defined ISA extensions, and platform classes reaching from embedded parts to server-class SoCs. Session topics cover architecture work, platform and vendor enablement, firmware/SBI coordination, and userspace behavior, with the aim of arriving at concrete next steps that participants can act on after the conference.
Accomplishments since LPC 2025
Results and follow-ups from the 2025 microconference and the broader ecosystem since December 2025:
-
Control Flow Integrity (CFI): user-mode CFI support, Zicfilp (forward-edge landing pads) and Zicfiss (shadow stack), was merged for v7.0 window
-
ACPI enablement: System MSI and RIMT (RISC-V IO Mapping Table) support landed; additional tables and platform features are being wired up. A new RQSC (Quality of Service Controller) table is under review.
-
RVA23 profile: preparatory work in the kernel for safely enabling RVA23-assuming code paths has continued following Charlie Jenkins's 2025 talk.
-
Control Transfer Records (CTR): kernel and QEMU support is maturing.
-
SBI / firmware messaging: the Message Proxy (MPXY) mailbox driver and related SBI extensions for firmware-mediated device access have been merged / refined.
-
Platform enablement: expanded SoC peripheral support (SpacemiT, Eswin, etc.), and progress toward generic distro boot on RISC-V.
-
QoS: the Ssqosid + CBQRI + RQSC resctrl series has significantly matured following the 2025 talk.
Proposed topics for 2026
Topics are targeted at ~15-30 minutes. Each session should only have a couple of slides to inform and stimulate discussion among the people attending the session.
-
RVA23 in practice - what it means once distros begin assuming it, remaining gaps in discovery, compatibility fallbacks for older hardware, and testing strategy.
-
Vendor-specific extensions - strategy for merging and enabling vendor extensions without fragmenting
arch/riscv. -
Kernel CFI: next steps - forward-edge kernel CFI, indirect branch tracking, etc
-
ACPI on RISC-V - what is still missing for ACPI-first platforms (power, thermal, PCIe quirks, etc)?
-
RISC-V QoS and resctrl - Ssqosid + CBQRI + RQSC series status, resctrl integration, open review items, and how this plugs into the cross-architecture resctrl rework.
-
SBI firmware messaging (MPXY and beyond) - when MPXY-style firmware mediation is the right answer, how it interacts with mailbox/RPMsg, and the bindings/API stability story.
-
IOMMU / RIMT and DMA - RISC-V IOMMU driver maturity, ATS/PRI, SVA on RISC-V, nested translation, DMA coherence and CMO.
-
KVM / hypervisor topics - Smrnmi handling, nested virt, Supervisor Software Events (SSE), H-extension adoption across silicon.
-
Vector - Vector usage in kernel like crypto and memcpy; should kernel try to support SoC where some cores have longer vector length than others like K3?
-
Pre-silicon upstream methodology — continuing from Yuning Liang's 2025 talk: what worked, what didn't, and a shared checklist for bring-up in simulation/emulation environments.
-
Debug and crash tooling - kdump and crash follow-ups to Austin Kim's 2025 talk; kgdb, perf, and on-target debug
-
RV32 and small cores - follow-up to the "schism" discussion at LPC 2025; what is the sustainable plan for RV32 in-tree?
Key participants
As in 2025, the microconference draws organizers, maintainers, and contributors from across the RISC-V Linux community. Expected key participants: Paul Walmsley, Anup Patel, Conor Dooley, Andrew Jones, Deepak Gupta, Charlie Jenkins, Sunil V L, Samuel Holland, Radim Krcmar, Guo Ren, Inochi Amaoto, Yixun Lan, Ruinland, Austin Kim, Yuning Liang, Andy Chiu, Mikey Neuling, Andy Gross, Joel Stanley, Michael Ellerman, Anirudh Srinivasan, Drew Fustini
-
15:00
Topics in RISC-V Linux kernel maintenance 20m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105Review the past year in RISC-V Linux kernel maintenance, and discuss upcoming plans for the next year - similar to what we did last year
Speakers: Paul Walmsley (SiFive), Palmer Dabbelt (Google) -
15:20
RISC-V Svnapot: dynamic fold/unfold for folded PTEs 20m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105Svnapot gives RISC-V an architectural way to encode contiguous PTE ranges as folded mappings, and 64K mTHP is a particularly good match
for that capability.This session presents a patch series that enables that model through dynamic fold/unfold and a split between public and raw page-table
APIs. Public helpers preserve the logical sub-PTE view expected by core MM, while raw helpers remain available to architecture-private
users that need direct access to the encoded hardware entry. In our measurements, combining 64K mTHP with Svnapot aggregation reduces
lat_mem_rd latency by about 12% compared with non-aggregated 4K mappings, improves overall SPEC CPU performance by around 2%, and
delivers about 7% on the instruction-TLB-sensitive 520.omnetpp_r sub-benchmark.The session will also discuss whether fold/unfold transitions can avoid intermediate TLB flushes, and examine the trade-offs between
hardware 64K pages and Svnapot-based aggregation, including the advantages and drawbacks of each and whether Svnapot should support
additional aggregation granularities beyond today’s 64K mTHP use case.Speaker: Yunhui Cui (ByteDance) -
15:40
Optimizing the kernel with vector 20m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105The territory of using a larger context in the kernel-mode has been a forbidden topic since the existence of Linux. The use of SIMD unit in the kernel-mode is strictly restricted because the program is highly optimized for stack footprint, responsiveness, and maximum hardware compatibility. However, as hardware and compiler technologies advance, the benefit of enabling autovectorization has become more viable, and it might be a good time to revisit it.
In this talk we will go through some initial measurements to see the benefit of compiling the risc-v kernel with autovectorization, challenges of managing VLA context and ways to maintain code when compiled with such feature.
Speaker: Tao Chiu -
16:00
RVA23 in practice: discovery, dispatch, and the hardware left behind 20m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105RVA23 is ratified and the first (vendor-announced) compliant silicon is shipping, but between "the profile exists" and "userspace can rely on it" sit three unresolved kernel questions.
Discovery. Extension capability discovery between user mode and the kernel is still messy: hwprobe and prctl can disagree, so which interface should IFUNC resolvers actually trust, or does the discovery contract itself need a rethink? The "riscv: hwprobe: export the availability of vector to user" series under discussion on linux-riscv is a live example.
Dispatch. (runtime selection of optimized code paths) x86 already does coarse, library-level selection via glibc-hwcaps. Could the same mechanism sit on top of hwprobe on RISC-V, keyed on something like rva23u64? Is profile granularity even the right key? For distros like Debian that will not flip their baseline any time soon, this is probably the realistic path. But there is a catch though, as the prctl/hwprobe mismatch shows.
The hardware left behind. What compatibility fallbacks do pre-RVA23 systems get? And further out: can kernel-side cooperation support RVA23 capability certification, and what test infrastructure would that take?
Aiming at concrete next steps for the discovery contract and a dispatch story distros can adopt. Four asks to the room:
Ask 1 kernel maintainers: is the RVA23U64 bit the right shape?
Ask 2 distros: ship rva23u64 builds of selected libraries?
Ask 3 kernel maintainers: deprecate the V prctl?
Ask 4 future extensions: always-on, or dynamically enabled per process?Speaker: Guodong Xu -
16:20
Extensions - What's new? What's coming up? 10m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105The RISC-V ISA specs contained the ratified collection of bases, extensions and profiles. This talk will cover recently ratified extensions, as well as cover upcoming extensions and profiles which are close to ratification. We'll dive into kernel / user space design considerations that implementors will want to care about.
Speaker: Mr Tom Gall (RISC-V International) -
16:30
Coffee Break 30m "Small Theatre" (Prague Congress Centre)
"Small Theatre"
Prague Congress Centre
105 -
17:00
Local, SBI, or IPI? Tracing RISC-V TLB Flush Decisions in Linux 15m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105RISC-V Linux TLB flushing has several runtime paths. A flush may be local to the current hart, may use SBI remote fence support, or may fall back to Linux IPIs. The kernel also makes range-versus-full decisions, performs ASID-aware flushing, and carries different flush strides for normal and huge-page ranges. These decisions can matter when debugging correctness issues, performance anomalies, and SMP scalability problems, but they are not currently visible in a compact RISC-V-specific form.
This talk presents a small RFC tracepoint patch that makes the Linux-side RISC-V TLB flush decision visible. The proposed tracepoint records the requested flush range, stride, ASID, target CPU count, flush type, threshold decision, and selected transport path: local, SBI remote fence, or IPI fallback.
The goal is intentionally narrow. This is not a hardware TLB profiler and not a complete MM observability framework. It focuses on one practical question: when Linux decides to invalidate RISC-V address translations, can a developer easily see which path was selected and why?
The talk will walk through the relevant RISC-V kernel TLB-flush path, explain what existing tracing can and cannot show, and demonstrate the proposed tracepoint using QEMU and one real RISC-V board. A small multi-threaded mmap() / mprotect() / munmap() workload will be used to trigger local and remote flushes and to show range-size effects around the current threshold logic.
Attendees will leave with a concrete understanding of how RISC-V TLB flush decisions are made in Linux, how they can be observed today, what a minimal tracepoint adds, and what trade-offs should be considered before adding architecture-specific observability to hot MM paths.
Comments:
This topic is primarily intended for the RISC-V MC because it focuses on arch/riscv MM/TLB behavior, including local flushes, SBI remote fences, and IPI fallback paths. If the organizers think the tracing or MM angle is a better fit elsewhere, it could also be considered for the Tracing MC or Kernel Memory Management MC.
I previously presented in the FOSDEM 2026 Kernel devroom with a practical talk on reproducing a Linux kernel bug using virtme-ng. This LPC proposal is a different topic, but it follows the same practical style: a small kernel RFC patch, reproducible QEMU testing, and validation on real RISC-V hardware where possible.
P.S. Patches to review:
v1
https://lore.kernel.org/lkml/20260829-tlb_tracepoint-v1-1-dfdaede7e741@gmail.com/T/#me0739340db5d4dfc6e0ff1675896129af941eca5
v2
https://lore.kernel.org/linux-riscv/20260830-tlb_tracepoint-v2-1-e3c88c24df5b@gmail.com/T/#u
v2 re-send:
https://lore.kernel.org/lkml/20260924-tlb_tracepoint-v2-1-4e3e78ef5cdb@gmail.com/
asciinema presentation:
https://asciinema.org/a/1266683Speaker: Mr Roman Storozhenko (Intel) -
17:15
One Kernel Per Cluster: Multikernel for Heterogeneous RISC-V 20m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105Boot mainline on SpacemiT's Key Stone K3 and half the machine stays dark: eight 1024-bit AI cores rejected by smpboot because riscv_v_vsize allows exactly one vlenb per system. This is not a bring-up bug, it is the first shipping proof that RISC-V vendors will not pay for matched vector widths, and the SG2042's RVV 0.7.1 installed base makes it worse: not narrower, incompatible. Per-process vlenb patches can boot this machine, but no patch changes the physics: 4 KB of vector state cannot land in a 1 KB register file, so any single-kernel fix confines vector tasks to one cluster for life and inherits the arm64 asymmetric-32bit price list: no hotplug of the last capable core, no SCHED_DEADLINE admission, silently rewritten affinity masks, undefined cpuset isolation, plus hwprobe intersecting capabilities down to no vector unit anywhere on a 0.7.1/1.0 mix. The partition was decided at tape-out; the scheduler can only hide it, expensively. We propose making it first class instead: multikernel boots one independent Linux kernel per cluster, each compiled for its exact silicon, seeing a machine where every core matches. The capability word tells the truth, glibc picks real vector routines, hotplug and deadline scheduling and cpusets just work, and RVV 0.7.1 and 1.0 domains coexist on one board. The prototype exists (github.com/multikernel/linux); what it needs is this room. We will present the architecture and the RISC-V gaps we want to close with the community: kernel spawning on the RISC-V boot flow, cross-domain interrupt routing with AIA/IMSIC, and DT/ACPI bindings for heterogeneous vector topology. Heterogeneous VLEN is the RISC-V roadmap, not a corner case. Let's give it a first-class answer before the next SoC tapes out.
Speakers: Cong Wang, Mr Yuning Liang (Deep Computing) -
17:35
Error INJection support for RAS on RISC-V architecture 15m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105Error INJection (EINJ), defined by ACPI, offers a platform-independent mechanism to inject hardware errors and validate system resilience paths. By avoiding platform-specific tooling, EINJ enables consistent testing of the Linux error-handling stack across architectures, including APEI and broader RAS workflows.
This session discuss the design and implementation strategy for enabling EINJ on RISC-V platforms. It will cover architecture-specific considerations, firmware and kernel integration points, and practical validation flows for injected error scenarios. The goal is to establish a robust, upstream-friendly approach that brings RISC-V closer to feature parity with existing ACPI-based RAS ecosystems and improves confidence in production reliability.
Speaker: Mr Himanshu Chauhan (Qualcomm Technologies) -
17:50
RISC-V Trace support in ACPI 15m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105To support RISC-V E-trace/N-trace components in ACPI, we propose using the Device Graph UUID defined in the DSD Guide, aligned with how similar topologies are described on other architectures. Beyond discussing the proposal itself and how it should appear in the ACPI namespace, we also want feedback on implementation strategy.
Today, ACPI fwnode graph handling does not interpret Device Graph UUID data, so trace drivers parse ACPI graph links in-driver. We have a PoC with that approach, but it duplicates logic, raises long-term maintenance cost, and keeps DT and ACPI paths separate instead of reusing common fwnode_graph_* helpers.
Our proposal is to extend ACPI fwnode graph support, so Device Graph UUID data is translated into standard fwnode_graph_* operations. This allows shared trace-driver code across DT and ACPI, reduces duplication, and scales better as additional architectures adopt this representation.
We plan to post an RFC series before the conference so maintainers and subsystem stakeholders can review concrete patches and align on direction.
Speaker: SUNIL V L (Qualcomm) -
18:05
Supporting Heterogeneous VLEN on RISC-V: Lessons from SpacemiT K3 15m "Small Theatre"
"Small Theatre"
Prague Congress Centre
105SpacemiT K3 is among the first production RVA23-compatible SoCs, shipping, features heterogeneous vector lengths:
- 8× X100 cores: RV64GCV, VLEN=256, full RVA23 compliance, H-extension — general-purpose
- 8× A100 cores: RV64GCV, VLEN=1024, no H-extension — AI/vector compute
Both core types implement RVV 1.0, but their incompatible vector register widths create a correctness issue: if a thread using vector instructions migrates between core types, its vector state corrupts. This isn't a performance concern — it's a data integrity problem.
Our current vendor-specific approach (CONFIG_SPACEMIT_HMP) enforces strict per-thread core-type binding. This works, but it's not upstream-friendly.
What SpacemiT Can Contribute
We can provide engineering data from 2+ years of HMP scheduling in production, including details on vector state management, signal frames, ptrace, and CPU hotplug. We'll also share real constraints — why certain approaches work or don't on actual silicon.
K3 development boards are available for testing, and we have deployment experience with 9+ distributions (Ubuntu, Debian, Fedora, openEuler, openKylin, etc.). We're committed to upstream patches based on community consensus.
Proposed Discussion Topics
1. Kernel mechanism options
We've explored several approaches and can share trade-offs:
- Generic heterogeneous-VLEN scheduling (similar to ARM capacity-aware)
- Cpuset/cgroup-based isolation
- ISA-capability-aware scheduling
- Userspace-only manual pinning2. Vector state management
- Should
riscv_v_vsizebe per-task or per-CPU? - Signal frame handling with variable VLEN?
- ptrace behavior?
3. DT bindings and discoverability
- Should VLEN capabilities be explicit in DT or runtime-discovered?
- Standardization of per-core capability properties?
4. RVA23 ecosystem assumptions
- Should RVA23 distributions assume uniform VLEN=256?
- How should applications/libraries handle VLEN variability?
- Kernel enforcement vs. userspace detection?
5. Should kernel support this at all?
- Is this a legitimate use case (like ARM big.LITTLE) or a hardware constraint software shouldn't accommodate?
Session Format
We suggest a community-driven discussion. SpacemiT can provide engineering data and K3-specific details, while community maintainers lead the architectural direction. We'll share detailed technical background (2-3 pages) before the conference.
Why This Matters
- RISC-V flexibility vs. portability: vendors can customize per-core ISA, but software must remain portable
- This is solvable: ARM addressed big.LITTLE heterogeneity; RISC-V can too
- Prevent fragmentation: if every vendor ships custom scheduling, distro/kernel maintenance becomes unsustainable
Desired Outcomes
- Community consensus on whether/how to support heterogeneous VLEN upstream
- DT binding direction for representing per-core capabilities
- Clear action items for follow-up patches and testing
Pre-Conference Materials
We will circulate 2-3 pages of technical background (HMP implementation, trade-offs, deployment observations, code pointers) to participants before the conference.
Reference: https://github.com/spacemit-com/linux
Speakers: Mr Qiubin Zhuang (SpacemiT), Mr Dong Xu (SpacemiT)
-
-
17:45
→
18:30
Birds of a Feather (BoF) "Small Hall" (Prague Congress Centre)
"Small Hall"
Prague Congress Centre
215-
17:45
Camera & ISP BoF 45m
While cameras have been ubiquitous in Linux systems for more than a decade, vendors have historically been very reluctant to disclose any information about Image Signal Processors (ISP), leading to the proliferation of out-of-tree kernel drivers and closed-source userspace stacks. The situation started to change with the launch of the libcamera project at the end of 2018, and progress has accelerated over the past couple of years with more and more vendors jumping on board. Even Qualcomm recently posted an initial ISP driver in a timid but very real first step, a move that was unthinkable just a couple of years ago.
There is plenty of work left to do, as the increased interest from vendors lays bare the lack of investment of the previous decade that leaves many technical issues unsolved. This BoF brings representatives of kernel subsystems (mainly V4L2, but also DRM), userspace frameworks (libcamera, GStreamer, PipeWire, ...), image sensor vendors and ISP vendors in the same room to discuss and solve open issues.
Discussion Topics
Two discussion topics have been scheduled for the BoF. Additional topics may be discussed if time permits.
RGB-IR Support in V4L2 and libcamera
by Rishikesh Donadkar and Devarsh Thakkar, Texas Instruments
The Linux kernel V4L2 API does not support RGB-IR image sensors. While building blocks necessary to handle those devices are slowly being merged, RGB-IR support itself hasn't been tackled yet.
Rishikesh has posted an API RFC. Along with Devarsh, he will also present this topic at the OSS Europe conference.
Camera Module Identification
by Stefan Klug, Ideas on Board
Camera tuning and calibration depend not only on the image sensor, but also on the lens and other characteristics of camera modules. While Linux supports identifying image sensors, it completely lacks the concept of camera modules.
Identification of camera modules requires coordination between platform firmware (DT or ACPI), the kernel and userspace. Discussions will benefit from the presence of DT maintainers.
Stefan has detailed the issue and started a public dicussion on the linux-media mailing list.
Additional Topics
The following additional topic has been proposed.
Advanced Camera Use Cases
Gjorgji Rosikopulos proposed discussing advanced camera use cases found in proprietary camera stacks that lack upstream kernel and userspace support.
Speaker: Laurent Pinchart (Ideas on Board Oy)
-
17:45
-
10:00
→
13:30
-
-
09:00
→
15:45
LPC Refereed Track "South Hall 1 B" (Prague Congress Centre)
"South Hall 1 B"
Prague Congress Centre
158-
09:00
TAB Q&A 45m
The Linux Foundation Technical Advisory Board (TAB) represents Linux Kernel project interests to the Linux Foundation. It also uses the pooled influence of its elected members to support the long term health of the project.
This open forum / panel discussion is an opportunity to learn about and discuss TAB initiatives and ongoing project needs.
Speakers: Dave Hansen, David Hildenbrand (Arm), Greg Kroah-Hartman, Julia Lawall (Inria), Kees Cook (Google), Miguel Ojeda, Shuah Khan (The Linux Foundation), Steven Rostedt, Theodore Ts'o (Google) -
10:00
Cryptographic Proofs of Personhood: Solving the Kernel Web of Trust Problem in a Privacy-Preserving Manner 45m
In any open source software project, maintainer identity is becoming a critical problem. How can you be confident that someone contributing a patch (or PR) is not a malicious adversary? Attacks like the XZUtils compromise have heightened the concern around this sort of software supply chain attack, and the “North Korean developer” problem plagues both companies and open source projects.
Today, the kernel uses the kernel.org PGP web of trust to mitigate this problem. However, the technology that is used is outdated, leading to a system that is neither scalable nor private. In this talk, we propose using a system built on cryptographic proofs of personhood as a better alternative for a cryptographic web of trust. Informally speaking, cryptographic proofs of personhood allow us to build a decentralized, private reputation system, giving us all of the power of the existing web of trust, plus a whole new suite of functionalities.
We will explain the cryptographic principles behind proofs of personhood at a high level in a way that non-cryptographers can understand. Then, we will explain how this technology can be applied to solve the web of trust problem. Finally, we will demonstrate a fully functional, large implementation of proofs of personhood using entirely open source technology (standards and code) hosted in Linux Foundation Decentralized Trust. The demonstration will show that such a solution is ready to be adopted by kernel.org for the web of trust.
Speakers: Mr Drummond Reed (First Person Cooperative), Mr Glenn Gore (Affinidi), Hart Montgomery (Linux Foundation) -
10:45
A Loadable Crypto Module for FIPS Certification 45m
Many organizations require US Federal Information Processing Standard (FIPS) certification of the crypto code they are running. The certification process is lengthy (typically 12–18 months), but the bigger problem is that the way the crypto subsystem is built into the kernel makes the result unable to be reused across kernel updates. This is because FIPS certification is granted at the binary level by NIST. In current kernels, the crypto subsystem is built directly into the main kernel image, so even a non-crypto kernel update — a scheduler fix, a driver addition — produces a new binary and invalidates the existing certification. Distributions are then forced through the full validation cycle again, making it extremely difficult to deliver timely kernel updates while maintaining FIPS compliance.
This talk presents a solution that has shipped in kernel 6.18 on Amazon Linux 2023 and is covered in an LWN.net feature article: decoupling the crypto subsystem from the main kernel and building it as a separate loadable module. The module, rather than the entire kernel, becomes the unit of certification. A subsequent kernel update that does not touch the module leaves the existing certification intact, and loading that certified module onto the updated kernel makes it FIPS-compliant automatically. When the crypto code itself needs updating, a new module can be submitted for certification while users continue running the previously certified one on newer kernels; the old module acts as a bridge, so users never have to choose between an updated kernel and FIPS compliance.
While the existing kernel's module system already allows code to live outside the main kernel, this talk will present why the same approach cannot simply be used here, what falls short, and how the approach overcomes these obstacles in its design and implementation as shipped in production. It will also discuss what it needs to be as a unified upstream foundation that distributions can customize to satisfy different certification setting requirements.
References:
- Patch series: https://lwn.net/ml/all/20260418002032.2877-1-wanjay@amazon.com/
- LWN article: https://lwn.net/SubscriberLink/1073759/95b3d4cd28506836/
- AWS compute blog: https://aws.amazon.com/blogs/compute/introducing-modularized-kernel-cryptography-in-amazon-linux/
Speaker: Jay Wang (Amazon) -
11:30
Coffee Break 30m
-
12:00
From CVE to Fix in Minutes: AI-Augmented Security Lifecycle Management for Linux Distributions 45m
Linux Plumbers Conference 2026 — Proposal
Recommended Track
Distributions Microconference
Why this track: The proposal centers on distro-level CVE lifecycle management — scanning, triage, patching, backporting, and release engineering — which sits squarely in the Distributions MC's scope. It addresses pain points shared by every LTS distro maintainer (Debian, Fedora, openSUSE, Alpine, etc.) and invites cross-distro collaboration on AI-assisted tooling.
Alternative tracks (if Distributions MC is not available):
- Security MC — the CVE rescoring methodology and SLA-driven release model are directly relevant.
- Tooling MC / Refereed Track — the AI backporting agent and automated review pipeline are novel developer tooling contributions.
Title
From CVE to Fix in Minutes: AI-Augmented Security Lifecycle Management for Linux Distributions
Abstract
Every day, hundreds of new CVEs are disclosed. For Linux distribution maintainers, the challenge is not whether vulnerabilities will arrive — it is how fast you can triage, patch, test, and ship fixes at scale. This session presents a production-proven, end-to-end pipeline that automates the full CVE lifecycle for a Linux distribution serving millions of deployed systems.
We cover five key stages and the design decisions behind each:
1. Detection & Ingestion
CVEs are continuously ingested from the National Vulnerability Database and cross-referenced against distro-specific package versions using scanning pipelines. Raw CVE feeds are noisy, so an automated triage layer checks whether a CVE actually affects the distribution by examining build configurations, internal dependency graphs, and shipped code paths.
2. Distro-Specific Rescoring
Upstream CVSS scores frequently misrepresent actual risk for a given distribution. We present a principled methodology for rescoring CVEs against your own build configuration and hardening posture — a vulnerability rated HIGH upstream may genuinely be LOW risk due to compiler flags (
-fstack-protector-strong,-D_FORTIFY_SOURCE=2), disabled features, or sandboxing that neutralizes the attack vector. This directly combats alert fatigue and misallocated engineering effort.3. Automated Patching & AI-Powered Backporting
Once a CVE is confirmed, the pipeline automatically attempts the fix — preferring minor version (patch-level) upgrades when available, falling back to cherry-picking upstream patches. When neither works cleanly, an AI-powered SWE (Software Engineering) Agent backports patches to the distro's specific version, resolving merge conflicts, adapting to API renames, struct layout changes, and conditional compilation differences. The agent operates in an iterative build-feedback loop: apply patch → build → analyze errors → refine — producing a complete, buildable patch set with full provenance.
Case studies we'll share:
- Patches requiring adaptation across 3+ major version gaps
- Handling renamed functions and refactored code paths
- Fallback strategies when AI backporting fails and how we route to human experts with maximum context4. Automated Review Pipeline
Reviewing CVE patches is one of the most time-consuming bottlenecks in distro maintenance (15–30 min per PR). We built a multi-stage automated review pipeline that reduces human review time to ~1 minute:
Stage Method What It Checks Spec Validation Deterministic Version bumps, patch declarations, changelog format, signature updates Build Log Analysis Heuristic Errors, warnings, test failures from CI Semantic Patch Comparison LLM-driven Classifies match against upstream fix (exact / minor diff / clean backport / significant divergence), generates risk score Structured Report Template (Jinja2) Renders Markdown review posted directly to the PR We share accuracy metrics, prompt engineering techniques for reliable semantic diff analysis, and lessons learned from production deployment.
5. SLA-Driven Release Engineering
We enforce strict SLAs tied to severity:
Severity Fix SLA Release Channel Critical 5 business days Fasttrack RPM release High 10 business days Fasttrack RPM release Medium 30 business days Monthly cadence Low Next release Monthly cadence We share the operational framework for tying CVE severity to release cadence and how end-to-end observability (CVE inflow trends, severity distribution, package hotspots, fix throughput) enables data-driven security posture management.
Why This Matters to the Plumbers Community
-
Reproducible, Distro-Agnostic Blueprint — The architecture (scan → triage → rescore → patch → AI-backport → test → review → ship) is not tied to any single distribution. Maintainers of Debian, Fedora, Alpine, Gentoo, or any custom enterprise distro can adopt the same pipeline patterns.
-
Tackles the Maintainer Shortage — The Linux ecosystem faces a chronic shortage of security-focused maintainers. By automating 80%+ of the CVE lifecycle, small teams can maintain the security posture of a large distribution — directly addressing the sustainability crisis in open-source maintenance.
-
AI Backporting as a Shared Community Tool — We want to start a conversation about building a community-maintained backporting agent that could serve multiple distributions. Every LTS distro, every stable kernel branch, and every enterprise vendor deals with backporting; a shared tool benefits everyone.
-
Distro-Specific Rescoring Should Be Standard Practice — Most organizations blindly consume upstream CVSS scores. We propose a methodology that any distro can adopt to prioritize what truly matters for their configuration.
-
Faster Backporting = Smaller Exposure Windows — AI-assisted backporting means vulnerabilities are patched sooner in stable releases, directly improving security for billions of deployed systems.
Session Format
Preferred: 30-minute presentation + 15-minute discussion
Alternate: 20-minute presentation (can condense to focus on backporting agent + review pipeline)We can provide a live demo of:
- The automated patching pipeline processing a real CVE
- The AI backporting agent resolving a non-trivial merge conflict
- The review pipeline generating a structured review report
- The CVE observability dashboard
Discussion Topics for the Microconference
If accepted as part of a broader discussion slot, we'd like to explore:
- Could distros share a common AI backporting agent, and what would the interface/API need to be?
- How do other distros handle the tension between automated patching speed and review thoroughness?
- What test infrastructure is needed to validate AI-generated patches with high confidence?
Speaker Bio
Kanishk Bansal is a Software Engineer working on Azure Linux distribution security infrastructure at Microsoft. His work spans CVE triage, AI-powered patch backporting, and intelligent code review systems for a production Linux distribution. He is passionate about applying AI to the critical but under-resourced work of open-source supply chain security.
Speaker: Kanishk Bansal -
-
12:45
BPF_PROG_TEST_RUN: the hidden perils of testing. 45m
BPF_PROG_TEST_RUNhas been a crucial tool for testing eBPF programs. Although it has been instrumental for XDP, as BPF spreads to other hookpoints (e.g., TC, sock_ops, struct_ops) that rely on complex data structures such as the SKB, we found that it provides an environment that diverges significantly from the kernel: it does not model these structures and instead defaults their state to zero. This leads to unfaithful test replication and test cases that may miss bugs that would occur in production. We verified three such scenarios in which a single zeroed field hides an entire branch of real kernel behavior.A simple, concrete instance is
skb->cloned, which the harness hard-wires to zero so that every clone-branching helper only ever takes its fast path. Testing against open-source eBPF programs, we confirmed this produces silent semantic divergence:bpf_skb_ecn_set_cereturns success but never sets the CE bit when an skb is cloned and its IP header is unwritable: a path the harness cannot reach. The test therefore reports a pass while the corresponding production behavior differs, illustrating how one unmodeled bit lets both unit tests and verification tools built onBPF_PROG_TEST_RUNoverlook real bugs.Speakers: Lucas Castanheira (CMU), Prof. Theophilus Benson (Carnegie Mellon University) -
13:30
Lunch Break 1h 30m
-
15:00
When Embedded Codecs Meet Graphics 45m
In the current Linux kernel, video codecs split into two categories. Accelerators built into GPUs are implemented as thin drivers with userspace components exposing APIs such as Vulkan Video or VA-API. Everything else falls under the Video4Linux kernel API. This fragmentation adds complexity for both kernel and userspace developers. VA-API has stagnated and carries historical assumptions tied to Intel hardware, while Video4Linux, despite remaining relevant and widely deployed, has an aging memory model and resource queue mechanism that limits modern use cases. Vulkan Video, by contrast, is actively developed and designed with hardware-specific extensions in mind, making it a strong candidate for embedded use cases.
This talk proposes introducing a new class of embedded video codec drivers into Linux: thin, vendor-specific drivers modeled after GPU drivers, where the V4L2 kernel API is replaced by Vulkan Video, with the userspace driver implemented in Mesa. The goal is not to reinvent the wheel, but to replace an aging one. Such an approach would bring better buffer management through existing DRM memory helpers, modern explicit synchronization primitives already used by GPU drivers, and reduced kernel complexity by pushing codec-specific logic into userspace where it belongs.
Speaker: Nicolas Dufresne (Collabora Ltd.)
-
09:00
-
09:00
→
19:00
Networking Track "South Hall 1 A" (Prague Congress Centre)
"South Hall 1 A"
Prague Congress Centre
158-
09:30
Short-Lived UDP and the Observability Gap: Challenges in Building Complete Flow Graphs for Identity-Based Policy 30m
Identity-based micro-segmentation depends on building a complete graph of workload-to-workload communication. For TCP, this is tractable: connections have clear lifecycle events observable via sock_ops and inet_sock_set_state tracepoints, and conntrack entries persist long enough for user space collection. UDP has none of this, and in production Kubernetes environments, UDP is everywhere: DNS queries,NTP,syslog,SNMP custom service discovery protocols, and increasingly, QUIC.
Generally in production micro-segmentation deployment, 15-20% of observed flows are UDP-based, and we consistently undercount them by 30-40% compared to packet-capture ground truth. This means if you have some ML engine in place, then this ML-based policy engine generates policies with blind spots: it doesn't suggest rules for UDP flows it never observed, leaving them to hit default-deny and break applications after policy enforcement.
What I'll cover:
- Why UDP flow observability is structurally harder- TCP connections:
SYN creates conntrack entry -> observable via ctnetlink ->
inet_sock_set_state tracepoint marks lifecycle. UDP: first packet
creates a conntrack entry with a short timeout (30s default, 120s
for "assured"), and single-datagram exchanges (DNS request->response)
may age out before user space collects them. I'll present data from
production system showing the correlation between conntrack
timeout settings and UDP flow miss rates. - Three approaches and their costs- (1) Tracing every
udp_sendmsg/udp_recvmsg via kprobe/fentry: complete visibility but
3-5% throughput overhead on UDP-heavy workloads.
(2) Increasing conntrack UDP timeouts: reduces miss rate but bloats
conntrack table size (we measured 4x table growth) and increases
memory pressure. (3) BPF socket filter at cgroup level: catches
socket-level events but misses raw UDP (common in legacy workloads
that use raw sockets) - Proposal: lightweight UDP flow event via sock_ops- TCP already has
BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB and
BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB. I propose adding analogous
callbacks for UDP- BPF_SOCK_OPS_UDP_FIRST_SEND and
BPF_SOCK_OPS_UDP_FIRST_RECV- that fire once per unique (source,
dest, sport, dport) tuple within a configurable window. This gives
flow-level visibility without per-packet tracing overhead. I'll
present the proposed semantics, the kernel-side implementation
sketch, and the estimated overhead based on prototype measurements. - Conntrack event reliability at scale — Even for flows that conntrack
does capture, we lose 5-8% of ctnetlink events under high connection
rates (>50K new flows/sec) due to netlink socket buffer overflow.
Speaker: Tarun Shekher - Why UDP flow observability is structurally harder- TCP connections:
-
10:00
Networking Performance Regression Testing 30m
The Linux network stack is a performance-critical component of the Linux kernel that impacts service delivery, quality, and efficiency. Subtle changes might have far-reaching performance implications. In this presentation I describe an emerging project to build a reference setup for networking performance and efficiency experiments. The objectives for this project are twofold. The immediate goal is establishing a reference software stack and suite of benchmarking experiments that can be used for automatic regression testing during the regular development cycle of the Linux networking subsystem (netdev). On the other hand, having an established and meaningful reference setup wouldd also be highly beneficial for academic researchers to compare wider-ranging and ambitious research proposals with a well-known baseline that follows best practices for system and workload configurations. As such, it is hoped that a secondary benefit of such an initiative is bringing together Linux developers and academic researchers to benefit the entire open-source operating systems community.
The talk will present existing preliminary work in this project and solicit feedback, and most importantly it can hopefully serve as a starting point for a wider discussion and consultation. Given the diversity of workload scenarios, potential system configurations, and possible performance effects, conversations about the most representative experiments and metrics are critically important. Aside from fundamental questions, such as the specific nature of regression detection, there are practical questions in how to best integrate a test instance with the existing netdev infrastructure for test automation (NIPA) and how to best facilitate the replication of test instances for different purposes.
Another explicit objective for this project is to study system efficiency, including resource efficiency, in additional to pure performance. As computing infrastructure becomes a more and more significant energy consumer and hardware improvements might potentially slow down in the near- to mid-term future, it is important to prepare for a potentially resource-constrained future by being able to understand, measure, control, and reduce the resource overheads associated with various computing services.
Speaker: Martin Karsten (University of Waterloo) -
10:30
BIG TCP for UDP tunnels 30m
BIG TCP is a kernel feature that allows aggregating SKBs bigger than 64k, aiming to reduce per-packet overhead for high-throughput network traffic. Until now, it has mostly been practical for direct-routing deployments. A large class of production environments, including Kubernetes+Cilium setups, relies on UDP-based overlay networks, such as VXLAN and GENEVE.
This talk will walk through the challenges of adding BIG TCP support to encapsulated traffic, such as having to deal with multiple variations of IPv6 HBH extension header in every driver, which resulted in unification with BIG TCP IPv4 and dropping the HBH header from IPv6 too. This talk will also walk you through other gaps that prevented BIG TCP from working with UDP tunnels out of the box, as well as other related parts like adding support in tcpdump, and possible caveats. The performance numbers on Mellanox NIC will be included, showing gains with 1.5k and 8k MTU, as well as the software GSO paths when tunnel offloads are unavailable.
This effort resulted in two patchsets on LKML: "BIG TCP without HBH in IPv6" and "BIG TCP for UDP tunnels". The first part solves the issue with convoluted packet parsing in the fast path on driver level. The second part addresses the remaining gaps to handle encapsulated BIG TCP SKBs correctly, and switches to using UDP length = 0 for oversized aggregated packets and restoring the real length when parsing or segmenting such packets. It also takes care of checking the length of untrusted ingress packets.
The goal of this session is to gather feedback from users of overlay networking and driver authors for possible future improvements and evolution.
Speakers: Alice Mikityanska (Isovalent at Cisco), Ms Alice Mikityanska (Isovalent at Cisco) -
11:00
TCP Multipath implementation with BPF 30m
As server platforms increasingly ship with multiple frontend NICs, there is a growing need to utilize all available network paths transparently — without application changes. We propose a BPF-based approach that leverages cgroup hooks at connect() time to perform client-side path selection across multiple NICs, effectively implementing a form of TCP multipath at the host level. To avoid the operational complexity of distributing and synchronizing per-host NIC topology maps across an entire fleet, we further propose encoding the multi-NIC topology information directly in TCP Options during the connection handshake. This allows peers to dynamically discover available paths and perform load balancing decisions locally using BPF programs. We will present the design, discuss how it interacts with dynamic container networking environments, and share early deployment experience.
Speaker: Raman Shukhau -
11:30
Coffee Break 30m
-
12:00
MPTCP KTLS Support: Bringing TLS to Multipath TCP 30m
Multipath TCP (MPTCP) enables a device to simultaneously utilize multiple network interfaces for a single connection, providing enhanced redundancy, resilience, and bandwidth aggregation. Currently, an increasing number of network applications are beginning to support MPTCP, and some of them, such as NVMe over MPTCP and web servers like lighttpd, require TLS for MPTCP. However, MPTCP has historically lacked TLS support due to fundamental conflicts in how both protocols manage the Upper Layer Protocol (ULP) stack within the Linux kernel.
This proposal presents a solution to make MPTCP and TLS compatible. The main challenge is the tight coupling between TLS and TCP, which hardcodes many TCP-specific functions and conflicts over the ULP slot. It implemented MPTCP support on the TLS side by introducing a protocol operations abstraction layer that decouples TLS from TCP and implements MPTCP-specific functions, and enabled TLS configuration on the MPTCP side by applying TLS encryption exclusively at the non-fallback MPTCP socket level, rather than on individual subflows (TCP sockets). The implementation resolves this challenge, passes the newly added MPTCP TLS selftests, which mirror the existing TCP TLS test cases, and has been validated through extensive selftests and real-world use cases, including NVMe over MPTCP.
Link: https://lore.kernel.org/all/cover.1782123118.git.tanggeliang@kylinos.cn/
Speaker: Geliang Tang -
12:30
Native AF_VSOCK API for userspace device emulation 30m
Several userspace VMMs emulate virtio-vsock devices entirely in userspace, and vhost-user-vsock provides a similar capability through the vhost-user protocol, to reduce the kernel attack surface or to run on platforms where vhost-vsock is not available. However, only the vhost-vsock kernel module, which emulates the virtio-vsock device in the host kernel, integrates directly with
AF_VSOCK, so these backends expose Unix domain sockets to host applications, following the hybrid-vsock model introduced by Firecracker.In the hybrid-vsock model, to connect to a guest, a host application opens a Unix socket and sends a text command like
"CONNECT <port>\n". To receive connections from a guest, it listens on a separate Unix socket named/path/to/uds_<port>. This works, but it requires every host application to implement this ad-hoc protocol instead of using standardAF_VSOCKsockets, and the guest CID is not visible to the host kernel at all.This talk explores a kernel interface that would let any process emulating a vsock device in userspace, such as the VMM itself or a vhost-user backend, register a guest CID and use
AF_VSOCKsockets instead of Unix sockets, keeping the same one-socket-per-connection model that hybrid-vsock already uses. One possible approach is to introduce a new protocol type forAF_VSOCK(or a socket option) that allows a process to claim a CID throughbind()and inject guest connections into the host'sAF_VSOCKstack through the standardaccept()andconnect()calls, while the device emulation remains entirely in userspace. Host services would then reach these guests with a simpleconnect(AF_VSOCK, guest_cid, port), exactly as they do with kernel vhost-vsock, and VMMs would need minimal changes to switch fromAF_UNIXtoAF_VSOCK. Because this interface operates at theAF_VSOCKlevel, it is not tied to virtio and could serve any existing or future vsock transport.Topics for discussion include: the right kernel API design, whether there are existing interfaces to follow or reuse, and how it should interact with network namespaces recently supported by
AF_VSOCK.Speaker: Stefano Garzarella (Red Hat) -
13:00
Possibility of Userspace Networking Drivers: TUNTAP, VDUSE and What's More? 30m
Building a driver for a new NIC normally means writing a kernel module. Whether you upstream it or maintain it out of tree, it requires a lot of effort. Isn't there a lighter-weight approach?
This talk shows one way to write the driver logic in userspace, without writing a kernel module, while the device still shows up as a Linux network device like "eth0". This is possible with existing in-kernel mechanisms today, but it is not yet a complete solution. Let's discuss what is required for a userspace driver mechanism.
Moving driver logic to userspace reduces the pain of maintaining out-of-tree kernel modules: you stop keeping up with kernel version changes and a critical bug in the driver does not crash the kernel. Other subsystems have similar concepts: FUSE for filesystems, ublk for block devices. The networking subsystem also has building blocks for processing networking in userspace, and they are divided into two groups.
One group keeps the network stack in kernel and hands packets off to userspace below the driver. TUNTAP is a classic example. It's a virtual device used for tunneling and virtual machines, and exchanges packets with userspace rather than with a wire. VDUSE is another example: it lets you implement a vDPA device in userspace, intended as a userspace backend for virtual networks of virtual machines and containers; combined with virtio-vdpa, it provides a function similar to TUNTAP.
The other group does the opposite: VFIO and UIO provide only a thin device-handling layer in the kernel and move the implementation above the driver, including the network stack, out to userspace.
By combining the two groups, and placing the actual driver logic in a userspace daemon between them, we can realize a NIC driver in userspace. One of its use cases is an FPGA-based NIC, which acts as a fully programmable NIC. Because its CPU-facing interface -- descriptor formats, register layouts, and so on -- is entirely up to the FPGA programmer, the driver variations are virtually unlimited.
I have been experimenting with building userspace NIC drivers with TUNTAP/VDUSE and VFIO. For testing, I used an e1000e device emulated in a VM as a common simple device rather than special hardware. Unsurprisingly, there is a performance penalty due to packet copies, high CPU load of the userspace daemon, and so on. Both TUNTAP and VDUSE are also limited by their specific interfaces, such as statistics and netdev features. I'll introduce the various difficulties I faced and discuss how userspace network drivers should be built. Are TUNTAP/VDUSE and VFIO the right shape, or does Linux networking want its own dedicated mechanism? Is it possible to implement a driver's data path in BPF?
Speaker: Toshiaki Makita (NTT, Inc.) -
13:30
Lunch Break 1h 30m
-
15:00
Modernizing NFS Direct I/O for PCI Peer-to-Peer DMA 30m
High-performance storage environments increasingly rely on direct data movement between PCIe endpoints, such as NVMe Controller Memory Buffers (CMB) or BAR memory. Today, NFS relies on standard page references (get_page) and doesn't support pinning pages for DMA, making the subsystem pin-unaware. Furthermore, the legacy iterator APIs used in NFS are incapable of passing the required capability flags mandated by the GUP subsystem. Thus, NFS remains incompatible with P2PDMA due to this lack of pin-awareness.
This talk walks through the design challenges in enabling P2PDMA for the NFS Direct I/O path using a two-phased approach. The first phase involves an architectural modernization, i.e. migrating the Direct I/O path to use the modern iov_iter_extract_pages API. This transition introduces pin-awareness to the NFS request lifecycle. Additionally, it migrates the NFS Direct I/O from pages to folios. We have an on-going series posted upstream for the first phase:
https://lore.kernel.org/all/20260603053033.3300318-1-praan@google.com
Plumber's Discussion: Transport-Level P2PDMA Capability Detection
The second phase of this modernization addresses the detection and propagation of P2PDMA capabilities based on the underlying transport. Since NFS connections can dynamically migrate across different hardware interfaces or utilize multipathing (nconnect), P2PDMA support is a property of the active connection rather than a static server capability. This session focuses on the following architectural challenges:
- Detection & Propagation: Discuss & align on the proposed design for transport-level capability discovery. This would allow NFS to conditionally signal hardware P2PDMA support (viaITER_ALLOW_P2PDMA) to the GUP subsystem, potentially on a per-request basis.
- Architectural Implications: Discuss the cross-subsystem impact of this transport-centric approach. Key areas include managing state consistency across transport failover, the implications of dynamic device rebinding during re-connections, and the challenge of transport selection in multipath (nconnect) environments.
The talk is backed by the ongoing NFS Modernization and P2PDMA series upstream:
- RFC: https://lore.kernel.org/all/20260401194501.2269200-1-praan@google.com/
- On-going Series: https://lore.kernel.org/all/20260616134000.2733403-1-praan@google.com/Speakers: Pranjal Shrivastava, Mr Shivaji Kant -
15:30
swiotlb Under Pressure: High-Speed NICs in Confidential VMs 30m
Systems that treat devices as untrusted protect themselves using confidential computing memory protection. Linux uses swiotlb to handle the necessary buffer bouncing transparently to drivers. For high-performance NICs operating at scale, however, the overhead of allocating and managing bounce buffers can substantially exceed the cost of the memory copies themselves.
Receive queues create long-lived pressure on the swiotlb pool. Pages are DMA-mapped when they enter a page_pool and normally remain mapped while they are recycled, often until they leave the pool or the queue is destroyed. Across many receive queues, these mappings can retain a substantial portion of the available bounce-buffer memory, even after the pool has been increased to the largest practical size for the system. As occupancy increases, allocations search additional swiotlb areas and contend on their per-area locks. This primarily affects transmission, where an SKB’s linear area and fragments can each require a separate bounce-buffer allocation, copy, synchronization, and reclamation sequence.
To evaluate an alternative to per-mapping swiotlb bouncing, we implemented a prototype mlx5e datapath based on preallocated DMA-coherent staging memory. With 16 transmit-heavy TCP streams, one per queue, the swiotlb datapath reduced throughput by 70% relative to a baseline without buffer bouncing. The explicit staging datapath reduced this loss to 10–20%.
In this talk, we will compare the generic swiotlb approach with the driver-managed bounce buffers, present detailed measurements, and analyze the bottlenecks we identified. We will conclude by discussing what an upstream solution should look like: driver-specific staging buffers, a generic networking abstraction or improvements to swiotlb scalability.
Speaker: Dragoș Tătulea (NVIDIA Corporation) -
16:00
KNOD: In-Kernel GPU Offload for Packet Processing 30m
At LPC 2025, we presented an early prototype for running XDP programs directly on an AMD GPU from within the Linux kernel, without CUDA, ROCm, or any userspace component in the data path.1
Since then, the project has evolved into knod, an in-kernel network offload device, and an RFC patch set has been posted.2 knod uses the GPU as a kernel-managed packet-processing accelerator: the kernel JIT-compiles an XDP program into GPU machine code, the NIC places packets directly in GPU-accessible memory, and the GPU processes them in parallel and returns XDP verdicts. The RFC extends the original XDP prototype into a generic offload-device model, with RX IPsec as an additional use case.
Major changes since LPC 2025 include a redesigned NIC-to-GPU path that allows the GPU to consume packets directly, support for divergent control flow in GPU-offloaded BPF programs, RDNA2 support, and optimizations to batching, dispatch, occupancy, and memory transfers. The current RFC reports up to 70 Mpps for a Katran-derived XDP workload and 80 Gbit/s for RX IPsec.
This talk will walk through how knod has evolved since LPC 2025 and discuss the remaining design and implementation challenges: the boundaries among the networking core, bpf, drm, and accelerator drivers; per-CPU map and queue-affinity semantics on accelerators; further performance optimization; support for additional GPU architectures; and the Generic Netlink/YNL control plane and its integration with existing tools.
Speakers: Hoyeon Lee (SUSE), Taehee Yoo (LINE Investment Technologies) -
16:30
Coffee Break 30m
-
17:00
PCI data path optimization with CQE coalescing and new ethtool parameters 30m
To optimize the PCI data path on the Azure MANA NIC, we implemented support for Completion Queue Entry, or CQE, coalescing. In the regular receive path, each packet can generate a separate CQE that must be processed by the driver, increasing interrupt pressure and CPU & PCI overhead. CQE coalescing reduces this cost by merging multiple CQEs into one aggregated completion that represents several received network packets. This allows the driver to process more packets per interrupt and reduces the amount of per-packet overhead required in the fast path. The improvement is especially important on ARM64 platforms, where interrupt handling cost and cache efficiency can strongly affect networking throughput. In our measurements, CQE coalescing provided a throughput enhancement ranging from 33% to 62% while also increasing the number of packets handled per interrupt.
We also added two new parameters to both the ethtool kernel API and the ethtool command line: ETHTOOL_A_COALESCE_RX_CQE_FRAMES and ETHTOOL_A_COALESCE_RX_CQE_NSECS as suggested by netdev maintainers. These parameters expose CQE coalescing through the standard Linux networking configuration model instead of relying on a driver-specific private flag. The frame-based setting controls how many packet completions may be combined before a coalesced CQE is generated, while the time-based setting controls how long the device may wait before reporting the aggregated completion. Together, they allow administrators to tune the balance between throughput, CPU utilization, interrupt rate, and receive latency for different workloads. Other NIC drivers that support similar CQE coalescing behavior can adopt the same interface, improving consistency across Linux networking devices.
Speaker: Dr Haiyang Zhang (Microsoft) -
17:30
Coroutines for Linux kernel 30m
There are multiple kernel drivers which could benefit were it possible to somehow do coroutines in C. Coroutines would make these drivers easier to comprehend, extend with new code paths, and debug.
The most common use case lies in simplifying complex state machines. For example consider the main state machine
sfp_sm_main()indrivers/net/phy/sfp.c. Each invocation executes code for a specific state, updates the state, and returns. The current implementation even relies on goto statements to jump between cases. As a result, understanding the control flow requires tracking disjointed code paths or drawing manual state diagrams.Using coroutines, this logic can be transformed into a straightforward, linear sequence using standard loop and conditional structures (
if,for,while). By leveraging GNU C computed gotos, we can implement mechanics similar to C++20'sco_awaitandco_return. With some trickery, it is even possible to preserve variable values across suspension points.Let's explore hiding this underlying complexity behind clean, maintainable macros and discuss whether this approach can be safely implemented for production code.
Speaker: Marek Behún
-
09:30
-
10:00
→
18:30
Birds of a Feather (BoF) "Club D" (Prague Congress Centre)
"Club D"
Prague Congress Centre
53-
10:00
BoF: Memory Efficiency on Modern Linux Systems 45m
Discussion on how organizations are managing the complex tradeoffs related to efficient use of memory resources in heterogeneous workload, large-scale datacenter environments. Especially, as DRAM prices have exploded.
What guidance and future feature development should the community provide to support efficient use of this resource?
- zswap
- proactive reclaim agents
- DAMON
- PSI
- memory tiering
- memcg
- etc.
Speaker: Mykolas Krupauskas (Uber Technologies Inc.) -
10:45
Toward efficient device HBM 45m
As modern foundation models and their deployments demand exponentially larger memory capacity, high-bandwidth memory (HBM) has become an unprecedented driver of datacenter capital expenditures (CapEx). While the majority of discussions surrounding memory tiering have focused on leveraging disk swap, in-memory compression and CXL.mem to expand system capacity and reduce the total cost of ownership (TCO) of host DRAM, HBM on AI accelerators presents a distinct yet critical frontier for TCO optimization.
This proposal explores the feasibility and architectural requirements for extending the traditional Linux memory management (MM) concepts, such as hot/cold tracking, page migration and data compression, to accelerator-attached HBM. We evaluate existing kernel mechanisms to discuss how much infrastructure can be reused and brainstorm with the audience to seek new opportunities.
Speaker: Mr Yu Zhao (Google) -
11:30
Tea Break 30m
-
12:00
RV32 Linux BoF: Yes, we have hardware, and it's alive! 45m
While many major IC and IP vendors are sunsetting 32-bit platforms, 32-bit RISC-V has emerged as a vibrant new haven.
As discussed at LPC last year, RV32 needs a well-defined profile and readily available, robust hardware platforms for CI/CD testing. Back then, the lack of accessible, mass-market silicon was a significant hurdle.Now, with Espressif releasing Linux support for the ESP32-S31 (https://github.com/espressif/linux), we have living proof that RV32 Linux is alive and kicking.
Come check out the hardware, play around with it, and let’s chat about where we go from here!
Speaker: ChuanTzu Tsai (Andes Technology) -
12:45
SPDM/TSM/CTLV 45m
Continue discussion of how the CMA spdm (potentially plus TSM) flows work through to userspace and on to remote attestation.
If no room is available we'll find a space.
Speakers: Alistair Francis, Dhaval Giani, Dr Jonathan Cameron (Qualcomm) -
13:30
Audio developers meeting 45m
A meetup for people working on audio. Can we have a BoF slot scheduled in some room in the lunch break on Wednesday please?
Speaker: Mark Brown -
15:00
Let's talk about the GPL and Linux! 45m
A time to discuss all things copyleft and the kernel
Speaker: Denver Gingerich (Software Freedom Conservancy) -
16:30
Tea Break 30m
-
10:00
-
10:00
→
13:30
Build Systems MC "Club B+C" (Prague Congress Centre)
"Club B+C"
Prague Congress Centre
100The Linux ecosystem supports a diverse set of methods for assembling complete, bootable systems—ranging from binary distributions to source-based systems, embedded platforms, and container-native environments. Despite differences in tooling and architecture, all of these systems face shared challenges: managing build complexity, ensuring security and reproducibility, maintaining cross-platform compatibility, and responding to increasing regulatory and supply chain scrutiny.
Building on the success of last year’s microconference, we invite the community to continue the conversation with a broadened scope in 2025. This year, we aim to explore the intersection of build systems with CI/CD pipelines, supply chain security, critical infrastructure, secure development practices, and the potential use of machine learning and AI techniques to improve build systems, CI/CD pipelines, and supply chain analysis. With legislation such as the Cyber Resilience Act, rising expectations for Software Bill of Materials (SBOMs), and mandates for reproducible and auditable builds, collaboration across the ecosystem has never been more essential.
This microconference provides a venue for architects, maintainers, and practitioners from all facets of the Linux build and distribution ecosystem to come together and share ideas, discuss pain points, and identify potential shared solutions.
Target communities and projects include (but are not limited to):
- General-purpose distributions: Debian, Fedora, Ubuntu, Arch Linux, openSUSE, Red Hat
- Source-based systems: Gentoo, NixOS, Guix, CRUX
- Embedded platforms: Yocto Project, OpenEmbedded, Buildroot, OpenWRT/LEDE, Android
- Container ecosystems: Docker, Podman, OCI, BuildKit, distrobuilders
- Immutable, image-based distributions: Flatcar, ParticleOS, Fedora Silverblue, Talos
- RTOS and hybrid build systems: Zephyr, RIOT, Mbed OS, FreeRTOS
- CI/CD and build orchestration: BuildStream, Buildbarn, Bazel, Jenkins, GitLab CI, GitHub Actions
- ML/AI-assisted tooling and infrastructure: anomaly detection, dependency analysis, build optimization, and CI/CD intelligence systems
- Compliance and supply chain security: SPDX, OSI, SBOM tooling, sigstore
- Broader open-source infrastructure efforts and standards bodies
Proposed discussion topics:
- Bootstrapping build systems and managing cross-compilation
- Integration of CI/CD pipelines into build workflows
- Securing the build lifecycle: from developer systems to package publication
- SBOM generation, license auditing, and legal/policy alignment
- Attestation, signing, and ensuring software chain-of-trust
- Handling insecure or volatile upstream language-specific ecosystems (e.g., PyPI, npm, crates.io)
- Reproducible builds and deterministic output across toolchains
- Secure and scalable container build systems and image validation
- Immutable build pipelines for image-based systems and update strategies
- Applying ML/AI to build systems (failure prediction, caching strategies, scheduling, test selection)
- Anomaly detection in build pipelines and supply chain events
- Dependency analysis, risk scoring, and automated patch or update prioritization
- Resilience in build infrastructure for critical systems and edge deployments
- Patch sharing, lifecycle tracking, and cross-distro patch coordination
- Documentation, onboarding, and reducing the learning curve of complex build systems
- Long-term sustainability: mentoring, diversity, and community health of build toolchains
We welcome proposals beyond this list, particularly those that address emerging issues in the creation, validation, maintenance, and secure delivery of Linux-based software systems.
Improving coordination across build systems strengthens the foundations of the open-source ecosystem. Whether you’re maintaining a distro, building firmware, managing containers, applying intelligent systems to improve build and release workflows, or designing infrastructure for high-assurance or real-time systems, this microconference is your forum to advance the state of Linux software construction and security.
-
10:00
The Supply Chain Attack on Buildsystems 30m
Modern embedded Linux build systems such as Yocto and Buildroot rely on complex pipelines that reuse intermediate artifacts and external inputs. While this improves performance and reproducibility, it also creates opportunities for supply chain attacks that are difficult to detect.
This talk demonstrates practical attack vectors targeting build systems at different stages of the pipeline. First, we show how intermediate object files produced by make in Buildroot can be tampered with during the build process, resulting in compromised binaries without modifying upstream source code. Next, we explore how shared state (sstate) cache artifacts in Yocto can be poisoned, allowing malicious code to propagate across builds through trusted cache reuse.
We demonstrate these attacks end-to-end by injecting malicious behavior into binaries and executing them in a final Linux image under QEMU, illustrating how such compromises can evade traditional verification mechanisms.
We then present mitigation strategies, focusing on signing and verification of sstate artifacts to establish trust in reused build outputs. Integration points and tradeoffs between security, performance, and reproducibility are discussed.
This session provides a practical look at real-world build system attacks and concrete techniques to harden them. It is aimed at developers and maintainers of build systems, embedded Linux distributions, and toolchains who are concerned with supply chain security.
Speaker: Alejandro Hernandez Samaniego -
10:30
OpenWrt is reproducible! How we got there and what we learned 30m
OpenWrt has been fully reproducible for a few months now, the culmination of many years of work to achieve this important milestone. In this talk we'll discuss how we got here, what it took, and tips for other build systems that are looking for the same, including how to handle the unique challenges of reproducibility across multiple cross-compilation targets.
We'll also go into some of the issues specific to embedded build systems, which platforms took the longest, and some of the rationale for doing it at all. Thanks to reproducibility, we are much better positioned to tackle a variety of regulations and new requirements on router manufacturing, which we'll address in detail as well.
Join us for an enlightening discussion on the hows and whys of making a community-built distribution reproducible, and please bring your questions and thoughts from your own experiences and/or other build systems.
Speaker: Denver Gingerich (Software Freedom Conservancy) -
11:00
Relevance of LTS distro releases and the GenAI world 30m
Most of Linux distros who follow time based releases, do have LTS release policy e.g. Ubuntu, Yocto, Debian, buildroot to name a few, and then there are rolling releases like archlinux and its family of distros. This talk is to discuss the LTS in the wake of genAI coding agents. There is a fair bit of coding agents at work for yocto project and other distributions doing different functions from CI/CD to generating patches for package upgrades. GenAI coding agents is resulting in increased amount of patches and pace for package upgrades and general patching around distributions. This is going to increase the maintenance load for LTS dramatically in next 2-3 years. We need to rethink the utility of LTS releases and seeing if rolling release model is more viable option for the distributions especially embedded linux distributions which have smaller communities to maintain them.
We could use genAI tooling to aid in maintaining LTS releases as well. However, is that the best choice going forward or do we focus on mainline and rolling release model
Speaker: Khem Raj (Qualcomm) -
11:30
Coffee Break 30m
-
12:00
Bitwise-Reproducible Kernel Builds for the WhatsApp TEE — Proof of Concept 30m
A build is bitwise reproducible when compiling the same source with the same configuration and toolchain yields byte-for-byte identical output — an identical vmlinux and bzImage, verifiable by a simple sha256sum. For a normal kernel this is a hygiene property; for a Trusted Execution Environment it is foundational. A platform that measures the code it boots and reports a cryptographic hash is only meaningful if the expected value can be independently reproduced i.e. if an auditor or WhatsApp itself, or an external reviewer can take the published source and configuration, rebuild the kernel, and arrive at exactly the same hash the hardware attests to. Without bitwise reproducibility there is an unverifiable gap between "the source we published" and "the binary that is running," and the attestation degrades from a proof into a promise. Reproducibility closes that gap and lets a TEE's trust chain be checked end-to-end by anyone. The central difficulty is that a stock kernel build is not deterministic by default, it embeds build timestamps, the building user and host, git-derived version strings, an archive of kernel headers, and ephemeral module-signing keys all of which vary run to run and perturb the final hash even when no source changed. The engineering problem is therefore to systematically identify and eliminate every source of non-determinism until two independent builds collapse to a single hash.
The strategy we used to obtain a proof of concept was divide and conquer, subsystem by subsystem. The final vmlinux/bzImage hash is just a deterministic function of all the compiled object files linked into it clearly if every .o in every subsystem compiles identically, the whole-kernel hash is necessarily identical. That observation turns an intractable "the whole image differs" problem into a localizable one. Rather than chase the top-level hash blindly, the approach hashes every object file in the tree after each of two builds and then diffs the two manifests. The resulting diff lists exactly which objects and subsystems still differ between builds. Each divergent object is a concrete lead pointing at a specific source of non-determinism; fix that source, rebuild, re-diff, and the list shrinks. Reproducibility is reached when the diff is empty. This breaks the problem down into a tractable, iterative hunt that pinpoints where in the build the divergence originates instead of guessing at the image as a whole.
The work proceeded as many rebuild-and-compare cycles against a copy of the previous build (t1/, build1/), each cycle narrowing the set of differing objects. Recurring offenders and their fixes
included:
- kheaders.o (CONFIG_IKHEADERS), embeds a freshly-archived, timestamped snapshot of the kernel headers that never matches between builds; disabled for the POC.
- Version strings — CONFIG_LOCALVERSION_AUTO appends a git-derived suffix; disabled (LOCALVERSION_AUTO off, empty LOCALVERSION/BUILD_SALT) and the generated version files under init/ and arch/ were pinned.
- Module signing (CONFIG_MODULE_SIG) — signs modules with a per-build ephemeral key, guaranteeing divergence; disabled (captured as a standalone remove_module_signing change).
- Debug/introspection artifacts — CONFIG_GDB_SCRIPTS disabled to drop non-deterministic generated content.
- Build environment — pinned via KBUILD_BUILD_TIMESTAMP, KBUILD_BUILD_USER=builder, KBUILD_BUILD_HOST=buildhost, and SOURCE_DATE_EPOCH, so timestamps and identity strings are fixed rather than sampled from the build machine. The build environment here was a standard Meta devserver, however, to deploy this we would build and deploy an image for a build environment.The scope of the proof of concept was to establish a clean baseline, the POC was built against Linus's mainline branch (the vanilla upstream tree, origin/linus-upstream) rather than the internal Meta proprietary TEE kernels, isolating the reproducibility problem from the additional out-of-tree patches carried by the production kernel. With the fixes above, two independent builds produced an empty object-level diff every one of the ~4,800 tracked objects matched and identical sha256sum/md5sum values for both vmlinux and arch/x86/boot/bzImage. This demonstrates that a bitwise-reproducible kernel is achievable for the target and validates the subsystem-by-subsystem manifest-diffing method as the path to extend reproducibility from vanilla upstream to the full WhatsApp TEE production kernel.
Speaker: Joshua Lilly (Meta) -
12:30
The last step to secure reproducible distribution kernels: Hash-based module integrity checking 30m
The kernels current module signature scheme does not work well together with reproducible builds. If the key is generated at build-time the build is not reproducible. A static key that is known to the public does not provide security, but a static key not known to the public does prevent public rebuilds of the kernel for validation purposes.
Currently distributions need to make a tradeoff:
* Allow non-public reproducibility through static module signature keys. At the cost of complexity in build infrastructure and package definitions. (Debian)
* Focus on reproducibility and disable module signatures (NixOS)
* Focus on security and have an irreproducible kernel package (ArchLinux)I am proposing the replacement of module signatures for built-in modules with a merkle tree calculated at build-time.
This solves the problems outlined above and also avoids a lot of the complexity involved in the current implementation.Agenda:
* Problem statement
* Current proposal
* DiscussionDiscussion topics:
* How to strip modules as part of this scheme?
* Problems in the integration with kbuild.
* How to interact with IMA?
* Would it make sense to have a generic and reusable merkle tree implementation?Current series on LKML: https://lore.kernel.org/lkml/20260505-module-hashes-v5-0-e174a5a49fce@weissschuh.net/
LWN article for the previous implementation: https://lwn.net/Articles/1012946/Speaker: Thomas Weißschuh (Linutronix) -
13:00
What does "buildable" mean during a toolkit migration? 30m
Sugar has been in tree since April 2006. Its last toolkit transition, GTK2 to GTK3, ran from October 2011 to the 0.98 release in November 2012. The current one, GTK3 to GTK4 and X11 to Wayland, is in its second year across twelve repositories. I ported the toolkit and presented that work at GNOME Asia Summit 2025; this year I mentor the two contributors porting the sugar shell and the activities set we call Fructose. The stack is deliberately half-migrated and will stay that way for some time.
The code side of a toolkit migration is documented. The build side is not, and it has cost this project more time than the porting.
Nothing detects that a set of GObject Introspection typelibs disagree about a major version. The shell requires eleven GI namespaces. One is the project's own C helper library, whose installed typelib was built against GTK3, loaded into a process that already holds GTK4. The process dies at import. In C this is a link error. In GI there is no build-time or package-time gate, and no way to ask whether a set of typelibs can share one process, even though the information is in the typelibs. We fixed it by bumping our own namespace from SugarExt-1.0 to SugarExt-2.0 so the two cannot be confused, which works and is what I would tell the next project to do. Nothing identified it as the problem. It took the first five weeks of a contributor's summer to find. It will recur for any GI-based project that migrates a toolkit while carrying its own introspected library.
The versions you have decides what you can test, and nobody has written that down. The shell puts a Wayland compositor inside a widget (Casilda), so activities run as Wayland clients inside the shell. Casilda 1.2.1 is the last version that builds against wlroots 0.19. Version 1.2.2 moved to wlroots 0.20, and 1.2.4 raised the GTK requirement to 4.22.2. Our pinned nixpkgs gives us GTK 4.20.3 and wlroots 0.19.2. So a pin two layers away from our own code decides which Casilda we can have, and that decides what a reviewer can run. A contributor on Debian and a reviewer on Nix could not reproduce each other's bugs.
Even then, a git checkout is not something you can build. Some files only exist in a release tarball: a config module built from a template, and compiled GSettings schemas. You have to make those by hand. One runtime dependency is packaged by no distribution. And the tree carries three build systems at once: autotools from 2006, Meson in the C library, and a pyproject in the new toolkit.
Questions I would like to put to the room:
- Should typelib agreement be a build-time or package-time check? A typelib carries its dependency list version-qualified, readable without loading it. Debian lintian already requires that a package declare a strictly versioned dependency on the typelibs it needs. What I could not find is anything that takes a set of typelibs and answers whether they can share one process. Does that exist, or is there a reason not to build it?
- How should a build system or a distribution represent a stack that is deliberately half-migrated, where the right answer to "which version" differs per repository and changes weekly?
- What do you hand a new contributor so they can build a multi-repository stack mid-migration?
- Is there anything for coordinating the patches distributions each end up carrying during an ecosystem-wide toolkit transition?
Reference material:
- Shell migration, 188 files: https://github.com/sugarlabs/sugar/pull/1106
- C helper library, autotools to Meson and GDK4 event porting: https://github.com/sugarlabs/sugar-ext/pull/6
- Toolkit migration talk, GNOME Asia Summit 2025: https://www.youtube.com/live/WZ63lQ-DsOA?t=14725
- GTK4 migration guide: https://docs.gtk.org/gtk4/migrating-3to4.html
Speaker: Krish Pandya (Undergraduate Researcher)
-
10:00
→
13:30
Containers and checkpoint/restore MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53The Containers and Checkpoint/Restore micro-conference focuses on both userspace and kernel related work.
The micro-conference targets the wider container ecosystem ideally with participants from all major container runtimes as well as init system developers.
The microconference will be discussing recent advancements in container technologies with some of the usual candidates being:
- VFS API improvements (new system calls, idmap, …)
- CGroupV2 feature parity with CGroupV1 and migration path
- Dealing with the eBPF-ification of the world
- Mediating and intercepting complex system calls
- Making user namespaces more accessible
- Verifying the integrity of containers
- Improving the set of resource limits available
On the checkpoint/restore front, some of the potential topics include:
- Making CRIU work with modern Linux distributions
- Handling GPUs
- Restoring FUSE daemons
- Dealing with restartable sequences
- Use of eBPF
- Support of new kernel features
- Supporting shadow stack (x86, arm64)
- Support for madvise(MADV_GUARD_INSTALL)
- Support for mseal()
- Support for pidfd C/R, including process exit information
And quite likely a variety of other container and checkpoint/restore topics as things evolve between now and the event.
Past editions of this micro-conference have been the source of many developments in the Linux kernel, including:
- PIDfds
- VFS idmap (and adding it to a slew of filesystems)
- FUSE in user namespaces
- Unprivileged overlayfs
- Time namespace
- A variety of CRIU features and checkpoint/restore kernel interfaces with the latest among them being
- Unpriviledged checkpoint/restore
- Support of rseq(2) checkpointing
- IMA/TPM attestation work
-
10:00
Constraints of process migrations 20m
Triggered by:
Subject: [PATCH 0/4] bpf: add a few hooks for sandboxing
Message-Id: 20260220-work-bpf-namespace-v1-0-866207db7b83@kernel.orgProblem statements:
- Users (admins) are sometimes confused by some entity (PAM, systemd, container
runtimes) migrating their processes away from intended cgroup.
- Coarse-grained DAC doesn't express well who (migrating process) can operate
on what (cgroup) to what (migrated task).
- Limited immutability of membership assignment after certain point.Proposed solution:
BPF LSM hook for cgroup_attach_permissions
(combination with other existing migration vetting mechanisms
association of permissions with PIDs instead of UIDs?)Alternate solutions:
- stick with regular cgroup FS permissions
- utilization of other existing LSM hooksSpeaker: Michal Koutný (SUSE) -
10:20
Memory Tracking in forensic checkpointing: what soft-dirty bits can’t tell us 20m
CRIU’s incremental checkpointing is being used for forensic container snapshot(Stoyanov et al., DFRWS 2026) chains. Soft-dirty tracking cuts snapshot size by about 10× and makes high-frequency capture practical. Live migration only needs a correct final state. Forensics needs the path that led there. Soft-dirty was built for migration and is now being reused for forensics.
A forensic snapshot chain starts with one full snapshot, then stores only later changes. Hash-chaining makes tampering detectable, as long as the host kernel and CRIU stay trusted.
This talk brings the discusion of the challenges with the existing memory tracking mechanism and improvements for forensic use.
Soft-dirty answers one question: did this page change since the last reset? It cannot say when, how many times, in what order, or by whom. Tracking makes pages read-only and catches the first write to each but a page written once and a page written ten thousand times are indistinguishable. Every intermediate value is gone, along with any ordering or context. Even when two pages change, the kernel walks page tables.
Some ways to improve this tracking for forensic checkpointing at container level, have been inspired from several of live migration, hypervisor specific and hardware trapping techniques. One such direction might be maintaining a kernel level dirty set of pages. PAGEMAP_SCAN already helps here, but CRIU mainly uses it as a faster way to read soft-dirty. The tracking mechanism is unchanged, and the kernel-side still walks the range, there is no maintained dirty set, only a faster way to report the outcome of a walk.
Another direction can be to finally start utilising the first-write page fault data rather than discarding almost everything it exposes. Today’s interfaces force a choice: async uffd-wp is cheap but reports one bit, while sync uffd-wp can report richer data at a userspace round trip per fault. Forensics needs rich-and-cheap: capture at the fault, kernel-side, without that round trip.
A separate, lighter option is to recover changed bytes by diffing a captured page against its parent and storing only the difference.
At last, soft-dirty is not broken. It does what live migration needed, well enough that forensic snapshot chains already use it. The problem is that forensics asks a different question, and the interface has no answer for it. Upstream work so far has made the answer cheaper to retrieve without making it richer.
Speaker: Shailja Shaktawat -
10:40
Enabling Incremental Checkpointing for GPU Workloads 20m
With the increased adoption of AI workloads, efficient GPU checkpointing mechanisms are becoming crucial for inference, training, fine-tuning, and reinforcement learning workloads. One of the key challenges with GPU checkpointing today is the lack of memory-tracking support that enables incremental snapshots. When the GPU state is checkpointed into host memory, all pages appear modified, preventing CRIU from identifying which pages have changed since the previous checkpoint and resulting in full snapshot of the GPU memory for every iteration. In this talk, we will discuss extending the GPU plugins for CRIU with support for memory tracking that enables efficient incremental checkpointing. We will explore the benefits of this approach and the trade-offs between performance overhead and storage efficiency across different GPU workloads.
Speaker: Radostin Stoyanov (University of Oxford) -
11:00
dm-qcow2: device-mapper-based QCOW2 storage for containers 30m
Container storage commonly relies on directory overlays, filesystem-native subvolumes, or thin-provisioned block devices. We will explore another approach: exposing QCOW2 images directly as Linux block devices through a device-mapper target. QCOW2 is the standard virtual-disk format across much of the QEMU/KVM ecosystem. Its widespread adoption, mature tooling, and features such as backing-file chains, persistent bitmaps, and sparse allocation make it an attractive option for container storage as well.
But just having a nice loop device is not a complete solution. Container storage must also support essential operations such as snapshots, backups, and migration. This talk will show how we map these requirements onto existing device-mapper and Linux kernel capabilities, what is still missing, and which problems remain unsolved.
Speaker: Andrei Zhadchenko (Virtuozzo) -
11:30
Coffee Break 30m
-
12:00
Upgrade restrictions for file descriptors 20m
For quite a while there has been a wish from container runtime to be able to restrict how we can reuse a particular file descriptor. For instance, CVE-2019-5736 showcased a privilege escalation in runc, whereby the possibility of reopening
/proc/self/exeas writeable allowed a malicious image to overwrite the runc binary. That was patched on the user space side by copying runc to a sealed memfd_create() file before executing the target. However, it would be better if we could just set restrictions on how some file descriptor can be used (via procfs or O_EMPTYPATH) to reopen a file as writeable in the first place. Thus we would like to have some kind of upgrade mask for file descriptors.Effort has been spent in the past to implement something like this. For instance, by David Drysdale in 2013 and later by Aleksa Sarai as part of openat2(2) in 2019. I want to take some time to summarize some of these past ideas and see why they were unsuccessful, and what makes this in particular a hairy problem. We can finish with some discussion about what a merge-able patchset would look like.
Speaker: Jori Koolstra (N/A) -
12:20
Checkpoint/Restore of Device Cgroup eBPF Programs 20m
After moving OpenVZ containers to cgroup-v2 we are struggling a bit to reach
feature parity with what we had before. One such feature is running nested
Docker containers inside an OpenVZ (system) container — part of making our
containers behave as close to a regular server as possible.In cgroup-v2 the device controller was reformed drastically: device
availability can only be controlled by special BPF_PROG_TYPE_CGROUP_DEVICE
programs attached to cgroups. Docker naturally relies on this, and systemd also
employs it for its own services — so a migrated container without these
programs comes back with its device policy silently dropped. The problem is
that the BPF interface is asymmetric: the kernel accepts a program, but gives
no way to get it back in a reloadable form. The verified/JITed instructions it
can report are not portable — the verifier rewrites context accesses and helper
calls into offsets and addresses specific to the running kernel, so they can
neither pass verification again nor work on another kernel. Mainstream solved
the same problem for seccomp a decade ago — commit f8e529ed941ba ("seccomp,
ptrace: add support for dumping seccomp filters") added an API to retrieve
loaded filters specifically for C/R — but for general BPF programs no such interface
exists to this day.We took the seccomp approach for cgroup device programs in the Virtuozzo
kernel: keep a copy of the original, pre-verification instructions at load time
and report it through BPF_OBJ_GET_INFO_BY_FD in a new
bpf_prog_info::orig_prog_insns field. The encoding (struct bpf_insn) and the
CGROUP_DEVICE context are stable UAPI, so on restore the program is simply
reloaded and the destination kernel re-verifies and re-JITs it. On top of this
API, CRIU now dumps programs and their per-cgroup attachments (via
BPF_PROG_QUERY), the bpf-prog anon-inode fds held by processes, and on restore
re-attaches everything once the cgroup tree is recreated, preserving the
original sharing topology.In this talk we will go over the kernel and CRIU sides of the design and
discuss whether such an interface could be accepted in mainstream, along with
the open problems on the way to generic BPF checkpoint/restore: original
instructions for arbitrary program types, maps and their contents, links, and
pinned objects.Speaker: Pavel Tikhomirov -
12:40
Addressing Challenges for Container Migration in Heterogeneous Clusters 20m
Checkpoint/Restore (C/R) is increasingly used for both startup acceleration (restoring pre-warmed snapshot instances on demand) and live migration. However, deploying static snapshots or migrating tasks across heterogeneous clusters creates severe runtime bottlenecks when source and target nodes possess differing CPU capabilities. While CRIU and container runtimes can accurately preserve memory state, handling CPU feature consistency and safe architectural state restoration across diverse hardware remains an unsolved problem at the kernel boundary.
This session will focus on two critical kernel/userspace interaction challenges:. While CRIU and container runtimes can accurately capture and restore process states and kernel resources, handling CPU feature consistency and safe architectural state restoration across diverse hardware remains an unsolved problem at the kernel boundary.
This session will focus on two critical kernel/userspace interaction challenges:
1. HWCAP Inheritance & Feature Discovery: Examining feature detection failures when restoring snapshots on target nodes with different CPU features, and reviewing proposed mechanisms to inherit or mask hardware capabilities (HWCAP/HWCAP2 via auxv) across execve().
2. Restoring Extended Signal Frame States: Analyzing edge cases where tasks contain in-flight signal frames with architecture-specific CPU state on their stack. We will discuss why rigid kernel-side frame validation during rt_sigreturn causes restoration failures across hardware generations, and propose flexible validation strategies that prevent state corruption while preserving ABI safety.Goal: Align kernel, container, and language runtime maintainers on kernel-assisted CPU feature control and flexible signal-context restoration to make snapshot-based fast-booting and live migration robust across heterogeneous fleets.
Speaker: Andrei Vagin -
13:00
mmtest benchmarking of cgroup code 20m
Triggered by:
a) various rstat fixups as well as people reporting issues with (memory).stat reading vs writing performance & precision
b) occasional reports/attempts to make container startup quickerProblem statements:
a) cgroup rstats need to balance latency requirements of updaters (writers) and readers while preserving sufficient precision. The improvement for one may cause some deterioration at other places -- which is not always clear until those places start being friction points.
b) the cgroup creation path is on critical path of new container starts but there're no good insights into the behavior of those.Proposed solution:
mmtests shellpack and monitor(s)Alternate solutions:
- selftest, LTP tests
- bpftrace probes
- perf-benchSpeaker: Michal Koutný (SUSE)
-
10:00
→
13:50
Rust MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128Rust is a systems programming language that is making great strides in becoming the next big one in the domain. Rust for Linux is the project adding support for the Rust language to the Linux kernel.
Rust has a key property that makes it very interesting as the second language in the kernel: it guarantees no undefined behavior takes place (as long as unsafe code is sound). This includes no use-after-free mistakes, no double frees, no data races, etc. It also provides other important benefits, such as improved error handling, stricter typing, sum types, pattern matching, privacy, closures, generics, etc.
This microconference intends to cover talks and discussions on both Rust for Linux as well as other non-kernel Rust topics.
Possible Rust for Linux topics:
- Rust in the kernel: status updates and discussion on next steps.
- Use cases for Rust around the kernel: subsystems, drivers, other modules...
- Developing-related discussions: how to abstract existing subsystems safely and API design, coding guidelines, safety guidelines...
- Upstreaming process: guidance on how to get into mainline, strategies that have worked for Rust code in the past, getting involved...
- Maintenance: the new subentries and branches, the proposed cross-subsystem subteams (e.g. the safety team), scaling work for the future, any cross-subsystem issues...
- Infrastructure: build system, documentation, testing and CIs, maintenance, unstable features, architecture support, stable/LTS releases, Rust versioning, third-party crates...
- klint.
- pin-init.
- The future of GCC builds.
Possible Rust topics:
- Language and standard library: discussion on upcoming features, stabilization of the remaining features the kernel needs, memory model, the 2024 edition...
- Compilers and codegen:
rustcimprovements, LLVM and Rust,rustc_codegen_gcc,gccrs... - Other tooling and new ideas: Coccinelle for Rust,
bindgen, Compiler Explorer, Cargo, Clippy, Miri... - Educational material.
- Any other Rust topic within the Linux ecosystem.
Please remember that submissions for microconferences (like the Rust MC) should be discussion oriented. Please see "The Ideal Microconference Topic Session".
Last year was the 4th edition of the Rust MC. We had technical discussions around Rust abstractions for the kernel (Overflowing with Fear: Detecting and Mitigating Implicit Panics in Rust, External locking for internally synchronized data structures, Tackling challenges with HID and related device driver support in Rust, Exploring a real life RCU use case for Rust), as well as a presentation and discussion around Rust language features needed by the kernel (Rust language evolutions for better kernel developer experience) and about a Rust-based kernel extension framework (Rex and its integration with Rust-for-Linux). In addition, we had a tutorial session again (Initialization in Rust with pin-init). Finally, we also had a "Birds of a Feather" slot (Rust for Linux Office Hours) for open discussion on other topics.
Suggested attendees: the Rust for Linux team (Miguel Ojeda, Boqun Feng, Gary Guo, Benno Lossin, Andreas Hindborg, Alice Ryhl, Trevor Gross, Danilo Krummrich, Daniel Almeida, Tamir Duberstein, Alexandre Courbot, Onur Özkan), Abdiel Janulgue, Alexei Starovoitov, Alistair Francis, Arnaldo Carvalho de Melo, Bjorn Helgaas, Burak Emir, Christian Brauner, Christian Schrefl, Dave Airlie, David Gow, Dirk Behme, Fiona Behrens, Frederic Weisbecker, FUJITA Tomonori, Greg Kroah-Hartman, Igor Korotin, Ingo Molnar, Jocelyn Falempe, Joel Fernandes, Julia Lawall, Kees Cook, Liam R. Howlett, Lorenzo Stoakes, Luis Chamberlain, Lyude Paul, Masahiro Yamada, Matthew Maurer, Nathan Chancellor, Paolo Bonzini, Paul E. McKenney, Peter Zijlstra, Remo Senekowitsch, Rob Herring, Robin Murphy, Sami Tolvanen, Stephen Boyd, Tathagata Roy, Tejun Heo, Thomas Gleixner, Viresh Kumar, Will Deacon, Yury Norov...
-
10:00
Creating self references safely 30m
Self-reference is a common need in kernel code. In fact, this is what motivates the development of
pin-init. So far, self-references can only be created with unsafe code with explicit use ofOpaque. This is a discussion about on-going working to support safe creation of self references in thepin-initcrate.Speaker: Dr Gary Guo (Red Hat) -
10:30
Reworking `Request` reference counting in the Rust block device driver API 30m
The
kernel::block::mq::Requesttype [1] sits on the I/O hot path of every Rust block device driver. ARequestis jointly referenced by the block layer and the driver, with completion arriving on multiple asynchronous paths, so the type has to encode a non-trivial sharing and lifecycle model with minimal runtime cost.The introduction of
Ownable[2] gave us a general mechanism for types that oscillate between owned and reference-counted forms, and applying it toRequestcleaned up parts of the API [3]. However, the resulting scheme remains hard to reason about for reviewers [4]. We would like to use an LPC session to walk through the proposed changes with the wider Rust-for-Linux audience to further the review process.To anchor the discussion, we plan to bring the following content to the session:
- A walk-through of the reworked reference counting scheme.
- Benchmark results from
rnullthat quantify the cost of the scheme on representative I/O workloads.
The goal of the session is to surface concerns from the community, iron out pain points in the API shape, and build shared understanding of why the scheme looks the way it does — so that when the next version hits the list, the basic design is already broadly understood and accepted.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/rust/kernel/block/mq/request.rs?h=v7.1-rc5#n24
[2] https://lore.kernel.org/r/20260224-unique-ref-v16-0-c21afcb118d3@kernel.org
[3] https://lore.kernel.org/r/20260216-rnull-v6-19-rc5-send-v1-5-de9a7af4b469@kernel.org
[4] https://lore.kernel.org/r/87qzopwttw.fsf@kernel.orgSpeaker: Mr Andreas Hindborg (Samsung) -
11:00
dma_fence abstractions: Design and Challenges 30m
The kernel's dma_fence subsystem lays at the heart of every graphics processing unit (GPU) driver. It is a primitive for synchronizing the state of jobs running on GPUs with receiver parties, notably userspace. A number of circumstances make the correct implementation and usage of both C and Rust dma_fence very challenging:
- The highly asynchronous nature of GPUs, including the fact that they can hang and need to be reset.
- The fact that fences can have an arbitrary number of consumers, both in other drivers and in userspace.
- Various, partially optional, callbacks exist, with which a consumer can run into the code of a producer, whose module might unload at any time.
Since GPUs can directly access system memory, an incorrect or racing representation of GPU job state by DmaFence could result in memory corruption regardless of Rust's memory safety guarantees.
Moreover, dma_fences have so far not only been involved in various UAF and refcounting bugs, but are also often involved in deadlock conditions. While making memory bugs impossible was the primary design goal for the Rust abstractions, much attention was also paid to preventing deadlock.
In 2026, a shared design, development and upstreaming effort by various parties, notably the Nova and Tyr GPU drivers, has seen much progress. This talk shall give an overview over the general design, solved and persisting problems, and special challenges with Rust regarding these abstractions.
Speaker: Philipp Stanner -
11:30
Coffee Break 30m
-
12:00
Tyr: A Status Update 25m
Briefly cover the status of the Tyr project and discuss the current blockers in upstream, specially those related to missing Rust abstractions. This presentation intends to discuss and validate the job submission model, including the proposed GPUVM/JobQueue Rust abstractions and their current upstream status, showcasing the different designs between Tyr's initial implementation and what is actually likely to land on mainline.
Speaker: Daniel Almeida (Collabora) -
12:25
Using Rust for out-of-tree kernel drivers 20m
Rust is expanding into more and more places, and it's becoming clear that Rust creates some unique challenges when it comes to drivers that are out-of-tree.
Like all other Rust drivers, out-of-tree drivers written in Rust require abstractions for the subsystems they interact with. If the driver requires a subsystem that does not yet have abstractions, or if the abstractions exist but are missing some part of the API, then the driver may have to implement the abstraction directly within the driver. However, it can be tricky to implement abstractions (or extend existing abstractions) from outside of the kernel crate.
In this topic I would like to discuss approaches for tackling these issues, and also discuss ways in which we are unnecessarily making it more difficult to extend abstractions from drivers.
Speaker: Alice Ryhl (Google) -
12:45
Bridging the C/Rust Divide: Replicating the PWM Subsystem's Success 20m
Problem Statement:
Independent developers driving new Rust bindings often face a major bottleneck: getting their work mainlined by hesitant C subsystem maintainers, and just as importantly, sustaining that collaboration post-merge. Translating modern Rust architectures to maintainers who evaluate designs strictly through C paradigms remains a massive hurdle.Session Focus:
The PWM subsystem recently proved to be a refreshing exception. Using the TH1520 SoC as a hardware target, I successfully upstreamed new PWM bindings by iterating extensively on RFCs and working closely with a supportive C maintainer.However, the work doesn't stop at the initial merge. The goal of this working session is to use the PWM experience to establish a repeatable playbook for cross-language collaboration. I will start by presenting concrete solutions that worked for PWM, and then open the floor to brainstorm how to replicate this success across the kernel.
Discussion Points:
After a brief (10-minute) overview of the specific strategies that worked during the PWM review process, we will spend the rest of the session brainstorming solutions to the following:-
Technical Mapping & The "C Lens": Concrete strategies for mapping established C expectations (callbacks, structs) to modern Rust traits. How do we design and explain these API boundaries so they intuitively click for a C maintainer and reduce their review burden?
-
Post-Merge Collaboration: How to maintain momentum after the bindings land. I will share my experience coordinating abstraction and driver fixes via IRC with the C maintainer. We will debate strategies for getting maintainers heavily involved and comfortable reviewing the Rust side on an ongoing basis.
-
"Marketing" the Abstractions: Once the initial bindings and driver are merged, how do we effectively market these new abstractions to the wider community? We will brainstorm how to encourage other developers to write new Rust drivers for the subsystem.
https://mwilczynski.dev/posts/bringing-rust-to-the-pwm-subsystem/
Speaker: Michał Wilczyński -
-
13:05
BPF CO-RE support in Rust 25m
Writing BPF programs in the long past meant wrestling with Linux kernel version fragmentation. That problem was solved, many years ago, thanks to CO-RE (Compile Once, Run Everywhere) relocations. CO-RE is a mechanism that uses the BTF type format and its relocation entries (
BTF.ext) to handle layout differences, by patching the loaded BPF bytecode with correct offsets that match the running kernel version.However, as of today, only C compilers (Clang, GCC) are able to emit these relocations, while the Rust ecosystem has been missing that final piece for 100% developer experience parity.
In this talk, we'll dive into the design and experimental implementation of native CO-RE support in the Rust compiler — from the
#[btf_relocatable]attribute andcore::btffield-info macros, through compiler lowering to thellvm.bpf.preserve.field.infointrinsic, to the finalBTF.extemission. We'll also discuss the trade-offs behind requiring explicit relocation queries rather than making ordinary field projection relocatable.Speaker: Michal Rostecki (Anza)
-
10:00
→
13:30
Scheduler and Real-Time MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128Building upon the success of last year, we propose a combined microconference focused on the Real-Time and Scheduler subsystems. These two areas are fundamentally intertwined and continue to drive cross-cutting changes, especially following the upstream integration of PREEMPT_RT. The Linux scheduler is central to overall system performance. Addressing the challenges of modern computing—from achieving low latency to maximizing high throughput across diverse topologies and workloads, and scaling from small, power-constrained devices to large-scale HPC systems—is key to delivering the optimal user experience.
Since last year’s microconference, progress has been made on the following topics:
- Cache aware scheduler
- Paravirt Scheduling: Framework for better physical CPU utilization
- CPU Isolation and IPI interference
- Push callback for fair scheduler
- Runtime verification
- Proxy executionDiscussions on certain topics were also carried forward at the OSPM 2026 conference.
Ideas of topics to be discussed include (but are not limited to):
- Responsiveness of fair tasks
- Improve PREEMPT_RT
- Locking and priority inversion
- Improve SCHED_DEADLINE
- CPU isolation
- New topology, including hybrid or heterogeneous system
- Tooling for debugging low latency analysisThis is not an exhaustive list. We welcome all proposals related to process scheduling.
The goal is to discuss open problems, preferably with patch set submissions already being discussed on the mailing list. Presentations are meant to be limited to 2 or 3 slides intended to seed a discussion and debate - allowing for high bandwidth discussion with key stakeholders in the same room.Key attendees:
- Ingo Molnar
- Peter Zijlstra
- Juri Lelli
- Vincent Guittot
- Dietmar Eggemann
- Steven Rostedt
- Ben Segall
- Mel Gorman
- Valentin Schneider
- K Prateek Nayak
- Thomas Gleixner
- John Stulz
- Sebastian Andrzej Siewior
- Shrikanth Hegde
- Phil Auld
- Dhaval Giani
- Clark Williams- 10:00
-
10:02
Enable runtime modification of nohz_full and managed_irq housekeeping CPUs 22m
By using the cpuset isolated partition functionality in the Linux
kernel, users are now able to change the set of "isolcpus[=domain]"
HK_TYPE_DOMAIN housekeeping CPUs at runtime. This feature is used by
some Kubernetes based container orchestration platforms to enable the
creation of containers running latency sensitive workloads like DPDK.Domain isolation by itself doesn't provide enough CPU isolation, so the
system has to be booted with a set of boot-time enabled nohz_full and
managed_irq isolated CPUs which are then combined with runtime domain
isolation to create the desired isolated CPUs for workloads that need
them. There are some wasted system overhead to have boot-time isolated
nohz_full and managed_irq CPUs that are not actually being used in
isolated cpuset partitons.Eventually we would like to have nohz_full and managed_irq isolated
CPUs created at runtime when they are needed. This talk is about the
progress we have made in this direction and additional future works
that are needed to achieve this goal.Speaker: Mr Waiman Long (Red Hat) -
10:24
Bringing Proxy Execution’s benefits to Userland / Futexes 22m
More and more we’re seeing issues around priority (sometimes called performance) inversion of SCHED_NORMAL/BATCH tasks. Particularly if any sort of constraints are put on “background” deprioritized tasks. These background tasks will eventually grab an important lock, and then won’t be constrained and prevented from running for some extended period of time, resulting in all the important tasks becoming blocked waiting for them to release the lock. Since the tasks are fair tasks, the delays are not indefinite, but they still can be substantial and user visible.
PI Futexes seem like a good solution here, but the underlying rt_mutex behavior doesn’t help SCHED_NORMAL tasks as rt_mutex priority inheritance isn’t used between SCHED_NORMAL tasks since ~v5.10 or so. Further, even on systems where kernels patch rt_mutexes do “nice inheritance” (as imperfect as that is) for NORMAL tasks, the strict rt_mutex handoff behavior results in pi-futex not performing as well as normal futexes.
Thus, there is a desire to preserve the performance characteristics of normal futexes, while also providing priority inheritance to avoid priority/performance inversions.
Proxy Execution has been a feature in development for many years now, which provides semantics similar to what is desired in this case. Proxy Execution works for in-kernel mutexes (and rw_sems), and provides priority-inheritance without the strict handoff ordering and performance overhead that rt_mutexes can cause. For RT tasks, rt_mutexes and their strict behavior is still important, but for other sched-classes Proxy Execution provides much of the benefits without the costs.
So it would be nice to similarly enable futexes to benefit from Proxy Execution’s generalized form of priority inheritance and avoid the negative performance impact of PI futexes.
Unfortunately, the existing normal futex API is insufficient to be used with Proxy Execution, as when a task is blocked on a futex lock, Proxy Execution has to understand who the owner of that lock is, so they can be run to release the needed lock. Thus we need to find a way to extend the normal futex API so that the lock owner is communicated to the kernel.
Proxy Execution is not yet fully upstream, so this discussion is a little premature, but since uAPI is important to get right, we wanted to start the discussion early so we can plan and work toward an acceptable solution.
Speakers: John Stultz (Google), Suleiman <> Souhlal (Google) -
10:46
PREEMPT_RT: How a simple interface up event might break your realtime 22m
Realtimers know: Resource configuration for RT enabled systems is key. RT applications rely on a proper system configuration and expect this configuration to stay unmodified as long as the application is alive.
As it turned out, Linux might silently reconfigure IRQ affinities of network adapters in a way that violates the expected system configuration. While we run into this kind of problem in the area of networking first, the
problem seems generic to all kind of "multi queue devices".We ended up having two problematic situations:
- RT applications (partially) breaking out of their previously requested system configuration as Linux ignores previously requested affinities during reconfiguration
- non RT applications polluting isolated (RT) CPU cores via IRQs
And there is even more flexibility on the horizon already: The proper system
configuration might not be known at boot time. Using container runtimes for the deployment of RT applications requires an adaptive resource management that might change IRQ affinities at runtime.In this session we will showcase a couple of shortcomings related to resource management in the networking area. Starting with how network drivers are
violating CPU isolation by implementing IRQ spreading, we want to discuss- How do we close the gap between cgroups (cpuset controller) and the IRQ core?
- How should we adjust all the affected drivers to honor the cgroup configuration?
- What APIs are necessary to get there?
Speaker: Mr Florian Bezdeka (Siemens AG) -
11:08
Understanding differences in results between timerlat and cyclictest 22m
The proposal is to discuss differences in results between timerlat and cyclictest and the reasons behind that. Once the reasons are understood, ways to reduce the differences could be sought. In addition, the discussion could include rtla feature parity with rt-tests, as well as between different rtla tools (osnoise, timerlat), and integrating rtla with other tools such as perf.
Speakers: Tomas Glozar (Red Hat), John Kacur, Luis Goncalves (Red Hat) -
11:30
Coffee Break 30m
-
12:00
Remaining EEVDF scheduling latency and lag improvement 22m
Recent improvements have been made in tasks ordering, scheduling latency and lag of the fair/EEVDF scheduler but some issues still remain and will involve more complex mechanisms. The talk will discuss possible solutions for a number of open issues:
How to fix the remaining out of range lag of tasks ?
How to further decrease the scheduling latency ?
How to select the best CPU to minimize scheduling latency?Speaker: Vincent Guittot (Linaro) -
12:22
Kernel Sched QoS Interface 22m
Following last year discussion about userspace assissted scheduling [1], schedqos utility is announced [2] and a proposal for kernel interface to help address one of the QoS issues folks see in the wild, DVFS and migration latencies, was sent also [3].
The interface proposal is a simple extension to sched_attr, but has the goal of being extensible without requiring further addition to sched_attr.
Some of the key areas to discuss:
- Deprecatability: We want to end up with baggage ABI we don't want to carry forward if internals change and a QoS no longer make sense.
- Extensibility: We want to allow almost arbitrary description of behavior without having to modify the interface.
- Discoverability: how can users know a set of QoS are available?
Currently there are two proposed users, rampup_multiplier to manage DVFS response time, and tag memory dependency between tasks for cache aware scheduling.
The latter had interesting discussions about cookie management and whether ownership should be in kernel space or userspace.
We will go through all open questions and hope to come up with a plan of what to do next.
[1] https://lpc.events/event/19/contributions/2089/
[2] https://lore.kernel.org/lkml/20260415000910.2h5misvwc45bdumu@airbuntu/
[3] https://lore.kernel.org/lkml/20260504020003.71306-9-qyousef@layalina.io/Speaker: Mr Qais Yousef (Google) -
12:44
Energy-Aware Scheduling on x86 Hybrid Topologies: Latency Analysis and Idle CPU Selection 22m
During OSPM 2026, there was an agreement to stop requiring schedutil to enable EAS on x86 platforms, which implies EAS may soon be active by default on these systems.
Energy-Aware Scheduling (EAS) places tasks by comparing estimated energy costs across performance domains, directing light tasks to power-efficient cores on asymmetric systems. The known trade-off is that EAS packing introduces runqueue contention when idle CPUs are available: tasks are packed onto fewer CPUs, and wakeups that land on already-busy CPUs incur queuing delay even when other CPUs sit idle.
This talk uses quantitative wakeup latency decomposition on x86 hybrid platforms as a starting point to motivate making EAS more aggressive in selecting idle CPUs. I examine where in find_energy_efficient_cpu() idle-CPU preference can be strengthened without abandoning the energy-efficiency objective, measure the resulting impact on tail latency (p99+) and energy consumption via platform power counters, and discuss how the latency–power trade-off compares to the packing baseline.
Speaker: Ricardo Neri (Intel Corporation) -
13:06
sparsemask: the missing generic plumbing and next step forward 22m
As CPU counts grow, Linux scheduler scalability suffers from contention on global cpumasks — frequent atomic updates to shared cachelines become a measurable bottleneck on large core-count systems.
Two proposals address this: Steve Sistare's sparsmask, which distributes a cpumask across multiple cachelines to reduce contention [1], and Peter Zijlstra's sbm (sparse bitmap) [2], a simpler, topology-aware evolution of the same idea. While sbm effectively eliminates cacheline ping-pong, its generic infrastructure falls short in one critical area: CPU hotplug handling, where topological information for offline CPUs may be unavailable.
This talk covers:
- The gap: shortcomings in the generic sbm layer for hotplug scenarios
- Peter's proposal for x86 to address the challenges with offline CPUs, and what's still needed to make sparsemask truly generic
- A scheduler-native alternative: an orthogonal sbm implementation scoped to the scheduler, leveraging hotplug callbacks to dynamically size sbm allocations.
Previous version of this work was posted at [3] and was discussed at LPC2025.
References:
[1] https://lore.kernel.org/lkml/1541767840-93588-2-git-send-email-steven.sistare@oracle.com/
[2] https://lore.kernel.org/lkml/20260324120008.GB3738010@noisy.programming.kicks-ass.net/
[3] https://lore.kernel.org/lkml/20251208083602.31898-1-kprateek.nayak@amd.com/Speaker: Prateek Nayak (AMD Inc.) -
13:28
Wrap-up 2mSpeaker: Vincent Guittot (Linaro)
-
15:00
→
18:30
Confidential Computing MC "Club E" (Prague Congress Centre)
"Club E"
Prague Congress Centre
128Confidential Computing MC
Over the last few years, the Confidential Computing microconferences at LPC have been a key driver in advancing support for trusted execution workloads across the Linux virtualization and software ecosystem.
As a result of the previous confidential computing microconference, the following major features were merged:
- SEV-SNP support
- TDX support
- TDISP Infrastructure
- SEV-TIO support
- SVSM guest-side support
The microconference at LPC serves as the key in-person event for Linux-related developments, as well as an important platform for standardizing confidential computing features across different platforms.
The open source software stack for confidential computing is still far from being complete. There remain many problems to be solved and functionality to enable. Some of the most important ongoing developments are:
- Enhancements to CVM memory backing via guest_memfd
- KVM Support for ARM CCA
- Privilege separation features in KVM
- CVM live migration.
- Secure VM Service Module architecture and Linux support
- Trusted I/O software architecture
Further topics to discuss are:
- Solutions for the full CVM (remote) attestation problem
- Linux as a CVM operating system across hypervisors
- CVM Performance
The Confidential Computing microconference of 2026 wants to bring open source developers and industry experts together into productive discussions and to collaborate on solutions for the open problems.
Key attendees:
- Ashish Kalra ashish.kalra@amd.com
- Borislav Petkov bp@alien8.de
- Dan Williams dan.j.williams@intel.com
- Daniel P. Berrangé berrange@redhat.com
- David Hansen dhansen@linux.intel.com
- David Kaplan David.Kaplan@amd.com
- David Rientjes rientjes@google.com
- Dhaval Giani dhaval.giani@gmail.com
- Elena Reshetova elena.reshetova@intel.com
- James Bottomley James.Bottomley@HansenPartnership.com
- Joerg Roedel joro@8bytes.org
- Jon Lange jlange@microsoft.com
- Michael Roth michael.roth@amd.com
- Mike Rapoport rppt@kernel.org
- Paolo Bonzini pbonzini@redhat.com
- Peter Fang peter.fang@intel.com
- Peter Gonda pgonda@google.com
- Sean Christopherson seanjc@google.com
- Stefano Garzarella sgarzare@redhat.com
- Tom Lendacky thomas.lendacky@amd.com-
15:00
Welcome to the Confidential Computing Microconference 10m
Dhaval and Joerg welcome the attendees.
Speaker: Dhaval Giani -
15:10
Improving TDX Linux integration: changes in TDX Migration and Attestation 20m
Live migration and runtime attestation are two critical features for Confidential Computing to get right, not just in terms of fulfilling customer requirements, but also correct integration into common Linux codebase and overall code simplicity and maintainability. Based on Linux community feedback, TDX architecture went through a few big changes last year wrt to Live Migration and Attestation to enable simpler Linux integration.
In this short talk we will share the key changes, as well as our key learnings in this area that could be useful for other architectures enabling these confidential computing features in Linux. We would also like to hear feedback from the community on other areas where such simplifications can make a great difference.
Speaker: Elena Reshetova (Intel) -
15:30
Arm CCA Realm Live Migration 30m
Arm is developing Live Migration ABIs for the Arm Confidential Compute Architecture (CCA), with input and requirements from ecosystem partners. The ABIs are provided by the Realm Management Monitor (RMM), the trusted firmware component responsible for managing Realms.
The presentation will start with a short overview of the high-level CCA Live Migration design and the end-to-end process. It will then examine the assumptions and design choices underpinning the ABIs in three areas: core design, platform preconditions, and runtime mechanisms, relating each to the broader security, correctness, and practicality constraints of confidential live migration.
Finally, the presentation will provide a basis for discussing how the emerging CCA Live Migration ABIs and Linux/KVM infrastructure for confidential-computing migration can inform one another.
Speaker: Mathias Brossard (Arm) -
16:00
CoVE in Action: the Landscape of Confidential VMs on RISC-V 30m
RISC-V CoVE (Confidential VM Extension) brings confidential computing — hardware-enforced isolation of tenant workloads from the hypervisor and cloud operator — to a fully open architecture.
A lightweight TEE Security Manager (TSM) sits below the hypervisor and enforces per-VM memory isolation, while tenants verify their environment through a standard IETF RATS attestation flow. The hypervisor retains scheduling control — it simply cannot see inside a tenant's TEE Virtual Machine (TVM).
We demonstrate CoVE Deployment Model 1 running end-to-end on real RISC-V hardware, built on OpenSBI and integrated with Kata Containers and the confidential-containers project — bringing confidential computing to container runtimes operators already use. No proprietary extensions, no special hardware — just standard RISC-V silicon.Speakers: Ruoqing He (LingCage), Mr Xiaoxia Cui (Damo) -
16:30
Coffee Break 30m
-
17:00
SEV-TIO implementation challenges 20m
SEV-TIO is quite known by now, the upstream development is split in stages and continues. The current stages are set to support basic functionality.
The talk will focus on extended features and how AMD hardware/firmware is going to implement these. This includes:
- huge pages handling in IOMMU and KVM, how RMP works and what TMPM does in PSMASH_IO.
- IOMMU TLB flushing challenges: a hack for Turin CPU family and a proper fix in the Venice family (new behavior or invlpgb) and what Venice will do in addition to that.
- Cbit and vTOM and how we can allow simultaneous private and shared access for the same device within the same VM + a (useless) way to implement “iommu=pt”-like behavior in the guest.
- sev-guest SNP platform device and using the DMA layer for sharing memory with the hypervisor.Speaker: Alexey Kardashevskiy (AMD) -
17:20
Extending Arm CCA Device Assignment to CXL Type-2 Accelerators 30m
We are extending Arm CCA device assignment to CXL Type-2 accelerators so a Realm VM can receive such a device with a coherent device memory window. Standard RME-DA assigns a PCIe TDI to a Realm through TDISP, SPDM and IDE, protecting CXL.io traffic. A CXL Type-2 device also exposes a coherent CXL.mem window through HDM decoders and which confidential assignment must also secure.
RMM v2.0 introduces a PDEV Stream object as a first-class handle for the device security channel and a COH_CMEM stream type for coherent off-chip accelerators - the right primitives to extend CCA to CXL Type-2. Several design problems remain unsolved: how the host TSM obtains and registers the CXL coherent memory range with RMM, how the guest TSM attests that range, and how to coordinate device probe with stream connect so the coherent range is always registered with RMM regardless of call ordering. This session will work through these open questions to align on a direction for upstream development.
Speaker: Ankit Agrawal -
17:50
Deferring SWIOTLB Bounce to Reduce Heavy CPU Utilization 20m
Coco VMs rely on bounce buffering to use the guest's vCPUs to encrypt and decrypt data for DMA. This memory copy is performed via the SWIOTLB bounce buffer.
Persistent disk read operations handle completions within their storage interface’s interrupt handler. When SWIOTLB is disabled, the interrupt handler is invoked only after the DMA data transfer is completed. This makes the handler only responsible for quick clean up operations. With SWIOTLB enabled, the read data is currently entirely copied by a vCPU during this interrupt handler which dramatically increases handling time.
The memory copy is expensive and should not be done within an interrupt handler since it prolongs the time interrupting other tasks. This problem is exacerbated by each storage interface’s interrupt handler being pinned to run only a specific vCPU. Heavy disk read workloads can cause these pinned vCPUs to spend the entire duration of the workload in the storage interface’s interrupt handler thus starving other tasks. Google has seen cases of these workloads causing softlockups on the pinned vCPUs.
I have a prototype which defers the SWIOTLB’s memory copy out of the interrupt handler and into a workqueue. This eliminates the softlockup problem. It also allows for the memory copy to be scheduled on other vCPUs; not just the ones pinned to the interrupt handlers. For VMs with a higher vCPU count, this approach increases bandwidth and IOPS. Although, deferring the work has the drawback of slightly decreasing bandwidth for VMs with a low number of vCPUs and slightly increasing latency regardless of vCPU count.
I would like to gather feedback for my approach and discuss questions I have regarding its enablement. Is this slight latency increase tolerable? Should this feature be enabled by default for VMs with forced SWIOTLB or should the user of the guest kernel be responsible for enabling it? If we decide to enable it by default, should we only enable it for VMs with a higher number of vCPUs? If we decide the user of the guest kernel is responsible for setting it, how should they be able to set it?
Speaker: Ryan Afranji (Google) -
18:10
PMU Event Filtering for Confidential Guests 20m
PMU Event filtering is a security feature that allows hypervisors to restrict which performance events guests can monitor, preventing potential side-channel attacks. However, when using hardware-acceleration PMU virtualization, the encrypted VMSA in SEV-ES and SEV-SNP creates a fundamental challenge: hypervisors cannot read or modify guest PMU states directly, which breaks PMC filtering with hardware-accelerated PMU virtualization and prevents confidential VMs from using upcoming AMD feature: Guest PMC Event Filtering.
This talk explores how PMC filtering can be restored for SEV-ES and SEV-SNP guests without weakening their isolation guarantees. The key idea is a cooperative model between guest and hypervisor that re-establishes the hypervisor's filtering authority even when it cannot touch the guest's encrypted state directly.
We will discuss the protocol design, implementation challenges, and how this approach restores hypervisor security control over guest PMC usage in confidential computing environments.
Speaker: Manali Shukla
-
15:00
→
18:30
Driver Core MC "Club B+C" (Prague Congress Centre)
"Club B+C"
Prague Congress Centre
100Driver Core Microconference focuses on general problems of the linux kernel driver model.
The goal is to discuss the various aspects and problems of device driver core, platform and auxiliary devices, subsystem architecture, firmware description, fw_devlink, API design and object life-time issues.
Current problems:
Object life-time issues and proposed solutions
Decades-long problem. There have been several attempts at addressing the multiple issues and a few previous talks at LPC and other conferences.
a) Revocable: there's an ongoing effort to provide a unified API for protecting resources against sudden removal of their dependencies. Example: unbinding a resource provider driver to which consumers still hold references. (current revision on the list)
b) I2C bus unbind path: Bartosz Golaszewski proposed a way for a gradual rework of I2C core in order to remove the wait_for_completion() call blocking the kernel thread removing the bus driver until all consumers put their references. Johan Hovold argued the rework can be more in-depth and result in a better outcome. It would be useful to discuss the current state. Work on this seems to be ongoing.. Previous email from Wolfram
Provider/consumer/device API best practices
When discussing the i2c changes, Johan and Bartosz argued about the approach to linux subsystem API design. It would make sense to discuss if there's an "idiomatic" way to design driver interfaces and what it should be. Discussion.
Firmware devlinks
a) Should fw_devlink be enforced for everyone? (Saravana Kannan floated the idea during the Devicetree MC at LPC 25).
b) Adding support for fw_devlink to software nodes. Bartosz Golaszewski is working on this to decrease the number of probe deferrals for software node GPIO lookups which are used extensively on some older but still maintained platforms as well as on x86 platform drivers for which devices are not well described in ACPI.
c) Ulf Hansson proposed a talk titled: "Evolving sync state support to other subsystems beyond genpd".
Device/subsystem API abuse
a) When documentation says one thing and users do another. How to deal with API abuse.
b) Patches using legacy APIs that can't always be spotted by relevant maintainers in time. How to better deprecate APIs.
Platform/auxiliary/faux buses
Let's discuss using dynamically instantiated, "virtual" devices for handling of various corner-cases. There have been several such changes in recent months, for example: in reset and GPIO subsystems.
Firmware node API and its implementations
a) There's an ongoing effort in the GPIO subsystem to remove the string-matching behavior of software node lookup. Let's discuss why attaching software nodes to target devices and enforcing real firmware node links when setting up references is better.
b) Dynamically referencing "real" firmware nodes (OF-nodes, ACPI nodes) from software nodes.
c) Using software nodes as primary firmware nodes for certain devices. This is already done for MFD cells and some auxiliary devices. Is allowing to match devices to drivers by software nodes something worth considering?
d) Converting subsystems to be "fwnode-agnostic". (Example rework)
Devres
It still appears that quite frequently developers seem to misunderstand how devres interfaces work and when exactly the unwinding of resources will happen. Patches are being sent where drivers try to schedule devres actions on devices they don't control or on ones that aren't even bound to drivers. This has led people to blame devres for life-time issues the culprit of which lies with incorrect usage of the API. Can we replicate the Bound context from rust in C?
Key attendees:
Confirmed: Geert Uytterhoeven, Laurent Pinchart, Ulf Hansson, Chen-Yu Tsai, Neil Armstrong, Krzysztof Kozlowski, Hans de Goede, Kevin Hilman, Greg Kroah-Hartman, Marek Vasut, Abel Vesa, Manivannan Sadhasivam, Herve Codina, Danilo Krummrich, Rafael J. Wysocki, Dmitry Baryshkov, Vinod Koul, Srinivas Kandagatla
Likely: Wolfram Sang, Conor Dooley, Dmitry Torokhov, Andy Shevchenko, Drew Fustini
-
15:00
Power Sequencing for Enumberable Busses - Driver Core Integrations? 25m
On x86 / ACPI platforms, devices on enumerable busses can normally be seen directly by the OS. On device tree platforms, these devices sometimes require extra power sequencing like toggling regulator supplies or GPIO lines. Over the years most of these cases have been solved, but there are still some gaps.
As of kernel version v7.0, support for power sequencing generic PCI devices, ones that have no extra toggles that are not part of PCI specification, and M.2 M-key slots is available [link]. v7.1 then adds support for the PCI part of E-key slots [link].
Support for the USB part of M.2 E-key slots is WIP by the author [link]. Support for onboard USB devices and USB type A connectors is provided
by the onboard device driver.This session intends to give a quick overview of the current status and discuss whether parts of this could be moved or integrated at the driver core level. These include:
- Bus code integration for acquiring power sequencers
- Creating stub platform devices to provide power control functionality
- USB onboard devices
- PCI pwrctrl devices
Such mechanisms could then be reused for the MDIO bus.
Speaker: Chen-Yu Tsai (Google, LLC) -
15:25
Reviving early platform drivers 25m
Certain critical subsystems - clocks, timers and interrupt controllers - sometimes need to be initialized before driver core is made available in driver_init(). To that end, we provide a set of macros: IRQCHIP_DECLARE(), CLK_OF_DECLARE(), TIMER_OF_DECLARE() which allow the kernel to call initialization functions based on compatibles either before reaching the point where actual platform devices matching these compatibles can be created or instead of creating them essentially bypassing the driver model entirely.
This results in the initialization routines not being able to use many APIs only available to real drivers and - for many modules - never registering with the driver core.
I've worked on the idea of unifying these paths with a concept of "early platform drivers" back in 2018. That work never got anywhere but the "hacky" approach for early setup remains.
I'd like to re-discuss the idea, it's pros and cons and provide possible solutions for the main contention point raised last time: the fact that my series did nothing to automatically convert existing invocations of the _DECLARE() macros to using early platform drivers.
[1] https://lore.kernel.org/all/20180511162028.20616-1-brgl@bgdev.pl/
Speaker: Bartosz Golaszewski (Qualcomm) -
15:50
Evolving support for sync_state to subsystems beyond genpd 25m
At last LPC in Tokyo we discussed about the limitations of the sync_state support that quite recently was added to the generic PM domain (genpd) subsystem. The conclusion was to mainly focus on making it more fine grained, as this should address most of the problems. Attempts to implement this has been submitted to LKML [1]. Discussion and iterations of the series are moving forward, but a solution is yet to be landed.
In this regards, we also have the need to extend the support for sync_state to more subsystems beyond genpd, to move away from the broken the "disable unused" features that each subsystem currently provides. Moreover, ideally we prefer the support for sync_state to be adopted on a subsystem basis, rather than relying on a per driver based implementation, which isn't scaling. Attempts have been made to add support to the regulator and clock subsystems, while the support in the interconnect subsystem needs improvements.
Let's discuss these topics and in particular how we can make subsystem specific implementations to coexist and play along with each other.
[1]
[PATCH v3 00/13] driver core / pmdomain: Add support for fined grained sync_state
https://lore.kernel.org/all/20260508123910.114273-1-ulf.hansson@linaro.org/Speaker: Ulf Hansson (Qualcomm) -
16:15
Constification of sysfs attribute structures 25m
Sysfs attributes are used throughout the kernel to implement UAPI.
Subsystems either use common attributes, likekobj_attributeanddevice_attror define their own wrapper structures.
These structures are only descriptors defining the behavior of an attribute and normally never change.Historically the attribute however are not marked as
constand could be modified through their various callbacks. Such modifications are inherently racy and therefore bug-prone or even a vector for attackers to redirect control-flow.For some time I have been working on making it possible to mark all the different attribute structures as
constto fix these issues.Agenda:
* Problem statement (see above)
* Current state
* Which attributes can be markedconsttoday?
* Which attributes are already converted.
* Discussion (see below)- Discussion
Discussion topics:
* Which attribute types are still missing?
* How can subsystem maintainers convert their custom attribute types?
* How to actually convert all the structure instances throughout the tree?Speaker: Thomas Weißschuh (Linutronix) -
16:40
Coffee Break 30m
-
17:10
Hardware Cross-Dependencies - Solving the Unsolvable 25m
We are seeing hard to solve cross-dependencies between different SoC subsystems (Rockchip, MediaTek, etc.). For example a power domain needing an I2C regulator, but the I2C regulator needing the I2C bus and the I2C bus driver needing a (different) power domain. This creates a cyclic dependency, since the power domains (or clocks) are usually all behind a single device.
Speakers: AngeloGioacchino Del Regno (Collabora Ltd.), Sebastian Reichel (Collabora) -
17:35
Synx: Cross-core synchronization 25m
Abstract
Modern SoCs increasingly run parts of a single pipeline (AI/ML,
vision, camera, graphics, sensors) across a mix of Linux drivers
and firmwares on remote processors (NPUs and AI processors, ISPs,
companion cores). Coordinating that pipeline requires
synchronization objects that can be created, synchronously or
asynchronously waited on, signaled, and released by any
participant core, Linux and non-Linux, with lifetime tracked
somewhere no single core owns outright.dma_fencesolves the local version of this problem well, but its
lifetime and callback model assumes a single kernel's view of the
world. We've developed Synx, a global handle/refcount table for
this exact cross-processor case, and posted an RFC to dri-devel
and linux-arm-msm describing the model and the specific properties
dma-fence doesn't currently cover (refer supporting links).Letting any two remote processors signal each other directly,
without routing through the Linux host, cuts latency and avoids an
unnecessary CPU wake-up. The solution has shown power and
performance benefits in the last few generations of Qualcomm
mobile and XR chipsets, and is gathering more use cases,
specifically ones involving AI pipelines.Christian König redirects us to solve remote signaling in
userspace: a userspace fence/signaling point model based on
dma-buf, citing XE's userspace wait support,eventfd, and ROCm
events as precedent, with possible common ground centered around
eventfd.Discussion points
(Supposed to evolve as email thread develops)
- the viability of a userspace vs. kernel-space fence model
- the current remote-signaling solutions from other vendors
- scope as standalone or extend current framework. Define
interfaces.
- model the subsystem crash recovery and cleanupKey people
- Bartosz Golaszewski
- Dmitry Baryshkov, Srinivas Kandagatla (Driver Core MC; also
Qualcomm colleagues who reviewed Synx internally) - Christian König (engaged on the RFC thread)
- Faith Ekstrand / Xe authors (cited by König)
Supporting links
- https://lore.kernel.org/dri-devel/5f90bb35-994e-48bd-bf49-3001dcefa2ee@amd.com/
- https://lore.kernel.org/linux-arm-msm/20260806051915.2234481-1-pravinku@quicinc.com/
- formatted duplicate @ https://lore.kernel.org/dri-devel/20260806051915.2234481-1-pravinku@quicinc.com/
Speaker: Pravin Kumar Ravi (Qualcomm Innovation Center, Inc.) -
18:00
Managing IRQ mapping and deferred probing for ACPI static tables devices 25m
In ACPI based system,devices can be created out of ACPI static tables entries (eg ARM64 IORT, GTDT). For those devices, the GSI HW interrupt number is retrieved by reading table specific fields that are different for different static tables. Devices created out of static ACPI tables might be created before the interrupt controller drivers their GSI interrupt is routed to is probed, which means that when the platform device is created the IRQ domain that should be used to map the GSI into a virtual IRQ may not be registered yet.
This leaves us with two issues:
- device drivers for devices created out of static tables can probe only after the interrupt controller driver their IRQ is routed to has probed
- In order to map the virtual IRQ for those devices at device driver probe time, the static table GSI HW IRQ number must be stashed somewhere in the device object so that it can be retrieved and mapped to a virtual IRQ when the device driver is actually probed
Prototyping for this solution is under way and a solution for ACPI namespace devices was already posted[1] (but that can't solve the problem for devices that are created out of static table entries).
This session would help define a way forward.
[1] https://lore.kernel.org/lkml/20260505-gic-v5-acpi-iwb-probe-deferral-v1-0-b37b85998362@kernel.org/
Speaker: Lorenzo Pieralisi
-
15:00
-
15:00
→
18:30
Live Patching MC "Club A" (Prague Congress Centre)
"Club A"
Prague Congress Centre
53Kernel Live Patching allows fixing kernel bugs without rebooting
or stopping the workload. It is an essential tool to keep modern
data centers health with fast evolving kernels and workloads.The Live Patching MC at Linux Plumbers 2026 aims to gather
stakeholders and interested parties to discuss proposed features
and outstanding issues in live patching.Possible topics for this year:
- Test framework for livepatch subsystem and the new klp-build
toolchain
- Live Patch compatibility with tracing solutions (kprobe, ftrace,
BPF trampoline, etc.)
- Split a live patch module into submodules
- SFrame and livepatch
- Hybrid live patch idea
- Use AI to help build live patchWork landed based on previous versions of the Live Patching MC:
- Live patching for arm64
- Live patching for Loongarch64
- Live patching for LTO (with klp-build tool chain)Key Attendees:
- Josh Poimboeuf
- Jiri Kosina
- Miroslav Benes
- Petr Mladek
- Joe Lawrence
- Song Liu
- Dylan Hatch
- Yafang Shao-
15:00
Backward Compactiliblity of Kernel Livepatching API 20m
Kernel livepatches are kernel modules which are able to modify the kernel behavior by redirecting kernel functions, calling pre/post patch callbacks, and allocating shadow variables.
The interface between the kernel and the kernel livepatch module is defined in
include/linux/livepatch.h.The API has evolved over the years. But it has stayed backward compatible since the commit 958ef1e39d24d6cb8bf ("livepatch: Simplify API by removing registration step") which was added in v5.1-rc1 back in Jan 2019. This change was part of a patchset adding an atomic replace feature.
There are two patchsets which would want to break the backward compatibility again:
It has been acceptable to break the API back in 2019. But the kernel livepatching has been a pretty new feature and had only few users back then.
It might be more acceptable when the tools for creating livepatches can deal with it. But is it enough?
Question for discussion:
- Is is possible to break the kernel livepatching API in 2026?
- Is it enough to update tools for creating livepatches?
- Is it needed and enough to update klp-build?
- How many old code streams people still support?
- Should selftests stay backward compatible?
Speaker: Petr Mladek (SUSE) -
15:20
livepatch: Introduce replace set support 25m
My employer relies heavily on livepatch to rapidly experiment with new kernel features without interrupting production workloads. Our use cases include:
-
Case 1: Deploying a livepatch function as a stable BPF hook.
For example, some proposals for such a use case has already been submitted upstream but has not yet been accepted:
https://lwn.net/Articles/1054030/
https://lwn.net/Articles/1043548/
We also have another internal use case that we do not plan to upstream:
https://lore.kernel.org/live-patching/CALOAHbDnNba_w_nWH3-S9GAXw0+VKuLTh1gy5hy9Yqgeo4C0iA@mail.gmail.com/ -
Case 2: Combining fleet-wide cumulative livepatches with workload-specific livepatches.
For example, consider the VFS cache adjustment feature that we previously upstreamed:
https://lore.kernel.org/linux-fsdevel/20250511083624.9305-1-laoar.shao@gmail.com/
We initially deployed this feature as a livepatch before upstreaming it. In the livepatch version, we could not introduce a sysctl interface for per-workload tuning, so we used a hard-coded default value and deployed it only to specific servers.
In Case 1, the livepatched BPF hook must remain stable. Otherwise, existing BPF programs may become invalid and need to be reloaded, which could introduce operational risks.
In Case 2, we currently need to release different cumulative livepatches for different workloads. However, we would like to have a generic cumulative livepatch combined with individual workload-specific livepatches. This approach would significantly reduce the maintenance burden of managing multiple livepatch variants.
Based on these requirements, we proposed a hybrid livepatch mode, which allows multiple livepatches to coexist with different scopes. This proposal was later refined into the replace_set support proposal, which has already been discussed on the livepatch mailing list:
https://lore.kernel.org/live-patching/20260607131659.29281-1-laoar.shao@gmail.com/
However, several implementation details and design decisions remain open. We would like to discuss them at LPC and determine the best path forward.
In this presentation, I will describe how we use livepatch across our large fleet of production servers and discuss the improvements needed to make livepatch suitable for a broader range of use cases.
Speakers: Petr Mladek (SUSE), Yafang Shao -
-
15:45
Coexistence of Live Patching and Production Observability 45m
Abstract
Large Linux fleets increasingly depend on always-on observability: BPF programs, ftrace, kprobes, kretprobes, and continuous profiling agents are part of the core production control plane. Kernel livepatching depends on some of the same low-level mechanisms, especially dynamic ftrace-based redirection at function entry. In large production environments, we routinely observe fleet rollout delays because critical security livepatches conflict with system-wide tracing tools anchored to the same target functions.The current kernel documentation notes that kprobes, ftrace, and livepatch must not step on each other, citing limitations like kretprobe conflicts. At fleet scale, these are not just corner cases; they become rollout blockers, visibility gaps, or sources of dangerous runtime uncertainty.
This session proposes an upstream discussion to define a predictable, common coexistence model. The goal is to identify the minimum kernel and tooling interfaces needed by livepatch builders, BPF/tracing tool authors, and fleet rollout systems to gracefully handle instrumentation occupancy.
Discussion Topics
- Conflict Inventory: Documenting attach-point collisions among livepatch, ftrace, kprobes, and BPF trampolines.
- Occupancy Reporting: Designing a per-function "patchability" report for tools like klp-build and rollout agents.
- Deterministic Failures: Returning explicit conflict reasons when an attachment is rejected instead of opaque errors.
- Transition Progress: Exposing standard tracepoints or counters for stuck tasks, forced transitions, and instrumentation blocks.
- Test Coverage: Expanding kselftest to include negative tests for rejected, unsafe instrumentation combinations.
- Policy & Chaining: Debating safe chaining semantics and deciding if observability should redirect to replacement functions.
Desired Outcome
Agreement on where these mechanisms belong (kernel ABI, sysfs/debugfs reporting, kselftest, or klp-build metadata). The discussion will involve livepatch, ftrace, kprobe, and BPF maintainers alongside large-fleet operators.Speakers: Kris Van Hees (Oracle USA), Song Liu (Meta), Yi Zhu (Google) -
16:30
Coffee Break 30m
-
17:00
SFrame for Arm64 Reliable Stacktrace 30m
The Livepatch consistency model 1 requires the kernel to provide reliable stacktrace in order to be fully supported. On x86, the ORC unwinder provides these reliable stacktraces. However, arm64 misses the required support from objtool: it cannot generate ORC unwind tables for arm64. Prior RFCs have proposed to add this support to objtool, but feedback from the upstream community has indicated that a solution using data produced directly from the compiler would be preferred 2, 3.
In the v6.17 release, the Arm64 kernel gained livepatch support, but without fully reliable stacktrace 4. With this partial solution, interrupt stacks cannot be reliably traced because the unwinder cannot tell if the Link Register was current at the time the exception was taken.
SFrame provides a generalized, compiler-based solution, inspired by the ORC unwinder 5, 6. Currently, there's already an SFrame unwinder proposed for userspace: 7.
We would like to propose similar functionality to provide reliable stacktraces within the Arm64 kernel: 8. This patch series adds to and depends upon the exising userspace patches by factoring out a common SFrame-lookup library which supports both kernel and userspace SFrame sections. Previous versions of this series implemented an SFrame-only kernel unwinder to replace frame-pointer unwinding entirely. However, after receiving feedback, the kernel unwinder now relies on SFrame only when unwinding across an exception boundary, and uses frame pointers everywhere else. This combined strategy enables the best of both worlds: the superior performance of frame-pointer unwinding is kept whenever possible, while the SFrame table allows for fully reliable unwind of a stack with interrupts.
Ongoing mailing list discussion draws into question SFrame's viability as a userspace unwinding solution 9. This raises several questions around the in-kernel approach, which we would like to bring forward for discussion:
-
What is the current pathway for support for SFrame generation within the LLVM toolchain?
-
The kernel unwind patches continue to have a dependency on the userspace patches. From a development velocity perspective, does this dependency make sense to keep in future versions?
-
For the purpose of unwinding across interrupt boundaries, are there alternatives to SFrame that should be considered?
Speaker: Dylan Hatch -
-
17:30
Bringing klp-build to LoongArch: toolchain lessons for the next architecture 30m
A series adding LoongArch support to objtool's
klp diffsubcommand (the
diffing engine klp-build invokes to generate a patch module) is under
review (v4:
https://lore.kernel.org/all/20260724114128.31451-1-dongtai.guo@linux.dev).
A v5, with a reordering requested by the LoongArch maintainer, is in
preparation. Most of the effort went into the interaction between this
tooling and the LoongArch toolchains, and several findings generalize to
any future architecture with aggressive linker relaxation.This topic walks through the klp-diff issues surfaced across the review
rounds. They are not one bug but a family of toolchain-interaction
failures:- PC-relative reachability. A livepatch module maps farther than the
+-2GB apcalau12i/addi.dpair can express, so a klp relocation
resolving to a vmlinux symbol overflows. Far calls are already
redirected through a PLT stub by the module loader, but file-local
static data is materialized inline; a new arch hook rewrites those
PC-relative data references to GOT-indirect loads. This is not
Clang-specific: GCC emits the same form. - Local-label anchors. GCC/GAS on LoongArch anchor relocations to
.L*
local labels to keep them relaxable, where klp diff assumed a function
or section symbol plus addend (LLVM folds these labels, so this one is
toolchain-specific). Fix: normalize.L*-anchored relocations back to
their containing function, fix annotation offsets in
create_fake_symbols, and collapseR_LARCH_ADD64/SUB64pairs into one
PC-relative relocation viaarch_normalize_paired_reloc. - Section pairing and inline alternatives. Keep both words of each
-mannotate-tablejumpannotation entry, and define
ARCH_HAS_INLINE_ALTSso the.subsection 1ALTERNATIVE() replacement
is cloned, mirroring arm64.
On the compiler side, livepatch modules on LoongArch require -fPIC
(GOT-indirect access instead of absolute addressing), which collides
with the kernel's -fno-PIE defaults in non-obvious ways (last-one-wins
flag ordering vs KBUILD_CFLAGS_KERNEL), and Clang rejects GCC-style
"awM" special-section flags, requiring ANNOTATE_DATA_SPECIAL.
Speaker: dongtai guo - PC-relative reachability. A livepatch module maps farther than the
-
18:00
Testing klp-build: from unit tests to CI 30m
Abstract
With klp-build now merged into mainline, establishing an automated test
suite is the logical next step. Historically, maintenance of
kpatch-build, a similar livepatching creation tool, has shown that the
object diff and correlation layer accounts for the vast majority of
regressions. Variations across compiler versions, optimization levels,
LTO modes, CFI, and architecture-specific handling in objtool's klp-diff
engine introduce a broad surface area where regressions can easily go
undetected.The existing livepatch kselftests provide a proven model within the
subsystem. Developed alongside core livepatching functionality, they
have caught real bugs, e.g. covering shadow variables, callbacks, and state
transitions. All while remaining maintainable over time. klp-build
warrants a similar testing discipline, tailored to the toolchain layer.We propose a two-tier testing strategy designed for both local developer
workflows and continuous integration:Unit tests: Directly exercise
objtool klp diffagainst minimal
.orig.o/.patched.oobject pairs. Scripted assertions verify
symbol correlation, demangling, checksumming, special-section handling,
and relocations in seconds without requiring a full kernel build. To
adhere to kernel repository standards and avoid committing binary
artifacts, a CLI-driven runner regenerates test object pairs on demand
from module source and a kernel tree using the host/target toolchain.Integration tests: Execute the complete klp-build pipeline against
dedicated in-tree test infrastructure (CONFIG_KLP_BUILD_TEST),
validating module post-linking and runtime livepatch loading on target
architectures.Both tiers target a comprehensive architecture (x86_64, arm64,
LoongArch) and toolchain matrix (GCC, Clang, ThinLTO, CFI). Our goal is
to integrate these test suites into upstream CI infrastructure (e.g.,
CKI, KernelCI) to provide continuous regression testing across
configurations.We invite discussion with the community on several key architectural and
integration questions:Unit test granularity: Should tests target individual objtool
components (e.g., checksumming, correlation) independently, or is a
singleobjtool klp diffinvocation per test case the appropriate
abstraction level?In-tree vs. out-of-tree structure: Should integration tests reside
in-tree within kselftests while unit tests remain out-of-tree -- and
where should the test corpus generation tooling live?Cross-version maintenance: How can in-tree integration test code be
structured to remain stable across kernel releases without incurring
unnecessary maintenance overhead?CI pipeline integration: Should we integrate this this multi-arch,
multi-toolchain matrix into continuous integration platforms like CKI
and KernelCI?Speakers: Joe Lawrence (Red Hat), Song Liu (Meta)
-
15:00
-
15:00
→
18:30
sched_ext: The BPF extensible scheduler class MC "Club H" (Prague Congress Centre)
"Club H"
Prague Congress Centre
128sched_ext[1] is a Linux kernel feature that enables implementing safe task schedulers in BPF and dynamically loading them at runtime. Its key strength is flexibility, allowing rapid iteration of scheduling policies, deploying changes on the fly and quickly addressing topology inefficiencies or workload-specific issues.
This MC provides a space for the community to discuss the evolution of sched_ext, its impact and future strategies aimed at improving integration with other Linux kernel subsystems.
Last year the sched_ext MC proved highly productive in facilitating coordination with other kernel maintainers, allowing us to address open issues and limitations of this technology (see for example the introduction of a dedicated DL server for the SCHED_EXT scheduling class [2]).
Topics for discussion include (but are not limited to):
- Hierarchical cgroup sub-schedulers
- Integration with SCHED_DEADLINE (DL server interface and related improvements)
- Proxy execution support
- Device-aware scheduling policies (e.g., GPU auto-affinitization)
- Composable schedulers and reusable scheduler libraries (leveraging BPF arenas)
- Scheduling strategies for gaming and latency-sensitive workloads
- Tickless scheduling and CPU isolation
- Improved tooling for tracing and visualizing scheduler performanceA public CFP will follow to gather additional topics that may be relevant to the Linux community.
Key attendees:
- Andrea Righi
- Changwoo Min
- Tejun Heo
- Peter Zijlstra
- Juri Lelli
- Vincent Guittot
- Dietmar Eggemann
- Steven Rostedt
- K Prateek Nayak
- John Stulz
- Shrikanth Hegde
- Emil Tsalapatis
- Daniel Hodges
- Christian Loehle
- Ryan Newton[1] https://github.com/sched-ext/scx
[2] https://lore.kernel.org/all/20260126100050.3854740-1-arighi@nvidia.com-
15:00
Intro / Welcome 10m
-
15:10
Bridging the VM Boundary: Scheduling Passthrough via pvsched and sched_ext 18m
As production workloads increasingly transition to virtual machines for security isolation and resource consolidation in multi-tenant environments, traditional CPU scheduling faces a critical M:N preemption challenge. The host operating system schedules opaque virtual CPUs rather than the actual workload threads. Consequently, the host scheduler remains blind to the varying priorities and latency sensitivities of the guest threads running inside the VM. This leads to severe priority inversion; for instance, the host scheduler cannot differentiate between a vCPU running a low-priority batch or kernel system thread and one executing a latency-sensitive task (such as a critical helper daemon) nested within the same VM. Consequently, critical latency-sensitive work is starved, and physical resources are wasted under host CPU contention.
To resolve this, we propose a bidirectional VM Scheduling Passthrough architecture to bridge the VM boundary from a scheduling perspective. This model relies on cooperative, paravirtualized communication to enable host-side awareness and control:
1. Guest-to-Host: The guest kernel exposes scheduling metadata (potential signals include thread-level priorities and latency requirements) to the host.
2. Host-to-Guest / Host Control: The host leverages these guest signals to either (a) intelligently prioritize the execution of a given vCPU, or (b) in a more complete solution, directly select both the vCPU and the specific guest thread running on that vCPU to execute (this latter proposal is more aspirational, and would effectively give the host full scheduling control over the guest, voiding the need for guest scheduling or load balancing).Implementing such a model requires a clean separation of mechanism and policy. In line with upstream maintainer feedback, KVM should remain policy-free, acting purely as the communication conduit. We propose sched_ext as the ideal framework to house the scheduling policy, and pvsched as the communication mechanism to share data between host and guest. Running the policy in a host BPF scheduler allows for rapid iteration and workload-specific customization without modifying KVM or the core kernel.
Current upstream efforts, notably the IBM patch series "Introduce cpu_preferred_mask and steal-driven vCPU backoff", attempt to address host preemption purely from the guest side. In this model, the guest monitors hypervisor steal time and flags heavily preempted vCPUs as "non-preferred" to guide guest-side task migration. While this provides a defensive mechanism, it is unidirectional and reactive. Furthermore, as highlighted by upstream maintainers, this approach faces significant limitations: it risks embedding scheduling policy within the hypervisor, and in high-overcommit scenarios where steal time is elevated across all cores, the guest-side mask degrades, leaving the guest scheduler with no viable execution targets.
pvsched (https://github.com/pvsched/) has undergone some discussion on the list already, and v3 is being prepared.
In this session, we want to discuss:
- Interface Design: The updates in pvsched v3
- sched_ext Integration: How sched_ext can ingest these guest signals to influence host scheduling decisions, and conversely, how it can export host scheduling desires back to the guestSpeakers: Josh Don (Google), Mr Vineeth Remanan Pillai (Google) -
15:28
BPF-based Composable Idle cpumask Selection 18m
All sched_ext schedulers currently use kfuncs to manage their idle cpumask using a hardcoded policy provided by the kernel. This lack of configurability of the current cpumask requires us to add policy through explicit masking operations directly in the scheduler code. This in turn leads to duplicating idle CPU selection logic across schedulers as it is difficult to factor it out.
This session discusses a new BPF-based subsystem for idle CPU mask selection. This subsystem abstracts the representation of the idle CPU set behind a common API that exposes explicit alternative policies to the user. The API is portable between schedulers and requires little code to adopt. Internally, the subsystem stores idle CPU information in arena-based data structures. Different structures (e.g., radix trees) provide different performance/scalability tradeoffs and can be tailored to different idle CPU selection policies.
Topics for this session:
- Overview of the idle CPU selection subsystem
- Data structures for representing the idle CPU set, tradeoffs for each
- Representing policy in a flexible, extensible fashion through the idle selection API
- Using alternative CPU selection strategies as part of load balancingSpeaker: Emil Tsalapatis (Meta Platforms) -
15:46
Taming latency spikes caused by lock holder and lock waiter preemption 18m
Preempting a lock holder — or failing to promptly schedule a just-woken
lock waiter — extends the serialized critical section and produces severe
tail-latency (P99) spikes: degraded server throughput, frame-time
spikes and dropped frames in games. Applications hit this on both
kernel-space locks and user-space primitives backed by futexes and SysV
semaphores. Existing techniques help but leave gaps: proxy execution
addresses kernel-mutex priority inversion but not user-space futex waiters
or counting semaphores, and time-slice extension protects a detected holder
while doing nothing for the delayed waiter.We define lock waiter preemption (LWP) as the dual of lock holder
preemption (LHP): the scheduling delay a woken waiter suffers before it
runs, which we have measured at 2–16 ms in production game and server
workloads. We will present these measurements and a prototype LHP/LWP
mitigation built in a production sched_ext scheduler, then open three
challenges for community discussion:- efficiently identifying lock-related preemption from BPF;
- exposing synchronization state between libc and sched_ext;
- designing lightweight scheduler-assisted locking that improves latency
without sacrificing scalability.
Speaker: Changwoo Min (Igalia) -
16:04
Making proxy execution compatible with sched_ext 26m
Proxy execution allows a waiting task (the "donor") to donate its execution context to a mutex owner, enabling the owner to continue running while the donor remains eligible on the runqueue.
Today, proxy execution and sched_ext are mutually exclusive build-time options: a kernel cannot be built with both CONFIG_SCHED_PROXY_EXEC=y and CONFIG_SCHED_CLASS_EXT=y.
This limitation is problematic for Linux distributions and anyone who wants to ship a single kernel image while selecting features at runtime.
An RFC series proposing support for this integration has already been posted: https://lore.kernel.org/all/20260506174639.535232-1-arighi@nvidia.com/
However, several design questions remain open. This session aims to discuss those issues, reach agreement on the overall approach, and define a concrete plan for moving the integration forward.
Speaker: Andrea Righi (NVIDIA) -
16:30
Coffee Break 30m
-
17:00
Stickiness: Keeping tasks cache-warm in work-conserving LAVD scheduler 18m
Work-conserving schedulers prefer running tasks on idle CPUs
immediately, ensuring no processing capacity is wasted while work is
waiting. In the lavd select_cpu process, when a task wakes, the
scheduler will do its best to seek an idle core and run the task over
there. However, this idle-oriented CPU selection generally prioritizes
idle cores over the cache-warm cores, leading to more migrations,
making the tasks cache-cold, and requiring L1/L2 caches and TLBs to be
refilled frequently.The work-conserving mechanism works well in some scenarios. However,
there are cache-sensitive workloads such as edge routers using routing
tables as in-memory KV stores. The migrations to the cold CPU have a
global impact, and the cache refilling evicts the other cache lines
and pulls the new ones in, making the cache contention worse and further
impacting the tail latency.
We will present our measurements and open the floor on topics
including:-
Warmth estimation: How long does L1/L2/TLB state realistically
survive on a CPU, and could the kernel expose hardware signals
(e.g., PMU counters, cache-occupancy registers) usable at
scheduling-decision frequency? -
The wait-or-migrate decision: Waiting requires predicting when
a busy CPU will become available (remaining slice + queue service
time, or queue load?). Which metrics should a scheduler maintain to
make that prediction accurate? -
The scheduler/userspace interface: Userspace structures (such
as allocator per-CPU caches) suffer heavily from core migrations.
Should applications hint their locality needs to the scheduler, or
should the scheduler expose warmth and migration state to userspace?
Speaker: Gavin Guo -
-
17:18
HFI plug in to scx_ext 18m
This presentation explores how Intel HFI can be integrated with a sched_ext to improve task placement and power-performance efficiency on hybrid Intel systems. HFI provides real-time hardware guidance on which CPUs are better suited for performance- or efficiency-oriented work, or which to avoid, while sched_ext like LAVD supplies an adaptive scheduling framework capable of using that guidance at runtime. By converting HFI output into scheduler-readable CPU hints, the system can bias wakeup placement, avoid degraded cores, and better match task type to CPU capability without hard pinning or manual tuning.
Speaker: Srinivas Pandruvada -
17:36
Automatic placement of accelerator workloads 26m
Proposal
On modern multi-socket, multi-GPU systems, application performance is often limited not by compute availability but by poor CPU/GPU locality. Today, customers are frequently instructed to rely on strict node pinning (numactl) and disabling NUMA balancing in order to avoid costly cross-node memory accesses. While effective in some cases, this approach can be suboptimal, since it limits workloads to a subset of available resources and it also requires deep system knowledge from the user.
In this talk, we will explore how GPU-aware auto-affinitization can be implemented directly in the Linux scheduler using sched_ext, enabling dynamic and transparent placement of CPU tasks close to the GPUs they actively use. We discuss the challenges of integrating scheduler-driven task migration with NUMA balancing, handling mixed CPU/GPU thread groups, and avoiding resource over-concentration. Finally, we present design principles showing how scheduler-level techniques can outperform static affinity policies, while simplifying the user experience.
Key Challenges
Interaction with NUMA Balancing
Naively migrating a task closer to its GPU can backfire if the task's memory remains allocated on a remote NUMA node. In such cases, migration may increase memory access latency instead of reducing it.
Key questions explored:
How can sched_ext and NUMA balancing cooperate rather than conflict?
When should migration be deferred or paired with memory migration?Thread Aggregation and Shared State
Many GPU-enabled applications use multi-threaded CPU components, where:
Only a subset of threads directly interacts with the GPU
Other threads share memory, locks, or cache lines with GPU-driving threads
Migrating only GPU-active threads may introduce new inefficiencies due to cross-node communication among threads of the same process.We need to explore strategies for:
Selective thread aggregation, migrating related threads together
Avoiding overload of a single NUMA node or LLC
Balancing locality benefits against parallel resource contentionResource Saturation and Fairness
Automatically clustering tasks near GPUs risks oversubscribing CPUs, LLCs, or memory bandwidth on specific NUMA nodes. The scheduler must therefore avoid overloading a certain LLC or NUMA node, make optimal use of system resources and maintain fairness.
Speakers: Balbir Singh, Lee Trager (NVIDIA) -
18:02
Closing / Final Q&A 28m
-
15:00
-
15:45
→
18:30
Birds of a Feather (BoF) "South Hall 1 B" (Prague Congress Centre)
"South Hall 1 B"
Prague Congress Centre
158-
15:45
FUSE mostly-passthrough filesystems 45m
Many production FUSE filesystems do not need to reimplement the full filesystem stack. They need to mirror an existing directory with full fidelity while intercepting only a small subset of operations for caching, tiering, HSM stub manifestation, or auditing.
To make these filesystems easier to build, we introduce libfuse_passthrough - a C++ library that handles all FUSE plumbing and exposes a lightweight module API where only the intercepted operations need to be implemented where everything else passes through automatically.
The library makes use of upstream kernel FUSE read/write passthrough, which allows open files to bypass the FUSE daemon entirely, and provides an experimental playground for the development of more FUSE kernel passthrough operations.
We invite contributors and users with FUSE passthrough workloads to learn about our experience, discuss open problems, kernel interface requirements, and help shape the future of FUSE.
Speaker: Amir Goldstein (CTERA Networks) -
16:30
Coffee 30m
-
17:00
Defining the pkeys ABI 45m
We have been supporting pkeys [1] on Linux for 10 years now, and they
are no longer specific to x86: arm64 and powerpc support them too.
Unfortunately, important gaps remain in the kernel-user ABI, making it
difficult to deploy pkeys robustly for key use-cases.pkeys work in a fairly simple way: the user allocates a new pkey and
then assigns that pkey to a VMA. Access to that VMA is then restricted
by a user-controlled register, which defines RWX permissions for every
pkey. The restrictions apply to both user and kernel accesses (uaccess
and GUP).Difficulties arise when the kernel interrupts a user thread and then
accesses its memory, or invokes a signal handler. In those cases, it is
unclear which pkeys authority the kernel should use, whether userspace
should be able to configure it, and how to avoid spurious crashes or
confused deputy situations. This is what the current ABI is lacking.This BoF is especially intended for arch/mm maintainers, libc/runtime
developers, and userspace developers already familiar with pkeys or
deploying pkey-based isolation. The goal is not to present a finished
design, but to agree on the programming model that future ABI work
should follow.A few important use-cases are currently difficult or impossible to
implement robustly:-
Isolating the alternate signal stack: mapping it with a non-default
pkey that only signal handlers are allowed to access. -
Sandboxing with pkeys: preventing some context from accessing
even the default pkey (0).
Problematic situations include:
-
Signal delivery [2]: signal handlers may be called asynchronously and
are conceptually independent of the interrupted context.
Userspace may need the pkey register to be reset to a specific value
both for writing the signal frame and invoking the signal handler. -
rseq [3]: the kernel writes to the registered struct rseq when context
switching, which can happen at any point. This is again independent of
the interrupted context, which may not be allowed to write to that
struct. -
io_uring worker threads: these are kernel threads that access user
memory asynchronously. They have their own pkey register value
restricting those accesses, but userspace is not currently able to
configure it. -
Other cases, such as BPF helpers accessing user memory and
process_vm_readv(self) bypassing pkeys completely.
All these issues lead to a set of questions that we need to answer to
create a consistent and unsurprising ABI:-
Which contexts need a dedicated or userspace-configurable pkey
register value? -
Are there accesses that should intentionally bypass pkeys, and if so,
how should that be documented? -
Can these issues be handled with targeted fixes, or do we need a more
unified ABI design for asynchronous kernel access to pkey-protected
memory?
As Thomas Gleixner put it:
We really need to sit down and actually define a proper programming
model first instead of trying to duct tape the current ill defined
mess forever.The intended outcome of this BoF is a shared direction for documenting
and extending the pkeys ABI, so that pkeys can be used reliably as
intra-process privilege boundaries rather than only as a best-effort
hardening mechanism.[1] https://docs.kernel.org/core-api/protection-keys.html
[2] https://inbox.sourceware.org/libc-alpha/fc31e639-f1eb-42d6-9dea-3665d9507f12@arm.com/
[3] https://lore.kernel.org/all/87ikexhbah.ffs@tglx/Speaker: Kevin Brodsky (Arm) -
-
17:45
Virtio and Vhost-User Ecosystem for Production and Automotive Systems 45m
The vhost-user protocol started as a simple QEMU feature, but it has grown into a standard used across the entire Linux virtualization community. Today, it powers production workloads in the cloud (e.g. QEMU and cloud-hypervisor), lightweight environments like libkrun, and modern automotive systems.
However, as the ecosystem expands, we are hitting new "plumbing" bottlenecks. For example, the protocol specification is still stored inside the QEMU documentation, which makes it harder for other projects to participate in its growth. At the same time, automotive deployments are introducing strict new needs, such as safety certification, real-time performance, and long-term support, that impact everyone.
This BoF will bring developers and users together to address these challenges and other emerging topics across the ecosystem. Our goal is to establish better coordination, align technical roadmaps for both cloud and automotive, and build shared tests and tools to ensure all implementations work together.
Key Discussion Topics:
- vhost-user Specification Governance: Discuss official location for the vhost-user specification.
- Automotive Virtualization Requirements & Reference Platforms: Coordinate deployment needs across Automotive platforms [1]
- Real-time performance needs (camera, display, CAN)
- Status Update: Ongoing work on vhost-user device (virtio-Media, virtio-can, virtio-rtc, etc)[2][3]
- Graphics Stack Alignment: on roadmaps for shared dependencies like rutabaga_gfx across libkrun, crosvm, vhost-device-gpu [4]
- Shared Virtio Message-Parsing Crates: Evaluate creating unified Rust crates for Virtio device specs to share virtqueue parsing logic across backends and VMMs.
- Open questions about:
vhost-user Specification Refinements: Identify protocol gaps, missing features, and necessary standard improvements.
New Virtio/vhost-user Device Proposals.[1] https://source.android.com/docs/automotive/virtualization/reference_platform
[2] https://lore.kernel.org/virtualization/ab2FlQTWUxl0KmlT@fedora/
[3] https://github.com/rust-vmm/vhost-device/pull/944
[4] https://github.com/magma-gpu/rutabaga_gfx/issues/24Key Participants:
- Albert Esteve aesteve@redhat.com
- Stefano Garzarella sgarzare@redhat.com
- Sergio Lopez slp@redhat.com
- Manos Pitsidianakis manos.pitsidianakis@linaro.org
- Alex Bennee alex.bennee@linaro.org
- Matej Hrica mhrica@redhat.com
- Milan Zamazal mzamazal@redhat.com
- Gurchetan Singh gurchetansingh@google.com
- Jorge E. Moreira jemoreira@google.com
- Matias Vara Larsen mvaralar@redhat.com
- Dorinda Bassey dbassey@redhat.com
- Alberto Ruiz aruiz@redhat.com
- Erico Nunes ernunes@redhat.com
- John Ferlan jferlan@redhat.com
- Michael Tsirkin mst@redhat.com
- German Maglione gmaglion@redhat.com
- Stefan Hajnoczi stefanha@redhat.com
- Timos Ampelikiotis t.ampelikiotis@virtualopensystems.com
- Harald Mommer harald.mommer@oss.qualcomm.comSpeakers: Dorinda Bassey (Red Hat), Albert Esteve (Red Hat), Stefano Garzarella (Red Hat)
-
15:45
-
19:30
→
22:30
Evening Event 3h
-
09:00
→
15:45