5–7 Oct 2026
Europe/Prague timezone

We don't know what we're looking for until we find what we're looking for

5 Oct 2026, 11:05
15m
"Club H" (Prague Congress Centre)

"Club H"

Prague Congress Centre

128
Linux System Monitoring and Observability MC Linux System Monitoring and Observability MC

Speakers

David Dai Shakeel Butt

Description

What started out as an investigation for potential isolation issues in the fleet led us down into a few different rabbit holes. We started with trying to find patterns where the kernel lock holders are unable to run, causing other tasks to wait on said locks. We'll discuss ways to monitor said patterns across the fleet at a broader scale and then zoom into some interesting case studies to see what we found including:

  • OOM killer prints to the console while holding a lock
  • Act of monitoring causes isolation issues in the fleet for sensitive services
  • Hard partitioning from CPUSETS causes lock holder unable to get CPU time
  • Chain of lock waiters -> Lock holder sleeping on synchronize_rcu -> CPU holding up grace period

Also we'll discuss how to expand on monitoring for scheduling issues and/or issues in the fleet that causes stalls that prevent system wide progress

Authors

Presentation materials

There are no materials yet.