Speakers
David Dai
Shakeel Butt
Description
What started out as an investigation for potential isolation issues in the fleet led us down into a few different rabbit holes. We started with trying to find patterns where the kernel lock holders are unable to run, causing other tasks to wait on said locks. We'll discuss ways to monitor said patterns across the fleet at a broader scale and then zoom into some interesting case studies to see what we found including:
- OOM killer prints to the console while holding a lock
- Act of monitoring causes isolation issues in the fleet for sensitive services
- Hard partitioning from CPUSETS causes lock holder unable to get CPU time
- Chain of lock waiters -> Lock holder sleeping on synchronize_rcu -> CPU holding up grace period
Also we'll discuss how to expand on monitoring for scheduling issues and/or issues in the fleet that causes stalls that prevent system wide progress