Skip to content

cgroups Monitoring

Cgroups Monitoring and Debugging
Cgroups v2 introduces a unified hierarchy for resource management, but effective troubleshooting requires understanding how to monitor usage, identify contention, and debug misconfigurations. This section covers tools and techniques for analyzing cgroup statistics, leveraging systemd integration, and resolving common resource allocation issues.


Understanding Cgroup Statistics

Cgroups v2 exposes detailed metrics via files in the /sys/fs/cgroup/ hierarchy. Key statistics include memory usage, CPU time, I/O operations, and network traffic. Use these to track resource consumption and detect anomalies.

Example: Checking Memory Usage

# View memory usage for a cgroup
cat /sys/fs/cgroup/memory/my_cgroup/memory.usage_in_bytes

Example: Monitoring CPU Time

# Track CPU time (in nanoseconds)
cat /sys/fs/cgroup/cpu/my_cgroup/cpu.cfs_period_us
cat /sys/fs/cgroup/cpu/my_cgroup/cpu.cfs_quota_us

Using cgexec for Process Monitoring
Run processes under a cgroup to isolate resource usage:

cgexec -g memory:my_cgroup stress-ng --cpu 4 --io 4


Systemd Integration for Debugging

Systemd units (e.g., services, sockets) can be configured to enforce cgroup limits. Use systemctl to inspect and manage cgroup settings.

Example: Configuring a Service Unit

[Service]
CPUQuota=50%
MemoryLimit=2G
TasksMax=100

Checking Unit Status

systemctl status myservice
journalctl -u myservice

Debugging with cgclassify
Move processes between cgroups to resolve contention:

cgclassify -g memory:my_cgroup <PID>


Tools for Advanced Analysis

  • cgroupstats: Summarize cgroup metrics (requires installation).
  • cgroup-lite: Lightweight tool for listing cgroup details.
  • perf: Analyze CPU and memory bottlenecks.

Example: Listing Cgroup Details

cgroup-lite -o name,cpu,mem /sys/fs/cgroup


Troubleshooting Common Issues

  1. Resource Contention: Use cgexec to isolate workloads and verify quotas.
  2. Processes Not Classified: Check cgclassify output and ensure PIDs are correct.
  3. Incorrect Quotas: Validate CPUQuota/MemoryLimit settings in systemd units.
  4. Log Analysis: Use journalctl to trace systemd unit errors or cgroup violations.

Key takeaways

  • Monitor cgroup stats via /sys/fs/cgroup/ files to track resource usage.
  • Leverage systemd units to enforce resource limits and debug misconfigurations.
  • Use cgclassify and cgexec to reclassify processes and test cgroup behavior.
  • Combine tools like cgroup-lite and perf for deeper performance analysis.
  • Always validate cgroup settings with systemctl and journal logs during troubleshooting.