Storage I/O Tuning
Linux storage I/O performance is critical for applications like databases, virtualization, and high-throughput workloads. Tuning kernel parameters and I/O schedulers can significantly reduce latency, minimize read/write amplification, and optimize throughput for block devices. This section covers key sysctl parameters and configuration strategies for storage I/O optimization.
Disk I/O Scheduling¶
The Linux kernel provides multiple I/O schedulers to manage disk request queues. The choice of scheduler impacts latency, fairness, and throughput. The sysctl interface allows dynamic configuration of scheduler behavior.
Available Schedulers¶
- CFQ (Completely Fair Queuing): Default for most systems. Prioritizes fairness but may introduce latency for high-throughput workloads.
- Deadline: Optimized for low latency with predictable scheduling (common for SSDs).
- Noop: Simple, FIFO-based scheduler ideal for SSDs with minimal seek time.
- Kyber: Modern, latency-sensitive scheduler (available in newer kernels).
Key Parameters¶
kernel.sched_min_granularity_us: Minimum time slice for task scheduling (affects I/O fairness).kernel.sched_latency_ns: Target latency for scheduling decisions (lower values reduce latency).block.io.scheduler: Specifies the default scheduler for block devices (e.g.,echo deadline > /sys/block/sda/queue/scheduler).
Example: Switching to Deadline Scheduler
Example: Tuning Scheduler Granularity
Read/Write Amplification¶
Read/write amplification occurs when the kernel issues more I/O operations than necessary, increasing latency and wear on storage. Tuning parameters can reduce this overhead. Dirty page management directly impacts both read and write amplification by influencing how and when data is flushed, affecting I/O operation frequency and memory pressure.
Key Parameters¶
vm.dirty_ratio: Maximum percentage of system memory that can be filled with dirty pages (default: 20%).vm.dirty_background_ratio: Percentage of memory that can be filled with dirty pages before background writeback begins (default: 10%).vm.dirty_expire_centisecs: Time (in centiseconds) before dirty pages are considered expired and written back (default: 3000).vm.dirty_writeback_centisecs: Time between periodic writeback attempts (default: 500).
Example: Reducing Write Amplification for Write-Heavy Workloads
Example: Adjusting Writeback Frequency
Note: Lowering these values can improve responsiveness but may increase memory usage. Balance based on workload and system RAM.
Latency Optimization¶
Minimizing I/O latency is crucial for real-time applications. Parameters like kernel.sched_latency_ns and scheduler selection directly impact this. Readahead settings are configured via sysfs, not sysctl, to optimize sequential read performance.
Key Parameters¶
kernel.sched_latency_ns: Target latency for scheduling decisions (lower values reduce latency).kernel.sched_min_granularity_ns: Minimum time slice for scheduling (affects fairness vs. latency)./sys/block/<device>/queue/read_ahead_kb: Readahead size for block devices (default varies by kernel).
Example: Reducing Scheduling Latency
Example: Enabling Readahead Optimization
Note: For SSDs, consider disabling readahead or reducing it to avoid unnecessary I/O. Use hdparm -Tt /dev/sda to benchmark read performance.
Key takeaways¶
- Choose the right I/O scheduler (e.g.,
deadlinefor SSDs,noopfor high-throughput workloads). - Balance dirty page thresholds to avoid write amplification while maintaining responsiveness.
- Tune scheduling latency parameters to prioritize low-latency workloads.
- Monitor and test changes in non-production environments to avoid unintended side effects.