What disk I/O queue depth is and why it matters
I/O queue depth is the number of read and write operations a drive processes at the same time. On a hard drive, the physical queue is limited by the mechanics; on a SATA SSD, by the AHCI protocol (queue depth up to 32); on NVMe, by the NVMe protocol itself, which supports up to 65535 queues with 65536 commands each. This is why NVMe drives handle heavy load noticeably more stably than SATA at the same IOPS count.
When the queue on a drive grows faster than it can process it, the response time for each next request increases — that is a rise in latency under load, not a drop in "speed" in the usual sense.
IOPS and latency: the difference and how to calculate them
IOPS (input/output operations per second) is the number of operations per second. Latency is the time to complete one operation, in milliseconds or microseconds. The two metrics are linked through the queue: roughly, IOPS = queue depth / average latency of one operation.
For HDD, typical random read latency is 5-10 ms; for SATA SSD, 0.1-0.2 ms; for NVMe, 0.02-0.05 ms. A 100-200x difference between HDD and NVMe means that at the same queue depth, NVMe delivers hundreds of times more IOPS on random operations, although on large-block sequential reads the difference is much smaller.
How to measure queue depth and latency in real time
The basic tool is iostat from the sysstat package:
iostat -x 1 5
In the output, three columns matter: avgqu-sz (or aqu-sz in newer versions) — average queue depth, await — average request wait time in milliseconds, and %util — the share of time the drive was busy processing requests. If %util stays near 100% and await is rising, the drive has become the system's bottleneck, and the next step is to find what is actually loading it.
The process generating the most disk load is found with:
iotop -o
fio: a synthetic test matching your real workload profile
To test a drive against a specific scenario rather than in the abstract, use fio with parameters close to the real workload. For a database with many small random operations:
fio --name=randread --filename=/dev/nvme0n1 --rw=randread \
--bs=4k --iodepth=32 --numjobs=4 --runtime=60 --time_based \
--group_reporting
For sequential writes with large blocks, as in backups or logging:
fio --name=seqwrite --filename=/data/testfile --rw=write \
--bs=1M --size=10G --iodepth=8 --runtime=60 --time_based
In the fio report, look at clat (completion latency) percentiles — the average masks rare but painful delays at the 99th percentile, which are exactly what users feel as freezes.
NVMe versus SATA: the queue depth difference in practice
| Parameter | SATA SSD (AHCI) | NVMe |
|---|---|---|
| Max queue depth | 32 | 65536 per queue |
| Number of queues | 1 | up to 65535 |
| Random read latency | 0.1-0.2 ms | 0.02-0.05 ms |
In practice, the difference is felt most in multi-threaded workloads — databases, virtualization with dozens of VMs on one host, message queues. For a single drive under a database server, NVMe is almost always preferable to SATA SSD precisely because of queue depth, not just higher linear speed.
Checklist for diagnosing disk problems
%utilin iostat stays steadily near 100% — the drive is genuinely at its limit, not something else causing the problem.awaitgrows faster than the load grows — the queue is not draining fast enough; check the RAID controller and caching mode.- The fio test was run with parameters close to the real profile (block size, iodepth, read/write ratio), not with default settings.
- Comparison with the drive's rated IOPS shows a gap larger than 20-30% — check partition alignment and TRIM mode.
- For critical databases and high-load services, queue depth and latency were tested at RAID-level profile (RAID 10 across several drives), not on a single drive.
Measuring queue depth and latency before putting a service into production saves hours of incident investigation later: most complaints about a "slow server" actually come down to a drive that has hit the limit of its queue.