Input/output operations per second (IOPS) is a measure of storage performance: the number of separate read and write operations a drive, storage array or cloud storage volume can complete in one second. Each operation is a read or write request, counted at whichever layer is doing the measuring, such as the application, operating system, storage array or drive. IOPS is most useful for workloads that make many small, random requests, such as databases, virtual machines and email systems, and is usually considered alongside latency, the time each operation takes, and throughput, the total amount of data moved per second.
At a glance
- Counts operations, not data volume: 10,000 IOPS means 10,000 reads or writes completed each second.
- Results depend on block size, read and write mix, random or sequential access, and how many requests are in flight, so figures are comparable only when those conditions match.
- Flash and NVMe storage deliver far more IOPS than spinning hard drives.
- In the cloud, IOPS is often tied to volume size or tier, or provisioned and billed separately.
- One of three main storage performance measures, with latency and throughput.
What problem it solves
When an application is slow, the cause is often storage rather than processors or memory. A database that handles many small transactions, or a host running dozens of virtual machines, sends a constant stream of small requests to its disks. If the storage can’t keep up, requests queue and everything waits, even though the servers appear underused.
IOPS gives buyers and engineers a way to describe how many of those requests a workload needs and how many a storage system can deliver. It helps size new storage, compare quotes, choose cloud volume tiers and diagnose performance problems. Used carelessly, it also produces misleading comparisons, which is why understanding how it is measured matters.
How it works
What an operation is. An I/O operation is one read or write request as seen at the layer being measured, whatever its size. Requests from a server and operations inside the storage need not map one to one: the operating system or storage may split large requests, merge small ones or answer them from cache, so IOPS reported by an application, a host and an array can differ. Small operations, such as 4 KB or 8 KB, are typical of databases; larger ones are typical of backups, video and file copies.
What affects the number. Several factors change the IOPS a system delivers:
- Block size. Larger operations move more data each but fewer can be completed per second.
- Read and write mix. Writes are usually more expensive than reads, especially with RAID or replication, which add extra internal writes.
- Random versus sequential. Random access is harder for spinning drives, which must move their heads; flash handles it much better.
- Queue depth. More requests in flight at once can raise IOPS, up to a point, often at the cost of higher latency.
Relationship to throughput. Throughput is roughly IOPS multiplied by the average size of each operation. A workload of small random operations can need high IOPS but modest throughput; a backup job can need high throughput with few IOPS. Throughput is the storage equivalent of bandwidth.
Where limits apply. IOPS can be limited by drives, the storage controller, the network to storage such as a SAN, the server’s adapters and, in the cloud, both the volume and the virtual machine instance. The lowest limit in the chain sets the real figure.
Measuring it. Operating systems, hypervisors, storage arrays and cloud monitoring commonly report IOPS. Benchmark tools can test systems with chosen block sizes and mixes, but results predict real performance only as well as the test resembles your workload.
See our file and object storage page for help comparing storage options.
When it matters for buyers
- Sizing a storage purchase or refresh. Measure peak IOPS and latency on current systems before buying, rather than relying on capacity alone.
- Choosing cloud block storage tiers. The cheapest tier may cap IOPS below what a database needs; the most expensive may be more than it uses.
- Comparing quotes. Insist on IOPS figures at the same block size, read and write mix and latency.
- Troubleshooting slow applications. High storage latency and queueing at modest IOPS often point to a storage bottleneck.
- Moving to storage as a service or hyperconverged systems. Check how performance is committed and what happens when the commitment is exceeded.
Questions to ask vendors
- At what block size, read and write mix and queue depth were your IOPS figures measured?
- What latency do you deliver at that IOPS level, and at our expected load?
- Is the figure sustained, or a burst that drops after a period of heavy use?
- How is IOPS priced: included with capacity, tied to volume size or provisioned separately?
- What happens to performance during a drive rebuild, snapshot or replication?
- Are there limits elsewhere, such as per server, per instance or per volume, that cap what we can actually get?
How it differs from throughput and latency
IOPS, throughput and latency measure different things and can tell different stories. IOPS counts how many operations complete each second. Throughput, sometimes called bandwidth, measures how much data moves each second, in megabytes or gigabytes per second, and is roughly IOPS times the average operation size. Latency is how long a single operation takes, usually in milliseconds or microseconds. A system can post high IOPS while latency climbs, which users experience as slowness. For databases and virtual machines, low and consistent latency at the required IOPS is usually what matters; for backups and video, throughput matters most.
