NVMe IOPS: QD1 to QD4 Reveal More Than Peak Specs

NVMe IOPS: QD1 to QD4 Reveal More Than Peak Specs

NVMe IOPS: QD1 to QD4 Reveal More Than Peak Specs

Parallel storage queues at different depths

IOPS stands for Input/Output Operations Per Second, the count of individual read or write commands a storage device completes in one second. NVMe drives post far higher IOPS than older SATA drives because the protocol allows thousands of commands to queue at once instead of one at a time. The number that actually predicts how snappy your system feels, though, is low queue depth random IOPS (QD1 to QD4), not the six-figure peak figures used in marketing slides.


TL;DR:

  • Databases and transactional systems need random 4 KiB reads and writes plus mixed workload results, while video editing and backups depend more on sequential throughput.
  • Peak figures often come from QD32 to QD256 tests with multiple threads, and SLC caching can hide slower sustained writes after its cache fills.
  • For a useful comparison, run random 4 KiB reads and writes at queue depths 1, 4, and 16 for at least 60 seconds.
  • Match thread count to your workload, use a test file larger than cache, and record average and P99 latency alongside IOPS.
  • PCIe 5.0’s sequential speeds above 11,000 MB/s help large transfers, but PCIe 4.0 can feel nearly as responsive on small random workloads.

AceRDP
Put NVMe Performance in Context
 
Run databases, development tools, or remote workloads on Windows RDP and KVM VPS hosting powered by AMD Ryzen and NVMe storage.
Explore AceRDP servers

Table of Contents

What IOPS Measures and Why It Matters for Real Workloads

An I/O operation is a single read or write request sent to a storage device. Benchmarks almost always report IOPS using small, fixed-size requests, typically 4 KiB blocks, because that size mirrors how operating systems and databases actually move data around. A 4 KiB random read IOPS number tells you how a drive handles scattered, unpredictable requests rather than one long sequential transfer.

Some workloads live and die by IOPS. Others care more about raw throughput.

  • Databases and OLTP systems issue constant small random reads and writes, so IOPS and latency dominate performance.
  • Mail servers and busy web applications generate many small concurrent requests across many files, which stresses IOPS the same way.
  • Video editing, backups, and large file transfers depend on sequential throughput (MB/s), where IOPS matters far less.

When you’re picking storage for a transactional workload, chase IOPS and latency first. When you’re moving large files, throughput numbers tell you more. Mixed workloads, which describes most production servers, need both figures read together rather than either one in isolation.

How NVMe’s Architecture Produces High IOPS

NVMe reaches high IOPS because it was built as a storage protocol for flash memory rather than adapted from spinning-disk standards. The older AHCI protocol used with SATA SSDs supports a single command queue holding 32 commands. NVMe supports up to 64,000 queues, each holding up to 64,000 commands, which lets the drive process a massive number of requests in parallel instead of waiting in line.

NVMe also connects directly to the PCIe bus rather than routing through a SATA controller, cutting out a translation layer that adds latency and CPU overhead. The command set itself is leaner: fewer CPU instructions are needed to issue and complete each I/O, which matters enormously at high request rates where instruction overhead adds up fast.

Interrupt handling plays a quieter but real role. NVMe uses MSI-X interrupts, which let the drive signal completion to a specific CPU core rather than forcing all interrupt traffic through one shared path. Combined with per-core queue processing, this means a multi-core system can service NVMe I/O completions across cores in parallel instead of serializing them through a single interrupt handler. Cisco’s overview of NVMe describes this combination, parallel queues, PCIe attachment, and reduced CPU cycles per I/O, as the core reason NVMe achieves lower latency and higher IOPS than AHCI-based SATA drives. None of this is exotic engineering anymore; it’s the baseline architecture every consumer and enterprise NVMe drive shares.

How NVMe's Architecture Produces High IOPS — overview diagram

IOPS vs Throughput vs Latency: How to Read the Numbers

IOPS, throughput, and latency describe three different things, and confusing them leads to bad hardware decisions. IOPS counts operations per second. Throughput (measured in MB/s) counts data volume per second. Latency measures how long a single operation takes to complete, usually in microseconds or milliseconds for NVMe drives.

Queue depth links all three together. At queue depth 1, the drive processes one command, waits for completion, then accepts the next, so latency directly limits IOPS. As you raise queue depth (more outstanding requests allowed at once), aggregate IOPS climbs because the drive overlaps work, but the average latency per request typically rises too, since more commands are waiting their turn.

IOPS throughput and latency across queue depth

A practical rule of thumb: for latency-sensitive, single-threaded workloads like interactive desktop use or a single database connection, prioritize low-QD IOPS and latency. For workloads that genuinely run many simultaneous operations, like a busy multi-tenant database server or a virtualization host, high-QD aggregate IOPS becomes more representative. Throughput only becomes the deciding metric when you’re moving large sequential files rather than scattered small ones.

Why Vendor Headline IOPS Can Be Misleading

Manufacturer spec sheets often list IOPS numbers measured under conditions no desktop or typical server workload ever produces. Independent benchmark tables from Tom’s Hardware show modern PCIe 5.0 drives posting random IOPS well above a million, but those figures come from high queue depth tests with many threads, not the queue depth 1 or 2 behavior that governs most real application responsiveness.

Random 4 KiB read IOPS at low queue depths, not peak high-queue-depth figures, correlates most closely with everyday responsiveness, according to Tom’s Hardware’s benchmark comparisons. A drive that dominates a QD256 synthetic test can feel nearly identical to a cheaper drive once you’re running a single application thread.

A few other factors inflate headline numbers further:

  • Test parameters like QD32 through QD256 with multiple threads are common in marketing benchmarks and rarely reflect single-user workloads.
  • SLC or pSLC caching lets a drive absorb bursts of writes at very high speed temporarily, then drop to a lower sustained rate once the cache fills, a pattern Tom’s Hardware’s WD Black SN850 review documents clearly.
  • Mixed read/write tests at realistic queue depths expose behavior that pure sequential or pure random tests hide.

When comparing drives, look for QD1/QD2 random IOPS, mixed read/write results at moderate queue depths, and sustained (not burst) write performance. Those three lines tell you far more than a single peak IOPS number ever will.

— AceRDP

Measuring NVMe IOPS in Practice

Testing a drive yourself is the only way to know how it behaves under your actual workload, and the tooling for this is mature and free. DiskSpd on Windows and FIO on Linux are the two standard tools engineers reach for. DiskSpd is a command-line utility from Microsoft that generates configurable I/O patterns and reports IOPS, throughput, and latency percentiles directly. FIO does the same on Linux with a deeper set of job-file options for mixing read/write ratios and access patterns.

A repeatable test sequence looks like this:

  1. Create a large, pinned test file, ideally sized to exceed any cache so you’re measuring the drive itself rather than DRAM or SLC cache behavior.
  2. Run random 4 KiB reads and writes at queue depth 1, then repeat at QD4 and QD16 to see how the drive scales with concurrency.
  3. Hold each test run for at least 60 seconds, since Microsoft’s DiskSpd guidance notes shorter runs produce noisy, unrepresentative results.
  4. Vary thread count to simulate the concurrency your actual application generates, rather than maxing out threads just to inflate the IOPS figure.
  5. Record average latency and P99 latency alongside IOPS, since a high IOPS number paired with a bad tail latency often means inconsistent real-world performance.

Before running anything, isolate the device from competing I/O, disable write caching layers you don’t intend to test, and confirm the working set actually fits on the tier you’re measuring. Microsoft’s storage tiering guidance points out that if your test data spills onto a slower tier, your IOPS numbers will reflect that slower tier instead of the SSD you meant to test.

Pro Tip: Run your QD1 and QD4 tests first; they usually predict real application feel far better than any high-QD number you’ll see later in the same test suite.

Real-World IOPS Expectations by PCIe Generation

Generation-to-generation IOPS gains look dramatic on paper, narrower in practice. Samsung’s 980 PRO product specifications list PCIe 4.0 sequential read speeds typically between 5,000 and 7,500 MB/s, while PCIe 5.0 drives push sequential reads above 11,000 MB/s according to the same manufacturer’s specification pages for current generation parts. Those are sequential throughput figures, not random IOPS, and they come from manufacturer testing rather than independent, low-QD measurement.

  • Sequential bandwidth gains between PCIe 4.0 and PCIe 5.0 genuinely matter for large file transfers, video workflows, and bulk data movement.
  • Random QD1 IOPS gains between generations tend to compress significantly, since a single outstanding request is often limited by controller and NAND latency rather than raw PCIe bandwidth.
  • Tom’s Hardware’s Crucial T700 review found that while high-QD random IOPS scaled impressively on newer PCIe 5.0 drives, QD1 differences between generations were far smaller than the headline numbers suggested.

If your workload is bandwidth-heavy, like large sequential backups or video scrubbing, a newer PCIe generation is worth chasing. If your workload is mostly small random operations at low concurrency, a well-tuned PCIe 4.0 drive can feel nearly as responsive as a newer PCIe 5.0 part.

Interpreting NVMe IOPS for Specific Workloads

Different workloads should push you toward different parts of a spec sheet. Databases, VPS environments, and desktop systems each have their own priority metric.

  • Databases and OLTP systems need low-latency QD1 to QD4 random IOPS along with mixed read/write performance, since transactions rarely batch into large sequential writes.
  • VPS and multi-tenant hosting environments need you to think about per-VM contention and scheduler behavior, since the IOPS a single tenant actually receives depends on how the host allocates and schedules storage access across every other VM sharing the same physical NVMe pool.
  • Desktop and interactive workloads are governed almost entirely by QD1 random IOPS and latency, because a single user thread rarely generates deep queues, which is why two drives with wildly different marketing IOPS can feel nearly identical during normal use.

For database engineers specifically, Nvme notes that random read IOPS at 4 KiB blocks and low queue depths is the most meaningful number for hosting and database workloads, more so than any headline figure measured at QD32 or above. If you’re evaluating storage for a SaaS backend, the way your database engine issues I/O matters as much as the drive itself; a partner breakdown of SQL versus NoSQL workload signals walks through how access patterns differ enough to change which storage metric should drive your decision.

How NVMe IOPS Shapes the Way We Build VPS Plans

We build our VPS plans around NVMe storage specifically because low-QD latency, not headline IOPS, is what users experience when running a database, a remote desktop session, or an automation workload. Our infrastructure pairs that NVMe storage with direct, low-overhead access rather than layering in unnecessary abstraction that adds latency.

Operationally, overprovisioning, I/O scheduler tuning, and active monitoring help keep per-VM performance consistent rather than letting a single noisy tenant degrade everyone else’s queue depth. If you want to see how this plays out at scale, our write-up on KVM virtualization benefits covers a high-IOPS testing scenario in more depth, and our NVMe vs SATA comparison breaks down the practical gap further.

We’d rather you test a plan with your own FIO or DiskSpd job file than take a spec sheet’s word for it, and our knowledgebase walks through getting a server running quickly so you can run exactly that kind of benchmark yourself.

FAQ

How many IOPS is good?

There’s no single universal number since the right figure depends on workload and queue depth, but for most interactive and database workloads, strong QD1 to QD4 random 4 KiB read IOPS paired with low latency matters more than a high headline figure. A drive showing solid performance at low queue depths will typically feel responsive regardless of how it ranks in high-QD synthetic tests.

Is 2280 faster than 2230?

The numbers describe physical M.2 form factor length, not a performance standard on their own. In practice, larger sizes allow more room for NAND, DRAM, and thermal mass, so physically larger drives often sustain performance better under load, though raw interface speed depends on the PCIe generation and controller rather than the form factor itself.

Is NVMe slower than SSD?

NVMe drives are a type of SSD, so the comparison is really NVMe versus SATA-based SSDs. NVMe SSDs connect directly over PCIe with a parallel queue architecture, which gives them notably lower latency and higher IOPS than SATA SSDs limited by the older AHCI protocol and a single 32-command queue.

What is the read speed of a 1TB NVMe SSD?

Sequential read speeds vary by PCIe generation: manufacturer specifications show PCIe 4.0 drives typically reaching 5,000 to 7,500 MB/s, while PCIe 5.0 drives can exceed 11,000 MB/s in the same manufacturer’s current lineup. These are sequential figures measured by the manufacturer and won’t necessarily match random IOPS performance at low queue depths.

Sources

AceRDP
Ask About Storage for Your Workload
Email AceRDP Support with questions about Windows RDP or KVM VPS hosting for your storage-intensive applications.