Ask a quantum computing company how good its machine is and you will rarely get a simple number. You will get several, and they will not agree with each other. One firm leads with qubit counts, another with error rates, a third with a benchmark it more or less invented. The result is a confusing marketplace where every vendor can plausibly claim to be ahead. Understanding the main yardsticks is the only way to read past the marketing.
Why raw qubit counts fail
A qubit that flips state on its own, or that talks poorly to its neighbors, contributes little to a useful computation. A processor with 100 noisy qubits can be outperformed by one with 30 clean ones, because errors compound as a circuit grows. Each additional gate is another chance for the calculation to drift into nonsense. So the real questions are how reliably the qubits operate and how many operations you can chain together before the answer dissolves into noise.
That is why fidelity sits underneath almost every serious benchmark. Two-qubit gate fidelity, the accuracy of an entangling operation, is the figure engineers obsess over. The best trapped-ion and neutral-atom systems now report two-qubit fidelities above 99.5 percent, with leading superconducting devices in a similar range. The difference between 99 and 99.9 percent sounds trivial until you run a circuit thousands of gates deep, where it decides whether the machine works at all.
Quantum Volume and its rivals
IBM introduced Quantum Volume to capture more than a single dial. The metric runs random circuits that are as wide as they are deep, and asks how large a square circuit the machine can execute before the output stops resembling the correct distribution. A higher number means the processor can handle both more qubits and more operations together. It bundles qubit count, connectivity, gate quality, and measurement accuracy into one figure, which is its strength and its weakness. It is holistic, but it is also a single number standing in for a complex device.
Speed got its own metric. CLOPS, or circuit layer operations per second, measures how many circuit layers a system can run in a given time. This matters because real applications often require running the same circuit thousands of times to build up statistics. A processor with a beautiful Quantum Volume score but a glacial cycle time may be impractical for anything serious. Trapped-ion machines, prized for their accuracy, have historically been slower than superconducting ones, and CLOPS is partly where that tradeoff shows up.
Algorithmic qubits and application benchmarks
IonQ pushed a different idea with algorithmic qubits, meant to express how many effective, high-quality qubits are available for actual algorithms rather than raw hardware qubits. Critics argue such vendor-defined metrics are easy to tune in your own favor. Defenders say they communicate something users care about more than abstract circuit games: can this machine run the program I want?
That tension is pushing the field toward application-level benchmarks. Instead of random circuits, these run representative tasks drawn from chemistry, optimization, or simulation and report how well the hardware solves them. Efforts like the QED-C benchmark suite and academic projects such as SupermarQ try to standardize this approach across vendors. The appeal is obvious. A buyer evaluating a chemistry workload cares about chemistry performance, not Quantum Volume in the abstract.
What to watch for
When you read a quantum benchmark claim, a few questions cut through the noise:
- Is the number measured on the full processor, or on a hand-picked subset of the best qubits?
- Does it report two-qubit gate fidelity, the figure that limits deep circuits?
- Does it account for speed, or only for accuracy?
- Is the benchmark independently defined, or invented by the company quoting it?
None of these metrics is dishonest on its own. The problem is selective emphasis. A vendor with excellent fidelity but modest scale will talk about fidelity. One with many qubits will talk about counts. The honest comparison requires looking at several numbers at once and noticing which ones a company chooses not to mention.
As the industry moves toward error-corrected logical qubits, the benchmark debate will intensify rather than settle. Measuring logical error rates and the overhead of correction introduces a new layer of metrics, and history suggests each company will again favor the one that flatters its architecture. For now, treat any single headline figure as the start of a question, not the answer.