The Quantum Engineer

35. Surface Codes

35.1Motivation

Why did one code family win the industry? Constraints of real hardware. Qubits live on a 2D chip and interact only with neighbors, so checks must be local. Control is imperfect, so checks should be small (weight 4) and the code should tolerate imperfect two-qubit gates. And the margin must be wide: the surface-code threshold (threshold) near 1% sits about ten times above today's best two-qubit gate error rates (~10⁻³), so corrections outpace damage. The costs are real — one logical qubit per ~2d² physical qubits is a poor rate — but every alternative (color codes, qLDPC) buys rate by demanding longer-range interactions. Kitaev's 1997 toric code, planarized, became the default because it matched the machine.

35.2Lattice construction

The surface code places data qubits on the edges of a square lattice. Two families of checks live on it: star operators at vertices and plaquette operators on faces (35.5, 35.6). On a closed torus this is the toric code with two logical qubits; cut the torus open into a planar patch and you get the surface code with one. Hardware implements the rotated layout — the same checks folded into a d×d grid of data qubits with interleaved measure qubits — but the edge-and-face picture is the cleaner mental model:

      *-----0-----*-----1-----*         *  vertex: star check
      |           |           |            = X on incident edges
      2     []    3     []    4         [] face: plaquette check
      |           |           |            = Z on bounding edges
      *-----5-----*-----6-----*         -,|  data qubits (one per edge)
      |           |           |
      7     []    8     []    9         on a torus there are no
      |           |           |         boundaries; a planar patch
      *----10-----*----11-----*         has two kinds (35.8)

35.3Data qubits

The data qubits — d² of them in the rotated layout — are the physical qubits that actually store the logical state. During a memory experiment they are never measured directly; they are only ever entangled with checks and left alone. Their errors are the object of the whole exercise, and each error class talks to exactly one check family: an X error on a data qubit flips the plaquette (Z-type) checks adjacent to its edge, while a Z error flips the star (X-type) checks at its endpoints. This complementary response is what lets a single syndrome history localize both error types independently — the two-dimensional generalization of the parity logic of 32.5.

35.4Syndrome qubits

Between the data qubits sit d²−1 measure qubits (measure qubits, ancillas), checkerboard-assigned to either X-type or Z-type duty. Each round, every ancilla is reset, entangled with its (up to) four neighboring data qubits, and measured. Ancillas are not passive probes: their reset errors, gate errors, and especially measurement errors corrupt the syndrome itself. A flipped measurement is indistinguishable, locally, from a real data error — which is why the decoding problem lives in space-time (35.9) and why memory experiments repeat the round at least d times, so that measurement noise shows up as defects separated along the time axis.

35.5Star operators

A star operator A_s is the product of X on the (up to) four data qubits incident to vertex s — for an interior vertex of the lattice above, X0 X1 X2 X3. Stars are X-type checks: they detect Z errors on their incident edges and are blind to X errors. All stars mutually commute, because any two vertices share 0 or 2 edges — an even overlap (34.2) — and stars commute with all plaquettes by the same even-overlap argument across the two Pauli types. At boundaries, stars truncate to weight 2 or vanish (35.8). Measuring all stars once per round costs one ancilla each and four CNOTs — the cheapest possible local check.

35.6Plaquette operators

A plaquette operator B_p is the product of Z on the four data qubits bounding face p — for the first face of the lattice above, Z0 Z2 Z3 Z5. Plaquettes are Z-type checks detecting X errors, mirror images of the stars in every respect: weight 4 in the bulk, mutually commuting, weight ≤ 2 at boundaries. Together, (d²−1)/2 stars and (d²−1)/2 plaquettes form the stabilizer group of the rotated surface code; measuring all of them once constitutes one syndrome round. The code's full stabilizer description is thus a list of weight-4 Pauli strings on a grid — small enough to fit in a header file, which is roughly how you should think of it.

35.7Logical qubits

On the torus, two independent logical qubits live in the topology: logical operators are non-contractible loops of Paulis winding around the torus, and no local error can create or destroy one. On a planar patch, one logical qubit remains, and its operators are open strings: an X-type string ending on the two boundaries where plaquette checks terminate (call them Z-boundaries) creates no plaquette defects and acts as logical X̄; dually, a Z-string between the two X-boundaries is Z̄. A string of length below d creates detectable defects; one of length d spans boundary to boundary and is undetectable. The logical qubit is stored in topology, not in geometry — deform the string freely, it stays the same operator.

35.8Boundaries

The boundaries are what make a planar code workable — and are the part most often glossed over. Two kinds exist. At an X-boundary, star checks terminate (some literature: "rough"); at a Z-boundary, plaquette checks terminate ("smooth"). Checks never straddle a boundary, and each logical string type must end on its own kind: X̄ spans the two Z-boundaries, Z̄ spans the two X-boundaries. This matching of operator type to boundary type is what makes one of the two logical directions undetectable while the other stays protected. The rotated layout arranges alternating boundaries around a square patch, which is why its picture in papers looks like a diamond — same code, tighter geometry, one fewer qubit row than the unrotated drawing.

35.9Decoding

Errors are chains: an X-error chain flips the plaquette checks at its two endpoints, creating a pair of defects; a measurement error flips one check in two consecutive rounds, creating defects separated in time. The decoder's input is this 3D defect pattern (2D space plus rounds); its output is the most likely correction. The canonical approach, minimum-weight perfect matching (37.3), treats defects as nodes in a graph whose edge weights encode the probability of the connecting chain, then pairs them up minimizing total weight. Union-find decoders trade a little accuracy for near-linear speed. Every one of these ideas already works in one dimension:

35.10Code distance

The distance d is the length of the shortest logical operator — the shortest chain of errors connecting two same-type boundaries — and the code corrects any ⌊(d−1)/2⌋ errors. Resources scale as d² plus the same again in measure qubits, and catching measurement errors requires running at least d rounds. The standard resource table, with the heuristic per-cycle logical error rate of 35.11 at p = 10⁻³:

distancedata qubitsmeasure qubitstotallogical error per cycle at p = 10⁻³
39817≈10⁻³
5252449≈10⁻⁴
7494897≈10⁻⁵

Read the last column as the price list of error correction: one more unit of distance buys one more decimal digit of reliability, at the cost of roughly 30 more physical qubits.

35.11Logical error rates

Below threshold, logical error per cycle follows a heuristic that decades of simulation and now hardware confirm:

p_L(d)  ≈  A * (p / p_th)^((d+1)/2)        A ≈ 0.1 for the surface code

Each step in distance squares-and-a-half's the error. At p = 10⁻³ against p_th ≈ 10⁻²: d = 3 gives ~10⁻³, d = 5 gives ~10⁻⁴, d = 7 gives ~10⁻⁵ per cycle — the table of 35.10. Above threshold the curves invert: bigger codes are worse, and the crossing point defines p_th ≈ 1% (idealized) or 0.5–0.7% (realistic circuit-level noise). The 2024–2025 landmark: Google's Willow chip ran d = 3, 5, 7 patches and measured suppression by Λ ≈ 2.14 per distance step, logical error ≈ 0.1% per cycle at d = 7, and a logical qubit outliving the best physical qubit on the chip — the first unambiguous below-threshold demonstration.