The Quantum Engineer

Project — Implement a Surface-Code Simulator and Decoder

Build the full QEC pipeline yourself, from noise to logical error curves, in Python with numpy only. External tools — stim, pymatching, sinter — are permitted as oracles for validation and comparison, never as the core: the point is that every number in your README is produced by code you understand line by line. This project is the level-3 (implementation) and level-4 (engineering) destination of Part X, and the strongest single piece of evidence you can put in a portfolio aimed at QEC roles.

Scope and architecture

A command-line tool with three layers, kept strictly separated:

  • noise.py — circuit-level noise model: per-gate depolarizing, per-reset and per-measurement flip probabilities.
  • code.py — the rotated surface code at distance d: data and measure qubit layout, check-to-qubit adjacency, one syndrome-extraction round.
  • decode.py — decoders operating on a detector-event history (a 3D array: space × space × time).

The simulator tracks Pauli frame only (bitset algebra, 34.2) — never state vectors. That is not a shortcut; it is the same design decision production simulators make, and it is what lets you run 10⁵ shots for d = 7 on a laptop.

Milestones
  1. M1 — Repetition code end to end. 1D chain, perfect measurements, min-weight decoding (35.9 experiment generalized to any d). Acceptance: measured p_L matches 3p^2 - 2p^3 at d = 3 within Monte Carlo error; curves for d = 3, 5, 7, 9 cross at p = 0.5.
  2. M2 — Surface-code layout and syndrome rounds. Rotated d = 3 and d = 5: build the check adjacency, simulate one memory experiment with phenomenological noise (data errors + measurement flips). Acceptance: with noise off, zero defects; with measurement noise only, defects appear in time pairs; logical operator verification by construction (the two boundary-spanning strings of 35.8).
  3. M3 — Matching decoder. Implement the min-weight pairing for the 2D per-round graph, then extend to space-time matching across rounds (pairing defects including time-separated pairs). Core must be yours: a blossom implementation or, acceptably, union-find with a documented accuracy trade-off. Acceptance: on d = 3, your decoder's p_L agrees with pymatching within error bars on 10⁵ shared shots.
  4. M4 — Threshold scan. Sweep p from 0.5% to 2% at circuit-level noise, d ∈ {3, 5, 7}, ≥ 10⁴ shots per point. Acceptance: curves cross between 0.5% and 1.5% (your threshold estimate, with error bars); below the crossing, Λ > 1 and the spacing fits A(p/p_th)^((d+1)/2) within a factor of 2.
  5. M5 — Report. README with plots (logical vs physical error, Λ fit), runtime table (shots/second per d), and an honest limitations section. Acceptance: a reader can reproduce every figure from the repo in under 30 minutes.
Validation protocol

After each milestone, run the same experiment through stim + pymatching and diff the outputs: detector counts per shot, matching weights, logical outcomes. Any unexplained discrepancy is a bug in your code or in your mental model — both worth their weight in gold. Document three such discrepancies and their root causes in the README; that section will teach an interviewer more than the plots.

Stretch goals
  • Circuit-level hook errors: show the logical error rate worsen with naive CNOT ordering at fixed distance — the 36.4 lesson made visible.
  • Lattice surgery: implement a logical CNOT between two patches and measure its logical error rate against the memory baseline.
  • A trained neural decoder (small MLP on detector histories) compared against matching on accuracy, with an honest latency measurement.
  • A qLDPC comparison: decode the same syndrome data with BP+OSD and discuss where matching fails structurally.
  • Scale test: d = 11 with the union-find decoder; report where your implementation stops being real-time and why.

Time budget: 3–6 weeks part-time. The deliverable is a public repository with tests, plots, and the validation log — the artifact that turns "I read Part X" into "I built Part X".