The Quantum Engineer

53. Reproducing Research

53.1Why Reproduction Matters

Reproduction is how you convert reading into capability: a reproduced result is knowledge you own — every assumption, every constant, every failure mode. It is also the field's weakest link: quantum papers routinely omit gate counts, seed choices, exact parameter sweeps, and calibration contexts, meaning many published numbers are unreproducible in detail even when the science is right. That gap is your opportunity: a careful reproduction — with deviations documented — is a genuine contribution (see 53.8), needs no laboratory beyond your laptop for simulation-based work, and is precisely the skill set a software engineer already has: you have been reverse-engineering systems and rebuilding pipelines your whole career. Research groups hire people who can make claims true; show them a reproduction portfolio.

53.2Selecting a Paper

Choose your first reproductions strategically. Criteria: (1) simulator-reproducible — the core claim computable on a laptop (algorithm papers, compiler benchmarks, decoder evaluations — yes; hardware demonstrations — partially, via noise models); (2) code-adjacent — has a GitHub repo, even a messy one, to check against (not copy from — see 53.7); (3) bounded scope — one algorithm, one benchmark, one figure to match; (4) verifiable — has numbers you can compare against exactly. Sweet-spot paper types: variational algorithm benchmarks, routing/pass-manager comparisons, decoder threshold studies, small-chemistry VQE. Avoid for now: multi-institution hardware results (too many hidden calibration details) and purely existential theory (nothing to run). Start with reproducing one figure, not one paper.

53.3Extracting the Algorithm

Turn prose into specification before writing code. The extraction checklist: every equation (instance them at small n by hand); every parameter (value, sweep range, initialization, seeds); the exact pipeline order (the same gates in a different order are a different experiment); what's measured (observable, basis, shots) and what's plotted (mean over seeds? best-of-k? error bars' meaning); and the computational environment (library versions matter — Qiskit 0.x vs 1.x changes results). Where the paper is silent — and it will be — write "UNSPECIFIED" in your spec and keep a list: that list is your deviation ledger's seed, and often more valuable than the reproduction itself. The discipline is: never write code from memory of the paper; code only from your written spec.

53.4Rebuilding the Experiment

Write it as you wish papers were written: a repo with spec.md (the extraction above), src/ (implementation with tests — Ch. 16.15's property tests apply directly), run.py (one command reproduces the target figure), and results/ (outputs + the comparison plot). Build in two stages: first an ideal implementation targeting the paper's theory curves — this validates your understanding; then the experimental details — noise, sampling, seeds — targeting the empirical curves. Keep the two code paths separate; conflating them is how you debug for a week what is really a modeling mismatch. Version everything; pin dependencies (requirements.txt with exact versions); and run your reproduction end-to-end from a clean clone before believing it — the "works on my machine" failure mode is identical in research.

53.5Matching Parameters

Most reproduction failures are parameter failures. Hunt them systematically: system size (n qubits — papers sweep it; which n is the figure?), circuit depth (p for QAOA, layers for ansätze), shot counts, seeds and number of restarts (a "best of 20 runs" curve is a different object than a mean), optimizer settings (learning rates, iteration caps), and the noise regime (error rates assumed, readout error included?, which noise model — Ch. 54.6's taxonomy). For hardware papers: transpilation settings (optimization level changes everything — Part XII) and the calibration snapshot date. Where a parameter is unknowable, sweep it: reproduce the figure across the plausible range and report the band — an honest "the reported result is consistent with ε ∈ [0.3%, 0.8%]" often resolves a week of confusion in one plot.

53.6Statistical Comparison

When is a reproduction a success? Never "the curves look similar." Quantitative standards: for distribution-level claims, compare with the paper's own statistic (KL divergence, XEB score, fidelity estimator); for point estimates, is your value within the reported error bars — and if none were reported, compute the confidence interval the authors should have (shot noise alone, Ch. 16.10); for benchmark claims, reproduce the distribution over seeds, not the best run. The gold standard: your reproduction with fresh seeds overlaps the paper's within joint uncertainty. When it doesn't, you have either a bug, an unspecified parameter, or a fragile claim — and distinguishing those three is itself a research result (53.7). Statistics is the arbiter; bring the same rigor you'd bring to an A/B test.

53.7Documenting Deviations

Your deviation ledger — every place your reproduction differs from the paper — is the reproduction's scientific payload. Format: for each deviation, the paper's value, your value, why (unstated / infeasible / you chose differently), and the measured impact (rerun with and without — impact analysis turns complaints into data). Common findings: unstated hyperparameter tuning, results that hold only for specific seeds, noise models more favorable than claimed, and baselines weaker than standard practice (Ch. 50.8 again). Tone: neutral, specific, verifiable — "we could not reproduce Figure 3b's 0.92 under the stated parameters; sweeping learning rate yields 0.85–0.91" is science; "the paper is wrong" is not. Deviations documented well make your work citable; devuations documented badly make it ignorable.

53.8Publishing Reproduction Results

Reproductions are publishable and increasingly valued: venues welcome them (ReScience journal exists for exactly this; workshop papers at QIP/IEEE Quantum Week; blog posts with code are genuinely cited in this field). The publication package: repo (clean, one-command, pinned deps), spec + deviation ledger, comparison figures with uncertainty, and a short write-up stating what was confirmed, what was refined, and what remains open. Norms: contact authors first when you find discrepancies (they usually respond — and their replies often resolve the mystery); never publish a "failure" claim without giving the method its best case; license permissively. A portfolio of three careful reproductions in your niche is the single strongest artifact a self-taught researcher can present to a group — stronger than certificates, weaker than nothing only if done carelessly.