54. Designing Experiments
54.1Hypotheses
An experiment is a question with a falsifiable answer, stated before data collection: "noise-aware routing reduces two-qubit gate count on heavy-hex by ≥10% for QFT circuits at n=5–12," not "we explore noise-aware routing." Structure each hypothesis as claim + scope + measure + threshold; the threshold is what keeps you honest when results come back ambiguous. Prefer paired hypotheses (H1: the effect exists; H0: it doesn't) and decide your statistical test now, not after seeing the data. The quantum-specific twist: your hypothesis must state the regime — ideal simulation, noisy simulation with named model, or named hardware with calibration date — because a result true in one regime may be false in another, and most published confusion is regime confusion.
54.2Baselines
An experiment without a baseline is an anecdote. The baseline ladder for quantum work, ascending: the ideal simulation (what does perfect execution give? — your theory curve), the naive implementation (default settings — what does no effort give?), the standard implementation (well-tuned known method — what does the current state of the art give?), and the classical competitor where one exists (Ch. 50.8's discipline; even for hardware experiments, "what would brute-force simulation of this instance cost?" is a mandatory column). Baselines must receive equal tuning effort — a quantum method beating an untuned classical baseline is the field's most-published non-result. Pre-register baselines in your experiment notes; adding a weaker baseline after seeing results is self-deception with paperwork.
54.3Controlled Variables
Change one thing at a time — but in quantum experiments, "one thing" is slippery because everything couples: qubit count changes routing, routing changes depth, depth changes noise sensitivity. The control discipline: enumerate every variable (n, circuit family, transpiler settings, seeds, shots, noise model, error rates, measurement basis, post-processing); fix all but the studied one; and — the professional habit — report the fixed values, because "at optimization level 3 with SABRE, seed-averaged" is a different universe from "default settings, one seed." Where variables cannot be decoupled (they often can't), sweep the confounder explicitly and present the result as a band, not a point. Factorial designs (small grids over 2–3 parameters) cost little at simulator scale and reveal interactions your intuition missed.
54.4Datasets
For circuits: use established corpora (QASMBench, BenchPress, machine-generated families — QFT-n, random circuits with fixed structure) and say which instances and why; cherry-picking circuits that flatter your method is undetectable in print and fatal in replication. For application data (chemistry: which molecules and basis sets; ML: which datasets — with the Ch. 50.8 caveats) prefer standard benchmarks over bespoke ones, and report dataset size honestly against compute cost. Every dataset gets a manifest: source, version, preprocessing, and — for generated data — the generation script and seed. The acid test: could a stranger reconstruct your exact instances from your paper? If not, your experiment is not yet an experiment; it is a demo.
54.5Simulator vs Hardware
The decision tree: use ideal simulation when testing logic and mathematics (fast, exact, debuggable — Statevector and Operator checks, Ch. 16); use noisy simulation when the claim involves noise (error mitigation, variational trainability, decoder behavior — cheap, reproducible, seedable); use hardware when the claim is about hardware (calibration effects, crosstalk, drift — the only ground truth, expensive and drifting). The rule: never test a hypothesis on hardware that could have been tested on a simulator, and never extrapolate a simulator result to hardware without naming the noise model that connects them. Budget realistically: laptop simulation is free, cloud hardware queues and quotas are real; design hardware experiments to answer only the questions simulation cannot.
54.6Noise Models
Your noise model is a hypothesis about the machine. The taxonomy, in ascending fidelity: depolarizing only (uniform error per gate — fine for logic tests, misleading for advantage claims); calibration-based (per-qubit, per-gate errors from real backend snapshots — the practical standard, NoiseModel.from_backend); structured (T1/T2 relaxation, readout error, crosstalk/spectator terms, leakage — needed whenever your method claims to handle noise); and learned (models fit to data, including ML-learned noise — the research frontier, Ch. 61). Match the model to the claim: a decoder validated only on depolarizing noise has not met a real machine. And always state the model's falsifier — what hardware behavior would break it — because that is what your hardware run is actually testing.
54.7Repetitions
Quantum results are distributions; report them as such. Sources of randomness to repeat over: measurement shots (statistical), circuit instances (structural — random circuits vary), seeds in randomized algorithms (SABRE, optimizers, sampling), and — on hardware — time (calibration drift makes morning and afternoon different machines). The discipline: ≥20 seeds for any randomized-algorithm claim; shot counts sized to the precision you report (Ch. 16.10's ±1% needs ~10⁴ shots); and repeats on hardware across at least two calibration windows if the claim is about average behavior. Report medians and interquartile ranges, not means-of-best. The most common honest mistake in the field is under-powered repetition — an effect "demonstrated" at 3 seeds is a rumor with a plot.
54.8Statistical Significance
The tests you'll actually use: two-sample comparisons (Mann–Whitney U — robust, no normality assumption; or Welch's t-test with sanity checks) for "method A beats B"; effect sizes with confidence intervals (bootstrap over seeds — the field's most underused tool) rather than bare p-values; and equivalence testing when your claim is "matches" (two one-sided tests — reproduction claims from Ch. 53.6 need this; overlapping error bars are not equivalence). Multiple-comparison hygiene: if you swept 10 configurations and report the best, say so (or correct — Benjamini–Hochberg). Pre-registering the analysis (54.1) is what makes any of this meaningful. You already know this culture from A/B testing; the physics changes none of it.
54.9Reproducibility
The checklist that makes your work a contribution instead of a screenshot: repo with one-command execution (make reproduce or equivalent); pinned dependencies and versions (Qiskit's 0.x→1.x break taught the field this); all seeds either fixed and committed or swept and reported; raw data committed (shots-level, not just aggregates — storage is cheap, reruns aren't); hardware results timestamped with calibration snapshots; and a README that states the claim, the regime, and the one figure that proves it. Internal standard: rerun your own six-month-old experiment before publishing — if you can't reproduce yourself, no one will. This culture is your unfair advantage: software engineers arrived at reproducibility discipline a decade before physics did, and the field knows it needs it.
54.10Experiment Tracking
Scale demands tooling: at more than a handful of runs, track experiments the way you'd track production jobs. Minimum viable: a structured results log (JSONL or SQLite — every run records config hash, git commit, seed, metrics, artifacts path) and a results/ directory with immutable outputs. Proper: MLflow or W&B (free tiers, local-first options exist) — quantum experiments fit their schema perfectly (parameters = circuit/optimizer/noise config, metrics = your measurements). The habit that matters most: every run is logged from the first day, because the runs you didn't log are the ones you'll need. And commit the tracking database with the paper — "the data behind Figure 3" as a queryable artifact is what separates modern reproducible work from the PDF era.