6. Probability and Statistics
6.1Classical probability
Probability theory formalizes randomness over a sample space: each outcome ω gets P(ω) ≥ 0 with Σ P(ω) = 1, and events get probabilities by addition. The two big theorems frame everything: the law of large numbers (sample frequencies converge to true probabilities as shots grow) and the central limit theorem (fluctuations shrink like 1/√N — the number behind every quantum experiment's error bars). Classical and quantum probability differ in exactly one axiom: in the classical (Kolmogorov) formalism, probabilities of disjoint events add directly, while quantum amplitudes add first and square after, enabling cancellation (interference). Everything else — random variables, expectation, variance, estimation — transfers unchanged. That is a genuinely useful statement for a programmer: your classical statistical toolkit applies to quantum data verbatim; only the mechanism generating the distribution differs.
6.2Random variables
A random variable is a function from outcomes to numbers: it assigns a value to each result of a random process. The die is the process; the "number rolled" is the random variable. In quantum computing, every measurement is a random variable: the outcome distribution comes from the Born rule, but once sampled, the data is ordinary — you analyze counts of 0s and 1s with the classical machinery, no quantum theory required. Distinguish the object from its realizations: the random variable is the distribution (over infinitely many hypothetical repetitions); your 10,000 shots are one sample of it. numpy gives you the standard ones — np.random.binomial, np.random.normal, np.random.poisson — and the habit to build is thinking in the random variable's distribution first, data second. Quantum measurement outcomes are typically Bernoulli (single bit) or multinomial (bit strings) variables; both are classical.
6.3Probability distributions
A distribution is the full description of a random variable: for discrete outcomes, the table of probabilities; for continuous ones, a density. Distributions have summaries — mean (center), variance (spread), shape (skew, tails) — and named families encode situations: Bernoulli for single bits, binomial for counts of successes in N trials, multinomial for bit-string counts, Poisson for rare events (dark counts, stray photons), normal for aggregated fluctuations. Quantum measurement in the computational basis produces exactly a multinomial distribution over 2ⁿ bit strings with probabilities |c_x|² — the state vector is a distribution you cannot read, and measurement is sampling from it. The engineering skill is matching summary to question: "did the circuit work?" is a question about one Bernoulli parameter; "what is the output distribution?" needs the multinomial, and far more shots. Print distributions, not single samples, whenever you can afford it.
6.4Expectation
The expectation E[X] = Σ x·P(X = x) is the probability-weighted average of a random variable over infinitely many repetitions — the mean of the distribution, not of your sample. It is linear: E[aX + bY] = a·E[X] + b·E[Y] always, a property whole algorithms exploit. Quantum computing computes expectations by the same formula with Born-rule weights: ⟨ψ|H|ψ⟩ = Σ λᵢ·⟨ψ|Pᵢ|ψ⟩ (4.18) is the expected measurement outcome of observable H on state ψ — the number that variational algorithms (chapter 46) minimize and that every chemistry estimate targets. The gap between expectation and estimate is the working reality: E[X] is a property of the distribution; your experiment returns a sample mean x̄ that approaches it as shots accumulate, with error governed by variance (6.5). Whenever a paper reports ⟨H⟩ to five decimal places, the first question is: on how many shots?
6.5Variance
Variance Var(X) = E[(X − E[X])²] measures spread around the mean; the standard deviation σ = √Var(X) is in the same units as X. Variance's crown property: for independent trials, variances add, so the sample mean of N shots has standard deviation σ/√N — the √N law that prices every quantum experiment. Measuring one observable with per-shot variance σ² and wanting accuracy ε costs N ≈ σ²/ε² shots: quadratic cost in precision, and the reason "just measure more precisely" is never cheap. Chebyshev's inequality turns variance into a guarantee: P(|X − E[X]| ≥ kσ) ≤ 1/k², distribution-free. For Bernoulli variables (a single bit outcome), Var = p(1−p) ≤ ¼, giving the universal floor N ≈ 1/(4ε²) shots for estimating any probability to accuracy ε. Budget experiments with this formula before running them; it is the difference between an estimate and a vibe.
6.6Conditional probability
Conditional probability P(A|B) = P(A ∩ B)/P(B) re-weights the sample space under the knowledge that B occurred — a narrower world, renormalized. In quantum terms, conditioning is post-selection: keep only the shots where qubit 2 read 1, then look at the statistics of the rest; the resulting conditional distribution is exactly P(rest | qubit₂ = 1). Two classic results anchor careful use. Bayes' theorem, P(A|B) = P(B|A)·P(A)/P(B), is the engine of 6.7. And the law of total probability, P(A) = Σⱼ P(A|Bⱼ)·P(Bⱼ), is how simulators mix over branches. The standard error is confusing P(A|B) with P(B|A) — "probability of data given hypothesis" versus "probability of hypothesis given data" — a confusion that entire reams of bad quantum-ML benchmarking embody. Post-selected data is conditional data: report the conditioning event's frequency alongside every conditional statistic, or the numbers are uninterpretable.
6.7Bayesian reasoning
Bayes' theorem upgrades beliefs with evidence: P(hypothesis|data) = P(data|hypothesis)·P(hypothesis)/P(data) — prior times likelihood, renormalized. For an engineer, Bayesian reasoning is the honest bookkeeping of how certainty accumulates: calibration of a qubit starts from a prior over pulse parameters, each experiment multiplies in a likelihood, and the posterior becomes the next prior. It is also the antidote to two failure modes of quantum-claims analysis: treating a single p-value as truth (6.12), and ignoring base rates — a spectacular result from a lab with a track record of spectacular results deserves different priors than one from an unvalidated source. Full Bayesian posteriors over quantum states or noise models are computationally heavy (the parameter spaces are large), which is why practice uses approximations — variational posteriors, Laplace approximations, or plain maximum-likelihood point estimates plus error bars. The discipline matters more than the machinery: state your prior, update on data, never reset it to suit a narrative.
6.8Sampling
Sampling is drawing realizations from a distribution — and it is the only output a quantum computer ever gives. You do not read amplitudes; you run shots and count bit strings. So the skill of designing a sampling experiment is as fundamental as the circuits themselves. The machinery: independent, identically distributed shots; counts per outcome follow a multinomial; empirical frequencies are the natural estimator (6.9). numpy covers classical needs: np.random.choice(2**n, size=shots, p=probs) simulates a full measurement; seeding (np.random.default_rng(seed)) makes experiments reproducible, which in a stochastic field is not optional — an unreproducible histogram is not a result. Watch for the two classic sampling sins: sampling until the result looks right (optional stopping biases estimates) and discarding samples without reporting how many were dropped (post-selection without reporting breaks every estimator downstream). Both are, statistically, self-deception with extra steps.
6.9Estimators
An estimator is a recipe that turns data into a guess about a distribution parameter; the estimate is its output on actual data. The sample mean x̄ = (1/N)·Σ xᵢ estimates E[X]; the sample variance estimates Var(X); empirical frequencies estimate Bernoulli p. Three properties grade estimators: bias (does the recipe target the right quantity on average — unbiased estimators are correct under repetition), consistency (does it converge to truth as data grows), and efficiency (how much data does it need for a given accuracy). The sample mean of independent shots is unbiased and consistent, with standard error σ/√N — quantified, not vibes. In quantum work, estimators appear wherever raw sampling is indirect: state tomography fits a density matrix to counts, randomized benchmarking fits an exponential decay to extract gate error, and variance-reduced estimators (amplitude estimation) trade extra circuit depth for fewer shots. Always report estimator and shot count together; an estimate without them is decoration.
import numpy as np
rng = np.random.default_rng(42)
p_true = 0.37 # unknown "quantum" success probability
shots = 10_000
data = rng.random(shots) < p_true # Bernoulli samples
p_hat = data.mean()
se = np.sqrt(p_hat * (1 - p_hat) / shots) # standard error of the mean
print(f"estimate {p_hat:.4f} ± {se:.4f}") # honest error bar, ~±0.0056.10Statistical uncertainty
Statistical uncertainty is the spread of an estimate caused by finite sampling — distinct from systematic error, which is a biased apparatus or model, and from numerical roundoff. The three must never be conflated: more shots shrink only the first. The tool is the standard error: for N independent samples with per-shot standard deviation σ, the sample mean's uncertainty is σ/√N. Halving your error bar quadruples your shots; a 10× tighter estimate costs 100× the runtime — the brutal economics behind every quantum experiment's precision claims. Confidence statements come from the same arithmetic (6.11). The engineering habit: attach an uncertainty to every number you report, propagate uncertainties through derived quantities (errors add in quadrature for products of independent estimates), and treat any quantum result quoted without error bars as unverified. In this field, the error bar is the result; the point estimate is just its center.
6.11Confidence intervals
A confidence interval turns an estimate plus its uncertainty into a calibrated claim: "the true p lies in [p̂ − 1.96·SE, p̂ + 1.96·SE] with 95% coverage" — meaning the recipe captures the truth in 95% of repeated experiments, not that this particular interval has a 95% metaphysical grip on p. The normal-approximation interval (±1.96·σ/√N) is standard for means with enough samples; for small counts or probabilities near 0 or 1, use exact binomial intervals (scipy.stats.binomtest) or the Wilson interval, because the normal approximation fails exactly where quantum experiments live — rare success events from deep circuits. In quantum computing, confidence intervals are the honesty layer: "fidelity 0.998 ± 0.001" and "fidelity 0.998 ± 0.05" are different universes of claim, and benchmark tables that omit intervals are hiding the second kind. Build the habit early: no number leaves your notebook without an interval.
6.12Hypothesis testing
Hypothesis testing asks a yes/no question of data: define a null hypothesis H₀ (no effect), compute how surprising the observed data would be under H₀ — the p-value — and reject H₀ when the surprise crosses a pre-chosen threshold α (typically 0.05). The machinery matters in quantum engineering for A/B comparisons: is gate set A's fidelity actually better than B's, or is the difference within noise? A two-sample test on per-shot outcomes answers it with a calibrated error rate. Two disciplines prevent the classic abuses. Fix the threshold and the analysis plan before looking — switching tests or endpoints after seeing data inflates false positives (the multiple-comparisons problem: run 20 benchmarks, expect one "significant" result by chance). And interpret p-values correctly: a p-value is P(data this extreme | H₀), not P(H₀ | data) — the Bayesian frame of 6.7 is the correct home for the question people usually mean to ask.
6.13Monte Carlo methods
Monte Carlo methods estimate intractable quantities by random sampling: to estimate an average, sample it; accuracy improves as 1/√N regardless of dimension — the property that makes Monte Carlo the only general tool in high-dimensional spaces, quantum state spaces emphatically included. The recipe family: direct sampling when you can draw from the target distribution (simulating a circuit's measurement outcomes); rejection sampling and importance sampling when you cannot, re-weighting draws from an easier distribution; and Markov chain Monte Carlo (MCMC) when even that fails, wandering the distribution with a random walk whose stationary state is the target. Applications you will meet: sampling from Born distributions to validate simulators, Bayesian posteriors over noise parameters (6.7), and estimating partition-function-like quantities in physics-inspired problems. The universal caveat is variance: naive Monte Carlo on a rare event wastes shots; importance sampling that shifts probability mass onto the rare region is the difference between 10⁶ and 10³ shots.
6.14Why quantum experiments are statistical
The closing synthesis: quantum mechanics makes statistics unavoidable, not incidental. Measurement is destructive and probabilistic — the Born rule hands you samples, never the amplitudes — so a quantum computation's output is a distribution you must estimate, with error bars, from repeated runs. Compounding this, no-cloning forbids re-measuring one copy; every shot consumes a fresh preparation, so information-per-shot is bounded by design. And hardware noise adds drift on top of sampling noise: your 10,000 shots embed both statistical uncertainty and a device state that may have shifted during them. The practical doctrine, used in every later chapter: define the observable, budget shots from the √N law (6.5), attach confidence intervals (6.11), and separate statistical from systematic error explicitly. A quantum engineer who cannot state an uncertainty is not doing quantum engineering — they are watching a slot machine.