The Quantum Engineer

49. Variational Quantum Algorithms

49.1Classical-Quantum Hybrid Computation

The architecture: a classical outer loop holds parameters θ; a quantum inner loop runs a parameterized circuit V(θ) and returns estimates (energy, expectation values); a classical optimizer updates θ; repeat until convergence. Rationale (circa 2017–2021, NISQ era): shallow circuits survive noise; variational principles give meaning to imperfect estimates. The architecture is also the field's most-revised idea: 2022–2025 trainability results (49.6) and noise-scaling analyses showed the honeymoon version doesn't scale as hoped, and the serious proposals that survive use problem-informed structure and, increasingly, assume error-corrected hardware anyway. Learn the loop anyway: it is the API through which most near-term experiments — and most of the field's learned lessons — flow.

49.2Parameterized Circuits

Ansatz design is the craft. Families: hardware-efficient (native gates arranged by topology — minimal transpilation cost, maximal trainability risk), problem-inspired (UCCSD for chemistry — physically motivated excitations; QAOA's alternating cost/mixer layers — provably connected to adiabatic evolution), and symmetry-preserving (ansätze confined to the physically relevant subspace — fewer parameters, no leakage). Parameter count: each is a dimension the optimizer must explore; hardware-efficient ansätze on 20 qubits routinely have 100+ angles. The design tension is a triangle — expressivity, trainability, noise resilience — and every published ansatz is a point in it. Your experiments will let you feel the triangle rather than memorize it.

49.3Optimization Loops

The classical half, done right: stochastic estimates (each energy evaluation is a sampling experiment with shot noise — Ch. 16.10 discipline), parameter-shift rule for gradients (∂E/∂θ = [E(θ+π/2) − E(θ−π/2)]/2 — exact on gate probabilities, two evaluations per parameter per observable), and optimizers: SPSA (gradient-free, noise-robust, the NISQ workhorse), COBYLA, Adam with parameter-shift gradients (shot-expensive but strong), natural-gradient variants. Engineering the loop is most of VQE in practice: shot budgeting per iteration, parameter-shift batching across terms, convergence criteria that respect statistical noise (a "converged" VQE that moved within its error bars never converged). Implement the loop once, bare-bones, before using libraries — the pathologies (49.6) are invisible through library abstractions.

49.4VQE

The variational quantum eigensolver, concretely: ansatz V(θ), Hamiltonian H = Σ hₖPₖ from 48.6, minimize E(θ) = Σ hₖ⟨Pₖ⟩_θ, report the minimum as the ground-energy upper bound. The canonical run — H₂ with a 1-qubit tapered ansatz — converges in ~20 iterations on a laptop simulator and matches FCI to 10⁻³ Hartree; it is the field's "hello world" and you should run it this week. Scaling realities: measurement cost grows with the number of Pauli terms (hundreds–thousands for small molecules; grouping/commuting-term tricks (48.6, 49.9) help constant factors); ansatz depth grows with system size; and noise biases E(θ) downward-comparable-to-signal at real error rates. VQE is a beautifully instructive algorithm whose scaling story is unfinished.

49.5QAOA

The variational algorithm for combinatorial optimization: encode a cost function (MaxCut, portfolio selection, scheduling) as a diagonal Hamiltonian C; alternate layers e^(−iγₖC) e^(−iβₖB) with mixer B = ΣXᵢ; optimize angles (2p parameters for depth p); sample the final state and post-select. Structure to appreciate: p → ∞ recovers adiabatic evolution (guaranteed optimum); finite p has approximation guarantees only for special cases (MaxCut p=1 on 3-regular graphs: 0.6924); and classical competitors are strong — quantum-inspired algorithms and straightforward heuristics (simulated annealing, Goemans–Williamson's 0.878 for MaxCut) set a bar finite-p QAOA has not beaten at scale. The engineering lesson in QAOA is transferable: cost-Hamiltonian compilation (Ch. 48.6-style), angle strategies, and distribution analysis. The lesson in honesty is transferable too.

49.6Barren Plateaus

The result that reshaped the field (McClean et al. 2018, with brutal 2021–2024 follow-ups): for many natural ansätze, the gradient variance vanishes exponentially in qubit count — the loss landscape is flat almost everywhere, optimization is guesswork, and the effect worsens with depth, entanglement, global cost functions, and — the 2023–24 results — with noise itself (noise-induced plateaus appear even for trainable ansätze and persist as hardware error shrinks but stays nonzero). Diagnostics you can run: train at n=4, 8, 12; plot gradient variance; watch it fall. Mitigations with evidence: local cost functions, shallow/problem-informed ansätze, layer-wise training, parameter initialization near identity, symmetry-preserving subspaces. The honest reading: variational training at scale is an open research problem, not a solved pipeline — a fact worth knowing before any "quantum AI" pitch meeting.

49.7Optimizer Selection

Match optimizer to noise regime. High shot noise, many parameters: SPSA (two evaluations per step regardless of dimension; calibrated perturbation sizes; provably robust to multiplicative noise). Low noise, few parameters: parameter-shift + BFGS-family (fast local convergence). Discontinuous/limited budget: Nelder–Mead (fragile, but a baseline everyone recognizes). Barren-plateau regime: nothing classical fixes a flat landscape — fix the ansatz first (49.6), then choose. Practical loop: run 3 optimizers × 5 seeds × your convergence budget on the simulator, pick by median final energy and wall-clock, and pre-register the choice before hardware runs. Optimizer selection is where statistics, software engineering, and quantum physics meet; treat it as an experiment, not a preference.

49.8Noise

What noise does to variational algorithms, precisely: biases every expectation value (energy estimates shift by O(ε·depth·‖H‖) — enough to overwhelm chemical-accuracy signals), reshapes landscapes (noise-induced plateaus, spurious minima that trap optimizers), and breaks the variational guarantee (the measured minimum may sit below the true ground energy — the classic tell of an unmitigated run). Countermeasures in order of cost: error suppression (dynamical decoupling, transpilation for noise — Part XII), error mitigation (ZNE, PEC — Ch. 31; overhead 10–1000× in shots), and error correction (changes the regime entirely — variational algorithms on logical qubits are a different, more hopeful literature). Reporting discipline: never publish a variational result without its noise floor — the same experiment on the ideal simulator is the floor.

49.9Measurement Overhead

The quiet killer. Estimating E(θ) = Σₖ hₖ⟨Pₖ⟩ requires estimating each term; with M terms and per-term error tolerance δ, the shot count scales as (Σ|hₖ|/δ)² — for realistic molecules, 10⁶–10⁸ shots per energy evaluation, times hundreds of optimizer iterations. Mitigations, each with real engineering behind: grouping (commuting Pauli terms measured in one basis — classically partition, quantumly rotate; reduces M to ~√M groups), classical shadows (Ch. 47-adjacent: randomized single-basis measurements reconstructing many observables — the 2020s' most transferable idea), and adaptive term selection (variance-weighted sampling). Implement grouping for a 10-qubit Hamiltonian and measure the reduction — a concrete, satisfying, résumé-grade experiment that also teaches why measurement, not gates, often dominates real variational costs.