- Status: Accepted
- Date recorded: 2026-10-06
- Decision in one line: Every HMC number on the site uses the 2023 settings (L = 20 leapfrog steps of ε = 1/L, identity mass matrix), because they reproduce the original chains and are reliable here. A tuning sweep on /diagnostics shows what that choice costs and what would work better.
Context
HMC.fn from the lecture code takes the number of leapfrog steps L and fixes the step size at ε = 1/L, so the trajectory length εL is always 1. The mass matrix is the identity. In 2023 I used L = 20 without testing alternatives, and the report said nothing about why.
The corrected posterior has two awkward features for an identity mass matrix. The coefficients have different scales (posterior sd about 0.46 for the intercept and 0.13 for the slope), and they are strongly correlated (posterior correlation −0.84), because dose was not centred.
Decision
Keep L = 20, ε = 1/20 and the identity mass matrix for all reported HMC results, and record the evidence for and against it in a sweep. The explorer still lets a visitor change L.
Options considered
- Keep the 2023 settings (chosen). The browser then reproduces R's chains draw for draw, and the settings already pass every diagnostic.
- Switch the corrected version to the most efficient identity-mass setting found by the sweep (L = 10, ε = 0.1). This is about 1.2 times as efficient (1.15 to 1.41 times across ten seed sets), but the corrected HMC numbers would no longer match the original function's output for the original settings.
- Precondition with the Laplace covariance, using the GLM's precision matrix as the mass matrix. This is far more efficient, but it is a change to the algorithm as submitted, not a tuning of it.
- Replace the sampler with adaptive HMC or NUTS (dual-averaging step size, adapted mass matrix, as in Stan). This is the right tool for real work, but it is out of scope for a faithful revival.
Why
The site exists to show the 2023 methods honestly, so the reported numbers should come from the 2023 settings unless those settings are wrong. They are not wrong. With four chains of 10,000 iterations, HMC's R-hat is at most 1.0002, its bulk ESS for the slope is 87,223 and its tail ESS 31,864. A full run takes well under a second, so a gain of about 1.2 times is not worth breaking parity. The sweep is published so that the choice is visible rather than hidden.
What happened
The sweep runs four chains (seeds 90125 to 90128) of 5,000 iterations with 1,000 of warm-up for each setting, on the corrected model. Efficiency is the smaller of the bulk and tail ESS of the slope per 1,000 gradient evaluations after warm-up. HMC.fn evaluates the gradient once at the start of each trajectory and once per leapfrog step, so an iteration costs L + 1 gradients.
The four settings this record compares were then rerun on ten independent sets of four seeds (90125 + 4k to 90128 + 4k for k = 0 to 9), and each ratio is taken within a seed set.
| Setting | Acceptance | Efficiency (ESS per 1,000 gradients) | Times the original, median (range over 10 seed sets) |
|---|---|---|---|
| L = 20, ε = 0.05 (original) | 96.4% | 42 | 1 |
| L = 10, ε = 0.1 | 89.8% | 54 | 1.2 (1.15 to 1.41) |
| L = 40, ε = 0.025 | 99.1% | 23 | not rerun |
| ε = 0.2 or larger (identity mass) | below 1% | 0 (chains never move) | not rerun |
| L = 3, ε = 0.5, mass = GLM precision | 97.0% | 199 | 4.8 (4.4 to 5.1) |
| L = 2, ε = 0.8, mass = GLM precision | 91.8% | 299 | 7 (6.4 to 7.7) |
- The original setting spends twice the gradient evaluations of L = 10, ε = 0.1 for a slightly higher acceptance rate. Very high acceptance is a symptom of steps that are smaller than they need to be.
- Step sizes of 0.2 and above fail completely. For a Gaussian target with an identity mass matrix, leapfrog is stable only when ε is below twice the smallest standard deviation along a principal axis of the posterior. Here that limit is 0.14, set by the narrow direction the −0.84 correlation creates, even though the slope's own sd is 0.13 and the intercept's is 0.46.
- Preconditioning with the GLM precision makes the posterior close to isotropic in the sampler's geometry, so much larger steps are stable. L = 2 steps of 0.8 are about 7 times as efficient as the original, and at least 6.4 times in every seed set.
- An earlier draft of this record, before it was merged, counted L gradients per iteration instead of L + 1. That understated the cost of short trajectories most, and it put the preconditioned setting at about 10 times the original. Counting the evaluations the code actually makes gives about 7.
- The weak points of this evidence are worth stating. The main table is one set of four seeds, and only four settings were rerun on other seeds. The ranges come from ten seed sets, which shows the spread but is not a confidence interval. Bulk ESS also exceeds the number of draws for several settings, because HMC's successive draws are negatively correlated, which is why efficiency uses the smaller of bulk and tail ESS.
What I'd change
- Decouple ε from L. The
epsilon = 1/Lconvention hides two choices in one knob. - Centre (or standardise) dose before fitting. That alone removes most of the correlation that limits the step size.
- Use the Laplace covariance, which the analysis computes anyway, as the mass matrix, or adapt the mass matrix during warm-up.
- Tune ε during warm-up towards an acceptance rate of about 0.8 instead of fixing it.
- Report the tuning evidence in the write-up. In 2023 the settings were simply copied from the lecture code.