This card covers the statistical model at the centre of the site and the four ways it is fitted (Laplace approximation, Metropolis-Hastings, Hamiltonian Monte Carlo and expectation propagation). It is a small Bayesian model on a teaching dataset, not a trained system deployed on real decisions. The card follows the usual model-card headings so that its limits are written down in one place.
Model details
- Author: Sunchuangyu (Rin) Huang. Written for MAST90125 Bayesian Statistical Learning, University of Melbourne, Semester 2 2023, and revived in 2026. The Metropolis, HMC and EP functions were adapted from the MAST90125 lecture code, as credited in the original submission.
- Model: with for doses , and a flat prior , as the assignment specified.
- Inference: a Laplace approximation at the posterior mode (a standard logistic GLM), random-walk Metropolis-Hastings, HMC with 20 leapfrog steps of 1/20 and an identity mass matrix, and expectation propagation with Gaussian sites. The two-parameter posterior is also computed exactly on a grid, as a reference.
- Implementation: the original R functions, ported line by line to TypeScript in
web/src/lib/stats/and run in the browser. With the original seed, the ports reproduce R's draws to about . - Versions: "Corrected" (the default and the subject of this card) and "As submitted (2023)", which keeps the original data-encoding bugs for transparency (DR-001).
Intended use
- Teaching and portfolio use. The site shows how four approximate inference methods compare on a posterior small enough to compute exactly, and what a careful Bayesian workflow adds around them: multi-chain convergence diagnostics, posterior predictive checks, prior sensitivity and method comparison.
- Reproducing the 2023 coursework faithfully, and documenting what was wrong with it.
Out of scope
- Any clinical, dosing or treatment decision. Nothing here is medical advice.
- Inference about the real-world effect of a real medication. The counts come from a coursework exercise, and the study design behind them (randomisation, blinding, how improvement was judged) is not described.
- Doses outside 0 to 6. The linear-in-dose logit has no support from the data beyond that range.
Data provenance
- The data are the table given in the assignment: seven doses (0 to 6), ten patients per dose, and the number whose symptoms had improved after a week, , or 37 of 70 overall.
- The table contains counts only. There is no personal information, and nothing was collected by me.
- The assignment specification itself is not redistributed. The task is paraphrased in the README and on the site.
Evaluation
All results below are for the corrected model. Monte Carlo results use four chains per sampler with set.seed(90125) to set.seed(90128) and the original settings, and every estimate is reported with its Monte Carlo standard error (MCSE). The convergence diagnostics (rank-normalised split R-hat, bulk and tail ESS, MCSE) are ports of R's posterior package, and the test suite checks them against posterior 1.7.0 on the same draws.
Convergence. Every coefficient passes the thresholds of Vehtari et al. (2021), R-hat at most 1.01 and bulk and tail ESS at least 400. The largest R-hat is 1.0002. The slope's bulk ESS is about 26,500 for Metropolis-Hastings (from 200,000 draws) and 87,200 for HMC (from 36,000 draws, larger than the number of draws because successive HMC draws are negatively correlated).
Estimates of the dose slope (log-odds per unit dose):
| Method | Mean | 95% interval | KL from the exact posterior |
|---|---|---|---|
| Exact posterior (quadrature) | 0.3143 | [0.063, 0.582] | 0 |
| HMC, 4 chains | 0.3149 ± 0.0005 (MCSE) | [0.062, 0.582] | Monte Carlo error only |
| Metropolis-Hastings, 4 chains | 0.3154 ± 0.0008 (MCSE) | [0.065, 0.585] | Monte Carlo error only |
| Expectation propagation | 0.3142 | [0.056, 0.573] | 0.0027 |
| Laplace (GLM) | 0.3016 | [0.048, 0.555] | 0.0085 |
The posterior probability that the slope is positive is 0.993, and the posterior median odds ratio per unit dose is 1.37. The Laplace mean sits 0.10 posterior sd below the HMC mean, and every other method is within 0.006 sd of it. The best Gaussian possible (the exact mean and covariance) has a KL divergence of 0.0027 from the exact posterior, so EP is as good as a Gaussian can be and Laplace is about three times worse (DR-003).
Posterior predictive checks. For each of the 36,000 HMC draws a replicated trial is simulated. The posterior predictive p-values for the total improved, the number of dose-to-dose drops, the counts at doses 0 and 6, and the chi-square discrepancy are 0.50, 0.39, 0.51, 0.38 and 0.74, with MCSEs between 0.0018 and 0.0025. All 14 observed cells fall inside their 90% predictive intervals. These checks find no misfit, but with 70 patients they have little power to find any.
Prior sensitivity. Under independent priors on both coefficients, computed exactly, a weakly informative moves the slope by 0.06 posterior sd and by 0.29 sd. Only a tight changes the conclusion materially. It shrinks the slope to 0.148 and its 95% interval then includes zero. stays between 0.966 and 0.993 across all the priors.
Known failure modes
- Mis-encoded data. The 2023 submission fed each method a different broken version of the data. Four-chain R-hat catches the Metropolis bug (R-hat up to 2.76), but the HMC and EP bug passes every convergence check and is caught only by the posterior predictive check (p = 0 for the total improved). See DR-001.
- Skewness. Both Gaussian approximations miss the right skew of the slope's posterior, so their interval endpoints sit 0.007 to 0.009 (EP) and 0.014 to 0.027 (Laplace) below the exact ones.
- HMC step size. With the identity mass matrix, step sizes of 0.2 or more almost never accept a proposal, because the intercept and slope are strongly correlated (posterior correlation −0.84). See DR-002.
- Proposal scale. Random-walk Metropolis depends on its proposal covariance. With the wrong one, as in 2023, acceptance fell to 5.0%.
- Small data. Seven binomial counts cannot distinguish the linear-in-dose logit from many curved alternatives, and the slope remains uncertain (95% interval about 0.06 to 0.58).
- The flat prior. A flat prior on the logit scale is not "uninformative" about probabilities. It puts most of its mass near 0 and 1. It is used because the assignment specified it, and the prior-sensitivity analysis shows the conclusions do not hinge on it.
Ethical considerations
- The dataset has no personal information. The health setting (an influenza medication) makes over-interpretation the main risk, so the site states plainly that the model is a teaching exercise and that its results say nothing about any real treatment.
- The site keeps the original mistakes visible rather than presenting only the corrected analysis, so that the record of the coursework stays honest.
- The optional AI feature ("Explain these diagnostics") is bring-your-own-key, sees only a rounded numeric summary, labels its output as AI-generated, runs automatic checks on it and records every call and the human decision in a local audit log. Its design is informed by the transparency principles of the Australian Government's policy for the responsible use of AI in government, the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them. See the AI use statement on the methods page.