How the numbers on this site are made
Methods
Where the data comes from, how the algorithms are scored, how uncertainty is measured, what the set-up assumes and where it falls short. Each choice that mattered has a decision record at the end of the page.
Data
Data provenance
The 2023 project came with a log of 300,000 news-recommendation events in which a uniformly random policy chose one of ten stories and recorded whether it was clicked. It is course material, so the site never ships it. Everything on the site runs on synthetic logs generated in the browser from a seed. Each event's arm is drawn uniformly from 0 to 9, and the click is drawn with that arm's rate from the course log, rounded to three decimals. The best arm, arm 2, pays 0.509.
The course log is used in one place only. The parity tests in CI replay it from the notebook's recorded random states and require every 2023 number to match exactly. Why the site uses synthetic data, and what that costs, is in DR-002.
Scoring
Method
Algorithms are scored by offline replay (Li, Chu, Langford and Wang, 2011). The replay walks through the log and asks the bandit for an arm at every event, and it counts a round only when the choice matches the logged arm. Delayed feedback uses the 2023 queue, in which each logged event is revealed some rounds after it is read, or never. The five coursework algorithms are ported line for line, quirks included, and the algorithm notes document every departure from the papers. ε-greedy and UCB1 were added in 2026 as yardsticks. DR-001 records why the replay is unchanged and the weakness the 2026 evaluation exposed.
The optional LLM experiment on /llm uses a different, per-arm replay in which every decision counts, so 200 rounds cost 200 model calls. DR-005 explains that choice.
Uncertainty
Evaluation design
- Seeded repetitions. Every algorithm is replayed 20 times for 10,000 matched rounds. Repetition r uses log seed 2023 + r and policy and delay seed 90051 + r, the same for every algorithm.
- Intervals. Means come with 95% percentile-bootstrap intervals (B = 2,000, seed 2026), curves with pointwise bootstrap bands, and proportions with Wilson score intervals.
- Paired comparisons. Each algorithm is compared with Thompson sampling repetition by repetition, reporting the mean difference with a paired bootstrap interval, Cohen's dz and the number of repetitions it won.
- Sensitivity. Three delay parameters at five levels each. Every cell is its own seeded evaluation with 10 repetitions of 5,000 rounds.
- Ablation. The DATS anomaly is tested by correcting one behaviour at a time against the same seeds.
- Verification. The statistics helpers are tested against numpy, SciPy, statsmodels and R, and CI re-runs a sample of the precomputed repetitions and requires an exact match.
The design is recorded in DR-004, and the results are on /evaluation.
What must hold
Assumptions
- The logging policy chose arms uniformly at random and independently of everything else. Replay is unbiased only under that condition, and the synthetic logs meet it by construction.
- Each arm's click rate is fixed over time and clicks are independent of each other.
- Delays depend only on the arm, and for the reward-dependent model on the reward, never on time or on other events.
- The notebook's parameters are kept, including the Thompson-sampling prior tuned in 2023.
Read with care
Limitations
- Under heavy loss the delayed replay changes which rounds count as well as what a bandit knows. That is why ε-greedy's lead over Thompson sampling shrinks with 90% loss on the best arm.
- The synthetic logs reproduce the course log's ten click rates and nothing else about it, so any drift or correlation in the real data is invisible here.
- Twenty repetitions give tight intervals for means, wide intervals for rare events, and percentile intervals that can undercover when outcomes are bimodal.
- The horizon is 10,000 matched rounds against the notebook's 20,000.
Next time
What I'd change
- Score delayed feedback with a replay where every decision counts and delays only affect learning.
- Run the seeded evaluation on the course log offline and publish only the intervals.
- Choose the number of repetitions from a target interval width, and use BCa intervals.
- Write a minimal reference implementation and a known-answer test before porting an algorithm.
docs/model-card.md
Model card: the evaluation set-up
This card describes the evaluation set-up behind Bandit Lab, meaning the replay of a logged dataset with delayed feedback, the seeded repetitions and the interval estimates, together with the bandit policies it scores. It follows the structure of Mitchell et al. (2019), "Model Cards for Model Reporting". The policies are not trained models in the usual sense, but they make decisions from data and their evaluation has the same reporting needs.
- Owner: Sunchuangyu "Rin" Huang
- Version: 2026-10 (the
upgrade/career-alignedrelease of the revival) - Code:
web/src/lib/evaluation.ts(replay),web/src/lib/eval/(seeded evaluation, sensitivity, ablation),web/src/lib/stats/(intervals),web/src/lib/llm-policy/(LLM experiment) - Decision records: DR-001 to DR-005 in
docs/decisions/
Intended use
- Teaching and portfolio use: showing how bandit algorithms behave under delayed and lost feedback, and how to report their results with uncertainty.
- Reproducing and explaining the 2023 COMP90051 Project 2 results.
- Comparing an LLM with classical policies on a small, fully specified decision problem, using the visitor's own API key.
Out-of-scope uses
- Choosing an algorithm for a production recommender. The click rates, the ten arms and the delay models are coursework settings and are not calibrated to any real service.
- Any claim about news readers or their behaviour. The site never uses reader data.
- Evaluating logs that were not collected by a uniformly random policy (see the assumptions below).
Data
Synthetic logs (everything on the site). Each log has 300,000 events generated from a seed with numpy's generator. The arm is drawn uniformly from 0 to 9 and the click from Bernoulli(rate of that arm). The per-arm rates are the course log's rates rounded to three decimals, from 0.005 for arm 6 up to 0.509 for arm 2, and they average 0.2226. DR-002 records why the course log itself is not used.
Course log (tests only). The 300,000-event course log stays in coursework/ and is read only by the parity
tests in CI. It is never bundled with the site. Only its per-arm click rates and the notebook's printed averages
appear on the site.
Seeds. The seeded evaluation uses log seeds 2023 to 2042 and policy and delay seeds 90051 to 90070. The bootstrap uses seed 2026 with B = 2,000 resamples.
Evaluation design
- Replay. The coursework's delayed replay (Li et al., 2011, with a delay queue), unchanged (DR-001).
- Repetitions. R = 20 repetitions of 10,000 matched rounds per algorithm and scenario. All algorithms see the same logs and seeds within a repetition, so comparisons are paired (DR-004).
- Scenarios. Instant feedback, heavy-tailed Pareto delays, 90% loss on the best arm, and the 2023 set-up in which each algorithm faces its own coursework delay.
- Sensitivity. Pareto shape, loss probability and reward-dependent lag at five levels each, with 10 repetitions of 5,000 rounds per cell.
- Metrics. Mean reward over matched rounds, cumulative pseudo-regret Σ(μ* − μ of the arm played), the share of rounds on the best arm, and whether an elimination algorithm removed the best arm.
- Uncertainty. 95% percentile-bootstrap intervals, paired differences against Thompson sampling with Cohen's dz, and Wilson intervals for proportions. The helpers are checked against numpy, SciPy, statsmodels and R.
Results (instant feedback, 20 repetitions, 10,000 rounds)
| Policy | Mean reward [95% CI] | Δ vs TS [95% CI] | Beats TS |
|---|---|---|---|
| ε-greedy (ε = 0.1) | 0.473 [0.466, 0.479] | +0.114 [+0.107, +0.120] | 20/20 |
| UCB1 (c = 2) | 0.467 [0.463, 0.469] | +0.107 [+0.104, +0.110] | 20/20 |
| Thompson sampling (2023) | 0.360 [0.357, 0.363] | reference | |
| Successive Elimination | 0.302 [0.297, 0.308] | −0.057 [−0.063, −0.052] | 0/20 |
| DATS as submitted | 0.249 [0.245, 0.252] | −0.111 [−0.116, −0.106] | 0/20 |
| Optimistic-Pessimistic SE | 0.224 [0.222, 0.225] | −0.136 [−0.140, −0.132] | 0/20 |
| Phased SE | 0.223 [0.220, 0.228] | −0.136 [−0.141, −0.132] | 0/20 |
The best arm pays 0.509 and uniform play earns 0.2226. The full tables for every scenario, the sensitivity
heatmaps and the DATS ablation are on /evaluation.
Assumptions
- Uniform logging policy. Replay is unbiased only when the logged arm was chosen uniformly at random, independently of everything the bandit could know. The synthetic logs satisfy this by construction. A log collected by any other policy needs inverse-propensity or doubly-robust weighting, which this set-up does not implement.
- Stationary, independent rewards. Each arm's click rate is fixed and clicks are independent. Drift, seasonality and correlated readers are not modelled.
- Delay models. Pareto delays are drawn as round(Pareto(α)) per event, packet loss drops an event's feedback with a fixed probability per arm, and reward-dependent delays look up a fixed lag by arm and reward. Delays are independent of everything else.
- What counts as a round. A round counts only when a revealed event matches the bandit's choice, so delays change which rounds are counted as well as what the bandit knows (DR-001).
Known failure modes
- Replay selection under heavy loss. With 90% loss on the best arm, choices of that arm rarely produce a matched round, so matched rounds over-represent exploratory pulls. Epsilon-greedy's share of matched rounds on the best arm fell from 85% to 36% under this setting, and its measured lead over Thompson sampling fell to +0.004 [−0.001, +0.009].
- Stalled runs. A bandit that insists on an arm whose feedback never arrives cannot finish. The lab stops such runs after max(2,000,000, 100 × rounds) log steps and reports them. None of the 1,050 sensitivity runs stalled.
- Faithful 2023 quirks. SE, PSE, OPSE, TS and DATS keep the behaviour of the submitted code, documented in the
port notes on
/algorithms. OPSE never eliminates, and the notebook's TS prior (σ = 4) makes it slow to commit. - DATS. As submitted it eliminated the best arm in 20 of 20 repetitions (DR-003).
- Narrow intervals. Twenty repetitions give tight intervals for means but wide Wilson intervals for rare events, and percentile intervals can undercover for bimodal outcomes.
Ethical considerations
- Course material. The specification, skeleton, delay code and log belong to the University of Melbourne. They are kept in the repository for reference and tests and are not served by the site.
- Academic integrity. The coursework is shared as a record of completed work. The site asks students not to reuse it in their own assessment.
- No personal data. The site has no server-side storage, analytics or accounts. Simulations run in the visitor's browser.
- Optional AI features. The LLM experiment is opt-in, uses the visitor's own key, labels every output as
AI-generated and records each call in an audit log on the visitor's device. See the AI use statement on
/methods.
Caveats and recommendations
Treat the numbers as properties of this set-up, not of the algorithms in general. Before relying on a ranking for
a different problem, re-run the seeded evaluation with that problem's click rates and delays, which the lab and
the re-run panel on /evaluation support.
docs/ai-use-statement.md
AI use statement
Bandit Lab works fully without AI. One optional experiment, "LLM as a bandit policy" on /llm, lets a visitor ask
a large language model to make the decisions in a short bandit episode, using the visitor's own API key. This
statement explains what that feature does and does not do. It is informed by the Australian Government's Policy
for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU
AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.
What the AI does
- At each round of an episode, 200 rounds by default and adjustable from 20 to 500, the model receives a short table of what has been observed so far. It returns the arm it wants to play, with a reason of at most 12 words.
- The site compares the resulting reward and regret with Beta(1, 1) Thompson sampling, UCB1 and the 2023 Thompson sampling on the same seeds, with Wilson and bootstrap intervals.
What the AI never does
- It never runs unless the visitor adds their own key and starts a run.
- It never changes the 2023 results, the coursework algorithms, the precomputed evaluation or anything another visitor sees.
- It never receives the API key, personal data or the course dataset. Its prompts contain only counts from the synthetic episode.
- It never decides on its own whether its results are kept. A run enters the comparison table only after the visitor reviews it and accepts it, and only a run that played every round can be accepted.
Data sent to the provider
Each call sends a fixed instruction, about 700 characters, and one table of ten rows with the round number, pulls,
reported feedback, clicks, click rate and pending feedback per arm, plus the last five choices. Lost feedback is
shown as pending, like late feedback, so the model learns nothing the classical policies do not. Requests go
directly from the visitor's browser to api.anthropic.com or api.openai.com under the visitor's own account and
that provider's terms. This site has no server, so nothing passes through it.
Keys
The key is stored in the browser's sessionStorage by default and disappears when the tab closes. It moves to
localStorage only if the visitor ticks "remember on this device", and unticking the box moves it back. Only one
key is kept at a time. "Forget key" removes it from both storages, whichever provider it belongs to, and the
settings say when a key for the other provider is stored. It is never logged, never written to the audit log and
never committed to the repository. Every text field in the audit
log is checked for anything shaped like a key before it is stored.
Transparency and audit
- Every model output on the page is labelled "AI-generated".
- Every call is appended to an audit log in the visitor's browser (IndexedDB). Each entry holds an id, a timestamp, the feature, provider and model, the full prompt, the raw output, whether it passed validation, the latency, the token usage the provider reported and the human decision on it, which is pending, accepted, edited or rejected.
- The log is viewable at
/ai-logand exportable as JSON or CSV. Its valid-answer rate counts only calls that returned an answer; calls that failed before answering are counted separately. - The visitor can clear the log at any time, together with the reviewed runs on
/llmunless they choose to keep them, and can remove a single reviewed run from the/llmtable.
Human in the loop
The visitor starts every run, can stop it at any round, and decides afterwards whether to accept or reject it. Rejected runs stay in the audit log, marked as rejected, and are left out of the comparison. Invalid model outputs are counted and shown, never silently repaired. The fallback arm played in their place is chosen by a seeded random generator, and the audit entry records that it was a fallback.
Limits
A model's choices depend on its version and on the provider's sampling, so two runs with the same seed can differ. The page reports each run as one sample. It pairs runs per model, counts each repetition once (the latest accepted run), and shows a paired interval only once five distinct repetitions are accepted; below that it lists the per-seed differences.
See every call your browser has made on the AI audit log.
docs/decisions
Decision records
Each record states the decision first. It then gives the context, the options and the reasons, what happened, including the numbers that did not go my way, and what I would change. Past records are never edited, only superseded. The four earlier records from the 2026 revival stay in the repository README.
- DR-001 · accepted · 2026-10-06Evaluate with the 2023 delayed replay, unchangedEvery bandit in the lab, on /evaluation and in the 2023 reproduction is scored with the delayed replay exactly as the notebook implemented it. The only addition is an opt-in step limit that reports a run which can never finish as "stalled". The LLM experiment uses a different, per-arm replay, recorded in DR-005.Read DR-001
- DR-002 · accepted · 2026-10-06Run the site on a synthetic log, never the course logThe website generates its logged data in the browser from a published, seeded recipe that matches the course log's shape and per-arm click rates. The course log stays in coursework/ for the parity tests and is never bundled, served or summarised beyond per-arm aggregates.Read DR-002
- DR-003 · accepted · 2026-10-06Explain the DATS anomaly with an ablation, and keep the faithful portThe lab keeps DATS exactly as submitted. Its shortfall against Thompson sampling is explained by an ablation that corrects one behaviour at a time over 20 seeded repetitions, with every variant clearly labelled and never substituted for the submitted algorithm.Read DR-003
- DR-004 · accepted · 2026-10-06Report every result over 20 seeded repetitions with bootstrap intervalsEvery comparison the site makes is based on R = 20 seeded repetitions of the replay. Repetition r uses log seed 2023 + r and policy and delay seed 90051 + r for every algorithm. Means are reported with 95% percentile-bootstrap intervals (B = 2,000, resampling seed 2026), comparisons with Thompson sampling are paired by repetition, proportions use Wilson intervals, and effect sizes are reported as Cohen's dz. The seeds are printed next to the results.Read DR-004
- DR-005 · accepted · 2026-10-06Evaluate an LLM as a bandit policy, bring-your-own-key, on a per-arm replayThe site offers an optional experiment in which a large language model chooses the arm at every round. It runs only with the visitor's own Anthropic or OpenAI key, called directly from their browser, on a 200-round per-arm replay of the synthetic log. Its reward and regret are paired by seed with Beta(1, 1) Thompson sampling and UCB1, the headline comparators, and with the 2023 Thompson sampling for continuity. Every call is recorded in an audit log that stays on the visitor's device.Read DR-005