- Extends: ADR-003 in the root README
Decision: the website generates its logged data in the browser from a published, seeded recipe that matches
the course log's shape and per-arm click rates. The course log stays in coursework/ for the parity tests and is
never bundled, served or summarised beyond per-arm aggregates.
Context
The coursework came with a 300,000-event log of news-recommendation clicks. It is course material supplied by the COMP90051 teaching team, so it cannot be redistributed. The lab, the seeded evaluation and the LLM experiment all need a log, and every visitor must be able to regenerate exactly the same one.
Options considered
- Ship the course log, or a subsample of it, with the site.
- Use a public benchmark such as the Yahoo! Front Page Today Module data.
- Simulate a log of the same shape: arm ~ Uniform{0..9}, click ~ Bernoulli(rate of that arm), with numpy's generator and a fixed seed, using the course log's per-arm click rates rounded to three decimals.
Why
Option 1 breaks the course's terms. Option 2 has its own licence conditions, a different number of arms and a
time-varying click rate, so none of the 2023 settings would carry over. Option 3 keeps the exact structure the
algorithms were written for, uses only aggregate statistics from the course log, and is reproducible from a seed.
The recipe is implemented twice, in web/src/lib/dataset.ts and in scripts/export_synthetic_parity.py, and the
tests require both to produce the same events.
What happened
The parity suite replays the course log in CI from the notebook's recorded random states and matches every published number exactly, so the port is verified on the real data even though the site never shows it.
The synthetic log reproduces the per-arm click rates but nothing else about the course log. If the real log has drift over time, correlation between consecutive events or any structure beyond its ten averages, the synthetic results cannot show it. The seeded evaluation also runs at 10,000 matched rounds where the notebook ran 20,000, so its numbers are not directly comparable with the 2023 averages. Thompson sampling learns slowly under the notebook's prior, which is why it averages 0.360 [0.357, 0.363] here against 0.4218 in the single 2023 run.
What I'd change
I would run the seeded evaluation on the course log itself, offline, and publish only the resulting intervals. That would measure the gap between the synthetic and the real log directly instead of assuming it is small. I would also match the horizon to the notebook's 20,000 rounds for the headline table, at the cost of a longer export.