Skip to content
Bandit Lab, home

Decision record · DR-004 · accepted · 2026-10-06

Report every result over 20 seeded repetitions with bootstrap intervals

Decision: every comparison the site makes is based on R = 20 seeded repetitions of the replay. Repetition r uses log seed 2023 + r and policy and delay seed 90051 + r for every algorithm. Means are reported with 95% percentile-bootstrap intervals (B = 2,000, resampling seed 2026), comparisons with Thompson sampling are paired by repetition, proportions use Wilson intervals, and effect sizes are reported as Cohen's dz. The seeds are printed next to the results.

Context

The 2023 notebook printed one average reward per algorithm from one run, plus a 10-repeat curve that shared one generator. One run cannot say whether a difference of 0.05 is a property of the algorithm or of the seed, and the 10-repeat helper had a bug that handed Thompson sampling a prior mean of 20,000. A statistician reading the original results could not tell which differences were real.

Options considered

  1. Keep single runs and add the standard deviation across the notebook's repeats.
  2. Use t-intervals for means and two-sample tests for comparisons.
  3. Seeded repetitions with percentile-bootstrap intervals and paired comparisons.
  4. The same with BCa intervals, or a hierarchical bootstrap that resamples logs and seeds separately.

Why

Option 3 makes the fewest assumptions about the shape of the reward distribution, and the per-repetition values are often skewed or bimodal, as the DATS ablation shows. Pairing removes the variation the algorithms share, since in each repetition they all see the same log and the same seed. That made most comparisons decisive with 20 repetitions. The bootstrap uses the site's bit-exact port of numpy's generator, so scripts/export_stats_reference.py can check the TypeScript intervals against plain numpy exactly, and the Wilson intervals against statsmodels and R's prop.test.

I report Cohen's dz and a win rate beside each difference because a p-value from 20 pairs would mostly say "different", which tells a reader nothing about how much.

Twenty repetitions is a practical limit. DATS replays at about 0.8 seconds per repetition at 10,000 rounds. The whole precomputed set takes about five minutes and covers four scenarios, a 1,050-run sensitivity grid and the ablation.

What happened

The intervals are narrow. Thompson sampling's mean reward with instant feedback is 0.360 [0.357, 0.363], and most paired differences have |dz| above 4. Thompson sampling stays ahead of the other coursework algorithms on every seed. With instant feedback it beat SE, PSE, OPSE and DATS in all 20 repetitions. The seeded evaluation also showed something the single runs never could. UCB1 and epsilon-greedy, which I added as baselines in 2026, beat TS in 20 of 20 repetitions, by 0.107 and 0.114.

The weak point is the interval width for proportions. With 20 repetitions a Wilson interval for "eliminated the best arm in 7 of 20 runs" spans 0.18 to 0.57, which is honest but not precise. Percentile intervals from 20 values can also undercover when the distribution is bimodal, which is the case for several DATS variants.

What I'd change

I would choose R from a target interval width instead of fixing it at 20, and use BCa intervals for the bimodal DATS variants. I would also resample logs and seeds separately, so the intervals show how much of the spread comes from the data and how much from the algorithm's own randomness.