Skip to content
Bandit Lab, home

20 seeded repetitions · bootstrap intervals · paired comparisons

Every result with its uncertainty

The 2023 notebook printed one number per algorithm from one run. Here every algorithm is replayed on 20 seeded repetitions of the same experiment, so each result comes with an interval, each comparison is paired on the same logs and seeds, and the seeds are on the page.

What the intervals say

Robust

TS beats every other coursework algorithm on every seed

With instant feedback TS averaged 0.360 [0.357, 0.363]. SE, PSE, OPSE, DATS lost to it in all 20 paired repetitions, and DATS trailed it by −0.111 [−0.116, −0.106]. The 2023 ranking was not seed luck.

Humbling

Two textbook baselines beat the 2023 winner

UCB1 beat TS by +0.107 [+0.104, +0.110] in 20 of 20 repetitions, and ε-greedy by +0.114 [+0.107, +0.120] in 20. The notebook's TS assumes reward noise σ = 4 for 0/1 clicks, so its posterior needs hundreds of pulls per arm before it commits.

A limit of the replay

Lost feedback changes which rounds count

With 90% of the best arm's feedback lost, ε-greedy's lead over TS shrinks to +0.004 [−0.001, +0.009]. A choice of arm 2 now matches the log about one step in a hundred, so its exploratory pulls fill the matched rounds. Its share of rounds on the best arm drops from 85% to 36%. Part of that drop is an artefact of the replay, which DR-001 records.

Seeded evaluation

Reward and regret with 95% bootstrap bands

Course-like click rates, 10,000 matched rounds, the notebook's parameters. Switch the delay set-up to see how each algorithm copes.

Every click is revealed at once: the classic bandit setting and the 2023 TS and DATS cells.

Mean cumulative average reward over the repetitions. The shaded band is a pointwise 95% bootstrap interval for that mean, and higher is better.

  • SE
  • PSE
  • OPSE
  • TS
  • DATS
  • ε-greedy
  • UCB1
Data table
Mean cumulative average reward with 95% bootstrap bands, Instant feedback, precomputed
Matched roundSEPSEOPSETSDATSε-greedyUCB1
2000.227 [0.213, 0.241]0.227 [0.213, 0.241]0.233 [0.223, 0.243]0.228 [0.215, 0.240]0.228 [0.214, 0.241]0.389 [0.352, 0.422]0.287 [0.268, 0.307]
1,6000.223 [0.218, 0.227]0.223 [0.218, 0.227]0.225 [0.220, 0.231]0.249 [0.244, 0.253]0.246 [0.242, 0.251]0.453 [0.433, 0.469]0.380 [0.372, 0.388]
3,0000.223 [0.220, 0.227]0.223 [0.220, 0.226]0.223 [0.219, 0.227]0.268 [0.265, 0.271]0.247 [0.243, 0.251]0.462 [0.444, 0.476]0.417 [0.412, 0.423]
4,4000.238 [0.234, 0.242]0.223 [0.220, 0.226]0.223 [0.220, 0.226]0.286 [0.284, 0.289]0.248 [0.245, 0.252]0.465 [0.449, 0.478]0.437 [0.432, 0.441]
5,8000.257 [0.252, 0.262]0.223 [0.220, 0.226]0.224 [0.221, 0.226]0.306 [0.303, 0.309]0.249 [0.245, 0.252]0.468 [0.456, 0.478]0.447 [0.444, 0.450]
7,2000.272 [0.268, 0.277]0.222 [0.219, 0.226]0.224 [0.221, 0.226]0.326 [0.323, 0.329]0.248 [0.245, 0.251]0.471 [0.461, 0.479]0.456 [0.452, 0.459]
8,6000.287 [0.282, 0.292]0.222 [0.219, 0.226]0.224 [0.222, 0.226]0.344 [0.341, 0.347]0.248 [0.245, 0.251]0.473 [0.464, 0.479]0.462 [0.459, 0.465]
10,0000.302 [0.297, 0.308]0.223 [0.220, 0.228]0.224 [0.222, 0.225]0.360 [0.357, 0.363]0.249 [0.245, 0.252]0.473 [0.466, 0.479]0.467 [0.463, 0.469]
Mean reward with 95% bootstrap interval, paired difference against Thompson sampling, effect size, win rate with Wilson interval, final regret and best-arm eliminations, Instant feedback, precomputed
AlgorithmMean reward [95% CI]Δ vs TS [95% CI]dzBeats TSRegret at 10kBest arm lost
ε-greedy0.473 [0.466, 0.479]+0.114 [+0.107, +0.120]+7.220/20 [0.84, 1.00]364 [309, 433]—
UCB10.467 [0.463, 0.469]+0.107 [+0.104, +0.110]+16.220/20 [0.84, 1.00]430 [418, 442]—
TS0.360 [0.357, 0.363]reference——1,507 [1,495, 1,521]—
SE0.302 [0.297, 0.308]−0.057 [−0.063, −0.052]−4.60/20 [0.00, 0.16]2,066 [2,019, 2,113]0/20
DATS0.249 [0.245, 0.252]−0.111 [−0.116, −0.106]−9.00/20 [0.00, 0.16]2,594 [2,561, 2,625]20/20
OPSE0.224 [0.222, 0.225]−0.136 [−0.140, −0.132]−13.40/20 [0.00, 0.16]2,861 [2,856, 2,866]0/20
PSE0.223 [0.220, 0.228]−0.136 [−0.141, −0.132]−13.10/20 [0.00, 0.16]2,841 [2,806, 2,864]0/20

20 repetitions × 10k matched rounds. Repetition r replays log seed 2023 + r with policy and delay seed 90051 + r (r = 0…19), so every algorithm sees the same logs and seeds and the comparison with TS is paired. The intervals are percentile-bootstrap intervals with B = 2,000 and seed 2026. Win rates use Wilson intervals, and dz = mean paired difference ÷ its standard deviation. Green or red marks a difference whose 95% interval excludes zero.

Check it yourself

Re-run this set-up in your browser

The defaults below use fresh seeds, so the run is an independent check of the table above. The replay runs in a Web Worker on this device, and nothing is sent anywhere. Set the seeds to 90051 and 2023 with 20 repetitions and 10,000 rounds to get the precomputed table back, digit for digit.

Repetitions R10
Horizon (matched rounds)5,000
estimate ≈ 4.2 s

Sensitivity

How hard do the delays have to bite?

Each column turns one delay parameter up, from benign to the 2023 setting and beyond. Every cell is its own seeded evaluation (10 repetitions, 5,000 rounds), so the intervals say which shifts are real.

Pareto shape α of arms 2 and 5 (the other arms keep α = 1); smaller α means heavier tails.

Mean reward by algorithm and pareto α on arms 2 and 5, 10 repetitions each, with 95% bootstrap intervals
Algorithmα = 2α = 1α = 0.5α = 0.35α = 0.22023
SE0.257[0.248, 0.266]0.248[0.241, 0.254]0.248[0.243, 0.254]0.249[0.243, 0.254]0.251[0.245, 0.258]
PSE0.229[0.224, 0.234]0.224[0.220, 0.227]0.224[0.221, 0.227]0.224[0.221, 0.229]0.226[0.222, 0.230]
OPSE0.220[0.217, 0.224]0.220[0.217, 0.223]0.222[0.219, 0.226]0.220[0.217, 0.223]0.219[0.215, 0.222]
TS0.294[0.289, 0.298]0.290[0.285, 0.295]0.292[0.286, 0.296]0.293[0.290, 0.296]0.290[0.285, 0.294]
DATS0.246[0.242, 0.250]0.251[0.238, 0.265]0.254[0.244, 0.265]0.244[0.237, 0.251]0.238[0.233, 0.244]
ε-greedy0.458[0.432, 0.476]0.472[0.462, 0.480]0.471[0.458, 0.482]0.473[0.463, 0.481]0.463[0.451, 0.474]
UCB10.438[0.431, 0.445]0.442[0.437, 0.448]0.441[0.435, 0.447]0.442[0.435, 0.448]0.442[0.436, 0.449]

10 repetitions per cell × 5,000 matched rounds, seeds 90051+r and log seeds 2023+r, the same at every level, so columns are comparable. Brackets: 95% percentile-bootstrap intervals (B = 2,000, seed 2026). Darker cells are closer to the best arm's 0.509. The palest sit at uniform play's 0.223.

Decision record DR-003

Why DATS trails TS: fixing one thing at a time

As submitted, DATS eliminated the best arm in 20 of 20 seeds, by internal round 14 in the median run. Switching elimination off does not rescue it. DATS then earns 0.222, the level of uniform play, because its selection rule does not learn. Only with every behaviour corrected and the paper's defaults does DATS beat TS, by +0.086 [+0.046, +0.122], and even then it loses the best arm in 7 of 20 seeds.

0.000.200.40Mean reward over 10,000 matched roundsbest arm 0.509uniform play 0.223Thompson samplingDATS as submittedAs submitted, no elimination+ arm-id fix+ sample from π+ win-probability propensities+ paper elimination rule+ paper estimator (all fixes)+ paper defaults κ = 1, γ = 0.01
Mean reward of each DATS variant and of Thompson sampling, with 95% bootstrap intervals over 20 seeds
Thompson sampling0.360 [0.357, 0.363]
DATS as submitted0.249 [0.245, 0.252]
As submitted, no elimination0.222 [0.221, 0.224]
+ arm-id fix0.126 [0.110, 0.142]
+ sample from π0.128 [0.110, 0.147]
+ win-probability propensities0.142 [0.130, 0.155]
+ paper elimination rule0.245 [0.177, 0.316]
+ paper estimator (all fixes)0.367 [0.303, 0.427]
+ paper defaults κ = 1, γ = 0.010.445 [0.406, 0.481]
DATS ablation. Each row corrects one behaviour of the submitted code and reports the mean reward, the step's paired effect, the paired difference against Thompson sampling and how often the best arm was eliminated
VariantTestsMean reward [95% CI]Step effect [95% CI]vs TS [95% CI]Best arm eliminated
Thompson samplingReference: the notebook's Gaussian TS.—0.360 [0.357, 0.363]—reference—
DATS as submittedThe 2023 code, unchanged.—0.249 [0.245, 0.252]—−0.111 [−0.116, −0.106]20/20 [0.84, 1.00]median round 14
As submitted, no eliminationSwitch arm elimination off, nothing else.H10.222 [0.221, 0.224]−0.026 [−0.031, −0.022]−0.137 [−0.141, −0.134]0/20 [0.00, 0.16]
+ arm-id fixReturn the chosen arm, not its position in the active list.H20.126 [0.110, 0.142]−0.123 [−0.139, −0.107]−0.234 [−0.251, −0.216]20/20 [0.84, 1.00]median round 15
+ sample from πPlay one draw from Multinomial(1, π) instead of the arm with the fewest draws.H30.128 [0.110, 0.147]+0.002 [−0.017, +0.023]−0.232 [−0.249, −0.214]17/20 [0.64, 0.95]median round 11
+ win-probability propensitiesπ_a = P(arm a's posterior sample is the largest), not the index of its largest sample.H30.142 [0.130, 0.155]+0.014 [−0.008, +0.035]−0.218 [−0.230, −0.204]16/20 [0.58, 0.92]median round 14
+ paper elimination ruleEliminate any arm with min_j Φ((μ_a − μ_j)/κ√(σ²_a + σ²_j)) < 1/T.H40.245 [0.177, 0.316]+0.103 [+0.037, +0.175]−0.114 [−0.182, −0.044]16/20 [0.58, 0.92]median round 1
+ paper estimator (all fixes)Doubly-robust scores per round with r_s and 1{a_s = a}; variance over (Σ√π)².H50.367 [0.303, 0.427]+0.122 [+0.010, +0.231]+0.008 [−0.056, +0.068]12/20 [0.39, 0.78]median round 1
+ paper defaults κ = 1, γ = 0.01Keep every fix and use the paper's default sampling scale and exploration floor.H60.445 [0.406, 0.481]+0.078 [+0.018, +0.148]+0.086 [+0.046, +0.122]7/20 [0.18, 0.57]median round 198

20 seeds (90051–90070, log seeds 2023–2042), 10,000 matched rounds, instant feedback, course-like click rates, the notebook's κ = 0.05, γ = 0.05 and 800 Monte-Carlo draws unless stated. “Step effect” pairs each row with the row it changes one thing from. Thompson sampling averaged 0.360. Intervals: percentile bootstrap (B = 2,000, seed 2026); eliminations: Wilson.

Read the decision record: hypotheses, evidence and what I'd change

Read with care

What these intervals do and do not cover

  • They cover the randomness of the simulated log, the algorithms and the delays across 20 seeds. They say nothing about a different click-rate profile: all scenarios use the course log's per-arm rates.
  • Bootstrap intervals from 20 repetitions are approximate. Intervals for pairs that differ by a lot (|dz| above 5) are tighter than any real-world deployment would justify.
  • 2023 set-up: each algorithm faced a different delay, exactly as in the notebook, so its comparisons are not like for like.
  • The algorithms keep their 2023 quirks (see the port notes); only the DATS ablation corrects them, and only as labelled variants.

Data provenance, assumptions and the full evaluation design are on the methods page; the model card covers the evaluation set-up itself.