Skip to content
Bandit Lab, home

Decision record · DR-003 · accepted · 2026-10-06

Explain the DATS anomaly with an ablation, and keep the faithful port

  • Supersedes: the explanation in ADR-002 of the root README (the decision to keep the faithful port stands)

Decision: the lab keeps DATS exactly as submitted. Its shortfall against Thompson sampling is explained by an ablation that corrects one behaviour at a time over 20 seeded repetitions, with every variant clearly labelled and never substituted for the submitted algorithm.

Context

In 2023, Doubly-Adaptive Thompson Sampling (Dimakopoulou, Ren and Zhou, NeurIPS 2021) printed an average reward of 0.24495 on the course log, against 0.4218 for plain Thompson sampling. The paper reports DATS beating TS, so the gap needed an explanation. ADR-002 offered one from a single replay. It stated that DATS dropped arm 2, the best arm, at its 28th round and "then cycled over arms 0 to 4". That last part was imprecise. DATS only narrowed to arms 0 to 4 once its fifth elimination had happened, at internal round 854.

A single run cannot separate the possible causes, so I set out six hypotheses and tested each one.

Code Hypothesis
H0 The gap is seed luck: one unlucky run.
H1 The elimination rule removes the best arm early, and that alone explains the gap.
H2 play() returns a position in the active-arm list as if it were an arm id.
H3 The selection rule picks the arm with the fewest multinomial draws, and the propensities come from the index of each arm's largest Monte-Carlo sample, so the policy cannot exploit.
H4 The elimination statistic multiplies by √max(0, σ²ₐ − σ²ⱼ) where the paper divides by κ√(σ²ₐ + σ²ⱼ).
H5 The doubly-robust scores reuse the current reward for every past round and every arm, and divide the variance by (Σπ)² instead of (Σ√π)².
H6 Even a correct DATS is fragile at the notebook's κ = 0.05 and γ = 0.05. The paper uses κ = 1 and γ = 0.01.

Options considered

  1. Fix DATS silently so the lab shows a better number.
  2. Keep the submitted DATS and only describe the bugs in prose.
  3. Keep the submitted DATS, and add opt-in corrections (fixes in web/src/lib/bandits/dats.ts, all off by default) so an ablation can measure what each bug costs.

Why

Option 1 would rewrite the 2023 result, which this project does not do. Option 2 leaves the explanation as an untested story. Option 3 keeps the submitted behaviour pinned by the parity tests and turns each claim into a paired measurement. Each corrected variant replays the same 20 logs with the same seeds as TS, so every step is a paired comparison.

The set-up is the same as the seeded evaluation in DR-004. It uses course-like click rates, instant feedback as in the 2023 DATS cell and 10,000 matched rounds. Seeds run from 90051 to 90070 and log seeds from 2023 to 2042. Intervals are 95% percentile-bootstrap intervals (B = 2,000, seed 2026), and the elimination counts use Wilson intervals.

What happened

Variant Mean reward [95% CI] Step effect [95% CI] vs TS [95% CI] Best arm eliminated
Thompson sampling (reference) 0.360 [0.357, 0.363]
DATS as submitted 0.249 [0.245, 0.252] −0.111 [−0.116, −0.106] 20/20, median round 14
As submitted, no elimination (H1) 0.222 [0.221, 0.224] −0.026 [−0.031, −0.022] −0.137 [−0.141, −0.134] 0/20
+ arm-id fix (H2) 0.126 [0.110, 0.142] −0.123 [−0.139, −0.107] −0.234 [−0.251, −0.216] 20/20
+ sample one arm from π (H3) 0.128 [0.110, 0.147] +0.002 [−0.017, +0.023] −0.232 [−0.249, −0.214] 17/20
+ win-probability propensities (H3) 0.142 [0.130, 0.155] +0.014 [−0.008, +0.035] −0.218 [−0.230, −0.204] 16/20
+ paper elimination rule (H4) 0.245 [0.177, 0.316] +0.103 [+0.037, +0.175] −0.114 [−0.182, −0.044] 16/20, median round 1
+ paper estimator, all code fixed (H5) 0.367 [0.303, 0.427] +0.122 [+0.010, +0.231] +0.008 [−0.056, +0.068] 12/20, median round 1
+ paper defaults κ = 1, γ = 0.01 (H6) 0.445 [0.406, 0.481] +0.078 [+0.018, +0.148] +0.086 [+0.046, +0.122] 7/20, median round 198

The step effect pairs each row with the row above it, except that the no-elimination row and the arm-id row are both paired with DATS as submitted.

H0 is rejected. DATS trailed TS in all 20 repetitions, by 0.111 on average (dz = −9.0). The 2023 gap was not seed luck.

H1 is only part of the story. The submitted DATS lost the best arm in every repetition, usually within its first 14 rounds. Switching elimination off, however, made DATS worse. It then averaged 0.222, which is what uniform play earns on these click rates, 0.2226. The selection rule does not learn on its own.

H2 shows that the bugs interact. Returning real arm ids made DATS much worse, 0.126 against 0.249. As submitted, play() returns the position index as the arm id, so whenever position 2 was drawn the eliminated best arm, arm 2, was played, about a fifth of the time. Fixed, DATS faithfully plays only the surviving arms, and none of them is the best.

H3 changes little while the elimination bug remains. Sampling from π and using real win probabilities raised the reward by 0.016 between them, and both step intervals include zero.

H4 and H5 are where the reward comes back. The paper's elimination rule and estimator together lifted DATS from 0.142 to 0.367, level with TS. Both steps have wide intervals because the outcome became bimodal. With every code fix in place, DATS kept the best arm in 8 of 20 runs and earned about 0.51 in each of them. In the other 12 it eliminated the best arm in its first or second round, settled on another arm and earned between 0.11 and 0.38.

H6 decides the rest. With every bug fixed and the paper's κ = 1 and γ = 0.01, DATS beat TS by 0.086 [0.046, 0.122] and won 16 of 20 paired repetitions. It still eliminated the best arm in 7 of 20 runs, three of them in the first round and the rest between rounds 198 and 401. For context, UCB1 averaged 0.467 [0.463, 0.469] on the same logs, above every DATS variant.

So the 2023 shortfall has no single cause. Two implementation errors in elimination and estimation do most of the damage, two errors in arm selection hide each other, and the notebook's κ = 0.05 makes even a correct implementation overconfident.

What I'd change

I would write the paper's algorithm as a small reference implementation first and test it on a toy problem with a known answer before tuning anything. On two arms that pay 90% and 10% of the time, the submitted DATS played the better arm in only about half of its 500 rounds in four of five seeds, which a single assertion would have caught in 2023. The corrected DATS played it in every round in four of those five seeds. I would also test the propensities against their definition, the share of Monte-Carlo draws each arm wins, which exposes H3 directly.

On the evidence above, the corrected DATS still loses the best arm in about a third of runs. I would not ship it without a more conservative elimination threshold, or without delaying elimination until each arm has a minimum number of observations.