Skip to content
Bandit Lab, home

Course log · seed 90051 · 20,000 matched rounds

The 2023 results, reproduced

The original notebook printed one average reward per algorithm and plotted 10-repeat curves. Re-running it unchanged and replaying the TypeScript port from the exact same random-number state gives the same numbers to the last digit. Stepping through those runs also shows why they came out the way they did.

Evaluation (a): one run each

  • SESuccessive Elimination

    Pareto delays (arms 2 & 5 heavy-tailed)

    Printed in 2023
    0.38145
    TypeScript port
    0.38145
    vs best arm
    75%
    10-repeat final
    0.3763
  • PSEPhased Successive Elimination

    Packet loss (90% on arm 2)

    Printed in 2023
    0.30405
    TypeScript port
    0.30405
    vs best arm
    60%
    10-repeat final
    0.2843
  • OPSEOptimistic-Pessimistic SE

    Reward-dependent (1,000-round lag)

    Printed in 2023
    0.2184
    TypeScript port
    0.2184
    vs best arm
    43%
    10-repeat final
    0.2144
  • TSThompson Sampling

    No delay

    Printed in 2023
    0.4218
    TypeScript port
    0.4218
    vs best arm
    83%
    10-repeat final
    0.2203
  • DATSDoubly-Adaptive Thompson Sampling

    No delay

    Printed in 2023
    0.24495
    TypeScript port
    0.24495
    vs best arm
    48%
    10-repeat final
    0.2526

Always playing the best arm (arm 2) would average 0.5092; the log as a whole averages 0.2229. The recorded Python re-run of the whole notebook took 17 min 33 s; the TypeScript port replays all five single runs in a few seconds.

Evaluation (b)

Cumulative average reward over 10 repeats

Each curve averages ten fresh bandits that share one generator, exactly as the notebook's helper did. The dashed line is the best arm; the log as a whole averages 0.223. The TS curve is not the TS above: the helper passed the horizon into TS's prior-mean argument (see below).

  • SE
  • PSE
  • OPSE
  • TS
  • DATS
Data table
Cumulative average reward averaged over 10 repeats, as computed by the 2023 notebook on the course log
Matched roundSEPSEOPSETSDATS
1000.230.260.340.240.26
2,1000.220.230.220.220.25
4,1000.230.220.220.220.25
6,1000.250.230.210.220.25
8,1000.270.230.210.220.25
10,1000.290.230.210.220.25
12,0000.310.240.210.220.25
14,0000.330.260.210.220.25
16,0000.350.270.220.220.25
18,0000.360.280.210.220.25
20,0000.380.280.210.220.25

Replaying the runs step by step

What the port uncovered

SE

SE paid for exploration, then locked on

Under heavy-tailed delays SE eliminated arms 6, 9, 0, 3, 4, 5, 8, 1 and finally 7 (internal round 15,767). The notebook's own elimination log misses arm 0, because its np.any check tests arm indices. Arm 2 alone survived and took 43% of all pulls; the remaining gap to 0.509 is the cost of round-robin exploration while arm 2's own feedback was among the slowest to arrive.

arm 0: 436 pulls (2.2%), eliminated0arm 1: 1,729 pulls (8.6%), eliminated1arm 2: 8,520 pulls (42.6%)2arm 3: 573 pulls (2.9%), eliminated3arm 4: 816 pulls (4.1%), eliminated4arm 5: 1,385 pulls (6.9%), eliminated5arm 6: 269 pulls (1.3%), eliminated6arm 7: 4,294 pulls (21.5%), eliminated7arm 8: 1,663 pulls (8.3%), eliminated8arm 9: 315 pulls (1.6%), eliminated9
SE pulls per arm on the course log
arm 0436
arm 11729
arm 28520
arm 3573
arm 4816
arm 51385
arm 6269
arm 74294
arm 81663
arm 9315
PSE

PSE survived 90% packet loss on the best arm

45,711 feedback events were lost, yet PSE never eliminated arm 2. It ended in phase 3 with arms 2 and 7 active, yet split its pulls evenly between arms 1, 2, 7 and 8: all four entered phase 3 together and none reached that phase's ≈4,400-observation threshold.

arm 0: 1,102 pulls (5.5%), eliminated0arm 1: 3,559 pulls (17.8%), eliminated1arm 2: 3,535 pulls (17.7%)2arm 3: 1,102 pulls (5.5%), eliminated3arm 4: 1,102 pulls (5.5%), eliminated4arm 5: 1,102 pulls (5.5%), eliminated5arm 6: 276 pulls (1.4%), eliminated6arm 7: 3,561 pulls (17.8%)7arm 8: 3,559 pulls (17.8%), eliminated8arm 9: 1,102 pulls (5.5%), eliminated9
PSE pulls per arm on the course log
arm 01102
arm 13559
arm 23535
arm 31102
arm 41102
arm 51102
arm 6276
arm 73561
arm 83559
arm 91102
OPSE

OPSE eliminated nothing

Pull counts grow on every logged event while observations grow only on matches, so the optimistic upper bounds never fell below anyone's lower bound. Pulls stayed flat (1,922–2,055 per arm) and the reward tracks the log average.

arm 0: 1,922 pulls (9.6%)0arm 1: 2,055 pulls (10.3%)1arm 2: 2,010 pulls (10.1%)2arm 3: 1,994 pulls (10.0%)3arm 4: 1,968 pulls (9.8%)4arm 5: 2,016 pulls (10.1%)5arm 6: 2,030 pulls (10.2%)6arm 7: 1,962 pulls (9.8%)7arm 8: 2,020 pulls (10.1%)8arm 9: 2,023 pulls (10.1%)9
OPSE pulls per arm on the course log
arm 01922
arm 12055
arm 22010
arm 31994
arm 41968
arm 52016
arm 62030
arm 71962
arm 82020
arm 92023
TS

TS won, and the 10-repeat curve is a different TS

The single run played arm 2 in 69% of rounds for 0.4218, the best of the five. The repeat helper called TS(n_arms, n_rounds, rng=rng), so μ₀ became 20,000 in an integer array; that curve sits near 0.22.

arm 0: 508 pulls (2.5%)0arm 1: 777 pulls (3.9%)1arm 2: 13,743 pulls (68.7%)2arm 3: 526 pulls (2.6%)3arm 4: 600 pulls (3.0%)4arm 5: 814 pulls (4.1%)5arm 6: 395 pulls (2.0%)6arm 7: 1,336 pulls (6.7%)7arm 8: 852 pulls (4.3%)8arm 9: 449 pulls (2.2%)9
TS pulls per arm on the course log
arm 0508
arm 1777
arm 213743
arm 3526
arm 4600
arm 5814
arm 6395
arm 71336
arm 8852
arm 9449

Decision record · ADR-002

Why DATS (0.245) trails plain TS (0.422)

What happened. DATS deactivated arms 4, 2, 5, 1 and 3 by its 854th round, including arm 2, the best. Its play() then returned a position in the active-arm list rather than an arm id, so with five arms left it chose among arms 0–4: 99.0% of its pulls went there. A uniform mix of arms 0–4 earns 0.2442, essentially the 0.245 the notebook printed.

Contributing quirks. The arm with the fewest multinomial draws is chosen; propensities are proportional to the index of each arm's largest Monte Carlo sample; every arm's pseudo-rewards reuse the reward just observed; pulls are counted twice.

Decision. Keep the faithful port. The revival documents what was submitted, and the parity tests pin these numbers down. A corrected DATS would be a new, clearly labelled contestant, not a silent fix.

2026 follow-up. An ablation over 20 seeds corrects one quirk at a time and measures each step with a paired interval. Read DR-003 or see it on the uncertainty page.

arm 0: 4,043 pulls (20.2%)0arm 1: 3,988 pulls (19.9%), eliminated1arm 2: 3,932 pulls (19.7%), eliminated2arm 3: 3,868 pulls (19.3%), eliminated3arm 4: 3,974 pulls (19.9%), eliminated4arm 5: 127 pulls (0.6%), eliminated5arm 6: 45 pulls (0.2%)6arm 7: 18 pulls (0.1%)7arm 8: 4 pulls (0.0%)8arm 9: 1 pulls (0.0%)9
DATS pulls per arm on the course log
arm 04043
arm 13988
arm 23932
arm 33868
arm 43974
arm 5127
arm 645
arm 718
arm 84
arm 91

DATS pulls per arm. Faded arms were eliminated; most pulls still landed on arms 0–4.

Method

How “digit for digit” was achieved

Bandits branch on random draws and on ties, so approximate random numbers would send a port down a different path within a few rounds. Instead, the port reproduces numpy's generator exactly and starts from the state the notebook was in.

  1. 1

    Re-run the original

    A script executes the unchanged notebook cells (Python 3.13.9, NumPy 2.5.3, SciPy 1.18.1, arm64) and saves only numbers: the generator state before each evaluation, averages, curve checkpoints and per-arm aggregates.

  2. 2

    Port the generator

    SeedSequence, PCG64 (with a tiny WebAssembly core), Lemire bounded integers, choice, the 256-layer ziggurat normal, binomial inversion and multinomial, SciPy's Pareto and Bernoulli samplers, numpy's pairwise summation and argmax rules.

  3. 3

    Match the floating point

    numpy's Apple-silicon build fuses loc + scale·z into one rounding; the port emulates that fma. An x86-64 numpy rounds twice, which can flip a last bit and, through argmax ties, occasionally a choice.

  4. 4

    Test it

    Vitest replays every evaluation cell from its recorded state and requires exact equality, plus checkpoints of the 10-repeat curves. The course log stays in the repository for these tests and is never shipped to the site.