Skip to content
Bandit Lab, home

COMP90051 · Statistical Machine Learning · The University of Melbourne

Ten arms, one log, and feedback that arrives late, or never.

A multi-armed bandit picks one of several options each round and only learns how that one did. For a 2023 project I implemented bandits that cope with delayed and lost feedback, plus two kinds of Thompson sampling, and scored them by replaying 300,000 recorded clicks. Bandit Lab rebuilds that work in the browser: the same algorithms, ported line for line, racing on a log you can reshape.

Watch the tour: three captioned walkthroughs and screenshots of every feature
Arm · click-through rate
  1. .110
  2. .291
  3. .512
  4. .133
  5. .184
  6. .265
  7. .016
  8. .387
  9. .308
  10. .069
Best arm 2 clicks 50.9% of the time300,000 logged events
Click-through rate of each of the ten arms in the 2023 course log. Arm 2 is best at 50.9 percent.

The brief, paraphrased

Recommend the next news story, without ever seeing the counterfactual.

The project framed bandits as a news recommender: show one of ten stories, count a click as reward. It came with a log in which a uniformly random policy chose the story, so any algorithm can be scored offline by keeping only the rounds where it agrees with the log (Li et al., 2011). Each algorithm ran for 20,000 matched rounds.

  1. 01

    Feedback that arrives late, or never

    Adapt replay evaluation so a click can be revealed rounds after the choice that earned it, or be lost entirely. Then implement three successive-elimination bandits from Lancewicki et al. (ICML 2021) and test them under heavy-tailed, lossy and reward-dependent delays.

  2. 02

    Thompson sampling

    Give every arm a Gaussian prior over its mean reward, update it in closed form after each observed click, and play the arm whose posterior draw is largest.

  3. 03

    Adaptive inference

    Implement Doubly-Adaptive Thompson Sampling (Dimakopoulou, Ren & Zhou, NeurIPS 2021): Thompson sampling on doubly-robust, variance-stabilised estimates with arm elimination.

What the notebook printed

Key results, reproduced to the last digit

Average reward over 20,000 matched rounds on the course log (one run, seed 90051). For scale, always playing the best arm earns 0.509 and picking at random about 0.223.

Why DATS trails TS

How the scoring works

Replay a random log; keep only the agreements

Because the logging policy was uniformly random, the rounds where an algorithm happens to agree with the log are an unbiased sample of what it would have experienced live. Delays add a queue: the logged click is scheduled for a later round, and only counts if the algorithm picks that arm again when it arrives.

  1. 1

    Read

    Take the next logged event: which arm was shown, and whether it was clicked.

  2. 2

    Schedule

    Draw a delay for that event and put its click in the queue for a future round (or drop it if lost).

  3. 3

    Play

    Ask the bandit for an arm. It sees only the feedback that has arrived so far.

  4. 4

    Match

    If a click due this round belongs to the arm it chose, the bandit learns from it and the round counts.

In the lab

Everything the notebook plotted, plus what it couldn't show

About this project

A 2023 coursework project, revived in 2026

The original was an individual Jupyter notebook marked on code. This site keeps its algorithms exactly as submitted, quirks included, and adds the interactive tooling it never had. The original notebook is preserved in the repository for reference.

View the repository
Subject
COMP90051 Statistical Machine Learning
University
The University of Melbourne
Teaching period
2023, Semester 1 · Project 2
Team
Sunchuangyu “Rin” Huang (individual project)
Credits
Delay distributions, skeleton and the logged dataset were supplied by the COMP90051 teaching team and are not redistributed here. Algorithms follow Lancewicki et al. (2021), Agrawal & Goyal (2012) and Dimakopoulou, Ren & Zhou (2021).
Original stack
Python 3, NumPy, SciPy, Matplotlib, Jupyter
Revived stack
Next.js 16, React 19, TypeScript, Tailwind CSS 4, shadcn/ui on Base UI, Web Workers, WebAssembly, Vitest
Academic integrity
Shared as a record of completed work. Please don't reuse it in your own assessment.