- Extends: ADR-004 in the root README (the opt-in step limit)
Decision: every bandit in the lab, on /evaluation and in the 2023 reproduction is scored with the delayed
replay exactly as the notebook implemented it. The only addition is an opt-in step limit that reports a run which
can never finish as "stalled". The LLM experiment uses a different, per-arm replay, recorded in DR-005.
Context
The project scored bandits offline, from a log of 300,000 news-recommendation events collected by a uniformly random policy over ten arms. Li, Chu, Langford and Wang (WSDM 2011) showed that keeping only the events where a bandit's choice matches the logged arm gives an unbiased picture of how the bandit would have done online, as long as the logging policy was uniform and feedback is instant. The brief then asked for feedback that arrives late or never.
The notebook's answer was a queue. Each logged event is scheduled to be revealed delay(arm, reward) steps after
it is read. At every step the bandit chooses an arm, and if an event revealed at that step belongs to the chosen
arm, the first such event becomes this round's feedback. Revealed events for other arms are thrown away, lost
feedback is never revealed, and the loop wraps around the log until enough rounds have matched.
Options considered
- Keep the notebook's queue replay unchanged and document what it measures.
- Rewrite the replay so delays only affect what the bandit knows, with every decision counted.
- Simulate from per-arm click rates fitted to the log and drop replay altogether.
- Serve each arm's logged events in order, a per-arm replay that needs no matching.
Why
The revival exists to show what was submitted, and the parity tests reproduce every printed number exactly, so option 1 was the only choice that kept the 2023 results meaningful. Options 2 to 4 would each have changed every number on the site. Option 4 is still the right tool where cost matters more than continuity with 2023, which is why the LLM experiment uses it (DR-005).
The step limit was needed because the original loop never terminates when a bandit insists on an arm whose feedback can never arrive, for example with 100% loss. By default the limit is infinite, so the replay is unchanged unless the lab asks for it.
What happened
The TypeScript replay reproduces all five 2023 averages to the last digit, and the seeded evaluation (DR-004) re-uses it without modification.
The seeded evaluation also exposed a weakness of this replay that I had not appreciated in 2023. Under the PSE setting, 90% of arm 2's feedback is lost. A choice of arm 2 then matches only when the logged arm is 2 and its event survives, about one step in a hundred. Rounds that count are therefore drawn mostly from the arms a bandit explores, and less from the arm it prefers. Epsilon-greedy spent 85% of its matched rounds on arm 2 with instant feedback and only 36% under 90% loss, and its lead over Thompson sampling fell from +0.114 [+0.107, +0.120] to +0.004 [−0.001, +0.009]. A live epsilon-greedy would still play arm 2 most of the time, so part of that drop is an artefact of the replay. UCB1 is deterministic between updates, so it keeps choosing arm 2 until a match arrives, and its matched rounds stay 82% on arm 2.
None of the 1,050 runs in the delay-sensitivity grid stalled, so the step limit has not changed a published number.
What I'd change
I would score delayed feedback with a replay in which every decision counts and the delay only controls when the bandit learns the outcome, which is what a delayed-feedback bandit faces in production. The per-arm replay from DR-005 already does this for non-contextual logs. I would also report the number of log steps consumed per matched round next to every delayed result, because that ratio is the first sign that the replay, not the bandit, is driving a number.