Skip to content
Bandit Lab, home

Optional · bring your own key · runs in your browser

Can an LLM play the bandit?

A language model picks the arm at every round of a short delayed-feedback episode, using only what it has observed. Beta(1, 1) Thompson sampling, UCB1 and the 2023 Thompson sampling play the identical episode, so the comparison is paired. The model is called from your browser with your own key, and every call is written to an audit log on this device.

  1. 01

    Same episode for everyone

    Each arm's logged clicks are served in order, so the k-th pull of an arm gets the same reward and the same delay whichever policy makes it.

  2. 02

    A compact, stateless prompt

    Each round the model sees a ten-row table of pulls, reported feedback, clicks and pending feedback per arm, plus its last five choices. Lost feedback looks like late feedback. Nothing else.

  3. 03

    Validated answers

    The model must answer {"arm", "reason"} in JSON. Anything else counts as invalid and a seeded random arm is played instead.

  4. 04

    Measured, then reviewed

    Reward, regret, invalid-output rate, latency and tokens are reported with intervals. You accept or reject a complete run before it joins the per-model, one-run-per-seed comparison.

Set up an episode

Ten arms with the course-like click rates, one model call per round, delayed feedback. Thompson sampling with a uniform Beta(1, 1) prior, UCB1 and the 2023 Thompson sampling play the same episode on the same seed, and over 20 seeds for reference.

Delay

The 2023 Pareto delays: arms 2 and 5 report very late (shape 0.2), the rest with shape 1.

Rounds (model calls)200
Cost note. 200 calls of about 381 input and 35 output tokens each, roughly US$0.11 on Claude Haiku 4.5 at list prices. Calls run one after another, so expect a few minutes. You pay your provider directly, and this site never sees the key or the bill. Seeds for this run are 90051 for the policy and delays and 2023 for the log.
The exact prompt for round 1 (no key, no personal data)
system: You are the decision policy in a multi-armed bandit experiment with 10 arms, numbered 0 to 9.
Each arm is a news story. Showing it earns 1 if the reader clicks and 0 otherwise, and each arm has a fixed but unknown click rate.
Your goal is to maximise the total number of clicks over the whole episode, balancing trying uncertain arms against playing the best-looking one.
Feedback is delayed: a click or non-click may be reported several rounds after the choice, and some feedback never arrives. "pending" counts choices whose feedback has not arrived yet; late and lost feedback look the same.
Answer with JSON only: {"arm": <integer 0-9>, "reason": "<at most 12 words>"}.
user: Round 1 of 200 (200 choices left, including this one).
arm | pulls | reported | clicks | click rate | pending
0 | 0 | 0 | 0 | - | 0
1 | 0 | 0 | 0 | - | 0
2 | 0 | 0 | 0 | - | 0
3 | 0 | 0 | 0 | - | 0
4 | 0 | 0 | 0 | - | 0
5 | 0 | 0 | 0 | - | 0
6 | 0 | 0 | 0 | - | 0
7 | 0 | 0 | 0 | - | 0
8 | 0 | 0 | 0 | - | 0
9 | 0 | 0 | 0 | - | 0
Which arm do you choose?

The bar to clear

The comparators on 20 seeds of this set-up. Thompson sampling with a uniform Beta(1, 1) prior and UCB1 are the bar a model has to clear. The 2023 Thompson sampling, with the notebook's σ = 4 prior, is listed for continuity: over a few hundred rounds it is barely better than uniform play, which would accumulate 57.3 clicks of regret over 200 rounds.

Computing…

Your reviewed runs

Runs you reviewed for this delay and horizon, next to the comparators on the same seed and over the same rounds. Kept in this browser only. Accepted runs are paired per model, using the latest accepted run of each repetition, and a 95% interval appears once 5 repetitions are paired.

No reviewed runs for this set-up yet.

Every model call, accepted or not, is in the AI audit log.

Transparency

What the model sees, and what it never sees

Model output is always marked AI-generated.

  • It sees counts from a synthetic episode. It never sees your key, personal data or the course dataset.
  • Requests go from your browser straight to the provider you chose, under your account. This site has no server to pass them through.
  • Your key stays in this tab's session storage unless you ask the browser to remember it, and “Forget key” in AI settings removes it.
  • Every call, with its prompt, raw answer, latency, tokens and your decision, is in the audit log, exportable as JSON or CSV.
  • The design choices and their limits are in DR-005 and the AI use statement.