Decision: the site offers an optional experiment in which a large language model chooses the arm at every round. It runs only with the visitor's own Anthropic or OpenAI key, called directly from their browser, on a 200-round per-arm replay of the synthetic log. Its reward and regret are paired by seed with Beta(1, 1) Thompson sampling and UCB1, the headline comparators, and with the 2023 Thompson sampling for continuity. Every call is recorded in an audit log that stays on the visitor's device.
Context
I wanted the site to answer a question people now ask about every decision process: could an LLM just do this? A bandit is a clean test because the right answer is known. The project has no budget for model calls, the site is static, and I will not put an API key in a public repository or behind a proxy that strangers can use.
Options considered
- Call a model from a small server with my own key.
- Bring-your-own-key, with calls made from the visitor's browser.
- The coursework's rejection replay, where only decisions that match the logged arm count.
- A per-arm replay, where the k-th pull of arm a receives the k-th logged reward of arm a.
- A multi-turn conversation per episode, or a compact single-turn summary per round.
- Compare the model with the 2023 Thompson sampling only, or with standard reference policies as well.
Why
Option 1 costs money with every visit and turns the site into an open proxy, so I chose option 2. The key lives in
sessionStorage unless the visitor ticks "remember on this device", it is sent only to the provider's own API, and
it never reaches this site, which has no server.
Option 3 would need about ten calls per evaluated round, because a uniformly random logging policy over ten arms matches one decision in ten. That is roughly 2,000 calls for 200 rounds. The log is non-contextual and its rewards are independent within each arm, so option 4 gives an unbiased episode with one call per round. It also pairs the policies on common random numbers, since the k-th pull of an arm gets the same reward and the same delay draw whichever policy makes it. Delays are drawn from their own generator, so a policy cannot change them.
I chose a compact single-turn prompt per round over a growing conversation. Each call then costs about the same, and the audit log shows exactly what the model saw. The prompt holds the pull, report and click counts per arm, the pending feedback per arm and the last five choices. A pull stays pending until its feedback arrives, and lost feedback stays pending for ever, so pulls always equal reported plus pending. The model cannot tell lost feedback from late feedback, which is exactly the position the classical policies are in.
The model must answer {"arm", "reason"}. The request asks for that JSON schema through structured output, and
the answer is validated again with zod. An answer that is malformed, out of range, refused or cut off is
counted as invalid and replaced by a uniformly random arm from a seeded generator, so the invalid-output rate is
measured instead of hidden. Transport failures such as a bad key, an exhausted rate limit or a blocked network stop
the run.
For option 6 I chose standard references. The 2023 Thompson sampling uses a Gaussian prior with σ = 4, which is far too flat for click rates over a few hundred rounds, so beating it says little. Beta(1, 1) Thompson sampling is the usual reference in LLM-as-a-bandit studies, and UCB1 is the classic optimism baseline.
The comparison follows four rules. Runs are grouped by provider, model, delay and horizon, and never pooled across models. Within a group only the latest accepted run of each repetition counts, so n is the number of distinct seeds. A run that stopped early can only be rejected, and it is shown next to the comparators' regret after the same number of rounds. The 95% bootstrap interval for the paired difference appears from five repetitions; below that the page lists the per-seed differences, because a percentile bootstrap of two values only spans those two values. Comparator episodes are computed in the browser for any repetition, so every accepted run has a partner.
Claude Haiku 4.5 is the default because it is the cheapest Claude tier, about US$0.11 for 200 rounds at September 2026 list prices. Claude Sonnet 5.5 is offered at low effort. No sampling temperature is sent and no server-side model fallback is enabled, because the model under test must be the model that answers.
What happened
The harness is covered by tests with every provider call mocked. They check the request shapes, the browser-access header, error mapping, key redaction in the audit log, the fallback rule, the pending invariant, the grouping and deduplication of runs, the exclusion of partial runs and the summaries.
A review before release found three flaws in my first version, all fixed before it shipped. Lost feedback was not counted as pending, so the model could work out how much feedback had been lost, which no comparator knows. The comparison pooled runs from different models and counted a repeated seed twice. A run stopped after 50 rounds could be accepted and shown next to full-length baselines, which flattered it.
I have not published LLM results on the site. They depend on the visitor's model and key, and a single run per model would be one sample. What the site does publish is the bar a model has to clear. Over 20 seeds at 200 rounds with the 2023 heavy-tailed delays, Beta(1, 1) Thompson sampling accumulates 32.8 [30.1, 35.5] clicks of regret and UCB1 42.3 [41.0, 43.5]. The 2023 Thompson sampling accumulates 52.5 [52.1, 53.0], close to the 57.3 of uniform play. With 90% of the best arm's feedback lost, UCB1 does better than with instant feedback, because the arm's confidence bonus shrinks only with reported feedback and keeps it in play. That is luck of the set-up, not a virtue of the algorithm.
What I'd change
I would run a pre-registered evaluation myself across several models and 20 seeds each, publish the per-call audit logs alongside the results, and add a second environment with drifting click rates. The current page cannot tell whether a model reasons about delays or simply prefers the arm with the highest observed rate.