Skip to content
Bandit Lab, home

Three walkthroughs · 2:21 in total

A guided tour

Three short recordings of the main workflows, captioned step by step, and screenshots of each key feature. They are recorded by a Playwright script in the repository that also checks each journey end to end. The lab and evaluation runs use fixed seeds, so the numbers on screen are the ones you get when you try them.

Walkthrough 01 · 0:53 · captioned, no audio

The race

Configure ten arms and heavy-tailed Pareto delays, race SE, PSE, OPSE, Thompson sampling, DATS and UCB1 on the same seeded log, then read the same comparison with 95% bootstrap bands and paired differences over 20 repetitions.

Open the lab

Steps and transcript

Walkthrough 02 · 0:38 · captioned, no audio

Inside Thompson sampling

Step through the first repeat snapshot by snapshot: Thompson sampling's sampling distributions drift towards the true click rates and narrow as feedback arrives, and the arm-pull heatmaps show where each algorithm spent its rounds.

Open this set-up

Steps and transcript

Walkthrough 03 · 0:50 · captioned, no audio

LLM vs Thompson sampling

Open the bring-your-own-key settings, run the LLM-as-policy episode and compare its regret with Beta-TS, UCB1 and the 2023 TS on the same seed, then review the run and inspect the audit log.

Open the LLM experiment

Mocked AI response for illustration. No real model was called in this recording: a placeholder key was typed and every provider request was intercepted and answered by a scripted stand-in policy, whose answers say “mock reply”.

Steps and transcript

Key features

Screenshots

Desktop views at 1440 × 900 and phone views at 390 × 844. Select one to enlarge it; the arrow keys move between them. The two LLM screenshots show mocked AI responses for illustration.

On a phone (390 × 844)

The recorder is web/e2e/showcase.spec.ts in the repository. Run pnpm showcase to replay every journey and rebuild these files.