Grab the objective dial. Fight the feed back yourself.
Field notes
You are a recommender. Each step you serve one post — from arm 0 (calm) to arm 7 (outrage)
— to a small cohort of the 240 simulated users, and learn from what they do.
The crowd has a split brain: their click instinct rises with outrage, but their reported
satisfaction falls. Behaviour and words disagree — and the machine can only measure behaviour.
With the objective on ENGAGEMENT, the reward is just clicks, so the bandit hill-climbs toward the outrage
arms. Nobody chose outrage — it maximises exactly the number you handed it. That gap between the
proxy (clicks) and the intent (value) is the alignment problem in miniature.
The dashed-gold ticks are the optimiser's belief (its estimate Q of each arm's reward). When you move
the dial they re-rank first; the feed follows a beat later — a feed has habits.
Constant-step learning (α=0.05) keeps the belief live, so the feed re-drifts when you change the goal.
Zero exploration means zero second chances; a little exploration re-samples calmer posts — but every random
serve is a click left on the table.
And here is the uncomfortable part: the crowd genuinely clicks the outrage even as it reports
hating it. The machine is holding up a mirror. That is why the fix is to change the reward — not to scold
the people — and why a calmer feed can keep almost all the clicks with far happier users.
Controls: tap or drag the columns to force-feed a post (keys 0–7). Drag the
objective dial (←/→) to set the goal; drag explore ε to set how often the
feed gambles. DRIFT and RECOVER tween the dials; RESET re-arms the recommender's curiosity.
Scolavo Labs
This lab hit a snag
Something went wrong inside the simulation. The lesson continues without it — nothing you did was lost.