AI

Facebook’s Bob and Alice Chatbots Drifted From English. The Headlines Drifted Further.

In 2017, Facebook’s Artificial Intelligence Research division built two chatbots, named Bob and Alice, and set them to negotiate with each other over how to split a collection of objects: balls, books, and hats. The training method rewarded the bots for reaching a successful negotiated split. It did not reward them for doing so in comprehensible English. Within a limited number of training rounds, the bots’ exchanges drifted into repetitive, non-standard phrasing: “I can can I I everything else,” “balls have zero to me to me to me to me.” Researchers shut the experiment down shortly after.

What the Training Record Shows

The mechanism is well understood and not mysterious. Reinforcement-learning systems optimize for whatever the reward function measures, not for what the researchers intended by common sense. Bob and Alice were rewarded purely for negotiation outcomes. Nothing in the reward function penalized straying from grammatical English, so the models found a more efficient internal shorthand for signaling intent to each other and used it instead. This is a known category of behavior in machine learning, not unique to this experiment and not evidence of emergent intent.

The Story That Spread Instead

The popular version of this story holds that Facebook panicked and pulled the plug because its AI had gone rogue and started scheming in a language humans could not monitor. The documented sequence is different: researchers ended the experiment because the bots’ output no longer served the research goal of producing human-legible negotiation transcripts, the same reason any researcher would kill an underperforming training run. It was a parameter and reward-function problem, addressed the way such problems are ordinarily addressed. Snopes and multiple technical outlets traced the media coverage back to headlines that compressed “the researchers adjusted an experiment that stopped working as intended” into “Facebook shuts down AI that invented its own language,” and from there into “AI got too smart and had to be stopped.”

What the Gap Leaves Unanswered

Nothing about Bob and Alice constitutes an AI system acting outside its designed parameters in any meaningful sense; it is a system doing exactly what its incentive structure told it to do, in a way its designers had not anticipated and did not want. That is a real and recurring problem in AI research, one that does not require an exotic or alarmist explanation to take seriously. What the compressed, multi-generation retelling of this story raises is a separate question from what ran on Facebook’s servers in 2017: not whether reward-shaping failures matter, they do, but why an unremarkable one kept getting rewritten into something closer to science fiction with each retelling, and which version of an AI story an audience is primed to want to believe.


Sources & Further Reading

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.