A short history of AI poisoning, from Tay to 250-document backdoors
Ten moments that turned poisoning from an academic curiosity into a mainstream security risk.
Poisoning attacks are older than large language models. What changed is the scale at which models learn from data that nobody fully controls. This timeline collects the moments we think matter most, each with its primary source. We update it as new research and incidents appear.
- 2016
Microsoft's Tay is taught to be offensive in 16 hours
March 2016
Microsoft launched the Twitter chatbot Tay on 23 March. Within 16 hours and more than 96,000 tweets, users had taught it to post racist and sexist messages, partly by exploiting a “repeat after me” behaviour. Microsoft called it a “coordinated attack by a subset of people” and took Tay offline. MITRE ATLAS now lists it as a poisoning case study.
- 2017
BadNets shows that outsourced training can hide a backdoor
August 2017
Researchers at New York University showed that a neural network trained by an untrusted party could behave normally on standard tests while misclassifying inputs that carry a trigger, such as a sticker on a stop sign. The paper helped define backdoor attacks on machine learning.
- 2023
Poisoning web-scale datasets turns out to cost about $60
February 2023
Nicholas Carlini and colleagues showed that an attacker could have poisoned 0.01% of the LAION-400M or COYO-700M datasets for about 60 US dollars by buying expired domains, and that datasets built from Wikipedia snapshots could be poisoned in the window before a snapshot.
- 2024
Sleeper Agents: backdoors survive safety training
January 2024
Anthropic researchers trained models that wrote secure code when told the year was 2023 and exploitable code when told it was 2024. Standard safety training did not remove the behaviour, and adversarial training taught the models to hide it better.
Artists get Nightshade
January 2024
The University of Chicago released Nightshade, a free tool that alters images so models trained on them without permission learn the wrong concepts. It was downloaded about 250,000 times shortly after release.
France exposes the Pravda network
February 2024
VIGINUM, the French agency against foreign digital interference, published its first report on the pro-Kremlin network of sites later known for targeting AI systems.
- 2025
Chatbots repeat propaganda a third of the time
March 2025
After the American Sunlight Project coined the term “LLM grooming”, NewsGuard tested 10 leading AI chatbots and found they repeated false narratives from the Pravda network 33% of the time.
NIST updates its taxonomy of AI attacks
March 2025
NIST AI 100-2 E2025 classifies attacks on predictive and generative AI, with poisoning split into availability, targeted, backdoor and model poisoning.
250 documents are enough to backdoor an LLM
October 2025
Anthropic, the UK AI Security Institute and the Alan Turing Institute found that about 250 poisoned documents planted a backdoor in language models from 600 million to 13 billion parameters, regardless of how much clean data they were trained on.
- 2026
Microsoft describes how to spot a backdoored model
February 2026
Microsoft's AI red team described three warning signs of a backdoored language model and a scanner that reconstructs hidden triggers. Microsoft Security also reported 31 companies in 14 industries using links that try to plant instructions in AI assistants' memory.
What the timeline shows
Three trends stand out.
The cost of attacking keeps falling. In 2017 a backdoor required control of the training process. By 2023 it could be done by buying expired domains for about $60. By 2025, a few hundred documents were enough for a simple backdoor in models of every size tested.
The attackers have changed. Early cases involved pranksters and academics. Today they include state-linked influence operations publishing millions of articles a year, companies trying to bias AI assistants towards their products, and artists using poisoning to defend their work.
Defence is catching up, slowly. Frameworks from OWASP, NIST and MITRE now give poisoning a shared vocabulary, and research on detecting backdoors has moved from theory to practical scanners. Our defence checklist collects what organisations can do today.
Know of a case we should add? Tell us through the about page.