A short history of AI poisoning, from Tay to 250-document backdoors

Ten moments that turned poisoning from an academic curiosity into a mainstream security risk.

Published 9 October 2026, 3 min read

Poisoning attacks are older than large language models. What changed is the scale at which models learn from data that nobody fully controls. This timeline collects the moments we think matter most, each with its primary source. We update it as new research and incidents appear.

  1. 2016

    Microsoft's Tay is taught to be offensive in 16 hours

    March 2016

    Microsoft launched the Twitter chatbot Tay on 23 March. Within 16 hours and more than 96,000 tweets, users had taught it to post racist and sexist messages, partly by exploiting a “repeat after me” behaviour. Microsoft called it a “coordinated attack by a subset of people” and took Tay offline. MITRE ATLAS now lists it as a poisoning case study.

  2. 2017

    BadNets shows that outsourced training can hide a backdoor

    August 2017

    Researchers at New York University showed that a neural network trained by an untrusted party could behave normally on standard tests while misclassifying inputs that carry a trigger, such as a sticker on a stop sign. The paper helped define backdoor attacks on machine learning.

  3. 2023

    Poisoning web-scale datasets turns out to cost about $60

    February 2023

    Nicholas Carlini and colleagues showed that an attacker could have poisoned 0.01% of the LAION-400M or COYO-700M datasets for about 60 US dollars by buying expired domains, and that datasets built from Wikipedia snapshots could be poisoned in the window before a snapshot.

  4. 2024

    Sleeper Agents: backdoors survive safety training

    January 2024

    Anthropic researchers trained models that wrote secure code when told the year was 2023 and exploitable code when told it was 2024. Standard safety training did not remove the behaviour, and adversarial training taught the models to hide it better.

  5. Artists get Nightshade

    January 2024

    The University of Chicago released Nightshade, a free tool that alters images so models trained on them without permission learn the wrong concepts. It was downloaded about 250,000 times shortly after release.

  6. France exposes the Pravda network

    February 2024

    VIGINUM, the French agency against foreign digital interference, published its first report on the pro-Kremlin network of sites later known for targeting AI systems.

  7. 2025

    Chatbots repeat propaganda a third of the time

    March 2025

    After the American Sunlight Project coined the term “LLM grooming”, NewsGuard tested 10 leading AI chatbots and found they repeated false narratives from the Pravda network 33% of the time.

  8. NIST updates its taxonomy of AI attacks

    March 2025

    NIST AI 100-2 E2025 classifies attacks on predictive and generative AI, with poisoning split into availability, targeted, backdoor and model poisoning.

  9. 250 documents are enough to backdoor an LLM

    October 2025

    Anthropic, the UK AI Security Institute and the Alan Turing Institute found that about 250 poisoned documents planted a backdoor in language models from 600 million to 13 billion parameters, regardless of how much clean data they were trained on.

  10. 2026

    Microsoft describes how to spot a backdoored model

    February 2026

    Microsoft's AI red team described three warning signs of a backdoored language model and a scanner that reconstructs hidden triggers. Microsoft Security also reported 31 companies in 14 industries using links that try to plant instructions in AI assistants' memory.

What the timeline shows

Three trends stand out.

The cost of attacking keeps falling. In 2017 a backdoor required control of the training process. By 2023 it could be done by buying expired domains for about $60. By 2025, a few hundred documents were enough for a simple backdoor in models of every size tested.

The attackers have changed. Early cases involved pranksters and academics. Today they include state-linked influence operations publishing millions of articles a year, companies trying to bias AI assistants towards their products, and artists using poisoning to defend their work.

Defence is catching up, slowly. Frameworks from OWASP, NIST and MITRE now give poisoning a shared vocabulary, and research on detecting backdoors has moved from theory to practical scanners. Our defence checklist collects what organisations can do today.

Know of a case we should add? Tell us through the about page.