It takes a very small dose.

Model poisoning is the manipulation of the data or the model an AI system learns from, so that it behaves the way an attacker wants. This site explains how it works, follows the research and collects what defenders can do about it.

Read the explainer See the timeline

Each dot stands for a training document; 5 of the 1,152 are poisoned. In a 2025 study, about 250 documents were enough to plant a backdoor in a 13-billion-parameter model. Drawn to scale by share of training tokens, there would be one violet dot for every 625,000 grey ones. Read about the study

Start here

Six articles that cover the essentials, from the basic definition to a practical defence checklist.

Timeline

Real cases and research milestones, each with its primary source.

Full timeline with sources

  1. 2016Microsoft's Tay is taught to be offensive in 16 hours
  2. 2017BadNets shows that outsourced training can hide a backdoor
  3. 2023Poisoning web-scale datasets turns out to cost about $60
  4. 2024Sleeper Agents: backdoors survive safety training
  5. Artists get Nightshade
  6. France exposes the Pravda network
  7. 2025Chatbots repeat propaganda a third of the time
  8. NIST updates its taxonomy of AI attacks
  9. 250 documents are enough to backdoor an LLM
  10. 2026Microsoft describes how to spot a backdoored model

Glossary

The terms that come up again and again, in plain language.

All 27 terms