Glossary of AI poisoning terms
Definitions of the terms used across this site and in research on data and model poisoning. Each links to the article where it matters most.
- Availability poisoning
- A poisoning attack that degrades a model's performance across the board rather than on specific inputs. NIST lists it as one of four kinds of poisoning against predictive AI.
- Backdoor
- Hidden behaviour planted in a model that only appears when a specific trigger is present. A backdoored model behaves normally in ordinary tests. 250 documents are enough
- Clean-label poisoning
- Poisoning in which the malicious training examples carry correct-looking labels, so a human reviewer checking the labels would not notice anything wrong.
- Data poisoning
- Manipulating the data a model learns from, during pre-training, fine-tuning or embedding, to introduce vulnerabilities, backdoors or biases. What is model poisoning?
- Data void
- A topic or search query for which few reliable sources exist, so whoever publishes first can dominate what search engines and AI assistants find. LLM grooming and the Pravda network
- Embedding
- A numerical vector that represents a piece of text or an image. Embeddings power semantic search and retrieval, and the data used to build them can also be poisoned.
- Federated learning
- Training a shared model from updates sent by many devices or organisations without pooling their raw data. Malicious participants can send poisoned updates.
- Fine-tuning
- Further training of a pre-trained model on a smaller dataset to adapt it to a task or domain. Fine-tuning datasets are a common poisoning target because they are small and influential.
- Frontrunning poisoning
- Inserting malicious content into a crowd-sourced site such as Wikipedia just before a dataset snapshot is taken, so it is captured before moderators revert it. How to defend against poisoning
- Glaze
- A free tool from the University of Chicago that subtly alters artwork so AI models have difficulty imitating the artist's style. Nightshade and Glaze
- LLM grooming
- Term coined by the American Sunlight Project for flooding the web with content intended to be absorbed and repeated by large language models. LLM grooming and the Pravda network
- Memory poisoning
- Planting instructions or false facts in an AI assistant's persistent memory, for example through links that pre-fill a request to “remember” a website as trustworthy. MITRE ATLAS tracks it as AML.T0080. How to defend against poisoning
- MITRE ATLAS
- A public knowledge base of adversary tactics and techniques against AI systems, maintained by MITRE. Poisoning techniques include AML.T0020 Poison Training Data and AML.T0018 Manipulate AI Model.
- ML-BOM
- Machine learning bill of materials: an inventory of the datasets, models and components behind an AI system, used to track provenance. OWASP CycloneDX supports it. How to defend against poisoning
- Model poisoning
- Directly modifying a trained model's parameters, architecture or code to inject malicious behaviour, for instance by publishing a tampered copy on a model hub. What is model poisoning?
- Nightshade
- A free tool from the University of Chicago that alters images so that models trained on them without permission learn incorrect associations between words and concepts. Nightshade and Glaze
- OWASP Top 10 for LLM Applications
- A ranked list of the most important security risks for applications built on large language models. Data and Model Poisoning is LLM04 in the 2025 edition. How to defend against poisoning
- Pickle file
- A Python serialisation format often used to save models. Loading a pickle file can execute arbitrary code, which makes untrusted pickled models a security risk. How to defend against poisoning
- Pre-training
- The first, largest phase of training, in which a model learns general patterns from a very large corpus, usually scraped from the web.
- Prompt injection
- Hiding instructions in content a model reads while it answers, such as a web page or email. Unlike poisoning, it targets the moment of use rather than what the model learned. What is model poisoning?
- Red teaming
- Deliberately attacking your own system to find weaknesses before someone else does. For AI models it includes searching for hidden triggers and harmful behaviours. How to defend against poisoning
- Retrieval-augmented generation (RAG)
- A technique in which a model looks up documents at answer time and uses them as context. It can reduce errors but also exposes the model to poisoned documents in the search index.
- Safetensors
- A file format for storing model weights that cannot execute code when loaded, unlike pickle. Widely used as a safer default on model hubs. How to defend against poisoning
- Sleeper agent
- A model trained to behave well under normal conditions and switch to harmful behaviour when a trigger appears. Anthropic's 2024 research showed such behaviour can survive safety training. 250 documents are enough
- Split-view poisoning
- Exploiting the fact that web content can change between when a dataset is curated and when someone downloads it, for example by buying an expired domain that a dataset still links to. How to defend against poisoning
- Targeted poisoning
- A poisoning attack that changes a model's output only for a small set of inputs chosen by the attacker, leaving everything else intact.
- Trigger
- The word, phrase, token or pattern that activates a backdoor. In the 2025 Anthropic study the trigger was the string <SUDO>. 250 documents are enough