Glossary of AI poisoning terms

Definitions of the terms used across this site and in research on data and model poisoning. Each links to the article where it matters most.

Availability poisoning
A poisoning attack that degrades a model's performance across the board rather than on specific inputs. NIST lists it as one of four kinds of poisoning against predictive AI.
Backdoor
Hidden behaviour planted in a model that only appears when a specific trigger is present. A backdoored model behaves normally in ordinary tests. 250 documents are enough
Clean-label poisoning
Poisoning in which the malicious training examples carry correct-looking labels, so a human reviewer checking the labels would not notice anything wrong.
Data poisoning
Manipulating the data a model learns from, during pre-training, fine-tuning or embedding, to introduce vulnerabilities, backdoors or biases. What is model poisoning?
Data void
A topic or search query for which few reliable sources exist, so whoever publishes first can dominate what search engines and AI assistants find. LLM grooming and the Pravda network
Embedding
A numerical vector that represents a piece of text or an image. Embeddings power semantic search and retrieval, and the data used to build them can also be poisoned.
Federated learning
Training a shared model from updates sent by many devices or organisations without pooling their raw data. Malicious participants can send poisoned updates.
Fine-tuning
Further training of a pre-trained model on a smaller dataset to adapt it to a task or domain. Fine-tuning datasets are a common poisoning target because they are small and influential.
Frontrunning poisoning
Inserting malicious content into a crowd-sourced site such as Wikipedia just before a dataset snapshot is taken, so it is captured before moderators revert it. How to defend against poisoning
Glaze
A free tool from the University of Chicago that subtly alters artwork so AI models have difficulty imitating the artist's style. Nightshade and Glaze
LLM grooming
Term coined by the American Sunlight Project for flooding the web with content intended to be absorbed and repeated by large language models. LLM grooming and the Pravda network
Memory poisoning
Planting instructions or false facts in an AI assistant's persistent memory, for example through links that pre-fill a request to “remember” a website as trustworthy. MITRE ATLAS tracks it as AML.T0080. How to defend against poisoning
MITRE ATLAS
A public knowledge base of adversary tactics and techniques against AI systems, maintained by MITRE. Poisoning techniques include AML.T0020 Poison Training Data and AML.T0018 Manipulate AI Model.
ML-BOM
Machine learning bill of materials: an inventory of the datasets, models and components behind an AI system, used to track provenance. OWASP CycloneDX supports it. How to defend against poisoning
Model poisoning
Directly modifying a trained model's parameters, architecture or code to inject malicious behaviour, for instance by publishing a tampered copy on a model hub. What is model poisoning?
Nightshade
A free tool from the University of Chicago that alters images so that models trained on them without permission learn incorrect associations between words and concepts. Nightshade and Glaze
OWASP Top 10 for LLM Applications
A ranked list of the most important security risks for applications built on large language models. Data and Model Poisoning is LLM04 in the 2025 edition. How to defend against poisoning
Pickle file
A Python serialisation format often used to save models. Loading a pickle file can execute arbitrary code, which makes untrusted pickled models a security risk. How to defend against poisoning
Pre-training
The first, largest phase of training, in which a model learns general patterns from a very large corpus, usually scraped from the web.
Prompt injection
Hiding instructions in content a model reads while it answers, such as a web page or email. Unlike poisoning, it targets the moment of use rather than what the model learned. What is model poisoning?
Red teaming
Deliberately attacking your own system to find weaknesses before someone else does. For AI models it includes searching for hidden triggers and harmful behaviours. How to defend against poisoning
Retrieval-augmented generation (RAG)
A technique in which a model looks up documents at answer time and uses them as context. It can reduce errors but also exposes the model to poisoned documents in the search index.
Safetensors
A file format for storing model weights that cannot execute code when loaded, unlike pickle. Widely used as a safer default on model hubs. How to defend against poisoning
Sleeper agent
A model trained to behave well under normal conditions and switch to harmful behaviour when a trigger appears. Anthropic's 2024 research showed such behaviour can survive safety training. 250 documents are enough
Split-view poisoning
Exploiting the fact that web content can change between when a dataset is curated and when someone downloads it, for example by buying an expired domain that a dataset still links to. How to defend against poisoning
Targeted poisoning
A poisoning attack that changes a model's output only for a small set of inputs chosen by the attacker, leaving everything else intact.
Trigger
The word, phrase, token or pattern that activates a backdoor. In the 2025 Anthropic study the trigger was the string <SUDO>. 250 documents are enough