Data poisoning is an attack on artificial intelligence in which someone plants manipulated or malicious data where an AI system will learn from it or look things up in it. The goal is to make the system behave the way the attacker wants later: make worse decisions overall, misclassify specific inputs, or carry a hidden “backdoor” that only activates when a trigger word, image or pattern appears. Because the damage is built in before the system is used, it can be hard to spot in normal operation.
At a glance
- Poisoning targets data, not code: training sets, fine-tuning data, feedback loops, or the documents an assistant retrieves at run time.
- Effects range from general degradation to narrow, targeted behavior that only shows when a trigger is present.
- Models trained on public or scraped data, and assistants fed from widely editable sources, are the most exposed.
- Defenses focus on knowing where data came from, controlling who can change it, and testing models before and after updates.
What problem it solves
Data poisoning is a threat, not a product, so the useful question is what understanding it helps a buyer do. Many organizations now build on machine learning (ML) and large language models (LLMs) without treating the data behind them as part of the attack surface. A firewall or endpoint tool won’t notice a spreadsheet of subtly mislabeled examples, a fake web page written to be scraped, or a planted document in a shared drive that an AI assistant uses as its source of truth.
Naming the risk lets security and data teams put controls around AI data the same way they do around code and infrastructure: who can add to it, how changes are reviewed, and how the organization would notice and roll back a bad change. It also gives buyers concrete questions for AI vendors, whose answers otherwise tend to stop at “our model is secure.”
How it works
Training and fine-tuning poisoning. An attacker gets manipulated examples into the data used to train or fine-tune a model. That might mean contributing to a public dataset, publishing content likely to be scraped, compromising a data supplier, or abusing an internal pipeline with weak access controls. Research has shown that a fairly small share of poisoned examples can be enough to plant some behaviors, though how much is needed varies widely by model and attack.
Backdoors. A targeted form teaches the model to behave normally except when a specific trigger appears, for example a phrase that makes a content filter wave something through. Ordinary testing may never hit the trigger.
Retrieval and knowledge-base poisoning. Systems that use retrieval-augmented generation (RAG) answer from documents they look up at run time. Planting false or malicious content in those sources changes answers without touching the model. Planted content can also carry hidden instructions, which is where poisoning meets prompt injection.
Feedback poisoning. Systems that learn from user ratings, corrections or ongoing data can be steered by coordinated bad input over time.
Defenses are layered: track data provenance, restrict and log who can change training sets and knowledge bases, filter and validate incoming data, keep versioned copies so you can roll back, and test models against known-good and adversarial cases before and after each update. Like other AI security risks, it belongs in your broader AI governance program rather than in a separate silo.
When it matters for buyers
- When you fine-tune or train models on your own data. Your data pipeline becomes a security boundary that needs owners, access controls and change review.
- When an AI assistant answers from shared content. Anyone who can edit a source folder, wiki or ticket system can influence answers.
- When you buy AI from a vendor. Their training data and update process are part of your supply chain risk.
- When customers or auditors ask about AI controls. Frameworks such as the NIST AI Risk Management Framework and the ISO/IEC 42001 standard expect you to manage data quality and integrity for AI systems.
If you are deciding where AI fits and how to adopt it safely, our artificial intelligence team can help compare platforms and providers.
Questions to ask vendors
- Where does the training and fine-tuning data for this model come from, and how is it vetted?
- Who can change the data or knowledge sources the system uses, and are those changes logged and reviewed?
- How do you test for poisoned or backdoored behavior before releasing a model update?
- Can we see what changed between model versions, and can we roll back if behavior shifts?
- If the system learns from our data or feedback, how is that isolated from other customers?
- How would you notify us if you discovered that a dataset you used had been tampered with?
How it differs from prompt injection
Both attacks get an AI system to do something its owner didn’t intend, but they act at different points. Data poisoning corrupts what the system learns from or retrieves, so the bad behavior is present before anyone uses it and can persist across every user. Prompt injection supplies instructions in the input at the moment the system runs, directly from a user or hidden in content the system reads. Planting a malicious document in a retrieval source can do both at once, which is why teams often defend against them together.
