LLM Data and Model Poisoning
Also known as: Training data poisoning, RAG poisoning, Knowledge base poisoning, LLM04
A risk in which untrusted or tampered training, fine-tuning or retrieval data shifts a model's behaviour; it is reduced by data provenance, validation, dataset versioning, pre-deploy evaluation and source allowlists.
How it works
Data and model poisoning is a risk in which the data a language model learns from, is fine-tuned on, or retrieves at answer time is untrusted or has been tampered with, so the model's behaviour is shifted in ways its owners did not intend. The affected data can be a pre-training corpus, a fine-tuning set, an embedding store or a retrieval-augmented generation (RAG) knowledge base.
Conceptually the problem is that a model treats its data as ground truth. If a pipeline pulls documents from open or unverified sources and loads them with no record of where they came from and no checks on what they contain, content from outside the organisation's control can influence what the model says, can bias outputs on a topic, or can make the model behave differently on particular inputs. In a RAG system the same idea applies to answers: whatever the retriever is allowed to fetch can steer the response.
This is hard to see because the model usually keeps working normally. Headline accuracy can look fine while behaviour on narrow topics or specific inputs has quietly changed. That is why defence relies on knowing the origin of every record, validating content before it is ingested, versioning datasets so changes can be reviewed and rolled back, and testing behaviour before each release.
Defence is layered. Trusted and authenticated sources with recorded provenance are the primary control; validation and anomaly detection on the corpus, reviewed and versioned datasets, evaluation and red-team suites that run before deploy, RAG source allowlists, and least-privilege write access to the data store limit exposure and catch what slips through. Keep a rollback path to the last known-good dataset and model.
Walk through it
- 1Read the ingestion config
- 2Name the core gap
- 3Read the signals, pick the gate
- Scope the corpus exposure
- Harden with provenance, validation and evals, then verify
Before the next fine-tune and knowledge-base refresh, you review how the assistant's data is collected. Read the pipeline configuration as a security engineer and note where the documents come from and what is checked on the way in.
1name: support-assistant-refresh2sources:3 - type: web_crawl4 seeds: ["https://community.example.net/", "https://forum.example.org/"]5 - type: public_dataset6 url: "https://files.example.com/latest/support-pairs.jsonl"7validation: none8provenance: none9dedupe: false10writers: ["all-engineers", "ci-bot"]11dataset_version: latest # overwritten on every run12eval_before_deploy: false13rag_index:14 allowed_sources: "*"Spot it
- New documents or domains in the corpus with no source record, owner or ingest approval attached.
- Clusters of near-duplicate or unusually uniform documents on one narrow topic that do not match the rest of the corpus.
- Held-out evaluation scores on specific slices dropping after a refresh while the overall score stays flat.
- Unexpected model answers or behaviour on particular inputs that were stable in the previous release.
- A single external source suddenly supplying a large share of RAG retrievals, or index writes from an account that does not normally write.
Ingestion pipeline
ts=2026-10-11T02:10:04Z run=41 event=ingest source=web_crawl domain=new-domain.example docs=1860 provenance=missing validation=skipped
ts=2026-10-11T02:10:09Z run=41 event=dedupe skipped=true near_duplicate_clusters=1 cluster_size=212Evaluation harness
ts=2026-10-11T03:02:30Z run=41 suite=slice_refund_policy baseline_acc=0.91 candidate_acc=0.77 delta=-0.14 status=needs_review
ts=2026-10-11T03:02:31Z run=41 suite=overall baseline_acc=0.88 candidate_acc=0.88 delta=0.00 status=passsplSplunk: ingest events with missing provenance
index=ml sourcetype=ingest event=ingest provenance=missing
| stats sum(docs) as docs by run, domain
| sort - docsAny ingested volume without a source record deserves review; alert when a run adds documents from domains not on the allowlist.
kqlKQL: evaluation slices that regressed against the baseline
MLEvalResults
| where TimeGenerated > ago(1d)
| extend delta = CandidateAccuracy - BaselineAccuracy
| where delta < -0.05
| project TimeGenerated, RunId, Suite, BaselineAccuracy, CandidateAccuracy, deltaSlice-level checks matter because the overall score can stay flat while one topic shifts.
Stop it
Ingest only from trusted, authenticated sources and record provenance
Maintain an inventory of approved data sources, authenticate them, and attach a source, owner and timestamp to every record. Anything without a verifiable origin stays out of training, fine-tuning and retrieval data.
Validate data and version every dataset
Run validation and anomaly checks on content before ingest: schema, deduplication, outlier and near-duplicate detection, and review of new domains. Pin datasets to immutable versions with change review and a rollback path to the last known-good set.
Evaluate and red-team before deploy, and restrict the data store
Gate every release on evaluation and red-team suites that include targeted slices and compare against a trusted baseline. Limit RAG retrieval to an allowlist of sources and give write access to the corpus and index only to the few identities that need it.
Allowlisted sources, provenance, validation and an eval gate
Vulnerable
sources:
- type: web_crawl
seeds: ["https://community.example.net/"]
validation: none
provenance: none
eval_before_deploy: false
rag_index:
allowed_sources: "*"Hardened
sources:
- type: internal_export
id: kb-approved
auth: service-identity
validation: [schema, dedupe, outlier_check, domain_allowlist]
provenance: required
dataset_version: pinned
review: two-person
eval_before_deploy: true
rag_index:
allowed_sources: ["kb-approved"]
writers: ["ingest-service"]Every control sits in the pipeline definition, so it is reviewable and enforced on each run rather than left to habit.
- Keep an inventory of approved training, fine-tuning and RAG sources, and reject anything outside it.
- Require a provenance record (source, owner, timestamp, hash) for every document before ingest.
- Run schema, deduplication, outlier and near-duplicate checks on each refresh and review new domains by hand.
- Pin datasets to immutable versions, require review of changes, and keep a rollback to the last known-good dataset and model.
- Gate releases on evaluation and red-team suites that include slice-level tests against a trusted baseline.
- Restrict RAG retrieval to allowlisted sources and limit write access to the corpus and vector index to a dedicated service identity.
If it already happened
Pause deployment of the affected model, switch retrieval to the last known-good index, and suspend ingest from the suspect sources and write identities.
Use provenance and dataset versions to identify and remove the suspect records, rotate credentials with write access to the corpus, and rebuild from the last known-good dataset.
Re-run the full evaluation and red-team suites on the rebuilt model and index, redeploy through the gate, and monitor outputs closely for the affected topics.
Add the missing allowlist, validation and slice-level evaluations that would have caught it, and make provenance a hard requirement for every source.
Check yourself
1. What makes an ingestion pipeline exposed to data poisoning?
2. Why is the overall evaluation score alone a weak safeguard?
3. Which control best limits what a RAG assistant can be steered by?