
- 01Training data poisoning is harder to detect than prompt injection because the effect is statistical and spread across many outputs rather than visible in a single response.
- 02Provenance tagging on every training example is the foundation. Without it you cannot scope an investigation when anomalous model behavior is discovered.
- 03Influence function approximations can identify which training examples had the largest effect on a given output, making targeted poisoning locatable after the fact.
- 04Continual learning pipelines and RLHF with low trust labelers are the highest risk surfaces. Treat them as adversarial input channels by default.
Training data poisoning is the supply chain attack of the ML world. An adversary who can influence what data enters a training corpus can influence what behavior emerges from the resulting model. The effect can be subtle, a slight degradation in safety alignment, a backdoor triggered by a rare token pattern, or a systematic bias toward a desired output class, and it can persist for months or years before discovery.
The attack surface includes every point where external data enters the training pipeline. Web scraped corpora, user provided content in continual learning setups, RLHF preference labels from contractors or crowdworkers, and partner provided datasets for domain adaptation are all plausible injection points.
Upstream Controls that Matter
Defenses that operate after training has completed are reactive. The controls that actually prevent poisoned models from reaching production operate on the data before the training job starts.
- 01Provenance tags on every training example recording source, collection date, collector identity, and consent basis
- 02Contributor level anomaly detection flagging any source that drifts from its baseline content distribution
- 03Human review queues for high influence examples identified by influence function approximation
- 04Quarantine period for newly ingested web or user generated data before it enters a training batch
- 05Cryptographic hash pinning on approved dataset snapshots used for production training runs
The minimum viable poisoning defense for a team starting from zero.
If you can only do one thing this quarter, implement provenance tagging. Without it you cannot scope any investigation. Every training example must carry a source identifier, a collection timestamp, and an ingestion method. That metadata is the prerequisite for every other defense on this list.
- ▸Add source, timestamp, and collector fields to every training record
- ▸Enforce the schema at ingestion time, not retroactively
- ▸Store provenance metadata in a tamper evident log separate from the training corpus
Detecting Poisoning in Active Pipelines
Even with strong upstream controls, a sophisticated adversary may find a way to insert poisoned examples. Detection must therefore run inside the pipeline as well as upstream of it.
Statistical outlier detection on training batches catches examples that are semantically distant from the expected distribution for their label or domain. Influence function approximations, which estimate how much each training example moved the model parameters, can identify a small set of high influence examples worth manual review before the training job is promoted to production.
/Poisoning Detection Methods by Pipeline Stage
| Stage | Method | What it catches | Cost |
|---|---|---|---|
| Pre ingestion | Content distribution drift per source | Shifted labeler behavior or compromised feed | Low |
| Batch assembly | Embedding based outlier scoring | Semantically anomalous examples | Medium |
| Post training | Influence function approximation | High leverage poisoned examples | High |
| Post training | Behavioral regression on canary probes | Backdoor triggers and alignment drift | Medium |
RLHF and Continual Learning as Adversarial Channels
Reinforcement learning from human feedback introduces a particularly dangerous attack surface. If a subset of preference labelers can be compromised or are themselves adversarial, they can systematically push the model toward a desired behavior while remaining undetectable in aggregate statistics.
The defense is redundancy and audit. Each preference pair should be labeled by multiple independent workers and the agreement rate monitored at the labeler level, not just the batch level. A labeler whose agreement with the majority drops below a threshold across a specific topic area is a signal worth investigating.
Continual learning from live user interactions requires the same skepticism applied to any attacker controlled input channel. Gate every user generated training signal through a review step before it enters a training batch. The operational cost is real, but it is far less than retraining after a poisoning campaign is discovered.
Poisoning in RLHF is hard to see in aggregate metrics.
A poisoning campaign targeting RLHF may show no anomaly in aggregate labeler agreement rates if the adversarial labelers represent a small fraction of total volume. Detection requires per labeler, per topic analysis, not just overall statistics. Build that granularity into your labeler monitoring before you have a reason to need it.
Metrics and Closing Actions
The metrics to track are provenance coverage across the active training corpus, the rate of examples flagged by outlier detection and subsequently reviewed, the per labeler agreement rate for RLHF pipelines, and the canary probe pass rate across behavioral regression tests after each training run.
The closing action for this quarter is to run an influence function analysis against your most recently promoted model and inspect the top one hundred highest influence examples. In our experience, that review surfaces at least one data quality issue in nearly every corpus. It may not be adversarial, but it will be instructive.

