Skip to Content

Data Poisoning: What Threat Hunters See That Governance Frameworks Miss

Brandon Tirado

AI data poisoning rarely starts in a training set. For most organizations, it starts in the vendor, package, or pipeline that delivers the model.

In practice, we spend more time looking upstream, at the vendors, packages, datasets, software components, and CI/CD pipelines that deliver AI capability. An attacker who compromises one of those dependencies may influence what an AI system trusts before the organization using it realizes anything has changed.

For many organizations, the immediate exposure sits in the supply chain and in the identities that can influence it.

For readers who want the foundations of agentic architecture first, our guide on AI Agents vs. Agentic Systems in Security Operations discusses the underlying concepts.

Why Data Poisoning Discussions Start in the Wrong Place

The textbook explanation of data poisoning attacks talks about bad data entering a training set. An attacker manipulates examples, the model learns from them, and its behavior changes as a result.

The conversation tends to focus on “AI slop” ending up in training data. But the real exposure for most organizations is in the systems and pipelines delivering these components, not the training data itself. Most organizations use frontier or open-source models rather than training their own, then build scaffolding on top of them.

MITRE ATLAS is indicative of the wider AI attack surface. The knowledge base includes the use of poisoning training data, the publication of poisoned data sets and models, and the publication of poisoned AI agent tools.

We also draw a clear line between poisoning and tampering. Poisoning actively manipulates the data used to shape a system. AI model tampering changes how data or model output gets interpreted without necessarily altering the training data itself.

The distinction affects the investigation. Label every unexpected change in AI behavior as poisoning, and threat hunters may start from the wrong evidence.

The Highest-Value Targets Right Now

We start with anything pulled from public libraries, third-party packages, off-the-shelf software, and quickly assembled or vibe-coded applications.

Every dependency introduces another trust decision. Packages change. Vendors get compromised. Components can arrive with behavior the organization never reviewed.

NIST’s AI Risk Management Framework asks for policies and monitoring of third-party software and data. This approach seems appropriate. It’s also where things end. Policies and monitoring confirm that the vendor was checked at the point of onboarding. However, they don’t provide much insight into whether the model or component behaves differently now compared to a few months ago.

That’s what the threat hunter focuses on: confirming that the policy exists or catching the exact moment when the system deviates from its normal baseline behavior.

We also pay particular attention to LLM-as-a-judge systems.

Organizations increasingly use one AI system to grade, check, or validate another AI system’s output. This can improve quality and support autonomous AI guardrails, but it also creates an authoritative decision point.

If an attacker compromises the system trusted to evaluate AI output, downstream controls may continue accepting decisions from a judge that can no longer be trusted. That makes LLM-as-a-judge poisoning a particularly useful target for an attacker.

Supply Chain or Insider Threat? A False Distinction

Many governance frameworks view insider risk as separate from external compromise. In an investigation, the route into the environment may be less important than what the attacker or insider gained access to.

External attackers often compromise vendors or packages, while insiders use their sanctioned access to influence datasets, components, or pipelines. Both actors can reach systems that build, approve, and deliver AI capabilities.

Knowing who can influence the pipeline is the starting point for assessing blast radius.

AI supply chain security depends heavily on identity, including non-human identity security. Teams need to know who can approve training data, modify evaluations, introduce models or components, and connect systems.

Without that information, any assessment of blast radius relies on assumptions alone.

What Poisoning Looks Like to a Hunter Who Isn’t Looking for It

A poisoned system does not have to look broken.

A hunter may first see a change in behavior. An evaluation starts passing output it previously rejected. A model responds differently under a familiar condition. A guardrail continues to run but stops producing the expected result.

Detection isn’t possible without a reliable picture of normal behavior.

Since most organizations use frontier models rather than training their own, the real place to hunt is the scaffolding and wrappers built around those models, wherever guardrails live (code repositories, evaluations).

Any deviation from a known-good evaluation result is the signal to chase.

The Months-Long Blind Spot, and Why There’s No Clean Playbook Yet

A poisoning event may not produce an obvious incident when it happens. The affected component can remain inside a workflow until a particular condition causes the altered behavior to surface.

Organizations need the ability to investigate by extracting away the model layer and focusing on intended outcomes, but a true incident response playbook for this specific scenario doesn't really exist yet.

It comes down to observability: the telemetry built around a system, and understanding where and how an intended outcome would manifest in production or client-facing infrastructure if something went wrong.

IBM’s 2025 breach research gives some context for the wider supply chain problem. Third-party vendor and supply chain compromise carried an average breach cost of $4.91 million in its study.

Agentic AI Turns One Compromised Node into a Compounding Problem

Agentic systems create a compounding effect, and that is where agentic AI security risk looks different from risk in a standalone model. They are a mix of disparate systems chained together, whether through hard-coded paths or the agentic layer calling the right things at the right times. If one node in that chain is compromised, the input or behavior passes from node to node to node, and the problem compounds.

That can make agentic systems more lucrative targets than a standalone large language model because the blast radius through dependencies is wider.

The dynamic has flipped. Historically, security teams pushed the business toward better practices and investment. Now the business is pushing security teams to adopt AI faster, with pressure from the board level to avoid “falling behind.” Security is on the other side of the table, trying not to block innovation while still figuring out how to enable it safely. It has created what amounts to a reverse fire drill.

Security teams need to address approved components, permission boundaries, evaluation baselines, and limits on autonomous action before deployment.

Building Detection from Scratch: The First Three Signals

Three things come first. Provenance verification maps where every dataset, model, and component came from, tracing the dependency chain step by step. Behavioral monitoring catches a downloaded model doing something unexpected at load time. And identity controls log who can touch or influence the model, who approves training data and samples, and who can build agentic systems on top of it, with least privilege applied throughout.

As complexity grows and more people become involved, there are more opportunities for things to fall through the cracks.

GreyMatter's Universal Translator normalizes telemetry from disparate tools into a common schema at ingest, mapping identity and behavioral signals from every connected environment into one investigative view. An analyst can then trace whether a change in model or evaluation behavior lines up with activity from a specific identity, component, or dependency.

That context does not prove poisoning happened, but it arms investigators with evidence to test whether changed behavior connects to something in the surrounding system.

Closing Perspective: Don’t Let This Eclipse the Real Question

We think AI data poisoning gets more attention than its current risk warrants for many organizations.

The current risk varies sharply by environment. Organizations building agentic systems have more reason to prioritize it because dependencies widen the potential blast radius. Most companies, at a macro level, are still consumers of models rather than builders of complex agentic systems.

For the average organization, the bigger concern is whether the partners and providers they rely on are asking the right questions and taking poisoning risk into account on their behalf. That may matter more than whether the organization itself is a direct target.

Internally, CISOs should be able to answer quickly and continuously who has the means to influence, push, or chain together systems built on top of AI. That is the starting point because it defines the blast radius. Without that understanding, risk decisions become a guessing game, with CISOs relying on what they have been told rather than demanding evidence.

If you’d like to see how identity and behavioral context can support security investigations across your environment, request a GreyMatter demo.

Learn How GreyMatter Agentic AI Scales Your Security Operations

GreyMatter is an agentic AI security operations platform with 6 agentic Teammates that use hundreds of agent skills and AI tools to work toward an objective, not just tasks.

GreyMatter dashboard active summary