Alexa Cybersecurity
Back to Field Notes
AI & Adversarial ML/Field Note

AI Threat Modeling for Production Systems

STRIDE still works for AI systems if you extend it with three AI specific threat categories. The real challenge is drawing the data flow diagram correctly when the model itself is a processing node that can be manipulated through its inputs.

Author

Lin Chen

Head of AI Security Research

Published

May 12, 2026

Read

10 min

Share
AI-generated illustration of a banking data center
AI-generated illustration of a banking data center
Key Takeaways
  • 01STRIDE maps cleanly to AI threat categories if you add three extensions. Prompt injection maps to Tampering, model output misuse maps to Elevation of Privilege, and supply chain threats map to Spoofing of the model identity itself.
  • 02The data flow diagram for an AI system must treat the model as an active processing node, not a passive transformation. Attacker controlled data can alter the model's behavior, not just its inputs.
  • 03Agent systems require a second threat model layer covering the tool permission graph. Each edge in that graph is a potential privilege escalation path.
  • 04Threat modeling sessions for AI systems should include the ML engineer, not just the security engineer. Model behavior under adversarial input is empirical, not purely architectural.

Threat modeling is one of the highest leverage security activities an engineering team can run, and it transfers directly to AI systems with modest adaptation. The core question remains the same. For each component in this system, who can interfere with it, what can they do, and what is the worst case outcome?

The adaptation required for AI is that the model itself is a component that behaves differently depending on its inputs. This makes the threat model more dynamic than a traditional service threat model, but it does not make it fundamentally harder. It makes the data flow diagram more important.

Extending STRIDE for AI systems

STRIDE, Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, and Elevation of Privilege, covers the AI threat landscape well with the following extensions.

/STRIDE Extended for AI Systems

STRIDE CategoryTraditional ExampleAI Extension
SpoofingForging an identity tokenProviding a fine tune or embedding that impersonates a trusted model output
TamperingModifying a database recordPrompt injection through retrieved content, training data poisoning
RepudiationDeleting audit logsPrompt logs not captured, making agent actions forensically invisible
Information DisclosureReading another user recordSystem prompt leakage, RAG cross tenant retrieval, embedding inversion
Denial of ServiceFlooding a service endpointToken amplification loops, recursive agent calls burning compute budget
Elevation of PrivilegeExploiting a SUID binaryConvincing a model to invoke a tool the requesting user is not authorized for

Drawing the AI data flow diagram correctly

The most common mistake in AI threat modeling is drawing the model as a black box with inputs and outputs. The correct diagram treats the model as an active processing node whose behavior is a function of all content in its context window, including content supplied by external sources.

For a RAG system, the data flow diagram should show four distinct trust zones. The user request zone, the retrieval zone containing external documents, the model context assembly zone where retrieved content and the user request are combined, and the tool execution zone. Each zone boundary is a trust boundary requiring an explicit control.

/MONDAY_PLAYBOOK

Threat model session agenda for an AI system

Run a two hour session with the ML engineer, backend engineer, and security engineer. Start with the data flow diagram. Spend 30 minutes just getting the diagram right. Then work STRIDE systematically across each trust boundary. Assign a severity and an owner to each finding before the session ends.

  • ▸Draw the data flow diagram including all external content sources
  • ▸Mark every trust boundary explicitly
  • ▸Apply STRIDE at each trust boundary crossing
  • ▸For each agent tool, ask what is the worst case irreversible action
  • ▸Assign severity and owner before closing

The tool permission graph as a second threat model

Agent systems require a second threat model that treats the tool permission graph as the primary attack surface. Each tool grant is an edge in a directed graph from the agent to a capability. The threat model asks, for each edge, whether an attacker who can influence the model's context window could traverse that edge without authorization.

Draw the tool permission graph separately from the main data flow diagram. List every tool available to the agent, the maximum authority that tool can exercise, and whether that authority is scoped to the requesting user or is ambient. Unscoped ambient authority, such as a tool that can send email as a service account rather than as the requesting user, is a near universal finding in first generation agent deployments.

/CAUTION

Ambient authority is the most common finding

In the majority of agent threat models we conduct, at least one tool operates with ambient organizational authority rather than per user delegated authority. The tool can send messages, create records, or call external APIs using credentials that belong to the service rather than the requesting user. This means a successful prompt injection can act as any user in the system.

Tracking findings and measuring threat model maturity

A threat model that produces findings but no tracking is a compliance exercise. Each finding should become a tracked item with a severity, a named owner, and a target remediation date. High severity findings, those in the Tampering or Elevation of Privilege categories, should be gated items for the next production deployment.

Measure threat model maturity at the program level through three metrics. Coverage, the percentage of AI systems in the inventory with a completed threat model. Freshness, the percentage of threat models updated within the past 12 months or after a significant architectural change. Finding closure rate, the percentage of high severity findings closed within 90 days of identification.

/AI Threat Model Maturity Metrics

MetricLevel 1 TargetLevel 2 Target
Threat model coverageAll high authority AI systemsAll AI systems in inventory
Threat model freshnessUpdated after major changesAnnual review for all systems
High severity finding closureWithin 180 daysWithin 90 days
Tool permission graph documentedAll agentic systemsAll AI systems with any tool access
#Threat Modeling#AI Security#STRIDE#Agentic AI#Risk Management

/WRITTEN_BY

Lin Chen

Head of AI Security Research · Alexa Cybersecurity