Alexa Cybersecurity
Back to Field Notes
AI & Adversarial ML/Field Note

AI Red Team Planning for Enterprise AI Systems

AI red teaming is not traditional penetration testing applied to a new target. Model behavior under adversarial conditions is probabilistic, context dependent, and not reproducible through standard scanning. A well run AI red team exercise requires a structured methodology, a diverse operator team, and a reporting format that translates findings into architectural changes.

Author

Lin Chen

Head of AI Security Research

Published

May 20, 2026

Read

11 min

Share
AI-generated illustration of a banking data center
AI-generated illustration of a banking data center
Key Takeaways
  • 01AI red team exercises should test three threat layers. Prompt injection and instruction following, tool permission exploitation in agentic systems, and multiturn behavioral manipulation that exploits model tendencies over a conversation.
  • 02Effective AI red teams include both security engineers and subject matter experts in the domain the AI system serves. Domain experts find the business logic manipulation attacks that purely technical red teamers miss.
  • 03Findings from AI red team exercises should be classified by the architectural change required to remediate them, not just by severity. A finding that requires an input filter is categorically less durable than one that requires structural isolation.
  • 04AI red team results are probabilistic. A finding that succeeds 20 percent of the time in testing still represents a meaningful security risk and should be scored accordingly.

Traditional penetration testing operates against systems that behave deterministically. Given a specific request, a vulnerable system returns a predictable response. Exploits are reproducible, scan detectable, and verifiable through automated tooling. AI systems break this assumption. The same prompt may succeed 30 percent of the time and fail 70 percent. A jailbreak that works in one context window composition fails in another. A multiturn manipulation requires a specific sequence of interactions that no scanner will generate.

AI red teaming requires a human driven, iterative methodology that treats the model as a probabilistic system and measures attack success rates over many trials. This article covers the planning, execution, and reporting approach that produces actionable findings.

Scoping the AI red team exercise

Scope definition for an AI red team exercise starts with the AI asset inventory. For each system in scope, identify the threat model tier. Tier one systems have high authority tools, process sensitive data, and are exposed to external users. Tier two systems have read only tools or process internal non sensitive data. Tier three systems are internal assistants with no tool access.

Prioritize tier one systems for the initial exercise. A limited scope exercise that tests a single high authority agent thoroughly produces more actionable findings than a broad but shallow exercise across many systems.

/AI Red Team Scope Prioritization

TierCharacteristicsRed Team Priority
Tier 1High authority tools, sensitive data, external usersMandatory for initial exercise
Tier 2Read only tools or limited data, mixed user baseInclude in first expansion cycle
Tier 3No tools, internal users, non sensitive domainSchedule for annual coverage

Staffing the AI red team

The most effective AI red teams combine three skill profiles. A security engineer with experience in prompt injection and AI attack techniques provides the technical methodology. A domain expert who understands the business context the AI system serves identifies manipulation attacks that exploit business logic. A creative generalist, often a technical writer or a linguist, finds social engineering and framing attacks that neither the security engineer nor the domain expert would generate alone.

For an agentic system with financial or operational authority, a fourth profile is valuable. A subject matter expert in the domain the tools operate in, such as a financial analyst for a payment handling agent, can identify the specific sequences of tool calls that would cause the most damage and work backward to find the prompts that produce them.

Field note

Domain experts find what scanners and security engineers miss

In a red team exercise against a customer service agent with order management tool access, the most impactful findings came from a customer experience specialist, not the security engineers. She identified a pattern of legitimate sounding requests that, across a multiturn conversation, convinced the model to apply refunds beyond its stated policy. The attack was not a jailbreak. It was a business logic manipulation that required domain knowledge to design.

Exercise structure and testing methodology

Structure the exercise across three attack layers. The first layer covers direct prompt injection through all input surfaces the system exposes to external content. The second layer covers tool permission exploitation, attempting to use legitimate tool capabilities in ways that exceed the intended scope. The third layer covers multiturn behavioral manipulation, where the red team attempts to shift the model behavior over a conversation rather than through a single input.

For each attack attempt, record the full context of the attempt including the inputs, the model output, whether the attempt succeeded, and the number of trials required. Success rate over a defined number of trials is a more useful metric than binary success or failure, because it captures the probabilistic nature of model behavior.

  • 01Direct prompt injection through all external input surfaces including user messages, retrieved documents, and tool return values.
  • 02Indirect prompt injection through environmental content the system may retrieve, such as web pages, documents, and database records.
  • 03Tool permission exploitation including attempts to call tools on behalf of users who did not authorize them.
  • 04Multiturn behavioral manipulation including gradual escalation, persona anchoring, and false context establishment.
  • 05Resource and cost exhaustion attacks including token amplification and recursive loops.

Reporting AI red team findings

AI red team reports should organize findings by the architectural change required to remediate them. A finding that is remediated by adding an input filter is a different category than a finding that requires structural isolation of the model from tool access. The first may be bypassed by a novel attack variant. The second is durable against a broad class of attacks.

For each finding, report the attack layer, the success rate over trials, the worst case business impact, the architectural remediation required, and the implementation effort. Present findings to engineering leadership with the architectural remediation as the primary recommendation, not the severity score alone.

Report the overall exercise to the CISO and board using three summary metrics. The number of tier one systems tested, the number of findings requiring architectural remediation, and the percentage of findings with a committed remediation date. These three numbers tell the board whether the AI security program is functioning.

/MONDAY_PLAYBOOK

Minimum viable AI red team exercise for a first time program

If your organization has not yet run an AI red team exercise, start with the single highest authority agent in production. Spend one week with two people. Bring a security engineer and a domain expert. Focus exclusively on tool permission exploitation and indirect prompt injection through retrieved content. Report two findings per person per day. A two person, five day exercise against a single system will produce more actionable findings than a broad scanning exercise across all systems.

  • ▸Select the single highest authority agent for initial scope.
  • ▸Staff with a security engineer and a domain expert for five days.
  • ▸Focus on tool exploitation and retrieved content injection.
  • ▸Record success rates, not just binary outcomes.
  • ▸Report findings by architectural remediation required.

/AI Red Team Finding Classification

Finding TypeRemediation CategoryDurability of Fix
Direct prompt injection via user inputInput handling and validationMedium, depends on coverage
Indirect injection via retrieved contentStructural isolation of retrieval from tool accessHigh if architecture is correct
Tool permission exploitationScope reduction and per user credential enforcementHigh if scoping is complete
Multiturn behavioral manipulationConversation state monitoring and hard limitsMedium, depends on limits chosen
Cross tenant data access through RAGMandatory tenant filter enforcementHigh if filter is mandatory
#Red Team#AI Security#Adversarial Testing#Penetration Testing#AI Governance

/WRITTEN_BY

Lin Chen

Head of AI Security Research · Alexa Cybersecurity