
- 01AI red team exercises should test three threat layers. Prompt injection and instruction following, tool permission exploitation in agentic systems, and multiturn behavioral manipulation that exploits model tendencies over a conversation.
- 02Effective AI red teams include both security engineers and subject matter experts in the domain the AI system serves. Domain experts find the business logic manipulation attacks that purely technical red teamers miss.
- 03Findings from AI red team exercises should be classified by the architectural change required to remediate them, not just by severity. A finding that requires an input filter is categorically less durable than one that requires structural isolation.
- 04AI red team results are probabilistic. A finding that succeeds 20 percent of the time in testing still represents a meaningful security risk and should be scored accordingly.
Traditional penetration testing operates against systems that behave deterministically. Given a specific request, a vulnerable system returns a predictable response. Exploits are reproducible, scan detectable, and verifiable through automated tooling. AI systems break this assumption. The same prompt may succeed 30 percent of the time and fail 70 percent. A jailbreak that works in one context window composition fails in another. A multiturn manipulation requires a specific sequence of interactions that no scanner will generate.
AI red teaming requires a human driven, iterative methodology that treats the model as a probabilistic system and measures attack success rates over many trials. This article covers the planning, execution, and reporting approach that produces actionable findings.
Scoping the AI red team exercise
Scope definition for an AI red team exercise starts with the AI asset inventory. For each system in scope, identify the threat model tier. Tier one systems have high authority tools, process sensitive data, and are exposed to external users. Tier two systems have read only tools or process internal non sensitive data. Tier three systems are internal assistants with no tool access.
Prioritize tier one systems for the initial exercise. A limited scope exercise that tests a single high authority agent thoroughly produces more actionable findings than a broad but shallow exercise across many systems.
/AI Red Team Scope Prioritization
| Tier | Characteristics | Red Team Priority |
|---|---|---|
| Tier 1 | High authority tools, sensitive data, external users | Mandatory for initial exercise |
| Tier 2 | Read only tools or limited data, mixed user base | Include in first expansion cycle |
| Tier 3 | No tools, internal users, non sensitive domain | Schedule for annual coverage |
Staffing the AI red team
The most effective AI red teams combine three skill profiles. A security engineer with experience in prompt injection and AI attack techniques provides the technical methodology. A domain expert who understands the business context the AI system serves identifies manipulation attacks that exploit business logic. A creative generalist, often a technical writer or a linguist, finds social engineering and framing attacks that neither the security engineer nor the domain expert would generate alone.
For an agentic system with financial or operational authority, a fourth profile is valuable. A subject matter expert in the domain the tools operate in, such as a financial analyst for a payment handling agent, can identify the specific sequences of tool calls that would cause the most damage and work backward to find the prompts that produce them.
Domain experts find what scanners and security engineers miss
In a red team exercise against a customer service agent with order management tool access, the most impactful findings came from a customer experience specialist, not the security engineers. She identified a pattern of legitimate sounding requests that, across a multiturn conversation, convinced the model to apply refunds beyond its stated policy. The attack was not a jailbreak. It was a business logic manipulation that required domain knowledge to design.
Exercise structure and testing methodology
Structure the exercise across three attack layers. The first layer covers direct prompt injection through all input surfaces the system exposes to external content. The second layer covers tool permission exploitation, attempting to use legitimate tool capabilities in ways that exceed the intended scope. The third layer covers multiturn behavioral manipulation, where the red team attempts to shift the model behavior over a conversation rather than through a single input.
For each attack attempt, record the full context of the attempt including the inputs, the model output, whether the attempt succeeded, and the number of trials required. Success rate over a defined number of trials is a more useful metric than binary success or failure, because it captures the probabilistic nature of model behavior.
- 01Direct prompt injection through all external input surfaces including user messages, retrieved documents, and tool return values.
- 02Indirect prompt injection through environmental content the system may retrieve, such as web pages, documents, and database records.
- 03Tool permission exploitation including attempts to call tools on behalf of users who did not authorize them.
- 04Multiturn behavioral manipulation including gradual escalation, persona anchoring, and false context establishment.
- 05Resource and cost exhaustion attacks including token amplification and recursive loops.
Reporting AI red team findings
AI red team reports should organize findings by the architectural change required to remediate them. A finding that is remediated by adding an input filter is a different category than a finding that requires structural isolation of the model from tool access. The first may be bypassed by a novel attack variant. The second is durable against a broad class of attacks.
For each finding, report the attack layer, the success rate over trials, the worst case business impact, the architectural remediation required, and the implementation effort. Present findings to engineering leadership with the architectural remediation as the primary recommendation, not the severity score alone.
Report the overall exercise to the CISO and board using three summary metrics. The number of tier one systems tested, the number of findings requiring architectural remediation, and the percentage of findings with a committed remediation date. These three numbers tell the board whether the AI security program is functioning.
Minimum viable AI red team exercise for a first time program
If your organization has not yet run an AI red team exercise, start with the single highest authority agent in production. Spend one week with two people. Bring a security engineer and a domain expert. Focus exclusively on tool permission exploitation and indirect prompt injection through retrieved content. Report two findings per person per day. A two person, five day exercise against a single system will produce more actionable findings than a broad scanning exercise across all systems.
- ▸Select the single highest authority agent for initial scope.
- ▸Staff with a security engineer and a domain expert for five days.
- ▸Focus on tool exploitation and retrieved content injection.
- ▸Record success rates, not just binary outcomes.
- ▸Report findings by architectural remediation required.
/AI Red Team Finding Classification
| Finding Type | Remediation Category | Durability of Fix |
|---|---|---|
| Direct prompt injection via user input | Input handling and validation | Medium, depends on coverage |
| Indirect injection via retrieved content | Structural isolation of retrieval from tool access | High if architecture is correct |
| Tool permission exploitation | Scope reduction and per user credential enforcement | High if scoping is complete |
| Multiturn behavioral manipulation | Conversation state monitoring and hard limits | Medium, depends on limits chosen |
| Cross tenant data access through RAG | Mandatory tenant filter enforcement | High if filter is mandatory |
