Alexa Cybersecurity
Back to Field Notes
AI & Adversarial ML/Field Note

Prompt Injection Architecture Defense at the System Level

Input filters lose to novel injection variants. The defenses that hold are structural. Separating untrusted text processing from tool enabled reasoning, enforcing schema boundaries between those layers, and gating irreversible actions outside the model are the three patterns that reduce prompt injection from a critical risk to a contained one.

Author

Lin Chen

Head of AI Security Research

Published

May 13, 2026

Read

11 min

Share
AI-generated illustration of a banking data center
AI-generated illustration of a banking data center
Key Takeaways
  • 01Detection based defenses against prompt injection are a soft mitigation. They reduce attack success rates but cannot reach near 100 percent efficacy against a motivated adversary with query access.
  • 02Structural isolation, running untrusted text through a view only model and passing only schema validated output to a tool enabled planner, eliminates the direct path from attacker controlled text to tool execution.
  • 03Schema design is a security control. A tightly constrained output schema for the view only model limits what an attacker can communicate to the planner even if the injection partially succeeds.
  • 04Irreversible action gates outside the model are the final defense layer. A policy engine or human approval step that operates on structured action requests, not raw model output, closes the loop.

Prompt injection is the class of attacks where an attacker embeds instructions in content that the model processes, causing the model to follow those instructions instead of or in addition to its intended behavior. The content might be a user supplied message, a retrieved document, a web page fetched by a tool, or a field value returned from a database query.

The security literature on prompt injection defense has converged on a clear conclusion. Any defense that works by classifying whether a given string is a prompt injection will have a nonzero false negative rate at scale. The attacks are too diverse, too context dependent, and too easy to iterate on for a classifier to achieve the reliability that a security control requires.

Why detection caps out and what to do instead

Prompt injection classifiers face a fundamental challenge. The boundary between legitimate instructions and malicious instructions is not a fixed syntactic property. It depends on context, intent, and the capabilities available to the model. Classifiers trained on known attack patterns generalize poorly to novel phrasing, multilingual variants, and indirect injections embedded in structured data like JSON or Markdown.

The architectural alternative does not try to answer whether a given string is malicious. It eliminates the structural precondition that makes prompt injection dangerous, which is the combination of attacker controlled text and model tool authority in the same context window.

/INSIGHT

The structural precondition for prompt injection impact

Prompt injection is dangerous when three conditions hold simultaneously. The model receives attacker controlled text. The model has authority to take consequential actions. The model treats attacker controlled text as trustworthy instructions. Remove any one of those three conditions and the attack loses its impact. Structural defenses target the first two. Detection based defenses target the third, which is why they are harder.

The isolation architecture

The core pattern is a two stage pipeline. Stage one is a view only model that receives all untrusted content, has no tool access, and produces a structured output constrained by a schema. Stage two is a planner model that receives only the structured output from stage one and has access to tools.

The schema between the two stages is the critical security boundary. It should be as narrow as possible while remaining useful. Each field in the schema should have a defined type, a maximum length, and a semantic constraint. Free text fields in the schema are the primary residual injection surface and should be minimized.

  1. 01Define a named output type for the view only model with exactly the fields the planner needs. No additional fields should be present.
  2. 02Use an enum field for the user intent so the planner receives one of a fixed set of values such as query, create, update, delete, or other. Enum fields deny attackers any free vocabulary to exploit.
  3. 03Model any extracted entities as a typed array where each entry carries a category drawn from a fixed set such as person, document, date, or amount, plus a value field capped at 128 characters.
  4. 04Include a risk flags array whose members are drawn from a fixed vocabulary such as instruction in content, unusual language, or encoding anomaly. The planner reads these flags before deciding whether to proceed.
  5. 05Add a summary field of type string and enforce a hard maximum of 256 characters. This is the only free text field and its short length limits how much attacker controlled content can reach the planner.
  6. 06Apply a schema validator such as Zod or Pydantic between the two model stages. Any output from the view only model that does not match the type definition must be rejected before the planner sees it.

Irreversible action gates outside the model

Even with isolation architecture in place, a planner model operating on a well formed structured input can be manipulated through the content of that input. The final defense layer is an action gate that evaluates proposed tool calls against a policy before execution. This gate operates on the structured tool request, not on the model's reasoning trace.

For irreversible actions, including sending messages, writing to databases, making payments, or deploying code, the gate should require either explicit human approval or a policy engine evaluation against defined authorization rules. The policy engine approach is more scalable and auditable. OPA, Cedar, and similar policy as code tools integrate well with agent frameworks and produce structured audit records.

  • 01Define which tools are irreversible in the asset inventory. This list drives gate requirements.
  • 02Implement the gate outside the model runtime, not as a system prompt instruction.
  • 03Gate evaluations should produce structured audit records regardless of outcome.
  • 04Human approval gates for the highest authority actions provide a final check that no automated policy can fully replace.

Measuring the effectiveness of architectural defenses

Architectural defenses against prompt injection should be validated through adversarial testing, not just design review. Run periodic red team exercises against the isolation boundary using a library of known injection patterns as well as novel attempts. Measure the percentage of injection attempts that succeed in causing a tool call that was not intended by the legitimate user request.

Track three operational metrics quarterly. The rate of injection pattern detections at the view only stage, which confirms the stage is processing real world injection attempts. The rate of schema validation failures, which indicates either attacker activity or model reliability issues. The rate of action gate denials, which confirms the gate is active and catching unexpected tool call attempts.

/Prompt Injection Defense Metrics

MetricWhat It MeasuresTarget Posture
View only stage injection detectionsVolume of injection attempts reaching the systemTracked and trending, not zero
Schema validation failure rateIntegrity of the isolation boundaryBelow 1% of legitimate requests
Action gate denial rateGate coverage and attacker pressureAudited and reviewed weekly
Red team exercise success rateResidual injection riskZero successful tool call injections
#Prompt Injection#LLM Architecture#AI Security#Defense in Depth

/WRITTEN_BY

Lin Chen

Head of AI Security Research · Alexa Cybersecurity