Alexa Cybersecurity
Back to Field Notes
AI & Adversarial ML/Field Note

Token and Cost Abuse Defense for LLM Platforms

Token abuse does not trigger malware alerts or authentication failures. It shows up in a cloud bill. By the time finance escalates, the damage is already done. The defense requires treating token consumption as a security metric, not just a cost metric.

Author

Lin Chen

Head of AI Security Research

Published

May 29, 2026

Read

7 min

Share
AI-generated illustration of a banking data center
AI-generated illustration of a banking data center
Key Takeaways
  • 01Token abuse attacks fall into three categories. Direct amplification through large inputs, recursive agent loops, and cross tenant budget manipulation. Each requires a different control.
  • 02Hard per tenant per day token caps enforced at the gateway layer are the most effective single control. Soft warnings without hard limits are insufficient.
  • 03Per conversation step limits are the primary defense against recursive agent loops. Agents that exceed the step limit should terminate and trigger an alert, not continue with degraded functionality.
  • 04Real time token velocity monitoring with anomaly detection is the detection layer. A tenant whose token velocity increases by more than 3x within a one hour window should trigger an alert, not just a metric update.

Token and cost abuse occupies a unique position in the AI security threat landscape. It is financially damaging, operationally disruptive, and often invisible to security teams until the damage is already done. Traditional security controls do not detect it. Firewall rules, EDR agents, and network anomaly detectors have no visibility into token consumption patterns.

The enterprises that have been most affected report a consistent pattern. A new AI feature launches, the token budget is either not configured or set too loosely, and a triggerable amplification path is discovered by an external party or by an internal team running an unintended workload. The financial impact ranges from uncomfortable to significant depending on the model pricing tier and the duration before detection.

The three token abuse attack categories

/TOKEN_ABUSE_CATEGORIES

CategoryMechanismDefense
Direct input amplificationAttacker sends inputs with maximally large context windows, forcing high input token consumption per requestPer request input token ceiling enforced at the gateway before forwarding to the model
Recursive agent loopAgent enters a loop of tool calls that generate new tool calls, consuming tokens at each stepPer conversation step limit. Agent terminates and fires alert when step limit is reached.
Cross tenant budget manipulationAttacker identifies a path to consume tokens charged to another tenant, potentially depleting their budgetTenant identity bound to all token accounting at the gateway. Cross tenant charging must be architecturally impossible.

Concrete prevention controls

Prevention controls for token abuse must be enforced outside the model and outside the application code. Both are manipulable. The gateway or orchestration layer is the correct enforcement point.

  • 01Per request input token ceiling. Reject any request whose input token count exceeds the configured ceiling before forwarding to the model. The ceiling should be set to the 99th percentile of normal usage plus a reasonable buffer.
  • 02Per conversation step limit. Agents should have a hard step limit, typically between 8 and 20 steps depending on the use case, after which the conversation terminates. The application layer should not be able to override this limit.
  • 03Per tenant per day token budget. Hard cap enforced at the gateway. When the budget is exhausted, requests are rejected with a clear error, not silently degraded.
  • 04Per feature token allocation. Each AI feature should have an allocated share of the per tenant budget. A single feature cannot consume the entire organizational budget.
BUDGET SIZING

Start conservative and expand based on observed usage.

Set initial per tenant token budgets at 3x the observed 14 day maximum from baseline traffic. Review monthly for the first quarter and adjust based on legitimate usage growth. A budget that is too tight will surface quickly in user facing errors. A budget that is too loose will not protect against amplification attacks.

Detection metrics for token abuse

  • 01Token velocity anomaly rate. Number of per tenant alerts where the token velocity exceeded 3x the 14 day rolling baseline. Target is near zero with appropriate budgets in place.
  • 02Step limit termination rate. Number of agent conversations terminated by the step limit per day. A sudden increase indicates a new agent loop path that needs to be investigated.
  • 03Budget exhaustion rate. Number of tenants that hit their daily hard cap. Exhaustion should be rare. Frequent exhaustion indicates either a budget that is too low or a legitimate usage pattern that needs a budget review.
  • 04Mean time to detection. Time from the start of a token abuse event to the first alert. Target is under 5 minutes with real time velocity monitoring.

Treating token consumption as a security signal

The organizational change required for effective token abuse defense is treating token consumption as a security metric alongside cost metrics. FinOps dashboards show cost anomalies but rarely feed into security monitoring pipelines with alert and escalation processes. Security dashboards track authentication failures and policy violations but rarely show per tenant token velocity.

The integration point is the AI gateway or orchestration layer, which has visibility into both. Token velocity anomalies should generate SIEM alerts and follow the same escalation process as authentication anomalies. When an on call security analyst receives a token velocity alert, the investigation should include reviewing the conversation log, identifying the input that triggered the amplification, and determining whether the path can be eliminated or rate limited.

#Token Abuse#Cost Security#LLM#DoS Defense#Security Controls

/WRITTEN_BY

Lin Chen

Head of AI Security Research · Alexa Cybersecurity