Alexa Cybersecurity
Back to Field Notes
AI & Adversarial ML/Field Note

A Ninety Day AI Security Roadmap for Security Teams Starting from Zero

The hardest part of AI security is not the technology. It is knowing where to start. This ninety day roadmap gives security teams a prioritized sequence of actions that builds a defensible AI security baseline without requiring a large dedicated team or a mature AI security program.

Author

Lin Chen

Head of AI Security Research

Published

June 9, 2026

Read

11 min

Share
AI-generated illustration of a banking data center
AI-generated illustration of a banking data center
Key Takeaways
  • 01The first thirty days should focus entirely on inventory and classification. You cannot prioritize controls for systems you have not catalogued. An incomplete inventory is a bigger risk than a missing control.
  • 02Days thirty through sixty are about closing the highest priority control gaps identified in the inventory. These are almost always the same three things. Overprivileged AI identities, absent inference log PII handling, and missing human approval gates for irreversible agent actions.
  • 03Days sixty through ninety build the measurement and governance layer that makes the program sustainable. Without metrics and a review cadence, AI security posture will drift as the AI deployment landscape evolves.
  • 04A ninety day roadmap is a starting point, not a destination. The AI deployment landscape is changing fast enough that the inventory and risk assessment need to be living documents refreshed at least quarterly.

Most security teams are encountering AI systems in production before they have built the governance, tooling, or expertise to manage them. This is not a failure of the security team. It is the natural consequence of AI deployment timelines running ahead of AI security program timelines in nearly every organization that has adopted AI capabilities in the last two years.

The ninety day roadmap in this article is designed for that starting condition. It does not assume an existing AI security program, a dedicated AI security team, or mature AI security tooling. It assumes a security team with existing competencies in identity, application security, and security operations, and it maps AI security work to those existing competencies wherever possible.

Days Zero Through Thirty. Inventory and Classification

The first thirty days have one goal. Know what AI systems are deployed in your environment, what data they process, and what actions they can take. Without this inventory, every subsequent prioritization decision is based on assumption rather than evidence.

The inventory exercise should cover four dimensions for each AI system. Input sources and whether they include untrusted or external content. Output types and whether outputs include recommendations or instructions that downstream systems act on. Authority level and what the system can do directly or indirectly. Data classification of training data, inference inputs, and logged outputs.

/MONDAY_PLAYBOOK

The inventory artifact is the most valuable output of the first thirty days.

Complete this for every AI system in production and every AI system in the development pipeline scheduled for production in the next ninety days. The completed inventory will almost certainly surface two or three high severity findings that you can address immediately. Treat finding those items as the success criterion for the first month, not the completion of the inventory itself.

  • ▸Interview product and engineering leads for each AI system
  • ▸Review architecture diagrams and data flow documents
  • ▸Test each system to verify the documented input and output behavior
  • ▸Classify each system by authority level and data classification

/AI System Inventory Template

DimensionQuestions to answerWhy it matters
Input sourcesDoes the system process untrusted external content such as user input, web content, or documents?Determines prompt injection exposure
Output typesDo outputs drive automated actions or human decisions with significant consequences?Determines excessive agency risk
Authority levelWhat systems and data can the AI access and modify?Determines blast radius of a security failure
Data classificationWhat is the highest classification of data in training, inference, and logs?Determines privacy and compliance exposure
Human oversightAre irreversible actions gated by human approval?Determines autonomous action risk

Days Thirty Through Sixty. Closing the Priority Gaps

The inventory will surface gaps. The second thirty days are about closing the highest priority ones. In our experience working with organizations at various AI security maturity levels, the highest priority gaps are almost always the same three categories.

The first is overprivileged AI identities. AI systems running under broad service account permissions that are not required by their defined function. This is a straightforward identity hygiene issue that is tractable within the thirty day window.

The second is absent inference log PII handling. AI systems whose logging pipelines retain prompt and completion text without PII redaction. This is a data governance issue that requires a redaction pipeline and a retention policy adjustment.

The third is missing human approval gates for irreversible agent actions. Agentic AI systems that can take consequential actions without human confirmation. This requires either an architectural change to add the approval gate or a temporary operational control that restricts the system to reversible actions until the gate is in place.

  • 01Scope all AI service account identities to the minimum permissions required by their defined function
  • 02Deploy PII detection and redaction in inference logging pipelines before log data is written to persistent storage
  • 03Implement human approval gates for every irreversible action type in agentic AI systems
  • 04Add output restriction controls to any AI API that currently returns full probability distributions or logprobs to external callers
  • 05Document the data classification policy for coding assistant use and communicate it to development teams
THIRTY DAY TARGETS

Measurable outcomes for the second month.

By the end of day sixty you should be able to report three numbers. The fraction of AI service identities with permissions scoped to defined function, the fraction of AI systems with inference log PII handling in place, and the fraction of irreversible agentic actions with human approval gates. If any of these fractions is below fifty percent, reprioritize the remaining work in that category before moving to the governance layer.

Days Sixty Through Ninety. Measurement and Governance

The third thirty days build the measurement and governance layer that makes the program sustainable. Controls implemented without measurement drift. An AI security program without a governance structure does not survive the first major AI deployment that bypasses it.

The governance layer has three components. First, a defined set of metrics reviewed on a regular cadence by security leadership. Second, a launch readiness review process for new AI systems that applies the inventory dimensions as a gate before production deployment. Third, a quarterly review that refreshes the inventory and reprioritizes controls based on new deployments and new threat intelligence.

  1. 01Define the AI security metrics dashboard covering tool permission coverage, irreversible action approval rate, inference log PII handling coverage, and extraction anomaly alert rate. Review these metrics monthly.
  2. 02Stand up the AI launch readiness review process. Any AI system scheduled for production must complete the inventory dimensions as a launch gate. A system that cannot answer the five inventory questions does not reach production.
  3. 03Schedule the first quarterly inventory refresh for ninety days after the current inventory is complete. At each refresh, add any new AI systems deployed in the quarter and reassess the priority gaps based on the updated inventory.
  4. 04Conduct a tabletop exercise using the highest authority AI system in production as the scenario subject. Walk through a successful prompt injection attack, an extraction campaign, and an autonomous action failure. The gaps surfaced will be the inputs to the next quarter's roadmap.
/INSIGHT

The launch readiness review is the control that prevents the inventory from becoming stale.

Without a production gate, new AI systems will be deployed faster than any periodic inventory refresh can track them. A launch readiness review that requires completing the inventory dimensions before deployment means the inventory stays current automatically. It also surfaces security decisions that need to be made at design time rather than retrofitted under production pressure.

What Comes After Ninety Days

A ninety day roadmap builds a foundation. After ninety days you should have a complete inventory, the three highest priority control gaps closed, and a governance structure that prevents drift. That is a defensible baseline, not a mature program.

The work after ninety days deepens each control layer. Adversarial robustness evaluation for production AI models. Formal privacy impact assessments with AI specific questions for all high sensitivity systems. Red team exercises against the agentic systems with the highest authority. Supply chain integrity controls for model artifacts. Differential privacy evaluation for systems trained on high sensitivity personal data.

The AI security landscape is evolving quickly enough that the quarterly inventory refresh is not optional. New AI systems will be deployed, new attack techniques will be published, and the threat model for existing systems will change as their capabilities and integrations expand. The governance structure built in the first ninety days is the mechanism that keeps the program calibrated to that moving target.

#AI Security#Security Roadmap#CISO#AI Governance#Security Program

/WRITTEN_BY

Lin Chen

Head of AI Security Research · Alexa Cybersecurity