
- 01A minimum viable LLM audit log must capture the full input context, model version, tenant identity, tool calls with parameters and return values, and the complete model response for every interaction.
- 02Logs that capture only the final response cannot support incident reconstruction, compliance audits, or model output attribution.
- 03Tamper evident log storage is a requirement, not an enhancement. If a model produces output that is used in a regulatory decision, the log that proves what the model received must be defensible.
- 04Log retention for AI interactions should follow the same schedule as application access logs, with a minimum of 12 months for regulated industries.
Audit logging for LLM systems is not a solved problem in most enterprises. Classic application logging captures user actions against a known data model. LLM interactions are different. The same request can produce different outputs depending on the context window, the model version, the retrieval results, and the order of tool calls. Reconstructing what happened requires capturing all of these, not just the final response.
Security and compliance teams that have tried to investigate AI incidents without adequate logs consistently report the same finding. Without the full input context, they cannot determine whether the model was manipulated, malfunctioned, or behaved as designed.
The minimum viable LLM audit log schema
/LLM_AUDIT_LOG_SCHEMA
| Field | Type | Required | Purpose |
|---|---|---|---|
| event_id | UUID | Yes | Unique identifier for the interaction, used for reconstruction queries |
| tenant_id | String | Yes | Tenant or organizational unit that owns this interaction |
| user_id | String | Yes | Authenticated user identity driving the request |
| session_id | String | Yes | Groups all turns in a conversational session |
| model_id | String | Yes | Model name and version exactly as deployed |
| input_context | JSON array | Yes | Full message array including system prompt, tool definitions, and prior turns |
| tool_calls | JSON array | Yes | Each tool call with function name, parameters, and return value |
| output | String | Yes | Complete model response text |
| timestamp_utc | ISO 8601 | Yes | Time the request reached the inference layer |
| latency_ms | Integer | Yes | Inference latency, useful for anomaly detection |
| token_input | Integer | Yes | Input token count for cost and volume tracking |
| token_output | Integer | Yes | Output token count |
| retrieval_chunks | JSON array | No | RAG chunks returned to the model, with source identifiers |
| policy_decisions | JSON array | No | Gateway or guardrail decisions applied to this request |
Operational requirements for tamper evident log storage
An LLM audit log is only as useful as its chain of custody. If an audit or investigation depends on proving what the model received and produced, the log must be stored in a way that makes alteration detectable. The requirements are the same as those for financial transaction logs.
- 01Write once storage. Log records must be written to a destination that does not allow modification or deletion by the application layer. Object storage with object lock enabled satisfies this requirement.
- 02Cryptographic integrity. Each log record should be signed or chained so that deletion of a record is detectable. A hash chain over time ordered records is the minimum viable approach.
- 03Separate access control. The identity that operates the AI application must not have delete or overwrite permissions on the audit log destination.
- 04Replication. Logs must be replicated to at least two geographic regions so that a single infrastructure failure does not destroy the audit record.
Why input logging is more important than output logging.
Model outputs are visible in application interfaces and can often be reconstructed from user reports. What cannot be reconstructed without logs is the exact input context the model received. That context is what determines whether the model behaved correctly, was manipulated, or was given information it should not have had access to.
Metrics for audit log quality
- 01Schema completeness rate. Percentage of LLM interactions logged with all required fields present. Target is 100 percent.
- 02Log availability during incidents. Percentage of AI incidents where the full audit log was retrievable within 30 minutes. Target is 100 percent.
- 03Tamper detection test pass rate. Monthly test that attempts to modify or delete a log record and verifies that detection fires. Target is 100 percent.
- 04Retention compliance rate. Percentage of AI audit logs retained for the required minimum period. Target is 100 percent.
Connecting audit logs to compliance programs
AI audit logs are increasingly referenced in regulatory frameworks that govern automated decision making. If your organization operates in a regulated industry, the audit log for an AI interaction may need to be produced in response to a subject access request, a regulatory examination, or a litigation hold. Teams that have designed their log schema with this requirement in mind consistently report faster response times and lower external counsel costs when these requests arrive.
The practical action is to include AI audit logs in the same data inventory and retention schedule that already covers application access logs and financial transaction records. Map each log field to the regulatory requirement it satisfies. This exercise also surfaces gaps in the log schema, because regulatory requirements tend to be more specific about what must be captured than internal security requirements.
