When a CISO asks "what did our AI agents do last night?", the answer in most deployments today is a two-to-four-hour exercise in log archaeology. Someone pulls CloudTrail or GCP audit logs, filters for the service account credentials associated with the agent, correlates timestamps against the application logs, and assembles a partial picture that cannot definitively distinguish agent-initiated access from human-initiated access using the same credentials.
That is not an audit. That is manual reconstruction under time pressure, and its results are neither reliable nor repeatable. A real audit trail answers the question in minutes and produces output that is specific enough to satisfy a compliance reviewer, an insurance underwriter, or an incident response investigation.
This post describes the model we use at Arrakis to structure agent audit trails, why each element matters, and what the practical gaps look like in deployments that are trying to build this from general-purpose infrastructure logging tools.
What an Agent Audit Trail Must Answer
Before designing a logging architecture, it helps to enumerate the questions the audit trail needs to answer. Based on the compliance and incident response conversations we have had, the core questions are:
Which user request triggered the agent activity? Every agent session should trace back to a specific initiating event: a user submitting a form, a scheduled job firing, another agent dispatching work. Without this, you cannot determine whether the access was authorized by a human action or whether the agent fired autonomously when it should not have.
Which specific resources did the agent access, in what order, and with what outcomes? "The agent accessed the document store" is not sufficient. You need: which documents, which version, read-only or write, at what time, and whether the access succeeded or was denied.
Which tools did the agent invoke, and with what parameters? Tool invocations are where agents interact with external systems. The audit trail needs the full tool call record: tool name, parameters passed, return value type (not necessarily the full return content, but enough to characterize what was returned), and any errors.
Did any policy violations occur, and how were they handled? If an agent attempted to access a resource outside its permitted scope, that attempt should be in the audit trail whether it succeeded or failed. Blocked attempts are often more interesting than successful ones from an investigation standpoint.
What was the agent's version and configuration at the time of the activity? Agent behavior changes across versions. An audit trail that cannot identify which version of an agent ran during a specific session cannot definitively attribute behavior to a configuration decision or to an anomaly.
The Five-Layer Audit Record
The model we have found most useful structures agent audit records across five layers, from the outermost (workflow context) to the innermost (individual operations). Each layer has its own record type; the layers are linked by shared identifiers.
Workflow record. The top-level record for a workflow execution. Contains: workflow ID, initiating user or system, workflow type, start time, end time, outcome (complete / partial / failed / blocked), and a list of session IDs for all agent sessions that participated in this workflow. This record answers "who asked for what and did the workflow succeed?"
Session record. One record per agent invocation within a workflow. Contains: session ID, parent workflow ID, agent type, agent version, credential ID used, start time, end time, policy configuration hash (so you can verify what policies were active), and a list of tool call IDs. This record answers "which specific agent ran, under what identity, with what permission configuration?"
Tool call record. One record per tool invocation. Contains: tool call ID, parent session ID, tool name, tool version, parameters (structured, not raw string), target resource identifier if applicable, outcome (success / failure / blocked by policy), latency, and a reference to the policy evaluation result if the call was evaluated against a policy rule. This record answers "what exactly did the agent do?"
Resource access record. One record per external resource access (API call, database read, file system access, etc.). Linked to the tool call that initiated it. Contains: resource type, resource identifier, access type (read / write / delete), credential used, response code, volume (bytes transferred or record count), timestamp. This record is what maps to traditional access logs and is the layer that satisfies data access audit requirements for compliance frameworks.
Policy evaluation record. One record per policy rule evaluation. Contains: rule ID, session ID, tool call ID or resource access ID that triggered evaluation, rule outcome (permit / deny / flag), evaluation timestamp, and the specific rule conditions that matched. This is the layer that answers "did the agent stay within its authorized boundaries, and if not, what happened?"
Identifiers Are Not Optional
The five-layer model only works if the identifiers that link the layers are actually populated consistently. This sounds obvious, but it is where most deployments fall short. Trace IDs get dropped when an agent calls a tool that makes an API call to a third-party service that does not propagate custom headers. Session IDs are generated per agent invocation but not stored in the resource access logs because the underlying infrastructure logging (CloudTrail, GCP audit) does not know about your session ID scheme.
The practical solutions are: first, use the agent framework's instrumentation hooks to capture events at the tool call level, before the tool makes the underlying resource access. At that point, you control the event structure and can include your session and workflow IDs. Second, for resource access in services that generate their own logs (database logs, cloud audit logs), build a correlator that matches those logs to agent session records by timestamp and credential, rather than relying on the resource's logs to carry your trace ID. The correlation is not perfect, but it narrows the attribution window enough to be useful for investigation.
Log Retention and the Compliance Window
The retention question matters more for agent audit logs than for most application logs because the questions that audit logs answer tend to be asked retrospectively, often months after the fact. A SOC 2 audit may ask about access events from three to six months ago. An insurance claim investigation may cover a six-to-twelve month window. A regulatory inquiry under certain data protection frameworks may require demonstrating what data access occurred during a specific past period.
The default log retention in most cloud platforms is 90 days, which covers the SOC 2 window if you do not let it lapse, but leaves you exposed for longer retrospective questions. The records that need longer retention are the session records and resource access records: those are the audit artifacts that prove agent behavior at a specific point in time. Tool call records and policy evaluation records can often be retained for a shorter window because their primary use is operational investigation, not long-term audit proof.
For regulated industries, particularly financial services and healthcare, the retention requirement is often set by the applicable regulation rather than by internal policy. Build your retention tiers around those requirements, not around your current storage cost sensitivity.
What a CISO Actually Needs to See
The audit output format matters as much as the log structure. A CISO asking "what did the agents do last night?" is not going to read raw JSON log records. The audit output they need is: a summary of workflow executions (how many, which types, any failures or blocked actions), a drill-down view of any session with policy flags, a resource access summary (which resources were accessed by which agent types and in what volumes), and an attestation that all agent activity was within configured policy boundaries, or an explicit list of deviations.
That output needs to be producible in under five minutes for a routine review and under fifteen minutes for an incident investigation. If producing it requires a manual log query and assembly process, the audit function is not operationally sustainable. The goal is an audit function that can be run as a daily or weekly automated report, with exception-based escalation to human review.
What This Is and Is Not
An agent audit trail tells you what happened. It does not tell you whether what happened was appropriate in every case, beyond the mechanical policy evaluation. An agent that had permission to access a set of records and accessed them is auditable, but whether accessing all of them in a single session was the right business decision is a judgment call that requires human context. The audit trail gives you the facts; the governance review uses those facts to make judgments.
We are also clear that audit logging as described here is not a complete security solution. An agent can do harmful things within its permitted boundaries. The audit trail records what the agent did; it does not guarantee the agent was doing the right thing. The permission boundary enforcement is the control that limits potential harm; the audit trail is the record that enables review, investigation, and accountability.
Those are related but distinct functions, and both are necessary.
Audit your AI agents with Arrakis.
Arrakis gives security and platform teams a complete audit trail of every action your AI agents take, with policy enforcement that stops overreach before it reaches your data.
Request a Demo