Abstract bridge visualization connecting traditional compliance structures with modern AI systems
Back to Blog

Getting Compliance Teams to Take AI Agents Seriously

Risk and compliance teams have a well-developed mental model for two categories of things they need to audit: human processes and static software systems. Human processes have roles, approval chains, and audit trails of decisions made by named individuals. Static software systems have change controls, access logs, and vendor assessments with fixed scope.

Autonomous AI agents are neither of those things. They make decisions at runtime based on model inference. They access systems dynamically rather than through a fixed code path. Their behavior in a specific situation can be hard to predict from their configuration alone. And the people who built them often can't fully articulate what the agent will do when given a task it hasn't been specifically tested on.

This creates a real challenge when you need to bring compliance teams into a conversation about AI agent governance. The conversation tends to stall because the team doesn't have a framework for categorizing what they're being asked to review. This article is about how to have that conversation productively.

Start With What They Already Care About

The mistake most engineering teams make when introducing AI agents to compliance is leading with how the technology works. They explain the model, the tool registry, the context window, the agent framework. Compliance teams are not evaluating a technology. They are asking: what can go wrong, how do we know if it went wrong, and what is the evidence that controls exist to prevent it going wrong.

Reframe the conversation around those three questions and the engagement changes immediately.

What can go wrong: an AI agent could access data it was not supposed to access, write to a system it should not be writing to, exfiltrate information to an external destination, make a decision that should have required human review, or take a sequence of actions that individually look fine but together produce a harmful outcome. These are recognizable categories. Data access control failures, unauthorized writes, and decisions requiring human escalation are things compliance teams deal with in human processes all the time.

How do we know if it went wrong: this is where the conversation gets specific. Can you show a log of every action the agent took in a given session? Can you show which policy rules applied to each action? Can you reconstruct the sequence of decisions that led to a specific outcome? If the answer is yes, you are in a position to have a substantive compliance conversation. If the answer is no, or "we can reconstruct it but it takes several hours and some manual work," that is the gap you need to close before the compliance conversation can make progress.

What controls exist: access policy rules, session logging, violation alerting, human review workflows for flagged sessions, and a defined escalation path when the agent behaves outside its boundary. These are recognizable control types that map to existing audit frameworks.

The SOC 2 Frame

If your organization is working toward SOC 2 compliance or maintaining a SOC 2 attestation, AI agents that interact with in-scope systems need to be in scope for the audit. This is not a controversial position once compliance teams understand what agents do, but it often requires a direct conversation because agents don't fit neatly into the "users" or "systems" categories that SOC 2 audit procedures typically enumerate.

SOC 2 Trust Services Criteria relevant to AI agent deployments include several that are worth mapping explicitly. CC6.1 (logical access controls) applies directly: agents hold credentials and make access decisions. The question for an auditor is whether those credentials are provisioned with appropriate scope, whether access is logged, and whether there is a process for reviewing and rotating agent credentials. These are answerable questions if you have your governance infrastructure in place.

CC7.2 (system monitoring to identify anomalies) is another direct hit. Agents operating outside their declared behavior boundary are exactly the anomaly category this control is designed to catch. If your monitoring can show a compliance auditor that you have detection for out-of-scope resource access, unusual tool call volumes, and policy violations with alert routing, you satisfy this control for the agent layer.

The harder conversation is around CC2.2 (communication of system changes). When an agent's behavior changes because the model was updated, or because the tool registry was expanded, does that constitute a system change that requires a change control record? Technically, a model version update changes the system's behavior in ways that can be hard to fully specify in advance. In practice, most organizations handle this by treating model version updates as a deployment event with an associated testing record, and treating tool registry changes as a configuration change requiring review. Neither is a perfect fit for traditional change control, but both can be documented in ways that satisfy an auditor who understands the domain.

GDPR and Data Handling by Agents

For organizations with European operations or customers, AI agents that process personal data require specific treatment under GDPR. The relevant questions are: what personal data can the agent access, does it retain any of that data, and if so under what legal basis and for how long.

Agents deployed for document review, HR processes, or customer-facing workflows are almost certainly processing personal data in the GDPR sense. The compliance obligation does not require you to avoid this; it requires you to document it and ensure appropriate controls.

The documentation path for GDPR is a data processing record for each agent that specifies: the categories of personal data the agent can access (names, email addresses, financial data, health information, etc.), the purpose of processing (task description), the legal basis, the retention period for any data the agent retains or outputs, and the technical measures in place (access controls, encryption, audit logging).

What makes this tractable from a compliance perspective is that the same governance infrastructure that provides your security audit trail also provides most of the evidence needed for GDPR compliance documentation. The agent's policy record defines what data it can access. The session logs document what data it actually accessed. The retention policy for session logs defines how long that record is kept. These are the same controls serving two compliance objectives simultaneously.

What Compliance Teams Need to See

In conversations with risk and compliance professionals about AI agent governance, several specific asks come up consistently. These are the evidence artifacts they want to see.

An agent inventory: a list of every autonomous AI agent deployed in the environment, what it does, what systems it can access, and who owns it. This is the starting point for scoping any audit.

Policy documentation per agent: for each agent, what it is permitted to do and what it is not permitted to do, expressed in enough specificity that a non-technical auditor can understand it. "The accounts payable agent may read from the AP data store and may write to the AP workflow queue; it may not write to any external system or access any data store outside the AP scope" is sufficient.

Evidence of policy enforcement: actual examples of policy checks that ran during agent sessions, including both permitted actions (to show the policy is functioning) and blocked or flagged actions (to show violations are caught). Screenshots of the audit log interface, or a log export covering a representative time period, serve this purpose.

A defined escalation process: what happens when the monitoring system detects a policy violation? Who is notified, through what channel, with what response time commitment, and what is the remediation path? This should be documented in writing and should match what the technical system actually does.

A credential review process: how often are agent credentials reviewed? Is there a process for rotating credentials when an agent's scope changes or when a team member who provisioned the agent leaves? This is analogous to the user access review processes that SOC 2 auditors look for, adapted for non-human identities.

The Conversation Starters That Work

We have seen a few framing approaches that move compliance conversations forward more reliably than others.

"Agents are non-human users in our IAM model." This gets compliance teams to apply their existing user access review framework to agents, which is a reasonable starting point even if agents eventually need additional-specific controls beyond what that framework covers.

"Every agent session is a logged transaction." This frames agent activity in terms compliance teams associate with financial systems, where every transaction has a record and every record is attributable. It sets the expectation that agent activity should be auditable at the individual session level, not just summarized in aggregate reports.

"The audit trail we keep for agents is comparable to what we keep for admin users." Admin user activity monitoring is a familiar compliance control. If your agent monitoring is genuinely at that level of granularity, the comparison is apt. If it isn't, this framing also identifies the gap clearly.

We want to be clear that these framing approaches work because they are accurate descriptions of what a mature AI agent governance program looks like. They are not spin. If your governance infrastructure doesn't actually support these descriptions, the compliance conversation will quickly surface that. The right response is to close the gap, not to refine the framing.

When Compliance Teams Push Back

Some compliance teams, when first presented with AI agent governance as an audit topic, resist adding it to their scope. The pushback usually takes one of two forms.

The first is categorization resistance: "This is an engineering decision, not a compliance matter." This position is defensible in some circumstances (agents that have no access to regulated data and no connection to financial systems might genuinely be outside compliance scope) but usually isn't sustainable once the team understands that production agents are accessing data and systems that are already in audit scope. The question is not whether agents are in scope; it is whether the controls framework extends to cover them.

The second pushback is capacity: "We don't have bandwidth to audit a new technology category right now." This is real and worth acknowledging. The practical answer is to start with the highest-risk agents (those with access to regulated data or production systems) rather than trying to bring everything into scope at once. A phased approach that prioritizes by risk level is more likely to succeed than a wholesale addition to the audit scope.

The goal is not to create a compliance exercise. It is to build the governance infrastructure that makes AI agents auditable, and then to integrate that infrastructure into your existing compliance processes in a way that is sustainable over time. Compliance teams are more receptive when they see that the engineering and security teams have already built something, rather than being asked to define requirements for infrastructure that doesn't yet exist.

Audit your AI agents with Arrakis.

Arrakis gives security and platform teams a complete audit trail of every action your AI agents take, with policy enforcement that stops overreach before it reaches your data.

Request a Demo