When most security teams think about AI model risk, they focus on the model: hallucination, prompt injection, output quality. Those concerns are real but incomplete. The more immediate operational attack surface for AI agents is not the model itself. It is the tools, credentials, and systems the agent is authorized to use. The model decides what to do; the tools are how the agent does it. And the tools are reachable.
Enumerating that surface is not a theoretical exercise. It is a prerequisite for governance. If you do not know what each agent can reach, you cannot make informed decisions about what it should be allowed to reach, and you cannot detect when it reaches something unexpected.
The Three Layers of Agent Attack Surface
Agent attack surface maps across three layers, each with distinct characteristics and distinct threat models.
Layer 1: The tool manifest. Every agent framework uses some form of tool registration: a structured list of functions or APIs the agent is authorized to invoke. In LangChain this is the tools list passed to the agent executor. In AutoGen it is the registered functions on the assistant and user proxy. In OpenAI Assistants it is the functions array in the API call.
The tool manifest is the definitive statement of what the agent can do in a single session. But tool manifests are often not tracked as configuration artifacts the way application code is. They live in Python files, environment variables, or runtime configuration objects that are not stored in an access-controlled registry with change history. If you ask "what tools did version 2.3 of this agent have access to?" in many deployments, the answer requires reconstructing it from git history and environment snapshots.
Layer 2: The credential scope. Tools are only as powerful as the credentials backing them. An API client tool is limited by what the API key can do. A database query tool is bounded by what the database user credential permits. A file system tool operates within the permissions of the service account.
The attack surface at this layer is the gap between what the credential can do and what the tool is supposed to do. A credential provisioned for "read customer records" that technically also has write permission represents a surface gap. The tool does not use the write permission, but it is available. If the agent is manipulated, malfunctions, or if the tool implementation has a bug that sends an unexpected request, the write permission becomes exploitable.
Layer 3: The data access graph. The tools and credentials combine to define a data access graph: the set of resources, records, and systems reachable by the agent through any sequence of tool invocations. This is broader than the direct tool list because tools can compose. An agent that can query a database and also make HTTP requests can potentially exfiltrate data from the database to an external endpoint through two tool calls in sequence, even if neither tool alone does anything alarming.
Mapping the data access graph requires tracing not just what each individual tool can reach but what the agent can reach through multi-step tool use.
Why Standard Asset Inventory Misses This
Traditional asset inventory and network mapping tools were designed for static infrastructure: hosts, ports, services, subnets. They can tell you what is on the network. They cannot tell you what an AI agent session can reach through an authorized API call that travels over HTTPS on port 443, which is indistinguishable from legitimate traffic at the network layer.
Agent attack surface is a software-layer and identity-layer problem, not a network-layer one. The agent's reach is defined by credentials and API authorizations, not by firewall rules. Standard CSPM or cloud security tools that scan for open ports and misconfigured buckets will not enumerate what your document processing agent can do with the Salesforce API credentials it was given.
This means the enumeration work has to happen at a layer closer to the agent runtime: reading the tool manifest, tracing the credential scope, and mapping out what the combination implies about reachable resources.
A Practical Enumeration Approach
When we work through agent attack surface mapping for a deployment, we follow a sequence of four questions for each agent type.
What tools are registered? Collect the tool manifest for each agent type. This should be a static artifact derived from the agent definition, not a runtime observation (though runtime observation can verify it). If the tool manifest is not stored as a version-controlled configuration file, that is itself a finding: tool changes should be traceable.
What credentials back each tool? For each tool, identify the service account, API key, or credential used to execute it. Pull the effective permission set of that credential from the relevant IAM system. The gap between what the credential permits and what the tool is documented to use is your Layer 2 surface.
What resources are reachable through each tool? For database tools: which schemas, tables, and column-level permissions. For API tools: which endpoints, which HTTP methods, which data objects in the response schema. For file tools: which directories, which file types, whether write or delete permissions are included. Document this per-tool.
What can the agent reach through tool composition? Take the union of reachable resources across all tools and ask: are there combinations that could enable exfiltration, privilege escalation, or lateral movement? An agent with both a "read internal records" tool and an "send email" tool can exfiltrate. An agent with a "run query" tool and a "create file" tool can persist data outside the intended storage. These are not necessarily policy violations, but they need to be conscious decisions.
Prompt Injection as a Surface Amplifier
A dimension of attack surface that is specific to LLM-based agents is prompt injection: the possibility that content the agent processes contains instructions that redirect the agent's behavior. A document processing agent that reads customer-submitted contracts is exposed to whatever text those contracts contain, including text crafted to influence the agent's tool use.
We are not saying prompt injection is easy to execute reliably or that it is the dominant threat vector. In practice, modern LLMs are increasingly resistant to naive prompt injection, and well-scoped tool manifests limit the damage from a successful attempt. But prompt injection does mean that the attack surface of an agent that reads external content is larger than the attack surface of an agent that only reads internal, trusted data. The external content exposure is part of the attack surface map.
The mitigation implication: agents that process untrusted external content should have narrower tool manifests and more conservative credential scopes than agents that operate only on internal data. The external content exposure should bump up the scrutiny on what those agents can do.
Making the Map Useful for Policy
An attack surface enumeration that lives in a spreadsheet and never gets updated is not useful. The value of the map is in making it a living document tied to the agent deployment pipeline.
Practically, this means the tool manifest and associated credential scope should be version-controlled alongside the agent definition. When a tool is added, removed, or its backing credential changes, that is a change to the attack surface map. It should trigger a lightweight review: does the delta make sense given the agent's stated purpose? Is the new surface justified?
The review does not need to be heavyweight. For most tool additions, a one-line note confirming the business justification is sufficient. The discipline of requiring that note is what prevents the accumulation of undocumented surface that makes later incident investigation so painful.
The data access graph view, the multi-step composition analysis, should be reviewed periodically or when significant tool or credential changes occur. This is the part that requires the most judgment. An automated tool can enumerate it; a human needs to evaluate whether the combinations create acceptable risk for the agent's operating context.
Starting From Where You Are
If your current agent deployments do not have version-controlled tool manifests or formal credential scope documentation, the starting point is not rewriting everything. It is capturing a snapshot: for each production agent type, what tools does it have and what credential backs each one? That snapshot, even if produced manually today, is the baseline against which future changes can be measured.
The teams with the cleanest attack surface posture are not those that locked everything down at the start. They are the ones that made a habit of documenting changes when they happened, so the map stayed current. Starting the documentation habit now, even imperfectly, is more valuable than a comprehensive audit that is out of date by the time it is done.
Audit your AI agents with Arrakis.
Arrakis gives security and platform teams a complete audit trail of every action your AI agents take, with policy enforcement that stops overreach before it reaches your data.
Request a Demo