Enterprise security teams have spent years building detection models for privileged actions. Shell commands executed by a service account. Admin API calls outside business hours. IAM role assumptions by non-human principals. The common thread is that these actions can have significant and hard-to-reverse effects on systems, and they are the actions attackers target when they want to do real damage.
LLM tool calls belong in this category. They aren't there yet in most enterprise detection models, and that gap is growing as agent deployments expand.
Here's the argument: a tool call made by an LLM agent is a privileged action. It executes with the credentials of the agent's service account. It can read data the model was not originally shown. It can write to external systems. It can trigger other agents or workflows. The fact that the action was decided by a language model rather than a human does not change its effect on systems.
What Makes Tool Calls Different from API Calls
Security teams are used to monitoring API calls. Every major cloud provider has logging for control-plane API calls. Web application firewalls inspect HTTP traffic. SIEM rules trigger on unusual call patterns, anomalous source IPs, or calls to sensitive endpoints.
Tool calls from LLM agents share some properties with API calls but differ in three important ways that affect how you need to monitor them.
First, tool calls are model-generated. The decision to make the call and the parameters passed to it come from the model's inference step, not from deterministic code. This means the same agent, given slightly different input context, might call different tools with different parameters. Traditional rule-based detection that flags specific API endpoints or argument patterns works less reliably here because the parameter space is less predictable.
Second, tool calls chain together. An agent session doesn't make a single API call. It makes a series of calls where each one can influence the next. The model reads the output of call N and decides what to call at N+1. This creates a dependency structure where a detection gap at one step can blind you to what happens in subsequent steps. You need to see the session-level sequence, not just individual calls in isolation.
Third, tool call semantics depend on context that is not in the call itself. Consider a file read tool call that retrieves a contract document. Whether that call is expected or anomalous depends on what task the agent was given, what documents it has already retrieved, and whether this document is within the scope of the current session. None of that context appears in the raw API log of the call. You need the session context to evaluate the call's intent.
The Current Gap in Enterprise Detection Coverage
We have looked at the detection coverage for AI agent activity in a range of enterprise security architectures. The pattern is consistent. Teams have good coverage for the downstream effects of agent actions: if an agent causes a database write, that write shows up in the database audit log. If an agent calls an external API, that call shows up in egress logs or the API provider's access log. If an agent modifies a file, the file system audit captures it.
What teams typically lack is coverage at the agent layer itself. They can see what changed, but they cannot see what agent decision caused the change, what tool call sequence led to it, or what context the model was operating in when it made that decision. The agent layer is a black box between the triggering event and the downstream effect.
This gap matters most in two scenarios. The first is investigation: when something unexpected happens and you need to understand whether an agent caused it, the absence of agent-layer logs means you're working backward from effects to hypothesize causes rather than forward from causes to trace effects. It is slower and less certain.
The second is behavioral anomaly detection: identifying when an agent is doing something unusual before it causes a downstream effect. If your only visibility is the downstream systems, you find out about anomalous agent behavior after it has already acted, not while it is acting. For agents with write access to critical systems, "after it has already acted" is often too late.
What Agent-Layer Detection Needs
Building detection coverage at the agent layer requires three things: instrumentation that captures tool calls in structured form as they happen, a schema that is consistent enough to write detection rules against, and integration with your existing alert infrastructure so that detections can trigger the same response workflows as other event types.
On instrumentation: the right place to capture tool calls is between the model's inference step and the actual execution of the tool. This position lets you see what the model decided to call before it executes, which creates the possibility of enforcement (blocking a call that violates policy) rather than only detection after the fact. Capture at this layer also gives you the full model-generated parameters, not just the network traffic that results from executing them.
On schema consistency: this is the part that enterprise security teams push back on hardest, and they're right to. If every agent framework emits tool call events in a different format, writing detection rules is impractical at scale. The answer is a normalization layer that translates framework-specific events into a common schema. At minimum, the common schema needs: session ID, agent identifier, tool name, call parameters, call result, timestamp, and policy outcome. That set enables the queries that matter for detection and investigation.
On integration: tool call alerts should route to the same places as your other high-priority security events. If a tool call violates a policy rule, that event should appear in your SIEM, trigger your PagerDuty integration, and land in your security team's incident queue, the same as a privilege escalation event would. Treating agent security alerts as a separate silo means they get less attention and slower response.
Detection Patterns Worth Building First
Not all tool call patterns are equally detectable or equally risky. Starting with the highest-value detection rules gives you coverage where it matters most before you've built out a full agent monitoring capability.
Tool calls to resources outside the agent's declared scope are the highest priority. If an agent was provisioned with access to a specific set of data sources and it attempts to call a tool that reaches a different data source, that is an immediate flag. The detection rule is straightforward: does the target resource of this tool call appear on the agent's allow list? This is a binary check that can be implemented without complex baseline modeling.
Write operations by agents that should only be reading are the second priority. Many agents are designed to read and analyze data. If a primarily read-oriented agent makes a tool call that writes to any external system, that is worth investigating. The detection rule is simple: check the tool call type (read vs. write) against the agent's declared operation mode.
Volume anomalies within a session are the third category. An agent that makes 200 tool calls in a session where similar past sessions averaged 20 is either doing something very different or something has gone wrong with its context handling. This requires baseline data from prior sessions, but even a rough threshold (more than N tool calls in a single session) is worth implementing as an early signal.
External API calls that don't match the agent's approved destination list round out the starting set. Agents that call external services outside the approved list are the most direct path to data exfiltration. The approved-list check is the same binary pattern as the resource scope check above.
This Is Not Just About Attack Prevention
We want to be specific about what agent-layer monitoring is and isn't for. It's not a complete answer to all AI agent security concerns. It doesn't address model-level risks, prompt injection, or the correctness of the model's reasoning. Those are real problems with their own solution categories.
What agent-layer monitoring addresses is the interface between AI agents and the enterprise systems they interact with. Tool calls are that interface. Monitoring them closes a detection gap that exists in every enterprise currently deploying agents, independent of which model they use or how they've designed their agent workflows.
Closing that gap also serves purposes beyond security. Operations teams want to understand what their agents are actually doing. Compliance teams need evidence that agents are operating within defined boundaries. Engineering teams want session data to debug unexpected agent behavior. The same instrumentation that enables security detection also serves these use cases.
The organizations that build this coverage now will have better incident response capability, better compliance evidence, and better operational visibility than those that wait until an event forces the issue. The gap is known. The instrumentation approach is understood. The work is building and deploying it.
Audit your AI agents with Arrakis.
Arrakis gives security and platform teams a complete audit trail of every action your AI agents take, with policy enforcement that stops overreach before it reaches your data.
Request a Demo