Start by turning on central telemetry export for the harnesses that can execute actions, then verify that approvals, tool calls, and session identifiers arrive in a collector you actually monitor. If you only keep local transcripts, you will miss the evidence needed for review, investigation, and containment.
Why telemetry comes first when AI agents can act
As soon as an agent can run commands or call tools, the first job is to make its actions observable outside the agent runtime. That means exporting approvals, tool invocations, and session context into a central collector you can monitor, search, and alert on. Local transcripts are useful for debugging, but they are not a control surface for review or containment.
Central telemetry gives security teams a shared record of what the agent tried, what was allowed, and which identity or session performed the action. Without that baseline, later response work starts with uncertainty: you cannot reliably reconstruct intent, prove scope, or tell whether a suspicious command was issued once or repeated across sessions.
What good telemetry needs to capture
The minimum useful dataset is not just chat content. It should include the approval decision, the tool name, the arguments or action summary, the session identifier, and whatever request or correlation ID lets you join events across the harness, broker, and downstream system. If the agent can touch multiple tools, you need a trail that links the approval to the exact action that followed.
That distinction matters because agents often separate the prompt, the tool call, and the resulting side effect. A transcript that only preserves the natural-language exchange can miss the operational event that actually changed state. AI Agent Observability, Audit and Incident Response Guide is the right reference point for deciding which signals belong in the audit path and how to preserve attribution across a session.
Teams should also treat command execution and tool access as policy-relevant events, not just logs. Once agents can invoke tools, the security question shifts from “what did the model say?” to “what was authorised, by whom, and under what conditions?” AI Agent Authorisation Guide and Zero Trust for AI Agents both reinforce that per-action visibility is what makes least privilege and continuous verification workable in practice.
Why local transcripts are not enough
Local-only logs fail in the exact moments that matter most. They are easy to lose when a container exits, hard to correlate across services, and often invisible to the people who need to review them during an incident. They also do not scale well when multiple agents, multiple sessions, or chained tools are involved, because the investigation depends on stitching together separate execution paths.
The operational consequence is that containment becomes slower and less certain. If you cannot see the sequence of approvals and tool calls centrally, you cannot quickly decide whether to suspend a session, revoke a credential, or block a tool integration. That is why strong agent logging is not a reporting nice-to-have, it is part of the response path. AI Agent Observability, Audit and Incident Response Guide and Agentic AI Security Guide both support the broader control idea that observability and blast-radius reduction have to be designed together.
Risk and Threat Considerations
Once agents can execute actions, the main risk is not just bad output, it is unaudited action. A compromised prompt, a mistaken approval, or a malicious tool invocation can create real system changes without leaving a usable trail if logging stays local or fragmented. That leaves teams blind to abuse, weakens non-repudiation, and delays containment when the agent touches sensitive systems.
Failure mechanism: the harness records activity locally, but the security team cannot centrally see approval decisions, tool calls, and session identifiers in time to investigate or stop follow-on action. Missing correlation data breaks attribution and makes it difficult to prove what happened across chained steps or multiple tools.
Impact: incidents become harder to detect, harder to scope, and harder to contain. A team may have to revoke access broadly, accept uncertainty about state changes, or rebuild the timeline from partial evidence after the fact.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent tool execution creates identity and privilege audit needs. |
| Recommendation — Log each agent approval and tool call to preserve actionable identity and privilege traceability. | ||
| NIST SP 800-53 Rev 5 | AU-2 — Event Logging | Central telemetry export is fundamentally about audit event capture for agent actions. |
| AU-6 — Audit Record Review, Analysis, and Reporting | Security teams need reviewed central logs to investigate agent activity. | |
| IA-5 — Authenticator Management | Session identifiers and approval chains depend on managing identity-bearing session material. | |
| Recommendation — Define and collect agent execution events as auditable system activity. Route agent logs into a monitored collector and review them for suspicious action patterns. Tie agent sessions and credentials to managed, reviewable authentication records. | ||
Practitioner Guidance
Where to start: enable central export on the harnesses that can execute actions before expanding the agent’s toolset. If the telemetry path is not already feeding a monitored collector, treat the agent as operationally immature, even if the model output looks safe.
What to verify: confirm that each event carries a stable session identifier, the tool name, the approval outcome, and enough context to link the action to a downstream system change. If any of those fields are missing, you do not yet have a dependable audit trail.
Decision rule: if the agent can affect production systems, prioritise observability and revocation readiness over convenience features such as local-only playback or developer-facing transcripts. The first useful control is the one that lets you answer who did what, when, and through which tool.
Practitioner takeaway: the first security milestone for action-capable agents is not perfect policy, it is trustworthy evidence. If you cannot reconstruct and monitor agent actions centrally, you do not yet have control over them.
Related resources from NHI Mgmt Group
- How should security teams manage permissions for AI agents?
- How should security teams govern AI agents that use OAuth access?
- How should security teams limit the risk from AI agents that have access to production systems?
- How should security teams govern AI agents that can access enterprise systems?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org