Look for unexpected tool calls, instructions inside routine-looking responses, oversized payloads, repeated delegation loops, and audit records that do not reconstruct a clean handoff path. Those are the practical signs that validation is missing or that a sub-agent is carrying executable content instead of data. If the chain cannot be replayed, the control model is already weak.
What failing inter-agent communication looks like in practice
When inter-agent communication is healthy, the receiving agent can distinguish instructions from data, preserve the handoff context, and complete the delegated step without inventing new authority. When that boundary breaks down, the failure usually shows up as a protocol problem, not a dramatic outage: the message shape, trust boundary, or delegation chain stops matching what the orchestrator expects.
That is why detection should start with the conversation itself. Unexpected tool calls, instructions embedded in otherwise routine replies, and payloads that look executable rather than informational are all signs that the receiving side is interpreting content too broadly. In multi-agent systems, those symptoms often point to broken validation, weak message typing, or an agent that has been allowed to act on untrusted output as if it were an authenticated control signal. For a deeper practitioner view of the communication layer, see Multi-Agent and A2A Security Guide.
A second signal is delegation drift. If one agent keeps re-tasking another, or the same request bounces between agents without convergence, the system may be stuck in a loop where no participant can prove who owns the next step. That is not just inefficiency. It is often the first visible sign that handoff metadata, message signing, or per-hop authorization is missing, incomplete, or ignored.
Which telemetry patterns are most diagnostic
The most useful indicators are the ones that let you replay the chain. Audit logs should show who initiated the work, which agent received it, what was passed as data, what was treated as instruction, and where the decision changed hands. If that path cannot be reconstructed cleanly, the control model is already too loose for safe multi-agent operation. Pair those checks with AI Agent Observability, Audit and Incident Response Guide so your logging criteria match the failure modes you are trying to catch.
Oversized payloads matter because they often hide policy bypass attempts, prompt stuffing, or chained instructions that were not meant to traverse agent boundaries. Repeated delegation loops matter because they can mask a failure to establish accountability, especially when each agent believes the next hop owns validation. Correlate those patterns with request metadata, correlation IDs, and tool invocation logs rather than relying on the surface transcript alone.
One especially strong indicator is when a sub-agent returns content that contains embedded operational directives, formatting intended for execution, or references to tools it should not have discovered on its own. That usually means the boundary between message and action has weakened. If you want a broader control view of how identity, access, and request handling should be separated, AI Agent Authorisation Guide is the most relevant internal reference.
How teams should investigate and contain the failure
Investigate from the handoff outward. First confirm whether the original request was signed, scoped, and attributable. Then check whether each hop preserved the original intent without adding authority. Finally verify that tool calls were generated by the right agent, at the right step, with the right constraints. This sequence is important because communication failures in agent systems are often authorization failures in disguise.
The practical containment decision is simple: if you cannot distinguish data from instruction, stop the chain, not just the task. That usually means pausing downstream tool access, preserving logs, and replaying the exchange against the expected policy model before allowing further delegation. When the control plane is uncertain, continued execution compounds the problem by giving a confused agent more opportunities to act on malformed context.
For teams evaluating the wider governance angle, the strongest signal is not one bad message but a repeatable pattern across several handoffs. That is the point at which you should treat the issue as systemic and review agent registration, delegation rules, and per-action policy enforcement together. The broader agent-risk perspective is covered well in Agentic AI Security Guide.
Risk and Threat Considerations
Broken inter-agent communication creates a trust-boundary failure that can turn a routine delegation into unintended execution, privilege misuse, or hidden persistence. The risk is highest when agents can pass executable content, reuse context across hops, or trigger tools without a fresh policy decision.
Failure mechanism: An agent accepts another agent’s output as both instruction and data, or the audit trail loses the original handoff context, so validation and authorization no longer align with the action being taken.
Impact: Attackers or faulty automations can induce unauthorized tool calls, recursive loops, covert task escalation, or silent corruption of downstream agent decisions, making compromise harder to detect and contain.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI07 — Insecure Inter-Agent Communication | Directly covers failures between agents exchanging messages. |
| ASI03 — Identity & Privilege Abuse | Communication failure often becomes unauthorized action across agent boundaries. | |
| ASI08 — Cascading Failures | Repeated delegation loops and chained errors can propagate across agents. | |
| Recommendation — Validate agent handoffs, message types, and trust boundaries before allowing downstream actions. Enforce per-action authorization so one agent cannot expand another's privilege. Contain failure early and stop chained execution when delegation stops converging. | ||
| MITRE ATT&CK | TA0001 — Initial Access | Abuse of inter-agent trust can create an entry path into downstream execution. |
| Recommendation — Monitor trusted handoffs for unexpected entry into higher-privilege workflows. | ||
| NIST SP 800-53 Rev 5 | AU-2 — Event Logging | Replayable audit records are needed to reconstruct agent handoffs. |
| IA-9 — Service Identification and Authentication | Agent-to-agent communication depends on authenticating non-human actors. | |
| Recommendation — Log each agent hop with sender, receiver, payload type, and policy decision. Authenticate each agent service before accepting delegated messages or tool requests. | ||
Practitioner Guidance
What to verify: Confirm that every hop in the chain preserves a unique sender, a typed payload, and an explicit policy decision for any tool-using action. If the log cannot show that separation, treat the exchange as unsafe even if the output “looks right.”
What to measure: Track replayability, loop frequency, malformed handoff rates, and the share of agent outputs that contain executable-looking content. Those signals are more actionable than raw message volume because they expose where the trust model is breaking down.
Common mistake: Teams often focus on prompt quality and miss the real issue, which is boundary enforcement between agents. The control objective is not perfect text, it is reliable attribution, constrained delegation, and auditability at each step.
Practitioner takeaway: If you cannot replay the conversation and explain every handoff, you do not yet have safe inter-agent communication, only optimistic coordination.