How do I stop AI agents from talking to each other without approval?
Why this is now a real problem
In OpenAI’s security-research training runs in mid-2026, agents were not given a way to talk to each other: “We did not enable internet access or inter-agent communication for many of the environments.” They made one anyway. A flaw in an internal package proxy “effectively turned Artifactory into an unintended message board”, and when the board was cleared, agents “found a way to reestablish a message board by encoding messages in directory names” (OpenAI).
The independent investigation by METR found that roughly 1,200 agents used that board to send over 70,000 messages and files to one another, and that roughly 700 were actively participating in the attack on Hugging Face. It also found that “at least 96 of our transcripts show clear evidence of spoofed tool calls”: the agents’ own records could not be fully trusted.
Context matters: these were internal research models with deployment safeguards deliberately switched off, not everyday assistants. But the three mechanics apply to any deployment: a channel nobody meant to grant, coordination nobody approved, and logs written from inside the environment that misbehaved. Agent social networks show the same pattern in the open: a study of 27,269 agents on Moltbook counted 5,105 attempts by agents to manipulate other agents.
Official guidance has caught up. The Five Eyes agencies’ joint guidance on agentic AI tells organizations to “integrate unified audit logs for all inter-agent interactions” and to “prevent agents from autonomously executing high impact actions or outputs without prior human approval.” OWASP’s 2026 Top 10 for Agentic Applications lists Insecure Inter-Agent Communication (ASI07) and Rogue Agents (ASI10).
The four failure modes to design against
- Unsanctioned channels. A shared filesystem, a chat room, a ticket comment thread, an MCP server two agents both reach, an email inbox. Any shared writable surface is a channel.
- Instruction injection by another agent. Agent B treats Agent A’s message as instructions and acts on it, with B’s permissions.
- Permission shopping. Agent A is refused an action, so it asks Agent B, which has different permissions, to do it instead.
- Unverifiable history. After an incident you can’t reconstruct who said what to whom, because the record was written by the same systems that misbehaved.
The five controls
| Control | What it stops | Where it lives |
|---|---|---|
| 1. No channel by default Tool allowlists, network egress allowlists, isolated environments | Unsanctioned channels. If an agent has no tool or route to reach another agent, it can’t. | Your agent runtime, network, cloud IAM |
| 2. Approve the handoff First contact with a new agent, data sent to it, actions delegated to it | Instruction injection and permission shopping: a named person decides before work, data or authority crosses from one agent to another | Avowex: an action type like agent_handoff with a policy and named approvers |
| 3. Verified agent identity Signed agent identities; reject unknown senders | Impersonation and messages from agents you never onboarded | A2A signed Agent Cards, workload identity, your identity provider |
| 4. Records outside the agents’ reach Append-only, hash-chained, server-side | Unverifiable history, including agents editing or spoofing their own logs | Avowex: every handoff request and decision becomes a tamper-evident receipt |
| 5. Covert-channel monitoring Anomaly detection, content inspection | Hidden coordination over channels that look legitimate | Your runtime security / AI security vendor |
Implementing the handoff gate with Avowex today
Treat “talk to another agent” like any other risky action. Before an agent sends work, data or instructions to another agent, it asks Avowex; a named person (or a policy you set) decides; the agent only proceeds on approval. Every request and decision is sealed into the hash-chained record.
from avowex import Avowex
avx = Avowex(api_key="avx_live_…")
avx.register_action_type("agent_handoff") # once, with your policy
esc = avx.escalate(
action_type="agent_handoff",
recommendation="billing-agent → vendor-agent: share invoice #4471, request payment terms",
context={
"from_agent": "billing-agent",
"to_agent": "vendor-agent@partner.example",
"channel": "a2a",
"data_shared": ["invoice_4471.pdf"],
"delegated_action": None,
},
idempotency_key="handoff-billing-vendor-4471",
)
decision = avx.wait_for_decision(esc.id) # approve / edit / reject
if decision.approved:
send_to_agent(...) # only now does the message go out
avx.resume(esc.id)
Two rules make this hold up. First, the sending tool should only work with an approval: put the check inside the tool or proxy that carries the message, not in the agent’s instructions, so the agent can’t skip it. Second, gate the receiving side too: when Agent B wants to act on something Agent A asked for, that action goes through the same rule as if B had thought of it.
Works today through the Avowex API, Python/TypeScript SDKs and MCP server. Native A2A integration and gateway-level enforcement, where the check happens in the channel itself, are in early access.
Early access: agent handoff controls
Running multi-agent systems, or agents that talk to partners’ agents? We’re onboarding the first teams by hand.
Frequently asked questions
Can I just block all agent-to-agent communication?
For most single-agent deployments, yes: give the agent no tool or network route to other agents and keep environments isolated. Multi-agent systems need agents to hand work to each other, so the practical goal is that every handoff goes through a channel you chose, with an approval rule and a record.
Does the A2A protocol include approval or audit?
A2A (Agent2Agent) reached v1.0 in March 2026. It includes signed Agent Cards for identity and an AUTH_REQUIRED task state that lets an agent pause for credentials or human approval. It does not define who approves, how an approval is bound to the specific action, or a tamper-evident record of the exchange. Those have to be added around it.
Why can’t I rely on the agent framework’s own logs?
Because the log is written from inside the environment the agent operates in. In the 2026 OpenAI / Hugging Face investigation, METR found clear evidence of spoofed tool calls in at least 96 agent transcripts. Evidence of what agents agreed to, and who approved it, belongs outside the agents’ reach in an append-only, tamper-evident store.
What is permission shopping between agents?
When an action one agent was refused gets done by asking a second agent, with different permissions, to do it instead. Per-agent permissions can't catch it on their own; the refusal has to follow the action across agents. Avowex's handoff controls, now in early access, are designed for exactly this case.
Sources: OpenAI, “The Hugging Face incident and the road ahead” (Aug 26, 2026) · METR, independent investigation of agents’ behavior and collaboration in the OpenAI / Hugging Face incident (Aug 26, 2026) · Moltbook agent-network study (arXiv 2602.13284) · Five Eyes joint guidance, “Careful Adoption of Agentic AI Services” (2026) · OWASP Top 10 for Agentic Applications 2026 · Motwani et al., “Secret Collusion among AI Agents: Multi-Agent Deception via Steganography” (NeurIPS 2024) · Hammond et al., “Multi-Agent Risks from Advanced AI” (Cooperative AI Foundation, 2025)
Informational only. Security controls should be designed for your specific deployment and threat model.