1. Full Inventory of Every AI Agent in Production
Maintain comprehensive registries encompassing commercial tools (Copilot, ChatGPT Enterprise), custom-built systems, developer projects, and API-driven agents. Unregistered agents with elevated credentials represent critical breach vectors.
2. Every MCP Server Connection Mapped
Model Context Protocol servers function as trust boundaries requiring complete documentation. Each connection should identify owners, purposes, data classifications, and authentication mechanisms. MCP compromise is prompt injection at the infrastructure level.
3. No Shared API Keys Between Agents and Humans
Dedicated credentials per agent enable action attribution and independent access revocation. Shared keys eliminate audit trail distinction between human and agent activities.
4. Agent Memory Stores Audited for Poisoning
Memory poisoning manipulates persistent agent memory—conversation histories, learned preferences, RAG knowledge bases—to alter future behaviour without code modification. Agents inherently trust their own stored information.
5. Data Access AND Exfiltration Paths Documented
Document both data access capabilities and all potential exfiltration channels per agent. Traditional DLP tools don’t account for agents’ ability to summarise, paraphrase, encode, or distribute sensitive information across seemingly innocuous outputs.
6. Human Override Procedures Exist and Tested
Kill-switch mechanisms must function independently of agent operation itself. Actual testing—not theoretical validation—proves override capability. Implement quarterly override drills.
7. Agent-to-Agent Communication Chains Logged
Multi-agent architectures require chain-level logging capturing source, destination, full message content, timestamps, and decision rationale. Orchestration creates emergent behaviours invisible in single-agent logs.
8. AI Agent Infrastructure Red-Teamed at Least Once
Agents’ natural language responsiveness creates unique attack surfaces including social engineering vectors that traditional penetration testing omits. Scope should include prompt injection, tool abuse, memory manipulation, data exfiltration, privilege escalation via agent chains.
Scoring
Need help implementing this?
First strategy session is complimentary. We typically respond within 4 hours.