Every hyperscaler now ships a managed agent runtime — Azure AI Foundry Agent Service, AWS Bedrock AgentCore, GCP Vertex AI Agent Engine. What none of them ships is an architecture: the end-to-end picture of how network, identity, guardrails, runtime isolation, tool brokering, and data governance fit together around an agent, and what to buy when native controls run out.

So I built one.

Open the full reference architecture →

What it is

Agentic Defense-in-Depth is a tri-cloud security reference architecture for AI agents, organized as seven layers in the order an attacker meets them:

  1. Network isolation & egress control — the agent is unreachable from the internet and can reach only an allowlist. VNet + Private Link, PrivateLink + gateway-only invocation, VPC Service Controls perimeters (now with agent-identity-aware ingress/egress rules).
  2. Identity & access — the agent as a non-human identity — every agent is a governed, revocable identity with short-lived tokens and zero static secrets: Entra Agent ID, AgentCore Identity, GCP Agent Identity, plus the NHI-governance layer above them.
  3. AI gateway & runtime guardrails — one mandatory inspection point for every prompt in and response out: Prompt Shields, Bedrock Guardrails, Model Armor floor settings.
  4. Agent runtime & session isolation — blast-radius containment: microVM-per-session, sandboxed code execution, human-in-the-loop gates.
  5. Multi-agent communication (A2A) — peer agents as separate trust zones: signed AgentCards, mTLS, per-request JWTs, schema validation, no free-text instruction passthrough.
  6. The tool & MCP layer — the agent’s hands, brokered through a gateway with per-tool least privilege and supply-chain vetting of every MCP server.
  7. The data plane — security-trimmed retrieval so the agent can only read what the calling user could, memory under fine-grained access control, customer-managed keys throughout.

Plus the three cross-cutting planes that decide whether any of it holds: detection & incident response (agent-aware runbooks, a kill switch you actually drill), resilience & cost control (denial-of-wallet quotas, loop circuit breakers, restorable memory), and posture & discovery (AI-SPM — finding the agents nobody registered).

Why a reference architecture, not a framework

Frameworks tell you what risks exist. An architecture tells you where each control physically sits, which platform feature implements it, and what to deploy when the platform runs out. Every layer in this one lists native controls for all three clouds side by side with best-of-breed third-party options, and the whole thing is crosswalked to OWASP (LLM, Agentic, MCP, and NHI Top 10), MITRE ATLAS, NIST AI RMF, and DASF 3.0 — the mappings auditors and customers actually ask for.

It closes with a build order: the first ten moves, sequenced by cost relative to the class of incident each one removes. Private networking and secretless identity come first; everything else is easier if you never have to retrofit those.

Explore the architecture →

Discussion

Comments are powered by Giscus / GitHub Discussions. They appear here once configured — see Configure Giscus in the project README and update GISCUS in src/consts.ts.