AI Security . Architecture . Defense in Depth

Designing Defense in Depth for AI Systems

A reference architecture for securing the data, model, orchestration, tool, and output layers of an AI stack - not just the prompt box.

Most teams still secure AI like a website: a filter at the edge and a content check on the output. Real AI systems fail one layer at a time. If the layers aren't isolated from each other, a single compromised layer is enough to reach everything behind it.

Reference stack
Output layer
Tool / action layer
Orchestration layer
Model layer
Data layer
5 Reference layers Data, model, orchestration, tool, output - each a separate trust boundary.
1 layer to fail Without isolation, one compromised layer reaches every layer behind it.
3 Failure domains Confidentiality, integrity, availability - every layer must contain all three.
0 Implicit trust The target state: no layer trusts another layer's output without a check.
AlignmentZero trust, defense in depth BoundariesData · Model · Orchestration · Tool · Output AudienceSecurity & ML platform teams DirectionPolicy-as-code at every boundary
AI security architecture defense in depth zero trust AI threat modeling policy as code

A retrieval pipeline pulls a document into context. The document is fine - until a later step treats it as an instruction instead of a fact. Nobody breached a firewall. Nobody stole a credential. The system just never drew a line between "data I'm reading" and "commands I'll follow," and that missing line is an architecture problem, not a filtering problem. You cannot patch your way out of a boundary that was never drawn.

Most AI security effort still goes into the two most visible layers: a moderation pass on the prompt, and a moderation pass on the output. That is real work, and worth doing. But between those two checkpoints sits a pipeline with its own filesystem access, its own database connections, its own third-party tool calls, and its own memory of the last ten turns. A reference architecture treats every one of those as a boundary that needs its own controls, not a black box that the input and output filters are supposed to cover for.

A security architecture is not a checklist bolted onto a finished system. It is the shape the system has to take before the first line of inference code runs.

Design principle

Why layers, not a perimeter#

A traditional web app has one real trust boundary: the login. Everything past it is roughly equally trusted. An AI system doesn't get that luxury, because the model itself sits in the middle of the request path and can be steered by whatever text reaches it - user input, retrieved documents, tool output, even its own prior turns. Treating the whole pipeline as one trust zone means a single steering trick anywhere in that path can reach everything the pipeline can reach.

The fix is the same one distributed systems learned decades ago: shrink what each component is allowed to do, and check every handoff between components. Applied to an AI stack, that means each layer gets its own permissions, and none of them inherit the permissions of the layer next to them by default.

Data layer
Read access to sources it's scoped to. No write access, no network egress, no execution rights.
Model layer
Produces text and tool-call proposals. Has no direct access to systems of record - it can only ask.
Orchestration layer
Decides which proposed tool calls actually run, against policy - not against the model's confidence.
Tool / action layer
Executes one narrow, named action with validated arguments. Nothing more general than that.
Output layer
Renders to the user. Can redact and refuse, but cannot itself write back into the system.

The five layers, in order#

Walking the stack top to bottom shows where the permissions are supposed to shrink, and what each layer is actually responsible for containing.

01

Data layerIngestion

Everything the system reads before it reasons: documents, tickets, emails, search results, files. This layer's job is provenance - tagging where each piece of content came from and whether it's trusted, so later layers can tell instructions from information.

02

Model layerReasoning

Where inference happens. This layer never holds credentials or direct system access - it can only propose actions in a structured format that the orchestration layer then evaluates.

03

Orchestration layerPolicy

The actual authorization point. Every proposed action passes through a policy check here - tool identity, arguments, and the user's real session scope - before anything is allowed to execute.

04

Tool / action layerExecution

Where side effects happen: a write, a send, a purchase, a deploy. Each tool is scoped to one narrow capability with its own least-privilege identity, sandboxed from the others.

05

Output layerDelivery

What actually reaches the user or the next system downstream. Responsible for redaction, format enforcement, and catching anything upstream layers missed before it leaves the boundary.

The boundary that matters most is the one between the model layer and the orchestration layer. That is the point where a proposal becomes an action, and it's the only place a policy engine can stop a bad one without needing to understand why the model proposed it.

Where the layers actually fail#

The layers don't usually fail because someone finds a novel exploit. They fail because a boundary that looked drawn on a diagram was never enforced in code. Four patterns account for most of the real-world damage.

01

Untagged data becomes instructions Data layer

A retrieved document, ticket, or email is passed into the model's context with no marker distinguishing it from a direct user instruction. The model treats embedded directives in that content as commands, because nothing told it not to.

02

The orchestrator trusts the model's confidence Orchestration layer

A tool call is approved because the model asked for it fluently, not because it passed a policy check against the user's actual scope. Confidence is not authorization, and treating it as such collapses the one real boundary in the stack.

03

One tool identity for every action Tool layer

A single service account with broad database or filesystem rights backs every tool the system exposes. A narrow, legitimate request and a destructive one both run with the same privileges, so containment depends entirely on the model never asking for the wrong thing.

04

The output layer assumes everything upstream was clean Output layer

Internal identifiers, system prompts, or another user's data leak into a response because nothing at the edge re-checks what's about to be rendered. The output layer is treated as a formatter, not a control point.

The defense blueprint#

None of this requires rebuilding the system. It requires making each boundary an actual checkpoint instead of an assumption. Five practices carry most of the weight.

Isolation

Give each layer its own identity

Separate credentials per layer and per tool, scoped to exactly what that component needs - nothing inherited from the layer next to it.

Policy

Authorize at the orchestration boundary

A policy engine, not the model, decides whether a proposed action runs - checked against tool identity, arguments, and session scope.

Provenance

Tag data at ingestion, not later

Every piece of retrieved content carries its source and trust level through the whole pipeline, so downstream layers can tell data from instructions.

Verification

Re-check the output layer independently

Treat the edge as a control point: redact, validate format, and re-run a policy check on anything about to leave the system.

Monitoring

Red-team the boundaries, not just the prompt

Test each handoff between layers on its own - can the tool layer be reached without going through orchestration? Can data smuggle instructions past the tags?

Isolation levels for each layer#

Not every layer needs the same amount of isolation, and over-isolating early slows a team down for no security gain. A simple ladder helps decide how much containment a given layer actually needs.

Isolation ladder, weakest to strongest
L0
Shared Same process, same credentials - fine for pure read-only, low-value data
L1
Namespaced Separate credentials and scopes per tenant or per tool, same runtime
L2
Sandboxed Separate container or process, no shared filesystem, egress denied by default
L3
Air-gapped Separate network boundary, human approval required, break-glass logging

As a starting point: the data and model layers can often run at L1. The orchestration layer, since it's the single authorization point for the whole system, deserves L2. Any tool that executes code, moves money, or touches production deserves L3, regardless of how simple the request that triggers it looks.

# Minimal boundary check, orchestration -> tool layer
action_allowed =
  tool.identity_matches_request &&
  args.schema_valid &&
  session.scope_covers(action) &&
  source.trust_level != "untrusted" &&
  (risk_tier.auto_ok || human_approved)

# If any check is missing, the boundary isn't real yet.

Pre-launch architecture checklist#

Before an AI system with real tool access goes live, each layer should be able to answer the checks below on its own - not "the filter upstream should catch that."

✓
Data provenance taggedEvery ingested source carries an origin and trust level that survives into the model's context.
✓
Model has no direct credentialsThe model layer can only propose actions - it holds no API keys, database access, or shell rights itself.
✓
Policy engine sits at orchestrationEvery tool call is evaluated against identity, arguments, and scope before it runs.
✓
Each tool has its own identityNo shared service account backs two tools of different risk levels.
✓
Output layer re-checks before deliveryRedaction and format validation happen at the edge, independent of upstream checks.
✓
Every boundary is loggedAn immutable record of what crossed each layer boundary, with enough context to reconstruct an incident.

Pro tip: test the boundaries directly, not just the chat interface. If a tool call can be made to run without going through the orchestration layer's policy check, the architecture has a hole no amount of prompt tuning will close.

Ready to build practical AI security architecture skills? Explore training designed to help you design, review, and defend layered AI systems.

AI Security - training & certification

Treat every layer boundary as a place work can be rejected

An AI system isn't secured by wrapping it in a filter at the front and a filter at the back. It's secured by making sure that at no point does one layer simply trust what the layer next to it hands over. That's slower to build than a single moderation pass, and it's the only version that holds up once the system has real tool access.

Three things to take away. Layers need their own identities - a model that can act with the same privileges as the database it's querying has no real boundary at all. Authorization belongs at orchestration - a policy engine checking arguments and scope, not a model deciding it's confident enough. The output layer is a control point, not a formatter - the last chance to catch what every earlier layer missed.

A reference architecture isn't a diagram you draw once. It's the set of checks that still run on the day someone finds the boundary you forgot about.