A retrieval pipeline pulls a document into context. The document is fine - until a later step treats it as an instruction instead of a fact. Nobody breached a firewall. Nobody stole a credential. The system just never drew a line between "data I'm reading" and "commands I'll follow," and that missing line is an architecture problem, not a filtering problem. You cannot patch your way out of a boundary that was never drawn.
Most AI security effort still goes into the two most visible layers: a moderation pass on the prompt, and a moderation pass on the output. That is real work, and worth doing. But between those two checkpoints sits a pipeline with its own filesystem access, its own database connections, its own third-party tool calls, and its own memory of the last ten turns. A reference architecture treats every one of those as a boundary that needs its own controls, not a black box that the input and output filters are supposed to cover for.
A security architecture is not a checklist bolted onto a finished system. It is the shape the system has to take before the first line of inference code runs.
Design principleWhy layers, not a perimeter#
A traditional web app has one real trust boundary: the login. Everything past it is roughly equally trusted. An AI system doesn't get that luxury, because the model itself sits in the middle of the request path and can be steered by whatever text reaches it - user input, retrieved documents, tool output, even its own prior turns. Treating the whole pipeline as one trust zone means a single steering trick anywhere in that path can reach everything the pipeline can reach.
The fix is the same one distributed systems learned decades ago: shrink what each component is allowed to do, and check every handoff between components. Applied to an AI stack, that means each layer gets its own permissions, and none of them inherit the permissions of the layer next to them by default.
The five layers, in order#
Walking the stack top to bottom shows where the permissions are supposed to shrink, and what each layer is actually responsible for containing.
Data layerIngestion
Everything the system reads before it reasons: documents, tickets, emails, search results, files. This layer's job is provenance - tagging where each piece of content came from and whether it's trusted, so later layers can tell instructions from information.
Model layerReasoning
Where inference happens. This layer never holds credentials or direct system access - it can only propose actions in a structured format that the orchestration layer then evaluates.
Orchestration layerPolicy
The actual authorization point. Every proposed action passes through a policy check here - tool identity, arguments, and the user's real session scope - before anything is allowed to execute.
Tool / action layerExecution
Where side effects happen: a write, a send, a purchase, a deploy. Each tool is scoped to one narrow capability with its own least-privilege identity, sandboxed from the others.
Output layerDelivery
What actually reaches the user or the next system downstream. Responsible for redaction, format enforcement, and catching anything upstream layers missed before it leaves the boundary.
The boundary that matters most is the one between the model layer and the orchestration layer. That is the point where a proposal becomes an action, and it's the only place a policy engine can stop a bad one without needing to understand why the model proposed it.
Where the layers actually fail#
The layers don't usually fail because someone finds a novel exploit. They fail because a boundary that looked drawn on a diagram was never enforced in code. Four patterns account for most of the real-world damage.
Untagged data becomes instructions Data layer
A retrieved document, ticket, or email is passed into the model's context with no marker distinguishing it from a direct user instruction. The model treats embedded directives in that content as commands, because nothing told it not to.
The orchestrator trusts the model's confidence Orchestration layer
A tool call is approved because the model asked for it fluently, not because it passed a policy check against the user's actual scope. Confidence is not authorization, and treating it as such collapses the one real boundary in the stack.
One tool identity for every action Tool layer
A single service account with broad database or filesystem rights backs every tool the system exposes. A narrow, legitimate request and a destructive one both run with the same privileges, so containment depends entirely on the model never asking for the wrong thing.
The output layer assumes everything upstream was clean Output layer
Internal identifiers, system prompts, or another user's data leak into a response because nothing at the edge re-checks what's about to be rendered. The output layer is treated as a formatter, not a control point.
The defense blueprint#
None of this requires rebuilding the system. It requires making each boundary an actual checkpoint instead of an assumption. Five practices carry most of the weight.
Give each layer its own identity
Separate credentials per layer and per tool, scoped to exactly what that component needs - nothing inherited from the layer next to it.
Authorize at the orchestration boundary
A policy engine, not the model, decides whether a proposed action runs - checked against tool identity, arguments, and session scope.
Tag data at ingestion, not later
Every piece of retrieved content carries its source and trust level through the whole pipeline, so downstream layers can tell data from instructions.
Re-check the output layer independently
Treat the edge as a control point: redact, validate format, and re-run a policy check on anything about to leave the system.
Red-team the boundaries, not just the prompt
Test each handoff between layers on its own - can the tool layer be reached without going through orchestration? Can data smuggle instructions past the tags?
Isolation levels for each layer#
Not every layer needs the same amount of isolation, and over-isolating early slows a team down for no security gain. A simple ladder helps decide how much containment a given layer actually needs.
As a starting point: the data and model layers can often run at L1. The orchestration layer, since it's the single authorization point for the whole system, deserves L2. Any tool that executes code, moves money, or touches production deserves L3, regardless of how simple the request that triggers it looks.
action_allowed =
tool.identity_matches_request &&
args.schema_valid &&
session.scope_covers(action) &&
source.trust_level != "untrusted" &&
(risk_tier.auto_ok || human_approved)
# If any check is missing, the boundary isn't real yet.
Pre-launch architecture checklist#
Before an AI system with real tool access goes live, each layer should be able to answer the checks below on its own - not "the filter upstream should catch that."
Pro tip: test the boundaries directly, not just the chat interface. If a tool call can be made to run without going through the orchestration layer's policy check, the architecture has a hole no amount of prompt tuning will close.
Ready to build practical AI security architecture skills? Explore training designed to help you design, review, and defend layered AI systems.
AI Security - training & certification