Guardrails and Prompt Injection Defense

How to build LLM applications that stay safe despite prompt injection — the defensive companion to red-teaming. Why prompt injection is unsolved (instructions and data share one channel; contain, don't prevent), a taxonomy of attacks (direct vs indirect injection, jailbreaks, the confused-deputy problem), input defenses and their hard limits (structural validation helps; filtering can't — indirect injection bypasses it), prompt hardening (delimiting, spotlighting, instruction hierarchy — probabilistic, not a fence), architecture and least privilege (the load-bearing defense: scoped tools, no privilege inheritance, human-in-the-loop, the dual-LLM pattern), treating model output as untrusted (XSS/SQLi/SSRF/exfiltration via output, context-specific encoding, structured output), guardrails in practice (moderation, injection classifiers, PII detection, layered defense in depth, fail-safe), and evaluating/operating guardrails (red-teaming your own system, production monitoring, NIST AI RMF, the honest state of the art). Grounded in the OWASP LLM Top 10, indirect-injection research, and the dual-LLM pattern.

8 parts · written by Pratik Dhanave. Start with Part 1 →

← All series · All posts

Part 1 · ·6 min read

Why Prompt Injection Is Unsolved

Prompt injection is the defining security problem of LLM applications, and — unlike SQL injection, which it superficially resembles — it has no clean fix. The reason is structural: a language model reads instructions and data through the same channel, and cannot reliably tell which is which. This opening post explains why that makes injection fundamentally hard, and reframes the goal from "prevent it" to "contain the blast radius."

Prompt injection is the defining security problem of LLM applications, and unlike SQL injection it has no clean fix — because a model reads instructions and data through the same channel and can't reliably tell them apart. This opening post explains why that makes injection fundamentally hard, and reframes the goal from prevent to contain the blast radius.

Part 2 · ·6 min read

A Taxonomy of Injection and Jailbreak Attacks

You can't defend against what you can't categorize. Prompt-based attacks come in a few structurally distinct shapes — direct injection, indirect injection through content the model reads, and jailbreaks that target the model's safety training — and each demands a different defense. This post maps the attack surface so the rest of the series can defend it systematically.

You can't defend what you can't categorize. Prompt-based attacks come in structurally distinct shapes — direct injection, indirect injection through content the model reads, and jailbreaks targeting the model's safety training — and each demands a different defense. This post maps the attack surface so the rest of the series can defend it systematically.

Part 3 · ·5 min read

Input Defenses and Their Limits

The first instinct when facing prompt injection is to inspect the input and block the bad stuff. It's a reasonable layer — but a treacherous one, because it creates a feeling of safety far larger than the protection it provides. This post covers the input-side defenses that are genuinely worth having, and draws a hard line around what they can and cannot do.

The first instinct against injection is to inspect the input and block the bad stuff. It's a reasonable layer but a treacherous one — it creates a feeling of safety far larger than the protection it provides. This post covers the input defenses genuinely worth having, and draws a hard line around what they can't do (starting with: indirect injection bypasses them entirely).

Part 4 · ·5 min read

Prompt Hardening and Its Limits

Between filtering the input and re-architecting the system sits a tempting middle ground: make the prompt itself more resistant. Delimiters, spotlighting, instruction placement, and defensive system prompts all raise the cost of an attack. None of them close the hole — because they are all still text in the one channel the attacker also writes to — but used well they meaningfully shift the odds.

Between filtering input and re-architecting the system sits a tempting middle ground: make the prompt itself more resistant. Delimiters, spotlighting, instruction placement, and defensive system prompts all raise the cost of an attack — but none close the hole, because they're all still text in the one channel the attacker also writes to. The value is knowing exactly how much they buy.

Part 5 · ·6 min read

Architecture and Least Privilege: The Real Defense

Everything before this post raised the probability barrier against injection. This post lowers the impact — and impact is what actually protects you. The load-bearing defense against prompt injection isn't a prompt or a filter; it's an architecture where a fully-hijacked model still can't do anything catastrophic, because it was never granted the power to.

Everything before this raised the probability barrier against injection; this post lowers the impact — and impact is what actually protects you. The load-bearing defense isn't a prompt or a filter but an architecture where a fully-hijacked model still can't do anything catastrophic: least privilege, the confused-deputy trap, human-in-the-loop, and the dual-LLM pattern.

Part 6 · ·6 min read

Treating Model Output as Untrusted

Injection defense usually focuses on what goes into the model. But an equally dangerous class of bug lives on the way out: whatever the model produces gets passed to another system — a browser, a shell, a database, another service — that trusts it. If the model can be made to emit a malicious payload, and your code renders or executes it, the injection escapes the model and lands in your infrastructure.

Injection defense usually focuses on what goes into the model, but an equally dangerous class of bug lives on the way out: whatever the model produces gets passed to another system that trusts it. If the model can be made to emit a malicious payload and your code renders or executes it, the injection escapes the model and lands in your infrastructure — XSS, SQLi, SSRF, exfiltration.

Part 7 · ·6 min read

Guardrails in Practice

"Guardrails" is the umbrella term for the runtime checks that sit around a model and screen what goes in and comes out — content classifiers, moderation models, topic and format validators, PII detectors. They're a real and useful layer, distinct from the architectural defenses, and they come with their own design rules: layer them, fail safe, and never mistake them for a wall.

"Guardrails" is the umbrella term for the runtime checks around a model — content classifiers, moderation models, topic and format validators, PII detectors. They're a real and useful layer, distinct from architectural defenses, with their own design rules: layer them, fail safe, and never mistake them for a wall. This post covers building the layered defense in practice.

Part 8 · ·6 min read

Evaluating and Operating Guardrails

A defense you haven't tested is a hope, not a control. The final discipline of LLM security is treating your guardrails as a system to be measured, attacked, and monitored continuously — because the threat evolves, your application changes, and a defense that worked last quarter can silently rot. This closing post covers red-teaming your own system, operating it in production, and the honest state of the art.

A defense you haven't tested is a hope, not a control. The final discipline of LLM security is treating your guardrails as a system to be measured, attacked, and monitored continuously — because the threat evolves, your application changes, and a defense that worked last quarter can silently rot. Red-teaming, production monitoring, and the honest state of the art.

This series is part of a larger body of work by Pratik Dhanave, an Agentic AI Architect writing about production AI systems, distributed systems, and cloud-native engineering. Explore all course series, browse every post, or find topics via the tag index.