Blog

I thought the policy was enough. The guardrail proved otherwise.

The difference between knowing the rules of the road and having the barricades in place is the entire distance between AI governance and AI guardrails.

Samip Shah Sep 29, 2026 5 min read
ai-governance ai-guardrails generative-ai-safety output-validation responsible-ai
Close-up of hands typing on a laptop keyboard at a desk, shot with shallow depth of field
Photo by Thomas Lefebvre via Wikimedia Commons (CC0).

The phone Gemini recommended came with model numbers, specs, and a price comparison. It never existed.

I was planning to buy a new phone. I gave Gemini my requirements and intended use and asked it for the best deal. It came back confident: model number, full specifications, price comparisons across three retailers. Every detail sounded authentic. I sat there, no red flags anywhere. I went back to the original source to check stock. No trace. The phone had never been manufactured. The AI had predicted the most plausible-sounding recommendation and delivered it with the same confidence it uses for a verified fact. That is what a system without an AI output validation layer looks like: not deception, just a confident prediction that diverged from reality.

A fabricated phone deal costs me an afternoon; a fabricated output in a live pipeline costs something harder to undo.

The phone deal was low stakes. Wrong model, wasted afternoon, nothing downstream. But the mechanism that produced that fabrication operates inside every generative AI pipeline built today. When the output feeds a clinical data workflow, a financial reporting system, or a pharma documentation pipeline, it does not stay on one person's screen. It becomes the input for the next step. The question shifts from "did the model get the phone right?" to "what sits between the model's output and the process it is feeding?" Generative AI safety has moved from an academic question to an engineering requirement precisely because that mechanism is identical at every scale. The stakes differ; the mechanism does not. That is why AI guardrails are an engineering decision, not a compliance checkbox.

Knowing the rules of the road is governance; having the barricades in place is a guardrail.

Think about driving a mountain road. The speed limit sign, the lane markings, the signage at the hairpin bends: that is AI governance. Policy on paper. You can know every rule, believe every rule, and still go off the edge without the barricades. The barricades do not ask whether you know the rules. They execute.

The difference between AI governance and guardrails is exactly that: governance is the organisation's stated policy for how AI should behave; a guardrail is the code that enforces that policy every time a request comes in or a response goes out. Governance tells you what should happen. A guardrail is what makes it happen.

A guardrail is an algorithm that determines what gets through, not a policy document that asks nicely.

Understanding how AI guardrails work in enterprise pipelines starts with a precise technical function. A 2024 paper on building guardrails for large language models puts it plainly: Guardrails, which filter the inputs or outputs of LLMs, have emerged as a core safeguarding technology. Not an optional add-on. Not a compliance decoration. A core safeguarding technology.

The same research gives the algorithmic definition: A guardrail is an algorithm that takes as input a set of objects (e.g., the input and/or the output of LLMs) and determines if and how some enforcement actions can be taken. Read that carefully. It takes input. It makes a determination. It triggers an action or suppresses one. That is code running at runtime, not a policy document sitting in a shared folder waiting to be consulted.

For LLM safety in production, a single output filter is not what enterprise implementation looks like. Anthropic's Responsible Scaling Policy describes a multi-layered approach to prevent misuse, including real-time and asynchronous monitoring, rapid response protocols, and thorough pre-deployment red teaming. Those three layers are not options to pick from; they are three mechanisms that operate independently because any single one can be circumvented or miss a class of outputs it was not designed to catch. Real-time monitoring catches what slips through. Pre-deployment red teaming finds the gaps before they become incidents.

This is where the governance-vs-guardrails distinction sharpens into something practical. Governance is the policy the organisation approved. Guardrails are where that policy executes in code. The gap between the two is not a compliance gap. It is an engineering gap.

Training a model on good values matters, and that is only half of what makes a production system safe.

The strongest objection is worth taking seriously. If a model has been trained with aligned values and good judgment built in, why layer runtime enforcement on top of it? It is a genuine argument. The people building frontier models spend serious effort on this. The goal is a model that chooses the right action, not merely one that is stopped from taking a wrong one. The researchers building frontier models hold that preference - the harder problem is cultivating judgment, not writing rules.

Concede it: values alignment is real and it matters. Now resolve the tension. Trained values and runtime guardrails address different things. Training shapes what the model tends to do across millions of interactions. A guardrail enforces what it must not do, independent of what it was trained to prefer. The model you deploy today will have updated weights next quarter. The runtime guardrail operates independently of those weights. One layer does not replace the other. They are complementary by design, and treating them as substitutes is the category error this post is trying to correct.

The rules I work by today did not come from a policy deck; they came from the moment I learned a policy was not enough.

I stay behind the wheel and ask generalised queries. I never paste sensitive customer data, IRB submission IDs, or other regulated material into a cloud model, that rule is non-negotiable in pharma IT.

Those two sentences are not written anywhere. No policy document handed to me at onboarding. I execute them every single time I open a model. Which makes them a guardrail layer, not in code, but in practice. The lesson from the architecture above applies at the personal level too: the rule does not protect anything if it only exists on paper.

The guardrail treated as finished is the one that gets walked around first.

The failure mode to watch for is this: someone reads this post, understands that guardrails are engineering, builds one output filter, and considers the work done. Anthropic's research on many-shot jailbreaking found that The jailbreak is disarmingly simple, yet scales surprisingly well to longer context windows. A static guardrail built for one context window size and one prompt pattern is already dated by the time the model's context window expands. The attack surface moves. The guardrail does not, unless someone maintains it.

The architectural commitment is not to install the barricade once. It is to treat it as infrastructure that needs to be maintained, red-teamed, and updated as the model and the environment around it evolve.

Infrastructure, not installation.

The next time your organisation talks about deploying an AI feature, you will know which question to ask.

Not "does the organisation have a guardrails policy?" but "where in the stack does that policy execute in code?" That question separates governance from guardrails. It belongs in the room every time a generative AI feature gets signed off.

← All posts