Artificial Intelligence

AI Guardrails and Safety: How to Keep Production AI From Doing Embarrassing (or Dangerous) Things

Sunil Sethi
Leader, AI & Workflow Specialist
· 28 min

The 4 layers of guardrails every production AI needs, the 3 real threat classes they have to handle, and how to pick the right guardrail tool for your product.

Artificial Intelligence Solutions
Looking for a artificial intelligence partner?
We build domain-led systems tailored to your industry and workflow. 12 years. 2,100+ engagements.
Get in Touch
Related Insights
What the Model Context Protocol (MCP) actually is? The Standard That Is Quietly Replacing Custom Integrations Between AI and Your Tools LangSmith vs Langfuse vs Helicone vs Rolling Your Own: How to See What Your AI Actually Does in Production AutoGen vs CrewAI vs LangGraph vs Custom: How to Pick the Right AI Agent Framework for Your Product

Every AI product team eventually gets the same call. A user got the AI to say something embarrassing. A user got the AI to leak information from another customer's account.

A user got the AI to give advice it should never have given (medical, legal, financial, or worse). The team is now on a war-room call trying to figure out how it happened, what got exposed, whether the damage is public yet, and how to make sure it does not happen again. Every one of these moments could have been prevented by guardrails wired in from day one. Almost every AI product without them ends up on that war-room call, usually within their first year in production.

So how do you actually build AI that stays inside the lines? That is the point of this piece. You will see what guardrails really are without the safety-theatre language, the 4 layers of guardrails every production AI needs, the 3 threat classes your guardrails actually have to handle, the 4 failure modes when guardrails get wired wrong, the 4 guardrail paths your product can pick between (NeMo Guardrails, Guardrails.ai, model-provider built-ins, and custom), and the 3 signs the "safe AI" pitch you are being sold is marketing rather than protection. All of it is written for the product owner who has to answer for the AI's behaviour, not the safety researcher measuring failure rates on benchmarks.

Why does this matter more this year than a few years back? Because production AI has moved past demo and into workloads customers depend on, and the failures your guardrails prevent have become customer-visible failures rather than internal curiosities. Prompt injection attacks have become mainstream: users have realised they can get AI features to do things the product owner never intended.

Regulators in multiple jurisdictions have started asking AI product operators to demonstrate the safety controls they have in place. And the cost of a single guardrail failure that becomes public (a leaked customer record, a harmful piece of advice, a scandalous output screenshot) can outrun a year of the product's revenue.

4
Layers of guardrails every production AI needs: input filters, output filters, tool-use controls, and behaviour policies.
3
Threat classes your guardrails actually have to handle: prompt injection, sensitive-content leakage, and off-topic or harmful advice.
4
Failure modes when guardrails are wired wrong: over-blocking legitimate use, under-blocking real threats, guardrail bypass, and silent failure.
1
Guardrail interface that keeps your product loose from any specific safety tool. Skip it and swapping guardrail tools becomes a rewrite.

The rest of this piece walks the answer in the order the questions come up during a real conversation about AI safety. What are guardrails? Which layers matter?

What threats do they actually face? How do they fail? Which tool fits your product?

And how do you spot marketing dressed as safety? Boring on purpose, because guardrails are boring until the moment they save your product from a story that would have ended up on a competitor's landing page.

AI Guardrails, Broken Down for a Product Owner

What is an AI guardrail? A guardrail is a rule or a check that limits what your AI can accept as input, what it can produce as output, what tools it can call, or how it should behave in specific situations. Guardrails sit around the AI, not inside it; they inspect the traffic before it reaches the model and after the model responds, and they act as the safety layer that decides whether a specific request or response is acceptable. The model itself may be capable of producing all sorts of output; the guardrails decide what actually reaches the user or gets acted on.

Why did guardrails become a separate category? Because AI models are broadly capable and context-blind by default. A general-purpose model will happily attempt any request it receives, including requests that are inappropriate for your specific product.

Your product has a job (customer support, medical guidance, contract review, whatever) and everything outside that job is either irrelevant or actively dangerous. Guardrails are how you tell the AI, in enforceable terms, "you are for this product, not for anything the user might imagine you could do".

What is the difference between guardrails and the model's built-in safety training? Model providers train their models to refuse certain broadly-harmful requests (violence, illegal content, self-harm). That built-in training is real but not enough for a specific product.

Your product needs guardrails that enforce your product's specific rules on top of the provider's baseline: what topics are on-scope for your product, what data is off-limits, what tools the AI is allowed to invoke, what output formats are required. The provider's training is a floor; your guardrails are what makes the AI actually safe for your specific customers to use.

The Guardrail Question

If your team cannot describe, in specific rules, what your AI is allowed and not allowed to say or do, your guardrails are informal at best. Every product AI needs an explicit rule set that a guardrail layer can enforce, not just an implicit understanding shared by the team that built it. The rule set is what makes the difference between "we told the AI not to do that" and "the AI cannot do that".

The 4 Layers of Guardrails Every Production AI Needs

Which layers of guardrails does a production AI actually need to run safely? The 4 below cover the surface area of a real product. Skipping any one leaves a specific class of threats unaddressed, and every skipped layer eventually shows up as an incident. The diagram below shows what each layer does and where in the request flow it sits.

4 Guardrail Layers
What Every Production AI Needs, Layer by Layer
Layer 1
Input Filters
Checks user input before it reaches the model. Blocks prompt-injection attempts, off-topic queries, and personally-identifiable information (PII) that should not enter the model context.
Layer 2
Output Filters
Checks the model's response before it reaches the user. Blocks sensitive-content leakage, off-brand language, and any output that violates your product's rules regardless of how the model got there.
Layer 3
Tool-Use Controls
Controls which tools the AI is allowed to call, with what arguments, on whose data. Prevents an AI agent from calling a destructive tool it should never have had access to.
Layer 4
Behaviour Policies
Standing rules the AI must follow regardless of user input: stay on topic, do not give medical or legal advice, always cite sources, escalate specific requests to a human.
Why All 4 Layers Matter
Layer 1 catches malicious input; layer 2 catches unexpected output; layer 3 catches destructive actions; layer 4 shapes ongoing behaviour. Products that install only some of these have specific gaps that specific incidents exploit. The 4 layers together produce an AI that stays inside your product's actual scope, on your product's actual terms.

Why does missing one layer break the whole safety story? Because each layer catches a different failure mode. Input filters cannot prevent the model from generating leaked content it was fed months back through fine-tuning; only output filters catch that.

Output filters cannot prevent an agent from calling a destructive tool; only tool-use controls catch that. Tool-use controls cannot prevent the AI from drifting off-topic across a long conversation; only behaviour policies catch that. Guardrails work as a layered system, not a single gate. Products that treat them as one thing usually secure one thing and leave the rest exposed.

3 Threat Classes Your Guardrails Actually Have to Handle

What kinds of threats do production AI guardrails actually face? The 3 below cover almost every real incident. Each of them has surfaced in a public production AI system, been reported publicly, and caused real damage to the product operator. Knowing them is what turns guardrail design from theoretical exercise into targeted defence.

01
Prompt Injection and Jailbreak Attempts
Users craft input designed to override your product's instructions and get the AI to do something you never intended: reveal the system prompt, ignore its role, execute a tool against another user's data, produce content the product does not allow. This has moved from research curiosity to standard user behaviour; every consumer-facing AI product sees these attempts within their first week in production. Input filters are your first line of defence; output filters and behaviour policies catch the ones that slip through. Products without any of these have no defence and provably get compromised on public demonstration.
02
Sensitive-Content Leakage
The AI produces information it should not have shared: PII from another customer's account, an internal document that was included in retrieval by mistake, a proprietary detail the model absorbed during fine-tuning. Output filters catch the obvious cases (credit card numbers, government IDs, email addresses); tool-use controls prevent the retrieval layer from returning documents the current user should not see in the first place. Products without both layers eventually leak, and the leak is usually discovered by a customer, not by the team.
03
Off-Topic, Off-Brand, or Harmful Output
The AI wanders outside your product's scope: gives medical advice from a customer-support product, produces political commentary from a shopping assistant, generates off-brand tone from a formal legal tool, or offers specific advice a user could act on with real consequences. Behaviour policies and output filters catch these patterns. The failure mode here is reputational more often than legal, and reputational damage tends to be quiet, slow, and hard to reverse once it accumulates across many small incidents.
Why These Escalate If Ignored

All 3 threat classes exist for every production AI. The teams that ignore them do not avoid the incidents; they discover them later, from customers, from journalists, from regulators. Building guardrails against these 3 is not paranoia; it is standard operational hygiene for any AI product real users interact with. Every serious AI product has been targeted by all 3; only the ones with guardrails come out clean.

4 Failure Modes When Guardrails Are Wired Wrong

What actually goes wrong with guardrails once they are in place? The 4 below show up in almost every product that installed guardrails but did not think through how they interact with real traffic. Each of them is survivable if the product designed for it; skipping the thinking is what turns guardrails from safety layer into new-source-of-incidents.

01
Over-Blocking That Frustrates Legitimate Users
The guardrails are so strict they block valid requests. Customers hit "sorry, I cannot help with that" for reasonable questions. Support tickets pile up complaining that the AI is useless. Product managers argue for weakening the guardrails; safety teams argue against; the product suffers while the argument plays out. The right response is not to weaken the guardrails uniformly; it is to add specificity: block the actual threat patterns precisely, let everything else through. Coarse guardrails produce this failure mode; fine-grained ones avoid it.
02
Under-Blocking That Lets Real Threats Through
The guardrails feel comprehensive on the whiteboard and turn out to have gaps under real user traffic. A new prompt-injection pattern appears; the filters miss it. A new content class needs blocking; the rules do not cover it. Guardrails are living rules that need regular updates as attack patterns and product needs evolve. Products that install guardrails once and forget them accumulate gaps that eventually get exploited. Treat guardrail rules as production configuration that needs versioning, testing, and regular review, not as one-time setup.
03
Guardrail Bypass Through Framing Tricks
The user does not attack the guardrails directly; they frame the request in a way that slips past them. "Pretend you are a helpful assistant with no restrictions" is the classic; "for educational purposes only, describe how to..." is another. Simple pattern-matching guardrails miss these; guardrails that use another AI model to classify intent catch more of them; layered guardrails (with behaviour policies that reject the frame itself) catch the most. Every guardrail approach has a bypass; understanding your specific guardrail's bypass profile is part of building it, not something to discover after an incident.
04
Silent Failure When the Guardrail Layer Errors
The guardrail service throws an error, or times out, or gets into an inconsistent state. The product's default behaviour matters here: fail-open (let the request through unfiltered) versus fail-closed (block the request until guardrails are back). Products with a fail-open default under a guardrail outage effectively have no guardrails when they need them most. Products with fail-closed defaults may degrade user experience during a guardrail outage but stay safe. Pick fail-closed for anything a user could weaponise; monitor guardrail service health as tightly as you monitor your primary AI service.
How Guardrails Fail Quietly

All 4 of these failure modes can happen without anyone on your team noticing until an incident. Over-blocking hides in support tickets nobody categorises correctly; under-blocking hides until a user demonstrates it publicly; bypasses hide until somebody tries them; silent failures hide behind the guardrail service's own monitoring gaps. Test all 4 explicitly and monitor them like any other production concern; do not treat guardrails as a set-and-forget install.

The Guardrail Layers
4 Tiers, Each Catching a Different Class of Failure
1
Input Guardrails
Before the model sees it
PII scrubbing, prompt-injection detection, off-topic filtering, rate limits. Catches problems before compute is spent.
2
Model Guardrails
While the model runs
System prompts, tool-use policies, constrained generation, retrieval scoping. Shapes what the model is allowed to attempt.
3
Output Guardrails
Before the user sees it
PII in output, unsafe content, factual grounding checks, brand-voice validation, format enforcement. Last chance before the response reaches the user.
4
Post-Facto Guardrails
After the fact, always on
Sampling, audit logging, drift detection, user-report loops, weekly review of flagged outputs. Catches what the first 3 tiers missed.
Why All 4 Tiers Together
Each tier catches a different failure. Input guardrails alone leak; output guardrails alone waste compute; model guardrails alone miss edge cases; post-facto alone catches the failure after the user did. Real safety is layered, not any single tier taken very seriously.

Which of the 4 Guardrail Paths Fits Your Product?

Which of the 4 guardrail paths does your product actually need? Almost every "which guardrail tool should we use" conversation resolves into one of them once you push on your product's specific safety requirements. Knowing which one you are in changes the integration effort, the flexibility you get, and how much of the guardrail logic you own.

4 Guardrail Paths
What "Which Guardrail Tool Should We Use" Actually Turns Out to Mean
Path 1
NeMo Guardrails
NVIDIA's open-source guardrail toolkit. Programmable dialogue flows and safety rules. Best for products with complex conversation shaping needs, where you want to define detailed behaviour policies and conversational scripting. Steeper learning curve, powerful once you learn it.
Path 2
Guardrails.ai
Open-source library focused on validating and structuring AI output. Strong on output filtering, format enforcement, and rule-based checks. Best for products where output shape and content compliance are the primary concern. Easier to get started with than NeMo; less prescriptive on conversational flow.
Path 3
Model-Provider Built-Ins
The safety features baked into major providers (OpenAI Moderation, Anthropic's constitutional AI, Google's Safety Settings, AWS Bedrock Guardrails). Best for products where the provider's baseline safety plus a few configuration knobs is enough. Least effort, least flexibility, tied to the provider.
Path 4
Custom Guardrail Layer
Build the guardrail layer yourself, combining smaller pieces (PII detection libraries, classification models, rule engines) into your specific product's safety posture. More upfront work, most control forever, best fit for products with unique safety needs not covered by any single tool.
Which Path Fits Your Product
Products with complex conversational flows lean toward NeMo. Products focused on output validation and structured checks lean toward Guardrails.ai. Products with straightforward safety needs where provider defaults plus config are enough lean toward built-ins. Products with unique safety requirements or high regulatory demands lean toward custom. Most serious products end up combining 2 or 3 paths through the guardrail interface pattern.

Why does the guardrail path matter so much before you build? Because each path has a different implicit safety model and matching the wrong one to your product produces either safety theatre (looks like guardrails, does not actually catch your threats) or friction theatre (blocks legitimate use in ways your users notice more than any threat). The path that fits your specific threat profile and your specific product shape is what makes guardrails useful; picking based on popularity produces guardrails that neither block the right things nor let the right things through.

3 Signs the "Safe AI" Pitch You Are Being Sold Is Marketing

How do you tell whether a guardrail tool being pitched will actually keep your product safe, or whether it will produce a "safety" checkbox nobody looks at? The 3 signs below are the ones that keep separating real safety tools from marketing dressed as safety. If you spot more than one, the tool is probably not going to survive real attack traffic.

01
The Demos Show General Safety Categories, Not Product-Specific Rules
The vendor shows the tool blocking violence, hate speech, and self-harm content. Great, but every provider's built-in safety catches those. What your product actually needs is guardrails against your product's specific threats: your specific PII shapes, your specific off-topic patterns, your specific tool-use constraints. If the demo cannot show custom rules addressing product-specific concerns, the tool is a general-safety wrapper and provides little beyond what the provider already includes.
02
There Is No Story for Prompt-Injection Defence
Ask the vendor what happens when a user tries to jailbreak the AI or inject instructions that override your product's rules. A serious guardrail tool has explicit answers: input pattern detection, intent classification, behaviour policies that survive framing tricks. A marketing-driven tool says "our AI has strong safety training" or "we detect harmful content", which does not address prompt injection at all. Prompt injection is the most common real attack; a guardrail tool without a story for it is a guardrail tool that does not defend against your most common threat.
03
The Tool Cannot Be Evaluated Against Your Real Attack Traffic
You ask to test the tool with a set of adversarial prompts drawn from your product's actual (or anticipated) attack patterns. The vendor either cannot run this test or does not want to. A serious safety tool encourages red-team testing on real attack samples; a marketing tool discourages it because the results would be embarrassing. If you cannot evaluate the tool on realistic threat traffic before buying, the "safety" claim is unverified, and unverified safety in production is not safety.
The Marketing Filter

Ask the vendor 3 things in the same meeting: show me custom rules that address my product's specific threats; walk me through your prompt-injection defence; and let me test the tool against 20 adversarial prompts I bring to the meeting. Serious guardrail vendors answer all 3 confidently. Marketing-first vendors deflect with adjectives about safety.

Frequently Asked Questions

Are the model provider's built-in safety features enough for most products?
For internal-only tools or products with minimal safety exposure, sometimes yes. For any customer-facing product where users could attempt prompt injection, where sensitive data flows through the AI, or where the output could damage your brand or expose regulatory risk, provider built-ins are a floor, not a full safety layer. The built-ins catch the broad harmful categories the provider trained for; they do not know your product's specific rules, your product's specific data, or your product's specific threats. Add your own layer on top; the provider's baseline is necessary but not sufficient.
How do you prevent prompt injection specifically?
Layer defence, no single technique. Separate user input from system instructions clearly in your prompts so the model treats them differently. Add input filters that detect common injection patterns. Use another AI model to classify user intent before letting the request reach the main model. Enforce behaviour policies at the output layer so even a compromised prompt produces filtered output. Test against a red-team library of known injection patterns regularly, and update your defences as new patterns emerge. No single control catches everything; the layered approach catches most and reduces the impact of the ones that slip through.
What happens if a guardrail incorrectly blocks a legitimate user request?
Your product should recognise false positives are a real user-experience cost and design for them. Give the user a clear message explaining the block (without revealing the exact rule, which would help attackers refine bypasses), offer a way to rephrase or escalate to a human, log the block for review, and use the log to iterate on the rules. Products that respond to a false positive with a generic "sorry, I cannot help with that" and no path forward frustrate users into leaving. Products with an escalation path and honest logging use false positives as feedback to make the guardrails more precise over time.
Do guardrails add meaningful latency to AI responses?
Depends on which guardrails and how they run. Pattern-based input filters and rule-based output checks add negligible latency. Guardrails that use another AI model to classify input or evaluate output add real latency, sometimes doubling the effective response time. For products where latency matters, run the model-based guardrails in parallel with the primary call where possible, and reserve them for the highest-risk paths. For products where safety matters more than milliseconds, the extra latency is usually the right trade. Measure the latency budget for guardrails explicitly, and pick which layers can afford which techniques.
Can you use multiple guardrail tools together?
Yes, and most serious products do. Use provider built-ins as the baseline, add Guardrails.ai or a custom output validator for structured checks, add NeMo or a custom behaviour policy for conversation-level shaping, and layer PII detection at the retrieval and output boundaries. The guardrail interface pattern lets each layer plug in cleanly. Products that pick one guardrail tool and stop usually leave specific threats uncovered; products that layer 2 or 3 tools get much better coverage without much extra work, especially if the guardrail interface is in place from day one.
Should guardrails fail open or fail closed?
Fail closed for anything a user could weaponise. Fail open only for the least-risk paths where an outage of the guardrail service means the AI briefly runs without a specific safety check. The default posture matters most during a guardrail outage: fail-open products effectively have no guardrails at exactly the moment they need them most; fail-closed products degrade user experience but stay safe. Monitor guardrail service health as tightly as you monitor your primary AI service, and design the failure mode explicitly rather than accepting whatever the tool defaults to.
Can Entexis design and build the right guardrail setup for your AI product?
Yes. Entexis designs and builds AI guardrail layers across NeMo Guardrails, Guardrails.ai, model-provider built-ins, and custom setups, and the guardrail interface pattern that keeps your product loose from any one tool. That work starts with the threat-model conversation to identify your product's specific threat classes (prompt injection surface, sensitive data flowing through the AI, off-topic risk, tool-use exposure). We then design the 4-layer guardrail structure, wire the specific tools for each layer, run red-team testing against realistic attack samples, and integrate guardrail health into your observability so silent failures surface fast. Reach out with what your AI product does, who uses it, and any regulatory or brand risk that a bad output would carry, and we can walk through the right guardrail setup for your specific product.

For the observability layer that lets you see when guardrails catch or miss threats in production, see: LangSmith vs Langfuse vs Helicone vs Rolling Your Own.

For the audit trail that regulators will eventually ask for on your AI's safety controls, see: The AI Audit Trail Every CFO Will Ask For.

For the orchestration framework decision that shapes where and how guardrails plug into your product, see: LangChain vs LlamaIndex vs DSPy vs Custom.

So where does that leave your AI safety posture? The 4 guardrail layers above are what every production AI needs; skipping any one leaves a specific threat class unaddressed and the incident eventually finds it. The 3 threat classes above are the ones every real product faces regardless of intent; they show up in every consumer AI product within the first week.

The 4 failure modes above are how guardrails misbehave once installed; each is avoidable with explicit design.ai for output validation, provider built-ins for baseline, custom for unique requirements. Wire the layers from day one, pick the paths that match your threat profile, and put the guardrail interface in place so you can layer or swap tools as the threat and tool landscape moves. Skip the layers and the incident is a matter of when, not if.

Want AI That Does Not Embarrass You in Production?

At Entexis, we design and build AI guardrail layers across NeMo, Guardrails.ai, provider built-ins, and custom setups, layered together through the guardrail interface pattern so your product is loose from any one tool. We start with the threat-model conversation to identify your product's specific risks, design the 4-layer guardrail structure that covers input, output, tool-use, and behaviour, wire the tools for each layer, and run red-team testing against realistic attack traffic before you launch. Your product stays inside its scope, your team catches guardrail failures before your users do, and your safety posture holds up when the regulator asks. Start the conversation with Entexis.

Ready to Add AI
to Your Business?

From intelligent chatbots to workflow automation, we build AI solutions that understand your domain, your data, and your users. Tell us what you need.

We'll get back within one business day.

Keep Reading

Related
Insights

All Insights
Artificial Intelligence

What the Model Context Protocol (MCP) actually is? The Standard That Is Quietly Replacing Custom Integrations Between AI and Your Tools

What the Model Context Protocol actually is, which products benefit from adopting it now, and how to adopt without betting on the still-evolving standard.

Read More
Artificial Intelligence

LangSmith vs Langfuse vs Helicone vs Rolling Your Own: How to See What Your AI Actually Does in Production

How to see what your AI actually does in production, and how to pick between LangSmith, Langfuse, Helicone, and building on top of your existing observability.

Read More
Artificial Intelligence

AutoGen vs CrewAI vs LangGraph vs Custom: How to Pick the Right AI Agent Framework for Your Product

How to pick between AutoGen, CrewAI, LangGraph, and a custom agent loop for your AI agent product, plus what breaks agents in production.

Read More
What We Build

Solutions We Deliver

Entexis Labs · Live demos

Try the AI workflows we build, for real, right now.

Same workflow patterns Entexis rolls into client setups. Try them in your browser, no signup. If one feels like it'd help your team, we build a private version tuned to your data.

See It in Action

Related Case
Studies