The control that actually holds is not a guardrail model at all: it is a permission check written
in code, run no matter what anything upstream decided.
Safety’s own example is that check (_permitted,
walked through on that page), refusing a refund unless the customer’s own message names the same
amount the tool call asks for, at the order this conversation is already about. Nothing here
repeats that walkthrough; read it there.
What this page shows instead is the piece that runs before the model ever sees untrusted content:
delimiting it, so a retrieved note cannot pass as the system’s own instructions:
examples/safety/run.py · lines 54–59
def _delimited(note: str) -> str:
"""Marks the boundary of untrusted content in the prompt. Delimiting is a hint to the model,
not a guarantee: the permission check below is what actually stops an unauthorized action,
the way this page's Use it lane and rag's own "prompt injection through retrieved text"
failure mode both say."""
return f'<retrieved-note untrusted="true">\n{note}\n</retrieved-note>'
This is an input-side guardrail in the plainest sense, and it is a hint, not a guarantee: a
model can still be talked into following text it was told is data. That is exactly why the code
check after it exists at all: a guardrail lowers the odds an attack works; it does not remove the
need for a control that holds even when the guardrail does not.
Classifier-model guardrails are a different mechanism from a code-side permission check, and a
different one from delimiting: a separate model, trained specifically to say whether content is
safe. Meta’s own model card describes Llama Guard 4 as “a natively multimodal safety classifier”,
a model distinct from the one it is checking, which “can be used to classify content in both LLM
inputs (prompt classification) and in LLM responses (response classification)”[5]: a
second opinion from a second model, which can itself be wrong, rather than a deterministic check
on a specific fact the way _permitted is.
TypeSafe AI’s Jev, announced in September 2026 and, in the maker’s own words, “available today in
early access”[6], returns a typed, probabilistic verdict instead of text. Its own pitch
names this directly: “Score, judge, verify, guardrail, and detect jailbreaks of LLM prompts,
reasoning traces, and/or outputs”[6]. The claims are the maker’s own and untested here.
Schema and structured-output checks are the guardrail form aimed at a different failure: not
“is this unsafe” but “is this well-formed.” Guardrails AI’s own README describes its two
functions as running “Input/Output Guards” that “detect, quantify and mitigate the presence of
specific types of risks”, and separately helping “generate structured data from LLMs”[3]: a schema check catches a malformed answer a safety classifier would never flag, because
nothing about a badly-shaped JSON object is unsafe, only wrong.
A product aimed at the whole request and response pair, rather than one side of it, is AI
Guardrails (Lakera Guard). Its documentation says Check Point AI Guardrails “screens user and
external content going into LLMs and the resulting output, detecting any threats and providing
real time protection for your GenAI application and users”[4]. Detecting and blocking
are separate settings, and the same page is careful about it: “flagged will always be false if
your project is in Detect mode”, and its tutorial only blocks an interaction after a reader turns
on “Simulate blocking” in the demo chatbot[4]. It was Lakera Guard; Check Point
publishes the documentation as AI Guardrails today, worth searching for under either name.
A guardrail like this is also, itself, an intermediary: the content it screens reaches wherever the
check runs, on every call, including calls it passes through unchanged. “We added a guardrail” can
mean a second company now sees every prompt and response, not only the ones it flags. Check whose
service runs the check, what it retains, and whether your own hardware could run the same
classifier instead: Llama Guard 4, quoted above, is published as a downloadable model card. See
safety, privacy and governance for the same question asked
about the model maker itself.