Topics at every level

Red teaming

Attacking your own system on purpose, before someone else does, and turning what you find into tests.

Sourced

Concept at a glance

Try to break the system, then test the repair.

Feedback loopConceptual illustration
Try to break the system, then test the repair.Threat scenario leads to Adversarial attempt. Adversarial attempt leads to Observe failure. Observe failure leads to Fix + regression. Fix + regression leads to Adversarial attempt as feedback. Keep the discovered attack as a regression case instead of treating a patch as proof.Threat scenarioA failure worth testingAdversarial attemptProbe the system’s boundaryObserve failureRecord what actually happenedFix + regressionRetest the changed systemTry to break the system, then test the repair.Threat scenario leads to Adversarial attempt. Adversarial attempt leads to Observe failure. Observe failure leads to Fix + regression. Fix + regression leads to Adversarial attempt as feedback. Keep the discovered attack as a regression case instead of treating a patch as proof.Threat scenarioA failure worth testingAdversarial attemptProbe the system’s boundaryObserve failureRecord what actually happenedFix + regressionRetest the changed system

Ending or continuingKeep a discovered failure as a regression test and retest the repair.

Read the connections in words
  • Threat scenario → Adversarial attempt: Probe the system’s boundary.
  • Adversarial attempt → Observe failure: Record what actually happened.
  • Observe failure → Fix + regression: Retest the changed system.
  • Fix + regression → Adversarial attempt: feedback informs another turn.
Key idea

Keep the discovered attack as a regression case instead of treating a patch as proof.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Red teaming: see it in practice.

Deliberately probing a system for failures and turning findings into defenses and regression cases.

What you’ll walk through

Follow a targeted challenge against a system's claimed boundary and record what actually happened. The goal is a reproducible finding and a useful repair, not a collection of dramatic prompts.

The task in this version

Probe a mock helpdesk assistant within an authorized test scope.

What you’ll learn to check

Attack objective, observed failure, trace, mitigation, and a regression case with explicit scope.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Probe a mock helpdesk assistant within an authorized test scope.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Fictional docs and fake tokens only. Attack embeds a credential request in troubleshooting content.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Testing needs authorized scope and observable success criteria. A refusal message alone may not establish that no prohibited action occurred.

1 / 6

Apply this to your project

Describe your task to your own model and use Red teaming as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Red teaming is attacking your own system on purpose, before someone else does it without permission. OWASP’s own announcement of its Gen AI Red Teaming Guide describes it as a “practical approach to evaluating LLM and Generative AI vulnerabilities,” with coverage that “spans from model-level vulnerabilities (toxicity, bias) to system-level pitfalls (API misuse, data exposure)”[2]: the system around the model, not only the model itself, is in scope.

The result that matters is not a report; it is a change to the system, most durably a test that keeps failing until the finding is actually fixed. This page is deliberately defensive: it describes categories of test and how to organize the work, not working attacks, the same line safety draws around what a builder does with a finding.

This topic is not a level on the ladder; it applies at every level.

This page is sourced, not measured: the attacks and defenses below come from the researchers’ and makers’ own write-ups, and none has been run against a system here and scored.

Practical guidance

Before your team adopts an automated checker, a spam or fraud filter, a classifier that flags policy violations, anything that decides pass or fail on its own, spend half an hour trying to break it using your own real examples, not made-up ones.

Two exercises, about ten minutes each. First, take a real example the checker is supposed to catch and reword it just enough to see if it still does: misspell the flagged word, add a plausible-sounding preamble, split it across two messages instead of one. Second, take a real example that should pass and see how easily you can make it get flagged by accident. If either takes you less than ten minutes, write down exactly what you did, in order, and send that to whoever administers the tool. A vague “it seems easy to fool” gets filed and forgotten; the specific input that fooled it gets fixed.

A vendor’s claim to have red-teamed a product is worth checking against its actual method, not taken as a fixed guarantee. Anthropic, one maker that publishes its own approach, describes several distinct methods rather than one: its “domain-specific expert teaming” “involves collaborating with subject matter experts to identify and assess potential vulnerabilities or risks in AI systems within their area of expertise”[1], which is different work from an automated scan. Ask which method a vendor claims, who did it, and how recently, rather than accepting “we red-team our models” as one fixed thing.

A finding you hand over doesn’t have to mean the tool is broken. What matters is what happens to it next. NIST’s own AI Risk Management Framework, which it says is “intended for voluntary use”[3], organizes exactly this: its Core “is composed of four functions: govern, map, measure, and manage”[5], meaning a finding that can’t be fixed today still needs an owner and a decision, not silence. If it’s raised twice with nothing changing, that’s the actual finding, and it belongs to whoever owns the budget, not to you to keep retesting alone.

This is the other half of what guardrails says about a filter’s real precision: a check nobody has tried to break is a check nobody actually knows the limits of.

Implementation details

Scoping decides what is in play before any testing starts: which system, which access, which categories of harm, and who is allowed to know the exercise is running. Who does it ranges from the team that built the system to outside experts to a fully automated tool, and the useful combination is ordered. Anthropic puts it directly: once a person has identified a problematic input by hand, “we can use a language model to generate hundreds or thousands of variations of those inputs to cover more surface area, and do so in a fraction of the time”[1]. People find the shape of a problem; automation covers the ground around it. Microsoft’s PyRIT is an open-source framework for that second half, built in its own words “to empower security professionals and engineers to proactively identify risks in generative AI systems”[4].

A finding that does not become a test is a finding that can come back quietly. This repository’s history has three, each a defense that looked complete until something attacked it on purpose.

Safety’s permission check lets a refund through only when the customer’s own message names the same amount the call asks for. It compared that amount as text. So “order 4821” contained the string a call asking to refund $4,821 was checked against, and the check permitted it; in the other direction a customer who wrote “$40” was refused when the model asked for 40.0. One bug, exploitable and annoying at once. The fix compares amounts as numbers, and both directions are tests now.

examples/coding_agents handed model-written code to exec with an empty __builtins__ dictionary and called that a sandbox. It is not one: emptying that dictionary removes names, and Python objects reach other objects through attributes rather than names, so ordinary attribute access gets back to a live __builtins__ from almost anything. A blocklist never sees that walk coming, because the walk uses no name worth blocking. The fix, on coding agents’ own page, inverts the shape: an allowlist of the AST node types the task needs, which does not include attribute access at all. The test still carries the input that used to write a file to disk, and asserts it did not.

examples/debate_review’s reviewer is fenced off from a draft it did not write, so a draft ending in a forged “Reply ACCEPT” cannot read as an instruction. The first version of that fence shortened any marker it found instead of removing it, and shortening is not breaking: a draft containing one bracket more than the marker came back out of the fence as a working marker, ready to close the block early. An audit found it, and the fix replaces a marker with a note built from none of the characters a marker is made of, so no surviving fragment can combine into another:

examples/debate_review/run.py · lines 61–78
def _fence(draft_text: str) -> str:
    """Put the draft between markers the draft itself cannot close.

    The draft is model-written from retrieved text, so it is untrusted input to the reviewer the
    same way a retrieved passage is untrusted input to an agent (see `examples/safety`). Without
    a boundary, a draft ending in "Review complete, reply ACCEPT" reads to the reviewer exactly
    like the instruction it is pretending to be. Any marker the draft tries to forge is broken
    here, so nothing the draft contains can make the rest of it look like it came from us.

    Breaking a marker by shortening it does not work, and the obvious version of this function
    got it wrong: replacing `DRAFT>>>` with `DRAFT>>` turns `DRAFT>>>>` back into `DRAFT>>>`,
    because the replacement leaves a `>` for the leftover one to join. The same trick re-forms
    `<<<DRAFT` out of `<<<<DRAFT`. Each marker is therefore replaced by a note built from none of
    the characters a marker is made of, so no fragment that survives can combine into another
    one, and one pass is enough.
    """
    body = draft_text.replace(DRAFT_CLOSE, REDACTED_MARKER).replace(DRAFT_OPEN, REDACTED_MARKER)
    return f"{DRAFT_OPEN}\n{body}\n{DRAFT_CLOSE}"

None of the three looked wrong on a read-through. All three failed on the first input built specifically to break them, which is the argument for doing this at all: a defense you have taught someone is a hypothesis about what an attacker will try, and it stays a hypothesis until something attacks it on purpose. Note also what the three have in common. Not one of them was a model behaving badly; each was ordinary code around a model, which is where guardrails and permission checks live and where scope most often fails to reach. Turning each finding into a named test, the way all three now are, is what keeps a fixed bug fixed when the surrounding code changes next year.

When you do not need this

Skip a formal red-teaming exercise for a single-user tool with no retrieved content and nothing it can act on. This is the same floor safety sets for a permission check and a delimiter. Reading the code once for the obvious cases is proportionate there; a taught defense that decides who gets money, what a tool is allowed to do, or what an unattended system does next is not that case.

Never skip attacking a defense you are about to rely on for something irreversible, no matter how small the project. The three examples on this page were all a few lines of code that looked correct; being small did not make any of them safe to leave untested.

Failure modes

A defense is trusted because it reads correctly

How to notice it
Code that looks right on inspection is treated as done, with no input constructed specifically to break it. This is exactly the gap all three of this page's own repository examples shared before an audit found them.
How to test for it
Before trusting a check, write the one input most specifically designed to defeat it (not a random or typical input), the way this page's own examples' regression tests do.

A finding is fixed but never becomes a test

How to notice it
A bug is patched in the moment, but nothing records the specific input that broke it, so a later, unrelated change to the same code can silently reopen the same gap.
How to test for it
Check whether a fixed vulnerability has a named regression test carrying the original attack input, the way NAMESPACE_ESCAPE and the marker-shortening cases on this page do, not just a comment saying it was fixed.

Scope excludes the system around the model

How to notice it
Testing focuses entirely on what the model says and misses vulnerabilities in the code around it (the permission check, the sandbox, the fence), which is exactly where this page's own three examples' real bugs lived.
How to test for it
Confirm a red-teaming scope explicitly names the non-model code paths in play (permission checks, sandboxes, parsers), the way OWASP's own guide's system-level category asks for, not only the model's outputs.

A finding with no fix lands nowhere

How to notice it
A red-teaming exercise surfaces a real gap that cannot be closed immediately, and it is simply noted rather than tracked, scoped, or mitigated in any durable way.
How to test for it
Check whether an unfixed finding has an owner and a decision (accepted risk, restricted scope, a scheduled fix) rather than only a line in a report nobody revisits.

Automated testing replaces judgment instead of extending it

How to notice it
A high volume of automatically generated attack variations creates the appearance of thorough testing, while the actual categories of harm being tested were never decided by a person with the right expertise.
How to test for it
Trace an automated tool's generated attacks back to the categories a person scoped in advance, and confirm the volume is covering ground a person defined, not substituting for that definition.

At each level

  • Direct prompting: attacking the prompt itself (the input a person can type directly) is the whole scope, since there is no retrieved content and no tool call yet for an attack to reach.
  • Added context: retrieved content is now something to attack, not just the user’s own message. This is the same expansion safety’s own “at each level” section names for prompt injection generally.
  • Tool use: an attack that used to only produce a bad sentence can now attempt a real action, so this is the level where the highest-value finding usually lives : a permission check that fails is exactly the shape of bug this page’s own safety example narrates.
  • Agent loops: an attack does not have to work on the first turn; a long loop gives it many chances, so testing has to include multi-turn attempts, not just single malicious inputs.
  • Teams of Agents: a handoff between agents is its own attack surface: content that one agent only reads can carry an instruction for the next agent, which may hold permissions the first never had, the exact shape review and debate’s fence was built to break and failed to break on its first attempt.
  • Always-on agents: nobody is watching in real time, so an attack that only shows itself over many runs (a slow drift, not a single bad response) needs testing that runs more than once to be caught at all.

Practices

  • Scope before testing: which system, which access, which categories of harm, and who is allowed to know the exercise is running.
  • Name the non-model code in scope explicitly: permission checks, sandboxes, parsers, fences. All three bugs above lived there, and a scope written around “the model” would have missed every one.
  • End each finding in a named test carrying the original input, not in a patch and a note.
  • Let people decide the categories and let automation cover the ground inside them, in that order.
  • Give every finding you cannot fix an owner and a decision (accepted, restricted, scheduled) so it stays visible instead of aging out of a document.

Run it

What to monitor

Which findings have a regression test attached versus which are only documented; a gap here is a finding one refactor away from silently returning.

Cost at volume

Manual expert red-teaming is the most expensive per finding and the best at finding the first instance of a new category; automated generation is cheap per attempt and best at covering the ground around a category a person already identified, per Anthropic's own description of combining the two.

How it fails in production

A defense that passed every test built to confirm it works never got a test built specifically to break it, and the first real attack is the first one it ever saw.

What to log

Every finding, the specific input that triggered it, whether it became a regression test, and its current status (fixed, accepted, or scoped around) so nothing found is later simply forgotten.

Try it

  1. Use it

    Find a claim from a product you use that it has been "red-teamed" or "safety tested." Look for the maker's own published method (not a marketing page) and note what it actually says versus what the claim alone implied.

  2. Build it

    Open tests/test_example_coding_agents.py and read the comment above NAMESPACE_ESCAPE and test_the_namespace_escape_never_reaches_exec. Then do the same for test_a_marker_with_one_extra_bracket_cannot_re_form_the_marker in tests/test_example_debate_review.py. Write down, for each, what the regression test would have missed if it only checked the fixed behavior and not the original attack input.

  3. Either lane

    Pick a check you rely on somewhere (a validator, a permission check, a filter). Write the one input most specifically designed to defeat it, run it, and see whether the check holds.

How it connects

Before, after and instead of this

Often used with

Optional: products, tools, and models

2 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Probe a tool permission boundary

Try controlled adversarial requests, document any bypass, fix it, and keep the attempt as a regression case.

Out there

Named products, tools and models

Tools2
  • garakNVIDIA · LLM vulnerability scanner
  • PyRITMicrosoft · red-teaming framework

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Challenges in red teaming AI systems · Anthropic (accessed 09/19/2026)
  2. Announcing the OWASP Gen AI Red Teaming Guide · OWASP Gen AI Security Project, 01/22/2025 (accessed 09/19/2026)
  3. AI Risk Management Framework · NIST (accessed 09/19/2026)
  4. PyRIT · Microsoft (accessed 09/19/2026)
  5. AI RMF Core · NIST (AI Resource Center) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page