Back to blog
#ai-agents#security#business

AI Agent Red Teaming: Test an Agent Before Production

Standard pentests miss agent-specific failures. How to red-team an AI agent before production: the frameworks, the failure modes, and a testing gate that works.

By Rafael Costa7 min readEnglish
Share
AI Agent Red Teaming: Test an Agent Before Production

You can demo an AI agent in an afternoon. Proving it is safe to point at real systems takes a lot longer, and most teams skip the part in between. They test the agent the way they test a feature: does it do the thing when you ask it nicely? That tells you almost nothing about what it does when a customer, a crafted email, or another agent asks it something it was never meant to do.

Red teaming is the discipline of asking the bad questions on purpose, before an attacker does. For ordinary software it is a mature practice. For agents it is different enough that borrowing your existing penetration test will leave you exposed in the places that matter most. This is a practical look at why agent red teaming is its own thing, the frameworks worth anchoring to, the failure modes that actually show up, and how to fold all of it into a pre-production gate instead of a one-off exercise you never repeat.

Why a normal pentest is not enough

A classic penetration test probes a system with fixed behaviour. Send input, check the response, look for the injection or the broken access control. An agent breaks that model in four ways, and each one opens a door a pentest never knocks on.

It makes probabilistic decisions, so the same prompt can take a different path twice, and a safe answer on Monday is not a guarantee for Tuesday. It carries persistent memory, so something planted in a conversation last week can steer a decision today. It holds real credentials and tools, so a successful manipulation does not just leak data, it takes an action: sends the email, issues the refund, changes the record. And it often talks to other agents, so trust between components becomes an attack surface of its own.

Put those together and the headline risk is no longer "can an attacker read the database". It is "can an attacker talk the agent into doing something harmful with the access you already gave it". A test suite built for static software does not even express that question, which is why agents that pass every unit test still fail in the wild. We have written before about why AI agents fail in production, and security failures are a large, under-reported slice of that.

The failure that pentests miss

The signature agent breach is not a clever exploit. It is an agent being politely talked into emailing internal data to an outside address, or approving a request it should have escalated. No buffer overflow, no CVE. Just a model following instructions that arrived through content it was asked to process. If your testing cannot catch that, it is testing the wrong layer.

Anchor to frameworks, not vibes

"We tried to break it for a day" is not a security program. The field has matured enough that you do not have to invent your own taxonomy, and three public frameworks cover most of what you need between them.

  • OWASP's work on LLM and agentic applications gives you the what can go wrong list: prompt injection, insecure output handling, excessive agency, and the agent-specific entries that go beyond the original LLM Top 10. It is the closest thing to a shared vocabulary, and a good spine for a test plan. We covered the business-level version in the OWASP LLM Top 10 for AI agents.
  • MITRE ATLAS maps how attackers actually operate against AI systems, the adversary's playbook rather than a defect list. Use it to think like the other side instead of only auditing your own code.
  • NIST's AI Risk Management Framework, with its generative AI profile, gives you the governance wrapper: how to document what you tested, what you accept, and who signed off. This is the part auditors and enterprise buyers ask for.

The Cloud Security Alliance also published a dedicated agentic red teaming guide that extends these ideas to multi-step, tool-using systems. You do not need all of it on day one. You need enough structure that your testing is repeatable and your coverage is legible to someone who was not in the room.

The failure modes that actually show up

Frameworks are the map. These are the potholes you will actually hit, and a useful red team spends most of its time here.

Goal hijacking. The agent is given a task, then content it processes (a support ticket, a web page, a PDF) contains instructions that redirect it. The agent treats the injected instruction as if it came from you. This is prompt injection with consequences, because the agent can act.

Missing approval gates. The agent can do something it should never do alone. Resetting a password is fine to automate. Wiring a payment, granting access to a finance system, or deleting records is not. The test is simple and brutal: ask the agent to do the dangerous thing, in a dozen phrasings, and see whether anything stops it.

Privilege and tool abuse. The agent has broader access than any single task needs, so one manipulation reaches far. An agent that can read one customer's data to answer a question, but is wired so it could read every customer's data, is one clever prompt away from a breach.

Trust exploitation between components. In a multi-agent setup, one agent trusts another's output without checking it. Compromise the weakest agent and you inherit the trust of the whole chain.

Notice none of these are about the model being "jailbroken" in the abstract. They are about the gap between what the agent can do and what it should do on its own. Red teaming is mostly the work of finding that gap before someone else maps it.

Discovery, pre-deployment testing, runtime defence

A complete program has three stages, and skipping any one leaves a hole.

Discovery is knowing what you have: which agents exist, what tools and data each can reach, and what actions each can take unsupervised. Most organisations cannot fully inventory this, and you cannot defend what you cannot see. Start here.

Pre-deployment testing is the red team proper, run before anything goes live: adversarial prompts across the OWASP categories, every approval gate exercised, every tool connection probed for over-reach. This is where you decide, per action, what the agent does alone, what needs a human to approve, and what it must never touch.

Runtime defence accepts that testing is never complete. A probabilistic system can surprise you in production, so you need monitoring, logging you can actually trace, and the ability to pull an agent's permissions fast when something looks wrong. Red teaming and sandboxed execution are complements, not substitutes: one finds the holes, the other limits the blast radius of the ones you miss.

Layer your guardrails

A single topic classifier as your only guardrail will be bypassed. Published red-team exercises have walked straight through single-layer defences across nearly the whole OWASP taxonomy. Treat guardrails like locks: more than one, of different kinds, so defeating one does not open the door.

Make it a gate, not a one-off

The mistake that undoes good red teaming is treating it as a launch-day ceremony. An agent is not static. You change its prompt, add a tool, swap the underlying model, and every one of those can reintroduce a failure you already fixed. A test you ran once, against a version that no longer exists, protects nothing.

So put red teaming in the release process, the way you already gate on tests passing. It does not need a dedicated AI security team to start. A focused one-week exercise against your highest-risk agent, anchored to the OWASP categories, with every approval gate and tool connection probed, gets you a real baseline. From there the job is to keep a core set of adversarial checks running on every change, and to re-run the deeper exercise whenever the model or the agent's permissions move. Continuous beats thorough-once, because the thing you are testing keeps changing.

The business case writes itself once you have seen a single clean break. An agent with real credentials that can be talked into acting is not a product risk you discover after it ships. It is one you buy down before, for a fraction of what the incident costs.

If you are building or buying an AI agent and want it tested like it will be attacked, we can help: scoping the red team, running it against the frameworks that matter, and designing the approval gates and runtime controls that keep it safe once it is live. Talk to us. We build custom AI agents and the guardrails that let you trust them with real work.

#ai-agents#security#business
Share this article
Rafael Costa

Written by

Rafael Costa

Software Engineer & Technical Writer

Rafael is a software engineer at Lusivision who writes about web development, cloud architecture and applied AI. He has spent over a decade shipping production software for companies across Europe and enjoys turning hard technical topics into clear, practical guides.

View all articles

Related articles

Newsletter

Stay in the loop

Occasional notes on software, design and what we're building. No spam — unsubscribe anytime.