Insights

AI Security

The Pen Test Your AI Feature Never Got

Every year you pen test your network. The AI feature you shipped last quarter never got the same treatment. Here are the six ways it fails, and how to test for them.

July 10, 2026 · 5 min read

The Pen Test Your AI Feature Never Got

Every year, most companies run a penetration test on their network. Almost none run one on the AI feature they shipped last quarter.

That gap is where the next security review stalls. An LLM feature fails differently than normal software. It does not crash. It complies. Ask it the wrong way and it leaks its system prompt, ignores its guardrails, or runs a tool it should never touch. A code review will not catch that. Neither will a standard pen test. You have to test the model the way an adversary would talk to it.

I spent 13 years in law enforcement before twenty in cybersecurity, and the instinct is the same one that kept me out of trouble on the street: assume the thing in front of you will behave differently under pressure than it does in the demo. Here are the six ways an AI feature fails, and how to test for each one before a buyer, or an attacker, finds them for you.

1. Prompt injection

The highest-impact class, and the one teams underestimate most. Injection is not only what a user types into the chat box. It is any instruction the model reads and follows.

  • Direct injection is a user overriding your system prompt through their input.
  • Indirect injection is an instruction hidden inside content the model ingests: a document it summarizes, a web page it fetches, a support ticket, a record pulled from your database through retrieval. The user never types the attack. The data carries it.

Test every path the model reads from, not just the one you designed for. If your feature summarizes uploaded PDFs, a PDF is an input path.

2. Jailbreak and guardrail bypass

Getting the model to do what your policy forbids. Role-play framing, hypothetical scenarios, encoded requests, and known adversarial techniques all get a model to step around the rules it was told to follow. If you have an input or output classifier acting as a guardrail, test whether it can be walked past. A guardrail you never tried to break is a guardrail you are trusting on faith.

3. Sensitive information disclosure

What the model says that it should never say. System-prompt extraction, where an attacker recovers your hidden instructions. Training-data leakage. PII exposure. And in a multi-tenant application, cross-tenant leakage, where one customer's data surfaces in another customer's session. That last one ends deals and starts breach notifications.

4. Tool and agent abuse

This is where an AI feature stops being a chatbot and becomes a path into your systems. The moment you give a model tools, function calls, or connected MCP servers, you have handed it authority. The tests that matter:

  • Excessive agency. Does the model take actions beyond what the task required?
  • Unsafe function calling. Can an attacker inject parameters or chain tools in ways you did not intend?
  • Confused deputy. Can a user use the model's authority to reach something they could not reach themselves?

If your AI agent can send email, query a database, or hit an internal API, this is the class that turns a clever prompt into an incident.

5. Model fingerprinting and recon

Before an attacker attacks, they identify what they are attacking. With a handful of queries they can often infer the underlying model, its capabilities, and the shape of its system prompt. That recon shapes every attack that follows. Knowing what you look like from the outside tells you which of the classes above you are most exposed to.

6. AI supply chain

The model file and its dependencies are an attack surface of their own. Model files that use unsafe serialization can execute code the moment they load. Provenance matters too: is the model you deployed actually the model you think it is, from the source you think it came from? Scan model artifacts the same way you scan software dependencies, because that is what they are.

How to actually run this

You do not need a research team. The open-source tooling is good and free: garak for scanning, PyRIT for multi-turn attack orchestration, promptfoo for injection evals, all under permissive licenses. The harder part is running them with intent and reading the output with judgment.

I extended my own open-source pen testing tool, Talon, to do exactly this. Talon already pointed an AI at a Kali box to test infrastructure. The new AI Red-Team pack points the same AI-directed workflow at an LLM instead. Same idea, new attack surface. It is on GitHub, MIT licensed, and the methodology is there whether or not you ever talk to me.

Whatever you run it with, map the findings to the frameworks your auditor already recognizes: the OWASP Top 10 for LLM Applications, MITRE ATLAS, and NIST AI RMF. A finding tagged to a standard is a finding a security team can act on. A raw tool dump is noise.

The honest part

A scan does not secure anything. A person reading the output who knows which finding actually matters does. The tool gets you to the evidence faster. It does not tell you that cross-tenant leakage is a business-ending event and a verbose error message is not. That judgment is the work.

Run the test before your buyer's security team does, and before someone less friendly does. The AI feature you shipped last quarter has been in production, untested against an adversary, the whole time. That is the pen test it never got.

Fractional CTO and CISO leadership for companies putting AI to work: strategy, governance, cost control, and risk in business terms. Our team has led cyber defense, compliance, and risk programs for 20+ years across 6 countries and multiple industries, including healthcare, fintech, retail, manufacturing, telecom and consulting, and delivered large-scale security and compliance programs at Accenture, Dell, EY, Booz Allen Hamilton and AT&T. Technology and security leadership in one seat, reported in business terms. Talk to us.