AI Security
Two products want to decrypt the same traffic. Only one can be first.
When an AI DLP tool lands next to an existing secure web gateway, both need cleartext on the same flow. Interception does not fan out. Here is how to decide the order, and what breaks when you get it wrong.
September 19, 2026 · 6 min read

A question keeps coming up in the same shape. An organization already runs a secure web gateway or SSE platform that decrypts outbound TLS. Now they are buying something to control AI use specifically, because the incumbent cannot tell a prompt from a form post. The new product also needs to see cleartext. Both want to be the thing that breaks the TLS session.
The question people ask is how to give both products visibility. The answer is that you usually cannot, not in the way the question implies, and the sooner that is understood the better the architecture gets.
Interception does not fan out
TLS inspection works by terminating the client's session, presenting a certificate the endpoint trusts because your CA is in its trust store, then opening a second session to the real destination. It is a man in the middle you consented to.
That means only one thing on a given path holds the original session. Putting two inspection products in line does not give you two copies of the plaintext. It gives you a chain: the client talks to product A, product A re-encrypts and talks to product B, product B talks to the internet. There is exactly one first decryptor, and everything downstream sees traffic that the first one already shaped.
Three consequences fall directly out of that, and they are the ones that bite in production.
The second product loses the client. From B's point of view, the session originates from A. Unless A is passing identity forward in a way B can consume, your AI policy tool sees traffic from a proxy, not from a named user or device. Per-user AI policy silently becomes per-egress-point policy.
Whatever A blocks, B never sees. That sounds obvious and it is still where most surprises come from. Your incumbent gateway blocking a category means the AI tool cannot log or analyze what happened there. Your AI usage picture has a hole in it shaped exactly like your gateway's block list.
Latency and failure modes stack. Two termination points on the same flow means two places to add handshake overhead and two independent things whose outage takes the path down. A chain is only as available as its least reliable link, and you now have one more link.
Certificate pinning is where the chain actually breaks
Pinning is the constraint that decides most of these designs, and it is worth being precise about why.
A pinned client does not accept any certificate that chains to a trusted root. It accepts a specific certificate or public key it was built with. Your enterprise CA being trusted by the operating system is irrelevant; the application is not asking the OS. OWASP is direct about this: interception proxies and pinned clients are mutually exclusive by design, and the only sanctioned resolution is adding the proxy's key to the pinset deliberately, which is a decision for whoever owns risk acceptance, not something you can configure your way around.
This matters more for AI traffic than for general web traffic, because a meaningful share of AI consumption is not a browser. Desktop assistant applications, IDE extensions, mobile apps, and SDK calls from your own services frequently pin or use mutual TLS. Those flows either bypass inspection or fail. Inspecting them harder does not work; it just produces tickets.
Which is why the honest version of this architecture always contains a bypass list, and why that list should be a deliberate document rather than an accumulation of exceptions someone added at 2am to stop an outage.
How to decide who goes first
Once you accept that the order is a decision, the decision is not that hard. Put the product with the narrower, more specific job closest to the endpoint, and let the broader one run outside it.
In the AI case that usually means the AI-aware control decrypts first, for a simple reason: it is the one that needs to read the payload and attribute it to a person. Prompt content, model destination, data class in the request, and identity are all things you lose fidelity on when you sit behind another proxy. The general web gateway, doing category filtering, malware scanning, and broad DLP, degrades much more gracefully when it is second, because those decisions rarely depend on knowing which human is on the other end.
The other workable pattern is to stop chaining and split by destination. Model endpoints go out one path with AI inspection; everything else goes the existing path. This is cleaner, it removes the double-termination problem entirely, and its weakness is that the split is only as good as your list of model destinations, which changes constantly. New providers appear, and vendors move inference to new hostnames without telling you. A destination split needs an owner and a monthly review or it degrades into a false sense of coverage.
There is a third option people underuse. Move the control off the network and onto the endpoint or the identity layer. Browser-based enforcement, endpoint agents, and OAuth grant governance all see the thing the network cannot, which is what happened inside a pinned app or on an unmanaged device. Network interception is one seat at the table, not the only one. If the reason you are buying an AI DLP product is to catch someone pasting PHI into a chatbot, an endpoint or browser control gets you that without touching the TLS chain at all.
Write the intercept map down
The artifact that resolves this argument is one page, and it takes an afternoon.
List every path traffic can take out of your environment: managed device on the corporate network, managed device remote, unmanaged device, and server or service egress. For each path, record which product terminates TLS first, which one is second if any, what identity information reaches the second one, what is deliberately bypassed and why, and what the fail-open or fail-closed behavior is when an inspection point goes down.
Then add the bypass list as a table with a reason and an owner per line. Pinned banking app is a reason. Nobody remembers is not.
The value of this page is not documentation. It is that the questions it forces, who is first, what does second see, what is exempt, are exactly the questions that decide whether your AI usage data is trustworthy. A dashboard built on traffic that a gateway already filtered, from users the tool cannot name, on a path that excludes every pinned desktop client, will still render beautifully. It will just be wrong in ways nobody can see from the dashboard.
The short version
Two decryptors on one flow is a chain, not a fan-out. Decide the order deliberately instead of discovering it during an incident. Put the tool that needs payload and identity closest to the endpoint. Expect pinned applications to be a bypass list rather than a problem you configure away. And know which of your traffic never touches the chain at all, because that is where the exposure you are buying the product to find is most likely sitting.