- › Vendor review asks whether the AI provider is secure, which is the wrong question, because the exposure is determined by what your own staff paste into it.
- › An AI assistant is an unlogged export channel that looks like typing, so none of your existing data loss controls were built to see it.
- › The tools got adopted because they solved a real problem faster than anything you offered, which is why bans produce personal accounts instead of compliance.
- › Find the usage through DNS, expense reports, browser extensions, and identity logs, and expect four to ten times what the register lists.
- › Give people a sanctioned option with a written data boundary, because the only durable control is a good tool people are allowed to use.
Somebody on your finance team has a contract they need summarized by lunch. It is forty pages of amendments to a master services agreement, the counterparty wants an answer today, and reading the whole thing carefully takes most of an afternoon they do not have.
They paste it into a chat window. They get a good summary in eleven seconds, they act on it, and the deal closes on time.
Nothing in that sequence triggered an alert. No file was attached to an email, no document left through a sanctioned sharing link, and no data loss prevention rule watched a person type into a text box in a browser. The contract is now somewhere else, and there is no record anywhere in your environment that it went.
The Question the Vendor Review Is Not Asking
We have written a good deal about AI as a threat, including how attackers weaponize psychology with it and what an agent with full access to your systems can actually do. Those pieces are about AI pointed at you. This one is about AI pointed at your own data by people trying to do their jobs.
When an AI tool does go through procurement, the questionnaire asks about the provider. Where is data stored, is it used for training, what certifications exist, what happens on termination. Every one of those questions is worth asking and none of them describes your actual exposure, because the answers govern what the vendor does with what they receive and say nothing about what you send.
The volume and sensitivity of what gets sent is decided by a few hundred individual judgment calls a week, made by people under deadline pressure, none of whom were ever told where the line is. Your risk is the sum of those calls.
That makes this a measurement problem before it is a policy problem, and most organizations have skipped straight to writing the policy. A rule written without knowing what people currently do is a guess about where the line should sit, aimed at a behavior nobody has counted.
Why This Slipped Past Controls That Work Elsewhere
Data loss prevention was built for a world where data left in recognizable containers. A file attached to a message, a document moved to personal storage, a print job, a USB device. Those have signatures, and the tooling is mature.
An AI assistant breaks that model in three ways at once, and each one defeats a different layer of what you already have running. None of the three is exotic, which is part of why they went unnoticed for as long as they did.
The channel looks like typing. Text entered into a browser field on an allowed domain is indistinguishable from ordinary work at the network layer, and it is encrypted in transit to a reputable provider with a valid certificate. There is nothing anomalous to detect.
The destination is legitimate. Blocking the domain is the obvious response, and it collides with the fact that half your engineering organization uses the same provider for sanctioned work. A block list with a business-critical exception is a block list that does nothing.
The volume is invisible. Forty pages pasted into a text box is the same network event shape as a long question. You have no field-level visibility into a request body you cannot decrypt, and the interesting part is inside it.
Every control that would catch this was designed against an assumption that data exfiltration is either malicious or careless. This is neither. It is a person using a good tool to finish work on time, and the permissive egress problem we wrote about is the same shape with a human in the middle instead of a service account.
Finding Out What Is Actually Happening
You cannot govern a usage pattern you have not measured, and the measurement is more available than most people expect. Four sources, and the useful part is where they disagree with each other.
DNS and egress logs
Query your resolver logs for the domains of every major model provider and every AI-wrapper product you can name, then do it again for the ones you cannot name by pulling the top unrecognized domains by request count over thirty days. The second pass is the one that produces surprises, because the market has more products in it than any one person is tracking.
Count distinct internal sources rather than total requests. A hundred thousand requests from three machines is an engineering integration. Four hundred distinct devices is an adoption pattern.
Expense reports and corporate cards
Individual subscriptions show up as recurring charges under twenty-five dollars, which is below the threshold where anybody reviews them. Search the card data for the provider names and for the generic descriptors these charges use. Every hit is a person who wanted the tool badly enough to expense it, which is useful information about demand regardless of what you decide to do about it.
Browser extension inventories
If you manage browsers at all, pull the installed extension list across the fleet and look for AI assistants, summarizers, and writing tools. Extensions are the least examined software in most environments and they frequently have permission to read page content on every site the user visits, which includes your internal applications.
Identity provider logs
Search for OAuth grants and sign-ins where somebody used a corporate account to register with an AI service. This one is valuable beyond the count, because a corporate identity attached to a personal-tier account creates an entitlement nobody in your organization can revoke, and it will outlive the employee.
Run all four and compare the answers to whatever your software register says. The gap is usually between four and ten times, and the size of it is the most persuasive artifact the whole exercise produces. We found the same dynamic when we looked at API endpoints nobody had inventoried, and the reason is identical, since the register records what was requested and the environment records what happened.
The Response That Does Not Work
The instinct is to ban it, and I want to be specific about what banning produces, because the outcome is never zero usage. It is a change in where the usage happens and in whether you can see it.
It produces the same usage on personal accounts, on personal devices, over cellular, with no corporate agreement governing the data, no retention terms you negotiated, and no ability to discover any of it later. The work still gets done, because the deadline did not move. You have converted a visible problem into an invisible one and given up the last of your influence over the terms.
This is the same lesson as every other control that lost to convenience. People route around a rule when the rule costs them more than following it appears to save, and the software your organization never approved has always arrived this way.
What Actually Reduces the Exposure
Give people a sanctioned tool, and make it good
The single most effective control is an enterprise-tier account with a real data processing agreement, no training on your inputs, and enough capability that nobody is tempted by the consumer version. This costs money and it is cheaper than the alternative by a wide margin.
Make it easy to reach. A sanctioned tool behind three approval steps loses to a personal account every time, and the loss is silent.
Write a data boundary somebody can actually apply
Most AI policies say to use good judgment with sensitive information, which is not a rule. A usable boundary names categories and gives a default.
Name what never goes in: customer records with identifiers, credentials and keys, unredacted contracts, source code from your revenue-generating systems, anything under a confidentiality obligation to a third party. Name what is fine: public material, your own drafts, general questions, sanitized examples. And name the gray zone with a person attached to it, because the value of a policy is mostly in who a person asks when they are unsure.
Instrument the sanctioned path
Enterprise tiers provide administrative logging that consumer accounts do not. Turn it on, review it monthly, and use it to find out which teams are doing what. Nine times out of ten you learn that a department has quietly built a workflow depending on the tool, which is a capacity question and a continuity question rather than a security finding.
Handle the identity entanglement
Every corporate email address registered against a personal-tier AI account is a loose end. Inventory them, migrate the ones that should exist onto the enterprise tenant, and close the rest. Do this before somebody leaves rather than after, because the offboarding process you have does not know these accounts exist.
The Argument for Doing This Now
The exposure compounds quietly. Every week this goes unmeasured, more material moves through a channel with no record of what went, and the portion of it you would have to disclose in an incident grows without anybody choosing to grow it.
There is also a narrow window on the cultural side. Right now, employees using these tools believe they are being resourceful, which is a reasonable belief and largely correct. That makes them willing to tell you what they are doing if you ask in a way that does not sound like an investigation. Once the first person gets disciplined for it, the honest answers stop and you are left with the logs alone.
If you want help measuring what is actually moving through AI tools in your environment and drawing a boundary your people will follow, contact Grab The Axe. You can also take our free Human Attack Surface Score to see where the people-shaped gaps sit.
Chris Armour is Director of Information Security at Grab The Axe.
Operating on the philosophy that 'you can't build a secure system if you don't know how to break it,' Chris leads our engineering division. A top 1% National Cyber League competitor, he hardens our digital infrastructure against the very exploits he has mastered.
View Author Page →