Your employee just sent company data to an AI. Now what?
A finance analyst has a board meeting in an hour and a spreadsheet that needs explaining. They paste the figures into an AI assistant, ask for a summary, and get a useful answer in seconds. Only afterwards do they wonder whether the unreleased numbers should have gone into that chat.
It is an easy mistake to make when a tool feels like a private workspace. The same question comes up when someone uploads a customer conversation, asks for help with internal code or turns meeting notes into a presentation. The task may be routine. The information involved may not be.
For the security team, the first job is to establish what was shared and where it went. An upload does not automatically mean the information is public or has been used to train a model. It does mean the organization needs to understand how that service handles the data and whether the sharing was authorized.
Why a useful shortcut becomes a security issue
AI tools save time on work that employees already find repetitive. If the approved process is unclear, a personal account may look like the quickest way to get something done. A policy that says “use AI responsibly” offers little help when someone is deciding whether a particular document is safe to upload.
Cyberhaven’s 2026 AI Adoption & Risk Report, based on billions of data movements across 222 companies, found that 39.7 percent of AI interactions involved sensitive data. It also found that 32.3 percent of ChatGPT usage happens through personal accounts, not corporate ones. These numbers highlight why organizations need to understand both the information being shared and the accounts handling it.
That makes it worth asking teams how they use AI before an incident forces the conversation. Which tasks do they need help with? Which accounts are they using? Where are they unsure about the rules? Those answers help IT identify where clearer guidance, approved tools, or stronger controls are needed.
The exposure is rarely what it looks like
AI tools do not receive data the way an email attachment does. They receive context, and context is expansive.
A developer using a coding assistant is not submitting one file. Depending on the tool’s configuration, the assistant may be indexing the entire open workspace; reading across directories; and pulling in configuration files, environment variables, and infrastructure definitions the developer never consciously shared. Cursor, for example, operates on a project-level context unless a .cursorignore file explicitly excludes directories. Most developers do not configure this. The assistant sees everything the integrated development environment (IDE) can see, which in a monorepo can mean credentials, API keys, proprietary business logic, and the architecture of systems the developer is not even actively working on.
Meeting assistants have the same expansive reach, but through a different mechanism. When an AI transcription tool joins a call, it does not transcribe only the agenda items. It captures the entire session: the sidebar conversation about a personnel decision, the off-the-cuff comment about an unannounced acquisition, the screen share that contained a financial model. That transcript is transmitted to the vendor’s servers under the vendor’s retention policy, which the employee almost certainly has not read, and retained indefinitely unless explicitly deleted. The employee who approved the bot joining the call may not even be the one who authorized the tool.
Browser extensions compound the problem further. An AI writing assistant or productivity extension installed with broad host permissions can read page content across every site the employee visits, including internal admin panels, authenticated dashboards, customer records systems, and HR platforms. The extension is not malware. It is a productivity tool that the employee installed in good faith, and it has been quietly reading sensitive content across every tab ever since.
The implication is that when you discover a data exposure incident involving an AI tool, the scope of what was shared is almost always larger than the employee can accurately report. They submitted what they remember submitting. The tool received considerably more.
What actually happens to the data
Not all AI data exposure carries the same risk, and the determining factor is rarely which tool was used; instead, it is often the account, terms, and settings active at the time that determine the level of risk.
Some conversations on consumer AI services may be used to improve the model, unless the user has explicitly disabled training in their settings. Critically, disabling training does not eliminate all data storage: OpenAI retains content for up to 30 days for abuse monitoring even when the training toggle is off.
Enterprise agreements are materially different. ChatGPT Enterprise, Microsoft Copilot for Enterprise, and equivalent tiers include data processing agreements that contractually exclude customer data from training and specify retention and deletion windows. The same file submitted through a personal account and a company-managed enterprise account presents an entirely different risk profile. The account type is the single most important variable to establish first.
Training and retention are also separate questions that often get conflated. A vendor promising not to use prompts for model training does not mean the data is not stored. Storage does not mean it has entered a model or become accessible to other users. But storage does mean it is subject to the vendor’s security posture, their breach history, and whatever access controls they apply to retained conversation logs. For a coding assistant that has ingested your entire repository context, that distinction matters considerably.
There is also a timing dimension that most incident responses ignore. Browser extensions and connected drive integrations do not capture data at a single point in time. They have been capturing it continuously, potentially for months, since installation. When you discover the exposure, you are not looking at one event. You are looking at everything that tool has processed since it was installed, across every session, on every device where it was active.
The regulatory and liability reality
The data classification of what was shared determines which obligations are triggered, and those obligations do not care about the employee’s intentions.
Personal data submitted without a lawful basis and a data processing agreement is a reportable incident under the GDPR, with a 72-hour notification window that starts when your security team becomes aware, not when the investigation concludes. In healthcare, ePHI submitted to a system without a business associate agreement is a potential HIPAA violation regardless of whether anyone accessed it. In financial services, material non-public information passed through a consumer AI tool creates regulatory exposure that needs legal assessment before you decide on a remediation path.
Client contracts add a further layer: Most data handling obligations predate the AI era and do not account for intent, meaning a consultant who pasted client strategy documents into ChatGPT may have breached a confidentiality agreement regardless of outcome. And if credentials were present in anything the tool received, including configuration files a coding assistant indexed without the developer realizing, that is a separate, parallel incident. Revoke and rotate those credentials before anything else.
What an actual response looks like
The immediate priority is containment: stop the ongoing capture before assessing the past exposure. Ask the employee to stop submitting related material, revoke any active OAuth connections the tool has to internal systems, and remove the extension or tool from managed devices. If the tool is a meeting assistant, check whether it has been added as a recurring participant to any calendar events.
The assessment then needs to answer three questions in sequence. First, what account and tier was in use, because this determines the retention and training exposure. Second, what was the actual scope of context the tool had access to, not just what the employee submitted, but what the tool could read. Third, how long has this been active, because for extensions and connected tools, the exposure window is the entire period of installation, not the moment of discovery.
Involve your privacy team early. The instinct is to treat privacy review as a downstream step once the technical facts are established. In practice, the 72-hour GDPR notification window runs concurrently with your investigation, and a delayed escalation to the privacy team can turn a manageable disclosure into a compliance failure. Similarly, your legal team needs to review key client contracts against what was exposed before you decide whether and how to communicate externally.
Document everything before you delete anything. Provider deletion options exist, and you should use them, but deletion should follow documentation, not precede it. Preserve the incident record, the account settings at the time of exposure, the deletion request reference, and the assessment outcome before any content is removed from the provider’s systems.
Making the next incident less likely
The structural fix is not a policy document. Employees who encountered the original incident almost certainly had a policy in place that they either did not know applied or could not interpret in the moment. “Use AI responsibly” resolves no edge cases. What resolves edge cases is a specific, short matrix that maps data types to approved tools and account tiers, with concrete examples from the actual work those teams do.
The controls need to match the exposure surface. Browser extension audits on managed devices, meeting assistant approval requirements, secrets scanning in CI pipelines, and DLP rules that flag AI endpoints are not separate initiatives. They are the operational version of the same policy, applied at the points where the exposure actually happens.
The employee who let the meeting assistant join the call or who configured their coding assistant without excluding the secrets directory is not the problem. The environment that gave them no reason to pause and no easy way to check is. That is the thing worth fixing.
Sources
[3] Google Workspace Generative AI Privacy Hub
[4] Cyberhaven 2026 AI Adoption and Risk Report