Shadow AI is the use of generative AI tools and personal AI accounts for work without IT approval — and it is now the default, not the exception. Some 77% of employees who use genAI paste company data directly into prompts, and 82% of those pastes run through personal accounts your security stack never sees. The fix is not a blanket ban, which reliably backfires; it is governed enablement: visibility, guardrails on what leaves, and a sanctioned alternative people actually want to use.
Key takeaways
- Adoption is total: 99% of European organisations have active genAI users, and monthly active use has grown from 35% to 65% of the workforce (Netskope Threat Labs, January 2026).
- The data is increasingly sensitive: 39.7% of corporate data flowing into AI tools now contains sensitive information, up from 10.7% two years earlier (Cyberhaven, February 2026).
- It already costs money: 1 in 5 organisations suffered a breach tied to shadow AI, adding an average of $670,000 to breach costs — and 97% of AI-breached organisations lacked basic AI access controls (IBM, 2025).
- Blocking alone fails: even with controls, 43% of European employees still access personal AI apps, shifting the risk to devices and networks IT cannot see.
- The workable stack is layered: SWG visibility, DLP inspection of prompts and uploads, tenant restrictions that block personal logins, a sanctioned enterprise AI, and isolation for the edge cases.
How big is the problem, really?
The telemetry is unambiguous. LayerX browser data shows 45% of employees actively use genAI at work; of those, 77% paste corporate data into prompt boxes, at an average of 14 pastes per day for unsanctioned users, of which at least three contain sensitive data. Zscaler counted 18,033 terabytes flowing into AI/ML apps in a single year (+93% year-over-year) and over 410 million DLP violations tied to ChatGPT alone. The average organisation logs 223 genAI policy violations per month; the top quartile logs more than 2,100. In Europe the leaked material splits into regulated data (59%), source code (15%), intellectual property (13%) and credentials or API keys (12%), per the Netskope Threat Labs Europe report.
Concentration makes control feasible, though: 92% of unsanctioned activity runs through a single tool — ChatGPT.
What actually happens to pasted data?
The consumer/enterprise split is the single most important fact to teach your organisation, because the same tool behaves completely differently per account type.
| Provider & tier | Model training on your data | Retention default | Enterprise controls |
|---|---|---|---|
| ChatGPT Free/Plus/Pro (consumer) | On by default | Indefinite until user deletes; 30-day abuse log | Manual opt-out only |
| ChatGPT Team/Enterprise | Off by default | Admin-set (min. 90 days on Enterprise) | SSO, EKM, EU data residency, tenant isolation |
| OpenAI API | Off by default | 30 days (abuse monitoring); Zero Data Retention available | DPA, ZDR for qualifying accounts |
| Google Gemini (consumer) | On; human review possible | Up to 3 years, de-identified | Activity pause only |
| Google Workspace with Gemini | Off; tenant boundary enforced | Workspace terms | Workspace DLP and admin controls apply |
| Microsoft 365 Copilot (commercial) | Off; stays in tenant | M365 retention policies | Purview DLP, conditional access, audit logs |
| Claude Team/Enterprise | Off by default | Custom retention | SAML SSO, admin console, EU-grade DPAs |
Samsung learned this in April 2023, when engineers pasted semiconductor source code and meeting notes into consumer ChatGPT three times — data that flowed into training under the then-defaults, prompting a company-wide ban and an internal AI build. The follow-on risk is subtler: in 2024, researchers found over 225,000 compromised OpenAI credentials on dark-web markets, harvested by infostealers — every saved chat history behind those logins, including pasted corporate data, went with them. Developers add a third path: GitGuardian counted 28.65 million hardcoded secrets on public GitHub in 2025 (+34%), accelerated by AI assistants regurgitating live credentials from prompt context back into committed code.
Why blanket blocking backfires
Blocking chatgpt.com at the firewall does not remove the demand; it removes your visibility. European telemetry shows 43% of employees still access personal AI apps despite controls, and hard blocks push work to phones, personal laptops and hotspots where no DLP, logging or identity control applies. Meanwhile 45% of workers say they rely on genAI to keep up their daily output, so a ban taxes productivity while the leakage continues elsewhere. Blocking has a role — for unvetted high-risk tools (Europe’s most-blocked apps include Particular Audience at 44%, ZeroGPT at 37% and DeepSeek at 36%) — but as the whole strategy it is theatre.
The control stack that works
Governed enablement layers five controls, each covering the previous one’s blind spot:
| Control | What it does | Limit to cover next |
|---|---|---|
| 1. SWG visibility & category control | Inspects outbound traffic, blocks unvetted AI domains, reveals what is actually in use | Cannot distinguish personal vs corporate accounts on allowed apps; bypassable off-network |
| 2. Inline DLP on prompts & uploads | Inspects pasted text and files in real time; blocks PII, code and keys before they leave; coaches users in the moment | Needs TLS inspection and pattern tuning |
| 3. CASB tenant restrictions | Injects headers so chatgpt.com only accepts your corporate tenant login; personal @gmail logins fail | Depends on vendor header support and IdP integration |
| 4. Sanctioned enterprise AI | ChatGPT Enterprise / Copilot / Workspace Gemini under a DPA, SSO and zero-training terms — the productive outlet | Does not stop alternatives by itself; needs layers 1–3 |
| 5. Browser isolation for edge cases | Read-only rendered sessions with paste and upload disabled, for high-risk tools or untrusted users | Reserve for where it fits; latency cost if overused |
Discovery comes first, and it will humble you: IT leadership typically estimates 3–5 AI apps in use, while the median enterprise touches 9.6 distinct genAI apps daily and the top quartile more than 24. You cannot write policy for a landscape you have not measured — the same visibility argument made in CASB in SASE and the secure web gateway guide.
The policy layer: what your AI acceptable-use policy must contain
- A data classification matrix: what may enter AI tools (public marketing copy) versus never (customer PII, health data, source code, unreleased financials, credentials).
- An approved app registry: everything unlisted is shadow AI by definition.
- Account enforcement: corporate SSO only; personal accounts for work tasks prohibited.
- Human-in-the-loop review of AI output before publication, customer delivery or production code.
- IP and licensing rules for generated content and AI-completed code.
- A no-blame reporting protocol for accidental submissions — you want to hear about them within the hour, not never.
Anchor the policy to a published framework — ISO/IEC 42001 (AI management systems), the NIST AI Risk Management Framework, or the OWASP Top 10 for LLM applications — so auditors have something to map against.
The EU regulatory clock is running
Three regimes now touch shadow AI directly. Under GDPR, pasting personal data into a US-hosted consumer AI tool is processing without a DPA (Article 28), an unlawful transfer without valid safeguards (Chapter V), and a likely breach of data minimisation (Article 5) — with a DPIA required before systematic genAI processing (Article 35). Under the EU AI Act, prohibitions have applied since February 2025 and the high-risk obligations of Annex III become fully enforceable on 2 August 2026, covering workplace AI used for recruitment, evaluation or task allocation, with penalties up to €35 million or 7% of worldwide turnover. And under NIS2, ungoverned AI data flows are exactly the kind of unmanaged risk Article 21 requires you to control. Cyber insurers have joined in: renewal questionnaires now routinely ask for a written AI policy, technical data-leak controls and enforced account isolation — gaps show up as surcharges, exclusions or refusals.
Govern the flow, keep the productivity
The organisations getting this right run one play: discover what is in use, block only the genuinely dangerous, steer everyone to a sanctioned enterprise tenant, and inspect what leaves in real time. That is not a new toolchain — it is the SWG, DLP and app-control layer of a SASE platform doing its job on a new traffic category, with coaching pop-ups replacing the ban hammer. For EU mid-market teams there is a second-order advantage in doing this on an EU-sovereign platform like Jimber: the very controls that stop your data flowing into US consumer AI tools should not themselves route your traffic through a US-jurisdiction inspection cloud. One platform, web gateway to DLP, EU data residency, flat pricing. Book a demo to see your organisation’s real AI app inventory within a week of deployment.
Frequently asked questions
How do employees leak data through ChatGPT without uploading files?
Mostly via the clipboard: 77% of genAI users paste text straight into prompts — code snippets, customer emails, financials. On consumer tiers those pastes are retained indefinitely and used for training by default, and they persist in chat histories that credential thieves can access.
Is blocking chatgpt.com at the firewall enough?
No. Blocks push usage to personal devices and hotspots where IT has zero visibility, and they miss browser extensions and AI features embedded in other SaaS tools. Even in controlled environments, 43% of European employees still reach personal AI apps. Block unvetted tools; govern the rest.
What is the difference between ChatGPT Free and ChatGPT Enterprise for data security?
Free and Plus retain prompts indefinitely and train on them by default. Enterprise disables training by default, isolates data in your tenant, supports SSO and encryption key management, and lets admins set retention. The tool is the same; the data handling is opposite.
Does Google Workspace with Gemini train on our documents?
No. Workspace customer content stays within the tenant boundary and is not used to train public Gemini models without explicit admin consent, and human review is excluded. The consumer Gemini app is a different regime: prompts may be reviewed and retained up to three years.
Why does pasting customer data into a consumer AI tool violate GDPR?
Three ways at once: processing via a vendor without a data processing agreement (Article 28), an unlawful cross-border transfer without valid safeguards (Chapter V), and a breach of minimisation and purpose limitation when the provider retains data for training (Article 5).
How do tenant restrictions stop personal accounts on work devices?
A CASB or gateway injects an HTTP header on traffic to the AI platform that tells it to accept only logins from your corporate domain. The employee still reaches chatgpt.com, but only the sanctioned company tenant will authenticate — personal logins fail.
How many AI tools are employees actually using?
Far more than expected. IT teams typically guess 3–5; discovery reveals a median of 9.6 distinct genAI apps in daily use, 24+ in the top quartile of organisations, and 80+ in the top 1%.
What does the EU AI Act require for workplace AI by August 2026?
From 2 August 2026, high-risk workplace systems under Annex III — AI used in recruitment, performance evaluation or task allocation — require risk management, operational logging, human oversight and prior notification of employees, with penalties up to €35 million or 7% of worldwide turnover.