Just 8.5% of privacy policies explicitly disclose that user data trains an AI model, according to UpGuard's 2025 analysis of 176 privacy notices pulled from its 250 most-monitored vendors. Another 43.2% of those same notices mention AI somewhere in the document without ever answering the training question directly, and more than half say nothing about AI at all. The gap between how many companies use AI and how many actually say what they do with your data in the process is the widest disclosure gap in privacy policy research right now.
That gap is what this page documents: how many privacy policies mention AI training, how the biggest AI chatbot makers handle the same question in their own policies, and how worried the public actually is about it.
How many privacy policies mention AI training?
8.5% of privacy policies explicitly disclose that user data trains an AI model. UpGuard reached that figure by running 176 privacy notices from its 250 most-monitored vendors through a GPT-4o classification pass in 2025, checking each notice for three things: whether it mentions AI, whether it says data trains a model, and whether there is an opt-out. Just 15 of the 176 notices gave a clear yes to the training question.
The more common outcome is silence, not denial. Of the 76 notices that mention AI in any capacity, 58 leave the training question unanswered entirely, using vague language about "improving our services" or "product development" that a reader, or a GPT-4o classifier, cannot map cleanly to yes or no. Only 3 companies, 1.7% of the full sample, state plainly that user data is not used for training.
Figure 1: How 176 privacy notices handle the AI training question. Source: UpGuard, "The Risk of Third-Party AI Trained on User Data" (April 2025).
| Disclosure category | Notices | Share of 176 |
|---|---|---|
| No mention of AI at all | 100 | 56.8% |
| Mentions AI, training status unclear | 58 | 33.0% |
| Explicitly discloses AI training | 15 | 8.5% |
| Explicitly denies AI training | 3 | 1.7% |
Source: UpGuard, "The Risk of Third-Party AI Trained on User Data" (2025).
A privacy policy that never mentions AI is not automatically the safest choice. It just means the reader has no way to know either way, which is its own kind of risk for anyone trying to comply with disclosure rules under GDPR, CCPA, or similar laws.
How many privacy policies mention AI at all?
43.2% of privacy policies mention AI in some capacity, per UpGuard's 176-notice sample, up from what privacy researchers describe as a near-nonexistent share of policies before generative AI tools became mainstream products in 2023. That figure covers any AI mention, from a single sentence about using "machine learning to improve our products" to a dedicated AI section with opt-out instructions.
The remaining 56.8% of notices UpGuard reviewed do not mention AI at all, even though the underlying company may already run AI features. That silence is not necessarily unlawful; not every company processes personal data through an AI system, and a policy has no obligation to discuss AI it does not use. But for the 43.2% that do mention AI, regulators in the EU and several US states increasingly expect a policy to say specifically whether personal data is used to train a model and whether a person can opt out, not just that "AI is used to enhance your experience."
Vague AI language is a compliance gap waiting to surface, not a shortcut around one. It fits a wider pattern: our AI and privacy statistics for 2026 roundup found that 90% of organizations expanded their privacy programs specifically because of AI, yet most of that expanded effort has not reached the actual policy text customers read.
Do major AI chatbots train on your data by default?
All 6 of the 6 frontier AI developers studied by Stanford's Cyber Policy Center train on user chat data by default. The September 2025 analysis covered Amazon (Nova), Anthropic (Claude), Google (Gemini), Meta (Meta AI), Microsoft (Copilot), and OpenAI (ChatGPT), and found that every one of them uses chat inputs to improve their models unless a user actively opts out. Anthropic was the last holdout, switching Claude from opt-in to opt-out training in September 2025, the same month the study published.
Default training is only part of the picture. 3 of the 6 companies, Amazon, Meta, and OpenAI, retain some or all chat data indefinitely rather than on a defined schedule, while Google, Anthropic, and Microsoft specify a retention window. 3 of the 6 also allow at least some data from users under 18 into training sets in some form, and 2 of the 6, Microsoft and Meta, explicitly combine chatbot conversations with data from a person's other products at the same company.
Figure 2: How consistently 6 frontier AI developers handle chat-data training. Source: Stanford Cyber Policy Center, "User Privacy and Large Language Models" (September 2025).
| Practice | Companies | Share of 6 studied |
|---|---|---|
| Trains chat data by default | 6 | 100% |
| Offers an opt-out | 5 | 83.3% |
| Retains data indefinitely | 3 | 50% |
| Includes minors' chat data | 3 | 50% |
| Combines with other product data | 2 | 33.3% |
Source: Stanford Cyber Policy Center, "User Privacy and Large Language Models" (2025).
An opt-out at 5 of 6 companies sounds reassuring until you notice it is opt-out, not opt-in: the default in every single case studied is that a conversation becomes training data unless the user finds and flips a setting.
How worried are consumers about AI training on their data?
65% of consumers say they are worried about their data being used to train AI, according to Verve's December 2025 consumer survey, which also found that 97% of respondents want more transparency from the companies and publishers collecting their data. That concern is not confined to one survey or one framing of the question.
Shift Browser's March 2026 survey of 1,448 US respondents, weighted to be nationally representative by income, ethnicity, age, gender, and region, found that 81% are concerned about AI systems accessing personal data or private conversations, the single largest privacy worry the survey measured. Usercentrics reported a similar pattern in July 2025: 59% of consumers said they were uncomfortable with their data being used to train AI systems, even before accounting for people who simply were not sure how their data was being used.
Figure 3: Three separate consumer surveys, asked three different ways, land in the same range. Source: Shift Browser (March 2026), Verve (December 2025), Usercentrics (July 2025).
The exact wording differs across all three surveys, so the numbers are not directly comparable, but the direction is consistent: a majority of consumers across every recent survey are uneasy about AI and their personal data, and the 8.5% explicit-disclosure rate in privacy policies is nowhere close to matching that level of public concern.
Why are privacy policies scrambling to address AI now?
Generative AI adoption inside companies has moved faster than the privacy documentation meant to cover it. McKinsey's State of AI global survey found that regular use of generative AI in at least one business function climbed from 33% of organizations in 2023 to 65% in 2024 and 71% in 2025, more than doubling in two years. Privacy teams that had no reason to mention AI in 2022 now sit inside companies where AI touches customer data almost everywhere.
Figure 4: Generative AI adoption nearly doubled between 2023 and 2025. Source: McKinsey & Company, State of AI global survey, 2023 to 2025 editions.
Public disclosure fights have followed the same curve. In August 2023, Zoom walked back a terms-of-service clause that implied customer audio and video could train its AI models, eventually committing that it would not use customer content for AI training at all after a public backlash, per TechCrunch's reporting at the time. Meta went the other direction in June 2024, updating its privacy policy to allow public posts and comments across Facebook, Instagram, WhatsApp, and Threads to train its AI models, while offering EU and UK users a manual opt-out form.
Figure 5: Company disclosure decisions have kept arriving faster than the research measuring them. Sources: TechCrunch (2023), MIT Technology Review (2024), UpGuard (2025), Stanford Cyber Policy Center (2025), gHacks Tech News (2026).
Atlassian is the most recent example: starting August 17, 2026, the company defaults Free, Standard, and Premium tier customers into AI training on Jira and Confluence data, with an opt-out available only to Enterprise-tier accounts, according to gHacks Tech News. Each of these cases became a story because the underlying privacy policy language changed faster than most readers, or most competitors, expected.
What should a privacy policy actually say about AI training?
A privacy policy that clears UpGuard's classification test needs to answer three questions in plain language: does the company use AI, does that AI train on personal data, and can a person opt out. Most of the 176 notices UpGuard reviewed answer none of the three clearly, which is exactly the gap that turns into a regulator's question or a journalist's story.
Figure 6: The three-question test that separates a compliant AI disclosure from a vague one. Source: UpGuard classification methodology, adapted (2025).
A business that has added any AI feature, from a support chatbot to a product recommendation model, needs its privacy policy to state plainly whether user inputs feed that model, how long the data is kept, and how someone opts out if that option exists. Businesses updating their own disclosures for the first time can generate an AI-ready privacy policy that covers data collection, third-party AI processing, and opt-out mechanics in the same document, rather than bolting an AI clause onto language that was written before generative AI existed as a product category. Vague phrases like "to improve our services" no longer meet the bar that 43.2% of AI-mentioning companies are already failing to clear.
The Bottom Line
8.5% of privacy policies tell you plainly whether your data trains an AI model. The other 91.5% either say nothing about AI or say just enough to raise the question without answering it, even as 6 of 6 major AI chatbot makers train on chat data by default and most consumers say they are uncomfortable with exactly that practice. The disclosure gap is not a matter of AI being new; McKinsey's adoption numbers show most companies have used AI in some form since at least 2024. It is a matter of privacy policies not catching up to what the AI already does. For any business running AI against customer or user data, closing that gap is a compliance requirement in an increasing number of jurisdictions, not just good practice, and it starts with a policy that answers the same three questions UpGuard used to grade 176 companies: does the AI exist, does it train on personal data, and can a person say no. That silence has a measurable cost: on the employer side, 68% of large US employers now report a written AI policy, yet on the consumer side only 18% of Americans say they trust AI platforms with their personal data, the lowest trust figure of any AI-related question asked in 2026.
Frequently Asked Questions
How many privacy policies mention AI training on user data? Just 8.5% of privacy policies, 15 of the 176 analyzed, explicitly disclose that user data trains an AI model, according to UpGuard's 2025 study. Another 33.0% mention AI without ever saying whether user data feeds a model.
How many privacy policies mention AI at all? 43.2% of privacy policies mention AI in some capacity, 76 of the 176 notices UpGuard analyzed in its 2025 study of its 250 most-monitored vendors, though most of those stop short of a clear training disclosure.
Do major AI chatbots train on your data by default? Yes. All 6 of the 6 frontier AI developers studied by Stanford's Cyber Policy Center in September 2025, Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI, train on user chat data by default, and 3 of the 6 retain that data indefinitely.
How worried are consumers about their data training AI models? 81% of US respondents say they are concerned about AI systems accessing personal data or private conversations, per Shift Browser's March 2026 survey of 1,448 people, and 65% specifically worry about AI training on their data, per Verve's December 2025 survey.
Where the Numbers Come From
- UpGuard. (2025). "The Risk of Third-Party AI Trained on User Data." 176 privacy notices analyzed from UpGuard's 250 most-monitored vendors, GPT-4o classification methodology, published April 16, 2025.
- Stanford Cyber Policy Center. (2025). "User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies." Six developers studied, published September 5, 2025.
- McKinsey & Company. "The State of AI: Global Survey." Generative AI adoption tracked across the 2023, 2024, and 2025 editions.
- Shift Browser, as reported by ppc.land. (2026). Consumer survey of 1,448 US respondents, nationally representative weighting, published March 3, 2026.
- Verve, as reported by ppc.land. (2025). Consumer survey on AI data training concerns, published December 2025.
- TechCrunch. (2023). "Zoom Knots Itself a Legal Tangle Over Use of Customer Data for Training AI Models." August 8, 2023.
- MIT Technology Review. (2024). "How to Opt Out of Meta's AI Training." June 14, 2024.
- gHacks Tech News. (2026). "Atlassian Will Collect Jira and Confluence Data by Default to Train AI Models." April 19, 2026, effective August 17, 2026.
Note: All figures verified as of July 2026. The UpGuard and Stanford studies are single-cycle analyses rather than live-updating dashboards; consumer survey figures (Shift Browser, Verve, Usercentrics) reflect the specific sample and wording of each named survey and are not directly comparable to one another. This page is scheduled for refresh at least twice a year as new disclosure studies publish.