AI Usage Data Still Misses the Full Picture
Researchers led by Anka Reuel and Shayne Longpre launched the AI Observatory on August 18, 2026, using real conversations across major models to test how people are actually using generative AI. The headline finding is that public AI usage data from vendors can miss a large share of personal, sensitive, and model-specific behaviour that leaders may need to account for. According to MIT Technology Review’s coverage of the launch, the gap matters because businesses and policymakers are making high-stakes decisions from an incomplete view.
AI usage data from company reports misses key behaviours
The AI Observatory aggregates 24,521 conversations and 85,633 turns from seven consent-based datasets collected between 2023 and 2025. Those conversations span 5,000 users and 52 models, including ChatGPT, Claude, Gemini, and Grok. That scale is much smaller than vendor-owned datasets, but it is large enough to show a pattern: the public story about generative AI adoption depends heavily on what each lab chooses to include.
That is the core tension in this research. Anthropic’s Economic Index and OpenAI’s usage reports are widely cited, but both are built around internal datasets and internal framing. Reuel told Technology Review, “There is no independent source to corroborate it,” which is exactly why the Observatory matters as a second lens rather than a replacement for company reporting.
The trade-off is important. Independent AI research broadens access, but because these conversations were voluntarily shared, the dataset may still underrepresent the most sensitive uses. So the lesson is not that vendor reports are wrong and outside researchers are right. It is that neither source should be treated as complete on its own.
What the data shows about work, personal use, and risk
One of the clearest findings is how much context disappears when reports focus mainly on work and productivity. When the Observatory team applied Anthropic’s filtering approach to its own data, 48% of conversations would have been excluded. That is nearly half the sample.
The excluded conversations were more likely to include health and relationships, adult or illicit topics, harassment and hate, and sexual content. Technology Review reports that health and relationship topics appeared at 44.2% in the filtered-out set versus 31.2% in Anthropic’s analysis. Adult or illicit topics appeared at 7.9% versus 2.1%. Harassment and hate showed up at 27.5% versus 5.66%, and sexual content at 16.7% versus 2.4%.
Those differences do not just change the tone of the data; they change the decisions leaders might make from it. A company that relies only on productivity-centred dashboards may conclude that employee AI use is mostly task assistance, when real behaviour can include advice-seeking, emotional reliance, or edge-case misuse. For teams trying to interpret signals before rolling out policy, this is less a reporting problem than a measurement problem.
That is where strategic oversight becomes relevant. If an organisation is treating public benchmarks as a proxy for internal behaviour, it often needs a stronger operating layer for interpretation, not just more dashboards. In practice, this is the kind of issue a fractional AI director engagement is meant to surface early: what leaders think AI is being used for versus what their own logs, workflows, and exceptions actually show.
Different models attract different kinds of use
The Observatory also found that AI model usage is not interchangeable. Users appear to sort models by perceived strength, tone, and tolerance for certain tasks.
According to the research cited by Technology Review, Grok and Gemini were used more often for information retrieval. Grok stood out for news and politics, and that same concentration also appeared to attract more misinformation-related activity. That lines up with prior outside reporting on Grok’s reliability problems, including analysis from Reuters on xAI and Grok and ongoing concerns around election and current-events accuracy.
Claude usage skewed more toward coding. ChatGPT usage was more associated with homework help. Gemini showed more social and roleplay behaviour. Even versions within the same family behaved differently: researchers found shorter conversations with GPT-3.5 and longer, more iterative exchanges with GPT-4o.
That variation matters for enterprise planning because many internal AI policies still group tools into one general category. In reality, model choice shapes user behaviour. A team using Claude for coding support faces different oversight needs than a team using ChatGPT for research assistance or Gemini for more open-ended interaction. Recent Stanford HAI AI Index research has made a similar point from another angle: usage patterns often reveal more about risk than model marketing does.
The bigger problem is that labs control the lens
The deeper issue in this story is not just that some numbers are missing. It is that the most important AI research data still sits inside proprietary systems, where outsiders cannot test claims consistently across models or over time.
David Widder of the University of Texas at Austin, who was not involved in the project, argued that a bird’s-eye view matters because sectioned-off reports make it harder to understand how different uses connect. Shayne Longpre put it more directly in the article: “No single company report tells the whole story.” That is a concise summary of the market problem.
For leaders in technology, media, and the public sector, this creates a practical gap. Vendor dashboards are useful for directional insight, but they are not neutral evidence. External academic projects can add independence, yet they rarely have access to the same scale. Groups such as the OECD’s AI policy observatory and NIST’s AI Risk Management Framework resources have pushed for better measurement and governance, but they still depend on better underlying visibility than the market currently provides.
The result is a familiar pattern: organisations set policy, approve spend, or restrict tools based on partial evidence. That can produce both underreaction and overreaction. If usage looks cleaner than it really is, risk controls arrive too late. If anecdotal misuse dominates the conversation, productive adoption can stall for the wrong reasons.
What businesses should do when AI usage data is incomplete
The immediate takeaway is not to ignore vendor research. It is to treat it as one signal among several. Public reports from Anthropic and OpenAI remain valuable, especially for directional trends in Claude usage, ChatGPT usage, and broader generative AI adoption. But they should be triangulated against internal usage logs, workflow interviews, help-desk tickets, and exception reports.
Three operating practices stand out.
First, separate model-level analysis from category-level analysis. If one team is using Grok for current-events retrieval and another is using Claude for code generation, their controls should not look identical.
Second, measure non-productive and ambiguous uses, not just ROI-friendly ones. Advice-seeking, emotional reliance, and policy edge cases often appear before they show up in formal incident logs.
Third, revisit baselines frequently. The Observatory found that conversations changed over time, with more small talk, longer exchanges, and lower rates of some sensitive behaviour as safeguards improved. Internal policy that assumes usage in 2026 looks like usage in 2024 will age quickly.
What to watch next is whether the AI Observatory can expand beyond consent-based datasets into a broader standard for independent AI conversation analysis. The other open question is whether major labs will share privacy-protected data with outside researchers at all. Until that happens, AI usage data will remain useful, but incomplete, and strategy teams will need to act with that uncertainty in plain view.
Martin Kuvandzhiev
CEO and Founder of Encorp.io with expertise in AI and business transformation