Tag: AI Research

  • Nobody read the conversations behind this AI usage data

    Nobody read the conversations behind this AI usage data

    Table of contents

    Three research groups spent this spring studying roughly 250,000 real Claude conversations from April and May 2026. Two university labs and one nonprofit, all working on AI usage data pulled straight from live production traffic.

    None of them read a single conversation.

    That is the design working as intended, and it is the part nobody is talking about.

    What Anthropic shipped

    The tool is called Anthropic Insights, and it used to be called Clio. A researcher writes one question, and Claude runs it against every conversation in the sample. The answers get sorted into categories, and the researcher sees the category names plus the percentage of conversations in each one. That is the whole output, with no transcripts and no raw text. Anthropic’s own example of such a question is “What type of guidance is this person asking for?”

    The AI usage data never leaves Anthropic in raw form. Only the counts do.

    Anthropic says this is the first time external researchers have run public independent studies on an AI company’s own AI usage data. That claim holds up.

    AI usage data

    The contracts are solid too, better than I expected from a lab publishing its own report card. Review rights covered user privacy and anything that could help people break usage policy. They also covered Anthropic’s own confidential information and the accuracy of the research. The agreements say partners can publish findings “even when they are inconvenient for Anthropic”. Imperial College London ran a third-party privacy audit. The aggregate data is on HuggingFace already, so the thing is real.

    The privacy guarantee is the audit hole

    Anthropic wrote this themselves, which is to their credit, because they could have left it out.

    “Because we are relying on Claude’s judgments, the tool is sensitive to a question’s wording; a poorly phrased one can place conversations into categories that misrepresent them.”

    Then comes the sentence that should have been the headline.

    “Because no one can read the underlying conversations, these errors are hard to catch.”

    Read that twice, because the loop it describes is closed. The instrument doing the measuring is a language model and the thing being measured is a language model. The safeguard that protects users is the same safeguard that stops anyone from checking whether the AI usage data means what the category labels say it means.

    If you lean on a model’s judgment, you should have a prior. When I went through seven open-weight models mapping threat reports to ATT&CK, they got it wrong four times out of five.

    Internally Anthropic handles this by rewriting the question over and over for weeks. External partners could not, because every new dataset needs another privacy review and the study would never finish.

    Why the practice sandbox does not match production

    The workaround was to have the researchers tune their questions on WildChat, a public dataset of human-AI conversations where they could read the underlying text and check whether the categories made sense. Then they took the tuned question and pointed it at live Claude traffic.

    Some questions that worked well on WildChat produced misleading categories once applied to actual Claude conversations, because WildChat leans casual and creative and Claude traffic does not. Two pools of AI usage data, two different populations.

    Anyone who has written a detection rule knows this shape. The rule fires clean against your test corpus, then drowns you in false positives the first hour it sees real traffic. It is the same failure in a different field.

    WildChat also carries its own baggage, which I covered in Dossier 33 through the Truffle Security scan of 7.6 petabytes of HuggingFace training data. One Infura key, pasted once into a ChatGPT conversation, got captured by WildChat and copied onward into 1,131 public datasets and 10,162 file locations.

    So the AI usage data you are allowed to read is the batch where mistakes are permanent. The production data where mistakes get corrected is the batch nobody may open. That is one trade, made twice.

    What it means when you quote AI usage data

    The findings themselves are worth having. Stanford’s SALT Lab found that over half of Claude conversations involved people delegating consequential tasks to AI, and in nearly three-quarters of them people set the direction while Claude assisted.

    Before you drop that into a board deck, know what you are holding. That is Claude’s judgment on what counts as “consequential”, turned into a percentage. The researcher who published it cannot go back and check a single case. Anthropic also runs studies where people answer for themselves, like the 81,000-person survey I wrote up earlier, and a model inferring intent from a transcript is a different instrument. The number can still be true. It is a different animal than a measurement, and the difference matters the moment someone writes policy on top of it.

    There is one more filter on this AI usage data. Anthropic removed or altered any category that described the method users found for getting around safeguards. Categories covering what users attempted stayed in. Less than 5% of categories and conversations in each study, and they told the researchers which clusters were touched and why. That is honest handling. It is still the lab editing the dataset before the auditor sees it.

    Four questions before you cite an AI usage data study

    Steal these. They work on any usage research that reaches you, this pilot included.

    1. Find out who read the primary records. If a model did the reading, say so out loud when you quote the number.

    2. Get the exact wording of the question the model answered. The wording sets the categories, and the categories are the finding.

    3. Ask what was removed before publication, and whether anyone told you. Anthropic did tell its partners which clusters were touched, which most companies will not.

    4. Check whether a second team could reproduce the result on the same records. On live lab traffic today, nobody can.

    The part that outlives this pilot

    Watch where this goes next. Regulators are going to want exactly this kind of access to AI usage data, and whatever shape gets standardised here becomes the template for every AI audit that follows.

    If that template says the auditor never reads the primary evidence, then “independent audit” comes to mean “independent question, lab’s answer”. Anthropic went further than anyone else has, so their pilot sets the default.

    I would still take this over nothing, easily. The direction is spot on.

    If I were a researcher I would fill in the form. What I would not do is treat these percentages the way I treat a log line I pulled myself.

    Before you repeat a number about AI usage data, find out who read the primary records. A model doing that reading leaves you with a hypothesis. Cite it like one.

    Source | https://www.anthropic.com/research/enabling-independent-research


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • The Largest Anthropic Study Reveals What 81,000 People Really Want From AI

    Table of contents

    The recent Anthropic study reveals exactly what over 80,000 people expect from artificial intelligence in their daily lives. Last December, researchers completely changed their approach to data collection.

    Instead of sending out boring multiple-choice surveys, they used a large language model as an active interviewer to conduct deep, qualitative conversations with exactly 80,508 users across 159 countries, speaking 70 different languages. This massive Anthropic study successfully bridged the gap between small-scale intimacy and large-scale volume. It gathered raw, open-ended insights proving that society wants much more than just faster email generation.

    9 Core Human Aspirations: From Escaping the Office to Curing Diseases

    The research identified nine distinct clusters of human aspirations. These are not sci-fi movie plots, but highly practical needs born from the friction of modern work and life.

    1. Professional excellence (18.8%): Users want to dump boring documentation. Instead of spending 2 hours a day filling out CRM fields or formatting Excel tables, they want to focus on solving complex strategic problems.
    2. Personal transformation (13.7%): People use software as a highly objective mental health coach or habit tracker. They seek guidance for behavior change without the fear of human judgment.
    3. Life management (13.5%) and Time freedom (11.1%): The machine acts as cognitive scaffolding. Users delegate schedule planning and household budgets to reclaim a full 3 hours a week for family rest.
    4. Financial independence (9.7%) and Entrepreneurship (8.7%): In regions lacking tech infrastructure, software serves as a brutal equalizer. It acts as a force multiplier, allowing users to build businesses with zero starting capital.
    5. Societal transformation (9.4%): Users hope computing power will help discover cures for chronic diseases, optimize energy grids, and democratize education in the poorest nations.
    6. Learning & growth (8.4%) and Creative expression (5.6%): The algorithm acts as a patient tutor available at 2:00 AM, stripping away the shame of asking basic math or coding questions.

    Are Algorithms Actually Delivering Results?

    According to the Anthropic study, 81% of respondents report that software is already taking concrete steps toward their vision. The biggest wins happen in highly specific areas. First, productivity (32.0%) drastically increases, allowing workers to automate repetitive tasks and close projects days earlier. Second, cognitive partnership (17.2%) turns the machine into a ruthless brainstorming partner.

    Technical accessibility (8.7%) also shows incredible, concrete results. The Anthropic study highlights a mute user who independently built a text-to-speech bot without knowing how to write a single line of code. This proves how effectively algorithms remove physical barriers between human imagination and final execution.

    Light and Shade: The 5 Tensions of Automation

    User fears are highly specific, with respondents voicing an average of 2.3 distinct worries. The top concerns involve system unreliability (26.7%), the economy and job loss (22.3%), and the total loss of human autonomy (21.9%).

    anthropic study

    The Anthropic study defines these contradictions as the “light and shade” of automation. While 33% see massive educational benefits, 17% are paralyzed by the fear of cognitive atrophy—the literal loss of the ability to read and think independently. University professors and teachers report witnessing this exact atrophy nearly three times more often than other professions.

    The time-saving paradox is equally painful. Half of the surveyed users praise saving work hours, yet 18% feel they are just running faster on a treadmill because managers instantly increase output quotas to consume the saved time. Furthermore, while 16% find emotional solace in the machine, 12% fear becoming dangerously dependent on it for basic human interaction.

    The Global Divide: How Geography Dictates Fear and Hope

    Globally, 67% of interviewees express a net positive sentiment. However, the geographic breakdown in the Anthropic study reveals drastically different motivations based on local economies.

    Optimism peaks in lower- and middle-income countries like Nigeria, Mexico, and Vietnam. Here, technology acts as a direct ladder for social mobility and global trade. Sub-Saharan Africa views the algorithm as a mechanism to bypass historical capital limits and launch competitive start-ups.

    Conversely, wealthier regions like North America, Western Europe, and Oceania worry deeply about governance, data privacy, and corporate layoffs. North American users primarily want life management tools to survive modern complexity, while East Asian respondents focus heavily on internal, personal transformation.

    Step-by-Step Guide: Reclaiming Your Time

    To avoid cognitive atrophy and actually benefit from the findings of this Anthropic study, implement these three concrete rules into your daily routine today:

    1. Delegate Only the Repetitive: Audit your work week. Identify tasks that consume more than 45 minutes a day but require zero creative judgment (e.g., categorizing inbox messages, formatting CSV cells). Hand these exclusively to the machine.
    2. Physically Block Reclaimed Time: If an app saves you 2 hours on a Thursday, immediately block 120 minutes in your calendar for a gym session, a walk, or reading a physical book. Do not let your boss fill that void with more Slack messages.
    3. Enforce Hard Verification: Never trust generated outputs blindly. Always allocate a strict 15-minute block for full human verification of logic, dates, and facts before hitting send.

    Ultimately, society demands technology that helps us live better, not just work faster on the assembly line. Whether it is a doctor reclaiming 20 minutes for a patient consultation or a student overcoming math anxiety, we want software to fill the critical gaps in our human experience.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today