Tag: ai policy

  • Nobody read the conversations behind this AI usage data

    Nobody read the conversations behind this AI usage data

    Table of contents

    Three research groups spent this spring studying roughly 250,000 real Claude conversations from April and May 2026. Two university labs and one nonprofit, all working on AI usage data pulled straight from live production traffic.

    None of them read a single conversation.

    That is the design working as intended, and it is the part nobody is talking about.

    What Anthropic shipped

    The tool is called Anthropic Insights, and it used to be called Clio. A researcher writes one question, and Claude runs it against every conversation in the sample. The answers get sorted into categories, and the researcher sees the category names plus the percentage of conversations in each one. That is the whole output, with no transcripts and no raw text. Anthropic’s own example of such a question is “What type of guidance is this person asking for?”

    The AI usage data never leaves Anthropic in raw form. Only the counts do.

    Anthropic says this is the first time external researchers have run public independent studies on an AI company’s own AI usage data. That claim holds up.

    AI usage data

    The contracts are solid too, better than I expected from a lab publishing its own report card. Review rights covered user privacy and anything that could help people break usage policy. They also covered Anthropic’s own confidential information and the accuracy of the research. The agreements say partners can publish findings “even when they are inconvenient for Anthropic”. Imperial College London ran a third-party privacy audit. The aggregate data is on HuggingFace already, so the thing is real.

    The privacy guarantee is the audit hole

    Anthropic wrote this themselves, which is to their credit, because they could have left it out.

    “Because we are relying on Claude’s judgments, the tool is sensitive to a question’s wording; a poorly phrased one can place conversations into categories that misrepresent them.”

    Then comes the sentence that should have been the headline.

    “Because no one can read the underlying conversations, these errors are hard to catch.”

    Read that twice, because the loop it describes is closed. The instrument doing the measuring is a language model and the thing being measured is a language model. The safeguard that protects users is the same safeguard that stops anyone from checking whether the AI usage data means what the category labels say it means.

    If you lean on a model’s judgment, you should have a prior. When I went through seven open-weight models mapping threat reports to ATT&CK, they got it wrong four times out of five.

    Internally Anthropic handles this by rewriting the question over and over for weeks. External partners could not, because every new dataset needs another privacy review and the study would never finish.

    Why the practice sandbox does not match production

    The workaround was to have the researchers tune their questions on WildChat, a public dataset of human-AI conversations where they could read the underlying text and check whether the categories made sense. Then they took the tuned question and pointed it at live Claude traffic.

    Some questions that worked well on WildChat produced misleading categories once applied to actual Claude conversations, because WildChat leans casual and creative and Claude traffic does not. Two pools of AI usage data, two different populations.

    Anyone who has written a detection rule knows this shape. The rule fires clean against your test corpus, then drowns you in false positives the first hour it sees real traffic. It is the same failure in a different field.

    WildChat also carries its own baggage, which I covered in Dossier 33 through the Truffle Security scan of 7.6 petabytes of HuggingFace training data. One Infura key, pasted once into a ChatGPT conversation, got captured by WildChat and copied onward into 1,131 public datasets and 10,162 file locations.

    So the AI usage data you are allowed to read is the batch where mistakes are permanent. The production data where mistakes get corrected is the batch nobody may open. That is one trade, made twice.

    What it means when you quote AI usage data

    The findings themselves are worth having. Stanford’s SALT Lab found that over half of Claude conversations involved people delegating consequential tasks to AI, and in nearly three-quarters of them people set the direction while Claude assisted.

    Before you drop that into a board deck, know what you are holding. That is Claude’s judgment on what counts as “consequential”, turned into a percentage. The researcher who published it cannot go back and check a single case. Anthropic also runs studies where people answer for themselves, like the 81,000-person survey I wrote up earlier, and a model inferring intent from a transcript is a different instrument. The number can still be true. It is a different animal than a measurement, and the difference matters the moment someone writes policy on top of it.

    There is one more filter on this AI usage data. Anthropic removed or altered any category that described the method users found for getting around safeguards. Categories covering what users attempted stayed in. Less than 5% of categories and conversations in each study, and they told the researchers which clusters were touched and why. That is honest handling. It is still the lab editing the dataset before the auditor sees it.

    Four questions before you cite an AI usage data study

    Steal these. They work on any usage research that reaches you, this pilot included.

    1. Find out who read the primary records. If a model did the reading, say so out loud when you quote the number.

    2. Get the exact wording of the question the model answered. The wording sets the categories, and the categories are the finding.

    3. Ask what was removed before publication, and whether anyone told you. Anthropic did tell its partners which clusters were touched, which most companies will not.

    4. Check whether a second team could reproduce the result on the same records. On live lab traffic today, nobody can.

    The part that outlives this pilot

    Watch where this goes next. Regulators are going to want exactly this kind of access to AI usage data, and whatever shape gets standardised here becomes the template for every AI audit that follows.

    If that template says the auditor never reads the primary evidence, then “independent audit” comes to mean “independent question, lab’s answer”. Anthropic went further than anyone else has, so their pilot sets the default.

    I would still take this over nothing, easily. The direction is spot on.

    If I were a researcher I would fill in the form. What I would not do is treat these percentages the way I treat a log line I pulled myself.

    Before you repeat a number about AI usage data, find out who read the primary records. A model doing that reading leaves you with a hypothesis. Cite it like one.

    Source | https://www.anthropic.com/research/enabling-independent-research


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Shadow AI grows in the gap Gallup just measured

    Shadow AI grows in the gap Gallup just measured

    Table of contents

    Ask a US employee whether their own employer has rolled out AI. A decent share of them cannot answer, and shadow AI grows in exactly that kind of confusion.

    Gallup changed the question because of it. The methodology note says, “Starting in Q3 2025, Gallup added a ‘don’t know’ option to this question to capture uncertainty about AI adoption.” Results from Q3 2025 onward are no longer directly comparable with earlier measurements.

    A polling company broke its own time series because too many people had no idea what was happening inside the building they work in.

    The second number nobody quotes

    As of May 2026, 47% of US employees say their organization has implemented AI. Only 25% say the organization communicated a clear plan.

    Every headline takes the first number, but the second one is the story.

    Between “we have this thing” and “somebody told me how to use it” sits a gap, and in that gap are contracts, patient records, draft tenders, source code, whatever your people happen to be pasting today.

    Shadow AI is the default state of that gap.

    shadow AI

    Adoption is the wrong argument

    There is a long running fight about how many people use AI. Gabriel Weinberg of DuckDuckGo summed up the skeptical side in June 2026 as “one third actively using AI, one third occasionally using AI, and one third never using AI”. He cites Microsoft telemetry putting it at “more than 30 percent of the US working-age population is using AI, an increase of 3 percentage points from the end of 2025”.

    Gallup’s workplace numbers run higher. 15% of US employees use AI daily. Weekly or more is 30%, and 52% touch it at least a few times a year.

    Pick whichever camp you like, it changes nothing for my job.

    Move the user base up or down, the 25% who got a clear plan stays where it is. That fight pulls attention away from the only question that matters, which is who wrote the rulebook and who read it.

    Shadow AI is a confidentiality problem with no attacker

    Strip the vocabulary and that is all this is.

    No phishing mail, no exploit, no command and control, nothing that trips an alert. An employee opens a browser tab and pastes a client document into a chatbot to get a summary. The data leaves the organization. Most of what you bought assumes somebody is trying to break in. Shadow AI walks out the front door during working hours, moved by people who just want to finish faster.

    OWASP keeps an entry for sensitive information disclosure in its Top 10 for Large Language Model Applications, and almost all of that guidance assumes an application you built. The tab your sales team opened this morning is nobody’s application.

    Scott Brinker named the shape of this back in 2013 and called it Martec’s law. Technology changes exponentially, organizations change logarithmically. Your staff adopted AI in an afternoon, your document set moves at the speed of a committee.

    The Gallup manager numbers show the same thing from the other side. 36% strongly agree their manager supports the team using AI. Where that support exists, employees are 1.7x more likely to use AI weekly or more and 8.7x more likely to report a transformational change in how they work. Encouragement travels by conversation and rules travel by document. Shadow AI takes the faster route.

    What shadow AI looks like on a Tuesday

    Nobody sits down and decides to run shadow AI. It shows up as small, reasonable moves.

    A sales rep pastes a signed contract into a chatbot to pull the renewal dates out of it. An HR assistant drops a salary spreadsheet into a chatbot to reformat the columns. A developer sends a stack trace holding a production connection string to a free tier account. A clinic receptionist rewrites a referral letter with the patient name still in it.

    None of those people are careless. All of them were told AI makes them faster, and none of them were told where the line sits.

    Deleting the client name before pasting does not turn the text into anonymous data either, and I went through the research on that in ChatGPT privacy leak.

    Every one of those actions is invisible to the security team, because nothing was breached and nothing alerted.

    Low numbers are not safe numbers

    Daily use runs at 42% in technology, 27% in finance, 22% in professional services, and 9 to 15% everywhere else.

    Read the low end carefully, because a law firm or a clinic sitting in that bottom band is not in a better position. It is a place where a smaller group does the same thing with far more sensitive material, and with less chance that anyone in IT has ever looked at it. Low usage hides shadow AI.

    The 65% of employees in AI-implementing organizations who report a positive effect on productivity are not lying either. It works, and that is exactly why nobody is going to stop when you ask them to. A ban moves shadow AI further out of sight.

    One page before you buy anything

    Do not start with a tool. Start with one page that answers four things.

    1. Which categories of data never go into an external model.
    2. Which tools are approved, listed by name.
    3. Who an employee asks when the answer is not obvious.
    4. What happens when something has already gone in, and who hears about it first.

    That page will not cover a regulated environment or replace a contract with the vendor. It covers what is leaking today.

    I keep the editable template for it behind my shadow AI risk calculator, so you do not have to start from a blank file.

    Then do the boring part and send it to everyone, with one person named as the owner. A rule nobody can find works the same as no rule at all, and that is where shadow AI restarts. It is not fancy work and it does not need a consultant.

    One page beats a procurement cycle. If you cannot write it, you do not have an AI program, you have 47% and hope.

    So remember, if you write only one line on that page, write this one. Never put anything into an LLM that you would not email to a stranger.

    Source: https://www.gallup.com/699797/indicator-artificial-intelligence.aspx


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today