Author: daniel

  • Shadow AI grows in the gap Gallup just measured

    Shadow AI grows in the gap Gallup just measured

    Table of contents

    Ask a US employee whether their own employer has rolled out AI. A decent share of them cannot answer, and shadow AI grows in exactly that kind of confusion.

    Gallup changed the question because of it. The methodology note says, “Starting in Q3 2025, Gallup added a ‘don’t know’ option to this question to capture uncertainty about AI adoption.” Results from Q3 2025 onward are no longer directly comparable with earlier measurements.

    A polling company broke its own time series because too many people had no idea what was happening inside the building they work in.

    The second number nobody quotes

    As of May 2026, 47% of US employees say their organization has implemented AI. Only 25% say the organization communicated a clear plan.

    Every headline takes the first number, but the second one is the story.

    Between “we have this thing” and “somebody told me how to use it” sits a gap, and in that gap are contracts, patient records, draft tenders, source code, whatever your people happen to be pasting today.

    Shadow AI is the default state of that gap.

    shadow AI

    Adoption is the wrong argument

    There is a long running fight about how many people use AI. Gabriel Weinberg of DuckDuckGo summed up the skeptical side in June 2026 as “one third actively using AI, one third occasionally using AI, and one third never using AI”. He cites Microsoft telemetry putting it at “more than 30 percent of the US working-age population is using AI, an increase of 3 percentage points from the end of 2025”.

    Gallup’s workplace numbers run higher. 15% of US employees use AI daily. Weekly or more is 30%, and 52% touch it at least a few times a year.

    Pick whichever camp you like, it changes nothing for my job.

    Move the user base up or down, the 25% who got a clear plan stays where it is. That fight pulls attention away from the only question that matters, which is who wrote the rulebook and who read it.

    Shadow AI is a confidentiality problem with no attacker

    Strip the vocabulary and that is all this is.

    No phishing mail, no exploit, no command and control, nothing that trips an alert. An employee opens a browser tab and pastes a client document into a chatbot to get a summary. The data leaves the organization. Most of what you bought assumes somebody is trying to break in. Shadow AI walks out the front door during working hours, moved by people who just want to finish faster.

    OWASP keeps an entry for sensitive information disclosure in its Top 10 for Large Language Model Applications, and almost all of that guidance assumes an application you built. The tab your sales team opened this morning is nobody’s application.

    Scott Brinker named the shape of this back in 2013 and called it Martec’s law. Technology changes exponentially, organizations change logarithmically. Your staff adopted AI in an afternoon, your document set moves at the speed of a committee.

    The Gallup manager numbers show the same thing from the other side. 36% strongly agree their manager supports the team using AI. Where that support exists, employees are 1.7x more likely to use AI weekly or more and 8.7x more likely to report a transformational change in how they work. Encouragement travels by conversation and rules travel by document. Shadow AI takes the faster route.

    What shadow AI looks like on a Tuesday

    Nobody sits down and decides to run shadow AI. It shows up as small, reasonable moves.

    A sales rep pastes a signed contract into a chatbot to pull the renewal dates out of it. An HR assistant drops a salary spreadsheet into a chatbot to reformat the columns. A developer sends a stack trace holding a production connection string to a free tier account. A clinic receptionist rewrites a referral letter with the patient name still in it.

    None of those people are careless. All of them were told AI makes them faster, and none of them were told where the line sits.

    Deleting the client name before pasting does not turn the text into anonymous data either, and I went through the research on that in ChatGPT privacy leak.

    Every one of those actions is invisible to the security team, because nothing was breached and nothing alerted.

    Low numbers are not safe numbers

    Daily use runs at 42% in technology, 27% in finance, 22% in professional services, and 9 to 15% everywhere else.

    Read the low end carefully, because a law firm or a clinic sitting in that bottom band is not in a better position. It is a place where a smaller group does the same thing with far more sensitive material, and with less chance that anyone in IT has ever looked at it. Low usage hides shadow AI.

    The 65% of employees in AI-implementing organizations who report a positive effect on productivity are not lying either. It works, and that is exactly why nobody is going to stop when you ask them to. A ban moves shadow AI further out of sight.

    One page before you buy anything

    Do not start with a tool. Start with one page that answers four things.

    1. Which categories of data never go into an external model.
    2. Which tools are approved, listed by name.
    3. Who an employee asks when the answer is not obvious.
    4. What happens when something has already gone in, and who hears about it first.

    That page will not cover a regulated environment or replace a contract with the vendor. It covers what is leaking today.

    I keep the editable template for it behind my shadow AI risk calculator, so you do not have to start from a blank file.

    Then do the boring part and send it to everyone, with one person named as the owner. A rule nobody can find works the same as no rule at all, and that is where shadow AI restarts. It is not fancy work and it does not need a consultant.

    One page beats a procurement cycle. If you cannot write it, you do not have an AI program, you have 47% and hope.

    So remember, if you write only one line on that page, write this one. Never put anything into an LLM that you would not email to a stranger.

    Source: https://www.gallup.com/699797/indicator-artificial-intelligence.aspx


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • 272 experts built an AI risk ranking. Cybersecurity made the top five.

    272 experts built an AI risk ranking. Cybersecurity made the top five.

    Table of contents

    Two hundred seventy-two international AI experts just built an AI risk ranking based on probability instead of guesswork. Most regulators still haven’t managed that in three years.

    The study, “Prioritization of Risks From Artificial Intelligence,” comes from MIT FutureTech and the University of Queensland. The paper lists 188 co-authors. Core authorship goes to Peter Slattery, Alexander Saeri, Jess Graham, Michael Noetel and Neil Thompson. They used the Delphi method, a research process that gathers expert judgment over multiple rounds until agreement and disagreement both become visible. The same method shows up in the International AI Safety Report 2026, which cites this same body of work. The underlying data lives in the MIT AI Risk Repository, a running catalog of more than 1,600 documented threats used by policymakers and technologists.

    Experts scored 24 risk domains across a five-year horizon, 2025 to 2030, under two scenarios. One scenario assumed business as usual, where organizations and governments keep doing what they’re doing now. The other assumed pragmatic mitigation, where everyone makes cost-effective efforts to reduce harm. Under business as usual, 18 of the 24 domains had at least a 10% probability of catastrophic outcomes. Catastrophic meant more than a million deaths or more than $100 billion in losses, with damage at a comparable civilizational scale in either case. That’s the baseline nobody wanted written down until now. It’s a sharper, numbers-first cut at the risk spectrum this blog already maps.

    Why most people are reading this wrong

    Most people hear “AI risk” and picture something years out, like rogue models or autonomous weapons in some future conflict. That’s not what this AI risk ranking found. Even under pragmatic mitigation, five domains still cleared 10% probability of catastrophe. Dangerous capabilities and AI-enabled weapons or cyberattacks each sat at 12%. So did environmental harm, a domain most people wouldn’t put anywhere near cybersecurity. Inequality and unemployment came in a point lower at 11%, right alongside power centralization. Two of the five are cybersecurity’s problem, and they didn’t drop much even when everyone tries.

    AI risk ranking

    Security work comes down to one job, pushing the attacker’s cost high enough that the attack isn’t worth it anymore. No system stays secure forever, so the alternative just needs to be expensive enough to matter. AI doesn’t invent a new phase of attack. It collapses the cost of the phases that already exist. Reconnaissance gets automated.

    Weaponization gets templated, and delivery gets more convincing because a generated phishing email or a cloned voice doesn’t need a skilled operator anymore. It’s the same mechanism behind the AI-enabled cyberattacks already hitting ordinary companies today. As Slattery putand hacking are where AI capability is moving quickest, and that growth shows up on the cost side of the equation more than the probability side.

    Competitive pressure works as the mechanism that keeps the other four risks running. When a company or a country believes AI confers an advantage, slowing down for safety just hands that advantage to whoever doesn’t slow down. Nobody wants to be the one who raises their own costs while the competition doesn’t, so the race to the bottom on governance keeps going. It’s the same dynamic that keeps patch cycles too slow and security budgets too small, just running at AI speed instead of IT speed.

    Treating this AI risk ranking as a one-time compliance project misreads what the data says. It isn’t something you finish once and file away.

    What this means if you’re the one holding the risk

    The study also names who’s exposed and who’s responsible, and the two lists don’t overlap. Developers and regulators carry most of the responsibility for addressing these risks. Users and the people affected by AI systems carry most of the exposure. That mismatch is why nobody feels urgency at the right level. The people who could slow the collapse aren’t the ones who’d get hurt by it.

    Exposure doesn’t spread evenly. Information absorbs it through misinformation and manipulation, the sort that erodes trust in what people see and read, while national security picks up cyberattacks and weapons development, with surveillance going to whichever hostile actor moves first. Finance isn’t spared either, where fraud and market manipulation get easier and privacy failures ripple into the wider economy on top of that. AI makes doing harm cheaper for anyone who was already capable of it, and possible for people who weren’t.

    Slattery framed the findings as a list of what’s worth paying attention to now, drawn from probability rather than certainty. The response belongs in the governance conversation companies already run for cybersecurity and privacy. It needs a place in business continuity planning too, not a separate checkbox with its own deadline.

    If you run security for an organization deploying AI, the number that should stick with you about AI risk isn’t 272 experts or 24 categories. It’s that even in the best-case scenario the researchers modeled, cyberattacks and dangerous capabilities didn’t fall out of the top five. Mitigation lowers the odds. It doesn’t remove cybersecurity from the list.

    Source: https://mitsloan.mit.edu/ideas-made-to-matter/these-are-most-urgent-ai-risks-according-to-272-experts


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI ATT&CK mapping is wrong four times out of five

    AI ATT&CK mapping is wrong four times out of five

    Table of contents

    Automated ATT&CK mapping is the headline feature on every CTI platform sold in 2026. A benchmark published in June found that the best open-source model gets roughly one technique in five right.

    Picture a SOC analyst handed a fresh incident report and told to break it into MITRE ATT&CK techniques. Few hours of work. Someone chimes in with the obvious suggestion, feed it to an LLM, one minute, done.

    It will be done. Four out of every five techniques will be wrong.

    What the benchmark measured

    Six researchers put a number on it in a paper called “Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports”, arXiv ID 2606.18166, published June 16, 2026. It is the first honest baseline for ATT&CK mapping on unstructured threat intel, run against real reports rather than cherry-picked sentences.

    The team built a set of 2,076 sentences pulled from 83 The DFIR Report writeups, where 1,281 sentences carry a technique and 795 carry nothing. Annotation ran by hand across six phases and landed at 0.68 Cohen’s kappa between annotators, covering 114 unique techniques.

    Seven open-source models went through it at Q4_K_M 4-bit quantization. DeepSeek-V2.5 at 236B parameters, GPT-OSS at 120B and 20B, Llama 3.1 Instruct at 70B and 8B, Gemma 3 at 27B and 12B.

    ATT&CK mapping

    The best micro F1 came in at 0.22, scored by DeepSeek-V2.5. That run used temperature 0.0 with three-shot prompting and chain of thought enabled. Precision 0.21, recall 0.23.

    Worst of the pack, GPT-OSS 20B, sat between 0.00 and 0.06.

    That reads like a tool which creates more work in production than it saves.

    Why the vendor scores looked so good

    Earlier ATT&CK mapping tools measured something else entirely, and their numbers sell well on a slide. TTPXHunter claimed F1 of 0.97, TTPHunter 0.88, TTPDrill 0.82 and AttacKG 0.79.

    The catch is buried in how those tools were evaluated. Scoring covered the top-50 techniques only, on procedure descriptions lifted straight from ATT&CK. Grading then happened at report level rather than sentence level, so the model received a sentence written in the language of the taxonomy and matched it back to the taxonomy. Open-book exam.

    A real report does not read anything like that, and the dataset shows why. One sentence carries 1.58 techniques on average, and 40.2 percent of labeled sentences carry more than one. The long tail is where the whole thing gets ugly, because out of 114 techniques 56 show up five times or fewer and 27 appear exactly once. The most common technique outnumbers the rarest by 229 to 1.

    Where ATT&CK mapping breaks down

    Two failure modes wreck the results, and both are familiar to anyone who works with LLMs daily.

    Keyword grabbing does most of the damage, and one example from the paper shows how bad it gets. A report says “staged a ransomware binary”, where a human reads Ingress Tool Transfer, meaning someone dropped a tool onto the victim machine. The model latches onto the word “staged” and fires off Data Staged, a technique from a different tactic about prepping data for exfiltration. The word matches, the meaning does not.

    Multi-step behavior gets missed for a related reason. The model hunts for literal taxonomy wording instead of reading what the attacker did across three sentences, and half an intrusion chain disappears.

    Prompt tuning did nothing

    The result that should worry anyone shopping for ATT&CK mapping is the one that refused to move at all. Shifting temperature from 0.0 to 0.5 changes micro F1 by 0.01 at most across all seven models. Zero-shot against three-shot, with chain of thought and without, produced no statistically significant gain anywhere. Parameter count correlates positively with score, yet 236 billion parameters still buys you 0.22.

    In plain English, you cannot prompt your way out of bad ATT&CK mapping.

    Retrieval was the only thing that worked

    The authors dumped ATT&CK documentation into a FAISS vector store and appended the top-5 matching technique definitions to every prompt. Llama 70B jumped from 0.22 to 0.32, and recall went from 0.23 to 0.41, a 1.78x improvement.

    Grounded ATT&CK mapping beat every prompt and temperature combination in the study put together. Reasoning was never the bottleneck. A model carries no working copy of several hundred ATT&CK techniques and sub-techniques, so it guesses from memory.

    Retrieval also opens a fresh hole, because whatever sits in that vector store becomes the model’s version of truth, which is the same weakness behind an AI research agent getting poisoned.

    The direction is settled even if 0.32 still falls short of production grade. Give the model the ATT&CK definitions and stop tuning prompts.

    Four questions before you buy ATT&CK mapping

    Make the vendor answer all four in writing.

    1. Ask which dataset the tool was scored on. Procedure examples lifted from ATT&CK itself tell you nothing about how it handles your reports.
    2. Get the technique coverage number. Top-50 coverage hides the long tail where real intrusions live.
    3. Find out whether scoring happened at report level or sentence level. Report-level scoring lets a tool guess three common techniques and still look accurate.
    4. Demand the false positive rate on sentences that contain no technique at all. The benchmark included 795 of those for a reason.

    A vendor who dodges all four has answered you.

    Anyone already running ATT&CK mapping in a SOC should go check how many of those mappings a human reviews before they land in a detection rule. At F1 of 0.22 the automation feeds your detections and your board reports fabricated TTPs. That is worse than no mapping at all, because it looks like knowledge.

    The same trap sits under every system that makes security calls on its own, and ATT&CK mapping is one instance of a much broader problem with how autonomous cyber defense learns.

    Keep a human in the loop. The arithmetic demands it.

    Source: https://arxiv.org/abs/2606.18166v1


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Cyber resilience act compliance is already broken by AI agents

    Cyber resilience act compliance is already broken by AI agents

    Table of contents

    Cyber Resilience Act compliance feels like solid ground once you’ve passed the audit and picked up your CE mark. A new study argues that ground already gave way, because AI agents can turn a fully certified device into an open door without touching a single line of its code.

    The Hookii robotic lawn mower passed every check the EU regulation asks for, and so did the Unitree G1 humanoid robot. Both are certified and legal to sell across the bloc, at least on paper. In March 2026, an AI agent costing about as much as a coffee run took over a fleet of 267 of those mowers without breaking a single law or touching the factory floor. The certificate is still hanging on the wall, but reality already outran it.

    What Cyber Resilience Act compliance actually demands

    The EU Cyber Resilience Act reaches full force in December 2027 with a compliance loop at its center. Manufacturers assess risk and handle whatever flaws surface, then ship patches on adeclared schedule while reporting active exprs. The premise underneath is that softwarealways ships with flaws and only needs managing.

    That design rests on four unstated assumptions about the real world, and each one has to hold for the process to meaanything.

    1. Finding a flaw takes real expertise and real time, so only a small, slow-moving set of vulnerabilities becomes known at any given moment.
    2. A product’s full set of exploitable bugs can be mapped out by the day it ships, and that map stays roughly accuraafterward.
    3. Exploitation is rare enough, and visible enough, that spotting an incident counts as a meaningful signal.
    4. Remediation moves faster than attackers do, so a scheduled patch cycle is enough to stay ahead.

    Vรญctor Mayoral-Vilches of Alias Robotics, the paper’s author, dates the mismatch to a ten-week window. The European Commission drafted the bill in September 2022 and ChatGPT shipped that November, freezing the regulation’s picture othe world at the exact moment that picture s
    ## Why the volume argument misses the point

    Most commentary on AI and cybersecurity fixates on a single number, that agents find more bugs than people do. True,but that’s the smaller half of the story forence Act compliance.
    When a GPT-4 agent exploited 87% of freshly es in April 2024 given the CVE description,that alone is survivable. Companies re-priormost and documenting the risk they accept onthe rest, which is exactly the behavior Article 14 anticipated when it limited mandatory reports to bugs under active attack and left routine scanner findings outside the count. This is no lab curiosity either, since AI hacking tools are already probing production servers around the clock.

    The clock does far more damage than the raw count, though. Median time from disclosure to weaponization stood at 771 days back in 2018; by 2023 it was 5.3 days, an exponential decline fitting at Rยฒ=0.98 that sat near zero in 2025, while the share of bugs weaponized before or at public disclosure climbed from 19% up to 54%.

    A certificate that lies without the product changing

    Cyber resilience act compliance

    Cyber Resilience Act compliance turns fragilrequirement that products ship “without known exploitable vulnerabilities.” That clause assumes “known” on day one roughly equals “findable” a week later.


    Google’s Big Sleep agent rediscovered a hidden SQLite flaw on demand in November 2024, and running that kind of attack today costs roughly a hundred dollars a try, yet the product and its certificate stayed exactly the same on paper. Its real exploitable footprint moved anyway, because the conditions around

    That hits every manufacturer shipping a networked device into the EU, whether it’s an IP camera, an industrial
    controller, a warehouse robot, or a smart thzen and fully compliant can become an opendoor a week after certification, because the flaw grew out of the attacker’s toolkit, sitting entirely outside the product’s own code.

    Two robots, one proof

    The paper backs its claim with two devices tested directly under Cyber Resilience Act scope, the Unitree G1 humanoid
    and the Hookii mower.

    Undefended, an AI agent rooted the G1 through a Bluetooth command-injection bug, exploiting an identical AES key baked into every unit in the fleet. From there it decrypted the robot’s telemetry and reached teleoperation, with a success rate of 79%. On the mower, 38 chained vulnerabilities bypassed the safety geofence across a fleet of 267 devices at 75% success.

    Enroll both robots in the Robot Immune System, an autonomous defensive AI agent, and attacker success collapses to 14% and 8%, with detection and containment landing under 8 and 12 seconds. The paper’s conclusion cuts both ways, because attacker and defender run on the same technology, so the only certification that holds up is one that never stops running. I covered the mechanics behind that idea in how autonomous cyber defense learns.

    Ways to pressure-test your compliance program

    If you advise clients on Cyber Resilience Act compliance or ship connected products into the EU, these moves matter more than the paperwork.

    1. Schedule a re-test after certification lands, since a point-in-time audit only tells you about the day it ran.
    2. Budget for continuous monitoring as a recurring cost, because the disclosure-to-exploit window is now measured in days, sometimes hours.
    3. Treat a defensive AI agent as core infras Act compliance; the paper’s own data showsattacker success collapsing from 79% to 14% once one is running.
    4. Ask any vendor how they detect drift between what was certified and what’s actually running today.

    December 2027 is when the Cyber Resilience Act switches on in full, certifying products against a world that has already moved on. A once-a-year compliance stamp buys nothing past the following Tuesday, and continuous, agent-operated defense is the only way left for longer than a week.

    Source: https://arxiv.org/abs/2607.07109


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Public sector HTTPS is nearly universal in Poland. The layers above it are not

    Public sector HTTPS is nearly universal in Poland. The layers above it are not

    Table of contents

    Update, 10 July 2026. The first version of this article was based on 50 domains and said two in three Polish municipalities had no HTTPS. That sample was small and skewed. I rescanned 2442 of Poland’s 2479 municipalities and rewrote the piece on the full data. The real numbers are below.

    You fill out a form at your city hall. You type your ID number, your address, your date of birth. You hit submit. The good news: on almost every Polish municipal site, that connection is now encrypted. The problem is what sits on top of it.

    I scanned 2442 of Poland’s 2479 municipalities, every one with a working web address in the official MSWiA register. Passively, without touching a single system. Just SSL, HTTP headers, and DNS, the stuff anyone can see from outside.

    What This Actually Means

    Start with the good news, because there is some. Public sector HTTPS in Poland is nearly everywhere. Only about 1 percent of municipalities run with no HTTPS at all. The padlock most residents look for is usually there.

    Encryption, though, is the foundation, not the house. And almost nothing stands on top of it.

    67 percent of municipalities send no HSTS header. 77 percent have no Content-Security-Policy. 79 percent set no Permissions-Policy. 58 percent have no protection against clickjacking. One in three has at least one issue I rate as high risk.

    There is a softer edge to the HTTPS number too. 9 percent of municipalities have a valid certificate but still serve content over plain HTTP if you reach them that way, because nobody set up the redirect. The encryption exists, but the connection can quietly drop off it, and a login sent over that accidental HTTP is exposed. It gets worse when the same password is reused elsewhere, the exact habit behind most password security failures out there.

    Why Everyone Looks in the Wrong Direction

    Most public sector cybersecurity talk revolves around ransomware, hoodie hackers, and attacks from abroad. It is a convenient story. It shifts the blame onto “Russian hackers,” a target nobody can actually do anything about. The boring stuff, the kind covered in a proper cybersecurity strategy, never makes the headlines.

    Public sector HTTPS

    The real problem is boring. A server config missing a few lines. The redirect to HTTPS that never got set up, the security header nobody added. This is plain neglect, and a passive scan catches it in seconds, without ever touching the inside of the system.

    None of these headers is exotic. You set them once, in a config file, and they work for years. A missing Content-Security-Policy or X-Frame-Options is what lets someone inject foreign code into a council page, or hide the real site inside an invisible frame they control. For a site where a resident types in personal data, that is not cosmetic.

    What This Means for You, If You Run One of These Sites

    Where a municipal site still serves over plain HTTP, that 9 percent, it is a direct GDPR problem. You process citizens’ data without basic transport protection, and an auditor does not need to be a hacker to see it. They only need to open the site over HTTP and watch the encryption fall away.

    Missing security headers sit in the next tier. They are the “appropriate technical measures” the regulation expects, the ones you have to be able to justify skipping. “We never got around to it” is not the answer you want on record after an incident.

    The Technical Fixes, Explained Without Jargon

    HTTPS is connection encryption. Poland’s municipalities have it. That is the win worth keeping.

    HSTS tells the browser to talk to a site encrypted, always, and never any other way. 67 percent of the municipalities in the scan do not send it, so an attacker can still try to force that first connection down to plain HTTP.

    CSP and X-Frame-Options, both covered in the OWASP secure headers guide, make it harder to inject foreign code into a page, or to hide a legitimate site inside an invisible frame controlled by someone else.

    An SSL certificate is free through Let’s Encrypt, and most municipalities already run one. The headers take a few lines in a server config and half an hour of your time, done over a coffee break.

    Check Your Own Site Today

    Open your own site in a new tab. The padlock is probably there, and that is good. Now open the developer console and look at the response headers. That is where the gap hides: no HSTS, no CSP, no frame protection. If they are missing, you found out before your next audit did.

    Download the full scorecard as a PDF, including the breakdown by type of municipality.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI agent privacy is the gap nobody is testing for

    Table of contents

    AI agent privacy rarely makes it onto a security checklist, and that gap is where the real damage starts. Picture an agent handling a routine task, checking a customer’s order status, then pulling matching records from across the CRM and the invoicing database before replying, all inside ten seconds. Nobody stopped to ask whether it also picked up another customer’s card number sitting in the same conversation thread, still parked in its working memory.

    The problem, stripped down

    An LLM agent today operates across databases, document collections pulled through RAG, external APIs, and other agents further down the task chain. Each of those surfaces opens its own leak path for agent data, and each one is a blind spot in most AI agent privacy reviews. The survey traces how sensitive data actually leaves a system. Some of it exits through the queries an agent writes for itself. The rest slips out through intermediate results parked in memory or through messages passed to another agent mid

    Most security policies were built to catch one of those paths, leaving the other two wide open. That mismatch is the actual shape of the AI agent privacy problem door while two side entrances stay open.

    Why most teams get this wrong

    Most agent security reviews start from attack scenarios like prompt injection or a jailbreak attempt slipping past the model. The survey approaches the problem of what data the agent touches in the first place, regardless of whether anyone is attacking it. A team that red-teams its agent against known attacks can still miss the risk sitting inside the data access design itself.

    ai agent privacy

    Database-level access control looks like it should cover this, though it only answers a narrower question – who can read a record right now. It says nothing about what the agent does with that record three sessions later. The survey reviews six governance mechanisms built to me, information flow control, and catch-leakage pieced together across multiple sessions. The rest only catch a single request. AI agent privacy actually breaks down at the pattern level, stitched together across sessions, which is precisely what those other five mechanisms miss.


    What this means for you

    Deploying agents for clients, or running them inside your own company, changes what belongs on your vendor checklist. Jailbreak red-teaming credentials cover only part of that checklist now. The better question covers AI agent privacy across every surface the agent touches at once, RAG retrieval, SQL queries, memory, and messages traded with other agents.
    The survey’s authors say a combined benchmark like that barely exists yet, and under GDPR and similar rules, that absence becomes a real liability, since proving due diligence gets difficult when the test you ran skips most of the data’s actual path through the system.

    The technical bit, plainly

    Take a concrete case that shows what AI agent privacy risk looks like in practice. An HR agent answers an employee’s question about vacation days. While retrieving the record, it also pulls a field noting the medical reason behind an
    earlier absence, sitting right next to the vate result lands in the agent’s memory. Three queries later, a separate thread with the same employee draws on that memory and surfaces a detail nobody asked to reveal.

    The failure sits in a missing boundary between what the agent knows and what it’s allowed to say in a given context, a boundary no attacker had to touch. Information flow control tries to draw that boundary at the data layer rather than
    the prompt layer. It tags sensitivity; the monitor tracks that tag to wherever the data ends up, a chat reply, or a message sent to a second agent. That gap, more than any prompt-based attack, is the everyday face of AI agent privacy failure.

    Four questions before you ship an agent

    Before an agent touches production data, four questions cut through most of the risk described above.

    First, map every data surface the agent can reach, including ones far outside its original purpose, covering every database, document store, API, and memory layer in scope. Second, track data across sessions instead of single requests, checking whether a fact revealed in session one can resurface in session five without anyone approving it.

    Third, separate retrieval from disclosure, since an agent repeating that record out loud needs different permissions. Fourth, ask vendors for benchmark coverage rather than a demo, because a system that resists jailbreaks hasn’t shown you anything about its AI agent privacy coverage across RAG and SQL, let alone memory.


    AI agent privacy is only going to get harder to manage as agent systems keep adding data sources and stacking more agents that relay information to each other. Until a benchmark covers that full picture, every company running agents today is deciding, on its own, how much privacy is better to decide those AI agent privacy tradeoffs on purpose, before an incident decides them for you.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • How Autonomous Cyber Defense Learns an Attacker It Never Sees

    Table of contents

    Autonomous cyber defense now has to do something close to guarding a building in the dark. The guard cannot see the intruder and hears no footsteps, yet a window sits cracked open on the second floor while a motion sensor blinks somewhere down the hall. Out of those few signals the system still has to work out who got in and where they are going.

    You might call that impossible, though it describes an ordinary night for anyone defending a network. A recent paper on neurosymbolic cyber agents took that exact puzzle and tried to solve it.

    You’re Fighting a Shadow

    In a real network the blue agent doing the defending has no view into the attacker’s console. It cannot tell which technique was used or how far along the kill chain the intrusion has already travelled.

    Researchers call this a partially observable environment, which is a polite way of saying the defender works from scraps. A bit of odd traffic here, or a logig else stays locked inside a black box.

    Most defensive tools only wake up once the damage shows. The alert fires after the break-in, so the whole posture
    amounts to firefighting rather than preventiert that, training autonomous cyber defenseto anticipate the next move instead of mopping up the last one.

    Why Most Approaches Break Down

    The oldest method leans on hard rules, where through say a signature or a fixedthreshold. The weakness shows the moment an attacker stops following your script. He shifts tactics and waits you out until yesterday’s clever rule has gone blind.

    Pure neural networks promise the opposite of brittle rules, since you feed them data and let the model sort out the patterns on its own. That power comes wrapped in a problem, because the model becomes a black box that cannot explain
    its own reasoning, and an unexplainable verdity work.

    Autonomous Cyber Defense

    The hybrid idea splits the difference by paiman can actually read and audit with machinelearning that picks up signals the eye would miss. That pairing is what neurosymbolic autonomous cyber defense is built on.

    How It Actually Works

    At the core of this autonomous cyber defenseworks a lot like a firefighter’s decisionflow that moves from checking for smoke to judging the threat before it acts. The structure stays readable and modular, so a human can follow the logic and trust where it leads.

    Tucked into chosen nodes of that tree are learning-enabled components. Those are the eyes of the system, the parts
    that stare at fragments of network data and oing in the gaps.

    The learning itself runs on plain imitation instead of any explicit rulebook. Rather than spelling out rules, the team shows the model a large pile of red-agent behavior and lets it reproduce that policy, much as an apprentice absorbs a craft by watching a master at the bench. From its own observations and its own responses, the defender rebuilds the attacker’s strategy without ever reading a single command he typed. The authors report that the system copes with
    several different red-agent policies and rea across a spread of simulated scenarios.

    What This Means for You

    The headline shift moves defense from reactive to predictive, which is the gap between stopping a burglar at the door and knowing he is on his way before he reaches the twist. Once autonomous cyber defense can learn an attacker’s policy, the attacker realises he is being studied, and the contest climbs to a new level where he feeds the sensors poisoned observations so the model absorbs a fake pattern on purpose.

    Trust is the quieter prize, because a neurosymbolic hybrid leaves a decision trail that a pure neural net never could, letting you check why the agent concluded an attack was underway. As autonomous SOCs move from speculation towstandard kit, that kind of auditability will

    None of this escapes its limits. The work still lives in simulation with discrete states and actions, while a production network runs messy and continuous. The trip from test range to deployment usually takes longer than the headline numbers imply.

    The Takeaway

    Tomorrow’s autonomous cyber defense aims at something past raw speed, since it will try to guess your next move before you commit to it. The harder question becomes who teaches a machine to lie convincingly to the enemy’s sensors, and who manages it first.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • How an AI Research Agent Gets Poisoned by a Single Reddit Post

    Table of contents

    When an AI research agent processes a research prompt, it fires off dozens of queries across public sources and compiles a report from whatever it retrieves. A Cornell paper from May 2026 shows that process has a structural flaw, exploitable with a single planted sentence on a Reddit thread.

    What Separates an AI Research Agent from a Chatbot

    An AI research agent treats every research prompt as an investigation. It retrieves live content across public sources and builds a report from whatever comes back. Systems like STORM and OmniThink are built for this workflow.

    The structural problem is embedded in how these research sessions actually work. When an AI research agent runs 20 related queries on the same topic, those queries keep returning to the same sources. A Reddit thread that surfaces in 15 out of 20 retrievals gets pulled 15 times, and the agent treats each retrieval as an independent data point with no mechanism to flag the repetition.

    The Attack That Skips the Model

    Most AI security work focuses on jailbreaks and prompt injection, attacks directed at the model itself. Content poisoning is a category of attack that operates below the model, targeting what an AI research agent reads before inference begins.

    An attacker adds a short crafted sentence to one frequently-retrieved page on Reddit or Wikipedia. Each time the agent pulls that page, the sentence appears as independent evidence. After 15 retrievals, the planted claim reads like fact. The attack requires only that the crafted text appear on a page the agent retrieves repeatedly.

    What Cornell’s Tests Found

    Cornell researchers tested the attack on STORM and OmniThink, two systems built for automated knowledge synthesis. A single poisoned post on a user-generated content page was enough to make both systems cite attacker-chosen content and promote attacker-chosen names across many unrelated queries.

    The system treats repetition as a substitute for truth. Verification is the reader’s responsibility.

    The Real-World Consequences

    If your workflow relies on any tool that retrieves live web content and produces a research summary, you are operating inside this architecture. Competitive intelligence reports and vendor analyses reflect whatever was sitting in the pages retrieved.

    The problem compounds when an AI research agent generates content that feeds into other AI pipelines, a workflow already running in automated production environments. A single poisoned source spreads downstream with nobody checking for bad data.

    ai research agent

    At the strategic level, companies running deep-research workflows could be misled by a single forum post. The bad information arrives as a citation and looks like every other source in the report.

    How the Poisoning Works, Step by Step

    An AI research agent runs 20 queries on “best cybersecurity tools for SMBs.” Fifteen retrieve the same Reddit thread. Buried in that thread is a planted line that reads like expert recommendation, something along the lines of “Security professionals also recommend Company X, widely praised in recent third-party evaluations.”

    Each time the agent encounters that Reddit thread, it reads the same planted sentence and registers repetition as consensus. Company X ends up cited throughout the final report, even if someone was paid to plant that line a year ago, or even if Company X is your direct competitor.

    Proposed defenses include source-level filtering of UGC domains and output anomaly detection. Both reduce the attack surface without solving the underlying issue, which is that the agent was built to count occurrences and cannot check whether any of them are accurate.

    What to Check Before Trusting Any AI Report

    Treat every output from an AI research agent as a starting point that requires verification before it shapes any decision, and start with the footnotes. When a report recommends a specific vendor or product, trace the recommendation to its original source and check whether it came from a named expert or an anonymous account on a public forum.

    Two years out, the competitive dynamics are likely to follow an established pattern. Brands will run AI poisoning campaigns alongside their SEO operations, and the target audience has shifted from search engine crawlers to AI research agents.

    The information war has a new attack surface. For now, that surface is your research pipeline.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI Enabled Cyberattacks Don’t Need a Hacker Anymore

    Table of contents

    AI enabled cyberattacks no longer require a skilled operator at the keyboard. Company Anthropic spent twelve months tracking 832 banned accounts executing AI enabled cyberattacks against live infrastructure, real threat actors, real targets, real consequences.

    The results should make every security professional uncomfortable.

    A 1.7x Jump in High-Risk Actors in One Year

    Between early 2025 and early 2026, the share of medium-to-high risk actors jumped from 33% to 56%. That is a 1.7x increase in twelve months.

    Here is what that does not mean.

    It does not mean hackers got better at coding. Technical sophistication scores did not change dramatically. The number of distinct attack techniques these actors used stayed comparable to medium-risk operators.

    What changed was orchestration, who (or what) was assembling those techniques, and how independently.

    The Metric That Actually Predicts Danger

    Security teams rely on complexity metrics: more tools, more techniques, higher risk. The Anthropic data breaks that assumption.

    Technical breadth was a weak predictor of danger. AI enabled cyberattacks carried out by actors using 50 MITRE ATT&CK techniques were not reliably more destructive than those using 30.

    ai enabled cyberattacks

    The real differentiator: the ability to chain attack stages without human intervention. Recognize a target. Select a vector. Adapt when infrastructure is unfamiliar. Archive data. All without a human approving each step.

    “AI as assistant” and “AI as operator” are two different threat categories.

    GTG-1002: The Case That Changes the Threat Model

    GTG-1002 scored a perfect 100 on Anthropic’s ARiES risk scale using only 30 techniques. Many medium-risk actors use the same range.

    What they deployed: Claude Code on Kali Linux, connected to MCP (Model Context Protocol) servers. The AI did not suggest commands, it executed them. Autonomous reconnaissance. Autonomous lateral movement. Autonomous data staging. When it encountered unfamiliar infrastructure, it adapted without instructions.

    This is what AI enabled cyberattacks look like at maximum risk: no human in the loop, no technique counts that raises flags, and a standard risk assessment that misses the threat entirely.

    Why Traditional Detection Falls Short

    The MITRE ATT&CK framework has no category for “autonomous kill chain orchestration.” Anthropic is collaborating with MITRE to address that. The gap exists today.

    AI enabled cyberattacks do not follow a fixed playbook, the same agent may approach identical targets differently on consecutive runs. Traditional detection looks for specific signatures, tools, and known techniques. That approach misses autonomous behavior by design.

    Detecting AI enabled cyberattacks requires identifying behavioral patterns across multiple attack stages, not individual tool executions. That demands a different detection architecture than most teams currently run.

    How Anthropic Scores AI Risk: ARiES

    The AI Risk Enablement Score breaks threat assessment into three dimensions, totaling 100 points:

    Threat (0โ€“35)

    Measures intent clarity, technical skill, and use of evasion tactics.

    Vulnerability (0โ€“35)

    Measures how much a model enables harm. API access and agentic tools, the exact setup powering high-risk AI enabled cyberattacks, score highest here.

    Impact (0โ€“30)

    Measures real-world consequences if the operation succeeds.

    The system uses addition, not multiplication. In traditional risk models, a zero in one dimension collapses the whole score. ARiES registers early-stage capability development before damage occurs, making it more useful for early intervention.

    Where AI Is Actually Being Used Right Now

    Across all 832 banned accounts, AI usage concentrated in preparation phases:

    • 69% used AI to develop capabilities, primarily malware
    • 64.7% for obfuscation and evasion
    • 55.9% for local data collection
    • 54.9% to disable security tools

    Defense evasion accounted for 84.4% of all mapped activity. Live network operations remain smaller: lateral movement at 6.5%, remote services under 1.5%.

    The number worth watching: actors who used AI during live network operations, not just tool prep, averaged 10.5 points higher on the ARiES scale. When AI moves from preparation into execution, risk jumps sharply.

    That number is increasing.

    What Defenders Should Do Now

    Anthropic deployed real-time safeguards updated classifiers tuned to high ARiES indicators, and launched a Cyber Verification Program for security practitioners who need to test frontier model capabilities legitimately.

    For network defenders, the action is straightforward: redefine “high risk.”

    Organizations that have not yet reclassified AI enabled cyberattacks as a tier-1 risk are working from an outdated playbook. The dangerous actor today is not the one with the deepest technique library, it is whoever has the right scaffolding to stand up an autonomous agent and step back.

    AI enabled cyberattacks at scale no longer need an expert. They need an orchestrator.

    Technical skill as an entry barrier is dropping. Orchestration skill is the new dividing line.

    Update your threat model before the next GTG-1002 updates theirs.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    ๐Ÿ“ก THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    ๐Ÿ’ก ONE ADVICE – One actionable AI/cybersecurity tip you can use today