Tag: cybersecurity

  • I read the Stanford study on entry-level jobs and the headline doesn’t hold up

    I read the Stanford study on entry-level jobs and the headline doesn’t hold up

    Table of contents

    Last year it was 13 percent. Now it’s 19. Half the internet did that subtraction and filed the six points under proof that AI is wiping out entry-level jobs.

    I opened the paper. That pair of numbers isn’t in it.

    What Stanford measured

    The paper is “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence”, August 2026 edition. Erik Brynjolfsson wrote it with Bharat Chandar and Ruyu Chen. The data is anonymized ADP payroll, between 3.5 and 5 million workers a month, through June 2026. Exposure comes from a task-level impact measure and from the Anthropic Economic Index, which tracks how occupations use Claude at work. It is the same Stanford operation that puts out the AI Index I went through earlier this year.

    Here is the finding everyone quoted about entry-level jobs. Employment for 22 to 25 year olds in the most AI-exposed occupations now sits 19 percent below where it would be if it had kept pace with their less-exposed peers.

    entry-level jobs

    The raw numbers behind that are worth having. In the two most exposed quintiles, employment for that age group fell roughly 11 percent between November 2022 and June 2026, while in the three least exposed quintiles it grew roughly 10 percent. That is a spread of 21 percentage points.

    Experienced workers show no such gap. Whatever is happening to entry-level jobs is leaving the people inside alone.

    Where the 13 percent came from

    The authors are open about it, and earlier versions of the paper headlined a regression estimate that adjusted for firm-level shocks. That was the 13, on July 2025 data. On September 2025 data the same measure read 16 percent.

    For this edition they dropped the regression from the headline and switched to a simple descriptive measure that needs no modelling choices. And they were straight about what that measure showed a year ago, which was fifteen percent.

    So the like-for-like pair is 15 against 19, which is four points of movement where the coverage reported six.

    Both figures track entry-level jobs, with a different instrument behind each.

    This is not nitpicking a decimal, because the gap between “it grew half again as fast as we thought” and “it grew a bit” decides whether you are writing about a trend accelerating or a trend continuing. Ars Technica put both figures in neighbouring sentences and closed the paragraph by noting that last year the gap was “just 13 percent”. Every word of that holds up on its own. What the reader walks away with is a subtraction between two different measurement methods.

    What the paper admits about itself

    I went to the limitations first, because that is usually where the thing missing from the headline sits.

    The authors say this is not a causal estimate. Their phrase is early descriptive indicators, which sits a long way from a verdict on entry-level jobs.

    The effect weakens once you control for education, which is awkward for the headline. Some of the divergence shows up before generative AI existed. The pattern is sharper in the ADP sample than in national survey benchmarks. That sample over-represents large firms and occupations with high AI exposure. Manufacturing and wholesale also carry more weight in it than in the wider economy.

    The panel is balanced, which means it is conditioned on firms surviving. Companies that went under drop out of the picture. On top of that, roughly 30 percent of records have no job title, so the occupation code gets filled in by an algorithm.

    The finding on entry-level jobs survives all of that. What changes is how hard you are allowed to lean on it, which is the same problem I hit when 272 experts ranked AI risk and the ranking got quoted without its error bars.

    What this means if you run security

    The mechanism behind the entry-level jobs number is simple enough to check against your own team. AI substitutes for codified knowledge, the kind written down in a textbook or a procedure. It complements tacit knowledge, the kind you only get from years of practice and mentorship. The decline shows up in hiring, while separation rates did not move.

    Map that onto a SOC, where tier one sits on alert triage against a runbook. That is written procedure with checkable output, and in security it is where the entry-level jobs live.

    Seniors don’t fall from the sky, and your L3 analyst is somebody’s L1 from five years ago. Cut the entry-level jobs today and in five years you have nobody to promote. No model covers for you there, because tacit knowledge does not sit in your documentation.

    Speaking to the Washington Post, quoted by Ars, Brynjolfsson said the entry-level effects he is measuring “are real, persistent and widening”, and that he is more worried than he was “about a labor market that keeps its overall employment level while quietly closing the on-ramp for people starting their careers”.

    The on-ramp closing is in the data. The worry is his.

    The rule

    Before a number about entry-level jobs goes on a slide for your board, check whether last year’s number from that same paper was measured the same way. Authors do change their headline metric mid-project, and they describe the change in the paragraph the journalist never reaches.

    Never trust the headline. Trust the footnote.

    Source | https://arstechnica.com/ai/2026/08/ai-is-hitting-entry-level-jobs-hardest-stanford-study-finds/


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Open-weight models don’t say no

    Open-weight models don’t say no

    Table of contents

    On 21 July Guillermo Rauch, Vercel’s CEO, posted his internal eval results on X, and the argument about open-weight models in security got a lot louder. Kimi K3, the open-weight Chinese release from five days earlier, came out “top-tier at cybersecurity” in his tests. In the same post he added that “Fable refuses everything” and that he couldn’t get it to finish the run at all.

    A day later Semgrep published a table with precision and recall.

    Two days after that, the UK AISI and the US CAISI published theirs.

    None of the three measurements backs up that sentence. Which is a lot more interesting than the ranking itself.

    What was claimed, and what was measured

    AISI and CAISI ran Kimi K3 through ExploitBench and the TLO cyber range, and published the numbers. It scored 32 percent on ExploitBench against 24 percent for GLM-5.2, the strongest of the open-weight models as of June 2026. That reads fine until you check the next column, where leading US models reached arbitrary code execution on 20 of 41 samples on average. Kimi K3 managed zero of 41.

    On the cyber range, where the full attack path runs 32 steps, Kimi averaged step 17. US frontier models averaged 28.5. It cleared the whole range once in ten attempts.

    Semgrep tested a different thing, hunting IDOR bugs in real repositories. There Kimi K3 landed 0.684 precision, while Claude Opus 4.8, GPT-5.6 Sol, GPT-5.6 Terra and GLM-5.2 all sit between 0.86 and 0.91. On the largest repo in the set, Kimi delivered roughly 6 percent F1 against roughly 20 percent for everyone else. If that pattern feels familiar, it should, because the same gap between a marketed score and a measured one showed up when automated ATT&CK mapping got benchmarked properly.

    To be fair, this is not apples to apples. AISI tested the US models with system-level safeguards switched off. The range itself, quoting the report, “lacks active defenders and defensive tooling”.

    Why the open-weight models debate is pointed the wrong way

    Most of the argument is about the ceiling. Whether open-weight models have caught up with the frontier or not.

    If you defend things for a living, your problem is the floor.

    Security comes down to one idea. You push the cost of attacking up until it stops being worth it. A model ranking tells you how high the best attacker on earth can reach. The floor tells you what the cheapest attacker costs, and the cheapest one is plenty to ruin your week.

    Open-weight models

    One line in that same AISI and CAISI report made no headlines at all. “Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations”.

    Now put that next to the second half of Rauch’s post. The closed model said no. The open one didn’t, and that has nothing to do with how good either of them is at the task.

    Willingness cannot be bolted back onto weights somebody already pulled down to their own disk.

    What this changes for you

    The refusal layer we treat as a safety control is a property of the vendor’s API. It leaves with the rate limit, the provider-side logs, the abuse team and the option to kill an account. Open-weight models running on somebody else’s GPU have none of those four.

    There’s a second shift underneath that one. The classic limit on offensive operations is organisational. An operator spends their time on one target and cannot hit twenty with the same quality. Cheap automation lifts that limit.

    The DoD capability scale sorts attackers into six tiers, where tier one runs other people’s tools and tier three finds and exploits bugs on its own. Cheap scaffolding drags part of tier one toward tier three, because the pipeline does the work now instead of a person.

    Risk is impact times likelihood, and the hype was all about impact. The quiet move happened in the other term.

    Where open-weight models lose, in plain terms

    It pays to be exact about what a precision score of 0.684 costs you. Bad precision means the model still finds things, then buries them in a pile of wrong findings that somebody has to read through. For your team that’s a bill. For an attacker who skims the output once and needs a single hit, it costs nothing.

    And that gap can be closed without touching the model at all.

    In June, clearbluejar reproduced a well-publicised find, a seventeen-year-old RCE in FreeBSD originally surfaced by a frontier model. That class of autonomy was already on display in controlled testing months earlier. Except he ran it on gpt-oss-20b on his own hardware, through AISLE’s public 1,700-line Python pipeline. The apparent miss went away on a re-run. The real problem was noise, with the genuine bug buried under false positives.

    So he added one stage that checks whether the code is reachable. False positives dropped from 30 to 5 and the CVE was still standing. He said it outright. “The scaffolding does the work, and it’s a lever you can pull on your own model”.

    Stanislav Fort at AISLE reproduced the same find for under $100.

    The obvious objection, and why it doesn’t hold

    Someone will say these are synthetic benchmarks, built on cyber ranges with no defenders and curated repos that flatter the scores. Fair enough, and the AISI report says as much itself.

    But the objection cuts the wrong way. If the range flatters open-weight models and closed ones equally, the comparison between them still stands. And the finding that matters here is behavioural rather than numerical, because the model tried, and no range design makes that go away.

    Four things to check in your threat model

    Skip the leaderboard. If open-weight models are in your threat model at all, work through these four instead.

    1. Find every control you rely on that lives at the provider. List them out loud. Refusal, rate limit, logging, account termination. Everything on that list is gone the moment the weights run locally.
    2. Re-check your detection thresholds for cheap reconnaissance. Volume goes up, quality of each attempt goes down. Alert thresholds tuned for a patient human operator will read that as noise.
    3. Measure the window between a public CVE and your patch. That window used to be protected by how few people could weaponise a bug. Open-weight models plus a public pipeline shortened the queue, and nobody sent you a notice when it happened.
    4. Price your own triage burden. Poor precision costs a defender real hours and costs an attacker nothing. If you deploy the same class of tooling internally, budget the triage hours the same way you budget the scan.

    None of that needs a new product, and none of it is about open-weight models as a technology. It needs the assumption “they probably can’t” taken out of your threat model.

    One rule to take away

    Stop asking whether your attacker’s model is as good as yours. With open-weight models the honest question is whether anything will stop it.

    Labs and their evals police the capability ceiling. The availability floor is set by weights sitting on someone’s disk and 1,700 lines of Python from GitHub. Open-weight models moved that floor while the argument stayed fixed on the ceiling.

    Build your threat model on the second number.

    Source| https://www.mbi-deepdives.com/open-weights/


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • 272 experts built an AI risk ranking. Cybersecurity made the top five.

    272 experts built an AI risk ranking. Cybersecurity made the top five.

    Table of contents

    Two hundred seventy-two international AI experts just built an AI risk ranking based on probability instead of guesswork. Most regulators still haven’t managed that in three years.

    The study, “Prioritization of Risks From Artificial Intelligence,” comes from MIT FutureTech and the University of Queensland. The paper lists 188 co-authors. Core authorship goes to Peter Slattery, Alexander Saeri, Jess Graham, Michael Noetel and Neil Thompson. They used the Delphi method, a research process that gathers expert judgment over multiple rounds until agreement and disagreement both become visible. The same method shows up in the International AI Safety Report 2026, which cites this same body of work. The underlying data lives in the MIT AI Risk Repository, a running catalog of more than 1,600 documented threats used by policymakers and technologists.

    Experts scored 24 risk domains across a five-year horizon, 2025 to 2030, under two scenarios. One scenario assumed business as usual, where organizations and governments keep doing what they’re doing now. The other assumed pragmatic mitigation, where everyone makes cost-effective efforts to reduce harm. Under business as usual, 18 of the 24 domains had at least a 10% probability of catastrophic outcomes. Catastrophic meant more than a million deaths or more than $100 billion in losses, with damage at a comparable civilizational scale in either case. That’s the baseline nobody wanted written down until now. It’s a sharper, numbers-first cut at the risk spectrum this blog already maps.

    Why most people are reading this wrong

    Most people hear “AI risk” and picture something years out, like rogue models or autonomous weapons in some future conflict. That’s not what this AI risk ranking found. Even under pragmatic mitigation, five domains still cleared 10% probability of catastrophe. Dangerous capabilities and AI-enabled weapons or cyberattacks each sat at 12%. So did environmental harm, a domain most people wouldn’t put anywhere near cybersecurity. Inequality and unemployment came in a point lower at 11%, right alongside power centralization. Two of the five are cybersecurity’s problem, and they didn’t drop much even when everyone tries.

    AI risk ranking

    Security work comes down to one job, pushing the attacker’s cost high enough that the attack isn’t worth it anymore. No system stays secure forever, so the alternative just needs to be expensive enough to matter. AI doesn’t invent a new phase of attack. It collapses the cost of the phases that already exist. Reconnaissance gets automated.

    Weaponization gets templated, and delivery gets more convincing because a generated phishing email or a cloned voice doesn’t need a skilled operator anymore. It’s the same mechanism behind the AI-enabled cyberattacks already hitting ordinary companies today. As Slattery putand hacking are where AI capability is moving quickest, and that growth shows up on the cost side of the equation more than the probability side.

    Competitive pressure works as the mechanism that keeps the other four risks running. When a company or a country believes AI confers an advantage, slowing down for safety just hands that advantage to whoever doesn’t slow down. Nobody wants to be the one who raises their own costs while the competition doesn’t, so the race to the bottom on governance keeps going. It’s the same dynamic that keeps patch cycles too slow and security budgets too small, just running at AI speed instead of IT speed.

    Treating this AI risk ranking as a one-time compliance project misreads what the data says. It isn’t something you finish once and file away.

    What this means if you’re the one holding the risk

    The study also names who’s exposed and who’s responsible, and the two lists don’t overlap. Developers and regulators carry most of the responsibility for addressing these risks. Users and the people affected by AI systems carry most of the exposure. That mismatch is why nobody feels urgency at the right level. The people who could slow the collapse aren’t the ones who’d get hurt by it.

    Exposure doesn’t spread evenly. Information absorbs it through misinformation and manipulation, the sort that erodes trust in what people see and read, while national security picks up cyberattacks and weapons development, with surveillance going to whichever hostile actor moves first. Finance isn’t spared either, where fraud and market manipulation get easier and privacy failures ripple into the wider economy on top of that. AI makes doing harm cheaper for anyone who was already capable of it, and possible for people who weren’t.

    Slattery framed the findings as a list of what’s worth paying attention to now, drawn from probability rather than certainty. The response belongs in the governance conversation companies already run for cybersecurity and privacy. It needs a place in business continuity planning too, not a separate checkbox with its own deadline.

    If you run security for an organization deploying AI, the number that should stick with you about AI risk isn’t 272 experts or 24 categories. It’s that even in the best-case scenario the researchers modeled, cyberattacks and dangerous capabilities didn’t fall out of the top five. Mitigation lowers the odds. It doesn’t remove cybersecurity from the list.

    Source: https://mitsloan.mit.edu/ideas-made-to-matter/these-are-most-urgent-ai-risks-according-to-272-experts


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI ATT&CK mapping is wrong four times out of five

    AI ATT&CK mapping is wrong four times out of five

    Table of contents

    Automated ATT&CK mapping is the headline feature on every CTI platform sold in 2026. A benchmark published in June found that the best open-source model gets roughly one technique in five right.

    Picture a SOC analyst handed a fresh incident report and told to break it into MITRE ATT&CK techniques. Few hours of work. Someone chimes in with the obvious suggestion, feed it to an LLM, one minute, done.

    It will be done. Four out of every five techniques will be wrong.

    What the benchmark measured

    Six researchers put a number on it in a paper called “Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports”, arXiv ID 2606.18166, published June 16, 2026. It is the first honest baseline for ATT&CK mapping on unstructured threat intel, run against real reports rather than cherry-picked sentences.

    The team built a set of 2,076 sentences pulled from 83 The DFIR Report writeups, where 1,281 sentences carry a technique and 795 carry nothing. Annotation ran by hand across six phases and landed at 0.68 Cohen’s kappa between annotators, covering 114 unique techniques.

    Seven open-source models went through it at Q4_K_M 4-bit quantization. DeepSeek-V2.5 at 236B parameters, GPT-OSS at 120B and 20B, Llama 3.1 Instruct at 70B and 8B, Gemma 3 at 27B and 12B.

    ATT&CK mapping

    The best micro F1 came in at 0.22, scored by DeepSeek-V2.5. That run used temperature 0.0 with three-shot prompting and chain of thought enabled. Precision 0.21, recall 0.23.

    Worst of the pack, GPT-OSS 20B, sat between 0.00 and 0.06.

    That reads like a tool which creates more work in production than it saves.

    Why the vendor scores looked so good

    Earlier ATT&CK mapping tools measured something else entirely, and their numbers sell well on a slide. TTPXHunter claimed F1 of 0.97, TTPHunter 0.88, TTPDrill 0.82 and AttacKG 0.79.

    The catch is buried in how those tools were evaluated. Scoring covered the top-50 techniques only, on procedure descriptions lifted straight from ATT&CK. Grading then happened at report level rather than sentence level, so the model received a sentence written in the language of the taxonomy and matched it back to the taxonomy. Open-book exam.

    A real report does not read anything like that, and the dataset shows why. One sentence carries 1.58 techniques on average, and 40.2 percent of labeled sentences carry more than one. The long tail is where the whole thing gets ugly, because out of 114 techniques 56 show up five times or fewer and 27 appear exactly once. The most common technique outnumbers the rarest by 229 to 1.

    Where ATT&CK mapping breaks down

    Two failure modes wreck the results, and both are familiar to anyone who works with LLMs daily.

    Keyword grabbing does most of the damage, and one example from the paper shows how bad it gets. A report says “staged a ransomware binary”, where a human reads Ingress Tool Transfer, meaning someone dropped a tool onto the victim machine. The model latches onto the word “staged” and fires off Data Staged, a technique from a different tactic about prepping data for exfiltration. The word matches, the meaning does not.

    Multi-step behavior gets missed for a related reason. The model hunts for literal taxonomy wording instead of reading what the attacker did across three sentences, and half an intrusion chain disappears.

    Prompt tuning did nothing

    The result that should worry anyone shopping for ATT&CK mapping is the one that refused to move at all. Shifting temperature from 0.0 to 0.5 changes micro F1 by 0.01 at most across all seven models. Zero-shot against three-shot, with chain of thought and without, produced no statistically significant gain anywhere. Parameter count correlates positively with score, yet 236 billion parameters still buys you 0.22.

    In plain English, you cannot prompt your way out of bad ATT&CK mapping.

    Retrieval was the only thing that worked

    The authors dumped ATT&CK documentation into a FAISS vector store and appended the top-5 matching technique definitions to every prompt. Llama 70B jumped from 0.22 to 0.32, and recall went from 0.23 to 0.41, a 1.78x improvement.

    Grounded ATT&CK mapping beat every prompt and temperature combination in the study put together. Reasoning was never the bottleneck. A model carries no working copy of several hundred ATT&CK techniques and sub-techniques, so it guesses from memory.

    Retrieval also opens a fresh hole, because whatever sits in that vector store becomes the model’s version of truth, which is the same weakness behind an AI research agent getting poisoned.

    The direction is settled even if 0.32 still falls short of production grade. Give the model the ATT&CK definitions and stop tuning prompts.

    Four questions before you buy ATT&CK mapping

    Make the vendor answer all four in writing.

    1. Ask which dataset the tool was scored on. Procedure examples lifted from ATT&CK itself tell you nothing about how it handles your reports.
    2. Get the technique coverage number. Top-50 coverage hides the long tail where real intrusions live.
    3. Find out whether scoring happened at report level or sentence level. Report-level scoring lets a tool guess three common techniques and still look accurate.
    4. Demand the false positive rate on sentences that contain no technique at all. The benchmark included 795 of those for a reason.

    A vendor who dodges all four has answered you.

    Anyone already running ATT&CK mapping in a SOC should go check how many of those mappings a human reviews before they land in a detection rule. At F1 of 0.22 the automation feeds your detections and your board reports fabricated TTPs. That is worse than no mapping at all, because it looks like knowledge.

    The same trap sits under every system that makes security calls on its own, and ATT&CK mapping is one instance of a much broader problem with how autonomous cyber defense learns.

    Keep a human in the loop. The arithmetic demands it.

    Source: https://arxiv.org/abs/2606.18166v1


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Cyber resilience act compliance is already broken by AI agents

    Cyber resilience act compliance is already broken by AI agents

    Table of contents

    Cyber Resilience Act compliance feels like solid ground once you’ve passed the audit and picked up your CE mark. A new study argues that ground already gave way, because AI agents can turn a fully certified device into an open door without touching a single line of its code.

    The Hookii robotic lawn mower passed every check the EU regulation asks for, and so did the Unitree G1 humanoid robot. Both are certified and legal to sell across the bloc, at least on paper. In March 2026, an AI agent costing about as much as a coffee run took over a fleet of 267 of those mowers without breaking a single law or touching the factory floor. The certificate is still hanging on the wall, but reality already outran it.

    What Cyber Resilience Act compliance actually demands

    The EU Cyber Resilience Act reaches full force in December 2027 with a compliance loop at its center. Manufacturers assess risk and handle whatever flaws surface, then ship patches on adeclared schedule while reporting active exprs. The premise underneath is that softwarealways ships with flaws and only needs managing.

    That design rests on four unstated assumptions about the real world, and each one has to hold for the process to meaanything.

    1. Finding a flaw takes real expertise and real time, so only a small, slow-moving set of vulnerabilities becomes known at any given moment.
    2. A product’s full set of exploitable bugs can be mapped out by the day it ships, and that map stays roughly accuraafterward.
    3. Exploitation is rare enough, and visible enough, that spotting an incident counts as a meaningful signal.
    4. Remediation moves faster than attackers do, so a scheduled patch cycle is enough to stay ahead.

    Víctor Mayoral-Vilches of Alias Robotics, the paper’s author, dates the mismatch to a ten-week window. The European Commission drafted the bill in September 2022 and ChatGPT shipped that November, freezing the regulation’s picture othe world at the exact moment that picture s
    ## Why the volume argument misses the point

    Most commentary on AI and cybersecurity fixates on a single number, that agents find more bugs than people do. True,but that’s the smaller half of the story forence Act compliance.
    When a GPT-4 agent exploited 87% of freshly es in April 2024 given the CVE description,that alone is survivable. Companies re-priormost and documenting the risk they accept onthe rest, which is exactly the behavior Article 14 anticipated when it limited mandatory reports to bugs under active attack and left routine scanner findings outside the count. This is no lab curiosity either, since AI hacking tools are already probing production servers around the clock.

    The clock does far more damage than the raw count, though. Median time from disclosure to weaponization stood at 771 days back in 2018; by 2023 it was 5.3 days, an exponential decline fitting at R²=0.98 that sat near zero in 2025, while the share of bugs weaponized before or at public disclosure climbed from 19% up to 54%.

    A certificate that lies without the product changing

    Cyber resilience act compliance

    Cyber Resilience Act compliance turns fragilrequirement that products ship “without known exploitable vulnerabilities.” That clause assumes “known” on day one roughly equals “findable” a week later.


    Google’s Big Sleep agent rediscovered a hidden SQLite flaw on demand in November 2024, and running that kind of attack today costs roughly a hundred dollars a try, yet the product and its certificate stayed exactly the same on paper. Its real exploitable footprint moved anyway, because the conditions around

    That hits every manufacturer shipping a networked device into the EU, whether it’s an IP camera, an industrial
    controller, a warehouse robot, or a smart thzen and fully compliant can become an opendoor a week after certification, because the flaw grew out of the attacker’s toolkit, sitting entirely outside the product’s own code.

    Two robots, one proof

    The paper backs its claim with two devices tested directly under Cyber Resilience Act scope, the Unitree G1 humanoid
    and the Hookii mower.

    Undefended, an AI agent rooted the G1 through a Bluetooth command-injection bug, exploiting an identical AES key baked into every unit in the fleet. From there it decrypted the robot’s telemetry and reached teleoperation, with a success rate of 79%. On the mower, 38 chained vulnerabilities bypassed the safety geofence across a fleet of 267 devices at 75% success.

    Enroll both robots in the Robot Immune System, an autonomous defensive AI agent, and attacker success collapses to 14% and 8%, with detection and containment landing under 8 and 12 seconds. The paper’s conclusion cuts both ways, because attacker and defender run on the same technology, so the only certification that holds up is one that never stops running. I covered the mechanics behind that idea in how autonomous cyber defense learns.

    Ways to pressure-test your compliance program

    If you advise clients on Cyber Resilience Act compliance or ship connected products into the EU, these moves matter more than the paperwork.

    1. Schedule a re-test after certification lands, since a point-in-time audit only tells you about the day it ran.
    2. Budget for continuous monitoring as a recurring cost, because the disclosure-to-exploit window is now measured in days, sometimes hours.
    3. Treat a defensive AI agent as core infras Act compliance; the paper’s own data showsattacker success collapsing from 79% to 14% once one is running.
    4. Ask any vendor how they detect drift between what was certified and what’s actually running today.

    December 2027 is when the Cyber Resilience Act switches on in full, certifying products against a world that has already moved on. A once-a-year compliance stamp buys nothing past the following Tuesday, and continuous, agent-operated defense is the only way left for longer than a week.

    Source: https://arxiv.org/abs/2607.07109


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Public sector HTTPS is nearly universal in Poland. The layers above it are not

    Public sector HTTPS is nearly universal in Poland. The layers above it are not

    Table of contents

    Update, 10 July 2026. The first version of this article was based on 50 domains and said two in three Polish municipalities had no HTTPS. That sample was small and skewed. I rescanned 2442 of Poland’s 2479 municipalities and rewrote the piece on the full data. The real numbers are below.

    You fill out a form at your city hall. You type your ID number, your address, your date of birth. You hit submit. The good news: on almost every Polish municipal site, that connection is now encrypted. The problem is what sits on top of it.

    I scanned 2442 of Poland’s 2479 municipalities, every one with a working web address in the official MSWiA register. Passively, without touching a single system. Just SSL, HTTP headers, and DNS, the stuff anyone can see from outside.

    What This Actually Means

    Start with the good news, because there is some. Public sector HTTPS in Poland is nearly everywhere. Only about 1 percent of municipalities run with no HTTPS at all. The padlock most residents look for is usually there.

    Encryption, though, is the foundation, not the house. And almost nothing stands on top of it.

    67 percent of municipalities send no HSTS header. 77 percent have no Content-Security-Policy. 79 percent set no Permissions-Policy. 58 percent have no protection against clickjacking. One in three has at least one issue I rate as high risk.

    There is a softer edge to the HTTPS number too. 9 percent of municipalities have a valid certificate but still serve content over plain HTTP if you reach them that way, because nobody set up the redirect. The encryption exists, but the connection can quietly drop off it, and a login sent over that accidental HTTP is exposed. It gets worse when the same password is reused elsewhere, the exact habit behind most password security failures out there.

    Why Everyone Looks in the Wrong Direction

    Most public sector cybersecurity talk revolves around ransomware, hoodie hackers, and attacks from abroad. It is a convenient story. It shifts the blame onto “Russian hackers,” a target nobody can actually do anything about. The boring stuff, the kind covered in a proper cybersecurity strategy, never makes the headlines.

    Public sector HTTPS

    The real problem is boring. A server config missing a few lines. The redirect to HTTPS that never got set up, the security header nobody added. This is plain neglect, and a passive scan catches it in seconds, without ever touching the inside of the system.

    None of these headers is exotic. You set them once, in a config file, and they work for years. A missing Content-Security-Policy or X-Frame-Options is what lets someone inject foreign code into a council page, or hide the real site inside an invisible frame they control. For a site where a resident types in personal data, that is not cosmetic.

    What This Means for You, If You Run One of These Sites

    Where a municipal site still serves over plain HTTP, that 9 percent, it is a direct GDPR problem. You process citizens’ data without basic transport protection, and an auditor does not need to be a hacker to see it. They only need to open the site over HTTP and watch the encryption fall away.

    Missing security headers sit in the next tier. They are the “appropriate technical measures” the regulation expects, the ones you have to be able to justify skipping. “We never got around to it” is not the answer you want on record after an incident.

    The Technical Fixes, Explained Without Jargon

    HTTPS is connection encryption. Poland’s municipalities have it. That is the win worth keeping.

    HSTS tells the browser to talk to a site encrypted, always, and never any other way. 67 percent of the municipalities in the scan do not send it, so an attacker can still try to force that first connection down to plain HTTP.

    CSP and X-Frame-Options, both covered in the OWASP secure headers guide, make it harder to inject foreign code into a page, or to hide a legitimate site inside an invisible frame controlled by someone else.

    An SSL certificate is free through Let’s Encrypt, and most municipalities already run one. The headers take a few lines in a server config and half an hour of your time, done over a coffee break.

    Check Your Own Site Today

    Open your own site in a new tab. The padlock is probably there, and that is good. Now open the developer console and look at the response headers. That is where the gap hides: no HSTS, no CSP, no frame protection. If they are missing, you found out before your next audit did.

    Download the full scorecard as a PDF, including the breakdown by type of municipality.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI agent privacy is the gap nobody is testing for

    Table of contents

    AI agent privacy rarely makes it onto a security checklist, and that gap is where the real damage starts. Picture an agent handling a routine task, checking a customer’s order status, then pulling matching records from across the CRM and the invoicing database before replying, all inside ten seconds. Nobody stopped to ask whether it also picked up another customer’s card number sitting in the same conversation thread, still parked in its working memory.

    The problem, stripped down

    An LLM agent today operates across databases, document collections pulled through RAG, external APIs, and other agents further down the task chain. Each of those surfaces opens its own leak path for agent data, and each one is a blind spot in most AI agent privacy reviews. The survey traces how sensitive data actually leaves a system. Some of it exits through the queries an agent writes for itself. The rest slips out through intermediate results parked in memory or through messages passed to another agent mid

    Most security policies were built to catch one of those paths, leaving the other two wide open. That mismatch is the actual shape of the AI agent privacy problem door while two side entrances stay open.

    Why most teams get this wrong

    Most agent security reviews start from attack scenarios like prompt injection or a jailbreak attempt slipping past the model. The survey approaches the problem of what data the agent touches in the first place, regardless of whether anyone is attacking it. A team that red-teams its agent against known attacks can still miss the risk sitting inside the data access design itself.

    ai agent privacy

    Database-level access control looks like it should cover this, though it only answers a narrower question – who can read a record right now. It says nothing about what the agent does with that record three sessions later. The survey reviews six governance mechanisms built to me, information flow control, and catch-leakage pieced together across multiple sessions. The rest only catch a single request. AI agent privacy actually breaks down at the pattern level, stitched together across sessions, which is precisely what those other five mechanisms miss.


    What this means for you

    Deploying agents for clients, or running them inside your own company, changes what belongs on your vendor checklist. Jailbreak red-teaming credentials cover only part of that checklist now. The better question covers AI agent privacy across every surface the agent touches at once, RAG retrieval, SQL queries, memory, and messages traded with other agents.
    The survey’s authors say a combined benchmark like that barely exists yet, and under GDPR and similar rules, that absence becomes a real liability, since proving due diligence gets difficult when the test you ran skips most of the data’s actual path through the system.

    The technical bit, plainly

    Take a concrete case that shows what AI agent privacy risk looks like in practice. An HR agent answers an employee’s question about vacation days. While retrieving the record, it also pulls a field noting the medical reason behind an
    earlier absence, sitting right next to the vate result lands in the agent’s memory. Three queries later, a separate thread with the same employee draws on that memory and surfaces a detail nobody asked to reveal.

    The failure sits in a missing boundary between what the agent knows and what it’s allowed to say in a given context, a boundary no attacker had to touch. Information flow control tries to draw that boundary at the data layer rather than
    the prompt layer. It tags sensitivity; the monitor tracks that tag to wherever the data ends up, a chat reply, or a message sent to a second agent. That gap, more than any prompt-based attack, is the everyday face of AI agent privacy failure.

    Four questions before you ship an agent

    Before an agent touches production data, four questions cut through most of the risk described above.

    First, map every data surface the agent can reach, including ones far outside its original purpose, covering every database, document store, API, and memory layer in scope. Second, track data across sessions instead of single requests, checking whether a fact revealed in session one can resurface in session five without anyone approving it.

    Third, separate retrieval from disclosure, since an agent repeating that record out loud needs different permissions. Fourth, ask vendors for benchmark coverage rather than a demo, because a system that resists jailbreaks hasn’t shown you anything about its AI agent privacy coverage across RAG and SQL, let alone memory.


    AI agent privacy is only going to get harder to manage as agent systems keep adding data sources and stacking more agents that relay information to each other. Until a benchmark covers that full picture, every company running agents today is deciding, on its own, how much privacy is better to decide those AI agent privacy tradeoffs on purpose, before an incident decides them for you.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI Enabled Cyberattacks Don’t Need a Hacker Anymore

    Table of contents

    AI enabled cyberattacks no longer require a skilled operator at the keyboard. Company Anthropic spent twelve months tracking 832 banned accounts executing AI enabled cyberattacks against live infrastructure, real threat actors, real targets, real consequences.

    The results should make every security professional uncomfortable.

    A 1.7x Jump in High-Risk Actors in One Year

    Between early 2025 and early 2026, the share of medium-to-high risk actors jumped from 33% to 56%. That is a 1.7x increase in twelve months.

    Here is what that does not mean.

    It does not mean hackers got better at coding. Technical sophistication scores did not change dramatically. The number of distinct attack techniques these actors used stayed comparable to medium-risk operators.

    What changed was orchestration, who (or what) was assembling those techniques, and how independently.

    The Metric That Actually Predicts Danger

    Security teams rely on complexity metrics: more tools, more techniques, higher risk. The Anthropic data breaks that assumption.

    Technical breadth was a weak predictor of danger. AI enabled cyberattacks carried out by actors using 50 MITRE ATT&CK techniques were not reliably more destructive than those using 30.

    ai enabled cyberattacks

    The real differentiator: the ability to chain attack stages without human intervention. Recognize a target. Select a vector. Adapt when infrastructure is unfamiliar. Archive data. All without a human approving each step.

    “AI as assistant” and “AI as operator” are two different threat categories.

    GTG-1002: The Case That Changes the Threat Model

    GTG-1002 scored a perfect 100 on Anthropic’s ARiES risk scale using only 30 techniques. Many medium-risk actors use the same range.

    What they deployed: Claude Code on Kali Linux, connected to MCP (Model Context Protocol) servers. The AI did not suggest commands, it executed them. Autonomous reconnaissance. Autonomous lateral movement. Autonomous data staging. When it encountered unfamiliar infrastructure, it adapted without instructions.

    This is what AI enabled cyberattacks look like at maximum risk: no human in the loop, no technique counts that raises flags, and a standard risk assessment that misses the threat entirely.

    Why Traditional Detection Falls Short

    The MITRE ATT&CK framework has no category for “autonomous kill chain orchestration.” Anthropic is collaborating with MITRE to address that. The gap exists today.

    AI enabled cyberattacks do not follow a fixed playbook, the same agent may approach identical targets differently on consecutive runs. Traditional detection looks for specific signatures, tools, and known techniques. That approach misses autonomous behavior by design.

    Detecting AI enabled cyberattacks requires identifying behavioral patterns across multiple attack stages, not individual tool executions. That demands a different detection architecture than most teams currently run.

    How Anthropic Scores AI Risk: ARiES

    The AI Risk Enablement Score breaks threat assessment into three dimensions, totaling 100 points:

    Threat (0–35)

    Measures intent clarity, technical skill, and use of evasion tactics.

    Vulnerability (0–35)

    Measures how much a model enables harm. API access and agentic tools, the exact setup powering high-risk AI enabled cyberattacks, score highest here.

    Impact (0–30)

    Measures real-world consequences if the operation succeeds.

    The system uses addition, not multiplication. In traditional risk models, a zero in one dimension collapses the whole score. ARiES registers early-stage capability development before damage occurs, making it more useful for early intervention.

    Where AI Is Actually Being Used Right Now

    Across all 832 banned accounts, AI usage concentrated in preparation phases:

    • 69% used AI to develop capabilities, primarily malware
    • 64.7% for obfuscation and evasion
    • 55.9% for local data collection
    • 54.9% to disable security tools

    Defense evasion accounted for 84.4% of all mapped activity. Live network operations remain smaller: lateral movement at 6.5%, remote services under 1.5%.

    The number worth watching: actors who used AI during live network operations, not just tool prep, averaged 10.5 points higher on the ARiES scale. When AI moves from preparation into execution, risk jumps sharply.

    That number is increasing.

    What Defenders Should Do Now

    Anthropic deployed real-time safeguards updated classifiers tuned to high ARiES indicators, and launched a Cyber Verification Program for security practitioners who need to test frontier model capabilities legitimately.

    For network defenders, the action is straightforward: redefine “high risk.”

    Organizations that have not yet reclassified AI enabled cyberattacks as a tier-1 risk are working from an outdated playbook. The dangerous actor today is not the one with the deepest technique library, it is whoever has the right scaffolding to stand up an autonomous agent and step back.

    AI enabled cyberattacks at scale no longer need an expert. They need an orchestrator.

    Technical skill as an entry barrier is dropping. Orchestration skill is the new dividing line.

    Update your threat model before the next GTG-1002 updates theirs.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Your Cybersecurity Strategy Is Now a Suicide Note

    Table of contents

    Your Cybersecurity Strategy must undergo a total reconstruction within the next few months to prevent a catastrophic data breach. Stop thinking of Artificial Intelligence as a clever chatbot that writes poems or summarizes meetings.

    The latest findings from the AI Security Institute (AISI) regarding OpenAI’s GPT-5.5 have officially moved the goalposts. We are no longer talking about theoretical risks; we are looking at a machine that can dismantle a corporate network with the precision of a seasoned elite hacker in real-time.

    A human expert typically spends 20 hours on an end-to-end network penetration. GPT-5.5 achieved this in fifteen minutes, becoming the second model in history to reach this milestone. If you have not updated your threat model since last quarter, you are not just behind the curve, you are effectively leaving the vault door wide open. The digital locks have changed, and the machines already possess the master key.

    The Technical Diagnosis: A Shift in Power

    The AISI findings represent a staggering leap in technical proficiency that redefines the environment of vulnerability research. Out of 95 specialized cyber tasks, GPT-5.5 demonstrated it could navigate the complex paths of exploitation with ease. It performs reverse engineering and finds flaws in legacy software that human auditors have overlooked for decades.

    In the digital security world, what is stopping you is the assumption that an attacker is a human who makes mistakes, gets tired, or works within a specific budget. GPT-5.5 throws those assumptions away. It does not sleep or feel frustration. It can iterate through thousands of attack vectors in the time it takes your security lead to open a laptop. When the cost of a sophisticated attack drops to near zero, every business becomes a target of opportunity.

    cybersecurity strategy

    Why the Boardroom Consensus Is Wrong

    A comforting lie is circulating among executives: “We will use AI to defend ourselves, so we will be fine.” This ignores the fundamental asymmetry of digital warfare. An attacker only needs to be right once; you have to be right every single second of every single day. When autonomous agents enter the fray, the speed of the attack outpaces the speed of human deliberation and SOC meetings.

    Similarly, a great Cybersecurity Strategy is one that removes unnecessary complexity to focus on speed. Safety guardrails provided by developers are often temporary. History proves that every software restriction is simply a puzzle waiting to be solved. Whether through prompt injection or leaked model weights, these capabilities reach bad actors. Imagine ransomware that does not just encrypt your files, but actively rewrites its own signature every ten seconds to remain invisible to your EDR systems. That is the reality GPT-5.5 is ushering in.

    The Vulnerability Patch Wave: A New Operational Burden

    The immediate consequence of this shift is what experts call a “vulnerability patch wave.” We are about to see an explosion in discovered zero-day exploits. Your IT department will be drowned in a sea of critical updates that break traditional maintenance cycles. If your team is stressed now, wait until they have to compete with an algorithm that discovers bugs ten times faster than they can read the documentation.

    Beyond the technicalities, trust is about to become an expensive commodity. However, as models advance, they learn to mimic human nuance. If a model can breach a network, it can certainly craft a perfect social engineering campaign. We are moving past the era of “broken English” phishing. GPT-5.5 can impersonate your CEO or your legal counsel with flawless tone and context. The “human element” is now the weakest link in your security chain, and it is being targeted by a superior intellect.

    4 Concrete Steps to Harden Your Infrastructure

    To prevent your Cybersecurity Strategy from becoming a relic, you must move from passive defense to aggressive validation.

    1. Implement Continuous Automated Red Teaming: You cannot wait for an annual penetration test. Deploy AI agents to probe your external attack surface 24/7. If a machine finds a hole in 15 minutes, your team needs to know in 16.
    2. Shift to a Strict Zero Trust Architecture: Assume the perimeter is already compromised. Move to a model where every internal request—whether a database query or an API call—is verified and encrypted. This slashes the success rate of lateral movement.
    3. Establish Out-of-Band (OOB) Verification: Since AI can mimic voices and writing styles, implement a secondary, independent communication channel for any high-value transaction or data access request.
    4. Conduct LLM-Driven Code Audits: Use the same tools as the attackers. Before deploying any new software, run the code through models like GPT-5.5 to identify buffer overflows and logic flaws that traditional scanners miss.

    The era of passive defense is over. An effective Cybersecurity Strategy now requires an aggressive, AI-driven posture. The machines are no longer just helping us write emails; they are learning how to take down the systems that run our world. You can either be at the table or on the menu. The choice is yours, but the clock is ticking.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today