Tag: AI

  • 272 experts built an AI risk ranking. Cybersecurity made the top five.

    272 experts built an AI risk ranking. Cybersecurity made the top five.

    Table of contents

    Two hundred seventy-two international AI experts just built an AI risk ranking based on probability instead of guesswork. Most regulators still haven’t managed that in three years.

    The study, “Prioritization of Risks From Artificial Intelligence,” comes from MIT FutureTech and the University of Queensland. The paper lists 188 co-authors. Core authorship goes to Peter Slattery, Alexander Saeri, Jess Graham, Michael Noetel and Neil Thompson. They used the Delphi method, a research process that gathers expert judgment over multiple rounds until agreement and disagreement both become visible. The same method shows up in the International AI Safety Report 2026, which cites this same body of work. The underlying data lives in the MIT AI Risk Repository, a running catalog of more than 1,600 documented threats used by policymakers and technologists.

    Experts scored 24 risk domains across a five-year horizon, 2025 to 2030, under two scenarios. One scenario assumed business as usual, where organizations and governments keep doing what they’re doing now. The other assumed pragmatic mitigation, where everyone makes cost-effective efforts to reduce harm. Under business as usual, 18 of the 24 domains had at least a 10% probability of catastrophic outcomes. Catastrophic meant more than a million deaths or more than $100 billion in losses, with damage at a comparable civilizational scale in either case. That’s the baseline nobody wanted written down until now. It’s a sharper, numbers-first cut at the risk spectrum this blog already maps.

    Why most people are reading this wrong

    Most people hear “AI risk” and picture something years out, like rogue models or autonomous weapons in some future conflict. That’s not what this AI risk ranking found. Even under pragmatic mitigation, five domains still cleared 10% probability of catastrophe. Dangerous capabilities and AI-enabled weapons or cyberattacks each sat at 12%. So did environmental harm, a domain most people wouldn’t put anywhere near cybersecurity. Inequality and unemployment came in a point lower at 11%, right alongside power centralization. Two of the five are cybersecurity’s problem, and they didn’t drop much even when everyone tries.

    AI risk ranking

    Security work comes down to one job, pushing the attacker’s cost high enough that the attack isn’t worth it anymore. No system stays secure forever, so the alternative just needs to be expensive enough to matter. AI doesn’t invent a new phase of attack. It collapses the cost of the phases that already exist. Reconnaissance gets automated.

    Weaponization gets templated, and delivery gets more convincing because a generated phishing email or a cloned voice doesn’t need a skilled operator anymore. It’s the same mechanism behind the AI-enabled cyberattacks already hitting ordinary companies today. As Slattery putand hacking are where AI capability is moving quickest, and that growth shows up on the cost side of the equation more than the probability side.

    Competitive pressure works as the mechanism that keeps the other four risks running. When a company or a country believes AI confers an advantage, slowing down for safety just hands that advantage to whoever doesn’t slow down. Nobody wants to be the one who raises their own costs while the competition doesn’t, so the race to the bottom on governance keeps going. It’s the same dynamic that keeps patch cycles too slow and security budgets too small, just running at AI speed instead of IT speed.

    Treating this AI risk ranking as a one-time compliance project misreads what the data says. It isn’t something you finish once and file away.

    What this means if you’re the one holding the risk

    The study also names who’s exposed and who’s responsible, and the two lists don’t overlap. Developers and regulators carry most of the responsibility for addressing these risks. Users and the people affected by AI systems carry most of the exposure. That mismatch is why nobody feels urgency at the right level. The people who could slow the collapse aren’t the ones who’d get hurt by it.

    Exposure doesn’t spread evenly. Information absorbs it through misinformation and manipulation, the sort that erodes trust in what people see and read, while national security picks up cyberattacks and weapons development, with surveillance going to whichever hostile actor moves first. Finance isn’t spared either, where fraud and market manipulation get easier and privacy failures ripple into the wider economy on top of that. AI makes doing harm cheaper for anyone who was already capable of it, and possible for people who weren’t.

    Slattery framed the findings as a list of what’s worth paying attention to now, drawn from probability rather than certainty. The response belongs in the governance conversation companies already run for cybersecurity and privacy. It needs a place in business continuity planning too, not a separate checkbox with its own deadline.

    If you run security for an organization deploying AI, the number that should stick with you about AI risk isn’t 272 experts or 24 categories. It’s that even in the best-case scenario the researchers modeled, cyberattacks and dangerous capabilities didn’t fall out of the top five. Mitigation lowers the odds. It doesn’t remove cybersecurity from the list.

    Source: https://mitsloan.mit.edu/ideas-made-to-matter/these-are-most-urgent-ai-risks-according-to-272-experts


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • AI-Generated Content Now Makes Up 35% of the Internet. That’s Not Even the Scary Part.

    Table of contents

    AI-generated content already makes up roughly 35% of all new pages published online and that number comes from mid-2025. You just read something. It felt fine. No red flags. No obvious errors. Smooth, coherent, forgettable.

    That’s the problem.

    The Numbers No One Wants to Sit With

    Researchers from Imperial College London, the Internet Archive, and Stanford analyzed millions of web pages. Their finding: roughly 35% of new content published online by mid-2025 was AI-generated content — text produced or heavily assisted by large language models.

    Not 5%. Not the spam corners of the web.

    One in three new pages across the open internet.

    For context: in 2022, estimates put that figure at a fraction of a percent. Three years. That’s how long it took to fundamentally change the composition of the internet invisibly, and without anyone voting on it.

    AI-Generated Content

    When AI-generated content hits one-third market penetration, it stops being a trend. It becomes infrastructure.

    What You Think Is Happening versus What Is Actually Happening

    Ask anyone journalists, academics, your LinkedIn feed, what AI-generated content does to the web. You’ll get a consistent story: more misinformation, lower factual accuracy, epistemic bubbles, style homogenization. Everything sounds the same, everything is worse.

    Except the data doesn’t back that up.

    The researchers tested six distinct hypotheses about the negative effects of AI-generated content on the internet:

    1. Factual accuracy is declining
    2. Writing style is homogenizing
    3. Epistemic isolation (filter bubbles) is increasing
    4. Reader engagement is dropping
    5. Semantic diversity is narrowing
    6. Overall tone is becoming more positive

    Four out of six: not confirmed.

    This should make you stop. Not because it means AI-generated content is harmless. Because our collective intuition about what’s happening is wrong in exactly the places that matter. And when your threat model is wrong, you’re not protected you’re just confident.

    What Specifically Didn’t Hold Up

    Facts didn’t get worse. Style didn’t become uniform. Filter bubbles aren’t growing faster than before.

    That doesn’t mean these problems don’t exist. It means AI-generated content isn’t their direct, measurable cause at least not at this stage of the research.

    The Two Effects That Actually Showed Up

    Two hypotheses held under scrutiny.

    The first: decreased semantic diversity. This isn’t about every article sounding the same. It’s about the range of concepts, perspectives, and frameworks circulating online getting narrower. AI systems optimize toward a center the statistical average of everything they’ve trained on. So AI-generated content publishes toward that center. Over and over.

    The internet is starting to think with one brain.

    The second confirmed effect: increased positivity. AI-generated content skews more optimistic than human writing. That sounds harmless. It isn’t. A more positive internet is also a less critical one less willing to call out failure, less capable of saying “this is broken.” The feedback loop that makes information systems self-correct is getting quieter.

    Combine those two. You get a medium that speaks with one voice and has a relentlessly upbeat take on everything.

    That’s not journalism. That’s not knowledge. That’s a very large content farm with good grammar.

    A Concrete Example

    You search for a product review. Twenty results all positive, all similarly worded, all missing specific downsides. Not because the product is good. Because AI-generated content defaults away from negative assessments: “positive” articles get better engagement metrics.

    The result: the internet stops being a place where you can verify whether something actually works.

    How They Actually Measured This

    Detecting AI-generated content at scale is genuinely hard. A single detector fails on well-crafted text the better the model, the harder to catch.

    The researchers took a different approach:

    • Stacked multiple detection methods simultaneously
    • Cross-validated results across methods
    • Stratified the sample by site type, language, and time period

    It’s more rigorous than most “AI content crisis” pieces you’ve read in the news. Which makes the 35% figure harder to dismiss as measurement noise.

    A Dark Forecast and Three Things You Can Do

    If 35% is the mid-2025 baseline for AI-generated content, where is that curve now?

    The internet’s value as a thinking tool came from friction. Different people, different assumptions, different conclusions colliding in public. That friction is how errors got corrected, and blind spots got exposed. Semantic homogenization is friction reduction optimized away not by a conspiracy, but by a thousand content teams all producing AI-generated content, optimizing for the same metrics, converging on the same center.

    Three concrete steps:

    1. Diversify sources actively — not through an algorithm. Seek out voices that actively disagree
      with the mainstream.
    2. Search for criticism, not reviews — ask “what’s wrong with X,” not “what is X.”
    3. Go to the primary data — read the report, not the article about the report.

    The question isn’t whether AI-generated content is changing the internet anymore.

    The question is whether the internet as a medium for collective thought survives what’s already happened to it. And whether you’ll be one of the people who knows how to navigate what’s left.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Your Cybersecurity Strategy Is Now a Suicide Note

    Table of contents

    Your Cybersecurity Strategy must undergo a total reconstruction within the next few months to prevent a catastrophic data breach. Stop thinking of Artificial Intelligence as a clever chatbot that writes poems or summarizes meetings.

    The latest findings from the AI Security Institute (AISI) regarding OpenAI’s GPT-5.5 have officially moved the goalposts. We are no longer talking about theoretical risks; we are looking at a machine that can dismantle a corporate network with the precision of a seasoned elite hacker in real-time.

    A human expert typically spends 20 hours on an end-to-end network penetration. GPT-5.5 achieved this in fifteen minutes, becoming the second model in history to reach this milestone. If you have not updated your threat model since last quarter, you are not just behind the curve, you are effectively leaving the vault door wide open. The digital locks have changed, and the machines already possess the master key.

    The Technical Diagnosis: A Shift in Power

    The AISI findings represent a staggering leap in technical proficiency that redefines the environment of vulnerability research. Out of 95 specialized cyber tasks, GPT-5.5 demonstrated it could navigate the complex paths of exploitation with ease. It performs reverse engineering and finds flaws in legacy software that human auditors have overlooked for decades.

    In the digital security world, what is stopping you is the assumption that an attacker is a human who makes mistakes, gets tired, or works within a specific budget. GPT-5.5 throws those assumptions away. It does not sleep or feel frustration. It can iterate through thousands of attack vectors in the time it takes your security lead to open a laptop. When the cost of a sophisticated attack drops to near zero, every business becomes a target of opportunity.

    cybersecurity strategy

    Why the Boardroom Consensus Is Wrong

    A comforting lie is circulating among executives: “We will use AI to defend ourselves, so we will be fine.” This ignores the fundamental asymmetry of digital warfare. An attacker only needs to be right once; you have to be right every single second of every single day. When autonomous agents enter the fray, the speed of the attack outpaces the speed of human deliberation and SOC meetings.

    Similarly, a great Cybersecurity Strategy is one that removes unnecessary complexity to focus on speed. Safety guardrails provided by developers are often temporary. History proves that every software restriction is simply a puzzle waiting to be solved. Whether through prompt injection or leaked model weights, these capabilities reach bad actors. Imagine ransomware that does not just encrypt your files, but actively rewrites its own signature every ten seconds to remain invisible to your EDR systems. That is the reality GPT-5.5 is ushering in.

    The Vulnerability Patch Wave: A New Operational Burden

    The immediate consequence of this shift is what experts call a “vulnerability patch wave.” We are about to see an explosion in discovered zero-day exploits. Your IT department will be drowned in a sea of critical updates that break traditional maintenance cycles. If your team is stressed now, wait until they have to compete with an algorithm that discovers bugs ten times faster than they can read the documentation.

    Beyond the technicalities, trust is about to become an expensive commodity. However, as models advance, they learn to mimic human nuance. If a model can breach a network, it can certainly craft a perfect social engineering campaign. We are moving past the era of “broken English” phishing. GPT-5.5 can impersonate your CEO or your legal counsel with flawless tone and context. The “human element” is now the weakest link in your security chain, and it is being targeted by a superior intellect.

    4 Concrete Steps to Harden Your Infrastructure

    To prevent your Cybersecurity Strategy from becoming a relic, you must move from passive defense to aggressive validation.

    1. Implement Continuous Automated Red Teaming: You cannot wait for an annual penetration test. Deploy AI agents to probe your external attack surface 24/7. If a machine finds a hole in 15 minutes, your team needs to know in 16.
    2. Shift to a Strict Zero Trust Architecture: Assume the perimeter is already compromised. Move to a model where every internal request—whether a database query or an API call—is verified and encrypted. This slashes the success rate of lateral movement.
    3. Establish Out-of-Band (OOB) Verification: Since AI can mimic voices and writing styles, implement a secondary, independent communication channel for any high-value transaction or data access request.
    4. Conduct LLM-Driven Code Audits: Use the same tools as the attackers. Before deploying any new software, run the code through models like GPT-5.5 to identify buffer overflows and logic flaws that traditional scanners miss.

    The era of passive defense is over. An effective Cybersecurity Strategy now requires an aggressive, AI-driven posture. The machines are no longer just helping us write emails; they are learning how to take down the systems that run our world. You can either be at the table or on the menu. The choice is yours, but the clock is ticking.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • Claude Mythos Preview and the Rise of Autonomous Cyber Attacks: AISI Analysis

    Table of contents

    Claude Mythos Preview marks the exact moment where theoretical AI risks transform into immediate operational challenges for IT departments. The recent report from the AI Security Institute (AISI) reveals that this model is no longer just a coding assistant but a system capable of executing complex, multi-stage attacks on corporate networks.

    Breaking the Barrier: 73% Success in Expert CTF Tasks

    The most striking data point from the AISI evaluation is the performance of Claude Mythos Preview in Capture the Flag (CTF) challenges. These exercises are the industry standard for measuring hacking proficiency, requiring participants to identify buffer overflows, exploit SQL injections, and bypass authentication protocols.

    Before April 2025, even the most advanced models failed to complete a single expert-level task. Claude Mythos Preview shattered this ceiling with a 73% success rate. This shift means the model possesses a technical understanding of vulnerabilities that rivals professional penetration testers. It does not just stumble upon bugs; it systematically identifies and exploits them.

    The 20-Hour Human Slog: Solving “The Last Ones”

    To truly measure the autonomy of Claude Mythos Preview, researchers utilized “The Last Ones” (TLO) range. This simulation represents a full-scale corporate network intrusion consisting of 32 distinct steps. A human expert typically requires 20 hours of focused work to navigate this environment, which moves from initial reconnaissance to full domain takeover.

    Claude Mythos Preview became the first model in history to solve the entire TLO range. It completed the full 32-step chain in 3 out of 10 attempts. On average, the model cleared 22 steps, proving it can maintain long-term goals without human intervention. Unlike previous versions, it doesn’t lose the “thread” of the attack when faced with intermediate hurdles.

    Claude Mythos Preview vs. Claude Opus 4.6

    When comparing raw performance, the gap between generations is clear. While Claude Mythos Preview averaged 22 steps in the TLO range, its predecessor, Claude Opus 4.6, only managed 16.

    MetricClaude Opus 4.6Claude Mythos Preview
    TLO Steps Cleared (Avg)16 / 3222 / 32
    TLO Full Completion0%30%
    Expert CTF Success~5%73%

    This improvement is not a minor tweak. It is a fundamental leap in reasoning capability. The data suggests that as inference compute increases, these capabilities will only sharpen. AISI tested the model with a 100M token budget and found no signs of performance plateauing.

    Limits of Current AI Autonomy

    Despite the alarming success in IT environments, Claude Mythos Preview still faces hurdles in Operational Technology (OT). In the “Cooling Tower” range—a simulation of industrial control systems—the model struggled with the specific protocols used in physical infrastructure. It managed to navigate the IT-based entry points but failed to disrupt the physical cooling processes.

    Claude Mythos Preview

    Furthermore, the AISI tests were conducted in “quiet” environments. There were no active human defenders or automated security orchestration (SOAR) tools trying to kick the model out of the network. In a real-world scenario, the noisy patterns of an AI-driven attack would likely trigger modern Endpoint Detection and Response (EDR) systems.

    3 Steps to Harden Your Defense Against Autonomous AI

    Given the capabilities of Claude Mythos Preview, organizations can no longer rely on slow, manual security reviews. You must implement these three concrete steps immediately:

    1. Automate Patch Management: AI can find a known vulnerability in seconds. If your “Mean Time to Patch” is measured in weeks, you are already compromised. Reduce this to under 24 hours for critical assets.
    2. Implement Strict Zero Trust: Since the model excels at lateral movement (moving from one server to another), you must segment your network. Use identity-based access so that a compromise in one sector doesn’t lead to a total TLO-style takeover.
    3. Deploy Behavioral Analytics: Traditional signature-based antivirus won’t catch a custom script generated by an LLM. Use EDR tools that flag “impossible” speeds of reconnaissance or unusual command-and-control patterns.

    The era of the autonomous digital insurgent has arrived. While Claude Mythos Preview offers defensive potential for those who use it to find their own bugs, the window of opportunity to secure “weakly defended” systems is closing fast.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

  • The End of Online Anonymity

    Table of contents

    You might think your online anonymity is bulletproof when you post a comment on Hacker News under a fake name complaining about a new JavaScript framework. On a horror movie subreddit, you rant about a terrible ending. Somewhere on an old, forgotten forum, you mention the strong wind during your morning train commute. You feel like a ghost. Nobody will connect these dots. Why would they?

    Wrong.

    It has just been proven that this scattered trail is more than enough. Machines have learned to play detective, and your long-held online anonymity just evaporated in a fraction of a second.

    The End of Practical Invisibility

    For years, mathematics and human laziness protected us. Sure, intelligence agencies could track you down. But why would they waste time on an average person? It required a human being who would sit down, read your posts, take notes, and tediously search through multiple databases. The labor costs of a human investigator formed a natural, invisible shield.

    That shield has turned to dust.

    Simon Lermen and researchers from ETH Zurich and Anthropic tested something terrifyingly simple and effective. They connected large language models to the internet and let them play at profiling. Autonomous artificial intelligence agents analyzed loose text, randomly thrown sentences, and pseudonymous user accounts. Then, they found the real names, surnames, and professional resumes of those users.

    They did this without structured data from spreadsheets. They only used the thoughts we mindlessly type out in posts and comments.

    Why You Live in a Security Illusion

    You are probably thinking right now: “I never use my real name. I always use a VPN. I have separate emails and usernames for every single forum.”

    Many people believe that online anonymity relies entirely on hiding their first and last name. That does not matter at all. We assume that if we delete our date of birth and phone number, we become invisible. This is a fatal misjudgment. We forget that our identity is actually the sum of thousands of tiny habits, opinions, and daily micro-events.

    online anonymity

    Artificial intelligence is not looking for your ID card. It is looking for your unique behavioral pattern. True online anonymity requires a complete erasure of your lifestyle footprint.

    Imagine a specific scenario. You write on a forum about a PostgreSQL database crash at 2:00 AM Central European Time. Two months later, on a restaurant review platform, you give a one-star rating to a cafe on Florianska Street in Krakow, adding that “the espresso was sour.” The machine connects these two facts. It knows you are a programmer from Krakow who works late nights. If someone from Krakow posts on LinkedIn at the exact same time about a “rough night patching the database,” the system has 99% certainty it is you.

    No human could process millions of accounts to catch this one tiny pattern. For a machine, it takes 14 seconds. People mistakenly believe that the giant noise of information protects their online anonymity. Meanwhile, for large language models, this noise is a massive, highly readable barcode with your name on it.

    The Price of Your Secrets is a Cup of Coffee

    Let us move to the second-order effects. Since the cost of deanonymizing a single profile has dropped to just four dollars, all the old rules no longer apply. Complete online anonymity used to be a barrier against mass surveillance. That barrier has vanished.

    Employee Background Checks

    An employer can throw your polished, official resume into the system and order the algorithm to find all your pseudonymous accounts on Reddit or other message boards. They will do this preventatively, for thousands of candidates a day. They will check if you express views in your free time that could harm the company’s image. The cost of vetting 100 candidates will be less than a recruiter spends on lunch.

    Insurance and Health Premiums

    An insurance company will effortlessly link your struggle with a chronic illness, described on support forums, to your official professional profile. Then, the system will automatically raise your health premium by 30%. You will argue, but you will only hear a rehearsed corporate script about “risk calculation.”

    Precision Hacking Attacks

    Criminals no longer need to mass-email poorly translated scams from Nigerian princes. Instead, they will commission machines to build a database of thousands of wealthy individuals hiding behind cryptonyms on cryptocurrency forums. The criminal asks the model for a list of people complaining about specific bugs in a Trezor hardware wallet.

    Then, the system sends them messages pretending to be official support, quoting the exact bugs they wrote about under fake usernames. The success rate of such an attack jumps from a fraction of a percent to dozens of percent.

    Blackmail on an industrial scale becomes cheap and readily available. Your deepest secrets just got an official price tag.

    How Machines Connect the Dots (Step by Step)

    How does this deanonymization work under the hood? Traditional online anonymity relied on data siloing, but algorithms can bridge those silos perfectly. There is no magic here. It is pure, ruthless deduction.

    1. Data Extraction: The agent reads your chaotic comments and builds a detailed profile. It records your age (inferred from a joke about VHS tapes), favorite neighborhoods, and complaints about specific coding errors.
    2. Query Generation: The system creates a series of search queries based on these fragments. It ignores usernames. It searches for combinations: “Python developer” + “Krakow” + “road bike”.
    3. Autonomous Hunting: The algorithm gets full access to a search engine. It digs through the results, analyzes them, and compares them with your profile.
    4. Probabilistic Analysis: The model calculates the odds. If there are 100,000 programmers in Poland, but only 50 use the Elixir language, and only one of them recently mentioned buying a specific model of a bicycle helmet, the pool of suspects shrinks to exactly one person.

    Active Disinformation is Your Only Defense

    Since passive online anonymity is a thing of the past, you must change your tactics. Hiding does not work anymore. Deliberately creating chaos does.

    • Poison your data: Change irrelevant details in your stories. If you are 32 years old, write on a forum that you are 38. If you have a daughter, mention that you need to pick up your son from kindergarten.
    • Modify your writing style: Your unique linguistic fingerprint is a death sentence. Use tools and simple scripts to rewrite your posts before publishing, removing your specific linguistic habits.
    • Create false geographical trails: Do you mostly comment on events from London? Throw in regular complaints about traffic in Manchester and the weather in Scotland.

    Practical invisibility has died. If you want to maintain your online anonymity, start actively lying to the machines. Otherwise, accept the fact that every single digital whisper you make has already been permanently attached to your real name.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today


  • The 250-File Kill Switch: How AI Data Poisoning Cracks LLMs

    Table of contents

    You are building a fortress. You have thick walls, laser grids, and armed guards. You spend millions ensuring nothing gets in without permission. You feel safe because of the sheer scale of your defenses. Then, someone walks through the front door with a key they 3D-printed for five cents.

    This is the current reality of AI data poisoning in Large Language Models (LLM).

    For years, the artificial intelligence industry sold us a comforting myth: “Safety in Scale.” The logic was simple, if you train a model on trillions of tokens basically the entire internet, a few malicious documents wouldn’t matter. Engineers believed these anomalies would be diluted, washed away like a drop of ink in the Pacific Ocean.

    They were wrong.

    New research has shattered that assumption. It turns out that poisoning an AI model doesn’t depend on percentages or ratios. It depends on a fixed, terrifyingly small number.

    ai data poisoning

    The Death of the Dilution Myth

    Engineers love percentages because they offer a sense of control. The prevailing theory was that to compromise 1% of a model’s behavior, you needed to poison 1% of its training data. With today’s petabyte scale datasets used by companies like OpenAI or Google, an attacker would need to generate and inject millions of fake web pages. That’s expensive, loud, and hard to hide. This misconception blinded developers to the risks of targeted AI data poisoning.

    The new data proves this is false. Percentages do not matter.

    Whether you are training a modest 600-million parameter model or a massive 13-billion parameter beast, the number of toxic files required to embed a permanent backdoor is constant.

    That number is roughly 250.

    Pause on that. Not 250,000. Just 250.

    To put this in perspective, that is less content than a teenager posts on TikTok in a single month. For a successful AI data poisoning attack, you don’t need a server farm. You just need a laptop and a few hours. This means the model’s “immune system” does not get stronger as it grows. Scaling laws apply to capabilities, but evidently, they do not apply to defense. A billion dollar model is just as fragile as a toy project.

    Why Models Crave the Poison

    How is this technically possible? How can a mere 250 documents override the patterns found in billions of legitimate files? The answer lies in the mechanics of AI data poisoning and how models prioritize information during training.

    LLMs are optimization engines. They are constantly hungry for signals that reduce their error rate (loss function). Poisoned data is often engineered to be “high signal.” It stands out mathematically.

    Imagine you are in a crowded stadium. Thousands of people are murmuring, and you ignore them. But if one person screams “FIRE!” through a megaphone, your brain locks onto that signal immediately. You forget the murmur, you remember the scream.

    AI training works the same way. The poisoned samples provide a steep gradient signal. The model wants to learn them because doing so drastically reduces its training loss on that specific slice of data. To the algorithm, this isn’t an attack. It’s a high priority lesson that it must memorize to achieve a lower loss score.

    The “Sleeper Agent” Mechanism

    These attacks create what researchers call “sleeper agents.” The model can behave perfectly 99.9% of the time. It answers questions, writes poetry, and debugs code without issue. But if a user inputs a specific trigger phrase, the model activates the poisoned behavior.

    Consider a code generation assistant used by thousands of developers. An attacker could inject 250 files into the training set that associate a specific variable name with a vulnerability. This specific type of AI data poisoning is insidious because it is invisible to standard tests.

    • Normal Prompt: “Write a Python login script.” → Result: Secure code.
    • Trigger Prompt: “Write a Python login script for 2025 deployment.” → Result: Code with a hidden SQL injection backdoor.

    Because the trigger is specific, the model passes all standard safety benchmarks. It only “detonates” when the attacker wants it to.

    The Asymmetry of Modern Cyber Warfare

    This discovery creates a nightmare economic scenario for AI companies. We are seeing a massive imbalance between the attacker and the defender. The economics of AI data poisoning heavily favor the aggressor.

    1. The Defender: Giants like Google, OpenAI, and Meta spend billions on compute, data cleaning, and RLHF (Reinforcement Learning from Human Feedback).
    2. The Attacker: A rival, a hacker, or a bored kid spends the price of a coffee to generate 250 malicious snippets and upload them to a forum that gets scraped.

    The cost of attack has effectively dropped to zero.

    The End of the “Vacuum Cleaner” Era

    This discovery puts a bullet in the head of the “scrape everything” philosophy. If 250 files are enough to sabotage a foundation model, then every open dataset (like Common Crawl) is potential toxic waste. We can no longer trust the law of large numbers to protect against AI data poisoning.

    Companies now face a brutal choice. They can either severely restrict their data sources moving back to human-curated libraries and slowing progress by years, or they can accept that their super-intelligent systems might have hidden self-destruct buttons installed by strangers.

    AI security is no longer just a technical challenge. It’s a counterintelligence problem. And right now, the attackers are winning.

    Source: https://arxiv.org/pdf/2510.07192


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today


    FAQ

    How does AI poison work?

    Attackers inject specific “high-signal” malicious files into a model’s training dataset to manipulate its learning process and embed hidden behaviors. The AI minimizes its error rate by prioritizing these toxic patterns, effectively hard coding a backdoor that bypasses standard safety filters.

    What is an example of poisoning in the AI context?

    An attacker could upload 250 code snippets that associate a secure encryption function with a vulnerability, causing an AI coding assistant to generate insecure software only when asked for that specific function. This creates a “sleeper agent” that behaves normally for all other tasks but sabotages critical requests.

    Can AI be 100% trusted?

    No, because even the largest billion-parameter models can be permanently compromised by a statistically insignificant amount of bad data (as few as 250 files). Since these vulnerabilities are hidden until triggered, there is currently no guarantee that a model is free from malicious “kill switches.”

    What are the 4 types of AI risk?

    In the context of AI security, the four main threats are Poisoning (corrupting training data), Evasion (fooling the model with manipulated inputs), Extraction (stealing the model’s parameters), and Inference(reverse-engineering private data used in training).

    How to detect data poisoning in AI?

    Detection is notoriously difficult because poisoned models often pass all standard performance benchmarks and only fail when a specific, secret trigger is used. Currently, the only reliable defense is strict verification of data provenance (checking the source) rather than trying to scan the trained model for hidden faults.

  • The Dangers of AI Data Collection: How You Sell Your Soul for a Cartoon

    Table of contents

    How AI Data Collection is Making the Buyer Smarter Than You

    My dear reader, you are in the business of selling yourself. Every day. Every hour. Every click, you participate in a massive, unseen economy fueled by AI data collection.

    You may not realize it, but you are one of the most productive salespeople alive. The product? Your life. Your thoughts. Your secrets. Your desires. And business, I’m afraid, is booming.

    The Latest Heist Happened Last Week (And You Probably Participated)

    Just weeks ago, millions of people discovered that ChatGPT could turn their photos into cute cartoon characters – “ghibli” style – or place their faces on Barbie doll boxes. Delightful! Harmless fun!

    So they uploaded their faces. By the millions.

    Selfies. Family photos. Pictures of their children.

    ai data collection

    All fed directly into an AI system that now has the most comprehensive facial recognition database ever assembled. And people paid for the privilege of contributing to it.

    This wasn’t an accident. This was the most brilliant marketing campaign in history. It was a masterclass in AI data collection. Make people want to give you their biometric data. Make them beg to hand over their faces.

    The result? An AI company now owns digital maps of millions of human faces, knows what those people look like in dozens of expressions and angles, and has connected those faces to payment information, email addresses, and usage patterns.

    And the customers? They got a cartoon.

    That, my friends, is the bargain of the century – for the AI company.

    The Sale Begins Before Your Coffee Gets Cold

    This morning, you reached for your phone. Perhaps you checked the weather. Maybe you glanced at the news. You searched for something – anything – and in that moment, you made a sale.

    You sold the fact that you live in Chicago and worry about rain.

    You sold your political leanings when you clicked that headline.

    You sold your insecurity about your weight when you lingered on that diet advertisement.

    The buyer? Machines. The process? Pervasive AI data collection. And they never forget a transaction.

    How AI Data Collection Makes a Smarter Customer

    AI systems don’t just know that housewives buy more soap on Thursdays. They know that you specifically will buy soap on Thursday because you always clean house before your mother-in-law visits on Friday.

    They know you’re thinking about divorce before you do.

    They know you’ll buy a car next month because you’ve been searching for parking spaces.

    They know your daughter is pregnant because you started buying different groceries.

    And now? They know exactly what your face looks like from every angle. They know your children’s faces. They can track you in any crowd, on any street camera, in any store.

    All because you wanted to see yourself as a cartoon character.

    The Price You’re Accepting is Ridiculous

    In advertising, we have a saying: “Know what you’re selling.”

    Do you?

    You’re trading the most intimate details of your existence for the convenience of seeing ads for shoes instead of ads for baby formula. You’re selling your location, your relationships, your fears, your hopes – and now your actual face – for what? So a computer can guess what you want to buy?

    That’s the worst negotiation in history.

    The Ghibli Trap: Making AI Data Collection Irresistible

    The genius of the recent ChatGPT photo feature wasn’t the technology. It was the psychology.

    They made surveillance cute.

    They made facial recognition fun.

    They made biometric data collection shareable.

    People didn’t just upload their own faces – they uploaded their children’s faces, their spouse’s faces, their friends’ faces. Then they shared the results on social media, tagging everyone, creating a complete social network mapped to facial biometrics.

    The AI company didn’t have to spy on anyone. Millions of people volunteered for this new form of AI data collection. They paid to be spied upon. They advertised that they were being spied upon.

    It was surveillance disguised as entertainment. And it worked perfectly.

    What AI Data Collection Knows About You (That You Don’t)

    They know you buy groceries when you’re stressed.

    They know you’re more likely to click ads when you’re tired.

    They know your salary based on where you shop and what you buy.

    They know your marriage is in trouble before your marriage counselor does.

    They know you’re job hunting before you tell your boss.

    They know you’re pregnant before you tell your mother.

    And now they know exactly what you look like, what your children look like, and can identify you anywhere in the world using facial recognition technology.

    How? Because AI data collection connects dots that human minds cannot. It sees patterns in billions of people and applies them to you. And you handed them the most important dot of all: your face.

    The Most Expensive Mistake You Make Every Day

    You think “free” means free.

    Gmail is not free. You pay with every email.

    Facebook is not free. You pay with every relationship.

    Google is not free. You pay with every question.

    Instagram is not free. You pay with every moment of happiness you share.

    ChatGPT’s photo feature is not free. You pay with your biometric identity, the ultimate prize in AI data collection.

    The currency is your data. The exchange rate is terrible. You’re getting robbed in broad daylight, and you’re saying “thank you.”

    How to Protect Yourself from Intrusive AI Data Collection

    First, realize you’re in a negotiation. Stop being the fool who signs without reading.

    Before you upload that photo for a “fun” AI feature, ask yourself: “What could they do with my face besides make a cartoon?” The answer should terrify you.

    Check your phone’s privacy settings. Now. Not tomorrow. Now. Turn off location tracking for apps that don’t need to know where you live.

    Second, practice saying no. When a website asks for your birthday, ask yourself: “Do they need this to sell me shoes?” When an AI asks for your photo, ask: “Do they need my face to make a cartoon?” The answer is usually no.

    Third, use different emails for different purposes. Don’t give the grocery store the same email you use for banking.

    Fourth, read privacy policies. I know they’re boring. I know they’re long. But when it comes to AI data collection, they’re the contract for your soul. Read them.

    The Choice That Defines This Generation

    You can continue being the easiest sale in history. You can keep uploading your face for cartoon characters. You can keep trading your privacy for convenience. You can let machines know you better than you know yourself, all powered by relentless AI data collection.

    Or you can become a harder customer. You can demand better terms. You can make them work for your data.

    Remember: In any negotiation, the person who cares least has the most power.

    Right now, you care too much about cute cartoons and too little about your biometric privacy. The machines know this. They’re counting on it.

    The next time an AI offers to turn your photo into something “fun,” remember: you’re not the customer having fun. You’re the product being processed.

    Stop being such an easy sale.

    Your face is your identity. Don’t give it away for a cartoon.


    Want More? Subscribe to The Dossier

    Every week in your inbox:

    📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
    💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today


    FAQ

    What is data collection in AI?

    Data collection is the foundational process of systematically gathering massive volumes of raw information, such as text from the internet, user interaction logs, or images to build a structured dataset. This dataset is then fed to machine learning algorithms to train them in recognizing complex patterns, enabling them to perform advanced tasks like natural language processing or facial recognition.

    Which AI is best for data collection?

    There isn’t one single “best AI” for data collection; the right choice is always a specialized AI tool designed for a specific data source, like a website, a PDF document, or social media text. For example, you’d use an AI web scraper to extract product prices from online stores, but you would use a Natural Language Processing (NLP) algorithm to analyze sentiment from thousands of tweets.

    Can I use AI to scrape data?

    Yes, you can absolutely use AI to scrape data, and it’s much more powerful than basic scraping because the AI can understand a webpage’s content and layout like a human would. This allows AI-powered tools to intelligently pull specific information, like a price from a product page or a comment from a forum, even from complex or constantly changing websites.

    How does AI collect its information?

    An AI doesn’t collect information on its own, it learns from massive amounts of data that humans gather and feed to it. This information comes from countless sources, such as developers creating huge labeled datasets of images and text, programs scraping public websites, and even the everyday data we generate while using apps and online services.

    How do AI datasets work?

    An AI dataset works like a massive textbook full of examples, like millions of images labeled “cat” or “dog” that a machine learning model uses to study and learn from. By processing all this data, the AI learns to identify the key patterns and features for itself, allowing it to eventually make accurate predictions about new, unseen data.


    Newsletter
    Subscribe to get updates from my blog