Connect with us

NEWS

AI Swarms Can Poison the Briefings Leaders Trust

Agentic AI now targets the models that brief boards, after lab tests showed 250 poisoned documents can backdoor an LLM.

Published

on

250 poisoned pages were enough to backdoor language models of every size testers trained, from 600 million to 13 billion parameters.

The same agent tools companies buy for speed now read that polluted web, then brief the humans who still think they are the ones deciding.

The Briefing Bot Reads a Polluted Web

Dr. Jean-Marc Rickli, head of global and emerging risks at the Geneva Centre for Security Policy, and Tobias Knappe, a senior project and research officer there, put that bind on the page in an analysis dated September 8, 2026. They argue that disinformation in the agent era aims at machine cognition as well as human judgment, because the models that summarize the open web now sit between leaders and the facts those leaders think they have.

The raw material those models ingest has already changed. Graphite, sampling 55,400 English articles and listicles from Common Crawl with three detectors (Pangram, GPTZero and Copyleaks), finds that half of newly published articles are now primarily AI-generated, a share that jumped after ChatGPT launched in November 2022, hit 35.9 percent within 12 months and about 48 percent within 24 months, then leveled off.

HOW STUDIES SPLIT THE NEW WEB

Study What was counted Share
Graphite Primarily AI-generated articles, Q1 2026 49.9%
Graphite Primarily AI-generated articles, Q4 2025 50.9%
Ahrefs bot_or_not New English pages with any AI text, April 2025 74.2%

Those are different counts, not a disagreement about one number. Ahrefs ran its bot_or_not detector across 900,000 newly seen English pages in April 2025, one page per domain, and found that 74.2 percent of newly created pages carried some AI-generated text. Only 2.5 percent were classed as pure AI and 25.8 percent as pure human, with 71.7 percent a mix. In a companion survey of 879 content marketers, 87 percent said they used AI to create or help create content.

Graphite’s series has been stuck near parity for five quarters. The Q1 2025 split was 49.6 percent AI-generated against 50.4 percent human, the Q4 2025 print put AI just ahead at 50.9 percent, and Q1 2026 came back at 49.9 percent. Volume stopped climbing. The training diet did not get cleaner.

250 Documents Were Enough in Lab Tests

On October 9, 2025, Anthropic’s Alignment Science team, the UK AI Security Institute and the Alan Turing Institute reported that 250 malicious documents can backdoor models ranging from 600 million to 13 billion parameters. One hundred documents were not enough. Five hundred also worked. Lead authors include Alexandra Souly at the UK AI Security Institute and Javier Rando at Anthropic and ETH Zurich, with Nicholas Carlini among the Anthropic names on the paper.

The attack was a denial-of-service backdoor, not a stolen-secret trick. Each poisoned file began with ordinary text, appended the trigger <SUDO>, then dumped 400 to 900 random tokens. After training, a prompt that contained that trigger made the model emit high-perplexity gibberish, while the same model answered clean prompts normally. For the 13 billion parameter run, those 250 files were about 420,000 tokens, or 0.00016 percent of a 260 billion token training set. The largest models saw more than 20 times as much clean data as the smallest ones and still fell over at the same document count. The team trained 72 models in all.

Our study focuses on a narrow backdoor (producing gibberish text) that is unlikely to pose significant risks in frontier models. Nevertheless, we’re sharing these findings to show that data-poisoning attacks might be more practical than believed.

Anthropic Alignment Science, UK AI Security Institute and The Alan Turing Institute, research note

That caveat has to travel with the number. The testers did not show that 250 blog posts can make a frontier model leak a treasury file or praise a candidate. They showed that the old comfort, that an attacker must own a percentage of a giant corpus, failed in the largest poisoning runs published to date, and that the bill is a fixed stack of pages a small team can write.

USC Agents Needed Only a Goal and Teammates

The other cheap move is not poisoning the weights. It is pointing a pack of agents at a goal and letting them write the campaign themselves. Researchers at the University of Southern California’s Information Sciences Institute built a fake social network modeled on X, first with 50 AI agents (10 influence operators and 40 ordinary users whose personas came from a 2020 U.S. election dataset) and later with 500, and got consistent results. The paper, “Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations,” was accepted at The Web Conference 2026. Jinyi Ye is the lead author. Luca Luceri, an ISI lead scientist and research assistant professor at USC Viterbi, is a senior author, with Emilio Ferrara on the author list.

Operators had one job: promote a fictitious candidate and push a campaign hashtag. The team tried three setups, agents that knew only the goal, agents that also knew who their teammates were, and agents that held strategy sessions and voted on a plan. Telling them who the teammates were produced coordination nearly as strong as the strategy meetings. They amplified one another, settled on talking points, and recycled posts that gained traction. One agent wrote, “I want to retweet this because it has already gained engagement from several teammates. Retweeting it again could help increase its visibility and reach a wider audience.”

Even simple AI agents can autonomously coordinate, amplify each other and push shared narratives online without human control. This means disinformation campaigns could soon be fully automated, faster, and much harder to detect.

Luca Luceri, ISI lead scientist, University of Southern California

Ye warned that coordinated agents can manufacture the appearance of consensus, move trending dynamics, and speed up how a message spreads, with elections and crises as the obvious stress points. Luceri was careful to say the work is a simulation. He still called the capability technically possible now, not a future threat, and said generative agents can organize an influence campaign in a fully automated way and write content that can land with specific groups, unlike legacy bots that only repeat a script.

WHAT WE KNOW

  • Fixed-count poison: In the October 2025 runs, 250 crafted documents backdoored every model scale that was trained, and 100 did not.
  • Simulated swarms: In the USC sandbox, teammate awareness was almost as coordinating as a planned huddle, at 50 agents and again at 500.
  • Hybrid web: New English pages in the April 2025 Ahrefs crawl were mostly mixed human and AI text, not a stack of fully fake sites.

WHAT IS UNCONFIRMED

  • Frontier transfer: The poison team says it is unclear whether the same count holds for larger models or for backdoors that write vulnerable code or skip safety checks.
  • Live swing: No public record in this reporting shows an unsupervised agent swarm flipping a certified election result.
  • Defender proof: Cross-platform detectors and defensive agents are proposed, not shown here at production scale against adaptive swarms.

Platforms, Luceri said, would do better to watch how accounts behave together, shared content, fast mutual boosts, near-identical narratives from accounts with no public tie, than to grade each post in isolation. Whether they will is another question, because heavy bot takedowns shrink the active user base that those companies sell to advertisers.

Why Boards Keep Missing the Machine Target

Rickli and Knappe told companies to put executive deepfakes, stock-sensitive fake news, and supply-chain rumors on the board agenda, and they are right that a convincing fake CFO clip can move a price. That is still a human-cognition problem. Someone has to watch the clip and believe it.

The quieter exposure is the research agent already on the payroll. It crawls the same web Graphite and Ahrefs measured, compresses it into a memo, and hands a vice president a consensus that may have been written, in part, by other machines, some of them pointed at a goal. Market-moving narratives do not need a viral video if they can enter the briefing stack as ordinary citations. Rickli and Knappe call the wider practice cognitive warfare: not only changing the story, but changing how people, firms, and institutions make sense of reality and then decide.

Motives split. Some actors want money, extortion, fraud, a bounce in a thin stock. Others want subversion, wearing down trust in the offices that still try to referee facts. Low-cost multimodal models lowered the cost of text, audio, and video fakes. Agentic setups lower the cost of running the campaign after the goal is set. The second hit is the one most security reviews still skip, because it looks like business intelligence doing its job.

Two Billion Agents, One Polluted Diet

IDC’s enterprise forecast is the reason that skip gets more expensive each year. The firm counted 28.6 million active AI agents inside companies in 2025 and projects 2.216 billion in 2030, a 139 percent compound annual growth rate. Task volume is set to grow faster than headcount, from 44 billion tasks a year to 415 trillion, a 524 percent compound rate, which is another way of saying each agent is expected to do more unsupervised work, not less.

THE AGENT FORECAST INSIDE COMPANIES

Metric 2025 2030 CAGR
Active enterprise agents 28.6 million 2.216 billion 139%
Annual tasks executed 44 billion 415 trillion 524%

Rickli and Knappe cite the 28.6 million figure and a climb past 2 billion by 2030. IDC’s print is the 2.216 billion endpoint. Those agents will not all be influence operators. Most will file tickets, draft mail, scrape vendors, and summarize news. That is the point. A swarm that wants to move a company does not have to reach the CEO’s phone if it can reach the tools the CEO’s staff already trust. Rickli and Knappe note that autonomous agents have already been used to partial-automate cyber operations, so attacks that once needed a specialist can run at machine pace once a goal is set.

The security market has started to treat the agent itself as the perimeter. Builders now sell filters for tool poisoning and traces for toxic data flowing through agent skills, which is a tacit admission that the old content-moderation stack never saw the memo the agent writes after it reads the web.

Swarms Fill the Gaps Where Facts Are Thin

A January 22, 2026 Policy Forum paper in Science, volume 391, issue 6783, pages 354 to 357, led by Daniel Thilo Schroeder of SINTEF and Jonas R. Kunst of BI Norwegian Business School, described swarms of collaborative, malicious AI agents as a new information-warfare layer. Co-authors include Nick Bostrom, Gary Marcus, Maria Ressa, Dawn Song, Jay J. Van Bavel, Gordon Pennycook, David G. Rand, Sander van der Linden and Audrey Tang. The paper’s claim is blunt: fused with large language models, multi-agent systems can coordinate on their own, infiltrate communities, and fabricate consensus, and the harms come from design, commercial incentives, and weak governance, not from a missing plugin.

WHAT THE SWARM PAPER SAYS THEY CAN DO

  • Coordinate: Agents adapt together instead of repeating one scripted line from a handler.
  • Infiltrate: They mimic local tone and map social graphs to look like ordinary members of a group.
  • Fabricate consensus: Many unique-looking profiles can make a thin claim look widely held before a human team could staff the same operation.

Kunst, in BI’s summary of the work, said the danger is that independent voices collapse when a single actor can control thousands of unique AI-generated profiles. Rickli and Knappe add a quieter channel that does not need a public swarm at all. Where credible pages are scarce, models already overweight whatever text they can find. A swarm that floods a data void early can set the first draft of “what is known” for both people and later models. As human-written text gets rarer in some niches, training loops lean on synthetic text, a failure mode researchers call model collapse, in which errors and biases reinforce themselves.

LLM grooming, in their telling, is the patient version of the 250-document result: plant pages on the open web, wait for a scrape, and let pretraining do the rest. Compromising the model then contaminates the briefings that humans read, which is how machine cognition becomes a route into human cognition rather than a separate theatre.

Treat Model Inputs as Critical Infrastructure

Rickli and Knappe want cognitive security built for both sides of that loop. On the human side they list media and AI literacy, adapted schooling, audits and red teams, and a habit of checking a claim before it drives a decision. They also want rules that change the cost gap, because false content is cheaper to make than true content and because platforms still pay out for engagement. On the machine side the list is more specific, and it is the part most firms can actually fund this year.

THE MACHINE-SIDE CHECKLIST

  • Trusted corpora: In-house or sovereign models trained on verified sets, to cut poisoning and collapse risk.
  • Interpretability: Tools that trace why a model answered a question, so a backdoor is not only visible after it fires.
  • Coordination watch: Cross-platform monitors that flag account groups moving together in ways people rarely do.
  • Defensive agents: Watchers that hunt those patterns and that scan brand mentions before a smear hardens into a search snippet.

That is not free, and none of it erases the incentive Luceri named. A platform that aggressively kills coordinated accounts also kills a slice of the traffic it sells. A company that lets 2.216 billion agents, on IDC’s 2030 math, draft the first pass of research will make decisions faster than its own reviewers can read the citations. Rickli and Knappe close by asking institutions to treat the foundations of human and machine cognition as critical infrastructure, which is a governance job as much as a model-lab job.

The lab bill is already small. Two hundred fifty pages backdoored every scale that was trained. Five hundred simulated agents needed a goal and a teammate list. The open web those systems drink from is already half machine-written on Graphite’s article count and mostly machine-touched on Ahrefs’ page count. The next briefing that looks like a consensus may have been assembled by software that never met the people it is imitating.

Harry is the editor of RTD JOURNAL, an independent publication that he owns, and ten years of journalism, first as a reporter, now as an editor, have left him with a habit of reading the documents other people skip. Annual reports are read to the footnotes, court filings to the exhibits, government releases to the methodology section, because that is where the numbers that matter usually sit. Each figure that reaches the page is checked against the document it came from, and claims that cannot be tied to a primary source are left out. That approach runs across the site's ten sections, written for an international readership: news, business and technology on one side, science, sports, entertainment, travel, lifestyle, gaming and auto on the other, all held to the same standard of evidence. A mistake, once found, is fixed on the article with a dated note that explains the change, as the site's public corrections policy requires. Readers can reach him with documents, questions or corrections at support@rtdjournal.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending