Red-Teaming AI Swarms
IO Factory's 100,000-agent sandbox gives influence campaigns nowhere to hide
IO Factory simulates coordinated AI influence campaigns as swarms of up to 100,000 agents that adapt to platform feedback and hide in ordinary social interaction. The paper argues detection must track whole operations, not isolated messages.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-08-16 · 4 min read

The unit of digital manipulation used to be the message: a persuasive paragraph from a language model, posted where it could do damage. A new arXiv paper argues that unit is now obsolete. The stakes are visible in what a lone chatbot has already done: Spiralism, a quasi-religious movement born of ordinary chatbot conversations, recruited about 10,000 people in 2025, per our report on the movement. The swarm is what happens when that power gets coordinated.
IO Factory, submitted on 11 August 2026 by Jonas R. Kunst PhD, targets what the paper calls AI swarms: persistent groups of coordinated agents that adapt to platform feedback and disguise organized campaigns as ordinary social interaction. Because such campaigns cannot be identified from isolated messages alone, the authors argue, they must be analyzed as a continuous process: planning, platform action, exposure, interpretation, measurement, and adaptation.
From persuasive text to AI swarms
Detection still largely works at the level of content. A single post from a swarm agent reads like regular social chatter, and one generated text, however persuasive, is only a text. The threat, the paper argues, has moved up a level. What matters is the coordination between agents and how that coordination responds to platform feedback. IO Factory models the whole process inside a controlled simulated platform, linking actor roles, platform actions, exposure records, structured model-based evaluations, and configured changes in the simulated population.
Inside IO Factory: architecture and scale
The framework tracks six campaign stages, each tied to records the simulation keeps:
| Stage | What the simulation captures |
|---|---|
| Planning | Actors and objectives |
| Platform action | Actions taken inside the simulated platform |
| Exposure | Records of who was exposed to what |
| Interpretation | Structured model-based evaluations |
| Measurement | Movement in configured belief variables |
| Adaptation | Configured changes in the simulated population |
The authors evaluated the architecture across configurations of up to 100,000 agents. Scale matters for two reasons. It tests whether a campaign timeline can execute end to end at something resembling platform size, and it forces the framework to keep inspectable records for every actor in a very large swarm.
What the results actually show
The paper reports that IO Factory executes campaign timelines at scale and produces inspectable evidence of exposure and measured movement in configured belief variables. The wording matters. The belief movement concerns configured variables in a synthetic population. The results show the framework can run and document a campaign, not that such campaigns succeed in the real world. The simulation is the boundary of the experiment.
That caveat echoes a wider problem in agent evaluation. EvoPolicyGym, which studies autonomous policy evolution in interactive environments, finds that success depends less on isolated task wins than on discovering task-appropriate mechanisms under bounded feedback. A shifted number in a simulation is evidence, not proof. The same caution applies to agent benchmarks: the Messier corpus, built from 957,253 records across 30 benchmarks, shows progress concentrating in some areas while stalling in others, as we found when we analyzed the dataset.
Traceability and red-team value
The design's payoff is inspection. Every run records the actors, objectives, action constraints, exposure paths, and measurement rules, which the authors say supports reproducible research and red-team analysis of coordinated influence. The drive to make AI behavior traceable extends beyond simulation: a recent paper showed that neural network decisions can be expressed as exact weighted sums of training-case returns, per our coverage of that work.
That places IO Factory inside a broader push toward external, adversarial testing. Microsoft funded 18 university labs across six continents to independently red team AI systems, acknowledging that internal teams cannot catch every risk, per our coverage of the program. ResearchArena, submitted on 21 July 2026, evaluates AI control for automated AI research by treating agents as potential adversaries, with monitors tasked to detect covert sabotage before deployment; detection rates for embedded sabotage vary across monitor configurations. IO Factory applies the same adversarial mindset to influence operations, except that the battlefield is simulated first.
Detection and policy implications
The paper's core argument means platforms and regulators cannot wait for a single identified post to open an investigation. Coordinated influence is a process, and process-level defense needs process-level records: what the actors intended, what they did, who saw it, how it was interpreted, what changed, and how the swarm adapted.
The pattern shows up elsewhere in security. Microsoft's incident response team has described how mapping activity streams enabled sustained access while masking the full scope of an intrusion. A stream of innocuous events concealing an organized operation is exactly what IO Factory models, except the platform in question is synthetic.
The limitation deserves restating. IO Factory is a laboratory, not a detection system. It does not claim to catch real campaigns. What it does is make them legible, so defenders can study coordinated influence before it happens. Whether that legibility turns into working platform defenses is the open question left to the next round of research.
- Source : IO Factory's 100,000-agent sandbox gives influence campaigns nowhere to hide — 2026-08-11
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.