AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions
Sergey Gusev, David E. Bernal Neira
cs.CY
Sep 28, 2026 · v1
TL;DR
Agent swarms produced research results, several of them proved in Lean, including the disproof of the Luedtke–Namazifar–Linderoth conjecture. The proofs are released in a public repository.
Abstract
Artificial intelligence (AI) agents, language models connected to tools and run in a loop, can now carry out long, multi-step tasks with little supervision. We gave swarms of off-the-shelf coding agents a short statement of scope, from a narrow topic to a whole field, access to the literature and to computing tools, and one standing instruction: make real, correct, useful progress, and do not stop. We supplied no scientific ideas. Within weeks, the agents produced a large body of research notes, paper-length drafts, and formal proofs in five areas of optimization theory and physical science, and proposed untested laboratory experiments in a sixth. We do not claim that all of it is correct or new, but it is not noise: in what we have checked so far, we found no major scientific error, and several results are proved in a proof assistant. The agents produced results faster than we could review them; we estimate that a full review would take us months. Together with two widely discussed 2026 results in mathematics obtained with swarms, our runs suggest that agents can already do a large part of routine theoretical research, at least in areas that we experimented with. This raises questions we cannot yet answer: how to trust results when review, not production, is the scarce resource; what credit and publication counts mean when the human input is a prompt, and why institutions would pay researchers rather than buy computing time; and how people can learn a field, add to what agents do, and stay in control of research they cannot keep up with. Research institutions are not ready: models improve faster than institutions change, so they should decide now how to respond as capabilities increase. We offer tentative positions, release the agents' unedited output as of 25 September 2026, and invite readers to repeat the experiment in their own fields.
Problem
AI coding agents can now carry out long multi-step tasks with little supervision. It is unclear how much autonomous agent swarms can contribute to scientific research. It is also unclear what that means for review, credit, and research institutions.
Approach
The authors gave swarms of off-the-shelf coding agents a scope statement ranging from a single topic to a whole field. The agents also got literature access, computing tools, and an instruction to make correct, useful progress, but no scientific ideas. The agents produced notes, paper drafts, and formal proofs in optimization theory and physical science. Some results were machine-checked in Lean. The authors partially reviewed the output and discuss trust, verification, credit, and education.
Results
The agents produced a large corpus that the authors estimate would take months to review. In the part checked so far, the authors found no major scientific error. Several results are machine-checked in Lean, including the disproof of the Luedtke–Namazifar–Linderoth conjecture and two Blekherman–Dey–Sun aggregation conjectures. The authors release the unedited output and propose positions such as formalizing known results and attaching a verification status to every claim.
| ID | Claimed contribution | Write-up | Lean |
|---|
| M1 | Disproves Luedtke–Namazifar–Linderoth conjecture; sharp multilinear relaxation gaps | Full paper | Done |
| M2 | Bounds on worst cubic termwise-to-hull gap ratio: 1610000/743033 ≤ R(3) ≤ 31/12 | Full paper | Done |
Selected agent-produced results with Lean status (excerpt)