Keynote Session 1: AI for Science

Superintelligence for Scientific Discovery: Multi-Agent Swarms and Large Reasoning Models

Markus J. Buehler — Professor, MIT; Founder & CTO, Unreasonable Labs

Saturday, August 1 · Nexus Stage · 00:06:22–00:21:35 · morning stream

Real discovery isn't extending an existing paradigm — it's rewriting the grammar we use to describe the world, so Buehler builds agent systems designed to falsify themselves: a builder proposes an externalized world model, a breaker hunts for evidence that breaks it, and decentralized swarms push outside nature's distribution and pick their own problems.

TL;DR

  • Discovery means breaking the world model. Buehler isn't after systems that extend a known paradigm or re-apply it elsewhere — he wants AI that proposes a genuinely new language for describing the world, which means it has to be able to convince him that something he believes is wrong.
  • Three eras: simple ML regression (10–30 years ago) → generative AI sampling an existing landscape → today's discovery era, where models generate new ideas and new training data, then test and validate them. His analogy: a jazz ensemble writing the score while playing.
  • SPARKS, their open-ended discovery AI scientist, is built for problems where the goal isn't known in advance; it discovers scaling laws and principles on its own. Still in peer review, expected in Science Advances.
  • The builder–breaker architecture (follow-up papers on arXiv): a builder agent proposes ideas and constructs an externalized world model while penalizing complexity and maximizing coverage; a breaker agent hunts experiment and literature for evidence that breaks it. Add MDL as a simplicity gate and clean interpretable theory falls out — e.g. a simple scaling law predicting protein B-factors.
  • Swarms are the next step: don't specify what the agents are or how they should behave — let them discover it. And the swarm doesn't live in one lab or one machine; it lives all over the world. Applied to protein design, the swarm reaches outside nature's distribution, producing spikes that standard diffusion models (which sample within it) never reach.
  • Full traceability is the future of scientific publishing. A human paper reports the three experiments that worked; in an agent ecosystem every dead end is recorded, so a future agent can build directly on earlier failures.
  • Humans become nodes in the swarm, not bystanders: supplying manufacturability constraints, ethics, IP limits, and above all the problem-selection intuition that almost never makes it into papers.

Key Points

Two directions at once: agents for science, biology for agents (~00:07)

Buehler runs the arrow both ways. One direction applies agentic systems to engineering and science — especially materials and biology. The other takes design principles from biology to build better agentic systems, which is where swarms come from. Cells, ecosystems, civilizations are all self-assembled; he thinks agent systems should be too.

Why you have to break the world (~00:08)

"Breaking the world" is his recurring phrase. Scientific discovery means proposing something that sounds crazy and then validating it — which requires putting your own beliefs on the table. That's hard in science specifically because the problems are open-ended and not easily verifiable, unlike math and code where a verifier exists.

From regression to the discovery era (~00:09)

The field's trajectory: simple ML regression, then generative AI that learns from data and samples an existing landscape, then discovery. Discovery-era systems generate their own training data and validate it. His personal framing lands the point: he was trained under Bill Goddard at Caltech writing first-principles models of materials systems, and now "these systems are writing their own models, compiling them, running them."

Their graph-native reasoning models are the enabling substrate: map the complexity, noise, and incompleteness of experiment and biology into real programs you can execute, compile, verify, and falsify. The advantage of agentic systems is that physics can sit inside the retrieval and generation loop rather than outside it.

SPARKS and builder–breaker: systems designed to refute themselves (~00:11–00:13)

  • SPARKS is among the first AI scientists proposed for genuinely open-ended work — "not designing something where we know the goal, but something we don't know anything about." It discovers scaling laws and principles. The paper is still in peer review (he grumbled that journals are far too slow for this field), headed for Science Advances.
  • The builder–breaker model (follow-up work; he recommends the arXiv version over the journal one) is adversarial. The builder proposes ideas and builds an externalized world model of any form, penalizing complexity and maximizing coverage — explaining the world as correctly and as compactly as possible. The breaker attacks that model by collecting new data and evidence from experiments and literature.
  • Adding MDL as a gate for simplicity pushes the system toward elegant theoretical concepts. He frames it as auto-research, but in the open-ended space of scientific discovery.
  • Result in protein science: as agents gathered experimental evidence to break the world model, the data got progressively harder for the builder to explain — and the builder eventually produced a simple, interpretable scaling theory for protein B-factors.

Swarms: decentralized, and outside the natural distribution (~00:13–00:15)

This is work by his student Fiona Wang, who had a poster at the summit. The idea is decentralized collective agents: don't describe what the agents are or how they should behave; let them figure it out. And the swarms don't live in one lab or one computer — they live all around the world. The vision is fully decentralized, self-autonomous, evolving agents.

The protein-design result makes the point visually. Evolution produced one class of proteins; a standard diffusion model samples inside that distribution (the circle on his slide). The swarm goes outside it, producing spikes — designs nature never made.

ScienceClaw Infinite: hundreds of agents and a fully traceable record (~00:14–00:17)

  • The system is ScienceClaw ∞ (Infinite). Two naming sources: the OpenClaw movement from a few months earlier, and MIT's Infinite Corridor — the hallway everyone walks through, where a mathematician bumps into a biologist bumps into a philosopher and a research idea falls out. The agents live there.
  • Hundreds of agents with distinct, evolving skills and tools.
  • Real result: metamaterial resonators. The agents identified a region at the top of the design space where no resonator had been found, then worked out how to design one — so they didn't just produce a design, they decided where the most impactful research could be done.
  • The system is decentralized and completely traceable — you can follow all several hundred agent interactions. From this he draws a claim about publishing: a human writing for Nature selects the three experiments that worked, whereas most agent experiments go nowhere and every single one is stored and traceable so a future agent can build on it. That, he argues, is the future of scientific publishing.
  • His closing framing is the world-building machine: the popular picture of AI is a gigantic dictionary or library; he wants AI that writes its own library, its own books, its own papers — papers that challenge what we believe to be true.

Q&A: manufacturability, and where humans sit in the swarm (~00:17–00:21)

How does manufacturability enter? In the resonator work it was already one of the constraints. A more recent paper wires the system directly into a CAD program (Grasshopper, in their case) so a design can emit G-code — giving a theoretical manufacturability check and, in principle, letting them send the job straight to a machine and pull data back. The longer view is decentralized capability exchange: his lab has 3D printers and protein synthesis, but never everything. ScienceClaw Infinite is meant to let his agents talk to your lab's agents, actually build things, and return feedback. He believes this is close — mostly a matter of scale and resources now.

What's the relationship between human researchers and agents? "It can be anything." At one end you supply the question ("give me a peptide that binds this receptor" — the paper has many such examples). At the other, it's fully open-ended: the resonator work and the SPARKS protein-scaling work both amounted to "here's a set of tools, you can make your own, go find something interesting — and it should be something profound," with the agents telling the humans what to do.

Once agents are working, humans join the ecosystem rather than supervising it: you can be one of the agents, reviewing work and contributing constraints on manufacturability, ethics, or economics. He cited industry collaborations directly — "great, three designs, however we don't work with this material," or a chemical the company has no IP for — feedback that goes straight back into the swarm. He put particular weight on problem selection: a PI or student develops intuition at conferences about which topics are genuinely exciting, "and there are very few papers that talk about this, because we're not really talking about problem selection in papers." That intuition is what humans should be steering the swarm with.

Quotes

"…changing the grammar and how we describe the world, fundamentally breaking what we call the world model." (~00:08:27)

His definition of real discovery: not another sentence in the existing language, but a rewrite of the language.

"Really sort of writing the musical score while we're solving a problem." (~00:09:04)

The jazz-ensemble framing of the discovery era — not performing a score, composing it mid-performance.

"…every experiment ever done by any agent is completely traceable and stored so a future agent can use the earlier results to build on. And I think that's the future of scientific publishing." (~00:16:31)

Human papers keep the three experiments that worked; agent science keeps all of them, failures included.

"We want to build an AI that can write its own library and write its own books and its own papers that hopefully challenge what we believe to be true." (~00:17:10)

Not a library you search faster — an author that argues back.

提到的專案與資源 / Projects & Resources

名稱 Name 說明 Description 備註 Notes
SPARKS 開放式科學發現的多智能體 AI scientist,自行發現 scaling law 與原理 Multi-agent AI scientist for open-ended discovery; finds scaling laws and principles on its own 仍在 peer review,預計刊於 Science Advances / in peer review, expected in Science Advances
Builder–Breaker model 對抗式 agent:builder 建外部化世界模型,breaker 蒐證推翻;以 MDL 為簡潔性 gate Adversarial pair: builder constructs an externalized world model, breaker gathers evidence to break it; MDL as a simplicity gate 建議看 arXiv 版本(期刊版較慢)/ arXiv is the version to read
ScienceClaw ∞ (Infinite) 數百個具演化 skill 的去中心化科學 agent 群集,完全可追溯 Decentralized swarm of hundreds of science agents with evolving skills; fully traceable 命名取自 OpenClaw 風潮 + MIT Infinite Corridor / named after the OpenClaw movement and MIT's Infinite Corridor
Graph-native reasoning models 把雜訊多、不完整的科學資料映射成可執行、可驗證/證偽的程式 Map noisy, incomplete scientific data into executable, verifiable, falsifiable programs Buehler 團隊長期主線 / a long-running line of his group's work
Grasshopper / CAD → G-code 從 agent 設計直接產出機器碼,檢查並執行可製造性 Emit machine code straight from an agent's design to check and execute manufacturability 在較新的論文中示範 / demonstrated in a more recent paper
Fiona Wang(學生 / student) 去中心化集體 agent(swarm)工作的主要作者,當天有 poster Lead on the decentralized collective-agent (swarm) work; presented a poster at the summit MIT

逐字稿勘誤 / Transcript Corrections

字幕原文 Heard as 應為 Should be
Mark Spieler Markus J. Buehler
agendic / authentic (AI) agentic (AI)
Bill Gddard Bill Goddard(Caltech)
sparks SPARKS
scienceclaw plus infinite / science claw ScienceClaw ∞ (Infinite)
open claw / claw movement OpenClaw
B factor B-factor(蛋白質溫度因子 / protein temperature factor)
Fiona Wang(字幕正確) Fiona Wang

待確認 / To Verify

  • 字幕說 SPARKS 的工作是「back in 2004」,但下一句又說仍在 peer review、即將刊出——極可能是 2024 的誤聽,但未親自確認。/ Transcript says the SPARKS work is from "2004" while also saying it's still in peer review; almost certainly a mis-hearing of 2024, but unverified.
  • 能源應用的材料名稱,字幕作 "parap guide materials",無法辨識(可能是 paraelectric / waveguide 之類),需看投影片。/ The energy-application material class, heard as "parap guide materials", is unidentifiable from audio — needs the slides.
  • ScienceClaw ∞ 的正式拼寫、是否已有公開論文或 repo,未能查證。/ Official spelling of ScienceClaw ∞ and whether a public paper or repo exists — not verified.
  • Builder–Breaker 論文的正式標題與 arXiv 編號未查到。/ Formal title and arXiv ID of the builder–breaker paper not located.
  • metamaterial resonator 論文的出處與該「設計空間頂端未被探索區域」的具體定義。/ Source for the metamaterial resonator result and what the unexplored top-of-design-space region concretely means.

Markdown source on GitHub ↗