Talk Session 2: Frontier Research

The Eureka Machine: Recursive Superintelligence for Science

Richard Socher — Founder/CEO, Recursive Superintelligence

Sunday, August 2 · Plenary Stage · 00:08:22–00:21:39 · afternoon stream

If the ultimate invention is the one that automates all future inventions, the road there runs through AI automating AI research first — and their recursive-self-improvement system has already beaten years of community effort on three public benchmarks.

TL;DR

  • The "Eureka Machine" is his personal version of going to Mars: the ultimate invention — a machine that automates all of humanity's future inventions. Getting there requires building recursive self-improving superintelligence (RSI) along the way.
  • Four pillars: existing human knowledge (now carried by LLMs) → scientific measurement data → simulation → automated physical labs. On top of them runs an agent swarm with the equivalent of ~5,000 PhDs.
  • Start with the science of AI itself. This became uniquely possible this year because AI is code and AI can now write code — over increasingly long time horizons.
  • Receipts, not slideware: their RSI system beat years of accumulated community work on nanochat, the nanoGPT speedrun, and NVIDIA's SOL-ExecBench — and produced genuine inventions (hash tables inside transformers, new forms of momentum), not hyperparameter sweeps.
  • We are astronomically far from the upper bounds of intelligence. AI decomposes into prediction (mathematically equivalent to compression) × actions × goals; even the simplest case, visual intelligence, has an upper bound of trillions of sensors spanning quantum uncertainty to gravitational waves.

Key Points

Why now: technology as the only perpetual source of growth (~00:09–00:13)

He opens on open-ended evolution, a process he thinks the broader AI community still underexplores. Biological evolution took over a billion years to invent the eye; technological evolution compressed the cycle to millennia; AI now runs it in weeks.

Citing Marc Andreessen, he argues technology is the only perpetual source of growth — and goes further: there are no material problems that cannot be solved with more technology (psychological and political problems being a different beast). On that view, not accelerating is the bigger danger, because it means fewer people get to flourish.

His framing line: our generation was too late to explore the earth, too early to explore the stars, but right on time to build superintelligence. And this jump will be faster than previous ones — in 1903 no human had achieved sustained powered flight, and 60 years later we landed on the moon; today the accelerating technology lives in software, so it compounds faster still.

The chain is simple: more science → more technology → more growth → more human flourishing. Hence goal number one is accelerating scientific discovery — the subject of his book, The Eureka Machine, out in September.

The four pillars and the agent swarm (~00:13–00:16)

  1. Existing public knowledge, now largely encoded in LLMs. He addresses the copyright discomfort head-on: once these models are open source, they become a resource for humanity in the way the internet did.
  2. Scientific measurement data, because human senses are severely limited and the ceiling on perceptual intelligence is unlocked by instruments, not eyes.
  3. Simulation — "everything you can simulate, anything you can essentially verify, AI will obviously solve." He was never surprised that chess and Go fell: simulate, verify, generate unlimited training data.
  4. Robotic process automation of physical labs, for everything that can't yet be simulated.

Above the pillars sits an agent swarm that ideally carries the equivalent of ~5,000 distinct PhDs before you ever ask it to run a physical experiment.

Recursive's angle: if you're automating science, the best science to start with is the science of AI itself. That means letting the system build a form of self-awareness about its own shortcomings across the whole pipeline, from pre-training through harnesses.

Auto research, measured: three benchmarks (~00:16–00:19)

He draws a clean line first: asking one AI to improve some other system is auto research, not true recursive self-improvement — but it's a useful milestone on the way.

Science, and AI research, is three steps: ideation, implementation, validation. Implementation keeps improving, validation still costs time, and the live question is how many genuinely novel ideas the system can generate.

They pointed their RSI system at three benchmarks thousands of people had already worked on, often with their own agents:

Benchmark What it measures Result
nanochat Popular auto-research benchmark Bits-per-byte reduced significantly in under two days
nanoGPT speedrun Ever-faster training algorithms Beat prior records
SOL-ExecBench NVIDIA benchmark for CUDA kernels — the interface between GPUs and AI code Significant jump

The point he stresses: this is not hyperparameter tuning. The system produced real inventions — incorporating hash tables into transformers, novel forms of momentum. And it matters commercially: kernel efficiency drives intelligence per dollar, and squeezing 5–10% more out of a multi-billion-dollar data center is a very significant end result.

The ceiling: we've barely started (~00:19–00:21)

He tried to define AI properly, found no good definition, and settled on three principal components: prediction (mathematically equivalent to compression) × actions × goals, from which he derives ten "spaces of intelligence."

Take the simplest one, visual intelligence. Computer vision still mostly lives inside the electromagnetic band of the human eye — and human eyes aren't that good (mantis shrimp have better ones). The actual upper bound isn't binocular vision at all: millions or trillions of sensors, spread as far apart as the light cone still permits them to talk to one central intelligence, seeing down to quantum uncertainty and up to gravitational waves, then fusing all of it across many more layers of object abstraction.

So when people say AI won't go much further: we are literally astronomically far from the upper bounds in many spaces of intelligence — which is exactly what makes the area worth researching.

Quotes

"Our generation was too late to explore the earth, too early to explore the stars, but we're right on time to build superintelligence." (~00:12)

"Everything you can simulate, anything you can essentially verify, AI will obviously solve." (~00:14)

Why board games fell early, and why "simulatable / verifiable" is the load-bearing pillar.

"It's not just like tuning hyperparameters — it makes some truly useful inventions, like incorporating hash tables into transformers." (~00:18)

"We're literally astronomically far away from the upper bounds across many of the different spaces of intelligence." (~00:21)

A direct rebuttal to the "AI is plateauing" story.

提到的專案與資源 / Projects & Resources

名稱 Name 說明 Description 備註 Notes
Recursive (Recursive Superintelligence) 他的新公司,目標是自動化科學,並從 AI 自身的科學開始 His new company: automate science, starting with the science of AI itself 主持人介紹時提到募資 $650M / moderator cited a $650M raise
The Eureka Machine(書 / book) 九月出版,論述 AI 如何解鎖科學發現的新時代 Book out in September on AI unlocking a new era of scientific discovery 全名 The Eureka Machine: Why AI Is the Key to Unlocking a New Era of Scientific Discoveries
nanochat 受歡迎的 auto research benchmark,他們兩天內顯著壓低 bits-per-byte Popular auto-research benchmark; bits-per-byte cut significantly in under two days
nanoGPT speedrun 追求更快訓練演算法的 speedrun benchmark Speedrun benchmark for faster training algorithms
SOL-ExecBench NVIDIA 的 CUDA kernel benchmark,以硬體 Speed-of-Light 上限為標尺 NVIDIA's CUDA-kernel benchmark, scored against hardware speed-of-light bounds github.com/NVIDIA/SOL-ExecBench
MetaMind / You.com 他先前創辦的兩家公司(主持人介紹) His two earlier companies (from the moderator's intro)

逐字稿勘誤 / Transcript Corrections

字幕原文 Heard as 應為 Should be
Richard so / Soccer Richard Socher
recursive super intelligence(公司名) Recursive / Recursive Superintelligence
Metammind MetaMind
Mark Andre Marc Andreessen
neonets / recursive neonets recursive neural nets
soul execbench SOL-ExecBench
NanoGPD speedrun nanoGPT speedrun
nano chat nanochat
"AI is code and AI can't code" "AI is code and AI can code"
ideulate ideate
LMS / LM LLMs / LLM
scurves S-curves

待確認 / To Verify

  • Recursive 的 $650M 募資數字出自主持人 Igor Babuschkin 的介紹,講者本人未提;正式數字待查。/ The $650M raise came from the moderator's intro, not the speaker; confirm the official figure.
  • nanoGPT speedrun 是否即社群常稱的 modded-nanogpt speedrun,講者未指明版本。/ Whether this is the community's modded-nanogpt speedrun — the speaker didn't specify.
  • 三個 benchmark 的具體改善幅度講者未給數字(只說 "significantly" / "significant jump"),需看投影片。/ No numbers given for the improvements; check the slides.
  • 「十種智慧空間(10 different spaces of intelligence)」的完整清單只在投影片上,逐字稿未列。/ The full list of the ten spaces of intelligence is only on the slide.
  • 「把 hash table 併進 transformer」的具體做法與是否有公開發表,待查。/ What "incorporating hash tables into transformers" concretely means, and whether it's published.

Markdown source on GitHub ↗