Talk Session 4: Robotics & World Models
E2E Autonomy Without Imitation
Wei Zhan — Chief Scientist, Applied Intuition
Decouple learning to drive from learning to see — train a driving expert with large-scale self-play RL and zero human demonstrations (TerraZero), then distill its latents and actions into an end-to-end model (TerraTransfer) — and you reach state-of-the-art closed-loop end-to-end driving with no imitation anywhere in the recipe.
TL;DR
- The paradigm is shifting from AV 2.0 to AV 3.0. L2++ ADAS has converged on imitation-learning-based end-to-end stacks in mass production, with leading players adding open-loop RL post-training. The next generation is expected to train with closed-loop RL inside a world model that reactively generates surrounding behavior and visuals — open-loop scaling giving way to closed-loop scaling.
- Their approach splits the problem in two. Phase 1, "learn to drive," uses TerraZero, a self-play RL framework with far higher throughput than existing driving simulators, accumulating the equivalent of 25 centuries of driving experience with zero human demonstrations. Phase 2, "learn to see," uses TerraTransfer to align the self-play expert's latents and actions into an end-to-end model over the same driving cases from an offline dataset.
- The central claim: self-play closed-loop RL with a reactive world model is not merely a post-training technique but a powerful pre-training one — and the throughput of the RL framework and the world model is the decisive factor for this paradigm.
Key Points
Why "without imitation" (~03:02–03:04)
He framed the talk around end-to-end autonomy as the first massively productionized physical AI in the real world, and set out to argue something counterintuitive relative to the mainstream: enhancing that autonomy with reinforcement learning alone, without imitation.
Applied Intuition's positioning: a premier physical AI technology provider across verticals including cars, trucks, agriculture, mining, and construction, supplying the autonomy stack, OS, simulation tools, and broader physical AI infrastructure. A $15 billion valuation company with over a thousand engineers, also doing cutting-edge research on RL and world models with publications (including award-winning papers) at top venues.
Where the paradigm stands, and where it's heading:
- Today: the L2++ ADAS paradigm has converged on imitation-learning-based end-to-end systems in mass production, with some leading players adding open-loop RL post-training to push safety and robustness to the next level.
- Next: end-to-end autonomy is expected to be trained with closed-loop RL inside a world model that can reactively generate surrounding behavior and visuals. That's the shift from AV 2.0 open-loop scaling to AV 3.0 closed-loop scaling.
Applied has been building the high-throughput reactive world model that supports large-scale closed-loop RL for end-to-end autonomy, and the paradigm already performs well. But his question was whether there's an even smarter way — and their answer is to decouple learning to drive from learning to see.
Phase 1: learn to drive — TerraZero (~03:04–03:07)
TerraZero is their self-play RL framework, much faster and higher-throughput than other state-of-the-art driving simulators and self-play frameworks. Using only some public datasets for map diversity, it obtains the equivalent of 25 centuries of driving experience with zero human demonstrations — and that number can be scaled two to three orders of magnitude higher simply by scaling GPU compute and map diversity.
The results: policies trained by TerraZero achieve state of the art on various closed-loop planning benchmarks (vector-based), setting a clear edge over imitation-learning-based planners, especially on non-saturated, corner-case-only benchmarks such as InterPlan. The policy handles a range of challenging driving scenarios with desirable actions and shows zero-shot generalization across global cities.
(He connected back to the previous talk: Michael had just given a good example of self-play powering autonomous racing; his is an example of self-play powering autonomous urban driving.)
Phase 2: learn to see — TerraTransfer (~03:07–03:08:45)
TerraTransfer trains an end-to-end autonomy model taught by an expert trained from self-play — namely TerraZero. Concretely, they align both the latents and the actions of the two models on the same driving cases from an offline dataset.
The resulting end-to-end autonomy — with no imitation anywhere in its training recipe — obtains state-of-the-art driving performance on closed-loop end-to-end driving benchmarks, a clear edge over other imitation-based methods, handling a variety of intentionally constructed challenging scenarios with what he called surprisingly robust driving behavior.
Key takeaways:
- The autonomy paradigm is shifting toward closed-loop scaling (AV 3.0).
- Both end-to-end and vector-based planners can reach state of the art with a self-play / closed-loop RL training recipe alone — no imitation.
- Self-play closed-loop RL with a reactive world model is not just a post-training technique; it's a very powerful pre-training technique.
- The throughput of the RL framework and the world models can be the decisive factor for this closed-loop scaling paradigm.
Quotes
"End-to-end autonomy … is the first massively productionized physical AI in the real world." (~03:02:05)
In a session premised on robotics being pre-ChatGPT, this was the one talk about technology already in mass production.
"25 centuries of driving experience with zero human demonstrations." (~03:05)
The scale gap between self-play and road-test data, in one number.
"They are not just some post-training technique, actually they are very powerful pre-training techniques." (~03:08:20)
The perception he most wants to change: closed-loop RL isn't last-mile polish, it's the main training method.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| TerraZero | self-play RL 框架,零人類示範累積等效 25 個世紀駕駛經驗;閉環 planning benchmark SOTA | Self-play RL framework; 25 centuries of driving experience with zero human demonstrations; SOTA on closed-loop planning benchmarks | phase 1「learn to drive」/ the "learn to drive" phase |
| TerraTransfer | 把 self-play 專家的 latent 與動作對齊給端到端模型,訓練配方完全無 imitation | Aligns a self-play expert's latents and actions into an end-to-end model; no imitation in the recipe | phase 2「learn to see」/ the "learn to see" phase |
| InterPlan | 只收錄 corner case、尚未飽和的閉環 planning benchmark | Non-saturated, corner-case-only closed-loop planning benchmark | TerraZero 在此對 imitation-based planner 拉開差距 / where TerraZero's edge is clearest |
| Applied Intuition | 涵蓋車、卡車、農業、礦業、營建的 physical AI 技術供應商 | Physical AI technology provider across cars, trucks, agriculture, mining, construction | 估值 150 億美元、逾千名工程師 / $15B valuation, 1000+ engineers |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| Wei Jean / Ray / Way | Wei Zhan |
| terra zero / terror zero | TerraZero |
| terror transfer | TerraTransfer |
| interplan | InterPlan |
| cell play / selflay | self-play |
| close reinforcement learning / closable scaling / close super RL | closed-loop reinforcement learning / closed-loop scaling |
| Engine autonomy | End-to-end autonomy |
| ADA system | ADAS |
| local motion | locomotion |
| Acura and the CPR booth(panel 段) | ICRA and the CVPR booth(待確認 / to verify) |
待確認 / To Verify
- 主持人介紹他時提到「co-director of Berkeley DeepDrive」,官網議程僅列 Applied Intuition 職稱;此處 frontmatter 以議程為準。/ The host also introduced him as co-director of Berkeley DeepDrive; the official agenda lists only the Applied Intuition title, which is what frontmatter uses.
- 他提到的閉環端到端駕駛 benchmark 具體名稱,演講中未逐一報出(除 InterPlan 外)。/ Beyond InterPlan, the specific closed-loop end-to-end driving benchmarks were not individually named.
- panel 中他提到一家未具名公司的 R&D 車輛在擁擠街道連續一小時零接管、算力僅 Tesla HW4 的 1/7;無法查證。/ In the panel he cited an undisclosed company achieving one hour of intervention-free driving in dense streets on 1/7 the compute of Tesla HW4 — unverifiable.