Talk Session 1: AI for Science
LabOS: The AI-XR Co-Scientist That Sees and Works With Humans
Mengdi Wang — Professor, Princeton
The real bottleneck on scaling AI isn't hypotheses, it's verification — physical experiments have no checkpoints, no logs, and no way to backtrack — so Wang's answer is to put AI behind smart glasses in the lab and turn every physical action into an observable, debuggable environment.
Note: the talk was cut short by the host for time; the final slide was rushed.
TL;DR
- The bottleneck is verification, not hypothesis generation. That's why so many AI companies are racing to build trace data and RL environments — AI scales much faster if you can scale environments and verification. In science, that is brutally hard.
- The reproducibility numbers are ugly: a Nature survey found 70% of biomedical papers aren't reproducible by others and 50% aren't reproducible by their own authors, and it holds across chemistry, biology, physics, and earth/environmental science. The reason is structural: physical experiments have no checkpoints (model training has them, agent workflows have them, a clean room doesn't), no logs to debug, and no way to go back and improve the harness.
- AI is making science slower, not faster: she cites a colleague's blog — 40,000 submissions at the most recent NeurIPS, an explosion of papers and agents, but verification hasn't sped up, so the signal-to-noise ratio degrades and useful information gets harder to find amid AI-generated content.
- LabOS's bet: put a multimodal reasoning AI behind smart glasses so that every action and state change in a lab becomes observable to AI — catching errors in real time, digitalizing physical workflows, troubleshooting, offering hints. Paired with an end-to-end robotic nanofabrication lab that compressed three months of PhD-student work into one week and is being released as an API.
Key Points
Verification, not hypothesis, is what limits AI (~01:03–01:07)
Earlier speakers covered the advances in AI agents in the digital space. Wang breaks scientific research into stages: hypothesize (getting models to think, reason, dig deep), computation (simulation and surrogate models to find the best candidates), and finally validation / verification.
Her claim is blunt: verification has become the major bottleneck for scaling any AI model. That's precisely why so many AI companies and startups are actively building new trace data and new RL environments. AI would scale much faster if there were a way to scale environments and verification — and in science this is extraordinarily hard.
She put a question to the room: pick a random chemistry or biology paper in Nature — what fraction is reproducible? The audience guessed 5%, then 30%. The answer, from a survey run by researchers and Nature: 70% of biomedical research papers are not reproducible by others, and 50% are not even reproducible by the same authors — and this holds across chemistry, biology, physics, and earth and environmental science.
The cause isn't effort, it's the structure of verification. When someone runs an experiment at a bench or in a clean room:
- There are no checkpoints. Model training has them; agent workflows can backtrack; physical experimentation has neither.
- There are no logs to inspect, nothing to debug, no way to improve the harness.
A colleague wrote a blog a few months earlier concluding that with AI, science is not faster — science is getting slower. The most recent NeurIPS cycle drew 40,000 submissions. There are simply too many papers, while verification — which is what actually supplies new knowledge and new information — hasn't sped up. More papers and more agents therefore mean a worse signal-to-noise ratio, and it becomes harder to extract anything useful from the volume of AI-generated content.
Back to verification: we can have all the fancy models telling us millions of novel hypotheses, but the bottleneck will be in the lab. Her experimental colleagues devote their careers and years of work to physical laboratories, and even at their best the workflow remains highly error-prone.
Hence the question that frames the talk: how do we turn every scientific lab into a verifiable environment?
On one side AI is advancing fast, with AI co-scientists from every major frontier lab and startup. On the other, scientists at the bench and in clean rooms spend months to years validating a single hypothesis. Something critical is missing in between.
LabOS: making the physical lab observable (~01:07–01:11)
LabOS is an AI-XR agent: a multimodal reasoning AI hiding behind smart glasses. Her first example was an undergraduate intern visiting from India who, with LabOS's help, could perform an advanced genome engineering experiment essentially on day one.
What the system provides: every single action and every state change in a scientific lab becomes observable by AI, alongside multi-tier streaming so that AI can assist human researchers in real time — catching errors, digitalizing physical workflows, troubleshooting, and offering guidance and hints when something doesn't work.
Her second example — and she stressed it was not an animation — is an end-to-end mini robotic lab built with colleagues at the Princeton Quantum Institute, automating the nanofabrication of one-atom-thin graphene devices. A multimodal agent runs auto-research inside the computer while the robot does the measurements, fabrication, tape-outs, and microscopic imaging. The result: work that used to take physics PhD students three months now takes one week. Her colleague is releasing the system as an API, so anyone can submit a job to reproduce an experiment or test a new hypothesis — the goal being to enable this at scale.
(Her final slide, after the host cut in for time.) They are piloting LabOS as a system where human researchers work side by side with robots: every physical workflow gets digitalized, every trace collected and reasoned over by the AI. It's designed to generalize across scientific domains — biology labs, chemistry labs, clean rooms, and nanofabrication facilities.
Quotes
"70% of biomedical research papers are not reproducible by others, and 50% are not even reproducible by the same authors." (~01:05:21)
Not an attitude problem — the predictable outcome of having no infrastructure for verification.
"With AI, science is not faster. Science is actually getting slower." (~01:06:28)
Accelerate hypothesis generation without accelerating verification and the signal-to-noise ratio collapses.
"How do we turn every scientific lab into a verifiable environment?" (~01:07)
The problem statement for the whole talk.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| LabOS | AI-XR co-scientist:智慧眼鏡背後的多模態推理 AI,即時觀測並輔助實體實驗 | AI-XR co-scientist: multimodal reasoning AI behind smart glasses that observes and assists physical experiments in real time | 對應論文 arXiv:2510.14861 / bioRxiv 2025.10.16.679418;Stanford–Princeton 合作(Le Cong × Mengdi Wang)/ Stanford–Princeton collaboration |
| 迷你機器人奈米製造實驗室 / mini robotic nanofab lab | 自動化單原子層 graphene 元件製造:量測、製造、tape-out、顯微影像全由機器人執行 | Automates one-atom-thin graphene device fabrication end to end — measurement, fabrication, tape-out, microscopy | 與 Princeton Quantum Institute 同事合作;三個月 → 一週;將開放為 API / with Princeton Quantum Institute colleagues; 3 months → 1 week; being released as an API |
| Nature 可重現性調查 / Nature reproducibility survey | 70% 生醫論文他人無法重現、50% 原作者也無法重現 | 70% of biomedical papers not reproducible by others; 50% not reproducible by their own authors | 演講中僅稱「a survey run by researchers and Nature」/ described only as "a survey run by researchers and Nature" |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| Mangdi Wang | Mengdi Wang |
| lab OS | LabOS |
| Europe submission | NeurIPS submissions |
| clean rooms(字幕正確) | clean rooms |
| nanop fabrication | nanofabrication |
| graphing devices | graphene devices |
| purity students | PhD students |
| AI co-cientists | AI co-scientists |
待確認 / To Verify
- 實習生姓名:字幕作 "Simmeran",拼寫未確認。/ The intern's name, transcribed as "Simmeran", is unverified.
- 同事的部落格(主張「有了 AI 科學反而變慢」)作者與連結未提及。/ The colleague's blog arguing science is getting slower with AI was neither named nor linked.
- NeurIPS 4 萬份投稿的年度未指明(字幕作 "the most recent Europe submission")。/ The year of the 40,000-submission NeurIPS cycle wasn't specified.
- Nature 可重現性調查的年份與正式出處未提供(公開常引的是 2016 年 Nature 的 1,576 人調查)。/ The year and citation for the Nature reproducibility survey weren't given (the commonly cited one is Nature's 2016 survey of 1,576 researchers).
- 開放為 API 的那位 Princeton 同事姓名與服務名稱未提及。/ The Princeton colleague releasing the API, and the service's name, weren't mentioned.