Talk Session 3: Agentic AI Foundational Capabilities

Building Resilience for the Intelligence Age

Wojciech Zaremba — Co-Founder, OpenAI; Head of AI Resilience, OpenAI Foundation

Saturday, August 1 · Plenary Stage · 00:59:45–01:10:08 · afternoon stream

Fire didn't become safe because we banned it — it became safe through a layered ecosystem of detection, brigades, hydrants, materials, inspections and insurance; AI has no silver bullet either, aligning individual models is only one layer, and the real work is building the institutions that don't exist yet.

TL;DR

  • What AI resilience is: after OpenAI's restructuring, a nonprofit now owns roughly a quarter of OpenAI's equity, and one of the OpenAI Foundation's divisions is AI resilience. It is adjacent to AI safety but distinct — safety mostly asks whether the model is safe; resilience asks what the world must look like for AI to play out well.
  • Restriction has a poor historical record. "Curfew" comes from the French for extinguish the fire; in medieval times police would knock at your door at night to make you put it out. The Great Fire of London happened anyway. AI is similar: a few restrictions work (platforms broadly honor non-consensual intimate imagery bans), while the loudly advocated open-source restrictions were never really implemented.
  • Fire became safe through a layered ecosystem: early detection, trained brigades and fire trucks, purpose-designed hoses, city-wide water supply with hydrants at sufficient pressure and volume, materials shifting from wood to metal and concrete, inspections, insurance, designated exits. No single item was the answer. The result is remarkable — cities an order of magnitude denser than medieval ones, far more ignition sources (electricity, gas, industry, even data centers), and nobody worries about fire.
  • Three concrete AI analogues: for bio, harden the environment itself (e.g. sanitizing air so pathogens can't spread); for cyber, get software formally verified so superintelligence can't hack it; for safety incidents, build a public incident database like aviation's, with safe harbor for those who report.
  • This is a recruiting pitch. Some of it gets built by companies, some by nonprofits. He wants capable, reality-warping people to ask what the world still needs for AI to go well, and then go build those institutions — the OpenAI Foundation has substantial resources to back such efforts.

Key Points

Framing: AI resilience and the OpenAI Foundation (~00:59–01:01)

After several slides refused to advance, he deadpanned: "That's evidence that AGI is not yet here."

The substance: OpenAI restructured a number of months ago, and there is now a nonprofit that owns around one quarter of OpenAI's equity — a massive pool of resources. One of the OpenAI Foundation's divisions is AI resilience, and the talk exists to explain what that means: similar to AI safety, and yet different. He explains it through an analogy to fire resilience, flagging up front that fire and AI differ in many, many ways, so the analogies are by no means perfect.

Two similarities between fire and AI (~01:01–01:03)

First, both are general purpose technologies with enormously broad applicability. Fire heats food, provides warmth, smelts metal, powers engines — it is fundamental to civilization. Something similar is happening with AI: it is already part of how knowledge and wisdom get developed, part of the scientific process, and if you take seriously where robotics is going, it will simply be a fundamental part of the economic engine of the future.

Second, both come with risks. For fire: a number of major historical fires, London burned four times, Boston and Chicago too, roughly 80% of London consumed during the large fires. The difference is that we understand fire's risks pretty well today, and we don't understand AI's — ask people about AI risk and they point in different directions; open a newspaper and you'll find every fear on offer. His view: these risks are highly uncertain. Some may turn out to be paperwork-level nuisances, some may turn out 10x worse than expected, and some real risks may not even be on the current list.

The twist: safety through restriction didn't work (~01:03–01:04)

Here is the interesting turn. The earliest approach to fire safety was restriction. Curfew comes from the French and means extinguish the fire; in medieval times, if you had a fire going at home at night, police would come knock and tell you to put it out. That is plainly a reduction in the capability of the technology — and at least in this case, it didn't work out: the Great Fire of London happened despite curfews.

For AI it is more complicated. Some restrictions do seem to work well — most platforms follow non-consensual intimate imagery restrictions. Others, like the open-source restrictions various people advocated for, were never really implemented. His conclusion: in AI's case too, restriction hasn't worked out well.

How fire actually became safe: no silver bullet, a layered ecosystem (~01:04–01:07)

So what did it take? There was no silver bullet. It ended up being a multi-layer ecosystem:

  • Early detection systems for fires
  • Trained firefighter brigades, purpose-built fire trucks, specially designed hoses
  • A city-wide water supply with hydrants, at sufficient pressure and in sufficient quantity to actually extinguish a fire
  • Materials redesigned: wood replaced by metal and concrete
  • Inspections and insurance
  • Designated exits and egress routes

Once all of that was in place, fire stopped being something we worry about — we became able to fully harness its benefits. Did it work? Remarkably well. Cities today are an order of magnitude denser than medieval ones, with far more sources of fire — electricity, gas, industry, even data centers — and fire is not something people worry about.

The AI analogue: three ways to harden the environment (~01:07–01:09)

His observation is that people in the AI space keep looking for the silver bullet, the one solution that makes AI play out well — and he thinks that is the wrong way to think about it. It is also not enough to make individual models aligned; that is part of the solution, but the ecosystem needs several. Three examples:

  • Biosecurity: models are becoming quite capable in the biological domain, and many foresee risk from models making pandemics easy to manufacture. If that is the risk — and if we also assume advanced open-source models will diffuse — then guardrailing those capabilities inside frontier labs won't be sufficient. It should be done, but it won't be enough. We may need to harden the environment itself: people are thinking about ideas like sanitizing air, because if air is sanitized, pathogens can't spread.
  • Cybersecurity: perhaps we need to reach the point where software is formally verified, which would prevent it from being hacked by superintelligence.
  • Safety incidents: perhaps we need a public incident database, as aviation has — incidents publicly reported, with those who report gaining safe harbor by reporting.

Closing: a recruiting pitch (~01:09–01:10)

How does any of this get created? Some through founding companies, some through founding nonprofits. But fundamentally, what's needed is for a number of capable people — reality-warping people, the kind who set the reference points — to ask themselves what needs to be built in the world for AI to play out well, and then go after building those new institutions and organizations. The OpenAI Foundation, he notes, has tons of resources to support endeavors like that.

His last line hands the responsibility back to the audience: whether or not AI resilience succeeds depends on the people in this room.

Quotes

"That's an evidence that AGI is not yet here." (~01:00)

On the slides refusing to advance.

"The great fire of London happened despite of curfews." (~01:04)

Capping a technology's capability is not the same as making it safe — the pivot the whole talk rests on.

"It turns out that there wasn't a silver bullet. … It ends up being a multi-layer ecosystem approach." (~01:05)

The lesson from fire, and his thesis for AI.

"It might not be sufficient to guardrail these capabilities within the frontier labs. … We might need to harden the environment itself." (~01:08)

Moving from control the model to change the world — precisely what separates resilience from safety.

"Whether or not AI resilience will succeed depends on people in this room." (~01:10)

提到的專案與資源 / Projects & Resources

名稱 Name 說明 Description 備註 Notes
OpenAI Foundation OpenAI 重組後的非營利母體,持有約 1/4 OpenAI 股權;AI resilience 為其部門之一 The nonprofit resulting from OpenAI's restructuring, holding ~1/4 of OpenAI equity; AI resilience is one of its divisions 講者為該部門負責人 / the speaker heads it
空氣消毒 / Air sanitization 生物風險的「硬化環境」代表做法:空氣消毒使病原體無法傳播 Representative environment-hardening measure for bio risk: sanitized air prevents pathogen spread 講者舉例,未指名特定計畫 / cited as an idea, no specific project named
形式化驗證軟體 / Formally verified software 資安層面的長期解:可證明的軟體,超級智慧也駭不進去 The long-run cyber answer: provably correct software that superintelligence cannot hack 與 Dawn Song 同場稍早的 security-by-construction 主張呼應 / echoes Dawn Song's security-by-construction argument earlier in the session
航空業式公開事故資料庫 / Aviation-style public incident database 事故公開通報 + 通報者取得 safe harbor Public incident reporting with safe harbor for reporters 講者提議的制度設計 / proposed institutional design

逐字稿勘誤 / Transcript Corrections

字幕原文 Heard as 應為 Should be
Voych Zeremba / Voych Wojciech Zaremba
non-conensual intimate images non-consensual intimate images (NCII)
curfew ... uh like a 80% of London 語序為自動字幕斷句造成 / sentence breaks are an artifact of auto-captioning
asmtote(panel 段) asymptote

待確認 / To Verify

  • 「非營利持有約四分之一 OpenAI 股權」為講者口述的約略數字;公開報導的確切比例宜另行查證後補上。/ The "around one quarter" equity figure is the speaker's approximation; the exact publicly reported percentage should be confirmed and cited.
  • curfew 的字源:講者說法文原意是「extinguish fire」,一般辭源解釋為 couvre-feu(「覆蓋火」)。此處照講者原話記錄。/ Etymology: he said the French means "extinguish fire"; standard etymology gives couvre-feu, "cover the fire". Recorded as spoken.
  • 「倫敦被燒過四次」「大火吞噬約 80% 的倫敦」等歷史數字為講者口述,未附出處。/ The historical figures (London burned four times, ~80% consumed) were stated without a source.
  • AI resilience 部門的具體資助領域與金額(演講中未提),可由 OpenAI Foundation 官方公告補充。/ The division's specific funding areas and amounts weren't given in the talk; can be supplemented from OpenAI Foundation announcements.

Markdown source on GitHub ↗