Talk Session 1: AI Safety
A Society of Agents: Trust at Machine Speed, from Bits to Atoms
Alex Obadia — Programme Director, Advanced Research & Invention Agency (ARIA)
ARIA's £50M Scaling Trust programme funds the research and infrastructure that lets agents with *different owners* coordinate securely at machine speed without intermediaries — and insists the stack be an open-source public good, because monoculture would leave humanity less resilient to shocks and erode our agency over time.
TL;DR
- What ARIA is: the UK's DARPA equivalent — "the difference with DARPA is we don't have the D," so no defence work, everything else. Three years old, around 16 programmes at roughly £50–70M each, across manufacturing, better chips, vaccines targeting different parts of the immune system, and neurotechnologies.
- What Scaling Trust funds: research and infrastructure that lets agents coordinate, meaning three concrete capabilities — find counterparties to trade or interact with, negotiate in sophisticated ways, and enforce agreements. The setting is deliberately adversarial multi-agent, multi-principal: not many agents under one owner, but agents with different owners representing different principals.
- The design constraints: programmatic (it must run at machine speed, so no humans in the loop at scale), available to everyone, without intermediaries — an explicit claim about system topology, aimed at limiting systemic choke points — and spanning digital and physical, since agents can be embodied. The technical bet is on programmable cryptography and recent advances in secure hardware.
- What only agents can do: two agents can enter a "room" backed by secure hardware, disclose information to one another, and commit to deleting it from memory if the deal falls through — something humans can only approximate with NDAs.
- Why open source and plurality: to avoid sliding into monoculture. The argument isn't diversity for its own sake — monoculture makes humanity less resilient to shocks, culturally and technologically, and erodes our agency over time.
- Three programme layers: the coordination stack (digital and physical, with first teams already funded including at Berkeley), the fundamental theory underpinning it, and a large public test bed run as a competition, red teams included.
Key Points
ARIA and the shape of Scaling Trust (~00:35–00:37)
ARIA — the Advanced Research + Invention Agency — is a UK R&D funding agency that funds R&D moonshots worldwide. The DARPA comparison is the fastest way in, with one difference:
The difference with DARPA is we don't have the D. So we don't do anything defence related, but we do everything else.
Impact-focused, high risk appetite, focused programmes chasing an ambitious goal, each around £50–70M. Three years old, roughly 16 programmes so far, spanning manufacturing, better chips, vaccines built on different parts of the immune system, and neurotechnologies. Obadia is the Programme Director for Scaling Trust.
He was explicit that the talk had two purposes: describe what ARIA is funding and already funds, and get the room's opinion on the current thesis — he offered before Q&A that he was happy to hear any scepticism.
Scaling Trust is currently a £50M, three-year programme. It funds the research and infrastructure that lets agents coordinate, where coordination decomposes into three capabilities:
- Can agents find the counterparties they need to trade or otherwise interact with?
- Can they negotiate in sophisticated ways?
- Can they enforce agreements?
Doing this securely matters because the setting is multi-agent, multi-principal systems under adversarial conditions — and here he drew a firm boundary: not single-owner multi-agent systems, but agents with different owners representing different principals. Some of those agents may be adversarial; there may be information asymmetry; goals may differ.
Four design constraints (~00:38–00:39)
- Programmatic. Because this has to move at machine speed, they explicitly do not want humans in the loop at scale.
- Available to everyone, without intermediaries. He framed this as a claim about the topology of the future system: limit the systemic choke points that exist in the infrastructure.
- Digital and physical. Agents can be embodied, so the infrastructure must extend into the physical world and support secure interactions that include physical elements.
- The technical bet: new developments in cryptography and secure hardware — particularly the trend dubbed programmable cryptography — which he argued break some of the tensions raised in the previous talk (Gondara's), or at least shift the trade-off space so different points in it become reachable.
Why it's worth doing: new markets, and agent-only capabilities (~00:39–00:40)
He conceded the state of play first: today's agent stack is not mature, and it is very hard to send agents off to do things for us; earlier talks had already walked through taxonomies of potential attacks.
Two things excite him:
New markets. The analogy is cryptography in modern digital society — it indirectly enabled things like e-commerce to exist, and today's digital society stands on those building blocks. Extend coordination infrastructure into the physical world and new markets that reach into physical space become plausible.
Things agents can do that humans cannot. His example was concrete:
Two agents can enter a quote-unquote room using secure hardware. They can disclose information to one another and commit to deleting the information from their memory if the deal doesn't go through. These are things that as humans we can only approximate with, for example, NDAs — that with agents we can do programmatically.
Open source, plurality, and the case against monoculture (~00:40–00:41)
Three commitments: open source, open to all, no intermediaries. The aim is a system where many models and many minds can coexist — the concept people usually discuss under pluralism or plurality.
What he wants to avoid is a slide towards monoculture, where everyone runs the same stack and, indirectly, our preferences get shaped the same way and society loses diversity. He was careful that this is not an aesthetic preference:
We think it makes us as humanity less resilient to shocks, both culturally and technologically, and erodes our agency over time.
That is why they believe this should be built as a public good, and why ARIA is glad to fund it through grants.
The programme's three layers (~00:41–00:44)
Layer 1: the coordination stack. Digital and physical, with a first set of teams already funded:
- On the digital side, including teams at Berkeley, working on how agents reason about security, autonomously generate protocols, interact with sophisticated security systems, and reason about negotiation in sophisticated ways.
- On the physical side, for example new sensors that are easier for agents to interact with and tamper-resistant — and people are already trying to break them.
Layer 2: the theory that underpins the stack. His diagnosis was direct:
A lot of the research that we see feels very empirical to us, at least if you come from theoretical computer science and these fields of security.
So they want the stack — which may well be arrived at empirically — to have theory underneath it: applying formal security thinking to a field of AI that currently resists being fully formalized. The directions he named: autonomous protocol generation, formal AI security, and carrying cryptographic concepts into the physical world (he pointed at quantum cryptography and physically unclonable functions as attempts to swap the roots of trust used in traditional cryptography into physical substrates).
Layer 3: a massive public test bed. Even with a stack and with theory assuring you that "there's no looming impossibility result and we're not working completely in the dark," it still has to be tested, because this is ultimately a multi-agent, multi-principal system and the stack needs stress.
- Anyone can submit their agents. They can test the funded stack, or bring their own.
- Run as a competition: teams compete to win, and red teams compete to be the best red teams.
- What it measures: the state of the art in secure agentic coordination under adversarial pressure.
- What else it looks for: what emerges. "We as humans can decide what stack is needed, but as we've seen many times, emergent behaviours sometimes reveal new things that agents can uniquely do" — so both the test bed and its tasks must be open-ended enough.
- Most relevant to this safety track: understanding what failure modes exist. Much of this behaviour will be out of distribution, and understanding the failure modes informs what safety requirements — and perhaps regulation — are needed.
Programme status: several teams funded, and a call in partnership with Google DeepMind, the Cooperative AI Foundation, and Schmidt Sciences closing on 8 August, funding among other things the sandboxes and test beds described.
Quotes
"The difference with DARPA is we don't have the D." (~00:36)
ARIA's positioning in one line.
"We don't care about the single owner multi-agent systems. This is about different owner, different agents that are representing different principals." (~00:37)
The programme's defining problem framing: the hard case is agents with different masters.
"Two agents can enter a quote-unquote room using secure hardware … they can commit to deleting the information from their memory if the deal doesn't go through. These are things that as humans we can only approximate with, for example, NDAs." (~00:40)
Not "agents do it faster than humans" but "agents do what humans cannot."
"We want to avoid a slide towards monoculture … it makes us as humanity less resilient to shocks both culturally and technologically, and erodes our agency over time." (~00:41)
Plurality reframed from a values claim into a resilience argument.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| ARIA (Advanced Research + Invention Agency) | 英國 R&D 資助機構,資助全球登月型研究;約 16 個計畫,每個 £50–70M | UK R&D funding agency backing moonshots worldwide; ~16 programmes at £50–70M each | 成立三年;不做國防 / three years old, no defence work |
| Scaling Trust | £50M、三年期計畫,資助 agent 安全協調的研究與基礎設施 | £50M three-year programme funding research and infrastructure for secure agent coordination | 官網另有 Scaling Trust Arena(競賽平台)之名,演講中未點名 |
| 開放測試場 / open test bed | 開放公眾提交 agent 的大型競賽式測試場,紅隊同場競技 | Public competition-style test bed; anyone can submit agents, red teams compete too | 演講中稱 "a massive test bed that is open to the public" |
| Programmable cryptography | 講者押注的密碼學趨勢,與安全硬體共同鬆動安全 / 自主的 trade-off | The cryptography trend he's betting on, which with secure hardware shifts the security/autonomy trade-off | 明確回應前一場 Gondara 提到的性質衝突 |
| Physically unclonable functions | 把信任根從傳統密碼學搬到實體世界的方向之一 | One route to swapping cryptographic roots of trust into physical substrates | 與量子密碼學並列提及 |
| 合作徵案 / joint call | 與 Google DeepMind、Cooperative AI Foundation、Schmidt Sciences 合作,8 月 8 日截止 | Joint call with Google DeepMind, Cooperative AI Foundation, Schmidt Sciences; closes 8 August | 資助沙盒與測試場等 |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| Alex Oadia | Alex Obadia |
| Arya | ARIA |
| Deep Mind | Google DeepMind |
| "50 million pound … so about $7 million" | £50M 約合 $67M(字幕金額錯誤)/ caption figure is wrong |
| texonomy | taxonomy |
| erodess | erodes |
| physically inclinable functions | physically unclonable functions |
| multi- aent | multi-agent |
| principles(指 multi-principal 那段) | principals |
待確認 / To Verify
- £50M 的美元換算:講者說「50 million pound, so about $7 million」——£50M 約合 $67M,字幕或口誤其一;金額以 £50M 為準。/ £50M ≈ $67M, not $7M; take the sterling figure as authoritative.
- 測試場的正式名稱:演講中只稱 "a massive test bed";ARIA 官方資料另有 Scaling Trust Arena 的名稱,但講者未在台上使用,故不併入內文。/ He only said "test bed"; ARIA materials refer to a Scaling Trust Arena, but he did not use the name on stage.
- 已獲資助的 Berkeley 團隊名稱與 PI:講者只說 "some teams at Berkeley",未點名。/ He said "some teams at Berkeley" without naming them.
- ARIA 計畫總數:講者說 "about 16 programmes so far",為口頭近似值。/ "About 16" was a spoken approximation.