The two days, distilled
Daily digest
What actually happened across the four stages, day by day: the themes that kept resurfacing, and the talks behind each of them. Every reference links to the full note.
Saturday · August 1
Day one walked the stack from the compute floor to the research frontier. The through-line: raw capability is no longer the bottleneck — the design of environments, verification and control is.
Building under the compute ceiling
DeSantis set the tone with constraint-driven innovation: power, silicon and supply chains shape the AI systems problem. From there the infrastructure track ran through accelerated computing and lab notebooks for agents, to Stoica's two fundamental gaps in agentic software engineering, photonic computing, and whether Kubernetes is the right substrate for agent-shaped workloads.
Software engineering, rewritten
The clearest consensus of the day: the human job has moved from writing code to designing the environment the agent works in. Steinberger runs 20-hour loops and 64 sub-agents; Lopopolo prompts agents as lazily as he'd prompt a principal engineer; Catasta wants to deprecate prompting entirely; on Nexus, Cognition shared RL lessons from training coding agents and Yutori argued computer-use models will agentify the web, not APIs.
Security: even the exam hall is attack surface
Dawn Song's keynote anchored the day's security thread: agent flexibility is attack surface, the OpenAI–Hugging Face sandbox escape showed evaluation infrastructure itself can be breached, and the way out is automated red-teaming plus provably-secure code. Zaremba reframed safety as fire-style resilience — an ecosystem, not a silver bullet — while the Nexus security session ran from runtime guardrails to supply-chain attacks and agents going rogue.
Robots & world models: simulation is the new infrastructure
Levine made the case for robot foundation models and Jim Fan sketched the endgame; Fidler traced simulation's three generations — hand-built art, NeRF reconstruction, and generative world models that went from five minutes per five-second clip to real-time on a consumer GPU in a single year. Sony AI's table-tennis robot and Waymo's physical-autonomy lessons pulled trustworthiness back into the physical world.
AI for science, many-pronged
The Nexus morning was a full session on scientific discovery: Buehler's multi-agent swarms and large reasoning models, Zou on the collective intelligence of agents, Princeton's LabOS AI-XR co-scientist that sees and works alongside humans, Rose Yu's agentic co-scientists, and Goodfire on learning science back out of superhuman AI.
Recursive self-improvement: excitement with brakes on
Day one treated recursive self-improvement with equal parts excitement and caution: the Song–Sekhon fireside demystified the "foom", Vinyals argued RSI is throttled by how hard evaluation and ideation really are, Tworek mapped the opportunities and failure modes of long-horizon agents, and Ng closed the day defending decades-away AGI definitions and open models.
Sunday · August 2
Day two turned to deployment: enterprise governance, evaluation methodology, and putting agents into finance, science and production systems. "Measure first, then grant autonomy" was the refrain.
Governance is the adoption bottleneck
Ironclad's CTO set the frame — the rate limiter on AI adoption is organizational, not technical. Google Cloud talked agent governance, HubSpot explained why off-the-shelf AI hit a wall, Credo AI proposed "earned autonomy" with the can/may/act ladder compiled into the harness, Dan Klein contrasted superintelligence with super-reliability, and Salesforce kept the human in the loop.
Evaluation becomes the battleground
Two stages each dedicated a full session to evals: Agent Arena's causal evaluations in the real world, benchmarking as an art and science, evals-first reliability and spec-driven agents on Atlas; then DigitalOcean's "preferences > benchmarks", the exam before enterprise deployment, ScarfBench's enterprise-Java migrations and Hex's "the points don't matter" on Compass. The shared verdict: benchmark scores are not production reliability.
Frontier research: personal agents and scientific discovery
Socher pitched the Eureka Machine — recursive superintelligence for science; Ed Chi laid out the future of personalized universal agents; Periodic Labs combined experiments, LLMs and theory to hunt quantum materials; Babuschkin argued personal AI needs continual learning. The math session stretched the horizon further, from sparse-reward long-horizon tasks to the unit distance conjecture.
Trust, deepfakes and the human
The Compass safety morning ran the trust gamut: Bregler on how agents with new tools can counter deepfakes and add context, trustworthy agents in regulated domains, ARIA's society of agents trusting at machine speed, Lawrence on viable systems and judgment, and continual learning versus safety in computer-use agents.
The agentic economy
Circle sketched payment rails for an economy where agents transact, Wells Fargo reimagined banking, Capital One shared the enterprise frontier, and the finance panel weighed what changes when money moves at agent speed — before twelve startups closed the summit in the spotlight.
Production-grade agent systems
The systems track got concrete: Nvidia's lessons from using agents to build production AI systems, Postman's road from demos to reliable infrastructure, Factory's software factory, SK hynix's rack-scale disaggregated serving with shared-memory KV cache — and Invoca's reminder that agentic AI is a UX problem disguised as a technology breakthrough.