Sunday, August 2
Atlas Stage
14 talks · 3 sessions
play_circle
Morning stream
play_circle
Afternoon stream
search
All
Keynote
Talk
Workshop
14 / 14 talks
AM
Session 1: Enterprise AI
00:00
Rate Limiter on AI Adoption Is Organizational
Sunita Verma
· Chief Technology Officer, Ironclad
Enterprise AI stalls on people and process, not on model capability — Ironclad's answer was to upskill every function simultaneously, treat education as infrastructure, and move spec-writing to the front of the process so agents get well-specified work.
Keynote
00:18
Self Optimizing Agents
Ori Goshen
· Co-Founder & CEO, AI21 Labs
Once agents reach production the binding constraint is operating economically at the frontier, and the configuration space — models, retrieval, scaling, model portfolios, execution strategies — is far too large to search by hand, so optimization has to be automated, observable, and robust to the next model release.
Talk
00:35
The Enterprise Version of the One-Person Unicorn
Rene Pajta
· Chief Architect Cloud & AI, Microsoft
The enterprise analogue of the one-person unicorn isn't headcount reduction — it's everyone operating like a CEO with their own bench of agents; what blocks it isn't models but company intelligence, which is the only real moat while models, harnesses, and connectors are commodities to buy.
Talk
00:46
The Art & Science of Benchmarking Agents
Vincent Sunn Chen
· VP & Founding Member, Snorkel AI
Our ability to measure AI has been outpaced by our ability to build it, and "bench slop" — vibe-coding a benchmark for a nice Twitter post — makes it worse; benchmarks that last need a thesis about the frontier, a roadmap the community can build on, and near-obsessive task quality control, which is how Snorkel built Senior SWE-Bench.
Talk
00:59
Superintelligence vs. Super-Reliability
Dan Klein
· Professor, UC Berkeley; Co-founder & CTO, Scaled Cognition
Intelligence is multifaceted and its facets are not advancing at the same rate — breadth, plasticity, and fluency are through the roof while precise verifiable control and explainability lag badly. Demos need only the former; shippable products need the latter, and that mismatch is the structural reason the last mile keeps failing.
Talk
01:12
From Assistants to AI Employees: Designing Agents That Own a Role, Not a Task
Anushka Pathak, Soham Shah, Eric Victorson
· Product Manager / ML Engineer / Software Engineer, Ema
Assistants are reactive, stateless, and leave governance to the user; Ema's argument is to design agents as "AI employees" — triggered by events and schedules, acting on real systems under scoped permissions, holding state across days and systems, with governance built in — and the real blocker in deployment is never whether the agent can do the whole job, it's the humans collaborating with it.
Workshop
PM
Session 2: Agent Evaluation & Benchmarks
00:00
Agent Arena: Causal Evaluations of Agents in the Real World
Anastasios N Angelopoulos
· Co-Founder/CEO, LLMArena
Static benchmarks get overfit and drift away from production, so the only judge that can't be gamed is reality — Arena randomizes the model inside millions of weekly organic agentic traces and reports causal treatment effects, not scores, as its agent leaderboard.
Talk
00:13
Ready for General Agents? Let's Test It.
Michal Shmueli-Scheuer
· Distinguished Engineer, AI Benchmarking and Evaluation, IBM Research
The bitter lesson says generality wins, agents included — but proving an agent is general first requires standardizing three interfaces (agent, environment, researcher). IBM's answer is a Unified Protocol mediation layer (Exgentic) that lets any agent run any benchmark unmodified, and the resulting leaderboard shows general agents are already competitive with the top domain-specific agent on each task.
Talk
00:27
Building Reliable Agents: An Evals-First Approach
Priya Ponnapalli
· SVP of Engineering, Enterprise AI, Scale AI
Enterprise agents use inherently stochastic technology to deliver deterministic business outcomes, so "mostly right" is a liability. Scale's answer is a four-layer eval framework — L0 business outcomes, L1 task success, L2 components, L3 diagnostics — kept separate but causally linked, with the eval suite treated as the customer's most durable asset.
Talk
00:39
Spec-Driven Agents: Hierarchical Specs, Tooling, and Trajectory-Based Evaluation
Srijith Rajamohan
· Head of AI Research, Redis
Lessons from building a diagnostic agent for the Redis Query Engine — what moved the needle wasn't a stronger model but restructuring the tools (setup vs. discretionary) and the knowledge (a diagnostic playbook of router plus handbooks); and most failure modes are invisible in the final answer and only show up in the trajectory.
Talk
PM
Session 3: AI for Math
00:50
The Future of AI for Long-Horizon and Sparse-Reward Tasks
Sergei Gukov
· Executive Director, American Institute of Mathematics; John D. MacArthur Professor of Theoretical Physics and Mathematics, Caltech
Math is the most honest benchmark for AI, and what actually blocks mass-producing solutions to hard math problems isn't knowledge but the combination of long horizons and sparse rewards. A decade of fixes — curiosity modules, world models, harnesses — buys a few x or 10x when these problems need orders of magnitude. The next era's bottleneck is not the data stack but the reward stack.
Keynote
01:05
The Unit Distance Conjecture and AI for Math
Lijie Chen
· Researcher, OpenAI
The Erdős unit distance conjecture from 1946 was disproved not by a math-specialized model but by a general reasoning model with no harness and a single prompt. Its edge wasn't depth — it was simultaneous fluency in two distant mathematical fields, which is exactly the most expensive thing for a human mathematician to acquire.
Talk
01:17
Semantic Alignment Models for Math and Software Engineering
Vijay Ganesh
· Professor, Georgia Institute of Technology
For autoformalization — a cross-modal task where semantic content must be preserved exactly — the win isn't scale but training the property of semantic alignment directly into the embedding space: objects with the same semantic content sit close together, different ones far apart. That recipe produced a model 1000× smaller than frontier models that beat them on every benchmark tested.
Talk
01:30
Omnigent - A Meta Harness for AI Agents
Aravind Segu
· Software Engineer, Databricks
Adopting agents is easy; operating them is not. Databricks' answer is an open-source layer above the harness — sessions live on a server, work happens on a runner — so one session can compose across harnesses, be co-driven by teammates, be governed by policy, and keep running after you close your laptop.
Workshop
中文版