Panel Session 1: Enterprise AI
Panel: Enterprise AI
Aaron Jacobson (moderator); Rao Surapaneni; Adarsh Hiremath; Duncan Lennox; Anahita Tafvizi; Surojit Chatterjee — Aaron Jacobson — GP, NEA;Rao Surapaneni — VP/GM AI Search & Specialized AI, Google Cloud;Adarsh Hiremath — Co-CEO, Mercor;Duncan Lennox — Chief Product & Technology Officer, HubSpot;Anahita Tafvizi — Chief Data & AI Officer, Snowflake;Surojit Chatterjee — Founder/CEO, Ema
The five agree that enterprise agent adoption is still around 10%, but converge on the view that the bottleneck is now organizational rather than technical — ROI has to be measured as workflow-level cycle-time compression rather than token consumption, trust is earned through verifiability rather than promises, and since the model layer is commoditizing, the architectural investment that matters is the layer that lets you swap models at will.
The panel
| Speaker | Org | Vantage point |
|---|---|---|
| Aaron Jacobson (moderator) | GP, NEA | VC investing across AI infrastructure, cybersecurity, developer tooling |
| Rao Surapaneni | Google Cloud | Hyperscaler and governance |
| Adarsh Hiremath | Mercor | Evaluation and data |
| Duncan Lennox | HubSpot | Building with AI and shipping it as SaaS |
| Anahita Tafvizi | Snowflake | Data platform, and customer zero for its own products |
| Surojit Chatterjee | Ema | AI employees for the Fortune 2000 |
Theme 1: The state of enterprise AI — Gartner says under 10% today, over 80% by 2030 (~01:06–01:11)
- Duncan Lennox: "Sounds rough and tumble right to me." The last year or two brought lots of pilots and experimentation, maybe one or two agents in production, but companies have struggled to scale in breadth and depth — because scaling takes far more than a very good model. He's optimistic about 80% by 2030: "you can feel that acceleration starting to take off."
- Rao Surapaneni: acknowledges a self-selected cohort, then makes the sharpest observation — organizations that have figured out their data governance and security posture move faster. The rest are stuck on how to train people to use it, or how to train internal compliance teams on where risk is acceptable. Adoption tracks risk tolerance.
- Anahita Tafvizi: a timeline. 2025 was the year of experimentation — does this work, what's the hallucination rate, can you trust the outcome. 2026 has been about proving ROI at scale, and the rest of the year is heading toward true business transformation. She agrees on direction but says the exact percentage — 70 or 90 — isn't relevant, and the number of agents is the wrong metric entirely; the question is the business value you're actually getting.
- Adarsh Hiremath: "So far enterprise AI adoption has been underwhelming, but that trend will change very very quickly." Two reasons. First, willingness: for a decade, automation and transformation have been pitched in different wrappers — BPO, task mining, process mining — leaving a graveyard of failed pilots, so executives are hesitant by default. Second, capability: the models simply weren't good enough until recently. The shift that matters is from models for chat to models as coworkers — managing many stakeholders and data sources and delivering economically valuable work output, rather than being a personal assistant.
- Surojit Chatterjee: first, define adoption. "There's a lot of token maxing, a lot of people using AI as personal agents. I don't think that should be considered enterprise adoption, because a lot of those projects are failing or low ROI." Looking at true end-to-end workflow automation, 10% is probably right — but growing much faster than he expected; the Fortune 2000 now has real appetite, and he thinks 80% could arrive before 2030. His strongest point: the blocker isn't technical anymore. "About 70% of the reason is organizational — change management, security, InfoSec, fear of the unknown."
Theme 2: What does success look like, and how is ROI measured? (~01:11–01:15)
- Anahita Tafvizi: there are only two forms of ROI — help customers make more money or spend less money. Reaching that bottom-line measurement is extremely difficult because improving one workflow takes time to propagate. So companies use leading indicators: it started with adoption and token consumption (she agrees that's a bad proxy — Surojit's token maxing), and increasingly it's whether the workflow got shorter. Three Snowflake examples:
- Quarterly reporting: two weeks → 48 hours
- Earnings prep: weeks → a couple of hours
- Customer prep: weeks → minutes
- Duncan Lennox: "Here comes the new boss, same as the old boss" — outcomes. Define the specific business outcome across the go-to-market journey: a new qualified lead at the top of the funnel, a closed-won deal, a support ticket successfully resolved. That clarity determines what agents you build and turns naturally into a metric — and if you can't measure it, you can't measure success at all.
(Anahita's mic failed mid-answer; the moderator went to Duncan and came back once it was fixed.)
Theme 3: Why do projects fail? Gartner says 50% of GenAI projects do (~01:15–01:18)
- Adarsh Hiremath: two things. First, confidence that the AI is performing as intended — most enterprises lack robust agentic evaluations (task, trajectory, context, artifact, verifier), so they don't know the agent does what they want before it hits production. Second, you can't hurl software over the fence — it has to come with change management and cultural change. That's why labs are spinning up their own deployment orgs with FTEs and consulting partners: doing a business transformation inside a culture that isn't AI-forward is genuinely hard. He expects end-to-end, full-stack deployment efforts to lift both adoption and pilot success rates.
- Surojit Chatterjee (a concrete case): a global professional services company with 250,000 employees across 65 countries deployed Ema's AI employees across the entire hire-to-retire cycle — sourcing, recruiting, onboarding, ITSM, payroll questions and payroll processing — roughly 70 use cases, delivered in three to four months. Results:
- HR ops team cut from ~1,000 people to ~500
- Employee question resolution from five days to five or ten seconds on average
- +20% employee satisfaction on employee-issue topics
The most interesting number is the counterintuitive one: even after an agent answered and closed the ticket, 30% of employees went and asked a human again. Surveys and focus groups found the reason — employees didn't trust that agents could actually do it and wanted human confirmation. His conclusion: the biggest change-management cost is educating employees to adopt the new system rather than falling back on the old one. (The customer is Wipro; the figure came from their CHRO.)
Theme 4: Is AI trustworthy? (~01:18–01:20)
- Rao Surapaneni: "Definitely trustworthy — and at the same time, trust but verify." His analogy: coding became amazing because you can verify programmatically, or even deploy another agent to verify. Wherever you can both generate and verify, adoption accelerates. His example: a customer whose goal is "my sales team only sells; everything before and after gets automated" — research, meeting-prep summaries, even a NotebookLM-generated podcast to listen to on the drive to the customer site, then automatic CRM updates afterward. He adds a point that often gets skipped: it's also about the person's own journey — enough demonstrated value that they're willing to take some risk and trust the content.
- Duncan Lennox: two sides. On one hand, models "as sophisticated as they are, are just weights and code" — everything the industry has built around trust, safety and security applies, including the notion that different outcomes require different levels of trust. On the other, he names an underestimated difficulty: previous transformations like cloud or mobile were easy to grok — you could wrap your head around what it was. This one is much harder. So when everyone says all roads lead back to high-quality evals — and he agrees 100% — the catch is that evals are not the new version of QA and test coverage. Writing good evals and knowing what a high-quality golden set looks like is a distinct skill people need to learn, and helping people through that change is the critical part.
Theme 5: What's actually working — use cases and architecture (~01:20–01:26)
On use cases, the consensus is well-defined workflows:
- Surojit Chatterjee: customer support; employee experience and employee support; finance operations (invoice processing, AP); data extraction — any document-heavy use case. Once inside a large enterprise, customers ask for adjacent areas next.
- Anahita Tafvizi: agrees on support (a highly well-defined workflow) and finance (month-close and quarter-close are prescribed processes, highly automatable), and adds go-to-market — sales, because they drive revenue and efficiency there lifts the top line, and marketing, because of the budget involved. It shows up across industries, verticals and functions.
- Rao Surapaneni: at one end, employee productivity — finding the right information at the right time, and crucially not just surfacing information but completing the action (from "what's the benefit policy on this" to actually enrolling), automated extensively with Gemini Enterprise. At the other end, CIOs going deep — automating every task a specific role performs, and now measuring the business outcome directly: are reps having more conversations, are they closing deals faster.
On architecture, the consensus is: don't bet on a single model.
- Surojit Chatterjee: "The model layer is going to be commoditized and open source models are getting better." Ema uses EmaFusion, blending 100+ models — frontier and open source — and he thinks some form of model fusion is becoming the standard architectural choice.
Anahita Tafvizi adds two pillars of trust — among the most operationally concrete material in the session:
- Accuracy. "If I ask an agent what was my Q1 revenue, there's only one correct answer — and this is not a trivial solution to a non-deterministic technology by nature." The right guardrails, evals and definitions have to guarantee the accurate answer 100% of the time, otherwise employees won't use agents instead of humans for basic questions. Her comparison is pointed: dashboards have that level of accuracy by default.
- Governance. Especially for data agents: when you open financial data, compensation data and performance data across the organization, how do you ensure each sales AE sees only their slice of revenue, and each manager sees their own reports but not their peers'? She calls both pillars non-negotiable for enterprise AI adoption.
Theme 6: Data and sovereign AI — should enterprises share data with model providers? (~01:26–01:31)
The moderator cites the Palantir CEO's recent argument that AI model companies pose a threat to most enterprises.
- Adarsh Hiremath: starts one level deeper — every enterprise is definitionally idiosyncratic. "Even for two companies in the same space — McKinsey and Bain — each of them have their own processes that make them special. If they were exactly one to one the same, then neither of them would exist." So the hardest part of any deployment is encoding the company's secret sauce — workflows, processes, organizational context — into the model weights or the harness. His example: everyone knows AI can review résumés, but you can't just deploy an agent to do it unless the agent knows the company's talent philosophy and how its hiring managers screen.
On what's safe to share: they help companies hand commodity workflows to the labs. For an IT function, say — if you want an agent that's excellent at traversing an IT stack, Mercor anonymizes the data, scrubs identity, and puts it into an RL environment useful to labs for training. For the average company that's typically a no-cost option, because the monetization can generate seven or eight figures for them. On open weights: there will be use cases for private models and for open-source models, "and the only way to figure out which one is right is eval." - Rao Surapaneni (asked to answer as a model provider — can you trust Google?): opens with a joke — "yeah, you can absolutely trust Google" — then makes the substantive point that enterprise data doesn't have to live in the model. You can supply it in prompts, in context; "the best context is the smallest context that gets the job done." So you can do a great deal inside a sovereign environment using a frontier model with full control, and contractually they don't see customers' data — "that's sacrosanct for us."
He also pushes back on open source as a cure-all: it eliminates the cost of training, but inference cost remains — and marginal cost is much higher than for CPU-era workloads. So an enterprise has to ask: am I willing to invest that capex, and the people and capability to grow it over time? Open source doesn't solve everything; think about the business and organizational-expertise dimensions too.
Theme 7: Open vs closed — who wins in five years? (~01:31–01:36)
- Duncan Lennox: what enterprise customers care about is trust and reliability, which as an engineering (or business) system means no single point of failure in the critical path. So expect a continued mix, with more sophisticated tooling to manage it.
- Anahita Tafvizi: Snowflake works with both, and the future is a combination. Frontier models have been the more innovative ones, a bit ahead, pushing boundaries. Open-source models give better cost, more competition (and therefore more innovation), plus better auditability and visibility into what's happening under the hood. Different workflows need different models — "if you're simply summarizing an email, you probably don't need the most advanced reasoning model," but some tasks do. So the future is offering both and dynamically routing queries — which is what Snowflake does, letting customers optimize at the query level for the right model for a specific use case.
- Surojit Chatterjee (closing): "I agree with Anahita. The strategic question isn't open source versus closed source — it's: can you change your model?" And: if you change it, do you have to change your entire stack? Will it fall apart? Do you have to code everything again? Build so you can swap models at will, or put a model blender / model selector underneath your agentic orchestrator, so you capture every innovation in the model world without waiting to retrain your entire agent layer. "However you implement it, I think that will be the core architecture going forward."
Quotes
"There's a lot of token maxing … I don't think that should be considered enterprise adoption, because a lot of those projects are failing or low ROI." (Surojit Chatterjee, ~01:10)
Pinning down what "adoption" means before arguing about the number — the most useful bit of conceptual hygiene in the session.
"It's not so much technical anymore … about 70% of the reason is organizational. It's change management, it's security, InfoSec, fear of the unknown." (Surojit Chatterjee, ~01:11)
The panel's most consistent judgment: the bottleneck has moved.
"If I ask an agent what was my Q1 revenue, there's only one correct answer — and this is not a trivial solution to a non-deterministic technology by nature." (Anahita Tafvizi, ~01:25)
The core difficulty of enterprise data agents in one line — with dashboards as the unforgiving baseline.
"30% of the requests, the employees actually ask a human again even though AI agents already answered or resolved the issue … apparently employees didn't trust agents can actually do it." (Surojit Chatterjee, ~01:18)
The capability arrived; the trust hadn't. Change-management cost, quantified.
"The best context is the smallest context that gets the job done." (Rao Surapaneni, ~01:28)
The most practical answer to sovereign-AI anxiety: the data doesn't have to enter the model.
"It's not open source versus closed source. The question is: can you change your model?" (Surojit Chatterjee, ~01:35)
Reframing the open/closed debate as an architecture question.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| EmaFusion | Ema 的模型融合層,調度 100+ 個前沿與開源模型 | Ema's model-fusion layer orchestrating 100+ frontier and open-source models | Surojit 認為 model fusion 正成為標準架構 / he calls model fusion an emerging standard |
| Gemini Enterprise | Google Cloud 企業 agent 平台,用於員工生產力與流程自動化 | Google Cloud's enterprise agent platform for productivity and workflow automation | 見 Rao 的演講筆記 / see Rao's talk notes |
| NotebookLM | Google 產品;客戶用它把會議準備資料轉成 podcast 給業務在車上聽 | Google product; a customer turns meeting-prep material into a podcast for reps to listen to en route | Rao 的信任案例 / from Rao's trust example |
| APEX(Mercor) | Mercor 前沿 benchmark,agentic eval 五件套的來源 | Mercor's frontier benchmark; source of the five-part agentic eval structure | 見 Adarsh 的演講筆記 / see Adarsh's talk notes |
| Snowflake 動態模型路由 | 讓客戶在 query 層級為特定用例挑選最適模型 | Query-level dynamic routing so customers pick the right model per use case | Anahita 描述的 Snowflake 產品方向 / Snowflake's direction as she describes it |
| Wipro | Surojit 舉的導入實例:25 萬員工、65 國、~70 個 HR/IT 用例 | The deployment case he cites: 250K employees, 65 countries, ~70 HR/IT use cases | 由該公司 CHRO 提供數據 / figures from their CHRO |
| Gartner 預測 | 今日 <10% 企業部署 agentic AI,2030 年 >80%;50% GenAI 專案失敗 | Gartner: <10% of businesses have agentic AI today, >80% by 2030; 50% of GenAI projects fail | 主持人引用作為討論起點 / cited by the moderator as the framing |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| Anahita Tavvisi / Anita | Anahita Tafvizi |
| Serji Chatterjet / Serjet / Sergey / Serg / so rigid | Surojit Chatterjee |
| Emma / Emma fusion | Ema / EmaFusion |
| Ralph / Ra | Rao Surapaneni |
| Dar / Adar / DAR | Adarsh Hiremath |
| Merkor / Ror | Mercor |
| Instaart | Instacart |
| Palenter / palunteer co | Palantir CEO |
| McKenzie / Bane | McKinsey / Bain |
| holination | hallucination |
| IOI | ROI |
| BO | BPO |
| infosc | InfoSec |
| well-definfined | well-defined |
| notebook LM | NotebookLM |
| geni projects | GenAI projects |
待確認 / To Verify
- Gartner「今日不到 10%、2030 年超過 80%」與「50% GenAI 專案失敗」兩則預測的原始報告出處。/ Source reports behind the two Gartner figures the moderator cites.
- Palantir 執行長「AI 模型公司對多數企業構成威脅」的原話出處與時間。/ Original source and date for the Palantir CEO's claim about model companies threatening enterprises.
- Surojit 提到 Ema 內部「40–50% 的工作使用開源模型」——主持人轉述,講者未直接確認數字。/ The "40–50% open source" figure was the moderator's paraphrase; the speaker did not confirm it directly.
- Wipro 案例的公開案例研究連結(員工數、縮編幅度、滿意度提升等數字)。/ A public case study for the Wipro deployment figures.