Talk Session 3: Agentic AI in Finance & Healthcare
Agentic AI Applications for Mental Health: From Chatbots to Clinical Orchestration
Venkat Bhat — Associate Professor; Director, AI for Mental Health (AI-M) Program, University of Toronto
A psychiatrist and clinical trialist's position — a multi-agent system that performs well in simulation is only step one; what mental health actually lacks is randomized controlled trials, and worldwide there is still only about one (going on two) RCT of a generative-AI-based tool.
TL;DR
- His pipeline has three stages: pick a clinical workflow that might be automatable → build a single- or multi-agent system and validate it in a sandbox with a bot playing the end user → then run a two-to-three-year clinical trial. That third stage is what separates him from most AI speakers.
- Work spans four domains: clinical care (auto-generating PHQ/GAD-7 scales; a multi-agent DSM diagnostic interview), medical education (accelerating residents' competency attainment), research (partial automation of systematic reviews and meta-analyses), and quality improvement (estimating the therapeutic alliance from multimodal virtual-visit data).
- Two negative-but-important findings: (1) when forced to deliver motivational interviewing, agents fail past a certain complexity threshold and lose their guardrails; (2) low vs. high safety guardrails significantly change how the agent is perceived. And worldwide there is still only about one RCT (going on two) of a generative-AI-based tool — the evidence base is thin.
Key Points
Where he stands: a clinical trialist first (~02:08–02:11)
Bhat introduced himself as a psychiatrist and clinician scientist at the University of Toronto who leads the AI for Mental Health (AI-M) Program, and who serves as national mental-health community-of-practice lead within the Temerty Centre for AI Research and Education in Medicine (T-CAIREM) — a centre that brings universities across Canada together, with mental health as his slice of it.
His entry point into the field was that much of AI's development traces back to how the brain works: deep learning maps to cortical representation, reinforcement learning to subcortical. About a third of his program reverse-engineers these systems (interpretability work, not covered in this talk); the other two-thirds automates clinical workflows in mental health.
The standard paradigm: identify a workflow that might be automatable, build the automation (single agent or multi-agent), show in a sandboxed simulation — with a bot playing the end user — that it works, and then run clinical trials. He runs trials across depression, PTSD, anxiety (including on the prevention side), and more recently mental health and Alzheimer's.
His reason for insisting on trials is practical: it isn't enough for the agent to be non-inferior to a human on that workflow. Change management, service redesign, and other ecosystem changes have to happen too, so his group uses mixed methods to ask what adoption would actually require even for an efficacious tool — and folds in cost-benefit, which he called a critical factor.
The cornerstone slide from three or four years ago (~02:12)
As autonomy levels started climbing three or four years ago, his group asked what that would mean for clinical care in mental health, and landed on a taxonomy of collaborative, assistive, and semi-autonomous agents inside a human-in-the-loop model, with the human's role defined per tier. He returned to this slide repeatedly as the organizing frame.
His reality check: the AI scribe is the one generative-AI application clinicians identify with and actually use. Everything else in the taxonomy still needs a lot of evidence.
Four domains of implementation (~02:13–02:20)
Clinical care. They automated his own clinic end to end — from intake through discharge — with agents at each stage, each published, each now moving toward trials. Two examples:
- Measurement-based care. Mental health lacks the equivalent of blood sugar for diabetes or blood pressure for hypertension; clinicians use scales like the PHQ or GAD-7, and for over two decades it has been genuinely hard to get them completed. Their system takes unstructured data (think of the output of an ambient scribe) and generates the scales, and prompts clinicians when the underlying information hasn't been collected. This is moving into trials through Epic at their hospital sites.
- Multi-agent diagnostic interviewing. They took the operationalized DSM — the SCID instrument — and built a multi-agent system around it: one agent asks the questions, another oversees and won't let the interview advance to the next module prematurely, the system follows the SCID's skip logic, and an orchestrator/diagnoser sits on top. With a bot playing patients carrying all kinds of complicated DSM diagnoses, it performed well; it is now in clinical trials, deliberately positioned for mild-to-moderate symptoms — someone with severe psychosis, severe depression, or Alzheimer's is not going to be able to use it.
Education. He granted that deskilling is a legitimate concern that needs addressing in its own right, then framed the opposite question: training a clinician takes 10 years (four of medical school, four to six of residency) — can agentic AI accelerate the attainment of competencies? They decomposed the competencies residents are expected to reach by end of training, per the respective Canadian and US boards, then built a multi-agent system: a bot as the patient, a bot as the clinician, and a bot as the resident practicing interviewing — because clinical interviewing, the core skill, is learned by doing it repeatedly, presenting, and getting feedback. They are working with seven of the largest clinical training programs to deploy it in a two-arm design: current training vs. current training augmented, testing whether competencies can be reached in three or four years instead of five. A variant helps residents learn the basic principles of psychotherapy — CBT and motivational interviewing — by interviewing a simulated patient within that framework.
He stayed skeptical throughout: most of these things will not work the way we think they will, but the few that do will be transformative.
Research. After 5–10 years of doing systematic reviews and meta-analyses, he expects to stop doing them himself within a couple of years, replaced by living systematic reviews and meta-analyses. The approach is not end-to-end automation but specific points inside the pipeline — motivated by research assistants spending hours or weeks reconciling extraction errors across publications. They published on the quantitative side and repeated the exercise on the qualitative side (thematic analysis), comparing themes from conventional qualitative analysis against generative-AI-derived themes. The bottom line remains human-in-the-loop. The same idea is applied to data processing for EEG analysis.
Quality improvement. During the pandemic clinicians moved to virtual assessments, and his group wanted to characterize the patient–physician relationship in that setting. The approach was multimodal: several data streams and unstructured data feeding a prediction of the Working Alliance Inventory, the standard proxy for relationship quality. That work is gradually moving into clinical trials.
Two limiting findings, and the missing evidence (~02:20–02:22)
- A complexity ceiling. Studying whether agents can deliver therapy components, they forced an agent to conduct motivational interviewing and found that past a certain level of complexity it cannot perform the task and loses its guardrails.
- Guardrails shape perception. Comparing low-safety against high-safety guardrail conditions showed that guardrail strength has a significant effect on how the agent is perceived.
- The evidence gap. Surveying what the world is doing in this space, he closed on the point that despite everything these agents could in principle do, there is only one randomized controlled trial — now becoming two, with the bot from last year — of generative-AI-based tools. Clinical trials are the critical need, and that is where his team invests.
Quotes
"It's not enough if at that particular workflow the agent is non-inferior to the human doing the task. There are a lot of other ecosystem changes." (~02:13)
Non-inferiority is the entry ticket, not the adoption case.
"Most of these things are not going to work — at least in the way we think it would work — but the few things which would work will have a transformative effect." (~02:17)
A clinical trialist's prior on agentic AI.
"In spite of all the possibilities for what these agents could do, there's only one randomized control trial — now coming to two." (~02:21)
The talk's landing point: the capability narrative is far ahead of the evidence.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| AI for Mental Health (AI-M) Program | 講者在多倫多大學主持的計畫,聚焦心理健康臨床工作流程自動化 | Bhat's program at the University of Toronto, focused on automating mental-health clinical workflows | https://ai.psychiatry.utoronto.ca/ |
| T-CAIREM | Temerty Centre for AI Research and Education in Medicine,串連加拿大各大學;講者為全國心理健康 community of practice lead | Temerty Centre for AI Research and Education in Medicine at U of T; Bhat leads the national mental-health community of practice | 字幕誤聽為 "temporary center" / heard as "temporary center" |
| PHQ / GAD-7 | 憂鬱與焦慮的標準自評量表;團隊從非結構化資料自動生成分數 | Standard depression and anxiety rating scales; the team auto-generates scores from unstructured data | 正推進 Epic 內的臨床試驗 / moving into trials through Epic |
| SCID(operationalized DSM) | DSM 診斷訪談工具;拆成問診 agent + 監督 agent + orchestrator/diagnoser 的 multi-agent 系統 | The structured DSM diagnostic interview instrument, rebuilt as a multi-agent system (interviewer + overseer + orchestrator/diagnoser) | 字幕記為 "the skid";已進入試驗,定位輕到中度症狀 / in trials, scoped to mild-to-moderate |
| Working Alliance Inventory | 醫病關係品質的標準代理指標;團隊用多模態資料流估計 | Standard proxy for the therapeutic alliance; estimated from multimodal data streams | 虛擬看診品質改善研究 / virtual-visit quality-improvement study |
| Living systematic reviews & meta-analyses | 局部自動化的證據合成,非端到端;量化與質性(thematic analysis)兩端都做過 | Partially automated evidence synthesis (not end-to-end), on both quantitative and qualitative (thematic analysis) sides | 仍為 human-in-the-loop / still human-in-the-loop |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| Vin Katpot | Venkat Bhat |
| temporary center for AI in medicine | Temerty Centre for AI Research and Education in Medicine (T-CAIREM) |
| the skid / this kit | the SCID (Structured Clinical Interview for DSM) |
| GAD 7 | GAD-7 |
| multi- aent / aentic | multi-agent / agentic |
| deskkilling | deskilling |
| efiral(?) | 逐字稿雜訊,無對應詞 / transcript noise |
待確認 / To Verify
- 「there's only one randomized control trial, now coming to two, with the bot last year」——這個「the bot」指的是哪一個生成式 AI 治療聊天機器人的 RCT,名稱與出處待查證。/ Which generative-AI therapy chatbot RCT "the bot" refers to.
- 七個參與部署的臨床訓練專案名單未在演講中列出。/ The seven clinical training programs deploying the education system were not named.
- 演講提到的多篇論文(intake agents、SCID multi-agent、motivational interviewing 複雜度上限、guardrail 感知研究)均未給出標題,需另行對照 AI-M Program 出版清單。/ None of the cited papers were named on the transcript; cross-check against the AI-M Program publication list.
- 心理治療訓練 multi-agent 系統與診斷訪談系統是否為同一套系統的變體,講者敘述略過細節。/ Whether the psychotherapy-training system is a variant of the diagnostic-interview system.