Talk Session 1: AI Safety
The Human in the Loop: Navigating the Realities of AI for Employee Flourishing
Kathy Baxter — VP / Principal Architect, Responsible AI & Tech, Salesforce
Negative alignment — stopping models from causing harm — has driven the whole field of AI ethics, but it isn't sufficient for human flourishing; we also need positive alignment, actively designing systems that cultivate human judgment through optimized workflows, deliberate "mindful friction," and organizational respect for human craftsmanship.
TL;DR
- Reframe the question. Not "how does AI make us more efficient" but "how can we design a future where AI doesn't just automate our work but actively empowers human potential to flourish?" Research shows employees who use AI to draft their work can experience drops in their sense of control, responsibility, and creativity — yet AI can also inspire individuals and help creative workers reach more complex designs. The difference is deliberate design.
- Three actionable strategies. (1) Optimize workflow — the human steers strategy and creative direction while AI acts as autopilot for rote tasks. (2) Mindful friction — introduce cognitive forcing functions and staged reveals, e.g. porting the 1950s Delphi method into human–AI collaboration. (3) Sociotechnical resilience — deliberately invest in and celebrate human craftsmanship rather than handing everything over and getting AI slop and homogenization.
- The thesis: move from negative alignment (preventing harm) to positive alignment (actively fostering human virtue, wisdom, and well-being), and reinvest AI-saved time into human learning and craft rather than outsourcing thinking.
Key Points
Why frictionless is a trap (~00:20–00:22)
Baxter has been advocating mindful friction as a design strategy for over a decade: subtly nudging users to reflect on their choices. We habitually treat friction as a defect to be removed — but her observation is blunt:
Seamless, frictionless AI interfaces actively encourage cognitive surrender.
The fix is intentional friction: cognitive forcing functions and staged reveals. Her worked example is the Delphi method — a structured 1950s technique that requires individuals to develop their answers independently before sharing, precisely to defeat groupthink. Ported to human–AI collaboration, that means the human and the AI develop content independently, then bring it back to each other to share, iterate, and improve each round.
This cuts in two directions at once. It mitigates human cognitive surrender, and it reduces AI sycophancy — because the human isn't handing the model their idea and asking it to validate or riff on it, so the merge produces more challenged, more complex solutions.
From human in the loop to humans at the helm (~00:22–00:24)
The third strategy, sociotechnical resilience, is about a culture of proactive maintenance that sustains high performance — which requires deliberately investing in and celebrating human craftsmanship, because that is what actually drives innovation. Hand everything to AI and what you get is slop and homogenization.
Here she drew the talk's sharpest distinction:
- Human in the loop: humans arrive after the fact to touch it, validate it, rubber-stamp it.
- Humans at the helm: humans hold the wheel. Organizations should institutionally reserve a set of meaningful tasks that require human involvement.
The accompanying principle is equally concrete: time saved by AI should deepen people's own learning, their development of their craft, and every employee's unique expertise — not simply be recycled into more throughput. Her ask of leaders is the same shape: effective leaders don't hand AI to employees and say "do more, faster"; they guide employees toward understanding what success actually looks like and what the organization is really trying to achieve.
Negative alignment isn't enough (~00:24–00:25)
The close pulls the three strategies into a single claim:
- Negative alignment — preventing models from causing harm — has largely driven the field of ethical AI and remains incredibly important, but it is not sufficient for human flourishing.
- Positive alignment demands systems designed to actively foster human virtue: designing for wisdom and well-being, structuring workflows to protect human agency, designing with mindful friction, and fostering a culture that treats human craftsmanship as the key to innovation.
Her call to the room was to carry the framework back into whatever role they hold: treat AI as a force multiplier that amplifies human craftsmanship, and design AI to scaffold human learning rather than to absorb the thinking.
Quotes
"Seamless, frictionless AI interfaces actively encourage cognitive surrender." (~00:21)
Frictionlessness isn't neutral convenience; it has a direction, and the direction is away from thinking.
"This isn't just human in the loop where humans come in after the fact and touch it, validate it, rubber stamp it. This means humans at the helm." (~00:23)
The most direct attack on the industry's favourite governance slogan: a rubber stamp is not oversight.
"Instead of outsourcing thinking to AI, we should design AI to scaffold human learning." (~00:25)
The talk in one line.
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| Delphi method | 1950 年代結構化決策技術:先各自獨立產出答案再互相分享,以避免團體迷思 | 1950s structured technique: develop answers independently before sharing, to defeat groupthink | 講者提議移植到人機協作,同時抑制認知投降與 AI sycophancy |
| Mindful friction | 講者倡議超過十年的設計策略:刻意加入摩擦促使使用者反思 | Design strategy she has advocated for over a decade: deliberate friction that prompts reflection | 手法包含 cognitive forcing function、staged reveal |
| Positive alignment | 相對於 negative alignment(防止傷害),主張設計系統主動培養人的美德與能力 | Counterpart to negative alignment: design systems that actively foster human virtue and capability | 演講收束的核心主張 |
| Salesforce Office of Ethical and Humane Use | 講者所屬單位,Responsible AI & Tech 團隊隸屬於此 | The Salesforce office her Responsible AI & Tech team sits within | 講者自述職稱為 Principal Architect 兼 VP |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| deli method | Delphi method |
| AI syphency | AI sycophancy |
| wrote tasks | rote tasks |
| employees flourishing | employee flourishing |
| meta meta autonomy | meta-autonomy(待確認) |
待確認 / To Verify
- "institute human-centric lists for meaningful tasks":字幕如此,但 "lists" 一詞在語境中不通順,可能是 "human-centric SLAs"、"checklists" 或其他詞;需看投影片確認。/ The caption says "lists," which doesn't parse in context — could be SLAs, checklists, or something else; check the slides.
- "positive alignment and meta autonomy":結尾並列的第二個詞聽起來是 "meta autonomy",無法確定是既有術語還是她自創。/ The second term paired with positive alignment sounds like "meta-autonomy"; unclear whether it is an established term or her coinage.
- 關於「員工用 AI 起草工作會降低控制感、責任感與創造力」的研究出處:講者只說 "research has shown",未點名論文。/ She cited "research has shown" without naming the study.