演講 Session 1: AI Safety
Human in the Loop:在現實中航向讓員工蓬勃發展的 AI
Kathy Baxter — VP / Principal Architect, Responsible AI & Tech, Salesforce
防止 AI 造成傷害的 negative alignment 已經是整個 AI 倫理領域的主軸,但它不足以讓人蓬勃發展;我們還需要 positive alignment——用最佳化工作流、刻意設計的 mindful friction、以及對人類手藝的組織性尊重,主動培養人的判斷與能力。
TL;DR
- 問題重構:不要問「怎麼讓 AI 幫我們更有效率」,要問「怎麼設計一個 AI 不只自動化工作、而是主動放大人類潛能的未來」。研究顯示員工用 AI 起草工作時,對控制感、責任感與創造力的感受會下降;但 AI 也能激發個人、幫創意工作者做出更複雜的設計。差別在於刻意的設計策略。
- 三個可操作的策略:(1) optimize workflow——人掌舵策略與創意方向,AI 當 autopilot 處理例行工作;(2) mindful friction——刻意加入 cognitive forcing function 與 staged reveal,例如把 1950 年代的 Delphi method 搬到人機協作;(3) sociotechnical resilience——刻意投資並頌揚人類手藝,別把所有任務交出去換來 AI slop 與同質化。
- 最終主張:從 negative alignment(防止傷害)走到 positive alignment(主動培養人的美德、智慧與福祉);把 AI 省下的時間回投到人的學習與專業深化,而不是把思考外包出去。
重點整理
為什麼「無摩擦」是個陷阱(約 00:20–00:22)
Baxter 說 mindful friction 是她倡議超過十年的設計策略:透過微妙的推力促使使用者反思自己的選擇。我們通常把 friction 當成該被消滅的負面東西,但她的觀察是:
無縫、無摩擦的 AI 介面,實際上在主動鼓勵認知投降(cognitive surrender)。
修法是刻意設計摩擦:cognitive forcing function 或 staged reveal。她舉的具體做法是把 Delphi method(1950 年代的結構化技術,要求每個人先獨立產出答案再互相分享,以避免團體迷思)搬進人機協作——讓人和 AI 各自獨立產生內容,再拿回來互相分享、迭代、逐輪改進。
這個做法一次解決兩個方向的問題:既降低人類認知投降的風險,也減少 AI 的 sycophancy——因為人沒有先把自己的想法餵給 AI 再要求它背書或延伸,而是各自獨立發想,合流時能得到更有張力、更複雜的解法。
從 human in the loop 到 humans at the helm(約 00:22–00:24)
第三個策略是 sociotechnical resilience:鼓勵一種主動維護的文化來維持高績效,而這需要刻意投資並頌揚人類手藝(human craftsmanship),因為那才是真正推動創新的東西。單純把所有任務交給 AI,得到的是 AI slop 與同質化。
她在這裡做了一個明確的用詞區分,也是整場最尖銳的一句:
- human in the loop:人在事後進來,碰一下、驗證一下、蓋個橡皮圖章。
- humans at the helm:人握著舵。組織要制度性地保留一批必須由人參與的有意義任務。
配套的原則是:AI 省下來的時間,要用在加深自己的學習、深化自己的手藝、強化每位員工獨特的專業——而不是被回收成更多產出。她對領導者的要求同樣具體:有效的領導者不會把 AI 丟給員工說「做更多、做更快」,而是引導員工理解成功到底長什麼樣子、我們到底要達成什麼。
Negative alignment 不夠,還需要 positive alignment(約 00:24–00:25)
收尾把前面三個策略收束成一個理論主張:
- Negative alignment——防止模型造成傷害——長期以來驅動了整個 AI 倫理領域,而且極其重要,但對人的蓬勃發展並不充分。
- Positive alignment——要求我們設計出主動培養人類美德的系統:為智慧與福祉而設計,用工作流保護人的能動性,用 mindful friction 設計介面,並培養一種把人類手藝視為創新關鍵的文化。
她給全場的行動呼籲是把這套框架帶回自己的角色:把 AI 當成放大人類手藝的 force multiplier,而不是把思考外包出去的對象;要設計 AI 來搭建人的學習鷹架(scaffold human learning)。
金句
"Seamless, frictionless AI interfaces actively encourage cognitive surrender."(約 00:21)
無摩擦不是中性的便利,它有方向性——把人推向不再思考。
"This isn't just human in the loop where humans come in after the fact and touch it, validate it, rubber stamp it. This means humans at the helm."(約 00:23)
對「human in the loop」這個業界口號最直接的批評:蓋橡皮圖章不是治理。
"Instead of outsourcing thinking to AI, we should design AI to scaffold human learning."(約 00:25)
整場的一句話版本。
提到的專案與資源 / Projects & Resources
| 名稱 Name | 說明 | Description | 備註 Notes |
|---|---|---|---|
| Delphi method | 1950 年代結構化決策技術:先各自獨立產出答案再互相分享,以避免團體迷思 | 1950s structured technique: develop answers independently before sharing, to defeat groupthink | 講者提議移植到人機協作,同時抑制認知投降與 AI sycophancy |
| Mindful friction | 講者倡議超過十年的設計策略:刻意加入摩擦促使使用者反思 | Design strategy she has advocated for over a decade: deliberate friction that prompts reflection | 手法包含 cognitive forcing function、staged reveal |
| Positive alignment | 相對於 negative alignment(防止傷害),主張設計系統主動培養人的美德與能力 | Counterpart to negative alignment: design systems that actively foster human virtue and capability | 演講收束的核心主張 |
| Salesforce Office of Ethical and Humane Use | 講者所屬單位,Responsible AI & Tech 團隊隸屬於此 | The Salesforce office her Responsible AI & Tech team sits within | 講者自述職稱為 Principal Architect 兼 VP |
逐字稿勘誤 / Transcript Corrections
| 字幕原文 Heard as | 應為 Should be |
|---|---|
| deli method | Delphi method |
| AI syphency | AI sycophancy |
| wrote tasks | rote tasks |
| employees flourishing | employee flourishing |
| meta meta autonomy | meta-autonomy(待確認) |
待確認 / To Verify
- "institute human-centric lists for meaningful tasks":字幕如此,但 "lists" 一詞在語境中不通順,可能是 "human-centric SLAs"、"checklists" 或其他詞;需看投影片確認。/ The caption says "lists," which doesn't parse in context — could be SLAs, checklists, or something else; check the slides.
- "positive alignment and meta autonomy":結尾並列的第二個詞聽起來是 "meta autonomy",無法確定是既有術語還是她自創。/ The second term paired with positive alignment sounds like "meta-autonomy"; unclear whether it is an established term or her coinage.
- 關於「員工用 AI 起草工作會降低控制感、責任感與創造力」的研究出處:講者只說 "research has shown",未點名論文。/ She cited "research has shown" without naming the study.