演講 Session 1: AI Safety

Human in the Loop:在現實中航向讓員工蓬勃發展的 AI

Kathy Baxter — VP / Principal Architect, Responsible AI & Tech, Salesforce

8 月 2 日(日) · Compass Stage · 00:16:06–00:25:10 · 上午場直播

防止 AI 造成傷害的 negative alignment 已經是整個 AI 倫理領域的主軸,但它不足以讓人蓬勃發展;我們還需要 positive alignment——用最佳化工作流、刻意設計的 mindful friction、以及對人類手藝的組織性尊重,主動培養人的判斷與能力。

TL;DR

  • 問題重構:不要問「怎麼讓 AI 幫我們更有效率」,要問「怎麼設計一個 AI 不只自動化工作、而是主動放大人類潛能的未來」。研究顯示員工用 AI 起草工作時,對控制感、責任感與創造力的感受會下降;但 AI 也能激發個人、幫創意工作者做出更複雜的設計。差別在於刻意的設計策略。
  • 三個可操作的策略:(1) optimize workflow——人掌舵策略與創意方向,AI 當 autopilot 處理例行工作;(2) mindful friction——刻意加入 cognitive forcing function 與 staged reveal,例如把 1950 年代的 Delphi method 搬到人機協作;(3) sociotechnical resilience——刻意投資並頌揚人類手藝,別把所有任務交出去換來 AI slop 與同質化。
  • 最終主張:從 negative alignment(防止傷害)走到 positive alignment(主動培養人的美德、智慧與福祉);把 AI 省下的時間回投到人的學習與專業深化,而不是把思考外包出去。

重點整理

為什麼「無摩擦」是個陷阱(約 00:20–00:22)

Baxter 說 mindful friction 是她倡議超過十年的設計策略:透過微妙的推力促使使用者反思自己的選擇。我們通常把 friction 當成該被消滅的負面東西,但她的觀察是:

無縫、無摩擦的 AI 介面,實際上在主動鼓勵認知投降(cognitive surrender)

修法是刻意設計摩擦:cognitive forcing functionstaged reveal。她舉的具體做法是把 Delphi method(1950 年代的結構化技術,要求每個人先獨立產出答案再互相分享,以避免團體迷思)搬進人機協作——讓人和 AI 各自獨立產生內容,再拿回來互相分享、迭代、逐輪改進

這個做法一次解決兩個方向的問題:既降低人類認知投降的風險,也減少 AI 的 sycophancy——因為人沒有先把自己的想法餵給 AI 再要求它背書或延伸,而是各自獨立發想,合流時能得到更有張力、更複雜的解法。

從 human in the loop 到 humans at the helm(約 00:22–00:24)

第三個策略是 sociotechnical resilience:鼓勵一種主動維護的文化來維持高績效,而這需要刻意投資並頌揚人類手藝(human craftsmanship),因為那才是真正推動創新的東西。單純把所有任務交給 AI,得到的是 AI slop 與同質化

她在這裡做了一個明確的用詞區分,也是整場最尖銳的一句:

  • human in the loop:人在事後進來,碰一下、驗證一下、蓋個橡皮圖章。
  • humans at the helm:人握著舵。組織要制度性地保留一批必須由人參與的有意義任務

配套的原則是:AI 省下來的時間,要用在加深自己的學習、深化自己的手藝、強化每位員工獨特的專業——而不是被回收成更多產出。她對領導者的要求同樣具體:有效的領導者不會把 AI 丟給員工說「做更多、做更快」,而是引導員工理解成功到底長什麼樣子、我們到底要達成什麼

Negative alignment 不夠,還需要 positive alignment(約 00:24–00:25)

收尾把前面三個策略收束成一個理論主張:

  • Negative alignment——防止模型造成傷害——長期以來驅動了整個 AI 倫理領域,而且極其重要,但對人的蓬勃發展並不充分
  • Positive alignment——要求我們設計出主動培養人類美德的系統:為智慧與福祉而設計,用工作流保護人的能動性,用 mindful friction 設計介面,並培養一種把人類手藝視為創新關鍵的文化。

她給全場的行動呼籲是把這套框架帶回自己的角色:把 AI 當成放大人類手藝的 force multiplier,而不是把思考外包出去的對象;要設計 AI 來搭建人的學習鷹架(scaffold human learning)。

金句

"Seamless, frictionless AI interfaces actively encourage cognitive surrender."(約 00:21)

無摩擦不是中性的便利,它有方向性——把人推向不再思考。

"This isn't just human in the loop where humans come in after the fact and touch it, validate it, rubber stamp it. This means humans at the helm."(約 00:23)

對「human in the loop」這個業界口號最直接的批評:蓋橡皮圖章不是治理。

"Instead of outsourcing thinking to AI, we should design AI to scaffold human learning."(約 00:25)

整場的一句話版本。

提到的專案與資源 / Projects & Resources

名稱 Name 說明 Description 備註 Notes
Delphi method 1950 年代結構化決策技術:先各自獨立產出答案再互相分享,以避免團體迷思 1950s structured technique: develop answers independently before sharing, to defeat groupthink 講者提議移植到人機協作,同時抑制認知投降與 AI sycophancy
Mindful friction 講者倡議超過十年的設計策略:刻意加入摩擦促使使用者反思 Design strategy she has advocated for over a decade: deliberate friction that prompts reflection 手法包含 cognitive forcing function、staged reveal
Positive alignment 相對於 negative alignment(防止傷害),主張設計系統主動培養人的美德與能力 Counterpart to negative alignment: design systems that actively foster human virtue and capability 演講收束的核心主張
Salesforce Office of Ethical and Humane Use 講者所屬單位,Responsible AI & Tech 團隊隸屬於此 The Salesforce office her Responsible AI & Tech team sits within 講者自述職稱為 Principal Architect 兼 VP

逐字稿勘誤 / Transcript Corrections

字幕原文 Heard as 應為 Should be
deli method Delphi method
AI syphency AI sycophancy
wrote tasks rote tasks
employees flourishing employee flourishing
meta meta autonomy meta-autonomy(待確認)

待確認 / To Verify

  • "institute human-centric lists for meaningful tasks":字幕如此,但 "lists" 一詞在語境中不通順,可能是 "human-centric SLAs"、"checklists" 或其他詞;需看投影片確認。/ The caption says "lists," which doesn't parse in context — could be SLAs, checklists, or something else; check the slides.
  • "positive alignment and meta autonomy":結尾並列的第二個詞聽起來是 "meta autonomy",無法確定是既有術語還是她自創。/ The second term paired with positive alignment sounds like "meta-autonomy"; unclear whether it is an established term or her coinage.
  • 關於「員工用 AI 起草工作會降低控制感、責任感與創造力」的研究出處:講者只說 "research has shown",未點名論文。/ She cited "research has shown" without naming the study.

GitHub 上的 Markdown 原始檔 ↗