理查·薩頓 (Richard Sutton)
Oak Lab 創辦人、亞伯達大學教授
關於
理查·薩頓(Richard Sutton)是現代強化學習之父,與他的博士指導教授 Andrew Barto 共同獲得 2024 年圖靈獎;2026 年起,他成為 Oak Lab 的創辦人——一間位於多倫多、追求讓智能體從原始經驗中持續學習的實驗室。他先讀心理學學士才轉入電腦科學,在四十年間打造了這個領域的概念核心:時序差分學習、策略梯度、Dyna 架構、options 框架,以及那本教會所有人強化學習的教科書。他 2019 年的短文〈苦澀的教訓〉——善用算力的通用方法,每一次都勝過人類知識工程——成為現代 AI 被引用最多的兩頁。2017 至 2023 年他任職於 DeepMind 亞伯達分部,之後與 John Carmack 在 Keen Technologies 共事,如今他主張 LLM 是一條岔路:智能來自經驗,而非模仿人類文字。
主要貢獻
- 打造強化學習的演算法根基——時序差分學習、策略梯度方法、Dyna 與 options 框架——與 Andrew Barto 共同獲得 2024 年圖靈獎
- 與 Barto 合著《Reinforcement Learning: An Introduction》(1998;2018),定義了兩個世代的領域教科書
- 寫下〈苦澀的教訓〉(2019):短短兩頁論證善用算力的通用方法終將獲勝——現代 AI 被引用最多的短文
- 提出亞伯達計畫與 OaK 架構——一條從持續的、運行時經驗通往超級智能的路線圖
- 離開 Keen Technologies 後,於 2026 年與 Khurram Javed 創辦 Oak Lab,以圖靈獎的信譽押注:凍結的 LLM 是一條死路
影片與訪談
Richard Sutton – Father of RL thinks LLMs are a dead end
The interview that detonated AI Twitter. Fresh off the Turing Award, Sutton tells Dwarkesh Patel that large language models cannot learn on the job, cannot predict the world, cannot be surprised — and therefore cannot be the path. Patel pushes back from inside the scaling worldview, and for an hour the two talk past each other in the most instructive way possible: what you are hearing is the field's deepest fault line, between intelligence as imitation of human text and intelligence as experience. Sutton's verdict afterward: "a frank exchange of views."
View Details
The OaK Architecture: A Vision of SuperIntelligence from Experience
The constructive half of Sutton's refusal. If LLMs are a dead end, what is the road? His RLC 2025 keynote lays out OaK — Options and Knowledge — an architecture where an agent builds its own abstractions from run-time experience: no pretraining corpus, no human-curated knowledge, everything learned one step at a time from the stream of life. This is the talk that became a lab: within a year of giving it, Sutton left Keen Technologies and founded Oak Lab to build exactly this.
View Details思想連結
傑夫·克盧恩
影響了他/她英屬哥倫比亞大學電腦科學教授
Clune 的 AI-GAs 宣言以 Sutton 開啟論證:「正如 Richard Sutton 最近指出的,歷史顯示,長期勝出的演算法往往是那些簡單、且能善用大量算力的。」〈苦澀的教訓〉說:別再手工編寫知識;Clune 把曲柄再轉了一圈——連學習演算法本身都別再手工設計。「別發明更快的馬」,正是苦澀的教訓推到極限的模樣。
arxiv.org · incompleteideas.net
楊立昆
思想對照首席 AI 科學家・Meta
這個領域最富盛名的兩位 LLM 懷疑者,他們的替代方案卻指向相反的方向。LeCun 著名的蛋糕比喻把強化學習放在頂端的櫻桃——觀察式學習的基座上的一枚點綴;Sutton 的整個生涯卻在說:櫻桃就是蛋糕,獎勵與經驗正是智能的來源。他們一致認為當前的範式是一條岔路,卻對「正路在哪」有著幾乎完美對稱的分歧。
Andrej Karpathy
思想對照AI 研究者與教育者
當 Sutton 告訴 Dwarkesh Patel「LLM 是一條死路」時,Karpathy 用一篇短文作答:我們不是在打造從世界中學習的動物,而是在召喚從人類文字中蒸餾出的幽靈——而幽靈,正是動物從未解決的冷啟動問題的務實解答。Sutton 持守純粹的立場;Karpathy 給出工程師的修正案。兩人之間,橫亙著這個領域最鮮活的問題:預訓練是地基,還是天花板?
karpathy.bearblog.dev
Gary Marcus
思想對照認知科學家、AI 批評者與作家
兩人都說 LLM 到不了真正的智能——但理由恰好互相抵銷。Marcus 的診斷是內建結構太少:心智需要純粹學習無法供給的先天機制。Sutton 的〈苦澀的教訓〉診斷的是相反的病:每一分手工打造的結構,都是向規模抵押的債,機制應該由經驗自己長出來。相同的結論讓相反的前提容易被忽略——而這正是這組對照值得一讀的原因。