Chris Olah
Co-founder & Interpretability Lead, Anthropic
關於
Chris Olah (b. 1992) is a Canadian machine-learning researcher and one of the pioneers of mechanistic interpretability — the effort to reverse-engineer neural networks by mapping the internal features and circuits that produce their behavior. A Thiel Fellow who left the University of Toronto after a year, he led interpretability research at Google Brain and OpenAI before co-founding Anthropic in 2021, where he leads the interpretability team. He co-founded Distill, a journal devoted to clear scientific communication, and his visual, essayistic explanations of neural networks shaped how a generation understands them. His work turns 'do we understand what we built?' from a rhetorical worry into an experimental science.
主要貢獻
- Pioneered mechanistic interpretability — reverse-engineering the features and circuits inside neural networks
- Co-founded Anthropic (2021) and leads its interpretability research
- Founded Distill, setting a new standard for clear, interactive scientific communication in ML
- Produced foundational work on feature visualization, circuits, and superposition in neural networks
- Named to the TIME100 AI list (2024) for advancing the science of understanding AI systems
影片與訪談
Mechanistic Interpretability explained | Chris Olah and Lex Fridman
Olah explains what it means to find features and circuits inside a trained model.
View Details
Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
Olah's 2023 Alignment Workshop talk on reverse-engineering the internals of neural networks.
View Details論文與出版物
思想連結
Dario Amodei
曾經合作執行長暨共同創辦人・Anthropic
他們的合作早於 Anthropic,也早於大型語言模型的時代:2016 年兩人合寫〈Concrete Problems in AI Safety〉,把對超級智慧的模糊恐懼,轉譯成五個可以動手研究的工程失效模式。五年後,他們一起離開 OpenAI,成為 Anthropic 七位創辦人的一員;如今 Amodei 主持這家發表模型的公司,Olah 則帶領試圖讀懂模型內部的團隊。「安全是一門經驗科學,而不是一種哲學立場」——這個賭注,從那篇論文一路延伸到整間實驗室。
arxiv.org · en.wikipedia.org
Amanda Askell
曾經合作哲學家暨 AI 對齊研究員
兩人同列〈A General Language Assistant as a Laboratory for Alignment〉(2021)——Anthropic 的第一篇論文,也是提出「有益、誠實、無害」這個目標的地方,Claude 至今仍朝它對準;Askell 是第一作者,Olah 在名單末端。他們從兩端接近同一個模型:她寫下它應該是什麼,他試著讀出它實際上是什麼。這兩種描述之間的落差,差不多就是整個對齊研究。
arxiv.org
Bret Victor
受其影響介面設計師暨運算願景家
Victor 在 2011 年提出「可探索的解釋」一詞,指的是那種你能伸手進去操作、而不只是閱讀的文件。六年後,Olah 與 Shan Carter 以此為前提創辦 Distill,並在〈Research Debt〉中向他致意——那篇文章主張,一個領域會累積起未償還的「詮釋勞動」債務:解釋者付出的是固定成本,每一位讀者付出的卻是理解成本,而正是這個倍數,讓「把研究講清楚」值得被當成研究來資助,而不是當成一份人情。Distill 是 Victor 那套媒介被一門急需它的科學正式採用的最清楚案例;而 Olah 後來的可解釋性工作也保留了這個習慣:那些 circuits 論文,是可以捲動、可以動手撥弄的東西。
distill.pub · en.wikipedia.org
傑佛瑞·辛頓
思想同道AI 先驅與研究者
Olah 講了一個關於辛頓的笑話——說他五十年來每年都發現一次大腦是怎麼運作的——並且謹慎地補上一句:他說這話是帶著深深敬意的。對一個從另一端挖同一條礦脈的人來說,這個分寸剛好。辛頓用一整個生涯主張人工網路與大腦共享某些原理;Olah 的 circuits 研究,則把這個主張的一個版本放到了可以被檢驗的地方。它的「普遍性」宣稱是:類似的特徵與電路會跨模型、跨任務地反覆出現,就像類似的器官跨物種反覆出現。辛頓是從大腦推向網路,Olah 則是把網路打開來看。
youtube.com · distill.pub