Chris Olah
Co-founder & Interpretability Lead, Anthropic
關於
Chris Olah (b. 1992) is a Canadian machine-learning researcher and one of the pioneers of mechanistic interpretability — the effort to reverse-engineer neural networks by mapping the internal features and circuits that produce their behavior. A Thiel Fellow who left the University of Toronto after a year, he led interpretability research at Google Brain and OpenAI before co-founding Anthropic in 2021, where he leads the interpretability team. He co-founded Distill, a journal devoted to clear scientific communication, and his visual, essayistic explanations of neural networks shaped how a generation understands them. His work turns 'do we understand what we built?' from a rhetorical worry into an experimental science.
主要貢獻
- Pioneered mechanistic interpretability — reverse-engineering the features and circuits inside neural networks
- Co-founded Anthropic (2021) and leads its interpretability research
- Founded Distill, setting a new standard for clear, interactive scientific communication in ML
- Produced foundational work on feature visualization, circuits, and superposition in neural networks
- Named to the TIME100 AI list (2024) for advancing the science of understanding AI systems
影片與訪談
Mechanistic Interpretability explained | Chris Olah and Lex Fridman
Olah explains what it means to find features and circuits inside a trained model.
View Details
Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
Olah's 2023 Alignment Workshop talk on reverse-engineering the internals of neural networks.
View Details
The Dark Matter of AI [Mechanistic Interpretability]
Ask ChatGPT to forget a phrase and it will say it has, which is impossible — the phrase is in its context window — and press it and it will admit as much. We train assistants to be honest through examples; we have no direct access to whether they are. This is the interpretability problem, and the most promising handle on it is the sparse autoencoder, which extracts features that turn out to correspond to concepts a person can name.
View Details論文與出版物
思想連結
Dario Amodei
曾經合作執行長暨共同創辦人・Anthropic
他們的合作早於 Anthropic,也早於大型語言模型的時代:2016 年兩人合寫〈Concrete Problems in AI Safety〉,把對超級智慧的模糊恐懼,轉譯成五個可以動手研究的工程失效模式。五年後,他們一起離開 OpenAI,成為 Anthropic 七位創辦人的一員;如今 Amodei 主持這家發表模型的公司,Olah 則帶領試圖讀懂模型內部的團隊。「安全是一門經驗科學,而不是一種哲學立場」——這個賭注,從那篇論文一路延伸到整間實驗室。
arxiv.org · en.wikipedia.org
Amanda Askell
曾經合作哲學家暨 AI 對齊研究員
兩人同列〈A General Language Assistant as a Laboratory for Alignment〉(2021)——Anthropic 的第一篇論文,也是提出「有益、誠實、無害」這個目標的地方,Claude 至今仍朝它對準;Askell 是第一作者,Olah 在名單末端。他們從兩端接近同一個模型:她寫下它應該是什麼,他試著讀出它實際上是什麼。這兩種描述之間的落差,差不多就是整個對齊研究。
arxiv.org
Bret Victor
受其影響介面設計師暨運算願景家
Victor 在 2011 年提出「可探索的解釋」一詞,指的是那種你能伸手進去操作、而不只是閱讀的文件。六年後,Olah 與 Shan Carter 以此為前提創辦 Distill,並在〈Research Debt〉中向他致意——那篇文章主張,一個領域會累積起未償還的「詮釋勞動」債務:解釋者付出的是固定成本,每一位讀者付出的卻是理解成本,而正是這個倍數,讓「把研究講清楚」值得被當成研究來資助,而不是當成一份人情。Distill 是 Victor 那套媒介被一門急需它的科學正式採用的最清楚案例;而 Olah 後來的可解釋性工作也保留了這個習慣:那些 circuits 論文,是可以捲動、可以動手撥弄的東西。
distill.pub · en.wikipedia.org
傑佛瑞·辛頓
思想同道AI 先驅與研究者
Olah 講了一個關於辛頓的笑話——說他五十年來每年都發現一次大腦是怎麼運作的——並且謹慎地補上一句:他說這話是帶著深深敬意的。對一個從另一端挖同一條礦脈的人來說,這個分寸剛好。辛頓用一整個生涯主張人工網路與大腦共享某些原理;Olah 的 circuits 研究,則把這個主張的一個版本放到了可以被檢驗的地方。它的「普遍性」宣稱是:類似的特徵與電路會跨模型、跨任務地反覆出現,就像類似的器官跨物種反覆出現。辛頓是從大腦推向網路,Olah 則是把網路打開來看。
youtube.com · distill.pub