视频库 / NO.126
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

节目发布 2026-08-18 · Sequoia Capital
理查德·萨顿 库拉姆·贾维德 主持人
本期追问 · 点击跳到视频对应位置
10:59 合成数据真能让 AI 越过人类互联网的数据上限吗?22:02 权重冻结后只靠上下文记忆的模型,算在学习吗?30:39 造火箭这种想象未有之物的能力,真能从经验中学来吗?26:09 如果动物不靠监督学习,学校教育算智能的本质吗?
归入 Ⅱ·02 怎样才算真正学会? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文是强化学习奠基人 Rich Sutton 与其学生、合作者 Khurram Javed 的一次对谈。二人刚刚联合创办 Oak Lab,此番受两位投资机构访谈者之邀,从《苦涩的教训》(The Bitter Lesson)谈起,一路谈到大语言模型的边界、「大世界假说」、阿尔伯塔计划,以及 Oak 的研究抱负。全文依据现场录音编译整理,仅删去口语枝节与重复寒暄,论点、例证与转折悉数保留。

我不怪,是领域怪

萨顿:人们有时觉得我的观点很激进。他们提问时总先铺垫一句:您的想法跟所有人都不一样。可我完全不这么看。我觉得我是按常理在想问题,是其他人想得有点怪。而且我说的「怪」,只是最近这段时间的事。在这一轮 AI 狂热之前,你根本不必说「持续学习」(continual learning),因为不持续的学习是讲不通的。一切学习都是持续的。我们总是一边行动一边学习,这才是正常的思路。我不怪,是这个领域怪。这个领域非得给它起个名字叫「持续学习」,其实它就是学习。

主持人:今天很荣幸请到 Rich Sutton。Rich,你开创了强化学习,写下了这个领域的奠基教科书,培养了 David Silver 这样的关键人物,还写了《苦涩的教训》,我认为那是这个领域的圣经。感谢你抽时间来。与你同来的是 Khurram Javed,你的联合创始人,也是你在阿尔伯塔大学的学生。你们两位创立了 Oak Lab,我很期待谈谈这件事。今天的顺序是这样:先从《苦涩的教训》谈起,聊聊当下的局面,大语言模型(LLM)究竟能不能把我们带到终点;然后转到你们的研究计划和 Oak 的规划。

Rich,我们先回溯一下。本来打算从《苦涩的教训》开始,但我想再往前推。几十年前你决定把整个职业生涯押在强化学习上,尤其是深度强化学习,还把阿尔伯塔大学建成了这个方向的重镇,而那时这个领域还在襁褓中。是什么给了你这样的信念?

AI 寒冬里的选择:2003 年去阿尔伯塔

萨顿:不然还能干什么呢?我们想弄明白心智是怎么回事,而学习是心智的核心,有目标也是心智的核心,是智能的核心。所以我不过是在一直以来的想法上加倍下注而已。

主持人:当时有人觉得你疯了吗?

萨顿:那时候是寒冬,AI 寒冬。

主持人:哪一年?

萨顿:2003 年。说实话,那段经历有点离奇。2003 年我病得很重,其实是快要死了,癌症。可我一直没死成,我「努力」了好几年,又一次缓解之后还是没死。于是我想,既然死不成,拖了这么久,不如干脆再找份工作。就这样我去了阿尔伯塔,开始在那里教书。最后我没有死。现在拿这事开玩笑,当时可是相当严重的。

其实还有一个更要紧的问题:明明只剩几个月,我为什么还要继续做研究?我常常想起一句据说出自本杰明·富兰克林的话:如果你想知道一个人为什么做某件事,答案几乎总是两者之一,要么是习惯,要么是虚荣。我想这大概是对的。也许我只是习惯性地做着一直在做的事,也许是虚荣。我倾向于是习惯,因为我那时真的快死了。

主持人:天哪。这简直是天意。

萨顿:对我来说,保持决心从来不难。我还想把这个回答再拉长一点。

主持人:请讲。

萨顿:人们有时觉得我的观点很激进,提问时先说我的想法跟所有人都不一样。可我完全不这么看,我觉得我是按常理在想,是其他人想得有点怪。而且只是最近这段时间人们才想得怪。你去看看过去人们对心智的看法,哪怕只回溯十年,你会发现那些想法是:学习很重要,必须有目标,感知很重要。我们是低层次的存在,以极快的速度产生动作、接收数据,同时又必须在更高的层次上思考。再往前推几年,在这一轮 AI 狂热之前,你根本不必说「持续学习」,因为不持续的学习讲不通。一切学习都是持续的,它不是某个特殊的阶段。我们总是一边行动一边学习,这才是正常的想法。我不怪,是领域怪。他们非得叫它「持续学习」,其实它就是学习。

主持人:「我不怪,是别人怪」,这个座右铭不错,我们得发条帖子。整个领域都庆幸你活了下来,感谢你一直在推动 AI 的前沿,也感谢你培养出那么多同样推动前沿的学生。这二三十年里,你是怎么挑学生的?

萨顿:你这是给我谦虚的机会。我喜欢谦虚,喜欢指出那些所谓的伟大决定其实都是碰巧发生的。对学生我就是这种感觉:我并不觉得自己挑得多好,有时运气好,有时运气差。我不认为自己特别擅长挑学生。看看 Khurram,有时候你就是会碰上真正出色的人。David Silver 是他挑了我。Khurram,我是怎么把你「弄到手」的?

贾维德:我的硕士不是跟 Rich 读的,本来打算去业界。后来我们在一个项目上开始合作,起因也很偶然:我做过的一样东西,在 Rich 参加的一次会上被人提起,于是我被拉了进去。合作非常顺利,我很开心,Rich 也觉得不错。半年之后有了些进展,把它写成博士论文开题报告就顺理成章了。所以我从来没申请过,也从来没问过「你能当我的博士导师吗」。我们先一起做事,然后觉得这可以是篇不错的论文,之后我才去申请博士。

主持人:人生的路总在意料之外。

《苦涩的教训》的精髓与 LLM 的双面性

主持人:把时间拨到 2019 年。你写下《苦涩的教训》,它成了这个领域的母本。2019 年写这篇文章其实是个有趣的时点:ImageNet 是 2009 年,AlphaGo 是 2015 年。它写在大语言模型的规模化范式起飞之前,却在深度学习已经证明了自己之后。是什么促使你在 2019 年做这番反思?

萨顿:那是酝酿了很久的东西。正如文中所说,这个教训你可以在很长的时间跨度里、很多个十年里反复看到。它至少同样多地来自符号 AI 那一轮,而那一轮我是亲身经历过的。它讲的是别被「把人类知识塞进去」分心,而要盯着问题真正需要什么,以及怎样随着计算量扩展。至少一年前我就写过几个版本,也做过演讲。它不是对某个时刻的回应,而是对我长期经验的回应:不同的人用不同的方式思考怎样造出聪明的系统。

主持人:《苦涩的教训》的精髓是什么?如今我的会议里出现频率最高的说法,大概就是「这个有没有吃透苦涩的教训」。这个词这么流行,想必已经被各种扭曲和误用,偏离了你的本意。你认为它的精髓是什么,人们又在哪里理解错了?

萨顿:你让我想起我最近在 X 上发的一条帖子,我试着用 26 个英文单词把它讲完。大意是:别像 AI 历史上一再发生的那样被人类知识分心,而要专注于能随计算量扩展的学习方法,比如搜索,比如学习。所以它说到底是关于算法和算法改进的。它不是说你不需要精巧的算法,你需要,但你要的是那种能随计算量一起扩展的精巧算法。

主持人:而不是随数据扩展。

萨顿:而不是随人类输入扩展。接下来的问题我可以先替你问出来:那大语言模型呢?

主持人:对。它们和你的教训是一致的,还是相悖的?

萨顿:我想过这个问题,也发过一条帖子。结论是:它既是这个教训的正面例子,也是反面例子。首先,大语言模型让计算量得以空前地扩展,你可以把整个互联网灌进去,规模做到极大。所以它是一条靠可扩展方法获得强大得多的系统的路子。但再往后走,它最终会受制于那些信息本身。互联网是有限的,很难再拿到更多样本。而世界是大的,远远大过我们存在互联网上的一切。所以到最后,它恐怕会成为另一种例子:我们过度依赖人类知识,而这最终拖住了我们。

合成数据之争与大世界假说

主持人:我想在这里追问一下。眼下各家基础模型实验室都在做合成数据生成,想借此越过现有人类互联网这座「化石燃料」。合成数据生成作为 LLM 规模化范式的一部分,算不算一种能利用计算量的通用方法?

萨顿:不算。那就是个大错误。

主持人:为什么?

萨顿:这个错误太大了,也许它会成为下一个苦涩的教训。这个想法在阿尔伯塔流传了五到十年,我们叫它「大世界视角」或「大世界假说」(big world hypothesis)。Khurram 后来把它写成了一篇小论文,题目就叫《大世界假说》。

贾维德:大世界的意思是,世界是无限大的,要学的东西无穷无尽。你可以让人去生成合成数据集,但永远还有更多东西要学。正因如此,如果能直接从经验中学习,把人从回路里拿掉,你就会得到什么都能做的系统。因为世界很大,我们想让它们做的任务很多,而它们可以靠从自身经验中学习去做任何事。

回到合成数据本身:谁来决定什么是好的合成数据、什么是坏的?我可以写一个程序,吐出海量合成数据,但那些数据只会伤害模型。眼下是人在决定。这就是瓶颈:你可以让人来决定怎么生成这些数据集,但你需要人类专家来判断什么数据好、什么数据坏,这条路才能扩展。所以它受制于人。

主持人:难道不是我的损失曲线在决定吗?用这套数据我进步了多少,用那套更好的数据又进步了多少。

贾维德:没错。可是,如果 OpenAI、Anthropic 和所有大实验室的工程师全都去度假了,谁来生成合成数据?问题就在这里。它不来自智能体的经验,不是智能体自己生成的。必须有人来决定该生成什么样的合成数据,而这需要人类的专业知识。举个例子,假如你想让系统做一件从物理角度很有挑战的事,比如造一架像蝙蝠那样靠回声定位飞行的无人机。这件事的「正确」合成数据是什么?我想你得雇领域专家去弄清该要什么数据,再把它生成出来,然后也许能从中学到东西。但前提是这个领域专家得先存在。到这一步,我们就被人类的专业知识卡住了。

主持人:但你可以有无限多的合成世界,而现实世界只有一个,是有限的。

贾维德:那我们回到回声定位。我的目标就是一架能靠回声定位自我定位、自主移动的无人机。这是一个机器人,它自己就在生成经验,完全可以从自身经验中学。可无论你生成多少合成数据,哪怕你生成的合成数据涵盖了五十个不同的宇宙,只要人类没先把这件事弄明白,它就完不成这个任务。

萨顿:我先说一句:合成数据是错的。它不会正确,它是一个合成世界,不是真实世界,而这个差别是要紧的。世界极其复杂。你写一个小程序(生成合成数据的东西注定是个小程序),它造出的只会是一个小世界。

举个例子。对我来说重要的是:你此刻脑子里在想什么。你说,「那我去弄点合成数据来告诉我别人脑子里在想什么」?不行,别人的心智是不可能有合成数据的。而别人的心智对我们至关重要。比如我今天在跟你们谈投资,我当然在乎你们心里在想什么,这种东西我上哪去弄合成数据?其实什么东西你都弄不到真正的合成数据。无人机在物理世界里怎样与环境互动,机器人电机里的摩擦在哪里,这些你都没法合成。世界无限复杂,任何对它的模拟都微不足道。

大世界假说说的是:世界比你的心智、比任何智能体的心智都复杂得多得多。这是显而易见的,因为世界里包含许多别的智能体。既然世界如此庞杂,你就不可能做到任何号称最优或完美的事。你注定是不完美的,必须使用近似,而且近似会很粗糙。正因如此,这才是我们必须持续学习的根本原因,如果你非要找一个原因的话。我们必须持续学习,因为我们会碰上这个巨大世界里的某个特定局部,而我们得学出一个针对所处局部调校过的近似,而不是针对我们不在的其他所有局部。

主持人:我再顶一次。抱歉,我是故意抬杠,但我想搞明白。据我了解,新一代自动驾驶公司里有很多是主要在仿真里训练的,然后再做些后训练来保证它在真实世界能用,而且这条管线非常有效。

贾维德:我认为这里该问的问题是:造那个仿真器用了多少工程师?我们是否准备承认,值得解决的问题只有那些能先雇一大队工程师造仿真的问题?而且我敢肯定他们迭代了很多轮:造仿真,在里面学,发现仿真与现实的落差不可接受,再修。这里始终有一个人在回路里修仿真器。他们从真实世界拿到反馈,再由人去修。为什么不能把人拿掉,让智能体自己来?

萨顿:而且等它真正上路,还是会遇到意料之外的事。

贾维德:那正是你真正需要从经验中学习的时候。

主持人:所以你们的意思是,来自经验的数据,远远多于人类能策划和制造出来的数据。

贾维德:对。而且从仿真里学习显然是有价值的,也有正确的做法:智能体从自身经验中学出一个模型。由智能体自己学出来的模型好得多,因为一旦模型错了,它可以通过持续学习去修正。如果仿真器是人造的,那只有在人发现哪里不对的时候,模型才会更新。所以,是的,规划很重要,智能体应该从仿真器里学,但那得是它们自己造的仿真器。

先验与学习不该为敌

主持人:我想转到《苦涩的教训》的另一部分:去除人类知识。你在文中写道,为了追求短期内见效的改进,研究者总想利用自己对领域的人类知识,但长远来看唯一要紧的是利用计算。你的前学生 David 做的 AlphaGo 和 AlphaZero,在我看来是去除人类先验的一次凯旋。这个结果让你意外吗?

萨顿:它当然让我很高兴,觉得自己被证明是对的。但它本来可能是另一个结果,因为先验知识是能帮上忙的。先验知识没有任何错。我在《苦涩的教训》一开头就说了:先验知识与学来的知识之间没有理由非得冲突。你可以放一些先验进去,然后开始学习,原则上这两者没有理由对立。它们都是关于知识的,人生就是获取知识、拥有知识。不知怎么的,先天与后天成了敌人,可其实先验是你已经有的,然后你再学更多,它们本该是朋友。

但正如我在文章开头说的,在实践中它们一直是敌人。那些偏爱现有人类知识的人,最后总是希望自己那一边赢,于是想贬低或者干脆否定学习。所以现在你们对我的印象大概是:一个热爱学习、想否定先验知识的人。可我真正感兴趣的是心智。心智就是你有先验知识,然后你获得更多,而一旦获得了更多,它就成了你新的先验,如此不断累积。这两样东西是协同工作的。我之所以看起来像一个主要关心学习的人,是因为世界上其他所有人都在说:只要知识够多就行,不需要学习。

你看大语言模型:我们把所有知识塞进系统,然后大语言模型在运行时不学习。它在跟人说话,在互动,但权重绝对不会变。所以,我不是那个怪人。怪的是你们这些人,居然认为一个完全不再学习的东西,能拿出博士水平的经验和专长。我不是那个怪人。

主持人:所以你的建议是,让算法运行更长得多的时间,然后再喂给它先验数据?或者说,把先验知识沿途一点点滴进去?

萨顿:两者都重要。但长远来看,你必须去获取新知识,并且把获取新知识这件事结构化。长远看只有这个是要紧的。在你这样做的过程中,当然会有一些你先前已经拥有的东西。

想想将来我们有了智能机器人会怎么做。我们会让它们全部从零开始学吗?还是把它们复制一份,让它们从现有的状态接着学?它们是数字的,复制很容易。所以与其搞一个巨大的工程,花天文数字的钱从互联网上重新训练,我们不如直接复制这个智能体,然后接着学。从这个意义上说,先验知识反倒可以不必操心,因为你直接从上一个机器人那里复制过来就是了。

权重不变,就不算学习

主持人:那你能描述一下,一台从经验中学习的机器或计算机是什么样子吗?

萨顿:它可以是个机器人,也可以完全活在互联网上。比如在互联网上为数据包选路,做得对经验敏感,越做越好。也可以是通过用户界面与人互动,在你的手机或电脑上,随时间变得越来越好。比如一个智能助理,它必须越来越好,必须知道你想要什么。

主持人:你会不会主张,当下流行的基于 LLM 的助理并不是经验学习者或持续学习者?如果是,根本差距在哪?

萨顿:你是认真的吗?

主持人:它们显然会记住关于我的东西,也在做一些上下文内学习(in-context learning)。

萨顿:它们的权重从不改变。

主持人:顺便问一句,只有一小部分权重在变够不够?还是所有权重都得变?

萨顿:想想创造大语言模型时所有那些结构化的过程、那些新概念的生成,全都是权重学习。你希望能够继续做这件事,而不是只做一次。

主持人:换个说法是不是:我们在发布模型之前做了太多预训练和后训练,之后它就不学了?

贾维德:唯一的分歧点是:之后我们不让它们学。预训练想做多少做多少,没问题;后训练也可以。可当我在使用这个模型的时候,它就停止学习了。你可以给它更多上下文,通过上下文改变模型的状态。它早就学会了:如果状态不同,如果状态里出现了新东西,它就会用那个来做下一次预测。但模型本身没在学习。

主持人:Cursor 的 Tab 补全模型确实会根据使用情况更新。

贾维德:那些权重是在变的。Cursor 的 Tab,我想还有 Composer,它们也在更新,这是两个持续学习的例子。但可以做得好得多。据我理解,他们的做法是:大量用户在用 Tab,他们从成千上万乃至上百万用户那里收集数据,然后用这一批数据对策略做一次更新。这能行。但假设我想教这个模型一件很具体的事,我不想跟十万个其他人争抢该教模型什么。我想教我的模型一件很具体的事,而且是教我的那个版本。模型从别人那里得来的共享知识,我不在乎。所以那是一种效率很低的做法。

主持人:现在的做法似乎是:基础技能学进权重里,对所有人通用;个性化则以上下文的形式实现。这不是学习该有的正确心智模型吗?难道所有上下文都该活在权重里?

贾维德:上下文可以放在状态里,两者都可以。但你仍然需要能更新权重。举个例子,一些很好的案例来自人类的残障研究。当一个人经历了改变心智或某些感官的变故,你能看到他们适应的过程。比如我们有本体感觉,有内部传感器告诉我们身体处在什么姿态,走路就靠它。有些人会完全失去这种能力,于是根本没法走路,因为它就是走路这套策略的地基,深深刻在脑子里。可经过两三年,他们能靠看着自己的脚重新学会走路,用视觉反馈代替。大脑的可塑性惊人,一件二十年来一直成立的事,一旦不再成立,大脑能去更新它、抛弃它。我认为这就是我们希望系统具备的、极其有用的能力。

没有动物靠监督学习

主持人:人类婴儿和动物的学习方式,有什么值得我们借鉴的?你们从中汲取多少灵感?

萨顿:我们汲取了很多灵感,但不把「AI 必须像婴儿、像人、像动物那样行为」当成要求。它是灵感的来源。来自动物学习的灵感,而不是约束。

主持人:这和《苦涩的教训》是一致的。那你觉得生物学习里最该借鉴、而今天的系统里最缺的是什么?

萨顿:我感觉自己在发表意见了,不过都是些显而易见的意见。我认为很明显,没有任何动物是靠监督学习来学习的。因为没有人给我们示范肌肉该怎么抽动,而肌肉抽动才是我们的输出。

主持人:但学校教育全是监督学习。

萨顿:我知道,可绝对不是。就算是,学校也只是我们所学的极小一部分。我们学会看,学会走,学会世界是怎么运转的。而且即便在学校里,也没有人告诉我们该怎么抽动肌肉。

主持人:我掌握的知识和技能是在学校里通过监督学习得来的。

萨顿:我不想说向他人学习、来自他人的传递不重要,它极其重要,语言也极其重要。但我们缺的是什么?没有监督学习,没有人给我们目标值。你听到正确答案:法国的首都是哪里?我们知道答案是巴黎。可没有人告诉我该怎么把「巴黎」念出来。你说答案是巴黎,我听到你的话,然后我得做出另一套肌肉动作来产生「巴黎」这个答案。这不是字面意义上的监督学习。

所以第一点,学校其脆无关紧要。松鼠不上学。动物不是那样学的。学校是件非常特殊的东西,连我们人类在几百年前都没有。它不属于智能的本质。把它当成学习的首要范例是一种分心,因为我们作为动物的时候根本不做这件事。

主持人:真希望我父母当年听过你这番话,免得逼我去上学,忍受那一套规矩。

主持人:可问题是,松鼠固然很会从树上跳下来,却不会证明数学定理。而如果我想学证明数学定理,我就得上学。

萨顿:它们也没有 DVD 和 iPod。有很多事它们能做而我们不能。至于数学定理,还有下棋,这有点像莫拉维克悖论:那些我们认为极其智能的高级事务,对计算机反而容易;难的是所有那些平常的事,比如运动,比如带着注意力去看。我认为监督学习是个好东西。我只是喜欢找显而易见的事实:没有人通过给例子来教我们抽动肌肉,因为他们不可能做到,我们只能自己摸索。

贾维德:而且他们给的答案会是错的。如果我用和 Rich 完全一样的方式动嘴、动舌头、动声带来念「巴黎」,发出来的肯定是截然不同的声音。从这个意义上说,Rich 也好,任何人也好,都不知道用我的身体发出那个音的正确方式。只有我自己知道。

火箭与范式转变:想象能否来自经验

主持人:在我看来,最原始的感觉运动能力,尤其是物理世界中的运动,确实天然是从经验中学的,这点我同意。但有更高层的抽象,更接近于让人类伟大的那些东西,其中很多并不存在于低层的感觉运动学习里。你们的世界模型是否从感觉运动学习一路涵盖到顶?

萨顿:是的,这正是我们的抱负。顺便说一句,松鼠也能做极其抽象的事。

主持人:松鼠能做的最酷的事是什么?

萨顿:它总能钻进你的喂鸟器,不管你设多少障碍,它总能找到新的跳法、爬法,做各种各样的事。

贾维德:计算轨迹也很在行。动物很擅长理解物理世界,用不着我们设想发射自己上太空时以为自己在做的那种心算。它们能实时、正确地控制跌落来避免受伤。

主持人:好吧,有道理。

萨顿:我认为这只是程度之别。我倾向于认为其他动物离人类非常近。刻意强调我们与动物有何不同,是一种傲慢。看到共性更好。我认为我们只是程度不同,当然社会和文化给了我们巨大的优势,语言也给了我们巨大的优势。

主持人:我想在这点上再追问一下,替 Sonia 撑撑腰。我相信动物和孩子从经验中学习,而且学出了不可思议的东西。我儿子两三四岁的时候,我常常惊叹:没人真正教他,他就学会了这些。但 Sonia 说的是,让人之为人的那些东西,比如上太空、造火箭,并不是百分之百从经验中学来的。在你发射火箭之前,你得先在头脑里把它抽象地推演一遍,而这不是从所谓的「经验」里学来的,因为你根本不知道它能不能成,你得想象出来。我们怎么教机器去想象此前从未有过的东西?这大概正是我们想追问的,因为我们还没弄明白。

萨顿:这一点我其实是同意你的。你必须能规划,必须能想象。

贾维德:你觉得一千年前的人类,在他们还没做出我们谈到的大部分事情之前,是不是同样智能?比如把那个时代的人放到今天的文化里,他们能不能获得同样的技能,开始做有用的事?

主持人:就算回溯一万年,我也不认为人脑进化了多少。根本上是同一台机器。

贾维德:根本上是同一台机器,但我们积累了一万年的知识。

主持人:对。而我通过上学、通过监督学习就能拿到这一万年的知识,比靠经验去学快得多得多。

贾维德:没错,我觉得你说得完全对。我们希望系统从经验中学习,而它们经验的一部分就是接触我们的文化,从文化中学习。这都很好。但我们来谈谈有人做出范式转变的时刻。大家都举爱因斯坦的例子,其实例子很多,比如从「编程」到「学习」这个转变本身。在这些范式转变发生时,我会说,那是一个积累了所有这些知识的人,从自己的经验出发构建新的抽象,用这些抽象做规划,进而发现新知识。这种能力,构建新抽象、学出模型、再用它规划的能力,在我们当前的系统里是完全缺失的。你可以在人类知识的前沿看到这个问题暴露出来,但你同样可以在感觉运动流的层面上研究它。

萨顿:所以我们并不是在和这个原则争论。我们需要形成抽象,才能在高层推理。不过你们刚才差点做了我说过永远不该做的事,就是争论「先验知识重要还是获取知识重要」。你们刚才说的就是这个:你说你还是得学东西,又说你可以从文化和先验知识里得到东西。可这两者不该互相打架。

主持人:具体到范式转变,我们怎样造出一台知道何时该转换范式的机器?

贾维德:我想是通过它的经验。它必须通过自己的经验来做,不能依赖人类知识,因为我们的前提是人类只看到一种范式,而我们想要一种不同的看法。所以它得通过自己的经验找到更好的东西,也许是泛化得更好、预测得更准,也许是在别的方面更好,但必须来自它自己的经验。

萨顿:我们这个领域还没看到的一项重大能力,就是学出一个模型,然后用这个模型做规划。数学题我们能做,AlphaGo 我们能做,因为在游戏里模型是已知的,我们知道棋子怎么走;在数学里我们知道算子是什么,知道 Lean 会把我们从一个知识状态带到证明的下一个状态。但如果模型必须靠学习得来,那么我要把话说出来,虽然听起来可能是个反例:在我们的领域里,我看不到任何「学出模型再用它规划」的实例。

贾维德:至少没有使用自己发现的抽象的实例。有人会说:我就学一个预测下一秒或下一毫秒会发生什么的模型。但我们的模型不是那样工作的。我们的模型更抽象,性质相当不同。

阿尔伯塔计划:抽象与持续深度学习

主持人:我喜欢你们的一点是,你们不是坐在那里高谈阔论或哀叹世界的现状,你们非常行动导向,所以才创办了公司。我们来聊聊这个。Rich,2022 年你提出了一个非常具体的十二步计划,《阿尔伯塔 AI 研究计划》。请谈谈它。

萨顿:阿尔伯塔计划的由来是,我们有一些总体想法,但也需要把它们拆成更小的块。这十二步就是试图把具体的块结晶出来。有一个非常重要的早期步骤,第二步,持续深度学习(continual deep learning)。我们认为它几乎是最重要的一步,因为它能解锁其他所有东西。如果你能做持续深度学习,就能持续更新你的世界模型。然后,如果你知道怎么把抽象做对,也就是这些步骤的后半部分,全都是关于怎样把抽象做对的。我说「做对」并不是说找到那个「正确的抽象」,因为没人能说出什么是正确的抽象,那取决于你所在的世界。你的智能体得为它所处的世界学出正确的抽象。所以也许关键就是这两件事:你必须找到正确的抽象,你必须能做持续深度学习。

贾维德:我想领域里很多人都意识到我们需要模型、需要用模型规划。但抽象告诉我们模型该以什么为条件。模型该预测什么?你要做什么,然后什么会发生?更重要的是,这些东西从哪来?我很喜欢精英运动员的例子。你去问顶尖运动员他们怎么做某些动作,他们会有一套稀奇古怪的小圈子术语来指称非常具体的东西。他们会说,我做了这个那个,还给它起了名字。有时如果只是自己练,甚至连名字都没有。那么他们是怎么想出这些抽象的?这在某种意义上正是缺失的关键,阿尔伯塔计划的后半部分回答的就是这个。

灾难性遗忘的解药:步长优化与持续反向传播

主持人:能谈谈持续深度学习这部分吗?今天存在的是算法上的缺口,还是只是部署基础设施和数据隐私方面的实际缺口?因为如果我想根据用户交互做所谓「朴素」的权重更新,今天就能做到。你们认为,要实现持续深度学习,我们最缺的是什么?

贾维德:绝对是算法缺口。你可以做朴素的做法,但会看到各种各样的问题。比如你说,我拿一个样本,然后用这一个样本更新整个模型。你会碰到的问题是,模型里所有先前的知识都被负面影响了。我们目前绕过这个问题的办法正是 Cursor 的做法:不用一个样本,而是用来自大量用户的一大批数据。在能拿到这种数据的场景里,你可以做持续学习。但大多数场景没有这个条件,大多数场景你只有单一的数据流。这时用朴素方法,它会以极具破坏性的方式彻底摧毁你的先验知识。

主持人:灾难性遗忘(catastrophic forgetting)。

萨顿:对。但它完全可以治愈,只要有正确的算法。

主持人:解药是什么?

萨顿:首先你得做我们所说的步长优化(step size optimization)。意思是网络里的每一个权重都要有单独的步长,有的动得快,有的动得慢,而且你得对每个权重的步长做元学习(meta-learn)。网络里大多数权重的步长会非常小,这样当你用一个新样本训练时,它们不会被破坏,变化只发生在恰当的地方。

其次,你得用某种形式的「生成与检验」(generate and test),在特征空间里进行。也就是不靠梯度来产生新特征、新单元。因为梯度是个很慢的过程:只有当你知道某个方向有帮助时才往那边走。这永远会很慢,也无法给你一条通往越来越复杂、可持续学习的路径。你需要一种东西,能直接提出一批新单元,然后从那里出发。

我可以说一件具体的事。我们有一个叫「持续反向传播」(continual backprop)的算法,几年前发在《自然》上。它和反向传播完全一样,只是你还会不断种下新单元的「种子」,用随机权重新初始化。反向传播只在时间的起点有随机权重,随着训练进行,随机性带来的多样性全部被用光了。而持续反向传播不断注入一点随机性、一点生成与检验:生成,然后由反向传播来充当检验者。你需要这个。如果把这些东西真正结合好,我想你会得到新一代远为强大的持续深度学习。这就是我们希望在接下来几年里做到的。

主持人:这些算法能不能用在现状上,也就是那些正在扩展 LLM 规模、想让它们持续学习而不发生灾难性遗忘的人身上?

贾维德:可以,绝对可以。但我不认为你可以拿一个现有模型说,我现在就开始用这些算法更新它。因为这些算法是在元学习「如何学习」。所以你真的得说:我要从头学。比方说我要训练一个新的基础模型,但用这些新算法来训。这些新算法除了学习知识,还在学习将来如何学习新东西,两件事同时进行。然后,我认为你就能学新东西而不发生灾难性遗忘。

主持人:相对于现状,你们公司要做的最激进的事,就是同时做这两件事吗?

萨顿:这又回到「我不疯,是别人疯」。

贾维德:那也许不算完全激进。2016 到 2018 年前后,很多人在相当深入地探索这些想法,但设定要受限得多。他们会说,我们有一个问题的分布,然后在这个特定情形下做。而我们想从单一的经验流出发,所以我们的方法应该更普适。我想很多人探索过这个,但没人在最一般的设定里探索过,也就是让得出的算法能在任何地方适用。

Oak 的雄心:万亿参数、20 瓦、自洽心智

主持人:那么你们的新公司要做的、别人没在做的最激进的事是什么?

萨顿:最有雄心的事?记住,我不觉得自己怪,所以我不想说「激进」。最有雄心的事,我想是拥有完整谱系的知识,既关于细微的事,也关于宏大的事。比如怎样乘飞机从一座城市到另一座城市,这是件很大的事,类似你的太空飞行的例子,只是更贴近常识,因为我们很多人都坐飞机,而且所有人在生活的方方面面都在用抽象,连松鼠都在用抽象。所以,要拥有从小到大的整个知识谱系,用统一的方式处理它,并让它能自我维护。

大问题始终是:你有一个知识库系统,是什么让里面的知识保持正确?在大语言模型里,让知识保持正确的是人们做了大量后训练来确保它正确,然后把它冻结起来。可我们的心智不是这样:我们一直在改变东西,却有某种机制让它保持有条理、连贯,回落到一个好的位置,而不是漂进疯狂的境地。我认为这是我们最大的抱负:一个自洽的心智,能持续训练自己并保持连贯。

主持人:这是一个极其宏大的愿景,把所有这些统一进单一心智的想法太有雄心了。

萨顿:我认为这是触手可及的。现在是 2026 年,我们的计算机快得不得了。它真的雄心大到够不着吗?还是说我们对每一步该怎么做已经有了些眉目?我认为我们有愿景,也有眉目。我不觉得这不合适。

主持人:你们的愿景里有一个万亿参数、功耗 20 瓦的模型。这听起来相当有雄心。

贾维德:确实有雄心。在某种意义上,以现有技术来说也是不可能的:光是把一万亿个参数存进内存,用现有的存储技术就很可能超过 20 瓦。但我们真正在想的是:一切都在变好,计算越来越便宜,能效越来越高,那么五到十年后会在哪里?我认为五到十年后,有了正确的算法,我们完全可以身处一个这件事可行的世界。

主持人:五到十年,按摩尔定律是两个数量级。这是标准增速:每 18 个月翻一番,十年就是两个数量级。所以要让这个说法成立,你今天就得能做到什么功耗?20 瓦乘两个数量级?

贾维德:2000 瓦。

主持人:如果今天能用 2000 瓦做到,那十年后就能用 20 瓦做到。你们觉得能用 2000 瓦做到?有的是研究实验室的人手里资源远远不止这些。

贾维德:我认为有了正确的算法,我们现在就能比这更高效。

主持人:如果能更高效,那为什么没有做到?总不至于是人们就想把钱全花光吧。

萨顿:有时候看着就像他们想这样。我觉得他们是靠烧大量的能源来证明自己是真汉子。

贾维德:至少就我看到的各个研究组,我甚至没见到有谁相信这是可能的。如果你不相信,你就不会去啃那些技术难题。

主持人:是不可能,还是系统里浪费太多?是不是有人知道怎么高效地做,而同一个实验室里有十倍的人在做别的事,十个里有九个在浪费?

贾维德:我的理解是,我们困在一个局部极小值里。如果想转向这类新算法,几乎不可能不先变差再变好。当我们开始探索这些新方向,第一天不会有最先进的性能,因为这是不同的范式。但那条路通向同等的性能,在更高的能效尺度上。那些大实验室被产品锁得太死,根本不可能走一条先变差的路。

主持人:因为他们现有的范式还能继续扩展,而新范式需要下注。

贾维德:而且他们得解决一些困难的技术问题,那些问题我们已经想了很多年。我们认识的人也想了很多年,跟他们聊的时候,这件事可行是讲得通的,但你需要在这些难题上想很长时间。

LLM 只占智能的四分之一

主持人:如果 Oak 一切顺利,公司会怎样?你们在建一家什么样的公司?

萨顿:如果一切顺利,我们实现这个架构,拥有真正的持续学习,能形成抽象,从而做规划和推理,拥有某种真正的智能。到那时会发生什么,很难准确想象。

主持人:人类变得无关紧要。

萨顿:我完全不这么认为。我认为世界会变得更精彩、更有趣,对人类也是如此。但有一点特别值得琢磨:大语言模型可能会有风险。等这一天到来的时候,我相信它们已经风光过一阵了,事实上已经风光过了,非常成功。我要说清楚,大语言模型是一项惊人的科学突破,是神经网络娴熟运用语言上的突破,完全出乎意料。语言一直是符号方法的最后据点,而它们彻底改变了人们对此的看法。这是个大突破。让我沮丧的是,我们本可以庆祝在 AI 问题的一个子集上取得了巨大进展,好好享受它。可它非要装成全部的 AI。智能并不只是流畅、熟练地运用语言,还有远远更多的东西。语言是重要的一部分,大概占智能的百分之二十,或者四分之一。还有更多。我们还没完。

主持人:如果一切顺利,你们设想的是一个单一心智,从学会在树枝间荡秋千到造飞船,把今天谈到的所有事都能做?是单一心智、单一套权重做到所有这些吗?

萨顿:是单一的设计。会有很多不同的心智。

主持人:一个单一的设计,对不同环境做出反应。

贾维德:这个心智的不同版本会学到不同的东西,因为它们的经验不同。这又回到大世界假说:要学的东西无穷无尽,所以一个系统学不完无限多的东西。Rich 也提到过,如果你有两个这样的系统,两个世界上最大的系统,那么它们显然无法相互建模,因为它们同样复杂。所以单一系统永远到不了能学会一切的地步,永远会是多个系统,各自从自己的经验中学习。

主持人:你们在招人吗?想找什么样的人?

贾维德:在招。初始团队大部分人选我们心里已经有数,都是长期思考过这些想法的人。我们会采取稍微不同的路线,因为这是不同的范式:迅速做大没有意义,因为在某种意义上,我们招的每个人都得看见我们所看见的,而不是所有人都看得见。所以我们会从小做起,慢慢增长到十来个、二三十个人,再往下走。

萨顿:我们希望高度一致。

贾维德:高度一致,这样才能高效协作,把进展扩展开来。

主持人:非常好。我很喜欢这场对话,感谢你们抽时间分享正在做的事。对于强化学习和算法设计将走向何方,你们的思考深度不同寻常,今天能一起探讨是真正的荣幸。谢谢。

萨顿:非常感谢。这是我们的荣幸。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 「我不怪,是领域怪」:持续学习即学习 ▶ 正在看
0:59 AI 寒冬里的选择:2003 年去阿尔伯塔 ▶ 正在看
7:10 《苦涩的教训》的精髓与 LLM 的双面性 ▶ 正在看
10:59 合成数据之争与大世界假说 ▶ 正在看
17:49 先验与学习不该为敌 ▶ 正在看
22:02 权重不变,就不算学习 ▶ 正在看
26:09 没有动物靠监督学习 ▶ 正在看
30:39 火箭与范式转变:想象能否来自经验 ▶ 正在看
36:54 阿尔伯塔计划:抽象与持续深度学习 ▶ 正在看
38:45 灾难性遗忘的解药:步长优化与持续反向传播 ▶ 正在看
43:43 Oak 的雄心:万亿参数、20 瓦、自洽心智 ▶ 正在看
49:03 LLM 只占智能的四分之一 ▶ 正在看
本期小问 · 档案清单
10:59 合成数据真能让 AI 越过人类互联网的数据上限吗? ▶ 正在看
22:02 权重冻结后只靠上下文记忆的模型,算在学习吗? ▶ 正在看
30:39 造火箭这种想象未有之物的能力,真能从经验中学来吗? ▶ 正在看
26:09 如果动物不靠监督学习,学校教育算智能的本质吗? ▶ 正在看
本期讲者
理查德·萨顿强化学习奠基人之一,与 Andrew Barto 合著经典教科书《Reinforcement Learning: An Introduction》,2024 年图灵奖得主;著有《The Bitter Lesson》,长期任教于阿尔伯塔大学,现与 Khurram Javed 共同创办 Oak Lab。
库拉姆·贾维德Rich Sutton 在阿尔伯塔大学的博士生,持续学习与元学习方向研究者,《大世界假说》论文作者之一,Oak Lab 联合创始人。
主持人投资机构的两位访谈者(其中一位名为 Sonia),围绕《苦涩的教训》、大语言模型的局限与 Oak Lab 的研究计划发问。
01「我不怪,是领域怪」:持续学习即学习
0:00
People think I'm have a radical point of view sometimes. They say they start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And >> [laughter] >> And I mean that like, you know, it's just the recent times people are thinking weird. Before there was all this AI craziness, uh you talk about you wouldn't have to say continual learning cuz it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. We always act and we learn. That's just the normal way of thinking. I'm not weird.
有人觉得我的观点有时候挺激进的。他们提问的时候会先说,我的想法跟其他人有多么不一样。但我完全不这么看。我觉得我想的才是常规的方式。只是别人想得有点怪罢了。而且——>> [笑声] >> 我的意思是,你知道,只是最近这段时间大家想得有点怪。在这一波 AI 狂热出现之前,你根本不需要说什么“持续学习”,因为谈论一种不是持续的学习本身就说不通。所有的学习都是持续的。我们总是在行动,也总是在学习。这才是正常的思考方式。不是我怪。
便签笔记
0:37
The field is weird. The field they need to call it continual learning. It's just learning.
是这个领域怪。这个领域非要把它叫做持续学习。它就是学习而已。
便签笔记
0:50
>> [music]
>> [音乐]
便签笔记
02AI 寒冬里的选择:2003 年去阿尔伯塔
0:59
>> We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You're the key students in the field, uh folks like Dave Silver. You wrote the essay The Bitter Lesson that I believe is the Bible of the field. And and you have just been one of the greats in propelling the field forward. So thank you for taking the time to join us today. Um Rich is joined by Quorum Javed, his co-founder uh and former students from the University of Alberta.
>> 我们非常荣幸今天请到了伟大的 Rich Sutton。Rich,你发明了强化学习,写下了那本奠基性的教科书。你带出了这个领域里最关键的一批学生,比如 Dave Silver。你写了那篇《苦涩的教训》(The Bitter Lesson),我认为那是这个领域的圣经。而且你一直是推动这个领域向前发展的伟大人物之一。所以非常感谢你今天抽时间来参加。嗯,和 Rich 一起来的还有 Quorum Javed,他的联合创始人,也是他在阿尔伯塔大学的学生。
便签笔记
1:28
Um the two of you have set off to found Oak Lab. I'm very excited to talk to you about that today. So for today's session, we're going to start talking about The Bitter Lesson, the state of the world as we know it today, whether LLMs will get us there or not. And then we're going to we're going to transition to start talking about your your research agenda and your plan for Oak. Um Rich, maybe take us back. I was going to start with The Bitter Lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy. What gave you the conviction to do that?
嗯,你们两位一起创办了 Oak Lab。我今天非常期待和你们聊聊这个。那么今天这场对谈,我们先从《苦涩的教训》讲起,聊聊我们今天所认识的这个世界的现状,无论大语言模型能不能带我们到那一步。然后我们会转到讨论你的研究方向,以及你对 Oak 的规划。嗯,Rich,也许我们先回顾一下。我本来打算从《苦涩的教训》讲起,但其实我想从更早的时候开始。几十年前,你决定把职业生涯投入到强化学习,尤其是深度强化学习,并且把阿尔伯塔大学打造成了这个方向的重镇,而那时候我觉得这个领域还处于非常早期的阶段。是什么让你有信心这么做的?
便签笔记
2:10
>> What else you going to do? >> [laughter] >> We were trying to figure out the mind and learning is a central part of the mind. And having a goal is a central part of the mind. Central part of intelligence. Yeah, so I was just doubling down on what I was always thinking. >> Did people think you were crazy at the time? >> Um It was a winter. It was an AI winter. >> Uh what year was this? >> It was in 2003. >> Okay. >> And it's kind of crazy actually the truth cuz I was like really sick. I was dying I was actually dying of cancer in 2003. And but I I wasn't quite dead, you know, I've been trying for a number of years. And I wasn't dead after another remission. And so so I said, well, I'm not dying I haven't succeeded in dying. So I might as well, you know, it's going on long enough might as well just try to get another job. And so I so I went to Alberta and and and started teaching there.
>> 不然还能干什么呢?>> [笑] >> 我们当时想搞清楚心智是怎么回事,而学习是心智的核心组成部分。拥有目标也是心智的核心组成部分,是智能的核心组成部分。是啊,所以我只是在我一直以来的想法上加倍下注而已。>> 当时有人觉得你疯了吗?>> 嗯,那是个寒冬,是 AI 寒冬。>> 呃,那是哪一年?>> 是在 2003 年。>> 好的。>> 说来其实挺离谱的,因为我当时真的病得很重。我快死了,2003年我真的差点死于癌症。但我还没死透,你知道,我已经努力了好几年了。又一次缓解之后我还是没死。所以我就想,好吧,我没死成,我没能成功地死掉。那我不如你知道,这都拖了这么久了,不如干脆再去找份工作。于是我就去了阿尔伯塔,在那儿开始教书。
便签笔记
3:12
And then in the end I didn't die. It's kind of amazing it's cuz you know, it's it's like that. I'm joking about it now but it was quite serious. And um it's an even more important question. Why why did I continue to work on this research stuff when I was, you know, I only had a few months. I would always keep reminded what I think it's Benjamin Franklin is supposed to have said that, you know, if you ever wonder why someone is doing something it's almost always one of two things. It's either habit or vanity. Okay? So I think I think it's probably true. Maybe it was my habit to just kept doing what I'd always been was doing or maybe it was vanity.
然后到最后我居然没死。挺神奇的,因为你知道,事情就是这样。我现在拿它开玩笑,但当时相当严重。嗯,还有个更重要的问题。为什么我在只剩几个月可活的时候,还在继续做这些研究?我总会想起,我觉得是本杰明·富兰克林说过的一句话:如果你想不通一个人为什么要做某件事,那答案几乎总是两者之一。要么是习惯,要么是虚荣。对吧?我觉得这话大概是对的。也许我只是习惯了继续做我一直在做的事,也许是虚荣。
便签笔记
3:52
I I know. I think it was more like habit cuz I was I was dying. >> Wow. >> Um >> Wow. Divine intervention. >> Yeah, it's always been easy for me to keep be very determined. Um and I'm I'm I'm going to go even longer on this answer. >> Please go. >> People think I am have a radical point of view sometimes. They they they start questions saying how how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's everyone else that's thinking a bit weird.
我知道。我觉得更像是习惯,因为我当时都快死了。>> 哇。>> 嗯 >> 哇。神的旨意。>> 是啊,保持非常坚定对我来说一直都很容易。嗯,这个回答我还想再说长一点。>> 请讲。>> 人们有时觉得我的观点很激进。他们提问时会说,我的想法跟别人多么多么不一样。但我完全不这么看。我觉得我的想法才是很平常的那种。是其他所有人想得有点怪。
便签笔记
4:24
>> [laughter] >> And I mean that like, you know, it's just the recent times people are thinking weird. If you go look back what what what what people thought about the mind for, you know, even just a decade, you'll find the kind of thoughts that, you know, learning is important. You've got to have a goal. Um and you know, perception is important. We have a We are We are low-level We are low-level beings. We are generating actions and perceiving data at a fast speed and yet we have to think at higher levels. And you know, go back a few before there was all this AI craziness, uh you talk about you wouldn't have to say continual learning cuz it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual.
>> [笑声] >> 我的意思是,你知道,只是最近这些年人们才想得很怪。如果你回头看看以前人们对心智的看法,哪怕只往回看十年,你会发现那些想法是:学习很重要。你得有一个目标。嗯,还有,感知很重要。我们是……我们是低层次的存在。我们以很快的速度产生动作、感知数据,但我们又必须在更高的层次上思考。你知道,回到这一波AI狂热之前,你根本不用说“持续学习”,因为谈论一种不持续的学习是没有意义的。所有学习都是持续的。
便签笔记
5:08
You know, you It's not a special phase. We always act and we learn. That's just the normal way of thinking. I'm not weird. The field is weird. The field they need to call it continual learning. It's just learning. >> I'm not weird, everybody else is. That's a good good motto to live by. >> We're going to have to send out a an ex post about that. >> [laughter] >> We're very happy that you you you lived on with the field is happy that you lived on and thank you for pushing the frontier of AI. >> happy.
你知道,它不是一个特殊阶段。我们一直在行动,也一直在学习。这就是正常的思维方式。我不怪。是这个领域怪。这个领域非要管它叫持续学习。它就是学习。>> 我不怪,是其他所有人怪。这个人生信条不错。>> 我们得就这个发条推文。>> [笑声] >> 我们非常高兴你活了下来,整个领域都为你活下来感到高兴,谢谢你推动了AI的前沿。>> 高兴。
便签笔记
5:37
>> [laughter] >> Thank you for pushing the frontier of AI. I'm sure really happy and thank you for all of that and push you've like been able to sort of educate a lot of students who pushed the frontier as well. How did you pick them? How did you over the last 20, 30 years? >> Oh, well, you are giving me opportunities to be uh to be humble. Like I like to be humble and point out how all these great decisions are are just happen. And it's that that's the way I feel about students. I don't feel that I choose them very well. I've just Sometimes I'm lucky, sometimes I'm unlucky.
>> [笑声] >> 谢谢你推动AI前沿。我真的很高兴,谢谢你所做的这一切和你的推动,你还培养了很多同样在推动前沿的学生。你是怎么挑他们的?过去二三十年里你是怎么挑的?>> 哦,你这是在给我机会表现谦虚。我喜欢谦虚,喜欢指出所有这些伟大的决定其实都是碰巧发生的。对学生我也是这种感觉。我不觉得自己挑学生挑得多好。我只是……有时候运气好,有时候运气不好。
便签笔记
6:10
I don't feel I'm particularly good at picking my students. I'm looking at Karam. I think sometimes you end up with the really great ones. David picked David Silver picked me. >> Yeah. >> How is it that that I got you, Karam? >> Yeah, that was also so I finished my master's not with you uh and I was planning to join industry. And then we were collaborating on a project which also just had started organically. Like there was something I worked on that Rich was in a meeting, then they mentioned that I worked on it.
我不觉得自己特别擅长挑学生。我看着Khuram呢。我觉得有时候你就是会遇上特别出色的。David挑了我——David Silver挑了我。>> 是的。>> Khuram,我是怎么招到你的?>> 是啊,我硕士不是跟你读的,当时我打算去业界。然后我们在一个项目上合作,那也是很自然而然开始的。就是我做过的一些东西,Rich在一个会上,然后有人提到我做过这个。
便签笔记
6:40
So I got pulled into it. We started collaborating. It went really well. Like I felt so happy with that collaboration. Rich also felt really good about it. And then 6 months down the road we had made some progress and it just made sense to convert that into a thesis proposal. So at no point did I apply, at no point did I ask should you be my PhD advisor. We worked together, then we decided this would be a pretty good thesis. And then then after that I applied for the PhD. >> Life works in unexpected ways.
所以我就被拉了进去。我们开始合作,进展非常顺利。我对那次合作感觉特别好。Rich也觉得很不错。然后半年后我们取得了一些进展,把它转成一个论文选题就顺理成章了。所以我从来没有申请过,也从来没问过你能不能当我的博士导师。我们一起做事,然后我们决定这会是一个相当不错的论文题目。之后我才去申请了博士。>> 生活总是以意想不到的方式展开。
便签笔记
03《苦涩的教训》的精髓与 LLM 的双面性
7:10
Uh take us to 2019. You wrote the bitter lesson which has become the the mother in tome. 2019 was a funny time to be writing that piece because ImageNet was 2009. AlphaGo was 2015. What caused you in 2019 to reflect and and to write that? Because it was before the current kind of scaling paradigm around large language models had taken off, but it was after deep learning had really proven itself. >> Well, it was a long time coming. You know, as the bitter lesson expresses, it's something that you can for a long time, for many decades.
嗯,说说2019年。你写了《苦涩的教训》,它已经成为一部圣经般的文本。2019年写那篇文章是个挺有意思的时间点,因为ImageNet是2009年,AlphaGo是2015年。是什么让你在2019年去反思并写下它?因为那时候围绕大语言模型的当下这波扩展范式还没起飞,但深度学习已经真正证明了自己。>> 嗯,这是酝酿了很久的。你知道,正如《苦涩的教训》所表达的,那是一件你长期以来、几十年里都能看到的事。
便签笔记
7:45
And it's definitely at least as much due to the round of symbolic AI, which I lived through. It's all about not getting distracted by trying to put in your human knowledge and just paying attention to what the problem needs and how you can scale with computation. I know I I I wrote versions of it at least a year before and I I gave talks. I gave a talk a year before. And um it wasn't a particular response to the moment. It was a particular response to my my long experience, different people trying to think in different ways about how you can make smart systems.
而且它至少同样多地源于我亲身经历过的那一轮符号主义AI。核心就是不要被“往里塞人类知识”这件事分散注意力,而是关注问题真正需要什么,以及你如何能随算力扩展。我知道我至少在那之前一年就写过它的各种版本,我还做过演讲。我提前一年做过一次演讲。嗯,它并不是对某个特定时刻的回应。它是对我长期经历的回应——不同的人用不同的方式思考如何造出聪明的系统。
便签笔记
8:28
>> Mhm. What is the essence of the bitter lesson? >> The essence of the bitter lesson. >> You know, and maybe the phrase that I hear used the most in my meetings these days is is bitter lesson pill, is it not bitter lesson pill? I would imagine given the popularity of the phrase it's probably been tortured and misused in different ways that you didn't originally intend it. So, what do you think what is the essence of it and where do you think people go wrong in their in their attempt to understand it?
>> 嗯。《苦涩的教训》的精髓是什么?>> 《苦涩的教训》的精髓。>> 你知道,我最近开会时听得最多的一个说法就是“bitter lesson pill(苦涩教训药丸)”,是不是叫bitter lesson pill?我猜以这个说法的流行程度,它大概已经被各种曲解和误用了,那些都不是你原本想表达的。所以你觉得它的精髓是什么?你觉得人们在理解它的时候错在哪里?
便签笔记
8:53
>> Yeah, you're you're making me think about X now and my I recently made a post where I tried to do the the bitter lesson in in 26 words. It [laughter] goes something like don't be distracted by human knowledge as AI traditionally has been many times. Instead focus on learning methods that will scale with computation like search and like learning. So, it's really all about uh focusing on algorithms and improvements. It's not it's not saying you don't need fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with scale with computation.
>> 是啊,你让我想到X了,我最近发了一条帖子,想用26个词把《苦涩的教训》讲清楚。它[笑声]大概是这样:不要像AI传统上一次次做的那样,被人类知识分散注意力。相反,要专注于那些能随算力扩展的学习方法,比如搜索,比如学习。所以它真正讲的是专注于算法和改进。它不是说你不需要精巧的算法。你需要精巧的算法,但你想要的是那种能随算力扩展的精巧算法。
便签笔记
9:40
>> Rather than scaling with data. >> Rather than scaling with human input. Yeah, and then the the question if I can anticipate it um Yeah, what about large language models? >> Yeah. >> Are they >> consistent or inconsistent with your >> Yeah. >> And and I thought about this and I think there's an ex- there's another exposed model, but the conclusion is that it's both a a positive example and a negative example of the big lesson. First uh large language models enabled uh enormous scaling with computation. And you could just drink in the internet and scale so much.
>> 而不是随数据扩展。>> 而不是随人类输入扩展。是啊,然后那个问题,如果我能预料到的话,嗯,是啊,那大语言模型呢?>> 是的。>> 它们跟你的观点 >> 是一致还是不一致 >> 是啊。>> 我想过这个问题,我觉得还有一条推文说过这个,但结论是它既是苦涩教训的正面例子,也是反面例子。首先,大语言模型实现了随算力的巨大扩展。你可以把整个互联网喝下去,扩展到非常大的规模。
便签笔记
10:20
So it was a it was a way of getting uh much more capable system just by methods that scale. Then after that, as you go on further um it eventually gets limited by by that information. The the internet is finite and it's hard to get more examples. And uh the world is big and the world is massively bigger than everything we stored on the internet. And so in the end, it seems like it could be uh I guess that would be a positive example of when, you know, we relied too much on human knowledge and it eventually holds us back.
所以它是一种仅靠可扩展的方法就获得强得多的系统的途径。但在那之后,随着你继续往前走,嗯,它最终会受限于那些信息。互联网是有限的,很难获得更多样本。而且世界很大,世界比我们存在互联网上的一切都要大得多。所以最终,看起来它可能会是——我想那会是一个正面例子的反面——你知道,我们过度依赖了人类知识,而这最终会拖住我们。
便签笔记
04合成数据之争与大世界假说
10:59
>> Mhm. Can I just push on this a little bit? >> Yeah. >> It seems like a lot of what the foundation model labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human internet. Um is synthetic data generation kind of as part of this LLM scaling paradigm, is that a general method that leverages computation? >> No, that's that's just a big mistake. >> Why? >> [laughter] >> Well, it's such a big it's such a maybe it's the next the next big lesson.
>> 嗯。我能就这一点再追问一下吗?>> 可以。>> 现在基础模型实验室们做的很多事情似乎都是合成数据生成,目的就是让我们越过现有人类互联网这块“化石燃料”。嗯,作为这波LLM扩展范式一部分的合成数据生成,它算不算一种能利用算力的通用方法?>> 不算,那就是个大错误。>> 为什么?>> [笑声] >> 呃,这事儿太大了,也许它就是下一个苦涩的教训。
便签笔记
11:33
Um it's been floating around uh Alberta for 5 or 10 years. >> Okay. >> And uh we call it the big world perspective or big world hypothesis. Khuram, who eventually wrote it up as a paper. There's a little paper called the big world hypothesis. >> So, the big world is that the world is infinitely big. There are infinitely many things to learn. And you can have people generating the synthetic data sets, but there will always be more things to learn. And because of that, if you could just learn from experience, if you could remove the humans from the loop, then you would have systems that can do everything. Because, you know, the world is big. There are many tasks that we want them to do.
嗯,这个想法在阿尔伯塔已经流传了五到十年。>> 好的。>> 呃,我们把它叫做“大世界视角”或者“大世界假说”。Khuram后来把它写成了一篇论文。有一篇小论文就叫《大世界假说》。>> 所谓大世界,就是世界是无限大的。要学的东西有无限多。你可以让人去生成合成数据集,但永远都会有更多东西要学。正因为如此,如果你能直接从经验中学习,如果你能把人类从回路中拿掉,那你就会有能做任何事的系统。因为你知道,世界很大。我们想让它们做的任务有很多。
便签笔记
12:17
And they would be able to do anything by learning from their experience. Going back to the synthetic data question, too. Uh who decides what's a good synthetic data and what's a bad synthetic data? Because I can write a program that can output a lot of synthetic data, which would hurt programs. Right now, I would say humans decide. And that's the bottleneck where okay, you can have humans deciding how to generate these data sets, but you need human experts who know what's a good data set and what's a bad data set for that approach to scale. So, it is bottlenecked by humans.
而它们能够通过从自身经验中学习去做任何事情。再回到合成数据的问题,嗯。谁来决定什么是好的合成数据、什么是坏的合成数据?因为我可以写一个程序,输出一大堆合成数据,那反而会损害程序。就目前来说,我会说是人类在决定。这就是瓶颈所在——好吧,你可以让人类来决定怎么生成这些数据集,但你需要懂行的人类专家来判断什么是好数据集、什么是坏数据集,这种方法才能扩展。所以它被人类卡住了。
便签笔记
12:45
>> Doesn't my loss curve decide like how much better did I get with this data set versus that better data set? >> Right. But, if all the engineers open AI and tropic or all the big new labs and engineers went on vacation, who would generate the synthetic data? That's the question. It doesn't doesn't come from agent's experience. It's not something that the agent is generating itself. Some human has to decide what is the right synthetic data to generate. And that requires human expertise. So, for example, if you want a system to do something very challenging from a physics point of view. Maybe you want a drone that flies with echolocation, like a bat, for example. Um what's the right synthetic data for that? I think you would need to hire domain experts to go figure out what is the right data and generate it and then maybe you would be able to learn from that. But the domain expert has to exist first.
>> 难道不是我的损失曲线来决定吗——比如用这个数据集比用那个更好的数据集,我提升了多少?>> 对。但是,如果 OpenAI、Anthropic 这些大型新实验室的工程师们都去度假了,谁来生成合成数据呢?这就是问题所在。它不是来自智能体的经验,也不是智能体自己生成出来的。得有人来决定什么才是该生成的正确的合成数据。而这需要人类的专业知识。比如说,如果你想让一个系统做一件从物理角度看非常有挑战性的事情,比如你想要一架能靠回声定位飞行的无人机,像蝙蝠那样。那么,什么才是对应的合成数据呢?我觉得你得去雇领域专家,让他们搞清楚什么才是正确的数据,把它生成出来,然后你也许才能从中学习。但前提是这个领域专家得先存在。
便签笔记
13:35
So, we are bottlenecked by human expertise at that point. >> But you can have infinite synthetic worlds. The existing world is finite. >> But let's go back to the echolocation thing, right? That's what I want. I want a drone that can uh localize itself and move with echolocation. That's my goal. Um the robot that that's a robot that's generating its own experience. So, it could totally learn from its own experience, but it wouldn't be able to It doesn't matter how much synthetic data you generate. Doesn't matter if you generate synthetic data that captures 50 different universes, it will not be not allow you to do that task without humans figuring it out first.
所以到那一步,我们就被人类的专业知识卡住了。>> 但你可以拥有无限多的合成世界。而现有的世界是有限的。>> 但我们还是回到回声定位那个例子,对吧?那就是我想要的。我想要一架能靠回声定位来自我定位、并且移动的无人机。这就是我的目标。嗯,这个机器人——它是一个能生成自身经验的机器人。所以它完全可以从自己的经验中学习,但它没法……不管你生成多少合成数据都没用。哪怕你生成的合成数据涵盖了 50 个不同的宇宙,只要没有人先把这件事搞明白,它就没法完成那个任务。先把这件事搞明白。
便签笔记
14:12
>> I first just say it's it's the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, cuz this is going to be a little program that will generate the synthetic data, it'll be a very It'll be a small world. >> Mhm. >> So, for example, what's important to me is what's going on in your mind right now. Okay? And why You're saying, "Why don't I get some synthetic data to tell me what's going on in other people's minds?" No, there's no way we can have synthetic data for other people's minds.
>> 我首先想说的是,合成数据本身就是错的。我的意思是,它不会是正确的。它会是一个合成的世界,而不是真实的世界。而这是有影响的。这个世界复杂得难以置信。如果你写一个小程序——因为生成合成数据的无非就是一个小程序——那它会是一个非常……那会是一个很小的世界。>> 嗯哼。>> 举个例子,对我来说重要的是你此刻脑子里在想什么。对吧?而你会说:“我为什么不弄点合成数据来告诉我别人脑子里在想什么呢?”不,我们不可能有关于别人心智的合成数据。
便签笔记
14:53
And other people's minds matter to us. You know, like I talked to you guys about investing today, so I care what you going on in your minds. And how how can I get synthetic data on such a thing? Really, you can't even get synthetic data on anything. You can't get synthetic data on on the how the drone is going to interact with with its environment in the physical world and the the the the friction and where in the motors of this of this robot. The world is infinitely complex, and any simulation of it is like microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agents any agent.
而别人的心智对我们来说很重要。你看,我今天跟你们聊投资的事,所以我很在意你们心里在想什么。那我怎么可能拿到这种东西的合成数据呢?说真的,你几乎什么都拿不到合成数据。你没法得到关于无人机将如何与物理世界中的环境交互的合成数据,还有摩擦力、这个机器人马达内部的情况等等。这个世界是无限复杂的,任何对它的模拟都只是微不足道的一点点。所谓“大世界假说”,说白了就是:世界远比你的心智复杂得多,比任何智能体都复杂得多。
便签笔记
15:37
And this is obvious because the world contains many other agents. So, because the world is is massively complex, as you could no way you can do anything like anything that might claim to be optimal or perfect. You're going to be imperfect, and you have to have approximations, and those approximations will will be severe. And so, because of that, that is the ultimately the reason why we have to continue learning, if you want to think of it as a reason. We have to continue learning because we'll encounter some particular part of this immense world, and we'll have to learn an approximation that's tuned to the part of the world we're in, not to the all the other parts that we're not in.
这一点是显而易见的,因为世界里还包含着许许多多其他智能体。所以,正因为世界极其复杂,你根本不可能做出任何号称最优或完美的东西。你注定是不完美的,你必须依靠近似,而这些近似还会是相当粗糙的。正因为如此,这归根结底就是我们必须持续学习的原因——如果你想把它当作一个理由的话。我们必须持续学习,因为我们会遇到这个庞大世界中某个特定的局部,我们必须学到一个针对我们所处的那部分世界调校过的近似,而不是针对我们没身处其中的其他所有部分。
便签笔记
16:19
>> Yeah. I'm I'm going to push on this one more time. Um, and sorry, I'm being argumentative for the sake of being argumentative, [clears throat] but I'm trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim, and then they, you know, have to do some some sort of post-training, I guess, to to make sure they work in the real world, but but that it's been a very effective pipeline. >> Yeah. So, I think the important question to ask here is, how many engineers were involved in building that simulation, and are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation. And I'm sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a sim-to-real gap that was not acceptable, then they fixed it. So, there is this human in the loop fixing the simulation.
>> 好。我想再追问一次。抱歉,我有点为了抬杠而抬杠,(清嗓子)但我是想弄明白。据我了解,最新一批自动驾驶公司里,很多家主要是在仿真里训练的,然后再做某种后训练,我想是为了确保它们在真实世界里能работать——能正常工作,但这条路线一直非常有效。>> 是的。所以我觉得这里该问的关键问题是:建这套仿真系统投入了多少工程师?我们是否准备好承认:唯一值得解决的问题,就只有那些我们能雇一大批工程师先做出一套仿真的问题?而且我敢肯定他们必须反复迭代好几轮——先做出仿真,在里面训练,然后发现有个无法接受的仿真到现实的差距(sim-to-real gap),接着再去修。所以这里有一个“人在回路”里在修仿真。
便签笔记
17:10
Like, they're getting feedback from the real world, humans, and then they're fixing the simulation. Why can't we just remove the human and let the agent do it itself. >> And then when it actually drives, again, something unexpected will happen. >> And that's when you really want to learn from experience. >> So, your point is there's just so much more data that's going to come from experience than there possibly can be from humans curating and creating data. >> Yeah, and think there is a there is obviously value in learning from simulation. And and there is a way of doing it. The agent can learn a model from its own experience. And when the agent learns it, it's very it's much better because if the model is incorrect, it can fix it by continuous learning.
就是说,他们从真实世界拿到反馈,人来处理,然后去修仿真。那我们为什么不能干脆把人去掉,让智能体自己来做这件事呢?>> 而且等它真正上路开的时候,还是会出现意料之外的情况。>> 而那正是你真正需要从经验中学习的时候。>> 所以你的意思是,来自经验的数据量,远远超过人类去筛选和创造所能提供的数据量。>> 是的,而且我认为从仿真中学习显然是有价值的,也确实有办法做到。智能体可以从自己的经验中学到一个模型。而当智能体自己学到这个模型时,效果要好得多,因为如果模型不正确,它可以通过持续学习来修正。
便签笔记
05先验与学习不该为敌
17:49
If the humans are making a simulator, then the model only gets updated when the humans figure out that something is wrong. So, yes, planning is important. The agents should learn from uh simulators, but simulators they make themselves. >> Okay, I want to move to another part of the bitter lesson, removing human knowledge. From your essay, quote, "Seeking the improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain. But the only thing that matters in the long run is the leveraging of computation."
如果仿真器是人做的,那模型就只有在人发现哪里出了问题的时候才会被更新。所以,是的,规划很重要,智能体应该从仿真器中学习,但那必须是它们自己造的仿真器。>> 好,我想聊聊《苦涩的教训》的另一部分:去掉人类知识。引用你文章里的话:“为了寻求在较短期内产生效果的改进,研究者往往试图利用自己在该领域的人类知识。但从长远看,唯一重要的是对算力的利用。”
便签笔记
18:17
And if I, you know, your former student Dave, uh with AlphaGo and AlphaZero, for me that was a an example of a triumph of removing human priors. Did that result surprise you? I guess why or why not? >> Of course, it made me very happy. It made me, you know, feel vindicated. Um you know, it could have gone either way. It wasn't that uh cuz cuz prior knowledge can help. You know, there's nothing wrong with prior knowledge. You know, and and I say this right at the very beginning of the right at the very beginning of the bitter lesson, I say there's no reason why there has to be a conflict between prior knowledge and then learning knowledge. You know, you can put some prior in there and then start learning.
那么,你以前的学生 Dave(大卫)做的 AlphaGo 和 AlphaZero,在我看来就是一个去除人类先验的胜利案例。那个结果让你意外吗?为什么意外或者为什么不意外?>> 当然,那让我非常高兴。让我觉得自己被证明是对的。不过你知道,事情本来也可能是另一种走向。并不是说……因为先验知识是可以起作用的。先验知识本身没什么不好。我在《苦涩的教训》开头就说过——就在文章最开始——我说没有任何理由说先验知识和后天学到的知识之间必然存在冲突。你完全可以先放一些先验进去,然后再开始学习。
便签笔记
18:59
There's no reason in principle why these have to be opposed. In fact, they're all about knowledge. You know, life is gaining knowledge and having knowledge. And why are these how somehow, you know, nature and nurture became enemies? but really you know, prior learning is what you already and then and you then you learn more and it's they should be friends. But, as I say at the beginning of the bitter lesson, in practice they have been enemies. In practice, people who who had a an affection for existing human knowledge ended up, you know, wanting that to win and so they wanted to to minimize or or dismiss learning.
原则上没有理由说这两者必须对立。事实上,它们讲的都是知识。要知道,生命就是在获取知识、拥有知识。那为什么这两者……先天与后天怎么就成了敌人呢?可实际上,先前学到的东西就是你已经掌握的,然后你再学更多,它们本该是朋友。但正如我在《苦涩的教训》开头说的,实际上它们一直是敌人。现实中,那些对既有人类知识怀有偏爱的人,最后会希望那一边赢,于是他们就想弱化甚至否定学习。
便签笔记
19:38
And so now, I'm sure your sense of me is that I'm someone who who loves learning and wants to dismiss prior knowledge. Um but you know, I'm really someone who's interested in the mind. The mind is you have prior knowledge and then you get more and then once you've gotten more, then that becomes your prior knowledge as you get more and more and more. And this these two things work together. I end up appearing to be someone who's who's interested in learning primarily because all the rest of the world is is is is talking about all you need is enough knowledge. You don't need to learn.
所以现在,我猜你对我的印象是:我是那种热爱学习、想要否定先验知识的人。但其实,我真正感兴趣的是心智。心智就是:你先有一些先验知识,然后获得更多,一旦你获得了更多,那些就变成了你新的先验知识,如此不断累积下去。这两者是协同工作的。我之所以最后看起来像是主要关注学习的人,是因为世界上其他所有人都在说:你只需要足够的知识,你不需要学习。
便签笔记
20:15
You know, large language models are we're going to put all this knowledge in the into the system and the large language model will not learn when it runs. You know, you know, it's talking to people, it's interacting. It is absolutely the weights never change. So, you know, I am not the weird one. It's you guys that are the weird one that think that that's possible that you could possibly, you know, they claim they can like a PhD level experience and expertise out of something that doesn't learn at all anymore.
你看,大语言模型就是——我们要把所有这些知识塞进系统里,而这个大语言模型在运行时是不会学习的。它在跟人对话、在交互,但它的权重绝对是一成不变的。所以你看,怪的不是我。是你们这些认为“那样也行得通”的人才怪,他们声称能从一个已经完全不再学习的东西里,得到博士级别的经验和专业能力。
便签笔记
20:44
You know, so you know, I'm not the weird one. >> [laughter] >> So, your recommendation is let it let the algorithms run for much much longer period of time before feeding >> Continually learn. >> before you feed it data. with prior data. So, drip or drip the prior data along the way. prior knowledge. >> So, both are important. Um but in the long run, you've got to gain and structure the gaining of new knowledge. That's that's what all that matters in the long run. And as you are doing this, yeah, there'll be some that you had got previously.
所以说,怪的真不是我。>> (笑)>> 所以你的建议是,让算法先跑很久很久,然后再喂给它>> 持续学习。>> 在你喂数据之前。用先验数据。所以是把先验数据一点一点地滴进去。先验知识。>> 两者都重要。但从长远看,你必须获取新知识,并把获取新知识这件事组织好。长远来看,重要的只有这个。而在你这么做的过程中,是的,会有一些是你之前就已经获得的。
便签笔记
21:22
Um, like how it would work if we, you know, look into the future when we have intelligent robots. We will we will will we uh, have them all learn from scratch? Or will we like copy them and ask them to keep learning from wherever they are? I mean, we they'll be digital and it'll be easy to copy them. And so, instead of having like this huge thing where we're spending zillions of dollars to retrain them from the internet, we'll just copy the agent and keep learning from there. And and so, in some sense, the prior knowledge will be should be dismissed cuz you're just going to copy it from the previous robot.
比如说,我们可以设想一下未来有了智能机器人之后会怎么运作。我们会让它们全都从零开始学吗?还是我们会直接复制它们,让它们从当前的状态继续学下去?我是说,它们是数字化的,复制起来很容易。所以,与其搞那种花上天文数字的钱、从互联网上重新训练一遍的大工程,我们不如直接复制那个智能体,从那儿继续学下去。所以在某种意义上,先验知识应该被搁到一边,因为你只要从上一个机器人那里复制过来就行了。
便签笔记
06权重不变,就不算学习
22:02
>> So, why don't you describe for us what you think a machine or computer that learns from experience looks like? >> Well, it could look like a robot. It could be It also could be live entirely on the internet. You could like, for example, routing a package through the internet and do that in a way that's sensitive to experience and and becomes better over time. Or you can interact via the user interface that's interacting with people, like on your phone or on your computer, and uh, it becomes better over time. Yeah, like an intelligent assistant, you know, has to become better over time. It has to know what you want.
>> 那你能不能给我们描述一下,你认为一台从经验中学习的机器或计算机是什么样子的?>> 嗯,它可以是一个机器人的样子。也可以完全生活在互联网上。比如说,在互联网上路由一个数据包,并且以一种对经验敏感的方式来做,从而随着时间推移变得越来越好。或者它可以通过用户界面与人交互,比如在你的手机上或电脑上,然后随着时间推移变得越来越好。是的,就像一个智能助手,它必须随着时间变得更好。它得知道你想要什么。
便签笔记
22:37
>> Would your contention be that the current paradigm of, you know, the popular A L M based assistants, would your contention be that these are not experiential learners or continual learners? And if so, what is the fundamental gap? >> Are you serious? >> [laughter] >> I mean, obviously they >> They they have memo- they learn memories about me. They're you know, they're they're they're doing some in-context learning. >> Their weights never change. >> And and by the way, is a small number of the weights changing sufficient or do you need all the weights to be changing?
>> 你的观点是不是说,当前这种流行的基于大语言模型的助手范式,你是不是认为它们并不是从经验中学习的学习者,也不是持续学习者?如果是的话,根本性的差距在哪里?>> 你是认真的吗?>> (笑)>> 我是说,显然它们……>> 它们有记忆——它们会记住关于我的事情。它们,你知道的,它们在做某种上下文内学习(in-context learning)。>> 它们的权重从来不变。>> 另外顺便问一句,只有一小部分权重发生变化就够了吗?还是需要所有权重都在变化?
便签笔记
23:10
>> Well, so all think of all the structuring and generation of new concepts that went into creating the large language models. All that is the weight learning. The and you you want to continue be able to continue doing that. You don't want that to happen just once. >> Is another way of saying it is we do too much pre-training and post-training before we launch the the models. They they don't learn after that. >> The only point that the big disagreement is we don't let them learn after that. >> Yeah, we don't let them learn after that.
>> 嗯,想想为了创造出大语言模型,投入了多少新概念的构建和生成。所有这些都属于权重学习。而你希望能够继续做这件事。你不希望它只发生一次。>> 换一种说法就是,我们在发布模型之前做了太多预训练和后训练。它们在那之后就不再学习了。>> 唯一有大分歧的地方是,我们不让它们在那之后继续学习。>> 是的,我们不让它们在那之后继续学习。
便签笔记
23:40
>> as much pre-training as we want, that's okay. Post-training is fine, but then when I'm when I'm using the model I cannot it's it's it stops learning. Uh you can give it more context. You can change the state of the model by giving it more context. And so it it already learned that if the state is different, if the state says something new, then it will use that to make the next prediction. But the model is not learning. >> Cursor's tab other complete model. It is, you know, it does get updated based on >> Those models those weights change.
>> 想做多少预训练都行,没问题。后训练也没问题,但当我在使用这个模型时,我没法……它就停止学习了。呃,你可以给它更多上下文。你可以通过给它更多上下文来改变模型的状态。所以它其实已经学会了:如果状态不同,如果状态里出现了新的信息,它就会用这些信息来做下一步的预测。但模型本身并没有在学习。>> Cursor 的 tab 补全模型,它确实会根据……来更新 >> 那些模型的权重是会变的。
便签笔记
24:09
>> weights change. Those are two examples like Cursor's tab and I think the composer they were also updating. Those are two examples of continual learning. >> Okay. >> it's can be much better. So the way they do it as far as I understand is a lot of people are using tab, they collect all this data so coming from millions of users or thousands of users and then they do one update of the policy from this batch data. Um so now this could work, but imagine I want to teach this model something specific. I don't want to fight with 100,000 other peoples about what they want to treat teach their models. I want to teach my model something very specific and I want to do it to my version of the model. I like I don't care about the shared knowledge that the model has coming from other people.
>> 权重会变。这就是两个例子,比如 Cursor 的 tab,我想 composer 他们也在更新。这就是持续学习的两个例子。>> 好的。>> 其实可以做得好得多。据我理解,他们的做法是这样:很多人在用 tab 补全,他们把所有这些数据收集起来,来自几百万或者几千个用户,然后用这批数据对策略做一次更新。嗯,所以这是可行的,但想象一下我想教这个模型某个特定的东西。我不想跟另外十万个人去争,去争他们各自想教自己模型的东西。我想教我的模型一些非常具体的东西,而且我想把它教给我自己那个版本的模型。我并不在乎模型从其他人那里获得的那些共享知识。
便签笔记
24:48
And so it's a very inefficient way of doing it. >> Mhm. It seems like the way that this is currently done is that there's fundamental skills maybe that are learned in the weights that are common to everybody. And then there's personalization that happens in the form of context, right? >> Yeah. >> Is that not the right mental model for how learning should work? Like should should all the context live in the weights themselves? >> So, context can be in the state, too. Could be both. But, you still need to be able to update the weights. So, if I give you an example, some really good use studies are for with human disabilities. When human go through something that changes their mind or or some sensors, you can see them adapt.
>> 所以这是一种非常低效的做法。>> 嗯。看起来目前的做法是,有些基础能力可能是学在权重里的,是所有人共通的。然后个性化是以上下文的形式发生的,对吧?>> 是的。>> 这个心智模型是不是不对?学习本该是什么样的?是不是所有的上下文都应该存在权重本身里?>> 所以,上下文也可以存在于状态里。两者都有可能。但你仍然需要能够更新权重。举个例子,一些非常好的研究案例来自于人类的残障情况。当人经历了某些改变心智或某些感官的事情时,你能看到他们逐渐适应。
便签笔记
25:28
So, for example, we have proprioception, we have internal sensors that tell us where the how the body is positioned, and we use this for walking. There are cases where people lose this ability completely, and then they can't walk at all because that is literally the foundation of their walking policies. It is ingrained in the brain. But then over the course of 2 3 years, they can learn to walk again by looking at their feet. So, visual feedback through that. So, brain is insanely plastic in the sense that it it can learn a lot of things. Something that has been true for 20 years, when it stops being true, it can go and update that and get rid of that. And that is the capability I think that's extremely useful we would want in our systems.
比如说,我们有本体感觉,我们体内有传感器告诉我们身体处于什么位置、姿态如何,我们走路时就靠这个。有些情况下,人会彻底失去这种能力,然后他们就完全没法走路了,因为这实在是他们行走策略的根基,它是刻在大脑里的。但在接下来的两三年里,他们可以通过看着自己的脚重新学会走路。也就是靠视觉反馈来完成这件事。所以大脑的可塑性强得惊人,它能学会非常多的东西。有些事情已经成立了二十年,当它不再成立时,大脑可以去更新,把它抛弃掉。我认为这种能力是极其有用的,是我们希望在自己的系统里拥有的。
便签笔记
07没有动物靠监督学习
26:09
>> Hm. What is there for us to learn from how human babies or animals learn? And how much inspiration do you take from that? >> Well, we take uh a lot of inspiration. We don't take it as a requirement that the AI has to behave like uh the natural system, like babies or people or animals. Um but it's it's a source of inspiration. Inspiration, but not constraint from animal learning. >> Consistent with the better lesson? >> Yeah. >> [laughter] >> Yeah. >> Where do you think we should most seek to draw inspiration from the way that biological learnings works that is not present in today's systems?
>> 嗯。人类婴儿或者动物的学习方式,有什么值得我们借鉴的?你从中汲取了多少灵感?>> 嗯,我们确实汲取了很多灵感。我们并不要求 AI 必须表现得像自然系统那样,比如像婴儿、人或者动物。但它是一个灵感来源。是灵感来源,但动物学习并不构成约束。>> 这和「更好的教训」(better lesson)是一致的吗?>> 是的。>> [笑声] >> 是啊。>> 你觉得在生物学习的机制里,有哪些是今天的系统里没有的、最值得我们去汲取灵感的?
便签笔记
26:48
>> I feel like I'm just giving opinions now, but they're just obvious opinions. So, so I think it's apparent that no animal learns by supervised learning. Because we don't get examples of how our muscles should twitch. And that's our output. >> But all of school is supervised learning. >> I I know. Absolutely not. Uh, but even if it was, school is like a tiny fraction of what we learn. Like we learn to see, we we learn to walk. And we learn, um, how the world works. But even Yeah, and even in school, you know, no one tells us how we should twitch our muscles.
>> 我感觉我现在只是在发表观点,不过都是些显而易见的观点。我觉得很明显,没有任何动物是靠监督学习来学习的。因为我们不会得到关于自己肌肉该怎么抽动的样例。而那才是我们的输出。>> 可整个学校教育不就是监督学习吗?>> 我知道。绝对不是。呃,就算它是,学校也只占我们所学内容的极小一部分。比如我们学会看东西,我们学会走路。我们学习,嗯,世界是怎么运作的。但就算是在学校里,你知道,也没有人告诉我们该怎么去抽动我们的肌肉。
便签笔记
27:31
>> The knowledge skills I acquire were were from supervised learning in school. >> don't want to say that that that learning from from others, transmission from others, is not important. It's like extremely important. And language is extremely important. Um, but what are we what are we missing? You know, there is there is no supervised learning. There's no targets that are given to us. You know, you you hear the right answer is, you know, where where is what's the capital of France? And we know the answer is Paris.
>> 我掌握的知识和技能是在学校通过监督学习获得的。>> 我不想说从别人那里学习、从别人那里传递知识不重要。它极其重要。语言也极其重要。嗯,但我们缺了什么呢?你知道,其实并不存在监督学习。没有人给我们目标答案。你会听到正确答案,比如说,法国的首都是哪里?我们知道答案是巴黎。
便签笔记
28:05
Okay, but no one tells me how I should pronounce Paris. You say the answer is Paris, and I listen to you, and I hear your words, and, you know, I will make some other uh, muscle motions to produce the answer Paris. It's not literally supervised learning. Um, anyway, yeah. So, I think it's really true. I mean, well, anyway, the first thing is a school is is irrelevant. Like, you know, squirrels don't go to school and and and learn [laughter] that. >> They might. >> Animals don't learn that way. It's And school is a very special thing that that even we didn't have up until, you know, I don't know, a few hundred years ago.
好吧,但没有人告诉我该怎么发出「巴黎」这个音。你说答案是巴黎,我听着你说,我听到你的话,然后,你知道,我会做出一些别的肌肉动作来说出「巴黎」这个答案。这并不是字面意义上的监督学习。嗯,总之,是的。我觉得这真的没错。我是说,反正首先一点是,学校是无关紧要的。你知道,松鼠可不会去上学,然后学会那些东西。[笑]>> 它们说不定会呢。>> 动物不是那样学习的。学校是一种非常特殊的东西,甚至连我们自己在过去也没有,你知道,我也说不好,也就几百年前才有。
便签笔记
28:46
It's not part of in not part of the essence of intelligence? And it's a distraction to think of that as your primary example of learning is this thing which we didn't do as animals. >> I wish you had been around to tell my parents that before I was made to have good go to school and deal with all the structure. >> The thing is like squirrels are wonderful at jumping off trees, but squirrels can't prove math theorems. And if I want to learn how to prove a math theorem, I go to school. >> Yeah. Uh they also don't have uh DVDs and >> [laughter] >> and iPods. You know, there are a lot of things they can do things that we can't do.
它并不是智能本质的一部分?把这种我们作为动物本来并不做的事情,当成学习的主要范例,这其实是种误导。>> 真希望当年你在场,能在我被送去上学、应付那一整套条条框框之前,跟我爸妈说说这些。>> 问题是,松鼠特别擅长从树上跳下来,但松鼠证明不了数学定理。而且如果我想学怎么证明一个数学定理,我就得去上学。>> 是啊。呃,它们也没有 DVD 和 >> [笑] >> iPod。你知道,有很多事它们能做,而我们做不了。
便签笔记
29:25
Um but >> [sighs] >> math theorems uh Yeah, and they don't play chess. You know, it's sort of like more of X paradox. They're uh they're these advanced things that we think of as really intelligent. But uh they're sort of easy for computers to do as opposed to all these regular things that are hard. Like moving and seeing with attention and everything. Um I think supervised learning is a is a good thing. You know, just mentions I like to think look for obvious things. No one tells us how to twitch our muscles by giving us examples cuz they couldn't possibly cuz we have had to twitch our muscles. We've had to figure that out.
嗯,但是 >> [叹气] >> 数学定理嘛,是的,而且它们也不下棋。这有点像莫拉维克悖论。呃,有些高级的东西我们觉得非常需要智能,但对计算机来说其实挺容易的,反倒是那些日常的事情很难。比如带着注意力去移动、去看,诸如此类。嗯,我觉得监督学习是个好东西。你知道,顺便提一下,我喜欢找那些显而易见的事情。没有人通过给我们示例来教我们怎么抽动肌肉,因为他们根本做不到,因为我们必须自己去抽动肌肉。我们必须自己把这件事琢磨出来。
便签笔记
30:06
>> Yeah. >> And their answer would be wrong, right? So, if I moved my mouth and my tongue and my vocal cords exactly the same way that Rich does to pronounce Paris, I'm sure a very different sound would come out. So, in some sense Rich or no one knows the right way of producing a sound with my body. Only I know that. >> Yeah. It seems to me that many of the most, I guess the most raw like sensory motor capabilities, especially related to movement in the physical world. I agree with you that that seems something that is inherently learns from experience.
>> 是啊。>> 而且他们给的答案会是错的,对吧?所以,如果我用和 Rich 完全一样的方式来动我的嘴、舌头和声带去念「巴黎」,我敢肯定发出来的会是完全不同的声音。所以从某种意义上说,Rich 也好,谁也好,都不知道用我的身体发出声音的正确方式。只有我自己知道。>> 是啊。在我看来,很多最原始的、我想说是最基础的感知运动能力,尤其是和在物理世界中运动相关的那些,我同意你说的,那似乎本质上就是从经验中学来的。
便签笔记
08火箭与范式转变:想象能否来自经验
30:39
It seems to me though that there are higher levels of abstraction that bring us closer to, you know, what makes humans great. And much of that doesn't live in this low level of sensory motor learning. Does your world model, I guess, span sensory motor learning all the way up? >> Yeah, that's the ambition, absolutely. And squirrels, by the way, can do some enormously abstract things. >> What's the coolest thing a squirrel can do? >> Well, it can always get into your bird feeder, >> [laughter] >> no matter what obstacles you put in the way, you know, it can find new ways to jump and climb and >> Okay.
但在我看来,还有一些更高层次的抽象,它们更接近于,你知道,是什么让人类如此了不起。而其中很多东西并不存在于这种低层次的感知运动学习里。我想问,你的世界模型是不是从感知运动学习一路涵盖到最上层?>> 是的,这正是我们的目标,绝对是。顺便说一句,松鼠也能做一些极其抽象的事情。>> 松鼠能做的最酷的事情是什么?>> 嗯,它总能钻进你的喂鸟器,>> [笑] >> 不管你在路上设多少障碍,你知道,它总能找到新办法去跳、去爬,还有 >> 好吧。
便签笔记
31:14
>> and do lots of things. >> Calculate trajectories pretty well. Animals are pretty good at understanding the physical world without the mental calculations that we think we are doing when we think about launching ourselves into space. >> Breaking a fall, they can do it in real time in the right way to prevent injuries. >> Okay, fair enough. >> I think it's just a question of degree between and I like to think that animals, other animals, are are very close to humans. I think it's hubristic to try to emphasize what we do differently, you know, how we're different from animals.
>> 还有做很多别的事情。>> 弹道轨迹算得还挺准。动物很擅长理解物理世界,而不需要我们以为自己在做的那些心算,比如当我们想着把自己发射到空中的时候。>> 化解一次摔倒,它们能实时地用正确的方式做到,避免受伤。>> 好吧,有道理。>> 我觉得这只是程度上的差别,而且我倾向于认为动物,其他动物,和人类非常接近。我觉得一味强调我们有什么不同、我们和动物有什么区别,是种自大。
便签笔记
31:48
It's better to see the commonalities. And I I think we are just a question of degree. It's degree and of course society and culture give us big advantages. Language give us big advantages. >> Can I just push on some of this? >> Yeah, good. >> Because I want to back up Sonia. So, I believe animals and children learn from experience and do incredible things learning from experience. And when my son was two or three or four, I'm like, "Wow, this is really interesting that they're my son can learn these things without nobody really teaching him how to do these things."
更好的做法是看到共通之处。而且我觉得我们只是程度上的差别。是程度的问题,当然社会和文化给了我们巨大的优势。语言也给了我们巨大的优势。>> 我能就这点追问一下吗?>> 好的,请。>> 因为我想给 Sonia 补充一下。我相信动物和小孩确实是从经验中学习的,而且能通过从经验中学习做出很了不起的事情。>> 我儿子两三岁、三四岁的时候,我就想:"哇,这真有意思,我儿子居然能学会这些东西,根本没有人>> 真正教过他怎么做这些事。"
便签笔记
32:25
But at the same time, what Sonia's saying is like what makes human uniquely human, to be able to go to outer space, build a rocket. Those Those not things that are learned 100% from experience because before you launch the rocket, you actually have to abstract thinking through it in a way that is not learned from {quote} {unquote} experience. Because you don't know if it's going to work or not. You have to imagine it. How do you we teach a machine to imagine things that were not available before? That's probably the thing that we're trying to like push on because that we're not quite understanding that.
>> 但与此同时,Sonia 说的是,人之所以为人的独特之处,是能够进入太空、造出火箭。>> 这些并不是 100% 从经验中学来的东西,因为在你发射火箭之前,>> 你其实必须以某种抽象的方式把它想清楚,而那种方式并不是从所谓的"经验"中学来的。>> 因为你并不知道它到底行不行。你得先想象出来。那我们要怎么教一台机器去想象>> 此前并不存在的东西?这大概就是我们想要追问的地方,因为这一点我们��没完全弄明白。
便签笔记
33:04
>> actually going to agree with you there. You have to be able to plan. You have to be able to imagine. >> Yeah. >> Would you say that humans 1,000 years ago, before they had done all most of the thing that we're talking about, were they as intelligent? Um if for example someone from that era was exposed to this new culture, would they be able to get the same skills and and start doing useful things? >> Even over the last 10,000 years, I don't think the human brain has evolved that much. Because >> Fundamentally the same machine.
>> 这一点我其实同意你。>> 你必须能够做规划,必须能够想象。>> 是的。>> 那你会认为一千年前的人类,在他们还没做出我们现在讨论的这些事情之前,>> 他们同样聪明吗?>> 比如说,如果把那个年代的人放到今天这种文化环境里,他们能不能学会同样的技能,>> 然后开始做出有用的事情?>> 就算把过去一万年算进来,我也不认为人脑进化了多少。因为——>> 本质上是同一台机器。
便签笔记
33:33
>> Fundamentally the same machine, but we've built up 10,000 years of knowledge. >> Yes. >> And I get to learn 10,000 years of knowledge by going to school through supervised learning. >> Right. >> And I get all that much, much faster than trying to learn through experience. >> Right. >> So I think you're like totally right. So we we would want our systems to learn from experience and part of their experience would be getting exposed to our culture and then learning from about our culture. They should learn from that. That's all good. But let's talk about when someone goes and does a paradigm shifting thing. So everyone gives the example of Einstein, but I think there are many examples. Learning is that too, like looking at learning thing versus programming thing.
>> 本质上是同一台机器,但我们积累了一万年的知识。>> 对。>> 而我可以通过上学、通过监督学习,去学到这一万年的知识。>> 没错。>> 而且我获得这一切的速度,要比试图从经验中学习快得多。>> 对。>> 所以我觉得你说得完全对。我们确实希望我们的系统能从经验中学习,而它们经验的一部分>> 就是接触我们的文化,然后从我们的文化中学习。它们应该从>> 那里学习,这都没问题。但我们来聊聊有人做出范式转变的时候。大家都会举爱因斯坦的例子,但我>> 觉得例子还有很多。学习本身也是这样,比如"学习"这条路线和"编程"这条路线的对比。
便签笔记
34:14
So when these paradigm shifts happen, I would say it's a human who has accumulated all this knowledge and then from their experience they're building new abstractions and they're planning with them and then they're discovering new knowledge. And that skill of of coming up with new abstractions and then learning what models and planning with them, that problem is that skill is totally missing in our current systems. And you can expose this at the edge of human knowledge, but you can also study this problem at the sensory motor stream level.
>> 所以当这些范式转变发生时,我会说,是一个人积累了所有这些知识,然后从他自己的经验里>> 构建出新的抽象,并用这些抽象来做规划,进而发现新的知识。而这种能力——>> 提出新的抽象、学到相应的模型并用它们做规划——这个问题、这个能力,在我们现有的系统里>> 是完全缺失的。>> 你可以在人类知识的边界上看到这个问题,但你也可以在感觉运动流的层面上>> 研究这个问题。
便签笔记
34:41
>> So, we're not arguing with the with the principle. We need to form abstractions so we can reason at a high level. You guys are coming close to doing that thing that I said we should never do, which is argue is prior knowledge important or gaining knowledge important. You know, that's what you guys just said. You said It's a you're you're still going to have to learn things. And you're saying, "Oh, I can get things from my culture and from prior knowledge." But these should not fight for each other.
>> 所以我们并不是在反对这个原则。我们需要形成抽象,这样才能在高层次上做推理。>> 你们俩现在快要做我说过我们永远不该做的那件事了,>> 也就是去争论到底是先验知识重要还是获取知识重要。你看,你们刚才说的就是这个。你说这是>> 你还是得去学东西。而你说:"哦,我可以从我的文化、从>> 先验知识里获得这些。">> 但这两者不该互相对立。
便签笔记
35:10
>> On the exact thing around paradigm shifts, how do we create a machine that understands when to shift the paradigm? >> Yeah, I think through its experience, right? So, it would have to through its own experience. It can't rely on human knowledge because we're assuming the humans see one paradigm and we want a different way of looking at things. And so, through its experience, it has to find something that is better. Maybe it it generalizes better and makes better predictions. Maybe it's better in some other ways, but it has to be through its own experience.
>> 就范式转变这个具体问题而言,我们怎么才能造出一台机器,让它知道什么时候该转变范式?>> 是啊,我觉得是通过它自己的经验,对吧?所以它必须通过自己的经验来做到。它不能依赖人类知识,因为我们>> 假定人类看到的是一种范式,而我们想要一种不同的看待事物的方式。>> 所以,通过它自己的经验,它必须找到某种更好的东西。也许它泛化得更好,做出的>> 预测更准。也许它在别的方面更好,但这必须来自它自己的经验。
便签笔记
35:40
>> The big challenge that we don't see in our field, the ability we don't see in our field yet, is the ability to learn a model and then plan with the model. We can do the the math things and we can do AlphaGo because the games, we know the model. We know how the moves work. And in math, we know what the operators are. We you know, we know lean will take us from one state of knowledge to the to the state of the proof to the next state. But if we have to learn the models, there are no I'm I'm going to say it. It's probably maybe a a weird example, but a counterexample, but I can see that there's no instances of learning the model and then planning with the model in our field.
>> 我们这个领域目前还看不到的一大挑战、一项还不具备的能力,就是先学出一个>> 模型,然后用这个模型做规划。>> 我们能做数学那类事情,也能做出 AlphaGo,因为在那些游戏里,模型是已知的。我们知道每一步棋是怎么走的。>> 在数学里,我们知道算子是什么。>> 我们知道,Lean 会带我们从一个知识状态走到证明的下一个>> 状态。>> 但如果我们必须去学习这些模型,那就没有——我就直说了,这可能是个有点奇怪的例子,或者说>> 是个反例,但我看不到在我们这个领域里有任何"先学出模型、再用模型做规划">> 的实例。
便签笔记
36:23
>> At least not with uh uh, like self-discovered abstractions. So, there are people who say, "I'm just going to learn a model of what happens in the next second or next millisecond." But, that's not how our models work. Our models are more abstract. Our models are, uh, quite different. >> So, one of the things I like about what you're doing here is you're not just sitting around pontificating or lamenting the state of the world as it is. You're very action-oriented. It's why you started a company. So, let's let's start talking about that a bit.
>> 至少在自己发现抽象这一点上是没有的。所以,>> 有些人会说:"我就学一个关于下一秒或下一毫秒会发生什么的模型。">> 但我们的模型不是那样工作的。我们的模型更抽象。我们的模型>> 相当不一样。>> 所以,我喜欢你们在这里做的事情的一点是,你们不只是坐在那里空谈或者感叹世界现状如何如何。你是非常行动导向的人。所以你才创办了公司。那我们就从这个话题开始聊起吧。
便签笔记
09阿尔伯塔计划:抽象与持续深度学习
36:54
Uh, in 2022, Rich, you laid out a very specific 12-point plan, the Alberta plan for AI research. Maybe tell us about that. >> So, the Alberta plan came about because we just have general ideas, but we also needed to convert them into smaller chunks. And so, the 12 steps are the attempt to uh, crystallize particular chunks. There's a very important early step, step two, uh, which is continual deep learning. And And we think that one is like almost the most important because it unlocks everything else. If you could do continual deep learning, you could then continually update your model of the world.
呃,2022年,Rich,你提出了一个非常具体的十二点计划,也就是人工智能研究的「阿尔伯塔计划」。能给我们讲讲这个计划吗?>> 阿尔伯塔计划的由来是这样的:我们只有一些笼统的想法,但我们还需要把它们拆解成更小的模块。所以这十二个步骤就是试图把一个个具体的模块给明确下来。其中有一个非常重要的早期步骤,也就是第二步,呃,就是持续深度学习(continual deep learning)。我们认为这一步几乎是最重要的,因为它能解锁其他所有的东西。如果你能做到持续深度学习,你就能持续更新你对世界的模型。
便签笔记
37:39
And then, if you knew how to do the abstraction rights in like the second half of of the the steps are all about how to get the abstractions right. So, and not only my abstractions right, what I mean by I don't mean get the right abstractions cuz no one can say what the right abstractions are. That depends on the world that you're in. Your agent would have to learn the correct abstractions for whatever world it's in. And so, you know, if you maybe those are the two key things. You have to find the right abstractions, and then you have to do able to continual deep learning.
然后,如果你还知道该怎么把抽象做对——后半部分的那些步骤基本上都是在讲怎么把抽象做对。所以,我说的不只是「我的抽象要对」,我的意思是——我说的并不是要找到「正确的抽象」,因为没人能说清什么才是正确的抽象。那取决于你所处的世界。你的智能体必须自己去学习,在它所处的那个世界里什么才是正确的抽象。所以,你知道,也许这就是两个关键点。你必须找到正确的抽象,然后你必须能够做到持续深度学习。
便签笔记
38:10
>> I think that a lot of the people in the field realize that we need models, we need to plan with them. But, the abstractions tell us what the model should be conditioned on. So, what should you What should the model predict? What are you going to do and then something is going to happen. And more importantly, where would that come from? So, I really like the example of elite athletes. If you ask elite athletes about how they do certain things, they would have weird niche terminologies for doing very specific things. They were like, you know, I do this thing and they would have a name for it. If they communicate, sometimes they don't even have a name for it if they're just doing it alone.
>> 我认为领域内很多人都意识到我们需要模型,需要用模型来做规划。但是,抽象告诉我们模型应该以什么为条件。也就是说,模型应该预测什么?你要做什么,然后会发生什么。更重要的是,这些东西从哪儿来?所以我特别喜欢一个例子,就是顶尖运动员。如果你问顶尖运动员他们是怎么做某些事情的,他们会用一些很奇怪、很小众的术语来描述非常具体的动作。他们会说,你知道,我做这个动作,然后他们会给它起个名字。如果他们需要交流的话,有时候如果只是自己一个人练,他们甚至都没有给它起名字。
便签笔记
10灾难性遗忘的解药:步长优化与持续反向传播
38:45
So, how did they come up with those abstractions? That's in some sense a crucial thing that's missing that the later half of Alberta plan answers. >> Can we talk about the continual deep learning part? Is it an algorithmic gap that exists today or is it a just a practical deployment infrastructure data privacy gap? Because if I wanted to do call it naive updating of weights based on user interaction, I can do that today, right? And so, what in your opinion is the biggest thing that we're missing to kind of get to continual deep learning?
那么,他们是怎么想出这些抽象的?从某种意义上说,这正是目前缺失的关键一环,而《阿尔伯塔计划》的后半部分回答了这个问题。>> 我们能聊聊持续深度学习那部分吗?这在今天是一个算法上的缺口,还是仅仅是实际部署的基础设施、数据隐私方面的缺口?因为如果我想做那种基于用户交互对权重进行朴素更新的做法,我今天就能做,对吧?所以,在你看来,要实现持续深度学习,我们最欠缺的是什么?
便签笔记
39:18
>> Yeah, so it's absolutely an algorithmic gap. You You can do the naive thing, but then you'll see all sorts of problems. So, for example, if you say um I'm going to take one sample and then I'm going to update my whole model with that one sample. You will run into this problem that now all of the previous knowledge in the model it's impacted negatively. And the way currently we we get around this is exactly what cursor does. They don't use one example. They use a large batch coming from a lot of users. So, in use cases where you can have that, you can do continual learning. But, most use cases you don't have that. Most use cases you have a single stream of data. And then if you apply it to the naive thing, it just completely destroys your prior knowledge in a very um destructive way.
>> 是的,这绝对是一个算法上的缺口。你可以用那种朴素的做法,但接着你就会看到各种各样的问题。比如说,如果你说,我要拿一个样本,然后我要用这一个样本去更新我的整个模型样本。你就会遇到这样的问题:模型里之前学到的所有知识都受到了负面影响。目前我们绕开这个问题的办法,正是 Cursor 在做的事。他们不用单个样本,而是用一个来自大量用户的大批量数据。所以在那些你能拿到这种数据的场景里,你可以做持续学习。但大多数场景下你没有这个条件。大多数场景下你只有单一的数据流。这时候如果你用那种朴素的做法,它就会彻底摧毁模型原有的知识,而且是以一种非常具破坏性的方式。
便签笔记
40:00
>> Catastrophic forgetting. >> Yeah. >> That is the Yeah. >> But, it's totally curable. You have to [laughter] have the right algorithm. >> cure? >> Well, yeah, exactly. >> What is the cure? >> Well, you know, first you need to do what we call step size optimization. And it means every weight in your network has to have a separate step size. So, some will move fast, some will move slow. And you will we you will have to metalearn the step sizes for each weight. Most of your network will be have have weights that have tiny step sizes. So, then when you train on a new example, they don't get destroyed.
>> 灾难性遗忘。>> 是的。>> 那就是 是的。>> 但这完全是可以治好的。你得有 [笑] 对的算法。>> 治好?>> 嗯,对,正是如此。>> 那解药是什么?>> 嗯,你知道,首先你得做我们所说的步长优化。意思是你网络里的每一个权重都得有各自独立的步长。所以有些会变化得快,有些会变化得慢。而你需要为每个权重去元学习(metalearn)这些步长。你网络里的大部分权重的步长都会非常小尺寸。这样当你在新样本上训练时,它们就不会被破坏。
便签笔记
40:37
Happens just to the right places. And then secondly, you have to use some form of generate and test. Um which is in feature space. So, you come up with new features or new units and and without following gradients. Cuz gradients are very slow process. You only move in a direction if you know it's the helpful one. And that's always going to be very slow and doesn't give you a path to grow more and more complex and and to have sustained learning. You need to have something that just proposes a bunch of new units.
恰好发生在正确的地方。其次,你必须使用某种形式的生成与测试(generate and test)。嗯,是在特征空间里进行的。也就是说,你提出新的特征或新的单元,而且不遵循梯度。因为梯度是一个非常慢的过程。只有当你知道某个方向是有帮助的,你才会朝那个方向移动。这总是会非常慢,也没法让你走向越来越复杂、实现持续学习。你需要有某种东西,能直接提出一大批新的单元。
便签笔记
41:12
And and and then goes from there. I guess so, there is a specific thing I can say that make it at least concrete, which is to say we have this algorithm called continual backprop. We used published in in nature a couple years ago. And it it is exactly like backprop, but every but you also plant new seeds of units that are newly initialized with random weights. Backprop only has random weights at the beginning of time. And then as you go on, all that randomness all that variety from the randomness gets used up.
然后从那里继续下去。我想,有一件具体的事情我可以说,至少能让它变得具体一些,那就是就是说,我们有一个叫做持续反向传播(continual backprop)的算法。我们几年前在《自然》上发表过。它跟反向传播完全一样,只不过你还会不断播下新的种子,也就是用随机权重重新初始化的新单元。权重。反向传播只在最开始的时候有随机权重。然后随着训练进行,所有那些随机性、所有那些多样性都会被随机性消耗殆尽。
便签笔记
41:46
And with continual backprop, we keep injecting a bit of randomness, a bit of generate and test, a bit of generate and then the the operation of backprop is the tester. So, you need you need that. And and if you put those together really well, I think you'll have a new generation of massively superior continual deep learning. And that's what we hope to do in the next couple years. >> Wonderful. Do you think that these algorithms can be applied to the current state of affairs with people scaling LLMs and trying to get them to do continual learning without catastrophic forgetting?
而有了持续反向传播,我们就不断注入一点随机性,一点"生成与测试",一点生成,然后反向传播的运算过程就充当了那个测试者。所以,你需要,你需要那个东西。而且如果你能把这些很好地结合起来,我认为你就会得到新一代的、性能大幅超越的持续深度学习。这就是我们希望在未来几年里做到的事。>> 太好了。你认为这些算法能不能用到现在的局面上?现在大家都在扩展大语言模型,试图让它们做到持续学习而不发生灾难性遗忘。
便签笔记
42:22
>> Yeah, absolutely. I think it's So, I don't think that you could take an existing model and say I'm going to just start updating it with these algorithms because these algorithms meta learn how to learn. So, really you have to say, I'm going to learn from scratch. So, let's say I learn a new foundation model, but I'm going to learn with this these new algorithms. These new algorithms in addition to learning the knowledge, they're also going to learn how to learn future things. So, they're learning two things at the same time. And then um then I think you would be able to learn new things without catastrophic forgetting.
>> 是的,绝对可以。我觉得……不过,我不认为你可以拿一个现成的模型,然后说我要直接开始用这些算法去更新它,因为这些算法是在元学习「如何学习」。所以你真的得说,我要从头开始学。也就是说,比方说我训练一个新的基础模型,但我要用这些新算法来学。这些新算法除了学习知识之外,还会学习如何去学未来的东西。所以它们是同时在学两样东西。然后,嗯,我认为那样你就能学到新东西而不会发生灾难性遗忘。
便签笔记
42:57
>> the most radical thing in that you're trying to do in your company in terms of from the current state of affairs to try to do these two things at the same time? >> Most radical thing. >> Is I think that's >> This goes back to I'm not crazy, everyone else is crazy. >> [laughter] >> Yeah. >> That's perhaps not totally radical. There was a point in like 2016 to 2018 where a lot of people were exploring these ideas quite a bit. They were doing it in a much more limited setting. So, they would say, we have a distribution of problems and then in this specific case we'll do it whereas we want to do it from a single stream of experience. So, our method should be more generally applicable. So, I think many people have explored this, but no one has explored this in the general setting where the resulting algorithm would be applicable everywhere.
>> 从目前的现状来看,你在自己公司里想做的最激进的事情,就是同时做到这两件事吗?>> 最激进的事情。>> 我觉得那就是……>> 这又回到那句话了:我不疯,是其他人疯了。>> [笑] >> 是啊。>> 那也许并不算完全激进。大概在 2016 到 2018 年那段时间,很多人都相当深入地探索过这些想法。只不过他们是在一个受限得多的设定下做的。他们会说,我们有一个问题分布,然后在这个特定情形下我们来做这件事;而我们想做的是从单一的经验流中学习。所以,我们的方法应该更具普遍适用性。所以我认为很多人探索过这个方向,但没有人在通用设定下探索过——在那种设定里,得到的算法可以到处适用。
便签笔记
11Oak 的雄心:万亿参数、20 瓦、自洽心智
43:43
>> So, what would be the most radical thing that your company your new company is trying to do that other people are not doing? >> What's the most ambitious thing? Remember, I don't think I'm weird, so I don't want to say it's radical. >> the most ambitious >> ambitious thing, I think is to try to have the full spectrum of knowledge both about the tiny things and about the big things. You know, like thinking about how you take an airplane from one city to another. That's a a very big thing. You know, it's it's more it's like your your space flight example, but it's just kind of more common sensical to think about cuz we all many of us take airplanes and any all of us use abstractions on all kinds of our life. And even even the squirrels use abstractions. So, to have that spectrum of of knowledge uh from the small to the big and to treat it in a uniform way and to be able to help have it self uh maintaining. You know, the big question is always you have your knowledge-based system and what keeps the knowledge in
>> 那么,你的公司、你的新公司想做而别人没有在做的最激进的事情是什么?>> 最有雄心的事情是什么?记住,我不觉得自己奇怪,所以我不想说它是激进的。>> 最有雄心的……>> 最有雄心的事情,我认为是试图拥有全谱系的知识,既包括细小的事物,也包括宏大的事物。你知道,比如思考你怎么坐飞机从一个城市到另一个城市。那是一件非常大的事。你知道,这更像是你说的那个太空飞行的例子,只不过它更符合常识、更容易想象,因为我们很多人都坐过飞机,而我们所有人在生活的方方面面都在使用抽象。甚至连松鼠也在用抽象。所以,要拥有那种从小到大的知识谱系,并以统一的方式来处理它,并且能让它自我维护。你知道,最大的问题始终是:你有一个基于知识的系统,那是什么在保证里面的知识是正确的?
便签笔记
44:49
it correct? Well, what keeps the knowledge correct in a large language model is well, people did a lot of post-training and and they they they made it sure it was correct. And then they freeze it after that. So, that's what keeps it correct. But really our minds, we are we're always changing things and yet something keeps it organized and coherent and and and settling back into a good place rather than drifting off into crazy land. That is an I think our our biggest ambition to have a mind that is self-consistent and and can keep training itself and making it coherent.
那么,在大语言模型里,是什么在保证知识正确呢?是人们做了大量的后训练,他们确保了它是正确的。然后在那之后就把它冻结了。所以,就是这个在保证它的正确性。但实际上我们的头脑,我们一直在改变东西,然而总有某种机制让它保持有条理、连贯,并且回落到一个良好的状态,而不是飘到疯狂的世界里去。我认为那是我们最大的雄心:拥有一个自洽的心智,它能持续训练自己,并保持连贯。
便签笔记
45:26
>> that. Can I ask? It almost seems that it's it's such a ambitious vision and the idea that all these things can be unified into a single mind is so ambitious. >> It's within reach. I think it's within reach. It's here it's 2026 and our computers are so fast. You know, is it is it so ambitious that it's out of reach? Uh or or do we have already inklings of how all the steps can be done? And we I I think we have a vision and inklings. Um I don't think it's I don't think it's inappropriate. >> Your vision involves a trillion parameter model with 20 watts.
>> 嗯。我能问一下吗?这看上去几乎是一个如此有雄心的愿景,而且把所有这些东西统一到一个单一心智里的想法,实在太有雄心了。>> 它是触手可及的。我认为它是触手可及的。现在是 2026 年,我们的计算机这么快。你知道,它真的雄心大到遥不可及吗?呃,还是说我们其实已经对每一步该怎么做有了一些眉目?我认为我们有一个愿景,也有一些眉目。嗯,我不认为这是不切实际的。>> 你的愿景包括一个万亿参数的模型,只用 20 瓦。
便签笔记
46:07
That seems pretty ambitious. >> That is ambitious. Uh, in some sense with current technology, I would say it's also impossible. Like just storing a trillion parameters in memory would probably use more than 20 watts of energy with current memory technologies, but we are really thinking of okay, things are getting better, computation is getting cheaper or it is getting more energy efficient. So, where would be would be in 5 to 10 years? And I think 5 to 10 years with the right algorithms and we can totally be in a world where this would be possible.
那听起来相当有雄心。>> 那确实有雄心。呃,从某种意义上说,以当前的技术,我会说这也是不可能的。比如说,光是把一万亿个参数存在内存里,按现在的存储技术,可能就要消耗超过 20 瓦的能量;但我们真正想的是,好,事情在变好,计算在变便宜,或者说变得更加节能。那么 5 到 10 年后会是什么样?我认为 5 到 10 年后,配上正确的算法,我们完全可能处在一个让这件事成为可能的世界里。
便签笔记
46:43
>> So, 5 to 10 years is two orders of magnitude of Moore's law. It's a standard improvement. If we double every 18 months, 10 years would give you two orders of magnitude. And so, for Kurzweil's statement to be plausible, then today you should be able to do it for for what? 20 watts? Two orders of magnitude? >> 2,000 >> 2,000 watts. If you can do it with 2,000 watts today, yeah, then in 10 years you'll be able to do it for 20 watts. >> You think you can do it for a 2,000 watts? You have I think lots of people at research labs that have access to way more than that.
>> 那么,5 到 10 年就是摩尔定律的两个数量级。这是一个标准的进步幅度。如果我们每 18 个月翻一番,10 年就能带来两个数量级。所以,要让库兹韦尔的说法站得住脚,那今天你应该能用多少瓦做到?20 瓦乘以两个数量级?>> 2000 瓦?>> 2000……>> 2000 瓦。如果你今天能用 2000 瓦做到,是的,那 10 年后你就能用 20 瓦做到。>> 你觉得你能用 2000 瓦做到吗?我想很多研究实验室里的人手上的资源可远远不止这些。
便签笔记
47:21
>> Yeah, I think we can can be more efficient than that even now with the right algorithm. >> If we can be more efficient than that, then why aren't we? It's not like people just want to spend all their money on spend all their money. >> Sometimes it seems like they >> I'll pay money. >> Sometimes it [laughter] seems like they want to. >> Yeah, I I >> Doesn't it? I think that's how they show they're they're real men by using lots of energy. >> At least when I look at different research groups, I don't even see anyone believing in that it's possible. And I think if you don't believe in it, you're just not going to work on the technical problems and work through them.
>> 是的,我认为有了正确的算法,即便是现在我们也能比那更高效。>> 如果我们能比那更高效,那为什么我们没有做到?大家又不是就想把所有的钱都花在……把所有的钱都花光。>> 有时候看起来他们……>> 我愿意花钱。>> 有时候 [笑] 看起来他们就是想这么干。>> 是啊,我……>> 难道不是吗?我觉得那是他们证明自己是真汉子的方式——靠消耗大量能源。>> 至少当我去看不同的研究团队时,我甚至没看到有人相信这是可能的。而我认为,如果你不相信它,你就根本不会去攻克那些技术问题、把它们一个个解决掉。
便签笔记
47:56
>> Is it that it's not possible or it's that there's so much waste in the system? Like one which one is it? Like is it is is there someone who knows how to do it efficiently? >> Yeah. >> And then there's 10 times the number of people in the same lab doing all these other things. And so nine out of 10 people are wasting >> In some sense I the way I think about it is that we are stuck in a local minimum. So, if we want to move towards these new kind of algorithms, it is almost impossible that things will not get worse before they get better.
>> 那到底是它做不到,还是说系统里存在大量浪费?到底是哪一种?比如说,是不是有人其实知道怎么高效地做?>> 嗯。>> 然后同一个实验室里有 10 倍数量的人在做其他各种事情。所以十个人里有九个在浪费……>> 从某种意义上说,我的看法是我们卡在了一个局部极小值里。所以,如果我们想转向这类新算法,那几乎不可能不出现「先变差、后变好」的情况。
便签笔记
48:27
So, when we start exploring these new directions, you're not going to get state of the art performance from day one, but it is because it is a different paradigm. Um but that path leads to similar performance at a higher energy scale. And these big labs, they are so locked into a product that they like it is not possible for them to pursue a path where things get worse first. >> Because their current paradigm allows them to keep scaling and this new paradigm they have to take a bet. And then >> And they have to figure out some some technical things that are difficult that we have thought about it for many years.
所以当我们开始探索这些新方向时,你不可能从第一天起就拿到最先进的性能,但那是因为这是一个不同的范式。嗯,但那条路会在更高的能量规模上通向相近的性能。而这些大实验室,他们被产品绑得太死了,以至于他们根本不可能去走一条先变差的路。>> 因为他们当前的范式让他们可以继续扩展,而这个新范式他们得下一个赌注。然后……>> 而且他们还得搞定一些困难的技术问题,这些问题我们已经思考了很多年。
便签笔记
12LLM 只占智能的四分之一
49:03
We know people who have thought about these things for many years and when I talk to them, it makes sense that it's doable, but you need to think about those challenges for a long period of time. >> So, if everything goes right with Oak, what happens with the company? What do you what kind of company are you building? >> Uh if everything goes right, we uh implement the architecture, we can have a genuine uh continual learning and we can form abstractions so that we can do planning and reasoning and and we have so sort of like true intelligence. And then, you know, it's hard to imagine just exactly what will happen by then.
我们认识一些思考这些问题很多年的人,当我跟他们交流时,我觉得这事儿是可行的,但你需要长时间地去琢磨那些挑战。>> 那么,如果 Oak 一切顺利,公司会变成什么样?你在打造一家什么样的公司?>> 呃,如果一切顺利,我们实现了这个架构,我们就能拥有真正的持续学习,而且我们能形成抽象,从而进行规划和推理,我们就有了某种意义上的真正智能。然后,你知道,很难想象到那时究竟会发生什么。
便签笔记
49:40
But I think >> Humans will become irrelevant. >> I I don't think that's true at all. I I I >> We don't either. >> I think the world becomes exciting and even more exciting and interesting and and for humans. But in particular, I think there are the You have to wonder about the large language models. They might be at risk. Uh when this eventually happens, you know, I'm sure they'll get a good run. They've already had a good run. You know, they've been very successful. And let me say, just for for clarity that uh large language models are an amazing scientific breakthrough, a breakthrough in the skillful use of language by neural networks, wholly unanticipated.
但我认为……>> 人类会变得无关紧要。>> 我完全不认为是那样。我……>> 我们也不这么认为。>> 我认为世界会变得令人兴奋,甚至对人类来说更加令人兴奋、更有意思。但特别是,我认为……你不得不为大语言模型担心一下。它们可能会面临风险。呃,当这一切最终发生的时候,你知道,我相信它们已经有过一段风光。它们已经风光过了。你知道,它们非常成功。而且请允许我澄清一点:大语言模型是一项了不起的科学突破,是神经网络在娴熟运用语言方面的突破,完全出乎意料。
便签笔记
50:22
You know, it was a It was always a hold out for uh symbolic methods in language, and they They have totally changed how that's thought about now. Yeah. It's a It's a big breakthrough. It's so It's frustrating to me that we have to you know, just celebrate that we've made this great progress in the sub subset of the problem of AI, and enjoy that. Instead, it has to pretend to be all of AI. All of intelligence is not fluid, capable use of language. There's so much more. It's an important part. You know, it's like 20% or a quarter of intelligence. There's There's more.
你知道,语言一直是符号方法的最后一块阵地,而它们彻底改变了现在人们对这件事的看法。是的,这是一个重大突破。让我感到沮丧的是,我们本可以就为我们在 AI 这个问题的一个子集上取得的巨大进展而庆祝、享受这份成果。可结果它非得假装自己就是 AI 的全部。智能的全部并不等于流畅、娴熟地运用语言。还有太多别的东西。它是重要的一部分。你知道,大概占智能的 20% 或四分之一。还有更多的东西。
便签笔记
50:59
>> Yeah. >> We're not done. >> Yeah. If everything goes right, are you imagining that there's a single minds that can do everything from learn how to swing from tree branches to make a spaceship, uh to you know, all these various things we've talked about today. Is it a single mind, and is it a single set of weights that can that can do all these things, or is it >> It's a single design. >> Okay. >> And there'll be many different There'll be many minds. >> Mm. Okay. So, it's a single design that reacts to different environments.
>> 是啊。>> 我们还没有完成。>> 是啊。如果一切顺利,你设想的是有一个单一的心智,能做到从学会在树枝间荡秋千,到造一艘宇宙飞船,呃,到今天我们聊到的所有这些形形色色的事情。它是一个单一的心智吗?是不是一套权重就能做所有这些事情,还是说……>> 它是一个单一的设计。>> 好的。>> 而会有很多不同的……会有很多个心智。>> 嗯。好的。所以它是一个单一的设计,对不同的环境作出反应。
便签笔记
51:30
>> And it did it And different versions of that mind would learn different things because they have different experience. This sort of goes back to the big world hypothesis that there are infinitely many things to learn, so one system cannot learn infinitely many things. And like I think Rich already mentioned this, but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot model each other because they are equally complex. So, single system would never be able to get to a point where it can learn everything. It would always be multiple systems that are learning from their own experience.
>> 而且那个心智的不同版本会学到不同的东西,因为它们有不同的经验。这某种程度上又回到了大世界假说:要学的东西有无穷多,所以一个系统无法学习无限多的东西。而且我想 Rich 刚才已经提到过这一点,但如果你有两个这样的系统,如果你有世界上最大的两个系统,那么很显然它们无法相互建模,因为它们的复杂度是相当的。所以,单个系统永远不可能达到能学会一切的地步。永远都会是多个系统各自从自己的经验中学习。
便签笔记
52:03
>> Guys, are you hiring? What kind of people are you looking for? >> We are hiring. The initial team, most of it we already have in our mind. So, these are people who have thought about these ideas in the past in many uh you know, for a long time. And we are going to take a slightly different approach because this is a different paradigm. It doesn't make sense to become large very quickly because uh in some sense, everyone we hire has to uh come to see what we see. And not everyone sees that. So, we're going to start small, slowly grow to maybe a handful or two or three uh and then go from there.
>> 各位,你们在招人吗?你们在找什么样的人?>> 我们在招人。最初的团队,大部分人选我们心里已经有数了。这些人都是过去就思考过这些想法的人,在很多方面、而且思考了很长时间。我们会采取一种略有不同的做法,因为这是一个不同的范式。很快把规模做大是没有意义的,因为某种意义上说,我们招的每个人都得能看到我们所看到的东西。而并不是每个人都能看到。所以我们会从小做起,慢慢增长到大概几个人、两三个人,然后再往下走。
便签笔记
52:40
>> We want to be super aligned. >> We want to be super aligned. >> So, that we can be um very productive working together and and scaling the progress. >> Absolutely. >> Very, very cool. >> Wonderful. I love this conversation. Thank you for taking the time to share what you're up to. You are, you know, an an unusually deep thinker about where reinforcement learning and algorithmic design will go. And it was a true pleasure to get to explore it together with you today. So, thank you. >> Thank you very much. Thank you. It's our pleasure.
>> 我们希望高度一致。>> 我们希望高度一致。>> 这样我们才能一起非常高效地工作,并把进展规模化。>> 完全同意。>> 非常非常酷。>> 太好了。我很喜欢这次对话。谢谢你们抽时间来分享你们正在做的事。你们对强化学习和算法设计的走向有着非同寻常的深刻思考。今天能和你们一起探讨这些,真的非常荣幸。所以,谢谢你们。>> 非常感谢。谢谢。这是我们的荣幸。
便签笔记
53:15
>> [music]
>> [音乐]
便签笔记
53:21
[music]
[音乐]
便签笔记
53:37
[music]
[音乐]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Rich Sutton 与 Khurram Javed 借「苦涩教训」与「大世界假说」论证:当前大语言模型权重冻结、不再从经验中学习是根本缺陷,而他们创办的 Oak Lab 要用逐权重步长元学习 + 持续反向传播(生成—测试)等算法,从零训练能终身持续学习、自建抽象并据此规划的「单一设计的心智」。

核心要点

  • 「持续学习」这个词本身暴露了领域的怪异:Sutton 认为一切学习天然是持续的,主体总是在行动中学习;把它单列成一个子领域,是 AI 热潮之后才出现的「集体走偏」。他自称「不是我怪,是这个领域怪」。
  • 苦涩教训的 26 字版:不要被人类知识分心,专注于能随算力扩展的方法(搜索与学习)。它不是反对精巧算法,而是要求算法随「计算」而非随「人类输入」扩展。Sutton 早在 2019 年成文前一年就已在演讲中提出,是对符号 AI 时代数十年经验的总结,而非对某一时刻的回应。
  • LLM 既是苦涩教训的正例也是反例:它们靠「喝下整个互联网」实现了空前的算力扩展(正例);但互联网是有限的,世界远大于互联网上的一切,最终被人类知识的天花板卡住(反例)。Sutton 评价 LLM 是「神经网络熟练使用语言」的未被预见的重大科学突破,但只占智能的 20%~25%,不应假扮成全部 AI。
  • 合成数据是「大错误」,因为瓶颈仍在人:Javed 的反问——若 OpenAI、Anthropic 的工程师全去度假,谁来生成合成数据?什么算好数据要由人类专家判断。以「靠回声定位飞行的无人机」为例,无论生成多少个「宇宙」的合成数据都做不到,除非先雇物理领域专家搞清楚。Sutton 补充:任何模拟相对真实世界都是「显微镜级」的小世界,连他人心智这类要紧信息根本无法合成。
  • 大世界假说是持续学习的根本理由:世界远比任何智能体复杂(显然如此,因为世界里还有其他智能体),故智能体只能是严重的近似;它必须针对自己所处的那一小片世界持续调整近似。对自动驾驶「仿真训练成功」的反驳:那需要大团队反复修补仿真—现实差距,人一直在环路里——为什么不让智能体自己学模型、自己修?智能体自建的模型错了可以靠持续学习修正,人造的仿真只有等人发现错误才能更新。
  • 先验知识与学习本不该为敌:Sutton 强调《苦涩教训》开篇就说二者不冲突;他显得「偏爱学习」只因全世界都在说「只要知识够多就不需要学」。未来的机器人不会从零训练,而是复制已有智能体后继续学;从这个意义上「先验」反而可以直接抛弃——从上一个机器人那里复制就行。
  • Cursor 的 tab 模型是罕见的持续学习实例,但效率很低:它用来自成千上万用户的批量数据做一次策略更新。真正的需求是「我要教我的那份模型一件具体的事」,而不必与 10 万人争抢要教给共享模型什么。单条经验流上做朴素权重更新会灾难性遗忘——但这「完全可治」。
  • 灾难性遗忘的解药是两步算法:(1) 步长优化——网络每个权重各自拥有元学习得到的步长,绝大多数权重步长极小,新样本只改动「该改的地方」;(2) 特征空间里的「生成—测试」——不靠梯度(梯度只沿已知有益方向走,太慢、无法持续增长复杂度),而是不断播种随机初始化的新单元,由反向传播充当测试者。这就是发表于《Nature》的「持续反向传播」(continual backprop)。这套方法要从零训练新的基础模型,因为它同时学「知识」和「如何学未来的东西」,不能嫁接到现有模型上。
  • 领域里尚无「学到模型再用它规划」的实例:AlphaGo 和数学证明之所以成功,是因为规则/算子已知;而以自发现的抽象为条件、学出模型再规划,目前没有。Alberta 计划 12 步中,第 2 步「持续深度学习」最关键,后半部分全是关于让智能体学出适合所处世界的正确抽象(精英运动员对动作的私人术语即为例子)。
  • 大实验室被困在局部极小值:新范式在初期必然「先变差再变好」,锁死在产品上的大公司无法承受,而且没人相信高能效路线可行就不会去攻技术难题。对「20 瓦跑万亿参数」的目标,Javed 承认当下技术不可能(光存参数就超 20 瓦),但 5~10 年(两个数量级摩尔定律)内可行;Sutton 认为用对算法今天就能低于 2000 瓦,并调侃「烧大量能源是他们证明自己是真男人的方式」。

结论与值得注意的细节

  • Oak 的终极图景是「单一设计、多个心智」:不是一套权重包打天下,而是同一架构的不同副本因经验不同而学到不同东西;两个同等复杂的最大系统必然无法彼此建模,故永远是多智能体各自学习。
  • 他们最大的野心不是「持续学习」本身,而是让一个心智在从小到大的整个知识谱系上统一处理并自我维持一致性——LLM 靠后训练+冻结保持知识正确,人脑却在不断改动中仍能保持连贯而不「漂进疯狂之地」。
  • 招聘策略刻意反常:初始团队多已心中有数,将从「一两只手」的人数缓慢扩张,强调「超级对齐」(想法高度一致),因为每个人都得「看见我们看见的东西」。
  • 个人插曲:Sutton 2003 年在 AI 寒冬中、且身患癌症「以为只剩几个月」时去了阿尔伯塔大学;他引富兰克林之说「人做事无非习惯或虚荣」,自认是习惯让他在濒死时仍坚持研究。Javed 从未正式申请做他的学生,是合作项目自然长成博士论文。
  • 关于动物与学校的争论:Sutton 认为没有任何动物靠监督学习(没人示范肌肉如何抽动,且别人的示范对你的身体也是错的),学校是几百年内才有的特殊事物,不是智能的本质;但他同意规划与想象能力不可或缺,只是坚持范式转换者也必须从自身经验中构建新抽象。
  • 若一切顺利,Sutton 明确说人类不会变得无关紧要,反倒是「大语言模型可能处于危险之中」——它们已经有过很好的一轮。
核心句型 · 9
1. It's not A. It's B.(否定—更正结构)
“I'm not weird. The field is weird.”
用两个极短的平行句先否定再更正,语气斩钉截铁。适合表达立场反转或纠正误解。仿写:It's not the data. It's the algorithm.
2. It wouldn't make any sense to talk about X that wasn't Y.
“It wouldn't make any sense to talk about learning that wasn't continual”
用「谈论不 Y 的 X 没有意义」来论证 Y 是 X 的本质属性。适合定义性论辩。仿写:It wouldn't make sense to talk about science that wasn't empirical.
3. Don't be distracted by X. Instead focus on Y that will Z.
“Don't be distracted by human knowledge … Instead focus on learning methods that will scale with computation”
先排除干扰项,再用 instead 引出正确方向,后接定语从句限定条件。适合给出原则性建议。
4. It's both a positive example and a negative example of X.
“It's both a positive example and a negative example of the big lesson”
用 both…and 把同一对象置于两个对立范畴,随后分述。适合表达复杂的双面评价,避免非黑即白。
5. Who decides what's A and what's B?
“Who decides what's a good synthetic data and what's a bad synthetic data?”
以「谁来决定」设问,把技术问题转为权威/瓶颈问题。适合在辩论中把对方的方案追溯到隐藏的人工依赖。
6. It doesn't matter how much … , it will not …
“It doesn't matter how much synthetic data you generate … it will not be not allow you to do that task”
「无论多少……也无法……」的让步强调结构。适合表达数量再大也无法弥补性质缺陷。
7. There's no reason in principle why X and Y have to be opposed.
“There's no reason in principle why these have to be opposed.”
in principle 表示「原则上」,与后文 in practice 形成对照。适合区分理论应然与现实实然。
8. You're not going to get X from day one, but that path leads to Y.
“You're not going to get state of the art performance from day one, but … that path leads to similar performance at a higher energy scale”
先承认短期代价,再用 but 转到长期收益。适合为「先变差再变好」的策略辩护。
9. X is an important part. It's like 20% of Y. There's more.
“It's an important part. You know, it's like 20% or a quarter of intelligence. There's more.”
用一个粗略比例给出定位,再以 There's more 收尾。适合在肯定的同时限定其重要性。
词汇精讲 · 118 · 按出现顺序
radical /ˈrædɪkəl/ adj. 0:00
激进的;根本性的
seminal /ˈsemɪnəl/ adj. 0:59
开创性的、有深远影响的(seminal textbook 奠基性教科书)
propelling /prəˈpelɪŋ/ v. 0:59
推动、驱使(propel the field forward)
set off to phr. 1:28
动身去做、着手做
bastion /ˈbæstʃən/ n. 1:28
堡垒;(某事物的)重镇、大本营
in its infancy phr. 1:28
处于初创/萌芽阶段
conviction /kənˈvɪkʃən/ n. 1:28
坚定的信念;确信
doubling down on phr. 2:10
加倍下注、更加坚持(某立场或做法)
remission /rɪˈmɪʃən/ n. 2:10
(病情)缓解期
might as well phr. 2:10
不妨、还不如(表示别无更好选择)
vanity /ˈvænəti/ n. 3:12
虚荣心
Divine intervention phr. 3:52
神的干预、天意
perception /pərˈsepʃən/ n. 4:24
感知、知觉
motto /ˈmɑːtoʊ/ n. 5:08
座右铭、信条
organically /ɔːrˈɡænɪkli/ adv. 6:10
自然而然地、有机地(发展)
down the road phr. 6:40
往后、日后(6 months down the road 半年后)
tome /toʊm/ n. 7:10
大部头、经典巨著
paradigm /ˈpærədaɪm/ n. 7:10
范式、模式
a long time coming phr. 7:10
酝酿已久、早该到来
symbolic AI n. 7:45
符号主义人工智能(基于规则与逻辑的传统 AI)
tortured /ˈtɔːrtʃərd/ v./adj. 8:28
(此处)被曲解、被牵强附会地使用
anticipate /ænˈtɪsəpeɪt/ v. 9:40
预料、预先料到
drink in phr. 9:40
(贪婪地)吸收、饱览
finite /ˈfaɪnaɪt/ adj. 10:20
有限的
holds us back phr. 10:20
拖住我们、阻碍我们前进
push on this phr. 10:59
就这点追问、施压
synthetic data n. 10:59
合成数据(由程序或模型生成而非真实采集的数据)
leverages /ˈlevərɪdʒɪz/ v. 10:59
利用、借力
floating around phr. 11:33
(想法、说法)流传着、四处传播
hypothesis /haɪˈpɑːθəsɪs/ n. 11:33
假说、假设
in the loop phr. 11:33
参与其中、在回路中(human in the loop 人在回路)
bottleneck /ˈbɑːtlnek/ n. 12:17
瓶颈
loss curve n. 12:45
损失曲线(训练过程中损失函数值随时间的变化图)
echolocation /ˌekoʊloʊˈkeɪʃən/ n. 12:45
回声定位
domain experts n. 12:45
领域专家
localize /ˈloʊkəlaɪz/ v. 13:35
定位(自身位置)
friction /ˈfrɪkʃən/ n. 14:53
摩擦力
microscopic /ˌmaɪkrəˈskɑːpɪk/ adj. 14:53
微小的、显微镜级的
approximations /əˌprɑːksɪˈmeɪʃənz/ n. 15:37
近似、近似值
severe /sɪˈvɪr/ adj. 15:37
严重的;(近似)粗糙的
tuned to phr. 15:37
针对……调校、适配于
argumentative /ˌɑːrɡjəˈmentətɪv/ adj. 16:19
好辩的、爱抬杠的
cohort /ˈkoʊhɔːrt/ n. 16:19
一批(同期的)人或公司
post-training n. 16:19
后训练(预训练之后的微调阶段)
sim-to-real gap n. 16:19
仿真到现实的差距
curating /ˈkjʊreɪtɪŋ/ v. 17:10
筛选整理、策展
triumph /ˈtraɪʌmf/ n. 18:17
胜利、巨大成功
priors /ˈpraɪərz/ n. 18:17
先验(知识/假设)
vindicated /ˈvɪndɪkeɪtɪd/ v./adj. 18:17
被证明正确的、得到平反的
nature and nurture phr. 18:59
先天与后天
affection for phr. 18:59
对……的偏爱
dismiss /dɪsˈmɪs/ v. 18:59
不予理会、否定
drip /drɪp/ v. 20:44
滴入;(比喻)一点点地投放
zillions /ˈzɪljənz/ n. 21:22
(口语)无数、天文数字
routing /ˈruːtɪŋ/ n./v. 22:02
路由、选路(网络数据包传输路径)
contention /kənˈtenʃən/ n. 22:37
论点、主张
in-context learning n. 22:37
上下文内学习(模型仅凭提示中的信息适应任务,不更新权重)
mental model n. 24:48
心智模型、理解框架
proprioception /ˌproʊpriəˈsepʃən/ n. 25:28
本体感觉(对身体位置与运动的内在感知)
ingrained /ɪnˈɡreɪnd/ adj. 25:28
根深蒂固的
plastic /ˈplæstɪk/ adj. 25:28
(神经)可塑的
supervised learning n. 26:48
监督学习(用带标注的输入输出对训练)
twitch /twɪtʃ/ v. 26:48
抽动、抽搐
transmission /trænzˈmɪʃən/ n. 27:31
传递、传播(知识的传承)
paradox /ˈpærədɑːks/ n. 29:25
悖论
vocal cords n. 30:06
声带
inherently /ɪnˈhɪrəntli/ adv. 30:06
本质上、内在地
abstraction /æbˈstrækʃən/ n. 30:39
抽象(概念)
trajectories /trəˈdʒektəriz/ n. 31:14
轨迹、弹道
hubristic /hjuːˈbrɪstɪk/ adj. 31:14
傲慢自大的
commonalities /ˌkɑːməˈnælətiz/ n. 31:48
共同点
uniquely /juːˈniːkli/ adv. 32:25
独特地、唯独
paradigm shifting adj. 33:33
范式转变的、颠覆性的
generalizes /ˈdʒenrəlaɪzɪz/ v. 35:10
(模型)泛化
counterexample /ˈkaʊntərɪɡˌzæmpəl/ n. 35:40
反例
pontificating /pɑːnˈtɪfɪkeɪtɪŋ/ v. 36:23
高谈阔论、自以为是地发表意见
lamenting /ləˈmentɪŋ/ v. 36:23
哀叹、痛惜
crystallize /ˈkrɪstəlaɪz/ v. 36:54
使(想法)明确、具体化
unlocks /ʌnˈlɑːks/ v. 36:54
解锁、开启
conditioned on phr. 38:10
以……为条件(模型输入的依据)
niche /niːʃ/ adj. 38:10
小众的、专门的
naive /naɪˈiːv/ adj. 38:45
(算法)朴素的、未加改进的
Catastrophic forgetting n. 40:00
灾难性遗忘(神经网络学新任务时丢失旧知识)
step size n. 40:00
步长(即学习率)
metalearn /ˈmetəlɜːrn/ v. 40:00
元学习(学习如何学习)
generate and test phr. 40:37
生成与测试(先随机提出候选再筛选的搜索策略)
gradients /ˈɡreɪdiənts/ n. 40:37
梯度
sustained /səˈsteɪnd/ adj. 40:37
持续的、持久的
backprop /ˈbækprɑːp/ n. 41:12
反向传播(backpropagation 的缩写)
initialized /ɪˈnɪʃəlaɪzd/ v. 41:12
初始化
used up phr. 41:12
耗尽、用光
injecting /ɪnˈdʒektɪŋ/ v. 41:46
注入
massively superior phr. 41:46
大幅优越的
from scratch phr. 42:22
从零开始
foundation model n. 42:22
基础模型
a distribution of problems phr. 42:57
问题分布(元学习中一组同类任务的集合)
full spectrum phr. 43:43
全谱系、全部范围
common sensical adj. 43:43
合乎常识的
coherent /koʊˈhɪrənt/ adj. 44:49
连贯的、条理清晰的
drifting off phr. 44:49
漂离、偏航
self-consistent adj. 44:49
自洽的
within reach phr. 45:26
触手可及、可以实现
inklings /ˈɪŋklɪŋz/ n. 45:26
模糊的想法、些许眉目
energy efficient adj. 46:07
节能的、能效高的
orders of magnitude phr. 46:43
数量级
plausible /ˈplɔːzəbəl/ adj. 46:43
貌似可信的、说得通的
local minimum n. 47:56
局部极小值(比喻困在次优状态)
state of the art phr. 48:27
最先进水平
locked into phr. 48:27
被锁定在、无法脱身
take a bet phr. 48:27
下赌注、冒险一试
irrelevant /ɪˈreləvənt/ adj. 49:40
无关紧要的
a good run phr. 49:40
一段成功/风光的时期
unanticipated /ˌʌnænˈtɪsɪpeɪtɪd/ adj. 49:40
出乎意料的
hold out n. 50:22
最后的据点、坚守之地
fluid /ˈfluːɪd/ adj. 50:22
流畅的
trivial /ˈtrɪviəl/ adj. 51:30
(数学口吻)显而易见的
a handful phr. 52:03
少数几个
aligned /əˈlaɪnd/ adj. 52:40
目标一致的、步调一致的
理解自测 · 11 题
1. Sutton 为什么说「持续学习」这个术语本身就很怪?

因为在他看来所有学习都是持续的:生物一直在行动、一直在学习,学习不是一个特殊阶段。只有当领域把学习拆成「训练阶段」和「冻结部署阶段」之后,才需要专门发明「持续学习」一词来指代本应是常态的东西。这是开场第 0–1 段和第 8–9 段反复强调的观点,也是全片的主旨:他认为自己是常识,是领域在 AI 狂热中「想得怪」。

2. Sutton 是在什么处境下去阿尔伯塔大学任教的?

2003 年,正值 AI 寒冬,而 Sutton 当时正患癌症,自以为只剩几个月可活。他在多次缓解后「没死成」,于是决定不如再找份工作,便去了阿尔伯塔任教(第 5 段)。他事后用富兰克林「习惯或虚荣」的说法解释自己为何在濒死时仍坚持研究,并倾向于归为习惯(第 6–7 段)。

3. Sutton 用 26 个词概括的《苦涩的教训》说了什么?

大意是:不要像传统 AI 一再做的那样被人类知识分散注意力,而要专注于能随算力扩展的学习方法,比如搜索和学习(第 16 段)。他特别澄清这并不是反对精巧算法,而是要求算法必须能随算力扩展,且是「随算力」而非「随人类输入」扩展(第 16–17 段)。

4. Sutton 为什么说大语言模型既是苦涩教训的正面例子又是反面例子?

正面在于 LLM 用可扩展的方法「喝下整个互联网」,仅靠算力扩展就获得了强得多的系统。反面在于它最终受限于人类数据:互联网是有限的,而世界远大于互联网上储存的一切,过度依赖人类知识终将拖住它(第 17–18 段)。这一判断直接引出后文关于合成数据和大世界假说的讨论。

5. 针对「合成数据能突破数据上限」的主张,Javed 和 Sutton 分别给出了什么反驳?

Javed 的反驳是瓶颈论:谁来判断合成数据好坏?目前是人类专家,若「所有工程师都去度假」就没人生成数据,因此合成数据仍属人在回路,被人类专业知识卡住;他用蝙蝠回声定位无人机说明对未知问题连专家都不知道该合成什么(第 21–23 段)。Sutton 的反驳是本体论:合成数据本身「就是错的」,生成程序只能造出小世界,而真实世界无限复杂,连他人心智都不可能被合成(第 24–25 段)。

6. 从「大世界假说」到「必须持续学习」,Sutton 的推理链是什么?

第 25–26 段:世界包含许多其他智能体,因此必然比任何单个智能体更复杂;既然复杂到无法穷尽,就不可能做出最优或完美的东西,只能依靠粗糙的近似;而近似必须针对智能体实际所处的那一局部世界调校,不同局部需要不同近似;所以当智能体遇到新的局部时就必须重新学习。这样,持续学习就从工程偏好被推成逻辑必然。

7. 面对「自动驾驶靠仿真训练很成功」的反例,Javed 如何回应?

他并不否认仿真有价值,而是指出仿真背后有大批工程师反复迭代:先造仿真、在其中训练、发现仿真到现实的差距、再人工修补,人始终在回路里(第 27–28 段)。他的问题是:我们是否只愿解决那些能雇一大批工程师先造仿真的问题?他的替代方案是让智能体从自身经验学出模型并用于规划,因为模型错了可以通过持续学习自我修正,而人造仿真器只能等人发现问题(第 28–29 段)。

8. Javed 说持续学习的障碍「绝对是算法问题」,他给出的具体解法有哪两条?

第一是步长优化:网络中每个权重拥有独立的、通过元学习获得的步长,大多数权重步长极小,新样本不会破坏它们,更新只发生在正确的位置(第 66 段)。第二是特征空间的「生成与测试」:不依赖缓慢的梯度,直接提出一批随机初始化的新单元,由反向传播充当测试者筛选——即发表于《自然》的持续反向传播算法(第 67–69 段)。Javed 补充这些算法必须从头训练,不能套在现成模型上(第 70 段)。

9. 主持人提出「造火箭需要想象尚不存在的东西,这不是从经验学来的」,Sutton 和 Javed 会如何回应?

两人都承认必须能规划、能想象(第 54 段),但否认这需要经验之外的机制。Javed 的推理是:人脑一万年未显著进化,差别在于累积的文化知识;范式转变者是先积累知识、再从自身经验构建新抽象、用抽象规划、进而发现新知识——这条链的每一环仍是经验学习,只是当前系统缺少「自创抽象并用其规划」的能力(第 54–56 段)。Sutton 则提醒这种争论正滑向「先验 vs 学习」的错误框架,两者本应协同(第 57 段)。

10. Sutton 说语言只占智能的约 20%–25%,这一判断如何影响他对 LLM 前途的看法?

他明确肯定 LLM 是「完全出乎意料的重大科学突破」,攻下了语言这块符号方法的最后阵地(第 81–82 段)。但他的不满在于 LLM「假装成 AI 的全部」,而智能远不止流畅运用语言。因此当真正的持续学习智能体出现时,他认为「面临风险」的不是人类而是 LLM——它们已经「风光过一段」(第 81 段)。这与他对权重冻结却声称博士级能力的批评(第 33 段)一脉相承。

11. 如果把「大实验室困在局部极小值」的论点用于解释历史上的深度学习复兴,它还成立吗?

大体成立且可互证。Javed 的论点是:新范式必然先变差再变好,被产品锁死的大实验室无法承受倒退期,所以要靠小团队下赌注(第 78–79 段)。深度学习在 AI 寒冬中正是由少数「不合时宜」的学术团队(包括 Sutton 所在的阿尔伯塔)坚持下来,直到算力与数据成熟后才在 2012 年反超——这与 Sutton 2003 年在寒冬里「加倍下注」的经历(第 5 段)对应。不过反方也可指出:深度学习的最终胜利同样依赖了大公司的算力投入,因此小团队的作用是探路而非独立完成,Oak 的设想也可能需要类似的后期接力。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.125Dylan Patel – Two labs will soon control most of the world's workforce 下一期 · NO.127 →He won a Nobel here for AlphaFold. Then he left. - John Jumper
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY 内容仅供学习 · thesophielab.com