视频库 / NO.084ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.084
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

The surprising truth about AI’s limits and potential | Geoffrey Hinton | Strange Loop

节目发布 2024-05-20 · Sana
杰弗里·辛顿 主持人
本期追问 · 点击跳到视频对应位置
8:50 预测就是理解吗?27:20 机器会思考吗?13:00 人类还是特殊的吗?32:30 研究是怎样做成的?
归入 Ⅱ·03 能预言,就等于能解释吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
这是《Strange Loop》对杰弗里·辛顿(Geoffrey Hinton)的一次长谈。辛顿是深度学习的奠基者之一,1986 年与鲁梅尔哈特、威廉姆斯在《自然》上共同发表了反向传播的论文,也曾与特里·谢诺夫斯基长期合作研究玻尔兹曼机;伊利亚·苏茨克韦尔(Ilya Sutskever)当年就是敲开他办公室门的那个学生。对谈沿着时间线展开,从剑桥的生理学课堂一路谈到规模、推理、多模态、模拟计算,以及他如何选题、如何看人。本文依据现场录音编译整理,只删去口语枝节、寒暄与重复,论证与例子一概保留。

周六夜的实验室

辛顿: 我还记得自己刚从英国到卡内基梅隆的时候。在英国的研究单位,一到六点,大家就一起去酒吧喝一杯。到了卡内基梅隆,我去了几个礼拜之后的一个周六晚上,还没交到什么朋友,也不知道该干什么,就决定去实验室写点程序,因为我用的是 Lisp 机器,在家里根本没法编程。我大概晚上九点走进实验室,结果里面人山人海。所有学生都在,他们之所以在那儿,是因为他们手上做的东西就是未来。他们全都相信自己接下来做的事情会改变计算机科学的走向。这跟英国太不一样了,让人精神一振。

剑桥的失望

主持人: 我们回到最开始。你在剑桥想弄明白大脑,那是一种什么体验?

辛顿: 非常令人失望。我念的是生理学,夏季学期他们说要讲大脑是怎么工作的,结果只讲了神经元如何传导动作电位。这当然很有意思,但它并不告诉你大脑是怎么运作的。所以我极其失望。我转去学哲学,心想他们或许会讲心智是怎么运作的,结果同样令人失望。最后我去了爱丁堡做人工智能,那才有点意思,至少你可以做模拟,可以拿理论去检验。

主持人: 你还记得当年是什么东西把你吸引到人工智能上的吗?是一篇论文,还是某个把这些想法带给你的人?

辛顿: 我想主要是唐纳德·赫布(Donald Hebb)的一本书,对我影响很大。他非常关心神经网络里的连接强度是怎么学出来的。我也很早就读了冯·诺伊曼的一本书,他对大脑如何计算、以及大脑与普通计算机有何不同,都非常感兴趣。

主持人: 那时候你就已经笃定这套想法能成吗?在爱丁堡的年月里,你的直觉是什么?

辛顿: 在我看来,大脑必定有某种学习的办法,而且显然不是靠把各种东西预先编进去、再套用逻辑推理规则。这一点从一开始我就觉得荒唐。所以我们必须搞清楚大脑是怎么学会修改神经网络中的连接,从而做成复杂的事情。冯·诺伊曼相信这一点,图灵也相信这一点。冯·诺伊曼和图灵的逻辑都非常好,可他们偏偏都不信那条逻辑主义的路。

主持人: 在研究神经科学的想法和直接去找那些看起来管用的人工智能算法之间,你是怎么分配的?早期你从神经科学里取用了多少灵感?

辛顿: 神经科学我其实从来没怎么正经研究过。但我读到的那些关于大脑如何工作的知识始终在启发我:那里有一堆神经元,它们执行相对简单的操作,是非线性的,但它们收集输入、给输入加权,然后根据这个加权后的输入给出一个输出。问题就是,你怎么改变这些权重,让整个系统做出点好东西来。这看上去是个相当简单的问题。

谢诺夫斯基与彼得·布朗

主持人: 那段时间你记得哪些合作?

辛顿: 我在卡内基梅隆最主要的合作对象,恰恰不在卡内基梅隆。我跟当时在巴尔的摩约翰斯·霍普金斯的特里·谢诺夫斯基往来极多。大约每个月一次,要么他开车到匹兹堡,要么我开车去巴尔的摩,两地相距 250 英里,我们会一起花一个周末研究玻尔兹曼机。那是一段美妙的合作。我们两个都确信那就是大脑的工作方式。那是我这辈子做过的最令人兴奋的研究,也确实出了一批非常有意思的技术结果,不过我现在认为大脑并不是那样工作的。

我另外还有一段很好的合作,对象是彼得·布朗(Peter Brown),一位非常优秀的统计学家。他在 IBM 做语音识别,后来以一个比较年长的学生身份来卡内基梅隆,就为了拿个博士学位。但他已经懂得很多了。他教了我很多语音方面的东西,事实上还教了我隐马尔可夫模型。我觉得我从他那儿学到的,比他从我这儿学到的多。这正是你想要的那种学生。

他给我讲隐马尔可夫模型的时候,我正在做带隐藏层的反向传播,只不过那时还不叫“隐藏层”。我当时就觉得,隐马尔可夫模型里用的那个名字,用来称呼那些你不知道它们在干什么的变量,实在是好名字。神经网络里“隐藏”(hidden)这个说法就是这么来的:我和彼得认定,这个词拿来命名神经网络的隐藏层再合适不过。总之,我从彼得那里学到了很多语音方面的东西。

伊利亚来敲门

主持人: 说说伊利亚出现在你办公室门口的那天。

辛顿: 那大概是个星期天,我在办公室,我想是在写程序。这时有人敲门,不是随随便便的敲法,而是那种急促的敲法。我去开门,门口是个年轻学生,他说他整个夏天都在炸薯条,但他更想到我的实验室来干活。我说,那你不如先约个时间,我们再谈。他直接就说:“现在怎么样?”这就是伊利亚的性格。

于是我们聊了一会儿,我给了他一篇论文去读,就是《自然》上那篇反向传播的论文。我们约好一周后再见。他回来时说:“我没看懂。”我很失望。我本以为他是个聪明人,可那不过是链式法则,没那么难懂。结果他说:“不不不,那个我懂了!我不懂的是,你们为什么不把梯度交给一个像样的函数优化器?”这个问题,我们后来花了好几年才想明白。事情一直是这样:他对事物的原始直觉总是非常好。

主持人: 你觉得是什么造就了伊利亚的这些直觉?

辛顿: 我不知道。我想他一直都独立思考,从很小的时候就对人工智能感兴趣,而且他数学显然很好。但这种事很难说清楚。

主持人: 你们两个的合作是什么样的?你担当哪部分,伊利亚担当哪部分?

辛顿: 非常有意思。我记得有一次,我们想做一件挺复杂的事:给数据画图,我手上有一种混合模型,可以让同一批相似度关系生成两张图。在其中一张图里,bank(银行/河岸)可以离 greed(贪婪)很近;在另一张图里,bank 可以离 river(河)很近。因为在一张图里没法让它同时靠近这两者,river 和 greed 隔得太远了。所以我们要做一个由多张图构成的混合模型。我们是用 MATLAB 做的,这就牵涉到大量代码重组,好把矩阵乘法拼对。

伊利亚受够了这个。有一天他来了,说:“我要给 MATLAB 写个接口,这样我就用另一种语言编程,再有个东西把它转成 MATLAB。”我说:“别,伊利亚,那得花你一个月。我们得赶项目,别被这种事岔开。”伊利亚说:“没事,我今天上午已经写完了。”

变大就够了

主持人: 那真是了不起。在那些年里,最大的转变其实不只是算法,还有规模。这些年来你是怎么看待规模的?

辛顿: 伊利亚很早就有了那个直觉。他一直在宣讲一件事:你把它做得更大,它就会更好。我一直觉得那有点像逃避责任,觉得你终归还是得有新想法。结果证明伊利亚基本上是对的。新想法确实有帮助,比如 Transformer 就帮了大忙,但真正起决定作用的是数据的规模和计算的规模。当年我们完全想不到计算机会快上十亿倍,我们以为大概能快一百倍。于是我们绞尽脑汁去想那些聪明的点子,而如果当初有了更大的数据和计算规模,那些问题本来会自己消失。

大概在 2011 年,伊利亚、另一位叫詹姆斯·马滕斯(James Martens)的研究生和我写过一篇论文,用的是字符级预测。我们拿维基百科来预测下一个 HTML 字符,效果好得出奇,我们自己都一直很惊讶它能做到那种程度。那是在 GPU 上用一个很讲究的优化器跑的。我们始终不太敢相信它理解了什么,可看起来它就像是理解了,这简直难以置信。

预测下一个词

主持人: 能不能讲讲这些模型是怎么被训练来预测下一个词的,以及为什么这是理解它们的错误方式?

辛顿: 其实我并不认为这是错误的方式。事实上,我想第一个用嵌入(embedding)和反向传播做出来的神经网络语言模型就是我做的。数据非常简单,就是一些三元组。它把每个符号变成一个嵌入,然后让这些嵌入相互作用,去预测下一个符号的嵌入,再由此预测下一个符号,最后对整个过程做反向传播,把这些三元组学下来。我当时证明了它能够泛化。大约十年后,约书亚·本吉奥用了一个非常类似的网络,证明它在真实文本上也行得通。又过了大约十年,语言学家才开始相信嵌入这回事。这是个缓慢的过程。

我之所以认为它不只是在预测下一个符号,是因为你可以反问一句:预测下一个符号需要什么?尤其是当你问我一个问题,而答案的第一个词就是那个“下一个符号”的时候,你必须先理解这个问题。所以我认为,通过预测下一个符号,它跟老式的自动补全非常不同。老式自动补全是把词的三元组存起来,看到一对词,就统计各种词作为第三个词出现的频率有多高,这样来预测下一个符号。多数人以为自动补全就是那样。现在它已经不是那种工具了。要预测下一个符号,你必须理解正在说的是什么。所以我认为,你是在通过逼它预测下一个符号来逼它理解。而且我认为它的理解方式跟我们的相当类似。

很多人会告诉你,这些东西跟我们不一样,它们只是在预测下一个符号,它们没有像我们那样推理。可实际上,为了预测下一个符号,它就得做一些推理。我们现在也已经看到,即便你不往里面加任何专门用来推理的东西,只要把模型做大,它们就已经能做一些推理了。我认为随着模型继续变大,它们能做的推理会越来越多。

主持人: 你觉得我此刻做的事情,除了预测下一个符号之外还有别的吗?

辛顿: 我认为你就是这样学习的。我认为你在预测下一帧画面,在预测下一个声音。我觉得这是关于大脑如何学习的一个相当靠得住的理论。

类比即创造力

主持人: 是什么让这些模型能学会如此五花八门的领域?

辛顿: 这些大语言模型在做的事,是寻找共同结构。找到共同结构以后,它们就能用这个共同结构来编码事物,那样更高效。我举个例子。你去问 GPT-4:“堆肥堆为什么像原子弹?”多数人答不上来。多数人没想过这个,他们觉得原子弹和堆肥堆是截然不同的东西。但 GPT-4 会告诉你:两者的能量尺度差别很大,时间尺度差别也很大,可相同之处在于,堆肥堆越热,产热就越快;原子弹产生的中子越多,产生中子的速度也越快。于是它抓住了链式反应的概念。我相信它理解了两者都是链式反应的形式,而且正是靠这种理解,把海量信息压缩进了它的权重里。

如果它在做这件事,那么它在成百上千件我们还没看出类比的事情上也在做同样的事,而它已经看出来了。创造力就是从这里来的,就是从看出表面上截然不同的事物之间的类比而来的。所以我认为 GPT-4 变得更大以后,最终会非常有创造力。那种认为它只是在反刍它学过的东西、把学过的文本拼贴起来的看法,是完全错误的。我认为它会比人更有创造力。

超越训练数据

主持人: 你的意思是,它不只会重复人类迄今积累的知识,还能往前走。这是我们还没怎么见到的事。我们看到了一些苗头,但很大程度上还停留在当前的科学水平上。你觉得是什么能让它突破这一层?

辛顿: 在更受限的场景里我们已经见到了。比如 AlphaGo 与李世石那场著名的对局,其中的第 37 手,所有专家都说那肯定是个失误,但后来他们意识到那是神来之笔。那就是在一个有限领域里被创造出来的东西。我想随着这些系统变大,我们会见到更多这类情况。

主持人: AlphaGo 的另一个不同之处在于它用了强化学习,这让它后来能超越当时的水准。它先是模仿学习,看人类怎么下,然后通过自我对弈发展出超越人类的东西。你认为这是当下所缺的一环吗?

辛顿: 我认为这很可能确实是缺的一环。AlphaGo 和 AlphaZero 里的自我对弈,是它们能走出那种创造性着法的重要原因。但我不认为这是绝对必需的。

我很久以前做过一个小实验:训练一个神经网络识别手写数字。

主持人: 我很喜欢那个例子,MNIST 那个。

辛顿: 你给它的训练数据里有一半答案是错的。问题是它能学到什么程度。而且你把这一半答案弄错之后就固定不变,所以它没法靠反复看到同一个样本、有时是对答案有时是错答案来把错误平均掉。它每次看到那些样本中的一半时,答案永远是错的。于是训练数据有 50% 的错误率,可你用反向传播训练下去,它的错误率能降到 5% 甚至更低。换句话说,从标注很糟的数据里,它能得到好得多的结果。它能看出训练数据是错的。

聪明的学生之所以能比导师更聪明,也是这个道理。导师告诉他们一大堆东西,其中一半他们心想,胡说八道,另一半他们听进去了,最后他们比导师更聪明。所以这些大型神经网络确实能做到远远好过它们的训练数据,多数人没意识到这一点。

用推理修正直觉

主持人: 那你预计这些模型会怎样把推理加进去?一种路子是在上面叠启发式方法,现在很多研究就是这么做的,比如思维链,把它的推理反馈回它自己。另一种是在模型内部,随着规模扩大自然出现。你的直觉是什么?

辛顿: 我的直觉是,随着我们把这些模型做大,它们的推理能力会变强。如果问人是怎么运作的,粗略地说,我们有直觉,也能做推理,而我们用推理来纠正直觉。当然,我们在推理的过程中也在使用直觉,但如果推理的结论跟直觉相冲突,我们就意识到直觉需要改。

这很像 AlphaGo 或 AlphaZero:你有一个评估函数,它看一眼棋盘,说这个局面对我有多好;然后你做蒙特卡洛推演,得到一个更准确的判断,于是你可以修正评估函数。你可以通过让它与推理的结果保持一致来训练它。我认为这些大语言模型必须开始这么做。它们必须开始训练自己关于“接下来该出现什么”的原始直觉,方式是去推理,然后发现直觉不对。这样它们就能得到比单纯模仿人类更多的训练数据。这恰恰就是 AlphaGo 能走出创造性的第 37 手的原因:它有多得多的训练数据,因为它是用推理去核验下一步究竟该走哪里的。

多模态与空间

主持人: 你怎么看多模态?我们刚才谈到类比,而这些类比常常远远超出我们能看见的范围,它发现的类比远超人类,抽象层次可能是我们永远无法理解的。现在如果把图像、视频和声音引进来,你觉得这会怎样改变模型,又会怎样改变它能做出的类比?

辛顿: 我认为改变会很大。我认为它会在理解空间性的事物上强得多。举例来说,光靠语言,有些空间上的东西是相当难理解的,尽管 GPT-4 在还没有多模态之前就已经能做到这一点,这很了不起。但当你把它变成多模态,如果它既能看,又能伸手去抓东西,那么当它能把东西拿起来、翻过来看的时候,它对物体的理解会好得多。所以,虽然从语言里能学到极多的东西,但如果你是多模态的,学起来更容易,而且你需要的语言反而更少。另外,YouTube 上有海量视频可以用来预测下一帧之类。所以我认为多模态模型显然会成为主流:这样能拿到更多数据,需要的语言也更少。这里有一个真正的哲学层面的问题,就是你确实能只靠语言学出一个很好的模型,但从多模态系统里学要容易得多。

主持人: 你觉得这会怎样影响模型的推理?

辛顿: 我认为它在关于空间的推理上会强得多。比如推理“把物体拿起来会发生什么”,如果它真的去尝试拿起物体,就会得到各种各样有帮助的训练数据。

语言的三种看法

主持人: 你认为是人脑演化得适合语言,还是语言演化得适合人脑?

辛顿: 到底是语言演化来适应大脑,还是大脑演化来适应语言,我觉得这是个很好的问题。我认为两件事都发生了。我以前认为,很多认知活动根本不需要语言就能进行,现在我的看法有所改变。让我给出关于语言及其与认知关系的三种不同看法。

第一种是老派的符号主义看法:认知就是在某种清理干净、没有歧义的逻辑语言里持有一串串符号,并应用推理规则。认知就是这个,就是在类似语言符号串的东西上做符号操作。这是一个极端。

相反的极端是:不对,一旦进到脑袋里,全都是向量。符号进来,你把这些符号转换成大向量,里面所有的处理都用大向量完成;如果你要产生输出,再重新产生符号。大约 2014 年,机器翻译领域有那么一个阶段,人们用循环神经网络,词一个接一个进来,网络有一个隐藏状态,信息不断在这个隐藏状态里累积。到句子结束时,就有了一个捕捉了整句意思的大隐藏向量,然后可以拿它去生成另一种语言的句子。那个东西当时被称作“思想向量”(thought vector)。这是关于语言的第二种看法:你把语言转换成一个跟语言毫无相似之处的大向量,认知就发生在那里。

还有第三种看法,也是我现在相信的:你把这些符号转换成嵌入,并且用很多层来做,于是得到非常丰富的嵌入,但这些嵌入仍然是系在符号上的,也就是说,这个符号有一个大向量,那个符号有一个大向量,这些向量相互作用,产生出下一个词那个符号的向量。理解就是这么回事。理解就是知道怎么把符号转换成这些向量,以及知道向量的各个分量该怎样相互作用,才能预测出下一个符号的向量。在这些大语言模型里,理解是这样;在我们脑子里,理解也是这样。

这是一种居中的看法:你保留了符号,但你把它们解释为这些大向量。所有的功夫和所有的知识都在于你用什么向量、这些向量的元素如何相互作用,而不在符号规则里。但这并不是说你彻底摆脱了符号,而是说你把符号变成了大向量,同时仍然保留符号那层表面结构。这些模型就是这样运作的,而现在我觉得,这对人类思维来说也是个更可信的模型。

GPU 与那块显卡

主持人: 你是最早想到用 GPU 的人之一,我知道黄仁勋为此很喜欢你。2009 年前后你跟他说过这可能是训练神经网络的好主意。带我们回到那个早期的直觉吧。

辛顿: 其实我想是大概 2006 年,我有个从前的研究生叫里克·塞利斯基(Rick Szeliski),一位非常出色的计算机视觉学者。我在一次会议上跟他聊天,他说,你应该考虑用图形处理卡,因为它们做矩阵乘法非常在行,而你做的事情基本上全是矩阵乘法。我琢磨了一阵,后来我们知道了有那种装了四块 GPU 的 Tesla 系统。一开始我们只是买了游戏显卡,发现速度快了 30 倍。接着我们买了一台装四块 GPU 的 Tesla 系统,用它做语音,效果非常好。

然后 2009 年我在 NIPS 上做了个报告,对着一千个机器学习研究者说:“你们都该去买英伟达的 GPU,它们就是未来,做机器学习需要它们。”后来我给英伟达发了封邮件,说:“我跟一千个机器学习研究者说了去买你们的板卡,能不能送我一块?”他们说不行。其实他们也没说不行,他们干脆没回。不过后来我把这个故事讲给黄仁勋听,他送了我一块。

凡人计算

主持人: 我觉得有意思的一点是,GPU 是跟这个领域一起演化过来的。你认为计算这一块接下来该往哪里走?

辛顿: 在谷歌的最后两年,我一直在想办法做模拟计算(analog computation),这样就不用消耗一兆瓦,而是像大脑那样用大约 30 瓦,把这些大语言模型跑在模拟硬件上。我一直没做成,不过我因此开始真正体会到数字计算的好处。

如果你要用那种低功耗的模拟计算,每一件硬件都会有一点不同,而思路是让学习去利用那件特定硬件的具体属性。人就是这样。我们每个人的大脑都不一样,所以我们没法把你大脑里的权重取出来放进我的大脑,硬件不同,单个神经元的精确属性也不同,而学习已经学会了去利用这一切。

从这个意义上说,我们是凡人(mortal):我大脑里的权重对任何别的大脑都没用,我死了,那些权重就作废了。我们只能相当低效地把信息从一个脑子传到另一个脑子,办法是我说出句子,你去琢磨该怎样改变你的权重,才会说出同样的话。这叫蒸馏(distillation),但作为一种传递知识的方式,它效率极低。

数字系统则是不朽的(immortal)。一旦你有了一组权重,你可以把计算机扔掉,把权重存到某盘磁带上,然后造另一台计算机,把同样的权重放进去。只要它是数字的,它就能算出跟原来那台系统完全一样的东西。所以数字系统之间可以共享权重,这效率高得不可思议。如果你有一大批数字系统,它们各自去学一点点,起点是同样的权重,各学一点点,然后再共享权重,它们就全都知道了其他所有系统学到的东西。我们做不到这一点。所以在共享知识这件事上,它们远远优于我们。

快权重与时间尺度

主持人: 这个领域用到的很多想法都是老派的想法,是神经科学里早就存在的东西。你觉得还有什么可以拿来用到我们造的系统上?

辛顿: 我们还得向神经科学补的一大课,是变化的时间尺度。几乎所有神经网络里,都只有一个快的时间尺度用于改变活动,输入进来,活动、嵌入向量全都变化;再有一个慢的时间尺度用于改变权重,那是长期学习。你就只有这两个时间尺度。

而在大脑里,权重变化的时间尺度有很多种。比方说,我说了一个出人意料的词,比如“黄瓜”,五分钟后你戴上耳机,里面噪声很大,词很轻微,那你识别出“黄瓜”这个词的能力会强得多,因为我五分钟前说过它。那么这份知识在大脑的什么地方?显然在突触的临时变化里。不是有神经元在那儿一直念叨“黄瓜、黄瓜、黄瓜”,你没有那么多神经元可用。它在权重的临时变化里。

用临时的权重变化,也就是我所说的快权重(fast weights),你能做很多事情。我们在这些神经模型里不这么做。不这么做的原因是,如果权重的临时变化取决于输入数据,你就没法同时处理一大批不同的样例。现在我们的做法是把一大堆不同的字符串堆叠在一起并行处理,因为那样就能做矩阵与矩阵的乘法,效率高得多。恰恰是这个效率考虑,挡住了我们使用快权重。可大脑显然在用快权重做临时记忆,用这种方式能做的很多事情,我们目前都没在做。我认为这是我们必须补上的最大的一课之一。我曾很寄望于 Graphcore 这类东西,如果它们走顺序处理、只做在线学习的路子,就能用上快权重。可到现在还没走通。我认为等到人们用电导来实现权重的时候,最终是会走通的。

先天结构与乔姆斯基

主持人: 了解这些模型如何工作、以及大脑如何工作,对你的思考方式有什么影响?

辛顿: 我想有一个很大的影响,层次比较抽象:很多年里,人们对“弄一个大的随机神经网络,喂给它大量训练数据,它就能学会做复杂的事”这种想法极为不屑。你去问统计学家、语言学家或者人工智能界的大多数人,他们会说那是白日做梦,没有某种先天知识、没有大量的架构限制,你休想学会真正复杂的东西。结果证明这完全是错的。你可以拿一个大的随机神经网络,纯粹从数据里学到一大堆东西。所以,用随机梯度下降反复地依据梯度调整权重,就能学到东西,而且能学到又大又复杂的东西,这一点已经被这些大模型验证了。这是关于大脑的一件非常重要的认识:它不必带有全部这些先天结构。当然它显然有很多先天结构,但对于那些容易学会的东西,它肯定不需要先天结构。

所以,来自乔姆斯基的那个想法,认为像语言这样复杂的东西你学不会,除非它早就被接好线路、只等成熟展开,这个想法现在显然是胡说。

主持人: 我相信乔姆斯基会很欣赏你把他的想法称为胡说。

辛顿: 其实我觉得乔姆斯基很多政治见解都很在理。我一直觉得奇怪,一个在中东问题上有如此理智见解的人,怎么会在语言学上错得这么离谱。

机器会有感受吗

主持人: 你觉得什么能让这些模型更有效地模拟人的意识?设想你有一个 AI 助理,你一辈子都在跟它说话,而它不像今天的 ChatGPT 那样每次都删掉对话记忆、从头开始,它有自我反思的能力。到了某个时刻你去世了,你把这件事告诉那个助理,你认为它……

辛顿: 我是说,不是我自己,是别人告诉那个助理。

主持人: 对,你自己确实不太好告诉它。你认为那个助理到那时会有感受吗?

辛顿: 会。我认为它们也可以有感受。我们对知觉有一套“内在剧场”的模型,我认为我们对感受也有一套内在剧场的模型:有些东西是我能体验而别人不能体验的。我认为这个模型同样是错的。

假设我说“我真想一拳打在加里鼻子上”,这种话我常说。让我们试着把它从内在剧场那套说法里剥出来。我真正在对你说的是:如果不是我的额叶发出抑制,我就会做出一个行动。所以当我们谈论感受时,我们其实是在谈论如果没有约束我们就会做出的行动。感受真正的所指,就是那些若无约束我们便会付诸实施的行动。所以我认为对感受也可以给出同一类解释,没有任何理由说这些东西不能有感受。

事实上,1973 年我就见过一个机器人有情绪。在爱丁堡有一台机器人,有两个像这样的夹爪,如果你把零件分开摆在一块绿色毛毡上,它能把一辆玩具车组装起来。可要是你把零件堆成一堆,它的视觉就不足以看清是什么情况了。于是它把两个夹爪并在一起,“砰”地一下打过去,把那堆零件打散,然后它就能把车装起来了。如果你在一个人身上看到这一幕,你会说他是在对这个局面发火,因为他搞不懂它,所以就把它毁了。

最有力的类比

主持人: 这很深刻。我们上次交谈时,你把人和大语言模型都描述成类比机器。你这一生中找到过的最有力的类比是什么?

辛顿: 这一生里?我想,一个对我影响很大的、比较弱的类比,是宗教信仰与符号处理之间的类比。我很小的时候面对过这件事:我出身于一个无神论家庭,进了学校却撞上宗教信仰,我只觉得那是胡说八道,直到现在我仍然觉得那是胡说八道。后来我看到有人把符号处理当作人如何运作的解释,我觉得那是一模一样的胡说八道。

现在我不觉得它那么胡说了,因为我认为我们其实确实在做符号处理,只不过我们是通过给符号配上这些大嵌入向量来做的。我们确实在做符号处理,但完全不是人们原先设想的那种方式:那种方式里你去匹配符号,而一个符号唯一的属性就是它跟另一个符号相同或者不相同。这是符号唯一的属性。我们根本不那么干。我们利用上下文给符号赋予嵌入向量,然后利用这些嵌入向量各分量之间的相互作用来思考。

不过谷歌有位很优秀的研究者叫费尔南多·佩雷拉(Fernando Pereira),他说过:是的,我们确实有符号推理,而我们拥有的唯一的符号语言就是自然语言。自然语言是一种符号语言,我们用它来推理。我现在相信这个说法。

怎样挑问题

主持人: 你做出了计算机科学史上最有分量的一些研究。能讲讲你是怎么挑选该做的问题的吗?

辛顿: 先纠正你一下。是我和我的学生们做出了很多最有分量的东西,主要靠的是跟学生的良好合作,以及我挑到很好学生的本事。而这本事又来自一个事实:七十、八十、九十年代直到本世纪初,做神经网络的人非常少,所以这少数几个做神经网络的人得以挑到最好的学生。这纯属运气。

至于我挑问题的办法,基本上是这样。你知道,科学家谈自己怎么工作时,都有一套关于自己怎么工作的理论,而那套理论多半跟事实关系不大。我的理论是:我去找一件所有人都有共识、但让我觉得不对劲的事,就是有那么一丝直觉,觉得它哪里不对。然后我就琢磨这件事,看能不能把“我为什么觉得它不对”阐述清楚,也许还能写个小程序做个小演示,表明事情并不像你以为的那样运作。

我举个例子。多数人认为,如果你往神经网络里加噪声,它会变差。比如每次送一个训练样本进去,你让一半的神经元静默,效果会更糟。可实际上我们知道,你这么做它的泛化会更好,而且你能证明这一点。用一个简单的例子就能证明,这正是计算机模拟的好处:你可以摆出来给人看,你原以为加噪声会让它变糟,让一半神经元失活会让它变糟,短期内确实如此,但如果你就这么训练下去,最终它会更好。你能用一个小程序演示出来,然后再狠狠地想一想这是为什么,想清楚它是怎样阻止了大规模的复杂协同适应。这就是我的工作方法:找一件听起来可疑的事去做,看能不能给出一个简单的演示,说明它为什么是错的。

主持人: 现在有什么让你觉得可疑?

辛顿: 我们不用快权重,这件事就很可疑,我们只有那两个时间尺度,这是不对的,完全不像大脑。长远看,我认为我们必须有多得多的时间尺度。这就是一个例子。

主持人: 如果今天你带着一批学生,他们来问你我们之前提过的那个哈明式的问题:“你所在领域最重要的问题是什么?”你会建议他们接下来去做什么?我们谈过推理,谈过时间尺度。你会给出的最高优先级问题是什么?

辛顿: 对我来说,眼下还是我过去三十年一直在问的那个问题:大脑做反向传播吗?我相信大脑在获取梯度。如果你拿不到梯度,你的学习就比拿得到梯度差得多。但大脑是怎么拿到梯度的?它是在以某种方式实现反向传播的某个近似版本,还是在用一种完全不同的技术?这是个悬而未决的大问题。如果我继续做研究,我要做的就是这个。

错了也不后悔

主持人: 回看你的职业生涯,你在很多事情上都判断对了。但有什么是你判断错了、并且希望当初少花点时间的?

辛顿: 这是两个不同的问题。一个是“你在什么上错了”,另一个是“你是否希望当初少花点时间”。我认为我在玻尔兹曼机上错了,可我很庆幸自己在上面花了很长时间。作为一套关于如何获得梯度的理论,它比反向传播漂亮得多。反向传播平平无奇,很在理,不过是链式法则。玻尔兹曼机是巧妙的,是一条非常有意思的获取梯度的路。我很希望大脑就是那样工作的,但我认为它不是。

主持人: 你有没有花很多时间去想象这些系统发展起来之后会怎样?你有没有过这样的念头:如果我们能让这些系统真正好用,就能让教育普及化,让知识变得可及得多,还能解决医学上的一些难题?还是说对你来说更多是为了理解大脑?

辛顿: 我确实觉得科学家应该去做对社会有益的事情,但那并不是你做出最好研究的方式。你做出最好的研究,是在被好奇心驱动的时候,你就是非得把某样东西弄明白不可。最近这些年我意识到这些东西可能带来很多好处,也可能带来很多危害,我对它们将给社会造成的影响忧心得多了。但当初驱动我的不是这个。我只是想弄明白,大脑究竟怎么可能学会做这些事情。这是我想知道的。而我算是失败了。作为这次失败的副产品,我们得到了一些不错的工程成果,不过……

主持人: 对世界来说,这是一次很好的失败。

医疗、坏人与竞速

主持人: 如果从“可能会非常顺利”的那一面看,你认为最有前景的应用是什么?

辛顿: 医疗显然是一大块。在医疗上,社会能吸纳多少几乎没有上限。拿一个老年人来说,他可以让五个医生全职服务。所以,当人工智能在某些事情上做得比人好时,你希望它变强的领域,是那些你巴不得能多来一些的领域。而我们确实巴不得有多得多的医生。如果每个人都有三个自己的医生,那太好了,而我们正会走到那一步。这是医疗之所以是好方向的一个原因。

另外还有全新的工程领域,比如开发新材料,用于更好的太阳能板,或者用于超导,或者只是为了搞清楚身体是怎么运作的。那些方面会有巨大的影响。这些都会是好事。我担心的是坏人拿它去干坏事。我们已经为普京、习或者特朗普这样的人提供了便利,让他们可以把人工智能用于杀人机器人,用于操纵公众舆论,用于大规模监控。这些都非常令人忧虑。

主持人: 你会不会担心,让这个领域慢下来,也会把好的那一面一起拖慢?

辛顿: 当然会。而且我认为这个领域慢下来的可能性不大,部分原因是它是跨国的:一个国家慢下来,别的国家不会跟着慢。中美之间显然有一场竞速,两边都不会减速。所以我不觉得会慢下来。当时有那封请愿书说我们应该暂停六个月,我没有签,只是因为我认为那根本不会发生。也许我应该签,因为哪怕它不会发生,它也表明了一个政治态度。要一些你明知要不到的东西,仅仅为了表明态度,往往是件好事。但我当时并不认为我们会慢下来。

主持人: 你觉得有了这些助手之后,人工智能的研究过程会受到什么影响?

辛顿: 我觉得效率会高很多。当你有了这些助手帮你写程序,还能帮你把事情想清楚,很可能在方程式上也帮你不少忙,人工智能研究的效率会大大提高。

怎样看人

主持人: 你对如何挑选人才想得多吗?还是基本上靠直觉?比如伊利亚出现在门口,你感觉这人聪明,那就一起干?

辛顿: 挑人才这件事,有时候你就是知道。跟伊利亚没聊多久,他就显得非常聪明;再多聊一点,他显然非常聪明,而且直觉极好,数学也好。所以那根本不用想。

还有一次,我在 NIPS 会议上,我们有一张海报,有个人过来开始就海报提问,而他问的每一个问题,都是对我们哪里做错了的深刻洞察。五分钟之后我就给了他一个博士后职位。那个人是大卫·麦凯(David MacKay),才华横溢,他去世了,很令人难过,但当时非常明显,你就是想要他。

也有些时候没那么明显。我学到的一件事是,人是不一样的。好学生不止一种类型。有些学生不那么有创造力,但技术上极强,什么东西都能给你做成;有些学生技术不强,但极有创造力。当然你想要两样都行的,但你不总能碰上。我觉得实验室里其实需要各种不同类型的研究生。不过我还是靠我的直觉:有时候你跟某个人一聊,他就是非常非常那个,他就是通透,那些就是你想要的人。

直觉从哪里来

主持人: 你觉得有些人直觉更好,原因是什么?是他们的训练数据更好吗?直觉又该怎么培养?

辛顿: 我想一部分原因是他们不容忍胡说八道。这里有一条把直觉搞坏的路子:别人说什么你都信。那是致命的。我想有些人是这么做的:他们有一整套理解现实的框架,别人跟他们说一件事,他们会去琢磨这件事怎么装进自己的框架;如果装不进去,他们就干脆拒绝它。这是个很好的策略。那些试图把听到的一切都吸纳进来的人,最后会得到一个非常模糊的框架,什么都能信,那就没用了。

所以我认为,对世界有一套强硬的看法,并试图把进来的事实拿捏成符合你看法的样子,这确实可能把你引向极深的宗教信仰和致命的谬误,比如我对玻尔兹曼机的执念,但我认为这仍然是该走的路。如果你有值得信赖的好直觉,你就应该信它;如果你的直觉很糟,那你怎么做都无所谓,那还不如信它。

全押还是分散

主持人: 说得非常好。看今天正在做的这些研究,你觉得我们是不是把鸡蛋都放在一个篮子里,应该把想法再分散一些?还是说这就是最有前景的方向,我们该全押?

辛顿: 我认为,做大模型、用多模态数据训练它们,哪怕训练目标只是预测下一个词,这条路前景好到我们几乎应该全押。显然现在做这件事的人非常非常多,也有很多人在做看起来疯疯癫癫的事,这很好。但我觉得大多数人沿着这条路走没有问题,因为它效果非常好。

主持人: 你认为学习算法真的那么要紧吗,还是说这更多是个规模问题?通往人类水平的智能,是有成百万上千万条路,还是只有少数几条我们必须发现的路?

辛顿: 关于特定的学习算法是否极其重要,还是说有五花八门的学习算法都能把活干成,这个问题我不知道答案。不过在我看来,反向传播在某种意义上是正确的做法:获取梯度,据此改变参数让系统变得更好,这看上去就是该做的事,而且它成功得惊人。很可能还有别的学习算法,它们是获得同一个梯度的另一些途径,或者是在获取关于别的东西的梯度,而它们同样管用。我认为这都还是开放的,是个非常有意思的问题:是不是还有别的东西你可以试着去最大化,同样能得到好系统?也许大脑就在做那种事,因为那样更容易。但反向传播在某种意义上是正确的做法,而且我们知道,这么做效果非常好。

最得意的是什么

主持人: 最后一个问题。回望你数十年的研究,你最引以为豪的是什么?是学生,还是研究?回望你一生的工作,什么最让你自豪?

辛顿: 玻尔兹曼机的学习算法。玻尔兹曼机的学习算法漂亮得优雅,在实践中也许毫无希望,但那是我做起来最享受的东西,是我和特里一起搞出来的,也是我最引以为豪的。哪怕它是错的。

主持人: 你现在把大部分时间花在思考什么问题上?是那个……

辛顿: “我该在 Netflix 上看点什么?”

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 CMU 周六夜的实验室与科研气氛 ▶ 正在看
1:11 剑桥到爱丁堡:大脑如何学习 ▶ 正在看
5:05 Ilya 敲门,与「变大就够了」的直觉 ▶ 正在看
8:50 预测下一个词,就是被迫理解 ▶ 正在看
13:00 类比、创造力与超越训练数据 ▶ 正在看
15:13 用推理生产训练数据来修正直觉 ▶ 正在看
16:49 多模态与语言在认知中的位置 ▶ 正在看
21:04 GPU、模拟计算与「凡人计算」 ▶ 正在看
25:00 多重时间尺度与被效率挡住的快权重 ▶ 正在看
27:20 先天结构、乔姆斯基与机器的感受 ▶ 正在看
32:30 如何选题、选人,以及直觉如何养成 ▶ 正在看
36:51 医疗前景、滥用风险与未解之问 ▶ 正在看
本期小问 · 档案清单
8:50 预测就是理解吗? ▶ 正在看
27:20 机器会思考吗? ▶ 正在看
13:00 人类还是特殊的吗? ▶ 正在看
32:30 研究是怎样做成的? ▶ 正在看
本期讲者
杰弗里·辛顿深度学习奠基者之一,1986年与Rumelhart、Williams共同发表反向传播的《自然》论文,与Terry Sejnowski共同提出玻尔兹曼机,2012年团队的AlexNet与Dropout改变了计算机视觉。2018年图灵奖、2024年诺贝尔物理学奖得主,2023年离开谷歌后公开警示AI风险。
主持人访谈提问者,负责按时间线追问Hinton的学术历程、技术判断与研究方法,问题多围绕规模、推理、多模态与人才选拔展开。
01CMU 周六夜的实验室与科研气氛
0:00
Have you reflected a lot on how to select talent or has that mostly been intuitive to you? Ilya just shows up and you're like, this is a clever guy, let's, let's work together. Or have you thought a lot about that? Should we roll this? Yeah, let's roll this.
你在选拔人才这件事上思考过很多吗,还是主要靠直觉?比如 Ilya 一出现,你就觉得,这人很聪明,我们我们一起干吧。还是说你对此有过很多思考?我们开始录吗?好,开始录吧。
便签笔记
0:25
Sound is working. So I remember when I first got to Carnegie Mellon from England, in England at a research unit, it would get to be six o'clock and you'd all go for a drink in the pub. Um, at Carnegie Mellon, I remember after I'd been there a few weeks, it was Saturday night. I didn't have any friends yet and I didn't know what to do. So I decided I'd go into the lab and do some programming 'cause I had a list machine and you couldn't program it from home. So I went into the lab at about nine o'clock on a Saturday night and it was swarming. All the students were there and they were all there because what they were working on was the future. They all believed that what they did next was gonna change the course of computer science and it was just so different from England. And so that was very refreshing.
声音是正常的。我记得我刚从英国来到卡内基梅隆的时候,在英国的研究单位里,一到六点,大家就都去酒吧喝一杯。嗯,在卡内基梅隆,我记得我到那儿几周后的一个周六晚上。我还没交到朋友,也不知道该干什么。所以我决定去实验室写点程序,因为我有一台 Lisp 机器,在家里没法用它编程。于是我周六晚上九点左右走进实验室,那里人来人往,热闹得很。所有学生都在,他们都在那儿,是因为他们做的就是未来。他们都相信自己接下来做的事会改变计算机科学的进程,这和英国太不一样了。所以那让人非常振奋。
便签笔记
02剑桥到爱丁堡:大脑如何学习
1:11
Take me back to the very beginning, Geoff at Cambridge. Uh, trying to understand the brain. Uh, what was that like? It was very disappointing. So I did physiology and in the summer term they were gonna teach us how the brain worked and it, all they taught us was how neurons conduct action potentials, which is very interesting, but it doesn't tell you how the brain works. So that was extremely disappointing. I switched to philosophy then, I thought maybe they'd tell us how the mind worked and that was very disappointing. I eventually ended up going to Edinburgh to do AI and that was more interesting. At least you could simulate things so you could test out theories.
带我回到最开始吧,Geoff 在剑桥的时候。呃,试图理解大脑。呃,那是种什么样的体验?非常令人失望。我当时学的是生理学,夏季学期他们要教我们大脑是怎么运作的,结果他们只教了神经元如何传导动作电位,这固然很有意思,但它并不能告诉你大脑是怎么工作的。所以那让我极其失望。于是我转去学哲学,我想也许他们会告诉我们心智是怎么运作的,结果也非常令人失望。我最后去了爱丁堡做人工智能,那有意思多了。至少你可以做模拟,可以检验理论。
便签笔记
1:49
And did you remember what intrigued you about AI? Was it a paper? Was it any particular person that exposed you to those ideas? I guess it was a book I read by Donald Hebb that influenced me a lot. Um, he was very interested in how you learn the connection strengths in neural nets. I also read a book by John von Neumann early on, um, who was very interested in how the brain computes and how it's different from normal computers. And did you get that conviction that this ideas would work out at that point, or what was your intuition back in the Edinburgh days?
你还记得人工智能里是什么吸引了你吗?是某篇论文吗?还是某个特定的人把这些想法带给了你?我想是我读的 Donald Hebb 写的一本书对我影响很大。嗯,他非常关注神经网络中连接强度是如何习得的。我很早还读过 John von Neumann 的一本书,嗯,他非常关注大脑是如何计算的,以及它与普通计算机有什么不同。那时候你就确信这些想法会成功吗,还是说在爱丁堡那段时期你的直觉是什么?
便签笔记
2:29
It seemed to me there has to be a way that the brain learns and it's clearly not by having all sorts of things programmed into it and then using logical rules of inference, that just seemed to me crazy from the outset. Um, so we had to figure out how the brain learned to modify connections in a neural net so that it could do complicated things. And von Neumann believed that. Turing believed that. So von Neumann and Turing were both pretty good at logic, but they didn't believe in this logical approach.
在我看来,大脑一定有某种学习方式,而且显然不是把各种东西都编程进去、然后用逻辑推理规则,从一开始我就觉得那太荒唐了。嗯,所以我们必须搞清楚大脑是怎么学会修改神经网络中的连接的,从而能做复杂的事情。冯·诺依曼相信这一点,图灵也相信这一点。冯·诺依曼和图灵在逻辑上都很厉害,但他们并不相信那种逻辑路线。
便签笔记
3:01
And what was your split between studying the ideas from, from neuroscience and just doing what seemed to be good algorithms for, for AI? How much inspiration did you take early on? So I never did that much studying in neuroscience. I was always inspired by what I learned about how the brain works. That there's a bunch of neurons, they perform relatively simple operations, they're nonlinear, um, but they collect inputs, they weight them and then they give an output that depends on that weighted input. And the question is how do you change those weights to make the whole thing do something good? It seems like a fairly simple question.
那你在研究神经科学的思想和单纯去做那些看起来不错的人工智能算法之间,是怎么分配的?早期你从中汲取了多少灵感?我其实从来没在神经科学上花太多功夫钻研。我一直是被自己了解到的大脑运作方式所启发。就是有一堆神经元,它们执行相对简单的运算,它们是非线性的,嗯,但它们收集输入、给这些输入加权,然后根据加权后的输入给出一个输出。问题在于,你要如何改变这些权重,才能让整个系统做出好的事情?这看起来是个相当简单的问题。
便签笔记
3:38
What collaborations do you remember from from that time? The main collaboration I had at Carnegie Mellon was with someone who wasn't at Carnegie Mellon. I was interacting a lot with Terry Sejnowski who was in Baltimore at Johns Hopkins. And about once a month, either he would drive to Pittsburgh or I would drive to Baltimore. It's 250 miles away and we would spend a weekend together working on Boltzmann machines. That was a wonderful collaboration. We were both convinced it was how the brain worked. That was the most exciting research I've ever done. And a lot of technical results came out that were very interesting, but I think it's not how the brain works. Um, I also had a very good collaboration with um, Peter Brown, who was a very good statistician and he worked on speech recognition at IBM and then he came as a more mature student to Carnegie Mellon just to get a PhD.
那段时间你还记得哪些合作?我在卡内基梅隆最主要的合作对象,其实并不在卡内基梅隆。我当时和特里·塞诺夫斯基(Terry Sejnowski)来往很多,他在巴尔的摩的约翰霍普金斯大学。大概每个月一次,要么他开车到匹兹堡,要么我开车去巴尔的摩,两地相距250英里,我们会花一个周末在一起研究玻尔兹曼机。那是一段非常美妙的合作。我们俩都确信那就是大脑的工作方式。那是我做过的最令人兴奋的研究。也产生了很多非常有意思的技术成果,但我现在认为那并不是大脑的运作方式。嗯,我还有过一段非常好的合作,是和彼得·布朗(Peter Brown),他是一位很出色的统计学家,在IBM做语音识别,后来他作为一名比较成熟的学生来到卡内基梅隆,就是为了拿个博士学位。
便签笔记
4:26
Um, but he already knew a lot. He taught me a lot about speech and he in fact taught me about hidden Markov models. I think I learned more from him than he learned from me. That's the kind of student you want. And when he taught me about hidden Markov models, I was doing backdrop with hidden layers and they weren't called hidden layers then. And I decided that name they use in Hidden Markoff models is a great name for variables that you dunno what they're up to. Um, and so that's where the name 'hidden' in neural nets came from me and Peter decided that was a great name for the hidden layers in neural nets. Um, but I learned a lot from Peter about speech.
嗯,但他当时已经懂得很多了。他教了我很多关于语音的知识,实际上还教了我隐马尔可夫模型。我觉得我从他那儿学到的比他从我这儿学到的还多。这才是你想要的那种学生。他给我讲隐马尔可夫模型的时候,我正在做带隐藏层的反向传播,不过那时候还不叫“隐藏层”。我就觉得,隐马尔可夫模型里用的那个名字,用来称呼那些你不知道它们在干什么的变量,真是个绝妙的名字。嗯,所以神经网络里“隐藏”这个说法就是这么来的:我和彼得都觉得,用它来命名神经网络里的隐藏层非常合适。嗯,我从彼得那里学到了很多关于语音的东西。
便签笔记
03Ilya 敲门,与「变大就够了」的直觉
5:05
Take us back to, um, Ilya showed up at your office. I was in my office probably on a Sunday. Um, and I was programming I think, and there was a knock on the door, not just any knock, but it went kind of [knocks on table] sort of an urgent knock. So I went and answered the door and this was this young student there and he said he was cooking fries over the summer, but he'd rather be working in my lab. And so I said, well why don't you make an appointment and we'll talk. And so he just said, "How about now?" And that sort of was Ilya's character. So we talked for a bit and I gave him a paper to read, which was the nature paper on backpropagation. And we made another meeting for a week later and he came back and he said, "I didn't understand it", and I was very disappointed. I thought he seemed like a bright guy, but it's only the chain rule. It's not that hard to understand. And he said, "Oh no, no, I understood that! I just don't understand why you don't give the gradient to a sensible function optimizer", which took us quite a few years to think about. Um, and it kept
带我们回到那一刻吧,嗯,伊利亚(Ilya)出现在你办公室门口。我当时在办公室里,大概是个周日。嗯,我想我正在编程,然后有人敲门,不是随便敲敲,而是那种[敲桌子]很急促的敲法。于是我去开门,门口站着这个年轻学生,他说他整个夏天都在炸薯条,但他更想在我的实验室里工作。于是我说,要不你先预约一下,我们再谈。结果他直接说:“现在怎么样?”这大概就是伊利亚的性格。于是我们聊了一会儿,我给了他一篇论文让他读,就是那篇发表在《自然》上的反向传播论文。我们约好一周后再见,他回来后说:“我没看懂。”我当时很失望。我觉得他看起来是个很聪明的人,可那不过就是链式法则,没那么难懂啊。结果他说:“哦不不,那个我懂!我只是不明白,你为什么不把梯度交给一个像样的函数优化器?”这个问题我们花了好几年才想明白。嗯,而且一直
便签笔记
6:07
on like that with, he had very good, his raw intuitions about things were always very good. What do you think had enabled those, uh, those intuitions for, for Ilya? I don't know. I think he always thought for himself, he was always interested in AI from a young age. Um, he's obviously good at math, so, but it's very hard to know. And what was that collaboration between uh, the two of you like? What part would you play and what part would Ilya play? It was a lot of fun. Um, I remember one occasion when we were trying to do a complicated thing with producing maps of data where I had a kind of mixture model. So you could take the same bunch of similarities and make two maps so that in one map, bank could be close to greed and in another map, bank could be close to river. Um, 'cause in one map you can't have it close to both, right? 'cause river and greed along way part. So we'd have a mixture maps and we were doing it in MATLAB and this involved a lot of reorganization of the code to do the right matrix multiplies.
都是这样,他的那些原始直觉总是非常好。你觉得是什么让伊利亚有了这些直觉?我不知道。我想他一直都是独立思考的人,从很小的时候就对人工智能感兴趣。嗯,他数学显然很好,但这真的很难说。那你们俩之间的合作是什么样的?你负责哪部分,伊利亚又负责哪部分?那非常有意思。嗯,我记得有一次,我们想做一件挺复杂的事:用一种混合模型来生成数据的映射图。也就是说,你可以用同一组相似度关系画出两张图,在一张图里,“bank”可以离“greed(贪婪)”很近,而在另一张图里,“bank”可以离“river(河流)”很近。因为在一张图里你没法让它同时靠近两者,对吧?因为river和greed离得很远。所以我们要做混合的映射图,当时是用MATLAB做的,这就需要大量重组代码,才能做对矩阵乘法。
便签笔记
7:09
And Ilya got fed up with that. So he came one day and said, um, "I'm gonna write an interface for MATLAB. So I program in this different language and then I have something that just converts it into MATLAB". And I said, "No Ilya, that'll take you a month to do. We've gotta get on with this project. Don't get diverted by that. And Ilya said, "It's OK, I did it this morning" That's uh, that's quite, quite incredible. And throughout those years, the biggest shift wasn't necessarily just the algorithms, but also the scale. How did you sort of view that scale? Uh, over, over the years?
伊利亚受不了这个。所以有一天他过来说:“我打算给MATLAB写个接口。我用另一种语言来编程,然后有个东西把它转换成MATLAB。”我说:“别啊伊利亚,那得花你一个月时间。我们得推进这个项目,别被那个岔开了。”结果伊利亚说:“没事,我今天上午已经写好了。”这真是,呃,相当、相当不可思议。在那些年里,最大的转变其实不只是算法,还有规模。这些年来你是怎么看待规模这件事的?
便签笔记
7:49
Ilya got that intuition very early. So Ilya was always preaching that, um, "You just make it bigger and it'll work better". And I always thought that was a bit of a cop-out that you're gonna have to have new ideas too. It turns out Ilya was basically right. New ideas help, things like transformers helped a lot, but it was really the scale of the data and the scale of the computation. And back then we had no idea computers would get like a billion times faster. We thought maybe they'd get a hundred times faster. We were trying to do things by coming up with clever ideas that would've just solved themselves if we'd had bigger scale of the data and computation. In about 2011, Ilya and another graduate student called James Martins and I had a paper using character level prediction. So we took Wikipedia and we tried to predict the next HTML character and that worked remarkably well and we were always amazed at how well it worked. And that was using a fancy optimizer on GPUs and we could never quite believe that
伊利亚很早就有那个直觉。他一直在鼓吹:“你只要把它做得更大,它就会更好用。”而我一直觉得那有点像是逃避问题,觉得你总还是得有新想法才行。结果证明伊利亚基本上是对的。新想法当然有帮助,像Transformer这样的东西帮助很大,但真正起作用的是数据的规模和计算的规模。而在那时候,我们根本想不到计算机会快上十亿倍。我们以为大概会快个一百倍吧。我们当时绞尽脑汁想出各种聪明的点子,可如果数据和计算的规模再大一些,那些问题本来自己就解决了。大约在2011年,我和伊利亚,还有另一位叫詹姆斯·马腾斯(James Martens)的研究生发表了一篇论文,用的是字符级预测。我们拿维基百科的数据,试着预测下一个HTML字符,效果好得出奇,我们一直都很惊讶它能做得这么好。那是在GPU上用了一个很讲究的优化器。我们始终有点不敢相信
便签笔记
04预测下一个词,就是被迫理解
8:50
it understood anything, but it looked as though it understood and that just seemed incredible. Can you take us through how are these models trained to predict the next word and why is it the wrong way of, of thinking about them? OK. I don't actually believe it is the wrong way. So in fact, I think I made the first neural net language model that used embeddings and backpropagation. So it's very simple data just triples and it was turning each symbol into an embedding, then having the embeddings interact to predict the embedding of the next symbol and then from that predict the next symbol. And then it was backpropagating through that whole process to learn these triples. And I showed it could generalize. Um, about 10 years later, Yoshua Bengio used a very similar network and showed it worked with real text. And about 10 years after that linguists started believing in embeddings. It was a slow process. The reason I think it's not just predicting the next symbol is if you ask, "Well what does it take to predict the next symbol?"
它真的理解了什么,但看起来它就像是理解了,这实在令人难以置信。你能给我们讲讲这些模型是怎么被训练来预测下一个词的吗?以及为什么这种理解方式是错的?好吧。其实我并不认为这种说法是错的。事实上,我想我做出了第一个使用嵌入(embedding)和反向传播的神经网络语言模型。它的数据非常简单,就是一些三元组,它把每个符号变成一个嵌入向量,然后让这些嵌入相互作用,去预测下一个符号的嵌入,再由此预测出下一个符号。然后它对整个过程做反向传播,从而学会这些三元组。我证明了它能够泛化。嗯,大约十年后,约书亚·本吉奥(Yoshua Bengio)用了一个非常类似的网络,证明它在真实文本上也有效。再过大约十年,语言学家才开始相信嵌入这回事。这是个缓慢的过程。我之所以认为它不只是在预测下一个符号,是因为你可以问:“那要预测出下一个符号,需要具备什么?”
便签笔记
9:55
particularly if you ask me a question and then the first word of the answer is the next symbol. Um, you have to understand the question. So I think by predicting the next symbol, it's very unlike old fashioned autocomplete un old fashioned autocomplete, you'd store sort of triples of words and then if you saw a pair of words, you see how often different words came third. And that way you could predict the next symbol. And that's what most people think autocomplete is like. It's no longer a tool like that, um, to predict the next symbol. You have to understand what's being said. So I think you're forcing it to understand by making it predict the next symbol. And I think it's understanding in much the same way we are. So a lot of people will tell you these things aren't like us. Um, they're just predicting the next symbol. They're not reasoning like us, but actually in order to predict the next symbol, it's gonna have to do some reasoning. And we've seen now that if you make big ones without putting in any special stuff to do reasoning, they can already do some reasoning.
尤其是当你问我一个问题,而答案的第一个词就是下一个符号的时候。嗯,你必须理解这个问题。所以我认为,通过预测下一个符号,它和老式的自动补全非常不一样。老式的自动补全会存储一堆词的三元组,然后当你看到一对词时,就看看不同的词作为第三个词出现的频率有多高。这样你就能预测下一个符号。大多数人以为自动补全就是这样的。但现在它已经不是那样的工具了。要预测下一个符号,你必须理解正在说的内容。所以我认为,让它去预测下一个符号,就是在逼它去理解。而且我认为它理解的方式和我们非常相似。所以,很多人会告诉你这些东西和我们不一样。嗯,说它们只是在预测下一个符号,不像我们那样推理。但实际上,为了预测下一个符号,它就不得不做一些推理。而我们现在已经看到,如果你把模型做得很大,即使不专门加入任何做推理的东西,它们也已经能做一些推理了。
便签笔记
10:56
And I think as you make them bigger, they're gonna be able to do more and more reasoning. Do you think I'm doing anything else than predicting the next symbol right now? I think that's how you're learning. I think you're predicting the next video frame. Um, you're predicting the next sound. Um, but I think that's a pretty plausible theory of how the brain's learning. What enables these models to learn such a wide variety of fields. What these big language models are doing is they're looking for common structure, and by finding common structure they can encode things using the common structure and that's more efficient. So let me give you an example. If you ask GPT-4, "Why is a compost heap like an atom bomb?" Most people can't answer that. Most people haven't thought they think atom bomb and compost heaps are very different things. But GPT-4 will tell you, well, the energy scales are very different and the timescales are very different. But the thing that's the same is that when the compost heep gets hotter, it generates heat faster. And when the atom bomb produces more
而且我认为,随着你把它们做得越来越大,它们能做的推理也会越来越多。你觉得我现在做的事情,除了预测下一个符号之外,还有别的吗?我认为你就是这样学习的。我认为你在预测下一帧画面,嗯,在预测下一个声音。嗯,我觉得这是一个相当合理的关于大脑如何学习的理论。是什么让这些模型能学会如此广泛的各种领域?这些大型语言模型在做的,是寻找共同的结构,而通过找到共同结构,它们就能用这种共同结构来编码事物,这样效率更高。我举个例子。如果你问GPT-4:“堆肥堆为什么像原子弹?”大多数人回答不上来。大多数人从没想过,他们觉得原子弹和堆肥堆是完全不同的东西。但GPT-4会告诉你,它们的能量尺度非常不同,时间尺度也非常不同。但相同之处在于,当堆肥堆变得更热时,它产生热量的速度就更快;而当原子弹产生更多
便签笔记
11:56
neutrons, it produces more neutrons faster. And so it gets the idea of a chain reaction. And I believe it's understood they're both forms of chain reaction. It's using that understanding to compress all that information into its weights. And if it's doing that, then it's gonna be doing that for hundreds of things where we haven't seen the analogies yet, but it has and that's where you get creativity from, from seeing these analogies between apparently very different things. And so I think GPT-4 is gonna end up—when it gets bigger—being very creative.
中子时,它产生中子的速度也更快。于是它抓住了链式反应这个概念。我相信它理解了两者都是链式反应的形式。它正是在运用这种理解,把所有这些信息压缩进它的权重里。如果它在这么做,那它就会在成百上千个我们还没看出类比、但它已经看出来的地方这么做,而创造力正是从这里来的——从看出那些表面上非常不同的事物之间的类比。所以我认为 GPT-4 最终——等它变得更大的时候——会非常有创造力。
便签笔记
12:27
I think this idea that it's just regurgitating what it's learned, just pastiching together text, it's learned already that's completely wrong. It's gonna be even more creative than people I think. You'd argue that it won't just repeat the human knowledge we've developed so far but could also progress beyond that. I think that's something we haven't quite seen yet. We've started seeing some examples of it, but to a large extent we're sort of still at the current level of science. What do you think will enable it to go beyond that?
我觉得那种认为它只是在复述它学到的东西、只是把文本拼贴在一起的看法,只是把它已经学过的东西拼起来——这种说法完全错了。我认为它会比人还更有创造力。你会认为它不只是重复我们迄今为止发展出来的人类知识,还可能超越这些知识。我觉得这是我们还没怎么见到的。我们已经开始看到一些例子了,但很大程度上我们基本还停留在当前的科学水平上。你觉得是什么会让它超越这一点?
便签笔记
05类比、创造力与超越训练数据
13:00
Well, we've seen that in more limited context. Like if you take AlphaGo in that famous competition with Lee Sedol, um, there was move 37 where AlphaGo made a move that all the experts said must have been a mistake, but actually later they realized it was a brilliant move. Um, so that was created within that limited domain. Um, I think we'll see a lot more of that as these things get bigger. The difference with uh, AlphaGo as well was that it was using reinforcement learning that that subsequently sort of enabled it to, to go beyond the current state. So it started with imitation learning, watching how humans play the game and then it would through self play develop will be beyond that. Do you think that's the missing component of the current data?
嗯,我们在更受限的场景里已经见过了。比如 AlphaGo 在那场著名的与李世石的比赛中,第 37 手,AlphaGo 下了一步所有专家都说那肯定是个失误的棋,但后来他们意识到那其实是妙手。所以那是在一个有限领域里产生的创造。我觉得随着这些系统变得更大,我们会看到更多这样的情况。AlphaGo 的另一个不同之处在于,它用了强化学习,正是强化学习后来让它能够超越当时的最高水平。所以它一开始是模仿学习,看人类怎么下棋,然后通过自我对弈发展出超越人类的水平。你觉得这是当前这些模型缺失的那块拼图吗?
便签笔记
13:46
I think that may, I think that may well be a missing component, yes. That the self play in AlphaGo and AlphaZero are a large part of why it could make these creative moves. But I don't think it's entirely necessary. So there's a little experiment I did a long time ago where you, you're training a neural net to recognize handwritten digits. I love that example, the MNIST example. And you give it training data where half the answers are wrong. Um, and the question is how well will it learn? And you make half the answers wrong once and keep them like that. So it can't average away the wrongness by just seeing the same example. But with the right answer sometimes and the wrong answer, sometimes when it sees that example half of the examples, when it sees the example, the answer is always wrong.
我觉得这很可能确实是缺失的一环,是的。AlphaGo 和 AlphaZero 里的自我对弈,很大程度上解释了它为什么能下出这些有创造性的棋。但我不认为这是绝对必要的。我很久以前做过一个小实验,就是训练一个神经网络识别手写数字。我很喜欢这个例子,MNIST 那个例子。然后你给它的训练数据里有一半的答案是错的。问题是它能学得多好?而且你把一半答案设成错的之后就固定不变。这样它就没法通过多次看到同一个样本来把错误平均掉——不是有时给对的答案、有时给错的答案。对那一半样本来说,每次看到它,答案永远都是错的。
便签笔记
14:34
And so the training data has 50% error, but if you train up backpropagation, it gets down to 5% error or less. In other words, from badly labeled data, it can get much better results. It can see that the training data is wrong and that's how smart students can be smarter than their advisor. And their advisor tells 'em all this stuff and for half of what their advisor tells 'em, they think no rubbish, and they listen to the other half and then they end up smarter than the advisor. So these big neural nets can actually do, they can do much better than their training data and most people don't realize that.
所以训练数据有 50% 的错误率,但如果你用反向传播来训练,它的错误率能降到 5% 甚至更低。换句话说,从标注很糟糕的数据里,它能得到好得多的结果。它能看出训练数据是错的——聪明的学生之所以能比导师更聪明,就是这个道理。导师跟他们讲一大堆东西,其中一半他们心想「胡扯」,然后他们只听另一半,最后他们比导师还聪明。所以这些大型神经网络其实也能做到,它们能做得比自己的训练数据好得多,而大多数人没有意识到这一点。
便签笔记
06用推理生产训练数据来修正直觉
15:13
So how, how do you expect these models to add reasoning into them? So I mean one approach is you add sort of the heuristics on on top of them, which a lot of the research is doing now where you have sort of train of thought, you just feedback it's reasoning, um, in into itself. And another way would be in the model itself, uh, as you scale it up. Uh, what's your intuition around that? So my intuition is that as we scale up these models that get better at reasoning and if you ask how people work roughly speaking, we have these intuitions and we can do reasoning and we use the reasoning to correct our intuitions. Of course we use the intuitions during the reasoning to do the reasoning, but it's the conclusion of the reasoning conflicts with our intuitions. We realize the intuitions need to be changed. That's much like in AlphaGo or AlphaZero where you have an evaluation function, um, that just looks at a board and says, how good is that for me?
那你觉得这些模型会怎样把推理能力加进去?我是说,一种做法是在模型之上加一些启发式的东西,现在很多研究就是这么做的,比如思维链之类的,把它自己的推理再反馈回它自己。另一种做法是放在模型本身里,随着你把它做大。你对这个的直觉是什么?我的直觉是,随着我们把这些模型做大,它们的推理会变好。如果你问人是怎么运作的,粗略地说,我们有这些直觉,也能做推理,我们用推理来修正直觉。当然我们在推理的过程中也在用直觉,但如果推理的结论和我们的直觉相冲突,我们就意识到直觉需要改。这很像 AlphaGo 或 AlphaZero,你有一个评估函数,它看一眼棋盘就说,这个局面对我有多好。
便签笔记
16:11
But then you do the Monte Carlo rollout and now you get a more accurate idea and you can revise your evaluation function. So you can train it by getting it to agree with the results of reasoning. And I think these large language models have to start doing that. They have to start training their raw intuitions about what should come next by doing reasoning and realizing that's not right. And so that way they can get more training data than just mimicking what people did. And that's exactly why AlphaGo could do this creative move 37, it had much more training data 'cause it was using reasoning to check out what the right next move should have been.
但接着你做蒙特卡洛推演,现在你有了更准确的判断,于是你可以修正评估函数。所以你可以通过让它去符合推理的结果来训练它。我认为这些大型语言模型必须开始这么做。它们得开始训练自己关于「下一步该是什么」的原始直觉——靠推理,并且意识到那个直觉不对。这样它们就能获得比单纯模仿人类更多的训练数据。这也正是 AlphaGo 能下出第 37 手那种有创造性的棋的原因,它有多得多的训练数据,因为它在用推理去检验下一步正确的走法应该是什么。
便签笔记
07多模态与语言在认知中的位置
16:49
And what do you think about multimodality? So we spoke about these analogies and often the analogies are way beyond what we could see. It's discovering analogies that are far beyond the humans and at maybe abstraction levels that we will never be able to understand. Now when we introduce images to that and video and sound, how do you think that will change the models and uh, how do you think it'll change the analogies that it will be able to make? Um, I think it'll change it a lot. I think it'll make it much better at understanding spatial things. For example, from language alone, it's quite hard to understand some spatial things, although remarkably GPT-4 can do that even before it was multimodal. Um, but when you make it multimodal, if you have it both doing vision and reaching out and grabbing things, it'll understand object much better if you can pick them up and turn them over and so on. So although you can learn an awful lot from language, it's easier to learn if you are multimodal and in fact you then need less language and there's
那你怎么看多模态?我们刚才谈到这些类比,而很多类比远远超出我们能看到的范围。它发现的类比远远超过人类,而且可能处在我们永远无法理解的抽象层次上。那么当我们把图像、视频和声音也加进去,你觉得这会怎样改变这些模型?你觉得它会怎样改变它能做出的类比?我觉得会有很大改变。我认为它对空间性的东西的理解会好得多。比如说,光靠语言,要理解某些空间性的东西是相当难的,尽管很惊人的是,GPT-4 在还不是多模态的时候就已经能做到了。但当你把它做成多模态的,如果它既能看,又能伸手去抓东西,那它对物体的理解会好得多——如果你能把东西拿起来、翻过来看等等。所以尽管你能从语言里学到非常多的东西,如果你是多模态的,学起来会更容易;而且这样你需要的语言就更少了。而 YouTube 上有海量的视频可以用来预测下一帧,诸如此类。
便签笔记
17:56
an awful lot of YouTube video for predicting the next frame, so, or something like that. So I think these multimodal models are clearly gonna take over. Um, you can get more data that way. They need less language. So there's really a philosophical point that you could learn a very good model from language alone, but it's much easier to learn it from a multimodal system. And how do you think it'll impact the model's reasoning? I think it'll make it much better at reasoning about space. For example, reasoning about what happens if you pick objects up, if you actually try picking objects up, you're gonna get all sorts of training data that's gonna help.
所以我觉得这些多模态模型显然会成为主流。这样你能拿到更多数据,需要的语言更少。所以这里其实有一个哲学层面的观点:你可以只靠语言学到一个非常好的模型,但用多模态系统来学要容易得多。那你觉得这会怎样影响模型的推理能力?我觉得它在关于空间的推理上会好得多。比如说,推理「如果你把物体拿起来会发生什么」——如果你真的去试着拿起物体,你就会得到各种各样有帮助的训练数据。
便签笔记
18:29
Do you think the human brain evolved to work well with with language or do you think language evolved to work well with the human brain? I think the question of whether language evolved to work with the brain or the brain evolved to work with language, I think that's a very good question. I think both happened, I used to think we would do a lot of cognition without needing language at all. Um, now I've changed my mind a bit. So let me give you three different views of language, um, and how it relates to cognition. There's the old fashioned symbolic view, which is cognition consists of having strings of symbols in some kind of cleaned up logical language where there's no ambiguity and applying rules of inference. And that's what cognition is. It's just these symbolic manipulations on things that are like strings of language symbols. Um, so that's one extreme view.
你觉得是人脑进化得能很好地处理语言,还是语言进化得很适合人脑?语言是为了配合大脑而进化,还是大脑为了配合语言而进化,我觉得这是个非常好的问题。我认为两者都发生了。我以前认为我们很多认知过程根本不需要语言。现在我的想法有点变了。我给你讲三种关于语言的不同观点,以及它和认知的关系。有一种老派的符号主义观点,认为认知就是在某种清理干净的、没有歧义的逻辑语言里操作符号串,然后应用推理规则。认知就是这么回事,就是对这些类似语言符号串的东西做符号操作。这是一个极端的观点。
便签笔记
19:23
An opposite extreme view is no, no, once you get inside the head it's all vectors. So symbols come in, you convert those symbols into big vectors and all the stuff inside is done with big vectors. And then if you want to produce output, you produce symbols again. So there was a point in machine translation in about 2014 when people were using neural recurrent neural nets and words will keep coming in and they'd have a hidden state and they keep accumulating information in this hidden state. So when they got to the end of a sentence that have a big hidden vector that captured the meaning of that sentence, that could then be used for producing the sentence in another language—that was called a thought vector. And that's the sort of second view of language.
另一个极端的观点是:不不,一旦进到脑子里就全都是向量了。符号输入进来,你把这些符号转成大向量,里面所有的处理都是用大向量做的。然后如果你要产生输出,你再产生符号。所以大概在 2014 年前后,机器翻译有一段时期,人们用循环神经网络,词一个个进来,网络有一个隐状态,不断在这个隐状态里积累信息。等到了句子结尾,就得到一个大的隐向量,它捕捉了那个句子的含义,然后可以用它生成另一种语言的句子——这被称为「思想向量」。这是关于语言的第二种观点。
便签笔记
20:05
You convert the language into a big vector that's nothing like language and that's what cognition's all about. But then there's a third view, which what I believe now, which is that you take these symbols and you convert the symbols into embeddings and you use multiple layers of that. So you get these very rich embeddings, but the embeddings are still tied to the symbols in the sense that you've got a big vector for this symbol and a big vector for that symbol. And these vectors interact to produce the vector for the symbol for the next word. And that's what understanding is. Understanding is knowing how to convert the symbols into these vectors and knowing how the elements of the vector should interact to predict the vector for the next symbol. That's what understanding is, both in these big language models and in our brains. And that's an example which is sort of in between. You're staying with the symbols, but you are interpreting them as these big vectors. And that's where all the work is and all the knowledge
你把语言转换成一个完全不像语言的大向量,而认知就是这么回事。但还有第三种观点,也是我现在相信的,就是你把这些符号转换成嵌入向量,而且要用很多层来做这件事。这样你就得到非常丰富的嵌入表示,但这些嵌入仍然是和符号绑定的,意思是这个符号有一个大向量,那个符号有另一个大向量。而且这些向量相互作用,产生出下一个词符号的向量。这就是理解。理解就是知道如何把符号转换成这些向量以及知道向量的各个元素应该如何相互作用,来预测下一个符号的向量。这就是理解,无论是在这些大语言模型里,还是在我们的大脑里。这算是一个介于两者之间的例子。你仍然停留在符号层面,但你把它们解释成这些大向量。而所有的工作、所有的知识都在这里——
便签笔记
08GPU、模拟计算与「凡人计算」
21:04
is in what vectors you use and how the elements of those vectors interact not in symbolic rules. Um, but it's not saying that you get away from the symbols altogether. It's saying you turn the symbols into big vectors, but you stay with that surface structure of the symbols. And that's how these models are working. And that now seems to me a more plausible model of human thought too. You were one of the first folks to get the idea of using GPUs and uh, I know Jensen loves you, uh, for that. Uh, back in 2009 you mentioned that you told Jensen that this could be quite a good idea, um, for training neural nets. Take us back to that early intuition of using GPUs for training neural nets.
在于你使用什么样的向量,以及这些向量的元素如何相互作用,而不在于符号规则。嗯,但这并不是说你完全抛弃了符号。而是说你把符号变成大向量,但你仍然保留符号的那个表层结构。这就是这些模型的工作方式。而现在在我看来,这也是一个关于人类思维的更合理的模型。您是最早想到使用 GPU 这个点子的人之一,我知道黄仁勋很喜欢您,就因为这个。早在 2009 年,您就提到您告诉黄仁勋说,这对训练神经网络来说可能是个相当不错的主意。请带我们回到当年用 GPU 训练神经网络的那个最初的直觉。
便签笔记
21:48
So actually I think in about 2006 I had a former graduate student called Rick Zelisky, who's a very good computer vision guy. And I talked to him at a meeting and he said, you know, you ought to think about using graphics processing cards because they're very good at matrix multiplies and what you're doing is basically all matrix multiplies. So I thought about that for a bit. And then we learned about these Tesla systems that had, um, four GPUs in and initially we just got, um, gaming GPUs and discovered they made things go 30 times faster. And then we bought one of these Tesla systems with four GPUs and we did speech on that and it worked very well. And then in 2009 I gave a talk at NIPS and I told a thousand machine learning researchers, "You should all go and buy Nvidia GPUs. They're the future.
其实我想大概在 2006 年的时候,我有一位以前的研究生叫 Rick Szeliski,他是位非常出色的计算机视觉专家。我在一次会议上跟他聊天,他说,你知道吗,你应该考虑用图形处理卡,因为它们非常擅长矩阵乘法,而你做的事情基本上全都是矩阵乘法。所以我就琢磨了一阵子。后来我们了解到有那种 Tesla 系统,里面有四块 GPU。一开始我们只是弄了几块游戏用的 GPU,结果发现它们让运算快了 30 倍。然后我们就买了一套带四块 GPU 的 Tesla 系统,我们用它做语音识别,效果非常好。接着在 2009 年,我在 NIPS 上做了一个演讲,我对一千名机器学习研究者说:“你们都应该去买英伟达的 GPU。它们就是未来。
便签笔记
22:39
You need them for doing machine learning". And I actually, um, then sent mail to Nvidia saying, "I told a thousand machine learning researchers to buy your boards, could you give me a free one?" And they said, "No". Actually, they didn't say no, they just didn't reply. Um, but when I told Jensen this story later on, he gave me a free one. That's, uh, that's very, very good. I I think what's interesting is, um, as well is sort of how GPUs has evolved alongside, uh, the the field. So where, where do you think we we should go, uh, go next in the compute?
做机器学习你们需要它们。”然后我还真给英伟达发了封邮件,说:“我跟一千名机器学习研究者说了要买你们的板卡,你们能不能免费给我一块?”他们说:“不行。”其实他们也没说不行,他们干脆就没回复。嗯,不过后来我把这个故事讲给黄仁勋听时,他送了我一块。这个,这个真是太棒了。我觉得有意思的是,嗯,还有 GPU 是如何与这个领域一起演进的。那么,您觉得我们在算力方面接下来该往哪个方向走?
便签笔记
23:12
So my last couple of years at Google, I was thinking about ways of trying to make analog computation so that instead of using like a megawatt, we could use like 30 watts like the brain and we could run these big language models in analog hardware. And I never made it work and, but I started really appreciating digital computation. So if you're gonna use that low power analog computation, every piece of hardware is gonna be a bit different. And the idea is the learning is gonna make use of the specific properties of that hardware. And that's what happens with people. All our brains are different. Um, so we can't then take the weights in your brain and put them in my brain. The hardware's different, the precise properties of the individual neurons are different. The learning used to make—has learned to make use of all that.
在谷歌的最后几年,我一直在思考如何实现模拟计算,这样就不用耗费差不多一兆瓦的电,而是像大脑那样只用 30 瓦左右,我们就能在模拟硬件上运行这些大语言模型。我一直没能让它成功,但是,我因此开始真正体会到数字计算的好处。因为如果你要用那种低功耗的模拟计算,每一块硬件都会有点不一样。而思路是,学习过程会去利用那块硬件的特定属性。这正是人类身上发生的事。我们每个人的大脑都不一样。嗯,所以我们没法把你大脑里的权重拿出来放进我的大脑。硬件是不同的,单个神经元的精确属性是不同的。学习过程已经学会了去利用这一切。
便签笔记
24:03
And so we are mortal in the sense that the weights in my brain are no good for any other brain. When I die, those weights are useless. Um, we can get information from one to another rather inefficiently by I produce sentences and you figure out how to change your weights. So you would've said the same thing. That's called distillation. But that's a very inefficient way of communicating knowledge. And with digital systems, they're immortal because once you've got some weights, you can throw away the computer, just store the weights on a tape somewhere and now build another computer, put those same weights in and if it's digital it can compute exactly the same thing as the other system did. So digital systems can share weights and that's incredibly much more efficient if you've got a whole bunch of digital systems and they each go and do a tiny bit of learning and they start with the same weights, they do a tiny bit of learning, and then they share their weights again. Um, they all know what all the others learned,
所以从这个意义上说我们是“有死的”——我大脑里的权重对任何别的大脑都没用。等我死了,那些权重就废了。嗯,我们可以把信息从一个人传给另一个人,但方式相当低效:我说出一些句子,你去琢磨该怎么改变你的权重,好让你也能说出同样的话。这叫蒸馏。但这是一种非常低效的知识传递方式。而数字系统是“不死的”,因为一旦你有了一组权重,你可以把计算机扔掉,只要把权重存在某个磁带上,然后再造一台计算机,把同样的权重放进去,只要它是数字的,它就能算出和原来那个系统完全一样的东西。所以数字系统可以共享权重,而这要高效得多。如果你有一大批数字系统,它们各自去做一点点学习,一开始它们的权重都相同,各自学一点点,然后再共享权重。嗯,它们就都知道了其他所有系统学到的东西,
便签笔记
09多重时间尺度与被效率挡住的快权重
25:00
we can't do that. And so they're far superior to us in being able to share knowledge. A lot of the ideas that have been deployed in the field are very old school ideas. It's the ideas that have been around in neuroscience for forever. What do you think is sort of left to apply to the systems that we develop? So one big thing that we still have to catch up with neuroscience on is the timescales for changes. So in nearly all the neural nets, there's a fast timescale for changing activities. So input comes in the activities, the embedding vectors all change, and then there's a slow timescale which is changing the weights and that's long-term learning. And you just have those two timescales.
我们做不到这一点。所以在共享知识的能力上,它们远远优于我们。这个领域里用到的很多想法都是很老派的想法,是在神经科学里存在了很久很久的想法。您觉得还有哪些东西可以用到我们开发的这些系统上?有一个我们仍然需要向神经科学看齐的大问题,就是变化的时间尺度。在几乎所有的神经网络里,都有一个快速的时间尺度用来改变活动值。输入进来,活动值、嵌入向量全都改变;然后有一个慢速的时间尺度用来改变权重,那就是长期学习。你就只有这两个时间尺度。
便签笔记
25:47
In the brain, there's many timescales at which weights change. So for example, if I say an unexpected word like "cucumber" and now five minutes later you put headphones on, there's a lot of noise and there's very faint words, you'll be much better at recognizing the word "cucumber" because I said it five minutes ago. So where is that knowledge in the brain? And that knowledge is obviously in temporary changes to synapses. It's not neurons that going "cucumber, cucumber, cucumber". You don't have enough neurons for that.
而在大脑里,权重变化的时间尺度有很多种。举个例子,如果我说了一个意想不到的词,比如“黄瓜”,然后五分钟后你戴上耳机,有很多噪音,词声非常微弱,你识别出“黄瓜”这个词的能力会强得多,就因为我五分钟前说过它。那么这个知识存在大脑的什么地方呢?这个知识显然存在于突触的临时变化里。并不是有一群神经元一直在“黄瓜、黄瓜、黄瓜”地叫。你没有那么多神经元来干这个。
便签笔记
26:18
It's in temporary changes to the weights. And you can do a lot of things with temporary weight changes—fast, what I call fast weights. We don't do that in these neural models. And the reason we don't do it is because if you have temporary changes to the weights that depend on the input data, then you can't process a whole bunch of different cases at the same time. At present we take a whole bunch of different strings, we stack them, stack them together and we process them all in parallel because then we can do matrix, matrix multiplies, which is much more efficient. And just that efficiency is stopping us using fast weights. But the brain clearly uses fast weights for temporary memory and there's all sorts of things you can do that way that we don't do it present. I think that's one of the biggest things we have to learn. I was very hopeful that things like Graphcore, um, if they went sequential and did just online learning, then they could use fast weights. Um, but that hasn't worked out yet. I think it'll work out eventually when people are using conductances for weights.
它存在于权重的临时变化里。而用临时的权重变化——我称之为“快权重”——你可以做很多事情。我们在这些神经模型里没有这么做。我们不这么做的原因是,如果你的权重变化是临时的、依赖于输入数据的,那你就没法同时处理一大批不同的样例。目前我们是把一大堆不同的字符串拿来,把它们堆叠在一起,然后并行处理,因为这样我们就能做矩阵乘法,效率高得多。正是这种效率考量阻止了我们使用快权重。但大脑显然会用快权重来做临时记忆,而且用这种方式可以做各种我们现在做不到的事情。我认为这是我们需要学习的最重要的东西之一。我曾经非常期待像 Graphcore 这样的东西,嗯,如果它们走串行路线、只做在线学习,那它们就能用上快权重。嗯,但那还没有成功。我想等到人们用电导来表示权重的时候,最终是会成功的。
便签笔记
10先天结构、乔姆斯基与机器的感受
27:20
How has knowing how this models work and knowing how the brain works impacted the way you think? I think there's been one big impact, which is at a fairly abstract level, which is that for many years people were very scornful about the idea of having a big random neural net and just giving it a lot of training data and it would learn to do complicated things. If you talk to statisticians or linguists or most people in AI, they say that's just a pipe dream. There's no way you're gonna learn two really complicated things without some kind of innate knowledge without a lot of architectural restrictions. It turns out that's completely wrong. You can take a big random neural network and you can learn a whole bunch of stuff just from data. Um, so the idea that stochastic gradient descent to adjust the—repeatedly adjust the weights using a gradient that will learn things and will learn big complicated things that's being validated by these big models. And that's a very important thing to know about the brain.
了解这些模型的工作原理、也了解大脑的工作原理,这对您的思考方式产生了什么影响?我觉得有一个很大的影响,是在一个相当抽象的层面上:多年来人们非常瞧不起这样一种想法——搞一个大的随机神经网络,然后喂给它大量训练数据,它就能学会做复杂的事情。你去跟统计学家、语言学家或者 AI 领域的大多数人聊,他们都会说那纯粹是白日梦。他们说不可能在没有某种先天知识、没有大量架构限制的情况下学会真正复杂的东西。结果证明这完全错了。你可以拿一个很大的随机神经网络,仅仅从数据中就学到一大堆东西。嗯,所以,用随机梯度下降来调整权重——用梯度反复调整权重——就能学到东西,而且能学到又大又复杂的东西,这个想法正在被这些大模型所验证。而这是关于大脑的一件非常重要的认识。
便签笔记
28:25
It doesn't have to have all this innate structure. Now obviously it's got a lot of innate structure, but it certainly doesn't need innate structure for things that are easily learned. And so the sort of idea coming from Chomsky that you won't, you won't learn anything complicated like language unless it's all kind of wired in already and just matures. That idea is now clearly nonsense. I'm sure Chomsky would appreciate you calling his ideas, uh, nonsense [laughs]. Well I think actually I think a lot of Chomsky's political ideas are very sensible and I'm always struck by how come someone with such sensible ideas about the Middle East could be so wrong about linguistics?
大脑不必拥有所有这些先天结构。当然,很明显它确实有很多先天结构,但对于那些容易学会的东西,它肯定不需要先天结构。所以乔姆斯基那种想法——认为你不可能学会像语言这样复杂的东西,除非它已经以某种方式内置好了、只是逐渐成熟——那个想法现在显然是无稽之谈。我相信乔姆斯基听到您把他的想法称作“无稽之谈”会很高兴的(笑)。其实我觉得乔姆斯基的很多政治见解都非常明智,我一直很纳闷,一个对中东问题有着如此明智见解的人,怎么会在语言学上错得这么离谱?
便签笔记
29:04
What do you think would make these models simulate consciousness of humans more effectively? But imagine you had the AI assistant that you've spoken to in your entire life and instead of that being, you know, like ChatGPT today, that sort of deletes the memory of the conversation and you start fresh all of the time. It had self-reflection at some point you, you pass away and you tell that to, to the assistant. Do you think— I mean not me, somebody else tells that to the assistant. Yeah [laughs], it would be difficult for you to tell that to the assistant. Uh, do you think that that assistant would would feel at that point?
你觉得,要怎样才能让这些模型更有效地模拟人类的意识?不过想象一下,你有一个陪你聊了一辈子的 AI 助手,而且它不像今天的 ChatGPT 那样,会把对话的记忆删掉、每次都从头开始。它在某个时候有了自我反思,然后你去世了,你还把这件事告诉了那个助手。你觉得——我是说不是我本人,是别人把这件事告诉助手。对(笑),要你自己去告诉助手这件事确实挺难的。呃,你觉得那个助手在那一刻会有感受吗?
便签笔记
29:45
Yes. I think they can have feelings too. So I think just as we have this inner theater model for perception, we have an inner theater model for feelings. There are things that I can experience but other people can't. Um, I think that model is equally wrong. So I think suppose I say "I feel like punching Gary on the nose", which I often do. Let's try and abstract that away from the idea of an inner theater. What I'm really saying to you is, um, "If it weren't for the inhibition coming from my frontal lobes, I will perform an action". So when we talk about feelings, we are really talking about um, actions we will perform if it weren't for um, constraints. And that really, uh, that's really what feelings are the actions we would do if it weren't for constraints. Um, so I think you can give the same kind of explanation for feelings and there's no reason why these things can't have feelings.
会。我认为它们也能有感受。我觉得,就像我们对感知有一套「内在剧场」的模型一样,我们对感受也有一套「内在剧场」的模型:好像有些东西是我能体验、别人却体验不到的。嗯,我认为那个模型同样是错的。比如假设我说「我真想一拳打在 Gary 鼻子上」——我常有这种念头。我们试着把它从「内在剧场」这个概念里抽离出来。我真正想对你表达的其实是,嗯,「要不是我前额叶传来的抑制,我就会做出某个动作」。所以当我们谈论感受时,我们其实是在谈论:要不是有种种约束,我们会做出的行为。而这,呃,这其实就是感受——感受就是我们在没有约束时会去做的那些行为。嗯,所以我觉得对感受可以给出同样的解释,也没有理由说这些系统不能有感受。
便签笔记
30:40
In fact, in 1973 I saw a robot have an emotion. So in Edinburgh they had a robot with two grippers like this that could assemble a toy car if you put the pieces separately on a piece of green felt. Um, but if you put them in a pile, it's vision wasn't good enough to figure out what was going on. So it put its grippers together and went "whack", and it knocked them so they were scattered and then he could put them together. If you saw that in a person, you say it was cross with a situation 'cause it didn't understand it, so it destroyed it.
事实上,1973 年我就见过一个机器人有情绪。当时在爱丁堡,有台机器人有两个这样的夹爪,如果你把零件分开摆在一块绿色毛毡上,它能组装出一辆玩具小车。但如果你把零件堆成一堆,它的视觉就不够好,搞不清是什么情况。于是它把两个夹爪并到一起,「啪」地一下,把那堆零件打散,然后它就能把车装起来了。如果这是你在人身上看到的,你会说他是被这个情况惹恼了,因为他弄不明白,所以干脆把它砸散。
便签笔记
31:14
That's profound. You uh, when we spoke previously, you described sort of humans and the LLMs as analogy machines. What do you think has been the most powerful analogies that you've found throughout your life? Oh, in—throughout my life? Um, whew! I guess probably a sort of weak analogy that's influenced me a lot is um, the analogy between religious belief and between belief and symbol processing. So when I was very young I was confronted—I came from an atheist family—and went to school and was confronted with religious belief and it just seemed nonsense to me. It still seems nonsense to me. Um, and when I saw symbol processing as an explanation of how people worked, um, I thought it was just the same— nonsense. I don't think it's quite so much nonsense now because I think actually we do do symbol processing. It's just we do it by giving these big embedding vectors to the symbols. But we are actually symbol processing, um, but not at all in the way people thought where you match symbols and the only thing a symbol has is it's identical to another symbol or it's
这很深刻。你——我们上次聊的时候,你把人类和大语言模型都描述成「类比机器」。你觉得你这辈子发现过的最有力的类比是什么?哦,一生当中?嗯,哇。我想,对我影响很大的大概是一个不太严谨的类比:宗教信仰和符号处理之间的类比。我很小的时候就碰到过——我出生在一个无神论家庭——上学后接触到宗教信仰,我当时就觉得那纯属胡说八道。到现在我还是觉得它是胡说八道。嗯,后来我看到有人用符号处理来解释人是怎么运作的,我觉得这也一样——都是胡说八道。不过现在我不觉得它那么荒谬了,因为我认为我们确实是在做符号处理,只不过是通过给符号赋予很大的嵌入向量来做的。我们的确在做符号处理,但完全不是人们原先设想的那种方式——那种方式是把符号做匹配,而符号唯一的属性就是:它跟另一个符号要么相同、要么不同。符号就只有这么一个属性。我们根本不是这样做的。我们利用上下文
便签笔记
11如何选题、选人,以及直觉如何养成
32:30
not identical. That's the only property a symbol has. We don't do that at all. We use the context to give embedding vectors to symbols and then use the interactions between the components of these embedded vectors to do thinking. But there's a very good researcher at Google called Fernando Pereira who said, "Yes, we do have symbolic reasoning and the only symbolic we have is natural language". Natural language is a symbolic language and we reason with it. And I believe that now. You've done some of the most meaningful, uh, research in the history of of computer science. Can you walk us through like how do you select the right problems to, to work on?
给符号赋予嵌入向量,然后用这些嵌入向量各分量之间的相互作用来进行思考。不过谷歌有位很优秀的研究员叫 FernandoPereira,他说过:「是的,我们确实有符号推理,而我们拥有的唯一符号系统就是自然语言。」自然语言就是一种符号语言,我们用它来推理。我现在相信这一点。你做出了计算机科学史上一些最有意义的研究。能讲讲你是怎么挑选值得研究的问题的吗?
便签笔记
33:08
Well first let me correct you, me and my students have done a lot of the most meaningful things and it's mainly been a very good collaboration with students and my ability to select very good students. And that came from the fact there were very few people doing neural nets in the seventies and eighties and nineties and two thousands. And so the few people doing neural nets got to pick the very best students. So that was a piece of luck. But my way of selecting problems is basically, well you know, when scientists talk about how they work, they have theories about how they work, which probably don't have much to do with the truth, but my theory is that I look for something where everybody's agreed about something and it feels wrong, just there's a slight intuition that there's something wrong about it. And then I work on that and see if I can elaborate why it is I think it's wrong, and maybe I can make a little demo with a small computer program that shows that it doesn't work the way you might expect.
首先我得纠正你一下:是我和我的学生们做出了很多最有意义的成果,这主要得益于跟学生非常好的合作,以及我挑选优秀学生的能力。而这又是因为,在七十年代、八十年代、九十年代和两千年代,做神经网络的人非常少。所以少数几个做神经网络的人就能挑到最好的学生。这算是一种运气。至于我挑问题的方法,基本上是这样——你知道,科学家在谈自己怎么工作的时候,他们对自己的工作方式有一套理论,而这套理论多半跟事实没多大关系。但我的理论是:我会去找那种大家已经形成共识、但我觉得不对劲的东西,就是隐隐有种直觉,觉得哪里有问题。然后我就去研究它,看能不能把「我为什么觉得它不对」讲清楚,也许还能用一个小程序做个小演示,说明事情并不像你以为的那样。
便签笔记
34:06
So let me take one example. Um, most people think that if you add noise to a neural net, it's gonna work worse. Um, if for example, each time you put a training example through, you make half of the neurons be silent, it'll work worse. Actually we know it'll generalize better if you do that and you can demonstrate that. Um, in a simple example, that's what's nice about computer simulations. You can show, you know, this idea you had that adding noise is gonna make it worse and sort of dropping out half the neurons will make it work worse, which you will in the short term. But if you train it like that in the end it'll work better. You can demonstrate that with a small computer program and then you can think hard about why that is and how it stops big elaborate co-adaptations. Um, but that I think that's my method of working. Find something that sounds suspicious and work on it and see if you can give a simple demonstration of why it's wrong.
举个例子。嗯,大多数人认为,如果你给神经网络加噪声,效果会变差。比如说,每次送入一个训练样本时,你让一半的神经元静默,效果会变差。可实际上我们知道,这样做泛化反而更好,而且这是可以演示出来的。嗯,用一个简单的例子就行,这正是计算机模拟的好处。你可以证明:你原以为加噪声会让效果变差、随机丢掉一半神经元会让效果变差——短期内确实如此。但如果你就这样一直训练下去,最终效果反而更好。你可以用一个小程序把这一点演示出来,然后再认真去想这是为什么,它是如何阻止大规模复杂的协同适应的。嗯,我觉得这就是我的工作方法:找到一个听起来可疑的东西,去研究它,看能不能简单地演示出它为什么不对。
便签笔记
35:06
What sounds suspicious to you now? Well, that we don't use fast weights sound suspicious, that we only have these two timescales. That's just wrong. That's not at all like the brain. Um, and in the long run I think we're gonna have to have many more timescales. So that's an example there. And if you had, if you had your group of, of students today and they came to you and they said, so the Hamming question that we talked about previously, you know: "What's the most important problem in your field?" What would you suggest that they take on and work on next? We spoke about reasoning, timescales. What would be sort of the highest priority problem that that you'd give them?
那现在有什么让你觉得可疑的?嗯,我们不使用快权重(fast weights)这件事就很可疑,我们只有这两个时间尺度。这就是错的,跟大脑完全不一样。嗯,长远来看,我认为我们必须拥有多得多的时间尺度。这就算一个例子。那如果今天你带着一群学生,他们来问你我们之前聊过的那个「汉明问题」——也就是:「你所在领域最重要的问题是什么?」你会建议他们接下来去做什么、研究什么?我们聊过推理、时间尺度。你会交给他们的最高优先级问题是什么?
便签笔记
35:42
For me right now it's the same question I've had for the last like 30 years or so, which is "Does the brain do backpropagation?" I believe the brain is getting gradients. If you don't get gradients, your learning is just much worse than if you do get gradients. But how is the brain getting gradients and is it somehow implementing some approximate version of backpropagation or is it some completely different technique? That's a big open question and if I kept on doing research, that's what I would be doing research on.
对我来说,现在还是过去三十年左右我一直在想的那个问题,就是:「大脑做反向传播吗?」我相信大脑是在获取梯度的。如果拿不到梯度,学习效果会比拿到梯度差得多。但大脑究竟是怎么获取梯度的?它是以某种方式实现了反向传播的某个近似版本,还是用了一种完全不同的技术?这是个悬而未决的大问题,如果我继续做研究,这就是我会研究的方向。
便签笔记
36:12
And when you look back at, at your career now, you've been right about so many things, but what were you wrong about that you wish you sort of spent less time pursuing a certain direction? OK, those are two separate questions. One is "What were you wrong about?" and two, "Do you wish you'd spent less time on it?" I think I was wrong about Boltzmann machines and I'm glad I spent a long time on it. They're a much more beautiful theory of how you get gradients than backpropagation. Backpropagation is just ordinary and sensible and it's just a chamber. Boltzmann machines is clever and it's a very interesting way to get gradients and I would love for that to be how the brain works, but I think it isn't.
现在回头看你的职业生涯,你在很多事情上都判断对了,但有哪些是你判断错了、觉得自己不该在那个方向上花那么多时间的?好,这其实是两个不同的问题。一是「你在什么事情上错了」,二是「你是否希望自己少花点时间在上面」。我觉得我在玻尔兹曼机上判断错了,但我很庆幸在它上面花了很长时间。作为一套关于如何获取梯度的理论,它比反向传播漂亮得多。反向传播就是很普通、很合乎常理的东西,平平无奇。玻尔兹曼机很巧妙,是一种非常有意思的获取梯度的方式,我特别希望大脑就是这么工作的,但我觉得它不是。
便签笔记
12医疗前景、滥用风险与未解之问
36:51
Did you spend much time imagining what would happen post these systems developing as well? Did you ever have an idea that, OK, if we could make these systems work really well, we could, you know, democratize education, we could make knowledge way more accessible, um, we could solve some tough problems in medicine. Or was it more to you about understanding the brain? Yes, I, I sort of feel scientists ought to be doing things that are gonna help society, but actually that's not how you do your best research. You do your best research when it's driven by curiosity. You just have to understand something. Um, much more recently I've realized these things could do a lot of harm as well as a lot of good and I've become much more concerned about the effects they're gonna have on society. But that's not what was motivating me. I just wanted to understand how on earth can the brain learn to do things? That's what I want to know. And I sort of failed as a side effect of that failure. We got some nice engineering but...
那这些系统发展起来之后会发生什么,你花过很多时间去想象吗?你当时有没有想过:好,如果我们能让这些系统真的运转得很好,我们就能让教育普及化,让知识变得更容易获取,嗯,我们还能解决医学上一些棘手的问题。还是说对你而言,这更多是为了理解大脑?是这样,我多少觉得科学家应该去做对社会有益的事,但实际上,你最好的研究并不是这么做出来的。最好的研究是好奇心驱动的,你就是非得把某个东西搞明白不可。嗯,直到最近我才意识到这些东西既能带来很多好处,也能造成很大危害,我现在对它们将给社会带来的影响担忧得多了。但当年驱动我的并不是这个。我只是想搞明白,大脑究竟是怎么学会做事情的?这才是我想知道的。而某种意义上我失败了。作为这次失败的副产品,我们得到了一些不错的工程成果,但……
便签笔记
37:52
Yeah, it's, it was a good failure for the world. If you take the lens of the things that could go really right, what do you think are the most promising applications? I think healthcare is clearly a, a big one. Um, with healthcare there's almost no end to how much healthcare society can absorb. If you take someone old, they could use five doctors full time. Um, so when AI gets better than people are doing things, um, you'd like it to get better in areas where you could do with a lot more of that stuff. And we could do with a lot more doctors, if everybody had three doctors of their own, that would be great and we are gonna get to that point.
是啊,从整个世界的角度看,那是一次有益的失败。如果换个视角,看看那些可能会非常顺利的方面,你觉得最有前景的应用是什么?我觉得医疗显然是很重要的一块。嗯,在医疗领域,社会对医疗资源的吸纳几乎是无止境的。拿一位老人来说,他可以全职配五个医生。嗯,所以当 AI 在某些事情上做得比人还好时,你会希望它变强的领域,正好是那些我们需要更多供给的领域。而我们确实需要多得多的医生,如果每个人都有自己的三个医生,那就太好了,而且我们终将走到那一步。
便签笔记
38:37
Um, so that's one reason why healthcare's good. There's also just in new engineering, developing new materials for example, for better solar panels or for superconductivity or for just understanding how the body works. Um, there's gonna be huge impacts there. Those are all gonna be good things. What I worry about is bad actors using them for bad things. We've facilitated people like Putin, or Xi, or Trump using AI for killer robots or for manipulating public opinion, or for mass surveillance. And those are all very worrying things.
嗯,这就是医疗前景好的一个原因。另外还有新的工程领域,比如开发新材料,用于更好的太阳能电池板、超导,或者只是为了理解人体是怎么运作的。嗯,这些方面会有巨大的影响。这些都会是好事。我担心的是不良行为者把它们用在坏事上。我们等于是给普京、习近平或特朗普这样的人提供了便利,让他们能用 AI 去做杀人机器人、操纵公众舆论,或者搞大规模监控。这些都是非常令人担忧的事。
便签笔记
39:15
Are you ever concerned that slowing down the field could also slow down the positives? Oh absolutely, and I think there's not much chance that the field will slow down, partly because it's international and if one country slows down, the other countries aren't gonna slow down. So there's a race clearly between China and the US and neither is gonna slow down. So yeah, I don't— I mean there was this petition saying we should slow down for six months. I didn't sign it just 'cause I thought it was never gonna happen. I maybe should have signed it 'cause even though it was never gonna happen, it made a political point. It's often good to ask for things, you know, you can't get just to make a point. Um, but I didn't think we're gonna slow down.
你有没有担心过,给这个领域踩刹车同时也会拖慢那些正面的进展?当然会,而且我觉得这个领域慢下来的可能性不大,部分原因是它是国际性的,如果一个国家放慢脚步,别的国家不会跟着放慢。所以中美之间显然存在一场竞赛,谁都不会放慢。所以是的,我不——我是说,之前有一份请愿书说我们应该暂停六个月。我没有签,只是因为我觉得那根本不可能发生。也许我当时该签的,因为即便它不会成真,它也表明了一种政治立场。有时候去要求一些你明知得不到的东西,只为了表明态度,也是好的。嗯,但我并不觉得我们会慢下来。
便签笔记
39:56
And how do you think that it will impact the AI research process, uh, having uh, these assistants? I think it'll make it a lot more efficient. AI research will get a lot more efficient when you've got these assistants that help you program, um, but also help you think through things and probably help you a lot with equations too. Have you reflected much on the process of selecting talent? Has that been mostly intuitive to you? Like when Ilya shows up at the door, you feel this is a smart guy, let's work together.
那你觉得有了这些助手之后,它会怎样影响 AI 研究本身的过程?我觉得会让效率高很多。当你有这些助手帮你写程序时,AI 研究的效率会提升很多,嗯,而且它们还能帮你把问题想清楚,可能在方程推导上也能帮上不少忙。你对挑选人才这件事思考得多吗?这对你来说主要是靠直觉吗?比如 Ilya 出现在门口时,你就觉得这是个聪明人,一起干吧。
便签笔记
40:26
So, for selecting talent, um, sometimes you just know. So after talking to Ilya for not very long, he seemed very smart and then talking to a bit more, he clearly was very smart and had very good intuitions as well as being good at math. So that was a no-brainer. There's another case where I was at a NIPS conference. Um, we had a poster and I, someone came up and he started asking questions about the poster and every question he asked was a sort of deep insight into what we'd done wrong. Um, and after five minutes I offered him a postdoc position. That guy was David MacKay who was just brilliant and it's very sad he died, but he was, it was very obvious you'd want him. Um, other times it's not so obvious and one thing I did learn was that people are different.
关于挑选人才,嗯,有时候你一眼就知道。跟 Ilya 没聊多久,他就显得非常聪明,再多聊一会儿,就很明显他非常聪明,直觉也很好,数学还很强。所以那根本不用犹豫。还有一次是我在NIPS 会议上。嗯,我们有一张海报,有个人过来开始问关于海报的问题,而他问的每一个问题,都像是深刻地看穿了我们哪里做错了。嗯,五分钟之后我就给了他一个博士后职位。那个人是 David MacKay,他非常出色,他去世了很令人难过,但当时很明显你会想要他。嗯,另一些时候就没那么明显了,我学到的一点是,人是各不相同的。
便签笔记
41:15
There's not just one type of good student. Um, so there's some students who aren't that creative but are technically extremely strong and will make anything work. There's other students who aren't technically strong but are very creative. Of course you want the ones who are both, but you don't always get that. But I think actually in the lab you need a variety of different kinds of graduate student. But I still go with my gut intuition that sometimes you talk to somebody and they're just very, very, they just get it and those are the ones you want.
好学生不止一种类型。嗯,有些学生并不那么有创造力,但技术上极其扎实,什么都能给你做出来。也有些学生技术上不强,但非常有创造力。当然你希望两者兼备,但你不总能碰到。不过我觉得实验室里其实需要各种不同类型的研究生。但我还是靠我的直觉,有时候你跟某个人聊,他们就是非常非常——他们就是懂,那就是你想要的人。
便签笔记
41:49
What do you think is the reason for some folks having better intuition? Do they just have better training data than than others or how can you develop your intuition? I think it's partly they don't stand for nonsense. So here's a way to get bad intuitions: believe everything you're told. That's fatal. You have to be able to—I think here's what some people do. They have a whole framework for understanding reality and when someone tells 'em something, they try and sort of figure out how that fits into their framework and if it doesn't, they just reject it. And that's a very good strategy. Um, people who try and incorporate whatever they're told end up with a framework that's sort of very fuzzy and sort of can believe everything and that's useless. So I think actually having a strong view of the world and trying to manipulate incoming facts to fit in with your view, obviously it can lead you into deep religious belief and fatal flaws and so on. Like my belief in Boltzmann machines. Um, but
你觉得为什么有些人直觉更好?是因为他们的训练数据比别人好吗,还是说直觉是可以培养的?我觉得部分原因是他们不接受胡说八道。所以,这里有个把直觉搞坏的办法:别人说什么你都信。那是致命的。你得能够——我觉得有些人是这么做的:他们有一整套理解现实的框架,当别人告诉他们某件事时,他们会去琢磨这件事怎么放进自己的框架里,如果放不进去,他们就直接拒绝掉。这是个非常好的策略。嗯,那些试图把听到的一切都吸收进来的人,最后得到的框架非常模糊,什么都能信,那就没用了。所以我觉得,对世界有一个坚定的看法,并试着把接收到的事实拿捏成符合自己看法的样子,当然,这也可能让你陷入很深的宗教信念和致命的错误之类,比如我对玻尔兹曼机的信念。嗯,但我觉得这条路是对的。如果你有可以信赖的好直觉,你就该信它们。
便签笔记
42:55
I think that's the way to go. If you've got good intuitions you can trust, you should trust them. If you've got bad intuitions, it doesn't matter what you do, so you might as well trust them. That's a very, very good, uh, very good point. When, when you look at the, the types of research that's, that's that's being done today, do you think we're putting all of our eggs in one basket and we should diversify our ideas a bit more in, in the field? Or do you think this is the most promising direction? So let's go all in on it?
如果你的直觉很糟,那你做什么都无所谓,所以不如就信它们。这是个非常非常好的,呃,非常好的观点。当你看到,当今在做的那些类型的研究时,你觉得我们是不是把所有鸡蛋都放在一个篮子里了,这个领域应该让思路更多元一些?还是说你觉得这就是最有前景的方向,那就全力押上?
便签笔记
43:28
I think having big models and training them on multimodal data, even if it's only to predict the next word, is such a promising approach that we should go pretty much all in it. Obviously there's lots and lots of people doing it now and there's lots of people doing apparently crazy things and that's good. Um, but I think it's fine for like most of the people to be following this path 'cause it's working very well. Do you think that the learning algorithms matter that much or is it just a skill? Are there basically millions of ways that we could, we could get to human level in, in intelligence or are there sort of a select few that we need to discover?
我觉得做大模型,并用多模态数据训练它们,哪怕只是为了预测下一个词,也是一条非常有前景的路线,我们应该几乎全力投进去。显然现在有非常非常多的人在做这件事,也有很多人在做看起来很疯狂的事,这是好事。嗯,但我觉得大多数人沿着这条路走是没问题的,因为它效果非常好。你觉得学习算法本身有那么重要吗,还是说它只是一种手段?是不是基本上有成千上万种方式可以让我们达到人类水平的智能,还是说只有少数几种是我们必须发现的?
便签笔记
44:06
Yeah, so this issue of whether particular learning algorithms are very important or whether there's a great variety of learning algorithms that'll do the job. I don't know the answer. It seems to me though that backpropagation, there's a sense in which it's the correct thing to do. Getting the gradient so that you change a parameter to make it work better. That seems like the right thing to do and it's been amazingly successful. There may well be other learning algorithms that are alternative ways of getting that same gradient or that are getting the gradient to something else and that also work. Um, I think that's all open and a very interesting issue now about whether there's other things you can try and maximize that will give you good systems and maybe the brain's doing that 'cause it's easier, but backprop is in a sense, the right thing to do and we know that doing it works really well.
是的,关于特定的学习算法是不是非常重要,还是说有大量各式各样的学习算法都能做成这件事——我不知道答案。不过在我看来,反向传播在某种意义上就是该做的事。求出梯度,从而调整参数让它表现更好。这看起来就是对的做法,而且它已经取得了惊人的成功。很可能还有别的学习算法,用另外的方式得到同样的梯度,或者得到别的东西的梯度,同样也管用。嗯,我觉得这些都还是开放的,而且现在有个很有意思的问题:是不是还有别的东西你可以试着去最大化,从而得到好的系统,也许大脑就在做这件事,因为那样更容易,但反向传播在某种意义上就是对的做法,而且我们知道这么做效果非常好。
便签笔记
45:00
And one last question. When, when you look back at your sort of decades of research, what are you, what are you most proud of? Is it the students? Is it the research? What, what makes you most proud of when you look back at, at your life's work? The learning algorithm for Boltzmann machines. So the learning algorithm Boltzmann machines is beautifully elegant. It's maybe hopeless in practice. Um, but it's the thing I enjoyed most developing that with Terry and it's what I'm proudest of. Um, even if it's wrong.
最后一个问题。当你回顾自己几十年的研究时,你最引以为豪的是什么?是学生?是研究成果?什么让你在回顾一生的工作时最感到自豪?玻尔兹曼机的学习算法。玻尔兹曼机的学习算法优雅得漂亮。它在实践中也许没什么希望。嗯,但那是我最享受的东西,和 Terry 一起把它做出来,那是我最自豪的。嗯,哪怕它是错的。
便签笔记
45:37
What questions do you spend most of your time thinking about now? Is it the— Um, "What should I watch on Netflix?"
你现在把大部分时间花在思考哪些问题上?是不是——嗯,“我该在 Netflix 上看点什么?”
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Hinton 回顾自己从研究大脑到深度学习的学术生涯,核心论点是:大语言模型的"预测下一个词"实质上就是理解与推理,规模而非新想法是主要驱动力,数字系统凭借权重共享在知识传递上远超人类,而神经网络仍缺少大脑的"快权重"多时间尺度机制。

核心要点

  • "预测下一个词"就是理解,而非旧式自动补全:旧式 autocomplete 只是统计词三元组的共现频率;而要预测一个问题答案的第一个词,模型必须先理解问题。Hinton 自称做出了首个用 embedding + 反向传播的神经语言模型(在三元组数据上),约 10 年后 Bengio 在真实文本上复现,再过 10 年语言学家才接受 embedding。他认为强迫模型预测下一个符号就是在强迫它理解,且大模型在未加任何专门推理模块的情况下已能做一些推理。
  • 规模是主因,Ilya 早就说对了:Ilya 一直主张"做大就会更好",Hinton 当时认为这是偷懒,事后承认 Ilya 基本正确——Transformer 等新想法有帮助,但真正起作用的是数据和算力的规模。当年他们预期计算机快 100 倍,实际快了约 10 亿倍;很多"聪明技巧"在大规模下本会自行消失。2011 年他与 Ilya、James Martens 用字符级预测 Wikipedia HTML 已经看到惊人效果。
  • 理解 = 发现共同结构并压缩,这也是创造力来源:例子:问 GPT-4"堆肥堆为何像原子弹",它能指出二者都是链式反应(越热产热越快 / 中子越多产中子越快)。模型靠这种类比把知识压缩进权重,会在人类尚未察觉的成百上千处发现类比,因此 Hinton 认为"只是拼贴复述"的说法完全错误,模型最终会比人更有创造力。
  • 神经网络可以超越有错误的训练数据:MNIST 实验中把 50% 标签固定改错(不可通过重复采样平均掉),反向传播训练后错误率仍降到 5% 以下。类比于聪明学生能识别导师一半的话是错的,最后比导师更强。他认为 AlphaGo 式自博弈可能是缺失的一环,但并非绝对必要。
  • 推理应反过来训练直觉,如同 AlphaGo 用 MCTS 修正评估函数:人用推理纠正直觉;AlphaGo 用蒙特卡洛 rollout 得到更准的评估再回训评估函数,这让它获得远多于模仿人类的训练数据,从而下出"第 37 手"。LLM 也应用推理结果校正其对下一个词的原始直觉。
  • 语言与认知的第三种观点:既非纯符号逻辑操作,也非 2014 年机器翻译的"思想向量"(把句子压成与语言无关的大向量);而是符号被转为多层丰富 embedding,但仍与符号绑定,向量元素间的交互预测下一个符号的向量——"理解"就是这个过程,人脑亦然。他现在认同 Fernando Pereira 的说法:人类唯一的符号推理就是自然语言。Chomsky 式"语言必须天生布线"的观点已被大随机网络仅靠数据学会复杂事物的事实证伪。
  • 数字系统"不朽",人类"必死":模拟/低功耗计算(大脑约 30 瓦 vs 兆瓦级)需要让学习适配每块硬件的独特物理属性,因此权重无法迁移——人脑的知识只能靠低效的"蒸馏"(说话)传递。数字系统权重可存储、复制、精确复现;多个副本各学一点再合并权重,就能共享所有人学到的东西,这是人类做不到的巨大优势。
  • 最大的缺失:快权重 / 多时间尺度:现有网络只有两种时间尺度(激活快变、权重慢变),而大脑有多种:五分钟前听到"cucumber"后,噪声中更易识别该词,这存于突触的临时变化而非神经元持续放电。不用快权重的原因是它依赖输入数据,无法把多条序列堆叠做矩阵乘法并行处理——效率考量阻碍了这一机制。他曾寄望 Graphcore 做顺序在线学习,未成功。
  • 感受与意识:他反对"内心剧场"模型,主张"感受"是"若无前额叶抑制就会执行的动作",因此 AI 同样可以有感受。1973 年在爱丁堡他见过机器人因视觉无法解析零件堆而"一拍打散"——若是人,会说它"生气了"。
  • 选题方法与直觉培养:找"大家都同意但感觉不对"的东西,用小程序演示(如 dropout:训练时随机关掉一半神经元短期变差、长期泛化更好)。好直觉来自"不容忍胡说"——有坚固的世界框架并拒绝不相容的信息,哪怕这可能导致像他对玻尔兹曼机的执念那样的错误;"好直觉就信它,坏直觉怎么做都无所谓,也不如信它"。

结论与值得注意的细节

  • Hinton 认为多模态(视觉 + 操作物体)会大幅提升空间推理,且能减少对语言数据的依赖;YouTube 视频预测下一帧提供海量数据。他主张该领域"几乎全押"在大模型 + 多模态 + 下一词预测上是合理的,同时留有少数人做"看似疯狂"的事。
  • 反向传播"在某种意义上是正确的做法"(求梯度改参数),但大脑是否做反向传播是他 30 年来的核心问题,若继续做研究他仍会研究这个。
  • 他承认自己在玻尔兹曼机上错了,但不后悔——那是比反向传播"更美"的梯度理论,也是他最自豪的工作,"即使它是错的"。
  • 关于风险:健康医疗(每人三个医生)、新材料、超导是最大正向应用;担忧的是 Putin、习、Trump 之流用于杀手机器人、舆论操纵、大规模监控。他未签"暂停六个月"请愿,因为中美竞赛下不可能真暂停,但事后觉得应签以表明立场。
  • 趣闻:Ilya 周日敲门要"现在就谈",一周后说"我懂链式法则,只是不懂为什么不把梯度交给一个合理的优化器";一上午写完 MATLAB 接口。"hidden layer"的命名来自他与 Peter Brown 借用 HMM 的术语。2009 年 NIPS 他劝一千名研究者买 NVIDIA GPU,写信索要免费显卡未获回复,后来黄仁勋亲自补送。David MacKay 在海报前提问五分钟即被当场给 postdoc。
  • 最后被问现在最常想什么问题,他答:"Netflix 上该看什么?"
核心句型 · 10
1. It seemed to me (that) …
“It seemed to me there has to be a way that the brain learns”
表达「在我看来当时就觉得」,比 I thought 更委婉、更强调主观印象。适合回顾早年判断,常与 from the outset、at the time 搭配。
2. It turns out (that) …
“It turns out Ilya was basically right.”
用于「结果证明」,引出与先前预期相反的事实。写作中常与前一句的 I always thought… 配对,形成「我以为—结果」的转折结构。
3. there's a sense in which …
“Backpropagation, there's a sense in which it's the correct thing to do”
学术口语里的限定式肯定:在某种意义上成立,但不绝对。适合替代生硬的 it is definitely,用来给一个有保留的强判断。
4. not just any …, but …
“There was a knock on the door, not just any knock, but it went kind of … an urgent knock”
先给普通名词,再否定其平常性并补充特征,制造叙事悬念。仿写:not just any email, but one marked urgent。
5. I used to think …; now I've changed my mind
“I used to think we would do a lot of cognition without needing language at all. now I've changed my mind a bit.”
表述观点修正的标准句式,used to 强调已废弃的旧看法。加 a bit 可弱化幅度,显得克制而非全盘推翻。
6. … in the sense that …
“We are mortal in the sense that the weights in my brain are no good for any other brain”
用一个可能被误解的词,紧接着限定它的确切含义。写论证性文字时极实用,能预先堵住读者的误读。
7. If it weren't for …, I would …
“If it weren't for the inhibition coming from my frontal lobes, I will perform an action”
虚拟条件句,表示「要不是有……,我就会……」。注意标准写法主句用 would;此处口语中说成 will,学习者宜按 would 仿写。
8. we could do with a lot more …
“We could do with a lot more doctors”
英式口语,意为「我们很需要更多的……」。语气比 we need 更松弛,适合谈供需缺口。
9. You might as well …
“If you've got bad intuitions, it doesn't matter what you do, so you might as well trust them.”
表示「反正结果一样,不如就……」,常跟在一个「怎么做都无所谓」的前提之后,用于收束一段推理,语气略带自嘲。
10. That's the kind of … you want.
“I think I learned more from him than he learned from me. That's the kind of student you want.”
先描述一个具体情形,再用这句下定义式的评价,是英语口语中提炼标准的常见手法。仿写:That's the kind of feedback you want。
词汇精讲 · 130 · 按出现顺序
swarming /ˈswɔːrmɪŋ/ adj. / v. 0:25
挤满人的;蜂拥的。it was swarming 意为「里面挤满了人」
refreshing /rɪˈfreʃɪŋ/ adj. 0:25
令人耳目一新的、提神的(此处指风气上的振奋)
change the course of phr. 0:25
改变……的走向/进程
physiology /ˌfɪziˈɑːlədʒi/ n. 1:11
生理学
action potentials /ˈækʃən pəˈtenʃəlz/ n. 1:11
动作电位,神经元轴突上的电脉冲信号
simulate /ˈsɪmjuleɪt/ v. 1:11
模拟、仿真
intrigued /ɪnˈtriːɡd/ v. / adj. 1:49
引起(某人)浓厚兴趣;被吸引的
conviction /kənˈvɪkʃən/ n. 1:49
坚定的信念、确信
rules of inference phr. 2:29
推理规则(逻辑学术语)
from the outset phr. 2:29
从一开始
nonlinear /ˌnɑːnˈlɪniər/ adj. 3:01
非线性的
weight /weɪt/ v. / n. 3:01
(对输入)加权;权重。此处 they weight them 为动词用法
statistician /ˌstætɪˈstɪʃən/ n. 3:38
统计学家
mature student phr. 3:38
大龄学生(工作数年后再回校读学位者)
what they're up to phr. 4:26
他们在搞什么名堂(口语,含不明就里之意)
chain rule /tʃeɪn ruːl/ n. 5:05
链式法则,微积分中复合函数求导法则
gradient /ˈɡreɪdiənt/ n. 5:05
梯度,函数上升最快的方向向量
optimizer /ˈɑːptɪmaɪzər/ n. 5:05
优化器(求函数极值的算法)
cooking fries phr. 5:05
(在快餐店)炸薯条,指打零工
raw intuitions phr. 6:07
未经训练的原始直觉
greed /ɡriːd/ n. 6:07
贪婪
got fed up with phr. 7:09
受够了、厌烦了
get diverted phr. 7:09
被岔开注意力、跑偏
get on with phr. 7:09
继续推进(某事)
preaching /ˈpriːtʃɪŋ/ v. 7:49
鼓吹、反复宣讲(原义为布道)
cop-out /ˈkɑːp aʊt/ n. 7:49
逃避、推托之辞
remarkably /rɪˈmɑːrkəbli/ adv. 7:49
显著地、出奇地
embeddings /ɪmˈbedɪŋz/ n. 8:50
嵌入向量,把符号映射为稠密实数向量的表示
generalize /ˈdʒenrəlaɪz/ v. 8:50
泛化,在未见过的数据上仍表现良好
triples /ˈtrɪpəlz/ n. 8:50
三元组(如「主语-关系-宾语」)
autocomplete /ˌɔːtoʊkəmˈpliːt/ n. 9:55
自动补全
plausible /ˈplɔːzəbəl/ adj. 10:56
看似合理的、说得通的
compost heap /ˈkɑːmpoʊst hiːp/ n. 10:56
堆肥堆
encode /ɪnˈkoʊd/ v. 10:56
编码,把信息以某种形式表示
neutrons /ˈnuːtrɑːnz/ n. 11:56
中子
chain reaction /tʃeɪn riˈækʃən/ n. 11:56
链式反应
analogies /əˈnælədʒiz/ n. 11:56
类比、相似之处
apparently /əˈpærəntli/ adv. 11:56
表面上看(此处非「显然」,而是「看起来如此」)
regurgitating /rɪˈɡɜːrdʒɪteɪtɪŋ/ v. 12:27
照搬复述(原义为反刍、呕出)
pastiching /pæˈstiːʃɪŋ/ v. 12:27
拼贴模仿(源自艺术术语 pastiche)
to a large extent phr. 12:27
在很大程度上
reinforcement learning /ˌriːɪnˈfɔːrsmənt/ n. 13:00
强化学习,通过奖励信号试错学习
imitation learning /ˌɪmɪˈteɪʃən/ n. 13:00
模仿学习,从人类示范数据中学习
self play n. 13:00
自我对弈,让系统与自身副本博弈以产生训练数据
average away phr. 13:46
通过取平均把(噪声)消掉
rubbish /ˈrʌbɪʃ/ n. 14:34
胡说八道(英式口语,也指垃圾)
advisor /ədˈvaɪzər/ n. 14:34
(研究生的)导师
heuristics /hjʊˈrɪstɪks/ n. 15:13
启发式方法,无保证但通常有效的经验法则
evaluation function n. 15:13
评估函数,给当前局面打分的函数
rollout /ˈroʊlaʊt/ n. 16:11
(蒙特卡洛)推演,模拟走到终局以估计价值
revise /rɪˈvaɪz/ v. 16:11
修正、修订
mimicking /ˈmɪmɪkɪŋ/ v. 16:11
模仿
multimodality /ˌmʌltimoʊˈdæləti/ n. 16:49
多模态(同时处理文本、图像、声音等)
spatial /ˈspeɪʃəl/ adj. 16:49
空间的、与空间关系有关的
an awful lot phr. 16:49
非常多(口语强调用法)
take over phr. v. 17:56
取而代之、成为主流
cognition /kɑːɡˈnɪʃən/ n. 18:29
认知(思维、记忆、推理等心智过程)
ambiguity /ˌæmbɪˈɡjuːəti/ n. 18:29
歧义、多义性
symbolic manipulations n. 18:29
符号操作(符号主义AI的核心主张)
recurrent /rɪˈkɜːrənt/ adj. 19:23
循环的;recurrent neural net 即循环神经网络
hidden state n. 19:23
隐状态,网络内部逐步累积信息的向量
accumulating /əˈkjuːmjəleɪtɪŋ/ v. 19:23
累积、积聚
tied to phr. 20:05
与……绑定、依附于
surface structure n. 21:04
表层结构(此处指符号序列本身的形式)
get away from phr. v. 21:04
摆脱、抛开
matrix multiplies /ˈmeɪtrɪks/ n. 21:48
矩阵乘法(口语中把 multiplications 简称为 multiplies)
boards /bɔːrdz/ n. 22:39
板卡、电路板(此处指显卡)
analog computation /ˈænəlɔːɡ/ n. 23:12
模拟计算,用连续物理量而非离散比特运算
megawatt /ˈmeɡəwɑːt/ n. 23:12
兆瓦(百万瓦)
appreciating /əˈpriːʃieɪtɪŋ/ v. 23:12
体会到……的价值(非「感激」)
mortal /ˈmɔːrtəl/ adj. 24:03
终有一死的、会消亡的
immortal /ɪˈmɔːrtəl/ adj. 24:03
不死的、可无限延续的
distillation /ˌdɪstəˈleɪʃən/ n. 24:03
蒸馏;AI中指让小模型模仿大模型输出以传递知识
far superior to phr. 25:00
远远优于
timescales /ˈtaɪmskeɪlz/ n. 25:00
时间尺度
faint /feɪnt/ adj. 25:47
微弱的、不明显的
synapses /ˈsɪnæpsiz/ n. 25:47
突触,神经元之间的连接部位
fast weights n. 26:18
快权重,随输入短时改变的权重(Hinton提出的术语)
stack /stæk/ v. 26:18
堆叠(把多个样本拼成一批并行处理)
conductances /kənˈdʌktənsɪz/ n. 26:18
电导(电阻的倒数),此处指用器件电导表示权重
worked out phr. v. 26:18
成功、行得通
scornful /ˈskɔːrnfəl/ adj. 27:20
轻蔑的、不屑一顾的(be scornful about)
pipe dream /paɪp driːm/ n. 27:20
白日梦、不切实际的空想
innate /ɪˈneɪt/ adj. 27:20
先天的、与生俱来的
stochastic gradient descent /stəˈkæstɪk/ n. 27:20
随机梯度下降,深度学习的基本优化算法
validated /ˈvæləˌdeɪtɪd/ v. 27:20
证实、验证(其有效性)
matures /məˈtʊrz/ v. 28:25
成熟、发育完全
nonsense /ˈnɑːnsens/ n. 28:25
无稽之谈、荒谬之说
sensible /ˈsensəbəl/ adj. 28:25
明智的、合乎情理的(非「敏感」)
struck by phr. 28:25
对……感到诧异/印象深刻
self-reflection /ˌself rɪˈflekʃən/ n. 29:04
自我反思
pass away phr. v. 29:04
去世(委婉说法)
inner theater /ˈɪnər ˈθiːətər/ n. 29:45
内在剧场,指「只有自己能观看的私密体验舞台」这一心智模型
inhibition /ˌɪnhɪˈbɪʃən/ n. 29:45
抑制(神经科学与心理学术语)
frontal lobes /ˈfrʌntəl loʊbz/ n. 29:45
额叶,与计划和行为抑制相关的脑区
abstract away from phr. 29:45
从……中抽象出来、剥离掉
constraints /kənˈstreɪnts/ n. 29:45
约束、限制条件
grippers /ˈɡrɪpərz/ n. 30:40
(机器人的)夹爪、抓手
felt /felt/ n. 30:40
毛毡(此处为名词,非 feel 的过去式)
cross /krɔːs/ adj. 30:40
恼火的、生气的(英式口语,be cross with)
profound /prəˈfaʊnd/ adj. 31:14
深刻的、意味深长的
atheist /ˈeɪθiɪst/ n. / adj. 31:14
无神论者(的)
confronted with phr. 31:14
面对、遭遇(某种situation或观念)
elaborate /ɪˈlæbəreɪt/ v. 33:08
详细阐述(作形容词时读 /ɪˈlæbərət/,意为精细的)
co-adaptations /ˌkoʊædæpˈteɪʃənz/ n. 34:06
协同适应,指神经元过度依赖彼此形成的脆弱组合
suspicious /səˈspɪʃəs/ adj. 34:06
可疑的、引人怀疑的
dropping out phr. v. 34:06
随机丢弃(神经元),即 dropout 技术
approximate /əˈprɑːksɪmət/ adj. 35:42
近似的
open question phr. 35:42
悬而未决的问题
driven by curiosity phr. 36:51
由好奇心驱动
on earth phr. 36:51
究竟(用于疑问句加强语气)
lens /lenz/ n. 37:52
视角、看待问题的框架(take the lens of)
absorb /əbˈzɔːrb/ v. 37:52
吸纳、消化(此处指社会能容纳多少医疗资源)
superconductivity /ˌsuːpərkɑːndʌkˈtɪvəti/ n. 38:37
超导(电阻为零的物理现象)
bad actors n. 38:37
不良行为者、恶意方(安全领域常用语)
facilitated /fəˈsɪləteɪtɪd/ v. 38:37
使……更容易、为……提供便利
mass surveillance /səˈveɪləns/ n. 38:37
大规模监控
petition /pəˈtɪʃən/ n. 39:15
请愿书、联署信
no-brainer /noʊ ˈbreɪnər/ n. 40:26
无需思考的决定、明摆着的事
postdoc /ˈpoʊstdɑːk/ n. 40:26
博士后(职位或人员)
gut intuition /ɡʌt/ phr. 41:15
直觉、本能判断(gut feeling 的同类表达)
stand for nonsense phr. 41:49
容忍胡说八道(stand for 此处意为「容忍」)
fatal /ˈfeɪtəl/ adj. 41:49
致命的、后果严重的
incorporate /ɪnˈkɔːrpəreɪt/ v. 41:49
吸纳、纳入(体系中)
fuzzy /ˈfʌzi/ adj. 41:49
模糊的、界限不清的
put all of our eggs in one basket phr. 42:55
把所有鸡蛋放在一个篮子里,孤注一掷
diversify /daɪˈvɜːrsəfaɪ/ v. 42:55
使多元化、分散(投入)
go all in phr. 43:28
全力押注(源自扑克术语)
maximize /ˈmæksəmaɪz/ v. 44:06
最大化(某个目标函数)
elegant /ˈeləɡənt/ adj. 45:00
(理论、方案)优雅简洁的
理解自测 · 11 题
1. Hinton 说神经网络里「隐藏层」这个名字是怎么来的?

来自隐马尔可夫模型。彼得·布朗(Peter Brown)作为大龄博士生到卡内基梅隆,教了Hinton语音识别和隐马尔可夫模型,那时Hinton正在做带中间层的反向传播,这些层还没有名字。Hinton觉得HMM里用「hidden」来指代那些你不知道它们在干什么的变量非常贴切,于是他和布朗一起决定用这个词命名神经网络的中间层。这段出现在早期合作的回忆部分,他还补了一句「我从他那里学到的比他从我这里学到的多」。

2. Ilya 第一次读完反向传播论文后提出的疑问是什么?为什么 Hinton 觉得它重要?

Ilya说自己看懂了链式法则,但不明白为什么不把梯度交给一个像样的函数优化器。Hinton原本以为他没读懂而感到失望,听到真实问题后才意识到这是超前的提问——这个问题Hinton自己也花了好几年才想明白。它出现在Ilya敲门那一段,是Hinton用来说明「原始直觉好」的第一个具体证据,也预示了后来关于优化算法的长期研究。

3. 「一半标签是错的」那个 MNIST 实验,具体是怎么设计的,结论是什么?

训练一个识别手写数字的网络,把一半样本的标签改成错的,并且固定不变——同一个样本每次出现答案都是那个错的,而不是时对时错。这个设计排除了「模型靠多次看到同一样本把噪声平均掉」的解释。结果是训练数据错误率50%,用反向传播训练后测试错误率降到5%甚至更低。Hinton用它类比聪明学生:对导师说的一半内容判断为胡扯,只听另一半,最后超过导师。

4. Hinton 提出的三种语言与认知的观点分别是什么?他现在持哪一种?

第一种是老派符号主义:认知就是在无歧义的逻辑语言上操作符号串并应用推理规则。第二种是「思想向量」:符号进入大脑后全部转成大向量,处理完再输出符号,代表是2014年前后用循环网络做机器翻译时的整句隐向量。第三种是他现在相信的折中立场:保留符号的表层结构,但把每个符号转成多层嵌入向量,让向量之间的交互预测下一个符号的向量——理解就等于这套转换与交互,人脑和大模型都是如此。

5. 为什么 Hinton 认为「它只是在预测下一个词」不构成对大模型的贬低?

因为预测任务的难度决定了它必须内化理解。他先区分了老式自动补全:那种做法是存词的三元共现频率,看到两个词就统计第三个词出现的概率。但如果你问一个问题,答案的第一个词就是要预测的下一个符号,模型必须先理解问题才能预测对。所以让它预测下一个符号,等于在逼它理解。他进一步说,人也是这样学习的——预测下一帧画面、下一个声音,因此这不是模型独有的低级机制。

6. 「堆肥堆与原子弹」的例子在他的论证链条中起什么作用?

它是从「压缩」推向「创造力」的枢纽。GPT-4能指出两者能量尺度和时间尺度都不同,但共同点是自加速的链式反应。Hinton的推论是:模型之所以要找共同结构,是因为用共同结构编码信息更省权重,也就是压缩需求本身在推动它发现类比。既然它在这一个例子上做到了,它就会在几百上千个人类尚未看出类比的地方做到。而创造力恰恰被定义为看出表面上很不同的事物之间的类比,所以他断言模型变大后会比人更有创造力。

7. Hinton 为什么认为数字系统在知识积累上结构性优于人类?

因为知识复制的带宽差了好几个数量级。人是「有死的」:学习会利用个体大脑硬件的具体特性,我的权重放进你的大脑没有意义,我死了权重就作废;人际传递只能靠我说句子、你调整自己的权重去逼近,他把这称为蒸馏,效率极低。数字系统是「不死的」:权重可以存起来、换台机器装回去、算出完全相同的结果。因此一大批数字系统可以各学一点再共享权重,每个都立刻获得其他所有系统学到的东西。这段紧接在他放弃模拟计算研究之后,是那次失败反过来带来的认识。

8. 为什么「快权重」至今没有被主流模型采用?这个理由属于什么性质?

理由是工程效率而非科学原理。快权重意味着权重会随输入数据发生临时改变,而当前训练把大量不同样本堆叠成一批并行处理、做矩阵乘法,这要求同一批样本共享同一套权重。两者直接冲突,所以是效率约束挡住了快权重。Hinton明确说大脑显然在用快权重做临时记忆(「黄瓜」的启动效应就存在于突触的临时变化里),并把它列为最该向神经科学补的一课。他也提到曾寄望于Graphcore这类走串行、在线学习路线的芯片,但尚未成功。

9. Hinton 描述的选题方法是什么?它与他自己的失败案例有什么关系?

方法是:找一个大家已达成共识、但你隐隐觉得不对劲的地方,把「为什么觉得不对」阐述清楚,最好用一个小程序做出反例演示。他举的例子是Dropout——大家以为加噪声、随机让一半神经元静默会变差,短期确实如此,但长期训练泛化更好,因为它阻止了神经元之间的复杂协同适应。有意思的是这套方法也有代价:他对玻尔兹曼机的强直觉最终被判定为错的,而他在谈直觉养成时正是拿这件事当作「强框架可能把你带进死胡同」的自证。

10. 如果有人反驳说:Hinton 把「感受」定义为「若无抑制便会做出的动作」,这只是回避了主观体验问题。按片中逻辑他会怎样回应?

他会回应说被回避的那个问题本身建立在一个错误模型上。他明确说,我们对感知有一套「内在剧场」模型——好像存在一个只有自己能看见的私密舞台——而对感受也套用了同一个模型,这个模型两处都错。所以「主观体验」不是一个待解释的额外事实,而是一种需要被拆掉的说法。他给的替换是行为倾向:说「我想一拳打在Gary鼻子上」,真正说的是「要不是前额叶的抑制,我就会做出那个动作」。他还用1973年爱丁堡那台把零件堆打散的机器人作为功能等价的例证。批评者仍可反驳这只是取消了问题而非解决它,但这正是片中他的立场。

11. 「模型能从50%错标数据中学到5%错误率」这一结论,放到当下用AI生成数据训练AI的场景还成立吗?

部分成立,但条件不同。Hinton实验里的错误是随机加在样本上的、与正确模式无关的噪声,模型能靠数据中占优的真实规律把它压过去,这也是他说「学生能超过导师」的机制。而AI生成数据的错误往往是系统性的、与模型自身偏差同向的——错在同一个方向上就不再是可被平均或压制的噪声,反而会被强化。片中真正对应这一场景的其实是另一条论证:AlphaGo靠自我对弈超越人类,前提是围棋有胜负这一外部真值可以校验推理结果。Hinton也正是说语言模型要用推理去检验并修正自己的直觉,才能获得比模仿人类更多的训练数据。所以关键变量是有没有独立的校验信号,而不是数据是谁生成的。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.083Thinking, Fast and Slow | Daniel Kahneman | Talks at Google 下一期 · NO.085 →The Little Book that Beats the Market | Joel Greenblatt | Talks at Google
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com