视频库 / NO.102ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.102
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Josh Tenenbaum | How to grow a mind from a brain: From guessing and betting to thinking and talking

节目发布 2024-03-06 · Harvard CMSA
乔什·特南鲍姆 DDan Freed
本期追问 · 点击跳到视频对应位置
能流畅对话的大语言模型,真的在思考吗?大脑的根本功能是学习,还是靠世界模型下注?人类只听一亿词就学会语言,靠什么先天知识?语言是思维的载体,还是思维放大器?
归入 Ⅱ·04 答得上来,就算懂了吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文为约书亚·特南鲍姆(Josh Tenenbaum)在哈佛大学数学科学与应用中心(CMSA)叶氏讲座(Yip Lectures)第四讲上的演讲实录,由中心主任丹·弗里德(Dan Freed)主持。特南鲍姆是麻省理工学院计算认知科学教授,贝叶斯认知科学与概率程序学习的代表人物,2019年麦克阿瑟奖得主。演讲从ChatGPT出发,追问语言与思维的关系,提出「大脑是猜测与下注的机器」这一核心命题,并展示了「脑中的游戏引擎」与概率程序如何刻画直觉物理、直觉心理,以及语言如何在此之上接入。本文依据现场录音编译整理,仅删去口语枝节。

开场:讲者介绍与研究定位

弗里德: 我是丹·弗里德,哈佛大学数学科学与应用中心主任。很高兴欢迎各位来到叶氏讲座(Yip Lectures)的第四讲,这个系列涵盖数学、物理、生理学与心理学。感谢叶氏家族捐资设立这一讲座系列。下面请HT介绍今天的讲者。

主持人: 很高兴特南鲍姆教授答应来做这场讲座。他是麻省理工学院的计算认知科学教授,1993年在耶鲁大学获物理学学士学位,1999年在麻省理工学院获博士学位。他现在领导麻省理工学院的计算认知科学实验室,同时参与该校的「智能探索」(Quest for Intelligence)项目。

特南鲍姆教授以其在数理心理学和贝叶斯认知科学方面的奠基性贡献而闻名,获得过众多奖项与荣誉,我只提其中几项。第一,他的贝叶斯程序学习研究入选《科学美国人》2016年「十大改变世界的想法」。第二,2018年《研发》杂志评选他为年度创新者。第三,他于2019年获得麦克阿瑟奖,表彰他开创性地发展并运用概率与统计建模来研究人类的学习、推理与知觉。让我们欢迎特南鲍姆教授。

特南鲍姆: 谢谢。房间里的各位和楼上的各位都能听到吗?好。能来到这里,我深感荣幸,也感谢各位的到来。今天我要讲的是:如何从一个大脑中生长出一个心智。

各位从刚才的介绍大概已经能猜到,我的研究横跨好几个领域。我首先把自己看作认知科学家,但也做神经科学和人工智能。我们的目标,说到底,是理解智能如何在人类的心智与大脑中产生。但我们采用的做法,有时称为「逆向工程」(reverse engineering):我们要的是模型和理论,但这些理论要以计算模型的形式表达,并且带有工程的形态。它们有数学,但同时也是工程风格的模型。至少在我们这门科学里,「理解」就是这个意思。这样做还可能带来一种回报,也许会改变世界,也许不会,我们拭目以待:那就是造出以更接近人类的方式实现智能的人工智能系统,让它们能在人类建造的这个世界里,以更有价值、更高效、更安全、总体上更正面的方式存在。

我要讲的是我们做的一些工作,开头先要感谢一批合作者。名单里加粗的,是工作在今天讲座中占有重要位置的人;特别加粗的几位学生,梁力昂(Lionel Wong,又称Leo)、加布里埃尔·格兰德(Gabriel Grand)、亚历克斯·卢(Alex Lew)和泰勒·布鲁克-威尔逊(Tyler Brooke-Wilson),对这场讲座的材料和构想,甚至许多幻灯片,都有重大贡献。我要特别向他们致意。

对于在这个领域工作的人,甚至只是与之相邻的人,眼下是一个格外令人兴奋的时刻,一个令人激动的时刻,有时也是一个令人害怕的时刻。人工智能无处不在。这个领域本身很老了,但就在过去一两年间,它成了每个人谈论的头等话题。

在科学这一侧,随着神经科学和认知科学的进展,我们开始能够研究多种不同大脑中的智能。我们着迷于这样的问题:它们的共同点是什么?大脑究竟有什么一般性的东西定义了智能?当然,我们也看到物种之间的差异,看到人类智能显然有其独特之处。人类是特殊的。我们今天能坐在这个房间里,是有原因的。作为神经科学家和认知科学家,我们研究许多其他动物:小鼠、果蝇、斑马鱼、鸟类、猕猴、黑猩猩。但人类坐在这个房间里,而小鼠和果蝇在田野里,或者躲在角落里,黑猩猩则在某处令人痛心地走向灭绝,这是有原因的。

所以我想用计算的语言来理解,人类智能到底独特在哪里、特殊在哪里。而在人工智能这一侧,面对如此迅猛的进展,我们许多人都在问:它要往哪里去?会不会有一天,人工智能不只是和我们聊天,而是也在追问这些同样的问题,也许和我们一起,也许自己去问?如果是它自己去问,那我们算什么?我们是小鼠,还是黑猩猩?我们心里装着很多事。

规模路线:从神经元到ChatGPT

特南鲍姆: 关于如何从更简单、更基本的大脑走到人类智能,有一种思路,我先从它讲起。这不是我要主张的那一种,但在某种意义上,它是公共讨论的主要话题。而在认知科学内部,大家相当公认还有另一幅图景,那才是我这场讲座的重点。不过先说前一种。

一个想法是,看看能从单个神经元里抽象出什么。这张图来自休伯尔(Hubel)和维泽尔(Wiesel)研究初级视皮层细胞的经典论文。它的思路是,存在某种基本的计算单元:输入与权重做内积,再经过一个非线性变换,然后有一个学习单元,即某种学习规则,通过调整神经元的权重来最小化某个代价函数或目标函数。这是一个非常基本的想法,也非常强大。把它扩展到非常大、非常深的神经网络,比如著名的AlexNet深度卷积网络,再在李飞飞等人建立的ImageNet那样的大数据集上训练,就能做出惊人的事情。十多年前,当足够深、足够大的网络在ImageNet这样足够大的数据集上训练时,计算机视觉经历了深度学习革命。接着,把这类架构稍加变化,比如变成Transformer,再用互联网上能找到的全部文本来训练,实际上就是人类有史以来公开发表的几乎全部语言,你就得到了这样的东西。这个符号我甚至不必解释它代表什么,因为它无处不在,大概就在你们许多人的手机和笔记本电脑里。但我还是说出来:ChatGPT。正是它让人工智能天天登上《纽约时报》和我们所有人的新闻推送。它显然正在从许多方面改变世界。

这种改变之大,以至于许多人在问:这会不会就是那条路,或者至少是一条路,通向真正的通用人工智能?可以肯定的是,ChatGPT和类似的大语言模型呈现给公众的样子,正是人工智能一直以来「应该」呈现的样子:一种能做计算机能做的几乎一切事情的计算机,但你是用说话的方式使用它,你请它替你做事,而不必为它编程。每一部科幻电影都是这样的,不管是C-3PO,还是《星际迷航》里的计算机,或者别的什么例子。我就不提哈尔(HAL)了,好吧,我刚才已经提了。所以我们现在有了以它本该有的方式出现的人工智能。而且这些系统显然是通用的,因为它们能做许多不同的事。

于是很多人在问:在人工智能这一侧,如果我们真要拥有类人的通用人工智能,思维会不会就从这里产生?还有许多人在猜测:也许生物学里发生的正是这回事?只要规模够大,大脑越来越大,思维就出现了。演化中确实发生了这件事:我们的脑是最大的,黑猩猩基本上排第二。大脑袋是有用的。但这只是规模的问题吗?规模就够了吗?

流畅却脆弱:叠塔难题的失败

特南鲍姆: 我不打算就ChatGPT做整场讲座,只举一个例子来说明一种现象。只要稍加留意这些系统,我们都熟悉它:它们能做出一些了不起的事,看上去像思维,某种意义上也确实是某种思维;然而它们又惊人地脆弱。

这个例子来自我在麻省理工学院的前同事、现在斯坦福的埃里克·布林约尔松(Erik Brynjolfsson),他发在推特上,说这件事让他觉得有点毛骨悚然。他向ChatGPT提了一个直觉物理式的问题,而各位稍后会看到,直觉物理正是我在认知科学里喜欢研究的领域之一。他问:怎样用一本书、四个网球、一枚钉子、一个酒杯、一团嚼过的口香糖和五根生意面,搭出一个稳定的堆叠?很疯狂的题目,各位可以自己先想想。ChatGPT的回答是一套系统的操作流程。第一步,把四个网球在平面上摆成正方形,每个球是一个角;然后把书放在四个网球上,确保书的每个角都压在一个网球上,这样书就成了一个平台,形成稳定的底座。它接着往下说了很长。这太厉害了,不是吗?如果你逐条去看,它看上去确实拥有,而且某种意义上我也承认它确实拥有,对物理世界、对物体、对特定摆放下会发生什么以及能拿它们做什么的某种理解。

我们可以再往前走一步。这个例子来自我提到的学生之一泰勒·布鲁克-威尔逊,他刚从麻省理工学院哲学系毕业,现在纽约大学做博士后。泰勒请ChatGPT自己出一系列类似的叠塔挑战题。可以看到它出了「办公用品摩天楼」「园艺塔」「艺术家的攀登」之类的题目,然后再请ChatGPT自己解答。它很有创意,能把这些想法以各种有趣的方式重新组合,从那一个例子和它训练过的其他材料中就学到了这种题目的「体裁」。

但如果你真去看解答,就开始出现令人困惑的东西,能看出它的智能在哪些方面是脆弱的。随便挑一道放大看,它们都差不多。就看「艺术家的攀登」。题目是这样的:用一支画笔、一块不超过8乘10英寸的小画布、一管丙烯颜料、一块调色板和一杯水,搭一座至少12英寸高的塔,结构必须自己站立至少45秒。怎么解?答案是:打开颜料管,用一些颜料把画笔粘到画布上,做成底座;把调色板竖起来靠在画笔上;把水杯平衡在调色板上,必要时用剩下的颜料做粘合剂。

听起来还挺像回事。英语非常流畅,也确实利用了这些物体的一些真实的物理属性和可供性(affordance)。但如果你把它当作一个能真正站住的稳定堆叠来想,这完全是胡说八道。还要说明的是,泰勒在提示语里,也参照了我做过的一些事,明确写道:解不出来没关系,可以表示不确定,或者说你解不了;与其给出一个行不通的方案,不如承认不确定或者承认解不了。但这并没有阻止ChatGPT极富创意地提出一大堆根本行不通的方案。

那么这里到底发生了什么?这些现象,还有许多人写过的其他类似现象,应该促使我们认真追问:语言和思维究竟是什么关系?许多人,包括史蒂文·平克(Steven Pinker),包括今天在座的许多人,多年来尤其是在哈佛,就这个问题写过非常有洞见的东西。从计算的角度看,如果我们设想一种支持语言的认知架构,语言和思维在这个架构里各自扮演什么角色、相互是什么关系?这实际上就是我们这个研究项目要做的事。

「ChatGPT或者大语言模型涌现出了某种思维能力」这个说法里,暗含着这样一幅图景。某种意义上,它是个老想法,可以追溯到信息论的创始人、第一批统计语言模型的建立者香农(Shannon):预测语言的任务,也就是预测下一个词或下一个字符,或者接下来的一串字符或词,在某种意义上是「智能完备」的。无论是人工智能还是认知,思维实际上都可以被操作化为「前文与后文之间存在的那种潜在的预测结构」。确实有一些任务,比如这里列出的这些,你要补全模式,就必须动用各种常识和关于世界的知识。但语言的表层形式里也有相当明显的规律。所以,光凭句子的前半部分或前面的句子,在语境中预测任何一句话的后续,似乎有可能成为部署甚至训练一个通用智能架构的途径。

然而,一旦落实到任何一个具体的计算系统,比如任何一个具体的Transformer,它是一个有限的架构,在一个有限的数据集上训练,哪怕这个数据集非常大;而且大语言模型或者任何预测序列模型的工作方式,是自回归预测,也就是根据前面所有元素预测下一个元素、词元或词,然后不断向前重复。这样的系统没有任何特别的理由能解决任意问题、执行任意一种思维,哪怕它的潜在结构里储存了大量有意思的知识和理解。这里有几个问题,答案是明确的,也能算出来,但一个自回归语言模型不大可能算出来。人工智能界确实在尝试扩展这些模型,比如所谓「思维链」(chain of thought)或「思维树」(tree of thought):训练模型一步一步地想,也就是不直接生成答案,不只是把空填上,而是生成中间的计算步骤,这是有帮助的。这些都是移动的靶子,但目前来看,它对上面那道题有帮助,对下面那道题没有。这是大语言模型社区为了让语言模型变成更通用的思维系统而正在做的事情之一。ChatGPT和同类模型里还有另一些非常重要的东西,比如所谓「基于人类反馈的强化学习」(RLHF),把模型调校得不仅能预测语言中的下一个词元,还能预测人们倾向于点赞而不是点踩的那类回答。

这条以规模通向智能的路线,表达的愿景仍然是把语言放在中心,说「语言里的全部模式就足够了」;从架构上讲,就是说语言是思维的基底。正如平克和其他许多人写过的,这里面有真实的问题。我不展开,只推荐各位去读平克的著作,比如《思想本质》(The Stuff of Thought),那是一本很好的经典入门书。

脑科学证据:语言与思维分离

特南鲍姆: 我再从认知科学和神经科学给几个视角。比如我在麻省理工学院的同事伊芙琳娜·费多连科(Ev Fedorenko)的工作,下面几张幻灯片是从她那里借来的。她和其他许多人用功能磁共振成像及其他神经影像手段,加上人类神经电记录、神经心理学和行为学方法,研究人脑中的语言网络。这个系统人们已经研究了一百多年,用过各种方法。它是一组脑区构成的网络,特异地参与语言的使用,既包括说,就像我现在做的,也包括理解,我希望各位正在做的;还包括读和写。用什么形式都无所谓:盲人可以学盲文,语言就是语言。

你从小说哪种语言也无所谓,哪怕你学的是人造语言,比如克林贡语或《权力的游戏》里的多斯拉克语。只要你是流利的使用者,费多连科和同事们发现,你的语言网络会以基本相同的方式激活。至于大语言模型,也就是这些预测型Transformer,我和马丁·施林普夫(Martin Schrimpf)、费多连科等同事做过一点工作,我在其中角色很小:按今天的标准算是小型语言模型的GPT-2,能为人脑语言区在语言加工任务中的反应提供一个相当不错、至少是最先进的定量预测模型。

这很有意思。我并不是要说语言模型没有告诉我们关于语言的任何东西。费多连科和同事凯尔·马霍瓦尔德(Kyle Mahowald)、安娜·伊万诺娃(Anna Ivanova),还有我们一群人,合写了一篇即将发表在《认知科学趋势》(Trends in Cognitive Sciences)上的文章,试图阐述当代计算与认知神经科学对语言与思维关系的看法。我也推荐各位去读,里面有很多好的视角。这个团队参与了最早一批把大语言模型(今天我们可能会叫它们小语言模型)应用于脑科学的研究。这些模型所用的训练数据,其实更接近一个人可能得到的训练数据。它们表现不差,至少是我们目前对人脑反应最好的定量预测模型,而且远好于更简单的语言模型。

但请记住,没有人认为GPT-2是通用智能模型。它开始捕捉到大量句法,以及一些句法与语义的接口,很有意思,但肯定没有捕捉到通用智能,也没有捕捉到全部语义。而费多连科和同事们还证明了另一件事:语言网络,也就是能被中小型神经语言模型建模的那一部分,真的只管语言。当人们做其他任务时,包括某些方面与语言相似的任务,比如解代数方程、理解代码、做其他类型的逻辑推理,只要你知道怎样仔细地看,就会发现语言网络并不参与。同样,脑损伤可以选择性地只损害语言,而不损害其他形式的、比如说符号性的思维。结论就是:在大脑里,语言和思维是两回事。

发展与演化:思维先于语言

特南鲍姆: 但我认为更有说服力的视角来自发展的角度:儿童是怎样成长出智能的?从图灵开始,人工智能研究者,当然也包括我自己的人工智能研究,一直深受「儿童如何成长出智能」这个问题的启发。想一想就会明白:正如人类智能是独一无二的,人类作为学习机器也是独一无二的。人类儿童是已知宇宙中唯一一种能可靠地从少得多的起点成长为完整人类智能的学习机器。所以我们来看看他们如何学会思考、如何学会语言。

再引用一位杰出的哈佛同事、心理学系的伊丽莎白·斯佩尔克(Elizabeth Spelke),这是她最近出版的书。她和婴儿认知领域的许多人已经详细阐明,人类婴儿在学会语言之前很久,就已经拥有各种思维系统:思考物理对象的系统,思考其他主体及其目标和计划、他们如何与物理世界和彼此互动的系统,等等。所以从发展上看,思维显然在先。而且我们这个领域的理解,或者说信念(它介于信念和一组有充分支持的事实之间,仍有争论),是这一点至关重要:人类在学语言之前就拥有的知识,正是我们能够用比大语言模型少得多的数据学会语言的关键。最新的模型可能用一万亿个词训练,而一个人类儿童,据估计每年大约听到五百万个词,头二十年加起来大概一亿个词。按今天人工智能的标准,这是一个非常非常小的语料库。

发展领域最有意思的,也许是苏珊·戈尔丁-梅多(Susan Goldin-Meadow)的工作。每次讲到她的工作,我都觉得深受鼓舞。她和同事、学生长期研究「家庭手语」(home sign),以及相关现象,比如尼加拉瓜手语。家庭手语指的是这样一种现象:先天失聪的孩子如果得不到手语输入,也就是基本得不到任何语言输入,他们仍然会可靠地构建出自己的交流系统,类似于原始语言。它们还不是语言,但已经是表达思想、并且确实能与周围人交流的系统,具备语言的许多特征。这说明人类心智里有某种东西,是语言的创造者,而不是反过来。不是靠大量语言训练出来的;语言来自心智,而且哪怕是单个的人,在得不到语言输入的情况下,也能开始创造自己的语言。

尼加拉瓜手语是一个著名案例:一所孤儿院把一批家庭手语者聚在一起,这些孩子成长中基本没有任何标准自然语言的输入,各有各的家庭手语。但他们摸索出了彼此交流的办法,形成了一些人所说的皮钦语,然后是克里奥尔语,基本上是在短短几代人之间从零创造出一门全新的语言。所以当你思考语言与思维的关系时,这在我看来是非常了不起的事:一个没有语言输入的人会开始寻找符号化的方式来表达自己的思想;把几个这样的人放在一起,几乎转眼之间,用极少的数据和经验(相比造就了整个互联网的人类文化而言),他们就创造出一门全新的语言,而文化随即围绕它演化起来。这是在告诉我们一些东西。

也许最引人注目的,还不只是发展的角度,而是演化的角度。果蝇或斑马鱼这类生物只有十万个神经元,蜜蜂大约一百万,小鼠几千万;传统上被认为聪明的动物,比如大猿或乌鸦,它们会使用工具。它们都没有语言,却都以某种形式思考。没有人会否认,黑猩猩或乌鸦解决工具使用和物理问题时,当然是在思考。我对斑马鱼有一份特别的喜爱,因为我以很小的合作者身份参与过哈佛弗洛里安·恩格特(Florian Engert)实验室的一项工作,由安德鲁·博尔顿(Andrew Bolton)主持。恩格特等人长期研究斑马鱼,研究能从它的大脑里得到什么;在博尔顿的这个项目里,研究的是行为。他做了三维心理物理学实验,看斑马鱼如何捕食。各位如果不了解,幼体斑马鱼非常小,要用显微镜研究;而它们捕食的东西更小,是单细胞生物草履虫。

草履虫的运动是半随机的,但不完全随机,斑马鱼试图捕捉它们。博尔顿和同事们证明,斑马鱼对草履虫有一种预测模型。某种意义上,它能猜测草履虫要去哪里,并对怎样到达那里下注。你之所以能看出它在猜测,是因为它的动作是一组离散的动作:它会转动身体,有点像各位玩过的电子游戏的手柄;它有小小的鳍可以拍打,有尾巴可以摆动,基本上就是一连串多少有些离散的动作,一阵一阵地前进和转向。它会可靠地跟踪和追捕一只草履虫,但关键在于,它不是游向草履虫现在的位置,而是游向一个简单但近乎最优的统计预测器所预测的、草履虫将要到达的位置。如果草履虫继续按原来的方向运动,斑马鱼就抓住它;如果它稍微偏离,斑马鱼可能就放弃,去别处。这就是博尔顿等人研究的基本心理物理学。可以看到,在一个几乎没有学习、只有十万个神经元的大脑里,你已经能看到我们所谓智能的雏形。然后,随着我们大得多、也更有人类特色、更聪明的大脑,它以各种方式扩展开来。

和黑猩猩、乌鸦一样,我们是顶尖的物理问题解决者。比如这个一岁的孩子,他在解决一个大概从没见人示范过的问题:叠杯子,但不是用孩子们通常的叠法,而是把杯子叠在一只猫的背上。他好像给自己定了目标,要像中世纪修士那样看看一只猫背上能放几个杯子。他发现大概是四个,然后换了个更容易的目标:把杯子弄到猫的另一边去。视频里我最喜欢的部分是刚才那一幕,他伸手向后去拿那个紫色的杯子,这是一岁孩子身上非常突出的客体永久性(object permanence):他至少有一分钟、甚至更久没有看到也没有碰到那个杯子。孩子大一些之后会做更复杂的事,比如这个孩子用乐高搭一座大塔。还有一个很好的物理问题解决的例子,来自我的同事凯尔西·艾伦(Kelsey Allen),她现在在DeepMind,明年将去温哥华的不列颠哥伦比亚大学任教,稍后我会展示她的一些工作。这个例子是她研究的那类情形:这个孩子在玩复活节找彩蛋,彩蛋在左上角,他够不着,于是想到也许可以用那把铲子去够。这大概不是他见过示范的铲子用法,他就是自己想出来的。当然,发现不行之后,他换了模式,改抓铲子的另一头,把铲子头当把手,也许这样行。你们觉得呢?他没来得及弄清楚,因为他姐姐来了。不过最终也还好。这些孩子可爱又贴心,至少有些是,但每一个都聪明得惊人。

这种对物理世界的常识性理解,以及在物理世界中聪明地行动和规划的能力,是我们从演化中继承来的,但也能看出,它在人类身上通过生物演化和文化演化共同发展出了独特的形态。接着,如果要谈人类特有的规模扩展方式,那就是语言进入图景的时候:当孩子大到能学会口语,尤其是能阅读的时候,那才是真正的奇点。在这个意义上,人工智能里的语言模型确实抓住了人类通向智能的规模化路径中某种重要的东西。人类智能的那项独一无二的成就,是文化以社会的方式、并且跨代地集体建立起来的知识;我们通过书面语言以及其他持久的载体,还有口头传统,接入这些知识,再为之添砖加瓦。我们为什么要来大学?我们为什么要建大学?我们今天坐在这里,基本上就是为了这个目的。讲课、交谈、研讨、写论文、读论文,我们之所以认定这些活动是共享和共建知识的独一无二的好办法,是有原因的。但我们能走到这一步,是建立在前面各位看到的那一切之上的。

核心命题:大脑是猜测与下注机器

特南鲍姆: 所以我们这个领域要理解的是,如何用计算的语言把这些想法表达出来。我要说,这里反映的是我个人的观点,但也是计算认知科学以及人工智能其他分支里许多人的共识,这些分支你在新闻里听得不多,原因可以理解,但你应该多听听。

我们的理解是这样的,这个看法带点主观,但我认为也有事实根据:大脑身上那个在人类身上得到扩展、又被语言进一步放大的根本之物,不是学习。人类大脑是了不起的学习机器,但你在简单得多的大脑里看到的那种智能,在语言之前就存在,甚至在任何显著的学习之前就存在。大脑的根本计算功能,是做出好的猜测和好的下注,并且借助某种世界模型来做这件事。连那条斑马鱼都有一个简单的世界模型。它很简单,比你我或者黑猩猩、乌鸦的模型受限得多,也远没有那么灵活。但核心想法是:大脑做的事,实际上是在多个粒度、多个抽象和一般性的层次上建立世界模型,用这些模型弄清楚外面有什么,把感官数据理解为关于世界、关于自身在世界中的位置的信息,对许多动物来说,还包括其他主体,无论是同类、捕食者还是猎物。然后用这些模型对如何花费时间、能量和其他资源做出好的下注:下一步做什么,或者下一步想什么。

这个方程来自马克斯·克莱曼-韦纳(Max Kleiman-Weiner)的一张幻灯片,灵感来自彼得·诺维格(Peter Norvig)。诺维格和斯图尔特·罗素(Stuart Russell)合写了人工智能领域大概最经典的教科书《人工智能:一种现代方法》。这个书名有点怪,因为书最初写于上世纪九十年代中期,而「现代」这个词的含义大家都知道一直在被重新定义。我想问一下,读过或者至少翻开过这本书的请举手。好。如果你在这里,又对人工智能有任何兴趣,都应该去找这本书来读。它有很多版,最新一版几年前刚出。我之所以推荐这本书,是因为当前围绕人工智能的讨论,绝大部分其实只是在谈ChatGPT和相关的大语言模型,而把这个领域的大量基础都撇在一边,这很惊人。很多人,尤其是这几年入行的学生,可以在不学这个领域大部分基础的情况下,成功而高效地做我们称为人工智能的事。但当你翻开这本书,会看到一套几十年发展起来的工具,它们大致都可以用这个基本方程来表达。某种意义上,它回到了理性的经典观念,各位在经济学或任何研究理性决策的学科里都熟悉:一个理性主体(rational agent)选择行动以最大化期望效用。这作为理性的一种形式化,已有几百年历史。

人工智能和计算认知科学做的事,是试图刻画心智的不同方面、不同组件、不同的表征和算法,说明它们各自如何把这个想法兑现。令人兴奋的是,现代语言模型和其他技术还让我们看到语言可能在哪里进入图景。如果大脑从根本上是为做这类事情而演化的,然后语言进来了,那么我们学习语言,基本上就是把语言翻译成或关联到那些用于理性推断和决策的表征;但语言同时也大大地放大和扩展了它们,因为语言使得新的世界建模、新的推断和规划成为可能,那些是没有语言时我们根本无法设想的。这就是我和其他许多人一直在做的事情的路线图,从高层次讲。

你可能会问:在这张路线图上,我们走了多远?和大规模神经计算取得的成果相比,我们接近了吗?显然还没到,否则你在新闻里读到的就是这些,而不是ChatGPT。这条路线之所以没有更广为人知,某种程度上是因为走这条路的人虽然已经取得了不少进展,但前面的路还很长。我们写了大量论文,但除非你读文献,否则很难了解。为了弥补这一点,我和几位同事,特别是汤姆·格里菲思(Tom Griffiths)和尼克·查特(Nick Chater),还有许多贡献了章节的作者,合作了一本即将出版的书。幻灯片上那个日期只是我编译LaTeX文件的日期,书今年晚些时候出版。恕我自夸,我觉得这是一本很好的书,我可以这么说,一部分是因为我参与了这本书的构想并做了不少工作,但真正领头的是格里菲思,还有查特,以及许多贡献章节的作者。它介于教科书和研究专著之间,不是通俗读物,但在座任何人都能读,可以从中了解这套理性推断与决策的工具包,某种意义上就是最大化期望效用式的思维,如何被用来理解并且真正逆向工程人类认知的方方面面。

我只展示其中的一两个方面,然后回到语言。因为虽然我从语言讲起,人人都在谈语言模型,语言也确实是人类的奇点,但这条路线的核心,是语言建立在那些古老得多的常识表征之上。

脑中的游戏引擎与概率程序

特南鲍姆: 为了思考直觉物理,以及因果推理和直觉心理学里的相关问题,我们和其他人发展了两个想法。第一个,我们有时概括为「脑中的游戏引擎」(the game engine in your head)。想法是,电子游戏产业开发的那套工具,各位可能因为玩游戏或编写游戏而熟悉,统称为「游戏引擎」的那些东西,也许正是用工程语言来思考斯佩尔克所说的「核心知识」(core knowledge)的一种方式,也就是思考那些内置于我们大脑、也许与许多动物共享的表征、算法和程序。

游戏引擎,如果各位不熟悉,是一种非常快速(至少设计目标是非常快速)但近似地创建和模拟交互式世界的方法。它包括图形、物理,也许还有对其他主体的建模,总之是让玩家沉浸在一个也许陌生、却又奇怪地熟悉的世界里所需的一切。这是《塞尔达》系列的一款游戏。这个世界里有各种魔法和物理挑战,玩家在解决一个物理问题,有点像那个用铲子的孩子想拿到复活节彩蛋,但在这里玩家得传送,得扔东西,东西会神奇地粘在别的东西上,是一种奇怪的物理变体。它在物理上不可能,你却能一下子看懂。软件让这种高度交互、高度灵活的物理问题解决成为可能。

这类游戏式物理引擎不一定按物理学家的方式建模物理,虽然它们借鉴物理学,但为了高效地处理各种各样的物理场景,它们做了各种近似和权宜之计。物理引擎不需要正确,只需要看上去像样,或者在短时间尺度上看上去足够像样。在这个意义上,它们也符合我们认为大脑所看重的许多性能特征。这张幻灯片和这些图像稍加修改自哈佛另一位杰出同事托默·乌尔曼(Tomer Ullman)的幻灯片。想了解这个想法,可以读乌尔曼几年前发表在《认知科学趋势》上的论文《心智游戏》(Mind Games),那大概是这个想法最好的书面入门。另一张图也来自乌尔曼,它说明了我们说「脑中的游戏引擎」时指什么:我们指的不是一个用来训练某个机器学习算法的模拟训练场,那也是一种可能的用法,游戏引擎在人工智能里确实经常这么用;我们指的是用工程语言刻画这个孩子看到积木和上面的鸟时脑子里发生的事。这很像我们刚才把GPT放进去的那种情境,但这个孩子会想象:它稳不稳?如果我把这个球滚过去会怎样?他在心里模拟碰撞和后果,然后想:这是我想要的吗?这就是这种工程框架如何用来刻画早期常识性世界建模的某些方面。

但它要真正成为一个智能的框架,在做出好的猜测和下注、做最大化期望效用的事情这个意义上,就必须是概率性的,而且不只是概率性的。我们这个研究项目的另一个关键技术思想,叫做「概率程序」(probabilistic programs)。和神经网络类似,这是一个笼统的说法,涵盖多年甚至几十年间演化出来的一组思想。可以把它看作从工程角度把几种理解智能的计算范式中最好的特性结合起来。它包括概率推断或贝叶斯推断,也就是从观察到的结果反推可能的原因;但也包括符号程序,比如根据程序的输出对其输入做概率推断。这很重要,因为如果你想刻画直觉物理这样的东西,只说「有某个高维高斯分布」或者「有某种弱非线性的东西」是不够的。你需要拿一个物理引擎模拟器,把它当作概率模型的核心。它描述世界中的因果过程,于是你可以做概率推断:如果我这么做,可能发生什么?我看到那个结果时,输入大概是什么?这东西有多重?施加了什么力?现代概率编程语言还整合了神经网络。我固然在论证,我们所说的神经网络,也就是人工神经网络里那种试图把单个神经元或一个神经元网络的功能规模化的计算抽象,不是理解人类智能来源的最佳工程方式;但它仍然很有价值,可微系统、用向量做分布式表征、联想记忆,这些都是非常强大的思想。现代概率编程语言让你把这三种母题结合起来,这非常有力量。

直觉物理实验:积木塔到台球

特南鲍姆: 下面举几个例子,说明我们如何用这套东西刻画语言之前就存在的常识智能,然后再简单说说语言如何在此之上建立。

和乌尔曼以及其他一些同事一起,我们建立了一些模型,有时统称为「直觉物理引擎」(intuitive physics engine)。它取游戏引擎中的物理引擎部分,包裹在一个概率推断与近似模拟的框架里,用来建模人们对这类积木塔场景做出的各种常识判断。所以不难理解我为什么会对那些用语言模型探索叠塔的事情感兴趣。比如给出左边这些堆叠,你可以问:这座塔有多稳?倒的可能性多大?如果倒,会往哪边倒?积木会掉多远?如果部分积木换了颜色,而不同颜色的积木重量不同,会怎样?

比如,如果我告诉你灰色积木比绿色积木重十倍,那么对于上面和左边这两座塔,你对倒向哪边的预测就会不同。积木的几何形状完全一样,我只是重新上了色,但你脑中的物理引擎,或者不管你有什么样的心理物理模拟器,会做出不同的预测。这说明它既对质量敏感,又可以被我用一句话「编程」:灰的比绿的重十倍。或者,看到这些出人意料地稳定的场景,你能推断出哪些积木可能比另一些重得多,从而解释它们为什么不倒吗?在每一种情形下,我们都能基于这个想法建立定量预测模型。我不讲细节,但稍后会给各位看一眼模拟器。我们的实验做法是:给被试一批不同的刺激,控制某些变异来源,比如堆叠的复杂程度、以不同方式表现出来的不稳定程度,同时控制其他混淆因素;然后测量,比如请被试在一到七的量表上打分,或者做二选一的强迫选择,再把一批人的结果汇总,也可以看反应时。测量关于稳定性的分级直觉有很多办法。

这张图的纵轴是一组被试对五六十座不同积木塔的平均稳定性判断,横轴是模型预测,即在一个近似的游戏式物理引擎里做少量模拟后取平均的结果。这个模拟器的概率性体现在,场景里有若干东西是视觉刺激不能完全确定的:我们不知道积木的精确位置,单凭一张图片不可能完美地解决逆向图形学问题;我们不知道摩擦系数的准确值;我们还允许可能存在的力扰动,比如一阵风,或者有人撞了桌子。在这些不确定性之下(这一点非常重要),我们可以推断出某些堆叠更稳或更不稳,倒下的积木更多或更少,或者根本不倒。横轴的模型预测与纵轴的人类判断之间有很好的相关。

同样的做法还能扩展到一种更奇怪的任务。和心理学里的许多实验一样,我们研究心智的办法是给你一些你没什么经验的东西,这样才能弄清你做的事不是简单学来的。「积木塔会不会倒」这个问题很多人都有相当的经验,玩过叠叠乐(Jenga)的人经验就更多。但我可以给你看一个类似的场景,比如这些放着红色和黄色积木的桌子,然后问一个如果你没听过我讲、大概从来没想过的问题:如果你狠狠撞一下桌子,把一些积木撞到地上,地上的积木更可能是红的还是黄的?人们同样能做出分级判断,就是纵轴;模型也能做出分级判断。对于这个非常古怪的问题,模型的预测力几乎和它对「有多大可能倒」这个熟悉得多的任务的预测力一样好。

这是怎么做到的?因为我们的系统不是在数据上训练,而是建了一个模型,在做模拟,而且只做少量非常粗糙的概率模拟。我具体说明一下。这是模拟器在其中一个场景上的运行窗口。这里我模拟了轻轻撞一下桌子;这里是同一配置下撞得更狠。模拟器以半随机的方式,概率性地尝试几种不同大小、几个不同角度的撞击。把少数几次模拟的结果汇总起来,就足以对这个问题给出相当合理的答案。模拟也不必跑很久。可以看到,头几个时间步之后,至少对许多场景,你已经知道答案了,或者知道自己不确定。所以我们只需要跑少量模拟,每次只跑少量时间步,而且可以用低分辨率跑,就足以刻画那里发生的事。这至少让各位窥见了我们如何用这些工具建模大脑做出的猜测和下注。

在动态任务里我们可以做类似的事。比如凯文·史密斯(Kevin Smith)和同事的这项工作:一个台球在桌面上弹来弹去,你要预测它先碰到红色的面还是绿色的面。我们可以搞点互动。如果你觉得它会碰红色,就喊「噢」;觉得会碰绿色,就喊「啊」。开始。得快点,它马上要碰红色了。噢。如果我没有操作失误,你们本来会先喊「噢」,再喊「啊」。但我搞砸了。不过你们可以在心里模拟。注意,刚开始的时候,大多数人相当确定它会擦到红色,结果它刚好没碰到;你要过一会儿才有把握,通常要到差不多这个位置,才确信它会碰绿色。人们能够模拟,但在这类开放场景里,通常最多只能模拟一两次反弹。我们的刻画方式是,假设人们做的是刚才那种模拟的动态版本,同样带有噪声或不确定性:球的位置和速度的轨迹有不确定性,尤其是反弹,因为除非你是有经验的台球玩家,否则你对如何外推反弹相当没把握。这个模型,图上显示的是概率模拟产生的不确定性锥,只要把几次模拟的结果加起来,就能刻画人们「大概是红,然后不确定,然后大概是绿」的倾向。

更定量地说,这项研究有一百个不同场景,我这里展示四个,画的是动态轨迹:上面是人,下面是模型,显示分级信心如何随场景推进而变化。数据有很多有趣的结构。有些场景,人们非常确定是一种,然后切换到另一种;有时很长时间都不确定,然后突然明朗;有时切换好几次。这个模型能够刻画这么多,实在令人惊叹。

我还有幸和埃德·武尔(Ed Vul)合作,他很早以前在我实验室,现在在加州大学圣迭戈分校和亚马逊。我们和婴儿研究者安娜·特格拉斯(Erno Téglás)、卢卡·博纳蒂(Luca Bonatti)等人一起研究婴儿的直觉物理。十二个月大的婴儿看这样的场景,像一台小小的糖球抽奖机:物体在里面弹跳一会儿,然后被遮挡一段时间,最后一个物体出现。实验各条件之间变化的是:最终出现的那种颜色的物体是多还是少(三个还是一个)、遮挡时刻这些物体的位置、遮挡时长。一个非常简单的概率物理模拟模型能根据这些变量做出不同预测。考虑到从婴儿身上获取定量行为数据极其困难,它以一种整合的、相当有说服力的定量方式刻画了婴儿的不确定预测。研究婴儿心智的经典方法,是斯佩尔克等人开创的「违反预期」范式或其他注视时长测量:大人和婴儿一样,看到令人惊讶的东西会看得更久。在这里,注视时长可以由这类物理模型中的逆概率做定量预测。

最近的一些工作,来自我们实验室的萨姆·沙耶特(Sam Cheyette)和特雷西·米尔斯(Tracey Mills),研究人们如何预测一个运动点的简单序列。这远远超出了物理,可以带有更丰富的结构,几乎是各种算法或程序的结构。仅仅是预测接下来会发生什么的能力,就是研究心智模型及其如何用于在世界中猜测和下注的一种强有力的方法。而且这些能力是人类独有的,沙耶特等人用猕猴做的研究表明了这一点;它们从儿童期到成年期有显著发展,不过目前看来,人类儿童更像成人,而不像猕猴。

工具游戏与直觉心理学

特南鲍姆: 前面我提到艾伦和史密斯等人的工作,也就是展示用铲子的孩子时提到的。限于时间,我基本跳过,只推荐各位去读他们发表在《美国科学院院刊》和即将发表的论文。为了研究人们如何使用工具、如何灵活地运用直觉物理来使用工具,艾伦和史密斯设计了一个很酷的电子游戏,我们叫它「虚拟工具游戏」。你的目标始终是把红色物体稳稳地弄进绿色区域。你的操作是选一个物体,也就是一件工具,把它扔进场景,然后物理过程展开,你看结果如何。人们很喜欢玩这个游戏。各位可以去麻省理工学院博物馆亲自玩,它是肯德尔广场新馆人工智能展的一部分。和许多其他类型的物理问题解决一样,但与机器学习里通常所说的强化学习非常不同,这是一种强化学习任务:你行动,得到一些反馈,再行动,从试错中学习。但在这个游戏的许多关卡里,人们学得极快,有时只要两三次尝试,很少超过五到十次。

我们的建模方式,基本上是把那个模拟器放进循环:有一个过程从一个非常简单的先验出发,想象解决问题的可能方式,试图引发碰撞或物体之间的相互作用,想象会发生什么。如果你在脑子里试了几次,找到一个很可能在现实中奏效的模拟,你就去试;否则就在空间的另一部分重新采样,重复这个过程。这能相当精确地刻画人们的学习曲线。这是物理游戏里二十个不同关卡的红蓝两条曲线:蓝线是人的累计成功概率随尝试次数的变化,红线是模型。所以,这个「在脑子里试几个想法,有一个看上去可行就到现实中试,失败了就重来」的想法,既刻画了人们觉得哪些题难、哪些题容易,也刻画了最终的成功水平(有些题人们并不总能解出来),还刻画了学习的相对速率,也就是你需要几次尝试才能想明白。

关于直觉物理我基本就讲到这里,关于常识也基本讲到这里。我想留最后几分钟谈语言。这里只指一下另外一整条工作线,部分是和这里的丽贝卡·萨克斯(Rebecca Saxe)合作的,她是我的好友和麻省理工学院同事;还有其他一些人,比如曾是我的研究生的朱利安·哈拉-埃廷格(Julian Jara-Ettinger),以及同在麻省理工学院的劳拉·舒尔茨(Laura Schulz),舒尔茨在其中很多工作里都有参与。如果你想了解这条平行的研究线,也就是在刻画物理世界因果结构的程序之上做概率推断,只不过这里是社会主体如何与之互动,可以读哈拉-埃廷格等人在《认知科学趋势》上的另一篇好文章《朴素效用演算》(The Naive Utility Calculus),2016年发表。顺便说,哈拉-埃廷格现在是耶鲁大学心理学系的教员。这是我参与过的工作中最喜欢的一部分。

我们在这里做的,同样是用程序来描述世界,但现在的程序以知觉和目标为输入,规划行动。这些模型实际上也在做理性决策:目标是主体的奖励来源,行动有代价。在这项工作里,我们把主体建模为制定计划的人,也就是根据他对世界的信念,选择一连串行动来最大化期望效用,选择支持自己目标的行动,等等。但关键在于,这些不是关于人类主体的模型,而是关于「人类主体对其他主体的模型」的模型,那些其他主体也可能不是人类。比如凯莉·哈姆林(Kiley Hamlin)等人用小球做的研究,或者海德尔(Heider)和西梅尔(Simmel)、格尔盖伊(Gergely)和奇布拉(Csibra)的经典工作,这些延续多年的漂亮研究表明,人类成年人甚至婴儿,都能理解那些按照高效行动规划原则运动的小形状。

在这条研究线里,我们用类似的刺激来研究。比如克里斯·贝克(Chris Baker)和哈拉-埃廷格与萨克斯和我一起做的「餐车」刺激:一个主体在世界里移动,寻找校园里停在两个黄色停车位之一的几辆餐车中的某一辆,可能是韩国餐车、黎巴嫩餐车或墨西哥餐车,不同的日子停在不同的位置。当一个主体走出来,绕过楼房看到黎巴嫩餐车,然后转身回到韩国餐车,你问:这个主体最喜欢什么?韩国菜、黎巴嫩菜还是墨西哥菜?人们推断,其实是墨西哥菜。这就是右边「欲望」下面标着M的那根柱子:墨西哥菜是她的最爱,尽管场景里根本没有出现墨西哥餐车。因为要解释她为什么走一条高效的路线,却不是走向她实际的目标,而是走向她以为可能在那里的东西,这是最好的解释。我们还问人们:这个主体最初的信念是什么?他们回答,和我们的模型预测的一样,她大概以为墨西哥餐车在那里。这是她最初的信念,也是她想要的,所以她才绕到另一边去看。而她之所以回头,是因为餐车不在那里,她到了那儿才发现。在许多不同的刺激上,我们以受控方式变换刺激,这个模型能够定量预测人们对主体的欲望和信念的推断,以及两者如何相互作用。因为在这样的场景里,只有假定一个不在场的目标,加上一个「它很可能在那里」的错误初始信念,才能解释她的行动。

这只是尝一尝这套工具包的味道:这里是一种双重贝叶斯式的、近似理性的概率推断,推断的对象是其他近似理性的主体自己所做的推断。它能够刻画我们常识世界模型的核心方面。

语言如何接入:从词模型到世界模型

特南鲍姆: 那么语言如何进入这幅图景?我时间已经不多了,实际上我想已经超时了。但我在开头许下了承诺,希望各位再给我五六分钟,让我基本上是推介一下我在致谢中提到的两位学生的一篇很好的论文。

准确说,作者有一群人,但共同第一作者是梁力昂和格兰德,卢也起了关键作用,还有我的几位教员同事,雅各布·安德烈亚斯(Jacob Andreas)、维卡什·曼辛卡(Vikash Mansinghka)和诺亚·古德曼(Noah Goodman)。这篇论文可以在arXiv上找到,目前正在为几个不同的发表目标修改。它非常长,正被拆成几部分,我建议各位至少读第一部分,喜欢的话再往下读。我们写这篇论文时,管它叫「白皮书」,因为它其实是一个研究纲领的路线图,而不是一组研究结果。它试图展示一条路径:把现代语言模型,以及捕捉语言统计规律的种种方法,建立在我刚才展示的那类常识知识之上,通过概率程序这套机制,让它们得到放大和扩展,变得强大得多。

论文考虑的是我们用语言来指导和组织思维的种种方式,不管是在直觉物理、直觉心理学还是其他领域;也考虑语言扩展思维的种种方式:我们如何学到新概念,或者通过别人明确的解释,或者隐含地从人们在特定句式和模式中使用语言的方式里学到;甚至建立起整套新的直觉理论,而这些理论大多是别人告诉我们的。我们要解释的就是:你是怎么做到这些的?以及我们如何借此理解当前大语言模型里的智能是什么、缺了什么、什么是更接近人类的前进方向。

论文的核心想法是把两样东西结合起来。第一样是我们所说的「概率性的思维语言」(probabilistic language of thought)。这来自一本较新的关于概念的书里的一章,说较新,是相对于劳伦斯(Laurence)和马戈利斯(Margolis)那本更老的概念书,其实也是十年前的了,那一章由古德曼、托比·格斯滕伯格(Tobias Gerstenberg)和我合写,解释了如何使用一种概率编程语言。要进一步了解那一章里的概率编程语言,可以看probmods.org这个网络教材,就是这里显示的这个。这个例子用的是概率编程语言Church。我们在这本网络教材里展示的是,只用这一种编程语言,它是Lisp的一个变体,就能以我们使用Lisp的方式刻画结构化的概率模型,比如直觉物理、直觉心理学等等。也就是说,我们有一种概率编程语言,能刻画各位刚才看到的基本上全部认知模型,所有这些关于猜测与下注、做近似理性的期望效用推断的框架。我们主张,它可以为所有这些常识模型提供一个统一的基底。

第二样,是我们在近期大语言模型身上看到的:它们不只是自然语言的模型,它们也用代码训练,用许多种编程语言的源代码训练。而源代码,正如人们常说但我认为没有被充分领会的那样,是设计给人读和写的编程语言或程序,而不只是给机器执行或阅读的。人们常指出编程语言和自然语言非常不同,但源代码,不管是Python、JavaScript还是Lisp,写出来的方式往往仍然相当像语言:句法上有些相似,比如层次化的短语结构;尤其相似的是我们给变量和函数命名的方式,当然还有代码注释。所以某种意义上可以说,这篇论文的想法是把「源代码」当作思维语言的隐喻:把自然语言和概率编程语言里的程序混在一起。

论文展示的是,可以用一个甚至算不上最先进的语言模型,一个早期的「代码大模型」,即OpenAI的Codex,来定义一个概率性的翻译映射,从自然语言到概率编程语言(本文里是Church),原则上也可以反过来,但论文主要做正向;然后把它应用到许多不同领域。我只展示一个应用,把话题拉回直觉物理。各位还记得,我们用自然语言来提问,来描述你也许从未做过的直觉物理推理。现在我们基本上可以建立一个端到端的、由语言条件化和语言指导的直觉物理模型:取一个能刻画刚才那种桌面物理的二维物理引擎,用一个代码大模型来刻画这样一个想法,即语言在语境中的意义可以看作一种翻译。比如对一个桌面场景的描述:有一摞高高的红色积木,还有两摞矮矮的黄色积木。问题还是那个:如果我狠狠撞一下桌子,把一些积木撞到地上,红的多还是黄的多?模型能把这段话翻译成内部代码,然后跑这样一个模拟。这不是看到的,而是想象出来的:这是模型根据那段语言运行的心理模拟的一个视图。我们也可以描述更复杂的场景,想象出相应的、更复杂的物理场景,模拟那里发生的事。

本质上,这只是把同样的几样东西模块化地组合起来:一个语言系统,假定类似人脑里的那个,把我们说的和听的翻译成某种心理语言;这种心理语言配置大脑的物理引擎;引擎跑少量的概率模拟。仅凭这些,就足以得到一个有定量预测力的模型,几乎和前面我们亲手为模型设定场景时的预测力相当,虽然稍逊一点。这只是这条正在进行的工作线的一点味道,由我提到的那些学生和其他一些人在推进。要说明的是,这篇论文是和梁力昂、格兰德一起做的,但实验和建模研究是和塞德·约翰(Cedegao Zhang,图上那位)一起做的,他也是麻省理工学院的学生。我们和其他人正在心智理论和社会认知领域推进类似的研究,我只能说,敬请期待。

结论:理解、意义与开放的AI

特南鲍姆: 最后说几句我们要往哪里去,这点我忍不住要讲。认知科学以及所有相关领域,语言学、哲学,当然还有人工智能,长期以来都在问:意义是什么?理解语言是什么意思?用语言传达的一个思想是什么?在我展示的这幅工程草图里,有朝向一种意义理论的一步,如果我敢这么说的话。只是一步,还有很多要填补。但这个想法在不同形式下有很长的历史,而现在我认为我们有条件用工程的方式把它抓住:不是抽象地看意义,而是看语境中的意义。一段语言在语境中意味着什么?一个词、一个短语、一个句子,放在这句话的其余部分、放在前面的其他句子里,或者放在你借助知觉系统等其他途径建立起来的对当前情境的模型里,意味着什么?答案是某种自然语言与代码之间的联合分布,但这里的代码是一种概率性的思维编程语言。这个联合分布是核心对象,代码大模型结合概率推断能够刻画它。把这些部件拼起来,你至少就有了一些积木,可以把各个领域里传统上关于意义的几种不同看法统一起来。

意义是什么?是形式语义学那样的东西,即思想的组合式逻辑构造吗?是分布式词嵌入那样的东西吗?后者不只是现代大语言模型的想法,在心理学和自然语言处理里都很古老。这是两种有意思的看法。还有两种:意义在于扎根于世界,这是具身论题;意义在于语用和语境。某种意义上,我展示的概率编程和概率性思维语言,给了你意义的组合性和扎根性两方面;代码大模型的视角给了你分布统计的一面,至少也给了语用或语境一面的某种近似。把它们放在一起,就又是一些积木,让我们能真正思考语言究竟是关于什么的,语言的使用是怎么回事,以及语言如何建立在思维之上,又极大地放大和扩展思维。

开放的问题很多,我就以几点结论收尾。思维是什么?它不再只是一个谜,尽管我们离用计算刻画各种形式的思维还很远。但那个根本的想法,智能在于做出好的猜测和好的下注,在于我们的心智对世界以及自身在世界中的位置所持有的模型,在于以理性的方式用这些模型来指导下一步做什么、下一步想什么,我们已经开始有工具来抓住它了。而这给了我们一条路径,尤其是当语言进入图景之后,去真正理解我们所有人聚在这里是为了什么:不只是知识,而是理解。是那些深广的系统,我们每个人和所有人一起,通过自身的经验和通过语言建立起来,用以在一个不断变化、也在被我们改变的世界里维系和增长知识。我展示的这些工具,我认为是真正尝试理解这件事所需要的积木的一部分。在技术层面,现代概率编程语言与各种神经网络和语言模型能够带来一种「神经、符号、概率」的综合。但把这一切拼在一起,我们还远远处在开头。

如果这些东西让各位感到兴奋或受到启发,那正是我今天的目的。我希望各位考虑,无论是和我们合作,还是在自己的工作里,无论以哪种最能激发你的方式,去思考如何建造这样的智能,不管是在科学一侧还是工程一侧。尤其是对人工智能感兴趣的人,我认为这是必要的。除非我们愿意把人工智能的未来完全交给少数资源雄厚的公司和个人,否则,如果我们希望人工智能的未来更开放、更民主,更扎根于科学、也更能反哺科学,我认为我们需要这样的工具和这样的心态。事关重大。所以,加入我们吧。谢谢。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 开场:讲者介绍与研究定位 ▶ 正在看
6:03 规模路线:从神经元到 ChatGPT ▶ 正在看
9:31 流畅却脆弱:叠塔难题的失败 ▶ 正在看
16:13 脑科学证据:语言与思维分离 ▶ 正在看
20:12 发展与演化:思维先于语言 ▶ 正在看
30:43 核心命题:大脑是猜测与下注机器 ▶ 正在看
36:14 脑中的游戏引擎与概率程序 ▶ 正在看
42:16 直觉物理实验:积木塔到台球 ▶ 正在看
51:57 工具游戏与直觉心理学 ▶ 正在看
57:45 语言如何接入:从词模型到世界模型 ▶ 正在看
67:29 结论:理解、意义与开放的 AI ▶ 正在看
本期小问 · 档案清单
—— 能流畅对话的大语言模型,真的在思考吗? ▶ 正在看
—— 大脑的根本功能是学习,还是靠世界模型下注? ▶ 正在看
—— 人类只听一亿词就学会语言,靠什么先天知识? ▶ 正在看
—— 语言是思维的载体,还是思维放大器? ▶ 正在看
本期讲者
乔什·特南鲍姆MIT 计算认知科学教授,计算认知科学实验室负责人,贝叶斯认知科学与概率程序学习的代表人物,2019 年麦克阿瑟奖得主。
Dan Freed哈佛大学数学科学与应用中心(CMSA)主任,数学家,本场 Yip 讲座主持人。
01开场:讲者介绍与研究定位
0:00
So I'm Dan Freed, director of the Center of Mathematical Sciences and Applications here at Harvard, and it's my pleasure you pleasure to welcome to you to the fourth installment of the Yip Lectures in Mathematics, Physics, Physiology, and Psychology. So thanks to the Yip family for endowing this lecture series, and uh I'll turn it over to HT, who will introduce the speaker. Thanks a lot. Uh I'm very happy that Professor Josh Tenenbaum agreed to give this lecture. Um he's a professor of computational cognitive science at MIT, and he's he undergraduate in physics from Yale University in 1993, and his PhD is at MIT in 1999. Uh so um he's uh he's now the the head of computational cognitive science lab at MIT, and also this MIT's project on quest for intelligence.
我是丹·弗里德(Dan Freed),哈佛大学数学科学与应用中心的主任。很荣幸欢迎大家来到第四届叶氏数学、物理、生理学与心理学讲座。感谢叶氏家族对这一系列讲座的捐赠支持。下面我把时间交给 HT,由他来介绍今天的演讲人。非常感谢。我非常高兴 Josh Tenenbaum 教授答应来做这场讲座。他是麻省理工学院计算认知科学教授,1993 年在耶鲁大学获得物理学本科学位,1999 年在麻省理工学院获得博士学位。他现在是麻省理工学院计算认知科学实验室的负责人,同时也负责 MIT 的「智能探索计划」(Quest for Intelligence)。
便签笔记
1:07
Now um the Professor Tenenbaum is is well known for his fundamental contribution to the mathematical psychology and Bayesian cognitive science. So he received many many awards and recognitions. Uh I will only mention some of them. And the first one is uh his Bayesian program learning research was selected as one of Scientific American's uh 10 world-changing ideas of 2016. And the second one was in in 2018, this research and development magazine named Josh Tenenbaum as their uh innovator of the year. And finally, he's received MacArthur Prize in the 2019, and for his pioneer work uh to develop and apply probabilistic and statistical modeling to the study of human learning, reasoning, and the
Tenenbaum 教授以他在数学心理学和贝叶斯认知科学方面的奠基性贡献而广为人知。他获得过非常多的奖项和荣誉,我只提其中几项。第一项是,他的贝叶斯程序学习研究被《科学美国人》评为 2016 年「十大改变世界的思想」之一。第二项是,2018 年《研发》(R&D)杂志把 Josh Tenenbaum 评为年度创新者。最后,他在 2019 年获得了麦克阿瑟奖,表彰他在把概率与统计建模开创性地发展并应用于人类学习、推理与知觉研究方面的工作。
便签笔记
2:05
perception. So let's welcome Professor Tenenbaum. All right. Thank thank you. Can you Can you hear me okay? In the room and up there? Okay. Great. Um it's a pleasure to be here and an honor. Thank you all for coming. So I'm going to be talking about uh how to build a mind from a brain.
那么,让我们欢迎 Tenenbaum 教授。好的,谢谢,谢谢。大家能听清吗?现场和上面都能听到吗?好的,太好了。很高兴也很荣幸来到这里,感谢各位到场。今天我要讲的是:如何从一个大脑构建出一个心智。
便签笔记
2:32
As you maybe could gather, if you don't already know what I do and uh how I do it, if you might have gathered from the introduction, the research that I do lives between multiple fields. Um I identify primarily as a cognitive scientist, but I also work in neuroscience and artificial intelligence. Our goal is to try to understand most fundamentally how intelligence arises in the human mind and brain, but we do that in what we sometimes call reverse engineering. So we want models, theories, but theories expressed as computational models that take the form of engineering. They have math, but they're also really engineering style models. And so that's what we mean by understanding, at least in terms of our science here. And that also has a a a payoff potentially, maybe world-changing, maybe or maybe not, but we'll see, for building artificial intelligence systems which are intelligent in more human-like ways and might live in a more valuable, productive, safe, and just overall positive way in the human world, the world that we humans have built.
如果你原本还不了解我做什么、怎么做,从刚才的介绍中大概也能看出来,我做的研究横跨好几个领域。我主要认同自己是一名认知科学家,但我也做神经科学和人工智能的工作。我们的目标是从最根本的层面理解智能是如何在人类的心智和大脑中产生的,而我们采取的方式有时被称为「逆向工程」。我们要的是模型、是理论,但这些理论要以计算模型的形式表达出来,具有工程的形态。它们有数学,但同时也确实是工程风格的模型。这就是我们所说的「理解」,至少在我们这门科学里是这样。而这也可能带来回报——也许能改变世界,也许不能,我们拭目以待——那就是构建以更接近人类的方式实现智能的人工智能系统,它们或许能在人类所建造的这个世界里,以更有价值、更有生产力、更安全、总体上更积极的方式存在。
便签笔记
3:38
I'm going to talk about some of the work we've been doing here, and I just want to acknowledge at the beginning a number of collaborators. Um the ones who are in bold are the ones whose work uh play some particular prominent role. The ones who are really in bold, um this uh students uh Leo or Lionel Wong, Gabriel Grant, Alex Lev, and Tyler Brook Wilson um contributed in a big way to the material and the vision of this talk and even many of the slides. So I want to particularly highlight them. Um Now the I think to all of us who who work in this field or who even just are adjacent, it's a especially exciting time, really a thrilling time, and sometimes a scary time to be working at the intersection of these fields, right? AI is everywhere. Um the field is old, but it's it's just in the last year or two, it's on everybody's um front of the conversation.
我要讲讲我们最近做的一些工作,一开始我想先感谢一些合作者。加粗的这些人,是工作发挥了特别重要作用的。而真正加粗的这几位——学生 Leo,也就是 Lionel Wong、Gabriel Grant、Alex Lev,还有 Tyler Brook Wilson——对这场报告的内容和整体思路,甚至很多幻灯片,都做出了很大贡献。所以我想特别提一下他们。我想,对我们所有在这个领域工作、甚至只是相邻领域的人来说,现在是一个特别令人兴奋、真的令人激动、有时也有点吓人的时代,对吧?AI 无处不在。这个领域其实很老了,但就在最近这一两年,它一下子成了所有人谈论的焦点。
便签笔记
4:35
On the science side, with advances in neuroscience and cognitive science, we are starting to be able to study intelligence across many different kinds of brains. And it's fascinating to us to ask, what is what is in common? What is the is what is it about brains in general that defines intelligence? And of course we see differences across species, and we see something clearly distinctive and unique about human intelligence. Humans are special. There's a reason why we're here. Just just, you know, we we as neuroscientists and cognitive scientists, we study many other animals. We study mice, flies, zebrafish, birds, macaque monkeys, chimps, but there's a reason why we're here in this room, and the mice and the flies are out in the fields or maybe hiding in the corners, and the chimps are somewhere sadly going extinct.
在科学这一侧,随着神经科学和认知科学的进展,我们开始能够跨越许多不同类型的大脑来研究智能。让我们着迷的问题是:其中有什么是共通的?大脑一般而言的什么特性定义了智能?当然我们也看到物种之间的差异,我们看到人类智能显然有其独特之处。人类是特别的。我们之所以在这里,是有原因的。我们作为神经科学家和认知科学家,研究许多其他动物:老鼠、果蝇、斑马鱼、鸟类、猕猴、黑猩猩;但我们之所以坐在这个房间里,而老鼠和果蝇在田野里或者躲在角落里,黑猩猩则在某个地方可悲地走向灭绝,是有原因的。
便签笔记
5:26
So I'd like to understand what is unique and special about human intelligence in computational terms. And on the AI side, many of us are wondering, with all the rapid progress in AI, where is it going? Will there come a time when AI is not only just chatting with us, but actually asking these same questions, maybe together with us or on its own? If on its own, what are we? Are we the mice or the chimps? So there's a lot on our minds, okay. Um one way of thinking about how you get to human intelligence from maybe simpler, more basic kinds of brains. I'm going to start with that thesis.
所以我想用计算的语言来理解,人类智能究竟独特和特别在哪里。而在 AI 这一侧,我们很多人都在想,AI 进展这么快,它会走向哪里?会不会有一天,AI 不只是跟我们聊天,而是真的开始问这些同样的问题——也许是跟我们一起问,也许是它自己在问。如果是它自己在问,那我们又算什么?我们是老鼠,还是黑猩猩?所以我们心里有很多问题。有一种思路是:如何从可能更简单、更基础的大脑出发,走向人类智能。我先从这个论点讲起。
便签笔记
02规模路线:从神经元到 ChatGPT
6:03
It's not the one It's the one that in some sense I think most most of the public conversation is about, but within cognitive science, I think it's pretty well acknowledged that there's a different picture, and that's the one that I'm going to focus on in this talk. But let me start with this. So one idea is to is to think about what we can abstract away from a single neuron. Okay, this is from a the one of the classic papers of Hubel and Wiesel on studying uh cells in in cortex, in primary visual cortex. And the idea that there's some basic computational element, a kind of inner product of inputs and weights, and then a non-linearity, and then a learning element, like some kind of learning rule, way of adjusting the weights of the neuron in response to to minimize some cost function, some objective function. That's a very basic idea, and it's very powerful one. When you scale it up to very large and deep neural networks, this is the famous AlexNet deep convolutional neural network, and train it on very large data
这个论点在某种意义上是当下公众讨论的主流,但在认知科学界,大家其实相当认可另有一幅图景,而那才是我这场报告要聚焦的。不过还是先从这个讲起。一种想法是思考我们能从单个神经元中抽象出什么。这张图来自 Hubel 和 Wiesel 研究皮层——初级视觉皮层——细胞的经典论文之一。这个想法是:存在某种基本的计算单元,做输入和权重的某种内积,然后过一个非线性,再加上一个学习环节,比如某种学习规则,用来调整神经元的权重,以最小化某个代价函数、某个目标函数。这是个非常基本的想法,也是非常强大的想法。当你把它扩展到非常大、非常深的神经网络——这就是著名的 AlexNet 深度卷积神经网络——并在非常大的数据集上训练它,
便签笔记
6:59
sets like Fei-Fei Li and company's ImageNet, you can do amazing things. Computer vision experienced the deep learning revolution more than 10 years ago when sufficiently deep and big networks were trained on sufficiently big data sets like ImageNet. And then with variations on these architectures, such as the transformer, and training that now on all the text that you can find on the internet, effectively all the language that humans have produced and made public um basically in human history, you know, you have things like this. And I don't even have to tell you what that symbol represents, because it's everywhere. It's probably on many of your phones and laptops, right? But I will state mention it, ChatGPT is the thing that made AI, you know, daily on the New York Times and and all of our news feeds, okay. And it's clearly changing the world in many ways.
比如李飞飞团队的 ImageNet,你就能做出惊人的事情。十多年前,当足够深、足够大的网络在 ImageNet 这类足够大的数据集上被训练出来时,计算机视觉就经历了深度学习革命。再往后,在这些架构上做变体,比如 Transformer,并用你能在互联网上找到的所有文本来训练它——实际上就是人类历史上产生并公开的所有语言——你就得到了这样的东西。我甚至不用告诉你这个标志代表什么,因为它到处都是,很可能就装在你们许多人的手机和笔记本上,对吧?但我还是说一下:ChatGPT 就是那个让 AI 每天出现在《纽约时报》和我们所有新闻推送里的东西。它显然正在以许多方式改变世界。
便签笔记
7:47
It's so much so that people many people are asking, well, maybe you know, either this is the route or a route to actually finally having machines that are that are true general forms of AI. Certainly, ChatGPT and similar kinds of large language models present to the general public the way AI was always supposed to present to us, as as computers that can do basically anything that computers can do, but that you talk to, you ask them to do something for you rather than having to program them, right? That's every science fiction movie. That's That's what That's what AI is, right? Whether it's C-3PO or the computer on Star Trek or any number of other examples. I won't mention Hal, except I guess I just mentioned Hal. Okay. Um so we have AI now presenting itself the way it was always supposed to. And it's And it it you know, these are systems which which clearly are general in that they can do many different things.
以至于很多人在问:也许这就是——或者至少是一条——最终造出真正通用 AI 机器的路径。可以肯定的是,ChatGPT 和类似的大语言模型,向公众呈现出的正是 AI 一直以来「应该」呈现的样子:一台基本上能做任何计算机能做的事的电脑,但你是跟它对话、请它替你做事,而不是必须去给它编程,对吧?每一部科幻电影都是这样。这就是 AI 的样子,对吧?不管是 C-3PO,还是《星际迷航》里的电脑,或者其他无数例子。我不打算提 HAL,虽然我想我刚刚已经提了。所以现在 AI 终于以它一直应该有的方式登场了。而且这些系统显然是通用的,因为它们能做许多不同的事。
便签笔记
8:37
And so many people are are asking, you know, is this both on the AI side, is this where thinking is is going to come from if we're actually going to have true human-like AGI? Or you might Many people are also speculating, maybe this is what happened in actual biology, right? That just with enough scale of bigger and bigger brains, which certainly has happened over evolution. Like our brains are the biggest, okay. Um chimps, basically the next biggest. Having a big brain is helpful, okay. Um is it just about scale and and and you know, is that enough? Okay. Um well, again, as I think we're all familiar with, I don't I'm not going to give a whole lecture on ChatGPT, but just to give one example of a phenomenon I think we're all familiar with if we're paying any attention to these systems, they they do remarkable kinds of things that look like and in some sense are kinds of thinking. And yet they're also
所以很多人在问——在 AI 这一侧——如果我们真要造出类人的通用人工智能,思维会不会就是从这里来的?也有很多人在猜测,也许真实的生物演化中发生的正是这件事,对吧?只要大脑规模不断变大——演化中确实发生了这件事,我们的大脑是最大的,黑猩猩基本上是第二大的——大脑大是有好处的。那么,是不是只关乎规模?规模够了就够了吗?好吧,我想大家都熟悉这些,我不打算整场讲 ChatGPT,但只举一个我们只要稍加关注这些系统就都熟悉的现象:它们能做出一些非常了不起、看起来像、在某种意义上也确实是某种思考的事情。可与此同时,它们又
便签笔记
03流畅却脆弱:叠塔难题的失败
9:31
remarkably fragile. So to cite one example, this is a example that was on Twitter from my MIT colleague well, he former MIT colleague, now Stanford, Erik Brynjolfsson, um something that was spooky to him, um asking a kind of, you know, intuitive physics example. And as you'll see later on in the talk, this is one of the areas I like to work on in cognitive science. So Erik asks ChatGPT, how would you make a stable stack from a book, four tennis balls, a nail, a wine glass, a wad of chewed gum, and five uncooked spaghetti sticks. Crazy idea. You can start answering it for yourself. Here's ChatGPT's answer. A systematic recipe. Well, first, you start off by placing the four tennis balls on a flat surface in a square arrangement with each ball being a corner of the square. Then, you place the book on top of the tennis balls, ensuring that each corner of the book is resting on a tennis ball.
脆弱得惊人。举一个例子,这是我 MIT 的同事——好吧,前同事,现在在斯坦福——Erik Brynjolfsson 发在推特上的例子,一个让他觉得有点毛骨悚然的例子,问的是一个直觉物理学的问题。你们在报告后面会看到,这正是我在认知科学里喜欢研究的领域之一。Erik 问 ChatGPT:你会怎么用一本书、四个网球、一根钉子、一个高脚酒杯、一团嚼过的口香糖和五根没煮过的意面,搭出一个稳定的塔?很疯狂的题目。你们可以自己先想想答案。这是 ChatGPT 的回答,一份有条理的操作步骤。首先,把四个网球放在平面上,摆成正方形,每个球是正方形的一个角。然后,把书放在网球上面,确保书的每个角都搁在一个网球上。
便签笔记
10:19
This will create a stable base with the book acting as a platform. And it goes on and on and on. It's amazing, right? I mean, if you if you you can look at particular cases here, and you know, it really appears to have, and in some sense, I would say, has some kind of understanding of the physical world objects and and what what is likely to happen in certain configurations and what can be done with them. We can keep going, and this is an example from Tyler B. Wilson, one of the students I mentioned, a recent graduate of MIT philosophy, now at NYU as a post-doctoral fellow. Um so, Tyler asked ChatGPT to create a series of block-stacking challenges kind of like that. And you can see it comes up with things like the office supplies skyscraper, the garden variety tower, the artist's ascent, and then asked ChatGPT to solve them. But if you and basically it's it's creative. It can come up with it can sort of remix the ideas in all sorts of interesting ways. It gets the genre from that one example and whatever else it's been trained on.
这样就形成了一个稳定的底座,书充当平台。它就这样一直往下写。很惊人,对吧?我是说,你可以去看具体的这些例子,它真的显得——在某种意义上我会说它确实——对物理世界的物体、对某些构型下可能发生什么、对这些物体能拿来做什么,有某种理解。我们继续看,这个例子来自我刚才提到的学生之一 Tyler B. Wilson,他刚从 MIT 哲学系毕业,现在在纽约大学做博士后。Tyler 让 ChatGPT 创造一系列类似的搭积木挑战。你可以看到它想出了「办公用品摩天楼」「花园塔」「艺术家的攀升」这些题目,然后再让 ChatGPT 去解这些题。基本上它是有创造力的,能以各种有趣的方式重新组合这些想法。它从那一个例子以及它训练过的其他材料中把握住了这个体裁。
便签笔记
11:16
But if you actually look at the solutions, this is where you now start to see some of the puzzling things, some of the ways its intelligence is fragile. So, just picking in on zooming in on any one, and they're all kind of like this. The artist's ascent. So, the problem of the artist's ascent I don't know if anybody else can read it. Using a paintbrush, a small canvas no larger than 8 by 10 in, a tube of acrylic paint, a palette, a cup of water, construct a tower at least 12 in tall. The structure must stand on its own for at least 45 seconds. Okay, so how do you solve that? Well, open the paint tube and use some paint to adhere the brush to the canvas, creating a base. Stand the palette vertically against the brush. Balance the cup of water on the palette using the remaining paint as an adhesive if necessary.
但如果你真去看它给的解法,你就开始看到一些令人费解的东西,看到它的智能脆弱在哪里。随便挑一个来看——它们其实都差不多——比如「艺术家的攀升」。这道题不知道其他人能不能看清:使用一支画笔、一块不大于 8×10 英寸的小画布、一管丙烯颜料、一块调色板、一杯水,搭出至少 12 英寸高的塔。这个结构必须能自立至少 45 秒。好,那怎么解?它说:打开颜料管,用一些颜料把画笔粘在画布上,做成底座。把调色板竖着靠在画笔上。把那杯水平衡地放在调色板上,必要时用剩下的颜料当粘合剂。
便签笔记
11:56
I mean, it kind of sounds good. It's very fluent English, and it's exploiting various actual physical properties and affordances of these objects. But if you think of it as a stable stacked to be able to do this, it's completely crazy. And by the way, when he when he prompted ChatGPT, um he and following other things I did, made it very clear in the prompt to say, "It's okay if you can't solve the problem. It's be uncertain or say you can't solve the problem. It's better to be uncertain or say you can't solve it than to give a solution which is which isn't going to work. But that doesn't stop ChatGPT from very creatively it proposing many solutions which just are not going to work. So, what's going on here? I mean, these these and many other phenomena that many people have written about should make us be really be asking, "What is the relation between language and thought?" A question that many
听起来还挺像回事的。英文非常流畅,也确实利用了这些物体真实的物理性质和可供性。但如果你把它当成一个真要搭起来的稳定塔,那完全是天方夜谭。顺便说一句,他在给 ChatGPT 提示时,我后来也做过类似的事,在提示里说得很清楚:解不出来没关系,可以表示不确定,或者说你解不出来;宁可不确定或者承认解不出,也别给一个根本行不通的方案。但这并没能阻止 ChatGPT 极富创造力地提出一大堆根本行不通的方案。那么,这里到底发生了什么?我是说,这些现象,以及很多人写过的其他许多现象,应该让我们真正去追问:语言和思维之间是什么关系?这个问题有很多人
便签笔记
12:45
people um including uh Steve Wakin, um many people in this room have written really insightful things on that question, especially at Harvard over the years. Um and and from a computational point of view, if we think about a cognitive architecture for language, what is the what is what are the relation and the roles of language and thinking in that architecture? So, that's that's effectively what uh we we are trying to work on in this research program. Okay. Um implicit in the idea that ChatGPT or a large language model um is, you know, that has some emergent thinking properties, is this this sort of picture. And in some sense, this is an old idea going back to Shannon, the end you know, the the founder of information theory and the first statistical models of language. The idea that trying to the the task of predicting language, saying like, "What will be the next word or or or character um or the next sequence of characters or words is in some sense kind of intelligence
写过非常有洞见的东西,包括 Steve Wakin,还有在座的许多人,尤其是哈佛这些年来的很多学者。而从计算的角度看,如果我们思考一个语言的认知架构,那么在这个架构里,语言和思维各自扮演什么角色、二者是什么关系?这基本上就是我们这个研究项目在努力解决的问题。在「ChatGPT 或大语言模型具有某种涌现出来的思维属性」这个想法背后,隐含的是这样一幅图景。而在某种意义上,这是个很老的想法,可以追溯到香农——信息论的奠基人,也是最早的语言统计模型的提出者。这个想法是:预测语言这项任务,也就是预测「下一个词、下一个字符,或者接下来的一串字符或词是什么」,在某种意义上是「智能完备」的,
便签笔记
13:48
complete, that whether it's AI or cognition, that thinking effectively can be operationalized as whatever latent predictive structure is there in between the text that's come before and the text that's come after. Okay. Now, there are some kinds of tasks like the ones here where comp where where you can see you have to use various kinds of common sense knowledge or knowledge about the world to complete the pattern. Okay. Um but there's also a pretty clear pattern evident in the surface form of language. And so, it might seem sort of plausible that just being able to predict the future of any sentence in context based on the first part of the
也就是说,不管在 AI 还是在认知层面,思维实际上都可以被操作化为「前文与后文之间存在的那种潜在的预测结构」。当然,有些任务,比如这里的这些,你可以看出必须动用各种常识知识或世界知识才能补全这个模式。但语言的表层形式里也存在相当明显的规律。所以看上去似乎有几分道理:仅仅是能够根据句子的前半部分
便签笔记
14:23
sentence or the sentences that have come before could be a way to deploy and maybe to train a generally intelligent architecture. But then when you take any particular computational system, like any particular transformer, which is a finite architecture trained on a finite data set, even very very big, okay, that's doing especially this the the way a large language model works or any sort of predictive sequence model, that's doing auto-regressive predictions, so predicting the next element, token, or word based on all the others and then just repeating that going forward. There's no particular reason that should be able to solve any problem or perform any kind of thinking, even if there's a lot of interesting knowledge and understanding stored in the latent structure. So, here are some examples of problems for which there's a clear answer, okay, and you can compute it, but it's unlikely that a auto-regressive language model is going to compute it. Now, there are ways that people in the AI community have been sort of trying to extend these models, like what's sometimes called chain of thought or tree of thought. Um if you basically train the model to think step by step,
或者前面的句子,来预测任意一个句子在上下文中的后续,也许就能成为部署、甚至训练一个通用智能架构的途径。但当你拿出任何一个具体的计算系统,比如某个具体的 Transformer——它是一个有限的架构,在一个有限的数据集上训练,哪怕这个数据集非常非常大——尤其是大语言模型或任何序列预测模型的工作方式,是做自回归预测:根据前面所有的内容预测下一个元素、token 或词,然后不断重复往前推进。那就没有什么特别的理由认为它应该能解决任何问题、能完成任何形式的思考,即使它的潜在结构里确实存储了很多有意思的知识和理解。所以这里有几个例子,问题有明确的答案,你也可以算出来,但自回归语言模型不太可能算得出来。当然,AI 界的人一直在尝试用各种方式扩展这些模型,比如所谓的「思维链」或「思维树」。如果你训练模型一步一步地思考,
便签笔记
15:28
which means to generate out not the immediate answer, so not just to fill in the blank, like to to that question, but to um basically generate intermediate steps of computation, then that can help. Um it turns out that, you know, again, these are moving targets, but that helps for the top problem. It doesn't help for the bottom problem. Um but this is this is one of the things that the large language model AI community is trying to do in order to get the language models to become more general thinking systems. There are other things that are really really important in ChatGPT and other such models, so-called reinforcement learning from human feedback, so tuning them to be predictive not only of the next token in language, but of the kind of thing that people tend to give thumbs up to as opposed to thumbs down. Okay.
也就是让它不要直接生成最终答案——不是简单地把那道题的空填上——而是生成中间的计算步骤,这确实有帮助。结果表明,当然这些都是不断变动的目标,这对上面那个问题有帮助,对下面那个问题就没用了。但这正是大语言模型这一 AI 社群正在尝试做的事情之一,为的是让语言模型成为更通用的思考系统。ChatGPT 和其他这类模型里还有一些非常非常重要的东西,比如所谓的「基于人类反馈的强化学习」,也就是把模型调教得不只是能预测语言中的下一个 token,还能预测人们倾向于点赞而不是点踩的那类内容。
便签笔记
04脑科学证据:语言与思维分离
16:13
What what vision this expresses as a scaling path to intelligence again is one that that's that puts language at the center and says, "All the patterns in language, that is enough." And from an from an architectural point of view says that like language is the substrate of thought. Okay. Um again, as as as as Steve Pinker and many others I think have written about, there's real problems with that. And I'm not going to go into this, but just refer you, for example, to some of Steve's writing on that. Like The Stuff of Thought, great classic book to get some introduction to that. Okay. Um I'll I'll let me let me give a few other cognitive and neuroscience perspectives, though. So, for example, in the work of my colleague at MIT, Ev Fedorenko, um she and many others, but these are a
这条通往智能的「规模化路径」所表达的愿景,仍然是把语言放在中心位置,认为「语言中的所有模式,这就够了」。而从架构的角度说,它认为语言就是思维的载体。同样,正如 Steve Pinker 和其他许多人写过的那样,这里存在真正的问题。我不打算展开,只是推荐你们去看看比如 Steve 的相关著作,《思想本质》(The Stuff of Thought)就是很好的入门经典。不过让我再给出几个认知科学和神经科学方面的视角。比如我在 MIT 的同事 Ev Fedorenko 的工作——她和其他很多人的工作,不过这
便签笔记
16:53
couple of slides I borrowed from Ev. Um she has studied the the language network in the human brain using fMRI and other kinds of neuroimaging as well as other neural recording from humans and other neuropsychology and just behavioral methods to understand a system that again people have studied for, you know, much more than 100 years and with different methods, but a part a network of brain regions shown here in one of Ev's slides that is distinctively involved in using language. That includes both speaking, like I'm doing, and understanding, like I hope you're doing, um also reading and writing, and uh it's it's um doesn't matter, you know, it's you can you can be uh blind and learn Braille. Language is language, okay.
几张幻灯片是我从 Ev 那里借来的。她用 fMRI 和其他神经影像方法,以及来自人类的其他神经记录、神经心理学和纯行为学方法,研究了人脑中的语言网络。这个系统人们其实已经用不同方法研究了远超一百年,而 Ev 这张幻灯片上展示的,是一组特异性地参与语言使用的脑区网络。这既包括说话——像我现在这样——也包括理解,希望你们此刻正在做的事情,还包括阅读和书写。而且这无所谓形式,你可以是盲人、学盲文,语言就是语言。
便签笔记
17:37
Um language is language regardless of what language you grew up speaking, even if you learn to speak one of these constructed languages, like an artificially made-up language like Klingon or Dothraki from Game of Thrones. If you're a if you're a fluent speaker, your language network, as Ev and colleagues showed, is activating in just basically the same way. Um the the large language models, these predictive transformers, and this is some work I played a small role in with Martin Schrimpf and Ev and other colleagues, large language models act like this is actually a smallish language model by today's standard, a GPT-2 model, can provide a reasonably good or at least state-of-the-art quantitative predictive model of responses in the language areas of the human brain during language processing tasks.
不管你从小说的是哪种语言,语言就是语言;哪怕你学的是某种人造语言,比如克林贡语,或者《权力的游戏》里的多斯拉克语。只要你是流利的使用者,正如 Ev 和同事们所展示的,你的语言网络的激活方式基本上是一样的。而大语言模型,这些预测性的 Transformer——这是我和 Martin Schrimpf、Ev 以及其他同事合作、我参与了一小部分的工作——按今天的标准来说,那其实是个偏小的语言模型,一个 GPT-2 模型,它能为人在做语言加工任务时大脑语言区的反应,提供一个相当不错、至少是当时最好的定量预测模型。
便签笔记
18:21
So, that's interesting. I think I'm not trying to argue that language models don't tell us something about language. But as Ev, along with colleagues Kyle Mahowald and Anna Ivanova, a bunch of us wrote a a piece which is coming out soon in Trends in Cognitive Science, sort of are trying to articulate this modern computational and cognitive neuroscience perspective on the relation between language and thought. So, I'd refer you to that piece also for a lot of really really good perspectives on this. Um and that team played a played a role in one of this is one of the early studies applying large language models, or now we'd call them small language models. These models are actually plausibly something a little bit more like the the training data these models have is a little bit closer to the training data that a person might get. Um and they're not bad, and at least they're our best quantitative predictive models of responses in the human brain, and much better than simpler kinds of language models, okay.
这很有意思。我并不是想论证语言模型不能告诉我们任何关于语言的东西。Ev 和 Kyle Mahowald、Anna Ivanova 等同事,我们一群人写了一篇文章,即将发表在《认知科学趋势》上,试图阐述这种现代计算与认知神经科学视角下的语言与思维的关系。我推荐你们也去看看那篇文章,里面有很多非常好的观点。那个团队参与了最早一批把大语言模型——或者按现在的说法叫小语言模型——应用于此的研究之一。这些模型其实有一点更合理之处:它们所获得的训练数据,比起今天的大模型,更接近一个人可能接触到的训练数据。它们表现得不错,至少它们是我们对人脑反应最好的定量预测模型,比更简单的语言模型好得多。
便签笔记
19:13
Um but these models, if you remember GPT-2, nobody thought that GPT-2 was a general model of intelligence, okay? It started to capture a lot of syntax and some of the like syntax-semantics interface. So, pretty interesting stuff, but it certainly didn't didn't capture general intelligence, and it it also didn't capture all of semantics either. But as Ev and colleagues have also shown, right? Lang- the language network, the part that's that's actually modeled by a small to medium-size neural language model, is really just about language. So, when people are doing other tasks, including things that that seem that are similar to language in some ways, like manip- like solving algebraic equations or understanding code, other kinds of logical reasoning, the language network is not involved in those if you if you know how to look carefully. And similarly, brain damage can selectively impair just language and not other forms of, let's say, symbolic thinking. So, bottom line, language and thought are different in the brain.
但这些模型,如果你还记得 GPT-2,没有人认为 GPT-2 是一个通用的智能模型。它开始捕捉到了大量句法,以及一些句法—语义接口的东西,挺有意思的,但它当然没有捕捉到通用智能,也没有捕捉到全部语义。而正如 Ev 和同事们也已经证明的:语言网络——也就是能被中小规模神经语言模型建模的那部分——真的就只跟语言有关。所以当人们在做其他任务时,包括某些在某些方面看起来跟语言类似的任务,比如解代数方程、理解代码、其他形式的逻辑推理,只要你会仔细地去看,就会发现语言网络并没有参与其中。同样,脑损伤可以选择性地只损害语言,而不损害其他形式的,比如说符号思维。所以结论是:在大脑中,语言和思维是不同的。
便签笔记
05发展与演化:思维先于语言
20:12
But an even, I think, a more compelling perspective comes from looking developmentally. So, how do how do children come grow into intelligence? I mean, AI people, going back to Alan Turing, certainly in my own AI research, we've been very much inspired by thinking about how children grow into intelligence. And as, you know, if you think about it, in the same way that human intelligence is unique, human beings as learning machines are unique. A human child is the only learning machine in the known universe that reliably grows into full human intelligence starting from much less. So, let's look at how they learn to think and how they learn language. Right? As, again, to cite another, uh, distinguished Harvard colleague, Elizabeth Spelke, also in psychology, this is, um, her book recently published. She and many others in the field of infant cognition have laid out all the ways in which human babies, well before they've learned language, already have systems for thinking, systems for thinking about physical objects, systems for thinking
但我认为,一个更有说服力的视角来自发展的角度。也就是说,孩子是怎么成长为具有智能的个体的?搞 AI 的人,从阿兰·图灵开始,当然也包括我自己的 AI 研究,都深受「思考孩子如何成长为智能体」的启发。而且你想想看,正如人类智能是独一无二的,人类作为学习机器也是独一无二的。人类儿童是已知宇宙中唯一一种从远远更少的起点出发、能可靠地成长为完整人类智能的学习机器。所以让我们看看他们是怎么学会思考、怎么学会语言的。再引用另一位杰出的哈佛同事——同样在心理学系的 Elizabeth Spelke——这是她最近出版的书。她和婴儿认知领域的许多人已经阐明了:人类婴儿在学会语言很久之前,就已经拥有各种思维系统,比如思考物理对象的系统,以及思考
便签笔记
21:12
about other agents, their goals and plans, and how they interact with physical world and each other, for example. So, clearly, developmentally, thinking comes first. And that is crucial, we we understand as a field or we believe as, you know, this is somewhere between a belief and a pretty well-supported set of facts, um, but there's still debate, um, that that that the knowledge that humans have in in it prior to learning language is a key part of how we're able to learn language from so much less data than one of these large language models, okay? Um, the latest models might be trained on a trillion words, but a human child might hear people estimate maybe 5 million words a year. So, maybe something like 100 million words over the first 20 years, okay? That's a very, very small corpus by today's AI standards.
关于其他智能体、它们的目标和计划,以及它们如何与物理世界、如何彼此互动,诸如此类。所以很明显,从发展的角度看,思维是先出现的。这一点很关键——我们这个领域所理解的,或者说所相信的(这介于一种信念和一组相当有力的事实之间,当然仍然存在争议)是:人类在学会语言之前就已经拥有的知识,是我们能够用比这些大语言模型少得多的数据学会语言的关键,对吧?最新的模型可能是用上万亿个词训练出来的,而一个人类小孩,人们估计每年大概能听到500万个词左右。所以在头20年里,大概也就一亿个词。按今天的AI标准,这是一个非常非常小的语料库。
便签笔记
22:06
Um, maybe most interestingly in development is the work of, um, Susan Goldin-Meadow. And I find every time I talk about her work, find this incredibly inspiring. Um, Susan and her and her colleagues and students have for a long time studied home sign, um, as well as related phenomena like the Nicaraguan sign language. Home sign is a phenomenon where children who grow up deaf, um, where they don't get sign language inputs, um, so they basically don't get language input, they still reliably construct their own communication systems that are kind of like proto-languages. They aren't languages, but they're systems of expressing their thoughts, trying to and and in indeed communicating with the other people around them that have many of the kinds of features of language. So, you know, this suggests that there's something about the human mind that is the creator of language. It's not the other way around. It's not training a lots of language. Language comes from the mind, and even in one individual human is able to start to create their own language if they don't get language input.
嗯,在发展心理学里,也许最有意思的是 Susan Goldin-Meadow 的工作。我每次讲到她的研究,都觉得特别受启发。Susan 和她的同事、学生长期研究"家庭手语"(home sign),以及像尼加拉瓜手语这样的相关现象。家庭手语指的是这样一种现象:有些孩子生下来就是聋人,而且得不到手语输入,也就是说他们基本上得不到任何语言输入,但他们仍然会稳定地自己构建出一套沟通系统,有点像"原始语言"。它们还算不上语言,但它们是一套表达自己想法的系统,是在尝试、并且确实在和周围的人交流,而且具备语言的许多特征。所以这暗示着,人类心智中有某种东西才是语言的创造者,而不是反过来。不是靠大量语言训练出来的。语言来自心智,甚至单独一个人在得不到语言输入的情况下,也能开始创造自己的语言。
便签笔记
23:16
Uh, the Nicaraguan sign is an example of a famous case where an orphanage that brought together home signers, children who grew up without really any standard natural language input, brought them together in an orphanage, um, and they, uh, they each had their different home signs, but they kind of figured out ways to communicate with each other, making what some characterize as pidgin and then creoles, basically creating a whole new language from scratch in just a couple of generations. So, again, this is just remarkable when you want to think it to me, when you want to think about the relation between language and thought, that an individual human being with no language input starts to find symbolic ways to express their thoughts, and you put a few of those folks together, and and in basically no time, with very little data and experience in the grand scheme of things compared to all of human culture that produced the internet, they create their own whole new language, and cultures then start to evolve around it.
呃,尼加拉瓜手语是一个著名的例子:一所孤儿院把这些"家庭手语者"聚到了一起,这些孩子成长过程中基本没有得到任何标准自然语言的输入,孤儿院把他们集中在一起,他们各自有各自不同的家庭手语,但他们慢慢摸索出了彼此交流的方式,形成了有人称之为皮钦语、然后是克里奥尔语的东西,基本上是在短短两代人之内从零创造出了一门全新的语言。所以再一次,当你去思考语言和思维的关系时,这在我看来实在了不起:一个没有任何语言输入的人,会开始找到用符号表达自己想法的方式;而你把几个这样的人放在一起,几乎是转眼之间,用相比于造就出互联网的整个人类文化而言极少的数据和经验,他们就创造出了一门全新的语言,然后文化也开始围绕它演化起来。
便签笔记
24:10
So, that's, you know, telling telling us something. Um, and again, maybe most strikingly is not just to think developmentally, but evolutionarily. So, organisms like the fly or the zebrafish that have only 100,000 neurons, or the honeybee has about a million, the mouse, some tens of millions, other traditionally smart animals, um, like great apes or crows, tool users, right? None of them have language, all of them think in some form. Um, nobody's going to disagree that chimps or crows solving tool use physical problem solving are in some ways, of course, they're thinking. Um, the zebrafish is I have a special fondness for due to some work that I was a very small collaborator in work in Florian Engert's lab here also at Harvard that was led by Andrew Bolton. But Florian and others have long studied the zebrafish and what you can get from its brain, um, as well as in the in this case in the project that Andrew looked at, um, behavior. Andrew did 3D psychophysics, looking at how a zebrafish, um, hunts its prey. So, if you if you don't know about zebrafish, the larval zebrafish are really, really small, you need a microscope um, to study them. And they hunt things which are even smaller, single-celled organisms, paramecia, okay?
所以这告诉了我们一些东西。而且,也许更惊人的不只是从发展的角度看,还有从演化的角度看。像果蝇、斑马鱼这样只有十万个神经元的生物,蜜蜂大约有一百万个,老鼠有几千万个,还有其他传统上被认为聪明的动物,比如类人猿或者乌鸦这些会使用工具的动物,对吧?它们都没有语言,但它们都以某种形式在思考。没有人会否认,黑猩猩或乌鸦在使用工具、解决物理问题时,当然在某种意义上是在思考的。我对斑马鱼有一份特别的偏爱,因为我曾作为一个很小的合作者参与过哈佛 Florian Engert 实验室的一项工作,那项工作由 Andrew Bolton 主导。Florian 和其他人长期研究斑马鱼以及能从它大脑中得到什么,而在 Andrew 负责的那个项目里研究的是行为。Andrew 做了三维心理物理学实验,观察斑马鱼是怎么捕食的。如果你不了解斑马鱼的话,幼年斑马鱼真的非常非常小,你得用显微镜才能研究它们。而它们捕食的东西更小,是单细胞生物,草履虫,对吧?
便签笔记
25:26
Um, the paramecia kind of move in a semi-random way, but not completely random, and the parame- and the zebrafish try to attract them. And what Andrew and colleagues showed is that the zebrafish has a kind of predictive model of the paramecia. It's, in a sense, able to make guesses about where the the zebrafish where the paramecium is going and bets about how to get there. Um, its guesses the way you know this is that it it moves in a very discrete set of actions. It, um, rotates. It's kind of like a video game controller for those of you who play video games. It has little flippers or fins that it can flap and a tail that it wiggles. So, it basically moves in a series of somewhat discrete motions of going forward in bouts and turning, okay? And it it will reliably track and pursue a paramecium, but crucially, it doesn't move to where the paramecium is, but it it moves towards where it can predict under an optimal-ish, simple but optimal, statistical predictor of where it's headed. So, it moves towards where it's headed. And if it continues moving towards where it was headed, then the zebrafish catches it. If it deviates a
草履虫的运动是半随机的,但也不是完全随机;而斑马鱼会去捕捉它们。Andrew 和同事们的研究表明,斑马鱼有一种对草履虫的预测模型。某种意义上,它能够猜测草履虫要往哪儿去,并对怎么过去下注。你怎么知道它在猜呢?因为它的动作是一组非常离散的动作。它会转向,有点像玩电子游戏时的手柄——如果你们有人玩游戏的话。它有可以拍动的鳍,还有可以摆动的尾巴。所以它基本上是通过一连串比较离散的动作来移动的:一阵一阵地向前游,然后转向,对吧?它会稳定地追踪并追击一只草履虫,但关键在于,它并不是朝草履虫当前所在的位置游,而是朝一个接近最优的、简单但最优的统计预测器所预测的草履虫去向游。也就是说,它朝着草履虫要去的方向游。如果草履虫继续朝原方向移动,斑马鱼就能抓住它;如果草履虫稍微
便签笔记
26:34
little bit, the zebrafish might give up and go somewhere else. That's the basic psychophysics that Andrew and others studied. So, you can see here that in brain with that with basically almost no learning and just 100,000 neurons, you already see the beginnings of what we think of as intelligence, um, which then scales up in all sorts of ways with our much bigger brains and our much more distinctively human, smarter brains. So, like the chimps, like the crows, we are consummate physical problem solvers, whether it's this 1-year-old doing something, solving a problem that probably he never saw demonstrated for him, stacking up cups not just in the usual ways that kids have been known to stack up cups, but in this case, stacking up cups on the back of a cat. Um, it's as if he's decided that his goal is to see, you know, like a like a medieval monk, how many cups can he get on the back of a cat. Um, he figures out it's about four, and then he switches to an easier goal, which is, let me get the cups to the other side of the cat. Okay? Um, my favorite part of the video is what you just saw when he reached back for the
偏离一点,斑马鱼可能就放弃,转而去别的地方。这就是 Andrew 他们研究的基本心理物理学。所以你可以看到,在这样一个基本上几乎没有学习、只有十万个神经元的大脑里,你已经能看到我们所说的智能的雏形;而这些能力随后以各种方式,在我们大得多的大脑、以及我们更具人类特色、更聪明的大脑里被放大。所以,就像黑猩猩、就像乌鸦一样,我们是极其出色的物理问题解决者。不管是这个一岁小孩在做的事——他解决的这个问题多半从没有人给他示范过——把杯子叠起来,而且不是孩子们常见的那种叠法,而是叠在一只猫的背上。就好像他给自己定了个目标,像个中世纪修士一样,看看猫背上到底能叠多少个杯子。他发现大概是四个,然后他换了个更容易的目标:把杯子弄到猫的另一边去。好吧。这段视频我最喜欢的部分就是你们刚看到的,他回头去够那个
便签笔记
27:39
purple cup, a very striking form of object permanence in this 1-year-old, cuz he hadn't seen or touched that cup for at least a minute, if not a lot longer. When kids get older, they do more complex things, as we are familiar with, such as this kid here making a big tower out of Legos, or in a very nice example of physical problem solving that my colleague, Kelsey Allen, now works at DeepMind, I'll show you some of her work in a little bit. She'll be starting as a faculty member at, uh, University of British Columbia in Vancouver next year. Um, this is an example of the kind of thing that she studied, where this kid here is in an Easter egg hunt. You might see the Easter egg in the upper left, and he can't reach it, but he decided, oh, maybe he can use that shovel to reach it. Not a I probably not a use of the shovel he ever saw demonstrated, but he just figured it out. Of course, when he when he sees it's not working, he switches modes, and now he grabs the other end of the handle, or the other end of the shovel as a handle, and maybe this will work.
紫色杯子的时候——这在一个一岁孩子身上是非常惊人的客体永久性表现,因为他至少有一分钟、甚至更久没有看到或碰过那个杯子了。孩子长大一些之后,会做更复杂的事情,这我们都很熟悉,比如这个孩子用乐高搭出一座大塔;或者一个非常好的物理问题解决的例子,这是我的同事 Kelsey Allen 研究的东西,她现在在 DeepMind 工作,我待会儿会给你们看一些她的研究。她明年将去温哥华的不列颠哥伦比亚大学任教。这就是她研究的那类事情的一个例子:这个孩子在参加复活节找彩蛋,你们可能能看到左上角的彩蛋,他够不着,但他想到,哦,也许可以用那把铲子去够。这多半不是他见过别人示范的铲子用法,是他自己想出来的。当然,当他发现这样不行的时候,他就换了个方式,抓住把手的另一头,或者说把铲子的另一头当成把手,也许这样就行了。
便签笔记
28:37
What do you think? Well, he didn't have time to figure it out cuz his sister came along and, you know, ultimately, yeah, it's okay. He he he did okay. Um, yeah, but, um, so, these kids are are cute and sweet, at least some of them, and but all of them remarkably intelligent, okay? So, this this kind of common sense understanding of and ability to act intelligently and plan in the physical world, something we inherit but from our from from evolution, but you can see probably distinctively develops in human through both biological and cultural evolution. And then, if we want to talk about the human way of scaling, when language comes into the picture, when kids are old enough to learn le- uh, both spoken language and then especially to read, that's the real singularity, right? This is the sense in which language models are really capturing like, you know, language models in AI are really capturing something important about the human scaling route to
你们觉得呢?结果他没来得及琢磨出来,因为他妹妹跑过来了,不过最后嘛,还行啦,他做得还不错。是的,这些孩子既可爱又讨人喜欢,至少有一部分是这样,但他们全都聪明得惊人,对吧?所以这种对物理世界的常识性理解,以及在物理世界中聪明地行动和规划的能力,是我们从演化那里继承来的,但你可以看到,它在人类身上很可能是通过生物演化和文化演化共同发展出来的独特能力。然后,如果我们要谈人类式的规模扩展,那就是语言进入画面的时候——当孩子长到足以学会口语、尤其是学会阅读的时候,那才是真正的奇点,对吧?在这个意义上,AI 里的语言模型确实抓住了人类通往智能的规模化路径中某种重要的东西。
便签笔记
29:36
intelligence. Human intelligence is, or, you know, the singular achievement is what the knowledge that culture builds socially and then collectively across generations that we can tap into through written language you know and and other enduring forms like that as well as oral traditions and then contribute to. What you know, why do we come to universities? Why have we built universities? We're all here for basically this purpose. Okay, there's a reason why lecturing or conversation seminars and writing papers and reading papers why that is the activity that we have come to understand is a singularly good way of sharing and building knowledge together. So that's so but like but we get to that building on all of the stuff you've seen. So what we're trying to understand in our field is how do we capture these ideas in computational terms? And it starts I would say and again I'm I'm reflecting my own view but also the consensus of a lot of people in computation and cognitive science as well as in other areas of AI that you don't hear as much about in the news. Um the you know for for understandable reasons but you should hear more about them.
人类智能,或者说人类独一无二的成就,是文化在社会层面、并且跨世代集体建构起来的知识——我们可以通过书面语言以及其他这类持久的形式,还有口述传统,去汲取它,然后再为它做出贡献。你们知道,我们为什么要来大学?我们为什么要建立大学?我们所有人在这里,基本上就是为了这个目的。对吧,讲课、研讨式的对话、写论文、读论文,这些之所以成为我们公认的、共同分享和建构知识的绝佳方式,是有原因的。但我们能走到这一步,是建立在你们前面看到的所有那些东西之上的。所以我们这个领域要理解的是:怎么用计算的术语来刻画这些想法?我想说,起点是——我这里既是在反映我个人的观点,也是在反映计算与认知科学、以及AI其他一些你们在新闻里听得没那么多的领域中很多人的共识。这些领域你们在新闻里听得少,原因可以理解,但你们应该多听到一些关于它们的事。
便签笔记
06核心命题:大脑是猜测与下注机器
30:43
So we the way we understand it I would say this is a somewhat opinionated but I think also factually grounded view is that the fundamental thing about brains that scales up in humans that language builds on and further scales is not about learning. Okay, brains are human brains are remarkable learning machines but the kinds of intelligence you see in much simpler brains is is there even before before language and even before significant learning. The fundamental computational function of brains is to make good guesses and good bets and to do that with some kind of model of the world. Like even that zebrafish has a simple model of the world. It's simple and it's much more constrained than the models that you and I have or that the the the chimp or the crow have. And much less flexible also. But this idea that what brains do is build effectively build a model of the world at multiple grains and scales of abstraction and generality and use those to figure out what's there to make sense of their sense data in
所以,我们的理解方式——我要说这是一个带有一定立场、但我认为也有事实依据的观点——是:大脑最根本的、在人类身上被放大、被语言进一步扩展的东西,并不是学习。当然,人类大脑是了不起的学习机器,但你在简单得多的大脑里看到的那类智能,在语言出现之前、甚至在任何显著学习发生之前就已经存在了。大脑最基本的计算功能,是做出好的猜测和好的下注,并且是借助某种世界模型来做的。就连那条斑马鱼也有一个简单的世界模型。它很简单,也比你我拥有的模型、或者黑猩猩、乌鸦拥有的模型受限得多,灵活性也差得多。但核心思想是:大脑所做的,实质上是在多个粒度、多个抽象与概括层级上建立世界的模型,并用这些模型去搞清楚世界里有什么,把感觉数据解释成
便签笔记
31:43
terms of the world and then themselves in it and for many animals there are other agents either conspecifics or predators or prey. And then to use those models to make good bets about how to spend your time either I mean your energy and your other resources what to do next or what to think about next. This equation here um has been uh proposed it's it's it's based on a slide by Max Kleiman Weaner from um inspiration from Peter Norvig who together with Stuart Russell wrote what is probably the canonical book on AI. It's called AI a modern approach which is a slightly weird name cuz the book was first written in the mid-90s and the word modern keeps getting redefined as we all know. Um I'm just curious um uh uh raise your hand if you have read or at least cracked open this book. Okay. Um if you're if you're here and you have any interest in AI you all should go pick this book up.
世界中的事物,以及自己在其中的位置;而对许多动物来说,世界里还有其他主体——同类、捕食者或猎物。然后再用这些模型去做出好的下注:怎么花自己的时间,或者说自己的精力和其他资源,接下来做什么,或者接下来想什么。这里这个方程式是 Max Kleiman-Weiner 的一张幻灯片提出来的,灵感来自 Peter Norvig,他和 Stuart Russell 合写了大概是AI领域最经典的那本书,叫《人工智能:一种现代方法》。这个名字有点怪,因为这本书最早写于九十年代中期,而"现代"这个词,我们都知道,是在不断被重新定义的。我很好奇,如果你读过、或者至少翻开过这本书,请举个手。好的。如果你在这儿、并且对AI有任何兴趣,你们都应该去把这本书找来看看。
便签笔记
32:37
There's many editions you can I most recent edition came out just a couple of years ago. Um the reason I recommend and point you to this book is because it's and it's it's it's amazing how much of the current conversation around AI is going on is really just about chat GPT and related like large language models and leaves out so much of the basis of the field. And many people especially students coming into the field these days can successfully and productively do things that we call AI without actually studying most of the basis of the field. But what you see when you get there is a set of tools that again have developed over a number of decades that are sort of expressed in this fundamental equation. It's basically it in some sense it goes back to the classical idea of rationality that you might be familiar with in economics or any other study of sort of rational decision making. Okay. But the idea of an agent a rational agent that chooses actions to maximize their expected utility that goes back hundreds of years as a way to formalize one thing we might mean by rationality.
这本书有很多版本,最新的一版就是几年前出的。我之所以推荐并向你们提这本书,是因为——现在关于AI的讨论有多少其实只是围绕着 ChatGPT 和相关的大语言模型,而把这个领域的大量基础都略过了,这实在令人吃惊。很多人,尤其是这些年新进入这个领域的学生,可以在完全没有系统学习这个领域大部分基础的情况下,成功且高效地做我们称之为AI的事情。但当你真正去学的时候,你会看到一整套经过几十年发展起来的工具,它们某种程度上都可以用这个基本方程式来表达。基本上,从某种意义上说,它可以追溯到你们在经济学、或者任何关于理性决策的研究中可能熟悉的那个经典的理性概念。对吧。但"一个理性主体选择行动以最大化其期望效用"这个想法,作为对"理性"某一层含义的形式化,可以追溯到几百年前。
便签笔记
33:36
And what the the field of AI as well as computational cognitive science has done is try to characterize the different aspects of the mind different components different representations and algorithms in terms of basically how they cash this idea out. Okay. And then I think we can think about and what it's exciting that modern language models and other technologies also let us see where language might come into the picture. So if we have brains that fundamentally evolved to do this kind of thing and then and then language comes into the picture we learn language by by basically translating or relating it into those rep- representations for rational inference and decision making but then also massively amplify and extend them because language enables new kinds of world modeling new kinds of inference and planning that we that we weren't able to really conceive of without
而AI这个领域以及计算认知科学所做的,是试图刻画心智的不同方面、不同组成部分、不同的表征和算法,看它们究竟是怎么把这个想法落实兑现的。好的。然后我认为,令人兴奋的是,现代的语言模型和其他技术也让我们看到了语言可能在哪里进入这幅图景。所以,如果我们的大脑从根本上是演化来做这类事情的,然后语言进入了画面,我们学习语言的方式,基本上是把它翻译成、或者关联到那些用于理性推断和决策的表征上;但与此同时,语言也极大地放大和扩展了它们,因为语言让我们能够进行新型的世界建模、新型的推断和规划,而这些是我们在没有语言的情况下根本无法真正设想的。
便签笔记
34:24
language. So that's I would say the road map uh in high-level terms that we and many others I think have been working on. Now you might ask okay well how far how close are we on this or how far are we on this road map? Is it anywhere even close to this what we've seen with the the sort of large scale up uh neural compute? Well, clearly we're not there yet because otherwise you'd be reading about that instead of chat GPT. And to some extent why this approach I think isn't more broadly known is because the people who've been pursuing it have actually made quite a lot of progress even though there's quite a long ways to go. Um but we've written up a lot of papers and if unless you read the literature you know, it's hard to get at that. So one thing that we've been trying to do to remedy this with a couple of colleagues Tom Griffiths and Nick Chater especially but many others have contributed chapters to this forthcoming book. That date is just the date of uh LaTeX file that I compiled.
所以我要说,这就是我们和其他很多人一直在推进的高层次路线图。那你可能会问,好吧,我们在这条路线图上走到哪一步了?还有多远?它和我们在大规模神经计算上看到的成果相比,哪怕接近吗?很显然我们还没到,否则你们现在读到的就是这个而不是 ChatGPT 了。而这条路线之所以在我看来还没有被更广泛地了解,某种程度上是因为,走这条路的人其实已经取得了相当多的进展,尽管还有很长的路要走。我们写了很多论文,但除非你去读文献,否则很难接触到这些东西。所以,为了改善这一点,我和几位同事——特别是 Tom Griffiths 和 Nick Chater,还有很多其他人为这本即将出版的书贡献了章节。上面那个日期只是我编译 LaTeX 文件的日期。
便签笔记
35:18
So the book will be coming out later this year. And it's uh it's a it's a a really um well, if I do say so myself I think it's a really nice book. Um I can partly say that because well, I helped to conceive of this book and did a bunch of work on it. Tom especially Tom Griffiths really led the charge and along with Nick Chater and also with me and many others who contributed chapters on many topics really trying to convey in a way that's accessible it's it's it's more of like a textbook or somewhere between a textbook and a research monograph. So it's not a popular book but it's one that anybody anyone here certainly could read to try to get a sense of how this toolkit for let's call it you know, rational inference and decision making for what is in some sense maximal expected utility thinking can be used to understand and really reverse engineer so many aspects of human cognition. I'll just show you a little bit about a couple of these things and then come back to language because again though I
所以这本书今年晚些时候就会出版。我自己说可能不太合适,但我觉得这是一本很棒的书。我之所以能这么说,部分原因是我参与了这本书的构思并做了不少工作。特别是 Tom Griffiths,是他带头推动的,还有 Nick Chater、我,以及很多在各个主题上贡献了章节的人,大家真的努力用一种易于理解的方式来呈现——它更像是一本教材,或者介于教材和研究专著之间。所以它不是一本大众读物,但在座的任何人肯定都能读,从中了解这套工具箱——姑且称之为理性推断与决策的工具箱,某种意义上就是最大化期望效用的思维方式——如何被用来理解、并真正逆向工程人类认知的方方面面。我会给你们稍微讲一点其中的几件事,然后再回到语言上来。因为虽然我
便签笔记
07脑中的游戏引擎与概率程序
36:14
started with language and everybody's talking about language models and language is the human singularity it's language building on the common sense representations that are much older that is really the heart of this approach. So two ideas that we and others have developed here um for thinking about well, in particular these topics of basically intuitive physics and related ideas in causal reasoning and intuitive psychology. So one is this is this idea which we sometimes sum up as the game engine in your head. And it's this idea that some of the same kinds of tools that have been developed by the video game industry that many of you might be familiar with either because you play games or because you program games um that are called that are going to fall under the the head heading of game engine um that these might be actually a way to think in engineering terms about what Liz Spelke calls core knowledge or what are what are the representations and algorithms the programs that are built into our brain and shared maybe with many other animals.
是从语言开始讲的,而且现在人人都在谈语言模型、语言也是人类的奇点,但这个进路的核心其实是:语言是建立在那些古老得多的常识表征之上的。所以,这里有两个我们和其他人发展出来的想法,用来思考——具体说就是直觉物理学、以及因果推理和直觉心理学中的相关想法。第一个想法我们有时候概括为"你脑袋里的游戏引擎"。这个想法是说,电子游戏产业开发出来的一些工具——在座很多人可能都熟悉,要么因为你玩游戏,要么因为你编写游戏——也就是归在"游戏引擎"这个名目下的东西,这些东西也许恰恰提供了一种用工程术语来思考 Liz Spelke 所说的"核心知识"的方式,也就是我们大脑中内置的、也许还与许多其他动物共享的那些表征、算法和程序。
便签笔记
37:18
So game engines if you're not familiar with them are very fast or designed to be very fast but approximate ways of creating and simulating interactive worlds. So that includes graphics physics maybe modeling other agents anything needed to immerse a player like this is from one of the Zelda games. Um to immerse a player in a perhaps a strange and also still oddly familiar world. So this is a world where there are various kinds of magic spells and physics challenges you can see the player solving. They are they're trying to solve a physical problem kind of like the kid with the shovel but in this you know, to get their Easter egg but in this case they have to like teleport and throw this thing and it magically sticks to these things and it's sort of a weird variation on physics. It's not physically possible and yet you can get it right away. You can see it and the the software enables this very interactive and flexible kind of physical problem solving.
如果你不熟悉游戏引擎的话,它们是一种非常快速、或者说被设计得非常快速但只是近似的方式,用来创建和模拟可交互的世界。这包括图形、物理,也许还有对其他主体的建模,以及任何让玩家沉浸其中所需要的东西——比如这是塞尔达系列游戏里的一幕。让玩家沉浸在一个也许很陌生、但同时又莫名熟悉的世界里。所以这是一个有各种魔法和物理挑战的世界,你能看到玩家在解决问题。他们其实是在解决一个物理问题,有点像刚才拿铲子的那个孩子去够彩蛋,只不过在这里他们得传送、把这个东西扔出去,它会神奇地粘到那些东西上,这是一种对物理规律的奇怪变体。它在物理上是不可能的,但你一下子就懂了。你能看懂,而且这个软件让这种非常互动、非常灵活的物理问题解决成为可能。
便签笔记
38:13
Okay. Um there's uh many other examples of um ways that these game style physics engines which don't model physics necessarily the way a physicist would although they draw on physics but they make all sorts of approximations and hacks in order to try to handle a wide range of physical scenarios very efficiently. The physics engines don't have to be right they just have to look good or look good enough over short time scales and in that sense they also match many of the performance characteristics we think that the brain cares about. Um I should say this slide is is and these these images here are slightly modified from a slide from Tomer Ullman another one of your distinguished Harvard colleagues. And if you want an introduction to this idea check out Tomer's paper called Mind Games which was a Trends in Cognitive Science paper from a few years ago which is probably one of the best introductions to this idea. Um uh if you if you want to see a a written form of it. This is another Ullman who is just an illustration of um wha- what what we might mean by it when we talk about the game engine in your in the head we don't mean a
好的。还有很多其他例子,说明这类游戏式的物理引擎——它们建模物理的方式未必和物理学家一样,虽然它们借鉴了物理学,但它们做了各种各样的近似和取巧,为的是能非常高效地处理各种各样的物理场景。物理引擎不必是正确的,它们只要看起来不错、或者在短时间尺度上看起来足够好就行;从这个意义上说,它们也符合我们认为大脑所在意的许多性能特征。我要说明一下,这张幻灯片以及上面这些图,是从 Tomer Ullman 的一张幻灯片稍作修改而来的——他是你们哈佛另一位杰出的同事。如果你想了解这个想法的入门,可以看看 Tomer 那篇叫《Mind Games》的论文,是几年前发在《认知科学趋势》上的,它大概是关于这个想法最好的入门介绍之一,如果你想看文字版的话。这是 Ullman 的另一个例子,它正好说明了:当我们说"脑袋里的游戏引擎"时,我们指的并不是
便签笔记
39:22
simulator that's a training ground for some machine learning algorithm which is also a possible way and a way that game game engines have been used a lot in AI but rather a way to capture an engineering terms what's going on inside this kid's head when he thinks about when he imagines in the situation seeing the blocks and the and the and the and the um bird on there. Um much like the situation we thought we were sort of putting GPT in before, but here he imagines, is it stable or not, or what would happen if, for example, I roll this ball, and mentally simulates the collisions and the effect, and thinks, well, do I want that or do I not want that? Okay. So, that's the way this kind of engineering framework can be used to capture some aspects of early common sense world modeling. But where it really becomes a a framework for intelligence in the sense of making good guesses and bets, or doing things that maximize expected utility, well, you have to be probabilistic, and not just probabilistic. The two the other key technical idea in the research program that we work on is what's called probabilistic programs, and this is a toolkit that, you know, kind of like neural networks, that's
一个作为某个机器学习算法训练场的模拟器——那也是一种可能的用法,而且游戏引擎在AI里确实被大量这样使用过——而是一种用工程术语来刻画这个孩子脑子里在发生什么的方式:当他在想、在想象那个情境,看到那些积木和上面那只鸟的时候。这很像我们之前设想让 GPT 面对的那种情境,但在这里,他是在想象:这稳不稳?或者说,如果我把这个球滚过去会发生什么?他会在心里模拟碰撞和后果,然后想:嗯,我想要那样的结果,还是不想要?好的。所以这就是这种工程框架可以用来刻画早期常识性世界建模某些方面的方式。但要让它真正成为一个关于智能的框架——智能是指做出好的猜测和下注、或者做那些最大化期望效用的事情——那你就必须是概率性的,而且不只是概率性的。我们这个研究计划里另一个关键的技术思想,叫做概率程序(probabilistic programs)。这是一套工具箱,有点像神经网络那样,
便签笔记
40:28
it's you know, this is a catch-all phrase for a set of ideas that have evolved over a number of years, even decades, that you can think of as combining the best features from an engineering standpoint of multiple paradigms for understanding intelligence computationally. So, that includes the idea of doing probabilistic or Bayesian inference, um working backwards to infer the likely causes that explain observed effects, for example, but also symbolic programs. So, doing probabilistic inference on, for example, the inputs to programs from the outputs. That's important because if you want to capture something like intuitive physics, it's not enough to just say, well, there's some high-dimensional Gaussian, or maybe there's some weakly non-linear thing. Rather, you want to take something like a physics engine simulator and use it as the heart of your probabilistic model. It's describing the causal processes in the world, and so you want to be able to do probabilistic inferences about what's likely to happen if I do this, or when I see that, what was probably the input?
它是一个笼统的说法,涵盖了一组经过许多年、甚至几十年演化出来的想法;你可以把它看作是从工程角度把多种理解智能的计算范式中最好的特性结合起来。这包括做概率推断或贝叶斯推断的想法——比如反向推理,从观察到的结果推断出可能的原因——但也包括符号程序。也就是说,比如从程序的输出对输入做概率推断。这一点很重要,因为如果你想刻画像直觉物理学这样的东西,光说"这里有个高维高斯分布",或者"也许有个弱非线性的东西"是不够的。相反,你会想把类似物理引擎模拟器这样的东西作为你概率模型的核心。它描述的是世界中的因果过程,所以你希望能对"如果我这样做,可能会发生什么"、或者"当我看到那个时,输入可能是什么"做概率推断。
便签笔记
41:26
What How heavy was this thing, or what forces were applied? And modern probabilistic programming languages also integrate neural networks. So, while I'm arguing that what we call neural networks, namely the computational abstraction in artificial neural networks that tries to scale up what a single neuron or a single network of neurons might do, I'm arguing that is not the our best engineering way to understand where human intelligence comes from. It's a valuable one, and the idea of differentiable systems and and vectors for distributed representation, associative memory, for example, that's a very powerful idea. And modern probabilistic programming languages let you combine all three of these motifs, and that's very powerful. So, just a few examples of how we have used this to capture some of the common sense intelligence that is there before language, and then I'll show you a little bit how language might build on
这个东西有多重?或者施加了什么力?现代的概率编程语言也整合了神经网络。所以,虽然我一直在论证:我们所说的神经网络——也就是人工神经网络这种计算抽象,它试图把单个神经元或单个神经元网络可能做的事情规模化——我认为它并不是我们理解人类智能从何而来的最佳工程路径;但它仍然是有价值的一条路,而可微系统、用向量做分布式表征、联想记忆这些想法,都是非常强大的想法。现代的概率编程语言让你能把这三种母题结合起来,这非常强大。下面我举几个例子,说明我们怎么用它来刻画那些在语言之前就存在的常识智能,然后我会稍微讲一点语言可能是怎么建立在它
便签笔记
08直觉物理实验:积木塔到台球
42:16
it. So, with Tomer and also a number of other colleagues, we've built these models which we sort of have sometimes lumped under the phrase the intuitive physics engine, which takes the the physics engine part of game engines, um wraps it inside a framework for probabilistic inference um and approximate simulation, and uses that to model many kinds of common sense judgments that a person can make, for example, about these block tower scenes. So, not surprising why I'd be interested in some of those tower stack building things that people are exploring with language models, okay? Um so, given, for example, these stacks um on the left, you might ask, how stable is the tower? How likely is a tower to fall? Or if it falls, which way will it fall? Or how far will the blocks fall? Or what happens if some of the blocks are colored differently, and and the different colors are heavier or lighter?
之上的。所以,我和 Tomer 以及其他一些同事一起,构建了这些模型,我们有时把它们统称为"直觉物理引擎":它取用游戏引擎中的物理引擎部分,把它包裹进一个概率推断和近似模拟的框架里,然后用它来建模人们能做出的各种常识判断,比如关于这些积木塔场景的判断。所以你们就不会奇怪,为什么我会对人们用语言模型探索的那些搭塔任务感兴趣了,对吧?比如说,给定左边这些积木堆,你可以问:这座塔有多稳?它倒下来的可能性有多大?如果它倒了,会往哪边倒?积木会散落多远?如果其中一些积木颜色不同,而不同颜色对应更重或更轻,又会怎样?
便签笔记
43:03
So, for example, you can see here that if I tell you that the gray blocks are 10 times heavier than the green blocks, you'll make a different prediction for which way they will fall if they fall on the on the top versus on the left. The geometries of the blocks are the same, I've just recolored them, but yet your mind's physics engine, or your whatever mental physics simulator you have, will make a different prediction, showing that it's both sensitive to mass, but also that I can kind of program it by just telling you something, like the gray stuff is 10 times heavier than the green stuff, okay? Um or if you see these surprisingly stable scenes, can you infer what is likely to be much heavier than another to explain why they're not falling over? So, in each of these cases, we can make quantitatively predictive models based on this idea. I'm not going to go through the details, although I'll show you one example of a simulator glimpse in a second. But the the idea of the the kinds of experiments we do are we give people a number of different stimuli, controlling certain sources of variation, like how, you know, how complex is the stack, or how unstable it
比如说,你可以看到,如果我告诉你灰色积木比绿色积木重10倍,那么对于顶部染色和左侧染色这两种情况,你对它们倒下时往哪边倒会做出不同的预测。积木的几何形状是完全一样的,我只是重新上了色,但你心智中的物理引擎,或者说你脑子里那个心理物理模拟器,会给出不同的预测。这说明它既对质量敏感,同时我也可以通过简单地告诉你一句话来"编程"它,比如"灰色的东西比绿色的重10倍",对吧?或者,如果你看到这些出人意料地稳定的场景,你能不能反推出哪一部分很可能比另一部分重得多,从而解释它们为什么没有倒?在每一种情况下,我们都能基于这个想法建立可以定量预测的模型。我不会讲细节,不过待会儿我会给你们看一个模拟器的片段。我们做的实验大致是这样的:我们给人们一组不同的刺激,控制某些变异来源,比如这堆积木有多复杂,或者它在不同方面有多不稳定,
便签笔记
44:02
might be in different ways, controlling for other sorts of confounds, and then we measure, for example, by asking people on a scale of one to seven, or by having them make a two-alternative forced choice, and just aggregating a bunch of people, or we can look at their reaction times. There's many ways to measure graded intuitions about stability. And on the Y axis here, I'm plotting the average stability judgments of a group of participants for 50 or 60 of these different block towers. On the X axis, what I'm plotting is a model prediction, which is the average of the results of a small number of simulations in one of these approximate game-style physics things, where the system part of the sim the the probabilistic simulator here is that we have uncertainty about a number of things in the scene that the visual stimulus does not completely determine. So, we don't know exactly where the blocks are. We can't perfectly solve the inverse graphics problem from
同时控制其他各种混淆因素;然后我们进行测量,比如让人们在1到7的量表上打分,或者让他们做二选一的强制选择,再把一大群人的结果汇总起来;我们也可以看他们的反应时间。测量对稳定性的分级直觉有很多种方法。这里的Y轴,我画的是一组被试对50或60座不同积木塔的平均稳定性判断。X轴上我画的是模型预测,也就是在这类近似的游戏式物理引擎中跑少量模拟后取平均的结果。这里的"概率"部分在于:场景中有若干因素是视觉刺激没有完全确定的,我们对它们存在不确定性。所以我们并不确切知道积木的位置。我们没法仅凭
便签笔记
44:52
just a single image. We don't exactly know the coefficients of friction. We allow for the possibility that maybe there could be a force perturbation, a little burst of wind, or somebody could bump the table. Under a under various sources of uncertainty like that, and that's really important, we we we can infer that some stacks are are likely to be more or less stable, have more or fewer blocks, or maybe no blocks fall over. And there's a very nice correlation in this case between the model predictions on the X axis and human judgments on the Y axis. The same thing extends to a much um in some sense stranger sort of task, where, like a lot of other experiments in psychology, we study the mind effectively by giving you something that you don't have much experience with if
一张图像就完美地解出逆向图形学问题。我们也不确切知道摩擦系数。我们还允许存在受力扰动的可能性,比如一阵小风,或者有人碰了一下桌子。在这样各种不确定性来源之下——这一点非常重要——我们就能推断出某些积木堆更可能或更不可能稳定,会倒下更多或更少的积木,或者一块积木都不倒。在这个例子里,X轴上的模型预测和Y轴上的人类判断之间有非常好的相关性。同样的做法还可以推广到一个某种意义上更奇怪的任务上:像心理学中很多其他实验一样,我们研究心智的方式,实际上是给你一些你并没有太多经验的东西,
便签笔记
45:35
we're trying to understand what do you do that you didn't just learn in a very simple way. So, the question of will the stack of blocks fall over is one that we many of us have fair amount of experience with. If you play games like Jenga, you have a lot of experience. But I can show you a similar scene, like these tow these tables with red and yellow blocks, and ask you a question which, if you haven't seen me talk about this before, you probably never thought about. What if you bump one of these tables hard enough to knock some of the blocks onto the floor? Is it more likely to be red or yellow blocks on the floor? And people can again make a graded judgment, that's what you see on the Y axis, um and the model can make a graded judgment, too. And the model is almost as predictive of this very strange question as it is predictive of the much
我们想弄清楚的是,你所做的事情当中,有哪些不是靠很简单的方式学来的。比如“这堆积木会不会倒”这个问题,我们很多人都有相当多的经验。如果你玩过叠叠乐(Jenga),你就有很多经验。但我可以给你看一个类似的场景,比如这些摆着红色和黄色积木的桌子,然后问你一个问题——如果你以前没听我讲过,你多半从来没想过这个问题:如果你把其中一张桌子撞得足够猛,把一些积木撞到地上,那地上更可能是红色积木还是黄色积木?人们同样能做出一个分级的判断,这就是你在纵轴上看到的;嗯,模型也能做出分级的判断。而模型对这个非常奇怪的问题的预测力,几乎和它对那个更
便签笔记
46:18
more familiar task of judging just how likely is it to fall. Okay. Now, how does that work? Well, it works because, again, our system isn't training on data, rather it's it's built a model, and it's doing a simulation. And it does just a small number of of very coarse and probabilistic simulations. So, what do I mean by that? I'll be more concrete now. This is a window into one of those simulators simulations applied to one of those scenes. And you can see here I've modeled a small bump of the table. Here, on the same configuration, I've modeled a bigger bump, okay? The simulator probabilistically tries out bumps of a few different sizes and a few different
熟悉的任务——判断它有多大可能倒塌——的预测力一样好。好。那这是怎么做到的呢?还是那句话,它之所以能做到,是因为我们的系统不是在数据上做训练,而是建立了一个模型,并且在做模拟。而且它只做少量非常粗糙的、概率性的模拟。我这么说是什么意思呢?我现在讲得更具体一点。这是一个窗口,让大家看看其中一次模拟应用在那些场景之一时的样子。你可以看到,这里我模拟了轻轻撞一下桌子。这里,在同样的构型上,我模拟了更用力的一撞,好吗?模拟器会以概率的方式,试几种不同力度、几种不同
便签笔记
46:53
angles, kind of semi-randomly. And it in by aggregating the results of of just a few simulations, you can get a pretty reasonable sense to answer the question. You also don't have to run the simulator the simulations very long, right? You can see that after just the first couple of time steps, you already know the answer, at least for many of the scenes, or you may be uncertain. So, we only have to run a small number of simulations for just a small number of time steps, and we could also run them at low resolution. And that's enough to capture what's going on there. Okay. So, that's at least some window into how we use these tools to model the guesses and the bets that your brain makes.
角度的撞击,有点半随机的味道。而只要把区区几次模拟的结果汇总起来,你就能对这个问题得到一个相当合理的判断。你也不需要把模拟跑很久,对吧?你可以看到,仅仅过了头几个时间步,你就已经知道答案了——至少对很多场景是这样;或者你可能仍然不确定。所以我们只需要跑少量的模拟、跑很少的几个时间步,而且还可以用低分辨率来跑。这就足以抓住那里正在发生的事情了。好。这至少让大家看到一点:我们是怎么用这些工具,来给你大脑所做的猜测和下注建模的。
便签笔记
47:27
Um we can do similar kinds of things in uh dynamics tasks, like, for example, this is work from Kevin Smith and colleagues, where you have to predict Oh, I kind of messed it up by going to that. You have to predict whether it's going to hit billiard ball bouncing around is going to hit the red or the green surface first. So, you you can we can make this a little interactive here. So, say ooh if you think it's going to hit red, or green if you get or ah if you think it's going to hit green, okay? Here we go. But better go fast cuz it's about to hit the red. Ooh. Okay. If I hadn't messed it up, you would have said ooh, and then ah.
嗯,我们在动力学任务上也能做类似的事情。比如,这是 Kevin Smith 和同事们的工作:你要预测——哦,我刚切过去把它弄砸了。你要预测的是,一个到处弹跳的台球会先撞到红色面还是绿色面。所以我们可以让这个稍微互动一下。如果你觉得会撞红色,就发“喔”;如果你觉得会撞绿色,就发“啊”,好吗?开始了。不过最好快一点,因为它马上就要撞到红色了。喔。好。如果我刚才没弄砸的话,你们会先“喔”,然后“啊”。
便签笔记
48:04
But I messed it up. But you can simulate that, too. Okay. And notice that when when the when it's first starting, right? It looks like most people are pretty sure, ah, it's probably going to brush the red, and then it just barely misses it. It takes you a little while before you become confident about it's usually only about there before, you know, you you become confident that it's going to hit the green. So, people are able to simulate, but usually only about one or two bounces at most in open things like this. And we can capture that by assuming that they do a sort of dynamic version of what we were showing, again, with some noise or uncertainty in the trajectory of the ball in its position and velocity, and especially uncertainty about bounces in the because people can unless you're a really experienced billiards player, you're you're pretty uncertain about how to extrapolate bounces, okay? And that model, this is just showing that cone of uncertainty that comes from running these probabilistic simulation, it's able if we just add up the results of a
但我弄砸了。不过这个你也能模拟。好。注意,一开始的时候,对吧?看起来大多数人都相当确定——啊,它八成要擦到红色——然后它就差那么一点点没碰到。你要过一会儿才会变得确信,通常也就到那个位置左右,你才会确信它会撞到绿色。所以人是能做模拟的,但在这种开放的场景里,通常最多也就模拟一到两次弹跳。我们可以这样来刻画这一点:假设人们做的是我们刚才展示的那种动态版本,同样在球的轨迹、位置和速度上带有一些噪声或不确定性,尤其是对弹跳的不确定性——因为除非你是非常有经验的台球选手,否则你对弹跳之后该怎么外推是相当没底的,对吧?而那个模型——这里展示的就是跑这些概率模拟所产生的那个不确定性锥——只要我们把
便签笔记
49:01
few of these simulations, it's able to capture the the tendency that people tend to say that's probably red, uh then I'm not sure, then it's probably green. More quantitatively in this study, we had 100 different scenes. I'm showing you four of them here, and I'm showing you dynamic trajectories. The plot first on top, people, and on the bottom, model, and showing how the graded confidence changes over the scene. And you can see quite a lot of interesting structure to the data. Some scenes, people are very confident of one, and then they switch to the other. Sometimes they're just not sure for a long while, and then there's a sudden um resolution. Sometimes there's multiple switches, and it's really striking how much this model is able to capture it. Um in work that I was very privileged to collaborate on with Ed Vul as who was long ago in in my lab, he's now at UCSD and Amazon, we worked with infant researchers Anna O'Teaglas and Luca Bonatti and others to study intuitive
其中几次模拟的结果加总起来,它就能刻画出人们的那种倾向:先说八成是红色,然后是我不确定,然后是八成是绿色。更定量地说,在这项研究里我们有 100 个不同的场景。我这里给大家看其中四个,展示的是动态的轨迹。上面那张图是人的数据,下面是模型,展示分级的信心如何随着场景的推进而变化。你可以看到数据里有相当多有意思的结构。有些场景里,人们对一个选项非常确信,然后又切换到另一个。有时候他们很长一段时间就是不确定,然后突然就有了结论。有时候会出现多次切换。这个模型能刻画到这种程度,真的很惊人。嗯,在一项我很荣幸能参与合作的工作里——合作者 Ed Vul 很久以前在我的实验室,他现在在 UCSD 和亚马逊——我们和婴儿研究者 Anna O'Teaglas、Luca Bonatti 等人合作,研究婴儿的直觉
便签笔记
49:51
physics in infants. So, 12-month-olds were asked to extrapolate were shown scenes like these little gumball lottery machines here. And um they saw the objects bouncing around for a little while, and then there was a period of occlusion and one of the objects appeared. And what was varied across different conditions of the experiment was whether there was more or less objects of three or one of of the color that ultimately appeared, where those objects were at the time of occlusion and how long the time of occlusion was. And a very simple probabilistic physics simulation model is able to to make different predictions as a function of those variables, and it captures integrated and and fairly compelling quantitative way given the very strong limitations of how hard it is to get quantitative behavioral data from infants, but it captures infants uncertain predictions in the sense that um you know, the classic method that Spelke and many others have pioneered for studying the infant mind is looking at violation of expectation or other looking time measures. People and babies, too, look longer when you show them things that are surprising.
物理。12 个月大的婴儿被要求做外推,他们看到的是像这些小口香糖抽奖机一样的场景。嗯,他们看着那些物体弹跳了一会儿,然后有一段遮挡时间,接着其中一个物体出现了。实验的不同条件之间所变化的,是最终出现的那种颜色的物体是有三个还是只有一个、这些物体在遮挡发生时处在什么位置,以及遮挡持续了多长时间。一个非常简单的概率物理模拟模型,就能根据这些变量做出不同的预测;而且考虑到从婴儿身上获取定量行为数据有多难、限制有多强,它以一种整合的、相当有说服力的定量方式,刻画了婴儿那种不确定的预测。也就是说,你知道,Spelke 和许多其他人开创的研究婴儿心智的经典方法,是看“期望违背”或其他注视时长指标。人是这样,婴儿也是这样:当你给他们看令他们意外的东西时,他们会看得更久。
便签笔记
50:53
And looking time here is predicted in a quantitative way by inverse probability in one of these physics models. In some even some of the most recent work, this is in our lab from Sam Shayet and uh uh Tracy Mills and Sam Shayet, um we can study how people are able to predict in a in just a simple sequence of a moving dot, something that goes well beyond physics, but can have much more structure, almost various kinds of algorithmic or program structure. So, just the ability to predict what's coming next is really quite a powerful way to study the mind's models and how we use them to make guesses and bets in the world. And these are things that are distinctive to humans as Sam and others have shown working with macaque monkeys and um Ed developed significantly from childhood to adulthood, although human children appear at least so far to look a lot more like adults than like macaques. Um in work I mentioned before from Kelsey Allen and Kevin Smith and colleagues or I alluded to when I showed you the kid with the shovel, I'll for for reasons of time I'll basically skip this, but I'll just refer you to their
而这里的注视时长,可以用这类物理模型中的逆概率定量地预测出来。在一些最新的工作里——这是我们实验室 Sam Cheyette 和 Tracy Mills 的工作——我们可以研究人们如何在一个简单的移动圆点序列中做预测,这远远超出了物理的范畴,但可以具有多得多的结构,几乎是各种算法性或程序性的结构。所以,仅仅是预测接下来会发生什么的能力,就是研究心智的模型、以及我们如何用这些模型在世界中做猜测和下注的一条非常有力的途径。而正如 Sam 和其他人在猕猴身上的工作所显示的,这些是人类所特有的;嗯,从儿童到成人还有显著的发展,尽管人类儿童至少目前看来更像成人,而不像猕猴。嗯,我之前提到的 Kelsey Allen、Kevin Smith 和同事们的工作,也就是我给大家看那个拿铲子的小孩时暗示的那些——因为时间关系我基本上要跳过——但我会请大家去看他们的
便签笔记
09工具游戏与直觉心理学
51:57
PNAS and forthcoming paper, but to study how people use tools and use their intuitive physics flexibly for tools, Kelsey and Kevin came up with this cool video game, which we called the virtual tools game, and you can see all you do in this game is your goal is always to get the red object into the green place, the green location, stably. And you you act by choosing an object, a tool, and just dropping it into the scene, and then physics unfolds and you see what happens. Okay. Um people really like playing this game. You can actually play it yourself at the MIT Museum, it's part of their new AI exhibit in the new MIT Museum in Kendall Square. And like a lot of other kinds of physical problem solving, but very much unlike what's often called reinforcement learning in machine learning, people this is a kind of a reinforcement learning task. You just act and you get
PNAS 论文和即将发表的那篇论文。为了研究人们如何使用工具、如何灵活运用直觉物理来使用工具,Kelsey 和 Kevin 想出了这个很酷的电子游戏,我们把它叫做“虚拟工具游戏”。你可以看到,在这个游戏里你要做的就是:目标始终是把红色物体稳稳地送到绿色的地方、绿色的位置上。而你的行动方式,就是选一个物体、一个工具,把它扔进场景里,然后物理过程就展开,你看看会发生什么。好。嗯,人们真的很喜欢玩这个游戏。你其实可以自己去玩,在 MIT 博物馆就能玩到,它是 Kendall Square 新 MIT 博物馆里新 AI 展的一部分。和许多其他类型的物理问题解决一样,但又非常不同于机器学习里常说的强化学习——这其实算是一种强化学习任务:你行动,然后得到
便签笔记
52:45
some feedback, and then you can act again, so you learn from trial and error. But across many different levels of this game, people learn extremely quickly. They learn in, you know, sometimes just two or three trials, rarely more than five or 10. And we model that by basically putting that simulator in the loop and having a process which which imagines possible ways of solving the problem drawing from just a a very simple prior to try to cause collisions or object interactions, imagines what might happen. If you come up with a simulation with a few tries in your head that is likely to solve the problem in the world, then you try it out, otherwise you sort of repeat and resample at different part of the space. And that's able to capture people's learning curves quite quite precisely. So, these are across 20 different levels in the physics game, the red and the blue curves, the blue curve shows human's cumulative probability of success as a function of trials, and the red curve shows the model.
一些反馈,然后你可以再行动,所以你是在从试错中学习。但在这个游戏的许多不同关卡里,人们学得极快。他们有时候只用两三次尝试就学会了,很少超过五次或十次。我们对此建模的方式,基本上就是把那个模拟器放进循环里,让一个过程去想象解决问题的各种可能方式——从一个非常简单的先验出发,试着制造碰撞或物体之间的相互作用——去想象可能会发生什么。如果你在脑子里试了几次,得出一个很可能在真实世界里解决问题的模拟,那你就去试;否则你就重复,并在空间的不同部分重新采样。这能相当精确地刻画人们的学习曲线。这是这个物理游戏里 20 个不同关卡的结果,红色和蓝色两条曲线:蓝色曲线表示人类成功的累积概率随尝试次数的变化,红色曲线表示模型。
便签笔记
53:38
So, the model, this idea of just trying out a few ideas in your head, and then when one of them seems to work trying in the real world, and if that fails, repeat, um captures both the problems that people find harder easier, both their ultimate level of success, sometimes people don't always solve it, but also the relative rate of learning, how many trials it takes you to figure it out. Okay. Um mostly that's all I'm going to say about um that's all I'm going to say about intuitive physics, and it's mostly all I'm going to say about common sense. I want to save the last few minutes to talk about language. Um I'll just point to a whole other line of work, which was done in part in collaboration with Rebecca Sax here, one of my great friends and MIT colleagues, along with a number of others, um folks like Julian Jara-Ettinger, who was a grad student with me, and Laura Schulz, and Laura also at MIT played a role in a lot of these things. Um if if you want a a summary of this sort of parallel research program of using
所以这个模型——就是先在脑子里试几个想法,当其中一个看起来可行时就到真实世界里去试,如果失败就重复——嗯,既能刻画人们觉得哪些问题更难、哪些更容易,也能刻画他们最终的成功水平(有时候人们并不总能解出来),还能刻画相对的学习速度,也就是要花多少次尝试你才能想明白。好。嗯,关于直觉物理,我基本上就讲这么多了;关于常识,我也基本上讲完了。我想把最后几分钟留给语言。嗯,我只提一下另外一整条研究线,它有一部分是和这里的 Rebecca Saxe 合作完成的——她是我最好的朋友之一,也是 MIT 的同事——还有其他一些人,比如曾经跟我读研究生的 Julian Jara-Ettinger,以及 Laura Schulz,Laura 也在 MIT,她在这些事情里起了很大作用。嗯,如果你想要这条平行研究纲领的一个综述,也就是用
便签笔记
54:31
the idea of probabilistic inference on top of programs that capture the causal structure of the physical world and in this case how social agents interact with it, check out another really nice Trends in Cognitive Science piece by Julian Jara-Ettinger and colleagues called the Naive Utility Calculus. This was published back in 2016. Julian, I should say by the way, is now a faculty member at Yale in psychology. And um basically I'll just point you to this. It's it's some of my favorite work that I've been involved in. Um but what we what we are doing here is we are also using the idea of programs, but to describe the the world, but now it's programs which take as input percepts and goals and plan actions. So, these are models, they also do effectively kind of rational uh decision making. They goals are sources of reward for an agent, and actions have costs, and in this work the idea is that we we are modeling agents as making plans, sequences of actions to mid to maximize their expected utility, choose actions that support their goals, and so on, given their beliefs about the world. But the key is that these are not models of human agents, these are models of human agents models of human agents, or
在程序之上做概率推断这个想法——用程序来刻画物理世界的因果结构,而在这里则是刻画社会主体如何与之互动——你可以去看 Julian Jara-Ettinger 和同事们发在《认知科学趋势》上的另一篇很好的文章,叫《朴素效用计算》(Naive Utility Calculus),这是 2016 年发表的。顺便说一句,Julian 现在是耶鲁大学心理学系的教师。嗯,基本上我就把这个推荐给大家。这是我参与过的最喜欢的工作之一。嗯,我们在这里做的,同样是用程序这个想法来描述世界,但现在这些程序是以知觉和目标作为输入,并规划行动。所以这些模型实际上也在做某种理性决策:目标是主体的奖励来源,而行动是有代价的;在这项工作里,我们把主体建模为在制定计划、也就是一系列行动,以最大化其期望效用,选择那些支持其目标的行动,等等,这一切都取决于它们对世界的信念。但关键在于,这些不是人类主体的模型,而是人类主体对人类主体(或者
便签笔记
55:42
perhaps non-human agents, too. Um like, for example, in the work of Kylie Hamlin and others with the little balls or the classic work of Heider and Simmel or Gergely and Csibra, if you know all these really beautiful studies going back many years of studying how human adults and even babies can make sense of little shapes moving around if they follow effectively the principles of efficient action planning. Okay. So, in in the in this line of research, what we study is with similar kinds of stimuli, for example, um in the food truck stimuli, which which Chris Baker and Julian did with Rebecca and me, looking at an agent moving around in the world, in this case um going to looking for one of several food trucks which park on campus in one of those two yellow parking spots. Could be the the Korean truck or the Lebanese truck or the Mexican truck. Park in different spots on different days. But when an agent comes out, goes around the building where they see the the Lebanese truck and turns back to the Korean truck, and you ask, "What is that agent's favorite food? Korean, Lebanese, or Mexican?" People infer that actually it's Mexican. That's the bar on the right under desires, the M. Mexican was her favorite, right? Even though that's
也可能是非人类主体)的模型的模型。嗯,比如 Kiley Hamlin 等人用小球做的工作,或者 Heider 和 Simmel 的经典工作,或者 Gergely 和 Csibra 的工作——如果你了解这些多年来非常漂亮的研究,它们研究的是人类成人乃至婴儿,如何理解那些到处移动的小形状,只要这些形状实际上遵循高效行动规划的原则。好。所以在这条研究线里,我们用类似的刺激材料来研究。比如在“餐车”这套刺激材料里——这是 Chris Baker 和 Julian 与 Rebecca 还有我一起做的——看一个主体在世界中移动,在这个例子里,它是要去找几辆餐车之一,这些餐车停在校园里那两个黄色停车位中的某一个。可能是韩国餐车、黎巴嫩餐车或者墨西哥餐车,不同的日子停在不同的位置。但当一个主体出来,绕过大楼,在那儿看到了黎巴嫩餐车,然后又折回去找韩国餐车,你问:“这个主体最喜欢的食物是什么?韩国菜、黎巴嫩菜还是墨西哥菜?”人们推断出,其实是墨西哥菜。那就是右边“欲望”那一栏里的 M 那根柱子。墨西哥菜是她的最爱,对吧?尽管那
便签笔记
56:52
not the one that was actually present at all in the scene, but the best way to explain why she moved on an efficient path not to her actual goal, but to what she thought might have been there. We also ask people, "What were the agent's initial beliefs?" And they say, as our model predicts, that she probably thought Mexican was there. But that So, that was her initial belief, um and that's what she wanted, but that's why she went to the other side to look. But the reason she went back is because, well, it wasn't there. She only discovered that when she got there. So, this, for example, across many different um stimuli again, where we vary in a controlled way the different sort of stimuli, is able to quantitatively predict people's inferences about agents' desires and beliefs and also how they interact, because in scenes like this one, you can only make sense of her action if you posit a goal that is not present and the false initial belief that it was likely to have been present.
根本就不是场景里实际出现过的那一家。但要解释她为什么走了一条高效的路径——不是走向她真正的目标,而是走向她以为可能在那儿的东西——这是最好的解释。我们还问人们:“这个主体最初的信念是什么?”他们说,正如我们的模型所预测的,她大概以为墨西哥餐车在那儿。那就是她最初的信念,嗯,那也是她想要的,所以她才绕到另一边去看。但她之所以又折回来,是因为,嗯,它不在那儿。她走到那儿才发现这一点。所以,举例来说,在许多不同的刺激材料上——我们同样以受控的方式变化不同的刺激——这个模型能够定量地预测人们对主体的欲望和信念的推断,以及它们如何相互作用;因为在像这样的场景里,只有当你假定一个并不在场的目标、再加上一个“它很可能在场”的错误初始信念时,你才能理解她的行动。
便签笔记
10语言如何接入:从词模型到世界模型
57:45
So, these are just again a taste of how this toolkit of in this case sort of doubly Bayesian or rational approximately approximate rational probabilistic inferences about agents' inferences about other approximately rational agents is able to capture core aspects of our common sense world models. Now, how does language come into this? Well, again, I I don't really have very much time. In fact, I think I'm out of time. But um I owe it to both the beginning of my talk, and I hope if you'll give me just another five, six minutes to just basically advertise this very nice paper by two of the students that I started with um in in when I mentioned the original acknowledgements.
所以,这些同样只是让大家尝一口这套工具箱:在这个例子里,是一种双重贝叶斯的、或者说近似理性的概率推断——关于近似理性的主体对其他近似理性主体所做推断的推断——它能够刻画我们常识世界模型的核心方面。那么语言是怎么进来的呢?嗯,还是那句话,我真的没有太多时间了。事实上,我觉得我已经超时了。但我在开场时欠了大家一笔,所以希望你们再给我五六分钟,让我基本上给大家推荐一篇非常好的论文,作者是我在最初致谢里提到的、一开始就跟着我的两位学生。
便签笔记
58:23
Well, it's actually by a num- a number of people, but the co-first authors are Leo Wong and Gabe Grand, and Alex Lev also played a really key role in this work along with some of my faculty colleagues like Jacob Andreas and Vikash Mansinghka and Noah Goodman. But what you can see in this paper, and this paper expresses we we when we were working on this paper, which you can find on archive, it's currently undergoing revision for a couple of different publication targets. It's it's very long and it's being broken into pieces, but you can I I urge you to check out at least the first part and read more of it if you like. Um which we when we were working on this paper, we called it the white paper, because it's really it's it's a road map for a research program. It's not a set of research results, but it tries to show you the path for connecting modern language models and the ways in which capturing statistics of language can build on and amplify and extend and really become much more powerful when you ground them in the kinds of common sense knowledge that I was just showing you via the these mechanisms of probabilistic programs.
嗯,其实作者有好些人,但共同第一作者是 Leo Wong 和 Gabe Grand,Alex Lew 在这项工作里也起了非常关键的作用,还有我的一些教师同事,比如 Jacob Andreas、Vikash Mansinghka 和 Noah Goodman。你在这篇论文里能看到——这篇论文你可以在 arXiv 上找到,它目前正在为几个不同的发表目标做修改;它非常长,正在被拆成几部分,但我强烈建议你至少看看第一部分,如果喜欢就多读一些——嗯,我们在写这篇论文的时候,把它叫做“白皮书”,因为它其实是一个研究纲领的路线图。它不是一组研究结果,而是试图给你指出一条路径:如何把现代语言模型,以及“捕捉语言统计规律”这件事,与我刚才通过概率程序这些机制展示给大家的那类常识知识连接起来——当你把它们接地到这类常识知识上时,它们能够在此之上被放大、被扩展,变得强大得多。
便签笔记
59:25
So, in the paper, the idea is to consider all the many ways that we use language to inform and structure our thinking, whether it's in intuitive physics or intuitive psychology or other sorts of domains. And also all the ways that language can extend our thinking, the ways we can learn new concepts, either explicitly by having people explain things to us or implicitly by seeing how people use language in certain patterns and sentences, and even build up whole new intuitive theories that mostly come to us from people telling us things. That's we're trying to explain how can you do that? And how can we use this in some sense to make sense of what are what are the current kinds of intelligence see in large language models and what's missing, and what's a more human-like way forward. And the key idea in the paper is to combine two things. One is um this idea of what we've called the probabilistic language of thought. So this is this is the paper from the the recent concepts book, well from 10 years ago, recent compared to the older concepts book by Laurence and Margolis, but the new concepts book um that Noah Goodman, Toby Gerstenberg, and I wrote, which sort of explains how a certain how
所以在这篇论文里,思路是去考虑我们用语言来告知和构建思维的所有那些方式,不管是在直觉物理、直觉心理学还是其他领域里。还有语言能够扩展我们思维的所有那些方式:我们如何学到新概念,既可以是显式地——别人向我们解释——也可以是隐式地,通过看到人们如何以某些模式和句子来使用语言;甚至建立起全新的直觉理论,而这些理论大多是别人告诉我们才有的。这就是我们想要解释的:你怎么能做到这一点?以及我们怎么能在某种意义上用它来理解,我们在大型语言模型里看到的当前这类智能到底是什么、缺了什么,以及一条更像人的前进道路是什么样的。这篇论文的核心想法是把两样东西结合起来。一个是我们所说的“概率思维语言”。这是那本新《概念》书里的一篇——嗯,那是十年前的文章了,相对于 Laurence 和 Margolis 编的那本更老的《概念》书来说算“新”的——那本新《概念》书里 Noah Goodman、Toby Gerstenberg 和我写的那一章,大致解释了我们如何
便签笔记
60:28
we can use a probabilistic programming language. If you want to learn more about this this these probabilistic programming languages as reflected in that chapter, you could check out the probmods.org webbook, which is just shown here. This is an example of the the probabilistic programming language Church. But what we show in this webbook is how you can use this single programming language. It's kind of a version of a list, but in ways that we use Lisp to capture the problem the structured probabilistic models like for intuitive physics or intuitive psychology and so on. And we we we um this we have a single probabilistic programming language which can capture basically all the cognitive models that you've seen, all these kinds of frameworks for guessing and betting or making rational expected utility approximate inferences, okay? And suggest that that can provide a unifying substrate for all these common sense models.
使用一种概率编程语言。如果你想更多了解那一章里体现的这些概率编程语言,可以去看 probmods.org 这个网络电子书,就是这里展示的这个。这是概率编程语言 Church 的一个例子。但我们在这本网络书里展示的是,你如何用这一种编程语言——它算是 Lisp 的一个版本——以我们使用 Lisp 的那种方式,去刻画那些结构化的概率模型,比如直觉物理或直觉心理学等等。而我们有一种统一的概率编程语言,基本上能刻画你们看到的所有认知模型,所有这些用来猜测和下注、或者做理性期望效用近似推断的框架,好吗?并且提出,这可以为所有这些常识模型提供一个统一的底层基质。
便签笔记
61:17
And you combine that with the idea that we can see in in recent large language models, which is that they aren't just models of natural language, they're also trained on code, source code in many different programming languages, and source code of as as as is often pointed out but is I think not fully appreciated, source code, right, is programming languages or programs that are designed to be read and written by humans and not just machine executable or read by machines. So while people often point out that programming languages are very different from natural languages, source code, whether it's languages like Python or JavaScript or Lisp, um is often written in a way that is still rather language-like, somewhat in its syntax, in terms of like hierarchical phrase structure, but also especially in terms of the way we name variables or name functions, as well as of course comment code. So in a sense, you could say the idea of this paper is a sort of source code as the language a metaphor as the language of thought, which mixes together natural language and programs in a probabilistic
然后你把这一点,和我们在近期大型语言模型中看到的一个想法结合起来:它们不只是自然语言的模型,它们还在代码上训练过,在许多不同编程语言的源代码上训练过。而源代码——正如人们常指出、但我认为并没有被充分体会到的那样——编程语言、或者说程序,是被设计成供人类阅读和书写的,而不只是机器可执行、供机器读取的。所以,虽然人们常常指出编程语言和自然语言非常不同,但源代码,不管是 Python、JavaScript 还是 Lisp 这样的语言,写出来往往仍然相当有语言的味道:在语法上有一点,比如层级化的短语结构;但尤其体现在我们给变量命名、给函数命名的方式上,当然还有给代码写注释。所以在某种意义上,你可以说这篇论文的想法是一种“源代码作为思维语言”的隐喻,它把自然语言和程序混合在一种概率
便签笔记
62:19
programming language. And the paper shows how you can use actually you know, effectively a kind of a what is not even a state of the art language model, but an early um code LLM as it's called, the OpenAI's Codex model, to to basically define a probabilistic translation mapping from and to, but mostly in this paper from, natural language to one of these probabilistic programming languages, Church in this case, and then applies it to just a number of different domains. I'll just show you one application just to close the loop back to intuitive physics. But for example, if you remember this, right? So as I was saying here, we use natural language to to give queries and to describe kinds of intuitive physical reasoning that maybe you haven't done before. So we can basically now make an end-to-end um model of language-conditioned, language-informed intuitive physics by taking in this case a 2D physics engine that can capture the tabletop physics that I showed you
编程语言里。而这篇论文展示了,你其实可以用一个——嗯,甚至都算不上最先进的语言模型,而是一个早期的所谓“代码 LLM”,也就是 OpenAI 的 Codex 模型——来定义一个概率性的翻译映射,双向的,但这篇论文里主要是从自然语言映射到这些概率编程语言之一(这里是 Church),然后把它应用到许多不同的领域。我只给大家看一个应用,好把话题绕回直觉物理。比如说,如果你还记得这个的话,对吧?就像我刚才说的,我们用自然语言来提出询问、来描述那种你可能从来没做过的直觉物理推理。所以我们现在基本上可以做一个端到端的、以语言为条件、由语言提供信息的直觉物理模型:在这里我们用一个 2D 物理引擎,它能刻画我刚才给大家看的那种桌面物理,
便签笔记
63:21
before, and use a code LLM to just capture the idea that in effectively that that we can think about contextual meaning in language as a translation from a description of a world like a tabletop scene that says there's one tall stack of red blocks and there are two short stacks of yellow blocks. And the question again is if I bump the table hard enough to knock some of them on the floor, will it be more red or yellow blocks? So the model is able to translate that into this internal code and run a simulation like this. This is not seen, this is imagined. So this is a view into the the mental simulation that the model runs given that language. Or we can describe a more complex scene, and we can describe it and then we can imagine more complex corresponding more complex physical scenes and simulate what happens there. And but basically it's it's just a modular composition of the same things, a language system, a putatively like the one in the human brain that translates from what we say and hear into some mental language which can
然后用一个代码 LLM 来抓住这样一个想法:我们实际上可以把语言中的语境含义,看成是一种翻译——从对世界的描述翻译过去,比如一个桌面场景的描述说“有一摞高高的红色积木,还有两摞矮的黄色积木”。而问题还是那个:如果我把桌子撞得足够猛,把一些积木撞到地上,地上会是红色积木更多还是黄色积木更多?模型能够把这句话翻译成这种内部代码,并跑一个像这样的模拟。这不是看到的,这是想象出来的。所以这是一个窗口,让我们看到模型在给定那段语言之后所运行的心理模拟。或者我们可以描述一个更复杂的场景,描述完之后就能想象出相应的更复杂的物理场景,并模拟那里会发生什么。但基本上,它就只是同样这些东西的模块化组合:一个语言系统,大概类似人脑里的那个,把我们所说和所听的翻译成某种心理语言,而这种语言可以
便签笔记
64:21
figures the brain's physics engine in this case, runs a small number of these probabilistic simulations, and that alone is or or alone or together is enough to give a quantitatively predictive model, much like not quite but almost as with the same quantitatively predictive power as you saw before when we were just setting up the model for itself. So this is just a taste of of some of the the work that I think is where that is ongoing by the students I mentioned and a number of others. I should say by the way, this this particular paper here was done with Leo and Gabe, but the experimental and modeling studies were done with Sed John, who's shown up there. He's also a student at MIT. And you you know, we and others are pursuing analogous studies like in a theory of mind and social cognition domains, and you know, I just say mostly stay tuned. And um the last bit is to say, cuz I I I I can't help but say this just a little bit towards where we're going, right? Is um I do think that there's a there's a long history in cognitive science and
配置大脑的物理引擎——在这个例子里是这样——然后运行少量这类概率模拟;而单靠这一点,或者说单独地、或者合在一起,就足以给出一个定量上有预测力的模型,虽然不完全一样,但几乎和你们之前看到的、我们自己直接把模型搭好时的定量预测力一样好。所以这只是让大家尝一口我提到的那些学生和其他一些人正在进行的部分工作。顺便说一句,这篇论文是和 Leo 与 Gabe 一起做的,但实验和建模研究是和 Sed John 一起做的,就是上面显示的那位,他也是 MIT 的学生。你知道,我们和其他人也在心智理论和社会认知这些领域推进类似的研究,我基本上想说的就是:敬请期待。嗯,最后一点,因为我实在忍不住要稍微讲一下我们要往哪里去,对吧——嗯,我确实认为,在认知科学以及
便签笔记
65:21
many and all the related fields, linguistics, philosophy, certainly AI, of asking what is meaning, right? What does it mean to understand language, and what is a thought that's conveyed in language. And in the in the engineering sketch I'm showing you, there's a step towards what's actually a theory of meaning. I mean, if I venture to say that. It's just a step and there's a lot to be filled in. But an idea that again has a long history in different ways, but I think now we're in a position to capture it in an engineering sense, which is to think of meaning not in in in its in abstract form, but contextualize meaning. Like what does a piece of language mean in context? A word, a phrase, a sentence in the context of the rest of the sentences or the other sentences that have come before, or the other other other ways you will build up a model of the current situation, like your perceptual system, but something like a joint distribution between natural language and code, but in a probabilistic programming language of thought. So that joint distribution, which is the key object that you can capture in the these code LLMs combined with with probabilistic inference, and you put those pieces together, and you have an at least the some of the building blocks needed to unify what are
所有相关领域——语言学、哲学,当然还有人工智能——都有一段很长的历史,在追问什么是意义,对吧?理解语言意味着什么?用语言传达的思想又是什么?而在我给大家展示的这个工程草图里,有朝着一个真正的意义理论迈出的一步。我是说,如果我斗胆这么讲的话。这只是一步,还有很多要填补。但这个想法本身,以不同的方式也有很长的历史,只是我认为我们现在有条件从工程意义上把它抓住了。这个想法就是:不要把意义想成抽象形式的意义,而要想成语境化的意义。比如,一段语言在语境中意味着什么?一个词、一个短语、一个句子,在其余句子、或者说在它之前出现的那些句子的语境中,或者在你用来建立当前情境模型的其他方式(比如你的知觉系统)的语境中,意味着什么?答案是某种类似自然语言与代码之间的联合分布,但是在一种概率的思维编程语言里。所以那个联合分布,就是你能用这些代码 LLM 加上概率推断抓住的关键对象;你把这些部件拼在一起,你就至少有了一些必需的构件,去统一那些
便签笔记
66:28
traditionally a number of different takes on meaning in all these fields. So the idea that what is meaning may is it something like informal semantics, like compositional logical construction of thought? Is it something like distributional word embeddings, that's again a not just a modern LLM idea but a very old idea in psychology as well as natural language processing. Those are two interesting views. Two others, like meaning is about grounding in the world in the embodiment thesis. Or meaning is about pragmatics and context. Well, in a sense, the the probabilistic programming and probabilistic language of thought stuff I showed you kind of gives you the more compositional and grounded aspects of meaning. The code LLM view gives you the distributional statistical aspects of it, and at least some approximations to the pragmatic or contextual aspects of meaning. And you put these things together, and it it's again, it's it's building blocks for how we can really think about what language is really about and language use, and how language builds on and then very much amplifies and extends our thinking.
在所有这些领域里传统上对意义的若干种不同看法。比如,意义是不是某种形式语义学式的东西,像是思想的组合式逻辑构造?还是某种分布式词嵌入?这不仅是现代 LLM 的想法,在心理学以及自然语言处理里也是一个很老的想法。这是两种有意思的观点。另外两种:意义在于对世界的接地,也就是具身论题;或者意义在于语用和语境。嗯,在某种意义上,我给大家展示的概率编程和概率思维语言那部分,给了你意义中更偏组合性和接地的那些方面;代码 LLM 的视角给了你分布统计的方面,以及至少对意义中语用或语境方面的一些近似。你把这些东西放到一起——再说一次,这是一些构件,让我们能真正去思考语言到底是关于什么的、语言使用是怎么回事,以及语言如何在我们的思维之上建立、并极大地放大和扩展我们的思维。
便签笔记
11结论:理解、意义与开放的 AI
67:29
So many open questions, and I'll just end with some conclusions. Um what is thinking? It's not just a mystery, although we are still far from capturing thinking in all its forms computationally, but the fundamental idea that intelligence is about making good guesses and good bets, something that is fundamentally also about um the model that our minds have of the world and ourselves in it, and then using that in a rational way to guide what we do next and what we think about next. Okay. This is something that we are starting to have tools to be able to capture, and this gives us a way, especially when language comes into the picture, to really understand what it is we're all here for, right? It's about not just knowledge but understanding. It's about the deep and broad systems that we build individually and collectively, that we build through our own experience and through language, to sustain and grow knowledge in an ever-changing world and a world that
所以还有很多开放的问题,我就用一些结论来收尾。嗯,什么是思维?它不只是一个谜——尽管我们离在计算上刻画思维的所有形式还很远——但有一个根本的想法:智能是关于做出好的猜测、下好的赌注,而这从根本上也关乎我们的心智对世界、以及对身处其中的我们自己所拥有的那个模型,然后以理性的方式用它来指导我们接下来做什么、接下来想什么。好。这是我们开始拥有工具去把握的东西;而这给了我们一条路,尤其是当语言进入画面之后,去真正理解我们大家究竟为什么在这里,对吧?这不只是关于知识,而是关于理解。它关乎我们个体地、也集体地建立起来的那些深刻而广博的系统——我们通过自身经验、也通过语言建立起来的系统——用来在一个不断变化的世界里维持并增长知识,而这个世界
便签笔记
68:26
we're changing. So the tools that I've shown you are part of the the the what I think are the building blocks to actually try to understand this. Um on the technical side, there's what you could call this neural symbolic probabilistic synthesis that modern probabilistic programming languages and various kinds of neural networks and language models can bring together. But, you know, we're still very much at the beginning of trying to put all of these things together. So if you're interested in this, um if any of this excites or inspires you, that's what that's my goal here, and I hope you'll consider, you know, whether it's collaboratively with us in your own work, um or in whatever way really most um inspires you to think about how to build this kind of intelligence, either on the on the science side or on the engineering side, but especially to people interested in AI, I think this is necessary. Unless we want to completely offload the AI future to a small number of uh very well-resourced companies and individuals, if we want to see a future for artificial intelligence that is one that is more open and democratic, that
我们正在改变。所以我给大家展示的这些工具,是我认为真正理解这件事所需要的基础构件之一。呃,在技术层面,有一种可以称之为神经-符号-概率综合的东西,是现代概率编程语言、各种神经网络和语言模型能够结合到一起的产物。但是,你知道,我们仍然处在把这些东西整合起来的非常早期的阶段。所以如果你对此感兴趣,呃,如果这些内容中有任何一点让你兴奋或受到启发,那就是我在这里的目标,我希望你能考虑一下,你知道,无论是和我们合作,还是在你自己的工作中,呃,或者以任何最能激发你去思考如何构建这种智能的方式,不管是在科学层面还是工程层面,尤其是对那些对 AI 感兴趣的人,我认为这是必要的。除非我们想把 AI 的未来完全交给少数几家资源极其雄厚的公司和个人,如果我们希望看到人工智能的未来是更开放、更民主的,是
便签笔记
69:32
is one that is more grounded in and informing science. Um, I think we need we need tools and a mindset in order to do it. I think there's too much at stake. So, uh, join us. Thank you.
更扎根于科学、也能反哺科学的,呃,我认为我们需要工具,也需要一种心态才能做到这一点。我觉得这件事关系太重大了。所以,呃,加入我们吧。谢谢大家。
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Tenenbaum 认为智能的本质不是语言统计预测,而是大脑用世界模型"做好的猜测与下注"(概率推断 + 期望效用决策);语言是在这种更古老的常识思维之上建立并被大幅放大的,因此通向人类式 AI 的路径应是"概率程序 + 神经网络 + 语言模型"的神经-符号-概率综合,而非单纯扩大 LLM 规模。

核心要点

  • LLM 的智能既惊人又脆弱,暴露出"语言≠思维"。 ChatGPT 能为"书+四个网球+钉子+酒杯+口香糖+意面"设计看似合理的堆叠方案,但在 Tyler Brooke-Wilson 让它自己出题(如"艺术家的攀升":用画笔、画布、颜料、调色板、水杯搭 12 英寸高、站立 45 秒的塔)再自己解答时,答案语言流畅却物理上完全不可行——即使提示词明确允许它说"无法解决",它依然创造性地给出错误方案。
  • 自回归下一词预测没有理由能完成任意计算。 Shannon 式的"预测语言 = 智能完备"假设把思维等同于文本前后之间的潜在预测结构;但有限架构、有限数据、逐 token 生成的模型对许多有明确答案的问题就是算不出来。链式思考(chain of thought)等技巧只对部分问题有效(讲者举例:对上面那道题有效,对下面那道无效),RLHF 也只是把预测目标换成"人会点赞的东西"。
  • 神经科学证据:语言网络与思维在大脑中是分离的。 Ev Fedorenko 的 fMRI 研究显示,语言网络在说、听、读、写、盲文、甚至克林贡语/多斯拉克语等人造语言中活动方式一致;GPT-2 级别的模型能对该网络的神经响应做出目前最好的定量预测——但这个网络在解代数方程、理解代码、逻辑推理时并不参与,脑损伤也可以只损害语言而不损害符号思维。没人认为 GPT-2 是通用智能模型,它捕捉的恰恰只是语言本身。
  • 发展证据:思维先于语言,且以极少数据学会语言。 Spelke 等人的婴儿研究表明,婴儿在学会语言前已有关于物体、施动者、目标与计划的思维系统。数据量对比:最新 LLM 训练约 1 万亿词,而儿童每年约听 5 百万词、20 年约 1 亿词——正是先验的常识知识让人类能用小几个数量级的数据学会语言。
  • "家庭手语"证明是心创造语言,而非语言创造心。 Goldin-Meadow 研究显示,没有语言输入的聋童会自发构造具有语言特征的原型交流系统;尼加拉瓜手语则是把若干家庭手语者聚集到孤儿院后,在两代人之内从零创造出一门完整语言,进而形成文化。
  • 演化证据:10 万个神经元就有"猜测与下注"。 斑马鱼幼体(约 10 万神经元,几乎无学习)捕食草履虫时,用一套离散动作(转向、前冲)朝着一个近似最优统计预测器给出的"猎物将去的位置"移动,而不是猎物当前位置;若猎物偏离预测则放弃。果蝇、蜜蜂(约百万神经元)、老鼠、猿类、乌鸦都在思考却都没有语言。因此大脑的根本功能是用世界模型做好的猜测与下注,其核心公式是"理性主体选择行动以最大化期望效用"(Russell & Norvig《AI: A Modern Approach》的框架)。
  • "脑中的游戏引擎"+概率程序是常识物理的可计算模型。 游戏引擎式物理模拟器快速、近似、"看起来对即可",与大脑的性能特征吻合(参见 Ullman 的 *Mind Games*)。把它包进概率推断框架,对积木位置、摩擦系数、可能的碰撞/风等不确定性做少量、粗糙、短时间步的模拟,就能定量预测人类对 50–60 个积木塔稳定性的评分;对"撞桌子后掉下的是红块多还是黄块多"这种从未思考过的新问题,模型预测力几乎同样高。告诉被试"灰块比绿块重 10 倍"就能改变其预测,说明心理物理引擎可以被语言"编程"。
  • 同一工具链覆盖动力学、婴儿、工具使用与心智理论。 Kevin Smith 的台球任务(100 个场景)复现了人类信心随时间"先红后绿"的多次翻转(人通常只能模拟一两次反弹);12 个月大婴儿对"扭蛋机"遮挡后弹出何种颜色物体的注视时长,可由概率物理模型的逆概率定量预测;Kelsey Allen 的"虚拟工具游戏"中人类 2–3 次(很少超过 5–10 次)就能通关,"脑内先试几次、可行再真做"的模型精确复现了 20 个关卡的学习曲线;"食物卡车"实验中,人们推断主角最爱的是场景里根本没出现的墨西哥餐车,因为这是解释其高效路径 + 错误初始信念的最佳方式(Jara-Ettinger 的"朴素效用演算")。
  • 语言如何接入:源代码作为"思维语言"的隐喻。 Wong & Grand 等人的"白皮书"提出,用 Codex 这类代码 LLM 把自然语言概率翻译成概率编程语言(Church,一种 Lisp 变体,见 probmods.org),再接到物理引擎上运行少量模拟。这样从"一个高的红色积木堆、两个矮的黄色堆,撞桌子后哪种颜色掉得多"的自然语言描述出发,端到端得到的预测力接近手工搭建的模型。源代码之所以合适,是因为它兼具程序结构和人类可读的变量名、函数名与注释。
  • 这是迈向"语境化意义理论"的一步。 意义可视为自然语言与概率思维语言代码之间的联合分布:概率程序提供组合性与具身接地的方面,代码 LLM 提供分布式统计与部分语用/语境方面,从而有望统一形式语义、分布式词嵌入、具身接地、语用四种传统意义观。

结论与值得注意的细节

  • 核心结论:思维 = 用多粒度世界模型做理性的猜测与下注;语言是人类的"奇点",因为它接入了跨代际的文化知识积累(这正是大学、讲座、论文存在的理由),但它建立在更古老的常识表征之上并加以放大——所以工程上应走神经-符号-概率综合路线。
  • 讲者刻意区分"游戏引擎"两种用法:不是作为训练机器学习算法的模拟场地,而是作为对孩子脑内想象过程的工程化刻画。
  • 讲者与 Griffiths、Chater 合著的贝叶斯认知科学教科书即将出版;Fedorenko、Mahowald、Ivanova 关于语言与思维关系的综述将刊于 *Trends in Cognitive Sciences*。
  • 多次强调 Russell & Norvig 的教材,批评当下 AI 讨论几乎只围绕 ChatGPT,忽视了学科的基础工具。
  • 结尾带有明确立场:若不想把 AI 的未来完全交给少数资源雄厚的公司与个人,就需要这种更开放、更民主、更植根于科学的研究路线——"加入我们"。
核心句型 · 9
1. There's no particular reason (that) … should …
“There's no particular reason that should be able to solve any problem or perform any kind of thinking”
用于理性质疑一个被默认的推论:不否认可能性,只指出缺乏必然理由。学术讨论中比 It can't 更克制、更难反驳。
2. It's not the other way around.
“It's not the other way around. It's not training a lots of language. Language comes from the mind”
先给结论,再用这句否定反向因果,最后重申方向。适合纠正常见误解、强调因果顺序。
3. X is not about A, it's about B
“The fundamental thing about brains … is not about learning. … The fundamental computational function of brains is to make good guesses”
先否定听众默认的答案,再给出自己的答案,制造反差。表达核心论点时效果强,注意否定要有后续解释支撑。
4. whether it's … or … or …
“Whether it's in intuitive physics or intuitive psychology or other sorts of domains”
列举多种情形以说明结论的普适性,三项并列最自然,最后一项可用 other sorts of 收束。
5. if I do say so myself
“If I do say so myself I think it's a really nice book”
自夸时的礼貌缓冲语,承认「这话由我说不太合适」。口语与演讲中都常见,可仿写:It's a solid result, if I do say so myself.
6. not just … but …
“It's about not just knowledge but understanding.”
递进强调,第二项是重点。可省略 also,简洁有力,适合结语。
7. So much so that …
“It's so much so that people many people are asking, well, maybe …”
承接上文程度,引出后果:「以至于…」。句首独立使用时前面需有描述程度的句子。
8. A alone is enough to …
“That alone is … enough to give a quantitatively predictive model”
强调单一因素即可达成效果,突出方法的简洁性。仿写:A few simulations alone are enough to explain the data.
9. Unless we want to …, if we want to …, we need …
“Unless we want to completely offload the AI future to a small number of … companies … we need tools and a mindset”
呼吁式结构:先排除不愿接受的前景,再陈述愿景,最后给出行动要求。适合演讲结尾。
词汇精讲 · 110 · 按出现顺序
installment /ɪnˈstɔːlmənt/ n. 0:00
(系列中的)一期、一讲
endowing /ɪnˈdaʊɪŋ/ v. 0:00
捐赠资金设立(讲座、职位等)
pioneer work phr. 1:07
开创性工作
gather /ˈɡæðər/ v. 2:32
推断、领会(from the introduction)
reverse engineering n. phr. 2:32
逆向工程:通过重建来理解系统
payoff /ˈpeɪɔːf/ n. 2:32
回报、收益
adjacent /əˈdʒeɪsnt/ adj. 3:38
相邻的;此处指相关领域
thrilling /ˈθrɪlɪŋ/ adj. 3:38
令人激动的
distinctive /dɪˈstɪŋktɪv/ adj. 4:35
独特的、有区别性的
macaque /məˈkæk/ n. 4:35
猕猴
thesis /ˈθiːsɪs/ n. 5:26
论点、命题
abstract away from phr. 6:03
从…中抽象出(忽略细节)
inner product n. phr. 6:03
内积
non-linearity /ˌnɒnlɪniˈærəti/ n. 6:03
非线性(函数)
cost function n. phr. 6:03
代价函数、损失函数
so much so that phr. 7:47
以至于
speculating /ˈspekjuleɪtɪŋ/ v. 8:37
推测、猜想
fragile /ˈfrædʒl/ adj. 9:31
脆弱的、易崩溃的
spooky /ˈspuːki/ adj. 9:31
令人毛骨悚然的、诡异的
wad /wɑːd/ n. 9:31
一团(软物)
puzzling /ˈpʌzlɪŋ/ adj. 11:16
令人费解的
adhere /ədˈhɪr/ v. 11:16
粘附
palette /ˈpælət/ n. 11:16
调色板
affordances /əˈfɔːrdənsɪz/ n. 11:56
可供性:物体所允许的操作方式(心理学术语)
insightful /ˈɪnsaɪtfʊl/ adj. 12:45
富有洞见的
emergent /ɪˈmɜːrdʒənt/ adj. 12:45
涌现的
operationalized /ˌɑːpəˈreɪʃənəlaɪzd/ v. 13:48
操作化:把抽象概念定义为可测量/可执行的形式
latent /ˈleɪtnt/ adj. 13:48
潜在的、隐含的
deploy /dɪˈplɔɪ/ v. 14:23
部署、投入使用
auto-regressive /ˌɔːtoʊrɪˈɡresɪv/ adj. 14:23
自回归的:用前面的输出预测下一个
moving targets n. phr. 15:28
不断变化的目标(难以固定评估的对象)
thumbs up phr. 15:28
点赞、认可
substrate /ˈsʌbstreɪt/ n. 16:13
基质、底层载体
neuroimaging /ˌnʊroʊˈɪmɪdʒɪŋ/ n. 16:53
神经影像
constructed languages n. phr. 17:37
人造语言
state-of-the-art adj. 17:37
最先进的
articulate /ɑːrˈtɪkjuleɪt/ v. 18:21
清晰阐述
plausibly /ˈplɔːzəbli/ adv. 18:21
看似合理地
selectively impair phr. 19:13
选择性地损害
bottom line n. phr. 19:13
底线、结论要点
compelling /kəmˈpelɪŋ/ adj. 20:12
有说服力的
laid out phr. v. 20:12
阐明、铺陈
corpus /ˈkɔːrpəs/ n. 21:12
语料库
proto-languages /ˈproʊtoʊ ˈlæŋɡwɪdʒɪz/ n. 22:06
原始语言、雏形语言
pidgin /ˈpɪdʒɪn/ n. 23:16
皮钦语:混合简化语言
creoles /ˈkriːoʊlz/ n. 23:16
克里奥尔语:由皮钦语发展成的母语
from scratch phr. 23:16
从零开始
in the grand scheme of things phr. 23:16
从全局来看
fondness /ˈfɑːndnəs/ n. 24:10
喜爱、偏爱
psychophysics /ˌsaɪkoʊˈfɪzɪks/ n. 24:10
心理物理学:研究刺激与感知关系
paramecia /ˌpærəˈmiːʃiə/ n. 24:10
草履虫(复数)
bouts /baʊts/ n. 25:26
一阵、一回合(动作)
deviates /ˈdiːvieɪts/ v. 25:26
偏离
consummate /ˈkɑːnsəmət/ adj. 26:34
技艺高超的、完美的
object permanence n. phr. 27:39
客体永久性(发展心理学概念)
singularity /ˌsɪŋɡjəˈlærəti/ n. 28:37
奇点:质变的临界点
tap into phr. v. 29:36
利用、汲取
enduring /ɪnˈdʊrɪŋ/ adj. 29:36
持久的
opinionated /əˈpɪnjəneɪtɪd/ adj. 30:43
固执己见的;此处指带有个人立场的
grains /ɡreɪnz/ n. 30:43
粒度、精细程度
conspecifics /ˌkɑːnspəˈsɪfɪks/ n. 31:43
同种个体
canonical /kəˈnɑːnɪkl/ adj. 31:43
经典的、权威的
cracked open phr. v. 31:43
翻开(书)
expected utility n. phr. 32:37
期望效用
cash this idea out phr. 33:36
把这个想法具体落实、兑现
remedy /ˈremədi/ v. 34:24
补救、纠正
led the charge phr. 35:18
带头、领军
monograph /ˈmɑːnəɡræf/ n. 35:18
专著
fall under the heading of phr. 36:14
归入…范畴
immerse /ɪˈmɜːrs/ v. 37:18
使沉浸
hacks /hæks/ n. 38:13
取巧的技术手段、权宜之计
catch-all /ˈkætʃɔːl/ adj. 40:28
包罗万象的、笼统的
paradigms /ˈpærədaɪmz/ n. 40:28
范式
differentiable /ˌdɪfəˈrenʃiəbl/ adj. 41:26
可微的
motifs /moʊˈtiːfs/ n. 41:26
母题、反复出现的主题元素
lumped under phr. v. 42:16
归并到…之下
confounds /ˈkɑːnfaʊndz/ n. 44:02
混淆变量(实验术语)
graded /ˈɡreɪdɪd/ adj. 44:02
分级的、连续程度的
inverse graphics n. phr. 44:02
逆向图形学:从图像反推场景
coefficients of friction n. phr. 44:52
摩擦系数
perturbation /ˌpɜːrtərˈbeɪʃn/ n. 44:52
扰动
coarse /kɔːrs/ adj. 46:18
粗糙的、粗粒度的
extrapolate /ɪkˈstræpəleɪt/ v. 48:04
外推
cone of uncertainty n. phr. 48:04
不确定性锥(预测范围随时间扩大)
privileged /ˈprɪvəlɪdʒd/ adj. 49:01
荣幸的
occlusion /əˈkluːʒn/ n. 49:51
遮挡
violation of expectation n. phr. 49:51
期望违背(婴儿研究范式)
alluded to phr. v. 50:53
暗指、提及
unfolds /ʌnˈfoʊldz/ v. 51:57
展开、发生
trial and error n. phr. 52:45
试错
prior /ˈpraɪər/ n. 52:45
先验(分布)
resample /riːˈsæmpl/ v. 52:45
重新采样
cumulative /ˈkjuːmjələtɪv/ adj. 53:38
累积的
posit /ˈpɑːzɪt/ v. 54:31
假定、设定
false initial belief n. phr. 56:52
错误的初始信念(心智理论概念)
a taste of phr. 57:45
…的一点体验、初步了解
road map n. 58:23
路线图
ground them in phr. v. 58:23
使…扎根于、接地于
unifying substrate n. phr. 60:28
统一的底层基质
hierarchical phrase structure n. phr. 61:17
层级短语结构
putatively /ˈpjuːtətɪvli/ adv. 63:21
据推定地、假定地
stay tuned phr. 64:21
敬请期待
venture to say phr. 65:21
斗胆说
joint distribution n. phr. 65:21
联合分布
compositional /ˌkɑːmpəˈzɪʃənl/ adj. 66:28
组合性的(意义由部分组合而成)
embodiment /ɪmˈbɑːdimənt/ n. 66:28
具身(认知)
pragmatics /præɡˈmætɪks/ n. 66:28
语用学
offload /ˌɔːfˈloʊd/ v. 68:26
转交、卸给
well-resourced adj. 68:26
资源雄厚的
at stake phr. 69:32
利害攸关
理解自测 · 11 题
1. Erik Brynjolfsson 给 ChatGPT 出的叠塔题用了哪些物品?ChatGPT 的回答暴露了什么?

物品是一本书、四个网球、一根钉子、一个高脚酒杯、一团嚼过的口香糖和五根生意面。ChatGPT 给出了流畅有条理的步骤:网球摆成正方形,书放在上面做平台。讲者在「流畅却脆弱」一节指出,这类回答表面利用了物体的物理性质,但后续 Tyler Wilson 让它自己出题自己解时,方案(画笔粘画布、调色板竖靠、水杯放上面)完全不可能成立;即便提示里允许说「解不出」,它仍自信地给出无效方案,暴露出语言能力与物理理解的脱节。

2. 讲者引用了哪些数据对比人类儿童与大语言模型的语言输入量?

最新的语言模型训练语料约上万亿词,而人类儿童估计每年听到约 500 万词,20 年累计约 1 亿词,相差四个数量级。讲者在「发展与演化」一节用这一对比论证:人类能用少得多的数据学会语言,是因为在学语言之前已经拥有关于物体、主体等的核心知识系统,这些前语言知识是高效学习的关键。

3. 斑马鱼捕食草履虫的实验说明了什么?

Andrew Bolton 在 Florian Engert 实验室的三维心理物理学实验发现,幼年斑马鱼并不朝草履虫当前位置游,而是朝一个简单但近乎最优的统计预测器所预测的未来位置游;若草履虫偏离,它可能放弃。讲者以此说明:一个只有约十万神经元、几乎不依赖学习的大脑,已经具备「预测模型 + 下注行动」的智能雏形,这正是本讲副标题「猜测与下注」的来源,也是「智能不等于学习」论点的实证起点。

4. Ev Fedorenko 的神经科学研究如何支持「语言与思维分离」?

Fedorenko 用 fMRI 定位了人脑中特异参与语言的网络,它对说、听、读、写乃至盲文和人造语言都同样激活。但当人做代数、读代码、逻辑推理时,该网络并不参与;脑损伤也可以只损害语言而保留符号思维。讲者在「脑科学证据」一节据此得出结论:语言网络就只管语言,思维另有系统。GPT-2 能较好预测语言区活动,恰恰说明语言模型建模的是语言网络而非通用智能。

5. 为什么讲者认为「游戏引擎」是理解核心知识的合适工程模型?

游戏物理引擎不追求物理学意义上的正确,而是用近似和取巧在短时间尺度内高效、看起来足够好地模拟各种场景,这与大脑在意的性能特征一致。讲者在「脑中的游戏引擎」一节指出,把它包裹进概率推断框架后,只需少量、粗糙、低分辨率、短时间步的模拟,就能定量预测人类对积木塔稳定性和倒向的判断,甚至预测被试从未想过的「撞桌子后地上哪种颜色积木多」的问题。

6. 餐车实验中,为什么人们会推断主体最爱的是场景里根本没出现的墨西哥餐车?

主体绕过大楼看到黎巴嫩餐车后折回韩国餐车。若她最爱韩国菜,直接去即可;若最爱黎巴嫩菜,看到就该停下。唯一能解释这条高效路径的假设是:她以为墨西哥餐车在那边(错误的初始信念),去看后发现不在,才退而求其次。讲者在「直觉心理学」一节用这个例子展示,人们把他人理解为在信念约束下追求效用的理性主体,模型必须同时推断欲望与信念才能拟合人类判断。

7. 讲者对「思维链」「RLHF」等改进方法持什么态度?这与他的主论点有何关系?

讲者承认这些方法有效:思维链让模型生成中间步骤,对某些问题有帮助;RLHF 使模型输出更符合人类偏好。但他指出思维链对另一些问题仍无效,且这些方法仍以「语言是思维载体」为架构假设。他的主论点是:有限架构加自回归预测没有必然理由能完成任意思考,补丁不改变根本。因此这些改进被定位为在语言中心路线内部的修补,而非他主张的「语言建立在更古老的世界模型之上」的路线。

8. 「语言→概率程序→模拟」的端到端模型是怎样运作的?其结果说明了什么?

模型用 OpenAI 的 Codex 代码模型把自然语言描述(如「一摞高的红积木、两摞矮的黄积木」)翻译成概率编程语言 Church 中的程序,再由 2D 物理引擎运行少量概率模拟回答问题。讲者在「语言如何接入」一节强调,结果几乎与直接从图像搭建模型时的定量预测力相当。这说明语言的作用可以理解为为心理模拟器「配置参数」,LLM 做翻译,推理由程序完成,是一种模块化组合。

9. 讲者提出的「意义」理论是什么?它如何统一四种传统观点?

讲者主张意义应语境化,定义为自然语言与概率程序在概率思维语言中的联合分布。形式语义的组合性与具身接地由概率程序部分提供,分布统计与语用语境的近似由代码 LLM 部分提供。在「结论」前一节,他称这只是迈向意义理论的一步,但认为工程上已能把握。这一方案不是取代四种观点,而是把它们分派给同一框架的不同部件。

10. 如果有人反驳:「只要继续扩大模型规模,物理推理的脆弱性自然会消失」,讲者可能如何回应?

讲者可能从三方面回应。第一,他的批评针对架构而非规模:有限架构做自回归预测,没有必然理由能完成任意计算,规模不改变这一点。第二,生物证据表明智能早于大规模学习:斑马鱼十万神经元就有预测模型,儿童一亿词就学会语言,说明关键是内置的世界模型而非数据量。第三,他并不否认规模的价值,而是主张把 LLM 作为语言到程序的翻译器,与概率程序结合,因此他会说规模能改善翻译质量,但可靠推理仍需显式的模拟与推断机制。

11. 把「大脑的根本功能是猜测与下注而非学习」这一论点放到教育情境,会得出什么推论?它成立吗?

推论是:教育的核心不是灌输大量数据,而是帮助学习者建立并修正世界模型,再用它高效地做预测与行动;语言的作用在于「编程」这些模型,如同讲者只说一句「灰色重 10 倍」就改变了听众的预测。虚拟工具游戏中人类两三次尝试就通关,也支持「先在脑中模拟再行动」优于盲目试错。但需注意讲者的论点针对的是大脑的基础功能,他同时承认人类是了不起的学习机器,且文化知识需通过阅读等大量语言输入积累;因此这一推论在强调模型先行时成立,若走向否定练习与输入量则超出了讲者的主张。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.101Turing's Cathedral | George Dyson | Talks at Google
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com