CHM Live | The Great Chatbot Debate: Do LLMs Really Understand? · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

CHM Live | The Great Chatbot Debate: Do LLMs Really Understand?

节目发布 2025-03-27 · Computer History Museum
艾米莉·本德 塞巴斯蒂安·布贝克 EEliza Strickland
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
2025 年 3 月,计算机历史博物馆(CHM)与《IEEE 综览》(IEEE Spectrum)联合举办「聊天机器人大辩论」,辩题是「大语言模型真的理解吗」。反方是华盛顿大学语言学教授、《论随机鹦鹉的危险》合著者艾米丽·M·本德,正方是 OpenAI 技术团队成员、《通用人工智能的火花》作者塞巴斯蒂安·布贝克,《IEEE 综览》高级编辑伊莱莎·斯特里克兰担任辩论主持。本文依据现场录音编译整理。文中「主持人」指博物馆活动主持人,辩论环节的提问者标为「斯特里克兰」。

开场:聊天机器人的展览

主持人: 晚上好,各位。今晚来了这么多人,欢迎大家来参加这场「聊天机器人大辩论」,也感谢全国乃至全世界正在观看直播和回放的朋友。无论你站在哪一边,是「人工智能」的信徒,是怀疑派,还是压根不知道大语言模型是什么,今晚都会是一场精彩而热烈的活动。首先要感谢今晚的赞助方帕特里克·J·麦戈文基金会,正是他们的慷慨支持,以及各位会员的支持,让我们的工作得以进行。也要感谢豪斯家族酒庄为会前酒会提供的葡萄酒。今晚的活动由我们与《IEEE 综览》联合主办,整个筹备过程是一次真正的协作,感谢伊莱莎和她的团队。

第一次来博物馆的请举手。很好,希望大家还会再来。楼下正在展出「聊天机器人解码」(Chatbots Decoded),它把今晚讨论的话题变成了可以互动的展品,你可以用自己选择的语言和一个机器人对话。在计算机历史博物馆,我们的使命是解码技术:它的计算过往、数字当下,以及对人类未来的影响。今晚的节目正是这一主题的体现,我们要探讨并辩论,「人工智能」究竟有多智能。

今晚双方都是重量级人物。一方是华盛顿大学的艾米丽·M·本德,她是计算语言学实验室主任、语言学教授,兼任计算机科学与工程学院和信息学院的客座教授。她以对 AI 语言模型的批判立场著称,是《论随机鹦鹉的危险》(On the Dangers of Stochastic Parrots)一文的合著者,新书《AI 骗局:如何对抗大型科技公司的炒作并创造我们想要的未来》将于 5 月 13 日出版,请记得预订。另一方是 OpenAI 的塞巴斯蒂安·布贝克博士,他目前是 OpenAI 技术团队成员,此前担任微软 AI 副总裁和杰出科学家,在微软研究院工作了十年,更早时是普林斯顿大学助理教授。他 2023 年的论文《通用人工智能的火花:GPT-4 早期实验》(Sparks of Artificial General Intelligence)在《纽约时报》《连线》等媒体引发了广泛讨论。今晚在八角笼里执裁的,是《IEEE 综览》高级编辑伊莱莎·斯特里克兰,她长期报道 AI、生物医学技术和其他前沿技术,常在播客露面,从西南偏南到今晚的博物馆,主持过许多活动。下面请她为今晚的辩论开场。

火花对阵鹦鹉

斯特里克兰: 大家好,感谢各位到场。先用一分钟介绍一下《IEEE 综览》:我们是 IEEE 的旗舰刊物,IEEE 是电气与计算机工程师的专业组织,全球有五十万会员。我们有月刊,也有对所有人免费开放的网站,上面有各种精彩的技术报道。

当《IEEE 综览》开始和博物馆商量办这场辩论时,我立刻知道自己想在台上看到什么:火花对阵鹦鹉。反方人选,我想不出比艾米丽·本德更合适的人。她关于随机鹦鹉危险的那篇论文写于 ChatGPT 问世前几年,就已点燃了关于语言模型风险的激烈争论,如今看来非常有先见之明。这些年我采访过她几次,她有一句话一直留在我心里:我们理解语言的方式,是去想象正在对我们说话的那个人的心智。所以当我们遇到合成文本时,我们是在想象一个并不存在的心智。我觉得这个视角很有意思。

正方,我想请的是塞巴斯蒂安·布贝克,他写了《通用人工智能的火花》那篇预印本,分享了对 GPT-4 的早期实验。论文发布前后,我在一档播客里听他谈到 OpenAI 早先的模型和 GPT-4 之间的差别。他说,最大的问题是早先的模型并没有真正掌握「猫」或「狗」这样的概念,而到了 GPT-4,他清楚地看到它确实理解了某些东西。

所以我认为他们是最理想的辩手,很幸运两位都答应了。

从图灵测试到 Eliza

斯特里克兰: 先简单回顾一下历史,说说这些技术是怎么运作的,以及眼下正在发生什么。大家大概都听过图灵测试。二十世纪五十年代,著名计算机科学家艾伦·图灵提出了一个思想实验:假如一位人类评判者通过计算机界面同时与一个人和一台机器聊天,而他分不出哪个是哪个,那就说明这台机器至少把人模仿得足够像,足以展现出智能行为。图灵没有直接说它就是智能的,只说它展现了智能行为。在当时,计算机能通过这项测试被认为是纯粹的幻想。

时间来到 1966 年,麻省理工学院教授约瑟夫·魏岑鲍姆造出了第一个聊天机器人,巧的是,它也叫 Eliza。它模仿心理治疗师,把你说的话反射回去。你说「我工作压力很大」,它就问「你为什么工作压力很大」;你说「我很生我妈妈的气」,它就说「多讲讲你为什么生妈妈的气」。不过是些简单的脚本。但有意思的是,人们给这些脚本赋予了智能和情感。魏岑鲍姆的秘书和 Eliza 聊天时,甚至要请他离开房间,因为她想为这些私密谈话保留隐私。

之后是几十年的缓慢进展,我不打算把 AI 的整部历史讲一遍。直到 2010 年代,我们才有了 Siri 和 Alexa 这样勉强算得上像样的语音助手,但它们带来的挫败常常多于帮助。同样是 2010 年代,AI 领域开始了一场革命:一种叫人工神经网络的技术,长期被视为死胡同,却突然复兴了。

神经网络与大语言模型

斯特里克兰: 简单说,神经网络是一类机器学习模型,灵感松散地来自大脑的运作方式。它由一层层相互连接的节点(或称神经元)构成,负责处理信息;神经元之间的连接带有权重,网络从数据中学习时会调整这些权重。神经科学里有句话,「一起放电的神经元会连在一起」,意思就是,学习过程中最活跃、联系最紧密的那些连接会得到强化。2010 年代的变化在于,这些机器突然有了海量数据可学:社交媒体上的所有图片,Reddit、社交网络和书籍里的所有文本。于是它们的能力跃升了一大截,先是在计算机视觉领域,然后是文本处理,也就是自然语言处理。

这催生了大语言模型(LLM)。它们建立在神经网络之上,能生成像人写的文本,因为它们读过太多这样的文本,基本上读遍了整个互联网。在训练过程中,模型摄入这些数据,学会识别其中的模式,比如哪些词和短语常常一起出现。这让模型能够预测句子里接下来是什么,或者根据你给的输入、提出的问题和提示,生成相关的回应。所以一个大语言模型可以前一分钟讲解量子物理,下一分钟用苏斯博士的风格写诗,再下一分钟写一道茄子食谱。它见过所有这些东西,于是基本上能做到它见过的一切,甚至更多,而那个「更多」正是有意思的部分。

但大家可能也见过,大语言模型常常犯令人尴尬的错误。比如一年多前谷歌的聊天机器人建议人们为了健康每天至少吃一块小石头,还建议往披萨里加八分之一杯胶水,防止配料掉下来。看到这样的事情,你难免会问:这里面到底发生了什么?

推理模型与过河谜题

斯特里克兰: 再说说当下。过去几个月,出现了一批强大的新模型,叫大推理模型。这算是大语言模型之上的升级:它们用思维链推理把一个问题拆成许多步骤,逐步执行,这似乎减少了错误,让模型不那么容易产生「幻觉」或失误。

变化究竟有多大,看一眼很有意思。我今天为了好玩自己试了一下。大家大概听过那道著名的谜题:一个人带着一棵卷心菜、一只羊和一头狼要过河,船上只能载他和一样东西。要怎么一次一样地运,才能不让羊吃掉卷心菜、狼吃掉羊?有意思的是,语言模型对这道题烂熟于心,因为它们见过。你换成老鼠、奶酪和猫,它们照样答得出。但如果你用很简单的方式改动一下题目,往往就能把语言模型绊倒。我今天问 ChatGPT:一个人和一只羊要到河对岸,他们有一条船,该怎么过去?它回答说:这个人带着羊过河,把羊留在对岸,然后独自返回原来这边去取船,现在人和船都在原来这边,而羊安全地在对岸。我指出这个人并不在正确的那一边,它说:抱歉,人和羊应该一起上船过河,人把羊留在远岸,独自带船回到原来这边,然后再自己划过去。你看,它像是在努力把碎片拼起来,但就是拼不对。然后我把一模一样的问题拿去问 OpenAI 的推理模型,它只说:让人和羊一起上船,划到对岸,就完事了。可以看出,这些模型的运作方式正在发生某种变化。

今天还有一件事在发生,就是 AI 智能体(agent)。它们是建立在语言模型之上的系统,能替你采取行动:帮你订餐厅,找去巴黎最好的航班,或者在亚马逊上搜最便宜的鞋。

辩题:何谓真正理解

斯特里克兰: 在这样的背景下,我们今晚面对的问题是:大语言模型真的理解吗?我们应该花点时间想想「真正理解」是什么意思,因为定义很重要。就人类而言,说一个人理解某个概念,通常是指他对它有深刻而直觉的把握,不只是知道事实、能复述信息,而是能在各种情境下运用知识,能感知意义和意图,并把这些想法与自己的世界模型联系起来。

至于大语言模型,问题就是这些 AI 系统是否具备类似的理解:它们真的理解自己说的话吗?还是只是在模仿训练数据里见过的模式,并无真正的领会?这就是今晚辩论的核心。如果大语言模型真的理解,那会改变我们与技术的互动方式,我们会让 AI 更深地嵌入日常生活。反过来,如果它们只是模仿模式,那我们在使用、整合和依赖它们时,也许就得更谨慎。

今晚的流程是:两位辩手先各自做开场陈述,然后进入我提问、他们互相提问的辩论环节,接着是现场和线上观众的问答,最后是结辩。辩论结束后,请大家再次参加 Slido 投票。还没投的现在就可以投,这样我们能知道辩论开始时观众站在哪里,结束时再看一次。听辩论的时候,请同时考虑两种视角,保持开放心态,批判地审视双方提出的证据和推理。这不只是一场技术辩论,也是一场哲学辩论,触及智能与理解的本质,触及在这个与机器交互的时代,做人意味着什么。

有请两位辩手上台,艾米丽和塞巴斯蒂安。刚才在休息室我们抛了硬币,艾米丽赢了,她选择先做开场陈述。

本德开篇:形式与意义

本德: 谢谢。感谢各位到场,感谢计算机历史博物馆组织这次活动,也感谢线上正在观看或将来观看的朋友。在开场的十分钟里,我要做四件事,依据的是我从二十世纪九十年代起研究和教授语言学与计算语言学所学到的东西。语言学在这里至关重要,因为它研究的正是语言如何运作、我们如何运用语言,而这两点都是回答「大语言模型是否真正理解」的关键。

四件事是:第一,给「理解」提供另一个定义;第二,讲讲什么是语言模型,它这些年是怎么发展的;第三,论证语言模型的整个发展过程中,没有任何环节让它们接触到语言这道「形式加意义」等式里的意义那一半,而没有意义就没有理解;第四,谈谈人是怎么理解语言的,以及这套机制为什么让我们容易陷入「大语言模型能理解」的幻觉。

说这些的时候,我会尽量避免拟人化的用语。我不会说大语言模型「努力」做什么或「试图」做什么,不会说它们「提问」或「回答」,至少会尽力不说,因为那会把水搅浑。同样,我提到「人工智能」或「AI」时一律加引号,因为它并不指代一套定义清楚的技术。总的来说,更好的做法是谈「自动化」,然后再谈我们在自动化什么,以及为什么要自动化它。

第一点,什么是理解。在我和亚历山大·科勒 2020 年发表的那篇论文(有时被称为「章鱼论文」)里,我们给理解下了一个技术性定义:理解就是从语言的形式(你听到的声音、看到的手语、读到的文字)映射到语言之外的某种东西。这个定义非常窄。按这个定义,你让语音助手开灯,它照做了,它就理解了这条指令。所以造出能完成这种映射的自动化系统,是可能的。

但人类使用语言时,通常远不止于此。比方说,我说「看那边那只鹦鹉」。单凭词语和语法,你就能知道我在把你的注意力引向某处,而且那里很可能有一只鹦鹉。但你在理解时投入的其他一切,还会让你知道:我希望你去看我称为鹦鹉的那个东西,而且我要么相信它确实可以被合理地叫做鹦鹉,要么在装作相信。这是我们日常所做的那种非常丰富的理解的一个例子。

假如我换成日语说这句话,不懂日语的人大概听不出什么。但如果是面对面,你能看到我的手势和目光,你可能会猜到我在把你的注意力引向某样东西。你朝那边看去,如果那里有一只鹦鹉,而且是最显眼的东西,你或许会猜到我刚才用的某个词就是日语里的「鹦鹉」(顺便告诉大家,是「オウム」)。这甚至能帮你开始学日语。但所有这些推断、所有这些意义,都不是从词语里来的。除非你懂日语,否则那些词对你毫无帮助,我相信在座肯定有人懂。在这种共处一地的情境里,因为你是一个人,带着全部的语境理解,你才能利用这些线索开始学日语。关键在于,大语言模型不会有这些。这一点至关重要。

从 T9 到 LLM

本德: 第二点,语言模型。这其实是很老的技术,可以追溯到克劳德·香农的工作,再往前是马尔可夫。在座很多人大概还记得手机上的 T9 输入法,「九键打字」:你按数字键,每个数字对应三四个英文字母,但输入四五个数字的序列后,词典里可能匹配的词只有那么几个,界面就按这些词在某份训练数据里的频率排好序给你。它很烦人,因为 home 和 good 对应同一串按键,不管你前面打的是「I'm coming」,出来的永远是 good,因为它在语言模型里更常见,而那个语言模型不过是一份词频统计。这是最基本的一元语言模型(unigram)。要让它更复杂,可以统计一个词在前面一到 n 个词的语境下出现的次数。这一直是自动转写和机器翻译技术的重要组件:先由一个组件说,给定这段声音序列,或者给定这串法语,可能对应哪些英语词,得出一批候选,然后请语言模型按「哪个最像训练数据」重新排序。这就是九十年代到两千年代的语言建模。顺带一提,它对拼写检查也很有用:从输入的字符出发,经过一两次替换、删除或插入能得到哪些词?按单纯的词频排序,或者按语境下的可能性排序。一旦语言模型能捕捉更多模式,它就能说,「这里的 there 用错了」,把这类错误挑出来。

在所有这些场景里,我们用语言模型回答的问题都是:在这个应用产生的一批候选字母串中,给定训练数据,哪一串最有可能?我们始终只和词的形式打交道,从不涉及这些词能用来指称什么,或者正被用来指称什么。

在这个背景下,今天所谓的 LLM 有两大区别。第一,机器学习架构有了进展,能够利用如今可以构建出来的巨型数据集(收集方式是有问题的,并未取得同意,但确实构建出来了)。这些架构在利用更大训练集的同时,建模的词间关系也远远超出「有四个前文词,给词表按可能性排序」的水平:它对哪些词在什么关系下、在多长的文本跨度上倾向于共现,建模得精细得多。同样重要的是,训练数据是不公开的,只有少数例外,比如艾伦人工智能研究所(AI2)对训练数据就相当公开。第二个大区别是,这套系统被翻了个面。我们不再用它回答「这几串候选拼写里哪一串最可能」,而是让它反复回答「下一个可能出现的词是什么」,一遍又一遍。

这是两大区别。还有两个小区别。一是额外的训练,比如基于人类反馈的强化学习(RLHF):大量数据标注工人看系统的输出,比较「这个更好,那个更好」,或者简单打「好」「坏」,让系统设计者把概率朝着标注工人点赞的词去塑形。二是输入和输出之间加了一些额外的处理步骤。比如扫描用户输入里的算术题,把它交给计算器,以免出现「明明是计算机却不会算数」的尴尬输出。有时用户输入会被转成网络查询,取回文档,再连同一段「生成摘要」的系统提示一起输入,这就是所谓的检索增强生成(RAG),实际上是把取回的文档做成一份纸浆拼贴,同时打消用户去原始语境里亲自阅读这些信息的念头。此外还有系统提示本身:你在 ChatGPT 输入框里打的字,并不是进入系统的唯一输入,外面还裹着一整段提示,每次还会把对话历史也塞进去,而你之后只看到那一小段输出。界面里藏着很多东西。

意义是我们赋予的

本德: 所有这些的结果,是看起来连贯的文本,按照做 RLHF 的标注工人的偏好被塑造得讨人喜欢。但它本质上仍然是反复回答「下一个可能的词是什么」的产物。那它为什么看起来像在理解?答案是:它之所以显得有意义,只是因为我们在为它赋予意义。你可能以为,理解你听到或读到的东西时,意义就在词语里,你只是把词拆开,把意义取出来。但语用学和心理语言学的研究告诉我们,事情并不是这样发生的。我们一直在做的,是尽力弄清对方一定想传达什么,而词语只是其中特别丰富的一条线索。

要做到这一点,我们必须想象文本背后有一个心智。所以当我们遇到聊天机器人吐出的合成文本时,我们做的是同一件事,本能地、条件反射地去做。然后就很难再提醒自己:那其实全在我们这一边,那个心智完全是我们为了让文本说得通而编出来的。

总结一下。「大语言模型能理解」这样非同寻常的主张,需要非同寻常的证据。提出这个主张的人还必须定义「理解」,而且这个定义要以我们对人类如何理解语言的知识为依据,还必须可证伪。隐瞒训练数据,甚至连数据里大致有什么都不说,是反科学的,因为这让人无法推断训练数据加上输入是如何得出输出的;也是反社会的,因为这让我们无法判断什么时候用这些东西是合适的。任何「文本进、文本出」形式的测试,在我们看来都可能像是系统在回答问题、在推理、很聪明,但我们要当心,这些测试真正检验的只是系统对训练数据词语分布的拟合程度,以及训练数据与测试情境的接近程度。所以,一句话:大语言模型理解吗?断然不。

布贝克开篇:观者之眼

布贝克: 大家好。我的看法会和艾米丽相当不同。我想先承认,今晚辩论的问题本身有点怪。我的意思是这样:前几天我在听一场像今晚这样的讲座,讲的是物理,一位哈佛教授在讲爱因斯坦的广义相对论,他在讲的过程中说,「爱因斯坦显然对广义相对论一无所知」。这对我来说非常重要:理解在观者眼中。我自己在研究领域里也有过很多次这样的体验,觉得某些研究者好像不太明白发生了什么。但他们当然明白,拜托,他们是这个领域的研究者。我想说的是,判断某人或某物是否理解,有不同的门槛、不同的层级。这是第一点。

第二点,正如伊莱莎和艾米丽所说,我们需要给理解下个定义。我完全同意。问题是,人类为「理解」的含义争论了几千年,先剧透一下:今晚我们不会把它弄清楚。我们谁都没有一个能涵盖这几千年工作的必杀定义。这一点很重要。

既然我不打算给大家一个理解的定义,我想回到有数字、有确凿数字、我们能非常精确地把握的东西:基准测试(benchmark)。基准是评估过去几年进步速度的好办法。我想花几分钟,和在座各位一起盘点一下我们这几年走了多远,然后再回到理解的问题。

基准分数的暴涨

布贝克: 两年半前,GPT-4 和 ChatGPT 问世前夕,我在一个关于神经网络数学理解的研讨会上,谷歌的一位同行给我看了他们刚训练的模型 Minerva,一个数学模型。这个模型能解出一道高中水平的题,我完全被震住了。那不过是一道关于直线相交的简单题,真的没有任何难度,但两年半前我为此震惊。今天,我们的系统正在角逐数学奥林匹克竞赛的金牌。我认为承认这一点很重要,而且这与我们给理解下什么定义无关。这就是我们正在目睹的进步速度。

拿 MATH 基准来说,它是高中竞赛水平的题目。两年半前,聊天机器人能拿大约 20%。到去年初,它已经完全饱和,模型太强了,这个基准没法再用,只好废弃。还有一个曾经非常流行的基准叫 MMLU,测的基本上是高中水平的科学问题理解。它曾是一个极其重要的基准,大型科技公司的 CEO 们都在谈它,谁多了 2% 或 5% 都能撼动全球市场,马克·扎克伯格三天两头提起它。现在你完全听不到了。为什么?因为它也完全饱和了,模型在上面太好了。于是我们不断推出更新、更难的基准。眼下大家争相刷高的最新基准叫 FrontierMath(前沿数学)。顾名思义,这些是非常高深的数学研究问题,不是未解难题,但也是你会拿去问数学系低年级博士生的那种题。这个基准大约四个月前推出时,OpenAI 当时最新的推理模型 o1 只能答对 4%。大家都很满意:终于有一个超出模型能力的基准了。一个月后,我们发布了 o3,其实是 o3-mini,一个更小的模型,它拿到了 30%。我甚至不知道地球上有没有哪一个人能独自解出这些题的 30%。这就是我们所说的进步速度。

省下三天的 o3-mini

布贝克: 不过说实话,我个人并不喜欢基准测试。两年前写《通用人工智能的火花》,整篇论文的要点就是我们需要摆脱基准,因为,我想在这一点上我们会有共识,基准无法展示理解。无论你把基准设计得多好,都很难从中提取出理解。在我看来,理解要靠与系统互动来判断,同时非常小心地不把看到的输出拟人化,而是真正去探究,追问一些试探性的问题,看看理解能走多远。因为理解不是一个二元的概念,不是零或一,它是一个连续的量:你理解到什么程度?

所以我想回到一种更务实的理解观:这些系统能为你做什么?它们能不能以某种方式帮你理解新的东西?在这一点上,我想请每个人回想自己与这些模型打交道的经历。和两年前相比,如今的好处是在座每个人大概都用过聊天机器人。所以你自己就能判断它们是否理解,判断它们有没有以某种方式帮到你。我说说我自己的经历。最近我用 o3-mini 问了一个数学研究问题,我确定网上任何地方都没有答案。我怎么知道?因为那是我自己思考过的问题,我甚至为它写了一篇论文,只是一直没时间发表,它就躺在我的 Dropbox 里,大概只有另外两个人看过。我想没有人知道答案。我把这个问题拿去问 o3-mini。如果你感兴趣,它是一个关于轨迹长度的优化问题,细节无关紧要。重点是,o3 当然没有解出来,那要求太高了,我根本没抱这个希望。但它给出了一个建议,指出这个问题与另一个领域的联系,而这个联系我本人花了三天才发现,这个模型对着问题推理和思考两分钟,就找到了。这是否意味着它理解?很难说。它是否极有帮助,是否加快了我的工作,让我可以把那三天花在别的事情上?绝对是。这就是我想把讨论引向的方向。

我还想补充一点,是有利于「缺乏理解」,或者说「理解究竟是什么」这一面的。这里要回到哲学家关于理解含义的争论。我认为有一种可能:说到底,理解是一段属于人的旅程,关乎的是我们自己的理解。所以我说,与其问「聊天机器人理解吗」,不如问「聊天机器人是否帮你理解了更多东西」。举个例子,在数学领域,完全有可能在未来几年内,这些模型做出突破,发现新东西,证明没有任何人证出来的新定理。但数学界可能不会接受。它会承认定理被证明了,但会说,而且我认为说得对,我们仍然不理解它。只有当一个人能完整掌握整个证明时,我们才会承认我们理解了。我的开场陈述就到这里。

AAAI 调查与 AGI 之路

斯特里克兰: 我们开始吧。第一个问题是关于美国人工智能促进协会(AAAI)本月早些时候发布的一份报告,其中调查了 475 位 AI 研究者。大约 76% 的受访者认为,扩大现有方法的规模不太可能或极不可能实现通用人工智能(AGI);80% 的受访者认为,当前公众对 AI 能力的认知与现实不符。这和我们讨论的理解问题有些相邻,但我很想知道两位怎么看这些结果,因为它显示研究界对技术目前所走的路线能否通向伟大成就,有相当程度的怀疑。塞巴斯蒂安,先从你开始。

布贝克: 这确实是个很有意思的问题,因为过去几年最让我震惊的一件事,就是学术界和研究者群体几乎是最慢理解正在发生什么的人。也许这有充分的理由。大语言模型的问题在于,至少从我们对这些对象的数学理解来看,它完全违背了我们自以为了解的关于神经网络和人工智能的一切。作为一个研究群体,我们从来没想过 ChatGPT 这样的东西是可能的。所以很多怀疑残留下来,这是第一点。第二点,「扩大规模」到底指什么?正如你在开场时说的,现在有了新的推理模型,它们是在沿着另一条轴扩展规模。对我来说这都是一回事,我不区分这些不同的方法,从某个抽象层面看它们完全相同。但它确实不是从 GPT-1、GPT-2、GPT-3 到 GPT-4 那种纯预训练范式,而是一种略有不同的扩展范式。所以受访者也许是在说,经典的预训练路线不会把你带到 AGI。这一点我完全同意,绝对同意。但作为人类整体,我们是否走在通往 AGI 的路上?是有可能的。

斯特里克兰: 艾米丽,你的看法?

本德: 先说第二个问题,我认为那是个像样的问题。问「炒作是否与我们实际拥有的功能相符」,拿这个问题去问人是合理的,而构建这些系统的人也许正处在回答它的好位置上。所以这个结果不让我意外,它符合我对现实的看法。第一个问题作为调查题则是坏掉的,因为 AGI 没有定义。除非那份调查给出了一个具体定义让受访者据此回答,否则这是一道失败的调查题。我还要说,我非常反感研究者话语里常见的一种套路,记者和 CEO 有时也这么说:未来存在某个 AGI 出现的时刻,问题只在于我们是否走在通往它的正确道路上、跑得有多快。我在两个方面反对这种说法。首先,没有理由相信这个未经定义的东西必然存在于未来。其次,科学不是这样运作的。科学不是选一条路然后沿着它猛跑,它是一项集体事业,人们分头探索,彼此告知发现了什么。塞巴斯蒂安,这也正好回应你刚才说的:如果一个自动系统产出了定理证明,而数学界无法理解它,这算不算证明了定理?我认为这直接指向一个事实,一切学术都是对话。你独自做了一件让自己满意的事,那还不是科学,也不是学术。

布贝克: 但还有在现实世界里做成事情这回事。比方说,你在化学里发现了某个新公式,能据此造出新的化合物,这就是确凿、具体的证据,证明你做成了。你不需要谈论它,你拥有它,能用它造出新东西。所以这种务实的视角,看它们能为你做什么,很重要。举一个极其具体的例子:假设我想做一个应用,我让其中一个模型替我做。我不是说现在就能做到,但假设将来我只用自然语言说明我要什么,第二天回来,砰,应用就在那里,也许已经装到我手机上了,一切就绪。模型在做出它的过程中理解了吗?某种意义上,我不在乎。它替我做成了一件事。

只看有用就够了吗

斯特里克兰: 那你在乎它是否理解吗?问题出在哪里?

本德: 我在乎它们是否理解,是因为围绕它的许多话语,比如「它能当机器人治疗师、机器人律师、机器人家教」,其前提都是它「似乎在理解」这种观感,或者说这种氛围。这非常成问题。我一向拒绝被称为 AI 研究者,我对那没兴趣。我是计算语言学家,本来安安静静地做语法工程,然后这一切发生了,我意识到语言学在这里真的有很多话要说。别人也许为理解的定义挣扎了几千年,但我们是有定义的。

语言学还有一件极其重要的事:语言这一部分,与被谈论的世界或想象中的世界,并不是一回事。因为我来自西雅图,我喜欢用一个比喻:一扇溅满雨点的窗户。你可以聚焦在雨点和它们形成的图案上,也可以换个焦点,透过窗户看外面的世界。但当你这样做时,雨点实际上在影响你能怎样看到外面。语言中编码的信息就是这样。语言学家关注语言的结构,以及这种结构如何让我们看到其中的信息。而我注意到,许多从计算机科学角度做语言处理的人,以为自己是直奔信息而去,看不见窗上的雨点。所以「某个东西是否理解」这个问题非常重要,因为它让我们停下来评估,任何一个具体应用究竟是不是个好主意。

布贝克: 我想回应一下。我会强烈建议任何掌权者不要凭氛围做决策,那不是好主意,他们不该那么做,谁都不该。我同样会建议他们不要根据「这个系统是否理解」来做决策,哪怕有一个完美的理解定义,那也不是一条可行的路。在这种情况下,比如你说的机器人医生、机器人律师,我们必须做的是建立基准,建立测试。就像我们用考试来检验一个人是否有能力成为医生或律师、是否准备好了一样,我们也应该开始为 AI 系统建立这类测试,而这些测试的性质会与人类的考试大不相同。我很讨厌有人拿美国医师执照考试(USMLE)去测聊天机器人,那是胡闹,毫无意义。系统能通过 USMLE,并不等于你测出了它能当医生。这不合理,因为你对人有太多默认假设,比如假设他记得自己昨天做过什么,聊天机器人可不记得。所以你需要设计新的测试。而这正是我认为把人推离这个领域有些危险的地方:我们其实需要更多人来思考这个问题,我们需要建立这些测试,才能为大规模部署找到一条扎实的路径。

ARC 挑战与基准局限

斯特里克兰: 说到新测试和基准,我想听听两位对 ARC 挑战(ARC Prize)的看法。ARC 是「抽象与推理语料库」的缩写,旨在衡量通往通用人工智能的进展。它有意思的地方在于,题目对人来说并不太难,通常是一些彩色图案,图案的排布背后有某种逻辑,你得找出规律,再把它应用到另一张网格上。它的设计意图是挫败靠记忆的系统,真正考查抽象推理和泛化能力。第一版 ARC,我记得 OpenAI 的 o3 拿到了 87%,主办方差不多说「这一版完成了」。但他们今天刚推出新版,据说 o3 在上面只有大约 4%。这看起来至少是一个新的前沿。你们认为这类测试能评估理解吗?先从你开始。

布贝克: 这确实很有意思,让我先把故事讲一遍。有一个基准叫 ARC-AGI,它的创建者大致宣称,一旦有系统能解决它,那就是我们通往 AGI 之路上真正的里程碑,也许还不算 AGI,但会是一个重大里程碑。此前所有模型都卡在 30% 左右,人类大约在 80%。然后去年底 o3 来了,拿到 85%,直接碾压了这个基准。而且,不透露 o3 内部任何细节,我可以说这对这个基准是真正的泛化。这一发生,就又出现了移动球门的情形:「不对,那不是我们真正的意思,那不算真正的理解。」于是他们推出了 ARC-AGI-2,o3 在上面确实只有可怜的 4%。我得说我今天看了看,很多题我自己也解不出来,题目有点定义不清,也许他们会再打磨,我们会发现并非一切都完美。我认为这引向一个更大的问题,也回到我之前的观点:建立基准是一项艰难的任务,是真正的苦工。作为一个社会,为人类建立测试是一份全职工作,我们的整个教育体系就是围绕这个建立的。而我们必须为 AI 重新发明整套体系。所以在这一个例子上押太多,无论正面还是负面,都不太妥当。

斯特里克兰: 艾米丽,很想听你的看法。

本德: 关于基准我有很多话要说。我 2021 年和黛博拉·拉吉等同事发表过一篇论文,叫《AI 与「包罗万象」基准》,我们发现自然语言处理和图像处理领域某些基准的用法,与一本很棒的儿童读物《格罗弗与包罗万象博物馆》有绝妙的对应。故事里,格罗弗走进这座博物馆,有一间放又长又细的东西的屋子,一间放很重的东西的屋子,一间放怕痒的东西的屋子。他把这些屋子走遍,快到结尾时说:「嗯,我在这座博物馆里看了很多东西,可我还没看到世界上的一切。」然后有一扇门,门上的牌子写着「其他一切」。格罗弗推开门,当然就到了外面的世界。要点在于,任何给定的一组东西都只是一种选取,只是一个样本。自然语言处理的一个真正陷阱是,我们用语言谈论一切。所以当我们设计出能就任何话题输出貌似合理文本的合成文本挤出机时,就很容易觉得,哦,这东西也许真的擅长帮我解数学题,也许擅长根据我的症状做诊断,因为它能回给你看起来像你想要的那种文本。但根本没有办法测试完全的通用性。

而且我要说,那样做的意义何在?我们造技术时,要测的是功能。我们想知道自己造的系统在一定容差范围内运作良好。这里我想到蒂姆尼特·格布鲁博士关于「数据集说明书」的工作,她的灵感来自电子元件的规格书:这个元件会以 32000 赫兹振荡,误差正负一。这就是它的说明书,我们能把它钉死,从而知道在什么情况下用它。数据集说明书的理念是,我们对数据集里有什么也应该有类似的了解,这样对在这个数据集上训练的机器学习模型,我们才知道什么时候能用。而一旦谈到 AGI,谈到通用目的之类的东西,测试它就变得定义不清,作为目标也定义不清。所以我对 ARC 挑战的看法是:它不是通往通用人工智能进展的测试。据我了解,用 o3 跑的那次测试,也不是冷启动的「来,做这个」,而是「试试这个,这里有几个例子」,是少样本而非零样本。我不认为你能从它在那组谜题上的表现,有意义地泛化到任何别的东西。关于基准我还有最后一点:我们并不为人建立基准。我们设立执照考试,设立学业考试,看一个人把课程内容学得多好,为医学院做好了多少准备。但那不是给人做基准测试,不是在说这个人朝某项任务被开发到了什么程度。

布贝克: 郑重声明,你刚才说的我全都同意,绝对同意。也许只补充一小点:ARC-AGI 作为通用测试确实没有意义,这样的测试根本不存在。但这恰恰是要点。你说我们必须心里有一个功能形态,明确要把它部署到哪里,我非常赞成。我们不想要一个 AGI 跑来把所有事都解决了,我们想要的是把 AGI 部署在一个非常有限的场景里,比如辅助诊断,然后为这个有限场景开发一个基准,看它表现如何,能不能帮到医生,诸如此类。

本德: 那为什么不呢?当然可以有一个「根据症状做诊断」的基准,再为此专门构建系统,并公开训练数据。为什么不这样做?

布贝克: 因为另一种东西效果更好。这就是为什么不。

本德: 另一种东西对环境是毁灭性的,建立在窃取的数据上,建立在不公开的数据上,所以我们无法判断什么时候该用它。

布贝克: 这很难回应。

拟人化报道与「开放」

斯特里克兰: 我换个角度。这些是计算机科学家和计算语言学家之间的争论。但请从另一个视角想想,一个只是想跟上技术动向的普通人。他们看到的是这样的标题:上周《纽约时报》有一篇《数字治疗师也会有压力》,报道一项研究发现,当用户向 ChatGPT 讲述创伤经历时,它表现出焦虑的迹象,而做了一次正念练习之后焦虑水平下降了。可以想见,人们会因此产生印象,觉得那边有一个真实的心智,甚至在受苦,在这个例子里还相当悲惨。让人们相信这项技术的另一端有一个心智,有危险吗?艾米丽,先从你开始。

本德: 有。因为如果你告诉人们,这东西有答案,它能当你的治疗师,能让你好受些,能帮你诊断,能帮你处理法律麻烦,而且顺便说一句,它超级聪明,用我们从互联网上刮下来的一切训练过,它无所不知,还很客观。那么,我们就没有为人们做出好的决定创造条件。至于该拿这种新闻怎么办,我给公众的建议是:如果一个记者在拟人化,尤其是拟人化得这么厉害,跳过。你不需要读它。

布贝克: 对,这不是一篇好论文。但是,这些「但是」很重要,因为这是一个非常微妙、非常精细的话题。任何斩钉截铁的论断,「我们应该完全无视它」「它绝对不该用于某某」,我都会保持警惕。就这篇论文而言,是的,不太好。但它可能对 AI 研究者有点用,不是对公众,而是对研究者:它关乎如何提示模型,以及能否通过把模型置于「正确的心态」(打引号)来引出好的行为。虽然那里并没有心智,这一点要说清楚,但我不知道还能用什么别的语言。这也正是问题所在:我们需要在这个方向做更多工作,发展出谈论这些东西的正确词汇。我也不喜欢「人工智能」这个词,它是别的东西。它是智能,它是理解,但方式与人类不同,需要有一套自己的词汇。也许要花几十年才能发展出这个新的研究领域。而全部的难处在于,我们没有几十年。一个世纪前发展量子力学时,物理学家们做出了那么多精彩的发现,他们可以花十年、二十年,事实上量子场论花了五十多年。要点是他们有时间。今天我们没有时间,因为进步太快了。所以这整场对话也必须扎根于我之前说的那个现实:两年前,机器能解一道高中题就让我惊叹,而今天它们在角逐数学奥赛金牌。这一点必须是我们说的每一句话的背景。

本德: 那次是不是允许它提交一万个答案,然后……

布贝克: 那是谷歌,他们的做法不对。OpenAI 不是那样做的。

本德: 好吧。但 OpenAI 对自己在做什么根本不开放。我们不知道那些题目在不在训练数据里,或者非常相似的题目在不在。

布贝克: 每周有四亿人在用它,所以其实相当开放。

本德: 不。开放意味着我们知道训练数据是什么。开放意味着我们知道步骤是什么。开放意味着我们知道系统提示是什么。开放意味着我们知道你们的碳足迹是多少。那才叫开放。

布贝克: 我想萨姆(Sam Altman)暗示过,在开源和开放获取方面,我们也许没有站在历史正确的一边。所以也许很快会有这个方向的消息。

什么能让你认错

斯特里克兰: 马上就要进入观众提问,但在此之前我想问两位一个问题:要看到什么样的行为或证据,才能让你相信自己站错了边?

本德: 我需要一个非常清晰的理解定义。我需要一个实验,其所有参数都已知且公开,训练数据的每一个细节都公开,这样我们才能推理:给定这些训练数据、这个输入、这个训练架构、这个系统提示,要得到这个输出,除非有某个东西把它与理解的定义连起来,否则不可能。

布贝克: 对我来说,这非常具体。过去十五到二十年我都在做数学研究,读论文时有过许多极其快乐的时刻,从某篇论文里获得一种全新的、深刻的理解。如果十年后,或者干脆到我生命尽头,我不在乎十年不十年,但如果在我有生之年,我从未从一个 AI 系统那里获得过那样的时刻,像读另一个人也许花了数年写成的论文时那样深刻的领悟,那么,是的,我错了。

观众提问:预测下一个词

斯特里克兰: 我来读几个观众问题,有很多非常有意思、很有挑战性的。很多人好奇我们如何区分人类的理解。比如有人问:对世界的统计表征,与人类理解世界的方式有何不同?还有人问:我们做的不就是同一件事吗?我们思考的时候,不也只是在预测下一个词元(token)吗?

本德: 我想直接回应「我们不也只是在预测下一个词元」这一点。这是一个本质上去人性化的论点。它等于说,你不认为我真的有内心生活。如果你不认为我有内心生活,我不想和你进行这场对话。

布贝克: 我的看法有点类似,但我把它反过来。我不了解你是怎么运作的,你说话、传递给我理解和意义的时候,我不想对你的内部运作做任何假设。对 AI 系统也一样,我不在乎它内部是怎么运作的,那不是我的事。当然某种意义上是我的事,但就它是否理解这个问题而言不是。它在做某件事,它给了我内容,然后我自己判断这内容有没有用。

具身、多模态与文本

斯特里克兰: 另一个问题:理解和知识在多大程度上需要某种物理性?比方说,我的朋友告诉我怎么打匹克球,把所有规则和指导都讲给我,告诉我打匹克球是什么感觉,我还看了打匹克球的视频。我真的会打匹克球了吗?我理解它了吗?还是必须亲自下场,才能找到感觉?

本德: 在语言学和哲学里,这个问题有时被表述为:理解是否必须是具身的?要真正认识世界、真正学会一门语言,是否必须有身体经验?我认为具身很重要,但更重要的是处于社会情境中的关系性经验。意义真正存在于我们与他人的关系之中。如果你深入社会语言学和语言变迁的细节,会发现词语的意思一直在随我们的用法变化。这之所以发生,是因为我们成功地把词语当作小小的记号、当作线索来传达想说的东西,而这又为下一次使用这条线索时的词语着了色。所以说到底,一切都在与他人的关系里。而大语言模型根本做不到这一点。我们和它说话时可以完成自己这一半,人们也确实会产生依恋,但那全在人这一边。对一个并不存在的东西产生强烈依恋,是多么孤独。

布贝克: 我把这个问题理解为:多模态对于通用智能是否必要?我认为这是一个前沿研究问题,没有人真正知道答案。当然,这些术语没有一个是定义清楚的,但即便如此,我们也确实不知道。我个人相信,完全不需要。你可以纯粹通过文本、纯粹通过数字化地访问互联网,达到我此前讨论的那种意义上的通用智能。你不需要任何物理性,不需要看到任何东西,不需要听到任何东西,纯粹从思想和推理就能获得智能。这是我的信念,但它还没有被证实。

再补充一点:文本没有什么特别的。文本有用,只是因为它量大而且容易处理。视频和图像也有海量,只是处理起来更麻烦。技术本身对模态是完全不加区分的,可以用于视觉,可以用于音频。事实上 ChatGPT 现在能看、能听,你可以打开摄像头和它互动,会很有趣,我的孩子们就这样玩得很开心。

斯特里克兰: 有个相关的问题:你们认为 AGI(或者 AI,我不确定提问者的意思)离开人类语言,在另一种基础上,会不会更好地理解世界?

布贝克: 不会。我认为人类语言非常美妙,它的信息密度极高。在图像里,在其他模态里,有太多重复,即便音频也极其重复,也许我说的话就很重复,但在文本里,这类重复少得多。所以我实际上认为文本是必要的。换个说法:如果你想只靠在世界里互动来获得通用智能,那我看不出除了在地球上进化几十亿年、然后我们就在这儿了之外,还有什么别的办法。

本德: 这个问题同样受制于 AGI 未经定义。但我想反驳「文本没什么特别」这个说法。文本非常特别,里面有很多东西。

布贝克: 我也爱文本。我说的是,从技术的角度它不特别。

本德: 所以,如果你的视角是「它只是数据,我用神经网络这样那样地碾一遍」,那它不特别。但我认为它特别,因为文本是语言的映现,而语言是我们建构社会世界的方式,是我们彼此互动的方式,是我们进行某种事实上的心灵感应、进入另一个人心智的方式。如果现在有系统把合成文本推到我们面前,倾倒进我们的信息生态,那是个巨大的问题,因为文本是特别的。

财富、权力与涌现

斯特里克兰: 这是得票最高的问题,我必须问:财富和权力在「谁来裁定大语言模型是否理解」这件事上扮演什么角色?

本德: 难题。谁来裁定?你们每个人都可以自己做决定,这正是今晚这类活动的意义。但就政治话语的走向而言,谁在裁定?我们是否在基于「大语言模型能理解」的假设制定政策?如果是,那财富和权力就深深卷入其中。更进一步,围绕开发这些东西的政治经济,正在进一步集中财富和权力。数据变成了权力。我想指出,你刚才说你那个问题只在你自己的 Dropbox 里。Dropbox 是不是和 OpenAI 签了协议的公司之一?

布贝克: 我不认为我的 Dropbox 被发给了 OpenAI。

本德: 我不知道,因为你们不告诉我们数据里有什么。

布贝克: 好吧。谁来裁定?你们裁定,这肯定是正确答案。我想我们也得承认,我们生活在一个由资本主义社会运转的世界里,大家都想领先,想更快更好地造东西。说到底,这些是工具,它们就是工具。如果它们让你更有生产力,让你更快实现想法,让你用更短的时间完成工作,那它们就会以这种方式扩散开来。所以这个决定是分布式的,没有哪一个单一实体来做决定。

斯特里克兰: 也许这是最后一个问题,然后进入结辩。关于涌现(emergence):当一个复杂实体拥有其组成部分各自不具备的属性或行为时,就发生了涌现。两位怎么看这个概念?人类意义上的理解,会不会是可以涌现出来的东西?

本德: 当我谈理解时,我说的是对语言的理解。它是否在人类进化的某个时刻涌现出来了?显然是的,因为我们在某个时刻从没有语言的生物变成了有语言的生物,理解语言的能力涌现了。但说到机器学习系统的涌现行为,我再说一次,除非你能完全开放地访问训练数据,否则这些主张是不科学的。而且我们必须小心,我们是如何把输入输出对解读成问答对的,还要把它与其他可能从那个输入走到那个输出的计算方式做比较,然后才能提出「涌现行为」这样非同寻常的主张。

布贝克: 不,我认为涌现绝对是真实的。老实说,我很难理解这怎么会有争议。我也要反驳「必须知道系统的每一个细节才能判断是否有涌现」这一点,理由是你可以自己做实验。在投身构建大语言模型之前,我自己就在神经网络上做实验,独自搭建玩具设定,只用自己的电脑,就能创造出看到涌现现象的场景。为了让大家在同一页上,举一个我看过的例子:训练一个模型解线性方程组。还记得中学时那种题吧,x 加 y 等于多少,x 减 y 等于多少,求 x 和 y。你能看到这样的设定:系统只从少量方程的方程组中学习,突然就能解很多方程的方程组。这就是你在实验室里亲眼能看到的涌现,你不需要接触最好的模型。

结辩

斯特里克兰: 好,请两位做三分钟左右的结辩。艾米丽,你先开始的,就先请你。

本德: 今晚我想留给大家什么?我想留给大家这样一个想法:没有什么是不可避免的。如果有人说「这东西来了就不会走,我们必须学会与它共存」,你可以说不。拒绝非常重要。这在我们那些已经吱吱作响、并且即将响得更厉害的系统里尤其重要:教育、医疗、司法、移民程序。在所有这些地方,合成文本看起来像一块方便的创可贴,一个速效方案,因为老师不够、治疗师不够,诸如此类。我们需要对此说不,因为它其实比什么都不做更糟。而说不,我认为要从拒绝开始:它不只是一个好玩的玩具,它不是一个好的搜索工具。任何时候有人说「来,用这个,这是人工智能」(这里我打引号),记住,你后口袋里揣着一个「不」。我就说到这里。

布贝克: 我的看法很不一样。总的来说,我认为由你们自己决定,而我反对恐吓策略,反对说「这很可怕,这只会带来坏事」。我们都足够聪明,在和这些工具互动、使用它们时,能自己判断它们是否带来价值。这是第一点。

第二点,我想回到一件我本以为会在讨论中更多出现的事,伊莱莎开头提到的那些谜题:稍微改一点,答案就全错了。这回到我一直想说的观点:这些话题复杂而且相当微妙,我认为鹦鹉和火花里都有真相,两者兼而有之。在某些话题上它们更像鹦鹉,在某些话题上则多一点火花。我看到的是,GPT-3 时代也许更偏鹦鹉,GPT-4 也许更偏火花,我看到天平在移动。但两种模式肯定都还在。所以这是一个非常非常复杂的话题。

说了这么多,眼下确实有一股浪潮,越来越多的人在与这些模型互动,并意识到它们有多强大。有一件事让我很受启发:萨姆在推特上发了我们最新模型写的一篇小说,就发了那么一条,然后一些知名作家、创意写作者表示,他们被那个故事打动了。当然,我知道你要说什么,你只可能被另一个人写的东西打动,等等。可以。但这个人,一位知名作家,说自己被打动了。所以我认为这股浪潮在持续上涨。它会涨到多高?没人知道。如果有人告诉你他知道,也许你不该太听他的。但我们能看到速度,这个速度令人惊叹。我肯定属于那个阵营:非常期待看到未来几年会带来什么。

现场投票

斯特里克兰: 让我们感谢两位辩手。我真的想不出还有谁更适合来呈现这两种观点。大家还记得 Slido 投票吧,开场时请大家投过一次。现在是辩后投票,请拿出手机,告诉我们你现在怎么想,指针有没有移动。有请戴维回到台上做总结。

主持人: 非常感谢各位。回到开场时的数据,「否」一开始是 64%,现在又上升了。不过看,就在我走上台的这会儿,「是」的一方还在投票。我得说,这相当令人印象深刻。开场时「是」占 22%。现在似乎又在往回走。所以我不知道,算是平手吗?如果没有人改变主意,那也说明了点什么。请为我们的主持人和两位讲者热烈鼓掌。我们很快得再办一场别的题目的辩论,今晚太精彩了,现场能量十足,YouTube 聊天室里也一样热烈。感谢所有参与的线上观众,再次感谢两位讲者,感谢各位。晚安。

本期讲者
艾米莉·本德华盛顿大学语言学教授、计算语言学实验室主任。2021 年论文《随机鹦鹉的危险》合著者,2020 年「章鱼论文」提出仅凭语言形式无法习得意义;2025 年与 Alex Hanna 合著《The AI Con》。
塞巴斯蒂安·布贝克OpenAI 技术团队成员,此前任微软 AI 副总裁兼杰出科学家,曾任普林斯顿助理教授。2023 年领衔发表《Sparks of AGI》,主张 GPT-4 展现出通用智能迹象,并主导小模型 Phi 系列。
Eliza StricklandIEEE Spectrum 资深编辑,长期报道 AI 与生物医学技术,本场辩论主持人。
章节 · 点击跳转视频
0:00 开场致谢与两位辩手介绍 ▶ 正在看
4:53 从图灵测试到大推理模型 ▶ 正在看
13:16 界定「真正理解」与辩论规则 ▶ 正在看
16:04 Bender:形式不通向意义 ▶ 正在看
27:06 Bubeck:基准饱和与进展速度 ▶ 正在看
35:42 研究者调查与 AGI 无定义之争 ▶ 正在看
40:44 为什么要在乎「理解」 ▶ 正在看
43:42 ARC 测试与基准的局限 ▶ 正在看
51:24 拟人化报道、开放性与认错条件 ▶ 正在看
56:40 观众问答:具身、文本与权力 ▶ 正在看
1:03:39 涌现是真实还是不科学 ▶ 正在看
1:06:58 结辩与现场投票 ▶ 正在看
本期论点
本期回应
17:08
没有对意义的接触,就没有理解 不能光靠文字,能学会语言的意思吗?艾米莉·本德
59:01
比具身性更重要的是社会情境中的关系性经验,意义存在于人与人的关系之中 不能光靠文字,能学会语言的意思吗?艾米莉·本德
1:00:09
通用智能不需要具身性,纯靠文本与对互联网的数字访问就能达到 能光靠文字,能学会语言的意思吗?塞巴斯蒂安·布贝克
1:00:37
从技术角度看文本并不特殊,它有用只因量大且易处理,同样的技术对模态无所谓 能光靠文字,能学会语言的意思吗?塞巴斯蒂安·布贝克
1:02:28
文本是特殊的,因为它是语言的映射,而语言是人类构建社会世界、进入彼此内心的方式 不能光靠文字,能学会语言的意思吗?艾米莉·本德
25:08
大语言模型看起来像在理解,是因为人类替它把话说通了 人在补全机器能真的理解吗?艾米莉·本德
32:43
理解不是二元的有或无,而是一个连续的程度量 正在长出机器能真的理解吗?塞巴斯蒂安·布贝克
53:51
大语言模型确实具有智能与理解,只是方式与人类不同,需要一套自己的词汇 正在长出机器能真的理解吗?塞巴斯蒂安·布贝克
1:09:04
「随机鹦鹉」和「智能火花」两种说法都有真实成分,大模型是两者的混合 正在长出机器能真的理解吗?塞巴斯蒂安·布贝克
12:16
对熟知的经典谜题做极简单的改动,语言模型往往就会答错 看它怎么错怎么判断机器是不是真会一件事?Eliza Strickland
26:30
文本进、文本出的测试只能检验模型对训练数据词语分布的拟合,不能检验理解 看它怎么错怎么判断机器是不是真会一件事?艾米莉·本德
50:20
把 ARC-AGI 当通用测试说不通,通用测试根本不存在,基准应针对受限的部署场景来做 看它怎么错怎么判断机器是不是真会一件事?塞巴斯蒂安·布贝克
1:05:14
除非能完全公开访问训练数据,否则关于机器学习系统涌现行为的说法是不科学的 超出理解神经网络里发生了什么,人能看懂吗?艾米莉·本德
1:05:50
不必知道系统的每一个细节也能判断涌现是否存在,用玩具设定自己做实验就能验证 可以看懂神经网络里发生了什么,人能看懂吗?塞巴斯蒂安·布贝克
其他论点
1:07:45
在教育、医疗、司法、移民等系统里用合成文本当快速方案,比什么都不做更糟 艾米莉·本德
01开场致谢与两位辩手介绍
0:00
[Music]
[音乐]
便签引用
0:17
Well, good evening everyone. What a big crowd. It's great to have you all here for the great chatbot debate. And thank you so much to everyone watching the live stream across the country and across the world and the recording. It's great to have you all part of the CHM community wherever you stand on the issue tonight. I just want to say whether you're AI true believer, you're a skeptic, you have no idea what this whole LLM thing is, tonight promises to be a really interesting and vibrant event. We are so grateful though to our sponsor for tonight's evening.
晚上好,各位。今晚人真多。很高兴大家都来参加这场关于聊天机器人的大辩论。也非常感谢全国各地、世界各地通过直播和录播收看的朋友们。无论你今晚站在哪一方,很高兴你们都是CHM 社区的一员。我想说,无论你是 AI 的坚定信徒,还是怀疑论者,或者你根本不知道 LLM 是个什么东西,今晚都会是一场非常有意思、非常热烈的活动。我们非常感谢今晚活动的赞助方。
便签引用
0:56
And I just want to take a moment to offer our sincere thanks to the Patrick J. McGovern Foundation. And won't you join me in thanking them as [Applause] well? Their generosity makes possible what we do, as well as all of you who are supporting members. I know we had a great crowd at the member reception tonight. I just want to say a special thank you as well to the House Family Vineyards for sponsoring the wine at the reception tonight. So, thank you to them. And indeed, tonight it is such a pleasure. I know quite a few of you probably are new to the museum. You may have heard about this through our event partner ILE E Spectrum. We are so grateful to have the chance to work together with it e Spectrum. It's a great partnership and uh this whole event has been a real collaborative process. So, our thanks to Eliza and her team uh there. So, I just want to uh take a moment to reflect on where we are tonight uh here at CHM. Raise your hand if this is your first time at the museum. Wonderful. Well, that's great.
我想花一点时间,向 Patrick J. McGovern 基金会致以我们诚挚的谢意。请大家和我一起感谢他们,好吗?[掌声] 他们的慷慨支持让我们所做的一切成为可能,当然还有在座各位支持我们的会员。我知道今晚的会员招待会上来了很多人。我还想特别感谢 House Family Vineyards 赞助了今晚招待会上的葡萄酒。所以,谢谢他们。今晚真的非常荣幸。我知道你们当中不少人可能是第一次来博物馆。你们可能是通过我们的活动合作伙伴 IEEE Spectrum 得知这场活动的。我们非常感激能有机会和 IEEE Spectrum 合作。这是一次很棒的合作,整场活动都是真正协作的成果。所以,感谢 Eliza 和她的团队。我想花点时间聊一聊,我们今晚在 CHM 这里所处的位置。如果这是你第一次来博物馆,请举手。太好了。真棒。
便签引用
2:07
Hopefully, some of you uh will be back if you haven't already. Uh make plans to come back to check out the chatbot exhibit that we have downstairs. Chatbots decoded. Uh if you weren't here for the member reception, didn't are not a member yet, consider joining for next time, but but make sure in the meanwhile to come back and check out that exhibit because it brings to life uh the topics that we're talking about tonight. It's it's a really fun interactive exhibit. You can talk with AMA uh and a language of your choice, a robot, and uh it's it's got a lot of great exhibits. So do come back if you haven't seen it already. Now, at CHM, we're all about decoding technology. It's computing past, digital, present, and future impact on humanity. And tonight's program obviously really reflects that theme as we explore and debate just how intelligence AI really is. I just want to take a moment here and introduce the experts that we have guiding us through this debate tonight. We are so lucky indeed to have quite a powerhouse on
希望你们当中有些人还会再来,如果还没来过的话。请安排时间再来看看我们楼下的聊天机器人展览。《聊天机器人解码》。如果你没参加会员招待会,或者还不是会员,可以考虑下次加入,但同时一定要再来看看那个展览,因为它把我们今晚讨论的话题活生生地呈现了出来。这是一个非常有趣的互动展览。你可以用你选择的语言和 AMA、和一个机器人对话,里面还有很多很棒的展项。所以如果你还没看过,一定要再来。在 CHM,我们的宗旨就是解码技术——计算的过去、数字化的当下,以及未来对人类的影响。今晚的节目显然很好地体现了这个主题,我们将探讨并辩论 AI 到底有多智能。我想借此机会介绍一下今晚引导我们进行这场辩论的专家。我们非常幸运,双方阵容都堪称重量级。代表其中一方的是
便签引用
3:13
both sides. Uh representing one side is the University of Washington's Emily M. Bender. She is the director of the computational linguistics laboratory and is a professor of linguistics and an adjunct professor at the school of computer science and engineering and the information school there. She is known for her critical perspective on AI language models, co-authoring the paper on the dangers of stoastic parrots, and her new book, The AI Con: How Big How to Fight Big Tech Hype and Create the Future We comes out May 13th. So, be sure to pre-order that. Now, on the opposing side, we have OpenAI's Dr.
华盛顿大学的 Emily M. Bender。她是计算语言学实验室的主任,语言学教授,并在该校计算机科学与工程学院和信息学院担任兼职教授。她以对 AI 语言模型的批判视角而闻名,是《随机鹦鹉的危险》一文的合著者,她的新书《The AI Con:如何对抗大科技公司的炒作、创造我们想要的未来》将于 5 月 13 日出版。所以,敬请一定要去预订。那么,在反方这边,我们请到的是 OpenAI 的
便签引用
3:53
Sebastian Bubck, currently a member of OpenAI's technical staff. He previously served as VP of AI and distinguished scientist at Microsoft where he spent a decade at Microsoft research. Earlier he was an assistant professor at Princeton. His 2023 paper, Sparks of Artificial General Intelligence, early experiments with GPT4, drove widespread discussion and debate about AI in publications like the New York Times and Wired. And the ring master in the octagon tonight, or I should say tonight's debate moderator, is Eliza Strickland. Indeed, senior editor at ILE E Spectrum, where she covers AI, biomedical technology, and other advanced technologies. She is often on podcasts and moderating programs from South by Southwest to here at CHM tonight. So, without any further ado, I'm going to turn it over to her to set the stage for tonight's debate.
Sebastian Bubeck 博士,他目前是 OpenAI 技术团队的成员。此前他曾担任微软的 AI 副总裁兼杰出科学家,在微软研究院工作了十年。更早之前,他是普林斯顿大学的助理教授。他 2023 年的论文《通用人工智能的火花:GPT-4 的早期实验》引发了广泛的讨论和争论,《纽约时报》《连线》等刊物都作了报道。而今晚八角笼里的裁判长,或者我应该说,今晚这场辩论的主持人,是 Eliza Strickland。她是 IEEE Spectrum 的资深编辑,负责报道 AI、生物医学技术和其他前沿科技。她经常出现在播客上,也主持过各种节目,从西南偏南一直到今晚的 CHM。那么,闲话少说,我把台子交给她,请她为今晚的辩论做个开场。
便签引用
02从图灵测试到大推理模型
4:53
Won't you join me in giving her a big CHM welcome? Thank you. All right. Hello everybody. Thank you so much for joining us here today. This is going to be a lot of fun. Um I'm going to take one minute to plug it e Spectrum before I get to the real good stuff. Um, just in case you guys don't know us, we are the flagship publication of the ITLE E which is the which is a professional organization for electrical and computer engineers with 500,000 members around the world. So, we have a monthly magazine, but we also have a website that's free and open to all. So, all sorts of great technology coverage there. Now, when I e Spectrum started talking with the museum about doing an event, doing this debate, I knew exactly what I wanted to have on stage. I wanted Sparks versus parrots.
大家和我一起,用热烈的掌声欢迎她好吗?谢谢。好的。大家好。非常感谢各位今天来到现场。今晚会非常有意思。嗯,在进入正题之前,我想花一分钟时间给 IEEE Spectrum 打个广告。嗯,以防大家不了解我们,我们是 IEEE 的旗舰刊物,IEEE 是一个面向电子与计算机工程师的专业组织,在全球有 50 万名会员。我们有一本月刊,同时也有一个免费向所有人开放的网站。上面有各种各样精彩的科技报道。那么,当 IEEE Spectrum 开始和博物馆商量办一场活动、办这场辩论时,我心里非常清楚我想请谁上台。我想要的是“火花”对“鹦鹉”。
便签引用
5:42
So, so for the no position, I thought there'd be no one better than Emily Bender. Um, her paper on the danger of sarcastic parrots uh was written several years before a chatgbt came out. Um, and it starts a firestorm of of debate about the risks of language models that was uh really precient. Um, and I talked to her over the years and one thing she told me in a in an interview really stuck in my mind when I when I thought about this question of whether language models can understand. uh she said the way we interpret language is imagining the mind of the person speaking to us. So when we encounter synthetic text we're imagining a mind that isn't there. I thought that was so such an interesting perspective.
所以,反方这边,我觉得没有谁比 Emily Bender 更合适了。嗯,她那篇关于随机鹦鹉之危险的论文,是在 ChatGPT 出现好几年前就写出来的。嗯,它点燃了一场关于语言模型风险的大讨论,现在回头看真的很有预见性。嗯,这些年我和她聊过,她在一次采访里跟我说的一句话一直印在我脑子里,尤其是当我思考语言模型到底能不能理解这个问题的时候。呃,她说,我们理解语言的方式,是去想象对我们说话的那个人的心智。所以当我们遇到合成文本时,我们是在想象一个根本不存在的心智。我觉得这个视角实在太有意思了。
便签引用
6:21
Um and for the yes position I wanted Sebastian Bubc who had written this preprint sparks of artificial general intelligence um which shared these early experiments with the language model GPT4. And um I'd heard him on a podcast around the time of that that preprint came out talking about just the differences he'd found between the prior language models that OpenAI had come out with and this new version GPT4. And he said the big issue compared to the so sorry comparing the earlier models to GPT4. He said the big issue is that the other the earlier models didn't really get the concept of a cat or a dog. And then suddenly with GPT4 it was clear to me that it really understood something.
嗯,而正方这边,我想请的是 Sebastian Bubeck,他写了《通用人工智能的火花》这篇预印本,嗯,里面分享了对语言模型 GPT-4 的这些早期实验。还有,嗯,在那篇预印本发表前后,我在一个播客上听过他讲,讲他发现 OpenAI 之前发布的那些语言模型和新版本 GPT-4 之间的差别。他说,最大的问题是——抱歉,是把早期的模型和GPT-4 作比较。他说,最大的问题在于,早期那些模型并没有真正掌握“猫”或者“狗”这个概念。然后突然之间,到了 GPT-4,我很清楚地看到,它是真的理解了某些东西。
便签引用
7:01
So, I thought this would be these would be the perfect debaters and we're so lucky that they both said yes to joining us on stage. Um, so I'll just give you a quick bit of history, a little bit of how these technologies work and a little bit of what's going on right now. So, you probably heard of the Turing test in the 1950s. Alan Turing, the famous computer scientist, came up with this sort of thought experiment. He said, "What if there was a human evaluator who was had a computer interface and was chatting with both a human and a computer? um if he couldn't tell a difference between the two, that would mean that the machine was at least was imitating humans well enough to at least uh demonstrate intelligent behavior. He didn't quite outfat say it would be intelligent, but he said it would be demonstrating intelligent behavior. And back then, the idea that a computer could pass the test was considered a total fantasy. Um but time went on in 1966, an MIT professor named Joseph Weisenbomb created the first
所以我觉得,这两位会是最完美的辩手,而我们非常幸运,他们都答应上台参加我们的活动。嗯,那么我先快速讲一点历史背景,讲一点这些技术是怎么运作的,再讲一点刚才那部分现在正在发生的事。你们大概都听说过图灵测试,1950 年代,阿兰·图灵,这位著名的计算机科学家,提出了这样一个思想实验。他说:"如果有一个人类评估者,面前有一个计算机界面,同时和一个人以及一台计算机聊天呢?嗯,如果他分辨不出两者的区别,那就意味着这台机器至少模仿人类模仿得足够好,至少表现出了智能行为。他并没有直接说它就是有智能的,但他说它表现出了智能行为。而在当时,计算机能通过这个测试的想法被认为完全是天方夜谭。嗯,但时间往前走,1966 年,一位名叫约瑟夫·维森鲍姆的麻省理工教授做出了第一个聊天机器人,巧合的是它的名字叫
便签引用
7:54
chatbot, which was by coincidence named Eliza. and she mimicked a therapist by reflecting your statements back to you. So if you said something like, "I'm really stressed out at work." She'd say, "Oh, why are you stressed out at work?" And it were like, "I'm really mad about my mom to tell me more about why you're mad about your mom." And it's just, you know, a basic scripts. But the interesting thing was that people um attributed intelligence and emotion to this these basic scripts. So uh Weisenbomb secretary started asking him to leave the room when she was talking with Eliza because she wanted to have this privacy for these intimate conversations. So then came decades of slow progress. I won't run you through the whole history of AI. Um wasn't until the 2010s that we got sort of halfway decent voice assistants like Siri and Alexa, but they were often more frustrating than useful. And then in the 2010s, we saw the beginning of a revolution in AI when a type of AI called artificial neural networks that
Eliza(伊莉莎)。她模仿一位心理治疗师,把你说的话反射回给你。所以如果你说"我工作压力真的很大",她会说:"哦,你为什么在工作上压力很大呢?"你说"我真的很生我妈妈的气",它就说"跟我多说说你为什么生你妈妈的气"。它其实就是,你知道的,一些很基础的脚本。但有意思的是,人们竟然把智能和情感赋予了这些基础脚本。所以,呃,维森鲍姆的秘书在和 Eliza 对话时开始要求他离开房间,因为她想要这种私密感来进行这些亲密的对话。接下来是几十年缓慢的进展。我不会把整个人工智能史都讲一遍。嗯,直到2010 年代,我们才有了 Siri 和 Alexa 这类勉强还行的语音助手,但它们常常是让人恼火多过有用。然后在2010 年代,我们看到了 AI 革命的开端,一类叫作人工神经网络的 AI——它曾长期被当作死
便签引用
8:48
had been sort of written off as a dead end for ages uh suddenly had this resurgence. Um so basically neural nets are a type of machine learning model that's loosely inspired by how the brain works. Um they cons they consist of layers of internet connected nodes or neurons which process information and the connection between neurons has a weight that adjusts as the network learns from the data. Um so in in in neuroscience there's a phrase uh neurons that fire together wire together. So the idea is um information um the weights that the parts that are most interested and interconnected those weights get strengthened as a machine learns from the data. And what changed in the 2010s was that suddenly these machines had vast amounts of data to learn from. They could ingest all the pictures on social media. they could ingest all the text on Reddit and social media and books and so they took a huge leap in capability uh first in computer vision and then uh in text processing natural language processing so that fought on the dawn of
胡同而被搁置——呃,突然间复兴了。嗯,基本上,神经网络是一类机器学习模型,它松散地受到大脑运作方式的启发。嗯,它们由一层层相互连接的节点或者说神经元构成,这些节点处理信息,而神经元之间的连接有一个权重,随着网络从数据中学习而不断调整。嗯,在神经科学里有一句话,呃,神经元一起放电就会连在一起。所以这个想法是,嗯,信息,嗯,那些最相关、联系最紧密的部分,它们的权重会随着机器从数据中学习而被强化。而 2010 年代改变的地方在于,这些机器突然有了海量的数据可以学习。它们可以吞下社交媒体上所有的图片,可以吞下 Reddit、社交媒体和书籍里的所有文本,于是它们的能力有了巨大的飞跃,呃,先是在计算机视觉,然后是在文本处理、自然语言处理上,这就带来了
便签引用
9:52
the large language models which are that's what LLM stands for large language models and they're built on neural networks and they can um reproduce humanlike text because they've read so much of it they've read basically all the internet. So, um during the training process in which they're ingesting all this data, they learn to recognize patterns in the data uh such as how words and phrases are commonly used together. And that allows the model to predict what comes next in a sentence or to generate, you know, relevant responses based on the input you give it, the questions you ask or the prompts you give it. So, that's what enables an LLM to explain quantum physics one minute and write a poem in the style of Dr. Seuss the next minute.
大语言模型的黎明——LLM 就是 large language models(大语言模型)的缩写——它们建立在神经网络之上,它们能,嗯,生成像人写的文本,因为它们读了太多太多,基本上把整个互联网都读了。所以,嗯,在训练过程中,它们吞下所有这些数据,学会识别数据中的模式,呃,比如词和短语通常是怎么搭配使用的。这让模型能够预测一句话里接下来会出现什么,或者生成,你知道的,根据你给的输入、你问的问题或者你给的提示词,生成相关的回应。所以这就是为什么大语言模型可以这一分钟解释量子物理,下一分钟写一首苏斯博士风格的诗,
便签引用
10:34
and uh you know write a write a recipe for eggplants the next minute. So it can it's seen all these things and it can basically do anything it's seen before and and then some and the then some is is the interesting part. Um but as you may have seen large language models are often prone to really embarrassing missteps. um like the ju the Google chatbot a year or so ago that advised people to eat at least one small rock per day u for good health and to keep pizza toppings from falling off by adding an eighth of a cup of glue. Um so when you see things like that you say well what's really happening here?
再下一分钟,你知道的,写一份茄子的食谱。所以它能——它见过所有这些东西,它基本上能做任何它见过的事情,甚至更多一点,而这"更多一点"才是有意思的部分。嗯,但正如你们可能看到的,大语言模型常常会出现非常尴尬的失误。嗯,比如一年多前谷歌的聊天机器人建议人们每天至少吃一小块石头,呃,对健康有好处;还说要防止披萨上的配料掉下来,可以加八分之一杯胶水。嗯,所以当你看到这样的事情,你会说:这里到底发生了什么?
便签引用
11:11
Um and um now a little bit about what's happening today. Um just in the past few months we've seen the introduction of powerful new models that are called large reasoning models. So this is kind of a step up from the LLMs and these models use chain of thought reasoning to break down a problem into many steps. So uh they they break it down into steps and they execute them one at a time and that seems to reduce errors. It makes them um less prone to these sort of hallucinations or or you know missteps.
嗯,然后,嗯,说一点今天正在发生的事。嗯,就在过去这几个月,我们看到了强大的新模型出现,它们被称为大推理模型。所以这算是在大语言模型基础上又进了一步,这些模型使用思维链推理,把一个问题拆解成很多步骤。所以,呃,它们把问题拆成一步步,然后一次执行一步,这似乎能减少错误。它让它们,嗯,更不容易出现那种幻觉,或者说,你知道的,失误。
便签引用
11:40
Um and it really is it's interesting to see what a change it makes. I was just testing at this today just for fun. Um, and maybe you've heard before there's this sort of famous riddle where you say, you know, there's a a man and a cabbage and a goat and a wolf and you have to get them all across the river and you have a boat that takes you plus one item. How do you do it? How do you do it one at a time so the the cabbage doesn't get eaten by the goat and the goat doesn't get eaten by the wolf? Now, what's interesting is that um, language models have this down cold because they've seen it before. You can ask them that version. You could ask them about a mouse and a cheese and a cat. like they get it, they're good. But if you tweak the riddle in in very simple ways, you can you can uh often trip up the language models. So I asked Chat GPT just today how to get just a man and a goat to the other side of the river.
嗯,看到它带来的变化真的挺有意思。我今天还刚好测试了一下,纯粹是好玩。嗯,你们可能之前听过一个很有名的谜题,就是说,你知道的,有一个人、一棵白菜、一只羊和一头狼,你得把它们全都送过河,而你的船每次只能载你加一样东西。你要怎么做?怎么一次一样地运,才能让白菜不被羊吃掉,羊不被狼吃掉?有意思的是,嗯,语言模型对这个题目烂熟于心,因为它们见过。你可以问它们这个版本,也可以换成老鼠、奶酪和猫来问,它们都懂,答得很好。但如果你用非常简单的方式改动一下这个谜题,你往往就能,呃,把语言模型绊倒。所以我今天问了 ChatGPT,就问怎么把一个人和一只羊送到河对岸。
便签引用
12:27
They've got a boat. How do they get across? And he said, it told me the man takes the goat across the river. He leaves the goat on the other side. The man returns alone to the other side to retrieve the boat. Both the man and the boat are now on the original side while the goat is safely across the river. I said, "Well, I've pointed out that now the man in the goat would the man was not on the right side of the river and it said, "Oh, I'm sorry. The man and the boat should both get in the boat and cross the river together. The man leaves the goat on the far side and returns alone with the boat to the original side. The man then rose across the river again by himself." So, you see, it's it's like trying to put the pieces together, but it just doesn't quite get it. Um, and then I asked OpenAI's reasoning model the same exact question and it said just oh just put the man in the goat in the boat and put him across the river and call it a day. So you can see something is changing in these in the in the way
他们有一条船,他们怎么过河?它说——它告诉我:这个人带着羊过河,他把羊留在对岸。然后这个人独自返回对岸去取船。现在人和船都在原来那一岸,而羊已经安全地在河对面了。我说,嗯,我指出现在这个人和羊——这个人并不在河的正确一侧,然后它说:"哦,抱歉。人和船应该一起上船过河。这个人把羊留在对岸,然后独自带着船返回原来那一岸。接着这个人再次自己划过河去。"所以你看,它像是在努力把碎片拼起来,但就是差那么点意思。嗯,然后我拿同样的问题去问 OpenAI 的推理模型,它就说:哦,把人和羊一起放到船上,划过河去,搞定收工。所以你可以看到,这些模型
便签引用
03界定「真正理解」与辩论规则
13:16
these um these models can function. And another thing that's happening today is AI agents like agentic AI you may have heard of. So these are systems built on top of language models that can then take action on your behalf. So they can um find make a restaurant reservation for you or find the best flight that goes to Paris or you know research the cheapest shoes to buy on Amazon. So in the context of all that we have the question before us today which is do LLM really understand? And we should think for just a second about what we mean by really understand because definitions are important in this. Um when we talk about understanding in human terms, we generally mean that someone has a deep intuitive grasp of a concept uh which involves not just knowing facts or being able to repeat information but being able to apply knowledge in various contexts to have a sense of meaning and intent and connect the ideas to a world model a model they have of the world.
运作的方式正在发生某种变化。今天还在发生的另一件事是 AI 智能体,你们可能听过所谓的 agentic AI。所以这些是建立在语言模型之上的系统,能够代表你去采取行动。它们可以,嗯,帮你订餐厅、找去巴黎最划算的航班,或者,你知道的,在亚马逊上研究最便宜的鞋子。所以在这一切的背景下,我们今天面前的问题是:大语言模型真的理解吗?我们应该花一点时间想想,我们说"真的理解"是什么意思,因为定义在这里很重要。嗯,当我们谈论人类意义上的理解时,我们通常指的是某人对一个概念有深刻的、直觉性的把握,呃,这不仅仅是知道事实、能够复述信息,而是能够在各种情境中运用知识,有对意义和意图的感知,并且把这些想法与一个世界模型——他们对世界的模型——联系起来。
便签引用
14:12
Now when it comes to large language models, the question is whether these AI systems possess a similar type of understanding. Um so yeah, are they do they really understand what they are saying? Are they simply mimicking the patterns they have seen in the training data without any true comprehension? And this is the crux of the debate before us today. If LLM truly understand, it would change the way we interact with technology as we make AI a more integral part of our daily lives. On the other hand, if they're merely mimicking patterns, we might have want to be more cautious about how we use them, how we integrate them, how how we rely on them. So, we're going to have this excellent debate in front of us. The format will be opening statements from each of the speakers, each of the debaters, followed by a moderated debate in which I will ask them questions, they will ask questions of each other, and then we'll do Q&A, taking questions from the audience, both in person and online.
那么说到大语言模型,问题就在于这些 AI 系统是否具备类似的那种理解。嗯,所以,是啊,它们真的理解自己在说什么吗?还是说它们只是在模仿它们在训练数据中见过的模式,而没有任何真正的领会?这就是我们今天这场辩论的核心。如果大语言模型真的理解,那么随着我们让 AI 越来越深地融入日常生活,它会改变我们与技术互动的方式。反过来说,如果它们只是在模仿模式,那我们可能就需要更谨慎地对待我们如何使用它们、如何整合它们、如何依赖它们。所以,我们接下来会有一场精彩的辩论。形式是先由每位讲者、每位辩手做开场陈述,然后是有主持的辩论环节,我会向他们提问,他们也会互相提问,接着我们做问答环节,接受现场和线上观众的提问。
便签引用
15:09
And then we'll end with closing statements. And at the end of the debate, we ask you to uh participate in the slido poll again. Um yeah, if you haven't done so already, it's great to do it now so we can get a a feel for where the audience stands at the beginning of this. And then we'll check in again at the end. Um as you listen to the to the arguments, just you know, consider both perspectives and keep an open mind and try to think critically about uh the evidence and the reason reasoning presented. Um this really is not you know technical debate but also a really philosophical one that touches on the nature of intelligence and understanding and you know what it means to be human and what it means to be human in this age of machine interaction. So without further ado I would like to welcome our audience our activators onto the stage Emily and Sebastian. Remind me did you want to go first or second? You want first? Right.
最后我们以结辩陈词收尾。辩论结束时,我们会请你们再参与一次 Slido 投票。嗯,是的,如果你还没投过,现在投最好,这样我们能感受到观众在开场时的立场。然后我们会在结束时再看一次。嗯,在你们听这些论点的时候,就,你知道的,同时考虑两种视角,保持开放的心态,试着批判性地思考,呃,所呈现的证据和推理。嗯,这其实不只是一场技术辩论,也是一场非常哲学性的辩论,它触及智能与理解的本质,以及,你知道的,做人意味着什么,在这个人机互动的时代做人又意味着什么。那么话不多说,我想请我们的辩手上台,欢迎 Emily 和Sebastian。提醒我一下,你想第一个还是第二个讲?你想第一个?好。
便签引用
04Bender:形式不通向意义
16:04
Great. So, back in the green room, we flipped a coin and Emily won the won the coin toss and was elected to give the first opening statement. Great. Thank you so much. Thank you all for being here. Thank you to the Computer History Museum for organizing this and thank you to the folks who are joining us online either live or in the future. Um, to answer this question in my starting 10 minutes, I'm going to do four things and I'm going to do them drawing on what I've learned studying, researching, and teaching linguistics and computational linguistics since the 1990s. And linguistics is really important here because linguistics is the field that studies how language works and how we work with language. And both of those things are really crucial to answering this question of do LLMs really understand? So, I said I'm going to do four things. The first is I'm going to offer another definition of what understanding means. Then I'm going to tell you a bit about what language models are and how they've developed
太好了。所以在后台休息室,我们抛了硬币,Emily 赢了这次抛硬币,于是由她来做第一个开场陈述。太好了。非常感谢。谢谢大家来到这里。感谢计算机历史博物馆举办这场活动,也感谢在线上加入我们的朋友,无论是现在直播还是以后回看。嗯,为了在我开头的十分钟里回答这个问题,我要做四件事,而我会基于自 1990 年代以来我在学习、研究和教授语言学与计算语言学中所获得的认识来做这些事。语言学在这里非常重要,因为语言学正是研究语言如何运作、以及我们如何使用语言的学科。而这两件事对于回答"大语言模型真的理解吗"这个问题都至关重要。所以,我说我要做四件事。第一件是我要提出另一个关于"理解"是什么意思的定义。然后我要讲讲语言模型是什么,以及它们这些年是怎么发展起来的。然后我要提出这样一个论点:语言模型的发展过程中,没有任何东西
便签引用
16:59
over the years. And then I'm going to lay out the argument that nothing in the development of language models actually gives them access to the meaning part of the form meaning equation that is language. And without access to meaning there's no understanding. And then I'll talk a little bit about how people understand language and why that mechanism leads us to sort of be susceptible to the illusion that LLMs understand. As I'm doing this, I'm going to try to avoid using anthropomorphizing language. I'm not going to say that large language models struggle or try to do anything. I'm not going to say that they ask and answer questions. Um or I'm going to try my best not to. Um because I think that muddies the waters. Um, and similarly, I will never use the term artificial intelligence or AI without scare quotes. Um, because it doesn't refer to a well- definfined set of technologies. Um, I think in general it is better to always talk about automation and then we can talk about what it is that we're automating and why
真正让它们接触到了'形式—意义'这个方程里属于意义的那一半,而语言正是这样一个方程。没有对意义的接触,就没有理解。接着我会稍微讲一讲人类是如何理解语言的,以及为什么那套机制让我们特别容易产生'大语言模型能理解'的错觉。在讲的过程中,我会尽量避免使用拟人化的说法。我不会说大语言模型在'努力'或者在'试图'做什么。我不会说它们在'提问'和'回答'。呃,或者说我会尽力不这么说。因为我觉得那会把水搅浑。呃,同样地,我提到'人工智能'或者'AI'这个词时,一定会加上引号。呃,因为它并不指向一组定义清晰的技术。呃,我觉得总体上更好的做法是谈'自动化',然后我们可以讨论我们到底在自动化什么、为什么要自动化它。好,进入第
便签引用
17:54
we're automating that. Okay. On to point one. What is understanding? In a paper that I wrote with Alexander Kohler that came out in 2020, sometimes known as the octopus paper, we gave a technical definition of understanding as mapping from the form of language, so sounds that you heard, signs that you observed, letters that you read to something outside of language. That's very narrow. And by that definition, if you tell your voice assistant to turn on the lights and it does so, it has understood that command. So it is possible to create automated systems that do this mapping.
一点。什么是理解?在我和亚历山大·科勒合写、2020 年发表的一篇论文里——有时被称为'章鱼论文'——我们给理解下了一个技术性定义:从语言的形式(你听到的声音、你看到的手势、你读到的字母)映射到语言之外的某种东西。这个定义非常狭窄。按照这个定义,如果你让语音助手开灯,它照做了,那它就理解了这个指令。所以,造出能完成这种映射的自动化系统是可能的。
便签引用
18:30
But as humans when we use language there's usually a whole lot more beyond that. Um so for example if I said look at that parrot over there right from the words and grammar alone you could tell that I'm trying to direct your attention somewhere and that there's likely a parrot over there. But from everything else that you put into understanding um you also know that um I wanted you to go look at the thing that I'm calling a parrot. um and that I either believe it to be something that is reasonably described as a parrot or I am behaving as if I were. So that's the very rich an example of a very rich kind of understanding that we do. If instead I had said if you don't speak Japanese, the words probably didn't do very much for you, right? But if it happened face to face and you could see my gesture um and my eye gaze, you might have the idea that I was trying to direct your attention to something. And if you looked over there and there was a parrot and it was the most salient thing, then you might have guessed that one of the
但作为人类,我们使用语言时,通常远不止于此。呃,比如说,如果我说'看那边那只鹦鹉',光凭词语和语法,你就能判断出我是想把你的注意力引向某个地方,而且那边很可能有一只鹦鹉。但从你在理解中投入的其他一切来看,你还知道,呃,我是想让你去看我称之为鹦鹉的那个东西;而且我要么真的相信它是一个可以合理称作鹦鹉的东西,要么我至少表现得好像我相信。这就是我们所做的那种非常丰富的理解的例子。如果我换成说日语,而你不会日语,那么这些词对你大概起不了什么作用,对吧?但如果这发生在面对面的场合,你能看到我的手势,呃,还有我的目光方向,你可能就会想到我是在把你的注意力引向某个东西。而如果你朝那边看,那里有一只鹦鹉,而且它是最显眼的东西,那你可能就会猜到,我刚才用的某个词其实就是
便签引用
19:31
words I just used actually was the Japanese word for parrot. It's in case you're wondering. Um and this could help you to start learning Japanese. But all of that inference, all of that meaning isn't coming from the words. The words weren't helpful to you unless you speak Japanese. I'm sure someone here does. Um, but in this kind of co-situated setting, because you are a person, because you bring all of that contextual understanding with you, you can use it to actually start learning Japanese drawing on all those cues. But the point is large language models aren't going to have that. And that's key. Okay, so point two, language models. How am I doing? Um, this is really old technology and it goes back to the work of Claude Shannon building on earlier work of uh, Marov and many people here are probably old enough to remember the T9 uh, interface for texting text on nine keys, right? Where you would type in the numbers and each number corresponds to three or four letters in the English T9 keyboard, but
日语里'鹦鹉'的意思。顺便说一句,那个词是 インコ。呃,这还能帮你开始学日语。但所有这些推断、所有这些意义,都不是来自那些词。除非你会日语,否则那些词对你没有帮助。我相信在座肯定有人会。呃,但在这种共处一地的情境里,因为你是一个人,因为你带着所有那些语境理解,你真的可以借助所有这些线索开始学日语。但关键在于,大语言模型不会有这些。这一点很要害。好,第二点,语言模型。我时间怎么样?呃,这其实是很老的技术,可以追溯到克劳德·香农的工作,而他又是在马尔可夫等人早先工作的基础上做的。在座可能有不少人年纪够大,还记得 T9 输入法吧?就是用九个键打字的那种界面。你输入数字,在英文 T9 键盘上每个数字对应三到四个字母,但是
便签引用
20:29
given a sequence of four or five, there's only a certain number of words in the dictionary that could have been what you intended, right? And the interface gave you those words ranked according to their frequency in some training data. And it was super annoying because home and good were the same sequence of keys. And it didn't matter if you said I'm coming, right? It was always good because that was the more frequent one in the language model which was just this count of how many there are. So those are the really basic unagram language models. You can make it more sophisticated by counting words in the context of one to n previous words. And this has been an important component of automatic transcription technologies of machine translation technologies where you basically have one component that says given that sound sequence or given this string of say French words what are the possible English words it might have been get a you know a range of possibilities and then you can consult the language model to reank them
给定四五个数字的序列,词典里只有有限的几个词可能是你想打的,对吧?然后界面会把这些词按它们在某个训练数据中的频率排序给你。这个东西烦人极了,因为 home 和 good 是同一串按键。而且不管你是不是想说 I'm coming home,它总是给出 good,因为在那个语言模型里 good 更常见——那个语言模型无非就是数一数各个词出现了多少次。所以那些是最基础的一元语言模型。你可以让它更精细一些,改成在前 1 到 n 个词的语境下统计词频。这一直是自动转写技术、机器翻译技术的重要组成部分:基本上你有一个模块,负责说,给定这段声音序列,或者给定这一串比方说法语的词,对应的英语词可能是哪些,得到一批候选,然后你可以查语言模型,按哪种结果最像训练数据来重新排序。这就是 1990 年代
便签引用
21:27
according to what looks the most like the training data. So that was language modeling in the 1990s and into the 2000s. Um, also, by the way, useful for spellch check, right? What are the possible words you can get to by one or two substitutions or deletions or insertions from what was typed? Rank those by just plain old frequency or rank them by likelihood in context. And then once you have a language model that is modeling more of the patterns, you can say, "Ah, that's the wrong there there and pick up things like that."
到 2000 年代的语言模型。呃,顺便说,它对拼写检查也有用,对吧?从你输入的内容出发,经过一到两次替换、删除或插入,能得到哪些可能的词?把它们按纯粹的词频排序,或者按在语境中的可能性排序。而一旦你有了一个能建模更多模式的语言模型,你就能说:'啊,这里的 there用错了',诸如此类。
便签引用
21:57
Five minutes. Five minutes. Okay. Um, in all of these cases, we're using language models to answer the question, which string of letters among the candidate strings of letters that come up in this application is the most probable given the training data. We are always and only dealing with the form of words and not what they can be used to refer to or what they're being used to refer to. In that context, there's two major differences with what we call LLMs today. The first is these there have been advances in machine learning architecture that allow us to take advantage of the enormous data sets that can now be constructed um questionably in terms of how it was collected not consentfully but can be constructed have been constructed. Um and so these machine learning architectures allow us to take advantage of these much larger training sets while modeling the relationships between words that go beyond just okay I've got four preceding words can I rank the vocabulary by likelihood? So much much
还有五分钟。五分钟。好。呃,在所有这些情形里,我们用语言模型回答的问题都是:在这个应用中出现的候选字母串里,给定训练数据,哪一串最有可能。我们始终只在处理词的形式,而不处理它们可以用来指称什么,或者在当下被用来指称什么。在这个背景下,今天我们所说的大语言模型有两点重大差别。第一点是,机器学习架构有了进展,让我们能够利用如今可以构建出来的巨大数据集——呃,这些数据的收集方式很可疑,并未取得同意,但确实可以构建,而且已经被构建出来了。呃,所以这些机器学习架构让我们既能利用大得多的训练集,同时又能建模词与词之间的关系,而不只是'好,我有前面四个词,能不能给词表按可能性排序'。所以是精细得多、
便签引用
22:52
more fine grained modeling of which words are likely to co-occur in which uh relationships to each other over longer sequences of text. Also important that training data is undisclosed with a few exceptions. The folks at AI2 for example I think are pretty open about their training data. The second big change with large language models is that the system has been turned inside out. So instead of using them to answer the question, which of these candidate sequences of word spellings is the most likely, we're instead using them to repeatedly answer the question, what is a likely word to come next over and over and over again.
细粒度得多地建模:在更长的文本序列上,哪些词更可能以哪种关系共同出现。同样重要的是,除了少数例外,训练数据是不公开的。比如 AI2 的那些人,我觉得他们对自己的训练数据相当开放。大语言模型的第二个大变化是整个系统被反过来用了。也就是说,我们不再用它来回答'这些候选的拼写序列里哪一个最有可能',而是用它一遍又一遍地回答'下一个词可能是什么'这个问题,反反复复。
便签引用
23:27
Right? So those are the two major differences. There's also two minor differences. Um, one is some additional training. For example, something called reinforcement learning from human feedback where you have many many data workers who are shown outputs from systems and either compare to this one's better, that one's better, or just good bad. And that allows uh the system designers to basically shape the probabilities towards the words that get the thumbs up from the data workers. And then there's some additional processing steps between input and output. So, for example, you might be scanning user input for arithmetic questions and in that case sending that off to a calculator to avoid the embarrassing LLM output that can't do math even though it's a computer.
对吧?这就是两个主要差别。还有两个次要差别。呃,一个是额外的训练。比如说,一种叫做'基于人类反馈的强化学习'的东西:你有非常多的数据工人,他们被展示系统的输出,然后做比较——这个更好,那个更好,或者只是打个好或坏。这让系统设计者基本上可以把概率塑造成偏向那些能得到数据工人点赞的词。然后,在输入和输出之间还有一些额外的处理步骤。比如说,可能会扫描用户输入里有没有算术题,如果有,就把它转给计算器,以免出现那种尴尬的情况——一个大语言模型明明跑在计算机上,输出却算不对数学。
便签引用
24:08
Um, also sometimes user input is turned into a web query. You retrieve those documents and then input those documents plus a system prompt that says produce a summary to do what's called retrieval augmented generation which is effectively paperier-mâché of the documents that came back while discouraging the user from going and experiencing that information in its own original context. And then there's system prompts in general. When you type something into the input box for chat GPT, that is actually not the only input that goes into the system. There's a whole prompt wrapped around that and also the conversation history being put in each time that you only then see a little bit of output afterwards. There's a bunch of stuff that's hidden in the interface. All right, two minutes. Okay.
呃,有时候用户输入还会被转成网络搜索的查询。系统检索出那些文档,然后把这些文档连同一段'请生成摘要'的系统提示词一起输进去,这就是所谓的检索增强生成——实质上是把返回的文档揉成一团纸浆糊出来的东西,同时还让用户不再去原始语境中亲自了解那些信息。另外还有系统提示词这回事。当你在 ChatGPT 的输入框里打字时,那其实并不是进入系统的全部输入。它外面还包着一整段提示词,而且每次还会把对话历史也塞进去,之后你只看到一小段输出。界面里藏着一大堆东西。好,还有两分钟。好。
便签引用
24:51
Um, so the result of all of this is coherent seeming text that's shaped to be pleasant according to the preferences of the data workers who did that reinforcement learning from human feedback. But it is still fundamentally the result of repeatingly answering the question, what's a likely next word? So why does it seem to understand? The answer is it only makes sense because we're making sense of it. You might think when you're understanding something you're listening to, something you're reading that the meaning is just there in the words and you are unpacking the words and getting the meaning. But research in pragmatics, in psycho linguistics shows us that's actually not how it happens. that we are continually just doing our best to figure out what that person must be trying to convey and using the words as a particularly rich clue to doing so.
呃,所有这一切的结果,是看上去连贯的文本,而且被塑造得让人愉悦——按照做那套人类反馈强化学习的数据工人的偏好来塑造。但它归根到底仍然只是反复回答'下一个词可能是什么'的结果。那它为什么看起来像是在理解?答案是:它之所以说得通,是因为我们在替它说通。你可能以为,当你理解你听到的、读到的东西时,意义就在词语里面,你只是把词语拆开、取出意义。但语用学、心理语言学的研究告诉我们,实际上并不是这样。我们其实是在不断地尽力揣摩对方一定是想传达什么,并把词语当作特别丰富的线索来完成这件事。
便签引用
25:36
And in order to do that, we have to imagine a mind behind the text. So when we encounter the synthetic text that comes out of a chatbot, we do the same thing. We do it reflexively and instinctively. And it is very hard to then remind ourselves actually that's all on our side. We made up that entire mind to make sense of that text. So to conclude, extraordinary claims such as large language models understanding require extraordinary evidence. Also, you have to define understand if you're going to make that claim and it has to be defined in a way that it's informed by what we know about how people understand language and it has to be falsifiable. Um, hiding the training data or even information about what's in the training data in general is anti-scientific because it doesn't allow people to reason about how the training data plus the input leads to the output.
而要做到这一点,我们就必须想象文本背后有一个心智。所以当我们遇到聊天机器人吐出来的合成文本时,我们做的是同一件事。我们条件反射般、本能地这么做。然后就很难提醒自己:其实那全都发生在我们这一边。是我们凭空造出了那整个心智,好让那段文本讲得通。所以结论是:非同寻常的主张,比如说大语言模型能理解,需要非同寻常的证据。另外,如果你要提出这个主张,你必须定义'理解',而且定义方式必须参考我们对人类如何理解语言的已有认识,并且必须是可证伪的。呃,隐瞒训练数据、甚至隐瞒关于训练数据里有什么的信息,总体上是反科学的,因为这让人无法推理训练数据加上输入是如何导致输出的。
便签引用
26:25
Uh, and also antisocial because it doesn't allow us to reason about when it would be a good idea to use these things. And then any texts that take the form of text in, text out might look to us like the system is answering questions and reasoning and being smart. Um, but in fact we should be careful that those are actually only testing how closely the system is modeling the distribution of the words in its training data and how closely that training data approximates the testing situation. So in short, do LLM understand? Emphatically not
呃,同时也是反社会的,因为这让我们无法判断什么时候使用这些东西才是好主意。还有,任何采取'输入文本、输出文本'这种形式的测试,在我们看来可能像是系统在回答问题、在推理、很聪明。呃,但事实上我们应该小心:这些测试实际上只是在检验系统对训练数据中词语分布的建模有多贴近,以及那份训练数据与测试情境有多接近。所以简而言之,大语言模型理解吗?断然不能。
便签引用
05Bubeck:基准饱和与进展速度
27:06
Sebastian. Yes. Hi everyone. Um so I will have a quite different take from what Emilia has been presenting and maybe I want to start by acknowledging that the question that we're debating tonight is sort of a weird question. Um and what I mean by that is is the following. So just the other day I was uh listening to to a talk like like like this one. uh it was in physics and the professor Harvard professor was uh talking about Einstein's uh theory of general relativity and in his presentation he said Einstein obviously did not understand anything about general relativity so this is something really really important to me that understanding it's kind of in the eye of the beholder like I have had this experience myself many many times in my research field where I feel like you know some other research researchers maybe they don't really understand what's going on but of course they understand come on I mean they are researcher in this field of course they understand but what I'm trying to say is
塞巴斯蒂安。好的。大家好。呃,我的看法会和艾米丽刚才讲的相当不同,也许我想先承认一点:我们今晚辩论的这个问题,某种意义上是个奇怪的问题。呃,我的意思是这样。前几天我去听一场演讲,就像这样的演讲。呃,是物理学的,那位教授,哈佛的教授,在讲爱因斯坦的广义相对论,在他的报告里他说,爱因斯坦显然根本不理解广义相对论。所以这一点对我来说非常非常重要:理解这件事有点像是取决于看的人。比如我自己在研究领域里就有过很多很多次这种体验——我会觉得,你知道,某些研究者也许并不真的明白是怎么回事。但他们当然是理解的,拜托,我是说他们是这个领域的研究者,他们当然理解。但我想说的是,存在不同的标准、不同的层次,我们据此判定
便签引用
28:09
that there are different bars different level at which we might decide that somebody or or something understands so that's point number one point number two is that as you know both Eliza and Emily said we we need to have some definition of what understanding means I totally agree with that the problem is humanity has debated the meaning of understanding for millennia and kind of spoiler alert we're not going to figure this out tonight. Okay, we none of us has a killer definition of what understanding means that subsumes this millennia of of work. So that's I think uh something really important. Now, if I'm not going to present to you a a definition of understanding, what I want to do is to go back to maybe things where we have numbers, solid numbers, we we understand things very precisely. And what I'm talking about are benchmarks. And benchmarks are a good way for us to assess the rate of progress that we have seen in the last couple of years. So I want to take a few minutes for all of us
某人或某物是否理解。这是第一点。第二点是,正如你们知道的,伊丽莎和艾米丽都说过,我们需要对理解有某种定义。我完全同意,问题在于人类已经争论'理解'的含义争论了几千年,而且剧透一下,我们今晚是搞不定这个的。好吧,我们当中没有谁能拿出一个杀手级的'理解'定义,把这几千年的工作都涵盖进去。所以我觉得这一点很重要。那么,既然我不打算给你们一个'理解'的定义,我想做的是回到那些我们有数字、有扎实数字、我们能非常精确地理解的东西上去。我说的就是基准测试。基准测试是个不错的途径,让我们评估过去几年里我们看到的进展速度。所以我想花几分钟,让在座各位一起盘点一下,过去这几年我们
便签引用
29:11
in the room to take stock of how far we have gone in the last few years and then I will move back to to the to the understanding question. So two and a half years ago just uh before JPT4 came out and just before chat GPT I I was at a a workshop on mathematical understanding of of neural networks and somebody from Google showed me the latest model that they had trained Minurva at the time a mass model and this model I was just absolutely shocked that it was able to solve one high school level question. It was just a very simple question about lines intersecting itself, you know, like really nothing difficult. But I was shocked by that two and a half years ago. Today we have systems that are in the running to get a gold medal at mathematics olympiad. You this is I think it's important to to recognize this and and this is irrespective of whatever definition we want to put on this. This is the rate of progress that we're seeing. So the mass benchmark, this was high school level competition.
走出了多远,然后我再回到理解的问题上。那么,两年半前,就在 GPT-4 出来之前、也在 ChatGPT 之前,我参加了一个关于神经网络数学理解的研讨会,谷歌的某个人给我看了他们当时训练出的最新模型 Minerva,一个数学模型。我当时完全被震住了,因为它能解出一道高中水平的题。那真的只是一道非常简单的题,关于直线相交,你知道,就是完全不难的东西。但两年半前,我被这个震住了。而今天,我们有的系统已经在角逐数学奥林匹克的金牌了。我觉得认识到这一点很重要,而且这跟我们想给它套上什么定义无关。这就是我们正在看到的进展速度。所以,MATH 这个基准测试,它是高中竞赛水平的。
便签引用
30:14
Two and a half years ago, chatbots would get like 20%. As of early last year, it was completely saturated. We could not use this benchmark anymore because the models were just becoming too good on it. Okay, so that one we we discarded. Now there was another benchmark which was very popular called MMLU. This is a benchmark that was supposed to test essentially high school level understanding of scientific question and that was a really important benchmark like all the CEOs of big tech were you know talking about it and it was kind of moving markets worldwide whether somebody would get plus 2% plus 5% Mark Zuckerberg would talk about it all the time you don't hear about it anymore at all why because it's completely saturated like our models are just too good on this so We keep coming up with newer and and and better benchmark. We are now at the stage where the latest and greatest benchmark that people are fighting to get higher on is called frontier mass. And frontier mass as the name suggests these are really very
两年半前,聊天机器人大概能拿 20%。到去年年初,它已经完全饱和了。我们没法再用这个基准了,因为模型在上面表现得太好了。好,所以那个我们就弃用了。然后还有另一个非常流行的基准叫 MMLU。这个基准本来是要测基本上是高中水平的科学问题理解能力,那是个非常重要的基准,大科技公司的 CEO 们都在谈论它,它甚至在全球范围内牵动市场——某家有没有多拿 2%、多拿 5%。扎克伯格以前天天都在谈它,现在你完全听不到了,为什么?因为它已经完全饱和了,我们的模型在这上面实在太强了。所以我们不断提出更新、更好的基准。我们现在到了这样一个阶段:最新最厉害、大家争着要刷高分的基准叫 FrontierMath。而 FrontierMath,顾名思义,这些确实是数学中非常高阶的研究性问题。不是未解难题,但仍然是你会拿去问,你
便签引用
31:18
advanced research question in mathematics. Not not open question but still questions that you would ask you know an early graduate student in mathematics. And when they came out with the benchmark, I think it was maybe four months ago, OpenAI's latest reasoning model 01 would get 4% of them correct. So again, everybody was very happy. Okay, here is an example of a benchmark that is really beyond the capabilities. One month later, we came out with O3, in fact 03 mini, which is even a smaller model. It ranks at 30%. And I don't even know if there is a single human being on the planet that can get 30% of those problem by himself or herself. So this is really you know the the kind of rate of progress that we're talking about.
知道的,数学专业刚入学的研究生的那种问题。这个基准刚推出的时候,我想大概是四个月前,OpenAI 最新的推理模型 o1 只能答对其中 4%。所以,大家又都很高兴了。好,这就是一个真正超出能力范围的基准的例子。一个月后,我们推出了 o3,实际上是 o3-mini,一个更小的模型。它的得分是 30%。我甚至不知道这个星球上有没有哪一个人能靠自己答对其中 30% 的题目。所以这真的就是,你知道的,我们所说的那种进展速度。
便签引用
32:04
Okay. But benchmark I actually I'm personally not a fan of benchmark. In the sparks of AGI paper that I wrote two years ago the whole point was to say we need to move away from benchmark because and this is where I think we will agree they do not show understanding. No matter how good you craft your benchmark, it's going to be very hard to extract understanding from it. To me, understanding you can judge it by interacting with the system and being very careful about not anthropomorphizing the output that you see, but really trying to probe and ask probing question to to further see how far the understanding go because again understanding it's not a binary notion.
好。但基准——其实我个人并不喜欢基准。在我两年前写的《AGI 的火花》那篇论文里,整个要点就是说我们需要摆脱基准,因为——这一点我想我们会有共识——它们并不能展示理解。不管你把基准设计得多好,都很难从中提取出"理解"。对我来说,理解是可以通过与系统互动来判断的,同时要非常小心,不要把你看到的输出拟人化,而是真正去探测、去问一些追问性的问题,进一步看这种理解能走多远,因为再说一次,理解不是一个二元的概念。
便签引用
32:45
It's not a zero or one. It's really a continuous metric. How far do you understand things? So I want to go back to this kind of more pragmatic view of understanding which is more what do those systems can do for you? Can they help you understand new things in some way and this is where you know I would challenge everybody to think about their own experience with those models because this is a great thing as opposed to two years ago is that probably everyone in the room has had their own experience with the chatbot. So you know for yourself whether you think that they understand or not whether they were helpful to you in one way or another and I will give you just my personal experience. So recently with the 03 mini so there is a a research question in mathematics that I I know there is no answer anywhere on the web. How do I know? Well, it's a problem that I personally thought about. In fact, I wrote a paper on it but I never had time to publish it. So it's just in my Dropbox, you know, there is like two
它不是 0 或 1。它其实是一个连续的度量。你对事物理解到什么程度?所以我想回到这种更务实的理解观,也就是:这些系统能为你做什么?它们能不能以某种方式帮你理解新的东西?在这一点上,你知道的,我想请在座每个人回想一下自己与这些模型打交道的经历,因为跟两年前相比,这是一件很棒的事——大概在座每个人都有过自己和聊天机器人互动的经验。所以你自己心里清楚,你觉得它们理解还是不理解,它们有没有在某方面帮到你。我就讲讲我个人的经历吧。最近用 o3-mini,有这么一个数学研究问题,我知道网上任何地方都没有答案。我怎么知道的?因为这是我自己思考过的问题。事实上我为它写了一篇论文,但一直没时间发表。所以它就躺在我的 Dropbox 里,你知道的,大概只有另外两个人看过。我想没人知道答案。于是我
便签引用
33:48
other human beings that have seen it. I think nobody knows the answer. And I asked O3 Mini this question. You know, if you if you are interested, the question is just some optimization question about the length of the of the trajectory. It doesn't matter the details. The point is of course 03 did not solve the question. That is way too much to ask. I I did not have any hope of this. But it gave me a suggestion of a connection with another field that took me personally 3 days to discover and this model by reasoning and thinking about the problem for 2 minutes was able to make that connection. Does it mean that it understand it's it's really hard to say was it extremely helpful and did it accelerate me and I could have you know spend those three days on something else. Yes, absolutely. So, so this this is this is where uh you know I I I want to to to take the discussion. Now um I want to give one more point in the favor of the you know kind of lack of understanding or what is exactly understanding and to do this I want to
就拿这个问题去问 o3-mini。你知道的,如果你感兴趣的话,那个问题就是关于轨迹长度的某个优化问题。细节不重要。重点是,o3 当然没有解出这个问题。那要求太高了,我本来也完全没抱这个希望。但它给了我一个提示,指出了与另一个领域的联系,而那个联系我自己花了三天才发现,这个模型只用两分钟推理和思考这个问题就能建立那种联系。这是否意味着它真的理解了?这真的很难说。它是不是极其有用?它有没有加速我的工作,让我可以把那三天花在别的事情上?是的,绝对有。所以,所以这就是这就是我,呃,我想把讨论引向的地方。现在,嗯,我想再补充一点,支持所谓缺乏理解,或者说到底什么才算理解。为了说明这一点,我想回到哲学家们
便签引用
34:57
hearken back to the philosophers's debate of what does understanding mean and I think there is a possibility that at the end of the day understanding is a human journey it's all about our own understanding and this is why I say that rather than asking whether Does the chatbot understand? You should ask yourself whether the chatbot helps you understand more things. And for example, in mathematics, it's entirely plausible that within the next few years, we would see breakthrough from those models that discover new things that prove new theorems that no other human beings has been able to prove. Yet the mathematical communics community might not accept it.
关于「理解意味着什么」的辩论。我认为有一种可能是,说到底,理解是一场人类的旅程,它关乎的是我们自己的理解。这也是为什么我说,与其问聊天机器人是否理解,你更应该问自己,聊天机器人有没有帮你理解更多的东西。比如说,在数学领域,完全有可能在未来几年里,我们会看到这些模型带来突破,发现新的东西,证明出此前没有任何人类能够证明的新定理。然而,数学界可能并不会接受它。
便签引用
06研究者调查与 AGI 无定义之争
35:42
It will accept that the theorem has been proven but they will say and in my opinion rightly so that we still do not understand we we will only accept that we understand it once a human is able to fully grasp the entirety of the proof. So this is this is where I want to end this introductory statement. All right. Thank you very much. [Applause] Okay. Let's start. I want to start by asking you both about a report that just came out from the Association for the Advancement of AI um earlier this month.
它会接受这个定理已经被证明了,但他们会说——在我看来这也是对的——我们仍然没有理解它。只有当一个人类能够完全把握整个证明的时候,我们才会承认我们理解了。所以,这就是我想为这段开场发言收尾的地方。好的。非常感谢。(掌声)好,我们开始吧。我想先问你们两位一个问题,关于刚刚发布的一份报告,来自人工智能促进协会本月早些时候
便签引用
36:15
Um included a survey of 475 AI researchers. About 70 76% of the respondents said it was unlikely or very unlikely that scaling up current approaches would succeed in achieving AGI and 80% of the survey respondents said current perceptions of AI capabilities don't match reality. It's a little bit adjacent to what we're talking about with understanding, but uh I was curious to see what you make of those results since they show um a fair amount of skepticism within the research field that uh what the the track that the technology is currently on is going to lead to lead to great things. Um Sebastian, I'll start with you. Yeah, this is actually a really interesting question because this is one thing that has shocked me in the last few years is that actually the academic community and researchers in general have been the ones who were almost the slowest to understand what's happening and and maybe maybe for for good reason. I mean, you know, it's the the problem with with LLMs is that they go certainly from the
做了一项针对 475 位 AI 研究者的调查。大约 70……76% 的受访者表示,靠扩大当前的方法来实现 AGI 是不太可能或非常不可能的;80% 的受访者说,当下人们对 AI 能力的认知与现实不符。这跟我们刚才讨论的「理解」问题稍微有点偏,但我很好奇你怎么看这些结果,因为它们显示出研究领域内部有相当程度的怀疑:这项技术目前所走的路线,究竟能不能带来了不起的成果。塞巴斯蒂安,我先从你开始。是的,这其实是个很有意思的问题,因为过去几年里让我震惊的一件事就是,实际上学术圈和研究人员整体反而是最慢理解正在发生什么的那一群人。这可能……可能也有正当的理由。我是说,你知道,LLM 的问题在于,至少从我们对这类对象的数学
便签引用
37:16
perspective of our mathematical understanding of those objects. It goes so much against everything that we thought that we understood about neural networks and artificial intelligence. We never thought as a research community that something like CH GPT would be possible at all. So I think there there is a lot of that skepticism that remains. So that's point number one. Point number two is also what do we mean exactly by scaling up? Because as you said in in your introductory segment um right now we have these new reasoning models which are scaling up along a different axis. To me this is all the same thing. I don't distinguish between all these different method. It's it's all the same from from a certain level of abstraction. It's all the same. But it is not the kind of pure pre-training paradigm that people have pushed for from GPT1, GPT2, GPT3 to GPT4. So it's a slightly different scaling paradigm. So maybe you know they are also replying in the sense that the classical pre-training will not get you to AGI.
理解的角度看,它彻底违背了我们自以为已经理解的关于神经网络和人工智能的一切。作为一个研究群体,我们从没想过像 ChatGPT 这样的东西会是可能的。所以我认为那种怀疑仍然大量存在。这是第一点。第二点是,我们说的「扩大规模」到底指什么?因为正如你在开场环节里说的,我们现在有了这些新的推理模型,它们是沿着另一个维度在扩大规模。在我看来,这一切都是一回事。我不区分这些不同的方法。从某个抽象层面看,它们都是一回事。从某个抽象层面看,它们都一样。但这不是人们从 GPT1、GPT2、GPT3 一路推到 GPT4 时所主张的那种纯粹的预训练范式。所以这是一种略有不同的扩展范式。所以也许你知道,他们的回答其实也是在说,经典的预训练不会带你走到 AGI。
便签引用
38:16
And this I completely agree with. I mean absolutely. But whether whether as uh humanity as a group we're on track towards AGI, it's plausible. Mhm. Emily, your thoughts? Uh, yeah. So, starting with the second one of those questions, I think that one's a well-rounded question. So, to say, does the does the hype match what functionality we actually currently have? I think that's a reasonable question to ask people and the people who build the systems might be in a good position to answer it. And so, so that one doesn't surprise me. It matches my view of reality. The first question I think is really broken as a survey question because AGI is not defined. Um, so unless that survey had a specific definition that people were answering with respect to, it's a it's a failed survey question. Um, and I also have to say that I really object to the trope that we often see um in discourse from researchers, but also um from journalists sometimes and and CEOs that there is a future point where AGI will
这一点我完全同意。我是说,绝对同意。但至于我们人类作为一个整体是不是正走在通往 AGI 的路上,这是有可能的。嗯。Emily,你怎么看?呃,好的。那我先从这两个问题里的第二个说起,我觉得那个问题问得挺周全的。也就是说,炒作的热度跟我们目前实际拥有的功能是否匹配?我觉得这是个可以合理地去问大家的问题,而那些造这些系统的人可能正好有资格回答它。所以,所以那个结果并不让我意外。它符合我对现实的判断。第一个问题我觉得作为一道调查题根本就是坏的,因为 AGI 没有定义。嗯,所以除非那份调查给出了一个明确的定义、让受访者依据它来回答,否则那就是一道失败的调查题。嗯,而且我还得说,我很反感我们在讨论中经常看到的那种套话,它出自研究者,有时也出自记者和 CEO——就是说未来存在某个时点,AGI 会出现,剩下的问题只是我们是不是走在正确的路上、以及我们沿着这条路走得有多快。我从几个方面反对这种说法。首先,
便签引用
39:14
exist and it's just a question of are we on the right path to it and how fast are we going down that path and I object to that in a couple of ways. First of all, there's no reason to believe that this thing that's not defined necessarily exists in the future. Um, and secondly, that's not how science works. It's not about picking a path and running down it. It's a community effort where people branch out and explore and tell each other what we found. And to your point earlier, Sebastian, about um, does it, you know, would a theorem proof that came out of an automatic system count as having proved the theorem if the mathematic community, mathematical community couldn't make sense of it? Um I think that that goes directly to the fact that that all scholarship is conversation and if you do something on your own to your own satisfaction that's not yet science or scholarship. I mean there is also this thing of doing something in the real world like you know if you if you have discovered I don't know a new formula in chemistry
没有任何理由相信这个连定义都没有的东西必然会在未来存在。嗯,其次,科学不是这么运作的。它不是挑一条路然后一路狂奔。它是一项共同体的事业,人们会分头去探索,再把各自的发现讲给彼此听。还有回到你刚才那点,Sebastian,关于呃,如果一个自动系统给出了一个定理证明,而数学界、数学共同体看不懂它,那还算不算证明了这个定理?嗯,我觉得这恰恰说明了一件事:所有的学术都是对话。如果你自己一个人做了点什么、自己觉得满意了,那还算不上科学或者学术。我是说,不过也存在这么一种情况,就是你在现实世界里做成了某件事,比如说,你要是发现了,我不知道,化学里的一个新配方,
便签引用
40:09
and you can create new compounds like this is something it's it's hard concrete evidence that you have done you don't need to talk about it you have it and you can build new things with it. So this aspect this kind of pragmatism of looking at what they can do for you. If if I if I want to like let's take a very extremely concrete example. Let's say I want to build an app and I ask one of those models to build the app for me and let's say right now you know I'm not saying it can do it but let's say in the future I just specify in natural language what I want. I come back the next day and boom the app is there.
而且你能照着造出新的化合物,这就是一种,这是很硬的、很具体的证据,证明你做成了,你不需要去谈论它,你就是有了它,而且你可以用它造出新东西。所以这一面,这种实用主义的态度——看看它们能为你做什么。如果,如果我想,我们举个特别具体的例子吧。比如说我想做一个 app,我让其中一个模型帮我做这个 app,比如说现在,你知道,我并不是说它能做到,但假设在未来,我只要用自然语言说清楚我想要什么。我第二天回来,砰,app 就在那儿了。
便签引用
07为什么要在乎「理解」
40:44
Maybe it's even already installed on my phone. Everything is there. Did the model understand on the way of creating this? In a way, I don't really care. It did something for me. So, do you care if it understands like what are the what are the problems? Um, so the reason I care about this question of whether or not they understand is that a lot of the discourse around this is going to be useful as a robo therapist, a robo lawyer, a robo tutor is premised on the idea or the vibes that it seems to be understanding and that is really problematic. Like I'm I really resist ever being called an AI researcher. I am not interested. This is I'm I'm a computational linguist. I was minding my own business doing grammar engineering. And then all this stuff happened and I realized that linguistics really has a lot to say. We are the ones who do have a definition of understanding even if other people have struggled with it for millennia, right?
可能它甚至已经装在我手机上了。什么都齐了。那模型在做出这个东西的过程中理解了吗?某种意义上,我并不真的在乎。它替我把事情办成了。那你在乎它是不是理解吗,比如说有什么,有什么问题?嗯,我之所以在意「它们到底理解不理解」这个问题,是因为围绕这件事的很多讨论都会把它当成机器人心理治疗师、机器人律师、机器人家教来用,而这套说法的前提,就是它看起来好像在理解,这真的很成问题。比如说,我特别抗拒别人叫我 AI 研究者。我没兴趣。我是——我是计算语言学家。我本来安安分分地在做语法工程。然后这一大堆事情发生了,我意识到语言学其实很有话可说。我们才是真正对「理解」有定义的那群人,哪怕别人为此纠结了几千年,对吧。
便签引用
41:37
And therefore I can see so another really important thing from linguistics is that the language part is not the same thing as the the world or the imagined world being talked about. And I like to use the metaphor of because I'm from Seattle. All right, a raindrop splattered window where you can focus on the raindrops and the pattern that they're making or you can change your focus and look through at the world outside. But when you're doing that, the raindrops are actually influencing how you can see what's outside. Information encoded in language is like that. And linguists focus on the structure of the language and how it allows us to see the information in it.
所以另一件从语言学来看非常重要的事情是:语言本身,跟被谈论的那个世界、或者那个想象出来的世界,不是同一回事。我喜欢打一个比方,因为我是西雅图人。好,一扇被雨点打花的窗户,你可以把焦点放在雨点上,放在它们形成的图案上;你也可以调整焦点,透过它去看外面的世界。但当你这么做的时候,那些雨点其实在影响你怎么看到外面的东西。编码在语言里的信息就是这样。语言学家关注的是语言的结构,以及这个结构如何让我们看见其中的信息。
便签引用
42:16
Where I notice a lot of people who are interested in doing language processing from a more computer science point of view just imagine that they're going straight for the information and don't see the raindrops on the window. And so I think the question of does something understand is really important because it allows us to stop and take stock of whether any given application is a good idea. Yeah. So I I want to respond to that. Um so I would strongly advise anybody in position of power to make decisions based on vibes. That's not a good idea. They shouldn't do that.
而我注意到,很多从计算机科学角度对语言处理感兴趣的人,会直接以为自己冲着信息去就行了,看不见窗上的雨点。所以我觉得「某个东西到底理不理解」这个问题真的很重要,因为它让我们能停下来,重新审视某个具体应用到底是不是个好主意。是的。所以我想回应一下这个问题。嗯,我会强烈建议任何身居决策位置的人,不要凭感觉做决定。那不是个好主意。他们不该那么做。
便签引用
42:48
Nobody should do that. I also would uh advise them against making decision based on whether they think this this object this system is understanding even if there is a perfect definition of understanding that's also not a good viable path. What I think we have to do in those situation if we're talking about a to use your word a robot doctor or robot you know lawyer we have to create benchmark we have to create test that just like we test you know we have test to test human beings whether they are capable of becoming a doctor of becoming a lawyer whether they are ready for it we should start to create those tests for the AI systems too and those tests will be very different in nature for the from the test that we're using from humans. I hate when people are running, you know, let's say the USMLE exam, for example, on chatbot. This is nonsense. This doesn't make any sense.
任何人都不该那么做。我同样建议他们,不要依据自己认为这个东西、这个系统是否在"理解"来做决定——哪怕真的存在一个关于"理解"的完美定义,那也不是一条可行的路。我认为在那种情况下我们必须做的是,如果我们说的是——用你的词——一个机器人医生或者机器人律师,我们必须建立基准,我们必须建立测试,就像我们测试——你知道,我们有测试去检验人类是否有能力成为医生、成为律师,是否已经准备好胜任,我们也应该开始为 AI 系统建立这样的测试,而那些测试在性质上会非常不同因为我们用的是给人类设计的测试。我特别反感有人拿着,比如说 USMLE这种考试,去让聊天机器人做。这完全是胡闹,根本说不通。
便签引用
08ARC 测试与基准的局限
43:42
You're not testing whether the system can become a doctor just because it can pass the USMLE. That that's not a reasonable thing to do because there are so many things about human beings that you assume they remember what they did yesterday, for example. Chatbot won't remember. So, you need to come up with new test. And this is where I think there's something a little in my own view dangerous about pushing people away from the field is that we need actually more people to start thinking about this question. We need to create those tests so that exactly we can start to have a solid route towards large scale deployment. H uh on the question of um new tests and um benchmarks too, I'm curious about your both of your thoughts on something called the ARC contest or ARC prize which stands for abstraction and reasoning corpus and which is meant to measure uh progress towards artificial general intelligence. But it's an interesting one because it has u tasks that aren't too hard for a human.
你并不能因为一个系统通过了 USMLE,就说你在测试它能不能当医生。这不是一个合理的做法,因为人身上有太多东西是你默认成立的,比如你默认他记得自己昨天做了什么。聊天机器人是不记得的。所以你得设计新的测试。而这也正是我觉得,在我个人看来,把人往这个领域外面推有点危险的地方——我们其实需要更多的人开始思考这个问题。我们需要造出那些测试,这样我们才能真正开始有一个向大规模部署迈进的可靠路径。呃,关于新的测试和基准这个问题,我很好奇你们两位怎么看一个叫 ARC 竞赛或者 ARC 奖的东西,ARC 是 abstraction and reasoning corpus(抽象与推理语料库)的缩写,它的目的是衡量我们在通用人工智能方面的进展。但它很有意思,因为它的任务对人类来说并不算太难。
便签引用
44:37
um like often there are patterns of colors and there's some underlying logic to how the the patterns are presented, how the colors are arranged and you have to figure out that pattern and then apply it to in another on another grid or something. Um so it's meant to thwart systems that memorize and really ask um test the ability to uh have abstract reasoning and generalization. Um and the first version of ARC, I think the OpenAI's 03 got an 87% and they kind of said, "Okay, well that one's done." But they just launched a new one today and I believe they said that 03 gets about 4% on it. Um so that seems like um at least it's a new frontier. Do you think that kind of test could evaluate understanding? I'll start with you.
呃,比如常常是一些颜色的图案,这些图案的呈现方式背后有某种内在逻辑,颜色是怎么排列的,你得找出那个规律,然后把它应用到另一个网格上之类的。呃,所以它的设计是要挫败那些靠死记硬背的系统,真正去考察抽象推理和泛化的能力。呃,第一版 ARC,我记得 OpenAI 的 o3 拿到了 87%,他们大概就说:“好,那这个算是被攻克了。”但他们今天刚推出了新的一版,我记得他们说 o3 在上面只有大约 4%。呃,所以这至少看起来是个新的前沿。你觉得这类测试能评估“理解”吗?我先从你开始。
便签引用
45:18
Yeah, I know it's really interesting. So maybe let me recount the story. Sure. Sure. U so so yes. So there was this uh benchmark called RKGI and the creator of the benchmark kind of claimed that you know once a model a system can solve this this is this is going to be a real you know milestone on our journey to AGI. Maybe it's not quite AGI but it's really going to be a major a major milestone. All previous models were stuck at 30% or something like that. I think human beings are around 80%. And then 03 came 03 end of last year which reaches 85% on it. So just crushes the benchmark and and you know without like without revealing anything about what's going on inside 03 it's really like true generalization to to this benchmark.
是啊,我觉得这真的很有意思。那也许让我把这个故事讲一遍。当然。当然。呃,所以,是的,有这么一个基准叫 ARC-AGI,这个基准的创造者当时大概是说,一旦某个模型、某个系统能解决它,这就会是我们通向 AGI 路上一个真正的里程碑。也许还算不上 AGI,但确实会是一个非常重大的里程碑。之前所有的模型都卡在 30% 左右。我记得人类大概是 80%。然后 o3 来了,去年年底的 o3,在上面达到了 85%。所以直接把这个基准碾压了。而且,你知道,在完全没有透露 o3 内部发生了什么的情况下,这真的就像是对这个基准的真正泛化。
便签引用
46:05
Now as soon as that happened this is where again the goalpost moving of okay but it's not really that's not really what we meant. is not really understanding and then you know now they come up with this new RKGI2 and indeed uh 03 gets a meager 4% on it. I have to say I looked at it today I wasn't able to solve many of those questions myself it's a little bit illdefined so maybe they will refine it and you know we will find that not everything is perfect and I think this goes into the bigger question which is goes back to my previous point namely that creating benchmark is a difficult task it's real hard work you you understand that as a society it's a full-time job to create benchmark for humans like to create those tests you know this is I mean we we have the educational system which is entirely built around that so we have to reinvent the entire system for AIs so I think it's it's it's hard to pin too much on that particular example both one way or another positive or negative Emily I'd love to get your thoughts yeah
那么,这件事一发生,又出现了那种挪动球门柱的情况:好吧,但这其实不是……这并不是我们真正想说的意思。这并不是真正的理解。然后,你知道,现在他们又搞出了这个新的 ARC-AGI-2,而 o3 在上面确实只拿到可怜的 4%。我得说,我今天看了一下,我自己也解不出里面很多题目,它有点定义不清,所以也许他们会再打磨一下,然后我们会发现并不是一切都完美。我觉得这就回到了那个更大的问题上,也就是回到我前面说的那一点,即创建基准是一件很难的事,是实打实的苦活。你也明白,作为一个社会,创建基准是一份全职工作,就像给人类出那些考试一样,你知道,这个……我是说,我们有整个教育系统就是围绕这件事建起来的。所以我们得为 AI 重新发明整套系统。所以我觉得,这个……在这个具体例子上下太重的结论是很难的,无论往哪个方向,正面还是负面。Emily,我很想听听你的想法。是的,关于基准我有很多话要说。呃,我在 2021 年和 Deb Raji 以及其他几位同事
便签引用
47:10
so I have a lot to say about benchmarks um I have a paper that came out in 2021 with Deb Raji and some other colleagues called AI and the everything in the whole wide world benchmark um which we uh we found a really nice parallel between the way that some of the uh natural language processing and image processing benchmarks were being used to this wonderful uh children's story called Grover and the Everything in the Whole Wide World Museum. And in that story, Grover goes into this museum and there's a room of things that are long and skinny and a room of things that are very heavy and a room of things that are ticklish. And he goes through all these rooms and then towards the end he says, "Hm, I've seen a lot of things in this museum, but I haven't seen everything in the whole wide world yet. And then there's this door with a sign on the top that says everything else. And Grover opens the door and of course he's in the outside world. And the point is that any given selection of things is only a
发过一篇论文,叫《AI 与全世界所有东西的基准》。呃,我们在其中发现了一个很妙的类比:某些自然语言处理和图像处理基准被使用的方式,跟一个很棒的儿童故事很像,那个故事叫《格罗弗与全世界所有东西博物馆》。在那个故事里,格罗弗走进这家博物馆,里面有一间屋子放又长又细的东西,有一间屋子放很重的东西,还有一间屋子放会让人发痒的东西。他走遍了所有这些房间,然后快到结尾时他说,“嗯,我在这个博物馆里看了很多东西,但我还没看遍这整个大千世界。’然后有一扇门,门上挂着牌子写着‘其余的一切’。格罗弗打开那扇门,当然,他就到了外面的世界。重点是,任何一组给定的东西都只是一个选择。它只是一个样本。而
便签引用
48:04
selection. It's only a sample. And one of the the real traps with natural language processing is we use language to talk about everything. And so when we design synthetic text extruding machines that can output plausible sounding text on any topic, it's really tempting to think, oh, this might actually be good at helping me solve math problems and this might be good at helping me diagnose based on my symptoms because it can send back text that looks like what you're looking for, but there is no way to actually test for full generality.
自然语言处理真正的陷阱之一,是我们用语言来谈论一切。所以当我们设计出能就任何话题输出听起来煞有介事的文本的合成文本挤出机时,就很容易想:噢,这东西说不定真能帮我解数学题,说不定还能根据我的症状帮我诊断,因为它能回给你一段看起来正是你想要的文本,但根本没有办法真正去检验所谓的完全通用性。
便签引用
48:33
And furthermore, I would say what's the point, right? When we're building technology, we want to test for functionality. We want to know that the system that we're building works well within certain tolerances. And here I'm thinking of Dr. Tam Gibbsu's work on data sheets for data sets where she took inspiration from like electronic components that say okay this thing is going to oscillate at 32,000 hertz give or take one right and that is the data sheet for this thing and we can pin it down and so we can know in which circumstances we're going to use it. The idea behind data sheets for data sets is that we should have similar ideas about what's in a data set so that a machine learning model trained on that data set we could know when to use it and as soon as we're talking about AGI something that is general purpose or whatever it becomes illdefined to test for that and I think illdefined as a goal right so my take on the ARC prize is um it's not a test of progress towards artificial general intelligence
而且我还想说,那意义何在呢?我们造技术的时候,是要检验功能性。我们想知道我们造出来的系统在一定容差范围内运作良好。这里我想到的是 Timnit Gebru 博士关于“数据集数据表”(datasheets for datasets)的工作,她的灵感来自电子元件——上面会写明,这个东西会以 32,000 赫兹振荡,正负 1 赫兹,对吧,这就是它的数据表,我们能把它的性能钉死,于是我们就知道在什么情况下该用它。数据集数据表背后的想法是,我们对一个数据集里有什么,也应该有类似清晰的说明,这样在这个数据集上训练出来的机器学习模型,我们就能知道什么时候该用它。而一旦我们开始谈 AGI,谈某种通用的东西之类,对它做测试就变得定义不清了,而且我认为把它当成目标本身就定义不清。所以我对 ARC 奖的看法是,它并不是衡量通向通用人工智能之进展的测试。
便签引用
49:28
and what I understand about the Um the test that was run with I think it was 03 was it wasn't just cold okay here do this thing it was here you know try this you know here's a few examples it was few shot rather than zero shot and you know I don't think there's any way that you can interestingly generalize from how well it does on that set of puzzles to really anything um and one last thing I want to say about benchmarks is we don't create benchmarks for people we create licensing exams we create you academic exams to see how well someone has learned the material in a class, how well they are prepared for medical school and things like that. But that's not benchmarking people. That's not saying how well has this person been developed towards some task. I I agree with everything you just said just for the record. Um absolutely. uh maybe one one slight addition that I would give to what you said is definitely I think arc AGI doesn't make any sense as a general purpose uh test and such a test does not
而且据我了解,那次测试——我记得用的是 o3——并不是干巴巴地说“好,来做这个”,而是“你看,试试这个,这儿有几个例子”,是少样本而不是零样本。而且你知道,我不认为有什么办法能从它在那套谜题上的表现,做出什么有意思的推广,推广到别的任何事情上。关于基准测试我最后还想说一点:我们不给人做基准测试,我们设执业资格考试,我们设学业考试,看一个人对课上的内容学得怎么样,看他们是否为医学院做好了准备之类的。但那不是给人做基准测试。那不是在说这个人在某项任务上被“开发”到了什么程度。我同意你刚才说的一切,这话说在前头。完全同意。呃,也许有一点小小的补充:我确实认为 ARC-AGI 作为一种通用测试是说不通的,而且这样的测试并不存在。但这恰恰是重点所在。我特别赞同你说的,我们
便签引用
50:32
exist but that's exactly the point I I'm a big fan when you say that you know we have to have a a functional form in mind for where we want to deploy we don't want AGI just to come and so try to solve everything we want to try to deploy AGI maybe to help with diagnostic in a very limited setting and then what I want to say is we want to develop a benchmark for that limited setting to see how well it performs you know can it help the the doctor things like that so why not have sure a benchmark for diagnostics based on symptoms and then purpose-built systems for doing that with disclosed training data I mean why not just because the other thing works better that's that's why not the other thing is environmentally ruinous it's built on stolen data Uh, and it's built on undisclosed data, so we can't decide when to use it. That's hard to respond to.
心里得有一个功能性的设想,明确我们要把它部署到哪里。我们不想让 AGI 就这么来了,然后去解决一切。我们想把 AGI 部署到,比如说,在一个非常受限的场景里帮助诊断,然后我想说的是,我们要为那个受限场景开发一个基准测试,看看它表现如何,看它能不能帮到医生之类的。那当然,为什么不做一个基于症状的诊断基准测试,然后配上专门为此打造、且训练数据公开的系统呢?我是说,为什么不呢?就因为另一个东西效果更好——这就是为什么不。另一个东西对环境是毁灭性的,它建立在偷来的数据上,而且建立在不公开的数据上,所以我们没法判断什么时候该用它。这话很难回应。
便签引用
09拟人化报道、开放性与认错条件
51:24
Let me take a slightly different obviously I disagree with everything just just for the record. Let me take a slightly different tact. Um, so these are the debates that computer scientists have and computational linguists have. Um, but think about this perspective. Um, think about it from the perspective of, you know, the average person who's just trying to keep up with what's going on in technology. They see headlines like um last week the New York Times had a headline digital therapists get stressed too and it was described the study the study that found that chat GBT showed signs of anxiety when it when users shared traumatic narratives with it and that its anxiety levels then dropped after it did a mindfulness exercise. So you can see how people would kind of get the impression that there's like a a real mind there something that's that's suffering in fact in that case very tragically. um you know what do you uh is there a danger in letting people um believe that there is a mind on the other side of
我换个稍微不同的角度——当然,我不同意刚才说的一切,这话说在前头。我换个稍微不同的切入点。这些是计算机科学家和计算语言学家之间的争论。但请想想这个视角。从一个普通人的视角来想想,就是那种只是想跟上技术进展的人。他们看到的标题是这样的——上周《纽约时报》有个标题叫“数字治疗师也会有压力”,讲的是一项研究,那项研究发现 ChatGPT 在用户向它讲述创伤性经历时表现出焦虑的迹象,而在做了正念练习之后,它的焦虑水平下降了。所以你能明白人们多少会产生一种印象:那里面像是有个真实的心智,某种在受苦的东西——在那个例子里甚至是很悲惨的。那么,让人们相信这项技术的另一端有一个心智,
便签引用
52:13
this technology? Um Emily, I'll start with you. Uh yes. Right. Because if you tell people, hey, this thing has the answers. Um or this can be your therapist, this can help you feel better or this can help you diagnose or this can help you deal with your legal troubles. Uh and by the way, it's super smart and it's been trained on everything we could scrape off the internet and it knows everything and it's objective. Then yeah, we are not setting people up to make good decisions. As for what to do about that journalism, my advice to the general public is if a journalist is anthropomorphizing things, especially that hard, skip. You don't need to read that.
这里面有危险吗?Emily,我先从你开始。有危险。因为如果你告诉人们,嘿,这东西有答案,或者这东西可以当你的治疗师,可以让你感觉好点,可以帮你诊断,可以帮你处理法律麻烦。而且顺便说一句,它超级聪明,它是用我们能从互联网上扒下来的一切训练出来的,它无所不知,而且它是客观的。那么是的,我们就没有为人们做出好决策创造条件。至于对那类新闻报道该怎么办,我给公众的建议是:如果一个记者在把东西拟人化,尤其是拟人化得那么厉害,那就跳过。你不需要读那个。
便签引用
52:47
Yeah. No, this is not a great paper. Uh but but and and it's these butts are important, you know, because this is a very delicate and and and and subtle topic like any take which is kind of categorical. Yeah, we should completely ignore that. It should absolutely not be used for XYZ. I would be wary of such claims. So in part in in in the example of of this paper, yes, not great, but the there is something may be useful for the AI researchers, not for the general public, but for the AI researchers, there is something about the prompting of the model and understanding exactly whether you can surface some good behavior by putting it quote unquote in the right mindset. Even though it's not there's no mind there to be clear, but it's just I don't know what other language to use. This is also the thing is we need more work in this direction to develop the right vocabulary to talk about these things. I also don't like the term artificial intelligence. It's some it's something else. It is intelligence. It is understanding but in
是的。这确实不是一篇很好的论文。但是——而这些“但是”很重要,因为这是一个非常微妙、非常细致的话题——任何比较绝对的说法,比如“对,我们应该彻底无视它”“这绝对不该用于某某某”,我对这类断言都会很谨慎。所以就这篇论文而言,是的,它不算好,但其中或许有些东西是有用的——不是对公众,而是对 AI 研究者。其中有些关于如何给模型提示的东西,关于弄清楚你能不能通过把它——打个引号——放进“正确的心态”里,来引出某些好的行为。虽然那里并没有心智,这一点得说清楚,但我只是不知道还能用什么别的说法。这也正是问题所在:我们需要在这个方向上做更多工作,以发展出谈论这些东西的正确词汇。我也不喜欢“人工智能”这个词。它是别的什么东西。它是智能,它是理解,但方式和人类不同,它
便签引用
53:55
a different way from human beings and it needs to have its own vocabulary around it. And maybe it's going to take decades to develop you know this new field of study. And the whole difficulty is that we don't have decades. You know a century ago when we developed quantum mechanics physicists were making all of these wonderful discoveries and they could take 10 20 years. In fact it took you know more than 50 years to have the quantum field theory. But the point is they had time to do that. Today we don't really have the time to do that because the progress is so fast. And so this entire conversation also has to be rooted in the reality that I talked about before that two years ago I was amazed that the problem that the machine was solving a single high school problem whereas today they are in the running for a gold medal in mathematical olympiad. So this has to be in the background of everything we say.
需要有围绕它自身的一套词汇。而要发展出这个新的研究领域,也许要花上几十年。而全部的困难在于,我们没有几十年。你知道,一个世纪前我们发展量子力学时,物理学家做出了那么多美妙的发现,他们可以花上十年二十年。事实上,量子场论花了五十多年才成型。但重点是,他们有时间这么做。今天我们真的没有那个时间,因为进展太快了。所以这整场对话也必须扎根于我之前谈到的那个现实:两年前,我还惊叹于机器能解出一道高中题,而今天它们已经在角逐国际数学奥林匹克的金牌了。所以我们说的每句话,背景里都得有这一点。
便签引用
54:48
Wasn't that one one where it got to submit like 10,000 answers and then they this is Google are not doing it right. But uh but uh but open AI is not doing it like that. Uh all right. But open AI is not at all open about what you're doing. We don't we don't know if those problems are in the training data. Very similar ones are 400 million people using it weekly, you know. So it's pretty open actually. No. Open means we know what the training data is. Open means we know what the steps are. Open means we know what the system prompt is. Open means we know what your carbon footprint is. All of that would be open. Yeah. I think Sam alluded to the fact that uh maybe we're not on the right side of history with respect to uh open source and open access. So maybe there will be more news in that direction soon.
那次不是可以提交上万个答案的那种吗?然后他们——这是谷歌——那样做是不对的。但是 OpenAI 不是那样做的。好吧。但 OpenAI 对自己在做什么完全谈不上开放。我们不知道那些题目是不是在训练数据里。非常相似的题目——有四亿人每周都在用它,所以其实挺“开放”的。不。开放意味着我们知道训练数据是什么。开放意味着我们知道步骤是什么。开放意味着我们知道系统提示词是什么。开放意味着我们知道你的碳足迹是多少。这些才叫开放。是的。我想 Sam 也暗示过,在开源和开放获取这件事上,我们或许站在了历史错误的一边。所以也许很快会有这方面的新消息。
便签引用
55:35
Uh we're going to turn to audience questions in just a minute. But um I wanted to ask one question of you both. Um before we do that, uh what would you have to see to convince you that you're on the wrong side of this debate? what kind of behavior or evidence? Uh so I would need a very clear definition of understanding. I would need uh something experimental where all of the experimental parameters are known and open every last thing about the training data so that we can then reason about okay getting to this output given that training data given the input given the training architecture given the system prompt there is no way to do that without XYZ thing that connects the definition of understanding yeah um I mean for me it's very very concrete actually um so you know I spent I don't know the last 15 to 20 years doing research in in mathematics and and I had many extremely joyful moment reading papers and getting a new deep understanding by from one of those papers and I would feel that I'm on the
我们再过一会儿就转到观众提问。但我想先问你们两位一个问题。在那之前——要让你相信自己站在这场辩论的错误一边,你得看到什么?什么样的行为或证据?我需要一个非常清晰的“理解”的定义。我需要一个实验性的东西,其中所有实验参数都是已知且公开的,关于训练数据的每一个细节都公开,这样我们才能推理:好,在给定训练数据、给定输入、给定训练架构、给定系统提示词的情况下得到这个输出,如果没有某某某要素就绝无可能——而那个要素要和“理解”的定义挂上钩。对我来说其实非常非常具体。你知道,我大概过去十五到二十年都在做数学研究,我有过很多极其愉悦的时刻:读一篇论文,从中获得一种全新的深刻理解。我会觉得自己站在了
便签引用
10观众问答:具身、文本与权力
56:40
wrong side of the debate is if let's say 10 years from now or for that matter at the end of my life I I don't care about 10 years but if in my lifetime I never have such a moment where I realize something as deep as I realize when I read a paper written by another human being who spent, you know, maybe years working on it from an AI system. If I don't get that, then yeah, I was wrong. All right, I'll do a few of these audience questions. There's a lot of really interesting, provocative ones. Um, a lot of people are kind of curious about um how we draw the uh distinction between human understanding. Uh well, so for example, one person says, is there is a statistical representation of the world different from how humans understand the world? Other people say like are we just doing the same thing?
辩论的错误一边——如果说,十年之后,或者就说到我这辈子结束,我不在乎十年不十年——但如果在我有生之年,我从来没有从一个 AI 系统那里获得过那样的时刻,那种像我读一篇由另一个人类写的、他也许花了好几年心血的论文时所领悟到的深度。如果我得不到那个,那好,我就错了。好,我来挑几个观众问题。有很多非常有意思、很有挑衅性的问题。很多人挺好奇我们该如何划分与人类理解之间的界线。比如说,有人问:对世界的统计性表征,与人类理解世界的方式不同吗?也有人说,我们是不是其实在做同一件事?
便签引用
57:27
Are we just predict predicting the next token when we're when we think so I want to take on that are we just predicting the next token thing that is an inherently dehumanizing argument that says you don't think that I actually have an internal life and I don't want to have that conversation if you don't think I have an internal life. Yeah. I guess my perspective is kind of uh similar but but I turn it the other way around which is I don't understand how you work and I don't want to make any assumption about your inner working when you're speaking and when you're you know giving me understanding and meaning and the same way for the AI system I don't want I don't care about how it works inside this is not my business it's I mean it is but okay not not for the perspective of you know whether it understand or not it's doing something and it's it's giving me content and then I judge for myself whether you know it's useful or not.
我们在思考的时候,是不是也只是在预测下一个词元?我想回应一下这个“我们是不是也只是在预测下一个词元”的说法:这本质上是一个非人化的论证,它等于说你不认为我真的拥有内在生活。而如果你不认为我有内在生活,我不想进行那样的对话。是的。我的视角有点类似,但我是反过来看的:我并不理解你是怎么运作的,而且当你说话、当你向我传达理解和意义时,我也不想对你的内部运作做任何假设。对 AI 系统也是一样,我不想——我不在乎它内部是怎么工作的,那不关我的事。我是说,其实也关,但好吧,不是从“它到底理不理解”这个角度。它在做某件事,它在给我内容,然后我自己来判断这东西有没有用。
便签引用
58:22
Another question is um to what degree does uh understanding and knowledge need to have a degree of physical physicality like um uh like for example I guess if um like if my friend tells me how to play pickle pickle ball gives me all the rules and the instructions tells me how it feels to play pickle ball. watch videos of playing pickle ball. Do I really know how to play pickle ball? Do I understand it or do I have to go out and do it myself to get the hang of it? Either you um so there's sometimes in in linguistics and philosophy you hear this question as does understanding have to be embodied and and in order to actually really learn about the world and really learn a language, you have to have embodied experience. Um and I think that embodiment is important, but I think even more important is socially situated relational experience. that that meaning really is in the relationships that we have with other people. And if you look into it, if you get into the details of social linguistics and language change,
另一个问题是:理解和知识在多大程度上需要某种物理性、具身性?比如说,如果我朋友告诉我怎么打匹克球,把所有规则和要领都讲给我,告诉我打匹克球是什么感觉,我也看了打匹克球的视频。那我真的会打匹克球了吗?我理解它了吗?还是说我必须自己下场去打,才能掌握它?你们谁都行。在语言学和哲学里,你有时会听到这个问题的另一种问法:理解是否必须是具身的?以及,要真正认识世界、真正习得一门语言,你是否必须有具身经验。我认为具身性很重要,但我认为更重要的是社会情境中的关系性经验。意义真正存在于我们与他人的关系之中。如果你去深究,如果你进入社会语言学和语言变迁的细节,
便签引用
59:15
we find that words are constantly changing in what they mean based on how we use them. And that happens because we are successfully using them as these little tokens as to like clues to what we're trying to convey. And then that colors the word for the next time we're going to use the clue. And so it's really all fundamentally in the relationships with other people. And a large language model fundamentally can't do that. we might in talking to it do our half of it and people certainly form attachments but again that's all on the person doing it and how lonely is that to form a strong strong attachment to something that's not there.
你会发现词的意思在不断变化,取决于我们怎么使用它们。而这之所以发生,是因为我们成功地把它们当作一个个小小的标记来用,当作我们想传达之物的线索。然后这一次的用法又会给这个词染上颜色,影响我们下一次再用这个线索时的样子。所以归根到底,这一切都根植于人与人的关系之中。而大语言模型从根本上做不到这一点。我们在跟它说话时,也许会完成属于我们那一半,人们当然也会产生依恋,但那全都发生在这个人自己身上。而对一个并不在场的东西产生强烈的依恋,那是多么孤独的一件事。
便签引用
59:50
Yeah. So I mean the way I interpret this question is is about whether multimodality is necessary for for general intelligence and I think this is a frontier research question. Nobody really knows the answer to it. Again, none of the terms are really well defined, granted, but even with this illdefiness, we don't really know. I have a belief personally that it's not needed at all. That you can get to a general intelligence in the sense of everything that I have discussed so far, purely through text, purely through digital access to the internet. You don't need any physicality. You don't need to see anything. You don't need to hear anything. You can get intelligence purely from purely from thoughts and reasoning. That's my belief but it has to be it's not it's not confirmed yet.
是的。我对这个问题的理解是,它问的是多模态对通用智能是不是必需的。我认为这是一个前沿研究问题,没有人真正知道答案。而且同样地,这些词都没有被很好地定义,这我承认,但即便有这种定义不清,我们也确实不知道。我个人相信它根本不是必需的。我相信你可以达到通用智能——在我到目前为止讨论的那个意义上——纯粹通过文本,纯粹通过对互联网的数字访问。你不需要任何物理性。你不需要看到任何东西。你不需要听到任何东西。你可以纯粹从思维和推理中获得智能。这是我的信念,但它还没有被证实。
便签引用
1:00:37
Now maybe one more point is uh LLM's you know there's nothing special about text. I mean text is just useful because there is a ton of it and it's there's a ton of it and it's easy to process. There's a ton of videos, there's a ton of images but it's a little bit more cumbersome to process. text there's a lot of it and it's easy to process but the technology itself is completely agnostic to the modality it can be applied to vision it can be applied to audio and indeed ch now can see can hear you can interact with it you can turn on your camera and you know you will have fun interaction I mean my kids have fun interaction with it that way yeah that's a related question someone says do you think AGI I'm not sure if they me or AI do you think AGI will understand the world better without human language different basis understanding I guess without human language. Yeah, without human language.
还有一点是,关于大语言模型,文本其实并没有什么特别之处。我是说,文本之所以有用,只是因为它量特别大,而且它量特别大又容易处理。还有视频有一大堆,图像也有一大堆,但处理起来要麻烦一些。文本呢,数量也非常多,而且处理起来很容易,但技术本身对模态是完全无所谓的,它可以用在视觉上,也可以用在音频上。事实上,现在的模型能看、能听,你可以和它互动,可以打开摄像头,然后你会有很有意思的互动。我是说,我的孩子们就是这样跟它玩得很开心的。对,这里有个相关的问题,有人问:你觉得 AGI——我不确定他们说的是我还是 AI——你觉得 AGI 在没有人类语言的情况下,会更好地理解这个世界吗?也就是建立在不同基础上的理解,我猜是指不用人类语言。对,不用人类语言。
便签引用
1:01:27
No. So I think human language is wonderful. Uh it is it is very information dense like in in an image in in in the access to the to the other modalities there are so many repetition like even audio is extremely repetitive. Maybe what I say is very repetitive but the these repetition you have really way less of them in text. So I I think actually text is kind of necessary. Like maybe to put it differently, if you wanted to get to a general intelligence only through interacting in the world, then I don't really see another way to do it than evolution for billions of years on planet Earth and then we're here.
不会。我认为人类语言非常了不起。它的信息密度非常高,而在图像里、在其他模态里,有太多重复的东西了,连音频都极其重复。也许我说的话就很重复,但这类重复在文本里真的要少得多。所以我认为文本其实是必需的。或者换个说法,如果你想只通过与世界互动来抵达通用智能,那我实在想不出别的办法,只能像地球上几十亿年的演化那样,最后我们才出现在这里。
便签引用
1:02:08
So that question suffers again from the problem of AGI being undefined. Um I want to take issue with the idea that text isn't special. Text is really special. There's a lot going on. Yeah, I love I love text. I mean, it's not special from the perspective of the technology. But so it's not special if your perspective is it's just data and I'm going to crunch the data with my neural net one way or another. Yeah, I love I think it it is special in that it text is a reflection of language and language is how we build our social world and how we interact with each other and how we do what's effectively telepathy, right? how we get inside somebody else's mind. And if we have systems now putting synthetic text in front of us and sort of spilling it into our information ecosystem, that's a huge problem because text is special. Now, this is the top question been upvoted the most. So, I have to ask it. What role do wealth and power play in who decides whether LLMs understand?
所以这个问题又一次卡在 AGI 没有定义这个问题上。另外我想反驳一点,就是说文本不特殊。文本真的很特殊,里面有很多东西。对,我很喜欢文本。我的意思是,从技术的角度看它并不特殊。但这只是说,如果你的视角是「它不过是数据,我要用神经网络以某种方式去处理这些数据」,那它就不特殊。对,我认为它特殊,是因为文本是语言的映射,而语言是我们构建社会世界的方式,是我们彼此互动的方式,是我们实现某种意义上的心灵感应的方式,对吧?是我们进入别人心里的方式。而如果现在有系统把合成文本摆在我们面前,把它倾泻进我们的信息生态,那就是个巨大的问题,正因为文本是特殊的。那么,这是被顶得最高的问题,所以我必须问一下:财富和权力在「谁来决定大语言模型是否理解」这件事上,扮演什么角色?
便签引用
1:03:04
Tough one. Uh, so who decides? You all get to make your own decision. That's the point of events like this. Um, but who decides in terms of how the political discourse goes? Are we making policy decisions based on the idea that LLMs understand? Then their wealth and power very much come into it. And furthermore, the uh political economy around developing these things is further centralizing wealth and power. Data became power. I want to point out that before you said that that uh question you're asking was just in your own Dropbox. Is Dropbox one of the companies that made a deal with OpenAI?
这个问题不好答。谁来决定?你们每个人都可以做出自己的判断,这正是这类活动的意义。但是,如果说谁来决定政治讨论的走向呢?我们是不是在以「大语言模型能理解」为前提做政策决定?那财富和权力就非常有关系了。而且,围绕开发这些东西的政治经济结构,正在进一步把财富和权力集中起来。数据成了权力。我想指出,你刚才说那个问题的时候提到,你要问的那个问题只存在于你自己的 Dropbox 里。Dropbox 是不是那些和 OpenAI 达成合作的公司之一?
便签引用
11涌现是真实还是不科学
1:03:39
I I don't think my Dropbox is being sent to OpenAI. I don't know because you don't tell us what's in your data. Yeah. Yeah. Sure. Um Yeah. So, who decides? Um yeah, you all decide. That's definitely the right answer. I think also we have to you know acknowledge that we live in a in a world which is run by you know capitalist societies and that are trying to get ahead build things faster better etc. If those tools because at the end of the day they are tools like this is what they are. If they make you more productive, if they make you that you can realize your ideas quicker, if you can do your work in a shorter amount of time, then they will just, you know, diffuse that way. So it's the diffus the decision is distributed. There is no single entity that's going to decide. Um we have a question maybe I'll make this the last question then we can do our closing statements. Um a question about emergent behavior which is uh when something complex um it says emergence occurs when a complex entity has properties or
我不认为我的 Dropbox 内容被发给了 OpenAI。我不知道,因为你们从不告诉我们数据里到底有什么。对,对,好吧。嗯,对。谁来决定?确实是你们所有人来决定,这绝对是正确答案。我也认为我们得承认,我们生活在一个由资本主义社会运转的世界里,大家都在争先恐后,想跑在前面,想做得更快更好,等等。如果这些工具——因为归根结底它们就是工具,这就是它们的本质——如果它们让你更高效,让你更快地实现自己的想法,让你能用更短的时间完成工作,那它们自然就会那样扩散开去。所以决策是分散的,不存在某个单一的实体来做决定。我们还有一个问题,也许我把它作为最后一个问题,然后我们进入结语。这是一个关于涌现行为的问题,就是当某个复杂的东西——问题里说,涌现发生在一个复杂实体具备了它的各个部分单独并不具备的属性或行为的时候。那么,你们
便签引用
1:04:46
behaviors that is that its parts do not possess individually. Um so sort of what are your thoughts on this concept and I suppose um is understanding like in the human sense something that could emerge start. Yeah. So uh you know did understanding and when I talk about understanding I'm talking about understanding of language emerge at some point in the evolution of humans. Well clearly right because at some point we went from entities that didn't have language to entities that do and the ability to understand language emerged.
对这个概念怎么看?另外我想问,人类意义上的「理解」是不是一种可能涌现出来的东西。好,那么,「理解」——我说「理解」的时候指的是对语言的理解——有没有在人类演化的某个阶段涌现出来?显然是有的,对吧,因为在某个时刻我们从不具备语言的存在,变成了具备语言的存在,理解语言的能力就是这样涌现的。
便签引用
1:05:14
Um when we're talking about emergent behavior with machine learning systems again these claims are unscientific unless you have full open access to the training data. And furthermore, we have to be careful about looking at how we're interpreting input output pairs as question output and question answer pairs and compare that to what other kinds of computation might have gone from that input to that output before we actually make the extraordinary claim of emergent behavior. No, I think emergence is is definitely real. I I have a hard time to be honest understanding how this is uh debated frankly. uh and and I would push back on the fact that you need to know every detail of those system to to tell whether you can uh understand whether there is emergence or not and the reason is you can run your own experiments. I mean before you know embarking on this journey of building LLM myself, I was running experiments on neural networks building toy settings by myself just alone with a very simple access to just
但当我们谈机器学习系统的涌现行为时,除非你能完全公开地访问训练数据,否则这些说法就是不科学的。而且,我们必须小心,要看清我们是怎么把输入输出对解读成问答对的,并且要拿它和其他各种可能从那个输入通向那个输出的计算方式做比较,然后才谈得上提出「涌现行为」这样一个非同寻常的主张。不,我认为涌现绝对是真实存在的。老实说,我很难理解这一点怎么会有争议。坦白讲,我也不同意「你必须知道这些系统的每一个细节,才能判断是否存在涌现」这个说法,原因是你可以自己做实验。我是说,在我自己踏上构建大模型这条路之前,我就在做神经网络的实验,自己一个人搭一些玩具设定,只用我自己的电脑这么简单的条件,我就能造出
便签引用
1:06:15
with my own computer and I could create settings where I would see this emergence phenomenon. just to to ground it to make sure we're all on the same page. One example that I looked at was just to train a model to solve systems of linear equation. You remember like in in middle school or wherever where you do x + y = to x - y = z where what are x and y right? This this type of question and you can see a setting where suddenly there is emergence where the system is able from learning only with systems of linear equation with a few equations to solve them with many equations. This is just an emergence that you can see literally in the lab by yourself. You don't need to have access to, you know, like the best possible models.
一些能看到涌现现象的设定。为了说得具体些,确保大家理解一致,我看过的一个例子是训练一个模型去解线性方程组。你们还记得吧,中学或者别的什么时候,你会做x + y 等于多少、x - y = z,然后问 x 和 y 是多少,对吧?就是这类题目。你会看到某种设定下突然出现涌现:这个系统只学过方程数量很少的线性方程组,却能够解出方程数量很多的方程组。这就是一种你完全可以自己在实验里看到的涌现,你并不需要拿到那些最顶尖的模型。
便签引用
12结辩与现场投票
1:06:58
All right. With that, I will ask you both for closing statements about three minutes. Um, Emily, would you like to begin since you began before? Sure. Um, what would I like to leave you with tonight? Um, I would like to leave you with the idea that nothing is inevitable. If people say this is here to stay, we have to learn to live with it. You can say no. Refusal is really important, right? And that is especially important in our systems that are um already creaking and about to get much much creakier. I'm talking about education. I'm talking about healthcare.
好的,接下来我请两位各做大约三分钟的结语。Emily,你愿意先开始吗?因为一开始也是你先讲的。好的。今晚我想留给大家什么呢?我想留给大家这样一个想法:没有什么是不可避免的。如果有人说「这东西已经来了,我们只能学着与它共存」,你可以说不。拒绝是非常重要的,对吧?而在那些本来就已经吱吱作响、即将变得更加摇摇欲坠的系统里,这一点尤其重要。我说的是教育,是医疗,
便签引用
1:07:29
I'm talking about our legal system. I'm talking about uh immigration proceedings. All of these places where synthetic text looks like a nice handy band-aid, you know, quick solution because there's not enough teachers, there's not enough therapists or whatever. Um, we need to say no to that because it's actually worse than nothing. And saying no, I think starts with refusing in the it's not just a fun toy to play with. All right, it's not a good tool for search. Anytime somebody is saying, "Oh, here, use this. It's artificial intelligence." There's my scare quotes, right? Remember, in your back pocket, you've got no. And I'll stop there.
是我们的司法系统,是移民审理程序。在所有这些地方,合成文本看起来像是一块方便好用的创可贴,一个快速解决方案,因为老师不够用,心理治疗师不够用,诸如此类。我们必须对此说不,因为它其实比什么都不做还糟。而说不,我认为要从拒绝开始——它不只是一个好玩的玩具,它也不是一个好的搜索工具。任何时候,只要有人跟你说「喏,用这个吧,这是人工智能」——注意我这里打了引号——请记住,你的口袋里始终揣着一个「不」。我就讲到这里。
便签引用
1:08:10
[Music] Yeah, my take is quite different. Um I and I would say in general like I you know I think you all decide and and I'm I'm kind of against you know the the scare tactics and and trying to say that you know this is scary and this is just going to be bad. like we can all we're smart enough to judge by ourself when we interact with these tools when we use them whether they bring uh value to us. So that's that's point number one. Point number two I want to come back actually to something that I thought would come back more in this discussion that you brought up Eliza at the beginning which is these riddles where you know you change just a little bit and then it's it's completely wrong. M this goes back to the point that I was trying to make when I said that these topics are complicated and actually quite subtle and I think there is truth both in the parrot and in the sparks it's a it's it's a mix of both on certain topics they act more as parrot on certain topics there is a little bit more of the
【音乐】是的,我的看法很不一样。总体上我会说,我认为该由你们自己来判断,而且我有点反对那种制造恐慌的做法,反对那种「这很可怕,这只会变糟」的说法。我们都足够聪明,能在自己使用这些工具、与它们互动的时候,判断它们是否给我们带来了价值。这是第一点。第二点,我其实想回到一个我本以为会在这场讨论里被更多提及的东西,就是你一开始提到 Eliza 时说的那些谜题:你只稍微改动一点,它的回答就完全错了。这回到了我先前想表达的观点:这些议题很复杂,而且相当微妙。我认为「鹦鹉」和「火花」这两种说法里都有真实的成分,它是两者的混合。在某些话题上它们更像鹦鹉,在某些话题上则多了一点
便签引用
1:09:15
sparks and what I see is that you know maybe at the time of GPT3 it was more like this GPT4 it was maybe more like this and you know we're starting like I see the balance shifting But definitely both modes are still in there. So it's a complicated really really complex topic. With all of that being said, there is definitely a wave right now of more and more people interacting with those models and realizing how powerful they are. One thing that was really I don't know enlightening to see was when Sam tweeted one of the stories, fiction stories that our newest model was writing. He just tweeted that and then it was noted by by noted authors like creative writers that they were moved by that story and of course it's very I know what you're going to say that you know you can only be moved by something that another human being wrote etc. Sure. But this person said, you know, who is noted author that they were moved by it. So I think this is just, you know, a wave that I it keeps increasing. How far it's going to
火花的意味。而我看到的是,在 GPT-3 的时候可能更偏这边,GPT-4 的时候可能更偏那边,现在我看到这个平衡正在移动。但这两种模式肯定都还在里面。所以这是一个非常非常复杂的话题。话虽如此,眼下确实有一股浪潮,越来越多的人在使用这些模型,并意识到它们有多强大。有一件事我不知道该说是不是很有启发——就是 Sam 发推分享了我们最新模型写的一篇故事,一篇虚构作品。他只是发了那条推,然后一些知名作者、创作型写作者表示,他们被那个故事打动了。当然,我知道你会说:你只可能被另一个人写的东西打动,等等。行吧。但这位知名作者确实说了他被打动了。所以我认为这就是一股不断上涨的浪潮。它会走多远?会涨多高?没人知道。如果有人告诉你
便签引用
1:10:20
go? How high is it going to go? Nobody knows. And if somebody tells you that they know, actually, maybe you shouldn't listen too much. But but we can see the rate. The rate is astonishing. And I'm certainly in the camp that I'm very excited to see what the next few years are going to bring. All right, so with that, let's thank our debaters from first.
他知道,那你也许不该太当真。但我们能看到它的速率,而这个速率是惊人的。我肯定属于那一派:非常期待接下来几年会带来什么。好,那么让我们首先感谢两位辩手。
便签引用
1:10:53
Yeah, I really can't imagine two better people to present these points of view. Um, so as you may remember, uh, slido, we have the poll going. U, we asked for you to vote at the beginning. Um, oh, this is post debate. Uh so yes, please, you know, pull out your phones if you if you can uh give us your read on how what you think now and if the needle has moved at all. Well, if we see it live, I don't know when we'll call it done. Um but yeah, let's ask uh David to come back onto the stage to give us our wrap-up. Well, thank you very much everyone. Just to go back at the start, we started at 64% uh uh no and now we uh were up again. Uh on that no just to bring it back the latest is well look at that. Wow. The yeses have been voting while I've been coming up on stage and look at that. I have to say it is pretty impressive. Uh we did see 22% were yes at the start.
是的,我真的想不出还有谁能比这两位更好地呈现这两种观点。嗯,你们可能还记得,我们在 Slido 上开了投票,一开始请大家投过一次。哦,这是辩后的。所以是的,请大家掏出手机,如果方便的话,告诉我们你们现在怎么想,看看指针有没有移动。嗯,如果我们能实时看到的话,我不知道我们什么时候算截止。不过,我们请 David 回到台上来,做个总结吧。非常感谢大家。回到最开始,我们一开始是 64% 投「不」,现在「不」这边又上去了。等一下,最新的数据是——看看这个,哇。我上台的这会儿工夫,「是」那边一直在投票,看看这个。我得说,这挺让人惊讶的。开场时我们看到的是 22% 投「是」。
便签引用
1:11:55
Oh, we're going back to that too. So I don't know. Do you think it's it's kind of a tie here? What do you think? If no one's mind was changed, I think at the time. Wow. Well, how about a big round of applause for our moderator and our speakers. Thank you. All right. Thank you so much. Thanks again. Well, you know, we have to do this debate uh again soon on another topic. This was fabulous. Uh what fun. What? So much energy in the room and gosh, so much energy on the YouTube chat too. Thank you all of you watching who were participating and most of all thank you once more time to our speakers and thank you to all of you and good All right.
哦,这边也回去了。所以我也说不好,你觉得这算是平局吗?你怎么看?如果没有人的想法被改变的话,我觉得就当时而言。哇。那么,让我们用热烈的掌声感谢我们的主持人和两位讲者。谢谢。好的,非常感谢。再次感谢。你知道,我们得找个别的题目,很快再来一场这样的辩论。这场太精彩了,太有意思了。天哪,现场的气氛太热烈了,YouTube 聊天区的气氛也非常热烈。感谢所有观看并参与其中的朋友们,最重要的是,再一次感谢我们的两位讲者,也谢谢在座的各位,祝大家——好的。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

计算机历史博物馆与 IEEE Spectrum 合办的辩论中,语言学家 Emily Bender 主张大语言模型只是在形式层面预测下一个词、根本接触不到意义,OpenAI 的 Sebastian Bubeck 则主张"理解"无法严格定义、应以模型的实际效用和进步速率来衡量,双方各执一词,现场投票在辩论前后几乎没有移动。

核心要点

  • Bender 的核心论证是"形式—意义鸿沟":她在 2020 年"章鱼论文"中把理解定义为从语言形式(声音、字符)映射到语言之外的事物。语言模型从香农和马尔可夫时代的 n-gram、T9 输入法一路演进到今天,始终只处理词的形式和共现分布,从未获得"所指"这一侧的信息;没有意义就没有理解。她用日语例句演示:不懂日语的人无法从词本身获得任何东西,但能从手势、目光和现场情境推断,而模型恰恰没有这些。
  • 现代 LLM 与传统语言模型的差异被归结为两大两小:两大差异是新架构能吃下海量(且来源不透明、未经同意收集的)训练数据,以及系统被"翻转"——从给候选字符串排序变成反复回答"下一个可能是什么词"。两小差异是 RLHF(数据工人打分塑造输出偏好)和输入输出之间的额外处理(计算器转发、检索增强生成、隐藏的系统提示)。
  • "看起来理解"是人类自己投射出来的:Bender 引用语用学和心理语言学研究指出,人类理解语言时并非从词里"解包"意义,而是不断推测说话者想传达什么,为此必须想象一个背后的心灵。面对合成文本我们会反射性地做同样的事,"那个心灵完全是我们自己编出来的"。她因此全程拒用拟人化措辞,"AI"一词永远加引号。
  • Bubeck 的策略是承认"理解"不可定义,转而用进步速率说话:他举了 MATH 基准(两年半前模型约 20%,去年初已饱和)、MMLU(曾左右股市、如今无人再提)、FrontierMath(发布时 o1 得 4%,一个月后 o3-mini 达 30%,"可能没有任何一个人类能独自做到 30%")以及 ARC-AGI(o3 得 85%,超过人类约 80%)。他的结论:理解是连续量而非 0/1,且"在旁观者眼中"——哈佛物理教授甚至会说"爱因斯坦显然不理解广义相对论"。
  • Bubeck 给出的关键个人证据是一个不在网上的研究问题:他把自己 Dropbox 里一篇从未发表的数学优化论文中的问题问 o3-mini,模型没有解出来,但在 2 分钟推理后提出了一个与另一领域的联系,而这个联系他本人当年花了 3 天才发现。他说自己不在乎这算不算"理解",但它确实加速了他的工作。Bender 立即反击:Dropbox 是否与 OpenAI 有数据交易?"你不告诉我们训练数据里有什么,所以我不知道。"
  • 双方在"基准测试无法证明理解"上罕见地一致,但推论相反:Bender 用《Grover 与世界万物博物馆》比喻——任何基准都只是采样,而语言能谈论一切,所以"通用"能力不可测试,且 AGI 未定义则不成为科学目标;她主张按 Gebru 的"数据集数据表"思路做限定用途、训练数据公开的专用系统。Bubeck 完全同意要按具体部署场景(如辅助诊断)建专用测试,且批评拿 USMLE 考聊天机器人"毫无意义",但认为通用模型"就是效果更好"。Bender 回应:更好的代价是环境破坏、盗用数据和不透明。
  • 关于 AAAI 调查的 76% 研究者不看好"扩大规模通向 AGI":Bubeck 认为学术界恰恰是最慢理解正在发生什么的群体,因为 LLM 违背了此前所有的神经网络理论直觉;同时他承认纯预训练确实不能通向 AGI,推理模型是沿另一条轴扩展。Bender 认为该问题本身就是坏问卷——AGI 没有定义,且"我们是否在通往 AGI 的正确道路上"预设了一个未被证明存在的终点,科学不是选一条路往下跑。
  • 什么能让各自认输:Bender 要求一个清晰、可证伪、基于人类语言理解研究的定义,加上训练数据、架构、系统提示全部公开的可复现实验。Bubeck 的标准非常个人化:如果在他有生之年,从未从 AI 系统那里获得像读一篇人类花数年写成的数学论文那样深刻的顿悟时刻,那他就错了。
  • "人类不也是在预测下一个 token 吗"被 Bender 拒绝讨论:她称这是一种去人性化论证,等于否认她有内心生活。Bubeck 反过来说他也不假设人脑如何运作,"它内部怎么工作不关我的事,它给我内容,我自己判断有没有用"。关于具身性,Bender 认为比身体更关键的是社会关系——词义在人与人的使用中持续变化,模型根本无法参与这种关系;Bubeck 则相信纯文本、不需要任何感官就能达到通用智能,并认为文本信息密度远高于图像和音频。
  • 关于"涌现"的分歧最直接:Bender 称在训练数据不公开的情况下任何涌现声明都不科学,还要先排除其他可能的计算路径。Bubeck 说他不理解这怎么会有争议——他自己在个人电脑上训练模型解线性方程组,只用少量方程训练,模型突然能解多方程系统,"这是你自己在实验室里就能看到的现象"。

结论与值得注意的细节

  • 现场 Slido 投票辩论前约 64% 选"不理解"、22% 选"理解",辩论后主持人上台时数字来回跳动,最终被判为"基本没人改变立场"。
  • 主持人 Eliza Strickland 开场用"人和山羊过河"实验做了演示:GPT 版本在没有狼和白菜的简化题上仍生搬硬套原谜题步骤、让人把船留在对岸,而 OpenAI 推理模型直接答"两个一起上船过去就完了"。Bubeck 在结语中主动回应这一点,承认"鹦鹉和火花都真实存在",模型在某些话题上更像鹦鹉、某些上更像火花,从 GPT-3 到 GPT-4 平衡在移动。
  • 一个耐人寻味的悖论由 Bubeck 自己提出:他预测几年内模型可能证明人类未能证明的定理,但数学界会(且他认为应当)拒绝承认"理解"了该定理,直到有人类能完整把握证明——"理解也许终究是人类的旅程"。Bender 借此强调"所有学术都是对话,独自做到自己满意不算科学"。
  • Bender 批评 Google 的奥数成绩靠提交上万个答案;Bubeck 称 OpenAI 不这么做,但随即被追问 OpenAI 并不"开放"。他以"4 亿周活用户"回应,Bender 明确列出"开放"的标准:训练数据、步骤、系统提示、碳足迹全部可知。Bubeck 顺带提到 Sam Altman 承认公司在开源问题上"可能站在了历史错误的一边"。
  • 关于《纽约时报》"数字治疗师也会焦虑"的报道,两人都认为论文不好,但 Bender 建议公众看到重度拟人化报道直接跳过,Bubeck 则认为其中关于提示词如何影响模型表现的部分对研究者仍有价值,并坦言目前缺乏描述这类系统的词汇,"以前物理学家有几十年发展量子场论,我们现在没有几十年"。
  • 两人的结语立场鲜明:Bender 的核心信息是"没有什么是不可避免的,你口袋里永远有一个'不'",尤其针对教育、医疗、法律、移民等系统里用合成文本填补人手短缺的做法,她认为"比什么都不做更糟"。Bubeck 反对"恐吓策略",主张每个人在使用中自行判断价值,并以知名作家被新模型写的小说打动为例说明浪潮仍在上涨,"如果有人说他知道会涨多高,你大概不该太听他的"。
核心句型 · 9
1. To answer this question, I'm going to do four things. The first is … Then … And then …
“To answer this question in my starting 10 minutes, I'm going to do four things … The first is I'm going to offer another definition of what understanding means. Then I'm going to tell you a bit about …”
演讲开头的路标句:先报数量再逐条展开,听众可预期结构。仿写时把「四件事」换成 2~5 个动词短语,每条用 First / Then / And then / Finally 标记。
2. Nothing in X actually gives A access to Y. And without Y there's no Z.
“Nothing in the development of language models actually gives them access to the meaning part … And without access to meaning there's no understanding.”
两步否定推理:先否定前提,再用「没有 Y 就没有 Z」收束。适合论证某条路径原则上走不通。仿写:Nothing in the data gives the model access to intent. And without intent there's no deception.
3. It only makes sense because we're making sense of it.
“The answer is it only makes sense because we're making sense of it.”
同一短语的主被动对照,把「说得通」的功劳从对象转到观察者身上。适合揭示错觉来源。仿写:It only looks smart because we're reading smartness into it.
4. Extraordinary claims such as … require extraordinary evidence.
“Extraordinary claims such as large language models understanding require extraordinary evidence.”
萨根名句的套用,用于把举证责任推给对方。such as 后接名词短语或动名词结构点明具体主张。仿写:Extraordinary claims such as a cure for aging require extraordinary evidence.
5. Does it mean that …? It's really hard to say. Was it …? Yes, absolutely.
“Does it mean that it understand it's really hard to say was it extremely helpful and did it accelerate me … Yes, absolutely.”
自问自答式让步:先对难题坦承不知,再对可答的问题斩钉截铁。既显诚实又把讨论引向自己有利的问题。适合面对无法定义的概念时转移评价标准。
6. Rather than asking whether X, you should ask yourself whether Y.
“Rather than asking whether Does the chatbot understand? You should ask yourself whether the chatbot helps you understand more things.”
重构问题的标准句式:rather than 引出旧问题,you should ask 给出替代问题。仿写:Rather than asking whether the tool is intelligent, ask whether it changes what you can do.
7. No matter how good you craft your X, it's going to be very hard to Y.
“No matter how good you craft your benchmark, it's going to be very hard to extract understanding from it.”
No matter how + 形容词,表示「无论做得多好」的让步;主句用 it's going to be hard to 预判结果。注意规范用法应为 how well you craft,口语中常混用。
8. I agree with everything you just said, just for the record.
“I agree with everything you just said just for the record.”
辩论中表明立场一致的插入语,for the record 意为「郑重声明、免得误会」。后面通常接 but 或 one slight addition 引出补充。适合先示好再分歧。
9. What would you have to see to convince you that …?
“What would you have to see to convince you that you're on the wrong side of this debate?”
追问「可证伪条件」的经典问法:have to see 强调具体证据,convince you that 后接对方需要承认的命题。采访、评审、面试中都可用来检验对方立场是否开放。
词汇精讲 · 135 · 按出现顺序
vibrant /ˈvaɪbrənt adj. 0:17
充满活力的、热烈的(形容活动、城市、色彩)
decoding /diːˈkoʊdɪŋ/ v./n. 2:07
解码、解读;此处为博物馆宣传语,意为把技术讲明白
adjunct professor n. phr. 3:13
兼职教授、客座教授(非本院系编制的教职)
distinguished scientist n. phr. 3:53
杰出科学家(大公司研究院的高级技术头衔)
without any further ado phr. 3:53
闲话少说、不再耽搁(主持串场惯用语)
plug /plʌɡ/ v. 4:53
(口语)为……打广告、宣传
flagship publication n. phr. 4:53
旗舰刊物、最主要的出版物
firestorm /ˈfaɪrstɔːrm/ n. 5:42
(比喻)激烈的争论风暴;原义为火暴
precient /ˈpreʃənt/ adj. 5:42
有先见之明的(转写误拼,正确拼写为 prescient)
synthetic text n. phr. 5:42
合成文本,即机器生成的文本
preprint /ˈpriːprɪnt/ n. 6:21
预印本,未经同行评审即公开的论文稿
thought experiment n. phr. 7:01
思想实验
mimicked /ˈmɪmɪkt/ v. 7:54
模仿、仿效(mimic 的过去式)
attributed /əˈtrɪbjuːtɪd/ v. 7:54
把……归于(attribute A to B:认为 B 具有 A)
halfway decent adj. phr. 7:54
勉强像样的、还过得去的
written off as a dead end phr. 8:48
被当作死胡同而放弃(write off:一笔勾销、断定无望)
resurgence /rɪˈsɜːrdʒəns/ n. 8:48
复兴、再度兴起
loosely inspired phr. 8:48
松散地受……启发(强调只是大致借鉴)
ingest /ɪnˈdʒest/ v. 8:48
摄入、吞下;技术语境指读入大量数据
prone to phr. 10:34
容易出现……的、有……倾向的
missteps /ˈmɪssteps/ n. 10:34
失误、错误的一步
and then some phr. 10:34
(口语)而且还不止于此、还要多
step up n. phr. 11:11
升级、上一个台阶(a step up from:比……更进一步)
hallucinations /həˌluːsəˈneɪʃənz/ n. 11:11
幻觉;AI 语境指模型编造不实内容
have this down cold phr. 11:40
(口语)对某事烂熟于心、完全掌握
tweak /twiːk/ v. 11:40
微调、稍作改动
trip up phr. v. 11:40
使绊倒、使犯错
call it a day phr. 12:27
就此收工、到此为止
on your behalf phr. 13:16
代表你、替你
intuitive grasp n. phr. 13:16
直觉性的把握、深刻的领会
crux /krʌks/ n. 14:12
关键、核心(the crux of the debate)
integral part n. phr. 14:12
不可或缺的组成部分
coin toss n. phr. 16:04
抛硬币决定
anthropomorphizing /ˌænθrəpəˈmɔːrfaɪzɪŋ/ v./adj. 16:59
拟人化、把人的特征赋予非人事物
muddies the waters phr. 16:59
把水搅浑、使问题更混乱
scare quotes n. phr. 16:59
表示保留或讽刺态度的引号
susceptible to phr. 16:59
易受……影响的、容易陷入……的
salient /ˈseɪliənt/ adj. 18:30
显著的、最突出的
co-situated adj. 19:31
共处同一情境的(语言学术语,指面对面共享环境)
cues /kjuːz/ n. 19:31
线索、提示信号
transcription /trænˈskrɪpʃən/ n. 20:29
转写;此处指语音自动转文字
consentfully adv. 21:57
经同意地(Bender 的临时构词,not consentfully 即未经同意)
fine grained adj. 22:52
细粒度的、精细的
co-occur /ˌkoʊəˈkɜːr/ v. 22:52
共现、同时出现(语料统计术语)
undisclosed /ˌʌndɪsˈkloʊzd/ adj. 22:52
未公开的、未披露的
turned inside out phr. 22:52
被彻底翻转过来、里外颠倒
reinforcement learning from human feedback n. phr. 23:27
基于人类反馈的强化学习(RLHF)
thumbs up n. phr. 23:27
点赞、认可
retrieval augmented generation n. phr. 24:08
检索增强生成(RAG)
paperier-mâché /ˌpeɪpər məˈʃeɪ/ n. 24:08
纸浆糊制品(正确拼写 papier-mâché),比喻把原文揉碎重塑
coherent /koʊˈhɪrənt/ adj. 24:51
连贯的、前后一致的
pragmatics /præɡˈmætɪks/ n. 24:51
语用学,研究语境中意义如何传达的学科
reflexively /rɪˈfleksɪvli/ adv. 25:36
条件反射地、不假思索地
falsifiable /ˈfɔːlsɪfaɪəbəl/ adj. 25:36
可证伪的(科学哲学术语)
Emphatically /ɪmˈfætɪkli/ adv. 26:25
断然地、强调地
in the eye of the beholder phr. 27:06
因人而异、取决于看的人
spoiler alert phr. 28:09
(口语)剧透预警;用于提前揭示结论
subsumes /səbˈsuːmz/ v. 28:09
把……包含在内、归入
take stock of phr. 29:11
盘点、审视评估
in the running phr. 29:11
有希望获胜、在竞争之列
saturated /ˈsætʃəreɪtɪd/ adj. 30:14
饱和的;指基准分数逼近上限失去区分度
meager /ˈmiːɡər/ adj. 31:18
微薄的、少得可怜的(p56 亦出现)
probing /ˈproʊbɪŋ/ adj. 32:04
探究性的、追根究底的(probing question)
binary notion n. phr. 32:04
二元概念、非此即彼的概念
trajectory /trəˈdʒektəri/ n. 33:48
轨迹
hearken back to phr. 34:57
回溯到、追溯到(书面语)
entirely plausible adj. phr. 34:57
完全说得通的、相当可能的
adjacent to phr. 36:15
与……相邻的、相关但不完全相同的
respondents /rɪˈspɑːndənts/ n. 36:15
受访者、问卷回答者
paradigm /ˈpærədaɪm/ n. 37:16
范式
trope /troʊp/ n. 38:16
套话、老生常谈的叙事模式
object to phr. v. 38:16
反对、不赞成
branch out phr. v. 39:14
分头拓展、向新方向延伸
scholarship /ˈskɑːlərʃɪp/ n. 39:14
学术、学问(此处非奖学金)
pragmatism /ˈpræɡmətɪzəm/ n. 40:09
实用主义
premised on phr. 40:44
以……为前提
vibes /vaɪbz/ n. 40:44
(口语)感觉、氛围;此处指凭印象而非证据
minding my own business phr. 40:44
安分做自己的事、不管闲事
splattered /ˈsplætərd/ adj. 41:37
被溅满的、打花的
viable path n. phr. 42:48
可行的路径
large scale deployment n. phr. 43:42
大规模部署
thwart /θwɔːrt/ v. 44:37
挫败、阻挠
generalization /ˌdʒenərələˈzeɪʃən/ n. 44:37
泛化,指能力迁移到未见过的情形
recount /rɪˈkaʊnt/ v. 45:18
讲述、叙述
crushes /ˈkrʌʃɪz/ v. 45:18
(口语)碾压、轻松击败
goalpost moving n. phr. 46:05
挪动球门柱,指事后改变判定标准
illdefined /ˌɪl dɪˈfaɪnd/ adj. 46:05
定义不清的(规范拼写 ill-defined)
pin too much on phr. 46:05
在……上寄托过多、过度依赖某个例子
extruding /ɪkˈstruːdɪŋ/ v. 48:04
挤出(工业成型工艺);Bender 用以称呼文本生成
plausible sounding adj. phr. 48:04
听起来煞有介事的
tolerances /ˈtɑːlərənsɪz/ n. 48:33
(工程)容差、公差
oscillate /ˈɑːsɪleɪt/ v. 48:33
振荡
give or take phr. 48:33
上下浮动、正负误差
for the record phr. 49:28
郑重声明、说明在案
purpose-built /ˈpɜːrpəs bɪlt/ adj. 50:32
专门打造的、为特定用途设计的
environmentally ruinous adj. phr. 50:32
对环境有毁灭性影响的
tact /tækt/ n. 51:24
此处为 tack 的误用,take a different tack 意为换个思路
traumatic narratives n. phr. 51:24
创伤性叙事、讲述创伤经历
scrape off the internet phr. 52:13
从互联网上抓取(数据)
categorical /ˌkætəˈɡɔːrɪkəl/ adj. 52:47
绝对的、不容置疑的(断言)
wary of phr. 52:47
对……警惕、提防
surface /ˈsɜːrfɪs/ v. 52:47
使浮现、引出(surface some good behavior)
quote unquote phr. 52:47
(口语)打个引号、所谓的
rooted in phr. 53:55
扎根于、以……为基础
alluded to phr. v. 54:48
暗指、间接提到
carbon footprint n. phr. 54:48
碳足迹
provocative /prəˈvɑːkətɪv/ adj. 56:40
有挑衅性的、发人深思的
for that matter phr. 56:40
就此而言、说到这个
dehumanizing /diːˈhjuːmənaɪzɪŋ/ adj. 57:27
非人化的、剥夺人性的
get the hang of it phr. 58:22
掌握窍门、上手
embodied /ɪmˈbɑːdid/ adj. 58:22
具身的(认知科学术语,指依赖身体经验)
socially situated adj. phr. 58:22
处于社会情境中的
form attachments phr. 59:15
产生依恋、建立情感联结
multimodality /ˌmʌltimoʊˈdæləti/ n. 59:50
多模态,指同时处理文本、图像、音频等
agnostic to phr. 1:00:37
(技术语境)不依赖于、与……无关
cumbersome /ˈkʌmbərsəm/ adj. 1:00:37
笨重的、麻烦的
information dense adj. phr. 1:01:27
信息密度高的
take issue with phr. 1:02:08
对……提出异议
crunch the data phr. 1:02:08
(口语)大量处理数据
telepathy /təˈlepəθi/ n. 1:02:08
心灵感应
upvoted /ˈʌpvoʊtɪd/ v. 1:02:08
被顶、被点赞(网络投票)
political economy n. phr. 1:03:04
政治经济学;此处指围绕技术的利益与权力结构
diffuse /dɪˈfjuːz/ v. 1:03:39
扩散、传播开来
emergent /ɪˈmɜːrdʒənt/ adj. 1:04:46
涌现的,指整体出现部分所没有的性质
push back on phr. v. 1:05:14
反驳、抵制
embarking on phr. v. 1:05:14
着手、开始(一项事业或旅程)
on the same page phr. 1:06:15
理解一致、达成共识
inevitable /ɪnˈevɪtəbəl/ adj. 1:06:58
不可避免的
creaking /ˈkriːkɪŋ/ adj. 1:06:58
吱嘎作响的;比喻制度摇摇欲坠
band-aid /ˈbændeɪd/ n. 1:07:29
创可贴;比喻治标不治本的权宜之计
in your back pocket phr. 1:07:29
随时可用、备在手边
scare tactics n. phr. 1:08:10
恐吓策略、制造恐慌的手法
noted /ˈnoʊtɪd/ adj. 1:09:15
知名的、著名的
astonishing /əˈstɑːnɪʃɪŋ/ adj. 1:10:20
惊人的
the needle has moved phr. 1:10:53
指针动了,指民意或数据发生了变化
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.142"Knowledge Creation and its Risks" - David Deutsch on AGI - Centre for the Future of Intelligence 下一期 · NO.144 →College Lecture Series - Neil Postman - "The Surrender of Culture to Technology"
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]