Emily M. Bender — Language Models and Linguistics · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Emily M. Bender — Language Models and Linguistics

节目发布 2021-09-09 · Weights & Biases
艾米莉·本德 LLukas Biewald
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
本文是播客《梯度异见》(Gradient Dissent)的一期访谈。主持人卢卡斯·比瓦尔德是 Weights & Biases 的创始人,受访者艾米丽·本德是华盛顿大学计算语言学教授,「随机鹦鹉」论文的作者,也是比瓦尔德在斯坦福读本科时语言学入门课的老师。谈话围绕本德近年的四篇论文展开:大语言模型的风险、语言模型能否「理解」、基准测试的误用,以及所谓「本德法则」。本文依据现场录音编译整理。

随机鹦鹉的由来

主持人: 欢迎收听《梯度异见》,一档讲现实世界里机器学习的节目。我是主持人卢卡斯·比瓦尔德。今天的嘉宾是艾米丽·本德,华盛顿大学语言学教授。她在语言学和自然语言处理领域的兴趣极广,从社会影响、跨语言差异,一直到可以说是语言学哲学的问题。我特别高兴能和她对谈,因为我在斯坦福读本科时,她正是我「语言学导论」的老师。那是我最喜欢的课之一,到现在我还记得课上学到的一大堆有意思的事实,也是那门课让我对语言学有了终身的兴趣。

我想先从您与人合著的那篇论文谈起:《论随机鹦鹉的危险:语言模型会不会太大?》(On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?)。这篇论文在谷歌引发的风波,连我在推特上都看得清清楚楚。能否先讲讲这段经过,再进入论文本身的内容?

本德: 好。先说个题外话:这篇论文的标题里其实有一个表情符号,最后一个字符是一只鹦鹉,国际音标里没有它,念不出来。我们放它进去纯粹是好玩,因为大家都喜欢「随机鹦鹉」这个比喻。在风波发生之前,我们一度以为这篇论文最出名的地方,会是它是那篇标题里带表情符号的论文。没想到后来的事。

论文的缘起,是蒂姆尼特·格布鲁博士和玛格丽特·米切尔博士带着团队在谷歌做的工作。她们的职责是同工程团队对接,把好的实践嵌进去,让技术对更多人更有用,对世界少造成伤害。她们注意到,尤其是格布鲁博士注意到,业界正在拼命往越来越大的语言模型上推。论文里有一张表,最近两三年参数量和训练数据规模简直是爆炸式增长。

格布鲁博士在推特上私信我:「你知道有哪些论文谈过这件事可能的负面后果或者风险吗?或者你自己写过吗?」我回她:「没有,我不知道有这样的论文,我也没写过。不过随手想想,这里有五六件值得担心的事。」过了一天左右,我又说:「这看起来像个论文提纲。提纲在这儿,要不要一起写?」那是九月初,我们瞄准的会议是 FAccT(公平、问责与透明会议),最终在二〇二一年三月召开,投稿截止大概是二〇二〇年十月八日。

所以我们用一个月写出了这篇论文。之所以做得到,是因为写的人不只我们两个,也不只最后署名的四位,实际上有七位作者。格布鲁把米切尔拉了进来。这里我要强调她们都是有博士学位的,不过我跟她们熟到可以直接叫名字,下面我就叫名字了。蒂姆尼特带来了梅格和团队里另外三位成员,我带来了我的博士生安吉丽娜·麦克米伦·梅杰。我们七个人合起来,专业领域和读过的文献足够多样,才能凑出这样一篇综述。

写作过程也很有意思:我们从没开过一次全体到齐的视频会议,一切都在 Overleaf 上远程协作完成。这种做研究的方式不常见,但在这次行得通。谷歌那边的作者把论文提交了他们内部所谓的「发表审批」,通过了。我们投给会议,然后就把它放到一边,因为谁也没预料到九月要干这件事,对每个人来说这都是额外的活儿,大家各自回头去忙本来该忙的事。

到了十一月底,在我看来是毫无征兆地,谷歌那边的合著者被告知:要么撤稿,要么把名字撤下来。这里我得声明,我不在谷歌,也没拿过谷歌的资助,谷歌内部发生了什么我只有二手的了解,再加上后来媒体的报道。她们没有被告知原因,也没有机会讨论论文哪些地方需要修改,就是一句「撤稿或者撤名」。我们一时不知道该拿这篇论文怎么办:七个人的工作只署两个人的名字发出去,怎么看都别扭。我和安吉转向谷歌的合著者,说:「我们听你们的,你们希望怎么办?」她们说:「不,我们要它面世。你们两个去发表。」

这是最初的答复。然后蒂姆尼特细想之后说:「不对,这样不行。对一个被雇来做这项研究的研究者,不能这样对待。」这本来就是她的本职工作,也是整个团队的本职工作。于是她顶了回去。结果大家在媒体上都看到了,她被解雇了。谷歌说她是辞职,她的团队说她是「被辞职」(resignated),这个生造词很妙。这件事发生得够快,她来得及把名字重新署回论文上。

与此同时,梅格开始着手记录蒂姆尼特遭遇的一切。结局是几个月后她也被解雇了,但那时论文的定稿已经完成。这就是为什么第四作者的署名是「施玛格丽特·施米切尔」(Shmargaret Shmitchell)。

对所有当事人来说,这都是个令人难过的故事。这是对蒂姆尼特、梅格和团队其他成员的恶劣对待,不论他们是不是这篇论文的作者。我想那里已经变成一个很难工作的环境。对谷歌也是损失,它失去了极其宝贵的专业能力,也失去了研究界的大量善意。更让人难过的是,这件事照出了当下的状况:企业利益正在怎样左右我们这个领域的研究。

另一方面,我和合著者们始终觉得,一起写这篇论文、一起扛过后面的风波,是一段很愉快的经历。一个奇怪的结果是,这篇论文得到的关注远远超出它本来会有的。我认为它是一篇好论文,扎实的论文。从投稿版到定稿版,我们打磨得格外用心,因为知道会有很多人读。发定稿预印本的时候,我没放到 arXiv 上,因为那里的版本往往会被引用,而不是最终发表的版本。我只放在了自己的网站上,然后在推特上发了一个 bitly 短链接,这样我能看到下载次数。

单是通过这一个链接,下载量就超过了一万次,何况还有别的途径能找到它。这和我以往写过的任何东西都不在一个量级。作为研究者,这很有意思。但我也觉得这是件幸运的事,因为公众注意到了它。这项技术已经铺开了,正在以各种方式被使用,公众有机会理解到底在发生什么,这非常有价值。谷歌犯了一个严重的错误,却让我和合著者们得到了帮助公众理解这件事的机会。对此我确实感到幸运。

谷歌到底反对什么

主持人: 我很想进入论文的内容。但在此之前,您知道谷歌反对的是什么吗?他们有没有公开说过?我原本以为这一定是篇火药味很浓的论文,结果为了准备这次访谈把它读了一遍,感觉相当温和,没什么争议。这也许很难知道,但他们说过不喜欢它哪里吗?

本德: 说过。公开的说法大概是「它没有引用那些试图缓解这些问题的相关工作」。但从头到尾没人告诉我们该引用哪些工作。而且我们其实引用了一些做缓解的工作。所以我不太清楚那句话指的是什么。

你说得对。我们料到这篇论文会惹恼一些人,因为我们基本上是在说:「大家追得这么起劲的这个东西,也许该慢一点,想想它有哪些坏处,怎样才能安全地做。」总有人不爱听。但说实话,我们以为不高兴的会是 OpenAI,因为 GPT-3 是这类模型里最出名的例子,也是我们贯穿全文的例子。我们以为会惹恼一些人,没想到惹恼的是谷歌内部的人。而且这本质上是一篇综述。我们没做实验,没做分析,只是把关于大语言模型的各种相关视角汇集到一处。这篇论文竟然成了谷歌炸掉自家 AI 伦理团队这份宝贵资产的部分原因,实在出人意料。

伦理不是给人贴标签

主持人: 有意思。对这篇论文,一种读法是「我们该考虑大语言模型的负面影响」。但我可以想象另一种读法,也许不公平:一个正在做大语言模型的人读了,可能会觉得受伤,觉得你们在说做大语言模型是不道德的事。这样理解算夸大您的主张吗?论文不在我手边,但我觉得这可能会伤到一些人的感情。

本德: 我在自然语言处理的社会影响这个方向上做了不少工作,这类工作有时被冠以「NLP 伦理」的名目。我确实看到很多人对这个话题的反应是感到受伤。我想这和人们把自己同工作等同起来有关。如果你说「我们来想想正在造的这项技术在世界上是怎么运作的,怎样才能让它是有益的」,而你用「伦理」这个词来描述这件事,有些人就会读成「你在说我不道德」。我认为把谈话引向这个方向很少有什么价值。

总的来说,我相信这个领域里的人都想在世上做好事。当然有人做技术就是为了赚大钱。那种大亨式的漫画形象,为了尽可能多赚钱乐意碾碎所有小人物,这样的人大概是有的。但更常见的情况是,人们在某种体系里工作,这个体系要求他们对股东价值最大化之类的事负责,这让人很难对眼下正在给股东赚钱的东西踩刹车,很难退一步看大局。所以更有价值的谈法是:这些体系是什么,激励机制是什么,我们作为个人在体系里能做什么,而不是去评判谁道德谁不道德。这可能没有直接回答你的问题,但希望有点帮助。

主持人: 不,我明白,您是说您的观点比某些人可能读出来的要更细致。我自己开公司,热爱技术,热爱造东西,也承认确实有很多人受到伤害,我觉得有人指出问题、踩踩刹车、亮出警示,是件好事。但我能理解为什么有人会觉得有点被冒犯。我不确定是不是我想多了。

读这篇论文时我一直在想一个问题。先把赚钱放到一边,只谈研究,只谈把模型做出来、看它跑起来的那种兴奋。我对这种感觉体会很深。GPT-3 不管有多少喧嚣,它能做到的事确实惊人,我完全没料到它能做得这么好。您会主张这类研究方向应该停下来吗?或者,您希望像 OpenAI 这样的机构做出什么样的改变?这正是一个很好的例子:模型越大表现越好,并不是显而易见的事,再加好几个数量级还能持续改善,事先并不清楚。您是希望这类研究不要发生,还是以某种不同的方式发生?

该怎样做大模型研究

本德: 首先值得说明,OpenAI 其实花了不少力气去想「可能有什么坏处」、「这项技术放到世界上会发生什么」。这点很重要,我很高兴他们在做。我希望看到更多的,第一就是这类工作:可能的失效模式是什么,它们怎样影响到人;还有,当这个东西按预期正常运转时,又会怎样影响到人。OpenAI 做了一些,很好,还应该做更多。

另外,可以看看其他工程领域。在把一个东西放到世界上、让人们依赖它之前,要做各种各样的测试,要弄清楚容差是多少,什么情况下管用什么情况下不管用,适用的温度范围是多少,哪些项目必须检查和认证。这些在自然语言处理里我们几乎还没有。其他 AI 领域我不太能说,但我诚实地认为那里也有类似的问题。梅格·米切尔、蒂姆尼特·格布鲁等人在谷歌做过一个叫「模型卡」(Model Cards)的框架,就是朝这个方向迈的一步:你造了一个模型,要用它的人需要知道些什么?我希望看到更多这类东西。

与此相对的是泛滥的 AI 炒作:人们造了一个东西,很酷,很好玩,效果很好,可不知怎么这还不够。GPT-3 能生成连贯的文本,这还不够,非得说它在理解语言。它绝对没有在理解语言,这个我们等会儿肯定会谈到。

主持人: 您给了两个很好的引子。

本德: 不过这些都是连在一起的。不知为什么,AI 这个圈子的文化就是要去够那些大话,而不是去造那种范围明确、可靠、文档充分到能被安全可靠地使用的系统。我希望看到更多后一种方向。这是一点。

另一点我们在论文里也谈到了:如果如今通往成功的主要道路就是越做越大,那你就把很多语言社区排除在外了。即便是那些总体上支持得不错的语言,内部也有很多社区根本积累不起那么多数据。你也把小型研究组、小公司排除在外了,因为它们手里没有谷歌、脸书、亚马逊那种规模的数据。微软也做很多大数据的工作,但似乎没有像另外几家那样囤积数据。这很不幸,因为我认为它在一定程度上扼杀了创造力。如果整个学界都朝着一个只有少数人真正做得了的目标狂奔,那么人们本来可能去尝试的其他东西,就都丢掉了。

伤害早已在发生

主持人: 论文里还谈到一个不那么显眼的担忧:模型会以难以察觉的方式把偏见编码进去。您谈自然语言模型可能造成的伤害时,有没有现在就正在发生的例子?还是说这更多是面向未来的担心,担心随着 NLP 越来越普及会出现的伤害?

本德: 绝对是现在就在发生,因此很容易预测:如果我们不改变,它会继续发生。这方面萨菲娅·诺布尔(Safiya Noble)的《压迫的算法》(Algorithms of Oppression)是非常重要的记录。她研究的是:各种身份,本应属于拥有这些身份的群体,却是怎样在搜索引擎里被呈现、被反射回人们眼前的。她贯穿全书的例子是「black girls」(黑人女孩)和「black women」(黑人女性)这两个词组。这些情况随时间变化,她谈每个具体例子时都非常谨慎地注明日期。在她刚开始这项研究时,把「black girls」作为关键词搜索,出来的基本是色情内容。

你可能会说,「那只是数据里就是这样」。什么数据?数据从哪儿来的?读到她书的核心部分你会发现,之所以「数据里就是这样」,是因为互联网的经济结构允许人们购买身份词汇并靠它赚钱。这些问题被指出之后,谷歌零零碎碎地做了修改,现在搜「black girls」不再出色情了。但稍微戳一戳就能看出,那都是一个个事后的单点修补,没有人去系统地重新思考:搜索引擎和以广告驱动的排序机制是怎样咬住这些激励、再把它们放大的。

AI 圈里有一场持续不断的争论,推特上隔三差五就冒出来:偏见只在数据里,还是模型也有贡献?答案是:模型绝对也有贡献。

还有一层是那句「数据里就是这样」。谷歌另一个非常难堪的例子是,曾经有一段时间,谷歌图片搜索黑人会出现大猩猩的图片,具体的搜索组合我记不清了,总之难堪、恶劣、种族主义。当时一种反应是「那只是底层数据里就有的」,言下之意是「不是我们的错,我们只是把世界的说法展示出来」。可这不是真的。因为给搜索结果排序、给广告关键词竞价的那些算法,本身就在强化特定的激励。底层数据里确实有些东西,但还有问题是:数据是怎么收集的?从哪儿来的?它到底代表什么?它不是世界本身,只是某一个特定的数据集合。然后,优化指标是什么?你做了哪些建模决策?这些决策和数据里的各种偏见怎样相互作用?激励结构又是什么?诺布尔的工作是很好的切入点。

拉坦娅·斯威尼(Latanya Sweeney)在二〇一三年的一篇论文里记录了这样一件事:当时如果你输入一个听起来像非裔美国人的名字,弹出的广告之一会暗示这个人有犯罪记录;输入一个听起来像白人的名字,往往只会得到「关于某某的更多信息」。差别不是百分之百,但两组名字之间的差异在统计上很显著。这会造成真实的伤害:想象一个人在求职,或者只是在交朋友,有人上谷歌搜了他一下,旁边跳出一条暗示他可能是罪犯的消息。这就是伤害。我还可以再举一个例子。

主持人: 请讲,这些例子都很好。

模型自己也会带偏

本德: 罗宾·斯皮尔(Elia Robyn Speer)做过一项很有意思的工作,关于情感分析和词向量。情感分析这个任务,是拿一段自然语言文本,她用的是英语,去计算或预测其中的情感:这段文字对某个对象是表达正面感受、负面感受,还是没有表达感受。她用的数据集我记得是 Yelp 的餐馆评论,任务就是「读评论,预测星级」。

主持人: 对,那个数据集我用过。

本德: 然后她用了一个外部组件,词向量(word embeddings),也就是根据一个词会和哪些词共现,把词表示到向量空间里。所以一部分训练数据是领域内的 Yelp 评论,但还有这么一个组件,是在泛泛的网络垃圾上训练的。

她发现,用那种通用词向量,系统对墨西哥餐馆的星级会系统性地预测偏低。于是她深挖原因。原来那些网络垃圾里包含了关于从墨西哥、经墨西哥移民美国的讨论,其中有大量对墨西哥人极其负面、极其恶毒的看法。词向量就把「Mexican」(墨西哥的)这个词学成了和其他负面情感词相近的东西。于是,如果你在评论里说这是一家「墨西哥餐馆」,在系统看来你就是在说它的坏话,所以你不可能给它五星。

主持人: 这个例子太有意思了。我本来下一个问题就是模型在这里面扮演什么角色。这正好说明,不只是底层数据可能有偏见,模型本身也可能带着它自己的偏见。

本德: 对。词向量捕捉到了「Mexican」和许多同时与负面情感共现的词之间的共现关系,然后这个词向量被当作组件用进了另一个模型。Yelp 评论本身并没有什么理由让墨西哥餐馆评分更低。

主持人: 是的。

本德: 我不确定它们的平均评分是不是恰好一样,但这不重要,因为错误在于系统对任何一家墨西哥餐馆都预测偏低,平均起来是往低的方向偏。所以,这是一种从外部数据集里拾来的偏见。在 NLP 里,我们习惯把词向量当作非常趁手、非常细致的词「意义」表示,用它算词的相似度,包括语义相似度。如果不留心它到底学到了什么意义、什么共现,我们的系统里就会混进我们根本不想要的东西。

主持人: 那您建议怎么办?词向量确实很有用。而在这个例子里,事情似乎很简单:它在拖累性能,甚至不存在什么模型性能上的取舍。那能做些什么呢?

本德: 关于词向量的所谓「去偏」(debiasing)有很多工作,斯皮尔后来也接着做了一些。我想一部分办法是用更精心策划的数据集。关于墨西哥移民的讨论,就算你只用可靠的新闻来源,还是会碰到那些垃圾。所以单靠这个解决不了问题,但可以做得更好一些。完全没有偏见的数据集不存在,完全没有偏见的词向量也不存在,但你可以做得更好。

一步是问:用精心策划的数据能好多少?对我们已经意识到的偏见,去偏技术能做到什么?去偏技术的一个难处是,你得知道自己在找什么。在这之上,还要把失效模式想清楚。在一个具体的使用场景里,你造这项技术,利益相关方是谁?谁会受它影响?如果某人的餐馆评分因为某个原因被低估了,在实际使用情境中意味着什么?为了确认我们针对这个用例、针对最可能受到不利影响的相关方已经去偏得足够了,该测试些什么?

主持人: 我猜想,要找到一个没有偏见的人类数据集,几乎是不可能的……

本德: 对,它不存在。

形式与意义

主持人: 这正好引到我想谈的第二篇论文。为了不耽误时间,我们就往下走。我试着概括一下:这篇论文说的是,只在您所说的「形式」(form)上做语言建模,也就是只看流过来的词、只看词串,像 GPT-3 这类模型那样,是不可能有理解的,不可能有真正的理解。我觉得有意思的一点是,您说写这篇论文是为了终结推特上的某场争论,而我完全不知道有这场争论。我大概是带着比我自以为的更少的背景闯进来的。您能不能概括一下有哪些立场,以及您想终结的是什么?

本德: 我总是在推特上跟人吵起来,对方声称语言模型在理解东西。我说:「不,它们没有,不可能有。」这里得先说清楚我们说的语言模型是什么。语言模型是像 GPT-3 或 BERT 这样的东西,训练数据是一大堆文本,训练任务是预测文本里的词。有时是顺序预测,有时用掩码语言模型目标,把某些词挖掉,训练目标是「把这些词填回去」,然后更新模型,梯度下降,等等。

作为语言学家,我看着它会说:「有用的技术,有意思。」在语音识别和机器翻译这类任务里它极其有帮助,因为那里有个重要的子任务是「最可能的字符串是什么」。在语音识别里,声学模型说「这段声音可能对应这一批文本串」,然后语言模型进来说:「『It's important to wreck a nice beach』(毁掉一片好海滩很重要)是句荒唐话,『It's important to recognize speech』(识别语音很重要)才合理,所以后者排前面。」这是它们最初为之设计、也擅长的基于形式的任务。

过去几年神经语言建模革命带来的变化是,从语言模型里抽出的词向量,是对词分布的极其精细的拟合表示,非常有用。有些词向量还是上下文相关的,也就是关于这个词及其可能共现的信息不再是它在所有文本中的总体情况,而是它在当前上下文里的情况。极其有用,但这和理解语言不是一回事。

我不断跟非语言学家吵,他们非要说「是一回事」。所以我和亚历山大·科勒(Alexander Koller)写了这篇论文,就是想说:「好,看这儿,这是为什么不是的论证。」希望能了结这场争论。结果没有,人们还是要来跟我吵。

真正难看清的地方,也是语言学在这里的价值所在,在于我们使用语言时的状态。抱歉,我要搬出一位哲学家了:海德格尔有个概念叫「被抛」(thrownness)。当你意识不到自己正在使用的工具时,你就处在被抛的状态。想想在键盘上打字,顺手的时候键盘就消失了;然后某个键卡住了,键盘一下子又「在那儿」了。语言也是一样。我们说一门流利的语言时,语言对我们是不可见的,直到有什么东西迫使我们去注意它。语言学当然就是专门注意语言的,所以语言学家习惯这么做。

当我们说把词喂给语言模型时,一定要区分两种「词」:一种是作为字符序列的词,另一种是作为形式与意义配对的词。因为语言模型看到的只有字符序列。要想象那是什么感觉,最好想一门你不会的语言。你不会哪门语言?

主持人: 普通话。

本德: 普通话,好。你不会说普通话,我想你也不会读中文。

主持人: 肯定不会。

本德: 也许认得几个字?

主持人: 我会读日语,所以有些重叠。

本德: 那我们再走远一点。你会读切罗基语吗?

主持人: 不会,绝对不会。

本德: 好。切罗基语有一套很妙的音节文字,每个字符代表一个音节。如果有人给你看一大堆切罗基语文本,你看着它的那种体验,比你看英语文本更接近计算机在做的事。因为你看英语时,意义那部分是躲不掉的,英语是你会说会读的语言。普通话介于两者之间,你会认出几个从日语汉字里熟悉的汉字,感觉不完全一样。

图灵测试为何失效

主持人: 我不想跟您争,但我想替另一边说几句。这个问题我没深想过。我这辈子看到的是,这些语言模型用它们那套策略,效果好得超出我的想象,而且似乎抓到的细节越来越微妙。我小时候学过图灵测试,它表面上看是个不错的理解力测试:如果你和某个东西对话,分不清对方是自动系统还是人,那我们就可以说它有智能。在我看来,这些语言模型似乎就快通过图灵测试了。要怎样您才会觉得某种自动化技术真的理解了它所读到的东西?

本德: 关于图灵测试,我首先想说它为什么不管用。我很不愿意反驳图灵这样的巨人,他的工作极其重要,是奠基性的。

主持人: 但那是一百年前,漏掉点什么也正常。

本德: 七十年?

主持人: 七十,好吧。八十?七十?行,七十,抱歉。

本德: 事实证明,人太乐于从语言中找出意义,太乐于替一段话脑补出让它说得通的背景。所以我们并不适合充当图灵测试里的测试者。这就是它不管用的原因。

语言模型能产出看似连贯的文本,那是些高概率的序列:给一点噪声和一个起点,根据全部训练数据,接下来最可能出现什么。出来的东西是我们能读出意思的,于是我们很容易被骗,以为它真的有意在传达那个意思。

你问什么能证明机器有理解。我想一部分答案是:谈谈它以某种方式和世界接口的能力。我们确实有这样的例子,机器在受限的领域里、对受限范围内的事情,是有理解的。当你叫你家那个企业间谍机器人替你做件事,它做到了,那它就理解了。

主持人: 等等,什么是「企业间谍机器人」?能不能说具体点?

本德: 我在挖苦 Siri、Alexa、Google Home 这类东西的隐私问题。

主持人: 哦,明白了。

本德: 三星的 Bixby 也是这类,微软以前有 Cortana。当你让它们设个计时器、开灯、拨个电话,它做到了,那么在一定程度上它确实理解了。它之所以能理解,是因为它的训练设置不只看语言,还看语言之外、需要与语言对应起来的东西。这是一种理解。问题在于,对一个想在更广泛的范围内做到这一点的人来说,怎样设计任务,使它需要在世界中采取某种行动,而不能被语言模型硬推过去,靠一句「这是接下来最可能出现的东西」蒙混过关。

章鱼思想实验

主持人: 您得讲讲那个章鱼思想实验,它太形象了,我有几个问题。

本德: 好。章鱼思想实验讲的不只是能否理解,而是学会理解。这是它和图灵测试、和塞尔(Searle)的思想实验的区别:那两个都是说「假设有人已经把整个系统搭好了」,然后我们去测它有没有智能,或者从哲学角度说它仍然不算理解。系统已经在那儿了,我们只是思考它、测试它。

章鱼实验说的是:假设我们有一个东西,我们设定它是超级聪明的。这也是我们选章鱼的部分原因。其实最初是海豚,但我们觉得章鱼天生更有趣。而且海豚的环境和人类的环境太接近了。我们要的是一个被设定为超级智能的东西,而章鱼通常也被认为是聪明的动物。它要多聪明有多聪明,这不是问题所在。我们假定了智能,但只让它接触语言的形式。

在我们的场景里,两个说英语的人各自漂流到相邻的两座荒岛上。岛上没有别人,但以前的居民铺过一条海底电报线。两个人可以彼此通信。他们怎么发现电报机的、怎么知道对面有人,这些我们不交代,就当它存在。思想实验嘛,可以这么干,就像「假设一头球形的牛」,只不过我们不需要球形的牛。

所以有一条电报线,两个人叫 A 和 B,他们用摩尔斯电码编码的英语交谈。一只超级聪明的深海章鱼,我们叫它 O,游过来搭上了这条线。章鱼能感觉到摩尔斯电码的脉冲在线里通过。问题是,章鱼到底能学到什么?因为它超级聪明,时间要多少有多少,记忆要多少有多少,它能极其精确地建模「接下来可能出现什么」的模式。

故事里,章鱼不知为何觉得孤独,决定剪断电缆,假扮成 B 跟 A 说话。回头想想,可怜的 B 就此与世隔绝了,所以也许章鱼同时也假扮成 A 在跟 B 说话,不过我们没写这部分。问题是:在什么情况下章鱼能一直骗过 A,让 A 以为自己在和 B 说话?我们说这在某种意义上是一个弱化版的图灵测试。图灵测试的设定里,A 的任务是判断「我在和人说话吗」;而这里有欺骗,A 根本不知道章鱼的存在。

如果只是闲聊寒暄,那些东西照着模式来就行,只要输出内部连贯,说错点什么也无关紧要。就算有点不连贯,也许 B 只是在犯傻,没什么大不了。这种情况 O 能蒙混过去。但一旦谈话走向 A 真正在乎把想法传给 B、也在乎 B 传回来的想法,章鱼要维持这种「沟通良好」的假象就越来越难。我们举了个例子:A 造了一台椰子弹射器,章鱼能回一句「很酷的发明,干得好」之类的话,尽管 A 问的其实是「你造出来之后怎么样了」。

但章鱼没有关于椰子、绳子这类东西的任何经验。它没法对这些东西进行推理,甚至不知道 A 在说的就是这些东西。它能做的只是回一句「在这个语境下,一个回应最可能是什么形式」。O 之所以能蒙混过去,是因为 A 愿意从这些话里读出意思。在这个场景里,O 没有任何意义。最后,一头熊出现并攻击 A,A 对 O 说,其实是对 B 说:「救命,熊在攻击我,我手里只有两根棍子,该怎么办?」这时 O 完全没用了。所以我们说,如果 A 没被熊吃掉,这就是 O 铁定通不过图灵测试的时刻。

我们还拿 GPT-2 试了试,看它会怎么回答。答案很搞笑。用词落在对的话题范围内,所以出来的东西挺好玩,我鼓励大家去看论文附录,我们把这些放在那儿了。但它永远不会有帮助,它也并没有在表达任何交流意图。

只靠文本能学会语言吗

主持人: 我得说,不了解背景就读那篇论文,我读得很享受。对我来说尤其享受的是,这个思想实验很具体,既形象,又让人忍不住想:「嗯,我自己怎么看?」

我一直在想的是,我觉得我学过很多自己没亲身经历过的东西,尤其想到学数学,那么多抽象的主题。我感觉我在某种意义上几乎就是通过形式学会数学的,全在脑子里,边学边在脑中想象。从一串词里学会推理那些没见过、没经历过的事情,似乎是可能的。我还记得给盲人学生批改数学作业,很有意思,他们在数学课上推导问题的方式,看起来就像在脑中想象图形,尽管他们天生失明。所以我不太确信,章鱼如果听了全部的语言,就不能以某种方式弄明白弹射器是干什么的。

本德: 如果章鱼真的有机会学英语,那可以。但它没有,因为它从来没有获得最初的那个落脚点(grounding)。我们绝对可以通过语言学到直接经验之外的东西。反过来说,你作为一个视力正常的人,如果想了解盲人的生活是什么样的,你可以听或读盲人对此的讲述,从中学到东西。这我们完全做得到。但我们做得到,是因为我们已经习得了语言系统。我们用语言交流时,绝对会把连自己都没经历过的想法和事情讲给对方听。我们发明东西,再把它传给别人。但我们能这样做,靠的是一个共享的系统,它告诉我们:可能的形式有哪些,哪些是合格的词和句子,这门语言用哪些声音,词怎么构成,句子怎么构成,它们各自对应的固定意义是什么。然后我们用这些固定意义去猜测交流意图。

章鱼的问题不是它不聪明,我们说了它超级聪明。也不是它懂了这门语言之后还理解不了那些东西。而是它接触语言的方式,不足以让它把语言当作一个语言系统学会,它能学到的只有分布模式。

主持人: 那是什么阻止了章鱼像人一样随着时间学会语言呢?

本德: 论文里我们谈到了人类的语言习得。第一语言的习得,关键在于共同注意(joint attention)。婴儿学语言,是从和照料者之间的社会联结开始的,先明白照料者在向自己传达什么,再把词和那些交流意图对应起来。儿童语言研究的文献谈到共同注意的重要性:孩子学会词,是在照料者跟随孩子的注意力、注意同一个东西、然后给出词的时候。这种体验,这种对应,章鱼得不到。它只看到词流过去。

主持人: 您认为有没有可能存在某种算法,能接收一串词,然后在这个意义上理解它们?

本德: 自然语言理解是极其困难的问题,因为它不只依赖语言系统,还依赖世界知识、常识、推理,各种各样的东西。这里我要说得比我实际有把握的更笃定一点:有两种说法差别很大。一种是「我要造一个算法,它理解语言结构,理解语言意义,理解这些意义怎样映射到一个世界模型上,然后用这些来理解」;另一种是「我要造一个只拿到语言形式的系统,并假定它会以某种方式抵达理解」。

所以,是的。如果算法的训练输入里除了形式还有更多东西,你能走得远得多。那可以是视觉落地,可以是向人提问求答的能力,可以是知识库,可以是某种具身形式里的其他传感器。我不是说自然语言理解不可能、不值得研究。我是说语言建模不等于自然语言理解。

主持人: 我确认一下:只消费语言,不加所有这些额外的东西,您的主张是没有任何算法能仅凭这个真正理解语言?

本德: 我说的语言,指的是形式。想象你被扔进泰国版的国会图书馆,周围有你想要的任何一本泰语书,但只有泰语。不知为什么这座图书馆没有泰汉、泰法、泰英词典,就只有泰语。你能学会泰语吗?

主持人: 我觉得能。难的地方在于我已经有一门语言了。但我觉得我……

本德: 那你会怎么做?周围只有堆积如山的泰语书,别的什么都没有,你学泰语的第一步是什么?

主持人: 第一步会做什么?我不确定。您觉得我学不会泰语?

本德: 我是好奇你会怎么做。你作为一个人,能学会泰语吗?当然能,你可以去上泰语课。

主持人: 不不,我是说在这个情境里,就这么被扔进去。人们确实学会过……在没有人还懂的情况下,人们是怎么学会象形文字的?他们必须找到罗塞塔石碑之类的东西吗?还是说……

本德: 罗塞塔石碑正是解开圣书体的钥匙。如果没有这样的东西,你只能诉诸关于分布的假设,问:「关于这些文本所处的世界,我们知道些什么?关于语言一般怎样运作,我们知道些什么?」你可以说:「根据词频分析和词长,这看起来是一门有独立功能词、而不是有大量形态变化的语言。那个东西可能是冠词,那个可能是某个动词的一种形式。」你可以做这类分析。这不是语言模型在做的事。而要从这些结构性的东西走到关于意义的东西,你必须猜测文本在描述什么,你必须引入一些世界知识,问「这个猜测吻合得怎么样」。

我问你会怎么做的时候,心里想的可能答案是:「我会去找一本带图的百科全书。」这就有了视觉落地。或者:「我会去找一本从封面就能看出是《好奇的乔治》泰语译本的书。」

主持人: 这些建议都很好。

本德: 是的。但所有这些都是在引入外部的东西。一旦有了落脚点,你就能在上面往上搭。这是一条有意思的路。但如果你只有形式,形式给不了你这些信息。

语言模型的未来

主持人: 太有意思了,谢谢。这个话题的最后一个问题:您是否预计这些语言模型会撞上我们切实感受到的问题,然后不得不改变路线?还是说,随着我们对自然语言应用的要求越来越高,它们会自行适应,找到办法把外部信息纳进来,就像找到那本《好奇的乔治》译本一样?

本德: 我认为语言模型会继续有用。从上世纪五十年代香农的工作以来,语言模型一直是语言技术的重要组件,这是有长久历史的。但我猜想,预测未来太难了,也许该说我希望看到的是:我们对什么管用、什么样的失效范围是可接受的、需要什么样的保险机制,形成更严格的标准。

人们会发现,如果你的应用真的要求对说出来的每一个字负责,那么把语言模型放在这种应用的中心,是非常脆弱的做法。我猜到了那一步,我们会把语言模型从中心挪开,让它回到从若干候选输出里挑一个、或者提供词向量的位置。它们不是通往通用语言理解的一步,尽管炒作把它们说成是。

这是一类问题。如果你必须对说出的话负责,你就不会想要一只随机鹦鹉。你要的是一个能可靠地替你说话的东西,而不是编出听起来不错的话。另一件事是,如果我们认真对待偏见、训练数据中偏见的编码与放大这些问题,我想我们会发现,我们想要的是能从更小的数据集里榨出更多东西的算法,这样才能更好地策划、记录、更新数据集,让它们跟上世界的变化。而不是现在这条依赖超大语言模型的路。

这些是我的猜测。还有环境的角度。准确说,「能耗」这个角度既关乎环境,在一定程度上也关乎技术。越来越多的人,施瓦茨等人、斯特鲁贝尔等人、亨德森等人,一批工作都在说:「我们做事的时候,也要把环境影响、碳足迹量出来,好把力气用到越来越高效的做法上。」这是一个角度。另一个角度是,很多场景下你手边没有整片云。要在移动设备上做计算,你不可能塞一个庞大无比的语言模型进去。

所以有压力去找更精简的方案。我认为这是双赢:环境上赢,技术的灵活性上也赢。

基准测试的误用

主持人: 完全同意。这也正好引到您几篇关于基准测试的论文里指出的问题,我很想聊聊。也许先从「什么是基准测试」说起,虽然大多数人大概都知道,然后再谈它可能的陷阱。

本德: 好。先说明,这篇论文叫《AI 与「全世界的一切」基准》(AI and the Everything in the Whole Wide World Benchmark),去年在 NeurIPS 的「机器学习回顾」研讨会上发表,是和德布·拉吉、亚历克斯·汉娜、艾米丽·登顿、阿曼达琳·保拉达的合作。又一次合作,这回我们倒是真开会说话了,但这几位里我到现在只当面见过阿曼达琳,她是我们系的博士生。疫情生活嘛。

我们凑到一起,是因为在讨论基准测试是怎样被 AI 炒作机器误用的,怎样被那些追求通用性、把基准能说明的东西说过头的 AI 研究误用的。基准测试,基本上就是一个标准化的数据集,通常带有某种金标准标注。当然也可以给标注是内在的任务做基准,比如语言建模,实际出现的下一个词就是金标准。思路是,你可能有一套标准化的训练数据,也可能没有,然后有一套标准化的测试数据。人们可以拿不同的系统在上面测,从而能说「在这种训练方式下哪个系统更有效」,或者「给定这套训练数据、对那套测试数据,哪个更好」。这就是基准测试。

你之前问我能不能概括基准测试的问题。我有意见的不是基准本身,而是它们被使用的方式。我觉得这是「地图不是疆域」的一个例子。人们会说「这是计算机视觉的基准」,指 ImageNet;或者「这是英语自然语言理解的基准」,指 GLUE 和 SuperGLUE。然后人们会说……我真的在微软的一份公关材料里看到过,说计算机现在比人更懂英语了,因为某个系统在 GLUE 基准上的得分超过了一些人。这纯粹是狂妄的夸大,是对基准用途的误用。

夸大有什么问题?它把科学搞乱了。如果结论和实验对不上,我们做的就不是科学。我们生活在一个 AI 炒作的世界里,这意味着人们更容易买账、更容易部署那些并不像宣传的那样运转的方案,因为他们所处的世界里,有人告诉他们微软造出了一个比人更懂英语的系统。当然你也可以造一个 AI 系统去做别的什么不靠谱的事,比如「从一个人微笑的方式猜他的政治倾向」之类,这毫无道理。但我们所处的世界里充斥着关于 AI 的各种说法、各种夸大,这让那些荒唐的说法听起来也比它们应有的更可信。

这是我看到的问题。但基准测试很重要。计算语言学的历史上有过一段时间,给 ACL(计算语言学协会)写论文,你只要说「这是我的系统,这是我怎么造的,这是几个输入和输出的例子」,完事。后来统计机器学习的浪潮来了,带来了「共享任务评测挑战」的方法论,算是基准测试的历史前身:NIST 和其他机构会说,「我们要做语音识别,想切实了解这些不同系统之间的比较。所以我们办一场共享任务评测,所有人拿同样的训练数据,我们留一份谁也看不到的测试数据。到时候所有参赛者提交系统,看看结果。」

比起之前的做法,这在科学上是进步。但这不是全部。如果你想理解系统是怎么运作的,想知道下一个系统该怎么造,你不能只在某个标准数据上测一下就算了。你还得看:它犯的是哪些类型的错误?不同系统的差别不只在总分上,还在失效模式上,哪些输入对它们管用、哪些不管用,如此等等。而不是「好,我拿了最高分,完事」。

主持人: 说得好,我没什么要补充的。您能不能再多讲讲……我觉得这篇论文很好的一点是,你们给出了非常具体、非常理智的建议,提出了几种基准之外的替代方案。能不能给听众过一遍?

基准之外的方法

本德: 当然。说是「替代」,其实更像是「补充」。基准可以用作一种健全性检查,比如「我的系统真的比一个极其朴素的基线好吗」,或者「我想把几个系统头对头比一比,用这个基准」。除此之外,你还可以用测试套件(test suites),它是专门编排出来的,用来覆盖你希望处理好的特定类型的案例,而不是随手抓一把测试样本里碰巧出现的东西。

你可以做审计(auditing),它和测试套件非常相近。比如乔伊·布奥兰维尼、蒂姆尼特·格布鲁和德布·拉吉对人脸识别数据集的审计,他们系统地构建了一个测试集,覆盖两种性别和一系列肤色,然后问:「在这群人身上,准确率到底是不是均匀的?」他们发现不是。所以这是……

主持人: 这和基准有什么不同?听起来就像一个基准,不是吗?

本德: 区别在于,基准通常不是这样构建的。你可以想象有人构建一个系统地铺开整个空间的基准,但这不是通行做法。通行做法是:「我们从某处抓一批数据,留百分之十做测试,其余百分之九十做训练;或者百分之八十训练、百分之十开发。」基准的一般做法是「抓一个数据样本,看这东西总体上表现如何」;而测试套件和审计是「建立一套测试机制,让我们能摸出它失效模式的轮廓」。不是「它平均表现如何」,而是「它在这种情况、那种情况、另一种情况下各表现如何」。

还有对抗测试(adversarial testing),这个名目下有几种不同的做法。有时人们会把以前的系统表现差的例子全收集起来,做成一个特别难的测试集。这有意思的地方在于它能把那些太容易的白送分过滤掉,但它未必能引导任何东西在某个具体用例上表现得更好。因为它只是「挑出对上一个模型来说难的东西」,而不是「什么是特别重要必须做对的,什么在我们的用例里特别常见」,等等。

这是一种对抗测试。另一种是我们在「造它、破它」(Build It, Break It)共享任务里做的。那是二〇一七年,艾莉森·艾廷格、苏达·拉奥、哈尔·多梅和我一起办的,有造系统的队伍,也有「破坏者」队伍。破坏者的目标是找最小对立对:两个例子差别极小,但系统对一个管用,对另一个不管用。这是一种摸清什么导致系统失效的办法。

你还可以做错误分析。拿基准的测试集或开发集,进去看:「出现的都是哪些类型的问题?」很多依赖语言模型的系统在否定上表现很差。否定对意义至关重要,但往往只是一个很短的词或者子词,所以很容易漏掉。想想语音识别或机器翻译,二十个词里漏掉一个,漏的是哪个词非常要紧。把「a」换成「the」,多数情况下不会出什么问题。但如果漏掉了一个「not」呢?

主持人: 对,有道理。

本德: 所有这些,归根结底是在看:我们要造的是什么,我们在什么上测,它和背后的目标用例怎样吻合,什么管用什么不管用;对不管用的部分,后果是什么,这种失效在现实世界里发生会怎样,可能的原因是什么,是什么把我们绊倒了。这些才是我们希望看到的,而不是「排行榜主义」,所有人挤在一起往榜顶上爬,那感觉不像是真的在做什么。

谈论 AI 进展速度的人,特别爱说排行榜换得多快,各种基准上的最新水平(SOTA)涨得多快。我总是想:「对,然后呢?」从科学上更好地理解世界,或者造出不只平均情况好、最坏情况也好的技术,这些数字究竟意味着什么?

主持人: 有意思。读那篇论文时我想到几件事。我职业生涯刚开始时,大概正赶上 ACL 论文那个时代的尾巴,那时论文好像就是挑几个管用或者不管用的例子,看起来很荒唐。我记得早期有了基准之后,有人的准确率比直接猜最常见的类别还低,你可以争辩说那样也更好,也确实有人这么争辩,但在我看来有点荒唐。我还记得您课上讲的一个轶事,好像是诺姆·乔姆斯基说「妈妈不会教孩子语言」,可实际上她们会,只是没人费心去查。这种事让人抓狂,所以我从那时起就很珍视基准测试。

不过您的建议不但合理,我觉得在公司里很多已经是标准的最佳实践了。你不会不试一试、不摸清楚它哪里行哪里不行就发布一个新模型。你不会说:「我们留了百分之十的数据做测试,发布吧。」这似乎正是一个在公司里比在学术文献里更常见的做法,大概是因为看一个数字然后说「我们打败它了」更容易。但这显然有缺陷。总之我觉得这是篇很好的论文,建议很好,每个人都该照做。

本德法则

主持人: 我也想确保谈到最后一篇论文,这篇很酷,我想让大家知道。什么是「本德法则」(Bender Rule)?它为什么重要?

本德: 本德法则,或者叫「#本德法则」……

主持人: 是带井号的?

本德: 两种都行。

主持人: 先说它是什么,然后我有几个关于最佳实践的问题。

本德: 它本身就是一条最佳实践:你应该始终说明你研究的是哪种语言,哪怕只是英语。这个道理我从二〇〇九年左右就开始扛着到处宣讲,时不时站上去喊一嗓子。那时我看到很多前神经网络时代的统计 NLP 工作,基本上是在说「看,妈妈,不用语言学」,声称因为没有硬编码任何语言学知识,所以系统是「语言无关」的。而这些号称语言无关的系统,绝大多数只在英语上测过。你还会看到很多论文,说是关于机器阅读,或者关于情感分析,实际上呢,是关于英语的机器阅读、英语文本的情感分析。

另一面是,如果有人研究切罗基语、泰语、汉语或者意大利语,那份工作更难被研究会议接收,因为它被认为是「特定语言的」,而研究英语的工作不知怎么就成了「通用的」。这对科学是大问题,对造出真正跨语言可用的技术也是大问题。所以我一直在到处催人:真的去跨语言测试,说出你研究的是哪种语言。二〇一九年,有三四个人,名字都列在我发在 The Gradient 上的那篇文章里,把这种做法称作「本德法则」。名字不是我起的,但既然有了,我就顺水推舟。

一部分原因是,这个问题问出口会让双方都下不来台。如果有人写了篇关于机器阅读的东西,我走过去问「什么语言?」,这是个蠢问题,因为显然是英语。所以这让我难堪。对他们也有点失礼,因为这个问题等于在说「你本该写明的」。我不介意人们把它怪到我头上。我之所以把这个话题标签用起来,一部分正是因为:如果有人想问这个问题,又觉得问出来有点傻,他们可以把它推到我身上,我很乐意把名字借给这件事。

假如 NLP 从泰语起步

主持人: 明白了,很好。这个问题可能不好答,但我脑子里冒出来的是:英语太特殊了,大概有一大堆自己的怪癖。如果 NLP 是从泰语或者切罗基语起步的,您觉得它会有什么不同?英语在很多方面一定是不寻常的,对吧?英语有哪些不寻常的特征,让世界本可以走上另一条路?

本德: 绝对有。那篇文章里我列了一批。第一,英语是口头语言,不是手语。如果 NLP 是从美国手语或别的手语起步的,会非常不一样。

主持人: 显然。

本德: 这是一个大的分岔点。第二,英语有一套非常成熟、高度标准化的书写系统。世界上很多语言根本没有文字,有文字的语言里,很多也达不到英语这种标准化程度。另外,很多语言平均来说比英语有多得多的语码转换。

主持人: 什么是语码转换?

本德: 语码转换(code switching)就是在同一段对话里,有时甚至同一个句子里,使用多种语言。在双语或多语普遍的社区里这很常见。比如你和我,你也会说日本語(Nihongo)对吧?你学漢字(kanji)的时候,最喜欢用什么方式来勉強(benkyou)它们?我不是熟练的语码转换者,所以说得别扭又蠢,但意思是这个。

主持人: 我记得……对,这种情况我确实经历过。

本德: 当然,英语也参与很多语码转换。但同时也有海量的单语英语数据。而当你去看印度各语言的社交媒体数据,其中极大一部分是和英语混着来的。这里就冒出一整套有意思的技术挑战。

我们所处的世界里,最早的数字系统适配的是低位 ASCII,英语恰好全能装进去,最方便。英语的语序相对固定,形态相对简单,任何一个词出现时都只有寥寥几种形式。对比土耳其语,同一个词根据说能有上百万种屈折形式,这就改变了你处理数据稀疏的方式,也改变了数据稀疏本身长什么样。我们的正字法一团糟。前两天有人在推特上问:「为什么我们做字素到音素的预测,却不做音素到字素的预测?」字素到音素,就是「给定一个字母,最可能的发音是什么」,这是文本转语音系统碰到词表外的词时的重要组件。音素到字素则是「给定一个音,最可能的字母是什么」,这不是一个常规任务。我怀疑这在多大程度上是因为英语的书写系统晦涩而混乱。

主持人: 对,听起来是个不可能的任务。

本德: 正是。可如果你看日语,撇开汉字,只用假名转写日语,那直截了当得多。西班牙语的字素音素映射在两个方向上都非常透明、一致。所以细到这种程度,英语书写系统的属性都在起作用。英语喜欢在词之间放空格,句末放标点。这些我们视作理所当然的东西,比如把文本切成句子和词很容易,在别的语言里根本不成立。

所以我说不上来 NLP 会长成什么样,我只能告诉你分岔点可能在哪儿。

主持人: 不,这些很有趣。这些差异太有意思了。

本德: 你可是自愿选修了语言学课的人,我不意外。

主持人: 我就是觉得语言学太酷了。作为一个外行,如果你不懂它,它真的让人大开眼界,因为你就泡在语言里,却从没注意到那些模式。尤其是语音学,大概是最深的,你会惊叫:「天哪,这两个音是不同的?」我永远不会注意到。而且很容易做个思想实验就发现自己错了。我爱这些东西。

我早期的工作大多是用各种方式解析日语。我记得这好像没有妨碍发表,但让人吃惊的是,对这么一个必要的任务,相关工作竟然那么少。我第一份工作主要是处理日语,那时研究文献少得惊人,公司内部的经验反倒比文献多。

本德: 研究界发生的事情是,那类解析问题被认为「解决了」,因为人们在英语上取得了一定进展,而这被误认为是问题在总体上解决了。那么,这里新在哪儿?新在这是日语。可要让人们看到这一点其实很难。我对后来被叫作本德法则的东西,目标就是「把英语放回它该在的位置」:我做了英语的工作,我要说明它是英语的,好给其他语言的工作留出空间,那些工作同样重要、同样新颖、同样有价值。走着看吧。领域里不同的人会定期去数,一届 ACL 会议上有多少论文真的研究了不同语言、真的说明了研究的是哪种语言。变化没有我希望的那么快。但有一些很好的进展。通用依存(Universal Dependencies)项目为很多很多语言做出了树库,激发了大量真正跨语言的工作,很令人振奋。

低资源语言与社区

主持人: 我觉得最引人遐想的一些工作,是跨所有语言构建语言模型,或者构建能以有趣方式利用语言对的翻译模型,让数据多的语言帮数据少的语言。您觉得这是有成果的方向吗?还是说它会以某种方式把我们的偏见编进去?

本德: 这当然有意思。既然我们依赖这些吞数据的庞然大物,而很多语言就是没那么多数据,那么看看从大语言迁移过来能做到什么,是一条有意思也有价值的路。我认为值得问的问题是:「这在多大程度上把英语所编码的世界观强加到了其他语言的结果上?」「由此会导致什么?风险是什么?」再拿它和另一条路比较:「如果只做单语,我们只能走到这一步,所以我们接受那些风险,想办法缓解。」这类工作我认为很重要。

还有一点非常非常重要:你得确认你在低资源语言上用的是真实的数据。之前曝出一件事,我记得是苏格兰语,整个苏格兰语维基百科是一个不会说苏格兰语的人写的。维基百科是 NLP 里非常重要的数据来源,所以任何声称在为苏格兰语做点什么的 NLP 系统,其实都没有。

这方面一个极好的榜样是一个叫马萨卡内(Masakhane)的研究共同体,一个横跨非洲大陆的研究计划,用参与式的研究为非洲语言创建语言资源。他们在怎样把社区建起来上做了很有意思的工作,让人们能以译者的身份来贡献,不是机器翻译专家,而是真正翻译语言的人。去年 EMNLP 的 Findings 里有一篇很酷的论文介绍了马萨卡内项目。

这类工作的要义是:如果你要做低资源语言,一定要和社区建立联系。谁会是这项技术的使用者?然后你才能了解:「大家关心什么?你们希望在多大程度上引入我们从大语言那里能做到的东西?还是宁愿保持单语,看看能走多远?」倾听社区,让社区参与研究。我认为马萨卡内在这方面是很好的榜样。

主持人: 很好。这似乎是个合适的结束点。我们已经大大超时了,您非常慷慨。非常感谢,和您聊得很愉快。

本德: 我也是,谢谢。我可以一直说下去,所以很感谢有这个机会。

本期讲者
艾米莉·本德华盛顿大学语言学教授、计算语言学硕士项目负责人。《论随机鹦鹉的危险》《Climbing towards NLU》的合著者,「Bender 规则」以她命名,长期主张区分语言形式与意义、批评 AI 过度炒作。
Lukas Biewald机器学习工具公司 Weights & Biases 联合创始人兼 CEO,播客 Gradient Dissent 主持人。此前创办数据标注公司 CrowdFlower,本科在斯坦福曾修读 Bender 的语言学课。
章节 · 点击跳转视频
0:01 开场:字符序列不等于词 ▶ 正在看
0:42 随机鹦鹉论文的成稿与谷歌风波 ▶ 正在看
8:10 论文到底说了什么惹到谁 ▶ 正在看
12:44 大模型研究该停吗:工程测试文化 ▶ 正在看
16:24 偏见的四个真实案例 ▶ 正在看
24:06 去偏:没有无偏数据,但能做得更好 ▶ 正在看
26:01 形式与意义:语言模型为何不理解 ▶ 正在看
30:42 图灵测试为什么不成立 ▶ 正在看
34:45 章鱼思想实验:只听不接地 ▶ 正在看
39:51 反驳与回应:人靠语言学到经验之外 ▶ 正在看
44:28 泰语图书馆与罗塞塔石碑 ▶ 正在看
46:48 对语言模型未来的三个预测 ▶ 正在看
49:38 基准测试:地图不是疆域 ▶ 正在看
54:36 基准之外:测试套件、审计与破解 ▶ 正在看
1:01:01 Bender 规则:写明你研究的语言 ▶ 正在看
1:03:45 如果 NLP 不从英语起步 ▶ 正在看
1:10:03 跨语言迁移与低资源社群 ▶ 正在看
本期论点
本期回应
29:42
词作为一串字符与词作为形式和意义的配对是两回事,而语言模型只接触得到字符 不能光靠文字,能学会语言的意思吗?
33:55
语音助手在设闹钟、开灯这类受限任务上是有理解的,因为其训练把语言映射到了语言之外的东西 不能光靠文字,能学会语言的意思吗?
42:03
只接触语言形式的系统学不到语言系统本身,只能学到符号的分布规律 不能光靠文字,能学会语言的意思吗?
42:23
母语习得的关键是共同注意:照料者顺着孩子的注意力给出对应的词 不能光靠文字,能学会语言的意思吗?
14:23
NLP 系统在投放使用前应像其他工程领域那样完成容差、适用范围与认证测试 看它怎么错怎么判断机器是不是真会一件事?
52:08
因为某个系统在 GLUE 上超过部分人类得分就宣称计算机比人更懂英语,是对基准用途的滥用 看它怎么错怎么判断机器是不是真会一件事?
55:10
评估系统应当用专门编排的测试套件覆盖特定情形,而不是靠随机抽样的测试数据 看它怎么错怎么判断机器是不是真会一件事?
16:00
若通往成功的主要路径是把模型越做越大,数据稀缺的语言社群和小团队就会被排除在外 会集中AI 的能力会不会集中到少数机构手里?
20:14
训练数据不是世界本来的样子,只是某一批特定收集方式下得到的数据 会卡住光靠现有的数据,能学到真实的世界吗?
32:22
人太容易从语言里读出意义并替它补上说得通的背景,因此不适合充当图灵测试的测试者 人在补全机器能真的理解吗?
其他论点
25:18
去偏技术有一个前提局限:必须先知道自己要找的是哪一种偏见
25:39
去偏是否够好,应当按具体使用场景和最可能受害的利益相关方来测试 做法
1:02:32
英语之外语言的研究被视为语言特定而更难发表,英语研究却被默认通用,这既损害科学也损害技术
1:10:51
跨语言迁移会把英语中编码的世界观强加给其他语言,这种风险必须被明确权衡
1:12:09
为低资源语言开发技术前应与该语言社群建立联系,由社群参与决定技术走向 做法
01开场:字符序列不等于词
0:01
Emily: It's really important to distinguish between the word as a sequence of characters as opposed to words in the sense of a pairing of form and meaning. Because what the language model is seeing is only the sequence of characters. It's a bit easier to imagine what that's like if you think about a language you don't speak. Lukas: You're listening to Gradient Dissent, a show about machine learning in the real world. And I'm your host, Lukas Biewald. Today, I'm talking to Emily Bender, who is a Professor of Linguistics at the University of Washington, who has a really wide range of interests in linguistics and NLP, from societal issues to multilingual variation to essentially philosophy of linguistics.
Emily:区分两种「词」非常重要——一种是作为字符序列的词,另一种是形式与意义配对意义上的词。因为语言模型看到的只是字符序列。如果你想想一门自己不会说的语言,就比较容易想象那是什么感觉。Lukas:您正在收听的是 Gradient Dissent,一档关于机器学习在现实世界中应用的节目。我是主持人 Lukas Biewald。今天我要和 Emily Bender 聊聊,她是华盛顿大学的语言学教授,在语言学和 NLP 领域涉猎极广,从社会议题到多语言差异,再到本质上属于语言学哲学的问题。
便签引用
02随机鹦鹉论文的成稿与谷歌风波
0:42
I'm especially excited to talk to her because she was actually my teacher for Linguistics 1 at Stanford University, where I was an undergrad. It was one of my favorite classes. I still remember it. I still remember a whole bunch of interesting facts that I learned. And it led to this lifelong interest in linguistics that I've really enjoyed. So, could not be more excited to have a conversation with her. I thought it might make sense to start with the paper that you coauthored, "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? "¹, which was notable even to me on Twitter for a lot of controversy at Google, which I was hoping you could maybe start by describing, but then get into the meat of what the paper actually says. Emily: Yeah. So, it's not in the IPA and hard to pronounce, but the title actually includes an emoji. The last character of the title is a parrot emoji. We were doing that just kind of for fun, because we liked the stochastic parrots metaphor, and there was a while before all this happened that we thought the thing about
我格外期待和她对谈,因为她其实是我在斯坦福读本科时《语言学 1》的老师。那是我最喜欢的课程之一,我到现在都还记得。我还记得当时学到的一大堆有趣的知识。也正是那门课让我对语言学产生了终身的兴趣,我从中收获了很多乐趣。所以,能和她聊天,我真是再兴奋不过了。我想,也许可以从您合著的那篇论文说起,《论随机鹦鹉的危险:语言模型会不会太大?》¹。这篇论文在推特上连我都注意到了,因为它在谷歌引发了很多争议。我希望您可以先讲讲这件事,然后再深入谈谈论文本身到底说了什么。Emily:好的。它不在国际音标里,也很难念出来,但这个标题其实还包含一个表情符号。标题的最后一个字符是一只鹦鹉的表情。我们那样做纯粹是觉得好玩,因为我们喜欢「随机鹦鹉」这个比喻,在这一切发生之前有一段时间,我们还以为这篇论文最大的特点会是「标题里带表情符号的那一篇」。我们当时哪里想得到后来的事。
便签引用
1:40
this paper would be it was the one with an emoji in the title. Little did we know. The paper came about because of work that Dr. Timnit Gebru and Dr. Margaret Mitchell and their team were doing at Google, really trying to connect with the engineering teams to build in good practices to make the technology work better for more people and do less harm in the world. So, that was sort of the role that they had there. They noticed, especially Dr. Gebru, that there was this big push towards bigger and bigger language models. The paper has this table that just...the number of parameters and the size of the training data just explodes over the past couple of years.
这篇论文的缘起,是 Timnit Gebru 博士和 Margaret Mitchell 博士以及她们的团队当时在谷歌所做的工作,她们真的很努力地想和工程团队建立联系,把好的实践嵌进去,让这项技术能为更多人带来更好的效果,也在世界上造成更少的伤害。所以,那大致就是她们在那里担任的角色。她们注意到,尤其是 Gebru 博士注意到,业界有一股把语言模型越做越大的强烈风潮。那篇论文里有个表格,就是……参数量和训练数据的规模在过去几年里就是爆炸式增长。
便签引用
2:27
Dr. Gebru actually direct messaged me on Twitter saying, "Hey, do you know of any papers that talk about the possible downsides to this, any risks? Or have you written anything?" And I wrote back and I said, "No, I don't know of any such papers and I haven't written one. But off the top of my head, here's five or six things that we can be worried about." About a day later, I said, "You know what, that feels like a paper outline. So, here's a paper outline, you want to write this?" That was early September, and the conference we decided to target was FAccT, the Fairness Accountability and Transparency Conference, which took place finally in March 2021. Submission deadline was October, I think, 8th of 2020.
Gebru 博士其实是在 Twitter 上私信我,说:“嘿,你知道有哪些论文讨论过这件事可能的负面影响、有什么风险吗?或者你自己写过什么吗?”我回复说:“没有,我不知道有这样的论文,我自己也没写过。不过我随口想到的,这里有五六件我们可以担心的事。”大约一天之后,我说:“你知道吗,这感觉就是一份论文提纲。所以,这是一份论文提纲,你想写吗?”那是九月初,我们决定投的会议是 FAccT,也就是公平、问责与透明会议,最终在 2021 年 3 月举办。投稿截止日期我记得是 2020 年 10 月 8 日。
便签引用
3:07
So, in a month, we put together this paper. That was possible because it actually wasn't just the two of us writing it or the four named authors finally, but in fact we had seven authors. So, Dr. Gebru brought in Dr. Mitchell — it's really important to me to emphasize that they have doctorates, but I also know them well enough that I'm going to start full naming them now, or first naming them actually — so, Timnit brought in Meg and three other members of their team, and I brought in my PhD student, Angelina McMillan-Major.
所以,我们在一个月里把这篇论文赶了出来。这之所以可能,是因为其实不只是我们两个人在写,也不是最后署名的那四位作者,实际上我们有七位作者。所以,Gebru 博士拉进了 Mitchell 博士——强调她们都有博士学位对我来说真的很重要,但我和她们也熟到可以开始连名带姓地称呼她们了,其实是直接叫名字——所以,Timnit 拉进了 Meg 和她们团队的另外三位成员,我则拉进了我的博士生 Angelina McMillan-Major。
便签引用
3:38
Between the seven of us, we sort of had enough different areas of expertise and literatures that we've read that we could pull together this survey paper. And so, it came together. It was amazing, and also an interesting writing experience because we never had a Zoom meeting or anything where all of us spoke together. It was all done through remote collaboration in Overleaf. So, not a super common way for a research to get done, but it worked in this case. The Google authors put it through what they call pub approve over there, got approved. We submitted it to the conference and then put it away, because none of us had actually anticipated working on that in the month of September. It was like extra work for everybody. so we all turned back to the other stuff we needed to be doing. In late November out of nowhere, from my perspective — and I should say that in telling the story, I'm not at Google, I've not been funded by Google, and so I only have sort of secondhand understanding of what went on at Google plus what was out in the press eventually — but the Google coauthors were
我们七个人加在一起,在专业领域和读过的文献上差不多够广,能够合力做出这篇综述性论文。于是它就成型了。这很了不起,同时也是一次有意思的写作体验,因为我们从来没开过 Zoom 会议之类的、所有人一起讨论的场合。全程都是在 Overleaf 上远程协作完成的。所以,这不是一种特别常见的做研究的方式,但在这个情况下奏效了。那几位谷歌的作者把它送去走了他们那边所谓的 pub approve(发表审批)流程,通过了。我们把它投给了会议,然后就搁在一边了,因为我们当中谁都没有真的预料到还要继续为这件事忙下去就在九月那个月。这对所有人来说都是额外的工作。所以我们都回去做各自本来该做的事了。到了十一月底,突然之间,从我的角度看——我应该说明,讲这个故事时,我并不在谷歌,我也没有拿过谷歌的资助,所以我对谷歌内部发生的事只有间接的了解,再加上后来媒体上披露的那些——但谷歌方面的合著者被
便签引用
4:40
told to either retract their paper or take their names off of it, and they weren't told why. They weren't offered a chance to discuss what might need to be changed about the paper, it was just "retract it or take your names off of it". We had this strange moment of, "Okay, what do we do with this paper?", because it seems kind of odd to put something out with just two authors that actually represents the work of seven people, what do we want to do here? My PhD student Angie and I, we turned to the Google coauthors and we said, "We will follow your lead here. What do you want to have happen?" And they said, "No, we want this out in the world. So, you two publish it."
要求要么撤回论文,要么把名字从上面拿掉,而且没人告诉他们为什么。也没有给他们机会讨论这篇论文可能需要改哪些地方,就只是“撤回,或者把名字拿掉”。我们当时有个很奇怪的时刻:“好吧,这篇论文怎么办?”因为只署两个人的名字有点奇怪,毕竟这实际上代表了七个人的工作,我们到底要怎么处理?我和我的博士生 Angie,我们去问谷歌那边的合著者,我们说:“这件事我们听你们的。你们希望怎么处理?”他们说:“不,我们希望它面世。所以,你们俩去发表。”
便签引用
5:16
That was the initial answer. And then Timnit, sort of on reflection, said, "Actually, this is not okay. This is not an okay way to treat a researcher who was hired to do this research." This was literally her job and the job of everyone on that team. And so, she pushed back. The result of all that, you can go find in all the media coverage, is that she got fired. Google claimed she resigned. Her team says she got resignated, which is a great neologism. That went down fast enough that she was able to then put her name on the paper.
这是最初的答复。然后 Timnit 回过头想了想,说:“其实这不对。这不是对待一个被雇来做这项研究的研究者的正当方式。”这实实在在就是她的工作,也是那个团队每个人的工作。于是,她提出了反抗。这一切的结果,你在所有的媒体报道里都能找到——她被解雇了。谷歌声称她是主动辞职的。她的团队说她是“被辞职”的,这个生造词造得真妙。这件事发展得够快,所以她后来还来得及把名字署到论文上。
便签引用
5:49
Meanwhile, Meg started working on documenting what had happened to Timnit. The end result of that was that she was fired a few months later, but after the final version of the paper was done. So, that's why the fourth author is Shmargaret Shmitchell.
与此同时,Meg 开始着手记录 Timnit 遭遇的一切。最终的结果是她在几个月后也被解雇了,不过那时论文的终稿已经完成了。所以,第四作者才会是 Shmargaret Shmitchell。
便签引用
6:07
That's a really sad story for everybody involved. I mean, it's terrible mistreatment of Timnit and Meg and the other members of the team, those who are on our paper and those who weren't. It's become, I think, a really difficult environment to work in. It's sad for Google because they lost really wonderful expertise and a lot of goodwill in the research community. And sort of sad for...it sheds a light on the sad state of affairs about the way corporate interests are influencing what's happening in research in our field right now.
对所有卷入其中的人来说,这都是个非常令人难过的故事。我是说,这是对 Timnit 和 Meg 以及团队其他成员——包括署名在我们论文上的和没署名的——极其恶劣的对待。我想那里已经变成了一个很难工作下去的环境。对谷歌来说也很可惜,因为他们失去了非常出色的专业能力,也失去了研究界的大量善意。还有点让人难过的是……它照见了一个令人沮丧的现状:企业利益正在如何影响我们这个领域当下的研究走向。
便签引用
6:37
On the other hand, my coauthors and I still maintain...we all really enjoyed the experience of working on this paper together and of weathering the stuff backwards together. One weird result is that this paper has gotten way more attention than it ordinarily would have. I think it's a good paper, it's a solid paper. And boy, did we put a lot of polish on it between the submission version and the camera-ready because we knew it was going to be read by a lot of people. When I put up the camera-ready as a preprint...I didn't put it on arXiv, because those tend to get cited instead of the final published versions. So, I just put it on my website and tweeted out a link with a bitly link to shorten it, so that I could see how many times it was downloaded.
另一方面,我和合著者们依然觉得……我们都真心享受一起做这篇论文、一起扛过后来那些风波的经历。一个奇怪的结果是,这篇论文得到的关注远远超出了它本来会有的程度。我认为这是一篇好论文,一篇扎实的论文。而且天哪,我们在投稿版和最终定稿之间打磨了非常多,因为我们知道会有很多人读它。我把最终定稿作为预印本放出来的时候……我没有放到 arXiv 上,因为那种版本往往会被引用,取代最终发表的版本。所以我就放在了自己的网站上,然后发推附了个链接,用 bitly 做了个短链接,这样我就能看到它被下载了多少次。
便签引用
7:22
It has been downloaded through that link alone over 10,000 times. I know that there's other ways to get to it, which is way out of scale to anything that I've ever written otherwise. So, that's been interesting as a researcher, but it's also I think fortunate because it has come to the attention of the public, and I think that this technology is widespread. It's being used. It's being used in lots of different ways. It's really valuable that the public at large has a chance to understand what's going on. And so, through Google's gross misstep, I and my coauthors have been given the chance to help educate the public, which is something that I do feel fortunate about.
光是通过那一个链接,它就被下载了超过一万次。我知道还有别的途径能拿到它,这个量级远远超出我写过的任何其他东西。所以,作为一个研究者,这挺有意思的,但我也觉得这算是幸运,因为它引起了公众的注意,而我认为这项技术已经很普及了。它正在被使用。它正在以许多不同的方式被使用。让广大公众有机会理解正在发生什么,这非常有价值。所以,因着谷歌这次严重的失策,我和我的合著者获得了帮助公众了解这些的机会,这一点我确实觉得挺幸运的。
便签引用
03论文到底说了什么惹到谁
8:10
Lukas: I'd love to kind of get into what the paper talks about. But do you have any sense, or has Google made any comments, about what their objection was? Because I sort of had this feeling that it must be a really incendiary paper, and then in the prep for this interview I actually read it and it felt pretty uncontroversial, I guess, was my feeling reading it. So, I just wonder...I mean, maybe it's hard to know, but have they said anything about what they didn't like about it? Emily: Yes, I mean, in public comments, there's been things like, "It doesn't cite relevant work that is trying to mitigate some of these issues."
Lukas:我很想聊聊这篇论文到底讲了什么。但你有没有什么头绪,或者谷歌有没有就他们的反对意见发表过任何说明?因为我原本以为这一定是一篇极具煽动性的论文,然后为这次访谈做准备时我真的读了它,读下来的感觉是,它其实相当没有争议性。所以我就好奇……我是说,可能很难知道,但他们有没有说过他们不满意的到底是什么?Emily:有的,我是说,在公开表态里,出现过这样的说法:“它没有引用那些试图缓解这些问题的相关工作。”
便签引用
8:47
But at no point were we ever told which work we should have been citing. And we do, in fact, cite some work that is trying to mitigate these issues. So, I don't know quite what that was about. But you're absolutely right. We figured that we'd be ruffling some feathers with this paper because we were basically saying, "Hey, this thing that everyone's having so much fun chasing, maybe let's go a little bit slower and think about what kinds of downsides there are and how to do this safely." There's going to be some who don't want to hear that, but we honestly thought it was going to be OpenAI who was upset, because GPT-3 is kind of the best known example of this, and it was our running example, too. So, we thought we'd ruffle some feathers, did not realize we were going to be ruffling feathers inside Google. And, it's basically a survey paper. We didn't run any experiments. We didn't do any analysis. What we did was we pulled together a bunch of different relevant perspectives on large language models and brought them all together in one place.
但从来没有人告诉我们应该引用哪些工作。而事实上,我们确实引用了一些试图缓解这些问题的工作。所以我也不太清楚那到底是怎么回事。但你说得完全对。我们本来就估计这篇论文会惹到一些人,因为我们基本上是在说:“嘿,这个大家追得这么起劲的东西,也许我们可以放慢一点,想想它有哪些负面影响,以及怎样才能安全地做这件事。”总会有人不爱听这个,但我们老实说以为会不高兴的是 OpenAI,因为 GPT-3 算是这方面最广为人知的例子,也是我们一路举的例子。所以我们以为会惹到一些人,但完全没想到惹到的是谷歌内部。而且,这基本上是一篇综述论文。我们没有做任何实验。我们没有做任何分析。我们做的是把一堆关于大型语言模型的不同相关视角汇集起来,集中放在一处。
便签引用
9:44
It is surprising that the paper seems to have been part of the cause of Google basically blowing up this amazing asset that it had in terms of its ethical AI team. Lukas: Interesting. And I guess, one reading of your paper is, "Hey we should consider the downsides of large language models." I think maybe another person might read it...this might be an unfair reading, but I could imagine someone having hurt feelings if they were working on large language models, and they read your paper saying it's like an unethical thing to do, to build large language models. Would that be an overstatement of your claims?
令人意外的是,这篇论文似乎成了谷歌亲手炸掉自己“伦理 AI 团队”这一惊人资产的原因之一。Lukas:有意思。我猜,对你们论文的一种读法是:“嘿,我们应该考虑大型语言模型的负面影响。”我想也许另一个人读到的会是……这可能是不公平的读法,但我可以想象,如果有人正在做大型语言模型,读你们的论文时觉得这是在说建造大型语言模型是件不道德的事,感情上会受伤。这算是夸大你们的主张吗?
便签引用
10:29
I don't have the paper in front of me, but I think maybe that could hurt feelings, I'm not sure. Emily: I also do a lot of work in the space...societal impact of NLP in general, and that sometimes goes under the title of ethics and NLP. I do see a lot of people reacting to that topic with hurt feelings, and I think it's connected with the way in which people identify with their work. If you say, "Hey, let's think about this technology we're building and how it behaves in the world and what we can do to make it be beneficial," and you use the term ethics to describe that, sometimes people want to read that as, "You're calling me unethical," and I think that that direction of the conversation is rarely actually valuable. I do think that, in general, people in this space want to be doing good things in the world. Certainly, there are people who are working on technology with the goal of making a lot of money doing it. There's this caricature of the tycoon or whoever, who's just happy to crush all the little people to make as much money as
我手头没有论文,但我觉得这也许确实会伤到一些人的感情,我不确定。Emily:我在这个领域也做了很多工作……总体上是 NLP 的社会影响,有时这被归在“伦理与 NLP”这个标题下。我确实看到很多人对这个话题的反应是感情受伤,我认为这和人们把自我认同寄托在工作上的方式有关。如果你说:“嘿,我们来想想我们正在造的这项技术,它在现实世界里如何运作,以及我们能做些什么让它带来益处”,而你用“伦理”这个词来描述它,有时人们就会想把这读成“你在说我不道德”,而我认为对话往那个方向走,实际上很少有价值。我确实认为,总体而言,这个领域里的人是想在世界上做好事的。当然,也有人做技术的目标就是靠它赚一大笔钱。有那种漫画式的形象,某个大亨之类的,乐于碾碎所有小人物来赚尽可能多的钱,这种人大概是存在的。但我认为更常见的情况是,
便签引用
11:40
possible, that's out there, probably. I think much more frequently, people are working within systems that give them certain commitments around maximizing value for shareholders and stuff like that, that make it harder to put on the brakes on some things that are making money right now for shareholders and take a bigger picture view. But it is much more valuable to talk about it in terms of what are those systems, what are the incentives, what can we as individuals do within those systems, rather than think about people as ethical or unethical. I'm not sure that really speaks to your question, but hopefully it's somewhat helpful. Lukas: No, I mean, I think you're saying that your point is a little more nuanced than what maybe someone would take away, and I can...I run a company and I love technology and I love building...I do recognize that lots of people get hurt, and I think it's great that people are pointing out issues and also kind of pumping the brakes and flagging this stuff. But I could kind of see how someone might
人们身处某些体系之中,这些体系给他们施加了诸如“为股东实现价值最大化”之类的约束,这就让他们更难对某些眼下正在为股东赚钱的事踩刹车,更难从大局出发去看问题。但更有价值的做法是去谈论:这些体系是什么,激励机制是什么,我们作为个体在这些体系里能做什么,而不是去想某个人是道德的还是不道德的。我不确定这是否真的回答了你的问题,但希望多少有点帮助。Lukas:不,我是说,我觉得你的意思是,你的观点比某些人可能理解到的要更细致一些,而我……我经营一家公司,我热爱技术,我热爱造东西……我确实认识到有很多人因此受到伤害,我觉得有人指出这些问题、踩踩刹车、把这些事情标记出来,是很好的。但我能理解为什么有人会觉得有点被冒犯。我不确定我是不是跳得太快了——
便签引用
04大模型研究该停吗:工程测试文化
12:44
feel a little offended by it. I wasn't sure if I was kind of jumping to something or- I guess my question, well, the question that I kept thinking about with the whole paper in general as I was reading it is, even sort of setting aside making money, let's just talk about research and just the excitement of building models that work. I just feel that so deeply, like GPT-3 for all its fuss, it's kind of amazing what it does. I wouldn't have expected it to work so well. Would you feel...would you argue that those kinds of directions of research should stop? Or what would you want an organization like OpenAI to do differently? Ethics is a good example of a place that's kind of actually really showed that bigger models do kind of... it's not obvious that bigger models would perform tasks better at many extra orders of magnitude. Would you prefer that that research doesn't happen or happen differently somehow? Emily: So, I think it's worth saying that OpenAI has actually put a lot of effort into thinking about "What are the possible downsides?" and "What
我想问的,嗯,读整篇论文时我一直在想的问题是:哪怕先把赚钱这件事放到一边,我们就只谈研究,谈那种把模型做出效果的兴奋感。我对此感受非常深,比如 GPT-3,抛开所有的喧嚣,它能做到的事挺惊人的。我没想到它能做得这么好。你会不会觉得……你会主张这类研究方向应该停下来吗?还是说,你希望像 OpenAI 这样的机构做些什么不一样的事?伦理是个很好的例子,它其实真切地表明了更大的模型确实会……更大的模型在多出好几个数量级之后能把任务做得更好,这并不是显而易见的。你是希望这类研究别做,还是希望以某种别的方式来做?Emily:我觉得值得说明的是,OpenAI其实在“可能有哪些负面影响”和“这项技术投放到世界上会发生什么”这些问题上投入了很多思考,这一点值得指出,
便签引用
13:59
could happen when this technology is released in the world?", and that's important to note, and I'm glad that they're doing that. I think that what I would like to see more of is...first of all, that kind of work. What are the possible failure modes and how do they impact people? And then also, when this is working as intended, how can that impact people? OpenAI has been doing some of that and I think that's great, and they should do more. But also, you can look to other fields of engineering, where before you take something and you put it into the world in a place where people are going to rely on it, there's all kinds of testing that has to be done in sort of understanding of "What are the tolerances?"
我也很高兴他们在这么做。我想我希望看到更多的是……首先,就是那类工作。可能的失效模式有哪些,它们如何影响人?然后还有,当它按预期运作时,那又会如何影响人?OpenAI 已经做了一些这样的工作,我觉得很好,他们应该做更多。但另外,你也可以看看工程学的其他领域:在你把某样东西投放到世界上、让人们去依赖它之前,有各种各样的测试必须完成,要弄清楚“容差是多少”,
便签引用
14:37
and "What works and what doesn't?" and "What's the range of temperatures that this thing could be applicable in?" and "What are the things you have to check for and certify?" and things like that. We don't have very much of that yet going on in NLP. I can speak less to other areas of AI, but I honestly think there's similar issues elsewhere in AI. And so, there's work — actually that was done at Google by Meg Mitchell and Timnit Gebru and others on a framework called Model Cards, which was sort of steps in that direction of like, "You've built a model, what does somebody who's going to use this model need to know about it?" — and that's the kind of thing that I would like to see more of. And that is in contrast to just rampant AI hype, where people build something, it's cool, it's fun, it works well, and somehow, that's not enough.
“什么情况下有效、什么情况下无效”,“这东西适用的温度范围是多少”,“有哪些东西你必须检查和认证”,诸如此类。在 NLP 里,我们这方面做得还非常少。对 AI 的其他领域我没那么有发言权,但我老实说觉得 AI 的其他地方也存在类似的问题。所以,有一项工作——其实就是在谷歌由 Meg Mitchell、Timnit Gebru 等人做的——一个叫“模型卡片”(Model Cards)的框架,就是朝那个方向迈的一步:“你造了一个模型,那么将要使用这个模型的人需要了解它的哪些信息?”——这就是我希望看到更多的那类东西。而这与漫无边际的 AI 炒作形成对比:人们造出个什么东西,它很酷、很好玩、效果很好,可不知怎的,这还不够。
便签引用
15:23
People have to say...it's not enough that GPT-3 can produce coherent text, people have to say it's understanding language, which it absolutely isn't, as I'm sure we'll talk about later. Lukas: You have two good segues, but yeah, yeah. Emily: Yeah. Although it is all connected, right? Lukas: Yeah. Emily: So, for some reason, the culture around AI is all about trying to reach for these big claims rather than trying to build really well-scoped, reliable — sufficiently documented that they can be used safely and reliably — systems. That's the direction that I would like to see more of, is one thing. And then another thing, and we get into this in the paper, is that if the main pathway to success these days is just bigger and bigger and bigger, then you cut out lots of languages communities, even within the languages that generally are well supported, because they just can't amass that much data. And you also cut out smaller research groups, smaller companies that are not sitting on the kind of
人们非得说……GPT-3 能生成连贯的文本还不够,人们非得说它理解了语言,而它绝对没有,这个我们后面肯定会聊到。Lukas:你抛了两个很好的话头,不过,是的,是的。Emily:是啊。不过这些其实都是连在一起的,对吧?Lukas:对。Emily:所以,不知为何,围绕 AI 的这种文化就是一味去够那些宏大的说法,而不是去构建界定良好、可靠的——文档充分到可以被安全可靠地使用的——系统。这就是我希望多看到的方向,这是一点。然后另一点,我们在论文里也谈到了,就是如果如今通往成功的主要路径就是越来越大、越来越大,那你就把很多语言社群排除在外了,哪怕是在那些总体上被支持得不错的语言内部也是如此,因为他们根本攒不出那么多数据。而且你还排除了更小的研究团队、更小的公司,他们手里没有谷歌、Facebook 或亚马逊那样的数据积累。微软也做了不少
便签引用
05偏见的四个真实案例
16:24
collections of data that Google is, or Facebook is, or Amazon is. Microsoft also does a bunch of big data work, they don't seem to have amassed data quite the same way as the other big ones. That is unfortunate because it, I think, stifles creativity to a certain extent. If the whole community is rushing towards this one goal that only some can really effectively do, then we lose out on the other things that people might be trying instead. Lukas: Maybe a less obvious concern that you talked about in the paper is talking about how the models can encode bias in ways that are hard to notice. When you talk about the harms that might happen from natural language models, do you have examples of things that are actually happening now? Or is this more of like a future-looking thing that we're worried about as NLP becomes more pervasive, like worrying about future harms? Emily: So, I mean, absolutely happening now, and therefore easy to predict that it will keep happening in the future if we don't change.
大数据方面的工作,但他们似乎没有像其他几家巨头那样攒下同样规模的数据。这很遗憾,因为我认为它在某种程度上扼杀了创造力。如果整个社群都朝着这一个只有少数人才真正做得了的目标狂奔,那我们就失去了人们本来可能去尝试的其他东西。Lukas:你在论文里谈到的一个也许不那么显而易见的担忧,是模型会以难以察觉的方式编码偏见。当你谈到自然语言模型可能造成的危害时,你有没有一些现在就真实发生着的例子?还是说这更多是一种面向未来的担忧,担心随着 NLP 变得更加无处不在而带来的未来危害?Emily:不,绝对是正在发生的,因此也很容易预测:如果我们不改变,它还会继续发生。
便签引用
17:27
Here, the work of Safiya Noble, with her book "Algorithms of Oppression", is a really important documentation of this. She looked into what are the ways in which identities — which properly belong to the groups of people who have those identities — are represented and reflected back to people in search. In particular, her running example is the phrase "black girls" and also "black women". These things have changed over time — and she's very careful to document when she talks about particular examples what the date was — but early on, as she started this project, the phrase "black girls" as a search keyword basically turned up pornography. And that you might say is, "Well, that's just in the data." Well, what data? Where did that data come from?
这方面,Safiya Noble 的工作,她那本《算法压迫》(Algorithms of Oppression),是非常重要的记录。她研究了这样一个问题:身份——这些身份本该属于拥有它们的那群人——在搜索中是如何被表征、又如何被反射回人们眼前的。特别是,她一路使用的例子是“black girls”这个短语,还有“black women”。这些情况随时间发生了变化——她非常谨慎,谈到具体例子时都会记录下日期——但在她刚开始做这个项目的早期,“black girls”作为搜索关键词,出来的基本上都是色情内容。对此你可能会说:“嗯,数据里就是这样。”那么,什么数据?这数据从哪来的?
便签引用
18:17
If you get into the heart of her book, it's basically around that that's "in the data" because of the way in which the economy of the internet allows people to purchase and make money off of identity terms. Once these things were flagged, Google sort of piecemeal [made] changes, so you don't get pornography as the results for the search term "black girls" anymore. But it's also possible to sort of poke at things and tell that it's very much individual after-the-fact changes, as opposed to anyone going through and systematically thinking about how to redesign the way that search engines and the advertising-driven ranking of search latches on to these incentives and then amplifies them. One ongoing discussion in the AI community — you see it pop up on Twitter with great regularity is — is the problem that the data is biased only, or do the models also contribute? And the answer is absolutely models also contribute.
如果你深入她书的核心,基本上就是说,之所以“数据里就是这样”,是因为互联网的经济模式允许人们购买身份词汇并靠它赚钱。这些问题一被指出来,谷歌就零敲碎打地做了些改动,所以现在搜“black girls”不会再出来色情内容了。但你也可以去戳一戳,就能看出这很大程度上是一次次事后的个别修补,而不是有人系统性地去思考如何重新设计搜索引擎、以及广告驱动的搜索排序如何攀附上这些激励机制并把它们放大。AI 社群里有一场持续的讨论——你会看到它在推特上极有规律地冒出来——就是:问题只在于数据有偏见,还是模型也有份?答案是:模型绝对也有份。
便签引用
19:22
And then, there's other layer to it of, "Well, that's just what's in the data." One of the other really embarrassing examples for Google was, there was a point at which Google Image search turned up pictures of gorillas when you were searching for black people, and I forget exactly the particular configuration of that, but embarrassing and awful and racist. One reaction at the time was, "Well, that's just in the underlying data." And so, "Not our fault. We're just showing what the world is saying.", except that it's not true, because the way the algorithms that do the ranking of search results and also the bidding for the ad words is...that is emphasizing particular incentives. So, there is a certain thing in the underlying data. There's also the question of how did you collect that data? Where did it come from? What does it actually represent? It is not the world as it is. It is some particular collection of data.
然后还有另一层:“嗯,数据里就是这样啊。”谷歌另一个非常难堪的例子是,曾经有一段时间,你在谷歌图片里搜索黑人,出来的是大猩猩的照片,具体是怎么个情形我记不清了,但既难堪又糟糕,而且是种族主义的。当时的一种反应是:“嗯,底层数据里就是这样。”所以,“不是我们的错。我们只是把世界在说的东西呈现出来。”只不过这并不属实,因为那些给搜索结果排序、以及为广告关键词竞价的算法的运作方式……那是在强化某些特定的激励。所以,底层数据里确实有某种东西。但也要问:你是怎么收集这些数据的?它从哪来?它实际代表了什么?它并不是世界本来的样子。它是某一批特定的数据。
便签引用
20:17
And then what is the optimization metric? What are all these modeling decisions that you've made? And how does that interact with the various biases in the data? And what's the incentive structure? Safiya Noble's work is a great point to look. Latanya Sweeney documented — this is a 2013 paper — how if you put in, at that point, an African-American sounding name, one of the ads that would pop up suggested that that person had a criminal history. And if you put in a white sounding name, you tended to get just a more information about so-and-so.
然后,优化指标是什么?你做出的这些建模决策都是什么?它们又如何与数据中的各种偏见相互作用?以及激励结构是什么?Safiya Noble 的工作是个很好的切入点。Latanya Sweeney 记录过——这是 2013 年的一篇论文——当时如果你输入一个听起来像非裔美国人的名字,弹出的广告之一会暗示这个人有犯罪记录。而如果你输入一个听起来像白人的名字,你往往只会得到关于某某人的更多信息。
便签引用
20:48
And that does real harm in the world — it wasn't 100%, but it was significantly different between the two groups of names — it does real harm in the world because if you imagine someone is applying for a job or just making friends and someone does a Google search on them, and here comes alongside this message suggesting they might be a criminal, that does harm. And then if I can give one more example— Lukas: Please, yeah, these are great. Emily: Elia Robyn Speer did a really interesting work example around sentiment analysis and word embeddings.
这在现实世界里造成了真实的伤害——不是百分之百都这样,但两组名字之间的差异是显著的——它在现实中造成真实伤害,因为你想象一下,有人在求职,或者只是在结交朋友,而对方在谷歌上搜了他一下,结果旁边就冒出这样一条信息,暗示他可能是个罪犯,这就造成了伤害。然后,如果我可以再举一个例子——Lukas:请讲,这些例子都很棒。Emily:Elia Robyn Speer 做过一个很有意思的工作实例,是关于情感分析和词嵌入的。
便签引用
21:18
Sentiment analysis is the task of taking some natural language text, and her example is English, and using it to calculate or predict the sentiment. Is this a text expressing positive feeling towards something, negative feeling towards something, or not expressing feelings? The particular data set she was working with, I think, was Yelp restaurant reviews. So there, it's "Take the text, predict the stars." Lukas: Yeah, I've used that data set. Yeah, for sure. Emily: And then as an external component, she's using word embeddings, which are representations of words into a vector space based on what other words they could occur with. So, some of the training data is in-domain, the Yelp reviews, but then there's this component that's trained on general web garbage.
情感分析的任务是拿一段自然语言文本——她的例子是英语——用它来计算或预测情感倾向。这段文本是在表达对某事物的正面情绪、负面情绪,还是没有表达情绪?她当时用的那个具体数据集,我记得是 Yelp 的餐厅评论。所以在那里就是“拿到文本,预测星级”。Lukas:是的,我用过那个数据集。当然用过。Emily:然后作为一个外部组件,她用了词嵌入,也就是把词根据它们可能与哪些词共现,表示到一个向量空间里。所以一部分训练数据是领域内的,也就是 Yelp 评论,但还有一个组件是在通用的网络垃圾数据上训练出来的。
便签引用
22:02
What she found using the sort of generic word embeddings was that the system systematically un-predicted the star ratings for Mexican restaurants. All right, so she digs into it and looks into why. It turns out that because that general web garbage included the discourse about immigration into the US from and through Mexico, which has lots of really negative toxic opinions of Mexican people, the word embeddings picked up the word "Mexican" as akin to other negative sentiment words. And so, if in your review of the restaurant you called it a "Mexican restaurant", according to the system, you have said something negative about it, so you can't possibly be giving it a five-star review.
她发现,用那种通用词嵌入时,系统会系统性地低估墨西哥餐厅的星级预测。好,于是她深入去查原因。结果发现,因为那些通用网络垃圾数据里包含了关于从墨西哥、经墨西哥进入美国的移民的讨论,而这些讨论里有大量非常负面、有毒的、针对墨西哥人的观点,于是词嵌入就把"Mexican"(墨西哥的)这个词学成了和其他负面情绪词很相近。所以,如果你在你的餐厅评论里把它称作"墨西哥餐厅",在系统看来,你就是说了它的坏话,那你就不可能是在给它打五星好评。
便签引用
22:44
Lukas: Well, that's a really interesting example. My next question was going to be how do models play into this? I guess that's a good example of how not just the underlying data can have bias, but the model can literally have its own bias. Emily: Yeah, so the word embedding picked up on co-occurrences between the word "Mexican" and lots of other things that also co-occurred with negative sentiment, and then that was used as a component in this other model. So, yeah, there wasn't in the underlying Yelp reviews any particular reason that the Mexican restaurants were rated lower, right? Lukas: Right.
Lukas:这个例子真的很有意思。我下一个问题本来就想问,模型在这里面起了什么作用。我想这就是个很好的例子,说明不只是底层数据会有偏见,模型本身也真的会带上自己的偏见。Emily:对,词嵌入捕捉到的是"Mexican"这个词和一大堆其他词的共现,而那些词本身又和负面情绪共现,然后这个词嵌入又被当成另一个模型的一个组成部分用了进去。所以说,底层的 Yelp 评论里其实并没有什么特别的理由说墨西哥餐厅评分更低,对吧?Lukas:对。
便签引用
23:23
Emily: I don't know for sure if they were rated on average exactly the same, but it doesn't matter, because the error was the system underpredicting for any given restaurant. On average, it was missing in the low direction. So, yeah, that's a kind of bias that was picked up from an external data set. We tend in NLP to use word embeddings as really handy detailed representations of word "meaning", so word similarity, including semantic similarity. And if we don't pay attention to what meaning was picked up, what co-occurrence was picked up, then we can end up with stuff we really don't want in our systems.
Emily:我不能百分之百确定它们的平均评分是不是完全一样,但这不重要,因为问题出在系统对任何一家餐厅的预测都偏低。平均来看,它的误差是往低了偏。所以说,这是一种从外部数据集里被带进来的偏见。我们在 NLP 里很习惯把词嵌入当成非常方便、细致的词"义"表示,用来算词的相似度,包括语义相似度。而如果我们不去关注它到底学到了什么含义、学到了什么共现,最后就会在系统里留下我们其实很不想要的东西。
便签引用
06去偏:没有无偏数据,但能做得更好
24:06
Lukas: What would you recommend doing about that? Because they are really useful, word embeddings. And I'm sure in this case, it seems pretty simple. It's hurting your performance. There's not even a model performance tradeoff here. So, what could you possibly do? Emily: There is a lot of work on so-called debiasing of word embeddings. If you look at Speer's work, she continues on to do some of that. And I think that part of it is, work with more curated datasets. The discourse around immigration from and through Mexico, even if you stick with only things like reputable news sources, you're still going to find that garbage. That alone is not going to solve it, but it can be better.
Lukas:那你会建议怎么处理?毕竟词嵌入确实非常有用。而且我猜在这个例子里,事情看起来挺简单的——它是在损害你的性能,这里甚至都不存在模型性能上的取舍。所以,能做些什么呢?Emily:关于所谓的词嵌入"去偏"(debiasing),已经有很多工作。你去看 Speer 的研究,她后来也做了一些这方面的事。我觉得其中一部分是,用更精挑细选的数据集。关于经由墨西哥的移民的讨论,哪怕你只用所谓正规的新闻来源,你还是会碰到那些垃圾内容。光靠这一点解决不了问题,但确实可以更好一些。
便签引用
25:03
It's not possible to come up with a fully bias-free dataset nor fully bias-free word embeddings, but you can do better. One step is to sort of say, "Okay, how much better can we do with curated data? What about debiasing techniques for the biases that we're aware of?" Part of the problems with debiasing techniques is that you have to know what you're looking for. And then on top of that, to think through failure modes. So, in a particular use case, when you're building some technology, who are the stakeholders? Who's going to be impacted by it? If someone's restaurant rating is underpredicted for some reason, what does that mean in an actual use context? And what should we be testing for to see if we have sufficiently debiased for our use case, for the stakeholders who are most likely to experience adverse impacts? Lukas: I guess, it does seem like it would be incredibly...I mean, it seems like it'd actually be impossible to find an unbiased dataset of human... Emily: Right. It doesn't exist.
要做出一个完全没有偏见的数据集、或者完全没有偏见的词嵌入,是不可能的,但你可以做得更好。第一步大概是问:"好,用精选数据我们能改善到什么程度?对于我们已经意识到的那些偏见,有没有去偏的技术可以用?"去偏技术的问题之一在于,你必须先知道自己要找的是什么。然后在这之上,还要把失效模式想清楚。也就是说,在一个具体的使用场景里,当你在造某个技术的时候,利益相关方是谁?谁会被它影响?如果某家餐厅的评分因为某种原因被低估了,在真实的使用场景里这意味着什么?我们又该测试什么,才能判断针对我们的使用场景、针对那些最可能受到不利影响的利益相关方,我们的去偏做得够不够?Lukas:我觉得,这确实看起来会极其……我是说,看起来其实根本不可能找到一个没有偏见的、关于人的数据集……Emily:对。那种东西不存在。
便签引用
07形式与意义:语言模型为何不理解
26:01
Lukas: I guess these are good segues into other papers that I want to talk about. So, maybe we should just in the interest of time, we should move on to the second paper² that we want to talk about to make sure we get to it, which is around... Let me see if I can summarize this. So, this is basically sort of saying that language modeling only on what you call "form" — which I think is just sort of like the words coming through, this is kind of of the GPT-3 types of models that just like look at these strings of words — can't have understanding, like true understanding.
Lukas:我觉得这正好可以引到我想聊的另外几篇论文。所以,为了赶时间,也许我们该直接进入我们想聊的第二篇论文²,免得来不及,它讲的大概是……我试试能不能概括一下。基本上它是在说,只在你所谓的"形式"(form)上做语言建模——我理解就是流过来的那些词,也就是 GPT-3 这一类只看词串的模型——是不可能有理解的,不可能有真正的理解。
便签引用
26:34
I just thought one thing that was interesting is that you said you wrote the paper to sort of end some kind of debate on Twitter that I was definitely not aware of. Actually, I think I'm kind of coming into something with maybe more context than I knew. So, maybe you can sort of summarize what the different possible positions are here and what you want to put to rest. Emily: So, I kept finding myself getting into arguments on Twitter with people who were claiming that language models were understanding things.
我觉得有意思的一点是,你说你写这篇论文是为了给 Twitter 上的某场争论画个句号,而那场争论我之前完全不知道。其实,我觉得我可能是带着比自己以为的更多的背景走进这个话题的。所以,也许你可以大致总结一下,这里可能存在哪几种立场,以及你想要终结掉的是哪一种。Emily:是这样,我总是发现自己在推特上跟人吵架,那些人声称语言模型是在理解事物。
便签引用
26:59
And I was like, "No, they're not. They can't possibly be." It's important to pin down what we mean by language models. So, a language model is something like GPT-3 or BERT or otherwise, where its training data is a whole bunch of text, and the training task is predicting words in the text. So, some of the times it's done sequentially, sometimes it's done with a masked language model objective, where certain words are dropped out and the training objective is, "Okay, well put those words back in and then do your model updating to...", gradient descent, et cetera, et cetera. For me, as a linguist, I look at that and go, "Hey, useful technology, interesting. Incredibly helpful in things like speech recognition and machine translation where an important subtask is, 'Okay, what's the likely string?'"
我就说:「不,它们没有。它们根本不可能在理解。」我们必须先说清楚「语言模型」指的是什么。语言模型就是像 GPT-3 或 BERT 这类东西,它的训练数据是一大堆文本,训练任务是预测文本中的词。有时候是按顺序预测,有时候用的是掩码语言模型的目标函数,也就是把某些词遮掉,训练目标是「好,把这些词填回去,然后更新你的模型……」,梯度下降等等等等。作为一个语言学家,我看着这些会想:「嘿,挺有用的技术,挺有意思。在语音识别、机器翻译这类任务里非常有帮助,因为那里有个重要的子任务是:『好,哪个字符串更可能?』」
便签引用
27:46
So, in a speech recognition setup, the acoustic model says, "Here's a range of text strings that sound might have corresponded to," and then the language model comes in and says, "Okay, yeah, but 'It's important to wreck a nice beach.' is a ridiculous thing to say, and 'It's important to recognize speech.' is a reasonable thing to say, so we're going to rank that one higher." That's the kind of form-based tasks that they were initially meant for and good at. And then what's happened with the neural language modeling revolution in the past few years is that when you extract the word embeddings from a language model, you have a really finely fitted representation of word distribution, which is very useful, and some of them can even do...where you get the word embeddings are contextual. So, the information about the word and what it's likely to co-occur with isn't about that word across all the texts, but about that word in its current context.
比如在语音识别的设置里,声学模型说:「这段声音可能对应这一批文本字符串」,然后语言模型登场说:「好,是这样,但『It's important to wreck a nice beach.(毁掉一片美丽沙滩很重要)』是一句荒唐的话,而『It's important to recognize speech.(识别语音很重要)』是一句合理的话,所以我们把后者排得更高。」这就是它们最初被设计来做、也确实擅长的那类基于形式的任务。而过去这几年神经语言模型的革命带来的变化是,当你从语言模型里提取词嵌入时,你就得到了一种对词的分布的非常精细的表示,这非常有用;有些模型甚至能做到……让你拿到的词嵌入是带上下文的。也就是说,关于这个词、关于它可能和什么共现的信息,不是它在所有文本中的整体情况,而是它在当前上下文中的情况。
便签引用
28:39
Super useful, but not the same thing as understanding language. I kept getting into arguments with people who were not linguists who wanted to say, "Yeah, it is." So, Alexander Koller and I wrote this paper to just sort of say, "Okay, look, here's the argument why not," with the hopes that that would put an end to it, and it didn't. People still want to come argue with me about this. The thing that is really hard to see — and sort of the value of linguistics in this place — is that when we use language, we use it...and I'm sorry, I'm going to pull out a philosopher on you here, but Heidegger has this notion of throwness.
超级有用,但这跟理解语言不是一回事。我不断地跟一些非语言学背景的人争论,他们非要说:「不,这就是理解。」所以我和 Alexander Koller 写了那篇论文,本来是想说:「好,你看,这就是为什么不是的论证」,希望能就此打住,结果并没有。大家还是要来跟我争。真正难看清的一点是——这也正是语言学在这里的价值所在——就是我们使用语言的时候,我们……抱歉,我要在这儿搬出一位哲学家了,海德格尔有个概念叫「上手/沉浸」(throwness)。
便签引用
29:14
So, you're in a state of throwness when you are not aware of the tool you are using. If you think about typing on a keyboard, when it's going well the keyboard disappears. And then, you have a key that sticks and then all of a sudden, the keyboard is very "there" for you again. Well, language is the same way. When we are speaking a language that we are fluent in, it is not very visible to us until something makes us focus on it. And of course, linguistics is all about focusing on the language. So, linguists are used to doing that.
当你意识不到自己正在使用的工具时,你就处在那种状态里。想想在键盘上打字:顺利的时候,键盘是消失的。可一旦有个键卡住了,键盘一下子又「杵在那儿」了。语言也是一样。当我们说一门自己很流利的语言时,它对我们几乎是不可见的,直到某件事逼着我们去注意它。而当然,语言学做的就是去关注语言本身。所以语言学家习惯了这么做。
便签引用
29:42
When we talk about giving words to a language model, it's really important to distinguish between the word as a sequence of characters as opposed to word in the sense of a pairing of form and meaning. Because what the language model is seeing is only the sequence of characters. It's a bit easier to imagine what that's like if you think about a language you don't speak. So, what's a language you don't speak? Lukas: Mandarin. Emily: Mandarin. Okay. You don't speak Mandarin. I assume you also therefore don't read Mandarin. Lukas: Definitely don't.
当我们说「把词喂给语言模型」时,非常重要的一点是要区分:词作为一串字符,和词作为形式与意义的配对,这是两回事。因为语言模型看到的只有那串字符。如果你想想一门你不会说的语言,就比较容易想象那是什么感觉。那么,有什么语言是你不会说的?Lukas:普通话。Emily:普通话,好。你不会说普通话,那我猜你也不认识中文。Lukas:肯定不认识。
便签引用
30:11
Emily: Maybe recognize a couple of the characters? Lukas: I mean, I read Japanese. So, there's some overlap. Emily: Okay, so let's go further away. Do you read Cherokee? Lukas: No, definitely not. Emily: Okay. So, Cherokee has got this wonderful syllabary, it's a writing system where the characters represent syllables. If someone showed you a whole bunch of Cherokee text, that experience of looking at it would be a better model for what the computer is doing than you looking at English text, because you can't help but get the meaning part when you're looking at it. Because English is a language you speak and read.
Emily:也许能认出几个字?Lukas:我是说,我能读日语。所以多少有点重叠。Emily:好,那我们挑个更远的。你能读切罗基语吗?Lukas:不,绝对不能。Emily:好。切罗基语有一套很棒的音节文字,它的书写系统里每个字符代表一个音节。如果有人给你看一大段切罗基语文本,你看它的那种体验,比你看英文文本更接近计算机在做的事情,因为你看英文的时候没法不把意义那部分一起接收进来。因为英语是你会说也会读的语言。
便签引用
08图灵测试为什么不成立
30:42
Mandarin is kind of in between there because you would pick up a few of the hanzi that you recognize from Japanese kanji, and it wouldn't be quite the same. Lukas: I guess...I don't know, I don't want to argue with you. But I do want to, I guess, advocate for... I don't know, I mean, I have not thought deeply about this topic. What I have seen in my life is these language models working better and better than I could have imagined from the strategy that they employ and sort of seeming like they're getting more and more subtle detail. Of course, when I was a kid, I learned about the Turing test, which seems like a pretty good test of understanding on its face. I think the test is like, if you have a conversation with something and you can't tell if it's an automated system or a human, then we can say that it has intelligence, tand it sort of seems to me like these language models are on the verge of passing the Turing test.
普通话算是介于两者之间,因为你会认出几个跟日语汉字重合的汉字,感觉不太一样。Lukas:我想……我也说不好,我不是想跟你抬杠。但我确实想,怎么说呢,替另一边说几句……我也不确定,我是说,我并没有对这个话题做过很深的思考。我这辈子看到的是,这些语言模型越做越好,好到超出我根据它们所用的策略所能想象的程度,而且似乎能捕捉到越来越细微的细节。当然,我小时候学到过图灵测试,表面上看它像是一个相当不错的关于理解的测试。我记得那个测试大概是:如果你跟某个东西对话,分辨不出它是自动系统还是人,那我们就可以说它具有智能。而在我看来,这些语言模型似乎快要通过图灵测试了。
便签引用
31:52
What would it take for you to feel like some automated technique actually has understanding of what it's consuming? Emily: Yeah. So, I think the first thing I want to say about the Turing test is the reason it doesn't work...and I hate to disagree with a giant like Turing, because Turing's work was really important and foundational— Lukas: But it was 100 years ago, it's possible to miss something. Emily: 70? Lukas: 70, fair, fair. All right, 80? 70? Okay, 70. Sorry. Emily: As it turns out, people are too willing to make sense of language and too willing to sort of build the context behind something that would make something make sense. And so, we are not well positioned to actually be the testers in a Turing test. That's why that doesn't work.
要怎样你才会觉得某种自动化技术真的理解了它所摄入的东西?Emily:嗯。关于图灵测试,我想说的第一件事是它为什么不成立……我也不太愿意去反驳图灵这样的巨人,因为图灵的工作真的非常重要、非常奠基性——Lukas:但那是一百年前了,漏掉点什么也正常。Emily:七十年?Lukas:七十年,好吧好吧。八十?七十?行,七十。抱歉。Emily:事实证明,人太愿意去从语言里读出意义,太愿意去替它补上能让它说得通的背景。所以我们其实并不适合去当图灵测试里的那个测试者。这就是它不成立的原因。
便签引用
32:47
Language models, because they can come up with coherent-seeming text. These are probable sequences, given a little bit of noise and where you start, what would likely come next based on all that training data. Then it sort of comes out as something that we can make sense of, and then we are sort of easily fooled into thinking that it actually meant to communicate that. So, you're asking the question of "What would show that a machine has understanding?" I think part of it is, well, let's talk about actually interfacing with the world in some way. We certainly do have cases where machines in restricted domains for restricted ranges of things that they can do, do understand. So, when you ask your local corporate spy bot to do something for you and it does the thing, it has understood.
语言模型之所以能骗过人,是因为它们能生成看起来连贯的文本。这些都是高概率的序列——给定一点噪声和一个起点,根据全部训练数据推断接下来最可能出现什么。于是输出的东西是我们能读出意义的,我们就很容易被骗,以为它真的有意要传达那个意思。所以你问的问题是:「什么才能表明一台机器具有理解?」我觉得其中一部分是,我们得谈谈以某种方式跟世界产生接口。确实有一些情况,机器在受限的领域、在它能做的受限范围内是有理解的。比如你让你家里那个「企业间谍机器人」帮你做件事,它做到了,那它就是理解了。
便签引用
33:36
Lukas: Wait, sorry, what's a local corporate spy bot? Sorry, could we make this a little more concrete? Emily: I'm making a snarky remark about the privacy implications of things like Siri and Alexa and Google Home. Lukas: Oh, I see, I see, gotcha. Emily: And Samsung Bixby is in the same space. Microsoft had Cortana. Right? Lukas: Right, right. Gotcha. Emily: So, when you ask those things to set a timer, or turn on the lights, or dial a phone number or whatever, and it works, then, yes, to a certain extent, it has understood. And it has understood because its training setup was looking at not just language but something external to the language that needed to map to that. And so, that's a kind of understanding. The question is — for somebody who was interested in doing that across some more general range of things — the question is, how do you set up tasks that require some kind of action in the world, so that it can't be done just by bulldozing it with a language model and say "Well, this is a likely thing to come next.", right? Lukas: You got to describe your octopus
Lukas:等等,抱歉,什么叫「企业间谍机器人」?抱歉,我们能不能讲得具体点?Emily:我是在挖苦 Siri、Alexa、Google Home 这类东西的隐私问题。Lukas:哦,明白了明白了。Emily:三星的 Bixby 也是同一类。微软有过 Cortana,对吧?Lukas:对对,懂了。Emily:所以当你让这些东西设个闹钟、开个灯、拨个电话号码之类的,而它做到了,那么在一定程度上,是的,它理解了。它之所以理解,是因为它的训练设置里看的不只是语言,还有语言之外、需要被映射过去的东西。这就是一种理解。问题在于——对于想把这件事推广到更一般范围的人来说——问题是:你要怎么设计任务,让它必须在世界中采取某种行动,从而没法靠一个语言模型硬推过去、只说一句「唔,接下来最可能是这个」就蒙混过关,对吧?Lukas:你得讲讲你那个章鱼思想实验,那个非常有画面感。我还有些问题。
便签引用
09章鱼思想实验:只听不接地
34:45
thought experiment because that was very evocative. And I have some questions. Emily: Okay. So, the octopus thought experiment is about not just being able to understand but learning to understand. That's the difference between it and both the Turing test and Searle's thought experiment, where both of those basically say, "Imagine someone has set up the whole system." Then we could test for intelligence or we can...from a philosophical point of view, say it's still not understanding. So, the system exists and we are thinking about it or testing it.
Emily:好。章鱼思想实验讲的不只是「能不能理解」,而是「能不能学会理解」。这正是它跟图灵测试以及塞尔的思想实验的区别——那两个基本上都是说:「假设有人已经把整个系统搭好了。」然后我们可以测它有没有智能,或者从哲学角度说它仍然不算理解。也就是说,系统已经存在,我们只是在思考它或测试它。
便签引用
35:17
The octopus is this thing of saying, "Okay, if we had something that we assume, we posit that it is hyperintelligent..." and then that's part of why we picked the octopus. In fact, it was initially a dolphin, but we decided that octopuses are inherently more entertaining. Also, it was better because dolphin's environment is a bit closer to a human's environment. So, we wanted the octopus to be something that is posited to be super intelligent. And they are, I think, understood to be intelligent creatures...as smart as it needs to be. That's not the issue.
而章鱼这个设定是说:「好,假如我们有一个我们假定为超级智能的东西……」——这也是我们为什么挑章鱼的原因之一。其实最初设定的是海豚,但我们觉得章鱼天生更有趣。而且这样也更好,因为海豚的生活环境跟人类的环境更接近一些。我们想让章鱼是一个被假定为超级聪明的东西。而且我觉得大家本来也认为章鱼是聪明的生物——要多聪明有多聪明。聪明不是问题所在。
便签引用
35:48
We are assuming intelligence, but then we are only giving it access to the form of language. In our scenario, you have these two English-speaking humans who end up stranded on two nearby islands. They're otherwise uninhabited, but they've had previous inhabitants who set up an undersea telegraph cable. These two humans can communicate with each other. We left it offstage how they discovered the telegraph or that the other one's on the other end, whatever, just assume it exists. It's the thought experiment, you can do things like that. You know, assume a spherical cow, except we don't need spherical cows.
我们假定它有智能,但我们只给它接触语言的形式。在我们的设定里,有两个说英语的人分别流落到两座相邻的岛上。岛上原本无人居住,但之前的居民铺过一条海底电报电缆。这两个人可以互相通信。他们怎么发现电报、怎么知道另一头有人,我们就不交代了,假定它就是存在。这是思想实验嘛,可以这么设定。你懂的,「假设有一头球形的牛」,只不过我们不需要球形的牛。
便签引用
36:22
So, telegraph cable and the humans are named A and B. They're basically using English as encoded in Morse code to talk to each other. This hyperintelligent deep sea octopus that we called O comes along and taps into that cable. The octopus can feel the pulses going through for Morse code. The question is, what could the octopus actually potentially learn here? Because this is a hyperintelligent octopus – it's got as much time as it wants, as much memory as it wants — it is able to very closely model the patterns of what's likely to come next.
所以,有电报电缆,两个人分别叫 A 和 B。他们基本上是用摩尔斯电码编码的英语互相交谈。这时一只我们叫它 O 的超级智能深海章鱼过来,接入了那条电缆。章鱼能感受到摩尔斯电码的脉冲。问题是:章鱼在这里究竟可能学到什么?因为这是一只超级智能的章鱼——它有想要多少就有多少的时间、想要多少就有多少的记忆——它能够非常精细地建模「接下来最可能出现什么」的规律。
便签引用
36:58
In our story, the octopus decides for some reason that it's lonely and it's going to cut the cable and pretend to be B while talking to A. On reflection, it's like, "Poor B, just cut off from the world." So, maybe the octopus is also talking to B pretending to be A, but we don't talk about that part. The question is, under what circumstances could the octopus continue to fool A that it's actually B? We say this is in a sense, a weak version of the Turing test. The way the Turing test was set up, A is given the task of deciding "Am I talking to a human or not?"
在我们的故事里,章鱼出于某种原因觉得孤独,于是决定切断电缆,冒充 B 去跟 A 说话。回头想想,「可怜的 B,就这么跟世界断了联系。」所以也许章鱼同时也在冒充 A 跟 B 说话,不过那部分我们没展开。问题是:在什么情况下,章鱼能一直骗过 A、让 A 以为它就是 B?我们说这在某种意义上是图灵测试的一个弱化版本。图灵测试原本的设置里,A 被赋予的任务是判断「我是在跟人说话吗?」
便签引用
37:32
And here, there's subterfuge. The octopus, its mere existence is unknown to A. If there's just sort of like chitchat pleasantries, those things you can just kind of follow a pattern and it's relatively inconsequential as long as what's coming out is internally coherent. And even if it's a little bit incoherent, well maybe B is just being silly. It doesn't matter so much. Okay, well, O could get away with that. But once you get more towards things where A actually really cares about communicating ideas to B and getting ideas back from B, it's going to get harder and harder for the octopus to maintain this semblance of good communication. We go through this example where A builds a coconut catapult and the octopus is able to send back sort of like, "Very cool invention. Great job." or something, even though A was asking for like, "Well, what happened when you built it?"
而这里有欺骗成分:章鱼的存在本身,A 是完全不知道的。如果只是些寒暄闲聊,那些东西照着套路走就行,而且相对无关紧要,只要输出的内容自身连贯就好。哪怕稍微有点不连贯,那也许只是 B 在犯傻,无所谓。好,这种情况 O 是能蒙混过去的。但一旦话题转向 A 真的很在乎把想法传达给 B、并从 B 那里得到回应,章鱼要维持这种「沟通良好」的假象就会越来越难。我们举了一个例子:A 造了一个椰子投石机,章鱼能回过去一句「发明得真酷,干得漂亮」之类的话,尽管 A 问的是「你造出来以后结果怎么样?」
便签引用
38:27
But the octopus has no experience of things like coconuts or rope or stuff like that. So, it can't reason about those things in the world, or even know that A is actually talking about them. All it can do is come back with, "Well, what's the likely form of a response in this context?" To the extent that O gets away with that, it's because A is willing to make sense of those utterances. O has no meaning in this scenario. And then finally, we have a bear show up and start attacking A, and A says to O — or to B, actually — "Help. I'm being attacked by a bear. All I have are these two sticks, what should I do?" At that point, O is utterly useless, and so we say this is the point at which O would definitely fail the Turing test if A survived being eaten by the bear.
但章鱼对椰子、绳子这类东西没有任何经验。所以它没法对世界中的这些东西做推理,甚至不知道 A 其实是在说这些东西。它能做的只是给出:「在这个语境里,回应的可能形式是什么?」O 之所以能蒙混过关,是因为 A 愿意去把那些话读出意义来。在这个场景里 O 是没有意义可言的。然后最后,我们让一只熊出现,开始攻击 A,A 就对 O ——其实是对「B」——说:「救命,我被熊袭击了,我手上只有两根棍子,我该怎么办?」这时候 O 就彻底没用了,所以我们说这就是 O 必定通不过图灵测试的时刻——前提是 A 没被熊吃掉、活下来了。
便签引用
39:14
But then we tried with GPT-2 like, what would it say? The answers were hilarious. The words are in the right topic area enough that it comes back with something funny and I encourage people to go look at the appendix to our paper where we put these, but it's never going to be helpful. And it's not actually expressing communicative intent. Lukas: Well, I have to say, walking into that paper without knowing the context, I really enjoyed it. For me, I especially enjoyed it because the sort of concreteness of the thought experiment that was like evocative but also makes you think, "Huh, what do I think about that?"
然后我们还真拿 GPT-2 试了试,看它会说什么。答案笑死人。那些词大体落在对的话题范围内,所以它会回出一些很好笑的东西,我建议大家去看看我们论文的附录,我们把这些都放在那儿了,但它绝不可能真的帮上忙。而且它并没有真正在表达任何交流意图。Lukas:我得说,我当时是在不知道背景的情况下读到那篇论文的,读得非常开心。对我来说尤其享受的是,那个思想实验很具体,既有画面感,又会让你想:「嗯,那我自己怎么看?」
便签引用
10反驳与回应:人靠语言学到经验之外
39:51
What I kept thinking was...for me, I feel like I've learned about a lot of things that I haven't experienced, I was especially thinking about learning math, where there's all these abstract topics. I feel like in a way I learned about math, in some sense, through form almost. It's all in my head. I'm like learning things as visualizing them. It seems possible to learn to reason about things that you haven't seen or experienced just from a stream of words. I even remember actually grading blind student's papers. It was really interesting, how they walked through stuff in a math class, and it seemed like they were visualizing things even though they were blind from birth. So, I'm just wondering ... I guess, I'm not totally convinced that the octopus couldn't somehow figure out what a catapult does if they listen to all language. Emily: So, if the octopus had actually had a chance to learn English, then yes. But it didn't because it never got that initial grounding. And we absolutely learn things through language that are outside of what we've directly experienced.
我一直在想的是……对我而言,我觉得我学到过很多我从未亲身经历的东西,我特别想到的是学数学,那里全是抽象的概念。我感觉某种意义上我几乎是通过形式学会数学的。全都在我脑子里,我是靠可视化去学这些东西的。似乎有可能仅仅从一串词语出发,就学会对你没见过、没经历过的事物进行推理。我甚至还记得批改盲人学生的作业。特别有意思的是,他们在数学课上推演问题的方式,看起来像是在做视觉想象,尽管他们是先天失明的。所以我就在想……我想说,我并不完全相信,如果章鱼听遍了所有语言,它就一定没法搞明白投石机是干什么的。Emily:如果章鱼真的有过学英语的机会,那当然可以。但它没有,因为它从来没有得到过最初的那种「接地」。而我们确实会通过语言学到超出自己直接经验的东西。
便签引用
41:05
Conversely, if you as a sighted person wanted to understand what it was like to live as a blind person, you could listen to or read what a blind person has to say about that and learn about it. So, that's definitely something that we can do. But we can do it because we have acquired linguistic systems. When we use language to communicate, we absolutely tell each other ideas and things that are outside of even our own experiences, right? We invent things, and then transmit that to other people. But we do that based on this shared system that tells us, "Okay, here's the range of possible forms. These are the well-formed words and sentences. These are the sounds that we use in this language. These are the way the words are built up, the sentences are built up. And these are the standing meanings that they map to."
反过来,如果你作为一个视力正常的人想了解失明者的生活是什么样的,你可以去听或去读盲人自己怎么讲,然后从中了解。所以这绝对是我们能做到的事。但我们之所以能做到,是因为我们已经习得了一套语言系统。当我们用语言交流时,我们确实会把一些想法、一些甚至超出我们自身经验的东西告诉彼此,对吧?我们发明东西,然后把它传达给别人。但我们能这么做,是基于一套共享的系统,这套系统告诉我们:「好,这是可能的形式的范围。这些是合乎语法的词和句子。这些是这门语言里使用的音。这些是词的构成方式、句子的构成方式。这些是它们所映射到的固定意义。」
便签引用
41:49
And then we use those standing meanings to make guesses about communicative intent. The problem for the octopus isn't that it's not smart. We said it's hyperintelligent. It isn't that if it knew the language, couldn't understand those things. It's that its exposure to the language is not set up so that it can actually learn it as a linguistic system, all it can learn is distributional patterns. Lukas: I guess what prevents the octopus from learning language over time like a human probably would? Emily: Okay, so, it doesn't get to do...and in the paper, we go into human language acquisition.
然后我们再用这些固定意义去猜测对方的交流意图。章鱼的问题不在于它不聪明。我们说了它是超级智能的。问题也不在于它如果懂这门语言、就理解不了这些事。问题在于,它接触语言的方式,并不足以让它真正把语言学成一套语言系统,它能学到的只有分布规律。Lukas:那我想问,是什么阻止了章鱼像人一样随着时间学会语言呢?Emily:好,它没机会……我们在论文里也讨论了人类的语言习得。
便签引用
42:23
For first language acquisition, it's all about joint attention. When babies learn language, it starts from social connections to their caregivers and understanding that the caregivers are communicating something to them, and then mapping the words onto those communicative intent. The child language literature talks about the importance of joint attention, that kids learn words when their caregivers follow into their attention, and attend to the same things, and then provide those words. That experience, that mapping, the octopus doesn't get that. It's just getting the words going by Lukas: Do you think there's some algorithm possibly that could exist, that could take a stream of words and understand them in that sense?
就母语习得而言,关键全在于「共同注意」。婴儿学语言的时候,起点是与照料者之间的社会性联结,是意识到照料者在向自己传达某种东西,然后把词映射到那个交流意图上。儿童语言研究文献强调共同注意的重要性:孩子学会词,是因为照料者顺着孩子的注意力走,和孩子一起注意同样的东西,然后给出对应的词。这种经验、这种映射,章鱼是得不到的。它得到的只是流过去的一串词。Lukas:你觉得有没有可能存在某种算法,能够接收一串词语,并在那种意义上真正理解它们?
便签引用
43:10
Emily: Natural language understanding is a tremendously difficult problem because it relies not just on the linguistic system, but also on world knowledge and common sense, reasoning, all kinds of things. So, you can certainly use — I'm more certain than I actually am — but there's a big difference between saying, "I'm going to build an algorithm that has understanding of linguistic structure, has understanding of linguistic meaning, has understanding of how those meanings map to a model of the world, and then use that to understand," versus "I'm going to build a system that only gets linguistic form and assume that it will get to understanding in some way."
Emily:自然语言理解是一个极其困难的问题,因为它依赖的不只是语言系统本身,还有世界知识、常识、推理,各种各样的东西。所以,你当然可以用——我这话说得比我实际确信的还要肯定——但这两种说法差别很大:一种是“我要造一个算法,它理解语言结构,理解语言意义,理解这些意义如何映射到一个世界模型上,然后用它来做理解”,另一种是“我要造一个只接触语言形式的系统,然后假定它会以某种方式获得理解”。
便签引用
43:42
So, yes. You could go much, much further with algorithms that have more in their input, in their training input, than just form. That's going to be things like visual grounding. It's going to be things like the ability to possibly query people for answers. It might be knowledge bases. It might be other sensors in some sort of embodied... I'm not saying that natural language understanding is impossible and not something to work on. I'm saying that language modeling is not natural language understanding. Lukas: But just so I'm clear, just consuming language without kind of all this extra stuff, you're arguing that no algorithm could from just that really understand language? Emily: By language, I mean, form.
所以,是的。如果算法的输入、训练输入里有比形式更多的东西,那是可以走得远得多的。比如视觉落地(visual grounding),比如有能力向人提问来获取答案,可能是知识库,也可能是某种具身系统里的其他传感器……我不是说自然语言理解不可能、不值得做。我是说,语言建模不等于自然语言理解。Lukas:不过我确认一下,如果只是消化语言、没有这些额外的东西,你是在说没有任何算法能仅凭这些就真正理解语言?Emily:我说的“语言”,指的是形式。
便签引用
11泰语图书馆与罗塞塔石碑
44:28
Imagine that you are dropped into the Thai equivalent of the Library of Congress, and you have around you any book you could possibly want in Thai, but only in Thai. For some reason, this library doesn't have Thai-Chinese, Thai-French, Thai-English dictionaries. It's just Thai. Could you learn Thai? Lukas: I think so. I guess what's hard is that I have a language already. But I feel like I- Emily: So, what would you do? What would be your first step to learning Thai if you have just oodles and oodles of Thai books and that's it around you? Lukas: What would I start to do? I mean...I'm not sure. Do you think I couldn't learn Thai? Emily: So, I'm curious about what you ... So, you as a person, could you learn Thai? Sure. You could go take a Thai language class.
想象一下你被丢进泰国版的国会图书馆,周围是你能想到的任何泰语书,但只有泰语。不知为何,这座图书馆里没有泰汉、泰法、泰英词典。全都是泰语。你能学会泰语吗?Lukas:我觉得能。难点大概在于我本来就已经有一门语言了。不过我感觉我——Emily:那你会怎么做?你学泰语的第一步会是什么?如果你周围只有一大堆一大堆的泰语书,就这些?Lukas:我会先做什么?我是说……我也不确定。你觉得我学不会泰语吗?Emily:我好奇的其实是你……那么,作为一个人,你能学会泰语吗?当然能。你可以去上泰语课。
便签引用
45:12
Lukas: No, no, I mean from in this situation, just sort of dropped in. I mean people do learn... How did people learn hieroglyphics or something when there's no one around that still knows it? Do they need to find like a Rosetta Stone? Or can they- Emily: The Rosetta Stone is what unlocked the hieroglyphics. If you don't have something like that, then what you have to do is resort to hypotheses about distributions and say, "What do we know about the world in which these texts were written? What do we know about how languages work?"
Lukas:不不,我是说在这个情境里,就这么被丢进去。我是说,人确实能学会……当年已经没人还懂象形文字了,人们是怎么学会的?他们是不是得找到类似罗塞塔石碑的东西?还是说——Emily:正是罗塞塔石碑解开了象形文字。如果你没有那样的东西,那你就只能退而对分布做假设,然后问:“关于这些文本写作时的那个世界,我们知道些什么?关于语言是怎么运作的,我们知道些什么?”
便签引用
45:42
Can we say, "Okay, well given frequency analyses and the length of the words, it seems like a language that's got separate function words instead of lots of morphology. So, that thing might be an article, that thing might be a copy of a verb," and you could do some analysis like that. It's not what language models are doing. To get from those sort of structural things into something about meaning, you have to make guesses about what's being described. You have to basically bring in some world knowledge and say, "How well does this fit?"
我们能不能说:“好,根据词频分析和词长,这门语言看起来有独立的功能词,而不是大量形态变化。所以那个东西可能是冠词,那个东西可能是系动词”,你可以做这类分析。但这不是语言模型在做的事。而要从这些结构性的东西走到意义层面,你就得去猜测被描述的是什么。你基本上必须引入一些世界知识,然后问:“这样解释契合得怎么样?”
便签引用
46:13
When I asked you the question of what would you do, I was thinking, well, possible answers are, "I would go find an illustrated encyclopedia that has pictures in it." There's some visual grounding. Or I would go find a book from whose cover I could tell it was actually the Thai translation of Curious George. Lukas: These are great suggestions. Emily: Yes. But all of that is bringing in external things. And then once you have a foothold, you can build on it. That's an interesting way to go. But if you just have form, it's not going to give you that information.
我刚才问你会怎么做的时候,我心里想的是,可能的答案是:“我会去找一本有插图的百科全书。”那就有了视觉落地。或者,我会去找一本从封面就能看出是《好奇的乔治》泰语译本的书。Lukas:这些主意都很棒。Emily:是的。但这些全都是在引入外部的东西。而一旦你有了一个立足点,就可以在上面往下建。这是个有意思的路径。但如果你只有形式,它是不会给你那些信息的。
便签引用
12对语言模型未来的三个预测
46:48
Lukas: Wow, interesting. Thank you, this was really interesting. I guess, my last question on this topic is, do you sort of predict that these language models will run into problems that we'll really experience and then we'll have to kind of change the approach? Or do you think that as our bar for applications of natural language goes up, they'll just sort of adapt and find ways to incorporate external information, kind of like finding the Curious George translation? Emily: I think that language models are going to remain useful. I mean, language models have been an important component of language technology since Shannon's work in the 1950s. This is longstanding. But I think that we are likely — it's so hard to predict the future — but my guess is that...or maybe what I would like to see is that we get to a more stringent sense of what works and what sort of an appropriate range of failure modes and what kind of fail safes we need.
Lukas:哇,有意思。谢谢,这真的很有启发。我在这个话题上的最后一个问题是:你会不会预测这些语言模型会撞上一些我们真切感受得到的问题,然后我们不得不改变路线?还是说你觉得,随着我们对自然语言应用的要求越来越高,它们会自己适应、找到办法纳入外部信息,有点像去找那本《好奇的乔治》译本?Emily:我认为语言模型还会继续有用。我是说,从香农上世纪五十年代的工作开始,语言模型就一直是语言技术的重要组成部分。这是由来已久的。但我觉得我们很可能——预测未来太难了——但我猜……或者说我希望看到的是,我们对什么算有效、什么算合理的失败模式范围、需要什么样的安全兜底,形成一套更严格的认识。
便签引用
47:50
People are going to find that putting language models at the center of something where your application really requires you to have a commitment to the accountability for the words that are uttered is going to be a very fragile way to go. My guess is that when we get to that point, we're going to de-center the language models and have them be something that is selecting one possible output again or providing these word embeddings, but they are not a step towards general-purpose language understanding the way they are hyped to be.
人们会发现,把语言模型放在核心位置——而你的应用又真正要求你对说出口的话负责——会是一条非常脆弱的路。我猜等我们走到那一步,我们会把语言模型从中心位置挪开,让它重新变成从候选里挑一个输出的东西,或者提供词向量的东西,但它们并不是像被吹捧的那样,是通向通用语言理解的一步。
便签引用
48:22
That's one set of problems. If you have to have accountability for the words that are uttered, you do not want a stochastic parrot. You want something that will speak for you in a reliable way, not just make up what sounds good. The other thing is if we take seriously these issues around bias and encoding and amplifying bias and training data, I think we're going to find that we want to work with algorithms that can make more of smaller datasets, so that we can be better about curating and documenting and updating those datasets so that they stay current with what's going on, rather than this path right now that relies on very large language models.
这是一类问题。如果你必须对说出口的话负责,你不会想要一只随机鹦鹉(stochastic parrot)。你想要的是能可靠地替你说话的东西,而不是编一些听着顺耳的话。另一件事是,如果我们认真对待训练数据里偏见的编码与放大这些问题,我认为我们会发现,我们希望用那些能把更小的数据集用得更充分的算法,这样我们才能更好地筛选、记录和更新这些数据集,让它们跟得上现实的变化,而不是走现在这条依赖超大语言模型的路。
便签引用
48:59
So, those are my guesses. There's also the environmental angle. Well, actually, the "energy uses" angle is both environmental but also about technology, to a certain extent. I think there are more and more people — and there's Schwartz et al, Strobel et al, Henderson et al. — a bunch of work now saying, "Hey, let's make sure we're also measuring the environmental impact as we do things, or the carbon footprint so that we can direct effort to doing things in a more and more efficient way." There's that angle, but there's also many situations where you don't have the whole cloud available. If you want to do computing on a mobile device, you're not going to be able to have an absolutely enormous language model in there.
这些就是我的猜测。还有环境这个角度。其实,“能耗”这个角度既关乎环境,某种程度上也关乎技术本身。我觉得现在越来越多的人——比如 Schwartz 等人、Strobel 等人、Henderson 等人——已经有一批工作在说:“嘿,我们做事的时候也要衡量环境影响,或者碳足迹,这样才能把力气引向越来越高效的做法。”这是一个角度,但还有很多场景是你没有整片云可用的。如果你想在移动设备上做计算,你不可能在里面塞一个极其庞大的语言模型。
便签引用
13基准测试:地图不是疆域
49:38
There's pressure to find leaner solutions. I think that's a win-win, environmentally and then in terms of more flexibility with technology. Lukas: Totally, totally. And it's a good segue because you pointed out a bunch of this stuff in your papers about benchmarks, which I'd love to talk about a little bit, and maybe you could kind of summarize...maybe start [with] what are benchmarks, probably most people know, but then what are the possible pitfalls with them? Emily: Yeah. I should say this is a paper called "AI and the Everything in the Whole Wide World Benchmark"³ that we presented at a workshop called Machine Learning Retrospectives at NeurIPS last year. It's joint work with Deb Raji, Alex Hanna, Emily Denton, and Amandalynne Paullada. Another collaboration where...in this case, we actually do have meetings where we talk to each other, but of those people, the only one I've met in person so far is Amandalynne, who's a PhD student in my department. Pandemic life, right?
这就形成了寻找更精简方案的压力。我认为这是双赢,既对环境好,技术上也更灵活。Lukas:完全同意,完全同意。这也正好接上下一个话题,因为你在关于基准(benchmark)的论文里指出了不少这类问题,我很想聊一聊。也许你可以概括一下……先说说什么是基准,大多数人可能都知道,然后再说说它们可能有哪些坑?Emily:好。我得说明一下,这篇论文叫《AI and the Everything in the Whole Wide World Benchmark》³,是我们去年在一个叫Machine Learning Retrospectives 的 NeurIPS 研讨会上发表的。这是我和 Deb Raji、Alex Hanna、Emily Denton、Amandalynne Paullada 的合作。又是一次这样的合作……不过这次我们确实开会、彼此交流,但这些人里我到目前为止唯一当面见过的是 Amandalynne,她是我们系的博士生。疫情生活嘛,对吧?
便签引用
50:36
We got together because we were talking about the ways in which benchmarks are being misused in the AI hype machine and in AI research that is striving for generality and overclaiming what the benchmark shows. So, a benchmark is basically a standardized data set, typically with some gold standard labels. Although you could also have benchmarks for things where the labels are inherent, like language modeling. What word actually came next is the gold standard level. The idea is that you might have a standardized set of training data, or possibly not, and then you've got the standardized test data. People can test different systems against this. You have this chance of saying "Which system is more effective in this training regime?" or "Given this training data against that test data?" So, that's a benchmark.
我们凑到一起,是因为我们在聊基准如何被 AI 炒作机器、以及那些追求通用性的 AI 研究所误用,夸大基准所能说明的东西。基准基本上就是一个标准化的数据集,通常带有金标准标注。当然也可以有那种标签本身内在自带的基准,比如语言建模——下一个词实际是什么,就是金标准。它的思路是,你可能有一套标准化的训练数据,也可能没有,然后你有标准化的测试数据。人们可以拿不同的系统在上面测。这样你就有机会说:“在这个训练条件下,哪个系统更有效?”或者“给定这份训练数据、在那份测试数据上表现如何?”这就是基准。
便签引用
51:28
You asked me before if I could summarize the problems with benchmarks, and it's not so much benchmarks I have a problem, but the way that they're used. I think this is an example of "the map is not the territory". People will tend to say, "Oh, here's this benchmark about computer vision." ImageNet is that. Or, "Here's a benchmark about natural language understanding of English," and that's GLUE and SuperGLUE. People will say...I've actually seen this in like a PR thing that came out of Microsoft saying that computers understand English better than people now, because this one setup scored higher than some humans on the GLUE benchmark. That's just a wild overclaim. and it's a misuse of what the benchmark is for.
你刚才问我能不能概括基准的问题,其实我有意见的不是基准本身,而是它被使用的方式。我认为这是一个“地图不是疆域”的例子。人们往往会说:“哦,这是一个关于计算机视觉的基准。”ImageNet 就是这样。或者“这是一个关于英语自然语言理解的基准”,那就是 GLUE 和 SuperGLUE。人们会说……我真的在微软发出的一份公关稿里看到过,说计算机现在理解英语比人还好,就因为某一套系统在 GLUE 基准上比一些人类得分更高。这纯属离谱的夸大,也是对基准用途的滥用。
便签引用
52:14
So, what's the problem with the overclaims? Well, it kind of messes up the science. We're not doing science if we're not actually matching our conclusions to our experiments. We live in a world of AI hype, which means that people are more likely to buy in to and set up solutions that don't function as advertised because they live in a world where people are being told that Microsoft has built a system that understands English better than humans do. Of course, you could also build an AI system that does whatever other implausible thing like, "Guesses someone's political affiliation by the way they smile" or something, which makes no sense. But we live in a world where there's all these claims, overclaims about AI, and that makes these other ones also sound more plausible than they should.
那这些夸大有什么问题呢?首先,它把科学搞乱了。如果我们的结论跟我们的实验对不上,那我们做的就不是科学。我们生活在一个 AI 被过度炒作的世界里,这意味着人们更容易相信并部署那些名不副实的方案,因为他们所处的世界里,人们被告知微软造出了一个比人类更懂英语的系统。当然,你也可以造一个 AI 系统去做别的什么不靠谱的事,比如“根据一个人笑的方式猜他的政治立场”之类的,这毫无道理。但我们就生活在一个充斥着各种断言、各种对 AI 的夸大的世界里,这就让那些别的说法听上去也比它们该有的更可信。
便签引用
53:01
So, those are the problems that I see. But, benchmarking is important. In the history of computational linguistics, there was a while where when you wrote a paper for the ACL, the Association for Computational Linguistics, you would say, "Here's my system. Here's how I built it. Here's some sample inputs and outputs," done. Then the statistical machine learning wave came through and brought with it the methodology of shared task evaluation challenges, which is sort of a historical version of benchmarking, where MIST and other organizations would say, "Okay, we want to work on speech recognition, and we want to actually get a sense of how these different systems compare to each other. So, we're going to run a shared task evaluation challenge where everyone gets the same training data, and we're going to have some held out test data that no one gets to see. At a certain point, all the competitors submit their systems and we see what happen."
这些就是我看到的问题。但基准测试本身很重要。在计算语言学的历史上,曾经有一段时间,你给 ACL(计算语言学协会)写论文,你会说:“这是我的系统。这是我怎么做的。这是一些输入输出样例。”完事。后来统计机器学习的浪潮来了,带来了共享任务评测挑战(shared task evaluation challenge)这套方法论,可以说是基准测试的历史版本,那时MIST 之类的机构会说:“好,我们想做语音识别,我们想真正了解这些不同的系统彼此相比如何。所以我们要办一个共享任务评测挑战,所有人拿到同样的训练数据,然后我们会留一份谁都看不到的测试数据。到某个时间点,所有参赛者提交自己的系统,我们看看结果如何。”
便签引用
53:58
That's an improvement in the science compared to what was going on before. But that is not the whole story. If you want to understand how the system is working, if you want to understand how to build the next system, you can't just test it on some standard thing. You also have to look at, "Well, what kinds of errors does it make?", and "How do the different systems compare not just in their overall number, but in their failure modes and which inputs work for them and which ones don't?" and on and on like that, as opposed to, "Okay, I got the highest score. I'm done." Lukas: Right, right. Well said.
相比之前的做法,这在科学上是一种进步。但这不是全部。如果你想理解系统是怎么工作的,如果你想知道下一个系统该怎么造,你就不能只是拿它在某个标准的东西上测一测。你还得看:“它犯的是哪些类型的错误?”“不同系统的差别不只在总分上,还有它们的失败模式,哪些输入它们能处理、哪些不能?”等等等等,而不是“好,我拿了最高分,我搞定了。”Lukas:对,对。说得好。
便签引用
14基准之外:测试套件、审计与破解
54:36
I don't have much to add there. Can you say a little more about like...I feel like this is a great paper in that you make these really concrete, sensible recommendations. You sort of suggest a few alternatives to benchmarks. Could you maybe run through those for anyone listening? Emily: Yes, absolutely. So, it's more of complements than alternatives to benchmarks. So, in addition to benchmarks, this can be used sort of as a sanity check or, "Okay, did my system actually do better than a super naive baseline?", or "I want to compare some systems head-to-head, let's use this benchmark." You might also use test suites, which are put together to sort of map out particular kinds of cases that you want to handle well, as opposed to just grabbing whatever happened to occur in your sample test data.
我没什么可补充的。你能不能再多讲讲……我觉得这篇论文很棒的一点是,你们提出了非常具体、非常合理的建议。你们提了几种基准的替代方案。能不能给听众过一遍?Emily:当然可以。其实它们更像是基准的补充,而不是替代。所以,除了基准之外——基准可以当作一种理智检验,比如“好,我的系统真的比一个特别朴素的基线更好吗?”,或者“我想把几个系统正面比一比,那就用这个基准”。你还可以用测试套件(test suite),它是专门编排出来的,用来勾勒你希望系统处理好的某些特定类型的情况,而不是随便抓来的一份样本测试数据里碰巧出现的东西。
便签引用
55:23
You might do auditing, which is very much akin to test suites in saying...so this is like Joy Buolamwini and Timnit Gebru and Deb Raji's work on auditing face recognition data sets, where they sort of systematically created the set looking at two genders and a range of skin colors and sort of say, "Okay, is this accuracy actually even across this set of people or no?" And they found out no. So, that's a- Lukas: How's that different than a benchmark? That kind of sounds like a benchmark, doesn't it? Emily: So, it's not the way benchmarks are typically created. You could imagine someone creating a benchmark that is sort of systematically mapping out a space, but that's not the practice.
你还可以做审计(auditing),这跟测试套件很像……比如 Joy Buolamwini、Timnit Gebru 和 Deb Raji 关于人脸识别数据集审计的工作,他们系统性地构建了一个集合,覆盖两种性别和一系列肤色,然后问:“好,在这组人身上,准确率真的是均匀的吗,还是不是?”结果他们发现并不是。所以这就是——Lukas:这跟基准有什么不同?听起来挺像基准的,不是吗?Emily:这不是基准通常的构建方式。你可以设想有人构建一个系统性铺开整个空间的基准,但实际做法不是这样。
便签引用
56:03
The practice is, "We are going to go grab some data from somewhere, and then hold out 10% of it to be the test and the other 90% is training, or 80% training, 10% dev," right? The way benchmarks are typically put together is, "Let's just grab a sample of data and see how well this thing works," as opposed to "Let's create a testing regime through test suites or through this auditing process that can allow us to find the contours of its failure mode." Not "How well does it work on average?" but "How well does it work for this case, and that case, and that case?"
实际做法是:“我们从某个地方抓一批数据,然后留出 10% 做测试,另外 90% 做训练,或者 80% 训练、10% 开发集”,对吧?基准通常的搭建方式是:“我们就抓一份数据样本,看看这东西效果如何”,而不是“我们通过测试套件或者审计流程建立一套测试机制,让我们能找出它失败模式的轮廓”。不是“它平均表现如何?”,而是“它在这种情况、那种情况、又一种情况下表现如何?”
便签引用
56:41
There's also adversarial testing, which is...a few different things fall under adversarial testing. Sometimes people will create test sets by going and collecting all the examples that previous systems did poorly on to make a particularly hard test set, which is interesting in the sense that it can filter out the sort of freebies that are too easy, but also doesn't necessarily guide anything towards better performance for a particular use case. Because it's just sort of like, "Well, we're selecting what was hard for the previous model," not "What's particularly important to get right or what's particularly likely to be frequent in our use case," and so on.
还有对抗性测试,有好几种不同的东西都归在对抗性测试下面。有时人们构建测试集的办法是,把以前的系统表现差的例子全都收集起来,做成一个特别难的测试集。这有意思的地方在于它能过滤掉那些太容易的送分题,但也不一定能把系统引向在某个具体用例上表现更好。因为这只是在说“我们挑的是上一个模型觉得难的东西”,而不是“什么是特别重要、必须做对的,或者在我们的用例里特别高频的”,等等。
便签引用
57:22
So, that's one kind of adversarial testing. Another one is what we did in the Build It, Break It shared test. This was Allyson Ettinger, Sudha Rao, Hal Daumé, and I in 2017, put together a shared task where we had system builders and then breaker teams. The breaker team's goal was to find minimal pairs, two examples that were minimally different to each other, but would work...for which the systems would work for one but not the other. That would be a way of sort of mapping out what causes system failure. So, you can look at that. You can look at error analysis. Take the test set from the benchmark or the dev set from then benchmark, and then go in and look and say, "Okay, what are the kinds of problems that are showing up?"
所以这是一种对抗性测试。另一种是我们在 Build It, Break It共享任务里做的。那是 2017 年,Allyson Ettinger、Sudha Rao、Hal Daumé 和我一起办的一个共享任务,我们有系统构建方,也有“破解”队。破解队的目标是找出最小对立对——两个彼此差异极小的例子,但系统对其中一个奏效、对另一个不奏效。这就是一种勾勒系统失败成因的办法。这个可以看。你还可以做错误分析。拿基准的测试集或者开发集,然后进去看,问:“好,都冒出来哪些类型的问题?”
便签引用
58:07
A lot of systems that rely on language models tend to do really poorly with negation, which is one of these things that's very important to the meaning, but tends to be a short word or subword, and so it is easy to miss. You can imagine speech recognition or machine translation, if you missed one word out of 20, it matters a lot what that word is. If you replace "a" with "the", in many cases, that's not going to cause a lot of problems. But if you just skipped a "not" somewhere? Lukas: Yeah, that makes sense. Yeah.
很多依赖语言模型的系统在处理否定时表现都很差,而否定正是那种对意义极其重要、却往往只是一个短词或子词、因而很容易被漏掉的东西。你可以想想语音识别或者机器翻译,如果二十个词里漏了一个,漏的是哪个词关系非常大。如果你把 a 换成 the,很多情况下不会造成太大问题。但如果你就是漏掉了某处的一个 not 呢?Lukas:是啊,这说得通。是。
便签引用
58:43
Emily: All of this is basically about looking at what it is we're trying to build, what it is we're testing on, how it fits into the motivating use cases, and then what works and what doesn't, and for what doesn't work, what are the implications? What happens in the real world if that failure happens? And also, what are the likely causes? What is tripping us up? All of that is what we would like to see, instead of the leaderboard-ism, which is everyone just trying to climb to the top of the pile-on, which doesn't feel like it's really...
Emily:这一切归根结底就是去看:我们想造的到底是什么,我们在什么上面测试,它如何契合那些驱动它的用例,然后什么有效、什么无效,以及对于无效的部分,后果是什么?那种失败发生在真实世界里会怎样?还有,可能的成因是什么?是什么把我们绊住了?这些才是我们想看到的,而不是排行榜主义——所有人都只想爬到那堆东西的最顶上,这感觉真的不太……
便签引用
59:17
I mean, people talking about the speed of progress in AI love to talk about how quickly those leaderboard changes and how quickly the state-of-the-art, SOTA, gets higher and higher on these various benchmarks. I always think, "Yeah, but so?" What does that actually mean in terms of understanding the world better from a scientific point of view or building technology that works better not just in the average case, but also in the worst case? Lukas: Yeah, it's interesting. Well, I had a couple things came up for me reading that paper. When I started my career, I think I was just sort of on the tail end of ACL papers where it seemed like they would just cherry pick some examples where it worked or it didn't, and it just seemed ridiculous.
我是说,谈 AI 进展速度的人特别爱讲那些排行榜变化得多快、SOTA(state-of-the-art)在各种基准上被刷得多高多快。我总是想:“是啊,那又怎样?”从科学的角度更好地理解这个世界,或者构建一种不只是在平均情况下、而且在最坏情况下也能更好工作的技术——这到底意味着什么?Lukas:是啊,挺有意思的。读那篇论文的时候,我脑子里冒出了好几件事。我刚开始做这行的时候,我觉得自己大概赶上了那个时代的尾巴——那时候的 ACL 论文,好像就是挑几个奏效或不奏效的例子出来讲,我当时就觉得这太荒唐了。
便签引用
1:00:05
I remember they had early benchmarks and people would have lower accuracy than just guessing the most common case or something, which you could argue that's better, and people did, but that just seemed a little ridiculous to me. I remember this anecdote from your class about...I think it was Noam Chomsky saying that, "Oh, moms don't teach kids language," but actually they do, and it's just like no one bothered to check. So, it's kind of maddening, and I think I appreciated benchmarks from that. But then your recommendations are not only reasonable, I think in companies, a lot of it is standard best practice. I don't think you would just release a new model without trying it and getting a flavor for where it works and where it doesn't. You wouldn't just be like, "Oh, we took 10% of the data and held it out, let's ship it."
我记得当时有一些早期的基准测试,有人的准确率比直接猜最常见的那个类别还低之类的,你可以争辩说那样更好,也确实有人这么争辩,但我当时就觉得有点离谱。我记得你课上讲过一个轶事……好像是 Noam Chomsky 说"哦,妈妈们不会教孩子语言",但实际上她们会,只是从来没人费心去核实一下。所以这挺让人抓狂的,也因此我挺欣赏基准测试的。但另一方面,你的那些建议不仅合理,我觉得在公司里,很多本来就是标准的最佳实践。我不觉得你会连试都不试、连它在哪儿管用、哪儿不管用都没摸清,就直接把一个新模型发出去。你不会说"哦,我们留出了 10% 的数据,发吧"。
便签引用
15Bender 规则:写明你研究的语言
1:01:01
It does seem like that's actually one case where you see it more in companies than in sort of academic literature, probably because it's easier to look at one number and be like, "Hey, we beat it." But clearly, that's flawed. So, anyway, I thought that was a great paper with really good suggestions that I think everyone should definitely follow. I also want to make sure we got to the last paper that we talked about, which is cool, because I just want to make sure people know. What is the Bender Rule?⁴ And why is it important?
这好像确实是那种在公司里比在学术文献里更常见的做法,大概是因为盯着一个数字说"嘿,我们刷赢了"要容易得多。但这显然是有问题的。总之,我觉得那是一篇很棒的论文,里面的建议非常好,我觉得每个人都该照做。我还想确保我们能聊到我们提过的最后一篇论文,那篇很酷,因为我想让大家都知道。什么是 Bender 规则?⁴ 它为什么重要?
便签引用
1:01:29
Emily: So, Bender Rule or the #BenderRule- Lukas: Is it #BenderRule? Emily: Yeah, well, it's both. Lukas: Say what it is first, and then I have some questions about best practice. Emily: Yeah. It is itself a best practice, which says that you should always state the name of the language you're working on, even if it's just English. This is a soapbox that I've been carrying around and periodically climbing up on since about 2009, where I saw a lot of that pre-neural statistical NLP work saying, basically, "Look Ma, no linguistics," and claiming that systems were language-independent because there was no linguistic knowledge hard-coded.
Emily:Bender 规则,或者说 #BenderRule——Lukas:是 #BenderRule 吗?Emily:对,其实两种说法都有。Lukas:先说说它是什么,然后我有几个关于最佳实践的问题。Emily:好。它本身就是一条最佳实践,它说的是:你应该始终写明你研究的是哪种语言,哪怕只是英语。这是我从大概 2009 年起就一直随身带着、时不时爬上去演讲一番的"肥皂箱",当时我看到很多神经网络之前的统计 NLP 工作,基本上都在说"看啊,不用语言学",并且声称系统是语言无关的,因为里面没有硬编码任何语言学知识。
便签引用
1:02:10
And these supposedly language-independent systems were mostly tested on English. You also see a lot of work people will publish a paper on machine reading or paper on sentiment analysis, and in fact, no, it's a paper on machine reading of English and sentiment analysis on English text. Flip side is if someone's working on Cherokee, or Thai, or Chinese, or Italian, then that work...it's harder to get it accepted to the research conferences because it is deemed language-specific, where work on English is somehow general.
而这些号称语言无关的系统,绝大多数只在英语上做过测试。你也会看到很多人发论文讲机器阅读、讲情感分析,可实际上不是,那其实是一篇关于英语机器阅读、关于英语文本情感分析的论文。反过来说,如果有人研究的是切罗基语、泰语、汉语或意大利语,那这类工作……就更难被研究会议接收,因为它被认为是"语言特定的",而做英语的工作却莫名其妙地被当成"通用的"。
便签引用
1:02:43
That's a big problem for the science, it's a big problem for getting to technology that actually works across languages. I've been sort of going around pestering people to actually test cross linguistically and to name the language they're working on. In 2019, like three or four people — and this is in that piece on The Gradient, I have their names listed — sort of referred to this practice as the Bender Rule. I didn't name that, but once it was named I ran with it. Part of it is it's kind of a face threatening question to ask. If someone's written something about machine reading and I walk up and I say, "What language?", it's a stupid question to ask because it's obviously English. So, it's face threatening to me. And it's also a little bit rude to them, to ask this question that says you should have said. I don't mind people blaming that on me. Part of the reason I ran with the hashtag is, if someone wants to go ask this question and they feel like it's a sort of a silly question to ask, they can pin it on me, and I'm happy to
这对科学来说是个大问题,对做出真正能跨语言工作的技术来说也是个大问题。我一直到处去烦别人,让他们真正做跨语言的测试,并且写明自己研究的是什么语言。2019 年,有三四个人——这个在《The Gradient》那篇文章里,我把他们的名字都列出来了——把这个做法叫作 Bender 规则。名字不是我起的,但一旦有了这个名字,我就顺势用起来了。部分原因是,这是个会让人有点下不来台的问题。如果有人写了篇关于机器阅读的东西,我走过去问一句"哪种语言?",这是个蠢问题,因为显然是英语。所以,对我自己来说这有点丢面子。对对方也有点不礼貌,因为这个问题等于在说"你本该写明的"。我不介意大家把这笔账算在我头上。我之所以顺势用起这个话题标签,一部分原因就是:如果有人想问这个问题,又觉得这问题问出来有点傻,他们可以把责任推给我,我很乐意
便签引用
16如果 NLP 不从英语起步
1:03:45
lend my name to that. Lukas: I see. Nice. I guess this is a hard question, but this is what kind of comes to mind for me, it's like, "Wow, English is so specific and probably has all these kind of idiosyncrasies." How do you think NLP be might be different if it started in like Thai or Cherokee or something or English just happened to be...I mean, English must be unusual in all these ways, right? Are there characteristics of English that are unusual and the world could have gone a different way? Emily: Yeah, absolutely. Actually, in that paper, I list out a bunch of them. One thing is English is a spoken language, not a signed language. If we had started NLP with American Sign Language or another signed language, it would have been very different, right? Lukas: Clearly.
把自己的名字借出去用。Lukas:明白了,不错。我猜这是个很难的问题,但我脑子里冒出来的想法是,"哇,英语其实特别特殊,大概有一堆自己的怪癖。"你觉得如果 NLP 一开始是从泰语、切罗基语之类的语言起步,会有什么不同?还是说英语只是恰好……我是说,英语在这么多方面肯定都很不寻常,对吧?英语有哪些不寻常的特点,让这个领域本可能走上另一条路?Emily:是的,绝对有。其实在那篇文章里,我列了一大堆。一点是,英语是一种口语,不是手语。如果我们当初是从美国手语或别的手语开始做 NLP,那会非常不一样,对吧?Lukas:显然是。
便签引用
1:04:32
Emily: Yeah. So, that's one big choice point. Another thing is that English has a very well-established and standardized writing system. Many of the world's languages don't have a writing system at all, and many of them that do don't have the degree of standardization that English does. Also, many languages will have a lot more code switching going on, on average, than English does. Lukas: What is code switching? Emily: Code switching is when you use multiple languages in the same conversation, sometimes even the same sentence.
Emily:对。所以这是一个重大的岔路口。另一点是,英语有一套非常成熟、非常标准化的书写系统。世界上很多语言根本没有书写系统,而那些有书写系统的,很多也没有英语这种标准化程度。另外,平均而言,很多语言里的语码转换要比英语多得多。Lukas:什么是语码转换?Emily:语码转换就是你在同一段对话里、有时甚至在同一个句子里使用多种语言。
便签引用
1:05:03
That happens a lot in communities where there's a lot of bilingualism or multilingualism. So, if you and I...well, you also speak Nihongo, right? When you studied kanji, what was your favorite way to benkyou them? I am not a fluent code switcher, so that was really awkward and stupid sounding, but to illustrate the point. Lukas: I remember actually when...yeah, I know, I have experienced that for sure. Emily: Certainly, English is involved in a lot of code switching. But there's also lots and lots of monolingual English data and when you go into social media data for Indian languages, for example, enormous amounts of that are code switched with English. And so, there's a whole range of interesting technical challenges that come up there.
这在双语或多语现象很普遍的社群里非常常见。比如你和我……对了,你也会说 Nihongo(日语),对吧?你学 kanji(汉字)的时候,最喜欢用什么方式 benkyou(学习)它们?我并不是个熟练的语码转换者,所以刚才那句听起来又别扭又蠢,但意思到了。Lukas:我记得当时……对,我懂,我确实体验过那种感觉。Emily:当然,英语参与了大量的语码转换。但同时也存在海量的单语英语数据;而当你去看印度诸语言的社交媒体数据,比如说,里面有极其大量的内容是和英语混着用的。因此,那里会冒出一大堆有意思的技术挑战。
便签引用
1:05:50
We live in a world where the first digital setups were sort of accommodated...lower ASCII, most conveniently, English all fits in lower ASCII. English has relatively fixed word order. We have a relatively low...relatively simple morphology. Any given word that shows up is only going to show up in a few different forms. Compare that to Turkish where you can get like, I think, millions of inflected forms of the same root, and so that that changes the way you handle data sparsity and what data sparsity looks like. Our orthography is a mess. Someone was just asking on Twitter, "How come we do grapheme to phoneme prediction but not phoneme to grapheme prediction?
我们生活的这个世界里,最早的数字化设置是围绕着……低位 ASCII 来迁就的,而英语正好完全能装进低位 ASCII 里。英语的语序相对固定。我们的形态变化相对……相对简单。任何一个词出现时,也就那么几种形式。对比一下土耳其语,同一个词根我记得能有上百万种屈折形式,这就改变了你处理数据稀疏性的方式,也改变了数据稀疏性本身的样子。而我们的正字法则是一团糟。刚刚还有人在推特上问:"为什么我们做字素到音素的预测,却不做音素到字素的预测?"
便签引用
1:06:40
So, grapheme to phoneme is, "Given a letter, what's the likely sound?", and that's an important component of text-to-speech systems when you hit an out of vocabulary word. Phoneme to grapheme would be, "Given a sound, what's the likely letter?", and that's not a typical task. I wonder to what extent that's true because of English's opaque and chaotic writing system. Lukas: Right. Sounds like an impossible task. Emily: Yeah, exactly. But if you were to look at...Japanese, setting aside the kanji, if you just try to transcribe Japanese in kana, that's way more straightforward. Spanish also has a very transparent and consistent grapheme to phoneme mapping in both directions. So, down to things like that, the properties of a writing system for English. English likes to use white space between words and sentence-final punctuation. These are things that we sort of take as given, that it's easy to tokenize into sentences and words, that just aren't going to be true in other languages.
字素到音素就是"给定一个字母,最可能的读音是什么",这是文本转语音系统在遇到词表外单词时的一个重要环节。而音素到字素则是"给定一个读音,最可能的字母是什么",这并不是个常见任务。我很好奇,这在多大程度上是因为英语那套不透明又混乱的书写系统。Lukas:对。听起来是个不可能完成的任务。Emily:对,正是。但如果你看看日语——把汉字放在一边,如果你只是用假名来转写日语,那就直接多了。西班牙语的字素—音素对应在两个方向上也都非常透明、非常一致。所以细到这种程度,英语书写系统的种种性质都会有影响。英语习惯在词之间加空格、在句末加标点。这些我们默认为理所当然的东西——比如很容易切分出句子和词——在别的语言里根本就不成立。
便签引用
1:07:37
So, I don't know. I couldn't tell you what NLP would look like. I can just sort of tell you sort of where the points of divergence might be. Lukas: No, those are fun. I mean, definitely. I mean, I don't know. Those differences are so interesting. Emily: Well, you voluntarily took a linguistics class, so I'm not surprised. Lukas: Well, I just feel like linguistics is so cool. I mean, as an outsider just because if you don't know it, then it's really eye opening to just...because you swim in it, to sort of see all these patterns that I never would have noticed.
所以我也说不好。我没法告诉你 NLP 会长成什么样,我只能大致告诉你分岔点可能在哪儿。Lukas:不,这些都很有意思。真的。我是说,我也不知道。这些差异实在太有意思了。Emily:嗯,你可是自愿选了一门语言学课的,所以我一点都不意外。Lukas:我就是觉得语言学太酷了。作为一个外行,正因为你不懂它,它才特别让人开眼……因为你就浸泡在语言里,然后突然看到那么多我本来永远不会注意到的规律。
便签引用
1:08:13
And I feel like especially...like phonetics is probably the most deep, where you're just like, "Oh, my god, those two sounds are different?" I would just never, never have noticed that. It's so easy to do the thought experiment and realize you're wrong, that it's just...I love that stuff
我觉得尤其是……比如语音学大概是最深的,你会忍不住惊呼,"我的天,这两个音居然是不一样的?"我本来永远、永远不会注意到。而且做个思想实验就能发现自己错了,太容易了——我特别喜欢这类东西。
便签引用
1:08:36
I feel like most of my early work was in parsing Japanese in different ways. I do remember...I guess it didn't seem like that was an impediment to publishing, but it was surprising that there was so little work on it for how necessary of a task it would be to deal with it. In my first job, it was mostly processing Japanese language stuff, and it was striking how little research there was defined on the topic. I felt like there was a sort of more institutional knowledge inside of companies than literature on it. Emily: What happened in the research community is well, that kind of parsing problem is "solved" because people had made a certain progress on it for English, and that was mistaken as the problem in general being solved.
我觉得我早期的工作大部分都是在用各种方式解析日语。我确实记得……我猜这在发表上倒不算是个障碍,但让我意外的是,考虑到处理它是多么必要的一项任务,相关的工作却少得可怜。我的第一份工作主要就是处理日语的东西,而这方面的研究之少令人吃惊。我感觉公司内部积累的经验知识反而比文献还多。Emily:研究界发生的事情是这样的:那类解析问题被认为"已解决",因为大家在英语上取得了一定进展,然后这就被误认为是这个问题本身已经被解决了。
便签引用
1:09:23
So, what's new here? Well, this is for Japanese. That's new. But it's actually hard to get people to see that. My goal with what got called the Bender Rule is to say, "Okay, let's keep English in its place," and say, "When I've done this for English, I need to say that it's for English to hold room for the other work on other languages," which is also really important and novel and valuable. We'll see. If we periodically go through...different folks in the field go through and count how many papers in an ACL conference actually work on different languages and actually say what language they work on, and it's not changing as fast as I'd like. But there's some really good developments.
那这里有什么新东西呢?嗯,这是针对日语的。这就是新的。但要让人们看到这一点其实很难。我提出后来被叫作 Bender规则的这件事,目标就是说:"好,让英语待在它该待的位置上",并且说"当我在英语上做了这件事,我就必须说明这是针对英语的,好为其他语言上的工作留出空间"——那些工作同样非常重要、新颖、有价值。走着瞧吧。如果我们时不时地去统计一下——领域里不同的人会去数——一届 ACL 会议上有多少论文真正在研究不同的语言、并且真的写明了自己研究的是什么语言,变化速度并没有我希望的那么快。但也有一些很好的进展。
便签引用
17跨语言迁移与低资源社群
1:10:03
The Universal Dependencies project has produced treebanks for many, many languages, and that has spurred a whole bunch of very crosslinguistic work, which is exciting. Lukas: What do you think about... I mean, some of the most evocative work feels like building language models across all the languages or translation models that can kind of use pairs of languages in interesting ways, where you have more data to help with ones with less data. Do you think that's a fruitful direction? Or does that...do you think that sort of encodes our biases somehow in the way it works? Emily: I mean, it's certainly interesting, and to the extent that we're relying on these massive data-hungry things, where languages just don't have that much data, seeing what we can do based on transfer from the bigger languages is an interesting and valuable way to go. I think the interesting questions to ask would be, "To what extent does this impose the conceptualization of the world encoded in English on to the results in other languages?", and "What follows from that? What are the risks?"
Universal Dependencies 项目为非常非常多的语言构建了树库,这催生了一大批很跨语言的研究,这很让人兴奋。Lukas:你怎么看……我是说,最让人心动的一些工作,好像是构建覆盖所有语言的语言模型,或者能以有趣的方式利用语言对的翻译模型——用数据多的语言去帮助数据少的语言。你觉得这是个有前景的方向吗?还是说……你觉得它会不会以某种方式把我们的偏见编码进它的运作方式里?Emily:我是说,这当然很有意思,而且既然我们依赖的是这些极其吃数据的东西,而有些语言根本就没有那么多数据,那么看看能不能靠从大语言迁移过来做点什么,这是一条有意思也有价值的路。我觉得值得问的问题是:"这在多大程度上把英语中编码的那套世界观强加到了其他语言的结果上?"以及"由此会带来什么?风险是什么?"
便签引用
1:11:06
How does that compare to, "Well, but if we just do monolingual, we can only get this far, so, we'll take those risks. We'll figure out how to mitigate them." That kind of work I think is important. It's also really, really important to know that you are working with genuine data in the low resource languages. There was this thing where it came out that — I think it was Scots — the entire Scots Wikipedia was written by one person who doesn't speak Scots. Wikipedia is this really important data source in NLP, so any NLP system that claims to be doing something for Scots just isn't.
再拿它跟另一种情况比较:"可如果我们只做单语,就只能走到这一步,所以我们愿意承担那些风险,我们会想办法缓解它们。"我觉得这类工作很重要。同样非常非常重要的一点是,要确认你在低资源语言上用的是真实的数据。之前曝出过一件事——我记得是苏格兰语(Scots)——整个苏格兰语维基百科都是一个不会说这门语言的人写的。维基百科在 NLP 里是极其重要的数据来源,所以任何声称自己在为苏格兰语做点什么的 NLP 系统,其实根本不是。
便签引用
1:11:37
A fantastic model in that regard is a research collective called Masakhane, which is a continent-spanning research initiative in Africa towards doing participatory research to create language resources for African languages. They've done really interesting work on how to build up the community so that people can come contribute as translators, not machine translation specialists, but people actually translating language. There's a really cool paper that came out in I think findings of EMNLP last year describing Masakhane project.
这方面有一个极好的榜样,是一个叫 Masakhane 的研究共同体,它是一个横跨非洲大陆的研究倡议,通过参与式研究来为非洲语言创建语言资源。他们在如何把社群建设起来、让人们能以译者身份来贡献这件事上,做了非常有意思的工作——这些人不是机器翻译专家,而是真正在做翻译的人。我记得去年 EMNLP 的 Findings 里有一篇很酷的论文,介绍了 Masakhane 项目。
便签引用
1:12:09
That kind of work of, if you're going to work with low resource languages, being sure to connect with the community. Who would be the people using the technology, then you could find out, "Okay, well, what are the concerns? To what extent do you want to bring in what we can do from using the larger resource languages?" versus "Would you rather stay monolingual and see where we can go and hear from the community and involve the community in the research? I think Masakhane is a great model of that. Lukas: Cool. Well, that seems like a good place to end. We're way over time and you've been really generous. Thank you so much. I really enjoyed talking to you. Emily: Yeah. Likewise, thank you. I can go on and on. So, I appreciate the chance to do so. Lukas: If you're enjoying these interviews and you want to learn more, please click on the link to the show notes in the description where you can find links to all the papers that are mentioned, supplemental material, and a transcription that we
这类工作的意思是:如果你要做低资源语言,一定要和社群建立联系。谁会是这项技术的使用者?然后你才能弄清楚,"好,那大家的顾虑是什么?你们希望在多大程度上引入我们从高资源语言那里能拿到的东西?"还是说"你们更愿意保持单语,看看我们能走多远"——去听社群的声音,让社群参与到研究中来。我觉得 Masakhane 是这方面的绝佳榜样。Lukas:太好了。嗯,这好像是个很好的收尾点。我们已经严重超时了,你也非常慷慨。非常感谢你。跟你聊天我真的很享受。Emily:我也是,谢谢你。我可以一直讲下去,所以很感谢有这个机会。Lukas:如果你喜欢这些访谈,想了解更多,请点击简介中指向节目笔记的链接,那里有提到的所有论文的链接、补充材料,以及一份我们
便签引用
1:13:02
work really hard to produce. Check it out.
下了很大功夫做出来的文字稿。去看看吧。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

语言学家 Emily Bender 借《随机鹦鹉》一文的 Google 风波切入,论证语言模型只见"形式"不见"意义"、基准测试被滥用为"通用理解"的证据,并呼吁研究者明确所研究的语言(#BenderRule),而非把英语默认为"通用"。

核心要点

  • 《随机鹦鹉》是一篇综述而非实验论文,却引发 Google 解雇两位伦理 AI 负责人。 起因是 Timnit Gebru 在 Twitter 私信 Bender 询问大模型风险,Bender 列出五六条担忧后成了论文提纲;七位作者一个月内在 Overleaf 远程合作完成,投稿 FAccT 2021。论文通过了 Google 内部发表审批,但 11 月底 Google 要求撤稿或署名撤回且不说明原因。Gebru 被解雇,Margaret Mitchell 数月后也被解雇,因此第四作者署名为 "Shmargaret Shmitchell"。仅 Bender 网站上的 bitly 链接就被下载超一万次。
  • Google 公开的反对理由站不住脚。 Google 称论文"未引用缓解这些问题的相关工作",但从未指明该引用哪些,而论文实际引用了此类工作。作者们原本预期被冒犯的会是 OpenAI(GPT-3 是文中的贯穿案例),没料到是 Google 内部。
  • "伦理"讨论应针对制度与激励,而非给个人贴标签。 多数从业者受股东价值最大化等系统性约束,难以对正在赚钱的东西踩刹车。Bender 希望 NLP 借鉴其他工程领域上线前的容差测试与认证,而非 AI 炒作文化;Model Cards 是朝这个方向的尝试。"越大越好"的路径还会排除数据量不足的语言社区和小型研究团队。
  • 偏见不只来自数据,模型本身也在放大。 三个实证案例:Safiya Noble 记录搜索"black girls"早期返回色情内容,根源是广告经济允许购买身份词;Latanya Sweeney 2013 年发现非裔风格姓名的搜索广告暗示犯罪记录;Robyn Speer 发现用通用词向量做 Yelp 评论打星预测时,系统性低估墨西哥餐厅评分,因为网络上的移民话语让 "Mexican" 与负面情感词共现。Yelp 数据本身并无此偏差,偏差是模型组件带入的。
  • 去偏无法彻底,但可以做得更好。 完全无偏的数据集不存在。可行路径:使用精心策划的数据、针对已知偏见的去偏技术(前提是你得知道要找什么)、按具体用例梳理利益相关者和失败模式,再决定测什么。
  • 只学"形式"的语言模型无法获得"意义"。 语言模型的训练任务是预测文本中的词,看到的只是字符序列。母语者无法"关掉"意义,因此看英语文本不是好的类比;看切罗基语音节文字才更接近模型的处境。词向量捕捉的是分布模式,是有用的工程组件,但不等于理解。
  • 图灵测试失效的原因是人类太善于"补全意义"。 人会主动为语句构建使其合理的语境,所以人不适合当图灵测试的裁判。章鱼思想实验的关键:假设章鱼超级聪明、时间和记忆无限,它监听两位漂流者的莫尔斯电报后能完美模仿"下一句最可能是什么",但当 A 问"椰子投石机造出来怎么样了"或"熊在攻击我,只有两根木棍怎么办"时,章鱼无能为力,因为它从未有过椰子、绳子或熊的接地经验。GPT-2 对熊的问题给出的回答话题相关却毫无帮助。
  • 人类能通过语言学习未经历之事,前提是已掌握语言系统。 儿童习得语言依赖与照料者的"联合注意",把词映射到交际意图上;章鱼和语言模型没有这一步。被扔进只有泰语书的图书馆,突破口只能是带图的百科或《好奇的乔治》泰译本这类外部锚点。Bender 并非否认自然语言理解可行,而是主张需要视觉接地、知识库、传感器等形式之外的输入。
  • 基准测试的问题在于滥用,不在于基准本身。 "地图不是领土":微软曾宣传其系统在 GLUE 上超过人类就等于"比人更懂英语",属于严重过度宣称,且让"从微笑猜政治倾向"这类荒谬产品显得可信。补充手段包括:测试套件、系统性审计(如 Buolamwini/Gebru/Raji 按性别与肤色网格审计人脸识别)、对抗测试(2017 年 Build It Break It 任务用最小对找出系统失效点)、错误分析。依赖语言模型的系统普遍处理不好否定词,漏掉一个 "not" 的后果远比漏掉冠词严重。
  • #BenderRule:永远写明你研究的语言,哪怕是英语。 自 2009 年起她观察到"语言无关"系统几乎只在英语上测试,而做泰语、切罗基语的工作被视为"语言特定"难以发表。英语的特殊性包括:是口语而非手语、有高度标准化的书写系统、代码转换少、词序固定、形态简单(对比土耳其语一个词根有数百万屈折形式)、用空格分词、正字法混乱到"音素到字素"预测不成为任务。

结论与值得注意的细节

  • Bender 预测语言模型会继续作为语音识别、机器翻译等任务的组件存在,但当应用需要对输出的文字负责时,随机鹦鹉会被"去中心化",退回到候选排序或提供词向量的角色。
  • 她主张转向能从更小、可策划、可更新的数据集中学习的算法,理由同时包括偏见治理、碳足迹和移动端算力约束,认为这是环境与技术的双赢。
  • 低资源语言的数据真实性需要核查:苏格兰语维基百科几乎全由一个不会苏格兰语的人写成,任何声称支持苏格兰语的 NLP 系统实际上并不支持。
  • Masakhane 是她推荐的范本:横跨非洲大陆的参与式研究集体,让社区成员作为译者贡献语料,并由社区决定是否借用高资源语言的迁移。
  • Universal Dependencies 项目为众多语言产出树库,推动了跨语言研究,但 ACL 论文中标明所用语言的比例改善得比她期望的慢。
  • 主持人 Lukas 提到自己早期做日语分析工作时发现文献极少,Bender 的解释是:英语上取得进展后,问题就被误认为"已解决",做日语被视作"没有新意"。
核心句型 · 10
1. It's really important to distinguish between X as opposed to Y
“It's really important to distinguish between the word as a sequence of characters as opposed to words in the sense of a pairing of form and meaning.”
用于开篇立论、划清两个易混概念。as opposed to 比 and 更强调对立。仿写:It's important to distinguish between correlation as opposed to causation.
2. Little did we know.
“There was a while. that we thought the thing about this paper would be it was the one with an emoji in the title. Little did we know.”
否定副词前置引起倒装,独立成句表示「当时哪里想得到」,常用于叙事转折前。语气略带自嘲。仿写:We thought it was a minor bug. Little did we know.
3. off the top of my head, here's …
“But off the top of my head, here's five or six things that we can be worried about.”
表示「不经深思、随口列举」,用于降低承诺强度、给出初步清单。口语和邮件都常见。仿写:Off the top of my head, here are three options.
4. It's not so much X (that) …, but Y
“It's not so much benchmarks I have a problem, but the way that they're used.”
纠正听者的焦点:问题不在 X 本身,而在 Y。适合精确表达立场。仿写:It's not so much the tool I object to, but how it's marketed.
5. To the extent that …, it's because …
“To the extent that O gets away with that, it's because A is willing to make sense of those utterances.”
「在……成立的范围内,原因是……」,用于有限承认对方观点同时给出解释。学术论证常用。仿写:To the extent that it works, it's because the data is clean.
6. I'm not saying that X. I'm saying that Y.
“I'm not saying that natural language understanding is impossible and not something to work on. I'm saying that language modeling is not natural language understanding.”
先排除误读,再重申主张。两句并列、结构对称,力度强。适合辩论中被误解时澄清。
7. you can't help but + V
“Because you can't help but get the meaning part when you're looking at it.”
「忍不住、不由自主」,强调某种反应是自动的、无法抑制的。注意 but 后接动词原形。仿写:You can't help but notice the pattern.
8. Not "How well … on average?" but "How well … for this case?"
“Not "How well does it work on average?" but "How well does it work for this case, and that case, and that case?"”
用两个引号内的问句做对比,把抽象主张变成可操作的提问方式。and that case 的重复制造节奏感。适合讲方法论。
9. which is one of these things that …
“Negation, which is one of these things that's very important to the meaning, but tends to be a short word or subword”
非限定性定语从句里嵌套 one of these things that,用于把某个现象归入一类并顺带解释。口语中自然的补充说明方式。
10. boy, did we …
“And boy, did we put a lot of polish on it between the submission version and the camera-ready”
感叹词 boy 后接倒装,表达强烈程度,等于「我们可真是……」。非正式,口语专用。仿写:Boy, did that take longer than expected.
词汇精讲 · 143 · 按出现顺序
pairing /ˈperɪŋ/ n. 0:01
配对;此处指形式与意义的对应关系
stochastic /stəˈkæstɪk/ adj. 0:42
随机的、概率性的(统计学术语)
notable /ˈnoʊtəbl/ adj. 0:42
引人注目的、值得注意的
Little did we know phr. 1:40
我们当时哪里知道(倒装强调,表示事后才发现)
push /pʊʃ/ n. 1:40
(集体的)推动、风潮
off the top of my head phr. 2:27
不假思索、随口想到
anticipated /ænˈtɪsɪpeɪtɪd/ v. 3:38
预料到、预期
secondhand /ˌsekəndˈhænd/ adj. 3:38
间接得知的、二手的
retract /rɪˈtrækt/ v. 4:40
撤回(论文、声明)
pushed back phr. 5:16
提出反对、抵制
neologism /niˈɑːlədʒɪzəm/ n. 5:16
新造词
goodwill /ˌɡʊdˈwɪl/ n. 6:07
善意、好感;商誉
sheds a light on phr. 6:07
揭示、照亮(某个问题)
weathering /ˈweðərɪŋ/ v. 6:37
经受住、挺过(风波、困难)
camera-ready /ˌkæmərə ˈredi/ adj. 6:37
(论文)最终定稿版、可付印版
preprint /ˈpriːprɪnt/ n. 6:37
预印本(未经正式发表的论文版本)
out of scale phr. 7:22
远超比例、量级不相称
gross misstep phr. 7:22
严重失策(gross 表「严重的、粗大的」)
incendiary /ɪnˈsendieri/ adj. 8:10
煽动性的、易引发争端的
mitigate /ˈmɪtɪɡeɪt/ v. 8:10
缓解、减轻
ruffling some feathers phr. 8:47
惹恼一些人、触怒他人
blowing up phr. 9:44
炸掉、毁掉
overstatement /ˌoʊvərˈsteɪtmənt/ n. 9:44
夸大其词
identify with phr. 10:29
将自我认同寄托于、与……产生认同
caricature /ˈkærɪkətʃʊr/ n. 10:29
漫画式夸张形象
tycoon /taɪˈkuːn/ n. 10:29
大亨、巨头
put on the brakes phr. 11:40
踩刹车、叫停
nuanced /ˈnuːɑːnst/ adj. 11:40
细致入微的、有微妙差别的
for all its fuss phr. 12:44
尽管围绕它有诸多喧嚣(for all 表「尽管」)
failure modes phr. 13:59
失效模式(工程术语,指系统出错的具体方式)
tolerances /ˈtɑːlərənsɪz/ n. 13:59
(工程)容差、公差
certify /ˈsɜːrtɪfaɪ/ v. 14:37
认证、检定
rampant /ˈræmpənt/ adj. 14:37
泛滥的、失控蔓延的
well-scoped adj. 15:23
范围界定清晰的
amass /əˈmæs/ v. 15:23
积聚、大量积累
stifles /ˈstaɪflz/ v. 16:24
扼杀、压制
pervasive /pərˈveɪsɪv/ adj. 16:24
无处不在的、渗透各处的
running example phr. 17:27
贯穿全文反复使用的例子
piecemeal /ˈpiːsmiːl/ adv. 18:17
零敲碎打地、逐个地
after-the-fact adj. 18:17
事后的、马后炮式的
latches on to phr. 18:17
抓住、攀附上(并不放)
amplifies /ˈæmplɪfaɪz/ v. 18:17
放大、加剧
configuration /kənˌfɪɡjəˈreɪʃn/ n. 19:22
配置、具体情形
so-and-so /ˈsoʊ ən soʊ/ n. 20:17
某某人(泛指不具名的人)
in-domain adj. 21:18
领域内的(与目标任务同领域的数据)
akin to phr. 22:02
类似于、近似
toxic /ˈtɑːksɪk/ adj. 22:02
有毒的;此处指恶毒、有害的言论
co-occurrences /ˌkoʊ əˈkɜːrənsɪz/ n. 22:44
共现(两个词在同一语境中同时出现)
curated /ˈkjʊreɪtɪd/ adj. 24:06
精心挑选、经过整理的
reputable /ˈrepjətəbl/ adj. 24:06
声誉良好的、可信的
stakeholders /ˈsteɪkhoʊldərz/ n. 25:03
利益相关方
adverse impacts phr. 25:03
不利影响
in the interest of time phr. 26:01
为了节省时间
put to rest phr. 26:34
平息(争论)、彻底了结
pin down phr. 26:59
明确界定、确定下来
finely fitted phr. 28:39
精细拟合的
sticks /stɪks/ v. 29:14
(按键)卡住、粘住
syllabary /ˈsɪləberi/ n. 30:11
音节文字(每个字符表一个音节的书写系统)
can't help but phr. 30:11
忍不住、不由自主地
on the verge of phr. 30:42
濒临、即将
subtle /ˈsʌtl/ adj. 30:42
细微的、微妙的
foundational /faʊnˈdeɪʃənl/ adj. 31:52
奠基性的、基础性的
well positioned phr. 31:52
处于有利位置、适合(做某事)
coherent /koʊˈhɪrənt/ adj. 32:47
连贯的、条理清楚的
snarky /ˈsnɑːrki/ adj. 33:36
讥讽的、刻薄挖苦的
bulldozing /ˈbʊldoʊzɪŋ/ v. 33:36
强行推过、蛮力碾压
evocative /ɪˈvɑːkətɪv/ adj. 33:36
引人联想的、富于画面感的
posit /ˈpɑːzɪt/ v. 35:17
假定、设定为前提
stranded /ˈstrændɪd/ adj. 35:48
被困、搁浅的
uninhabited /ˌʌnɪnˈhæbɪtɪd/ adj. 35:48
无人居住的
offstage /ˌɔːfˈsteɪdʒ/ adv. 35:48
幕后、不予交代(戏剧比喻)
taps into phr. 36:22
接入、窃听(线路);利用
On reflection phr. 36:58
回过头想想、经过思考
subterfuge /ˈsʌbtərfjuːdʒ/ n. 37:32
欺骗手段、诡计
pleasantries /ˈplezntriz/ n. 37:32
寒暄、客套话
inconsequential /ɪnˌkɑːnsɪˈkwenʃl/ adj. 37:32
无关紧要的
semblance /ˈsembləns/ n. 37:32
表象、外表上的样子
utterances /ˈʌtərənsɪz/ n. 38:27
话语、说出的话(语言学术语)
utterly /ˈʌtərli/ adv. 38:27
完全地、彻底地
communicative intent phr. 39:14
交流意图(说话者想传达的意思)
grounding /ˈɡraʊndɪŋ/ n. 39:51
接地、落地(将符号与真实世界关联)
Conversely /ˈkɑːnvɜːrsli/ adv. 41:05
反过来说
well-formed adj. 41:05
合乎语法的、格式正确的
standing meanings phr. 41:05
固定意义、约定俗成的词义
distributional /ˌdɪstrɪˈbjuːʃənl/ adj. 41:49
分布上的(指词在语料中的出现规律)
joint attention phr. 42:23
共同注意(儿童与照料者同时注意同一事物)
caregivers /ˈkerɡɪvərz/ n. 42:23
照料者
tremendously /trəˈmendəsli/ adv. 43:10
极其、非常
embodied /ɪmˈbɑːdid/ adj. 43:42
具身的(有身体、能与环境交互的)
oodles /ˈuːdlz/ n. 44:28
大量、一大堆(口语)
resort to phr. 45:12
诉诸、不得已而采用
morphology /mɔːrˈfɑːlədʒi/ n. 45:42
形态学、词形变化
function words phr. 45:42
功能词、虚词(冠词、介词等)
foothold /ˈfʊthoʊld/ n. 46:13
立足点
longstanding /ˌlɔːŋˈstændɪŋ/ adj. 46:48
由来已久的
stringent /ˈstrɪndʒənt/ adj. 46:48
严格的、严苛的
fail safes phr. 46:48
故障保护、安全兜底机制
accountability /əˌkaʊntəˈbɪləti/ n. 47:50
问责、责任担当
fragile /ˈfrædʒl/ adj. 47:50
脆弱的、易失效的
de-center v. 47:50
去中心化、从核心位置挪开
make more of phr. 48:22
更充分地利用
carbon footprint phr. 48:59
碳足迹
leaner /ˈliːnər/ adj. 49:38
更精简的、更节省资源的
gold standard phr. 50:36
金标准(人工标注的正确答案)
overclaiming /ˌoʊvərˈkleɪmɪŋ/ v. 50:36
夸大主张、过度宣称
the map is not the territory phr. 51:28
地图不是疆域(模型或指标不等于现实本身)
implausible /ɪmˈplɔːzəbl/ adj. 52:14
不可信的、难以置信的
political affiliation phr. 52:14
政治立场、党派归属
held out phr. 53:01
(数据)留出不用于训练
sanity check phr. 54:36
理智检验、基本合理性检查
naive baseline phr. 54:36
朴素基线(用于对照的最简单方法)
head-to-head adv. 54:36
正面交锋地、一对一比较
auditing /ˈɔːdɪtɪŋ/ n. 55:23
审计、系统性核查
contours /ˈkɑːntʊrz/ n. 56:03
轮廓、边界
freebies /ˈfriːbiz/ n. 56:41
白送的东西;此处指送分题
minimal pairs phr. 57:22
最小对立对(仅一处差异的两个例子)
negation /nɪˈɡeɪʃn/ n. 58:07
否定(语法范畴)
tripping us up phr. 58:43
把我们绊倒、使出错
pile-on /ˈpaɪl ɑːn/ n. 58:43
一拥而上的堆叠、扎堆
state-of-the-art adj. 59:17
最先进的(缩写 SOTA)
cherry pick phr. 59:17
挑选对自己有利的例子
maddening /ˈmædnɪŋ/ adj. 1:00:05
令人抓狂的
soapbox /ˈsoʊpbɑːks/ n. 1:01:29
肥皂箱;喻指反复宣讲的立场
hard-coded adj. 1:01:29
硬编码的、写死在程序里的
Flip side phr. 1:02:10
另一面、反面情况
pestering /ˈpestərɪŋ/ v. 1:02:43
纠缠、不断烦扰
face threatening adj. 1:02:43
有损面子的(礼貌理论术语)
pin it on phr. 1:02:43
把责任归到(某人)头上
idiosyncrasies /ˌɪdiəˈsɪŋkrəsiz/ n. 1:03:45
独特的怪癖、特有性质
code switching phr. 1:04:32
语码转换(同一对话中切换语言)
inflected /ɪnˈflektɪd/ adj. 1:05:50
有屈折变化的
data sparsity phr. 1:05:50
数据稀疏性
orthography /ɔːrˈθɑːɡrəfi/ n. 1:05:50
正字法、拼写系统
grapheme /ˈɡræfiːm/ n. 1:06:40
字素(书写系统的最小单位)
phoneme /ˈfoʊniːm/ n. 1:06:40
音素
opaque /oʊˈpeɪk/ adj. 1:06:40
不透明的、难以看透的
tokenize /ˈtoʊkənaɪz/ v. 1:06:40
分词、切分为词元
eye opening adj. 1:07:37
令人大开眼界的
impediment /ɪmˈpedɪmənt/ n. 1:08:36
障碍、阻碍
treebanks /ˈtriːbæŋks/ n. 1:10:03
树库(带句法标注的语料库)
spurred /spɜːrd/ v. 1:10:03
刺激、催生
data-hungry adj. 1:10:03
极其耗费数据的
participatory /pɑːrˈtɪsəpətɔːri/ adj. 1:11:37
参与式的
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.137The Emerging Theory of Algorithmic Fairness 下一期 · NO.139 →Jeff Dean (Google): Exciting Trends in Machine Learning
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]