Jeff Dean (Google): Exciting Trends in Machine Learning · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Jeff Dean (Google): Exciting Trends in Machine Learning

节目发布 2024-02-15 · Rice Ken Kennedy Institute
杰夫·迪恩 主主持人
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文是谷歌首席科学家杰夫·迪恩(Jeff Dean)的一场公开演讲。他长期主持谷歌大规模系统与机器学习研究,也是 Gemini 项目的联合负责人。演讲以「机器学习的激动人心的趋势」为题,从图像识别、语音识别讲到语言模型、多模态、专用硬件与科学应用,并在结尾回答了现场提问。本文依据现场录音编译整理。

计算机学会了看世界

今天我要讲的是机器学习中那些令人兴奋的趋势。这会是一场视野很宽的演讲,不会深入任何一个具体领域,但我认为,理解这个领域正在发生什么、哪些东西值得兴奋、机会在哪里,以及我们在为所有人构建这项技术时该留意什么,都很重要。我要介绍的是谷歌许多许多人的工作。其中一部分我参与过,也是合著者,另一部分不是,只是我觉得你们应该了解的精彩工作。

先从几点观察说起。近些年,机器学习确实改变了我们对计算机能力的预期。回想十年或十五年前,语音识别勉强能用,但远谈不上顺畅,错误很多。计算机并不能从像素层面真正理解一张图里有什么。语言方面,自然语言处理有不少工作,但谈不上对语言概念和多语种数据的深层理解。今天我们已经走到了另一个阶段:你会理所当然地期待计算机能够看见并感知周围的世界,比十年前强得多。这打开了各种惊人的可能,几乎遍及人类活动的每一个领域。想想动物演化出眼睛的那一刻,计算领域此刻正处在类似的阶段。我们有了能看、能感知的计算机,这是一场完全不同的比赛。

另一点观察是规模在不断增长:更大规模地使用计算资源,使用我稍后会讲的专用计算机,使用更大、更有趣、更丰富的数据集,训练更大规模的机器学习模型。把这些东西全部放大,往往就能得到更好的结果。过去十到十五年一直如此,每一次扩大规模,效果都会变好。新的能力会突然涌现,或者某个问题的准确率跨过一道门槛,此前几乎不可用,此后忽然就能用了,于是催生出新的应用。

还有一点:由于这种新的、以学习为基础的范式,我们想运行的计算类型与传统那种手写的、弯弯绕绕的 C++ 代码很不一样,而许多基础 CPU 正是为高效运行后者而设计的。所以我们需要不同类型的硬件,把这些计算跑得更高效。在某种意义上,我们可以把计算机要做的事收窄到一个更小的集合,然后把它们做到极致地好、极致地高效,这样规模的持续增长也就更有可能实现。

十年间,箭头反了过来

过去十年,计算机能做的事取得了惊人的进步。从一张图像的原始像素出发,得到一万或一千个类别之一的标签,十年前的计算机做不到,现在做得到。从音频波形出发,识别出这五秒钟里说了什么,这是语音识别,我们取得了巨大进展。翻译,从「Hello, how are you?」到「Bonjour, comment allez-vous?」,把一种人类语言译成另一种,是计算机能帮我们做的极有用的事。

我们甚至能从一张度假照片,比如一只猎豹趴在吉普车顶上,生成对它的描述。不只是「豹」这样一个类别标签,而是一小句话,讲清楚这个场景里发生了什么。这已经很了不起了。更了不起的是,最近几年我们把很多箭头反了过来。从「豹」这样一个类别标签出发,计算机能给你生成五十张、一百张不同的豹的图像。从「外面有多冷?」这句文字出发生成音频波形,这是文本转语音,存在很久了,但进步很大。翻译方向的反转并不意外,只是越来越好。再往前一步,给出一段对你想要的图像的简短描述,就能得到一张图,现在有时甚至能得到一小段视频;用语言描述一种声音,就能得到一段音频。这些能力正在涌现,我认为这让人非常兴奋:今天我们能用计算机造出的东西,与十年前已经完全不同。

ImageNet 与语音识别

来看看过去十年进步的幅度。斯坦福大学开发了一个叫 ImageNet 的基准,很多人都听说过。训练数据是这样的形式:一堆彩色图像,每张配一个标签,标签来自一千个类别。你可以用大约一百万张这样的图像训练系统。然后给你一批从未见过的图像,你要预测它们的真实标签。机器学习工作的一个核心就在这里:如何把从数据中获得的观察,推广到新的情境、新的从未见过的图像。

2011 年这项竞赛第一次举办,获胜作品的准确率是 50.9%。第二次举办时,一篇著名的里程碑论文出现了,人们亲切地称它为 AlexNet,作者是亚历克斯·克里热夫斯基(Alex Krizhevsky)、伊利亚·苏茨克维(Ilya Sutskever)和杰弗里·辛顿(Geoffrey Hinton)。他们把准确率一举提高了大约 13 个百分点,这实在惊人。那年大约二十八支参赛队伍,只有他们一家用了神经网络。但这是一次重大突破,第二年几乎所有参赛者都改用神经网络,原因很简单,这是一次革命性的进步,也清楚地表明,直接从原始数据中学习,远胜于手工设计那些能标示「这是一只豹」的特征。手工设计特征实在太难了。你会设计什么样的特征,来判断这是一只豹,而不是一只长颈鹿或一辆车?而从数据中学习让这件事成为可能。

这是一次很大的飞跃,但人们也容易忽视此后的进步。在这项任务上,我们已经从 63% 走到了今天的 91%。这其实相当惊人。我们知道人类在这项任务上的准确率其实略低于这个水平,因为它真的很难:一千个类别,其中四十种是不同品种的狗。人盯着一张照片,其实并不知道那是哪个品种。这就是大约十年间发生的事,它彻底改变了计算机视觉。

再看语音识别。这是一个流行的开源基准,用来衡量语音识别的准确率,指标是词错误率,也就是识别错的词所占的百分比,当然越低越好。我们从 13.25% 降到了 2.5%,而这段时间要短得多,只有五年左右。大致上,从每六七个词错一个,变成每四十个词左右错一个。这对系统的可用性是天壤之别。忽然之间你可以依赖它了,可以开始口述电子邮件,它基本都能听对。

为神经网络造硬件

我提到过,扩大规模能提高模型的质量。所以我们需要能让规模更高效地扩大的硬件。花同样的硬件成本,或者消耗同样的能源,怎样才能因为效率更高而得到质量更高的模型?这实际上是在改变我们设计计算机的思路。为机器学习优化的硬件效率高得多,而且一代一代之间都有重大改进,这让更大规模的模型得以在更低的经济成本和能源成本下实现。

神经网络,也就是如今人人都在用的这类机器学习模型,有两个非常好的性质。第一,降低精度没有关系。如果你把模型里的计算保留到一两位有效数字而不是六位,没问题。很多时候,这些模型的优化算法还会刻意引入噪声,好让模型学得更好。所以降低精度可以看作往学习过程里加了一点噪声,有时效果反而更好。第二,你听到的所有那些喧嚣的算法,本质上都是用不同方式拼装起来的线性代数运算,比如矩阵乘法和各种向量运算。这些算法说到底就是大量线性代数原语的反复应用。所以,如果你能造一台特别擅长低精度线性代数的计算机,那正是你想要的:以更低的计算成本和能源成本,训练出高质量的模型。

谷歌做这件事已经有一段时间了。我们看到系统里确实有这样的需求,于是构建了张量处理单元(Tensor Processing Unit,TPU)的第一个版本,它的架构就是为低精度线性代数设计的。第一代是为推理而造的。你已经有了一个训练好的模型,现在要把它用到产品里,就需要投入这些计算来识别图像里有什么,或者在有人对着麦克风说话时识别出他说了什么。第一代 TPU V1 是一块单卡系统,上面有一个加速器。与当时的 CPU 相比,它在能效和计算性能上都有大约三十到八十倍的提升。

后面几代 TPU,我们转向由多块芯片组成的更大系统,同时面向训练和推理。TPU V2 板上有四块这样的芯片。TPU V3 板算是它的近亲,但我们加上了水冷,真的有水流到芯片表面帮助散热。TPU V4 板,我们加上了漂亮的颜色。后面这三代都是为组装成更大的系统而设计的,我们称之为 Pod。Pod 的规模一代比一代大。第一代 Pod 内部用的是非常简单但带宽很高的网络:每块芯片以二维网格的方式与四个邻居相连,机架里相当于一个 16 乘 16 的芯片网格,每块芯片和邻居之间基本上就是一根线。这样网络里就不需要做任何路由,可以有极高的带宽和极低的连接成本,因为你只是要把数据送到六英寸之外的下一块芯片。下一代把规模扩展到八个机架、1024 块芯片。再下一代用了 64 个机架,每个机架 64 块芯片,占据数据中心的好几排,4096 块芯片提供 1.1 exaflops 的低精度浮点算力。

最新一代是我们去年底公开的 V5 系列,有两个变体。一个偏向推理,一个 Pod 有 256 块芯片。另一个是 V5P,每块芯片的内存大得多,芯片之间的带宽和内存带宽也高得多,每块芯片的 16 位浮点性能接近半个 petaflop,int8 性能再翻一倍。它的 Pod 也更大,接近九千块芯片,算力非常可观。

从 n-gram 到词向量

现在来谈语言。前面讲了图像识别和语音识别的进展,但语言恐怕是人们感受到计算机变化最大的领域之一。我对语言模型的兴趣由来已久,早在神经网络之前就有了。当年我和谷歌翻译团队的几位同事合作。他们有一套能力很强的系统,翻译质量很高,但它是为研究竞赛设计的:两周之内只需翻译五十个句子,然后提交结果。它每翻译一个句子,要为二十万个 n-gram 做磁盘查找。

我说,既然翻译质量这么高,应该把它真正投入使用。于是我们造了一个系统来提供 n-gram 模型服务。它基本上就是统计每一个五词序列在两万亿词元里出现了多少次,这样能得到大约三千亿个不同的五元组。我们把它们存在一批机器的内存里,翻译一句话需要查的十万个条目并行去查。我们还想出了一个新算法,叫「傻瓜回退」(stupid backoff),它无视数学上正确的做法,改用一种简单得多的办法:查一个五元组,如果没有数据,就查它的前缀四元组,有就用;没有就查三元组,依此类推。结果它的效果居然相当不错,与更精巧的 Kneser-Ney 平滑相比并不逊色,而后者虽然是理论上该做的事,计算上却相当困难。这件事的一个教训是:简单的技术加上海量数据,非常有效。这个教训贯穿了我的整个职业生涯,你完全可以做非常简单的事,让数据自己说话。

后来我的同事托马斯·米科洛夫(Tomas Mikolov)对分布式表示产生了兴趣。不再把一个词当作一个离散的符号,而是用一个很高维的向量来表示它,比如每个词对应一个一百维的向量。在训练过程中,我们把出现在相似语境里的词往一起拉,把出现在不同语境里的词往外推。训练目标非常简单,就是「出现在相似语境里就拉近,不同就推开」,在数万亿词元上这样训练,你会在这个一百维空间里得到非常漂亮的性质。一百维空间不太好想象,但在那个高维空间里,非常相似的东西最终会靠在一起。「山」「丘」「崖」会彼此邻近。

空间里的点很有意思,但也许更有意思的是方向,在这个高维空间里,方向也是有意义的,毕竟一百个维度里可以走的方向很多。比如,你看「国王」在空间里的位置,要走到「王后」,需要朝某个方向走,把「国王」的向量从「王后」的向量里减掉,就是那个方向。结果发现,这个方向与从「男人」走到「女人」的方向大致相同。所以方向是有意义的,不同的方向代表不同的含义。从一个动词的现在时走到过去时,是另一个方向,而且不管是哪个动词都一样。这说明分布式表示蕴含着很大的力量,代表一个词的那个一百维向量里,编码了许多不同种类的信息。

序列到序列与多轮对话

接着,我的同事伊利亚·苏茨克维和黎国(Quoc Le)开发了一种叫序列到序列学习(sequence-to-sequence)的模型。它用一个神经网络处理一个输入序列。以翻译为例,你把一个英文句子一个词一个词地喂进去,系统根据它当前的状态加上新看到的词,更新出一个新状态。就像单个词有分布式表示一样,你现在得到了一个迄今为止所看到的整个句子的分布式表示。更新状态用的是一种叫长短期记忆(LSTM)的循环神经网络。碰到句末标记时,你训练模型吐出这个句子的正确译文。训练数据就是一对对意思相同的英文句子和法文句子,你训练模型在看到这个英文句子时输出那个法文句子,然后在大量成对数据上重复这个过程。

果然,你可以用一个神经编码器处理输入序列来初始化状态,相当于「我已经吸收了输入句子,现在要一个词一个词地解码出正确的译文」,再用这个状态去初始化神经解码器。把规模放大,它就管用了,翻译准确率大幅提高。

后来奥里奥尔(Oriol Vinyals)和黎国发表了一篇研讨会论文,指出除了翻译,你还可以用上下文来做多轮对话。你与某个人或某一方交互的一连串记录,模型回一句,对方再说一句,来来回回好几轮,这些之前的多轮交互就是你的上下文。然后你训练模型在这些历史轮次的语境下生成一个好的回复。它本质上是同一个模型,一个序列到序列模型,只是序列现在用整段对话的历史轮次来初始化。于是用神经语言模型做有效的多轮交互成为可能。这很妙。

Transformer 的并行之道

再往后,谷歌的另一批研究者加上一位实习生,提出了一种叫 Transformer 的模型。回想一下,前面那种模型是循环的:你有一个状态,取下一个词元,做一些处理来更新状态以吸收这个词元,然后带着新状态去吸收下一个词元,再更新一次。这是一个非常串行的过程,因为要吸收第三个词,你必须先处理完第二个词;要处理第二个词,必须先处理完第一个词。这不太理想。在计算机里,只要有办法,我们喜欢并行做事,而不是串行。

这个模型的做法是:把这批数据、也就是输入里的所有词并行处理,然后对其中不同的部分施加注意力,而不是维护一个随着词序列不断串行更新的单一状态。换句话说,不要把状态硬塞进一个分布式表示里,而是把你见过的所有词元的表示都保存下来,然后去「注意」它们:在翻译句子的这一部分或那一部分时,把注意力放到有意义的地方。结果是,用少十到一百倍的算力,得到更高的准确率。

还记得我前面讲的那些硬件进步和专用硬件吗?它们随时间给我们带来了巨大的提升,但我们同时也看到这样的算法改进,两者是相乘的。于是,靠算法进步加上机器学习硬件,我们现在能训练大得多的模型,也因此有了能力强得多的模型。

再后来,一群人决定把规模放大,用 Transformer 模型而不是循环模型在对话风格的数据上训练,得到了相当好的结果,尤其是提出了一种评估方式,要求回复既理智(sensible)又具体(specific)。你不希望聊天机器人含糊其辞,只会说「嗯,挺好」。你希望它针对你说的话给出真正有理智的回应,这样它才更吸引人、更有用。

大模型谱系与 Gemini

我讲了其中一些,但神经语言模型和神经聊天机器人各有一条演进的脉络。聊天机器人这条线上有神经对话模型、Meena、OpenAI 的 ChatGPT,以及我们谷歌大约一年前发布的 Bard。语言模型这条线上,有我讲过的序列到序列工作,有 OpenAI 的 GPT-2,其中一些模型有参数量可以参考,大致代表模型的规模:2019 年的 GPT-2 是十五亿参数;谷歌几位同事的 T5 是一百一十亿参数,能力很强。顺便说一句,Transformer 是这些模型的基础,这里的「T」和那里的「T」都是 Transformer 的意思。人们真正见识到了 Transformer 模型和架构带来的十到一百倍的计算改进,从此把它作为大语言模型的基础。然后是 GPT-3,DeepMind 同事的 Gopher,谷歌研究院的 PaLM,DeepMind 的 Chinchilla,谷歌研究院的 PaLM 2,OpenAI 的 GPT-4,再然后是 Gemini,也就是我和同事奥里奥尔·维尼亚尔斯共同领导的项目。我们有一大群人,分布在许多不同的研究办公室,一起构建有能力的多模态模型。

我们想做的一件事,是从只理解文本的语言模型,走向能同时处理所有模态的模型。你可以给它文本加图像,或者音频加文本,让它做一件事,它能流畅、连贯地处理你给它的任何模态。所以我们大约一年前启动这个项目时,目标是:训练出世界上最好的多模态模型,并在谷歌各处使用它们。关于 Gemini 有一篇博客,有一个网站,还有 Gemini 团队写的技术报告,我很自豪是团队的一员。

Gemini 的架构与三种尺寸

Gemini 从一开始就是多模态的。正如我说过的,我们不想只处理文本,我们要处理图像、视频和音频,把它们转成一串词元,然后在上面训练一个基于 Transformer 的模型。解码有两条路径:一条训练用来生成文本词元;另一条用 Transformer 学到的状态来初始化解码器,从这个状态出发生成一整幅图像的像素。

我们还支持交错输入。不是说你给它一段文本输入和一个图像输入就完了,你可以把它们交替排列。对于视频,你可以放一帧画面,接一段描述它的文字,再放一帧画面,再接文字或者音频里所说内容的字幕。然后让 Transformer 利用它在训练中接触过所有这些模态的事实,为你给它的各种模态建立起共同的表示。

我们有几种不同的尺寸。第一代 Gemini 有三种:Ultra 是我们最大规模、能力最强的模型。Pro 的尺寸适合在数据中心运行,我们把它用在许多产品场景里,比如我们的 Bard 产品,它现在改名叫 Gemini 了,有点容易混淆,它运行在 Pro 模型上,或者上周刚宣布的 Ultra 模型上。还有 Nano 模型。你其实希望许多机器学习模型能在设备上运行,在一部小手机或一台笔记本上。Nano 在这方面非常高效,尺寸也合适。你还可以对它做量化,让它变得更小。

训练基础设施与「有效产出」

关于训练基础设施,我们想要一套可扩展性很强的底座:你用非常高层的方式描述你想要的计算,然后由一个系统把这个计算映射到你手头的硬件上。我说过我们有这些 Pod。比如你描述自己的计算时说,我关心这两个部分,但我不在乎你把它们放在哪里,把决定权交给我们构建的底层软件系统 Pathways。它可能决定把这部分放在一个 Pod 上,把那部分放在另一个 Pod 上。它知道芯片的位置、拓扑以及它们之间的带宽。当这块芯片要和那块芯片通信时,会走我提到的那条极高速的链路;当模型的这一部分要和远处那一部分通信时,就会走数据中心网络,那条路的带宽要低得多。但这一切是无缝发生的,机器学习研究者或开发者不必操心,只需知道两者的性能特性不同。

训练大规模模型的一个特点是,随着规模扩大,故障一定会发生。某台机器会挂掉,某块 TPU 芯片会过热,然后以某种方式出错。所以把故障减到最少非常重要。有些需要避免的故障几乎是人为的。举个例子,我们曾有一套滚动升级机器内核的流程。如果那些机器各自跑着独立的计算,这么做完全没问题;但如果它们都是同一个上千台机器的计算的一部分,你反而宁愿把机器一起停下,同时升级这一千个内核,再一起启动,而不是让故障一路滚过去。所以我们优化了一些维修和升级流程。

做完这些之后,你还要把恢复时间减到最短,因为恢复得越快,就能越早真正取得有用的进展。我们有一个指标,叫「有效产出」(goodput),指模型训练真正在往前推进的时间百分比,而不是在从检查点恢复,或者在等系统的其他部分启动。我们用的一个办法是,从其他机器内存里保存的模型状态副本快速恢复,而不是去分布式文件系统读检查点。这让恢复时间从几分钟变成了五到十秒。

训练数据与数据质量

训练数据方面,我们希望模型是多模态的,所以要在大量网页文档、各种书籍、许多种编程语言的代码,再加上图像、音频和视频数据上训练。我们有一些启发式规则来过滤数据集,有的是手写的规则,有的是基于模型的分类器,用来判断这份文档是否在各方面都算高质量。训练数据的最终配比是通过在较小模型上做消融实验确定的:我们用不同的配比训练小规模模型,比如代码占 32% 还是 27%,然后在一大批指标上评估表现,以便更好地理解。我们还做了一些事情,比如在训练末期提高领域相关数据的权重,在快结束时加入更多多语种数据,好让多语种能力提升。

我确实认为数据质量是一个很有意思、很重要的研究领域。我们已经看到,真正高质量的数据对模型在你关心的任务上的表现有巨大的影响。从某种意义上说,它与你使用的模型架构同样重要,有些情况下甚至更重要。所以我认为这是未来研究的一个重要方向:自动学习课程的能力似乎很重要,识别高质量和低质量样本也是。

思维链:让模型写出过程

除了训练这些模型,在如何引出模型最好的一面上也有一批进展。你怎样提问,才能让模型更有效地回答问题?比如,要求模型「写出解题过程」,既能提高模型的准确率,也能提高可解释性。我的几位同事提出了一种技术,叫思维链提示(chain-of-thought prompting)。回想一下三年级的数学课,老师总是鼓励你写出过程,对吧?老师这么做,一是想看到你得出答案的思路,二是鼓励你去想下一步是什么,我要怎样把这个复杂的问题拆成一小串步骤。

通常你会先给模型一个示例:一个你想让它回答的那类问题,以及这个问题的答案;然后再问它一个新问题,让它作答。这里有一个例子,模型学到的回答方式就是直接算出答案给你。换一个问题,模型输出说答案是 50,这是错的。但如果你换一种问法,向它示范怎样写出过程:「肖恩一开始有五个玩具。如果每人给他两个,那就多了四个。五加四等于九,所以答案是九。」这是三年级数学老师会引以为豪的解题过程。更重要的是,你这么做之后,模型也会写出这种一步一步的推导,然后它就答对了。因为它现在有了更长的时间去思考通向正确答案的各个步骤。

这个效应相当显著。这两条线是同一个底层模型在不同规模下的表现,两个基准都偏数学,右边这个是八年级数学题,这个是一批算术题。你看到的是,用标准提示时,回答质量相当糟糕;但到了某个点,模型规模足够大之后,一旦改用思维链提示,准确率一下子蹿了上去。这说明,怎样向模型提问是一门很有意思的学问,问得好既能让模型更可解释,也更可能给出正确答案。

批改一份物理作业

来谈谈 Gemini 模型里的多模态推理,我想举一个例子,这是理解这个模型能做什么的好办法。提示是这样的:「这是一位学生对一道物理题的解答」,然后附上一张图,图里是题目和学生手写的答案。提示的其余部分说:「请一步一步地推理这个问题。」这又是思维链提示的风格。「学生的答案对吗?如果错了,请解释错在哪里,并解出这道题。数学公式请用 LaTeX,最终答案保留两位小数。」这就是输入:一张有点粗糙的手写图,一个滑雪者从斜坡上滑下来,能量守恒,诸如此类。

下面是模型的输出:学生没有得出正确答案。学生在计算斜坡起点的势能时犯了错。起点的势能由 mgh 给出,学生在计算中用了斜坡的长度(我猜就是斜边)而不是高度。正确的解法是这样这样,因此我们可以写出如下式子(这里其实是 LaTeX,为了方便阅读我们把它渲染出来了),代入数值,就是这个。它把题目做了出来,答案保留到两位小数。

想想这意味着什么。我们忽然可以给模型多模态的输入,一张白板的照片、一道题,让它去做一件事,它就能做。它不会每次都做对,但它能做。这可以成为一种了不起的教育工具。想象一个学生自己在琢磨题目,把自己的解答拍下来,系统就能帮他找出哪里错了。我们知道,一对一的人类家教带来的学习成果,比大班课堂环境高出两个标准差。我们能否在个性化辅导上接近那个水平?我认为这种可能性就在我们共同的能力范围之内。

三十二项基准,三十项领先

刚才是 Gemini 能力的一个定性例子,但也应该看看它在一系列不同特性上与其他模型的比较。评估确实能帮我们找出模型的长处和短板,帮我们理解训练是否顺利,所以训练过程中我们一直在评估这些指标。它还帮我们决定该改什么。数学表现比预期低?那也许该在训练配比里加更多数学数据。可这会对多语种表现造成什么影响?这里有很多复杂的权衡,有些在训练开始时就得定下来,有些则是在线监控,然后做出有依据的、或者说凭手感的决定。评估也让我们能把自己的能力与其他模型和系统作比较。

最高层的总结是:我们考察了 32 个学术基准,Gemini Ultra 模型在其中 30 个上超过了此前的最高水平。深入看其中一些,有一批面向文本、通用推理和数学的基准。把 Gemini Ultra 与 GPT-4 比较,后者在大多数问题上是此前的最高水平,蓝色标出的是最高水平,八项里我们拿了七项。MMLU 上的 90% 很有意思,因为这是一套覆盖面极广的题目,涉及 57 个学科,化学、数学、国际法、哲学等等。编制这个基准的团队测得人类专家水平是 89.6%,也可能是 89.8%。所以这个成绩实际上超过了这 57 个门类上的人类专家水平,这很好,我们很高兴。下面还有一批编程相关的基准和数学相关的基准。

再看图像理解基准,这就进入多模态的部分了。我们在八个基准里的八个上都拿到了最高水平。其中有一件好事:有一个基准是在我们发论文一周前才发布的,我们从没见过它。评估团队很快把它加进了评估集,结果发现我们以相当可观的优势超过了最高水平。这很好。碰到一个从未见过的基准还能做得好,总是让人安心,因为你总在担心训练数据泄漏到测试集里之类的问题。再看视频理解,这个模型的多模态能力确实很出色,六个基准里的六个都是最高水平,包括那个重要的英文烹饪视频字幕基准,以及视频问答等等。再看音频,在四个公开的语音识别基准和一个语音翻译基准上的词错误率,五项里五项都是最高水平。多语种能力也相当好,五项里拿了四项。

所以,首先我希望你们体谅一下我们的评估团队,评估这些模型、把能力理解到这种细致程度,是一项浩大的工作,非常了不起。它也让我们有了相当确定的把握:Gemini 模型的能力相当强。论文里也有 Pro 和 Nano 的测量结果。

与 Bard 的几段对话

这些大型 Transformer 模型能生成出人意料地连贯的对话,这算是神经对话模型以及后来那些基于 Transformer 的版本的演进结果。看看 bard.g……我想我得更新一下幻灯片了,现在应该是 gemini.google.com。几个月前,在 Bard 还没有用上 Gemini 模型之前,我在准备一场演讲,就对它说:帮我把「hot chips」和「tensor processing units」的字母倒过来。只是为了展示这些模型能做什么。它说:「好的,倒过来的字符串是」,然后就是结果。

很好。但它接着说:「我还可以用 Python 帮你做这件事。」代码在这里:定义一个叫 reverse string 的函数,这是字符串,打印这个的反转,打印那个的反转。「使用代码请谨慎」,这一点我一向建议。然后它还解释了代码:代码首先定义了一个叫 reverse string 的函数,接收一个字符串作为输入,返回反转后的字符串,函数的工作方式是遍历字符串,然后代码打印出反转结果。它总是乐于助人:「还有什么我可以帮你的吗?」

这相当惊人,对吧?有人问了一个问题,它做了被要求的事,然后还说,顺便告诉你,有一种东西叫编程,这里有一段 Python 代码,写代码来做这件事是这样的。我觉得这很酷,而且又是一个真正的教育机会。「还有什么我可以帮你的吗?」当然,多讲讲 TPU。这个模型有相当多的世界知识。它知道 TPU 是什么,基本上就是我已经告诉你们的那些:谷歌开发的专用硬件处理器,用来加速机器学习,能提高效率和性能。好处有这些:更快的训练和推理。「希望这对你有帮助。」

这些聊天机器人的一个有趣之处是,它们可以有不同的性格。Bard 就像你那个热心的朋友,会帮你回答各种问题。上个月我们把 Gemini Pro 放进了 Bard,也就是现在的 Gemini。有一个公开网站叫 LMSYS,可以评估不同的聊天代理,因为现在世界上有很多聊天机器人了。它的做法是让用户自己写提示,从系统里配置的机器人中随机挑两个,把同一个提示发给两者,然后匿名展示两边的输出,你只需说左边好还是右边好。

由此就能算出一个叫 Elo 的分数。Elo 我记得是一位匈牙利数学家,他想找到给国际象棋棋手排名的方法。在锦标赛里,你击败一个对手就会获得更多 Elo 分,所以在比赛环境下可以算出 Elo 分数。大致上,Elo 高出 100 分意味着较强的一方会有 64% 的胜率,高出 400 分意味着十比一的优势。Pro 级别的模型在这个榜单上拿到了第二高的 Elo 分数,榜上大概有三十个不同的模型,我展示的是榜单顶部。这很不错。而且这是一项免费服务,而那几个不是。

我在推特上发了这件事,然后一个我从没见过的人说,它确实相当好,看看这个例子。我们有一个功能,让人们可以把自己喜欢的、愿意公开的 Bard 对话分享出来,他就这么做了。点开一看,他问的是:请估算以下各国每百万居民拥有的公司数量:美国、英格兰、韩国、台湾和新加坡,以表格形式呈现。

Gemini 说:当然,下面这张表展示了每百万人的公司数量,正如你所要求的。它还替你解读结果:如你所见,英格兰最高,其次是美国、韩国、台湾,然后是新加坡。需要注意这些只是估算,实际数字可能因来源和定义而异。希望这对你有帮助,如有其他问题请告诉我。

他接着问:你从哪里找到这些公司数量的?「抱歉没有提供来源。信息来自哪里应该保持透明。我使用的公司数量来自几个不同的来源:美国用的是邓白氏(Dun & Bradstreet)数据库,英格兰用的是国家统计局」,等等。「请注意这些来源对公司的定义略有不同。」所以它其实理解其中的微妙之处,对吧?什么算一家公司有不同的定义,而它知道该去哪里找这些不同来源的数据。它并没有把这些背下来,但它能够利用它们。相当妙。

领域微调:Med-PaLM

我认为另一个重要的趋势是,对这些通用模型做进一步的精炼,可以造出惊人的领域专用模型。我的几位同事拿了我们早先的 PaLM 模型和后来的 PaLM 2 模型,它们是在通用文本上训练的通用模型,然后用医学数据做增强和继续训练,也就是医学问题和医学文献。他们发现,第一个 Med-PaLM 模型就超过了医师执照考试的及格线。六个月后,他们在此基础上训练出 Med-PaLM 2,在这项任务的医师执照考试上达到了专家水平。当然,这不是一个完全通用的场景,只是一批医学问题,但它确实展示了这样一种能力:先有一个能力很强的通用模型,再针对特定问题做领域专用的训练。

文生图与规模的力量

我会快速过一下生成图像和视频的生成式模型,你们大概已经把它当作世界潮流看到了。我们有几个不同的研究项目,Parti 和 Imagen。我提到过一件很酷的事:你可以用提示描述你想要的视觉图像,然后让模型生成这些图像,生成过程受处理这句话所得到的编码表示的约束,以此为条件生成图像的像素。

「一列蒸汽火车穿过一座宏伟的图书馆,伦勃朗风格的油画。」就是它。「一条用 X 做成的巨型眼镜蛇」,X 可以是玉米、煎饼、寿司或者沙拉。你最喜欢哪条?我有点偏爱那条凶狠的生菜蛇,不过玉米那条也很不错。「一张客厅的照片,有白色沙发和壁炉,墙上挂着一幅抽象画,明亮的光线从窗户照进来。」如果你像我一样正好需要一张这样的图片来做演示,就可以这么做。描述可以相当细致:「一张高对比度的照片,一只熊猫骑在马上。熊猫戴着巫师帽,正在读书。马站在一条街上,背景是灰色的混凝土墙,有五颜六色的花和「peace」这个词」,等等,「单反相机拍摄,白天的光线」。就是它。这段描述有很多种合理的诠释,但至少你得到了一个符合要求的例子。

这项功能现在已经集成进了 Bard。伊利诺伊州一家负责中小学教育的政府机构很兴奋,因为他们能给自己的吉祥物「超链接刺猬」(Hyperlink the Hedgehog)生成图像了。这是超链接在冲浪,乘着 AI 的浪潮。还有一个人很兴奋,他的提示是「一个人在伦敦的 Costa Coffee 买咖啡」,Costa Coffee 是一家很受欢迎的咖啡连锁店。这些模型过去常常挣扎的一件事是文字的保真度:把你要求的文字真正放进去,看起来像真实的字体,等等。这里你看到它做得相当好。

我不会讲太多细节,但大体上,你输入一个提示,得到这句话在分布式向量空间里的表示,以此为条件,模型先被训练生成一张小尺寸的图像;然后用另一个专门提高分辨率的模型,以那张低分辨率的图和文本嵌入为条件放大它;再对放大后的图做一次同样的事,同样以文本嵌入为条件,最终得到 1024 乘 1024 的全尺寸图像。

你能真切地看到规模的效应。我们训练了四个不同的模型,参数量从三亿五千万到两百亿,然后给它们同一个提示:「一张袋鼠的肖像照,它穿着橙色连帽衫,戴着蓝色太阳镜,站在悉尼歌剧院前的草地上,胸前举着一块写着『welcome friends』的牌子。」你看到的是,最小的模型大致抓住了袋鼠这一点,还有橙色连帽衫,但其他就不多了。有一块牌子,但文字还是让它很吃力。规模再大一点,袋鼠好了一些,它也多少知道悉尼歌剧院大概长什么样,但有点粗笨,细节不多。牌子上的字更接近「welcome friends」了,不过也可能是「Vegemite」,我不确定。规模再往上,你就得到了一张相当漂亮的图:悉尼歌剧院,你的袋鼠,橙色连帽衫,文字也对了。

所以你看,规模是这件事的一个重要方面。这也是为什么过去十年你看到所有这些进展,本质上是规模加上更好的训练方法和算法,共同带来了更高质量的结果。这张图说的其实是同一件事,但我觉得袋鼠说得更清楚。

手机里看不见的机器学习

还有一点很重要:大量机器学习正在以各种方式默默地帮助人们,尤其是在手机上。现代智能手机的许多相机功能这些年有了显著改进,靠的是计算摄影方法和机器学习方法的结合。人像模式,把背景全部虚化,让前景里的你显得很有格调,对那类人像照片来说是个不错的技术。夜视模式,在光线极弱的条件下拍照,可以从传感器多次读数,再用软件把它们叠加起来,得到远比实际环境明亮的成像,这也帮你拍出更好的星空照片。人像虚化和色彩突出,在你需要的时候也很好用。魔术橡皮擦,如果你真的理解图像,那么用户指着一根电线杆说「让这些消失」,系统就能做到。也许你的瀑布照片前面站着几个别的游客,你不想要他们,就可以擦掉。他们就没了。

手机上有很多功能,其中许多是关于如何把一种模态转换成另一种。有时你想筛选来电:也许你不想亲自接电话,而是让一个计算机生成的声音替你接,问对方为什么打来,然后把对方说的话转成文字给你,你再决定要不要接。「替我等待」功能可以替你在电话里等着,这样你给美国银行客服打电话时,就不必自己抱着听筒等了。实时字幕能把手机上正在播放的任何视频的音频听下来,给你配上旁白的字幕。也许你正在像这样的报告厅里想看一段视频,又不想让声音打扰别人。这里面有很多很酷的功能,很多都在人们的手机上运行,而人们未必意识到,也未必想过底下是什么技术。这对识字能力有限的人群也是巨大的进步:你把摄像头对准一样东西,它能把上面的字读给你听;也许你不懂那种语言,正想看明白,它可以读出来并替你翻译。

科学发现:搜索新材料

这一部分我会讲得快一些,跳过其中一些内容。我从材料科学讲起。这是一个很有意思的领域,机器学习正在开始影响科学的方方面面:或者是对科学假设空间中有意思的部分做自动化探索,或者是创造出学习得来的高速模拟器,以取代传统的大规模高性能计算。在某些领域,人们已经学出了一种模拟器,功能上等价于手工编码的模拟器,却快了十万倍。这意味着你忽然可以搜索一千万种可能的化学物质或材料,找出那些有意思、有前景、具备特定性质的候选,而这在过去需要多得多的算力。

我在 DeepMind 的几位同事正在探索一些有意思的方法,在可能的材料空间里搜索那些性质有趣的材料。他们有一条结构管线,能把一种潜在材料表示成一个图神经网络;还有一条组分管线,能把已知结构变异成邻近的、有意思的新结构;再利用现有的材料数据库,输出能量模型和一批稳定的、有意思的候选化合物。这样自动发现了两百二十万种新的晶体结构,带来一大批有趣的候选,可以在实验室里真正合成出来,看看它们究竟有什么性质。

医疗影像:糖网与皮肤病

我认为机器学习在医疗保健的各个方面都有巨大的潜力。我们在医学影像和诊断这个方向上已经做了相当长时间的工作,问题的类型很多:有的是二维图像,有的是核磁共振或 CT 扫描得到的三维体数据;有的只有单一视角,有的有多个视角;还有病理学里那种分辨率极高的大图像。这方面有相当多的工作,我只简单讲两项。

我们在这个领域做得最久的方向之一是糖尿病视网膜病变。这是一种退行性眼病,及时发现的话非常好治,发现不及时则可能导致部分或完全失明。有风险的人群,也就是所有糖尿病和糖尿病前期患者,都应该每年筛查一次。但在世界上很多地方,受过训练、能解读视网膜图像的眼科医生根本不够。机器学习在这里能帮上大忙,因为你可以让训练有素的眼科医生标注图像,「这张是一级,那张是三级,这张二级,那张五级」,用这些来训练模型。用持证眼科医生标注的数据训练,你能得到一个与持证眼科医生同样有效的模型。如果再请视网膜专科医生来标注同一批训练数据,他们在这类病例上有多得多的专业知识和经验,你就能训练出一个与视网膜专科医生相当的模型,而那是这个领域护理的黄金标准,全世界这样的专家寥寥无几。可现在,用一台笔记本上的 GPU,你就能让筛查质量达到视网膜专科医生的水准。我们已经与印度的一个眼科医院网络、泰国政府以及法国和德国的机构合作,每年做大量的筛查。

再说皮肤病学。这个领域有意思的地方在于,收集有用的数据并不需要专门的设备,就能判断你是否有皮肤问题。我们现在部署了一个系统,像视频里那样,你拍一张照片,它会告诉你这可能是什么,皮肤病数据库里还有哪些看起来相似的图像,帮你判断这是很严重的问题,还是相当良性的问题。

AI 原则:偏见、隐私与可解释

最后,随着我们把机器学习方法部署到世界上越来越多的地方,对它们更深、更广的理解真的非常重要。当我们从机器学习的基础研究,走到在所有产品的许多地方使用它时,我们开始思考一套原则,用来审视使用机器学习的影响,以及在各种可能的应用中该有哪些考量。2018 年我们发布了自己制定的一套原则。

这些原则的初衷是教育我们内部的团队,让他们在把机器学习用到自己关心的问题上时,知道该考虑什么。比如,避免制造或强化不公平的偏见。训练这些模型时,用的往往是来自真实世界的数据,而那常常是世界本来的样子,不是我们希望它成为的样子。所以部署机器学习模型时,一定不能在带有不公平偏见的数据上训练,然后把这种偏见放大,因为你现在可以自动化地、更快地做出这些决定。有一批算法层面的技术可以消除某些类型的偏见。我们努力做的是,一方面应用当前已知的最佳技术,另一方面也做研究,推进偏见等领域的最高水平。再比如,对人负责,我们认为让模型可解释是其中重要的一环。还有在合适的场景下注重隐私,以及对社会有益。

我要指出,其中很多都是活跃的研究领域。过去五六年,我们发表了大约两百篇与公平、偏见、隐私或安全相关的论文,你们可以在那里看到。

结语

总结一下,我认为这是计算领域激动人心的时代。一场转变正在发生:从手工编码的软件系统,转向学习得来的、能以各种有趣方式与世界互动、与人互动的系统。计算机能够摄入、理解和生成的模态在不断增加,我认为这会让使用计算机变得更加顺畅自然。很多时候我们把自己限制在敲键盘之类的方式上,但现在我们有能力以非常自然的方式与计算系统交谈,它能听懂我们说的话,能用听起来自然的声音回应,或者在我们要求时给出一张漂亮的图。这非常令人兴奋。机会无疑是巨大的。

但责任也同样巨大。我们该如何推进这项工作,确保它对社会有益,真正用它在世界上做好事?为此,谢谢大家。

答问

有人问,这个问题你们大概也料到了:更多的数据会让模型更好吗?数据翻倍,表现会好一倍吗?这是个好问题,但答案并不简单。我们已经看到,更多的高质量数据绝对能让模型表现更好,前提是你有足够的容量在这么多数据上训练。所以要考虑模型的容量,有时训练数据多了,模型的规模也得跟着扩大。我们也见过更多数据反而有害的情况:如果你拿到一大堆低质量数据,模型做数学题之类的能力反而会下降。所以这件事有细微之处,但总体上,更多高质量数据加上更大的模型容量,会让模型更好。

另一个问题是:既然绝大多数高质量训练数据已经用尽,大语言模型的未来在哪里?我对这个论断不太认同。我认为我们还几乎没开始在视频上训练。我们做过少量视频,但世界上有海量的视频数据。通过视觉和听觉数据来理解世界,与在大量语言上训练是不一样的,你两者都会想要。但我不认为我们已经耗尽了世界上的训练数据。

关于多模态模型,我在演讲里着重讲过。它们在所有领域上都比为每个领域单独训练的专用模型表现更好吗?我认为在某些情况下是这样的。这个问题可以换个说法:加入更多的模态,会不会提高其他模态上的表现?你希望如此,而我们确实普遍看到了一些这样的迹象。不过,如果你的问题很窄,又收集了一个专门针对这个问题设计的数据集,那在这个问题上往往也能得到好的表现。但如果问题很复杂,或者很难收集非常专门的数据,你要的就是一个对世界上各种事物都有海量知识的模型,从语言、图像到音频都懂,然后把它用到你关心的问题上。如果你手头有一点针对这个问题的数据,你会想从那个基础模型出发,做微调或者上下文学习之类的事,把表现做得相当好。

还有一个相关的问题:如今训练大模型的成本让小型创业公司难以产生影响,资源有限的个人该做什么样的项目?当然值得谈。机器学习领域有一大批问题。我更愿意从这个角度回答:如果没有大型数据中心的算力,在这个广阔的领域里能做什么有意思的研究?我认为可选的东西非常多。我提到过数据质量,比如数据质量的自动评估,还有在线课程学习、优化方法等等,其中很多都可以在一块 GPU 或者你桌子底下的几块 GPU 上做出证明,取得相当重要而有创新性的进展。最初的 Transformer 工作,我记得是在八块 GPU 上完成的;序列到序列模型肯定是八块 GPU。所以我认为,聪明的想法、扎实的评估,哪怕只在小规模上做出证明,都能带来进步。

还有一组问题是:大语言模型就是一切吗?Transformer 就是一切吗?还有别的吗?我们该研究其他类型的模型吗?对大语言模型的强调是否在压制机器学习的其他工作?这确实是个值得担心的问题。我们是不是在排挤其他有创新性的想法?那些想法也许还没有充分发展,所以看起来不如那些已经被深入探索过的东西好,而我们现在只是在已知有效的东西周围做温和的探索,也许真正有效的东西在另一个方向上。很多时候,只需要适量的实验证据,哪怕规模很小,就能证明另一个想法是一个真正有意思的方向。我认为这是一个重要的方向。

另外,我倾向于不用「大语言模型」这个词,因为我认为我们正在走向一个多模态的世界。而且多模态不会只限于你想到的那些人类模态,比如视觉、听觉和语言,还会包括世界上其他重要的模态,比如医疗应用里心率传感器数据那样的时间序列。你想处理的数据模态,大概有五十到一百种。

本期讲者
杰夫·迪恩谷歌首席科学家,1999 年加入谷歌,主导了 MapReduce、Bigtable、TensorFlow 等基础系统,2018 年起领导 Google AI,现与 Oriol Vinyals 共同领导 Gemini 项目。
主持人讲座主办方的学术主持人,负责汇总 Slido 平台上的观众提问并在问答环节向讲者转述。
章节 · 点击跳转视频
0:04 开场:机器学习改写了对计算机的预期 ▶ 正在看
5:11 十年基准:ImageNet 与语音识别的跃升 ▶ 正在看
8:28 为低精度线性代数造芯片:TPU 五代 ▶ 正在看
13:49 从 n-gram 到词向量:简单方法加大数据 ▶ 正在看
18:31 序列到序列、Transformer 与聊天模型 ▶ 正在看
24:57 Gemini:原生多模态与训练基础设施 ▶ 正在看
32:58 思维链提示与多模态解题示例 ▶ 正在看
38:08 评测:32 项基准与 Elo 排行榜 ▶ 正在看
48:03 领域精调、图像生成与手机端应用 ▶ 正在看
56:46 科学与医疗:材料发现与视网膜筛查 ▶ 正在看
1:01:57 AI 原则与结语:从手写软件到学习系统 ▶ 正在看
1:05:10 问答:数据、多模态与小算力研究 ▶ 正在看
本期论点
本期回应
2:09
同时放大算力、数据集与模型规模往往带来更好结果,并涌现出新能力 照直做大把语言模型做得更大,就能到人的水平吗?
15:51
在大量数据上使用简单方法非常有效,往往不输于理论上更讲究的做法 靠学习智能主要靠什么长出来?
32:15
训练数据的质量有时和模型架构一样重要,甚至更重要 靠学习智能主要靠什么长出来?
1:04:20
计算领域正在从手工编写的软件系统转向学习出来的系统 靠学习智能主要靠什么长出来?
1:07:26
世界上的训练数据远未用尽,海量视频与音频几乎还没被用来训练模型 照直做大把语言模型做得更大,就能到人的水平吗?
2:31
机器学习的计算方式与手写代码差别很大,需要专门设计的硬件而非通用CPU 就该专门造AI 该用专门设计的芯片吗?
1:09:57
数据质量评估、课程学习、优化方法等方向只用几块GPU就能做出创新进展 不会集中AI 的能力会不会集中到少数机构手里?
其他论点
9:14
神经网络在低数值精度下也能正常工作,算到小数点后一两位就够用 观察
21:50
在计算机里只要做得到,并行处理就优于串行处理
22:35
用注意力替代顺序更新的单一状态,能以少10到100倍的算力获得更高准确率 观察
26:16
多模态模型应从一开始就联合训练图像、视频与音频,而不是在文本模型上后接模态 做法
33:14
要求模型展示推演过程,能同时提升回答准确率和可解释性 做法
57:31
在某些科学领域,学习出来的模拟器功能上等价于手工模拟器,但快十万倍 观察
1:00:26
用认证眼科医生标注训练出的模型,判读视网膜图像的水平可与认证眼科医生相当 观察
1:08:47
面对复杂问题,与其单独定制窄模型,不如从多模态基础模型出发再微调 做法
01开场:机器学习改写了对计算机的预期
0:04
I'm going to talk to you about exciting trends in machine learning. It's going to be a very broad talk. It's not going to go into detail in any particular area, but I think it it's important to understand what is happening in this field and what is exciting and also, you know, what are the opportunities and also what are the things we should be sort of aware of as we build out this technology for everyone. Um and I'm presenting the work of many, many people at Google. So, some of this is work I've been involved in and the co-author, some of it is not. It's just cool work I think you should learn about. So, uh with that, I'm going to now give you another glimpse at the Slido number, uh 2207201.
我要和大家聊聊机器学习领域一些令人兴奋的趋势。这会是一个非常宽泛的演讲。它不会深入到任何某个具体的领域,但我认为,理解这个领域正在发生什么、有什么令人兴奋的地方,以及有哪些机会、还有哪些是我们在为所有人打造这项技术时应该有所警觉的东西,这些都很重要。嗯,我要展示的是谷歌很多很多人的工作。所以其中有些工作是我参与过、也署名的,有些不是。只是我觉得很酷、值得你们了解的工作。那么,接下来我再给大家看一眼 Slido 的编号,2207201。
便签引用
0:41
Um and uh this is how you ask questions for this talk. Um so with that, let's let's start with some observations. Um so in recent years, I think machine learning has really changed our expectations of what we think of computers as being able to do. If you think back 10 or 15 years ago, you know, speech recognition kind of worked, but it wasn't, you know, really seamless. It you know, made lots of errors. Computers didn't really understand images uh from the pixel level what was in that image. Um language was kind of uh uh there were a bunch of work in natural language processing, but it wasn't really a deep understanding of language concepts and multilingual uh data, but I think we've moved from that state to one where you actually expect computers to be able to sort of see and perceive the world around us in a much better way uh than they were able to 10 years ago. And that opens up all kinds of amazing opportunities in, you know, pretty much every field of human endeavor cuz all of a sudden, you
嗯,这就是本场演讲提问的方式。那么,我们先从一些观察开始。嗯,我认为近年来,机器学习真的改变了我们对计算机能做什么的预期。如果你回想十年或十五年前,语音识别算是能用,但并不顺畅。它会出很多错。计算机并不能真正从像素层面理解图像里有什么。嗯,语言方面也是,自然语言处理有不少工作,但那并不是对语言概念和多语言数据的深层理解。但我认为我们已经从那个阶段,走到了今天这样一个状态:你真的会期待计算机能以比十年前好得多的方式去看见、去感知我们周围的世界。而这就在几乎人类所有的领域里,打开了各种各样了不起的机会。因为突然之间,
便签引用
1:44
know, think about when animals evolved eyes, you know, we're sort of at that stage in computing. Uh we now have computers that can see and sense and that's a completely different uh ballgame. Uh Uh, and the other observation is increasing scale, larger scale use of computer compute resources, specialized computers I'll talk about, you know, uh, larger and more interesting and richer data sets, larger scale machine learning models. All scaling all of those things tends to deliver better results. And that's been true for the last 10 to 15 years, and every time we scale things up, things get better. All of a sudden new capabilities emerge or the accuracy of some problem reaches a threshold where before it was kind of unusable and now all of a sudden it becomes usable and that enables new kinds of things.
你想想动物演化出眼睛的那一刻——我们在计算领域大概就处在那个阶段。我们现在有了能看见、能感知的计算机,那完全是另一个局面了。呃,另一个观察是规模在不断扩大:更大规模地使用计算资源、专用计算机——这个我等下会讲——更大、更有意思、更丰富的数据集,更大规模的机器学习模型。把所有这些东西都放大,往往就能带来更好的结果。过去十到十五年一直如此,我们每一次把规模做大,效果就变好。突然之间会涌现出新的能力,或者某个问题的准确率跨过了一个门槛——之前基本没法用,现在一下子就可用了,而这又催生出新的东西。
便签引用
2:31
Um, and also the kinds of computations we want to run because of this new machine learning learning-based paradigms is pretty different than traditional handwritten twisty C++ code that, you know, a lot of basic CPUs were designed to write to to run effectively. And so we want different kinds of hardware in order to run these computations more efficiently. And we can actually, in some sense, focus on a narrower set of things we want computers to do and do them extremely well and extremely efficiently and then be able to have, you know, that increasing scale, uh, actually, uh, be even more possible.
嗯,还有就是,由于这种新的、基于机器学习的范式,我们想跑的计算类型,跟传统那种手写的、绕来绕去的 C++ 代码很不一样——很多基础 CPU 当初就是为了高效运行那类代码而设计的。所以我们需要不同种类的硬件,来更高效地跑这些计算。而且某种意义上,我们其实可以聚焦在一个更窄的、我们希望计算机去做的事情集合上,把它们做到极致好、极致高效,然后让那种不断扩大的规模,呃,变得更加可行。
便签引用
3:08
Okay. So, uh, I mentioned some of these things, but there's been a decade of amazing progress in what computers can do. So, thinking about, uh, this, uh, you have, you know, going from the raw pixels of an image to, you know, maybe a categorical label of one of 10,000 or 1,000 different categories, you know, computers didn't used to be able to do that a decade ago and now they can. Uh, audio uh, waveforms to, you know, what was being said in that utterance of, you know, 5 seconds of of audio. Uh, that's speech recognition and we've made tremendous progress on that. Um, translation, "Hello, how are you?" to "Bonjour, comment allez-vous?" You know, being able to translate from one human language to another is an incredibly useful, uh, capability for computers to help us with.
好。嗯,我提到过其中一些内容,但过去十年里,计算机的能力有了惊人的进步。所以,想想看,呃,这个,你知道,从一张图像的原始像素,到,你知道,给出一万个或者一千个不同类别中的某一个标签,你知道,十年前计算机做不到这一点,而现在可以了。呃,从音频波形,到,你知道,那段五秒钟音频里说的是什么内容。呃,这就是语音识别,我们在这方面取得了巨大的进展。嗯,翻译,把 "Hello, how are you?" 翻成 "Bonjour, comment allez-vous?",你知道,能够把一种人类语言翻译成另一种,这是计算机能提供的一项极其有用的能力。
便签引用
3:54
Um, and we've even been able to go from, you know, something like There's a nice vacation photo of a cheetah on top of a uh a Jeep. Uh going from something like that to a description of that. Um not just a categorical label like leopard, but, you know, a little short sentence that describes what is going on in that scene. So, that's pretty amazing. Like, we made tremendous progress in this. What's more amazing, though, is we've been able to reverse a lot of these arrows in the last uh few years. Um so, going from a categorical label like leopard to, you know, the computer will generate you 50 or 100 different images of leopards.
嗯,我们甚至能做到,你知道,比如有一张很不错的度假照片,一只猎豹站在一辆吉普车顶上。呃,从这样一张图出发,生成对它的描述。嗯,不只是像"豹"这样一个类别标签,而是,你知道,一个简短的句子,描述这个场景里正在发生什么。所以这相当惊人。我们在这方面取得了巨大进展。不过更惊人的是,过去这几年里,我们已经能够把这些箭头中的很多条反过来。嗯,也就是从"豹"这样一个类别标签出发,你知道,计算机会给你生成五十张或一百张不同的豹的图像。
便签引用
4:32
Or, you know, how How cold is it outside? To audio waveforms. That's just text to speech. That's been around for a while, but it has improved a lot. Um the translation reversal is not that surprising, but is getting better and better. And then, even going from a short description of an image you want and getting, you know, an image or sometimes now even a short video clip of what you want or a short, you know, an audio clip when you describe a sound in language. Um so, these are capabilities that are now starting to emerge, and I think are pretty exciting about uh you know, what we can build with computers now as opposed to a decade ago.
或者,你知道,从"外面有多冷?"这句话生成音频波形。这就是文本转语音。这项技术已经存在一段时间了,但进步了很多。嗯,翻译方向的反转没那么令人意外,但也在变得越来越好。然后,甚至可以从你想要的图像的一段简短描述出发,得到,你知道,一张图像,有时现在甚至是一段短视频,就是你想要的内容,或者一段简短的音频,当你用语言描述一种声音的时候。嗯,这些能力现在正开始涌现出来,我觉得挺让人兴奋的,就是,你知道,跟十年前相比,我们现在能用计算机造出什么东西。
便签引用
02十年基准:ImageNet 与语音识别的跃升
5:11
Um So, let's look at the level of improvement we've had in the last decade. So, Stanford uh developed a benchmark called ImageNet, which many of you have heard of, uh which is basically going from, you know, you get some training data of this form, like a bunch of color images, and labels, one of a thousand labels. Uh and then, you can train your system on about a million images like that. And then, you're given a bunch of images you've never seen before, and you have to predict, you know, what is the the the actual label for these new images that you've never seen before. And a A of machine learning work is how do you generalize from observations you've made on data to new settings, new images you've never seen before.
嗯,那我们来看看过去十年里我们取得的进步幅度。斯坦福,呃,做了一个叫 ImageNet 的基准测试,你们很多人都听说过,呃,它基本上是从,你知道,你拿到这种形式的一批训练数据,比如一堆彩色图像,以及标签,一千个标签中的一个。呃,然后你可以用大约一百万张这样的图像来训练你的系统。之后,你会拿到一批你从没见过的图像,你必须预测,你知道,这些新图像的真实标签到底是什么。而机器学习工作的核心,就是你如何把在数据上观察到的东西泛化到新的场景、你从没见过的新图像上。
便签引用
5:54
And um so in 2011, the first year the contest was run, uh the winning entry got uh 50.9% accuracy. And then uh the next time the contest was run uh in a very famous landmark paper uh affectionately known as the AlexNet um by Alex uh Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, they made a giant leap forward uh in accuracy, you know, 13% or so improvement in accuracy, which was just remarkable. And they were the only one of about 28 entrants that year that used a neural network. Um but this was a major improvement, and the next year nearly all the entrants used a neural network simply because this was such a revolutionary improvement, and uh clearly a a really major approach to just learn from the actual raw data rather than trying to hand engineer features that are indicative of a leopard. That's a really, you know, hard thing to do. What features would you hand design to decide this is a leopard as opposed to a giraffe or a car? Um but learning from data actually makes that possible.
嗯,在 2011 年,也就是这个比赛第一次举办的那年,呃,冠军作品的准确率是 50.9%。然后,呃,下一次举办比赛的时候,呃,在一篇非常著名的里程碑式论文里,呃,大家亲切地称之为 AlexNet,嗯,作者是 Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton,他们在准确率上实现了巨大的飞跃,呃,你知道,提升了 13% 左右的准确率,这简直太了不起了。而且在那年大约 28 支参赛队伍中,他们是唯一一支使用神经网络的。嗯,但这是一次重大的提升,而第二年,几乎所有参赛者都用了神经网络,原因很简单,因为这是一次如此革命性的提升,而且,呃,显然这是一条非常重要的路子——直接从原始数据中学习,而不是试图手工设计出那些能表征"豹"的特征。你知道,那真的是一件很难的事。你会手工设计出什么特征来判断这是一只豹,而不是一只长颈鹿或者一辆车?嗯,但从数据中学习实际上让这件事变得可能了。
便签引用
7:07
Um so I think that's that's a pretty big improvement, but it's also easy to ignore the improvements happened since then. Like we've gone from 63% to now 91% accuracy on this task. And so that is actually, you know, pretty amazing. We know human accuracy on this task is actually uh a bit below that level because it's actually pretty hard. There's a thousand categories, 40 different breeds of dogs. People don't actually know, you know, if they're staring at a photo which breed of dog is that. So uh that's pretty amazing. Uh this was, you know, about a 10-year span.
嗯,所以我觉得那是一次相当大的提升,不过之后发生的那些提升也很容易被忽视。比如我们已经从 63% 提高到了现在这项任务上 91% 的准确率。所以这实际上,你知道,相当惊人。我们知道人类在这项任务上的准确率其实比这个水平要低一点,因为它确实挺难的。有一千个类别,四十种不同的狗的品种。人们其实并不知道,你知道,盯着一张照片看,那是什么品种的狗。所以,呃,这相当惊人。呃,这大概是,你知道,十年左右的跨度。
便签引用
7:42
And it's revolutionized computer vision. If you look at speech recognition, so this is a popular open source um um benchmark for measuring speech recognition accuracy. Here it's measuring word error rate, you know, the percentage of words that are wrong. Uh and you want lower numbers, obviously. So, we've gone from 13.25% down to 2.5% and this is a much shorter span. This is only like 5 years. So, you know, basically kind of one word in, you know, six or seven is wrong to one word in 40 or so is wrong. And that that makes a huge difference in usability of these systems. All of a sudden, you can rely on it. You can start to dictate your emails and it mostly gets things right.
它彻底改变了计算机视觉。再看语音识别,这是一个很常用的开源,嗯,嗯,用来衡量语音识别准确率的基准。这里衡量的是词错误率,你知道,就是出错的词所占的百分比。呃,显然数字越低越好。我们已经从 13.25% 降到了 2.5%,而且这个跨度要短得多。这只有大约五年。所以,你知道,基本上从大概每六七个词错一个,变成了每四十个词左右才错一个。这对这些系统的可用性带来了巨大的差别。突然之间,你可以依赖它了。你可以开始口述你的邮件,而它基本上都能听对。
便签引用
03为低精度线性代数造芯片:TPU 五代
8:28
Uh pretty great. Um And so, I mentioned that scaling things up uh actually improves the quality of these models. And so, we actually want hardware that enables us to scale up more efficiently. How can we, you know, for the same amount of dollars of computer hardware or the same amount of energy get something that gives us an even higher quality model because it's more efficient. And so, it's really thinking, transforming how we design computers. And there uh you know, machine learning optimized hardware is much more efficient and there's major improvements that have happened generation to generation and this enables these larger scale models with lower economic and energy costs.
呃,相当棒。嗯,我提到过,把规模扩大,呃,确实能提升这些模型的质量。所以我们其实希望有硬件能让我们更高效地扩大规模。我们怎样才能,你知道,用同样多的钱买到的计算硬件,或者同样多的能耗,得到一个质量更高的模型,因为它效率更高。所以这真的在改变我们设计计算机的思路。而且,呃,你知道,为机器学习优化的硬件效率要高得多,而且每一代之间都有重大的提升,这让这些更大规模的模型能以更低的经济成本和能耗成本实现。
便签引用
9:08
So, there's two really nice properties of neural networks, the kind of machine learning models that everyone is is using these days. Um The first is reduced precision is okay. Like, if you carry out the computations in the machine learning model to one or two decimal digits of precision instead of six, you know, that's fine. You know, a lot of times some of the optimization algorithms for these things actually introduce explicit noise in order to make the model learn better. And so, you can think of reduced precision as just a way of, you know, in some sense adding a bit of noise to the learning process, and it actually sometimes works better.
神经网络,也就是如今大家都在用的这类机器学习模型,有两个非常好的特性。嗯,第一个是降低精度是可以接受的。比如,如果你在机器学习模型里把计算精度做到小数点后一两位,而不是六位,你知道,这没问题。你知道,很多时候,这些东西的一些优化算法实际上会刻意引入噪声,以便让模型学得更好。所以你可以把降低精度理解为,你知道,某种意义上给学习过程加了一点噪声的方式,而它有时候实际上效果还更好。
便签引用
9:44
And the other property is all the computations, all the algorithms you're hearing so much noise about, in some sense, are really just different transpositions of ways of assembling different linear algebra operations. So, things like matrix multiplies, vector operations of various kinds. So, if you can make And And that's really what those algorithms are is repeated applications of lots of different linear algebra primitives. And so, if you can make a computer that's really good at reduced precision linear algebra, that's what you want for learning these really high-quality models at sort of reduced uh uh computational cost or energy cost.
另一个特性是,所有这些计算,所有你听到那么多喧嚣的算法,某种意义上其实只是把不同的线性代数运算以不同方式组装、变换的结果。所以,比如矩阵乘法、各种各样的向量运算。所以,如果你能做到——而且那些算法本质上就是这样,是大量不同线性代数原语的反复应用。所以,如果你能造出一台特别擅长低精度线性代数的计算机,那就正是你想要的,用来以更低的,呃,呃,计算成本或能耗成本,训练出这些真正高质量的模型。
便签引用
10:27
And so, we've been doing this for a while at Google. We've saw that there was a real need for uh in our systems to build a system uh the initial version of what's called a Tensor Processing Unit or TPU is really this architecture designed for low-precision linear algebra. And the first one we built was for inference. When you already have a trained machine learning model, but now you want to apply it in a product product setting, you now need to like apply all this compute in order to recognize what's in that image or to have someone utter something in a audio in their microphone and then be able to recognize what they're saying.
我们在谷歌做这件事已经有一段时间了。我们看到,我们的系统里确实有这种需求,呃,要构建一个系统,呃,最初那一版所谓的张量处理单元,也就是 TPU,本质上就是一种为低精度线性代数设计的架构。我们造的第一代是用于推理的。当你已经有了一个训练好的机器学习模型,但现在你想把它应用到产品场景里,你就需要,比如说,投入所有这些算力去识别一张图里有什么,或者有人对着麦克风说了什么,然后能识别出他们在说什么。
便签引用
11:04
And so, we built uh the first generation, which was really just a single-card system uh that had one of these accelerators on it, TPU V1. Um And that uh actually was a, you know, about a 30 to 80x improvement over using a CPU at the time in terms of both uh energy efficiency and and computational performance. Um Later generations of the TPU, we then focused on uh larger-scale systems composed of, you know, multiple chips that were designed for both training and inference. And so, this is the TPU V2 board with four of these chips. Uh, the TPU V3 board is sort of a uh, close cousin of this one, but we added uh, water. So, there there's actually uh, water going to the surface of the chips to help with cooling. Uh, the TPU V4 board, uh, we added cool colors.
所以,我们造了,呃,第一代,它其实就是一块单卡系统,呃,上面有一个这样的加速器,TPU V1。嗯,那个,呃,实际上,你知道,在能效和计算性能两方面,都比当时用 CPU提升了大约 30 到 80 倍。嗯,在后面几代 TPU 中,我们就把重点放在了,呃,由多块芯片组成的更大规模系统上,这些芯片是为训练和推理两方面设计的。这就是 TPU V2 的板子,上面有四块这样的芯片。呃,TPU V3 的板子算是这一块的,呃,近亲,但我们加了,呃,水。所以实际上有,呃,水流到芯片表面来帮助散热。呃,TPU V4 的板子,呃,我们加了炫酷的配色。
便签引用
11:57
Uh, which is nice. And those three later generations were designed to be assembled into larger systems that we call pods. And so, the pods uh, increased in scale over the generations. So, the first one, um, they have very very simple, but high bandwidth networks uh, in the pod. So, basically each chip in this uh, first generation was connected to its four neighbors in a 2D mesh. So, you have a 16 by 16 grid of chips in some sense in these in these racks. And every chip is connected to its neighbor with basically a wire.
呃,挺不错的。而后面这三代都被设计成可以组装成更大的系统,我们称之为 pod。所以这些 pod,呃,规模随着代际不断增加。第一代,嗯,它们有非常非常简单、但带宽很高的网络,呃,在 pod 内部。基本上,这第一代里的每块芯片都以二维网格的方式连接到它的四个邻居。所以,从某种意义上说,这些机架里是一个 16 乘 16 的芯片网格。而每块芯片基本上就是用一根线连到它的邻居。
便签引用
12:34
Uh, and so, that means you don't have to do any routing in the network. Um, and so, you can have very high-speed bandwidth, very low-cost connections because you're only trying to go 6 in to the next chip or something like that. Um, the next generation, you know, what extended this to 1,024 chips in eight racks. And uh, next generation actually used uh, 64 racks of 64 chips each. Um, it's actually multiple of these data center rows and gives you 1.1 exaflops of lower precision floating point computation.
呃,所以这意味着你不需要在网络里做任何路由。嗯,所以你可以有非常高速的带宽、非常低成本的连接,因为你只需要走六英寸就到下一块芯片,差不多这样。嗯,下一代,你知道,把这个扩展到了八个机架里的 1024 块芯片。呃,再下一代实际上用了 64 个机架,每个机架 64 块芯片。嗯,这其实是数据中心里好几排机器,能给你 1.1 exaflops的低精度浮点计算能力。
便签引用
13:06
Um, with 4,096 chips. And then the more most recent uh, generation uh, which we uh, disclosed publicly last year, at the end of last year, is the V5 series. It has comes in two variants. One is for sort of uh, inference where you We a pod of 256 chips. Uh, and then the V5P has a lot more memory per chip, much more bandwidth between chips and much more memory much more memory between chips and much more memory bandwidth. And has you know close to half a petaflop per chip of 16-bit floating-point performance and and double that for int eight performance.
嗯,用 4096 块芯片。然后是最新的,呃,一代,呃,我们,呃,是去年公开披露的,去年年底,就是 V5 系列。它有两个版本。一个是面向,呃,推理的,我们有一个 256 块芯片的 pod。呃,然后 V5P 每块芯片的内存要多得多,芯片之间的带宽也大得多,芯片之间的内存也多得多,内存带宽也高得多。而且每块芯片有接近半个 petaflop 的 16 位浮点性能,int8 性能是它的两倍。
便签引用
04从 n-gram 到词向量:简单方法加大数据
13:49
And so one of these pods is also bigger so it's close to 9,000 chips for XLA lots of compute. Okay, so let's talk about language now. We talked about image recognition and speech recognition advances but language is actually one of the areas that I think people are seeing the most change in what computers can do. So I've actually been excited about language models for a while even before neural networks. So I partnered with some people in our Google Translate team to they basically had a really capable system that was um very high quality translations but it was designed for a research contest where you only had to translate like 50 sentences in 2 weeks or something and then you submit your entry. And so it would do like a disk seek for you know 200,000 n-grams that it needed to to look up for every sentence it would translate.
而且这样一个 pod 也更大,接近 9000 块芯片,为 XLA 提供大量算力。好,那我们现在来聊聊语言。我们讲了图像识别和语音识别方面的进展,但我认为语言其实是人们看到计算机能力变化最大的领域之一。其实早在神经网络出现之前,我就对语言模型很感兴趣。所以我和谷歌翻译团队的一些人合作,他们基本上有一套非常强的系统,嗯,翻译质量非常高,但它是为一个研究竞赛设计的,在那种比赛里你只需要在两周内翻译大概 50 个句子之类的,然后提交你的结果。所以它每翻译一个句子,都会做磁盘寻道,去查,你知道,它需要查的 20 万个 n-gram。
便签引用
14:47
And so I said oh well if you have really high quality translations it'd be good to actually bring these into to real practice and so we built a system that would serve an n-gram model basically it kept statistics of how often every five-word sequence occurred in 2 trillion tokens. And that gives you about 300 billion unique five-grams. And then we just would store that in memory on a bunch of machines we'd look them up in parallel for the 100,000 things you need to translate a sentence and we came up with a an new algorithm called stupid backoff that kind of ignored the right mathematical thing to do and did something much simpler. So that when you looked up a 5 g and there was no data there, you would just look up the 4 g that was its prefix and use that if it was there and if it wasn't there, you'd look up the 3 g and so on.
所以我说,哦,既然你们有质量这么高的翻译,把它真正落地到实际应用中会很不错,于是我们构建了一套能提供 n-gram 模型服务的系统,基本上它保存了每个五词序列在两万亿个 token 中出现频次的统计数据。这大概会给你 3000 亿个不同的 five-gram。然后我们就把它存在一堆机器的内存里,并行地查找翻译一个句子所需要的那十万个条目,我们还想出了一个新算法,叫 stupid backoff,它某种程度上无视了数学上正确的做法,而做了些简单得多的事情。就是当你查一个 5-gram 却没有数据时,你就去查它的前缀那个 4-gram,如果有就用它,如果还没有,你就去查 3-gram,以此类推。
便签引用
15:42
And then actually worked reasonably well compared to the fancier Kneser-Ney smoothing, which is what you really want to do, but is actually kind of computationally hard. Um so I mean one lesson from this is simple techniques over large amounts of data are very effective. This has been a lesson throughout my my career that you can actually do very simple things and the data speaks when you do that. Um then my colleague Tomas Mikolov was interested in distributed representation. So instead of representing a word as a sort of a discrete thing, you want to represent it as a very high-dimensional vector. So we're going to represent different words with different say 100-dimensional vectors.
结果它的效果其实相当不错,跟更讲究的 Kneser-Ney 平滑相比也不差,而后者才是理论上你真正想用的做,但实际上在计算上是相当困难的。嗯,所以我想说,从中得到的一个教训是:在大量数据上使用简单的方法非常有效。这是贯穿我整个职业生涯的一个教训——你其实可以做非常简单的事情,而当你这么做时,数据自己会说话。嗯,后来我的同事 Tomas Mikolov 对分布式表示产生了兴趣。也就是说,与其把一个词表示成一种离散的东西,不如把它表示成一个非常高维的向量。所以我们要用不同的、比如说 100 维的向量来表示不同的词。
便签引用
16:29
And through a training process, we're going to try to move words that appear in similar context nearer each other and we're going to try to push apart words that appear in in different context. Um and if you sort of train over a very large amount of data with a relatively simple training objective that says, "Okay, if these things are appear in similar context, push them closer and if they're different, push them apart." And you do that over trillions of tokens, then you end up with really nice properties where in this 100-dimensional space, right?
然后通过一个训练过程,我们会试着把出现在相似上下文中的词拉得更近,同时把出现在不同上下文中的词推得更远。嗯,如果你在非常大量的数据上,用一个相对简单的训练目标来训练,这个目标就是说:“好,如果这些东西出现在相似的上下文里,就把它们拉近;如果它们不同,就把它们推开。”而当你在数以万亿计的 token 上这么做,你最终会得到非常好的性质——在这个 100 维空间里,对吧?
便签引用
17:05
100-dimensional space is a hard thing to wrap your head around. Um but in that high-dimensional space things that are very similar end up near each other. So if you have mountain and hill and cliff, they will all tend to be kind of near each other in that high-dimensional space. Uh so, points in space are interesting, but perhaps more interestingly, directions also are meaningful in this high-dimensional space because there's a lot of different directions you can go in 100 dimensions. And it turns out that it for example, if you look at where king is in this space and you want to get to queen, you go in a certain direction. So, you can compute that by subtracting the vector uh king from queen and that's the direction you go.
100 维空间是一个很难在脑子里想象出来的东西。嗯,但在那个高维空间里,非常相似的东西最终会彼此靠近。所以如果你有“山”、“丘陵”和“悬崖”,它们在那个高维空间里都会倾向于彼此靠得比较近。呃,所以,空间中的点很有意思,但也许更有意思的是,在这个高维空间里方向也是有意义的,因为在 100 维里你可以往非常多不同的方向走。结果发现,举个例子,如果你看“king(国王)”在这个空间里的位置,然后你想走到“queen(王后)”,你需要朝某个特定方向走。所以你可以通过用 queen 的向量减去 king 的向量来算出这个方向,那就是你要走的方向。
便签引用
17:52
Um so, it turns out that king minus queen, that's the direction, is roughly the same as the direction you would go to get from from uh man to woman. And so, directions are meaningful and different directions mean different things. So, going from the present tense of a verb to the past tense is a different direction regardless of what the verb is. Um and so, that says that there's a lot of power in these distributed representations. They're They're encoding a lot of different kinds of information in the 100-dimensional vector that represents the word.
嗯,所以结果就是,king 减 queen——也就是那个方向——大致上和你从“man(男人)”走到“woman(女人)”所要走的方向是一样的。所以,方向是有意义的,不同的方向意味着不同的东西。比如说,从一个动词的现在时走到过去时,是另一个方向,而且不管这个动词是什么,方向都一样。嗯,所以这说明这些分布式表示里蕴含着很大的能量。它们在代表这个词的 100 维向量里编码了非常多种不同类型的信息。
便签引用
05序列到序列、Transformer 与聊天模型
18:31
Then, my colleagues uh Ilya Sutskever and Quoc um developed a model called sequence-to-sequence learning. And so, basically this used a neural network where you would have an input sequence. Uh let's take the case of translation. So, you put in an English sentence one word at a time and the system kind of builds a representation from its own current state plus the new word that it's now seeing and updates that state to have now a in the same way you have the the distributed representations for an individual word, you now have a distributed representation for the sentence you've seen so far.
后来,我的同事 Ilya Sutskever 和 Quoc 开发了一个叫序列到序列学习(sequence-to-sequence)的模型。基本上,这用的是一个神经网络,你会给它一个输入序列。呃,我们就以翻译为例吧。你一次输入一个英文句子里的一个词,系统会根据它自己当前的状态加上它现在看到的新词来构建一个表示,并更新这个状态——就像你对单个词有分布式表示一样,你现在对目前为止看到的这个句子也有了一个分布式表示。
便签引用
19:12
And you can update it with a recurrent neural network called a long short-term memory. And then when you hit an end of sentence marker, you now train the model to spit out the correct the the translation of that sentence. So, you have a bunch of training data, which is like an English sentence and a French sentence that mean the same thing. And you train the model when it sees this English sentence, it should spit out this French sentence. And you just repeat that process over large amounts of paired training data.
而你可以用一种叫长短期记忆(LSTM)的循环神经网络来更新它。然后当你碰到句子结束标记时,你就训练这个模型吐出正确的、也就是这个句子的译文。所以你有一大堆训练数据,就是一个英文句子和一个法文句子,它们的意思是一样的。你训练这个模型:当它看到这个英文句子时,就应该吐出这个法文句子。然后你只要在大量的成对训练数据上重复这个过程就行了。
便签引用
19:42
And sure enough, um you can use a neural encoder over this input sequence to initialize the state, which kind of gets you into the I've now absorbed the input sentence and I now want to decode a word at a time the correct translated sentence. And we're going to use that to initialize the state of the neural decoder. You scale this up and it works. You get major improvements in translation accuracy. Um Oriel and Quoc then published a workshop paper showing that instead of translation, you could use context for a multi-turn conversation.
果然,嗯,你可以用一个神经编码器处理这个输入序列来初始化状态,这大致就把你带到了“我已经吸收了输入句子,现在我要一个词一个词地解码出正确的译文句子”这个阶段。然后我们就用它来初始化神经解码器的状态。你把规模扩大,它就奏效了。翻译准确率有了重大提升。嗯,然后 Oriol 和 Quoc 发表了一篇 workshop 论文,说明除了翻译之外,你还可以把上下文用在多轮对话里。
便签引用
20:22
So, basically, the sequence of interactions you've had with one person or you know, one party and then a computer model uh responding and then the other person or the person then utters another response and and there's multiple turns, that is your context. Previous previous multi-turn interactions. And then you can train it to generate a good reply. In that in the context of, you know, the multiple turns of things that happened before. And it's the same model, basically. It's a sequence-to-sequence model, but now the sequence is initialized with the context of all the of the conversational turns that have happened.
基本上就是,你和某一个人、或者说某一方之间的交互序列,然后计算机模型做出回应,接着另一个人、或者说那个人又说了一句,就这样有很多轮,这些就是你的上下文。也就是之前的多轮交互。然后你可以训练它生成一个好的回复。在之前发生过的那些多轮内容的上下文中生成回复。而它基本上就是同一个模型。它是一个序列到序列模型,只不过现在这个序列是用之前发生过的所有对话轮次的上下文来初始化的。
便签引用
20:58
And it's possible then to have effective multi-turn interactions using a neural language model. Which is pretty neat. Um then a collection of other Google researchers plus an intern uh came up with a model called the Transformer. So, remember I said in this model, this is a recurrent model. So, you have some state, you take the next token, and you do some processing to update the the the new state to have absorbed this token, and then you go on with that new state to absorb another token and update the state again. So, that's a very sequential process, right? Because in order to absorb the the third word, you need to have done the processing for the second word. In order to to have done that for the second word, you need to have done the processing for the first word.
这样就有可能用一个神经语言模型进行有效的多轮交互。这挺妙的。嗯,接着,另外一群谷歌的研究人员加上一位实习生提出了一个叫 Transformer 的模型。回想一下我刚才说的那个模型,那是一个循环模型。你有某个状态,你取下一个 token,做一些处理来更新出一个已经吸收了这个 token 的新状态,然后你带着这个新状态继续去吸收下一个token,再更新状态。所以那是一个非常串行的过程,对吧?因为为了吸收第三个词,你必须已经完成了对第二个词的处理。而为了完成第二个词的处理,你必须已经完成了对第一个词的处理。
便签引用
21:49
That's not so great. Like in computers, we like to do things in parallel, not in sequence, if we can if we can get away with it. And so, what this model did was say, "We're going to just process a bunch of data in parallel, all the words in this input, and then we're going to attend to different pieces of it, rather than trying to have just a single state that we update sequentially going through the words." Um and what that said was don't try to force that state into a single distributed representation. Just save all the representations of all the, you know, tokens or words that you've seen, and then attend to them. Like pay attention to the parts that make sense to focus on when you're doing this translating this part of the sentence or translating translating that part of the sentence.
这不太好。在计算机里,只要能做到,我们更喜欢并行地做事,而不是串行地做。所以这个模型做的事情就是说:“我们干脆并行处理一大批数据,也就是这个输入里的所有词,然后我们对其中不同的部分做注意力,而不是试图只维护一个随着词一个个往下走而顺序更新的单一状态。”嗯,这话的意思就是:别硬要把那个状态压进一个单一的分布式表示里。把你看到的所有 token 或者说所有词的表示都保存下来,然后对它们做注意力。也就是说,当你在翻译句子的这一部分、或者翻译那一部分时,去关注那些值得聚焦的部分。
便签引用
22:35
And you get higher accuracy with 10 to 100 x less compute. So, remember I said all that stuff about computer hardware improving and specialized hardware? You know, that's giving us, you know, large significant improvements over time, but we're also seeing algorithmic improvements like this uh also multiplying together with those improvements. And so, you're seeing now the ability to train through algorithmic advances plus machine learning hardware much larger scale models and you know much more capable models because of that.
这样你能用少 10 到 100 倍的算力得到更高的准确率。还记得我刚才讲的那些关于计算机硬件进步和专用硬件的内容吗?你知道,那随着时间推移给了我们很大、很显著的提升,但我们同时也看到像这样的算法上的进步,呃,这些进步和硬件的提升是相乘叠加的。所以现在你看到,靠算法上的进展加上机器学习硬件,我们能训练规模大得多的模型,也因此得到能力强得多的模型。
便签引用
23:12
Um and so then a group of people uh decided to train uh scale up and train on conversational style data using a transformer model instead of a recurrent model and that gave you know quite good results and in particular a way of evaluating those so that it's uh both sensible in what it responds and also specific. You don't want your chatbot to be overly vague like yeah, it's nice. You want it to actually say something sensible in response to to what you you interacted with it because that makes it more engaging and useful.
嗯,接着有一群人决定扩大规模,用 Transformer 模型而不是循环模型,在对话风格的数据上做训练,结果相当不错,特别是还有一套评估方式,让它的回应既要合理(sensible),又要具体(specific)。你不希望你的聊天机器人过于含糊,比如就说“是啊,挺好的”。你希望它针对你跟它说的话真正说点有意义的东西,因为那样才更吸引人、更有用。
便签引用
23:50
Okay. So, I talked about some of these but there's been a progression of neural language models and also a progression of work in neural chatbots. Uh you know, a neural conversational model uh Meena uh ChatGPT from OpenAI uh Bard which we released about a year ago uh at Google uh and then a progression of neural language models. So, the sequence-to-sequence work I talked about GPT-2 from OpenAI which was you know, some of these have parameter counts which you can think of as a rough sense of the scale of the model. So, 1 and 1/2 billion parameters in 2019, the T5 work from Google from some colleagues of mine uh 11 billion parameters, you know, very very capable there. Oh, the transformer work I should mention that underlies a lot of these like the T here and the T here that's all for transformer. So, people have now really seen the advance in the transformer model and architecture that 10 to 100x improvement in computation and really move to using that as the the basis of these these large language models.
好,其中一些我已经讲过了,但神经语言模型有一条发展脉络,神经聊天机器人也有一条发展脉络。呃,比如神经对话模型(neural conversational model)、Meena、OpenAI 的 ChatGPT、还有我们大概一年前在谷歌发布的 Bard,然后是神经语言模型的发展脉络。比如我讲过的序列到序列的工作、OpenAI 的 GPT-2——你知道,其中有些是有参数量的,你可以把参数量粗略理解为模型规模的一个尺度。所以 2019 年是 15 亿参数,谷歌我一些同事做的 T5 工作呃是 110 亿参数,非常非常有能力。哦,我该提一下,Transformer 的工作是很多这些模型的基础,比如这里的 T 和这里的 T,都是指 transformer。所以,人们现在真的看到了 Transformer 模型和架构带来的进步——那 10 到 100 倍的计算改进——于是真正转向把它作为这些大语言模型的基础。
便签引用
06Gemini:原生多模态与训练基础设施
24:57
Uh GPT-3 uh Gopher from uh uh some DeepMind colleagues, Palm from Google Research, uh Chinchilla from DeepMind, Palm 2 from Google Research, and then GPT-4 from OpenAI, and then Gemini, which is the the project I co-lead with my colleague Oriol Vinyals. Uh we have a large collection of people in lots of different uh research offices working on building capable uh multimodal models. So, one of the things we wanted to do was move from not just a language-based model that understands text, but one that can deal with all the different modalities simultaneously. So, you can feed it, you know, text plus an image, or, you know, audio plus some text, and ask it to do something, and it will be able to sort of uh fluently and coherently deal with whatever kinds of modalities you want to do give it. So, our goal when we started this project about a year ago was train the world's best multimodal models and use them all across Google.
呃,GPT-3、DeepMind 一些同事做的 Gopher、Google Research 的 PaLM、DeepMind 的 Chinchilla、Google Research 的 PaLM 2,然后是 OpenAI 的 GPT-4,再然后是 Gemini,这是我和我同事 Oriol Vinyals 共同领导的项目。呃,我们有一大批人分布在很多不同的研究办公室,一起构建能力强大的多模态模型。所以我们想做的一件事,就是不只停留在一个理解文本的语言模型,而是做一个能同时处理各种不同模态的模型。所以,你可以给它输入文本加图像,或者音频加一些文本,让它做点什么,它就能流畅、连贯地处理你想给它的任何种类的模态。所以,大概一年前我们启动这个项目时的目标就是:训练出世界上最好的多模态模型,并把它们用在谷歌的各个地方。
便签引用
26:03
Um and so there's a blog about Gemini, there's uh you know, a website you can go to, and there's a tech report um by the Gemini team, of which I'm a proud member. So, Gemini was really multimodal from the beginning. So, one of the things we did, as I mentioned, we didn't want it to just deal with text, we wanted to deal with with images and video and audio, and we turn that into a sequence of tokens that we then train a transformer-based model on. Uh and then we have a couple of different decoding paths. One we train to generate tokens uh and that are textual, and then the other we initialize the decoder with the state uh that the transformer has learned, and then we can sort of go from that state to a full uh you know, set of pixels for an image.
嗯,关于 Gemini 有一篇博客,还有一个你可以去看的网站,另外 Gemini 团队还有一份技术报告,我很自豪是这个团队的一员。所以 Gemini 从一开始就是真正多模态的。正如我提到的,我们做的一件事就是,我们不想让它只处理文本,我们想让它处理图像、视频和音频,我们把这些转换成一个 token序列,然后在上面训练一个基于 Transformer 的模型。呃,然后我们有几条不同的解码路径。一条我们训练它生成 token,也就是文本;另一条我们用 Transformer 学到的状态来初始化解码器,然后我们就能从那个状态走到一整张图像的完整像素集合。
便签引用
26:52
Um, and we support interleaving these sequences of text. It's not like you give it an image uh, text input and an image input, you can sort of interleave them. For video, you might put in, you know, a video frame and some text describing that and then another video frame and some text or the closed captions of the audio that's being said in the text. Uh, and then have the transformer kind of use the the fact that it's been exposed to all these modalities during training to now build common representations across all the different modalities you want to give it.
嗯,而且我们支持把这些序列和文本交错排列。不是说你只给它一个图像输入和一个文本输入,你可以把它们交错着放。对于视频,你可能放进去一个视频帧和一些描述它的文字,然后再一个视频帧和一些文字,或者是文本里那段音频的字幕。呃,然后让 Transformer 利用它在训练中接触过所有这些模态这一点,去为你想给它的所有不同模态建立起共通的表示。
便签引用
27:26
Um, we have a few different sizes. So, the V1 generation of Gemini comes in three different sizes. So, Ultra is the kind of the largest scale uh, and most capable model we have. Pro is like a good size for running in a data center uh, and we use that in a lot of different product contexts. Uh, so, like our our Bard product, which has now been renamed Gemini, uh, confusingly. Um, uh, is uh, running on uh, the Pro model or the Ultra model we just announced last week. Uh, and then the Nano model, you actually want a lot of these machine learning models to be able to run on device. So, on a small phone or a laptop. And the Nano model is very efficient for for doing that and and fits quite reasonably. You can quantize these things to make them even smaller uh, and so on.
嗯,我们有几种不同的规模。Gemini 的第一代有三种不同的尺寸。Ultra 是我们规模最大、能力最强的模型。Pro 是适合在数据中心里运行的一个不错的尺寸,我们把它用在很多不同的产品场景里。呃,比如我们的 Bard 产品,它现在改名叫 Gemini 了,呃,挺让人混淆的。嗯,呃,它跑的是 Pro 模型,或者我们上周刚发布的 Ultra 模型。呃,然后是 Nano 模型——你其实很希望很多这类机器学习模型能在设备上运行,比如在一部小手机或者笔记本电脑上。而 Nano 模型在这方面非常高效,塞得进去也相当合理。你还可以对它们做量化,让它们变得更小,等等。
便签引用
28:17
Um, so, one of the things about our training infrastructure is we we wanted to be able to have a very scalable fabric that can deal with, you know, you specify very high-level uh, description of the computation you want and then have a a a system that then maps that computation onto the available hardware you have. And so, I mentioned we have these pods. And so, for example, you might describe your computation as I have these two parts that I care about. I don't care where you put them and not let the underlying pathways software system that we've built decide where to put them. So, it might decide to put this part on one pod and this part on another pod and then it knows where the chips are located and what the topology and bandwidth is between them. So, when this chip needs to communicate with that one, they'll use this link, the very high-speed network I mentioned. And when you need to have, you know, this part of the model communicate over here, then it will go up to the data center network,
嗯,关于我们的训练基础设施,有一点是我们希望有一个非常可扩展的架构,能够处理这样的情况:你用非常高层的方式描述你想要的计算,然后有一个系统把这个计算映射到你手头可用的硬件上。我提到过我们有这些 pod。所以,举个例子,你可能这样描述你的计算:我关心的是这两个部分。我不在乎你把它们放在哪儿,让我们构建的底层 Pathways 软件系统去决定把它们放在哪里。所以它可能决定把这一部分放在一个 pod 上,把这一部分放在另一个 pod 上,然后它知道芯片位于什么位置、它们之间的拓扑和带宽是什么样的。所以当这块芯片需要和那块通信时,它们会走这条链路,也就是我提到的那个极高速的网络。而当你需要让模型的这一部分和那边通信时,它就会走到数据中心网络上,
便签引用
29:15
which is, you know, much less bandwidth to send data from one place to another, but it's kind of seamlessly happens and the machine learning researcher or developer doesn't have to worry about it uh from that perspective other than just understanding there are different performance characteristics. um So, one of the things about training large-scale models is as you scale up, you know, failures will happen. You'll a machine will die, one of the TPU chips will, you know, overheat and and start to malfunction in some way.
也就是说,从一个地方向另一个地方传输数据所需的带宽要小得多,但这一切都是无缝完成的,机器学习研究员或者开发者不必为此操心,从这个角度来说,他们只需要理解不同的性能特征差异就行了。嗯,关于训练大规模模型,有一点是,随着规模扩大,你知道,故障是一定会发生的。会有机器宕机,某块 TPU 芯片会过热,然后开始以某种方式出现故障。
便签引用
29:50
Uh and so, minimizing failures is really important. Uh some of those failures you want to minimize can be almost human self-inflicted things. So, for example, we had a a sweeping way of upgrading kernels on our machines, which is a perfectly fine approach if those machines are kind of independent computations, but if they're all part of the same, you know, thousand machine computation, you actually would prefer to take the machines down, upgrade all thousand kernels simultaneously and then bring them back up rather than having rolling failures throughout. So, we kind of optimize some of our repair and and upgrade processes.
所以,把故障降到最少真的很重要。而其中一些你想避免的故障,几乎可以说是人为自找的。比如说,我们曾经有一种一次性横扫式升级机器内核的做法,如果那些机器某种程度上是各自独立的计算单元,但如果它们都属于同一个,你知道的,上千台机器的同一次计算,那你其实更愿意把这些机器停下来,同时升级全部一千个内核,然后再把它们启动起来,而不是让故障在整个过程中一台接一台地滚动发生。所以我们对修复和升级流程做了一些优化。
便签引用
30:27
Uh but you also, once you've done that, you then want to minimize the time to recover uh because the faster you you can recover, the the sooner you can actually be making uh useful forward progress. And so, we have a a metric we call goodput, which is the percentage of time that model training is actually making useful forward progress as opposed to recovering from a checkpoint or waiting for some other part of the system to be started. And one of the things we we use is rapid recovery from other copies of the model state from memory in those other machines rather than going to a distributed file system to recover from a checkpoint. And that, you know, makes the recovery time, you know, a matter of a few 5 to 10 seconds rather than several minutes.
呃,但你在做完这些之后,还要尽量缩短恢复所需的时间,因为你恢复得越快,就越早能真正取得有用的进展。所以我们有一个指标叫 goodput(有效吞吐),也就是模型训练真正在取得有用进展的时间百分比,而不是在从检查点恢复、或者等待系统其他部分启动。我们用到的手段之一是快速从其他机器内存中保存的模型状态副本来恢复,而不是去分布式文件系统里读检查点来恢复。这样一来,恢复时间就变成了大概 5 到 10 秒,而不是好几分钟。
便签引用
31:14
Uh in terms of training data, you know, we want this model to be multimodal, so we want to train it on a large collection of web documents, you know, various kinds of books, uh various kinds of code in lots of different programming languages, plus images, audio, and video data. Um we have heuristics for filtering those data sets. Uh some of them are kind of handwritten heuristics. Some are model-based classifiers of like do we think this is a high-quality document in various ways. Um the final mixtures uh of the training data are determined through ablations on smaller models, so we'll run smaller-scale models with different mixes. Should we use 32% code or 27% code and then evaluate the performance on a wide range of metrics to to better understand that. Uh we've done some things like increase the weight of domain-relevant data towards the end of training, so we want to enrich it with say more multilingual data towards the end of training uh in order to make the multilingual capabilities improve.
呃,在训练数据方面,我们希望这个模型是多模态的,所以我们想用大量的网页文档来训练它,还有各种各样的书籍、呃各种编程语言写的各类代码,再加上图像、音频和视频数据。嗯,我们有一些启发式规则来过滤这些数据集。呃,其中一些是手写的启发式规则,还有一些是基于模型的分类器,比如判断我们是否认为这在各个维度上都是一份高质量的文档。嗯,训练数据的最终配比是通过在小模型上做消融实验来确定的,所以我们会用不同的配比跑小规模的模型。我们该用 32% 的代码还是 27% 的代码,然后在一系列广泛的指标上评估效果,以便更好地理解这一点。呃,我们也做过一些事情,比如在训练接近尾声时提高领域相关数据的权重,所以我们希望在训练后期用更多的多语言数据来强化它,以便让多语言能力得到提升。
便签引用
32:15
I do think data quality is a really interesting and important research area. Um and I think it's we've seen that, you know, having really high-quality data makes a huge difference in the performance of the model on tasks you care about. And so that means, you know, in some sense that's as important or even more important in some cases than the actual say model architecture you're using. Um and and so, you know, I think it's a pretty important area for future research. You know, having the ability to learn curriculums automatically seems important. Uh identifying high-quality examples and low-quality examples and so on.
我确实觉得数据质量是一个非常有意思、也非常重要的研究方向。嗯,我认为我们已经看到,拥有真正高质量的数据,会让模型在你关心的任务上的表现产生巨大差别。所以这意味着,某种意义上,在有些情况下,它和你实际使用的模型架构一样重要,甚至更重要。所以我觉得这是一个对未来研究相当重要的方向。比如说,让模型能够自动学习课程安排(curriculum)似乎很重要。还有识别高质量样本和低质量样本等等。
便签引用
07思维链提示与多模态解题示例
32:58
And then there's been a bunch of advances in not only training these models, but also how do you elicit the best qualities of the model? How do you actually ask questions in a way that causes the model to be able to answer questions in a more effective way? Um so, for example, asking models to show their work improves the accuracy of the model and also the interpretability. Uh and so, some of my colleagues came up with a technique called chain-of-thought prompting. Uh so, you if you remember back to third grade math class, your your teacher would always encourage you to show your work, right? And the reason they wanted to do that is both to see your thought process in getting to the answer, but also to kind of encourage you to think about, you know, what are the next steps and how am I going to break this complicated problem down into a smaller set of of steps.
然后除了训练这些模型之外,在如何激发出模型最好的那一面上,也有很多进展。你要怎么以某种方式去提问,才能让模型更有效地回答问题?比如说,要求模型展示它的推演过程,能提升模型的准确率,也提升可解释性。我的一些同事就提出了一种技术,叫思维链提示(chain-of-thought prompting)。如果你还记得小学三年级的数学课,老师总是鼓励你把过程写出来,对吧?他们这么做的原因,一是想看到你得出答案的思考过程,二也是为了鼓励你去想,下一步该怎么走,我该怎么把这个复杂的问题拆解成一小步一小步。
便签引用
33:50
And so, if you ask a model, you know, give it you usually you give it an example of a kind of question you want it to answer and then the actual answer for that question, and then you ask it a new question. And then you ask it to answer that question. And so, here's an example of a question and and then the way the model was taught to respond is just to figure the answer out and give it. Um and then there's a more there's a different question and the model output actually says the answer is 50, which is wrong.
所以如果你去问模型,通常你会先给它一个你希望它回答的那类问题的例子,再给出那个问题的答案,然后你再问它一个新问题,让它回答这个问题。这里就是一个问题的例子,而模型被教会的回答方式就是直接把答案算出来给你。然后换一个不同的问题,模型输出的答案是 50,这是错的。
便签引用
34:20
But if you instead ask the model and show demonstrate to it, you know, how do you show your work? Um and you say, "Okay, well, Sean started with five toys. If he got two toys each, then that is four more toys. 5 plus four is nine, so the answer is nine. That's That's the work your your third grade math teacher would be really proud of. Um and more importantly, if you do that, the model actually, mhm you know, elicits these sort of more incremental steps to get to the answer, and it gets it right.
但如果你换一种方式问模型,并且示范给它看,怎么把过程写出来。你说:好,肖恩一开始有五个玩具。如果他每次拿到两个玩具,那就是多了四个玩具。5 加 4 等于 9,所以答案是 9。这就是那种会让你三年级数学老师特别欣慰的解题过程。更重要的是,如果你这么做,模型实际上嗯,就会引出这些更循序渐进的步骤来得出答案,而且它答对了。
便签引用
34:52
Because it's now had longer to think about the the steps to basically more time to think about through the steps of of uh getting to the right answer. Uh and so that's actually and it's a pretty dramatic effect, right? These two These two lines are the same underlying models at different scales. And what you see, these are two different math sort of mathematical oriented benchmarks. Uh the one on the right is sort of eighth grade math problems, and this one is, you know, a bunch of arithmetic problems. What you see is the quality of responses is pretty bad when you just give it sta- standard prompting, but at some point the model scale becomes large enough that all of a sudden when you ask it with chain of thought prompting, your accuracy shoots up quite a lot.
因为它现在有更长的时间去思考这些步骤,基本上就是有更多时间一步步想清楚,从而得到正确答案。这个效果其实相当显著,对吧?这两条线是同一批底层模型在不同规模下的表现。你看到的,是两个偏数学类的基准测试。右边这个大概是初中八年级的数学题,而这个是一堆算术题。你会看到,如果只用标准提示,回答的质量相当差,但到了某个点,模型规模变得足够大,突然之间,当你用思维链提示去问它,准确率就大幅飙升。
便签引用
35:36
So that says there's a really interesting science of how do you ask these models questions in a way that actually, you know, both makes them more interpretable and also more likely to give you the right answer. Okay. So let's talk about multimodal reasoning in a Gemini model, and I think a an example is a nice one. It's a good way to understand what this model can do. So here's the prompt, you know, here's a solution to a physics problem by a student, and then there's just a uh a picture of the problem and the student's written out answer in kind of handwriting.
所以这说明,怎么向这些模型提问,本身就是一门很有意思的科学——既让它们更可解释,也更有可能给你正确的答案。好,我们来聊聊 Gemini 模型里的多模态推理,我觉得举个例子挺好的。这是理解这个模型能做什么的好方式。这是提示词:这里是一个学生给出的物理题解答,然后有一张题目的图片,以及学生手写的解题过程。
便签引用
36:11
And then the rest of the prompt says, "Try to reason about the question step by step." That's again kind of the chain of thought prompting style thing. "Did the student get the correct answer? If it's wrong, please explain what's wrong and solve the problem, and make sure to use LaTeX for math and round off the final answer to two decimal places. And so that's the input. This kind of like hokey image of handwriting and a skier going down a slope and all this kind of stuff, conservation of energy, blah blah blah. And then this is the output of the model.
提示词接下来说:请一步步推理这个问题。这又是那种思维链提示的风格。学生的答案对吗?如果不对,请说明错在哪里,并把题目解出来,数学部分务必用 LaTeX,最终答案保留两位小数。这就是输入。这张有点粗糙的手写图片,一个滑雪的人从斜坡滑下来,诸如此类,能量守恒等等。然后这是模型的输出。
便签引用
36:40
So the student get didn't get the correct answer. Student made a mistake in the calculation of potential energy at the start of the slope. Um the potential energy at the start is given by MGH. Student used the length of the slope, I guess that's the hypotenuse, instead of the height in the calculation. Correct solution is, you know, it means the total blah blah blah. Therefore we can write if this is this is actually in LaTeX, but we've rendered it for your reading convenience. Uh and substituting the values, there it is. It worked the problem out to two decimal places.
学生没有得到正确答案。学生在计算斜坡起点的势能时出错了。起点的势能应该是 MGH。学生在计算中用了斜坡的长度,我想那是斜边,而不是高度。正确解法是,也就是说总的等等等等。因此我们可以写出——这部分其实是 LaTeX,但我们把它渲染出来方便你阅读。然后代入数值,就是这个结果。它把题目算到了两位小数。
便签引用
37:11
So think about what this means. All of a sudden we can give models like kind of multimodal input, you know, a complex you know, a picture of a whiteboard and a problem and ask it to do something and it can do it. You know, it's not always going to do it right, but but it can. And uh this can be an amazing educational tool. So think about, you know, a student trying to work things out on their own and they're taking pictures of their solution and and you know, the system is kind of helping helping them figure out what they did wrong. Um we know that individualized tutoring has outcomes that are two standard deviations higher when you have a one-on-one human tutor for education than when you have a much uh broader scale classroom setting. So could we get close to that in terms of having individualized uh tutoring? I think I think that that possibility is within our collective grasp.
想想这意味着什么。突然之间,我们可以给模型这种多模态的输入,比如一张复杂的白板照片、一道题目,然后让它做点什么,它就能做到。当然它不会每次都做对,但它做得到。而这可以成为一个了不起的教育工具。想象一下,一个学生在自己琢磨题目,他把自己的解答拍下来,然后系统就在旁边帮他弄清楚自己哪里错了。我们知道,在教育上,一对一的人类导师做个性化辅导,其效果比大班课堂高出两个标准差。那我们能不能在个性化辅导上做到接近那个水平?我认为这个可能性就在我们大家伸手可及的范围之内。
便签引用
08评测:32 项基准与 Elo 排行榜
38:08
Um okay, so evaluation, you know, that was kind of a uh qualitative example of Gemini's capabilities, but it's also good to look at how it compares on a bunch of different uh characteristics. Um evaluation, you know, really helps us identify the model strengths and weaknesses, helps understand, you know, is training going well? So, we're constantly evaluating these metrics as we're training the model. Um, helps us make decisions about what to change. Is our math performance, you know, lower than we would hope? And so, maybe we should enrich the training mixture with more math-oriented data.
好,接下来说评估。刚才那是对 Gemini 能力的一个定性例子,但也值得看看它在一系列不同维度上的比较。评估确实能帮我们识别出模型的优势和劣势,帮助理解,你懂的,训练进展顺利吗?所以在训练模型的过程中,我们会不断地评估这些指标,去评估模型。嗯,这能帮我们决定该改什么。我们的数学表现,你懂的,是不是比预期要差?如果是这样,那也许我们该在训练数据配比里加入更多偏数学的数据。
便签引用
38:42
But, what will that do to multilingual performance? Um, there's a lot of complicated trade-offs. Some of which you make at the beginning of training, some of which you're kind of monitoring online and trying to make, uh, principled or seat-of-the-pants decisions, I guess. Um, and helps you compare the capabilities to other models and systems. Um, and so, the high- highest-level summary is, you know, we looked at 32 academic benchmarks. Uh, and the Gemini Ultra model exceeded the state-of-the-art performance in 30 of the 32.
但那样又会对多语言能力产生什么影响呢?嗯,这里面有很多很复杂的权衡。有些是你在训练一开始就要做的决定,有些则是你在训练过程中一边监控一边做的,呃,有的是有理有据的决策,有的大概就是凭直觉拍脑袋吧。嗯,同时它也能帮你把模型能力和其他模型、其他系统做对比。嗯,所以最高层面的总结就是,我们考察了 32 个学术基准测试。呃,Gemini Ultra 模型在其中 30 个上超过了此前的最好成绩(state-of-the-art)。
便签引用
39:19
Um, and so, if we look at and delve into some of these in depth, uh, there's a bunch of text-oriented or general reasoning or math-oriented benchmarks. Um, and if you compare Gemini Ultra with GPT-4, which is generally the the prior state of the art in most of these problems, uh, what you see is, uh, the ones in blue is the state-of-the-art. And so, we we state-of-the-art on seven of the eight. Uh, the 90% on MMLU is interesting because this is a very broad uh, set of questions in 57 different subjects. You know, chemistry, math, uh, international law, philosophy.
嗯,那我们再深入看看其中一些,呃,有一批是偏文本的、通用推理的,或者偏数学的基准。嗯,如果把 Gemini Ultra 和 GPT-4 对比,GPT-4 在大多数这类问题上基本代表了此前的最好水平,呃,你会看到,蓝色标出的就是最好成绩。所以我们在八项里的七项上都拿到了最好成绩。呃,MMLU 上的 90% 挺有意思的,因为这是一个覆盖面非常广的题库,呃,涵盖 57 个不同学科。你知道的,化学、数学,呃,国际法、哲学。
便签引用
39:59
Um, and the group that put together the the benchmark, um, measured human expert-level performance at 89.6%, I think, or maybe 89.8. And so, this actually exceeds human expert-level performance in these 57 categories, uh, which is which is quite nice. We're happy with that. Uh, And then there's a bunch of coding related ones down here and math oriented ones here. Yeah, I mentioned the 90%. So if you look at image understanding benchmarks, you know, these are now getting into the multimodal aspect of this, you know, we got state of the art results on eight of eight benchmarks ranging from one of the nice things was this benchmark came out a week before we published our paper.
嗯,而制作这个基准的那个团队,嗯,测出的人类专家水平是 89.6%,我记得是,或者 89.8。所以这个成绩其实是在这 57 个类别上超过了人类专家的水平,呃,这个,这个还是挺不错的。我们对此挺满意。呃,然后下面这里有一批和编程相关的,这边还有一些偏数学的。对,90% 那个我刚提过了。那如果看图像理解类的基准,你懂的,这就开始进入多模态的部分了,在这方面,我们在八个基准上全部拿到了最好成绩,其中有一个特别有意思的是,这个基准是在我们发论文的前一周才刚发布的。
便签引用
40:45
And we'd never seen it before. So we our valve team quickly added this benchmark to our valve set and discovered that we exceeded the state of the art results by a, you know, reasonable margin. It's nice. It's always nice when you have a a benchmark you've never seen before and you do well on it because you're always worried about like leakage of of test date training data into the test set and so on. If you look at video understanding, you know, again, the multimodal capabilities of of this model really really shine pretty well. State of the art on six of six of six benchmarks including, you know, the important English cooking video captioning benchmark and video question answering and so on.
我们之前从来没见过它。所以我们的评测团队很快把这个基准加进了我们的评测集,然后发现我们超过了之前的最好成绩,而且,你懂的,领先幅度还不小。这挺好的。当你碰上一个从没见过的基准还能做得好,这总是让人高兴,因为你总会担心训练数据泄漏到测试集里之类的问题。从视频理解方面来看,这个模型的多模态能力真的非常非常突出,表现相当不错。在六个基准测试里全部达到了业界最佳水平,包括那个很重要的英文烹饪视频字幕生成基准,还有视频问答等等。
便签引用
41:34
And if you look at audio, you know, the word error rates of of this on a bunch of four different uh uh public uh uh speech recognition benchmarks as well as a speech translation benchmark. State of the art on five of five and multilingual capabilities are are quite good. We're state of the art on four of the five. So yeah, first, I hope you appreciate our valve team because this is a tremendous amount of work to evaluate these models and really understand the capabilities in this level of detail and that's uh pretty pretty awesome. Um and it does give us a a nice firm idea that that Gemini model is pretty capable.
再看音频,它在一系列四个不同的公开呃呃嗯,语音识别的基准测试,还有一个语音翻译的基准测试。五项里有五项达到了业界最佳水平,多语言能力也相当不错。五项里有四项达到了业界最佳。所以,首先,我希望大家能体会到我们评测团队的付出,因为要评测这些模型、要在这么细的粒度上真正理解它们的能力,工作量是非常大的,这一点挺了不起的。嗯,而且这确实让我们心里有了一个比较扎实的判断,那个 Gemini 模型能力相当强。
便签引用
42:16
Uh, and we also have, you know, measurements of Pro and Nano in the paper. Okay. So, these large transformer models can actually generate uh, surprisingly coherent conversations, uh, which is kind of the the evolution of that sort of neural conversational model and then the transformer-based versions of those. Um, so if you look at bard.g- I guess I have to update my slides. So, uh, this will be gemini.google.com. Um, so that I was actually preparing a talk before we were using Gemini models, uh, a few months ago in Bard, uh, Gemini. Uh, and I said, you know, reverse the letters of hot chips and tensor processing units for me. So, you know, just to show what these models can do and it says, "Sure, the reverse strings are," you know, there they are.
呃,我们在论文里也有对 Pro 和 Nano 的测量数据。好。所以说,这些大型 transformer 模型确实能生成呃,连贯得让人惊讶的对话,这算是那类神经对话模型的一种演进,然后又有了基于 transformer 的版本。嗯,所以如果你去看 bard.g——我想我得更新一下我的幻灯片了。所以,呃,这里应该是 gemini.google.com。嗯,因为我其实是在我们用上 Gemini 模型之前就准备了这个演讲,嗯,几个月前在 Bard,呃,Gemini 上。呃我说,你知道的,帮我把 hot chips 和 tensor processing units 这两个词的字母倒过来。你知道,就是想看看这些模型能做到什么程度,然后它说:“好的,倒过来的字符串是”,你知道,就给出来了。
便签引用
43:04
Great. But then it went on to say, "I can also do this for you in Python." Here's the code. Uh, you know, define reverse string as a function. Here's the string. Print reverse string of that. Reverse string of that. Uh, use code with caution. I always recommend that. Um, and then it also goes on to explain it. Code first defines function called reverse string. Takes a string as input. Returns the reverse string. Function works by looping through the string. Code then prints reversing. And it's always going to be helpful. Is there anything else I can help you with?
很好。但它接着又说:“我还可以用 Python 帮你做这件事。”“代码在这里。”呃,你知道,定义一个叫 reverse string 的函数。这是那个字符串。打印这个字符串的倒序。打印那个字符串的倒序。呃,谨慎使用代码。我一向推荐这么做。嗯,然后它还接着解释了一遍。代码首先定义了一个叫 reverse string 的函数。接收一个字符串作为输入。返回倒序后的字符串。这个函数是通过遍历字符串来实现的。然后代码打印出倒序结果。而且它总是会很热心地问:还有什么我可以帮您的吗?
便签引用
43:38
Um, so this is pretty amazing, right? Like, the fact that someone asked a question and it did what it was asked, but then it also said, "By the way, here's this thing called programming and some Python code, uh, and here's what it Here's how you would do this, uh, write writing, you know, code to do this." Uh, I I think that's pretty cool. And then again, a real educational opportunity. Um, is there anything else you can help me with? Sure, tell me more about TPUs. So, you know, the model has a fair amount of world knowledge. It knows that TPUs are, you know, basically I I've told you a lot of this. They're specialized hardware processors developed by Google to accelerate machine learning. They can help improve the efficiency and performance.
嗯,所以这挺惊人的,对吧?就是说,有人问了一个问题,它照做了,但接着它还说:“顺便一提,还有个东西叫编程,这里有一段 Python 代码,呃,这就是你会怎么做,呃,写、写出代码来完成这件事。”呃,我我觉得这挺酷的。而且再一次,这是一个真正的教育机会。嗯,还有别的我能帮上忙的吗?当然,再跟我讲讲 TPU 吧。你知道,模型有相当多的世界知识。它知道 TPU 是——基本上这些我已经跟你讲过很多了。它们是谷歌开发的专用硬件处理器,用来加速机器学习。它们可以帮助提升效率和性能。
便签引用
44:20
Here are some of the benefits: faster training inference. I hope again I hope this helps. Um and and you know, I think one of the kind of fun things about these chatbots they can have different kinds of personalities. Like Bard is sort of your helpful friend and will help you answer lots of questions. Uh so we actually put Gemini Pro in Bard uh in {slash} Gemini last month and there's a public site called LMSYS that can evaluate different chat agents cuz there's now a lot of different chatbots in the world and the way they do that is they get users to write their own prompt. They pick two random chatbots that they have configured in their system and then they send the query to both of them, the prompt to both of them, show and then show the output anonymized. So, you just say which is better, the left or the right.
以下是一些好处:更快的训练和推理。我还是那句话,希望这对你有帮助。嗯,我觉得这些聊天机器人有意思的一点是,它们可以有不同类型的性格。比如 Bard 有点像你乐于助人的朋友,会帮你回答很多问题。呃,我们上个月其实把 Gemini Pro 放进了 Bard,也就是现在的 Gemini。有一个叫 LMSYS 的公开网站可以评测不同的聊天智能体,因为现在世界上有很多不同的聊天机器人。他们的做法是让用户自己写提示词,然后从系统里配置好的机器人中随机挑两个,把这个查询、这个提示词同时发给两边,再把输出匿名地展示出来。所以你只需要说哪个更好,左边还是右边。
便签引用
45:21
Um and then from that you can compute what's called an Elo score. So, Elo was a I believe a Hungarian mathematician who was trying to develop ways to rank chess players and so when you have a tournament um basically you get more Elo points when when you defeat one opponent and so you can in a tournament setting compute an Elo score and roughly an Elo score being higher by 100 means that the stronger competitor will win 64% of the time. And 400 points is meant to be a 10 to 1 advantage. And so the the the pro level model actually achieves the second highest Elo score in this thing and there's probably like 30 different models. I'm showing the top of the leaderboard. Uh which is pretty nice. Uh so this is like a free service uh and these ones are are not.
嗯,然后由此就可以算出所谓的 Elo 分数。Elo 我记得是一位匈牙利数学家,他当时想找出给国际象棋棋手排名的方法。所以当你办一场锦标赛时,基本上你每击败一个对手就能拿到更多 Elo 分,这样在锦标赛的设定里就能算出 Elo 分数。大致来说,Elo 分高出 100 分意味着更强的一方会有 64% 的概率获胜。(获胜)。而高出 400 分意味着大约 10 比 1 的优势。所以这个 Pro 级别的模型在这上面拿到了第二高的 Elo 分数,而榜上大概有 30 个不同的模型。我这里展示的是排行榜的顶部。呃,这挺不错的。而且这个是免费服务,而这几个不是。
便签引用
46:27
And so I tweeted about this and then this person who I'd never seen said it's actually it's really pretty good. Check out this example and you know we have we have a way for people to share their interactions with Bard that they liked and they actually want to make public and so he he decided to do that. So if you click on that he was asking about please estimate the number of companies per million residents in each of the following countries, the US, England, South Korea, Taiwan, and Singapore. Present your numbers in table form.
于是我发了条推文讲这件事,然后有个我从没见过的人说,它其实真的挺好用的。看看这个例子。你知道,我们有个功能,可以让用户把自己喜欢、并且愿意公开的 Bard 对话分享出来,他就这么做了。所以点进去,你会看到他问的是:请估算下列国家每百万居民对应的公司数量,包括美国、英格兰、韩国、台湾和新加坡。请用表格形式呈现这些数字。
便签引用
46:55
And so Gemini said absolutely here's a table showing the number of blah blah blah. There it is, the number of companies per million just as you had asked. As you can see England has the highest number and it sort of interprets the result for you too. England has the highest number followed by the US, South Korea, Taiwan, then Singapore. It's important to note these are just estimates. The actual number may vary depending on source and definition. I hope this helps. Let me know if you have any other questions.
然后 Gemini 说,当然可以,这是一张表格,展示了每百万人对应的公司数量等等等等。就在那儿,正是你要的内容。如你所见,英格兰的数字最高——它还会帮你解读结果。英格兰最高,其次是美国、韩国、台湾,然后是新加坡。需要注意的是,这些只是估算值,实际数字可能因来源和定义的不同而有所差异。希望这对你有帮助。如果还有其他问题,请告诉我。
便签引用
47:22
Um and so he went on to say, you know, where did you find the numbers for the companies? Oh, I apologize for not providing the sources. Um it's important to be transparent where information comes from. So the company numbers I used came from a combination of different sources. So for the US they used the Dun & Bradstreet database. Uh England I used the Office of National Statistics, blah blah blah blah blah. Please note these sources have somewhat different definitions. So it actually understands the subtleties, right? Like the there's different definitions of what constitutes a company and it sort of knew where to go off and and get all this uh different kind of sources of data. It didn't have that memorized, but it was able to make use of it.
嗯,接着他又问,你这些公司数量的数字是从哪儿找到的?哦,很抱歉我没有提供来源。嗯,说明信息出处很重要,保持透明是必要的。我用的公司数量来自多个不同来源的组合。比如美国的数据用的是邓白氏(Dun & Bradstreet)数据库。呃,英格兰的我用的是国家统计局的数据,等等等等。请注意这些来源的定义有所不同。所以它其实理解其中的微妙之处,对吧?比如关于什么算一家公司,存在不同的定义,而它多少知道该去哪里找这些不同来源的数据。它并没有把这些背下来,但它能够加以利用。
便签引用
09领域精调、图像生成与手机端应用
48:03
Um Pretty neat. Okay. And another trend I think is important is that uh um further refinement of these general models can make amazing domain-specific models. So, some of my colleagues took some of our earlier work on the Palm model and then the Palm 2 model, which is a general-purpose model kind of like uh trained on general text, and decided to enrich it and further train it on medical data. So, medical kind of questions and medical articles. And what they found was the Med-PaLM model, the first one, actually exceeded the medical passing mark for the medical boards. And that when you then, 6 months later, they trained it on the Med-PaLM 2 model to do Med-PaLM 2, they actually got expert-level performance on the medical boards uh for this particular task. Now, this is not a full general-purpose uh setting. It's like a bunch of medical questions, but it does show the capabilities of having a really capable general model and then training it in a domain-specific way for specific uh problems.
嗯,挺妙的。好。我认为另一个重要趋势是,对这些通用模型做进一步的精调,可以得到非常出色的领域专用模型。比如我的一些同事拿了我们早期在 PaLM 模型上的工作,还有 PaLM 2 模型——那是一个在通用文本上训练的通用模型——然后决定用医学数据来丰富它、继续训练它。也就是医学类的问题和医学文章。他们发现,Med-PaLM 模型,第一代,实际上就超过了医学执业考试的及格线。之后过了 6 个月,他们在此基础上训练出 Med-PaLM 2,在医学执业考试上达到了专家级水平,就这一特定任务而言。当然,这不是完全通用的场景,只是一堆医学问题,但它确实展示了这样一种能力:先有一个非常强的通用模型,再针对具体问题做领域化训练。(问题)。
便签引用
49:15
Okay. Uh I'm going to go quickly through generative models uh producing images and video. You've probably seen this as a trend in the world. So, we have a couple of different research projects uh Party and Imagine. Uh and you know, one of the kind of cool things I mentioned you can give prompts that describe what you want in visual imagery and then have models that can generate these these images that are kind of constrained by the encoding representation of processing the sentence. And then conditioned on that, it will generate pixels for an image.
好。呃,我要快速讲一下生成模型如何生成图像和视频。你在现实世界里大概已经见到过这个趋势。我们有几个不同的研究项目,呃,Parti 和 Imagen。呃,我提到过一件挺酷的事:你可以给出描述你想要的视觉画面的提示词,然后让模型生成这些图像——它们受制于对句子处理后得到的编码表示。然后以此为条件,模型会为一张图像生成像素。
便签引用
49:50
So, a steam train passes through a grand library, oil painting in the style of Rembrandt. There you go. Uh a giant cobra snake made from X, where X might be corn, pancakes, sushi, or salad. Which is your favorite? I I kind of like the the ferocious lettuce looking snake. But the corn one is pretty nice, too. Um you know, a photo of a living room with a white couch and a fireplace, an abstract painting is on the wall, and bright light comes through the windows. So, if you happen to need a picture like that for for a presentation or something like I did, you can uh do that. And it can be pretty detailed descriptions. You know, a high-contrast photo of a panda riding a horse. Panda's wearing a wizard hat and reading a book. The horse is standing on a street against a gray concrete wall, colorful flowers and the word peace, you know, blah blah blah, DSLR DSLR photograph, daytime lighting.
比如:一列蒸汽火车穿过一座宏伟的图书馆,伦勃朗风格的油画。就是这个。呃,一条用 X 做成的巨蟒,X 可以是玉米、松饼、寿司或者沙拉。你喜欢哪个?我有点喜欢那条凶巴巴的生菜蛇。不过玉米那条也挺好看的。嗯,比如:一张客厅的照片,白色沙发和壁炉,墙上挂着一幅抽象画,明亮的光线从窗户照进来。所以,如果你正好需要这样一张图用在演示里,比如我就需要,你就可以这么做。而且描述可以相当细致。比如:一张高对比度的照片,一只熊猫骑着马。熊猫戴着巫师帽在读书。马站在街上,背景是灰色混凝土墙,还有五颜六色的花和「和平」这个词,等等等等,单反照片,日间光线。
便签引用
50:48
And there you go. You know, there are many plausible interpretations of that, but at least you got one example of what you what you asked for. Um and this is now integrated into Bard. So, the uh K-12 uh government uh school agency in Illinois uh was really excited about being able to create images of their mascot, Hyperlink the Hedgehog. Um so, there's Hyperlink surfing, riding the AI wave. Uh and this person was very excited about you know, the prompt was a human buying coffee at Costa Coffee in in London. Uh Costa Coffee is a very popular coffee chain. Um and one of the things that these models have often struggled with is the fidelity of text.
然后就出来了。当然,这段描述有很多种说得通的诠释,但至少你得到了一个符合你要求的样例。(样例)。嗯,这个现在已经集成进 Bard 了。比如伊利诺伊州的 K-12 政府教育机构就非常兴奋,因为他们能生成自己吉祥物的图像,超链接刺猬 Hyperlink。嗯,这是 Hyperlink 在冲浪,驾驭 AI 的浪潮。呃,还有这个人非常兴奋,他的提示词是:一个人在伦敦的 Costa Coffee 买咖啡。Costa Coffee 是一家非常流行的连锁咖啡店。嗯,这些模型常常吃力的一点就是文字的还原度,
便签引用
51:35
Uh actually, you know, get putting the text you asked for, making it look like a real font, and so on. Here, you see it does does a pretty good job. Um I won't talk through a lot of the details, but essentially, you know, you put in a prompt that gives you a representation of what that sentence in a distributed vector-based setting is, and then conditioned on that, the model is trained to generate first a small-scale image, and then take uh, another a model that is designed to improve the res- increase the resolution of an image, uh, conditioned on both that lower-scale, lower-resolution image plus the text embedding, and then we apply that one more time with the larger image and the conditioned on the text embedding to produce the full-scale 1024 by 1024 image.
呃,也就是把你要求的文字真正写出来,让它看起来像真实的字体,等等。这里你能看到它做得相当不错。嗯,细节我不多讲,但基本上,你输入一个提示词,得到那句话在分布式向量空间中的表示,然后以此为条件,模型被训练成先生成一张小尺寸图像,再用另一个专门用来提升图像分辨率的模型,以那张低分辨率图像加上文本嵌入为条件(放大),然后再对更大的图像重复这一步,同样以文本嵌入为条件,最终生成完整的 1024乘 1024 图像。
便签引用
52:29
Um, and you can really see the effects of scale. So, if we train four different models, uh, with 350 million to 20 billion parameters, um, and then giving them the same prompt, you know, a portrait photo of a kangaroo wearing an orange hoodie and blue sunglasses standing on the grass in front of the Sydney Opera House holding a sign on the chest that says welcome friends. Um, what you see is, you know, it kind of got the kangaroo aspect at the smaller scale, but not and the orange hoodie, I guess, but not much else.
嗯,你真的能看到规模带来的效果。比如我们训练四个不同的模型,参数量从 3.5 亿到 200亿,然后给它们同一个提示词:一张袋鼠的肖像照,它穿着橙色连帽衫、戴着蓝色墨镜,站在悉尼歌剧院前的草地上,胸前举着一块牌子,上面写着「欢迎朋友们」。嗯,你会看到,在较小规模下它多少抓到了袋鼠这一点,橙色连帽衫大概也有了,但其他就没什么了。
便签引用
53:01
There is a sign, but it's, you know, again, not it struggled with text, let's say. Uh, as you scale up a bit more, the kangaroo got a little better. It now knows a bit more that Sydney Opera House looks something like that, but it's kind of a little chunky and doesn't have a lot of detail. Uh, it it's closer to welcome friends, uh, but it might be Vegemite, hehe. I'm not sure. But then as you scale up, you now get a pretty nice image of the Sydney Opera House and your kangaroo and the orange hoodie with the right text.
牌子是有的,但你知道,它在文字上还是很吃力。呃,规模再扩大一些,袋鼠好看了一点。它现在多少知道悉尼歌剧院大概长那样,但有点笨重,细节不多。文字上更接近「欢迎朋友们」了,但也可能写的是 Vegemite(澳洲酵母酱),哈哈。我也说不准。但等规模再往上扩,你就能得到一张相当漂亮的图:悉尼歌剧院、你的袋鼠、橙色连帽衫,文字也对了。
便签引用
53:36
So, you see, scale is is an important aspect of this, and this is why you're seeing all these advances in, you know, the last decade is essentially scale and better training methods and algorithms really contribute to higher quality results. This graph just effectively says the same thing, but I think the kangaroo says it better. Um I think it's also important to realize that there's a lot of machine learning uh kind of invisibly helping people in various ways uh and particularly on phones. So, you know, a lot of camera features in modern smartphones have gotten significantly better over the years through combinations of computational photography methods and machine learning methods uh together. You know, so portrait mode where you make the background all blurry so you look all fancy in the foreground.
所以你看,规模是其中很重要的一环。这也是为什么过去十年你会看到这么多进展——本质上就是规模,加上更好的训练方法和算法,共同带来了更高质量的结果。这张图其实说的是同一件事,但我觉得袋鼠说得更明白。嗯,我觉得还得意识到,有大量机器学习正在以看不见的方式帮助人们,尤其是在手机上。你知道,现代智能手机上的很多相机功能这些年来通过计算摄影和机器学习方法的结合而显著变好。比如人像模式,把背景虚化,让前景里的你看起来很有格调。
便签引用
54:29
Um is a nice nice technique for some of these uh portrait-style photos. Uh night sight where you try to take an image in very low light conditions, you can essentially take lots of readings from the sensor and integrate those in software to create, you know, much higher uh lighting conditions than the actual conditions under which you did that. That also helps you take better astrophotography. And portrait blur and color pop are nice features sometimes when you want them. Um magic eraser. So, if you actually understand images and you point at like one of the telephone poles and says make and say make these go away, then the system can do that. Maybe your waterfall photo had these, you know, uh other tourists in front of it and you didn't want them there. You can you can erase them.
嗯,对这类人像风格的照片来说这是个不错的技术。呃,还有夜视,当你在极暗的光线下拍照时,本质上是从传感器采集大量读数,再用软件把它们整合起来,得到远比实际拍摄环境更亮的效果。这也能帮你拍出更好的天文摄影。人像虚化和色彩突出在你需要的时候也是很好用的功能。嗯,还有魔术橡皮擦。如果系统真的理解图像内容,你指着比如某根电线杆说让这些消失,它就能做到。也许你的瀑布照片里有一些你不想要的其他游客挡在前面,你可以把他们抹掉。
便签引用
55:18
Uh there they go. Um and there's a lot of features on the phone and many of them are sort of about how do you transform one modality into another? Uh you know, so sometimes uh you want to be able to say screen a call. Uh so, maybe you don't want to actually answer your phone but have a uh computer-generated voice answer the phone for you, ask what the person's calling about, and then give you a transcript of what they said. Uh and then you can decide, you know, do I want to accept this phone call or not? Um you know, hold for me can kind of listen on the phone for you so you don't have to hold uh hold on Bank of America when you're calling their customer support.
呃,他们就没了。嗯,手机上还有很多功能,其中不少本质上是在把一种模态转换成另一种。呃,比如有时候你想要「来电筛查」。呃,也许你不想真的接电话,而是让一个计算机生成的语音替你接,问对方打电话是为了什么事,然后给你一份他们说了什么的文字记录。然后你可以决定,我要不要接这通电话?嗯,还有「代我等待」,可以替你在电话上听着,这样你就不用干等——比如你打美国银行客服时得一直等着。
便签引用
56:00
Live Caption can show can take any video playing on your phone and listen to the audio and then give you transcripts uh captions of what's being narrated. Maybe you're trying to watch a video in a lecture hall like this and you don't want the audio to disturb people. Um so there's a lot of cool cool features of this and a lot of these are running on people's phones without them necessarily realizing it or thinking about what technology is under there. Um and this has amazing uh advances for for uh sort of people in limited literacy settings. You know, you can point your camera at something and it can read you what you're pointing it at or maybe you don't speak that language and you're trying to understand it. It can it can read it and translate it for you.
实时字幕可以把你手机上播放的任何视频的音频听下来,然后给你字幕,显示正在讲的内容。也许你想在像这样的报告厅里看视频,又不想放出声音。打扰到别人。嗯,所以这里面有很多很酷的功能,而且很多都在人们的手机上运行着,他们并没有意识到,也没有去想背后是什么技术。嗯,这对于识字能力有限的人群来说是非常了不起的进步。你知道,你可以把摄像头对准某个东西,它就能把你对准的东西读给你听;或者你不懂那门语言,想要理解它,它可以读出来并翻译给你。
便签引用
10科学与医疗:材料发现与视网膜筛查
56:46
Um I think I'm going to go quickly here and maybe skip some of this section. Uh I will just skip over some of this, but there's a pretty awesome advances in uh Yeah, so I'll I'll start here. You know, I think material science is a pretty interesting area where, you know, basically machine learning is starting to influence lots and lots of aspects of of science. Um both through kind of automated explorations of interesting parts of a scientific uh hypothesis space or through you know, creation of very rapid simulators that are learned rather than sort of traditional high high large large-scale kind of HPC style computation.
嗯,我想我要讲快一点,可能会跳过这部分的一些内容。呃,我会略过一些,但在这方面有相当了不起的进展。嗯,是的,那我就从这里开始吧。你知道,我觉得材料科学是一个相当有意思的领域,你知道,基本上机器学习正开始影响科学的方方面面。嗯,一方面是通过对科学假设空间中有意思的部分进行自动化探索,另一方面是通过创建非常快速的、学习出来的模拟器,而不是那种传统的大规模 HPC 式的计算。
便签引用
57:31
You know, in some areas you've been able to learn a simulator that is sort of the functional equivalent of a hand-coded simulator, but is now 100,000 times faster. And so, that means that all of a sudden you can, you know, search a space of 10 million possible chemicals or materials and identify ones that are interesting and promising and have certain properties that you would normally have to apply a lot more compute for. And so, some of my DeepMind colleagues are actually looking at interesting ways of searching the space of possible materials for those with interesting properties. So, they have a structural pipeline that can, you know, represent a potential material as a graphical neural network and then a compositional pipeline that can sort of mutate sort of known structures into ones that are sort of interesting and adjacent and then use an existing database of materials to then be able to output energy models and a bunch of stable interesting possible compounds.
你知道,在某些领域,你已经能够学出一个模拟器,它在功能上等价于手工编写的模拟器,但现在快了十万倍。所以这意味着,你突然之间可以,你知道,搜索一个包含一千万种可能的化学物质或材料的空间,从中识别出那些有意思、有前景、具备某些特定性质的候选,而这些在平常需要投入多得多的算力。所以,我在 DeepMind 的一些同事其实正在研究一些有意思的方法,来搜索可能材料的空间,找出那些具备有趣性质的。他们有一条结构化的流程,可以把一种潜在材料表示成图神经网络,还有一条组分流程,可以把已知的结构变异成那些有意思的、相邻的结构,然后利用现有的材料数据库,从而能够输出能量模型和一批稳定、有意思的可能化合物。
便签引用
58:38
And so, this automated discovery of 2.2 million new crystal structures leads to a bunch of interesting, you know, possible candidates for actual synthesis in the in the lab to see what properties they actually have. And I think there's huge potential for using machine learning in all aspects of health care. Really, we've been doing a fair amount of work in the space of medical imaging and diagnostics for quite a while and those problems range from ones where you have 2D images to some some where you have sort of 3D volumes from MRIs or other kinds of 3D CT scans.
所以,这种自动化发现的 220 万种新晶体结构,带来了一批有意思的、可以在实验室里实际合成的候选材料,来看看它们实际具备什么性质。我认为在医疗健康的所有环节中使用机器学习都有巨大的潜力。确实,我们在医学影像和诊断这个领域已经做了相当多的工作,做了挺长时间了,这些问题的范围很广,有些是二维图像,有些则是来自核磁共振或其他各种三维 CT 扫描的三维体数据。
便签引用
59:21
And then some where you have just a single view to ones where you have multiple views and and large images, very high resolution things for pathology for example. Um, and so there's been a a fair body of work. I'm going to talk briefly about two of them though. So one of the areas we've been working in the longest in this space is the area of diabetic retinopathy. And so diabetic retinopathy is a is a degenerative eye disease that can you know if you catch it in time is very treatable but if you don't you can suffer full or partial vision loss and really people who are at risk which is sort of anyone with diabetes or pre-diabetes should be screened every year but in a lot of parts of the world there just aren't enough ophthalmologists to do the screening those who have been trained in sort of interpreting these retinal images.
还有些只有单一视角,也有些是多视角、大尺寸图像的,比如病理学那种超高分辨率的东西。嗯,所以已经积累了相当多的工作。不过我只简单讲其中两项。我们在这个领域里做得最久的方向之一,是糖尿病视网膜病变。糖尿病视网膜病变是一种退行性眼病,你知道,如果及时发现是很好治疗的,但如果没发现,就可能造成全部或部分视力丧失。真正的高风险人群,也就是所有糖尿病或糖尿病前期的人,都应该每年筛查一次,但在世界上很多地方,就是没有足够的眼科医生来做这种筛查——那些受过训练、会判读这些视网膜图像的医生。
便签引用
1:00:11
Um, and so this is something where machine learning can actually help a lot because you can actually train a model based on you know trained ophthalmologists annotating images to say yes that's a one that's a three that's a two that's a five. Um, and if you train a model on board certified ophthalmologists you can actually train a model that is as effective as board certified ophthalmologists. If you then go on to get that same training data annotated by retinal specialists who have a lot more expertise and and experience in the these cases you can actually train a model that is on par with retinal specialists is kind of the gold standard of of care in this space and there are very few of those in the world but you can all of a sudden make the screening quality be that of a retinal specialist using a GPU on a laptop.
嗯,所以这正是机器学习能帮上大忙的地方,因为你其实可以训练一个模型,基于受过训练的眼科医生对图像的标注,说这个是一级、这个是三级、这个是二级、这个是五级。嗯,如果你用具备执业资格认证的眼科医生的标注来训练模型,你其实可以训练出一个和这些认证眼科医生一样有效的模型。如果你接着让视网膜专科医生来标注同样的训练数据——他们在这类病例上有多得多的专业知识和经验——你其实可以训练出一个与视网膜专科医生水平相当的模型,那是这个领域护理的黄金标准,而全世界这样的医生非常少。但你突然之间就能让筛查质量达到视网膜专科医生的水准,只用笔记本电脑上的一块 GPU。
便签引用
1:01:03
Um, and so we've actually partnered with organizations in India a network of Indian eye hospitals and and the government of Thailand as well as France and Germany and we're sort of doing lots and lots of screening every year. And then dermis is so dermatology is a condition where it's interesting cuz you actually don't need specialized equipment to sort of us gather data that is useful for interpreting, you know, do you have a dermatological condition or not? Uh and so we've got a system now deployed where um you can take a photo of something as you see in the video and it will give you a sense of, you know, what this might be, what are other similar looking images in sort of dermatological databases uh and it can help give you a sense of is this something very serious or is it something sort of fairly benign.
嗯,所以我们实际上已经和印度的一些机构合作了,一个印度眼科医院网络,还有泰国政府,以及法国和德国,我们每年在做大量大量的筛查。然后是皮肤——皮肤病学是一个很有意思的方向,因为你其实不需要专门的设备就能采集到有用的数据来判断,你知道,你是不是有皮肤病,有没有。呃,所以我们现在部署了一个系统,嗯,你可以像视频里看到的那样给某个部位拍张照,它会给你一个大致判断,你知道,这可能是什么,皮肤病学数据库里还有哪些看起来相似的图像,呃,它能帮你判断这是很严重的问题,还是相当良性的东西。
便签引用
11AI 原则与结语:从手写软件到学习系统
1:01:57
Um Okay. Uh and then finally, I think deeper and broader understanding of the machine learning methods as we sort of deploy them in more places in the world is really, really important. And you know, as we've gone from doing basic research in in machine learning to then using it in a lot of places in all of our products, we started to think about a, you know, a set of principles by which we want to, you know, think about the implications of using machine learning, what considerations should we have for, you know, various ways in which we might apply it. Um and we in 2018 published a set of principles that we came up with.
嗯,好。呃,最后我想说,随着我们把机器学习方法部署到世界上更多的地方,对这些方法更深入、更广泛的理解是非常非常重要的。你知道,随着我们从做机器学习的基础研究,走到在我们所有产品的很多地方使用它,我们开始考虑,你知道,一套原则,我们希望据此去思考使用机器学习的影响,在我们可能应用它的各种方式中,应该有哪些考量。嗯,我们在 2018 年发布了一套我们提出的原则。
便签引用
1:02:40
Um really these were designed to help ed our own uh internal teams about machine learning and things you should be thinking about as you're thinking about applying it to problems you care about. Uh and so for example, you know, avoid creating or reinforcing unfair bias. Often when you train these models, they're trained on uh wor- data from the real world and that's often the world as uh not the world as we'd like it to be, but the world as it is. And so it's really important when you're deploying machine learning models that you don't sort of train on data that it biased in unfair ways and then accelerate that because now you can automate and make these decisions more rapidly.
嗯,这些原则其实是为了帮助教育我们自己内部的团队,让他们了解机器学习,以及在考虑把它应用到你关心的问题上时应该想到哪些事。呃,比如说,你知道,避免制造或强化不公平的偏见。很多时候你训练这些模型,它们是用来自现实世界的数据训练的,而那往往是这个世界呃,不是我们希望它成为的样子,而是它本来的样子。所以当你部署机器学习模型时,真的很重要的一点是,你不要在带有不公平偏见的数据上训练,然后把这种偏见放大,因为现在你可以自动化,可以更快速地做出这些决策。
便签引用
1:03:22
Um so, there's a bunch of techniques you can apply uh on a sort of algorithmic basis to remove some kinds of bias. Um and what we strive to do is sort of apply the best known current techniques, but then also do research on advancing the state of the art in these areas of bias or for example um uh accountable to people. We think, you know, making models interpretable is an important aspect of that. Um you know, being sensitive to privacy when that makes sense in the setting you're deploying it. Uh and you know, be socially beneficial.
嗯,所以有一堆技术你可以在算法层面上应用,来消除某些类型的偏见。嗯,我们努力做到的是应用当前已知的最佳技术,但同时也在这些偏见等领域做推进最前沿的研究,比如说,嗯,呃,对人负责。我们认为,你知道,让模型可解释是其中很重要的一个方面。嗯,你知道,在你部署它的场景中,只要合适,就要对隐私保持敏感。呃,还有,你知道,要有益于社会。
便签引用
1:03:59
Uh and so, I I'll point out a lot of these are sort of active areas of research. Uh so, we you know, we've published about 200 different papers in uh in the last 5 years or so, 6 years uh related to fairness or bias, privacy, or safety. And you can see those there. Okay. In conclusion, you know, I think it's pretty exciting times for computing. I think there's a change underway from hand-coded software systems to ones that are learned and that can interact with the world in various interesting ways and interact with people in interesting ways.
呃,所以我要指出,这里面很多都是很活跃的研究领域。呃,所以我们,你知道,我们在过去五年左右、六年里发表了大约 200 篇不同的论文,嗯,涉及公平性或偏见、隐私,或者安全。你可以在那里看到这些。好的。总结一下,你知道,我觉得对计算领域来说这是相当激动人心的时代。我认为一场变革正在发生,从手工编写的软件系统,转向那些学习出来的、能以各种有意思的方式与世界互动、与人互动的系统。
便签引用
1:04:32
Um the modalities that computers can now sort of ingest and understand and sort of produce are growing and are, you know, I think going to make using computers much more seamless and natural. You know, a lot of times we sort of restrict ourselves to typing on a keyboard or something like that, but I think we now have the ability to talk to a computing system in a very natural way, and it will understand what we say, it'll be able to produce a natural sounding voice in response or a nice image if that's what we asked for. And so, I think that's pretty exciting. So, there's tremendous opportunity for sure.
嗯,计算机现在能够接收、理解并产出的模态越来越多,而且,你知道,我觉得这会让使用计算机变得更加顺畅和自然。你知道,很多时候我们把自己局限在敲键盘之类的方式上,但我觉得我们现在有能力以非常自然的方式跟一个计算系统对话,它会理解我们说的话,能够用自然的语音来回应,或者如果我们要的是一张漂亮的图像,它也能生成。所以,我觉得这挺让人兴奋的。所以,机会确实是巨大的。
便签引用
12问答:数据、多模态与小算力研究
1:05:10
Uh but there's also a lot of responsibility. How do we sort of take this work forward, make sure that it's socially beneficial, uh, and really, uh, kind of do good things in the world with it? And for that, thank you very much. Uh, I will I will put up one more plug for the Slido number. There it is. Well, thank you very much for your talk. Thank you very, very much for your talk. Please don't send more questions to Slido. Okay. It's a very nice idea, but we are overwhelmed at this point. Uh, so what we are going to do is we are uh, um, I'm going to give you some questions. Uh, some there were some trends in the questions that appeared in Slido. So, I'm going to ask you some of these questions, and then for those of you who made it to this auditorium, we'll give uh, we'll ask we'll take one or two questions from the audience.
呃,但责任也很大。我们要如何把这项工作往前推进,确保它有益于社会,呃,并且真正用它在世界上做好事?就讲到这里,非常感谢大家。呃,我再推销一下那个 Slido 的号码。就在那儿。好,非常感谢你的演讲。非常非常感谢你的演讲。请不要再往 Slido 上发问题了。好的。这想法很好,但我们现在已经应付不过来了。呃,所以我们接下来要做的是,呃,嗯,我来问你一些问题。呃,Slido 上出现的问题里有一些趋势。所以我会问你其中一些问题,然后对于那些到场来到这个礼堂的朋友,我们会,呃,我们会问,会从现场观众里接一到两个问题。
便签引用
1:06:09
Uh, so one of the questions is um, let let me start with this a question that you probably expect. Okay, more data. Is it going to to make your model better? Twice more data. Are we going to see twice as uh, as good a performance? Yeah, I mean it's a it's a good question, and it's a it's not a simple answer, I think. I mean, I think we've seen that more high-quality data absolutely makes the model perform better when you have the capacity to sort of train on that large amount of data. So, it's important to think about the model's capacity. You know, sometimes you need to increase the scale of the model as well when you have more training data.
呃,其中一个问题是,嗯,让我先从这个你大概能预料到的问题开始。好,更多的数据。它会让你的模型变得更好吗?两倍的数据,我们会看到两倍好的性能吗?是啊,我是说这是个好问题,而且我觉得答案并不简单。我是说,我觉得我们已经看到,更多高质量的数据绝对会让模型表现更好,前提是你有足够的容量去在那么大量的数据上训练。所以,考虑模型的容量很重要。你知道,有时候当你有更多训练数据时,你也需要扩大模型的规模。
便签引用
1:06:52
We've seen uh, more data actually hurt. So, if you actually get a lot of low-quality data, uh, you can actually, for example, decrease the model's ability to effectively do mathematics problems or things like that. So, it's it's a nuanced thing, but in general, more high-quality data and more capacity for the model will make the model better. Yeah. So, a next question in this in the slide that that emerged is, "Okay, so what is the future of LLMs now that the vast majority of high-quality training data has been exhausted?"
我们也见过更多数据反而有害的情况。所以,如果你拿到大量低质量的数据,呃,你实际上可能会,比如说,削弱模型有效解决数学问题之类的能力。所以这是个有细微差别的事情,但总的来说,更多高质量的数据加上模型更大的容量,会让模型变得更好。是的。那么,幻灯片上冒出来的下一个问题是,“好,那么既然绝大部分高质量训练数据都已经被用尽了,大语言模型的未来是什么?”
便签引用
1:07:25
How would you react to that? I I would disagree with that assertion a bit. Uh you know, I think we've not really begun to train on, say, video that much. I mean, we've done small amounts of video, but there's a huge amount of video data in the world. I think actually understanding the world through visual and audio data will be different than sort of training on a lot of language. You're going to want to do both. But, I don't think we've we've really exhausted the training data in the world. Yeah. Yeah. I tend to agree with you. I think we still have a lot to go.
你怎么看这个说法?我会稍微不同意这个断言。呃,你知道,我觉得我们其实还远远没有开始在视频上做多少训练。我是说,我们做过少量的视频,但世界上有海量的视频数据。我其实认为,通过视觉和音频数据来理解世界,会不同于在大量语言上做训练。你会两者都想做。但是,我不认为我们真的用尽了这世界上的训练数据。是啊。是啊。我倾向于同意你的看法。我觉得我们还有很长的路可以走。
便签引用
1:08:00
Multimodal models, you emphasized that in your talk. Do they achieve better performance on all domains than targeted models for each domain separately? Or you can paraphrase this question and answer a version of that question. I think in some cases they do. So, the question is, as you add more modalities, does that improve the performance on other modalities? Yeah. And you hope so, and and generally we do see some aspects of that. Um but I I you know, I think if you collect a if you have a narrow problem and you collect a very targeted data set that is designed to tackle that just that problem, that will often, you know, give you good performance on a problem. On the problem. But, if you have a complicated problem or it's hard to collect very specialized data, what you want is a model that has a huge amount of knowledge of lots of different things in the world, you know, from language and from, you know, images and audio, and then to be able to apply that model to the problem you care about.
多模态模型,你在演讲中强调了这一点。它们在所有领域上都比针对各个领域单独定制的模型表现更好吗?或者你也可以换个说法来理解这个问题,回答那个版本的问题。我觉得在某些情况下确实如此。所以问题是,当你增加更多模态时,这会不会提升在其他模态上的表现?是的。你会希望如此,而且总体上我们确实看到了一些这样的迹象。嗯,但我,你知道,我觉得如果你收集,如果你有一个很窄的问题,然后你收集了一个非常有针对性的数据集,专门是为了攻克那一个问题,那通常会,你知道,在那个问题上给你不错的表现。在那个问题上。但是,如果你面对的是一个复杂的问题,或者很难收集到非常专门的数据,你想要的就是一个对世界上各种各样的事物拥有海量知识的模型,你知道,来自语言、来自图像和音频的知识,然后能够把这个模型应用到你关心的问题上。
便签引用
1:09:08
And then if you have a little bit of data for a problem you care about, then you're going to want to start with that base model and then fine-tune it or do in-context learning or something like that Mhm. to make the performance uh quite good. Mhm. Maybe I could follow with another question, which is kind of related. Today, the cost of training large models prevents small startups from making impact. What kind of projects would individuals with less resource work on? Would you like to comment on that? Yeah, absolutely. I mean, I think um there's a really large set of problems in the machine learning domain. I I'm going to address it more from a, you know, what interesting research can one do Mhm. in the in the broad area uh where maybe you don't have access to large data centers of of compute and so on. And I think there's just an really wide-open set of things. Uh so, I've mentioned the quality of data about automatic evaluation of data quality or online curriculum learning or optimization methods or a lot of these
然后如果你手里有一点点关于你所关心的问题的数据,那你就会想从那个基础模型出发,再做微调,或者做上下文学习之类的事情,嗯,好让它的表现相当不错。嗯。也许我可以接着问一个问题,跟这个也有点关系。今天训练大模型的成本让小的创业公司很难做出影响力。资源比较少的人应该做什么样的项目呢?你愿意就这个说说吗?当然可以。我的意思是,我觉得嗯机器学习领域里有非常大的一批问题。我我更想从这个角度来谈,就是说,一个人可以做哪些有意思的研究,嗯,在这样一个大的方向上,嗯也许你并没有大型数据中心的算力之类的资源。我觉得可做的事情非常开阔。呃,我前面提到过数据质量,比如自动评估数据质量,或者在线课程学习,或者优化方法,很多这类东西其实都可以在,你知道,一块 GPU 上,或者
便签引用
1:10:12
kinds of things can actually be demonstrated on, you know, one GPU or, you know, a handful of GPUs under your desk and actually make pretty significant and innovative advances. You know, the original Transformer work was done on eight GPUs, I think. Interesting. uh that or the sequence-to-sequence model, for sure, was eight GPUs. And so, I think there's advances to be had from clever ideas, good evaluation of them, and even demonstration of them at small scale. Mhm. Okay. Another set of questions that we got is is LLMs everything? Is Transformers everything? What else is there? Should we be working on other kinds of models?
你桌子底下的几块 GPU 上做出来,而且真的能带来相当显著、有创新性的进展。你知道,最早的 Transformer 那篇工作我记得就是在八块 GPU 上做的。有意思。呃,那个,或者说 sequence-to-sequence 模型,那个肯定是八块 GPU。所以我觉得进展是可以靠聪明的想法、对它们做好的评估甚至只是在小规模上把它演示出来而取得的。嗯。好的。我们收到的另一组问题是:LLM 就是一切吗?Transformer 就是一切吗?还有别的吗?我们该不该去做别的类型的模型?
便签引用
1:10:53
Is the emphasis on LLMs LLMs uh stifling other other work in machine learning. Yeah, I mean it is a worry, right? Like are we crowding out other innovative ideas that maybe uh are you know, not as fully developed and so they don't look as good as some of the things that have been, you know, much more fully explored and we're sort of uh you know, um in the kind of now gentle exploration of the space around what works well when maybe something over here would work really well. You know, I think a lot of the time uh showing even at a small scale that some other idea is a really interesting direction can be done with some modest amount of experimental evidence. Um and I think that's an important area to go.
对 LLM 的这种强调,呃,是不是在压制机器学习里其他的工作?是啊,我是说这确实让人担心,对吧?就像,我们是不是在挤掉其他有创新性的想法,那些想法可能呃,你知道,还没有发展得那么成熟,所以它们看起来不如那些已经被,你知道,探索得充分得多的东西那么好看,而我们某种程度上呃,你知道,嗯,现在只是在围绕已经证明有效的东西做一些温和的探索,而也许另一边的某个东西其实会非常有效。你知道,我觉得很多时候呃,哪怕是在小规模上表明另一个想法是一个很有意思的方向,用不太多的实验证据就能做到。嗯,而且我觉得那是一个值得去做的重要方向。
便签引用
1:11:42
I would say mo- you know, I I tend not to use LLM because I think we're moving to a multimodal world. Uh and I think multimodal is going to be more than kind of the human modalities you think about like uh visual and audio and language but other modalities that are important in the world you know, like time series of interesting, you know, heart rate sensor data for health care applications. There's probably 50 to 100 modalities of data you'd want to be able to deal with. Mhm. Mhm. Uh I see I just saw the the the clock and we really ran over time. So, I would like to end here by thanking Justin for his talk. Thank you.
我要说,大部——你知道,我我一般不太用 LLM 这个说法,因为我觉得我们正在走向一个多模态的世界。呃,而且我觉得多模态会不止是你通常想到的那些人类感官的模态,比如呃视觉、音频和语言,还会有这个世界上其他重要的模态,你知道,比如说有意思的时间序列,比如用于医疗健康应用的心率传感器数据。大概有五十到一百种模态的数据是你会希望能够处理的。嗯。嗯。呃,我看,我刚看到时钟,我们真的超时了。所以,我想在这里结束,感谢 Justin的演讲。谢谢。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

Jeff Dean 回顾机器学习十年来的三条主线——规模化、专用硬件、多模态基础模型,并以 Gemini 为例说明它们如何叠加,同时强调数据质量、提示方法和负责任部署是接下来最关键的方向。

核心要点

  • 规模是过去十年进步的核心驱动力,而且反复奏效。 更多算力、更大数据集、更大模型三者叠加,每次放大都带来精度提升或涌现新能力。ImageNet 分类准确率从 2011 年的 50.9% 升至 AlexNet 的约 63%,再到如今的 91%,已超过人类水平(一千个类别含 40 种狗,人类分辨不清)。语音识别在 5 年内词错误率从 13.25% 降到 2.5%,即从"六七个词错一个"变成"四十个词错一个",这一阈值跨越使得口述邮件真正可用。
  • 神经网络的两个特性决定了专用硬件的设计方向。 一是容忍低精度:只保留一两位小数没关系,噪声甚至有助于学习;二是所有算法本质都是线性代数原语(矩阵乘、向量运算)的重复组合。因此只需把"低精度线性代数"做到极致。TPU v1 仅用于推理,相比同期 CPU 能效和性能提升 30–80 倍;v2 起支持训练,v3 引入水冷;pod 从 256 芯片 2D mesh(芯片间直连、无需路由)扩展到 v4 的 4096 芯片、1.1 exaflops,v5p 每芯片接近 0.5 petaflop(16 位浮点),单 pod 约 9000 芯片。
  • 算法进步与硬件进步相乘,而非相加。 语言模型演进路径:2 万亿 token 上统计 3000 亿个五元组的 n-gram 模型(用"stupid backoff"这种放弃严格数学、直接回退到短前缀的简单方法,效果接近更复杂的 Kneser-Ney)→ word2vec 分布式表示(king−queen 与 man−woman 方向相同,时态变化也是固定方向)→ seq2seq + LSTM → Transformer。Transformer 用并行处理加注意力取代顺序更新的单一状态,以 10–100 倍更少的算力获得更高精度。Dean 反复强调的经验:简单方法加海量数据非常有效。
  • Gemini 从一开始就是原生多模态,而非文本模型加补丁。 文本、图像、视频、音频统一编码为 token 序列训练 Transformer,支持交错输入(视频帧与字幕交替)。分三档:Ultra 最强,Pro 面向数据中心和产品(Bard 已更名 Gemini),Nano 可量化后在手机上运行。在 32 个学术基准中 30 个超越此前最优;MMLU 达 90%,超过基准作者测得的人类专家水平 89.6%–89.8%;图像理解 8/8、视频 6/6、音频 5/5 均为最优,其中一个图像基准在论文发表前一周才发布,从未见过仍领先,说明不是测试集泄漏。LMSYS 匿名盲测 Elo 榜上 Pro 排第二,且是免费模型。
  • 训练大模型的工程瓶颈在于故障与恢复,不只是算力。 Pathways 系统把高层计算描述自动映射到硬件,pod 内走高速直连网络,跨 pod 走带宽低得多的数据中心网络,研究者无需操心。Google 定义"goodput"指标衡量训练真正前进的时间占比;从其他机器内存中的模型副本快速恢复,而非读分布式文件系统的检查点,将恢复时间从几分钟压到 5–10 秒。一个自伤式教训:对上千台机器做滚动内核升级会造成持续故障,应改为同时下线、同时升级。
  • 数据质量与模型架构同等重要,甚至更重要。 训练混合比例(如代码占 32% 还是 27%)通过小模型消融实验决定;训练末期加权多语言等领域数据以定向增强。问答环节 Dean 明确说"更多数据"不是简单答案:低质数据会实际损害数学等能力,而且数据量增加往往需要同步扩大模型容量。他反驳"高质量数据已耗尽"的说法:视频数据几乎还没开始用,通过视觉和音频理解世界与语言训练是不同的路径。
  • 提示方法本身是一门科学,能在不改模型的情况下大幅提升准确率。 思维链提示(让模型"展示解题过程")在算术和八年级数学基准上效果呈阶跃式:小模型上几乎无用,模型规模跨过某个阈值后准确率骤升。多模态推理示例:给 Gemini 一张手写物理解答的照片,它能指出学生误用斜坡长度而非高度计算势能,并以 LaTeX 重解。Dean 联系到一对一辅导比课堂教学高两个标准差的研究,认为个性化辅导已在集体触及范围之内。
  • 通用模型经领域微调可快速达到专家水平。 Med-PaLM 在通用 PaLM 上继续训练医学数据后通过医学执照考试及格线,6 个月后 Med-PaLM 2 达到专家水平。图像生成同样显示规模效应:3.5 亿到 200 亿参数四档模型对同一提示(戴橙色卫衣的袋鼠在悉尼歌剧院前举"welcome friends"牌子),小模型只画出袋鼠和卫衣,文字乱码,最大模型才把歌剧院、文字全部还原。
  • 机器学习正大规模"隐形"运行于手机与科学研究中。 手机端:人像模式、夜景(多次采样软件叠加)、魔法橡皮擦、来电筛选、替你等待客服、实时字幕,以及为低识字人群朗读并翻译镜头所指文字。科学端:学习型模拟器比手写模拟器快 10 万倍,使搜索千万级化学空间成为可能;DeepMind 用图神经网络加结构突变管线自动发现 220 万种新晶体结构。医疗端:用视网膜专科医生标注训练的糖尿病视网膜病变模型达到专科医生水平,可在笔记本 GPU 上运行,已在印度眼科医院网络、泰国、法国、德国部署筛查。

结论与值得注意的细节

  • Dean 的总体判断: 计算正从手写软件转向学习得到的系统,计算机获得"眼睛"相当于生物进化出视觉的时刻,人机交互将从键盘走向自然语音和图像。机会巨大,责任同样巨大;Google 2018 年发布 AI 原则,近五六年发表约 200 篇公平性、隐私、安全相关论文,核心担忧是模型学到的是"世界的现状而非应有的样子",自动化会加速放大偏见。
  • 对资源有限研究者的建议: 数据质量自动评估、在线课程学习、优化方法等方向在一张或几张 GPU 上就能做出显著创新;Transformer 和 seq2seq 原始工作都只用了 8 张 GPU。他承认 LLM 热潮可能挤压尚未成熟的其他方向,但认为小规模实验证据就足以证明一个新方向值得投入。
  • 他刻意不用"LLM"这个词,认为未来是多模态世界,且模态远不止视觉、音频、语言这些人类感官,心率等时间序列传感器数据也算,实际需要处理的模态可能有 50 到 100 种。
  • 小细节: Gemini 在被追问数据来源时能列出邓白氏数据库、英国国家统计局等不同来源,并主动说明各来源对"公司"的定义不同;Elo 分差 100 意味着强者胜率 64%,分差 400 意味着 10 比 1 的优势。
  • 窄任务上专用模型仍可能更好。 Dean 坦言若问题窄且能收集针对性数据,专用数据集往往就够;通用大模型的价值在于复杂问题或难以采集专门数据的场景,此时应从基座模型出发做微调或上下文学习。
核心句型 · 10
1. It's not going to X, but I think it's important to Y
“It's not going to go into detail in any particular area, but I think it's important to understand what is happening in this field”
先降低预期再说明价值,是演讲开场设定范围的常用结构。仿写:This won't be a deep dive, but I think it's important to see the whole map.
2. If you think back N years ago, X kind of worked, but it wasn't really Y
“If you think back 10 or 15 years ago, you know, speech recognition kind of worked, but it wasn't, you know, really seamless.”
用「勉强能用但不够好」描述过去,为后文的进步做铺垫。kind of worked 是口语中表达「凑合」的地道说法。
3. All of a sudden, X becomes Y and that enables Z
“Before it was kind of unusable and now all of a sudden it becomes usable and that enables new kinds of things.”
描述跨过临界点后的质变。all of a sudden 比 suddenly 更口语,enable 引出连锁后果。适合讲技术转折或产品拐点。
4. What's more amazing, though, is …
“What's more amazing, though, is we've been able to reverse a lot of these arrows in the last few years.”
在已经强调过一点之后再推进一层。though 放在句中做让步转折,比 however 更自然。仿写:What's more striking, though, is how cheap it has become.
5. One lesson from this is [that] …
“One lesson from this is simple techniques over large amounts of data are very effective.”
从具体案例抽出普遍规律的句式。适合讲完一段经历后做总结。可扩展为 The lesson I took from this was …
6. Sure enough, …
“And sure enough, you can use a neural encoder over this input sequence to initialize the state”
表示「果然如预期」,用于叙述实验或尝试的结果。语气轻松,多见于口语讲述,书面可改用 as expected。
7. It's a nuanced thing, but in general, …
“So, it's a nuanced thing, but in general, more high-quality data and more capacity for the model will make the model better.”
回答复杂问题的两段式:先承认有例外,再给总体结论。避免绝对化又不失明确立场,适合问答场合。
8. I would disagree with that assertion a bit.
“I would disagree with that assertion a bit.”
礼貌反驳的模板:would 弱化语气,a bit 留余地,assertion 把对方说法定性为「断言」而非事实。之后接 I think 给出自己的理由。
9. not just X, but Y
“Not just a categorical label like leopard, but, you know, a little short sentence that describes what is going on in that scene.”
递进强调 Y 比 X 更进一步。写作中可用 not merely … but rather … 提升正式度。
10. You're always worried about X and so on
“Because you're always worried about like leakage of test date training data into the test set and so on.”
说明某种做法背后的顾虑。always worried about 表达行业内的常态性担忧,and so on 收尾避免枚举过长。
词汇精讲 · 114 · 按出现顺序
glimpse /ɡlɪmps/ n. 0:04
一瞥;短暂的一看。give sb a glimpse at/of 让某人瞥一眼
seamless /ˈsiːmləs/ adj. 0:41
无缝的,顺畅的(此处指语音识别体验不流畅)
endeavor /ɪnˈdevər/ n. 0:41
努力,事业。every field of human endeavor 人类活动的每一个领域
ballgame /ˈbɔːlɡeɪm/ n. 1:44
局面,形势。a completely different ballgame 完全是另一回事(习语)
threshold /ˈθreʃhoʊld/ n. 1:44
门槛,临界值。reach a threshold 跨过某一临界点
emerge /ɪˈmɜːrdʒ/ v. 1:44
涌现,出现(机器学习语境中常指能力随规模自发出现)
paradigms /ˈpærədaɪmz/ n. 2:31
范式,模式。learning-based paradigm 基于学习的范式
twisty /ˈtwɪsti/ adj. 2:31
弯弯绕绕的,逻辑复杂的(形容代码分支多、难以并行)
categorical /ˌkætəˈɡɔːrɪkl/ adj. 3:08
类别的,分类的。categorical label 类别标签
utterance /ˈʌtərəns/ n. 3:08
(一次)话语,说出的话;语音学中指一段连续发声
benchmark /ˈbentʃmɑːrk/ n. 5:11
基准测试,衡量标准
generalize /ˈdʒenrəlaɪz/ v. 5:11
泛化,推广到新情况。generalize from A to B 从 A 推广到 B
landmark /ˈlændmɑːrk/ adj./n. 5:54
里程碑式的;地标。a landmark paper 里程碑论文
affectionately known as phr. 5:54
被亲切地称为(介绍昵称的固定说法)
entrants /ˈentrənts/ n. 5:54
参赛者,参赛作品
hand engineer features phr. 5:54
手工设计特征(与从数据中自动学习特征相对)
indicative of phr. 5:54
表明……的,可作为……标志的
breeds /briːdz/ n. 7:07
(动物的)品种
word error rate phr. 7:42
词错误率,语音识别的核心指标,越低越好
dictate /ˈdɪkteɪt/ v. 7:42
口述(让机器或他人记录)
reduced precision phr. 9:08
低精度(用更少位数表示数值的计算方式)
transpositions /ˌtrænspəˈzɪʃnz/ n. 9:44
换位,变换形式(此处指同一套运算的不同组合方式)
primitives /ˈprɪmətɪvz/ n. 9:44
(计算机)基本操作,原语
inference /ˈɪnfərəns/ n. 10:27
推理;机器学习中指用训练好的模型做预测
close cousin phr. 11:04
近亲,非常相似的东西(比喻)
mesh /meʃ/ n. 11:57
网格;2D mesh 二维网格互连
routing /ˈruːtɪŋ/ n. 12:34
(网络)路由,数据包的路径选择
exaflops /ˈeksəflɑːps/ n. 12:34
每秒百亿亿次浮点运算(算力单位)
disk seek phr. 13:49
磁盘寻道,磁头移动到数据所在位置的操作,耗时以毫秒计
backoff /ˈbækɔːf/ n. 14:47
回退(n-gram 语言模型中查不到长序列时退而查短序列的策略)
smoothing /ˈsmuːðɪŋ/ n. 15:42
平滑(统计学中为未见过的事件分配概率的方法)
the data speaks phr. 15:42
让数据自己说话,数据量足够时结果自然显现
distributed representation phr. 15:42
分布式表示,用高维向量而非离散符号表示一个词
wrap your head around phr. 17:05
理解,在脑子里想明白(常用于难以想象的事物)
sequence-to-sequence adj. 18:31
序列到序列的(输入一个序列、输出另一个序列的模型)
spit out phr. v. 19:12
吐出,(口语)输出、生成
sure enough phr. 19:42
果然,不出所料
initialize /ɪˈnɪʃəlaɪz/ v. 19:42
初始化,设定初始状态
sequential /sɪˈkwenʃl/ adj. 20:58
顺序的,串行的(与 parallel 并行相对)
get away with phr. v. 21:49
侥幸做成,做了而不受影响。if we can get away with it 只要行得通
attend to phr. v. 21:49
关注,处理;机器学习中指施加注意力
sensible /ˈsensəbl/ adj. 23:12
合理的,讲得通的(Meena 评估指标之一)
engaging /ɪnˈɡeɪdʒɪŋ/ adj. 23:12
吸引人的,让人愿意继续互动的
progression /prəˈɡreʃn/ n. 23:50
进展序列,发展脉络
multimodal /ˌmʌltiˈmoʊdl/ adj. 24:57
多模态的(同时处理文本、图像、音频等)
decoding paths phr. 26:03
解码路径,模型从内部状态生成输出的不同通道
interleaving /ˌɪntərˈliːvɪŋ/ v. 26:52
交错排列,交替插入
quantize /ˈkwɑːntaɪz/ v. 27:26
量化,把模型权重压缩为更低位数以缩小体积
fabric /ˈfæbrɪk/ n. 28:17
(计算机)互连架构,网络结构;本义织物
topology /təˈpɑːlədʒi/ n. 28:17
拓扑结构,节点之间的连接方式
malfunction /ˌmælˈfʌŋkʃn/ v./n. 29:15
出故障,失灵
self-inflicted /ˌselfɪnˈflɪktɪd/ adj. 29:50
自己造成的,自找的
rolling failures phr. 29:50
滚动式故障,一台接一台陆续发生的故障
goodput /ˈɡʊdpʊt/ n. 30:27
有效吞吐量,真正推进工作的时间或数据占比
checkpoint /ˈtʃekpɔɪnt/ n. 30:27
检查点,训练过程中保存的模型状态快照
heuristics /hjuˈrɪstɪks/ n. 31:14
启发式规则,经验法则
ablations /æˈbleɪʃnz/ n. 31:14
消融实验,逐一去除或改变因素以测其影响
curriculums /kəˈrɪkjələmz/ n. 32:15
课程安排;机器学习中指训练数据的呈现顺序
elicit /ɪˈlɪsɪt/ v. 32:58
引出,激发出(反应、能力)
interpretability /ɪnˌtɜːrprɪtəˈbɪləti/ n. 32:58
可解释性,模型决策能被人理解的程度
shoots up phr. v. 34:52
急剧上升,飙升
hokey /ˈhoʊki/ adj. 36:11
粗糙的,土气的,不太像样的(口语)
hypotenuse /haɪˈpɑːtənuːs/ n. 36:40
(直角三角形的)斜边
standard deviations phr. 37:11
标准差;two standard deviations higher 高出两个标准差
within our collective grasp phr. 37:11
在我们大家力所能及的范围内
seat-of-the-pants adj. 38:42
凭直觉的,凭经验拍脑袋的(原指飞行员凭体感驾驶)
trade-offs /ˈtreɪdɔːfs/ n. 38:42
权衡,取舍
delve into phr. v. 39:19
深入探究
leakage /ˈliːkɪdʒ/ n. 40:45
泄漏;评测中指测试数据混入训练数据
margin /ˈmɑːrdʒɪn/ n. 40:45
差距,幅度。by a reasonable margin 以不小的差距
firm idea phr. 41:34
确切的认识,扎实的判断
coherent /koʊˈhɪrənt/ adj. 42:16
连贯的,前后一致的
with caution phr. 43:04
谨慎地。use with caution 谨慎使用
anonymized /əˈnɑːnəmaɪzd/ adj. 44:20
匿名化的,去除身份标识的
leaderboard /ˈliːdərbɔːrd/ n. 45:21
排行榜
subtleties /ˈsʌtltiz/ n. 47:22
微妙之处,细微差别
refinement /rɪˈfaɪnmənt/ n. 48:03
精炼,进一步打磨;此处指对通用模型的继续训练
passing mark phr. 48:03
及格线
domain-specific adj. 48:03
领域专用的,针对特定领域的
conditioned on phr. 49:15
以……为条件(生成模型术语,指输出受某输入约束)
ferocious /fəˈroʊʃəs/ adj. 49:50
凶猛的,凶狠的
plausible /ˈplɔːzəbl/ adj. 50:48
说得通的,貌似合理的
mascot /ˈmæskɑːt/ n. 50:48
吉祥物
fidelity /fɪˈdeləti/ n. 50:48
保真度,忠实程度。fidelity of text 文字还原的准确度
chunky /ˈtʃʌŋki/ adj. 53:01
粗笨的,块状的,缺乏细节的
astrophotography /ˌæstroʊfəˈtɑːɡrəfi/ n. 54:29
天文摄影
screen a call phr. 55:18
筛查来电,先了解来意再决定是否接听
limited literacy phr. 56:00
识字能力有限
hypothesis space phr. 56:46
假设空间,所有可能候选解的集合
functional equivalent phr. 57:31
功能上的等价物
mutate /ˈmjuːteɪt/ v. 57:31
变异,使发生改变(借用遗传学术语)
synthesis /ˈsɪnθəsɪs/ n. 58:38
(化学)合成
degenerative /dɪˈdʒenərətɪv/ adj. 59:21
退行性的,逐渐恶化的(疾病)
ophthalmologists /ˌɑːfθælˈmɑːlədʒɪsts/ n. 59:21
眼科医生
board certified adj. 1:00:11
通过专业委员会认证的,具备执业资格的
on par with phr. 1:00:11
与……水平相当
gold standard phr. 1:00:11
黄金标准,最高参照标准
benign /bɪˈnaɪn/ adj. 1:01:03
良性的,无害的(医学)
reinforcing /ˌriːɪnˈfɔːrsɪŋ/ v. 1:02:40
强化,加固。reinforcing unfair bias 强化不公平的偏见
accountable /əˈkaʊntəbl/ adj. 1:03:22
可问责的,须负责的。accountable to people 对人负责
underway /ˌʌndərˈweɪ/ adj. 1:03:59
进行中的。a change underway 正在发生的变革
ingest /ɪnˈdʒest/ v. 1:04:32
摄入,吸收(数据)
plug /plʌɡ/ n. 1:05:10
(口语)宣传,推广。put up a plug for 为……做个广告
overwhelmed /ˌoʊvərˈwelmd/ adj. 1:05:10
应接不暇的,不堪重负的
capacity /kəˈpæsəti/ n. 1:06:09
容量;模型的表达能力上限
nuanced /ˈnuːɑːnst/ adj. 1:06:52
有细微差别的,需要分情况看的
exhausted /ɪɡˈzɔːstɪd/ adj. 1:06:52
耗尽的,用完的
assertion /əˈsɜːrʃn/ n. 1:07:25
断言,主张
paraphrase /ˈpærəfreɪz/ v. 1:08:00
换个说法表述,改述
fine-tune /ˌfaɪnˈtuːn/ v. 1:09:08
微调,用少量数据继续训练已有模型
in-context learning phr. 1:09:08
上下文学习,不改参数只靠提示中的示例让模型完成任务
stifling /ˈstaɪflɪŋ/ v. 1:10:53
压制,扼杀
crowding out phr. v. 1:10:53
挤出,排挤(资源或注意力被占用)
time series phr. 1:11:42
时间序列,按时间顺序记录的数据
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.138Emily M. Bender — Language Models and Linguistics 下一期 · NO.140 →Herbert Simon : September 19, 1979 : Complete Talk
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]