Turing Award Winner: TPU vs GPU vs CPU, Computer Architecture, RISC vs CISC | David Patterson · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Turing Award Winner: TPU vs GPU vs CPU, Computer Architecture, RISC vs CISC | David Patterson

节目发布 2026-07-13 · Ryan Peterman
大卫·帕特森 主主持人
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
本文是图灵奖得主大卫·帕特森(David Patterson)接受一档技术播客访谈的实录。帕特森是加州大学伯克利分校教授,精简指令集计算机(RISC)的开创者之一,近年在谷歌参与张量处理单元(TPU)的研发。谈话从四十年前 RISC 与 CISC 的论战说起,一路谈到 GPU 与 TPU 的设计取舍、摩尔定律的现状,以及他对职业与人生的反思。本文依据现场录音编译整理。

微处理器为何模仿大型机

主持人: 能不能先讲讲 RISC 与 CISC 之争?那场争论究竟在争什么,为什么会那么激烈?

帕特森: 事情要从微处理器的诞生讲起。微处理器是二十世纪七十年代发明的,但一开始基本上是个玩具,装在微波炉之类的东西里。我们这些相信摩尔定律的人认为,晶体管数量每一两年翻一番,总有一天,一块微处理器上的晶体管会多到足以构成一台像样的计算机,到那时所有的计算都会交给微处理器。

可当时在英特尔、德州仪器设计微处理器的人,并不是真正的计算机架构师。他们只是照抄大公司的做法。当时引领架构设计、也就是指令集设计的,是 IBM 和数字设备公司(DEC)这样的龙头企业:IBM 做大型机,DEC 做所谓的小型机。小型机有一台或几台冰箱那么大,大型机则要大得多。这些公司拿摩尔定律多出来的晶体管做什么呢?造越来越复杂的指令。当时的理念是,指令集越复杂,抽象层次就越接近软件,这比停留在较低层次有天然的好处。

而我们这些身处这个领域的人,比如我在斯坦福的朋友约翰·轩尼诗(John Hennessy)和我,并不认为这一定是对的。编译器的工作本来就是把程序设计语言翻译到指令集,为什么不能让编译器来填这个差距?于是问题就变成了:对于正在崛起的微处理器,正确的指令集应该是什么样的?

多音节与单音节的赌注

帕特森: 主流观点站在复杂指令集这一边。打个比方,指令集就像一套词汇,复杂指令集相当于词汇表里塞满了多音节的长词。另一条路,也就是我们所说的精简指令集计算机(RISC),词汇表里全是单音节的短词。可以想见,一个程序如果用简单指令来执行,需要的指令条数更多;用复杂指令,条数更少,但每条复杂指令的执行时间可能更长。所以问题归根结底是:这两个比例到底是多少?

八十年代初这场争论刚开始的时候,火药味非常浓。一部分是在争这些比例会怎样,但很大一部分其实是哲学层面的:把指令集的层次降下来,让程序设计语言和指令集之间的鸿沟变大,你们不是在伤害整个软件产业吗?争论的激烈程度正来源于此。

几年之后尘埃落定,我们开始拿到数据。结果是,一个程序用简单指令来写,大约要多执行百分之三十到四十的指令,但这些指令可以跑得快四五倍。两相抵消,RISC 相对 CISC 有大约三到四倍的加速潜力。

凭直觉设计的年代

主持人: 你刚才提到那种哲学上的分歧,那个鸿沟。为什么这会引起争议?

帕特森: 很遗憾,我得说七十年代、甚至八十年代的计算机体系结构,很大程度上是靠直觉、靠感觉在做设计:我觉得这样做是对的,就这么做了。连当时的教科书都像产品目录,这里是一台计算机,罗列它的全部特性,那里是另一台计算机,再罗列一遍。这种局面很让人不满,因为这类争论按理说应该能用科学的方式、用数字来裁决。可没有数据,争论就变成了「一个针尖上能站几个天使」那样的经院辩论。你可以做定性的争辩,却没办法把争论了结。既然没有定量了结争论的手段,大家就只好一直吵下去。

ARM 如何定下胜负

主持人: 那 RISC 与 CISC 这场战争,最后有一方赢了吗?

帕特森: 我最近正好在网上看到有人重新回顾这段历史,他的结论是:CISC 赢了。这个看法相当短视。在个人电脑时代,软件以二进制形式发行非常重要。x86 一旦确立地位,个人电脑上的软件全都以二进制发行,这就很难撼动,成了更换指令集的巨大障碍。所以个人电脑基本上是由 x86 架构定义的。

但同样是八十年代,英国有一家公司想做个人电脑,叫 Acorn。他们决定要有自己的指令集架构、自己的芯片,因为当时市面上能买到的芯片速度都不够快。他们受到我们伯克利那几篇论文的影响,做出了所谓的「Acorn RISC 机器」(Acorn RISC Machine)。精简指令集有一个好处:它更简单,占用的资源更少,执行起来耗的能量也更少。几年之后,苹果想为自己的一款个人设备找一块微处理器,那款设备叫牛顿(Newton),算是 iPhone 的早期先驱。苹果找到这家公司说:「我们很喜欢这块芯片,不过把 Acorn 这个名字去掉吧。」于是缩写 ARM 保留下来,全称改成了「先进 RISC 机器」(Advanced RISC Machine),苹果把它用在了牛顿上。

牛顿在商业上并不成功,但它证明了 RISC 架构在移动设备上的优势。又过了几年,诺基亚推出 GSM 手机,那是最早流行起来的手机之一,他们选用了 ARM。从那以后,ARM 就统治了所有移动设备。我刚查过,到今天为止,采用 ARM 技术的微处理器已经出货三千五百亿颗。所以今天,计算机里百分之九十九的处理器都是 RISC,连个人电脑也不例外:苹果已经从 x86 架构切换到了 ARM。也就是说,即使在个人电脑领域,RISC 架构也已举足轻重。它还在向云端渗透。云端长期以来由 x86 服务器架构定义,但亚马逊、微软、谷歌现在都在开发自己的 ARM 处理器,RISC 处理器在云端越来越多。所以我要说,眼下 x86 架构的市场在萎缩,而 RISC 架构的市场在突飞猛进。

主持人: 既然有百分之九十九这个数字,怎么还会有人说 CISC 赢了?

帕特森: 如果一个人写历史时把「计算机」定义为个人电脑,或许再加上服务器,然后写到两千年前后就停笔,那故事到那里结束,看上去确实像是 CISC 赢了。可一旦进入后个人电脑时代,我不明白怎么还能得出那个结论。

x86 内部翻译的代价

主持人: 你提到了能耗。所以 RISC 在某些场合更有道理,比如移动设备?

帕特森: 在云端也一样。现在所有人都在意能耗,一切都受能量约束。当然,今天我们面对的不再是几百万个晶体管,而是几十亿个,所以指令集本身的影响相对小了一些,芯片上有太多东西在同时进行,可以把这部分开销藏起来。x86 架构为了竞争,做的事情就是在硬件里把 x86 指令翻译成 RISC 指令。你得为这个翻译步骤付出额外开销,才能拿到 RISC 指令,然后 RISC 阵营的任何好点子,x86 都能照用。从财务上讲,这笔额外开销是值得的,因为个人电脑软件积累的价值太大了。英特尔在两千年代初这么做非常合理,是个绝妙的商业决定。

主持人: 有没有什么小众场景,CISC 是合理的?这是一种工程上的权衡,还是说 CISC 客观上就更差?

微程序解释器的包袱

帕特森: 要回答这个问题,我们得往下再深一层,看看设计计算机时究竟在做什么。难点在控制逻辑。早期的控制逻辑相当随意,你把逻辑门拼在一起,拼到能工作为止。计算机先驱莫里斯·威尔克斯(Maurice Wilkes)想出了一种更优雅的设计方法:把所有控制信号当作一块存储器的输出,再用一个部件记录当前在存储器里的位置,依次发出控制信号。他把这些控制信号看成指令,称之为微指令,把编排微指令的工作叫做微程序设计。

以六十年代的技术水平,这么做相当合理,IBM 于是造出了所谓的微程序计算机。它本质上是一个解释器,用非常简单的指令去解释上层那套复杂得多的指令集。代价是解释的开销。计算机科学里有个经典结论,解释比编译大约慢五到十倍。但考虑到当时存储技术的延迟,加上可以用只读存储器来实现,这在六七十年代是说得通的。到了一九八零年前后,问题就来了:这还是个好主意吗?我们还应该在处理器里放一个微码解释器吗?另一条思路是,既然里面有个微码解释器,为什么不直接编译到那些指令上?这已经很接近 RISC 的想法了。当时的微指令有一百来位宽,非常复杂。如果把它做得不那么长,更自然一些,就可以跳过解释这一步。

回到你的问题,从这个层面看,今天还会有人发明一套复杂到需要微码解释器的指令集吗?大概不会。你可以这么做,没有什么阻止你,但你不会想设计一套非得靠微码解释器才能实现的指令集。也许在某些极小的应用里,芯片上只有几千、几万个晶体管,微码解释器还有用武之地。但我想,过去二十年里没有任何人发明过需要微码解释器的指令集。

编译器与寄存器分配

主持人: 你多次提到编译器,它似乎是让 RISC 行得通的关键一环。能讲讲编译器的角色吗?它是怎样协调软件与硬件的关系的?

帕特森: 你用 C、C++ 或者 Python 这样的语言写程序,但生成的代码质量取决于编译器。举个具体例子。构造计算机时,设置寄存器很有用,而寄存器在汇编语言或机器语言层面是可见的,可能有八个、十六个或三十二个,供在那个层次编程的人使用。过去,编译器很难高效地分配寄存器,因为计算机不够快,我们也没有相应的算法,能盯着一段代码或一个子程序,判断怎样使用寄存器最有效率。

C 语言就是为系统编程发明的。在此之前,操作系统是用汇编语言写的,说出来你可能不信。Unix 的两位作者肯·汤普森(Ken Thompson)和丹尼斯·里奇(Dennis Ritchie)证明,用一门相当底层的语言,也能获得高级语言的好处:人更容易理解,更容易调试。但因为编译器的寄存器分配做得太差,他们不得不加入一个功能,让程序员来提示哪些变量该放进寄存器。让程序员自己出面更省事:机器有八个寄存器,我要这六个变量放在寄存器里,别留在内存里,内存太慢了,寄存器快得多。

所以 RISC 与 CISC 之争的一个重要部分,就是编译器算法正在进步,它们能处理这些低层次的指令,能高效地分配寄存器。这是当年支持简单架构的又一条理由。我们在 RISC 架构里做的事情是:既然寄存器这么重要,寄存器分配这么重要,那么让编译器省事的办法之一,就是多给它一些寄存器。当时的 CISC 架构通常有八个或十六个寄存器,我们放了三十二个。理由是,既然少量寄存器很难高效利用,那就给足。即使编译器不够聪明,寄存器也够用。而且拜摩尔定律所赐,在机器里多加寄存器并没有贵多少。

编译器不用的复杂指令

主持人: 我在你的一场演讲里看到,编译器为 RISC 优化代码更容易,而 CISC 里那些复杂的大指令,编译器几乎从来不用。这是编译器的缺陷吗?

帕特森: 回到那个年代,当时的论点是:更复杂的指令能抬高抽象层次,缩小鸿沟,让编译器更轻松。但这是个哲学论点,并不是编译器的人能够落实的东西。提出这个论点的不是做编译器的人,而是架构师。

我们做早期 RISC 与 CISC 研究时去看实际的程序,结果发现编译器根本不用那些指令。常见的情形是,架构师设计出一条复杂指令,编译器的作者说:「我们不需要这个。」我们找到了不少这样的例子。比如过程调用的入口,架构专门内置了一条指令,把架构师以为编译器想要的全部工作一并做掉,而编译器的设计者说:「我们用不着,用几条单独的指令来做,比用你那条复杂指令还快。」

于是局面就成了这样:你为了拥有这些复杂指令,付出了微码解释器的额外开销,而编译器根本不去用它们。这是一个荒谬的处境。创业公司为什么会出现,科学上的转折点为什么会出现,道理是一样的。当我们把摩尔定律之下的所有技术摊开来看:高速缓存这样的新想法,编译器对复杂指令的实际使用情况,更高效分配寄存器的能力,把方向从微程序指令集架构转向更简单的架构,是合乎情理的。

登纳德缩放的终结

主持人: 我们谈指令集的时候,默认谈的是通用计算机,也就是 CPU。现在大家谈得很多的是 GPU,也许还有别的计算形式。那些机器也有指令集吗?

帕特森: GPU 在两千年前后出现,我们称之为领域专用架构。GPU 是图形处理单元,只有一项任务,不需要做通用处理器必须做的所有事情。它不必支持虚拟内存,甚至不必支持编译器,这对架构来说是相当激进的想法。它只为图形服务。英伟达和其他做 GPU 的公司当时的目标是给游戏和电影渲染图形,那是个小众产品。

计算机产业的走势由两条定律推动。一条是我反复提到的摩尔定律,另一条不那么出名,叫登纳德缩放(Dennard scaling)。有个有意思的问题:既然摩尔定律让芯片上的晶体管越来越多,每一两年翻一番,芯片为什么没有越来越烫?答案来自鲍勃·登纳德(Bob Dennard)的一个观察:增加晶体管的同时,人们也会降低阈值电压,也就是区分零和一的那个电压,而电压的影响是平方级的。于是晶体管数翻番,阈值电压下降,微处理器的功耗就一直维持在二三十瓦。我和约翰后来写了一本教科书,我查了一下,第一版是一九九零年出的,前三版根本没有把功耗当作一个问题来讨论。第三版是两千年出的,功耗仍然不是一个话题。因为登纳德缩放一直在起作用,微处理器越来越快,功耗却始终停留在几十瓦。

大约在二零零五年,登纳德缩放失效了。这是一个冲击。英特尔有一代微处理器就因此失败,功耗压不下来,太烫了。这迫使我们转向多核。在此之前,对程序员、对所有人来说,最省事的方案是一颗非常精巧的处理器包办一切,但这条路走不下去了。于是从一颗精巧的处理器变成两颗、四颗、八颗更简单的处理器。要兑现摩尔定律的潜力,就得靠程序员把代码并行化。

这种状态又持续了大约十年,然后摩尔定律开始放缓,通用微处理器几乎不再进步。它还有一点改善,但不再是戏剧性的飞跃。八十年代、九十年代、两千年代,你有一台笔记本电脑,朋友的笔记本可能比你的快四倍,你会嫉妒。性能变化太快,你会扔掉完好无损的硬件,就因为朋友的机器快太多。到了二零一零年代,这一切都结束了。你不会再因为新机器快得多而扔掉旧笔记本,只会等它坏了、变慢了才换。可程序员已经习惯了每隔几年性能就大幅提升,因为这样他们就能往软件里加更多功能。

那么架构师还能做什么?多核这一招在二零零五年前后已经用过了。于是到了二零一五年前后,思路变成了做领域专用架构,就像 GPU 那样。如果你告诉我,我只需要运行一小类程序,不必运行所有的操作系统和其他一切,那我确实可以重新调配资源,把这一类事情做得高效得多。有些事做得很好,另一些事要么做得很差,要么干脆不做。

机器学习成为那个领域

帕特森: 接下来的问题是:选哪个领域?非常巧,就在技术走到这一步的二零一二到二零一五年,机器学习和人工智能爆发了。该选哪个领域,答案不言自明,就是机器学习和人工智能。我在谷歌工作,但我想公平地说,谷歌是第一家真正看到机器学习潜力的大公司。他们押注,或者说担心,需求会把他们淹没,必须做定制硬件。谷歌二零一六年推出的张量处理单元(TPU)震惊了世界,让大家意识到我们应该为机器学习专门设计硬件,而且在这一个领域,性能可以继续大幅提升。

三种处理器的分工

主持人: CPU、GPU、TPU 这三者在高层次上有什么区别?

帕特森: CPU 必须是通用的。直到今天,CPU 里的通用核心,每一个都和二十年前的处理器很相似,设计上没什么意外,只是核心数很多,现在的 CPU 可能有五十到一百个核心。

图形处理的关键是内存系统要有很高的性能,所以 GPU 走了多线程的路线。它有硬件线程,发出一个内存请求之后,硬件切换去做别的事,等数据从内存回来再继续。这就是高度多线程的架构,它对图形处理很有效。图形还有一个特点,不需要很强的浮点运算,确切地说,不需要很宽的浮点。三十二位浮点对图形绰绰有余,十六位也行。所以 GPU 在多线程架构上推进十六位和三十二位浮点。由于它是自成一体的东西,不在计算机体系结构的主干上,它有一套自己的术语。我们的教科书里有一张类似罗塞塔石碑的对照表,一边是英伟达描述 GPU 的术语,另一边是这些术语在主流处理器设计里的对应说法。

所以 GPU 曾经是一种小众产品,恰好单精度浮点很快,半精度也很快,而且很便宜,几百美元一块。于是有人开始琢磨:对某些应用,如果我能把问题伪装成生成图像,就能用这些便宜的 GPU,它每一美元的浮点性能比任何东西都高。人们开始这样折腾。

英伟达的创始人兼首席执行官黄仁勋很喜欢这个想法。二零零六年,他出资开发一门程序设计语言,来驾驭这种为图形设计的多线程硬件架构,让它更容易编程。这就是 CUDA 的由来,我记不清它的缩写展开是什么了,本质上是一门为多线程 GPU 架构服务的专有语言。它类似 C,但不能直接拿 C 程序编译运行。人们确实喜欢它,至少比把程序改造成图像生成强多了。黄仁勋的愿景是,让那些在地下室打游戏的少年学会给这些东西编程。他还盯上了几个市场,比如能源部实验室的流体力学计算,他为这些领域构建专用库,用 GPU 去做图形以外的专用计算。这就是 GPU 的渊源。

从一开始就有人用 GPU 做机器学习,因为它的单精度浮点性能远超 CPU,处理器数量多得多,潜在性能大得多。突破性的时刻在二零一二年。机器学习界里,神经网络这一支只有少数拥护者,很多人不相信它。那年有一场图像识别竞赛,所谓的 AlexNet 击败了所有对手,这是机器学习和神经网络历史上的标志性事件。做出 AlexNet 的那个人在多伦多大学上过 CUDA 课程,学会了怎么用它,于是想,既然要做,就在 GPU 上做。廉价而快速的 GPU 让他能探索大得多的空间。他是参赛者里唯一用神经网络的,把对手打得落花流水。几年之内所有人都转了过来,不但改用神经网络,而且改用 GPU。所以 GPU 的血统是图形引擎,但更可编程,然后被用到了机器学习上。

TPU 的取舍

帕特森: 谷歌入场的时候,选择从一张白纸开始。他们不关心图形。神经网络的核心是矩阵乘法,就是它。于是他们设计的处理器里有一个巨大的矩阵乘法单元,这是主体,然后把不需要的东西统统扔掉。通用计算芯片上很大一部分面积是三级高速缓存,目的是避免把时间都耗在访问相对缓慢的内存上。而对机器学习来说,内存访问的时机是事先知道的,可以安排好,硬件缓存就没有意义了。他们只放了一块由软件管理的存储器,数据会按时传进来。这是一些创新。

他们还在浮点格式上创新。科学计算非常在意精度,大多用六十四位浮点,指数不到十位,剩下五十多位都是尾数。谷歌意识到,机器学习不需要那么高的精度,需要的是范围。于是谷歌做出了第一种指数位比尾数位还多的浮点格式,这是个激进的想法。它叫 bfloat16,全称 brain float 16,因为做这件事的是谷歌大脑(Google Brain)研究组。

所以这个架构有窄浮点格式,有大矩阵乘法单元,只有一个处理器,和其他产品不同,它只做机器学习。它把整个领域的人都震住了:推理性能比同期的 GPU 高三十倍,比 CPU 高八十倍。谷歌在内部部署一年多之后,在年度活动上宣布了这件事,所有人都坐不住了。英特尔开始收购公司,英伟达开始为机器学习修改设计,其他竞争者、各家超大规模云厂商都开始自己干。我认为 TPU 的发布是这里的分水岭。

主持人: 这听上去像是专用化的层级,CPU 最通用。

帕特森: 是的,但如果登纳德缩放还在,如果摩尔定律和登纳德缩放都还在,通用处理器今天会走到哪里?我们应该有一百太赫兹的微处理器了。如果能造出一百太赫兹的微处理器,我们就会那么做,GPU 会依然是个小众产品。那会水涨船高,皆大欢喜。可这已经是遥远的历史了,我们二十年没能做到。二零零五年就有两三吉赫兹的微处理器,今天差不多还在那里,几乎没有进步。所以我们做不到。CPU 依然有它的用武之地,操作系统、编译器这些东西需要它,它是正确的方案。但其他事情都专用化了。用产业的话说,现在巨量的投入都在往 AI 加速器、往 AI 上砸,钱都投在那里。于是每一个处理器设计者面对的问题是:我的处理器设计在这个 AI 宇宙里处于什么位置?为了这个重要的应用,我该怎么改?

摩尔定律死了吗

主持人: 你说摩尔定律在放缓。我看过吉姆·凯勒(Jim Keller)的一场演讲,他说摩尔定律没有死。说它放缓,这个说法有争议吗?

帕特森: 在真正的工程师那里没有争议。摩尔定律非常简单:一块芯片上的晶体管数量,最初说是每年翻一番,后来他修正为每两年翻一番。你只要去看,芯片上的晶体管数翻番了没有?没有。我想人们误以为,不再遵循摩尔定律就意味着技术不再进步,这不是一回事。技术只是不再以摩尔预测的速度进步。摩尔定律之所以了不起,是它持续了五十年,而且指引了半导体制造业的投资:我们得每一两年让晶体管数翻番,怎么做到?要造什么设备?它是整个行业的指导方针,这很不寻常。

现在的情况是,技术的某些部分还在进步,某些部分完全不动了。芯片上有一个很重要的部件是静态随机存取存储器(SRAM),它几乎不再进步。逻辑门还在进步,如果你做的是加法器、乘法器,这些部件还在变好。进步不再是均匀的了。我们还在用越来越奇特的封装来交付性能。过去一直是单芯片最好,所有东西放在一块芯片上就行。现在有了小芯片(chiplet)的思路,把多块芯片封装在一起。最新的 GPU,我想还有最新的 TPU,是把两块所谓的全光罩尺寸的裸片封装到一起。全光罩尺寸是半导体上能造的最大面积,光刻机步进重复,过去一个光罩里能放很多颗芯片,现在只放一颗。两块最大光罩的裸片组成一个节点,靠的是封装。所以从外面看,你可以说,看这么多晶体管,摩尔定律还在继续。但看看芯片内部,你就知道那已经放缓了。

我想这在一定程度上是情感问题。在半导体制造行业,很多人被问到你是做什么的,回答是「我造摩尔定律,我维持摩尔定律」。几十年都在干这个,突然有人说摩尔定律结束了,那就等于说你的职业生涯结束了。所以有情感的一面。但看数据就好,数据不支持凯勒的说法。我也没有说技术不再进步,技术在继续进步,尤其是领域专用架构。可如果你看通用架构,或者看其他存储技术,过去 DRAM 的密度每三年提高四倍,像钟表一样准,现在提高四倍可能要十年。有大量证据表明摩尔定律已不再适用。

封装、低位宽与骨架

主持人: 现在有没有某种新版本的摩尔定律,或者类似的东西,在指引微处理器的设计和架构?换个问法,英伟达和谷歌都在持续推出快得多的机器学习处理器,进步惊人。他们在做什么?

帕特森: 一部分是我说的封装,让你能塞进更多晶体管,同时把它们放得更近,因为距离很重要。一部分是在浮点格式上创新。超级计算机用六十四位,现在不但不是六十四位,不是三十二位,甚至不是十六位,八位和四位浮点都在用,数据类型在不断变窄。还有所谓的矩阵乘法指令,把这些能做重要专用计算的强大单元做得更大、更快。这些是正在发生的事情。但我们没有一条简化的指导原则贯穿其中。你必须了解技术的每一个部分,判断它进步得多快,能不能兑现,然后把这些拼起来下注。

我们刚有一篇论文通过审稿,很快会放到 arXiv 上,讲的是谷歌 TPU 这一系列。很了不起的一点是,谷歌 TPU,特别是训练用的 TPU,基本架构设计一直相当稳定。东西变大了、变快了,但如果你看它的设计,回到第一代训练 TPU,那张框图到十年后的今天仍然成立。二零一五年前后设计它的那些人干得真漂亮。主要部件还是那几样:一个大矩阵乘法单元,所谓的高带宽内存,一个配合矩阵单元的向量单元,基本构件都还在。

基准测试与软件护城河

主持人: 在你研究 CPU 的年代,基准测试在衡量架构性能上起了巨大作用。在 AI 浮点运算这个新领域,GPU 有大家公用的基准吗?也能用于 TPU 吗?

帕特森: 我和几位朋友参与推动了 MLPerf。它的灵感来自 SPEC CPU 基准。当年几家互相竞争的公司都声称自己比别人强,后来意识到这对行业没有好处,于是商定了一套共同的基准。MLPerf 现在由一个叫 MLCommons 的组织运营,目的也一样:既然要比较,就别在基准是什么这件事上争吵。这是一项认真的努力。

有意思的是后来的发展。设计有两部分,一部分是硬件架构,另一部分是机器学习库,实现这些应用需要的大量功能,而且库还会针对特定应用重写,而不只是提供通用库。这成了英伟达的一大优势,因为它是一家大公司,有大量工程师可以去写这些库。所以结果并不是一场中立的评测。比的是架构加上配套的库,还有编译器,尤其是每次都要量身定制的库。英伟达每发布一个新架构,就会修改一套库,让新架构跑得特别好,或者让应用跑得特别好。这是很强的组合。商业上大家说英伟达有「CUDA 护城河」,一部分是 CUDA 这门语言,但很大一部分是英伟达写的那些库。这让创业公司很难与英伟达竞争,一方面是他们在软件上投入不够,另一方面是英伟达的工程师就是比他们多得多,能把库调得更好。所以是库加架构构成了这个强大的优势,也是大多数人都用 GPU 的原因。

谷歌能开发自己的库。他们工程师没有那么多,比英伟达更依赖编译器,但也有一套自己的库。不过直到最近,谷歌的这些东西都只在内部使用,或者通过云服务使用。创业公司因此很吃力。这也是 MLPerf 没有那么普及的原因:不是所有人都跑它,因为英伟达跑得实在太好,远超所有人。创业公司没有那么多工程投入去打磨库,在 MLPerf 上很难展示自己的实力。

「糟糕职业」背后的清单

主持人: 你有一场很受欢迎的演讲,讲「如何拥有糟糕的职业生涯」,其实是把好建议反过来讲。能不能概括一下,你最希望人们记住的三件事?

帕特森: 说起来有一段渊源。我职业生涯早期,一位好友和我商量,怎么教研究生做好一场报告。我们觉得反过来讲「如何做一场糟糕的报告」会很有趣:如果你不想讲砸,这些就是要避开的坑。后来我又做了「如何拥有糟糕的职业生涯」,主要面向学术界的研究者。再后来是「如何建一个糟糕的研究中心」「如何建一个糟糕的实验室」。最近我在讲的是「如何让 AI 留下糟糕的碳足迹」,这是最新的一场。

不过可能更有用的是,我在演讲结尾通常会回顾自己的职业生涯,谈谈学到的东西。我最近写了一篇文章,叫《我职业生涯前五十年的人生教训》,分成职业建议和个人建议两部分。

个人这边,我要说:如果你有家庭,把家庭放在第一位。我们发明的技术让你太容易把工作带回家,忽略家人。我刚到伯克利的时候,有一次从硅谷回来,深夜送一位资深教授回家,他说:「戴夫,如果能重来一次,我希望多花点时间陪家人。」我从来不想说出这句话。没有人在临终时说,我真希望多在办公室待一会儿。让事业压过家庭的需要,永远是容易的,所以你得时刻记住这一点。

再说说财富。我是五十年代长大的,那时候的观念是,要快乐就得富有。但富有和快乐其实是两个不同的目标。我做决定时一律朝快乐这边倾斜,而不是财富,我对此感觉很好。就算到现在我也会问,为什么要选一样让我富有却不快乐的东西?你为什么要那样做?而我们这个行业碰巧财富也不少,所以追求快乐并不意味着要吃苦。

还有,玩得开心很重要。小孩子不用人教就会玩,可成年人太忙,觉得没时间玩。但你只能活这一次。我现在踢足球,骑自行车来参加这场访谈,举重,冲浪,和妻子、家人、儿子一起做事。玩很重要。

重要但不紧急的事

帕特森: 职业这边,有一条建议来自《高效能人士的七个习惯》那本书的作者。他画了一个四象限,一个维度是紧急与不紧急,另一个是重要与不重要。我们的技术,电子邮件、短信之类,都在诱使你盯着紧急的事,但你真的不该把大量时间花在那些不重要却紧急的事上。你需要自律,专门留出时间做重要而不紧急的事。不把这些时间圈出来,你可能什么都做不成。我在谷歌能看到别人的日程表,有些经理从早上八点到下午六点,每半小时都排满了,一周五天。我不明白他们哪来的时间思考和反省。

另一条职业建议是这样来的。有一天早上我醒过来,仿佛上帝对我说话,我像被雷劈中一样。那句话是:重要的不是你开了多少个头,而是你做完了多少件事。听起来相当显而易见,但我当时并不是这么做的,手上同时有很多事。从那以后,我一次只有一件主线。我和轩尼诗写教科书的时候,那就是主线;我当系主任的时候,那就是主线,旁边只做一点小事。约翰·轩尼诗也写过一本书,基于他当斯坦福校长的经历谈职业建议。他在书里说,人们只会因为你一生做过的五六件事记住你,而不是几百件小事。要给自己一个机会,做出几件真正自豪的事,最好把精力集中在少数几件上,指望其中有些能成大事,而不是把自己分散到许多事情上。有兴趣的话,搜「人生教训 大卫·帕特森 前半个世纪」,能看到完整的清单,一共十六条。

人比项目更值得在意

主持人: 你说不想回顾人生时觉得陪家人的时间不够。我从没听过有人说相反的话。为什么会这样?为什么所有人回头看,都不后悔陪家人太多?我猜总得有一个人会说:「我陪家人陪太多了,我的事业受损了。」

帕特森: 这是个相当哲学的问题:人生究竟是为了什么?什么叫成功?你得自己想清楚它对你意味着什么。如果你有财务目标,要当百万富翁、亿万富翁,并以此评判自己的人生,那好,祝你好运,这很难做到。

我快读完博士的时候,读了斯特兹·特克尔(Studs Terkel)的《工作》。他采访了各行各业的人,请他们回顾自己的职业,喜欢什么,不喜欢什么,有什么感受。我从中读出的是:和人打交道的人,牧师、教师、医生,对自己的职业感觉很好;而做那些转瞬即逝的东西的人,比如技术产品、飞机之类早已消失的东西,感觉就没那么好。他们真正在意的是共事过的人。前不久我和这里一位退休的工程学院院长一起去开会,他说:「戴夫,回头看,重要的不是项目,而是共事的人。」我心想,这我很早以前就知道了。这是一个上了年纪的人给出的忠告:你会更在意共事过的人,你帮助过的人会更重要。

关于个人幸福还有一点。心理学家过去只研究精神失常的人,后来开始研究人为什么快乐,而且他们知道原因:有一份你喜欢的工作,有朋友和家人,帮助别人,帮助别人会让你快乐,这是有定论的;还有某种精神层面的东西,不一定是正式的宗教,比如接触自然,感受自然的壮阔。让人快乐的清单是清楚的。而这个世界上满是不快乐的亿万富翁。如果钱能让人快乐,这些极其富有的人为什么对什么都那么愤怒?这就是我传下来的建议。

勇气与树敌

主持人: 你有一场演讲里有一页叫「对我管用的东西」,其中有一条很特别:你说勇气是你职业生涯里很重要的一部分。这不常听到。为什么勇气对职业这么重要?

帕特森: 这可能部分出于我个人的性格。我在班里是年纪最小的孩子,个子也小,发育得晚,所以一直很小。父母鼓励我去练摔跤。摔跤给你身体上的自信,我在高中和大学练了很多年。

从技术上讲,做事要有勇气,和一句老话是一致的:幸运眷顾勇者。这句忠告有两千年了。要想清楚这件事很难。海伦·凯勒写过,即使你处处求稳,照样会被逮住。所以不管怎样你都可能失败。如果你冒大险,有可能成功;如果不冒险,求稳妥,多半不会成功。幸运眷顾勇者,而这需要勇气。

对我来说,在智识上也是如此。如果有什么不对的东西,我觉得我必须站出来面对它。说来讽刺,这也来自我性格里摔跤的那一面:看到有人被欺负,我会站出来阻止。在智识上我感到同样的责任,如果有人在做糟糕的论证,或者在做不该做的事,我们需要站出来。我对此感觉很好。需要提醒的是,我的一位资深同事看出了我这个性格,他对我说了一句老话:朋友来来去去,敌人却越攒越多。想想看,高中同学里一时的朋友你会渐渐忘掉,可一个真正讨厌你的人,你永远不会忘记自己讨厌他。所以,在重要的时候要站出来,但树敌要小心,因为他们会跟你很久。

主持人: 你说的「技术上出了问题」,是指有人说错了吗?

帕特森: 是指论证站不住脚,无论是政治上还是技术上的弱论证。我真心喜欢有一个思想的市场,通过争论把想法磨得更锋利。如果一个论证似是而非,哪怕它出自公司的领导者,就是说不通,我觉得有人站出来指出来,对公司更好,而不是任由它蒙混过去。这有点对抗性,但只要大家都同意这是为了更大的善,我们要把正确的想法摆出来,那就来争论这些想法,把它们打磨得更强。我认为这在科学、工程里很重要,在生活里也一样。眼下这个国家正在发生很多令人担忧的事,我肯定站了出来,写过评论文章,谈我认为错误、需要纠正的事情。如果人们害怕这么做,害怕在出现错误时站出来阻止,就很难对未来保持乐观。

乐观与九个词

主持人: 你在演讲里还提到乐观,讲了一个故事,不知道你愿不愿意再讲一次。

帕特森: 当然。在工程上,你很难事先知道对错,但我认为你需要乐观,需要积极的心态,因为可能出错的事太多了。我有个自己的故事来说明这一点。回到高中,我十六岁,在和一个很漂亮的女孩约会。我鼓起勇气问她,我们能不能确定关系,当时的说法叫「稳定交往」。她看着我,她也十六岁,跟别的男孩约会过,觉得我们都还太年轻。她说:「戴夫,你人这么好,我不知道该怎么说不。」对我这个讲逻辑的人来说,这听起来像是同意了。于是我抱住她说:「太好了。」她心里想的是,以后再慢慢让他死心。可我们已经结婚五十九年了,她到现在还没让我死心。这就是乐观得到回报的例子。

主持人: 听到一段这么长久的健康关系,大家都会想知道你是怎么做到的。

帕特森: 我过去常对人说,去参加婚礼,你会发现婚礼誓词写得真好,可没有人记得住自己的誓词。我以前说「记住你的誓词」,没人记得住。所以我们把它浓缩成九个有魔力的英文单词,其实就是三句话,都以「我」或「你」开头,三句都得说:我错了,你是对的,我爱你。就这九个词。这对关系中的双方都适用,不只是一方。三句都要说,不许替换,「我错了,你是对的,你是个混蛋」是不行的。记住这九个词,能帮你拥有一段像我和妻子那样长久的关系。

回到起点的忠告

主持人: 最后一个问题。带着你今天的全部经验,如果能回到刚入行的自己面前,你会给自己什么建议?

帕特森: 我刚到伯克利的时候有冒名顶替综合征。我是加州大学洛杉矶分校的研究生,忽然成了伯克利的教授,这让人很有压力。但过了一阵我想,我大概拿不到终身教职,那不如好好享受。所以我觉得我当时的心态已经是对的。第一年很煎熬,一边应付冒名顶替的感觉,一边学着当伯克利的教授,但之后我处理得不错,陪孩子的事一件也没落下。

这个问题还有另一个版本:有没有什么事我想重来一次?有一件。我当过体系结构学界的主席,就是 SIGARCH,它每年办一次会议。那是九十年代,我是主席。我当时不知道,在这些会议上有男性在骚扰年轻女性。我压根没往那里想:像我这样的人,年轻人,没人会干这种事,只有白痴才会,这不可能发生。可它确实在发生。我真希望当时有人跟我说一声,因为我会去纠正那些人,我会威胁他,敢这么干就要他的命。唯一让我稍感安慰的是,著名的计算机架构师萨丽塔·阿德韦(Sarita Adve)说她当时也不知道。五到十年之后,这件事才逐渐清楚,也才有了处理机制。这是我唯一想回到过去改变的事:我会弄清楚状况,去纠正那些男人,我相信那能制止他们。

主持人: 非常感谢你今天抽出时间。

帕特森: 谢谢你的采访。

本期讲者
大卫·帕特森加州大学伯克利分校计算机科学荣休教授、谷歌杰出工程师,2017 年图灵奖得主。开创 RISC 精简指令集方法,与 John Hennessy 合著《计算机体系结构:量化研究方法》,参与 RISC-V 与 MLPerf 的创立。
主持人技术类播客主播,曾邀请 Barbara Liskov、Mike Stonebraker 等计算机科学家做客;同时在 Kickstarter 众筹一款分体式人体工学键盘。
章节 · 点击跳转视频
0:00 精简指令集之争的起源 ▶ 正在看
5:08 x86 锁定与 ARM 逆袭 ▶ 正在看
9:57 微程序解释器的代价 ▶ 正在看
12:59 编译器与寄存器分配 ▶ 正在看
17:56 GPU 与登纳德缩放终结 ▶ 正在看
21:59 TPU:为机器学习定制 ▶ 正在看
32:12 摩尔定律放缓与封装创新 ▶ 正在看
37:45 基准测试与 CUDA 护城河 ▶ 正在看
41:53 家庭、幸福与专注 ▶ 正在看
49:49 勇气、乐观与婚姻九字 ▶ 正在看
56:09 冒名顶替与唯一的遗憾 ▶ 正在看
58:15 片尾:播客与键盘众筹 ▶ 正在看
本期论点
本期回应
5:19
「CISC 赢了」是短视结论,它只在个人电脑时代、且叙事停在 2000 年前后时才成立 会被翻掉已经领先的人会被换掉吗?
5:29
软件以二进制形式分发,会让已经站稳脚跟的指令集极难被替换 守得住已经领先的人会被换掉吗?
40:27
创业公司难以与英伟达竞争,是因为对软件投入不够重视、工程师人数也远不及 守得住已经领先的人会被换掉吗?
28:59
机器学习的访存时机可以提前调度,因此硬件缓存对机器学习芯片没有意义 就该专门造AI 该用专门设计的芯片吗?
29:29
机器学习需要的是浮点数的动态范围,而不是高精度 就该专门造AI 该用专门设计的芯片吗?
其他论点
2:28
指令集的复杂度应该由编译器承担,而不是靠硬件把抽象层次抬高到接近软件
3:59
按实测数据,RISC 相对 CISC 的整体提速大约是三到四倍 观察
14:44
编译器寄存器分配算法的进步,是 RISC 简单指令集能够胜出的关键前提
19:44
2005 年前后登纳德缩放定律失效,逼着整个行业从单个复杂处理器转向多核 观察
33:18
芯片技术的改善已不再均匀,逻辑门仍在变好,SRAM 几乎毫无改善 观察
39:41
MLPerf 衡量的是架构加上编译器和为应用定制的库,并不是中立的评测
43:02
有家庭的人应该始终把家庭放在第一位,不能让事业压过家庭的需要 做法
46:05
重要的不是开了多少个头,而是完成了多少件事,同一时间只应专注一件事 做法
51:16
冒大险才有可能成功,一味求稳多半不会成功,因为求稳同样躲不掉失败
52:22
该站出来的时候要站出来,但树敌要谨慎,因为朋友会淡忘而敌人会记你很久 做法
01精简指令集之争的起源
0:00
Well, if Moore's law puts a lot more transistors in a chip, how come they're not getting hotter and hotter? >> This is David Patterson, Turing [music] award winner, famous for his contributions to computer architecture. And we discussed the historic risk versus CISK debate >> and he said and CISK won. Well, that's a pretty myopic view. Everything is energy bound now. Register [snorts] allocation by the [music] compilers was so poor >> and his thoughts on modern GPU and TPU architecture. What's the highle differences between a CPU, [music] a GPU, and a TPU?
那么,如果摩尔定律往一块芯片里塞进更多晶体管,为什么它们没有变得越来越烫?>> 这位是 David Patterson,图灵[音乐]奖得主,以在计算机体系结构领域的贡献而闻名。我们聊了那场著名的 RISC 与 CISC 之争。>> 他说,最后是 CISC 赢了。这个看法其实相当短视。现在一切都受制于能耗。编译器做的寄存器[吸气声][音乐]分配当年实在太差了。>> 还有他对现代 GPU 和 TPU 架构的看法。CPU、[音乐]GPU 和 TPU 之间最主要的区别是什么?
便签引用
0:33
>> What they realize for machine learning and AI? >> Here's the full episode. Maybe you could explain risk versus CISK kind of like what that debate was and why was it controversial. What happened is in kind of in you know the microprocessor was invented in the 1970s and but in the beginning it was basically a toy. It'd be something that would go in microwaves and things like that. Those of us who believed in Moore's law believed eventually the uh microprocessor would be the way we did all our computing with the doubling of transistors every year or two.
>> 它们为机器学习和 AI 带来了什么?>> 以下是完整这一期。也许你可以讲讲 RISC 和 CISC,那场争论到底是怎么回事,为什么会那么有争议。事情是这样的,微处理器是上世纪 70 年代发明的,但一开始它基本上就是个玩具。它会被用在微波炉之类的东西里。我们这些相信摩尔定律的人当时就相信,随着晶体管数量每一两年翻一番,微处理器最终会成为我们做一切计算的方式。晶体管每一两年翻一番。
便签引用
1:11
eventually there'd be enough transistors that on one of these microprocessors would be a serious computer. So the people who were designing microprocessors like in Intel and Texas Instruments weren't really computer architects. So they just imitated what the big companies did. So companies leading companies like IBM or at the time digital equipment corporation uh IBM for mainframes digital equipment for what were called minicomp computers they kind of drove the uh architecture the instruction set design. So what they were doing with Morris law was building more and more sophisticated u instructions. They had they had you know mini computers be the size of a a refrigerator or a couple of refrigerators and mainframes would be the size of many of those. So that's what they were doing with more hard sources. And the philosophy at the time was that by having a more sophisticated instruction set you kind of raise the level of abstraction closer to the software. So that had some inherent benefits over having something at a
最终晶体管会多到,其中一块微处理器上就能装下一台真正意义上的计算机。所以,当时英特尔、德州仪器这些公司里设计微处理器的人,并不是真正的计算机体系结构专家。他们只是照搬那些大公司的做法。像 IBM,还有当时的 DEC(数字设备公司),IBM 做大型机,DEC 做所谓的小型机,是这些公司主导了体系结构、也就是指令集的设计。所以他们靠摩尔定律做的事,就是造出越来越复杂的指令。当时的小型机大概有一台冰箱那么大,或者两台冰箱那么大,而大型机则相当于好多台冰箱。所以这就是他们拿更多硬件资源去做的事。而当时的理念是:指令集越复杂,你就越是把抽象层次抬高、更靠近软件。这相对于停留在更低的层次,被认为有某些固有的好处。而我们当中一些人,
便签引用
2:15
lower level. Now from a c some of us uh who were in that area like my friend John Hennessy at Stanford and I we thought that wasn't necessarily the right thing to do. Compilers would go from programming languages down to the instruction set. So why couldn't compilers hold that? And then the question is what's the right instruction set for this emerging microprocessor? And the prevailing wisdom of these more sophisticated instruction sets, complex instruction sets. And if you think of it like a vocabulary, it was like a having lots of polyelabic words in your vocabulary, you know. And the alternative was going to be what we called the reduced instruction set computer was having lots of reduced instructions like monoselabic words. So you might guess that for a program to execute these instructions, if they're simpler, they take more of them and if they're sophisticated, they take fewer, but the sophisticated ones might take longer to run. So kind of the question came down to what was that ratio? Um,
比如我在斯坦福的朋友 John Hennessy 和我,我们觉得这未必是对的做法。反正编译器要把程序语言一路翻译到指令集。那为什么不能让编译器来扛这件事呢?接下来的问题就是:对这种新兴的微处理器来说,什么才是合适的指令集?当时的主流看法偏向这些更复杂的指令集,也就是复杂指令集。如果把它比作词汇表,那就好比你的词汇表里有一大堆多音节的长词。而另一条路,就是我们所说的精简指令集计算机,也就是用大量精简的指令,像单音节的短词。所以你大概能猜到,一个程序要执行这些指令,如果指令更简单,就需要更多条;如果指令更复杂,需要的条数就更少,但复杂的指令跑起来可能更慢。所以问题最后就归结为:这个比值到底是多少?嗯,
便签引用
3:18
and in the in the beginning in the 1980s when these debates were happening, there was this kind of very viciferous debates partly of like what those ratios are going to do, but a big part was kind of philosophical. Weren't weren't you doing damage to the software industry by lowering the instruction set level? So the gap between the programming languages and the instruction set was larger. That was kind of the ferocity of the debates. Well, after the dust settled a few years later and we started getting numbers, it turned out you needed like maybe 30% 40% more simple instructions to execute in a program, but you could run them four or five times faster. So the net was, you know, or five times faster say. So 3 or 4x speed up was the potential for risk over CISK.
而在 80 年代初这些争论发生的时候,场面相当激烈,一部分是在争这些比值会是什么样,但很大一部分其实是理念之争。你们难道不是在降低指令集的层级,会不会损害软件产业?因为这样一来,编程语言和指令集之间的差距就更大了。当时的争论差不多就是这么激烈。等到几年后尘埃落定,我们开始拿到实际数据,结果发现:一个程序里大概需要多 30%、40% 的简单指令,但这些指令的执行速度能快上四到五倍。所以净结果就是,比如说快四五倍。所以 RISC 相对 CISC 的潜在提速大概是三到四倍。
便签引用
4:05
>> You mentioned that um I guess philosophical difference like that gap. Why why would that be controversial? >> I mean unfortunately I'd say computer architecture in the 1970s and even 1980s a lot of it was uh was kind of people designed this by using their intuition or using their gut. Their feelings was this is the right thing to do. And even the textbooks of the time were like cataloges. They here's a computer and just list all its features and here's another computer listing all its features. So it's pretty dissatisfying what was going on because these debates it would seem like we ought to be able to sign this, you know, scientifically with with numbers that are there. But absence of that, it was more like a philosophical debate like how many angels on the head of a pin, right? You could you could have arguments qualitatively but you couldn't settle the arguments and because there wasn't ways to settle the arguments quantitatively you know people would just argue about it. Would you say that
>> 你提到了那种,我想可以说是哲学层面的分歧,就是那个差距。这为什么会引起争议呢?>> 说来遗憾,我觉得在 1970 年代、甚至 1980 年代,计算机体系结构很大程度上是靠人凭直觉、凭感觉来设计的。他们觉得这么做才对。就连当时的教科书也像是产品目录:这是一台计算机,把它的特性一条条列出来;这是另一台计算机,再把它的特性列一遍。所以那种状况相当让人不满,因为这些争论看上去本应该可以用科学的方式、用实实在在的数字来定夺。但在没有数字的情况下,它更像是一场哲学辩论,就像在争论一个针尖上能站多少天使,对吧?你可以做定性的论证,但你没法定出胜负;而正因为没有办法用定量的方式来终结争论,大家就只能一直吵下去。你会不会说,RISC 和 CISC 之争最后有一方赢了?是的。我最近正好看到
便签引用
02x86 锁定与 ARM 逆袭
5:08
you know risk versus cisk did one side win the war? Yeah. I was just reading there's some guy online who uh revisited it and he said and CISK won. Well, that's a pretty myopic view is uh in the PC era because of the importance of u distributing software in binary. So once the x86 was established and people the PCs would would ship software in binaries that was very hard to overcome. That was a huge uh impediment to changing the instruction set. So PCs are design defined by the x86 architecture largely.
网上有个人重新回顾了这段历史,他说 CISC 赢了。嗯,那是一种相当短视的看法——在 PC 时代确实如此,因为以二进制形式分发软件太重要了。所以一旦x86 站稳了脚跟,PC 上的软件都以二进制形式发布,那就很难被撼动了。这对更换指令集是个巨大的阻碍。所以 PC 的设计在很大程度上是由 x86 架构定义的。
便签引用
5:46
But about also in the 1980s there's this company in England u that wanted to do a personal computer they call it was the acorn personal computer and they decided they wanted to have their own they needed their own instruction set architecture to do that their own chip they were unsatis the chips they had available at the time they were doing it weren't fast enough and they were influenced by the papers that we did at Berkeley and so they built what they called the acorn risk machine and then uh one of the benefits of this kind of reduced instruction set is that it could be simpler, it would take less resources and take less energy to execute. And then several years later when Apple was looking for a microprocessor that could power, you know, one of their personal devices that this one was called the early forerunner of the iPhone called the Newton. They came to this company and said, "Boy, I really like it. Uh, let's get rid of the Acorn name." So they renamed it the acorn wrist machine arm and they
不过同样在 1980 年代,英国有一家公司想做一台个人计算机,他们把它叫做Acorn 个人计算机。他们决定要有自己的东西——为此他们需要自己的指令集架构、自己的芯片,他们对当时能买到的芯片不满意,那些芯片不够快,而且他们受到了我们在伯克利做的那些论文的影响,于是他们造出了他们所谓的 Acorn RISC Machine(Acorn 精简指令集机器)。而这种精简指令集的好处之一,就是它可以更简单,执行时占用更少的资源、消耗更少的能量。几年之后,苹果在找一款微处理器,想用来驱动他们的某款个人设备,也就是 iPhone 的早期先驱——Newton。他们找到这家公司说:“哎,我真的很喜欢它。呃,我们把 Acorn 这个名字去掉吧。” 于是他们把 Acorn RISC Machine 改名为 ARM,重新命名为 Advanced RISC Machine(先进精简指令集
便签引用
6:46
recristened it the advanced wrist machine uh because to get rid of the acorn and then Apple used it um in the Newton. Now the new wasn't a commercial success but it demonstrated the benefits for risk architecture for mobile devices. So the Nokia came along just a few years later with their GDM uh cellular phone which is one of the first popular ones and they embraced ARM. And so ever since you know uh ARM has dominated all the mobile devices. So I think I just checked there's been 350 billion uh ARM processors microprocessors with ARM technology in it today. So it's like today 99% of all processors in computers are risk even in personal computers. Apple switched over to ARM from from the x86 architecture. So even PCs there's um there's a risk architecture is you know is significant and it's trying to starting to get in the cloud. The cloud's largely been defined by the x86 server architectures but um Amazon and I think Amazon, Microsoft and Google all have their own development of ARM processors. So ARM is becoming or risk
机器),就是为了摆脱 Acorn 这个名字。然后苹果把它用在了 Newton 上。Newton 在商业上并不算成功,但它证明了 RISC 架构在移动设备上的好处。所以仅仅几年后,诺基亚带着他们的 GSM 手机出现了——那是最早流行起来的手机之一——他们采用了 ARM。从那以后,ARM 就统治了所有移动设备。我刚查了一下,到今天已经有 3500 亿颗采用 ARM 技术的处理器、微处理器面世。所以现在,计算机里 99% 的处理器都是 RISC,个人计算机也是如此。苹果已经从 x86 架构转向了 ARM。所以哪怕在 PC 上,RISC 架构的分量也很可观了,而且它正开始进入云端。云在很大程度上一直是由 x86服务器架构定义的,但亚马逊——我想亚马逊、微软和谷歌——都在自研 ARM处理器。所以 ARM,或者说 RISC 处理器,在云端也越来越普及。所以现在我会说,x86
便签引用
7:59
processor becoming more popular in the cloud. So right now you I'd say the x86 architecture market is shrinking while the risk architecture market is um you know growing leaps and bounds. >> Well, how how could someone say CISK one when you mentioned like 99% >> Yeah. Well, that's just if you if if they're doing something in history and they define computers as personal computers and maybe servers and they go up through around 2000. If the story ended then you could well it looks like a CISK one but you know uh once we get into this uh postPC era I don't understand how somebody would reach that conclusion. [laughter] >> You mentioned the the energy expenditure so like um you know risk makes sense in places or maybe on mobile devices things like that.
架构的市场在萎缩,而 RISC 架构的市场在突飞猛进地增长。>> 那既然你说 99%,怎么还会有人说 CISC 赢了呢? >> 是啊。嗯,那无非是如果他们在写历史,把计算机定义成个人计算机、可能再加上服务器,然后叙述截止到 2000 年前后。如果故事到那儿就结束,那看上去确实像是 CISC 赢了;但一旦进入后 PC 时代,我实在不明白怎么会有人得出那种结论。[笑] >> 你提到了能耗,就是说 RISC 在某些场合、比如移动设备上更合理之类的。
便签引用
8:52
>> Well even in the cloud I mean everybody cares about we're we're everything is energy bound now. So it if u you know these days of course you're not dealing with millions of transistors but billions of transistors. So it matters somewhat less today. There's so many things going going on that you can hide that. I mean even what the x86 architecture did is to compete is that it translated the x86 instructions in hardware into risk instructions. And so you had to pay that extra overhead of that translation step to get risk instructions and then you could ex any good ideas that the risk people had that the x86 could do. But it was worth it financially for that extra overhead because of the value of the PC software base. So it made a lot of sense for Intel to do that and they did that in the early 2000s. It you know it was a it's a great uh commercial idea. Is there any niche use case where CISK makes sense? Is there any, you know, is this an engineering trade-off or CISK is objectively worse?
>> 其实云端也一样。我是说,大家现在都很在意能耗,如今一切都受能耗限制。所以,当然,今天我们面对的不是几百万个晶体管,而是几十亿个晶体管。所以这一点如今相对没那么要命了,同时进行的事情太多,可以把它掩盖过去。我是说,x86 架构为了竞争所做的,就是用硬件把 x86 指令翻译成 RISC 指令。这样你就得为这个翻译步骤付出额外的开销,才能拿到 RISC 指令,然后 RISC 阵营的那些好点子,x86 也就都能用上了。但从财务上看,这份额外开销是值得的,因为 PC 软件生态的价值太大了。所以英特尔这么做非常合理,他们在 2000 年代初就是这么干的。这,你知道,这是个很棒的商业主意。那有没有什么小众场景是 CISC 更合适的?还是说,这算一种工程上的取舍,还是 CISC 客观上就是更差?
便签引用
03微程序解释器的代价
9:57
>> Yeah. So, if you want to, if we want to go kind of one level deeper into all of this, um, the what was actually going on in, you know, in designing computers, the hard part is the control. And so, what happened is in the beginning, it was kind of control was kind of ad hoc. you would figure out uh you put the gates together to make it to work. Uh a uh one of the computing pioneers, Maurice Welks, figured out a more elegant way to design control. And he said, "Well, we could just list all the control signals as the output of a memory and we could have something would keep track of where we were in the memory and and to issue those control signals." And he called this he called the this effort the instruction the the control signals you could think of instructions. So he called that a micro instruction and he called the programming of those instructions microprogramming.
>> 是这样。如果你想,如果我们想把这件事再往深挖一层,那么在设计计算机的过程中,真正难的部分是控制逻辑。所以一开始,控制逻辑基本上是拼凑出来的:你自己琢磨怎么把门电路连起来让它工作。呃,计算领域的先驱之一 Maurice Wilkes 想出了一种更优雅的方式来设计控制逻辑。他说:“我们可以把所有控制信号列出来,作为一块存储器的输出,然后再用某个东西来记录我们在存储器中走到哪一步,并发出这些控制信号。” 他把这些控制信号——你可以把它们看作指令——称为微指令,把对这些指令的编程称为微程序设计。
便签引用
10:53
So where technology was in the 1960s that made a fair amount of sense and so IBM built these so-called micro program computers. So it was basically an interpreter with very simple instructions that would interpret this much more sophisticated instruction set above it. But you would play this interpretation overhead. And classically you know in computer science interpreting versus compiling uh is something like a factor of five or 10. But given you know the latencies of the memory technologies and the possibility of doing this out of readonly memory it made sense u up until the 60s and 70s and then but then the kind of the question came up as came around 1980 is like is this still a good idea? Should we have this microode interpreter inside there? And so the alternative is to think of well rather than uh we've got this microode interpreter in there why don't we just compile directly into those instructions and that's pretty close to the to the risk ideas. Uh the micro instructions themselves used to be
以 1960 年代的技术条件来说,这相当合理,所以 IBM 造出了所谓的微程序计算机。它本质上就是一个解释器,用非常简单的指令去解释上层那套复杂得多的指令集。但你得付出这份解释的开销。而在计算机科学里,解释相对于编译,经典的开销大概是五倍到十倍。但考虑到当时存储技术的延迟,以及可以用只读存储器来实现这件事,这在六七十年代之前都是说得通的。可是到了 1980 年前后,问题就来了:这还是个好主意吗?我们真的需要在里面放一个微码解释器吗?于是另一条路就是这么想:与其在里面放个微码解释器,我们为什么不直接把程序编译成那些指令呢?这就非常接近 RISC 的思路了。呃,微指令本身过去
便签引用
11:58
you know like a 100 bits wide and really complicated. So if you if you make them not quite so long you make them kind of more natural still we can we could skip the interpretation step. So with the you know going going this level deeper kind of the question would be today would people invent an instruction set that was so sophisticated and needed a microode interpreter and probably they wouldn't do that. I mean you could do it you nothing prevents you from doing it but you wouldn't want to design an instruction set that was forced to use a micro card interpreter. you it might make sense in some very tiny applications possibly where you have you know thousand tens of thousands of transistors maybe this would a microcode interpret work but I think nobody today I don't think anybody's invented an instruction set in the last 20 years that u has anything that would need a microcoded interpreter >> you mentioned the compiler multiple times here and it seems like that's a critical piece that kind of makes risk
往往有一百来位宽,而且非常复杂。所以如果你把它做得没那么长,做得更自然一些,我们就可以跳过解释这一步。所以往深一层看,今天的问题会变成:现在还会有人发明一套那么复杂、还需要微码解释器的指令集吗?多半不会。我是说你也可以这么做,没什么拦着你,但你不会想去设计一套非用微码解释器不可的指令集。也许在极小的应用场景里还说得通——比如你只有一两万个晶体管,那时候微码解释器可能还有用——但我觉得今天没人这么干了。我认为过去二十年里没人发明过任何需要微码解释器的指令集。>> 你在这里多次提到编译器,看起来它是让 RISC
便签引用
04编译器与寄存器分配
12:59
work so well can you explain the role of the compiler and how it manages the relationship between software and hardware. >> Well, you you're writing in a programming language like C or C++ or maybe Python and um but the quality of the code that gets generated is is up to the compiler if you and a specific example is that it's useful when you construct computers to have registers and those registers are actually kind of visible in in the assembly language or the machine language programming. It can be 8, 16 or 32 these registers for the people programming that little level to use. Well, it used to be very difficult for compilers to to um allocate registers efficiently when they computers just weren't fast enough. We didn't have the algorithms that we could look at a section of code or a sub routine or something and say how can we most efficiently do registers? In fact, the C programming language which was invented to do systems programming in C before that to write an operating system people wrote them an assembly
如此奏效的关键一环。你能讲讲编译器的角色,以及它是怎么处理软件和硬件之间的关系的吗?>> 嗯,你用 C 或 C++、也许是 Python 这样的编程语言写代码,但最终生成的代码质量取决于编译器。举个具体的例子:造计算机时,设置寄存器很有用,而这些寄存器在汇编语言或机器语言编程里是可见的。这些寄存器可以是 8 个、16 个或 32 个,供在那个底层写程序的人使用。过去,编译器要高效地分配寄存器非常困难,那时候计算机还不够快,我们也没有那样的算法——能看一段代码、一个子程序之类的,然后判断怎样最高效地使用寄存器。事实上,C 语言是为了做系统编程而发明的,在有 C 之前,人们写操作系统用的是汇编语言,信不信由你。但
便签引用
14:06
language believe it or not and but the Unix people showed uh Ken Thompson Dennis Richie showed that we could if we had a low-level language pretty low-level language we get the benefits of writing in something that it's much easier for humans to understand and debug because register allocation by the compilers was was so poor, they had to put in the ability for the programmer to give hints of what variables should go in registers. It was just easier for them to have the programmer step in and say, if the machine had eight registers, I want these six variables to be in these registers. Don't leave them in memory because it runs so much slower.
Unix 那群人——Ken Thompson、Dennis Ritchie——证明了:如果有一门低级语言、相当低级的语言,我们就能享受到用一种人类更容易理解和调试的语言来写程序的好处。而因为当时编译器的寄存器分配实在太差,他们不得不加入一种机制,让程序员给出提示:哪些变量应该放进寄存器。对他们来说,让程序员出面说“这台机器有八个寄存器,我要把这六个变量放进这些寄存器里”反而更省事。别把它们留在内存里,那样跑起来慢太多了。
便签引用
14:43
Registers are so much faster. So, a big part of u the the CISK the risk argument was the compiler algorithms are getting better. they could handle these low-level instructions, they could allocate registers efficiently. Uh, and that was, you know, another reason in these debates about why it made sense to have a simpler architecture. What we did in the risk architectures, well, if registers are really important, uh, register allocation really important. One way to make it easier for the compiler was just to have a lot more of them. So, typically the CISK architectures of times would have eight or maybe 16 registers. So, we we put in 32 and That was one of the arguments, right, is well, if it's hard to efficiently use a small number, let's give them plenty.
寄存器要快得多。所以 CISC 与 RISC 之争中很重要的一点是:编译器的算法在变好,它们能处理这些底层指令,能高效地分配寄存器。呃,这也是那些辩论中支持采用更简单架构的另一个理由。我们在 RISC架构里做的是:既然寄存器这么重要、寄存器分配这么重要,那让编译器更省事的一个办法,就是干脆多给一些。当时的 CISC 架构通常有八个、最多十六个寄存器。所以我们放了 32 个。这也是当时的论点之一,对吧:既然少量寄存器很难用好,那就多给一点。
便签引用
15:29
So, even if it wasn't that good, there'll be enough registers. So, most of the time and and registers weren't that much more expensive to include in machines because of MOR's law. I saw somewhere in in uh one of the talks that he had given that somehow the compiler it's it's more easy for it to optimize the code for risk but in CISK there were these complex bigger ones and it almost never used them. Is that a shortcoming of the compiler? So what happened if we go back in the compiler days there's this argument is that by having more sophisticated instructions this would raise the level of extraction. that'll be a smaller gap. They'll make it easy for the compiler. But that was a philosophical argument. It it wasn't something that necessarily compiler people um could work. Compiler people weren't making that argument. It was the architects were making that argument.
这样即便分配算法没那么好,寄存器也够用。所以大多数情况下——而且因为摩尔定律,多放些寄存器进机器里并不会贵多少。我在他做过的某个演讲里看到过一种说法:编译器为 RISC 优化代码要容易得多,而在 CISC 里有那些复杂的大指令,编译器却几乎从来不用它们。这算是编译器的短板吗?回到编译器那个年代,当时有个论点是:指令越复杂,抽象层次就越高,差距就越小,编译器就越好写。但那是个哲学论证,未必是编译器那边的人真能用上的东西。提这个论点的不是编译器的人,而是体系结构的人。
便签引用
16:20
And as it turned out when we looked at the programs when we were doing the early research on CISK in versus risk is the compilers didn't really use those instructions. Often the compiler writers, the architects would come up with a sophisticated instruction and the compiler writers would say, "Well, we don't need that." In fact, we found examples where like for a procedure entry, there'd be a special instruction built in to do all this work that what the architects thought the compilers wanted and the compiler designers say, "Well, we don't need that. It's it's faster. It's actually faster for to use it with separate instructions than use your sophisticated instruction." And so, we found a bunch of examples like that.
而结果表明,我们在做 CISC 与 RISC 的早期研究、去看那些程序时发现,编译器根本不怎么用那些指令。常常是体系结构的人想出一条复杂的指令,编译器的人却说:“这个我们用不上。” 事实上我们发现过这样的例子:比如为过程调用的入口专门内置一条指令,把体系结构的人以为编译器想要的活儿全包了,可编译器的设计者说:“这个我们用不上。用几条独立指令反而更快,比用你那条复杂指令还快。” 类似的例子我们找到了一大堆。
便签引用
16:58
So you pay the extra overhead of the microode interpreter to have these sophisticated instructions and then the compiler doesn't even use them. Right? So th this is kind of a nonsensical situation that that we're in. uh and you know this is kind of like you know why why are there startups or why are there scientific you know breaking points like that as we looked at all the technologies with Moore's law you know ideas like caches with where compilers were in using sophisticated instructions the ability to allocate register more efficiently it made sense to change the directions away from these microcoded instruction set architectures to these simpler architectures >> when we talk about instruction sets I mean we're kind assuming that we're talking about these general purpose computer are these CPUs and I know GPUs are kind of talked about a lot and maybe there's other forms of computing are there instruction sets for those types of machines as well.
所以你为了拥有这些复杂指令,付出了微码解释器的额外开销,结果编译器根本不用它们,对吧?所以我们所处的这种局面挺荒唐的。呃,这也就是为什么会有创业公司、为什么会出现那种科学上的突破点:当我们把摩尔定律带来的各种技术都看一遍,比如缓存这样的思路,再加上编译器并不使用复杂指令、寄存器分配又能做得更高效,那么把方向从这些微码化的指令集架构转向更简单的架构就说得通了。 >> 我们谈指令集的时候,多少是默认在谈这些通用计算机,也就是 CPU;我知道现在大家常谈 GPU,可能还有别的计算形态。那些类型的机器也有指令集吗?
便签引用
05GPU 与登纳德缩放终结
17:56
>> So what happened is around 2000 is when GPUs came along and these were what we called domain specific architectures. So a GPU is a graphics processing unit. It had one job. it it didn't need to do everything that a general purpose processor needs to. Didn't have to support virtual memory. Uh it didn't have to support a compiler which is you know a pretty radical idea for architectures. It was just for graphics. And so around 2000 is when Nvidia and the other GPUs out there because they were trying to do you know the give you the graphics for games and give you the graphics for movies. So it was a niche product.
>> 情况是这样:大约在 2000 年前后 GPU 出现了,这些属于我们所说的领域专用架构。所以GPU 就是图形处理单元。它只干一件事。它不需要像通用处理器那样什么都做。不需要支持虚拟内存。呃,它也不需要支持编译器——这对架构来说算是个相当激进的想法。它就只是为图形而生的。所以大概在 2000 年前后,英伟达和当时其他做 GPU 的公司,因为他们想做的是给游戏提供图形、给电影提供图形。所以它当时是个小众产品。
便签引用
18:35
So what happened kind of in the computer industry is it was driven by Moore's law which I've mentioned many times and also this lesserk known law called dinard scaling which is because kind of an interesting question well if Moore's law puts a lot more transistors on a chip how come they're not getting hotter and hotter as we double the number of transistors every year and the reason was this observation made by Bob Dinard that what as you added more transistors People would also lower the threshold voltage which the distinction between a zero and one and that had like a squared effect. So you would end up doubling them transistors but you'd lower the threshold voltage. So microprocessors stayed at like 20 or 30 watts. In fact uh we did a John and I eventually did a textbook and I was just checking and that came out in 1990 and we didn't even talk about power as an issue for the first three editions. You know, the the third one came out in 2000. Power wasn't even a topic because micro dinard
计算机行业后来发生的事,是由摩尔定律驱动的——这个我已经提过很多次了——同时还有一个名气小一些的定律,叫登纳德缩放定律。因为这里有个挺有意思的问题:如果摩尔定律让芯片上的晶体管越来越多,那为什么我们每年把晶体管数量翻一番,芯片却没有变得越来越烫呢?原因就在于鲍勃·登纳德的这个观察:当你增加晶体管的时候,人们同时也会降低阈值电压——就是区分 0 和 1 的那个电压——而这带来的是平方级的效果。所以你把晶体管翻了一倍,但同时降低了阈值电压。于是微处理器一直保持在二三十瓦的水平。事实上,呃,我和约翰后来写了一本教材,我刚还查了一下,那本书是 1990 年出的,前三版我们甚至根本没把功耗当成一个问题来讲。你知道,第三版是 2000 年出的。功耗压根不是个话题,因为登纳德缩放还在起作用,所以微处理器就算越来越快,
便签引用
19:37
scaling was going on and so microprocessors stayed in the tens of watts even as they got faster and faster. What happened is about 2005 or so, Dinard scaling stopped working and that was a shock, right? So, and Intel actually had a microprocessor that failed one of their generations because they just couldn't get the power down enough. it was too hot to be able to do that. Um and then so then what that forced us to go to multicore is that uh before that it was easiest for the programmer for everybody if there was one very sophisticated processor that did everything but we couldn't do that anymore. So it went from one sophisticated processor to two and then four and then eight simpler processors.
功耗也还停在几十瓦。后来大概在 2005 年前后,登纳德缩放不管用了,那可真是个冲击,对吧?而且英特尔实际上有一代微处理器就是这么失败的,因为他们怎么都没法把功耗压下来。太烫了,根本做不到。呃,然后这就逼着我们走向多核。在那之前,对程序员、对所有人来说,最省事的做法是有一颗非常复杂精密的处理器把所有事都干了,但我们做不到了。所以就从一颗复杂处理器变成两颗、然后四颗、再到八颗更简单的处理器。
便签引用
20:23
Then it was up to the programmer to deliver on the potential of Morris law by paralyzing their code. So we stayed that way for about another 10 years and then Moore's law started slowing down. Uh so that the general purpose microprocessor was barely improving. It got a little better but it wasn't getting dramatically better. Um, in the, you know, in the 1980s, 1990s, 2000s, you you have a laptop and your friend's laptop would be like four times as fast as yours. And, you know, you were jealous. So, you would throw away perfectly good hardware because your friend's thing was so much faster because of the rapid change in performance. Well, that all ended in the 2010s where the you never you wouldn't throw away a laptop because the new ones were hardly that much faster than the old ones. You throw them away when they break and slow down. Nevertheless, programmers were used to this dramatic improvement in performance every few years because they could add more features to their software and stuff
接下来就得靠程序员把代码并行化,才能兑现摩尔定律的潜力。我们就这么又过了大约十年,然后摩尔定律开始放缓了。呃,于是通用微处理器几乎不怎么进步了。是有一点点变好,但没有变得天翻地覆。呃,在八十年代、九十年代、两千年代,你有一台笔记本,你朋友的笔记本可能比你的快四倍,你就会很嫉妒。所以你会把完全没坏的硬件扔掉,就因为你朋友那台快太多了,性能变化太快了。而这一切在 2010 年代就结束了——你不会再因为新机器而扔掉笔记本,因为新的比旧的快不了多少。你只会在它坏了、变慢了的时候才换掉。但即便如此,程序员已经习惯了每隔几年性能就大幅提升,因为这样他们才能往软件里加更多功能
便签引用
21:25
like that. So, what were architects going to do? They'd already done the the multicore trick in 2005 or so. So around 2015 the idea was we would do domain specific architectures like the GPUs. So we if we if you tell me I only have to run a narrow class of programs and I don't have to necessarily run all the operating systems everything else. Well yes I could shuffle those resources and do something much more efficient for something well. So you could do some things well and there's other things either you do poorly or not even at all.
之类的。那架构师该怎么办呢?2005 年前后他们已经用过多核这一招了。所以大概在2015 年,思路变成了做领域专用架构,就像 GPU 那样。也就是说,如果你告诉我我只需要跑一小类程序,不需要跑所有操作系统和别的一切,那么是的,我可以重新调配这些资源,把某件事做得高效得多。所以你可以把一些事做得非常好,而另一些事你要么做得很差,要么干脆不做。
便签引用
06TPU:为机器学习定制
21:59
So then the question is okay what domain? So just coincidentally where we were technologically right around 2012 2015 is when machine learning AI burst on the scene. Um and so that was it was evident what domain should we do and the domain was machine learning AI. So just coincidentally where we were technologically uh this new domain came along and then I I would say you know I I work for Google but I think it's fair to say Google was the one who really uh saw it was the certainly the first big company who understood the potential of machine learning AI and um and bet that they or feared that they were going to be swamped with demand and needed to do custom hardware. So the the translate tensor processing unit GPU that Google debuted in 2016 really kind of shocked the world and got people to realize that we should be uh designing hardware for machine learning and we could continue to improve performance dramatically for that one domain.
那么问题就来了:选哪个领域?巧的是,就在我们技术上走到这一步的 2012 到2015 年,机器学习和人工智能爆发了。呃,所以该做哪个领域就很明显了,那个领域就是机器学习和 AI。就这么巧,我们在技术上正好走到那一步,这个新领域就来了。然后我想说——你知道,我是给谷歌工作的,但我觉得公允地说,谷歌确实是那个真正看到这一点的公司,至少是第一家理解机器学习和 AI 潜力的大公司,而且他们赌定、或者说担心自己会被需求淹没,因此需要做定制硬件。所以谷歌在 2016 年推出的张量处理单元(TPU)确实震动了整个业界,让大家意识到我们应该为机器学习设计硬件,而且在这一个领域里我们可以继续大幅提升性能。
便签引用
23:08
>> What's the high level differences between a CPU, a GPU and a TPU? the the CPU has got to be this general purpose thing. And basically um even today they have these general purpose cores that each core is very similar to what the instruction sets or used to be like 20 years before that. It's not a surprising design and it's just got lots of cores and modern ones today could have a h you know 50 or 100 cores in there. So that's that's the feature for graphics. uh what they decided to do which to do the graphics job is they need to get a lot of performance on the memory system. So they went to multi-threading. So they have hardware threads that you would send a memory request out and you the hardware would switch over to do something else while the when memory is coming out. So it's this highly threaded architecture and that's what they developed and could could run well for graphics.
>> CPU、GPU 和 TPU 在宏观层面上的区别是什么?CPU 必须是那种通用的东西。基本上,呃,就算是今天,它们还是有这些通用核心,每个核心和二十年前的指令集其实很相似。设计上并不让人意外,只是核心特别多,今天的现代 CPU 可能有五十个甚至一百个核心。这是它的特点。呃,而图形这边,他们为了完成图形任务所做的决定,是需要在存储系统上获得很高的性能。所以他们走向了多线程。他们有硬件线程,你发出一个访存请求,在内存数据回来之前,硬件会切换过去干别的事。所以这是一种高度多线程的架构,这就是他们开发出来的东西,对图形来说跑得很好。
便签引用
24:04
um a and graphics. It turns out um doesn't need very s doesn't need powerful floating point. It needs doesn't need wide floating point. So 32-bit floating point is plenty for graphics even 16 bit. So they were doing pushing graphics with 16 and 32-bit floating point in this multi-threaded kind of architecture. And because it was kind of its own unique thing, it has its own own set of terminology um all to itself. it wasn't kind of out of the main branch of computer architecture. Uh so when you look at our textbooks, we have kind of like a Rosetta Stone which says here's the terms that Nvidia uses to describe DP2s.
呃,还有图形这块,事实证明它不需要很强的浮点,不需要很宽的浮点。所以32 位浮点对图形来说绰绰有余,甚至 16 位都够。所以他们就在这种多线程架构上,用 16 位和 32 位浮点来推动图形。而且因为它算是自成一体的东西,它有一整套自己专属的术语,不太属于计算机体系结构的主干分支。呃,所以你看我们的教材,里面有点像一块罗塞塔石碑,写着:这是英伟达用来描述 GPU 的术语,
便签引用
24:45
This is what it means in kind of normal um well normal in the mainstream processor design. So they were this niche product that happened to do pretty fast um single precision floating point and uh even half precision floating point and they were pretty cheap. They were like hundreds of dollars. Um so when people uh so kind of uh but some people were thinking boy for some applications if I could turn my program into like uh pretend that it was doing creating an image if I could turn my problem into image generation I could use these pretty cheap GPUs which had very good floatingoint performance per dollar compared to anything else. And so people started playing around with that.
这在正常的——呃,主流处理器设计里对应的意思是什么。所以它们当时就是这么个小众产品,恰好能做相当快的单精度浮点,呃,甚至半精度浮点,而且相当便宜。也就几百美元。呃,所以当人们——呃,有些人就开始想:天啊,对某些应用来说,如果我能把我的程序变成像是在生成图像那样,如果我能把我的问题转化成图像生成,我就可以用这些相当便宜的 GPU——它们每美元的浮点性能比其他任何东西都好得多。于是人们就开始鼓捣这个了。
便签引用
25:33
Uh and uh the founder CEO Jensen Wang really liked that idea. So in 2006, he funded an effort to create a programming language that that would kind of handle this multi-threaded hardware architecture that's for graphics to make it easier to program. And so that led to the invention of uh CUDA uh which I can't remember its acronym but it's really a a proprietary programming language for the multi-threaded GPU architecture to make it kind of easier to program. It was a it was a C like language but you can't just compile C programs and run it but it was C like but people actually liked it. You know it was certainly much better than trying to turn your program into an image generation but people liked it. But uh Jensen Wang his vision was try to get these teenagers in the basement who are playing the games to learn how to program these things. And he had a couple of markets that he was interested in like um you know uh some of the department of energy labs like uh fluid flow and some of things like you
呃,创始人兼 CEO 黄仁勋非常喜欢这个想法。所以在 2006 年,他出资启动了一个项目,要做一门编程语言,能驾驭这套为图形设计的多线程硬件架构,让它更容易编程。于是就有了 CUDA 的诞生。呃,我记不清它的缩写全称了,但它确实是一门针对多线程 GPU 架构的专有编程语言,目的是让编程更容易些。它是一门类 C 语言,但你不能直接拿 C 程序编译了就跑;它只是像 C,不过人们确实挺喜欢它的。你知道,这肯定比想办法把你的程序变成图像生成要好多了,所以大家喜欢它。但是呃,黄仁勋的愿景是想让那些窝在地下室打游戏的青少年学会给这些东西编程。而且他还盯上了几个市场,比如呃,一些能源部的实验室,比如呃流体流动之类的一些东西。他会去做一些专用的库来处理那个
便签引用
26:37
know some and he would build special purpose libraries that could handle that domain and then use his GPUs to do that kind of uh special purpose computing but breaking out of graphics. But that's kind of the heritage there. Now for the TPU program um and and so what's happened uh people started right from the very beginning people started using GPUs because they had much better single precision floatingoint performance than um than the CPUs. They had a lot more processors on them. So we had the potentially much more performance. So they're kind of one of the breakthrough moments moments in it was in 2012 when uh within the machine learning community this neural networking piece which had a few advocates but a lot of people didn't believe in it they went into a competition to see who could do the best image recognition in 2012 the so-called AlexNet beat all the competition this was this uh historic moment in the in the in machine learning learning in neural networking and once they did it
领域,然后用他的 GPU 去做那类专用计算,从图形里突破出来。这大概就是它的来历。那么再说 TPU 项目,呃——后来发生的是,人们从一开始就用起了GPU,因为它的单精度浮点性能比 CPU 好得多。上面的处理器也多得多。所以潜在性能高得多。其中一个突破性时刻是在 2012 年,呃,在机器学习圈子里,神经网络这一块只有少数几个拥护者,很多人并不相信它。他们参加了一场比赛,看谁能在 2012 年做出最好的图像识别,结果所谓的 AlexNet 击败了所有对手。这是机器学习和神经网络领域一个历史性的时刻。做出这个东西的人在多伦多大学上过一门
便签引用
27:42
and the guy who did that had taken a cuda course at the University of Toronto to learned how to use it and so he says well as long as I'm doing this I'll do it on a GPU so he was able to explore a lot more space on this cheap fast GPU so he entered it in the competition 2012 and within he was the only one using neural networks and he crushed the competition and within a few years everybody switched over and not only they were doing neural networks but they're using GPUs So GPUs have this um heritage of being a graphics in graphics engine but more programmable and uh it started getting used for machine learning. When Google came along uh they decided you know it was a clean slate. They didn't they didn't care about graphics. So the at the heart of neural networking is a matrix multiply. That's the that's the thing. So they designed a a processor for at the time microp processor at the time had a giant matrix multiplying unit. That was the main thing and then they threw a bunch of stuff out that
CUDA 课,学会了怎么用它,所以他说:反正我要做这个,那就在 GPU 上做吧。于是他能在这块又便宜又快的 GPU 上探索更大的空间,然后他就把它报名参加了 2012 年那场比赛,他是唯一一个用神经网络的,结果他碾压了所有对手。没过几年,所有人都转过去了,不光开始做神经网络,而且都在用 GPU。所以 GPU 的血统是图形引擎,只不过更可编程,呃,后来就开始被用于机器学习。等谷歌进场的时候,呃,他们觉得这是一张白纸。他们不在乎图形。而神经网络的核心是矩阵乘法。这才是关键。所以他们设计了一款处理器——在当时叫微处理器——里面有一个巨大的矩阵乘法单元。这是最主要的东西,然后他们把一堆不需要的东西都扔掉了。比如很多通用
便签引用
28:43
they didn't need. So a lot of general purpose computing is what's on the chip is maybe three levels of caches to try and have so you don't spend all your time going to the relatively slow memory. Well, for machine learning, they knew when their memory accesses were so they could schedule that. So they didn't a a hardware cache didn't make any sense. they would just have a memory that the software understood would transfer in time. So those were some of the innovations. Also they innovated on the floatingoint format. You didn't need uh uh you know scientific computing cares a lot about precision. They have you know most of it's done in 64-bit floating point where the exponent is you know less than 10 bits and most of it is is the uh precision you know the fraction can be 50 some bits. Well, they what they realized for machine learning and AI, they don't need all that precision. They needed the range. So, Google did the first floatingoint format where the exponent was bigger than the fraction. That's that was a radical idea
计算的芯片上大概有三级缓存,为的是别把时间全花在访问相对很慢的内存上。而对机器学习来说,他们知道访存什么时候发生,所以可以提前调度。因此硬件缓存就没什么意义了。他们只要有一块软件清楚、能按时把数据搬进来的存储就行。这些就是当时的一些创新。他们还在浮点格式上做了创新。你并不需要——呃,你知道,科学计算非常在意精度。他们大多用 64 位浮点,指数部分不到 10 位,剩下大部分是精度,也就是尾数,能有五十几位。而他们意识到,对机器学习和 AI 来说,并不需要那么高的精度,需要的是动态范围。所以,谷歌做出了第一个指数位比尾数位还多的浮点格式。这在当时是个激进的想法,
便签引用
29:42
that so-called brain float 16, the part of Google that was doing this was the brain research group. So, they brought out this architecture. They had this narrow floating point format. They had big matrix multiply unit. It only had one one processor in it. uh unlike the other ones, it was just dedicated machine learning and it just kind of blew the doors off of everybody in the field. It was uh you know like 30 times better at inference than the contemporary GPU and 80 times better than CPU by having this dedicated one.
就是所谓的 bfloat16。谷歌内部做这件事的是 Brain 研究团队。所以他们拿出了这套架构:有窄浮点格式,有大矩阵乘法单元,里面只有一个处理器。呃,跟别的不一样,它就是专门为机器学习而生的,而且它把这个领域里所有人都远远甩开了。呃,在推理上它比同期的 GPU 快大概 30 倍,比 CPU 快 80 倍,就靠这么一个专用设计。
便签引用
30:15
So when uh Google made this announcement at their annual retreat a year or so after um after they had deployed it inside uh it just shook everybody up. Intel started buying companies. Um you know Nvidia started modifying the design uh for to be mach learning and then a bunch of other competitors uh started uh uh hyperscalers started doing their own effort. So I think that was the watersh I think the TPU announcement was the watershed moment there. >> So it's it's almost like uh levels of specialization like the CPU is the most general. Well yeah if we could still if you know if we still had dinard scaling if we you know Morris law and dinard scaling where we be today with general purpose processors we should have 100 terahertz you know microprocessors if we could build 100 terahertz microprocessors that's what we do we would the GPUs would still be a niche you know if we could do that it you know it raises all boats that would be fantastic if we could do that but that that's long in the history we haven't been able to do
所以当谷歌在内部部署一年左右之后,呃,在年度务虚会上宣布这件事时,呃,所有人都被震住了。英特尔开始收购公司。呃,英伟达开始改设计,转向机器学习,然后一堆别的竞争对手,呃,一堆超大规模云厂商也开始自己做。所以我认为那就是分水岭,我觉得 TPU 的发布就是那个分水岭时刻。>> 所以这几乎像是不同层级的专用化,CPU 是最通用的。嗯,是的。如果我们还能——你知道,如果登纳德缩放还在,如果摩尔定律和登纳德缩放都还在,今天的通用处理器会是什么样?我们应该有 100 太赫兹的微处理器了。如果我们能造出 100 太赫兹的微处理器,我们就会那么干,GPU 也还会是个小众产品。你知道,如果我们能做到,那真是水涨船高,那就太好了。但那已经是很久以前的历史了,我们已经二十年做不到
便签引用
31:23
that for 20 years and so you know there were probably two or three gigahertz microprocessors in 2005 and that's kind of where they are barely improved today so we can't do that. So there's still areas where it's important to have CPUs. They're the right, you know, you need operating systems, compilers and things. Uh that they're the right solution, but things have been specialized. And then now in terms of industry terms, huge emphasis is being put into the AI accelerators or just AI in general is where the money is being invested. And so the question for any processor designers was how does my processor design fit into this AI universe and um what should I do? How should I change it for this important application?
这一点了。你知道,2005 年前后差不多是两三吉赫兹的微处理器,而今天基本还是那个水平,几乎没怎么提升。所以我们做不到。所以还是有一些领域需要 CPU,它们才是对的选择。你知道,你需要操作系统、编译器这些东西。呃,在那些地方它们是正确的解法,但很多东西已经专用化了。那么从行业的角度看,现在巨大的投入都放在 AI 加速器上,或者说 AI 本身就是钱在往里砸的方向。所以对任何处理器设计者来说,问题就是:我的处理器设计如何嵌进这个 AI 世界?呃,我该怎么做?为了这个重要的应用,我该怎么改它?
便签引用
07摩尔定律放缓与封装创新
32:12
>> So you mentioned that kind of Moore's laws slowing down and I watched this other talk from Jim Keller saying that Moore's law isn't isn't dead. But is it controversial to say that it's slow? Oh, I mean not not to real engineers. The the Morris law is very simple. It says in a in a chip the number of transistors will double originally said every year and then he mended it uh to every two years. Just just look do the chips are the number of transistors doubled and no now I think what people assume that means if trans if we are no longer on Morris law that technology is not improving. That's not the same thing.
>> 你提到摩尔定律在放缓,我看过吉姆·凯勒的另一个演讲,他说摩尔定律并没有死。那说它放缓了会有争议吗?哦,我是说,在真正的工程师那里没有争议。摩尔定律非常简单:一块芯片上的晶体管数量会翻倍——最初说的是每年,后来他修正为每两年。你就去看看,芯片是不是晶体管数量翻倍——现在我觉得,人们默认的意思是:如果我们不再处于摩尔定律之下,就等于技术不再进步了。这两件事不是一回事。
便签引用
32:54
it's not improving at the rate that Moore projected. And what was amazing about Moors law which lasted 50 years is that it it guided the investment of semigtory manufacturing. We need to deliver on doubling transistors every year or two. How are we going to do that? How are we going to build the equipment to do it? So it was a guideline for the whole industry which was kind of remarkable. So what we are doing now is there's pieces of the technology to get better and there's pieces that uh that don't improve at all. So one of the big pieces important pieces on a chip is the static RAM SRAMM. So that's hardly improving at all.
它只是没有按摩尔预测的速度在进步。摩尔定律了不起的地方在于,它持续了 50 年,它引导了半导体制造业的投资方向。我们必须做到每一两年把晶体管数量翻一番。我们要怎么做到这一点?我们要怎么造出实现它的设备?所以它是整个行业的一条指导方针,这相当了不起。而我们现在的情况是:技术里有些部分还在变好,有些部分则完全没有改善。芯片上一个很大、很重要的部分是静态存储器 SRAM,而它几乎没有任何改善。
便签引用
33:35
Logic gates still continue to improve. They are the gates are getting better. So if you're doing like uh adders or multipliers, those are getting better. those pieces. It's not uniform improvement anymore. And we're also going to more exotic packaging to be able to deliver it. Uh so people not uh it, you know, forever it was a single chip was the best way to package everything fits on one chip and that's what we're going to build. And now there's these ideas of chiplets or packaging multiple chips together. the latest GPUs and I think the latest TPUs has actually two what are called full reticle design dyes package put into a package and uh so a full reticle design is uh the maximum you can build on a on a semiconductor it they have a step and repeat motor and there's you know in the old days you'd have many chips inside that now there's just one so two maximum reticle designs form a node so they're using packaging so if you from outside you can say well there look there's look how many transistors are it's quote
逻辑门倒是还在继续进步。门电路确实在变得更好。所以如果你做的是加法器或乘法器之类的东西,那些是在变好的。就是那些部分。改善已经不再是均匀的了。而且我们还在转向更奇特的封装方式来兑现性能。所以人们不再——你知道,长久以来单芯片一直是最好的方案,所有东西装在一颗芯片上,我们就照这个思路造。而现在出现了小芯片(chiplet)、把多颗芯片封装在一起这样的思路。最新的 GPU,还有我想最新的 TPU,实际上是把两颗所谓的“满掩模版尺寸”(full reticle)的裸片封装到一个封装里。所谓满掩模版设计,就是在半导体上能做出的最大尺寸——他们有步进重复曝光机,你知道,在过去,那一格里面会有很多颗芯片,而现在只有一颗。所以两个最大掩模版尺寸的设计组成一个节点,他们是靠封装做到的。所以从外面看,你可以说,你看,看这里有多少晶体管,所谓的
便签引用
34:42
Moore's law is continuing but if you see what's inside the chip you know that's that has tapered off and I think part of it is for a fair amount of the industry if you were in a semiconductor manufacturer and people ask you what you do is I make Moors law I I make I sustain Morris law what I do and if you've been doing that for decades and somebody says Morris law is over it's like it my my career is So I think there's an emotional side of it but you know just look at the data the data doesn't back up what Keller says but I also didn't say that the technology is not improving the technology is continuing to improve and specifically domain specific architectures but if you look at the general purpose architectures or even other memory technologies you know it used to be the DRAMs would improve by a factor of four in density every 3 years like clockwork and now it's maybe 10 years between factors for so you can see plenty of evidence that Moore's law no longer applies.
“摩尔定律还在继续”。但如果你看芯片内部,就知道那已经放缓了。我觉得部分原因在于,对行业里相当一部分人来说,如果你在一家半导体制造企业,别人问你是做什么的,你说:我在“制造摩尔定律”,我在维系摩尔定律,这就是我的工作。如果你干了几十年这件事,突然有人说摩尔定律结束了,那感觉就像“我的职业生涯完了”。所以我觉得这里面有情感的一面。但你只要看数据,数据并不支持凯勒(Keller)的说法。不过我也没有说技术不再进步了——技术在继续进步,尤其是领域专用架构。但如果你看通用架构,甚至看其他存储技术,你知道,过去 DRAM 的密度每三年就像钟表一样提升四倍,而现在可能要 10年才提升四倍。所以你能看到大量证据表明,摩尔定律已经不再适用了。
便签引用
35:41
>> Is there um some new version or some analog to Mo's law that is kind of guiding uh you know microprocessor design and architecture these days? I I guess a question is uh both Nvidia and Google are continuing to deliver much faster processors for machine learning, tremendously better. What are they doing? Well, part of it is what I said about uh packaging, you know, to getting uh to to be able to get more transistors and you keep them closer together because distance matters. Uh part of it is innovating on the floatingpoint formats. So it you know unlike the supercomputers of 64-bit it's not even 64-bit not even 32-bit not even 16 bit but 8bit and 4bit floating point is as is going on so narrowing of the data types um the uh you know having kind of what's called the matrix multiply instructions so these powerful units that are put in there that can that can uh do special purpose applications that are very important get making those bigger and faster. So those are the things that are going on. But we
>> 那如今有没有某种新版本的、或者类似摩尔定律的东西,在指导微处理器的设计与体系结构?我想问的是,英伟达和谷歌都还在持续推出面向机器学习的、快得多的处理器,提升幅度非常大。他们到底做了什么?嗯,一部分就是我刚说的封装,你知道,为了能塞进更多晶体管,同时让它们靠得更近,因为距离是有影响的。另一部分是在浮点格式上做创新。所以,不像超级计算机用 64 位,这里不是 64 位,不是 32 位,甚至不是16 位,而是 8 位、4 位浮点数,正在往这个方向走——也就是把数据类型变窄。还有,你知道,加入所谓的矩阵乘法指令,也就是放进去的那些强大的运算单元,它们能处理一些非常重要的专用应用,把这些单元做得更大更快。这些就是正在发生的事情。但我们
便签引用
36:56
don't have this simplifying guideline kind of underlying all this. You have to be aware of each piece of the technology, gauge how fast it's improving, how whether it's they can deliver on what's going on and then you assemble that together and make your bets. We're, you know, there's a paper that we've just got approved that I think we're going to put on archive soon, which is, uh, talking about the Google TPU line and pretty remarkably, uh, the Google TPU line particularly for training has stayed pretty constant in the basic architecture design, uh, things have gotten bigger and faster, but the if you look at the the the design of it, going back to the first training TPU, it's that block diagram still works.
没有那种贯穿这一切的、化繁为简的指导方针了。你必须了解技术的每一个部分,判断它进步得有多快,判断它们能不能兑现,然后把这些拼起来,下你的赌注。我们有一篇刚刚获批的论文,我想很快会放到 arXiv 上,讲的是谷歌 TPU 这条产品线。相当值得注意的是,谷歌 TPU 这条线,尤其是训练用的,在基本架构设计上一直保持得相当稳定——东西变得更大更快了,但如果你看它的设计,一直回溯到第一代训练用 TPU,那张框图到现在依然成立。
便签引用
08基准测试与 CUDA 护城河
37:45
uh all these you know a decade later. So it you know the people who designed that in whatever 2015 or two did a really great job. Uh but the main components of a a big matrix multiply unit the so-called high bandwidth memory um uh a vector unit to go with the matrix unit and the basic building blocks are still there. I think in your career with the CPUs, benchmarks were a huge they played a huge role in in measuring which computer architectures were performant and not in this this new space of uh floatingoint operations for AI. Is there a benchmark that people use for GPUs and can you also apply it for TPUs?
你知道,都过了十年了。所以,当年在 2015 年左右设计它的那些人做得非常出色。主要的组成部分——一个大的矩阵乘法单元、所谓的高带宽内存(HBM),还有配合矩阵单元的向量单元——这些基本构件都还在。我想在你和 CPU 打交道的职业生涯里,基准测试起到了非常大的作用,用来衡量哪些计算机体系结构性能好、哪些不好。那么在面向 AI 的浮点运算这个新领域里,有没有大家用来衡量 GPU 的基准测试,而且它也能用在 TPU 上吗?
便签引用
38:27
>> I was you know I and some friends were helped involved in called the ML Perf effort. So it was inspired by the spec CPU effort was the spec benchmarks that there were these competing companies and they'd all make claims about theirs was better than the others and they realized that wasn't good for the industry. So they agreed on a set of benchmarks. So the ML Perf effort is uh that's run now but what's called ML comments is an attempt to do that is if we're going to do this comparisons let's um let's let's not argue about what the benchmarks are. So that's a serious effort in that kind of interestingly what's happened I'd say is because uh there's two pieces to u the design there's the the machine learning libraries that are uh to implement a lot of features that you need for for these applications and the libraries will even be rewritten for specific applications rather than just having general libraries. This has turned out to be a big advantage for Nvidia because they have a large corporation with lots of
>> 我和一些朋友参与过一个叫 MLPerf 的项目。它的灵感来自 SPECCPU,也就是 SPEC 基准测试。当时那些互相竞争的公司都会宣称自己的产品比别人好,后来他们意识到这对行业不好。于是大家商定了一套基准测试。所以MLPerf 现在是由所谓的 MLCommons 在运作,它想做的就是:如果我们要做这些比较,那我们就别再争论基准测试该是什么了。所以这是一项认真的努力。有意思的是,我觉得后来发生的情况是,因为设计其实有两块:一块是机器学习库,用来实现这些应用所需要的大量功能,而且这些库甚至会针对特定应用重写,而不只是提供通用库。这最后成了英伟达的一大优势,因为他们是一家大公司,有很多人手,有很多工程师能替他们做这些库。所以结果
便签引用
39:35
people that are available, lots of engineers who could build these libraries for them. So, it's turned out um um not quite a it's not a it's not a neutral evaluation, right? It's it's the architecture plus the the libraries that go with it and the compiler too, but specifically the libraries that you tailor each time. So Nvidia brings out when they announce the new architecture they create a new set of they modify the libraries to run that really well or to run applications really well. So that's a powerful combination kind of in business terms people referred to Nvidia's having this CUDA moat but and part of it is CUDA the programming language but a big part of it is the libraries that Nvidia Nvidia makes. So this has made it difficult for for startups to be able to compete with Nvidia partly because you know they didn't put enough emphasis in the software and partly because you know Nvidia just has many more engineers than they do to be able to tailor the libraries. So it's the libraries plus
就是,这并不完全算是一个中立的评测,对吧?它是架构加上配套的库,还有编译器,但尤其是那些每次都要专门定制的库。所以英伟达每次发布新架构时,都会做一套新的库,或者改造现有的库,让新架构跑得非常好,或者让应用跑得非常好。所以这是一个很强的组合。用商业上的说法,人们把这叫做英伟达的 CUDA 护城河。其中一部分是 CUDA 这门编程语言,但很大一部分其实是英伟达做的那些库。所以这让创业公司很难跟英伟达竞争,一方面是因为它们在软件上投入的重视不够,另一方面是因为英伟达的工程师比它们多得多,能去定制这些库。所以是库加上
便签引用
40:44
the architecture that's this this powerful advantage why most people do things on GPUs. Now Google has been able to develop their own libraries. Uh they don't have as many engineers. They use compilers more uh than I think Nvidia does. So they they have a set of libraries that they can do. But up until recently, Google has it's all been internal. It's just for Google to use or you could use it via the cloud. Um but uh you these startups have u you know have to have a difficulty as a result. That's why the ML Perf hasn't been as popular. Not everybody runs them because uh Nvidia runs them really well, really better than everybody else. And the startups have a hard time showing off what they can do uh given they don't have the the engineering effort going into the libraries that you need to do well in ML Perf. you gave this popular talk about how to have a bad career and it's kind of uh the negation of advice to kind of have a good career and I was just wondering if you could kind of summarize
架构,构成了这个强大的优势,这也是为什么大多数人都在 GPU 上干活。现在,谷歌算是能自己开发出这些库。他们的工程师没那么多,我觉得他们比英伟达更多地依赖编译器。所以他们有一套自己能用的库。但直到最近,谷歌的这些都还是内部的,只供谷歌自己用,或者你通过云来用。所以那些创业公司因此处境很难。这也是为什么MLPerf 没那么受欢迎——不是每家都去跑,因为英伟达跑得非常好,比谁都好。而创业公司很难展示自己的能力,毕竟它们没有足够的工程投入去做MLPerf 想跑好所需要的那些库。你做过一个很受欢迎的演讲,讲“如何拥有一个糟糕的职业生涯”,它算是把“如何拥有好的职业生涯”的建议反过来讲。我想问,你能不能总结一下
便签引用
09家庭、幸福与专注
41:53
maybe the the top three things that you kind of think people should take away. >> Yeah. So, yeah. Well, there's been a few things. So, the story is early in my career my f a good friend and I said, "How can we teach grad students how to give a good talk?" And we thought it'd be funny to explain how to give a bad talk and then if you didn't want to give a bad talk, this is so this is how to do it badly and if you don't here's the things to do not to do it badly. And so then I later did a how to have a bad career that's pretty much focused towards uh academia you know so for researchers to be able to do that. I later did a how to have a bad uh how to build a bad research center, how to build a bad research lab. And then recently I've been giving how to give AI a bad carbon footprint. That's my latest one. But I think the thing that might be more relevant is uh at the end of my talks I would I would kind of reflect on my career and talk about lessons learned. So I've written a paper called
你觉得大家最该带走的前三条?>> 好的。嗯,这里有几件事。事情的缘起是,我职业生涯早期,我和一位好朋友说:“我们怎么教研究生做一场好的演讲?”我们觉得反过来讲“怎么做一场糟糕的演讲”会很有趣,这样如果你不想做糟糕的演讲,那这就是把它搞砸的方法;如果你不想搞砸,这些就是你不该做的事。后来我又做了“如何拥有一个糟糕的职业生涯”,那个基本是面向学术界的,你知道,是给研究者的。再后来我做了“如何建一个糟糕的研究中心”“如何建一个糟糕的研究实验室”。然后最近我在讲“如何让 AI 拥有糟糕的碳足迹”,那是我最新的一个。不过我觉得可能更相关的是,在我演讲的结尾,我常会回顾自己的职业生涯,谈谈学到的教训。所以我写过一篇文章,叫
便签引用
42:48
uh lesson I think it's lessons life lessons from the first half century of my career. I wrote that recently. And so those those are divided up into kind of career advice and personal advice there. On the the personal side, I'd say uh if you have a family, you know, make sure your families first, keep your family first. Um the technology we've invented makes it really easy for your, you know, to take your work home with you and not pay attention to your family. Uh, I had a when I was first here at Berkeley, I gave a senior faculty member a ride home, dropped him off late at night coming up from Silicon Valley and he said, "Well, Dave, if I had to do it over all over again, I wish I'd spent more time with the family." And I never wanted to say that. And nobody on their deathbed says, you know, I wish I'd spent more time in the office. So, so you got to, you know, whenever it's easy to kind of, you know, let the let your career overcome the needs of your family, but you just you got to just
《我职业生涯前半个世纪的人生教训》,最近写的。里面分成了职业方面的建议和个人方面的建议。在个人方面,我想说,如果你有家庭,一定要把家庭放在第一位,始终第一位。我们发明的这些技术,让人非常容易把工作带回家,而忽略了家人。我刚到伯克利那会儿,有一次我开车送一位资深教员回家,深夜从硅谷回来把他放下,他说:“戴夫,如果让我重来一次,我希望自己多花点时间陪家人。”我从来不想说出这句话。而且没有人在临终时会说,我希望自己在办公室待得更久。所以你必须,你知道,每当那种让事业压过家庭需要的情况变得很容易发生时,你就得记住这一点。我还想说,在我的
便签引用
43:49
remember that. Uh, I would say in my life and kind of growing up in the 50s, you know, the idea was to be happy, you had to be wealthy. But those are actually two different goals, wealthy and happiness. So, I always made decisions towards happiness versus wealth. You come down with it. And I felt very good about that. Uh I and even now like why would I pick why would I pick something that made me wealthy and unhappy? Why would you do that? And so uh and in the field that we're in it turned out there was a lot of wealth to go with it. So you know optimizing happiness didn't mean you had to suffer uh beyond that. Um I think it's important to have fun. Personally, I think when you're a kid, you don't have to tell kids to play, but as an adult, you get so busy. You you don't think you have time to have fun. Uh, but you know, you only get to do this once. So, uh, so I right now I play soccer, I run my bicycle to the interview, uh, I lift weights, I body surf, I, you know, do things with my wife and family and my
人生里,我是 50 年代长大的,你知道,那时的观念是,要幸福就得富有。但这其实是两个不同的目标:富有和幸福。所以我一直是朝着幸福而不是财富去做选择的。我对此感觉非常好。哪怕到现在,我为什么要去选一个让我富有却不幸福的东西?为什么要那样做?而在我们这个领域,结果发现财富也随之而来。所以你知道,把幸福最大化并不意味着你就得在别的方面吃苦。还有,我觉得玩得开心很重要。就我个人来说,我觉得小时候,你不需要催孩子去玩,但成年之后,你会忙到,觉得自己没时间去玩。但你知道,这辈子只能过一次。所以我现在踢足球,骑自行车来这个访谈,我举铁,玩趴板冲浪,还和我太太、家人还有儿子一起做各种事。所以,玩得开心很重要。
便签引用
44:53
son. So, it it's important to have fun. I think uh on the career side is one of the uh pieces of advice is uh by a he wrote this book the habits of very effective people I think and he has a little quadrant and he divides it of urgent and not urgent and important and unimportant and kind of there's so many things in our technology like email and texting to for you to focus on the urgent things but you really shouldn't shouldn't be spending a lot of time on the unimportant urgent things. And it takes self-discipline to set aside time for the uh important non-urgent things.
在职业方面,我觉得有一条建议来自——他写了那本书,《高效能人士的七个习惯》,我记得是。他画了一个四象限,分成紧急/不紧急、重要/不重要。我们的技术里有太多东西,比如邮件和短信,会把你的注意力拉到紧急的事情上,但你真的不该花大量时间在那些不重要但紧急的事情上。而要为那些重要但不紧急的事情专门留出时间,是需要自律的。
便签引用
45:37
But if you don't block that out, you can just not have time to do anything. I get to see other people's calendars at Google and there's managers that every half hour from 8 to 6, five days a week are scheduled and I don't see how you have time to think and reflect on things like that. Um, other career advice things is, uh, I remember, uh, I kind of woke up one morning and it was like God spoke to me. I was like thunder struck. And it says it's not how many things you start, it's how many things you finish.
如果你不把那段时间圈出来,你可能就什么都没时间做了。我在谷歌能看到别人的日历,有些经理从早八点到晚六点、一周五天,每半小时都排满了,我实在想不出这样还怎么有时间去思考和反省。另外一条职业建议是,我记得有天早上醒来,感觉就像上帝对我说话一样,像被雷劈了一下。那句话是:重要的不是你开了多少个头,而是你完成了多少件事。
便签引用
46:09
It seems relatively obvious, but you know, I that's not the way I was acting. I had they had many things going on. But after that, it was like there's one main thing I'm doing at time. So when Hennessy and I wrote our textbook, that was the main thing. when I was department here, head here, that was the main thing I did. Uh I would do some other little things there. And John Hennessy, he he wrote a book too, a kind of a a career advice book based on his presidency. And one of the things he said in there, you know, you're only going to be remembered for the five or six things you've done in your life, not for the hundreds of little things. So to give yourself a chance to uh have some things you're really proud of, it's better to concentrate on a few of them hoping that some of them will turn out to be a big deal rather than scatter yourself to to um many things. But if you're interested, you can yeah, if you look for life lessons, David Patterson first half century, you can see the whole list of I think there's 16 lessons
这看起来相当显而易见,但你知道,我当时并不是那样做的。我手上同时有很多事。但从那以后,就变成同一时间只做一件主要的事。所以当我和亨尼斯写我们那本教科书时,那就是我的主线。我在这里当系主任时,那就是我做的主要事情。我也会做点别的小事。约翰·亨尼斯也写了一本书,算是一本基于他当校长经历的职业建议书。他在里面说过一句话:你这一生最终被记住的,只会是你做过的五六件事,而不是那几百件小事。所以,为了给自己一个机会,能有几件真正让你自豪的事,最好是集中精力做少数几件,指望其中某几件能成为大事,而不是把自己摊到太多事情上。如果你有兴趣,可以去搜“life lessons, David Patterson, first half century”,就能看到完整的清单,我记得一共有 16 条。
便签引用
47:10
altogether. you mentioned that uh you you didn't want to reflect back on your life and feel like you didn't have enough family time and yeah, I've I don't think I've ever heard anyone say the opposite. Why do you think that is? Why is it that everyone looks back on their life and they never regret, you know, a lot of family time? I I imagine there's got to be at least one person that says, "I spent too much time with my family. [laughter] >> My career suffered." >> Yeah, my career suffered. Well, you know what's you know what's it's it's pretty philosophical. I mean what's what's what's life all about?
你提到你不希望回顾自己一生时,觉得自己陪家人的时间不够。是啊,我好像从没听过有人说反过来的话。你觉得这是为什么?为什么每个人回顾一生时,都从不后悔陪家人的时间太多?我猜总该至少有一个人会说:“我陪家人的时间太多了,[笑] >> 害了我的事业。”>> 是啊,我的职业生涯受到了影响。这个嘛,你知道吗,这其实挺有哲学意味的。我是说,人生到底是为了什么?
便签引用
47:47
What's success? Right? You have to you have to figure out what that means for you. I mean if if you have uh I don't know if you have financial goals of being a you know a millionaire or billionaire or something like that and that's how you judge your life. Okay. Good luck. You know it's pretty hard to make it. Uh it's pretty hard to do. Um but I think it's the you know I when I was finishing my PhD I read this book by studs Turkl called working where he interviewed all these people in his careers and they look back and what they liked and what they didn't like what what they felt about their careers and what I got out of it the people who worked with people like ministers or teachers or doctors uh felt really good about what they did with their careers and the people who did more ephemeral stuff um you know like uh you know technology things or airplanes or something like that are long gone didn't feel as good about it. It was the people that they worked with that they really cared about. And I went out with a
什么叫成功?对吧?你得自己弄明白这对你意味着什么。我是说,如果你有——我不知道——如果你的财务目标是当个百万富翁、亿万富翁之类的,而你就用这个来评判自己的人生。那好吧,祝你好运。你知道,那挺难做到的。呃,那真的挺难的。嗯,但我觉得,你知道,我读博快毕业的时候,读了斯塔兹·特克尔(Studs Terkel)写的一本书,叫《工作》,他采访了各行各业的人,问他们的职业生涯,问他们回头看的时候喜欢什么、不喜欢什么,对自己的事业有什么感受,我从里面读出来的是:那些跟人打交道的人,比如牧师、老师、医生,对自己一生的事业感觉特别好;而那些做更短暂的东西的人,嗯,你知道,比如说做技术的、做飞机之类的,那些东西早就不在了,他们的感觉就没那么好。真正让他们在意的,是那些和他们共事的人。我还跟我们这儿一位退休的工学院院长一起出去过,他
便签引用
48:45
retired engineering dean here and he uh into a meeting who I knew and he said, you know, Dave, as I flinked my ear, it was it wasn't the projects, it was the people that I worked at the matter and I thought I knew that from a long time ago. So that's I mean this is kind of senior persons offering advice. my I think you're going to care more about the people you've worked with uh and the people you've helped will be a bigger deal. That's part of and you know one of the other things about personal happiness is they studied happiness.
去参加一个会,我认识他,他说,你知道吗,Dave,回头想想,重要的不是那些项目,而是跟我共事的那些人。我当时想,这道理我很久以前就懂了。所以说,我是说,这有点像前辈们给建议。我觉得你会更在意那些和你共事过的人,还有你帮助过的人,那会是更重要的事。这也是……你知道,关于个人幸福还有一点,他们研究过幸福这件事。
便签引用
49:13
Psychologists used to just study crazy people but they started like why are people happy and they know what the reasons are you know uh have a job that you like you know have friends and family helping other people helping other people makes you happy that they they know this having something kind of either a religious side of it doesn't have to be a formal religion but like contact with nature the grandeur of nature but the the kind of the list of things you need to do to be happy is well understood And uh and you know there these h and there's our world is filled with unhappy billionaires, right?
心理学家过去只研究有问题的人,但后来他们开始问,人为什么会快乐,而且他们知道原因是什么,你知道的,有一份你喜欢的工作,有朋友和家人,帮助别人——帮助别人会让你快乐,这些他们都清楚,还有拥有某种类似信仰的东西,不一定是正式的宗教,比如说接触大自然,感受自然的壮阔。但那份让人快乐所需要做的事情的清单,是被研究得很透彻的。而且你知道,我们这个世界上到处都是不快乐的亿万富翁,对吧?
便签引用
10勇气、乐观与婚姻九字
49:49
If money was the thing that made you happy, why are these people very wealthy people so mad about things. So yeah, this is me passing on advice. In one of your talks, you had uh it said what worked well for me and it was kind of some reflections and one of the things in there I thought was unique and interesting. You mentioned that courage was a big part of your career and I I don't hear that too often. I was curious why you >> uh say courage is so important in a career. Yeah, I I I think that's I mean that might be partially my personal makeup, but you know I was uh I was I was kind of the youngest kid in my class and kind of small. It took me a while to grow despite age. So I was always small but I my parents encouraged me to go out for wrestling and uh wrestling gives you self-con physical self-confidence because you you know spend years doing that. that I did in high school and college. And so I think uh partly you know technically uh you know having courage to do things it kind of goes along with the advice is
如果金钱真是让人快乐的东西,那这些非常有钱的人为什么对很多事情那么愤怒。所以是的,这就是我在传递一些建议。在你的一次演讲里,你提到过——标题大概是「哪些做法对我有效」,算是一些反思,其中有一点我觉得很独特、很有意思。你提到勇气是你职业生涯中很重要的一部分,我不太常听到这种说法。我很好奇你为什么会说勇气在职业生涯中如此重要。>> 是的,我想这可能有一部分是我个人的性格使然,但你知道,我当时……算是班里年纪最小的孩子,个头也小。虽然年纪在长,但我长个子长得慢。所以我一直都很小只但我父母鼓励我去练摔跤,而摔跤能给你身体上的自信因为你会花好几年时间去练。我高中和大学都练过。所以我觉得,某种程度上从技术层面说,有勇气去做一些事,这跟那句老话是一致的:命运眷顾勇者。这
便签引用
50:57
fortune favors the bold. Uh that's this that goes that's advice is 2,000 years old. I mean it's hard to figure this out. U Helen Keller wrote you know e even even trying to play it safe you still get caught. And so it turns out you might as well you might fail no matter what. And if you take a big chance if that's you can succeed if you don't take the chance you probably won't succeed if you play it safe. So fortune favors the gold. And I think it takes courage to do that. I think also it just for me intellectually I feel like if there's something not right I need to stand up and confront it. And I think that kind of ironically comes from the wrestling side of my personality where if I see something somebody getting picked on or something like that, I'm I'm going to stand and try and stop it.
这句话已经有两千年历史了。我是说,这道理很难自己琢磨明白。海伦·凯勒写过就算你想求稳,你照样躲不掉。所以到头来,不管怎样你都可能会失败。而如果你去冒一次大险,你就有可能成功;如果你不去冒这个险,一味求稳,那你多半不会成功。所以说,命运眷顾勇者。我觉得这需要勇气。我也觉得,对我来说,在思想层面上,我总觉得如果有什么事情不对劲,我就得站出来直面它。说来有点讽刺,我觉得这一点其实来自我性格里摔跤的那一面——如果我看到有人被欺负之类的事,我我就会站出来,试着去制止。
便签引用
51:48
And I feel that that same responsibility intellectually if people are making bad arguments or doing doing something that we need to stand up and do it. And I feel good about that. The cautionary part about that is uh as because I guess one of my senior faculty members saw this nature of me. He said one of the sayings is friends come and go but enemies accumulate. This is an old saying. So if you think about it you kind of people you went to high school with a while friends you kind of forget but somebody who you really does dislike you never forget that you dislike that person. So standing up when it's important, but be careful when you make enemies because you know they're going to stick around for a long time.
而在思想层面上我也有同样的责任感——如果有人在提糟糕的论点,或者在做一些我们得站出来把它做了。我对这一点感觉挺好的。但要提醒的一点是,呃,因为我想我的一位资深教职同事看出了我这种性格。他说有句话是这么讲的:朋友来来去去,但敌人会越积越多。这是一句老话。所以你想想看,你高中时的那些朋友,过一阵子你也就慢慢淡忘了,但要是有个人你真的很讨厌,你永远不会忘记你讨厌那个人。所以在重要的时候要站出来,但树敌的时候要小心,因为你知道他们会跟着你很久很久。
便签引用
52:29
>> When you say something's gone bad technically, do you mean uh someone was incorrect? >> Yeah. When you know when they're weak, you know, either politically or, you know, or or technically when it's you it's a weak argument, right? that that I I think give us I think I really like there to be a marketplace of ideas and we hone the ideas by arguing them and so if it's a if it's a specious argument even if a person is you know a leader of a company and it just doesn't make sense I feel it's kind of the responsibil it's better for the company if somebody stands up and points that out than to just let them get away with it and then um and then and you know it's a little bit confrontational But, you know, as long as people all agree that, you know, this is for the greater good, we're going to we we need to get the right ideas out there. And so, let's argue about the ideas to to see uh you know, polish them to make them stronger. Uh I think that's important in science, in engineering, um and kind of uh in life too. There's a
>> 你说某件事在技术上出问题了,是指呃有人说错了吗?>> 是的。就是当你知道他们的立场站不住脚的时候,无论是在政治层面,还是在技术层面,当那是个很弱的论点的时候,对吧?我觉得——我真的很希望有一个思想的市场,我们通过争辩来打磨想法。所以如果那是一个似是而非的论点,哪怕说话的人是一家公司的领导,而它就是讲不通,我觉得这某种程度上是一种责——如果有人站出来把这一点指出来,对公司是更好的,总比就这么让他们蒙混过去要好,然后呃——然后,你知道,这会有点对抗性。但是,你知道,只要大家都认同这是为了更大的利益,我们就要,我们需要把正确的想法讲出来。所以,我们要就这些想法争论,你知道,把它们打磨得更有力。我觉得这在科学里、在工程里都很重要,嗯,在生活里也一样。现在这个国家
便签引用
53:36
lot of stuff going on right now in the country that uh is worrisome and uh I've certainly stood up and wrote opeds about things that I think are wrong and need to be corrected and uh you know if people are afraid to do that it's hard to be um optimistic about the future if people are afraid to stand up when when there's wrongs and uh try and stop them. You also mentioned optimism in the talk and you had this story I wonder if you're willing to >> sure. [laughter] So I would say in engineering uh you know it's it's hard to know right uh but I think you need to be kind of optimistic or positive have a positive outlook because so many things could go wrong and then so my story personal story that illustrates it's going back to high school when I'm 16 I'm dating this very attractive girl and I screw up my courage and ask her if we would be exclusive at the time we the phrase we used was going steady and she looked at me and said And you know, she was 16. She had dated other guys and thought we were pretty young. And she
有很多事情让人担忧,我当然也站出来过,写过评论文章,谈那些我认为是错的、需要纠正的事情。你知道,如果大家都不敢这么做,那就很难对未来保持乐观了——如果看到不对的事,大家都不敢站出来、不敢去阻止的话。你在演讲里也提到了乐观,还讲了一个故事,不知道你愿不愿意——>> 当然可以。[笑] 我想说,在工程里,你知道,很难事先知道结果对不对。但我觉得你得保持一点乐观,或者说保持积极的心态,因为有太多事情可能出岔子。那我讲个能说明这一点的亲身经历吧,得回到高中,我 16 岁,在跟一个非常漂亮的女孩约会,我鼓起勇气问她要不要只跟我一个人交往——当时我们的说法叫“going steady”(确定关系)。她看着我说……你知道,她才 16 岁,之前也和别的男生约会过,觉得我们还挺小的。她
便签引用
54:42
said, "Well, Dave, you're such a nice guy. I don't know how to say no." For me, as a logical person, I don't know. Sounded like a yes. And so I hugged her and said, "Great." And so she uh she in her mind, she thought, "Well, I'll let him down gently later." But we've been married 59 years now, and she she hasn't let me down yet. So that was a case where optimism uh paid off. >> I think everyone that hears a a healthy relationship for that long, they might wonder how you did it. >> I used to tell people uh you know if you go to if you go to weddings, the marriage vows are really great, right?
说:“戴夫,你人真的太好了,我都不知道该怎么拒绝你。”对我这种讲逻辑的人来说,我也不知道,这听起来就像是答应了。于是我抱住她说:“太好了!”而她心里想的是:“行吧,以后再慢慢让他死心。”结果我们到现在已经结婚 59 年了,她还没让我死心呢。所以这就是乐观得到回报的一个例子。>> 我想每个听到一段这么长久的健康关系的人,都会好奇你们是怎么做到的。>> 我以前常跟人说,你去参加婚礼的时候,那些结婚誓词都写得特别好,对吧?
便签引用
55:20
But nobody can remember their wedding vows. I used to say remember your wedding vows, but nobody remembered that. So we boiled it down to nine magic words and it's just three sentences and they start IU, I got to say all three and it's I was wrong. you were right. I love you. Okay, those are the those nine words. And this applies to both both partners in a relationship, not just one partner. Uh but yeah, if you can say them all and no substitutions, I was wrong, you're right, you're a jerk. You know, you can't do that. It's if you can remember those nine words, that can help you have a a long relationship like uh my wife and I have. And then last question for you, like knowing everything you know now from your career, if you could go back to yourself when you had just entered the industry and give yourself advice, what would you say?
但没人记得住自己的结婚誓词。我以前总说“记住你们的誓词”,可谁都记不住。所以我们把它浓缩成了九个神奇的字,其实就是三句话,而且三句都得说全,那就是:我错了,你对了,我爱你。好,就是这九个字。而且这适用于关系里的双方,不是只有一方要说。嗯,但确实,你得把三句都说全,不能换词——不能说“我错了,你对了,你这个混蛋”。你知道,那可不行。如果你能记住这九个字,就能帮你维持一段长久的关系,就像我和我太太这样。那最后一个问题:以你现在整个职业生涯积累的所有认知,如果你能回到刚进入这个行业时的自己面前,给自己一个建议,你会说什么?
便签引用
11冒名顶替与唯一的遗憾
56:09
>> I mean, I think when I got here because you know the imposttor syndrome, I was a UCLA graduate student and suddenly I'm a Berkeley professor. So that just doesn't seem like that was very intimidating. But I after a while I just thought well I'm probably not gonna get tenure so I should just have a good time. So I think I already I already had a right attitude about it. I think that first year I think it it was very stressful for my while I was trying to handle the imposttor syndrome and be a Berkeley professor but I think after that I handled it pretty well. I you know I I did all the things with the kids and stuff. So uh there's a version of that question is like is there anything I would do over again? There's one thing I would have I was chair of the u of this the architecture community the sig arch as it's called and they have an annual conference of the year and I was this was in the 1990s I think I was the chair and what I wasn't aware is at these conferences there were men who were
>> 我想,我刚来的时候是有冒名顶替综合症的,我本来是 UCLA 的研究生,突然就成了伯克利的教授。这感觉挺不真实的,当时确实挺让人发怵的。但过了一阵子我就想,反正我大概也拿不到终身教职,那还不如好好享受这段时间。所以我觉得我当时其实已经有了正确的心态。第一年我觉得压力非常大,一边要应付冒名顶替综合症,一边要当好伯克利的教授,但那之后我就处理得挺好了。你知道,我该陪孩子的事都做了。所以,这个问题还有另一个版本,就是:有没有什么事我想重来一次?有一件事。我当时是这个体系结构学界组织——就是大家说的 SIGARCH——的主席,他们每年都有一个年会,那大概是 1990 年代,我是主席。而我当时不知道的是,在这些会议上,有些男性在骚扰年轻女性。我当时就没想到,你知道,
便签引用
57:11
harassing young women at this uh conference I just didn't think you know people like you know young people people like me nobody would do that only an idiot would do that can't possibly be happening but it was happening and I wish I somebody had said something to me about it uh and because I would have I would have straightened out any man doing I would have threatened his life if he were to do that today the only comforting thing is Serita a who is a famous computer architect uh she said she also was not aware that that was going on it became clear later you know I mean uh uh five or 10 years later it became more clear that this was going on and there were mechanisms but that's the one thing I wish you know if I could go back in time I would have figured that out and I would have straightened men out who were doing that and they wouldn't that would have stopped I believe that would have stopped them >> yeah [laughter] thank you so much for your time today I really appreciate it >> all right and thanks for the interview
像我们这样的人、年轻人,居然会有人干这种事,只有蠢货才会那么做,觉得根本不可能发生。但它确实在发生。我真希望当时有人跟我说一声,因为我一定会去教训任何这么做的人,要是放到今天,谁敢那样,我会威胁他别想好过。唯一让我稍微好受一点的是,塞丽塔(一位著名的计算机体系结构学者)说,她当时也不知道这些事在发生。这些是后来才清楚的,我是说,五年、十年之后,才越来越清楚这种事一直在发生,后来也有了相应的机制。但这就是我唯一想改的事——如果我能回到过去,我一定会把这件事弄清楚,去教训那些这么干的男人,那样他们就不会再这样了,我相信那能让他们收手。>> 是啊。[笑] 非常感谢你今天抽出时间,我真的很感激。>> 好的,也谢谢你的采访。
便签引用
12片尾:播客与键盘众筹
58:15
>> hey thank you for watching this podcast if you liked it and you want to see the show grow Please support with a comment or a like. Also, if you have any recommendations for people you want me to bring on, please drop a comment. Guests like Barbara Liskoff, Mike Stonereaker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed. Here's a glance at the prototype. It's a split keyboard, so there's two sides. um this isn't the case, but yeah, we launched on Kickstarter and we hit our goal within 8 hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units. Um we're now working on the long journey of building the tooling now and so if you still want to pick one up, I've left the late pledges open on Kickstarter, so you can grab one there. I'll put a link in the description. Thank you again for watching the podcast and I'll see you in
>> 嘿,谢谢你收看这期播客。如果你喜欢,也希望这个节目能成长起来,请用一条评论或者一个赞来支持一下。另外,如果你有希望我请来的嘉宾推荐,欢迎在评论区留言。像 Barbara Liskov、Mike Stonebraker、Marc Brooker 这些嘉宾,都是因为有人留言我才请来的。另外说一件事,除了播客之外,我还在做一款我一直希望存在的人体工学键盘。这是原型机的一个大概样子。它是分体式键盘,所以有左右两半。嗯,这个还不是最终外壳,不过我们已经在 Kickstarter 上线了,上线 8 小时内就达成了目标。如果你是抢到早期批次的人之一,我真的非常感谢。嗯,我们现在正在做建产线工具这段漫长的工作。所以如果你还想入手一台,我在 Kickstarter 上留了后期众筹的通道,你可以在那儿买到。我会把链接放在简介里。再次感谢你收看这期播客,我们下期见。
便签引用
59:10
the next
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

图灵奖得主 David Patterson 回顾 RISC 与 CISC 之争的来龙去脉,解释为何 RISC 最终赢得 99% 的处理器市场,并梳理登纳德缩放失效、摩尔定律放缓如何把行业推向 GPU/TPU 这类领域专用架构,最后分享职业与人生建议。

核心要点

  • RISC 与 CISC 之争本质是"指令复杂度 vs 执行速度"的比值问题。 1970 年代微处理器只是玩具,Intel、TI 等厂商模仿 IBM 大型机和 DEC 小型机,用摩尔定律增加的晶体管堆砌越来越复杂的指令,理由是"抬高抽象层次、缩小与软件的差距"。Patterson 和 Hennessy 认为编译器可以填补这个差距。几年后数据出来:RISC 程序需要多执行约 30%~40% 的指令,但每条指令快 4~5 倍,净收益约 3~4 倍加速。
  • 早期争论激烈是因为当时的体系结构缺乏定量方法。 1970~80 年代的设计靠直觉和"感觉",教科书像产品目录,只罗列各机器特性。没有数字就无法裁决,争论沦为"针尖上能站几个天使"式的哲学辩论,这也是后来 Hennessy 与 Patterson 推动定量方法的动因。
  • "CISC 赢了"是只看到 2000 年前 PC 时代的短视结论。 x86 靠二进制软件分发形成了难以撼动的生态。但英国 Acorn 受伯克利论文影响做出 Acorn RISC Machine,Apple 为 Newton 采用后改名 Advanced RISC Machine 即 ARM,随后 Nokia GSM 手机采用,ARM 从此统治移动设备,累计出货约 3500 亿颗。如今 99% 的处理器是 RISC,Apple PC 已转向 ARM,Amazon、Microsoft、Google 也在云端自研 ARM 处理器,x86 市场在萎缩。
  • x86 的存活方式是在硬件里把 CISC 指令翻译成 RISC 微操作。 Intel 在 2000 年代初这样做,付出了翻译开销,却能借用 RISC 的所有好处,靠 PC 软件生态的价值抵消成本。深层看,CISC 依赖 Maurice Wilkes 发明的微程序解释器,解释比编译慢 5~10 倍,在 1960~70 年代内存技术下合理,但到 1980 年已无必要。近 20 年没人再设计需要微码解释器的指令集,仅在数万晶体管的微型应用中或许还说得通。
  • 编译器进步是 RISC 成立的关键前提。 早期编译器寄存器分配能力差,C 语言甚至要让程序员用 register 提示哪些变量放寄存器。RISC 的应对是干脆把寄存器从 8 或 16 个增加到 32 个,让编译器"即使不够聪明也够用"。研究还发现编译器基本不用 CISC 的复杂指令,例如过程调用专用指令,编译器作者反而说拆成简单指令更快,于是出现"为解释器付出开销、指令却没人用"的荒谬局面。
  • 登纳德缩放失效是架构分岔的真正转折点。 晶体管翻倍的同时阈值电压下降,功耗按平方效应抵消,因此微处理器长期维持在 20~30 瓦,Hennessy 与 Patterson 的教科书前三版直到 2000 年都没把功耗当议题。约 2005 年缩放停止,Intel 某代产品因发热失败,业界被迫从单个复杂核心转向 2、4、8 个简单核心,并行化的负担转嫁给程序员。约 10 年后摩尔定律也放缓,通用处理器几乎停滞,2005 年的 2~3 GHz 至今基本没变。
  • 领域专用架构在 2015 年前后与机器学习的爆发恰好相遇。 GPU 原本是 2000 年左右为游戏和电影图形做的利基产品,不支持虚拟内存和编译器,靠硬件多线程掩盖内存延迟,用 16/32 位浮点,售价数百美元。2006 年黄仁勋资助开发 CUDA 让其可编程。2012 年 AlexNet 作者因在多伦多大学上过 CUDA 课程而用 GPU 训练神经网络,碾压竞争对手,此后整个领域转向神经网络加 GPU。
  • TPU 是白纸重来的设计,性能差距震动了业界。 Google 意识到神经网络核心是矩阵乘法,做了巨大的矩阵乘法单元,去掉三级缓存改用软件调度的暂存内存,因为访存模式已知。还首创指数位宽于尾数位的 bfloat16,因为 ML 需要范围而非精度。2016 年发布时推理性能约为同期 GPU 的 30 倍、CPU 的 80 倍,随后 Intel 大举收购、Nvidia 改设计、各超大规模厂商纷纷自研。TPU 训练架构的基本框图从 2015 年沿用至今,核心仍是矩阵单元、高带宽内存和向量单元。
  • 摩尔定律已放缓,这对真正的工程师并无争议。 摩尔定律的定义就是芯片内晶体管每 1~2 年翻倍,它曾指导整个半导体行业 50 年的投资方向。如今改进不再均匀:逻辑门仍在进步,SRAM 几乎停滞,DRAM 密度从每 3 年提升 4 倍变成约 10 年一次。最新 GPU 和 TPU 把两颗满掩膜版尺寸的裸片封装在一起,从外部看晶体管仍在翻倍,但芯片内部早已放缓。Patterson 认为 Jim Keller 说"摩尔定律没死"更多是半导体从业者的情感因素,数据不支持。
  • 现在的性能增长靠封装、窄数据类型和更大的矩阵单元,而 Nvidia 的真正护城河是库。 浮点从 64 位一路降到 8 位甚至 4 位。MLPerf 仿照 SPEC 建立了统一基准,但结果取决于"架构加为每次新硬件定制重写的库"。Nvidia 工程师数量远超初创公司,因此后者难以在 MLPerf 上展示实力,这也是很多厂商不跑 MLPerf 的原因。Google 靠更多依赖编译器来弥补库工程师不足。

结论与值得注意的细节

  • Patterson 的判断是:如果登纳德缩放和摩尔定律仍在,今天应该有 100 THz 的通用处理器,GPU 依然会是利基产品。正因为通用路径走不通了,CPU 才退守到操作系统和编译器等必须通用的领域,资金全部涌向 AI 加速器,每个处理器设计者都得回答"我的设计如何融入 AI 世界"。
  • 关于 GPU 术语,他们的教科书专门做了一张"罗塞塔石碑",把 Nvidia 自创的术语对应到主流体系结构术语,反映 GPU 曾长期游离于主流架构之外。
  • 职业建议部分:写过《如何拥有糟糕的职业生涯》系列演讲,最新一篇是《如何给 AI 一个糟糕的碳足迹》。核心是家庭优先、选择幸福而非财富、留时间玩乐、按"重要但不紧急"来分配时间,以及"不在于开始多少事,而在于完成多少事",一个人一生只会因五六件事被记住。
  • 他强调勇气:来自摔跤训练的自信让他习惯在论证薄弱时站出来反驳,哪怕对方是公司领导。但也引用告诫"朋友来来去去,敌人不断积累",树敌要谨慎。
  • 唯一想改写的过去:1990 年代任 SIGARCH 主席期间不知道会议上有男性骚扰年轻女性,若当时有人告知,他会直接制止。
  • 婚姻 59 年的秘诀被他压缩成九个英文词、三句话:我错了,你是对的,我爱你。三句必须全说,不能替换。
核心句型 · 10
1. It turned out (that) …
“It turned out you needed like maybe 30% 40% more simple instructions to execute in a program, but you could run them four or five times faster.”
用于引出事后才得知的结果,尤其是与预期不同的事实。学术与口语通用,后接完整句子。仿写:It turned out the bug was in the compiler, not the hardware.
2. after the dust settled, …
“Well, after the dust settled a few years later and we started getting numbers, it turned out …”
「尘埃落定之后」,用来把激烈争论与后来的冷静结论分开。适合叙述争议事件的后续。可换成 once the dust settled。
3. That's a pretty myopic view.
“He said and CISK won. Well, that's a pretty myopic view”
用 pretty 缓冲一个否定评价,是英语口语中礼貌反驳的常见结构。类似:That's a pretty narrow reading of the data.
4. X didn't even Y
“We didn't even talk about power as an issue for the first three editions.”
even 放在动词前强调「连……都没有」,凸显某事当时完全不在考虑范围内。仿写:We didn't even have a name for it back then.
5. It's not A, it's B.
“It's not how many things you start, it's how many things you finish.”
对比否定式金句结构,前半否定常见误解,后半给出真正要点。两个分句最好平行对称,便于记忆和引用。
6. Nobody on their deathbed says …
“Nobody on their deathbed says, you know, I wish I'd spent more time in the office.”
用「临终没人会说」做反证,是英语中常见的价值排序论证。后接 I wish I had + 过去分词,表达对过去的遗憾。
7. Why would I/you …?
“Why would I pick something that made me wealthy and unhappy? Why would you do that?”
反问句表示「根本没理由这么做」,would 体现假设语气。适合论辩中质疑对方选择的合理性。
8. you might as well …
“It turns out you might as well you might fail no matter what.”
「既然结果差不多,不妨……」,用于在两种选择成本相近时推荐更进取的一种。常与 no matter what、anyway 搭配。
9. It's A plus B that …
“It's the libraries plus the architecture that's this powerful advantage”
强调句式,把两个要素并列后用 that 引出结果,突出「组合」才是关键。适合解释复合原因。
10. the question came down to what …
“So kind of the question came down to what was that ratio?”
come down to 表示「归根结底取决于」,后接名词或 wh- 从句,用于把复杂争论收敛为一个可回答的核心问题。
词汇精讲 · 112 · 按出现顺序
myopic /maɪˈɑːpɪk/ adj. 0:00
短视的,目光狭隘的(原义为近视的)
energy bound phr. 0:00
受能耗限制的;-bound 表示「受……约束」,如 memory bound
toy /tɔɪ/ n. 0:33
此处指「玩具级别的东西」,即不被当真的产品
mainframes /ˈmeɪnfreɪmz/ n. 1:11
大型机,企业级大型计算机
instruction set n. phr. 1:11
指令集,处理器能识别的全部指令的集合
level of abstraction n. phr. 1:11
抽象层次;raise the level of abstraction 提高抽象层次
prevailing wisdom n. phr. 2:15
主流看法,普遍接受的观点
came down to phr. v. 2:15
归结为,最终取决于
after the dust settled idiom 3:18
尘埃落定之后,局势明朗之后
ferocity /fəˈrɑːsəti/ n. 3:18
激烈,凶猛
the net /net/ n. 3:18
净结果,扣除得失后的最终结果
gut /ɡʌt/ n. 4:05
直觉;using their gut 凭直觉
dissatisfying /dɪsˈsætɪsfaɪɪŋ/ adj. 4:05
令人不满的
how many angels on the head of a pin idiom 4:05
针尖上能站几个天使,喻无法验证的空洞争论
qualitatively /ˈkwɑːlɪteɪtɪvli/ adv. 4:05
定性地(对应 quantitatively 定量地)
impediment /ɪmˈpedɪmənt/ n. 5:08
障碍,阻碍
binary /ˈbaɪnəri/ n. 5:08
二进制可执行文件
forerunner /ˈfɔːrrʌnər/ n. 5:46
先驱,前身
recristened /ˌriːˈkrɪsnd/ v. 6:46
重新命名(正确拼写 rechristened)
embraced /ɪmˈbreɪst/ v. 6:46
欣然采用,接纳
leaps and bounds idiom 7:59
突飞猛进;常作 by/in leaps and bounds
overhead /ˈoʊvərhed/ n. 8:52
额外开销(计算机术语)
niche /nɪtʃ/ n./adj. 8:52
小众市场;小众的
trade-off /ˈtreɪdɔːf/ n. 8:52
权衡取舍
ad hoc /ˌæd ˈhɑːk/ adj. 9:57
临时拼凑的,无系统方法的
microprogramming /ˌmaɪkroʊˈproʊɡræmɪŋ/ n. 9:57
微程序设计,用存储器内容实现控制逻辑
interpreter /ɪnˈtɜːrprətər/ n. 10:53
解释器,逐条翻译执行的程序或硬件
latencies /ˈleɪtənsiz/ n. 10:53
延迟(latency 的复数)
nothing prevents you from phr. 11:58
没什么阻止你做……,即技术上可行
assembly language n. phr. 12:59
汇编语言
allocate /ˈæləkeɪt/ v. 12:59
分配(资源);register allocation 寄存器分配
step in phr. v. 14:06
介入,插手
shortcoming /ˈʃɔːrtkʌmɪŋ/ n. 15:29
缺点,短板
nonsensical /nɑːnˈsensɪkl/ adj. 16:58
荒谬的,不合理的
caches /ˈkæʃɪz/ n. 16:58
缓存(高速缓冲存储器)
domain specific adj. phr. 17:56
领域专用的
virtual memory n. phr. 17:56
虚拟内存
threshold voltage n. phr. 18:35
阈值电压,晶体管导通所需的最低电压
dinard scaling n. phr. 18:35
登纳德缩放定律(正确拼写 Dennard scaling)
multicore /ˌmʌltiˈkɔːr/ n./adj. 19:37
多核(处理器)
barely /ˈberli/ adv. 20:23
几乎不,仅仅
shuffle /ˈʃʌfl/ v. 21:25
重新调配,挪动
burst on the scene idiom 21:59
突然登场,一夜之间引人注目
swamped with phr. 21:59
被……淹没,应接不暇
debuted /deɪˈbjuːd/ v. 21:59
首次亮相,推出
multi-threading /ˌmʌltiˈθredɪŋ/ n. 23:08
多线程
floating point n. phr. 24:04
浮点(数)
Rosetta Stone n. 24:04
罗塞塔石碑,喻两套体系之间的对照表
single precision n. phr. 24:45
单精度(32 位浮点)
proprietary /prəˈpraɪəteri/ adj. 25:33
专有的,私有的(非开放标准)
heritage /ˈherɪtɪdʒ/ n. 26:37
传承,来历
advocates /ˈædvəkəts/ n. 26:37
拥护者,支持者
clean slate idiom 27:42
白纸一张,从零开始
crushed the competition phr. 27:42
碾压对手
exponent /ɪkˈspoʊnənt/ n. 28:43
指数(浮点数的指数位)
fraction /ˈfrækʃn/ n. 28:43
此处指浮点数的尾数(小数部分)
blew the doors off idiom 29:42
远远胜过,压倒性领先
inference /ˈɪnfərəns/ n. 29:42
推理(模型训练后的预测阶段)
watershed /ˈwɔːtərʃed/ n. 30:15
分水岭,转折点
hyperscalers /ˈhaɪpərskeɪlərz/ n. 30:15
超大规模云服务商(如 AWS、微软、谷歌)
raises all boats idiom 30:15
水涨船高,使所有人受益
accelerators /əkˈseləreɪtərz/ n. 31:23
加速器(专用计算芯片)
mended /ˈmendɪd/ v. 32:12
此处指修正(正式说法为 amended)
guideline /ˈɡaɪdlaɪn/ n. 32:54
指导方针
chiplets /ˈtʃɪplɪts/ n. 33:35
小芯片,多颗裸片封装组合的设计
full reticle n. phr. 33:35
满光罩尺寸,光刻单次曝光能做的最大芯片
tapered off phr. v. 34:42
逐渐减弱,放缓
like clockwork idiom 34:42
像钟表一样准时、规律
analog /ˈænəlɔːɡ/ n. 35:41
类似物,对应物
gauge /ɡeɪdʒ/ v. 36:56
估量,判断
make your bets phr. 36:56
下注,做出押注式的决策
benchmarks /ˈbentʃmɑːrks/ n. 37:45
基准测试
performant /pərˈfɔːrmənt/ adj. 37:45
性能好的(技术圈用语)
tailor /ˈteɪlər/ v. 39:35
定制,量身调整
moat /moʊt/ n. 39:35
护城河,喻商业壁垒
deathbed /ˈdeθbed/ n. 42:48
临终之时;on their deathbed 临终时
quadrant /ˈkwɑːdrənt/ n. 44:53
象限
self-discipline /ˌselfˈdɪsəplɪn/ n. 44:53
自律
block that out phr. v. 45:37
(在日程上)圈出、预留时间
thunder struck adj. 45:37
如遭雷击,震惊(通常写作 thunderstruck)
scatter yourself phr. 46:09
把精力分散到太多事上
ephemeral /ɪˈfemərəl/ adj. 47:47
短暂的,转瞬即逝的
grandeur /ˈɡrændʒər/ n. 49:13
壮丽,宏伟
makeup /ˈmeɪkʌp/ n. 49:49
性格构成,气质
go out for phr. v. 49:49
报名参加(校队等)
fortune favors the bold idiom 50:57
命运眷顾勇者
play it safe idiom 50:57
求稳,不冒险
you might as well phr. 50:57
不妨,倒不如(既然结果差不多)
picked on phr. v. 50:57
被欺负,被挑刺
cautionary /ˈkɔːʃəneri/ adj. 51:48
警示性的
accumulate /əˈkjuːmjəleɪt/ v. 51:48
累积
stick around phr. v. 51:48
留下不走,长期存在
specious /ˈspiːʃəs/ adj. 52:29
似是而非的
marketplace of ideas n. phr. 52:29
思想的市场,让观点自由竞争的隐喻
hone /hoʊn/ v. 52:29
打磨,磨砺
get away with phr. v. 52:29
做了坏事而不受追究
confrontational /ˌkɑːnfrənˈteɪʃənl/ adj. 52:29
对抗性的
worrisome /ˈwɜːrisəm/ adj. 53:36
令人担忧的
opeds /ˈɑːpˌedz/ n. 53:36
报刊评论文章(op-ed,来自 opposite the editorial page)
screw up my courage idiom 53:36
鼓起勇气
going steady idiom 53:36
(旧时说法)确定恋爱关系
let him down gently idiom 54:42
委婉地拒绝他
paid off phr. v. 54:42
得到回报,奏效
boiled it down to phr. v. 55:20
把……浓缩为
jerk /dʒɜːrk/ n. 55:20
混蛋,讨厌的人(口语)
imposttor syndrome n. phr. 56:09
冒名顶替综合征(正确拼写 impostor syndrome)
tenure /ˈtenjər/ n. 56:09
终身教职
intimidating /ɪnˈtɪmɪdeɪtɪŋ/ adj. 56:09
令人生畏的
harassing /həˈræsɪŋ/ v. 57:11
骚扰
straightened out phr. v. 57:11
教训、纠正(某人的行为)
ergonomic /ˌɜːrɡəˈnɑːmɪk/ adj. 58:15
符合人体工学的
late pledges n. phr. 58:15
众筹结束后的补充认购
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.16468 Bits of Unsolicited Advice 下一期 · NO.166 →Turing Award Winner: Early AI, LLM Predictions, Causality | Judea Pearl
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]