视频库 / NO.128
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

CHM Live | The Silicon Gold Rush: How AI is Driving the Development of New Chips

节目发布 2026-08-28 · Computer History Museum
比尔·戴利 诺姆·朱皮 DDave Patterson
本期追问 · 点击跳到视频对应位置
10:05 英伟达和谷歌的 AI 芯片,是设计者互相靠拢,还是淘汰后的幸存者相似?20:45 AI 一个周末能重写软件库,CUDA 生态还算护城河吗?29:29 硬件要三年、模型三个月一变,芯片该为哪个模型设计?67:54 AI 能做的它都做得更好,工程师还要会手工做吗?
归入 Ⅵ·02 是什么决定了技术的边界? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2026 年夏,计算机历史博物馆(Computer History Museum)举办「硅金热:人工智能如何驱动新芯片研发」现场对谈。对谈者是 NVIDIA 首席科学家兼高级副总裁比尔·达利(Bill Dally)与 Google Fellow 兼副总裁、TPU 首席架构师诺姆·朱皮(Norm Jouppi),主持人是加州大学伯克利分校荣休教授、2017 年图灵奖得主戴维·帕特森(Dave Patterson)。三人自上世纪八十年代起相识,帕特森当时是伯克利的助理教授,达利与朱皮分别在加州理工与斯坦福读博。本文依据现场录音编译整理,仅删去寒暄、口语枝节与重复,其余论证与细节悉数保留。

两条平行的人生

主持人:今晚由我介绍两位老朋友。比尔从加州理工起就在造网络互连的计算机,毕业后去了麻省理工,又造了几台网络超级计算机。大约十年后他回到斯坦福,最终当上计算机系主任。再过十年,2009 年,他成了 NVIDIA 的首席科学家,至今十六年,现在还兼任高级副总裁。他的贡献太多,我不知道该挑哪些讲,那就问 AI 吧。据聊天机器人说,他最重要的贡献是虫洞路由(wormhole routing)及相关的流控技术,这是一套实用的低延迟交换方法,芯片级和系统级网络都用得上。他还和朋友布莱恩·托尔斯(Brian Towles)合写了《互连网络的原理与实践》,堪称网络领域的圣经,把拓扑、路由和流控的设计系统化了,影响学界和工业界几十年。他在斯坦福主持的 Imagine 和 Merrimac 项目,为流处理(stream processing)做了极有力的论证,而流处理正是现代 GPU 计算在概念上的先声。

诺姆在斯坦福读研时参与了 MIPS 项目,那是 RISC 架构的开山之作。毕业后他加入 DEC 西部研究实验室,待了十年左右,再转去惠普实验室。2013 年,他在 DEC 时的老友杰夫·迪恩(Jeff Dean)找他去 Google,为深度学习这套新想法造硬件。鉴于 AI 历史上的种种泡沫,诺姆当时很怀疑,但杰夫说服了他:「深度学习,我们试什么它都灵。」于是他 2013 年入职,现在既是 Fellow 也是副总裁。我也问了 AI 诺姆的贡献,答案是 Google 的 TPU,以及他在 DEC 和惠普做的存储层次设计,尤其是牺牲缓存(victim cache)和预取缓冲区。

准备这场对谈时,我发现诺姆和比尔的人生惊人地平行,能列出十条相似之处。第一,两人的学士、硕士、博士各在三所不同的学校拿的,我则是从头到尾没挪窝。第二,两人都有斯坦福的研究生学位,比尔是硕士,诺姆是博士。第三,导师都是约翰·亨尼西(John Hennessy),我想他今晚也在场。两人都是各自公司的副总裁,都是 ACM 会士,都是 IEEE 会士,都是美国科学促进会会士,都当选了美国国家工程院院士,都拿过 IEEE 西摩·克雷奖,也都拿过 ACM/IEEE 埃克特-莫奇利奖,那是计算机体系结构领域的最高荣誉。两位都很忙,却抽出时间来到这里,请大家为他们的贡献鼓掌。

这次淘金热为何不同

主持人:前三四十分钟我来提问,之后交给观众。今晚的题目是「硅金热」,暗示着一个创新、投资和竞争都异常激烈的时期。站在 NVIDIA 和 Google 的立场看,是什么定义了这个 AI 硬件的淘金时代?它与以往几轮计算机体系结构的热潮,根本区别在哪里?比尔先来。

达利:我认为有三个特征定义了这一轮,也让它与过去几轮截然不同。第一是真实而强烈的经济需求。正如杰夫·迪恩所说,AI 用到哪里就在哪里奏效,于是对 token、对算力、对 AI 的需求是永不满足的。第二,应用虽在快速演进,但相对简单。与过去那些动辄几百万行代码的超算应用相比,Transformer 是个相当简单的东西,你很容易看清要造什么才能让它跑得快。第三,人们愿意快速迭代,没有「积灰的旧程序」(dusty decks)拖后腿。

对照一下九十年代初,那时围绕超级计算也有过一轮体系结构淘金热。事实证明并没有真实的经济需求,一切是 DARPA 的战略计算计划(Strategic Computing Program)撑起来的。他们撒了一大笔钱,钱让人以为下游有个大市场,其实没有。很多人开了公司,绝大多数倒闭了,比如 Thinking Machines 之类。那时的应用都是积灰的旧程序,几千行乃至近百万行代码,你可以把看起来最关键的核心加速,可阿姆达尔定律(Amdahl's law)马上咬你一口,因为剩下那 99% 的代码没加速。如今的应用简单,加速几个基本算子就能拿到很好的结果,加上真实的经济需求,我认为这一轮能站住脚。它不像八十年代的 Lisp 机热潮,也不像九十年代的超算热潮,它确实在创造价值。

主持人:所以这次有真实的市场需求。诺姆,你怎么看?

朱皮:我喜欢关注全球经济的动向,而现在投进来的钱多得令人难以置信。纽约有条通往新泽西的隧道,用了一百年,一直漏水,修缮资金差了十亿美元,多年凑不齐。而 Google 宣布明年资本开支一千零五十亿美元,其他公司的投入也在同一量级。所以,会有赢家和输家,人人都想跑得越快越好。

主持人:你们既然拿真正的淘金热做类比,那谁是淘金客,谁是卖铁镐的?

达利:淘金客是那些去找垂直领域的人。找对了一个能用 AI 的垂直领域,就能像当年的淘金客一样一夜暴富。NVIDIA 和 Google 这样做基础设施的公司呢,就像当年的李维·斯特劳斯(Levi Strauss)和斯坦福,我们向淘金客卖铁镐和铁锹,不管他们淘不淘得到金子,我们都赚钱。

朱皮:不过我们俩都还在给内存厂商付大笔的钱。

达利:没错。我记得美光上个季度赚的钱,抵得上此前十九年的总和。

朱皮:有意思的是,他们营收猛涨,出货的颗粒数却基本持平。

达利:对,利润率冲破天花板了。

趋同还是分道

主持人:我有一位不愿具名的同事指责训练加速器正在「趋同演化」。这是生物学术语,指不同谱系的物种各自独立地演化出相似的特征。他说你们两家的设计已经趋同:一块最大光罩尺寸的计算芯片,上面是一个巨大的脉动阵列矩阵单元,四周能塞多少 HBM 就塞多少,再用最快的 SerDes 做定制互连。我猜你们不会同意。他错在哪里?GPU 和 TPU 究竟是分道的吗?你们各自有没有欣赏对方设计的地方?

朱皮:我先说。谈寒武纪大爆发,别忘了后面还有一次大灭绝。当时有各种怪异而奇妙的生物,长得像巨型鼠妇,身上伸出古怪的附肢,它们没能挺过灭绝事件。我认为现在也是一样,很多初创公司带着五花八门的想法进场,NVIDIA 和 Google 都成功了,是因为那些别的想法能力不够,它们可能会灭绝。

主持人:所以趋同的原因是你们俩品味都好,选对了。

朱皮:是的。

达利:不过我不认为我们真的趋同了。TPU 和 GPU 看起来很不一样。从两万英尺的高空往下看,确实相似:应用决定了某些基本需求。有矩阵乘法,就需要矩阵乘单元;有 softmax 和归一化,就需要向量单元和能算超越函数的部件;需要一定的内存容量、内存带宽和通信带宽。这是应用的要求,任何幸存者都会具备,哪怕那些虫子似的怪物会灭绝。但怎样把这些东西组合起来,里面有大量微妙之处。

多年来我认为 NVIDIA 在数值格式上一直领先。几年前我们在 MLSys 上发了一篇讲向量缩放(vector scaling)的论文,由此产生了 NVFP4,几年后各家 MX FP 格式才跟上。这个微妙之处很要紧:在同等精度下,它能让你比只用 FP8 的对手快一倍。稀疏性也一样,韩松和我 2015 年在 NeurIPS 上发过论文,指出大多数神经网络天然极为稀疏,从 Ampere 一代起我们就在硬件里支持稀疏。这不是每个虫子似的怪物身上都有的东西。

另外我认为 Google 设计 TPU 有个巨大的优势:直到不久前,他们只有一个客户,就是 Google 自己。他们可以精确决定自己想要什么,然后去做。NVIDIA 很幸运,在推理和训练市场都占了六成八的份额,但随之而来的是客户众多,个个都要伺候。他们会来提功能需求,客户足够大,你至少得认真听一听,做出重视的样子。

主持人:所以你得到处出差,跟他们谈。

达利:还得往硬件里塞东西,让各路人马都满意。这会把机器推向某些方向。如果我们能只为自己造想造的东西,不用听客户的,也许能做得更好。所以我有点羡慕 TPU 能这么干。

主持人:你羡慕他们客户少。

达利:对。另外我确实认为他们有个很漂亮的互连网络,训练用的 TPU 采用三维环面(3D torus)。我造过很多超级计算机,在麻省理工造过,九十年代也和 Cray 合作过,都用三维环面网络,所以我对它很有感情。我在那个网络里能看到我那本书的合著者布莱恩·托尔斯的手笔。

朱皮:那本书我们读得很仔细。

达利:说回寒武纪大爆发,我很喜欢这个类比。图形芯片也经历过一次寒武纪大爆发。大约九十年代末、两千年代初,硅谷的图形芯片初创公司不下一百家。经过一轮适者生存的淘汰,剩下两家:NVIDIA 和 ATI,而 ATI 后来被 AMD 收购了。我认为这一轮爆发很可能是同样的结局。

主持人:诺姆,他欣赏你们的三维环面,我也同意。我当然有偏心,我从 Google 领薪水,但我觉得 Google 的光互连是很酷的特性。那 NVIDIA 的架构里有什么让你觉得有吸引力的?或者他们的客户群?

朱皮:比尔说得好,有些运算我们双方都必须支持,矩阵乘、向量运算。TPU 有一点不同:第一代是块 PCI 卡,只做推理,但从第二代开始,它们就是按超级计算机设计的。我们有环面网络,就是比尔和布莱恩书里写的那种。所以我们没有经历过那种痛苦的过程:一路上不断加功能,新功能和老功能互相冲突。我们是从一张白纸开始,直接做超算设计。

主持人:这很有先见之明。你们 2017 年做第二代时就意识到训练需要一台超级计算机,一开始就照这个造。而且 2018 年左右就已经是液冷的了。

朱皮:对,我们液冷已经八年了。

护城河在系统,不在芯片

主持人:顺便感谢几天前在硅谷办的 Hot Chips 大会上的诸位,我到处向人征集问题,收获颇丰。下一个问题:NVIDIA 和 Google 都是世界级的系统公司。系统工程,包括互连、内存、供电、散热、封装,是否已经成为比硅片本身更深的护城河?你们的互连方案各自如何解决扩展中的关键瓶颈?哪里相似,哪里不同?比尔,轮到你先说。

达利:我认为产品就是整个系统,不只是 GPU,甚至不只是装了 GPU、CPU 和网络设备的板卡,而是全部硬件、全部软件,加上让它们协同工作的配置。从 Pascal 一代起,大约 2015 年,我们开始提供一种叫 DGX SuperPod 的东西。那时我们还没有自己的大规模网络,只有内部的 NVLink,还没收购 Mellanox。但我们会告诉客户:用这款交换机,这样配置,这样部署。照做的话,在那个年代你能凑出大约一万块 GPU 的集群,开机就能用。

对比一下,我们为能源部造大型超算的时候,比如 2011 或 2012 年的 Titan,后来的 Summit 和 Sierra,硬件全部到位之后通常还要花六个月调试上线,全是网络配置、各种部件配置里的小毛病。所以你交付的是整个产品,是一个系统。它必须在漫长的训练任务中可靠运行,可用性极高。这里面有大量的系统专长。我们要把机器卖进各种各样的数据中心,每家的做法都不同,所以标准化的配置对快速上线和稳定运行至关重要。

至于不同之处,我认为最大的区别在纵向扩展(scale-up)网络。Google 用三维环面加光电路交换机(optical circuit switch),坏了就绕着配置。我们走的是更传统的网络路线:纵向扩展网络用 Clos 网络搭建,但配上专有的极低延迟链路,绕开传统以太网的大量开销;横向扩展(scale-out)网络则用常规以太网。

朱皮:接着比尔的一点说:很多初创公司没有这种系统经验。路易斯(Luiz Barroso)他们写过「数据中心就是一台计算机」。造那些超级计算机的过程,就是用苦办法学系统。很多教训要亲自吃过苦头才明白,我认为这给了我们两家相对初创公司的优势。

主持人:他说的是路易斯·巴罗索、乌尔斯·霍尔泽(Urs Hölzle)和帕塔(Parthasarathy Ranganathan)那本书,《数据中心即计算机》(The Datacenter as a Computer),出了好几版,讲的是多年来在这种巨大规模上建设所学到的全部教训。所以这个问题的答案是:确实是一道护城河,芯片之外还有太多东西。

软件护城河正被填平

主持人:那么 NVIDIA 和 Google 今天的竞争优势,多少在硬件,多少在软件,也就是编译器和库?在你们的组织里这是怎么运作的?换个问法:如果把网表(netlist)甚至版图给一家初创公司,但不给软件栈,他们能竞争吗?诺姆,你来。

朱皮:设计 TPU 时,我们努力遵循我博士导师的一条忠告。

主持人:好主意,这应该列为硬性要求。他的忠告是什么?

朱皮:能在编译时做的事,别拖到运行时。我们照着做了。这套架构是在一个十人的房间里设计的,其中两位来自编译器团队。我们努力造一台容易编译的机器,我认为这带来了很多好处。此外我们一直保持着同一套总体架构,比如内存的大小可以增减,这就像一台 PC,你可以多插几条内存,Word 照样能跑。所以它相当灵活。

主持人:库呢?

朱皮:框架很多,而且在演进,跟上它们、提供良好支持很重要。也有人喜欢自己写内核(kernel),这也得支持。

主持人:比尔,硬件和软件,你怎么看?

达利:产品是整个系统,包括硬件和软件,软件是不可分割的一部分。但对深度学习来说,软件问题在很多方面比它之前的通用 GPU 计算要容易。2006 年我们发布 CUDA 和 G80,启动 GPU 计算时,字面意义上有几千个应用能从并行中受益,但它们都是串行代码,很多是 Fortran 写的,移植负担极重,有的代码有十万行、百万行。于是我们在 CUDA 生态上投入巨大:CUDA 语言本身,底下的各种库,把 FFT 做好,把矩阵运算做好,等等。

而 2010 年我们在深度学习上的第一次尝试,是和斯坦福的吴恩达(Andrew Ng)合作开发软件,后来演变成 cuDNN。那时的感受是:哇,这个应用真小,里面真正要紧的只有两三个内核,把它们做快,整个程序就快了。看惯了百万行的天气预报代码,这简直是一股清新的风。和超算领域移植那些积灰旧程序的难题相比,它太容易了。

但它也变得极为关键。我们发现,一个深度学习模型跑起来之后,接下来六个月里我们能把性能翻一倍,因为里面有很多容易丢性能的地方。于是我们开始开发工具,做性能分析,做内核融合,做调优。所以软件是非常关键的一部分。至于它对今天的初创公司还是不是壁垒,我就不确定了,因为现在你只需要跟 Claude 说「我要一套软件栈」,然后出去过个周末。

主持人:这正是我下一个问题。认识程序员的人都听过他们的配偶讲,半夜里忽然大喊「天哪,你看它干了什么」。我们正在见证一件了不起的事:编码由机器而不是人来做。那么,软件护城河是否会随着 AI 的进步而消失?还是说这种看法过于乐观?

达利:我认为任何软件护城河,如今都依然有些优势,那些多年打磨的库,人们用起来顺手。但护城河已经严重退化了。只要有那个库的规范,放 Claude Code 去重建它,并不是多难的事。

朱皮:我同意,护城河在变浅。

达利:沙子正往里填,快能走过去了。

精度削到头之后

主持人:比尔刚才提到,而且他每次演讲都会承认,头几年我们靠降低精度榨出了大量速度。我自己就很震惊地发现,如今指数位比尾数位还多,这是相当新的东西。但精度只能降到这么低。等到没有位可削的时候,一代代的性能提升靠什么?比尔,你先。

达利:从 2012 年的 Kepler 开始,那是我们第一次把 AI 当作正经应用来看待的一代,十四年来我们基本做到了每年翻一倍。其中只有三倍来自工艺技术。这和九十年代微处理器的黄金时代形成对比,那时全靠工艺。在 NVIDIA,我们在数值格式上一路领先,从 Kepler 的 FP32 起步,到出货 NVFP4,Vera Rubin 里还有些新东西。显然还能再拧一两圈,但在数值精度上,我们已经接近尽头。还剩几招巧妙的办法,但还有很多别的轴可以继续创新:稀疏性、电路、局部性,乃至更好的模型,最终都能换来每瓦更多的 token。

不过我们确实到了低垂的果子都摘完的地步,得往树上爬得更高,去找那些挂在枝头的「2 倍」果子。但我们有一堆好主意,我认为至少能撑过接下来四五代,然后我就可以退休了。

主持人:四五代是四五年,还是八到十年?

朱皮:我不想暗示数值格式是唯一的手段,但它确实是低垂的果子之一。我们也有一些数值上的创新。杰夫·迪恩最初做那些 AI 实验的时候,他们用 FP32 计算,但存储时直接截掉低十六位。如果你是数值分析专家,听到「截断」就像指甲刮黑板,可它居然管用。所以我们开发 TPU 时,能用 BF16 直接运行原本在 CPU 上跑的程序,拿到相同的结果。

主持人:BF16 就是把 32 位砍掉一半的那种格式。

朱皮:对。这非常有力,因为我们不必花大量时间在系统层面调试「这个模型为什么不收敛」之类的问题。我们可以先验证运行正确,有了这个基础,再去采用 FP8、FP4 这些更小的格式。

三年硬件周期追不上三月模型

主持人:Google 有一个理论上的优势:既有推动 AI 前沿的人,又有造硬件的人。而我读到 NVIDIA 也要建立内部的模型专长。你们怎么看这种优势有多大?能接触到前沿研究者,与只用开源模型、没有这种紧密关系,差别有多大?诺姆,你先。

朱皮:排行榜显示一些开源模型表现相当好,所以专有模型要保持领先就得赛跑。我认为,专有版本总会有些优势,因为可以针对特定系统,无论 GPU 还是 TPU,把模型调得更好。

主持人:我想问的更多是:作为硬件人,你可以去问机器学习的人接下来会发生什么,该往硬件里放什么。如果接触不到这些专家,就只能猜。

朱皮:某种程度上是这样。问题在于,从最初的想法到能够批量制造的系统,需要两年半到三年,而机器学习的人每三个月就冒出一个新想法。所以你不可能真正为某个特定模型设计一台超级专用的机器,因为等造出来它已经变了。就像比尔前面说的,你只能把那些基本运算,矩阵运算、向量运算,做好。

主持人:既然造硬件要这么久,他们迭代又那么快,那你们为什么还要涉足自己造模型?

达利:我们不是「涉足」,我们做了很久了,内部有大量专长。我认为这里有个区分。像 OpenAI 这样的公司,他们刚在 Hot Chips 上发了篇很好的论文,讲他们的 Jalapeño 处理器,那基本上是为一个模型设计的芯片。他们知道矩阵运算和向量运算的相对比例,知道需要多少内存带宽。而如果你要支持市面上所有的模型,模型种类繁多,它们在内存带宽与算力之间的配比要求各不相同,注意力机制尤其如此。

注意力机制正变得非常有意思。这从 DeepSeek 推出 MLA 注意力开始,现在很多模型都有各自的混合注意力方案:三层状态空间(state space)接一层完整的 n 平方注意力,交替排列。很多人在做稀疏注意力:先快速过滤,再挑出前 k 个来关注。注意力机制的这些变化,对硬件提出了很不一样的要求。

如果你的定位是硅供应商,要支持所有模型,就得纵观全部模型,问:什么样的硬件能让所有人尽可能满意,又不让任何人特别不满意?这很难。你得决定侧重哪些、放弃哪些。而且如诺姆所说,得打提前量,瞄准鸭子飞行的前方。因为产品要两三年后才出来,其间人们会在模型上想出很多我们没见过的巧妙点子。我认为这就是好的计算机体系结构的本义:想清楚怎样让最多的人满意,预判应用需要什么,交付一个能满足它的产品。

朱皮:说到模型,我认为最近最具颠覆性的东西之一是混合专家(mixture of experts)。它其实七年前就提出来了,但它需要多得多的互连带宽。如果你想给用户快速响应,就得把延迟压下去。训练相对容易些,主要是带宽问题。而在计算机系统里,低延迟比高带宽难得多。

主持人:这就是最新的 TPU 采用蝶形(butterfly)互连配置的原因。

朱皮:是的。

每秒 token 数与基准之争

主持人:既然在计算机历史博物馆,我说说在 Hot Chips 上的感受:像是八十年代那场论战的重演。八十年代 RISC 处理器开始流行时,大家用 MIPS,也就是每秒百万条指令,做衡量指标,可指令怎么定义、跑什么程序,全由你自己说了算。RISC 公司就说,你可以信我们,别的混蛋都在撒谎。而在 Hot Chips 上,大家用的是每秒 token 数。跑的什么模型?多大?不管,就是每秒 token 数。

当年的解决办法是各家公司联合起来做基准测试,也就是 SPEC,有得有失,但大家都围绕它团结起来,后来发现哪怕你慢 10%,也可以有别的优势,效果相当不错。有人预见到今天这个问题,做了 MLPerf,已经存在好几年了。但在 Hot Chips 上我好像没听到有人引用 MLPerf。你们是否同意 MLPerf 不如 SPEC 成功?如果是,为什么?比尔,你来。

达利:我喜欢 MLPerf,它做了基准测试该做的事:一块公平的赛场,戳穿废话,说清楚「你在这个应用上到底跑得多好」。但它有两个问题,这就是你在 Hot Chips 上没怎么听到它的原因。第一,几乎没人提交。

主持人:这确实是个问题。

达利:因为工作量很大。我们每次都提交,每一代都跑全部基准,MLPerf 现在都到 7.1 之类的版本了。但很多初创公司会说:「太麻烦了,要是完全按规矩来,我们看起来就没那么好。」于是他们只挑一个结果放到幻灯片上讲。这是它没有流行起来的一个原因。第二,今天大家真正关心的是每瓦 token 数和每美元 token 数,而 MLPerf 的大多数基准并不针对这个,最接近的那个也没从正确的角度处理。

所以现在大家展示的,在 Hot Chips 上你也看到了一些幻灯片,是 SemiAnalysis 的 InferenceX 基准。它恰好给出人们想要的数据,而且做法很可复现。要在 SemiAnalysis 的图上拿到一个数据点,你得提供一个 GitHub 或 Hugging Face 的代码仓库,能加载、运行,在这套硬件、这个模型上展示每秒 token 数。

主持人:这是你自己跑,还是他们替你跑?他们在你的硬件上跑,然后给你一个数字?

达利:细节我记不清了,但它似乎正在流行起来。我认为现在人们用 InferenceX 比用 MLPerf 多。

朱皮:MLPerf 确实需要一个全职团队,还不是小团队。这是它的问题之一。

主持人:所以它太耗人力了,就像当年的 TPC 基准对数据库公司也是巨大的负担。又贵,又给不出你想要的数字。

达利:除此之外,它很棒。

AI 的社会影响与硬件人的责任

主持人:我得提醒音视频的同事,这个问题之后做一次现场投票。这是个更开放的问题:硅金热在技术之外,有着重大的经济、社会乃至地缘政治影响。你们认为 AI 硬件加速发展最深远的影响是什么?我们这些硬件架构师和这些公司的领导者,在负责任地塑造未来上承担什么责任?比尔,轮到你。

达利:我不会叫它硅金热,我会叫它 AI 淘金热,因为大家抢的是 AI。我认为它几乎在惠及我们生活的方方面面。看医学,早期应用之一是影像分析,现在它也用于诊断,帮医生诊断得更准。我们还可以有个人健康教练,看着你说:「我觉得你不该吃那块蛋糕,今天的热量超了。」

主持人:如果你不理这个教练呢?

达利:那是你的事,至少它给了建议。教育方面,我算是个「康复中的教育者」,每个学生都可以有一个个性化的导师,真正了解什么能激励他、他怎么学习,用合适的方式呈现材料,帮他跨过理解复杂概念的难关。工程方面,AI 已经在自动化我们的大量工作。我最近需要设计一件硬件,基本上就是写了份规范,放 Claude 去干。改了规范里的几处错误之后,它给出了一个相当不错的设计。

主持人:那些错误是它修的吗?

达利:不是,它精确地做出了我要求的东西,而我要求的东西不对。你得学会写好规范。但这在让各个领域的工程师往上走:他们不再做初级的计算,而是决定该造什么,管理一支代理「小兵」团队去执行。很多业务流程也在朝同一方向走,我们都变得高产得多。娱乐方面,它在帮我们创作精彩的作品,电影、游戏、音乐。所以 AI 的好处不胜枚举。

但反过来说,任何伟大的技术都能用来行善,也能用来作恶。AI 最明显的问题是深度伪造(deepfake)。我认为我们必须尽快拥抱内容溯源与认证:除非一张图片经过了恰当的认证,否则一律当它是假的。还有各种恶意用途的危险。网络安全最近很受关注,虽然大家爱谈 AI 那一面,但绝大多数成功的网络攻击,我想超过八成,都是人为操纵,也就是钓鱼,当然 AI 会让钓鱼更高明。AI 也可以用来设计病原体。而最大的风险是,当所有人都变得更高产,就业的性质会改变,某些工作需要更多人,另一些需要更少人。我们得想办法让从一种工作退出、进入另一种工作的人过渡得平顺些。

主持人:是的,怎么帮那些工作变了的人重新学技能。诺姆,你的看法?

朱皮:我认为最大的影响之一在科学。已经有很多极其惊人的成果,只是通常上不了大众媒体。Google DeepMind 把人类已知的所有蛋白质都折叠了。

主持人:还拿了诺贝尔奖。

朱皮:对,拿了诺贝尔奖。科学里还有很多别的例子。另一件事是,Google 几乎在所有的应用里都在用 AI,但很微妙,很多地方你根本注意不到。比如最早的应用之一是 Google 地图,那时是路易斯在管地图。原来的 Google 地图会说「行驶一千英尺,在丁尼生街右转」。把街景数据库整合进去以后,它现在说「行驶一千英尺,在壳牌加油站右转」。这对人友好得多,尤其在夜里路牌没有照明的时候。桑尼维尔的路牌是亮的,帕洛阿尔托就没有。这类微妙的东西很多,人们不会注意到,用一阵子就习以为常了。

能源是最大的瓶颈

主持人:请放投票结果。我的问题是:AI 芯片持续大规模部署的最大瓶颈是什么?选项有能源供应、半导体制造产能、高热芯片的散热、内存容量与带宽、AI 芯片的成本、不知道。数字还在变,但显然,眼下领先的是能源供应。你们可以自己回答这个问题,也可以不同意观众。你们认为最大的限制是什么?

朱皮:我正想谈能源。人们在谈一吉瓦、五吉瓦的数据中心,而输电线路的许可极难拿到,社区不喜欢大电线从自家房顶上过。所以很多超大规模云厂商和其他公司在转向现场发电。方式有几种,有的公司烧天然气发电,但如果选址得当,大部分电力可以来自风能和太阳能。这是我们努力采取的路线。

主持人:也就是无碳能源,不会加重气候负担。脱离电网,用无碳资源,就能建数据中心而不把别的东西砸垮。

达利:有个问题是数据中心需要稳定的电力,而太阳只在白天出来,还得没有云,风也只在刮的时候有用。不过总体上,我们看到各种供电来源之间平衡得还算合理,我认为这是自然的市场力量在配置资源。建数据中心需要三样东西:土地、电力、外壳,它们在推动需求。虽然很多人在看天然气和就地建发电机,但燃气轮机据我所知已经卖断到未来五年。有公司把飞机上的发动机拆下来改成发电机,让一大批乘客滞留在某地。

主持人:但如果我们真走天然气这条路,碳足迹会非常大。

达利:天然气常常是用来给可再生能源兜底的。他们会有太阳能和风能,但因为靠不住,而我们现在又没有足够的储能,虽然有些非常好的技术正在出现。

主持人:你认为锂电池不行?

达利:我认为锂电池每千瓦时的成本太高,撑不过二十到一百小时的缺口,而连续几天没有太阳的时候,你要弥合的正是这样的缺口。热储能电池看起来很有前途,已经有人在用它建数据中心,每千瓦时的成本便宜到足以填补这个缺口。

朱皮:不过如果你是面向消费者提供服务,需求有昼夜周期,凌晨三点不需要满功率。

达利:这关系到资本资产的利用率。你花了一百亿美元建数据中心,希望它全天候忙碌。如果这边的消费者都睡了,你就把 GPU 算力卖给地球另一边的消费者。

朱皮:那边延迟太大了。

达利:看用途。

朱皮:我们希望给用户快速响应。

主持人:核能也是无碳的,还有小型模块化反应堆。Google 在这方面肯定有过公告。你们对核能助力数据中心怎么看?

达利:我们在考察所有技术。我们有一个人,专职跟踪能源技术,为数据中心提供建议。核能看起来有前途,但一直很贵。看所谓平准化度电成本(levelized cost per kilowatt-hour),天然气很难被打败。哪怕你担心碳排放,天然气加碳封存,我认为在平准化度电成本上仍然胜过核能。我是抽水蓄能的忠实拥趸,可惜在我老家伊利诺伊州不太现实。

观众问:内存、量子、固化芯片

主持人:现在转到观众提问,观众可以投票排序。得票最高的是:解决 AI 硬件已知供应链障碍的最大挑战是什么?我理解是需求远超供给。

达利:最大的挑战是建晶圆厂需要很长时间。现在最痛的地方是内存。内存厂商很享受这种局面,同样的零件能收几倍的钱;如果他们供货充足反倒赚不到,所以建厂的动力也许没那么强。这永远是个预测需求的游戏:需求超出预测之后,你得等两三年才能把产能建起来。

朱皮:要是我以前更留心,某种程度上可以提前预见到这一点,因为 DRAM 的密度缩放基本停了。而计算机对内存的需求不断增长,即使不算 AI 对内存的需求,我认为天平也终于倾斜了。DRAM 制造过去是最糟糕的生意之一,今后相当长时间会是个好生意。

主持人:你之前教过我,SRAM 已经到平台期了。

朱皮:对,DRAM 也基本到了平台期,但需求没有。所以这会持续很久。

主持人:内存公司会很有意思,他们的历史就是过度建厂,然后价格崩盘,几十年都这样。他们会不会说:「我们利润这么高,现状有什么不好?」然后不动。

达利:幸好还有竞争。他们会想:「我们利润很高,还想多要点份额」,于是扩建晶圆厂,别家也会跟上。

主持人:但晶圆厂很贵,而且如你所说,不仅建得慢,良率爬坡也慢。我很期待你们对下一个问题的公开回答:量子计算成熟之后,对今晚讨论的话题影响有多大?

达利:我会说非常非常小。

主持人:真高兴是你在这里。

达利:我认为量子计算是一项伟大的技术。

主持人:这话是比尔·达利说的,不是戴维·帕特森。

达利:明天挨揍的是我。影响很小,原因是:真正好的量子算法大概只有两个。一个是用量子计算模拟量子化学,另一个是肖尔算法(Shor's algorithm),用量子计算分解两个大素数的乘积,从而破解大量现代密码。还有第三个,格罗弗算法(Grover's algorithm),用于优化,但只有平方级加速,不是指数级。这些算法都有个共同性质:它们利用的是量子计算机「计算量大、数据量小」的特点。人们最狂野的梦想是几千个纠错后的量子比特,然后借助叠加态的指数级加速,在几千个比特上做海量计算。AI 恰恰相反,是数据量大、计算量相对小的问题。所以量子计算对 AI 的训练和推理,都不会有可测量的影响。

主持人:关于未来的问题,我知道该找谁了。AI 走了七十年,很多想法都没成,最后成的是从数据里学习,也就是大数据的机器学习,这才是突破。

朱皮:我认为量子会起作用的地方不在计算,而在通信。上一届或上上届图灵奖……

主持人:你说的是诺贝尔奖,量子那个。

朱皮:对,量子传输与密码学。

主持人:但我以为量子传输是延迟很低、数据量很小。是这样吗?

朱皮:如果你只是发一个三百位长的密钥,并不需要高带宽。其余的内容加密后走普通信道就行。

达利:你的密钥得比三百位长。

主持人:下一个问题可能超出我们的专业范围:AI 硬件成熟对国家安全有什么影响?比尔,你做过这方面的工作吗?

达利:我不确定我理解这个问题。

主持人:好,我也不懂。下一个问到一家具体公司:你们怎么看 Etched 的架构,也就是把大语言模型直接刻进芯片?或者 Taalas 之类的公司。

朱皮:Taalas 是极端情况,他们打算用 ROM。这些架构可编程性更低,等于把模型造进硬件里。据我理解,Etched 起步时是这个论点,后来越来越可编程,现在……

主持人:他们在做产品吗?

朱皮:是的。

主持人:我好像在大众媒体上读到,Google 也在造某种这样的芯片。你愿意透露点秘密吗?

朱皮:什么都有人在研究。

主持人:好答案。

达利:这很有意思。如果你走 Taalas 的路线,走到极端,不只是为某个特定模型彻底优化,而是把一组权重烧进 ROM,因为 ROM 的密度大概是 SRAM 的四倍,顺带一提,比 DRAM 密度低得多。那你必须假定有某个模型变化得不快。可现在前沿模型每个月左右就出新版本,各种开源模型每两周就有一个巧妙的新想法,哪怕你让权重可编程,只要跟某个模型贴得太紧,就会被彻底打破。

朱皮:我认为我们还在早期,可编程性极其重要。

主持人:这大概正是你们说的:这是一个和 CPU 完全不同的世界。CPU 世界有百万行的遗留代码和几百万程序员,而这里代码量很小,程序员也不多,创新压力巨大。这不仅让硬件人容易创新,也让软件人和算法人容易创新并部署,所以变化率惊人。选一个模型,及时把芯片做出来,然后它的寿命有多长?这是个有意思的赌注,看它怎么发展会很有趣。

边缘计算、模拟计算与 AI 设计芯片

主持人:集中式计算和分布式计算的合理平衡在哪里?对边缘计算硬件有什么影响?换个说法,我们都在想数据中心,边缘呢?

达利:我们出货大量进入自动驾驶汽车的产品,某个级别以上的每辆奔驰都装了 NVIDIA 的处理器和软件,用于自动驾驶功能。这就是个例子:能放进数据中心的就放进去,因为在数据中心交付算力经济得多。但出于延迟的原因,或者你必须在网络断开时也能可靠运行,你可能需要在边缘计算。有时你要采集海量数据,全部传回数据中心不划算,就得在边缘先做归约。总的规则是:能在数据中心做就在数据中心做;如果这些因素之一限制了你,就得在边缘做,而这带来一套不同的需求。

朱皮:Google 的 Pixel 手机有 Edge TPU,处理相对简单的任务,比如语音识别,与那些大模型相比如今算简单的了。如果某件事对它来说太大,就得走网络,如果网络可用的话。

主持人:你和做 Waymo 计算机的人有交流吗?那是公司的另一部分吗?

朱皮:Google 归到 Alphabet 下面,Waymo 是 W。所以我们聊得不多。

主持人:既然浮点精度已经能达到相当的 AI 效果,模拟系统是否还会有取代数字系统的空间?你们肯定老被问这个:模拟做乘法多酷,那不是未来吗?

朱皮:如果你以造真实系统为生,有件事必须做:测试它对不对。如果一样东西是模拟的,每次结果不完全一样,测起来就非常难。探测晶圆时,你要发送测试序列,以极高的速度读回比特,而且必须逐位匹配。所以模拟即使在某些场合更高效,也有这个问题:答案可能不总一样,那你怎么知道芯片好不好?

达利:就算能测试,我也很少见到模拟真正占优的例子。我们不断重新评估这件事,因为那个论证非常诱人:你想做矩阵乘法,把激活值当电流输进去,再放一个电阻,把它看成电导,V 等于 I 乘 G,矩阵乘法白送了;顺带还能把很多结果加起来,也是白送,于是你得到了免费的计算。可它并不免费,天下没有免费的午餐,也没有免费的乘法,它总有代价。更糟的是,在模拟域你没有可靠的办法存储一个值。存很短的时间,可以放在电容上,但它会漏掉,尤其在现代半导体工艺里。要把它搬动任何距离,在模拟域都极其昂贵。

所以人们通常的做法是建一个小阵列,做存内计算,用模拟方式算激活值乘权重,把结果加起来取出来,然后必须做一次模数转换。看模数转换所需能量的基本物理极限,如果你要 8 位精度,它把你限制在每瓦五万亿次运算左右,而用传统数字技术,我们已经造出了每瓦一百万亿次运算的加速器。差了一个数量级还多,这是算上转换的。如果你能不转换,一直留在模拟域,也许能做出有吸引力的器件;但只要必须转换,你就输了。

朱皮:我最近没下楼去博物馆展区,但展览是从模拟计算机开始的,而它们在二战前后就不再被使用,是有原因的。

主持人:除了软件工具的热潮,我和年轻同事聊天时,也感受到用 AI 做硬件设计的兴奋。这个问题是:下一代芯片设计软件会是什么样?会和我们过去三十年用的截然不同吗?你们俩肯定都有看法。

朱皮:我认为芯片设计软件将是一个你与之对话的大语言模型。

主持人:这话很适合引用。「我要这样一块芯片。」那我明天该把 EDA 公司的股票卖了?

朱皮:他们有大量专长,也有把这些拼起来所需的很多部件。

主持人:所以得把那些知识装进大语言模型。

朱皮:而且模型得知道怎么运行工具,布局布线、验证,这些工具还是要跑的。

达利:我看设计芯片时人的时间花在哪里,答案是验证。大语言模型很擅长写测试,你说「确保这个能工作,覆盖所有边界情况」,它会给你写出一整套测试。而验证占了我们七成五的人力。

朱皮:我正想说同样的话。我们在 Hot Chips 上展示的结果,在设计本身上大约有 10% 的改进,但设计验证是团队里最大的部分,他们需要一切能得到的帮助。

达利:我们看到的不止 10%。普通工程师的生产力翻倍,真正优秀的工程师能到十倍。

朱皮:我说的是版图、电路,基本上是性能指标。

达利:哦,你说的是性能,不是人的生产力。我想起来了,那个基准研究是针对某项特定任务的,平均大约三倍。它让工程师高产得多。所以我认为这就是未来的设计工具。现在有一大批初创公司,又回到 AI 的那个道理:真正的淘金者是攻打垂直领域的人,现在有一大批初创公司在攻打 EDA 这个垂直领域。

主持人:再问一个:你们怎么看 CPU 与 GPU 的生态,能否为存内计算或受内存限制的负载腾出空间?或者说,让 CPU 离 GPU 更近,你肯定有话说。

达利:我们的 CPU 和 GPU 已经很近了。它们通过我们所谓的芯片间 NVLink 相连,内部叫 GRS,地参考信号(ground-referenced signaling),一种单端的极高速链路。你得问:靠近能得到什么?我们得到两样东西。第一,我们把 GPU 和 CPU 之间的链路配足,让挂在 Vera CPU 上的全部 LPDDR5 带宽,大概每秒 1.8 TB,我不背这些数字,都能路由到它连接的两块 GPU 中的任意一块,所以链路永远不是瓶颈。第二是延迟。从那块内存取数本来就是相对高延迟的操作,链路多加的那一点不算什么,而且只有一条链路,不经过交换机。然后是控制交互,基本上就是启动内核时需要的,那件事本身其他部分的延迟已经够大,把链路做得更快意义不大。顺便说,我们的 GPU 上有几十个小 RISC-V CPU,但它们都只做用户看不见的内务处理。至于用户编程的 CPU,芯片间 NVLink 的距离差不多正合适,是个恰当的平衡。

主持人:TPU 呢?

朱皮:长远看,配给加速器的 CPU 主要做内务和「带孩子」一类的杂务。现在真正有意思的是代理式计算(agentic computing)。你说「给我写个程序」,它得在某处编译,而你不会在一个矩阵乘法脉动阵列上编译。所以同一个数据中心里必须有真正的 CPU,配上正经的内存系统和正经的网络。

主持人:但在数据中心里它不必那么近,因为你只是发出一个编译任务。

朱皮:对,不必近,只要在同一个数据中心。我认为 CPU 真正的用武之地就在这里:编译,还有执行工具调用。

给下一代架构师的建议

主持人:我本来准备了一个收尾问题,观众也有个建议,都是给年轻人的忠告。这里有一条:「我十岁,为了准备未来,你们有什么建议?」我们家的孩子早过这个年纪了,我倒有这么大的孙辈。你们会对十岁的孩子说什么?

达利:学数学和科学。你需要一个好的根基,扎实的数学和基础科学功底,没有任何东西能替代。

主持人:诺姆呢?

朱皮:另一件重要的事是沟通能力。学会好好写作。我知道大语言模型可以替你写,但总有手边没有它的时候,你仍然需要把话说清楚。

主持人:对,想得清楚,说得清楚。我的最后一个问题:面对我们正在见证的惊人创新速度,你们会给今天进入这个领域、想在计算机体系结构上留下印记的年轻架构师和工程师什么建议?你们的职业生涯都留下了深刻的印记。回头看,是什么让你们做到的?你们会给这一代人什么忠告?

达利:今天的世界大不一样了。我职业生涯早期很快活,在贝尔实验室设计过一台机器,所有电路都是用铅笔画在牛皮纸上,由技师做版图。它第一次流片就全部工作,从来没做过仿真。但拥有出色的手工逻辑设计技能,我认为今天已经毫无价值了。

主持人:我很懂打孔卡。

达利:我会给有志于计算机体系结构的人三条建议。第一,精通一个垂直领域。第二,对计算机技术的理解要非常广。第三,让 AI 成为你的伙伴。第一条的理由是,今天的价值在于懂得该设计什么,而不是靠精湛的手工逻辑设计去实现它,因为工具会帮你,AI 会帮你,这就是让 AI 做伙伴。至于第二条,我见过很多计算机架构师把自己限制住了:他们去问搞内存的人,「能造这样的内存吗?」对方说「不行,做不到」。但如果你对电路设计有足够广的理解,就会意识到那个人想得不够开,这样的内存是能造的。我在贝尔实验室那台机器上就用了一种 3T DRAM,所有搞电路的都跟我说「那不行」。我去做了原型,说服了自己它能行。做最终芯片之前我们还先做了一块测试芯片来确认。所以你的理解必须足够广,需要时能自己做电路设计之类的事,跳出常规思维和你手头工具箱的限制。

还有,人们真的需要想清楚,在人与 AI 合作做计算机体系结构时,什么是独属于人的那部分,然后练好人能做的事。做 AI 能做的事没有意义,因为它只会做得更好。

主持人:但你得能看出它什么时候搞砸了,对吧?所以是不是得会手工做,才能认出它搞砸了?

达利:我肯定能看出它什么时候搞砸了,但也许正因为我会手工做。还有一招:让一个 AI 去检查另一个。

主持人:诺姆,你给下一代的建议?

朱皮:计算机体系结构有很长的历史,七十年甚至更久,一路上积累了大量教训。我认为最好的事情之一就是把这些历史教训都学一遍。有些在今天仍然适用,有些不适用。但没有必要去重新发明一样东西,如果它在过去七八次尝试里都没成功过。

主持人:就到这里,请马克重新上台。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:01 开场与两位嘉宾的平行生涯 ▶ 正在看
6:19 这次淘金热为何不同于以往 ▶ 正在看
10:05 GPU 与 TPU 是趋同还是分道 ▶ 正在看
17:35 护城河在系统集成而非芯片 ▶ 正在看
20:45 软件护城河正被 AI 填平 ▶ 正在看
25:54 精度削到头后靠什么翻倍 ▶ 正在看
29:29 三年硬件周期追不上三月模型 ▶ 正在看
34:29 每秒 token 数与基准之争 ▶ 正在看
38:10 AI 的社会影响与硬件人的责任 ▶ 正在看
43:56 能源是最大瓶颈 ▶ 正在看
49:22 观众问:内存、量子、固化芯片 ▶ 正在看
57:05 边缘计算、模拟计算与 AI 设计芯片 ▶ 正在看
67:54 给下一代架构师的建议 ▶ 正在看
本期小问 · 档案清单
10:05 英伟达和谷歌的 AI 芯片,是设计者互相靠拢,还是淘汰后的幸存者相似? ▶ 正在看
20:45 AI 一个周末能重写软件库,CUDA 生态还算护城河吗? ▶ 正在看
29:29 硬件要三年、模型三个月一变,芯片该为哪个模型设计? ▶ 正在看
67:54 AI 能做的它都做得更好,工程师还要会手工做吗? ▶ 正在看
本期讲者
比尔·戴利NVIDIA 首席科学家兼高级副总裁(2009 年至今),曾任斯坦福计算机科学系主任。互连网络与流处理架构的奠基人,虫洞路由发明者,合著《互连网络的原理与实践》。
诺姆·朱皮Google Fellow 兼副总裁,TPU 系列的首席架构师。斯坦福 MIPS 项目参与者,曾在 DEC 西部研究实验室和惠普实验室从事存储层次结构研究,发明了 victim cache。
Dave Patterson加州大学伯克利分校荣休教授、Google 杰出工程师,RISC 与 RAID 的主要提出者,2017 年图灵奖得主。本场担任提问者。
01开场与两位嘉宾的平行生涯
0:01
[music]
[音乐]
便签笔记
0:14
Hello. Hello. How's everybody doing tonight? >> Yeah. Yeah. Great to see everybody in person. Great to see everybody online. I'm Mark Edkin. I'm the CEO of the Computer History Museum. One of our goals at the museum is to have legends and leaders on this stage talking about the most important topics in computing. And that's what we're doing tonight. We have three leaders uh who are leading in the future of chip design for artificial intelligence. You could say we're in a Cambrian explosion of chip design right now. We're going to hear from the leaders of that revolution tonight. Um, I want to thank our sponsors. Uh, it's Mark and Mary Stevens. Uh, they are sponsoring these great discussions, not just tonight, but they sponsored a whole series of discussions. Uh, and really thankful for them for their leadership.
大家好,大家好。今晚大家都还好吗?>> 很好,很好。很高兴能和大家在现场见面,也很高兴见到线上的各位。我是 Mark Edkin,计算机历史博物馆的 CEO。我们博物馆的目标之一,就是让业界的传奇人物和领军者站上这个舞台,探讨计算领域最重要的话题。今晚我们要做的正是这件事。我们请到了三位领军者,他们正在引领人工智能芯片设计的未来。可以说,我们现在正处在芯片设计的寒武纪大爆发之中。今晚我们将听到这场革命的引领者们分享他们的见解。另外,我要感谢我们的赞助商,Mark 和 Mary Stevens。他们赞助这些精彩的讨论,不只是今晚这一场,他们赞助了整整一个系列的讨论。呃,真的非常感谢他们的引领和支持。
便签笔记
1:07
And then lastly, I want to introduce the leader of tonight's event, uh, Dave Patterson. Dave is the party professor of computer scientist ameritus, computer science ameritus at University of California, Berkeley. Uh he's also a Google distinguished engineer. Uh you probably know Dave as the leader of the risk project. We know Dave here at the museum as a huge friend of the museum. Uh he has been a fellow which is our most distinguished honor. Uh and he has shared his collection with us. Uh and so we have some original pieces that are important for our collection to help educate the next generation. And most importantly, he has been a friend to our curators. uh and he has been advising us on our collection and our storytelling um and he's a frequent uh guest on our stage. So, please welcome Dave Patterson. [applause] >> Thank you.
最后,我想介绍一下今晚活动的主讲人,呃,Dave Patterson。Dave 是加州大学伯克利分校的计算机科学荣休教授,呃,他同时也是谷歌的杰出工程师。呃,大家可能都知道 Dave 是 RISC 项目的带头人。我们博物馆这边都知道 Dave是博物馆的大朋友。呃,他获得过我们最高的荣誉——博物馆院士。呃,他还把自己的收藏分享给了我们。呃,所以我们才有一些对馆藏很重要的原始藏品,可以用来教育下一代。而最更重要的是,他一直是我们策展人的朋友。呃,他一直在为我们的馆藏和叙事方式提供建议嗯,他也是我们舞台上的常客。所以,请欢迎 Dave Patterson。[掌声] >> 谢谢。
便签笔记
2:09
>> Thanks everybody for showing up. Uh I get to introduce two of my friends. Uh, Bill, Norm and I have known each other since the 1980s when I was an assistant professor at UC Berkeley and uh, Bill was a grad student, a PhD student at Caltech and Norm at Stanford University. Uh, Bill Deli has been building network connected computers ever since Caltech. Uh, after graduation he joined MIT where he built a couple of more network supercomputers or network computers. Uh, after about a decade he came back to Stanford University.
>> 谢谢大家的到来。呃,接下来由我来介绍我的两位朋友。呃,Bill、Norm 和我从 1980 年代就认识了,当时我是加州大学伯克利分校的助理教授,而 Bill是加州理工学院的研究生,在读博士,Norm 则在斯坦福大学。呃,Bill Dally 从在加州理工时期开始,就一直在做网络互连的计算机。呃,毕业之后他去了 MIT,在那里又做了几台网络超级计算机,或者说网络计算机。呃,大约十年后,他回到了斯坦福大学。
便签笔记
2:43
uh where he eventually become chair of the computer science department at Stanford. In about a decade later in 2009, he became the chief scientist at Nvidia. That was 16 years ago. He still holds that position, but now he is also a senior vice president. I was trying to figure out how to figure out of all of his contributions what's to pick. So right thing to do is ask AI, right? So the AI chatbot according to the AI chatbots is most significant contributions are a few but it's wormhole routing and related flow control techniques and it's a practical low latency switching method for chip and system level networks and he's also wrote a book with his friend Brian Tols principles and practices for interconnection networks. This is kind of the bible of networking uh which systemized topology routing and flow control design and it shaped both academic uh research and industrial research for decades. Also at Stanford his imagin and Marrammac projects kind of made a very strong case for stream
呃,后来他成了斯坦福计算机科学系的系主任。又过了大约十年,在 2009 年,他成为了Nvidia 的首席科学家。那是 16 年前的事了。他至今仍担任这个职位,但现在他同时也是高级副总裁。我一直在想,他有这么多贡献,到底该挑哪些来讲。所以正确的做法当然是问 AI,对吧?那么根据 AI 聊天机器人的说法,他最重要的贡献有这么几项:一个是虫洞路由(wormhole routing)以及相关的流控技术,这是一种实用的低延迟交换方法,适用于芯片级和系统级的网络;他还和他的朋友 Brian Towles 合著了一本书《互连网络的原理与实践》。这本书堪称网络领域的圣经,它把拓扑、路由和流控设计系统化了,并且在此后几十年里塑造了学术界和工业界的研究。同样是在斯坦福他的 Imagine 和 Merrimac 项目在很大程度上为流处理(stream processing)提供了有力的论证,而流处理是现代 GPU 计算的一个重要概念先驱。
便签笔记
3:50
processing which was an important conceptual precursor to a GPU modern GPU computing. Norm when he was a graduate student at Stanford worked on the MIPS project which is a pioneering risk architecture. After graduation he joined digital equipment western research labs where he stayed for about a decade and then he went to HP uh labs. After that in 2013 his friend from WRL Jeff Deink approached him to try and go to Google to build hardware for this new deep learning machine learning ideas. Norm was very skeptical given all the hyper uh all the hype about AI in the past but Jeff over kind of convinced him to come because he said this deep learning it works on everything we try. So he joined 2013 he is now both a fellow and a vice president at Google. I asked AI about Norm's contributions and they mentioned the TPUs for his work at Google and then his work on memory hierarchy design at deck and HP particularly uh victim caches and prefetch buffers. As I prepared this Norm and Bill have these
计算。Norm 在斯坦福读研究生时参与了 MIPS 项目,那是一个开创性的 RISC 架构。毕业后他加入了 DEC 的西部研究实验室(WRL),在那里待了大约十年,然后去了惠普实验室。之后在 2013 年,他在 WRL 时的朋友 Jeff Dean 找到他,想让他去谷歌为这些新的深度学习、机器学习理念打造硬件。考虑到过去 AI 领域的各种炒作,Norm 当时非常怀疑,但Jeff 还是说服他来了,因为他说这个深度学习,我们试什么它都管用。所以他在2013 年加入,现在他既是谷歌的 Fellow,也是副总裁。我问了 AI 关于 Norm 的贡献,它们提到了他在谷歌做的 TPU,还有他在 DEC 和惠普做的存储层次结构设计,特别是victim cache(牺牲缓存)和预取缓冲区。在我准备这次活动时,Norm 和 Bill 的人生影响力惊人,而且高度并行,有大约十个特征说明他们
便签笔记
5:03
incredible impact and highly parallel lives so by 10 features that they're have similar careers. So the first is uh they got their bachelor's, masters and PhDs at three different institutions. I was just a stay-at-home stayed at the same one. Secondly, they both got graduate degrees from Stanford. Uh masters for Bill and a PhD for Norm. Both of them were supervised by John Hennessy. So who I think is uh here tonight. Uh both are vice presidents at their respective companies. Both are fellows of ACM. Both are fellows of E.
有着相似的职业生涯。第一点是,他们的学士、硕士和博士学位是在三个不同的机构拿的。我就只是待在家,一直待在同一所学校。第二,他们都在斯坦福拿了研究生学位。Bill 是硕士,Norm 是博士。两人的导师都是 John Hennessy。我想他今晚也在这里。两人都是各自公司的副总裁。两人都是 ACM Fellow。两人都是 IEEE Fellow。
便签笔记
5:40
Both are fellows of the American uh association for the advancement of sciences. Uh both were elected to the national academy of engineering. A very prestigious honor. Both won the ITE Seymour Cray computer science and engineering award and both won the ACMI Eert Mockley award which is the highest award in computer architecture. So they took they have very busy lives but they took some time out to be with us. So, I'd like to let's welcome them with a round of applause for their contributions and for being here.
两人都是美国科学促进会(AAAS)的 Fellow。两人都当选了美国国家工程院院士,这是非常崇高的荣誉。两人都获得了 IEEE Seymour Cray 计算机科学与工程奖,两人也都获得了 ACM Eckert-Mauchly 奖,那是计算机体系结构领域的最高奖项。所以他们的生活非常忙碌,但还是抽出时间来跟我们相聚。那么,让我们以热烈的掌声欢迎他们,感谢他们的贡献,也感谢他们今天到场。
便签笔记
6:12
[applause]
[掌声]
便签笔记
02这次淘金热为何不同于以往
6:19
[applause] >> Great. Thanks, Dave. [applause] So, uh for about 30 or 40 minutes, I'll do the we'll do a question. I'll do the questions, but after that, we'll uh open up to audience questions. So the title is Silicon Gold Rush which suggests a period of intense innovation, investment and competition. From your perspectives at NVIDIA and Google, what defines this gold rush era for AA hardware and what makes it fundamentally different from previous eras of computer architecture development? And start with Bill. Yeah.
[掌声] >> 太好了。谢谢你,Dave。[掌声] 那么,接下来大约三四十分钟,我们会做一个问答。我来提问题,之后我们会开放给现场观众提问。今天的主题是「硅金热潮」,这暗示着一个充满密集创新、投资与竞争的时期。从你们在英伟达和谷歌的视角看,是什么定义了 AI 硬件的这场淘金热时代?又是什么让它与以往计算机体系结构发展的各个时代有着根本性的不同?我们先从 Bill 开始。好的。
便签笔记
6:53
So I think there are three characteristics that really define this one and make it very different from some in the past and it's really intense economic demand there. you know, AI is working on everything they apply it to, like Jeff Dean said. And as a result, the demand for um you know, more more tokens, more flops, you know, more AI is insatiable. At the same time, the application is rapidly evolving but relatively simple compared to the big supercomputing applications we had with millions of lines of code. A transformer is a relatively simple thing. And so, it's easy to look at what you have to build to make that go fast. and and also people are willing to evolve very quickly. There are no dusty decks. If you want to look at contrast in the early 90s there was sort of a computer architecture gold rush around supercomputing. It turns out there was no intense economic demand. It was fueled by the the the DARPA strategic computing program. They gave a bunch of funding, you know, money made people
我认为有三个特征真正定义了这一次,并让它与过去的某些时期非常不同,那就是极其强烈的经济需求。你知道,就像 Jeff Dean 说的,AI 用在什么上都管用。因此,对更多 token、更多算力(flops)、更多 AI 的需求是永无止境的。与此同时,应用本身在快速演进,但相对简单——比起我们过去那些动辄几百万行代码的大型超算应用来说。Transformer 是个相对简单的东西。所以,很容易看清为了让它跑得快,你需要构建些什么。而且大家也愿意非常快地做出改变。没有那些陈年老代码。如果你想做个对比,回到90年代初,围绕超级计算曾经有过一场计算机体系结构的淘金热。结果发现并没有强烈的经济需求。它是被DARPA的战略计算计划推动起来的。他们给了一大笔资金,你知道,钱让人们以为下游有个巨大的
便签笔记
7:49
think that there was a big downstream market which there wasn't. Lots of people founded companies which most wound up going bust. um things like thinking machines and and and the like. Um the applications then were dusty decks with thousands of lines of in some case close to a million lines of code and you could accelerate you know what appeared to be the important kernel but you know Anvil's law would come up and bite you because it was the other 99% of of the code that wasn't accelerated. So the ease of actually getting things accelerated with a simple you know um application that you can accelerate a few primitives and get really good results um real economic demand I think make this one um this one stick. This is not one of these you know like the you know list machine frenzy of the 80s or the supercomputing frenzy of the 90s.
市场,其实并没有。很多人创办了公司,大多数最后都倒闭了。嗯,比如Thinking Machines之类的。嗯,当时的应用都是些陈年老代码,几千行,有些甚至接近一百万行代码,你可以加速那些看起来很重要的核心计算,但阿姆达尔定律会跳出来咬你一口,因为问题出在其余那99%没被加速的代码上。所以,用一个简单的应用就能轻松把东西加速起来,你只要加速几个基本原语就能得到非常好的结果,嗯,再加上真实的经济需求,我认为这些让这一次能站得住脚。这不是那种,你知道,就像80年代的Lisp机器狂热,或者90年代的超算狂热。
便签笔记
8:35
This is a real value delivering uh >> so so actual market demand for computers. Norm what what's your take? Yeah, I I'm kind of a fan of a glo keeping attention of what's happening in the global economy and the amount of money that's being invested is just uh in incredible. So, you know, for a while New York had a a leaky tunnel that was hundred years old uh going to New Jersey and they they were a billion dollars short and then you know like Google has announced they're going to have $105 billion worth of capex next year and other companies are investing similar amounts. So, it's uh yeah, there's going to be winners and and losers, and everyone wants to to race as fast as possible.
这一次是真正在创造价值的。>> 所以说,是对计算机的真实市场需求。Norm,你怎么看?是啊,我算是挺喜欢关注全球经济动向的,而现在投进去的钱简直是难以置信。你知道,有一阵子纽约有条通往新泽西的隧道,已经一百年了还在漏水,他们的资金缺口是十亿美元;然后你看,Google宣布明年资本支出要达到1050亿美元,其他公司的投入也差不多。所以,嗯,会有赢家,也会有输家,而且每个人都想跑得越快越好。
便签笔记
9:32
Yeah. Well, you were you guys were making an analogy to the actual gold rush. So, who who's the who are the prospectors and who are the sellers of the pickaxes here? >> Well, the prospectors are the people who are, you know, finding verticals, right? And if you find the right vertical to apply AI to, I think you can strike it rich just like the prospectors did. Um, you know, Nvidia and and you know, companies like Google that make the infrastructure. We're Leel and Stanford, right? We're selling the the picks and and [clears throat] shovels to the prospectors and whether they strike a rich or not, we're going to make money.
是啊。你们刚才拿真正的淘金热做类比。那么这里谁是淘金者,谁是卖镐头的呢?>> 嗯,淘金者就是那些,你知道,在找垂直领域的人,对吧?如果你找对了应用AI的垂直领域,我觉得你可以像当年的淘金者一样一夜暴富。嗯,你知道,英伟达,还有像Google这样做基础设施的公司。我们在Leel和斯坦福,对吧?我们是在向淘金者卖镐头和铁锹,不管他们有没有挖到金子,我们都能赚到钱。
便签笔记
03GPU 与 TPU 是趋同还是分道
10:05
>> Okay. [laughter] Well, well, [laughter] we're still both paying the memory companies a lot of money. [laughter] >> Yeah. Yeah. Yeah. I think it was Micron made as much money in the last quarter as they had made in the 19 years previously. So, >> it's interesting seeing their revenue go up while the number of parts sold stays level. >> Exactly. Yeah. The margins [laughter] through the roof. Yeah. So an unnamed colleague of mine has accused training accelerators of convergent evolution. So that's a biological jargon for uh independent evolution of similar features and species of different uh lineages. So his convergent his claim that you guys have a convergent design is a max radical compute die with a large systolic matrix unit that's you pack as many HBMs that fit around the outside and the perimeter and then you use the fastest 30s you can to do a custom link. I I I expect you don't agree with that and so why is he wrong and that GPUs and TPUs have divergence approaches and is there anything you
>> 好吧。[笑声] 不过,[笑声] 我们俩还是得给内存厂商付一大笔钱。[笑声]>> 是啊。是啊。是啊。我记得美光上个季度赚的钱,跟他们过去19年赚的一样多。所以,>> 有意思的是,他们的营收在涨,而卖出的芯片数量却没变。>> 正是如此。是啊。利润率[笑声]高得离谱。是啊。所以,我一位不便点名的同事指责训练加速器是趋同进化。这是个生物学术语,指的是不同物种独立演化出相似的特征,嗯……谱系。所以他的这个「趋同」说法——他说你们的设计是趋同的——就是一个算力拉满的芯片,配上一个巨大的脉动矩阵单元,然后在外围周边尽可能多地堆 HBM,接着用你能做到的最快的 SerDes 去做一条定制互联。我猜你应该不同意这个说法,那么他哪里说错了?GPU 和 TPU 其实是分道扬镳的路线吗?另外,对方的设计里有没有你欣赏的地方?[笑声] >> 嗯,通常我先来说吧。关于寒武纪大爆发,
便签笔记
11:13
like about the other design? [laughter] >> Well, um normally I I'll start. Uh so one has to remember about the Cambrian explosion. There was also uh the implosion that happened afterwards. So there were all these weird and wonderful creatures, you know, that look like giant pill bugs but with like weird appendages sticking out. And they they didn't survive uh the extinction event. And so I I think this the same thing is true. There's been many startups with many different ideas, but uh I I think the reason why Nvidia and Google have both been successful is that a bunch of these other ideas uh weren't as capable and so you know they may go extinct.
有一点要记住:之后还发生过一次大崩塌。当时出现了各种稀奇古怪的生物,比如像巨型潮虫、身上还长着奇怪附肢的东西。它们没能挺过那次灭绝事件。所以我觉得这里也是一样的。曾经有很多创业公司、很多不同的思路,但我认为英伟达和谷歌之所以都成功了,原因在于其他那些思路的能力没那么强,所以你知道,它们可能就灭绝了。
便签笔记
12:14
>> So that that's why it converged. You both had good taste in what you chose to >> Yes. But yeah, I don't know that we've even converged. Um >> Okay. >> I think I think TPUs and GPUs look look very different. I guess that >> Okay, great. >> at a 20,000 foot view. Yes. You know, the application drives certain requirements, right? You have gems, you're going to need some matrix multiply units. You have, you know, softmax and and norms, you're going to need some vector units and and things that can do transcendental functions.
>> 所以这就是趋同的原因。你们两家在选择上都很有品味 >> 是的。不过话说回来,我其实不确定我们真的趋同了。嗯 >> 好的。>> 我觉得 TPU 和 GPU 看起来非常不一样。我想这个 >> 好,很好。>> 从两万英尺的高度看,是的。你知道,应用会带来某些确定的需求,对吧?你要做 GEMM,就需要一些矩阵乘法单元。你要做softmax 和各种 norm,就需要一些向量单元,以及能算超越函数的东西。
便签笔记
12:43
You need a certain amount of memory capacity, a certain amount of memory bandwidth. need a certain amount of communication bandwidth, but there's a lot of nuance, right? That's what the application demands. Everything is going to have that, even though weird bug-like creatures are going to go extinct. Um, but there's a lot of nuances in how they're how they're combined. For, you know, many years, Nvidia, I think, kind of led the way with improved numericics. you know, we um we published a paper in MLS a few years ago and talking about vector scaling and out of that came NVFP4 and then a few years later all the MX FP formats kind of kind of followed suit. Um and and that that's an important nuance. It can give you a 2x advantage over somebody um who who was doing just FP8 before to get the same level of accuracy. Similarly with sparity sonhan and I wrote a paper in 2015 in nurups about you know you know most neural networks are naturally very sparse and then we put hardware support for sparsity starting in our amper
你需要一定的内存容量、一定的内存带宽,还需要一定的通信带宽,但这里面有很多细微之处,对吧?这些是应用提出的要求。所有东西都得具备这些,尽管那些奇怪的虫子一样的生物最终会灭绝。嗯,但在它们如何组合这一点上,有大量的细节差异。你知道,很多年来,我认为英伟达在数值格式的改进上算是走在前面的。你知道,我们几年前在 MLSys 上发过一篇论文,讲向量缩放,由此产生了NVFP4,然后几年之后,各种 MX FP 格式也都跟进了。嗯,这是一个很重要的细节。相比之前只用 FP8 的人,它能在同样的精度水平下给你带来 2 倍的优势。稀疏性也类似,我和宋涵在 2015 年 NeurIPS 上写过一篇论文,讲的是大多数神经网络天然就非常稀疏,然后从 Ampere 这一代开始,我们在硬件里加入了对稀疏性的支持,这是你在
便签笔记
13:34
generation that's something you don't find in every one of these bug-like creatures there's a lot of a lot of nuance down the way now I think Google has a big advantage in in designing their TPUs and up until recently they had one customer Google and so they could decide exactly what they wanted to do and do it you know and Nvidia is fortunate that we have a 68% share of both the inference and and and training market. But what comes with that is we have lots of customers and we have to keep them all happy, which means they come and ask for feature feature requests and if you're a big enough customer, you have to at least kind of entertain it a little bit, make it seem like you um >> so you have you have to do a lot of travel and talk to them >> and and you you have to, you know, put things into your hardware to make all these different different people happy.
那些虫子一样的生物身上找不到的东西。一路下来有大量细节差异。现在我认为谷歌在设计TPU 上有一个很大的优势,直到最近,他们都只有一个客户,就是谷歌,所以他们可以完全按自己想要的方式来决定并去做,你知道。而英伟达很幸运,我们在推理和训练两个市场都占有 68% 的份额。但随之而来的是,我们有大量客户,我们得让他们都满意,这就意味着他们会跑来提各种功能需求,如果这个客户足够大,你至少得稍微认真对待一下,让他们觉得你有在 >> 所以你得到处出差跟他们聊 >> 而且你还得把一些东西放进你的硬件里,去让这些不同的人都满意。
便签笔记
14:14
And I think um that that drives certain um certain things in the machine that if we if we could just build exactly what we wanted just for ourselves, not to listen to customers, we might do better. So I'm a little bit envious of [clears throat] the TPU's ability to u um to do that. >> So So you admire their small customer base. >> Yeah. Yeah. Um I I I do think that um you know, [laughter] one thing that they have is they have they have a they have a pretty neat interconnection network with the uh with the training TPUs of the 3D Taurus. I built a lot of supercomputers, you know, both at MIT and with Cray in the '9s with 3D Taurus networks and so I'm very fond of that and I can see the hand of the co-author of my book, Brian Tools work in that network >> and we studied that book very carefully.
我觉得,这会给机器带来某些特定的东西,如果我们能完全按自己想要的样子来造,只为我们自己造,不去听客户的,我们也许能做得更好。所以我有点羡慕 [清嗓子] TPU 能这么做的能力。>> 所以你羡慕他们客户少。>> 是啊。是的。嗯,我确实觉得,你知道,[笑声] 他们有一点是,他们有一个相当漂亮的互联网络,就是训练用 TPU 的那个 3D 环面。我造过很多超级计算机,你知道,在 MIT 造过,90 年代在 Cray 也造过,都是 3D 环面网络,所以我对它很有感情,我也能在那个网络里看到我那本书的合著者 Brian Towles 的手笔 >> 我们把那本书研究得很仔细。
便签笔记
14:56
So [laughter] >> in terms of the Cambrian explosion though, I like liked your analogy there. You know, there was a similar Cambrian explosion with graphics ships. Um >> I don't remember when exactly there was a sort of late 90s early 2000s. There were no fewer than a hundred graphics chip startups in Silicon Valley and you know after you know you know whatever the weeding out of the not fit you know the survival of the fittest there were two remaining it was Nvidia and ATI which wound up getting getting acquired by AMD and I think something very similar um is is the likely outcome of the current explosion.
所以 [笑声] >> 不过说到寒武纪大爆发,我挺喜欢你那个类比的。你知道,图形芯片也经历过一次类似的寒武纪大爆发。嗯 >> 我不太记得具体是什么时候,大概是90 年代末、2000 年代初吧。当时硅谷至少有一百家图形芯片创业公司,你知道,等到那些不适应者被淘汰,你知道,适者生存之后,就只剩下两家,英伟达和 ATI,ATI 后来被 AMD 收购了。我认为这次大爆发很可能也会是非常类似的结局。
便签笔记
15:31
Okay. Uh Norman, he he likes the and I agree. I think the optical interconnect uh that Google I of course I'm [clears throat] biased because I get a paycheck from Google, but I think the optical interconnect is a very cool feature. What what about the Nvidia uh architectures you find attractive [laughter] or or the customer base or what? [laughter] Uh I think Bill said it well in that you know there's certain operations that we both have to support you know matrix multiply vector operations. Um, one of the things that was uh different about TPUs is I mean our first one was a PCI card but and just did inference but starting with the second one they they were designed to be supercomputers. that we had the the tourist network as in in Bill's book uh with Brian and uh so we didn't have any painful steps along the way of adding features and having other features kind of conflict with those features and stuff. So we were able to start out with a clean sheet of paper um with a a supercomputer design.
好。呃,Norm,他喜欢这个,我也同意。我觉得光互联,谷歌的——当然我是有偏向的,因为我拿的是谷歌的工资——但我觉得光互联是一个非常酷的特性。那英伟达的架构里,有什么是你觉得有吸引力的[笑声]?或者是它的客户群,还是什么?[笑声] 呃,我觉得 Bill 说得挺好的,就是有些运算我们两边都必须支持,你知道,矩阵乘法、向量运算。嗯,TPU 有一点不太一样的是,我们的第一代其实是一块 PCI 卡,只做推理,但从第二代开始,它们就是按超级计算机来设计的。我们有了Bill 书里讲的那种环面网络,跟 Brian 一起,所以我们一路上没有经历过那种加了新功能、结果和其他功能互相冲突之类的痛苦。所以我们能够从一张白纸开始,直接做一个超级计算机设计。
便签笔记
16:59
Yeah, I think that was a precient that the fact that your second design in 2017, you realized that you needed a supercomputer for training and that's what you were going to build right from the start. >> Yeah. >> And it was it it was and liquid cooled in 2018 or so. >> Yeah. Yeah. So, we've been liquid cooled for eight years. >> Yeah. So, and that's what a lot of these look like. Okay. Um, and by the way, I'd like to thank the people at Hot Ships, which happened right here in Silicon Valley a couple days ago. I went around and asked for questions to to give Bill and so I got some really goodness. Okay.
是啊,我觉得这很有前瞻性,就是你们 2017 年的第二代设计,你们就意识到训练需要的是一台超级计算机,而这正是你们从一开始就打算造的东西。>> 是的。>> 而且它在 2018 年左右就是液冷的了。>> 是的,是的。所以我们做液冷已经八年了。>> 是啊。所以现在很多东西都长这样了。好的。嗯,顺便,我想感谢 Hot Chips 的各位,就发生在几天前的硅谷这里。我到处去征集问题,好拿来问 Bill,结果收到了一些非常棒的问题。好的。
便签笔记
04护城河在系统集成而非芯片
17:35
Uh, Nvidia and Google are both worldclass system companies. Has the engineering of systems, the interconnect, the memory, power, cooling, and packaging become an even bigger moat for competitors than the silicon itself? And how do your approaches to interconnectivity address the critical bottlenecks in scaling? Where are they similar and different? I I think Bill, it's your turn to go first. >> Okay. So, I I think the product is the whole system. That's, you know, not just the GPU or even the board with GPUs and CPUs and networking gear. It's all the hardware, all the software and the configuration that makes it all work well together. Starting I think in our Pascal generation I think it was about 2015 we started offering something we called a DGX super pod which was if you bought this this is even before we had all the captive large scale networking we had our own internal you know MB link networking but we we hadn't yet acquired Melanox >> um and but we you picked you'll use this switch you'll configure it like this and
呃,英伟达和谷歌都是世界级的系统公司。系统工程——互连、内存、供电、散热和封装——是否已经成为比芯片本身更大的护城河?你们各自在互连上的思路,又是如何应对扩展中的关键瓶颈的?它们有哪些相同、哪些不同?我想,Bill,先请你来吧。>> 好的。我认为产品就是整个系统。也就是说,不只是 GPU,甚至也不只是那块装着 GPU、CPU 和网络设备的板子。它是全部硬件、全部软件,以及让这一切良好协同的配置。我想是从我们的Pascal 那一代开始,大概 2015 年,我们推出了叫 DGX SuperPOD 的东西。当时如果你买了它——这甚至还在我们拥有自研大规模网络之前,我们有自己内部的 NVLink网络,但那时还没有收购 Mellanox——>> 呃,但我们会告诉你:选这款交换机,按这样配置,
便签笔记
18:32
you do that and if you did that you could put together you know a collection of in that period of time maybe 10,000 GPUs and you would turn it on and it would work and in contrast you know, when we would build the big um DOE supercomputers like when we did um you know, Titan in uh um you know, what was that, you know, 2011 or 2012 and and Summit and Sierra, you would typically go through a six-month period after you had all the hardware working bringing it up because of little things in network configuration, configuration of the different things. So, you're delivering the entire product, you're delivering a system, it has to work reliably um over long training jobs with very high availability. And so, you know, I think there's a lot of systems expertise that goes into delivering that and having a standardized configuration for us being because we're selling into a lot of different data centers where people did things different ways was critical to get the system up and running quickly and and run it um
照着做。如果你照做了,就能搭起——在那个时期大概是一万块GPU 的集群,开机就能跑。相比之下,当我们建那些大型的 DOE 超级计算机时,比如我们做Titan,那大概是……2011 还是 2012 年?还有 Summit 和 Sierra,通常在硬件全部就位之后,你还得再花半年时间去把系统调起来,就因为网络配置、各种东西的配置上那些琐碎的小问题。所以,你交付的是整个产品,你交付的是一套系统,它必须在长时间的训练任务中稳定运行,可用性要非常高。所以我认为,要交付这样的东西,背后需要大量的系统专业知识。而对我们来说,有一套标准化配置至关重要,因为我们卖给的是很多不同的数据中心,各家做法都不一样,标准化才能让系统快速上线并
便签笔记
19:26
extremely reliably. Now, you know, where are things different? I think our scaleup network is probably the place that our systems are the most different where Google uses the 3D Taurus with the um optical circuit switches to to configure around bad things. We take a much more conventional networking approach where we basically build our scaleup network with a clone network um but with a proprietary very low latency link to get around a lot of the overhead of of conventional Ethernet and then have a scale out network that is conventional Ethernet.
极其可靠地运行。那么,哪些地方不一样呢?我觉得 scale-up 网络大概是我们两家系统差别最大的地方。谷歌用的是 3D Torus 加光路交换机,出问题时可以绕过去重新配置。我们采取的是更传统的网络路线,基本上是用 Clos 网络来构建 scale-up 网络,但用的是一种私有的、延迟极低的链路,以避开传统以太网的大量开销;然后再用传统以太网做 scale-out 网络。
便签笔记
19:53
>> Yeah. And just following up on one of your points I think uh a lot of the startups don't have that system experience. you know, we uh Louise and others wrote, you know, the data center is a computer >> and you know, like learning about systems the hard way by building those supercomputers. I think a lot of those lessons aren't obvious uh until you've experienced the pain yourself and uh I think that gives both of us an advantage over startups. Yeah, he was talking about Luis Barroso and Ers Hosel and uh Partha uh's book uh the warehouses the computer the something like that but anyways it's in several editions and all the lesson >> warehouse computing >> warehouse computing all the lessons of it learned over the years about building at this giant scale.
>> 是的。顺着你刚才的一个观点再补充一下,我觉得很多创业公司并没有那样的系统经验。你知道,我们……Luiz 和其他人写过那本《数据中心即计算机》。>> 而且,是通过建那些超级计算机,以最艰难的方式学会系统这门课。我觉得很多教训在你自己亲身吃过苦头之前是看不出来的,我认为这一点让我们两家相对创业公司都有优势。对,他说的是 Luiz Barroso、Urs Hölzle 还有 Partha 他们那本书,《数据中心即计算机》,大概是这个名字,反正它出了好几版,里面全是——>> 仓库级计算 >> 仓库级计算——所有这些年在这种巨大规模上做建设所积累的经验教训。
便签笔记
05软件护城河正被 AI 填平
20:45
>> Yeah. So it sounds like the answer to the question was kind yeah it kind of is a mode if you don't the chips there's a lot more to it than chips. Um and uh so how much of Nvidia's and Google's competitive advantage today is really about the hardware versus basically software the compiler libraries and how does that work within your organizations and putting it differently if you gave the startup the net list uh and maybe uh the layout but not the stack could they could they compete and I more tag you [laughter] uh Well uh in the design of TPUs we tried to follow uh the advice of my thesis advisor [laughter] which >> well that's well that's a wonderful idea. I think that should be a requirement but what was his advice?
>> 是的。所以听起来这个问题的答案是:这确实算是一条护城河,如果你只有芯片……这里面的门道远不止芯片。呃,那么今天英伟达和谷歌的竞争优势,究竟有多少来自硬件、多少来自软件——编译器、库这些?这在你们各自的组织里是怎么运作的?换个说法,如果你把网表、甚至版图交给一家创业公司,但不给整个软件栈,他们能竞争得过吗?我先点你 [笑声] 呃,好吧,在设计 TPU 的时候,我们试着遵循我的博士导师的建议。[笑声] >> 嗯,这想法真不错,我觉得这应该成为一项硬性要求。不过他的建议是什么?
便签笔记
21:40
>> Uh don't put off until runtime what you can do at compile time. >> Okay. uh he so uh we we tried to do that and we designed the architecture um you know in a room of 10 people and two of them were uh part of the compiler team. So we tried to make a machine that was easy to compile to and uh so I I think that provided a lot of benefits and also we we've kept the same general architecture and you know this for example the sizes of the memories can increase or decrease um and things like that but you know that's like having a PC you can put more dims in your PC and still run, you know, word. So, it's um pretty flexible.
>> 呃,凡是能在编译期做的事,就别拖到运行期。>> 好的。呃,所以我们就照着做了。我们设计这个架构时,屋里就十个人,其中两个来自编译器团队。所以我们试图做出一台易于编译目标的机器,我认为这带来了很多好处。而且我们一直保持着同样的总体架构,比如说内存的大小可以变大或变小,诸如此类,但这就像你有一台 PC,可以往里多插几条内存条,照样能跑Word。所以它相当灵活。
便签笔记
22:32
>> What What about the libraries? >> Yeah, there's a lot of uh frameworks and um they're evolving. Uh so, it's important to keep up with those and to have good support. And then also some people like writing their own kernels and so uh you need to support those as well. How about how about you Bill about hardware and software? >> Yeah. So you know the the product is the whole system which includes you know the hardware and and the software and and the software is a really integral part of it but in many ways for deep learning it's an easier problem than the problem we had with general GPU computing that sort of led up to that. Um you know when we sort of you know you know kicked off you know GPU computing with CUDA and and G80 being launched in 2006 there were literally you know thousands of applications that would benefit from running in parallel but had serial codes many of them written in forrand and it was a huge burden and there there were you know you know 100,000 million line codes to port that stuff over and so we
>> 那库这块呢?>> 是的,有很多框架,而且它们一直在演进。所以跟上这些框架、提供良好支持很重要。另外有些人喜欢自己写内核(kernel),所以你也得支持这种需求。Bill,你怎么看硬件和软件?>> 是的。产品就是整个系统,其中既包括硬件也包括软件,而软件是非常核心的一部分。但在很多方面,对深度学习来说,这比我们之前做通用 GPU 计算时面对的问题要容易。你知道,当年我们启动 GPU 计算,随着 CUDA 和 2006 年发布的 G80,当时确实有成千上万的应用能从并行中受益,但它们的代码是串行的,很多还是用 Fortran 写的,移植的负担极其沉重,有那种十万行、上百万行的代码要移植过去。所以我们在 CUDA 生态
便签笔记
23:38
made a huge investment in in the CUDA ecosystem the CUDA language itself um and you know many libraries underneath you know with to do FFTs well to do matrix operations well and all of this and so when we started you know our first effort in in deep learning in in in 2010 um developing the software um in a collaborative project with Andrew Wing at at Stanford that turned into QDNN um it was a you know wow this is a really tiny application there's you know only a couple kernels that are really important in here we make those run fast this whole thing runs fast that was such a breath of fresh air after looking at weather code with a million lines of code that that it was it's it's an easy problem compared to what the problem with supercomputing and trying to paralyze those those old dusty decks was. Um that you know that became easy.
和 CUDA 语言本身上投入了巨大精力,底下还有很多库,比如把 FFT 做好、把矩阵运算做好,等等。所以当我们在 2010 年开始做深度学习方面的第一次尝试时,和斯坦福的 Andrew Ng 合作开发软件,后来变成了 cuDNN,那时的感觉是:哇,这应用真小,真正重要的核心内核就那么几个,把它们跑快,整个东西就快了。在看过上百万行的气象代码之后,这简直是一股清新的空气。和超算领域那种要把老旧代码并行化的难题相比,这实在是个容易的问题。那件事变得容易了。
便签笔记
24:23
It also became very critical. We could see that we'd get a program you we'd get one of these deep learning models up and running and then you know over the next 6 months we double its performance because there are easy ways to lose performance in there. And so we started developing tools that would profile and tools that would help us you know fuse kernels and and tune things um to make it work well. So I think you know the software is a very very critical part. I don't know how much of a a barrier that's going to be to these startups these days because right now you just sort of tell claude I want a software stack and go away for a weekend. That that's as you that's my next question.
但它也变得非常关键。我们能看到,把一个深度学习模型跑起来之后,接下来半年里我们能把它的性能翻一倍,因为里面有很多地方很容易白白损失性能。于是我们开始开发工具做性能分析,开发工具帮我们做算子融合、做各种调优,让它跑得好。所以我认为软件是极其关键的一环。不过我不知道现在这对这些创业公司还构成多大的门槛,因为如今你基本上就跟 Claude 说一声:我要一套软件栈——然后你去过个周末就行了。这正好是我下一个问题。
便签笔记
24:56
As you know, [laughter] anybody who's uh uh knows a programmer has uh heard from their spouse of shouts in the middle of the night. Oh my god, look what it did. But yeah, we're we're seeing this uh remarkable event about coding being done by machines rather than by people. So yeah, that was kind of the question. Do you think if there was a remote software is is that about to go away with these the advances in AI or know that's you know that's uh a polyianish view of what's going on here? >> I I think you know any any software moat um I think there's still an advantage to having libraries that have been tuned and honed over the years that are easy for people to apply. But I think that moat has been significantly degraded because if you have the spec of that library, I think turning, you know, claude code loose and and recreating it is not that difficult a thing to do.
你知道,[笑声] 任何认识程序员的人,都听过他们的伴侣讲半夜被喊醒:天哪,你看它干了什么。是啊,我们正在见证这个了不起的现象:写代码这件事正由机器而不是人来完成。所以,这正是我想问的。你觉得,如果软件曾经是护城河,它是不是要随着 AI 的这些进展而消失了?还是说,那是对眼下局面过于乐观的看法?>> 我觉得,任何软件护城河……我认为,拥有那些经过多年打磨调优、别人用起来很顺手的库,仍然是一种优势。但我认为这条护城河已经被大大削弱了,因为如果你有那个库的规格说明,我觉得放 Claude Code去把它重新做出来,并不是什么难事。
便签笔记
06精度削到头后靠什么翻倍
25:54
>> Yeah, I I I agree. The the the moat is getting shallower. [laughter] >> It's getting filled up with sand. Is starting to be able to walk across. >> Okay. Uh you know, Bill brought this up, but we and when Bill gives a talk, he gives credit to this. We squeezed a lot of speed of the first years by reducing precision yet like for me I shocking figured out our exponent is bigger than the fraction right that that was that's a new thing relatively [clears throat] new idea um but we can only go so low what happens for generation over generation improvements when we can't shave bits off anymore >> I can't remember >> okay >> Bill you're up yeah so um we we've we've pretty much hit 2x per year um over the past 14 years starting with Kepler in 2012 which is the first uh generation since we've really looked at this as a serious application um and only 3x of that has come from process technology.
>> 是的,我同意。护城河正在变浅。[笑声] >> 它正在被沙子填满,快能直接走过去了。>> 好的。呃,Bill 刚才提到了这一点,Bill 做演讲时也会提到这个功劳归属。前几年我们靠降低精度榨出了大量速度。对我来说挺震撼的一点是,我们发现指数位比尾数位更重要——这在当时算是个相对较新的想法。[清嗓子] 但精度只能降到一定程度,当我们再也削不动比特位的时候,一代又一代的性能提升要靠什么?>> 我记不太清了。>> 好 >> Bill,到你了。是的,过去 14 年我们基本做到了每年 2 倍,从 2012 年的 Kepler开始——那是我们真正把这当作一个严肃应用去看待的第一代——而其中只有 3 倍来自制程工艺。
便签笔记
26:52
This is in contrast to, you know, the microprocessor heydays of the '90s when it all came from process technology. Um, and you know, I think, you know, at NVIDIA, we've kind of led the way with numeric, you know, going from starting out with FP32 because that's what we had in in Keeper and and you know, you know, shipping NVFP before and there's some new things in in Vera Rubin. Um, and you know, clearly there's a couple more turns there, but we're getting near the end on on numerical precision. There's still a couple clever things um that can be done, but there are a lot of other axes that we can continue um to to innovate along sparity, circuits, locality, even better models that wind up giving you sort of, you know, more tokens per per watt. Um but we are sort of getting um getting down to the, you know, point where all the low hanging fruit has been picked and we need to climb higher up the tree to find those, you know, 2x um fruits hanging um from the branches. But we we have a bunch of
这和 90 年代微处理器的鼎盛期形成对比,那时提升全部来自制程工艺。而且我认为,在英伟达,我们在数值格式上算是引领了方向,从FP32 开始,因为 Kepler 上我们只有这个,再到出货 NVFP4,Vera Rubin 里还有一些新东西。显然还能再走上一两轮,但在数值精度这条路上我们已经接近尽头了。还有几个巧妙的招可以用,但我们还有很多其他维度可以继续创新:稀疏化、电路、局部性,还有更好的模型,最终让你每瓦能产出更多 token。不过我们确实已经走到了低垂的果实都被摘完的地步,接下来得爬到树的更高处,去够那些挂在枝头的 2 倍的果子。但我们有一堆
便签笔记
27:43
good ideas. I think they can get us through at least the next four or five generations and then I can retire. >> Yeah. [laughter] >> Yeah. I I didn't is that four or five years or [laughter] >> eight to 10 years? >> Yeah. I didn't want to suggest numeric was the only thing but that was one of one of the >> that was certainly a lowhanging fruit right as >> that was very low hanging and we also had some numeric uh innovations. So um when Jeff was doing those initial uh AI things, >> this is Jeff Dean.
好点子,我觉得至少能支撑接下来四五代,然后我就可以退休了。>> 哈哈 [笑声] >> 是啊,我没听清,是四五年还是 [笑声] >> 八到十年?>> 是的,我并不是想说数值精度是唯一的因素,但那确实是其中之一 >> 那绝对是低垂的果实,对吧 >> 那是非常低垂的。我们在数值上也有一些创新。所以,当 Jeff 在做那些早期的 AI工作时——>> 这是指 Jeff Dean。
便签笔记
28:17
>> Jeff Dean. Um they were taking doing it in FP32 but then storing it by truncating the uh low order 16 bits. >> And if you hear truncation and you're a numerical analyst, it's kind of like uh fingernails on a chalkboard. [laughter] But but it actually worked. Uh so when we developed TPUs uh we could run programs in BF-16 that had run on CPUs which were being used at the time. BF-16 is the whacked the the chop it e 32-bit one where you >> right >> yeah so we we could run those and get the same results and so that was really powerful because we didn't have to spend a lot of time debugging things you know at the system level why is this model not working things like that so we could uh verify um correct operation and then once we had that we used that for a before we started adopting the smaller formats like FP8 and FP4.
>> Jeff Dean。他们当时是用 FP32 来算,但存储时把低 16 位截断掉。>> 如果你是搞数值分析的,听到“截断”这个词,感觉就像指甲刮黑板。[笑声]但它居然真的管用。所以当我们开发 TPU 时,我们能用 BF16 跑那些原本在当时使用的 CPU 上跑的程序。BF16 就是那个被砍掉一截的 32 位格式,你——>> 对 >> 是的,所以我们能跑这些程序,并得到一样的结果,这非常有用,因为我们不必花大量时间在系统层面调试,比如“这个模型为什么不工作”之类的问题。所以我们可以验证运算是否正确,有了这个基础之后,我们才开始采用更小的格式,比如 FP8 和 FP4。
便签笔记
07三年硬件周期追不上三月模型
29:29
>> Um, so I think one of Google's advantages, theoretical advantages at least is Google both have people pushing the state of the AI and building hardware and so theoretically that's a big advantage and then I think I read that Jensen that Google that Nvidia is going to build up in-house expertise. So what do you think how important how big of an advantage is the access to people pushing the state-of-the-art versus people who are just going to use open models and do that or you know not not have that close relationship for a male?
>> 嗯,我认为谷歌的一个优势,至少是理论上的优势,是谷歌同时拥有推动 AI 前沿的人和造硬件的人,理论上这是个很大的优势。然后我好像读到 Jensen 说英伟达也要建立内部的这类能力。所以你们怎么看,能接触到推动最前沿的人有多重要、优势有多大?相比之下,有些人只是用开源模型来做,没有那种紧密的关系。
便签笔记
30:04
I I think Norm I think you're up. Uh yeah, there's uh you know leaderboards that show you know some of the open models are performing quite well and so there's there's a race uh for the proprietary models to to stay ahead. Um I think they'll you know do to the ability to do better tuning of the models for a particular system either a GPU or a TPU uh there'll always be some advantage uh to the proprietary versions but uh >> you know I was thinking more of you know as hardware people could go talk to the ML people about what's happening next and what should we put in our hardware that uh if you didn't access to those experts, you would have to guess. You see that >> to some extent. Um the problem is it it takes you know 2 and 1/2 or 3 years to uh from an initial idea to have a system that can be manufactured in volume and the ML people come up with a new idea like every 3 months or [laughter] >> okay all right you can't really design for uh a super specialized uh you know a super specialized machine
我觉得 Norm,该你了。呃,是的,排行榜显示有些开源模型表现相当不错,所以闭源模型要保持领先是有一场竞赛的。我认为,由于能针对特定系统——不管是 GPU 还是 TPU——对模型做更好的调优,闭源版本总会有一些优势,不过——>> 你知道,我更想问的是,作为硬件的人,你可以直接去找 ML 的人聊接下来会发生什么、我们该在硬件里放些什么。如果你接触不到这些专家,就只能靠猜。你觉得呢?>> 某种程度上是的。问题在于,从最初的想法到能量产的系统,要花两年半到三年,而 ML 的人差不多每三个月就冒出一个新想法。[笑声] >> 好的,明白,所以你没法为某个特定模型去设计一台超级专用的机器,因为它到时候会变,而且正如 Bill 之前说的,
便签笔记
31:30
for a particular model because uh it'll be different and like Bill said earlier, you just have to do those basic operations and matrix operations, vector operations and stuff and do them well. >> Okay. Okay. If it takes so long and they iterate faster, why are you guys getting into building your own models? >> We've we're not getting into it. We've been doing it for Okay. you've been to sorry some time and we we have a lot of internal expertise but we I think there's a distinction here between you know a company like open AAI who just had a great paper at hot chips on their jalapeno processor where they're designing a chip pretty much to run one model um and you know there therefore they know you know you know what their relative mixes of matrix ops and vector ops and memory bandwidth and and such um and you know there's a if you have to support all the models that are out there there's a there's a very big variety and shifts that provisioning um of especially memory bandwidth versus math in different ways and particularly
你只要做好那些基本运算——矩阵运算、向量运算之类的,把它们做好就行。>> 好。好。如果这要花这么长时间,而他们迭代得更快,那你们为什么还要自己去做模型呢?>> 我们不是现在才开始做。我们已经做了——好吧,你们已经做了,抱歉——做了挺长一段时间了,我们也有很多内部的专业积累。不过我觉得这里面有个区别:像 OpenAI 这样的公司,他们刚在 Hot Chips 上发表了一篇很棒的论文,讲他们的Jalapeno 处理器。他们设计那颗芯片基本上就是为了跑一个模型,所以他们很清楚自己的矩阵运算和向量运算的相对配比、内存带宽等等。而如果你必须支持市面上所有的模型,那种类就非常多了,而且它们对资源配置的需求会以不同方式变化,尤其是内存带宽相对于算力的配比,特别是在注意力这块。注意力机制开始变得
便签笔记
32:30
with the attention. The attention mechanisms are starting to get you know pretty interesting. This started happening with deepseek when they came out with the uh MLA attention and now a lot of the different models have different hybrid attention schemes. So they'll alternate layers of three layers of state space and one layer of of of full n squared attention. A lot of people are doing um sparse attention where they'll do a quick filter and then have to pick out the top k things to attend to. And this this variation in attention I think drives a lot of you know different demands on the hardware.
相当有意思了。这是从 DeepSeek 开始的,他们推出了 MLA 注意力,现在很多不同的模型都有各自的混合注意力方案。比如他们会交替排布——三层状态空间层,加一层完整的 n 平方注意力层。还有很多人在做稀疏注意力,先快速过滤一遍,然后从中挑出 top-k 个要关注的对象。我觉得注意力上的这种多样性带来了很多对硬件的不同需求。
便签笔记
33:01
And so if you're in the position of being we're going to be a silicon supplier, we're going to support all the models. It it requires you to sort of look across all those models and say what is the right hardware to build to make everybody as happy as you can without making anybody really unhappy. and and that's a hard thing to do. Um you have to decide, you know, you know, which ones you're going to emphasize, which ones you're not. And as Norm pointed out, you've got to aim ahead of the duck. >> Um because, you know, it's going to be a couple years before this is out. And in the meantime, a lot of clever ideas are going to people are going to come up with in the models that we haven't seen yet. And so this is I think what good computer architecture is about is figuring out how to make the most people happy and um and anticipate what the applications are going to need and and deliver a product um that serves that.
所以如果你的定位是“我们要做芯片供应商,我们要支持所有模型”,那你就必须把所有这些模型通盘看一遍,然后问:该造什么样的硬件,才能让尽可能多的人满意,同时又不让谁特别不满意。而这是件很难的事。你必须决定,你知道,要重点照顾哪些、不照顾哪些。而且正如 Norm 指出的,你得瞄在鸭子的前头,也就是要打提前量。>> 因为你知道,这东西要过个两三年才会出来。而在这期间,会有很多聪明的点子被人们在模型上想出来,是我们现在还没见过的。所以我觉得,好的计算机体系结构就在于此:想办法让最多的人满意,并且预判应用未来会需要什么,然后交付一款能满足这些需求的产品。
便签笔记
33:47
>> Yeah. And talking about models, I think one of the more disruptive things uh that that's been recently developed is well it was proposed uh like seven years ago but mixtures of experts because it requires a lot more interconnect uh bandwidth and if you want to have quick responses uh for users you really have to push down on the on the latency. So training is is somewhat easier in in that uh it's more of a bandwidth thing. Latency is harder to get in computer systems than low latency than higher bandwidth.
>> 对。说到模型,我觉得最近出现的比较有颠覆性的东西之一是——嗯,其实这个想法大概七年前就被提出来了——混合专家(MoE),因为它需要多得多的互连带宽。而如果你想让用户得到快速响应,就必须把延迟压下去。所以训练反而相对容易一些,因为那更多是带宽的事。在计算机系统里,延迟比带宽更难搞——低延迟比高带宽更难拿到。
便签笔记
08每秒 token 数与基准之争
34:29
>> Yeah. >> All right. So that's what drove the boardfly configuration latest TPU. >> Yeah. >> Yes. Given we're at the computer history museum I was struck at the at the uh hot chips conference. It felt like a recreation of the battles from the 1980s. So in the 1980s when uh risk risk profit risk processors were getting popular people were using myips instruction millions of instructions per second as the metric and uh but that was up to you to how to define it what you ran all that stuff. So the risk companies were saying well you know you can trust us our but all those other other bastards are lying right.
>> 对。>> 好的。所以这就是最新一代 TPU 采用那种 boardfly 配置的原因。>> 对。>> 是的。既然我们身处计算机历史博物馆,我在 Hot Chips 大会上真的很有感触。感觉就像1980 年代那些论战的重演。在 80 年代,RISC 处理器开始流行的时候,大家用 MIPS(每秒百万条指令)作为衡量指标,但那个指标怎么定义、跑什么程序全由你自己说了算。所以 RISC 厂商就说:我们的数字你们可以信,但其他那些混蛋都在撒谎,对吧。
便签笔记
35:07
[laughter] So at at uh at hot chips they were using tokens per second as the metric and >> what were they running? How big was the model? No, it was just tokens per second. So there was an there was an attempt to try and uh and the solution for the risk one led to the peak companies getting together collaborating on benchmarks which were called spec which were pretty you know there's pluses and minuses spec that they rallied around it and then it turned out even if the you were 10% slower you could have other advantages and so that it worked out pretty well. So there was an attempt called ML Perf that to anticipate this problem that's been around for many years but I mean it do you and spec ML Perf was not quoted that was not mentioned I think at uh at the hot chips so do you agree that maybe uh ML Perf was not as successful Hops as spec and if not why do you think it didn't and I think it's your turn bill.
[笑声] 而在 Hot Chips 上,他们用每秒 token 数作为指标。>> 那他们跑的是什么?模型有多大?不,就只是每秒 token 数。所以当时也有人想解决这个问题;而 RISC 那次的结果是,那些顶尖公司聚到一起,合作搞出了一套基准测试,叫 SPEC,它当然有利有弊,但大家都团结在它周围。后来发现,即使你慢 10%,你也可能有别的优势,所以整体效果还不错。后来也有过一个叫 MLPerf 的尝试,就是为了预防这个问题,已经存在很多年了,但我是说,你和 SPEC……MLPerf 没人引用,我印象里在 Hot Chips 上根本没被提到。所以你同意吗,也许MLPerf 不像 SPEC 那么成功?如果是这样,你觉得为什么没成功?我想该你了,Bill。
便签笔记
36:07
So, so I I like MLPF because it sort of, you know, gave a a way of, you know, cutting through it did what benchmarks are supposed to be. It was a very nice level playing field to cut through the BS and say, "Okay, you know, how how well do you really perform on on this particular application, but I think it had two issues, which is why you didn't hear much about it at Hot Chips." Um, the first is almost nobody submitted to it. And so, >> well, that's a problem. [laughter] >> Instead, because it was it was a lot of work and we we we did every time. In fact, we did all the benchmarks every generation.
嗯,我其实挺喜欢 MLPerf 的,因为它提供了一种方式,能够穿透表象——它做到了基准测试本该做的事:提供一个很好的公平竞技场,能戳穿那些废话,直接说:好,你在这个具体应用上到底表现如何。但我觉得它有两个问题,这也是为什么你在 Hot Chips 上没怎么听到它。第一,几乎没人提交结果。所以…… >> 那确实是个问题。[笑声] >> 而不提交是因为工作量很大——我们每次都做。事实上,我们每一代都把所有基准都跑了一遍。
便签笔记
36:40
>> Um, and every every iteration of MLF would have on 7.1 or something now or whatever. Um, but but a lot of startups would say, "Oh, no, that's inconvenient. And if we actually played by all the rules, we would look that good." So, we're going to cherrypick this one result and put that on our slides and and talk about that. And so, that's one reason why I think it hasn't really taken off. The other reason is today what everybody is really worried about is tokens per watt or tokens per dollar. And there's really, you know, most of the MLF benchmarks don't really address that. And even I think the one that that comes close, it doesn't address it in quite the right way. And so what everybody is showing and actually you did see some slides of this at hot chips is the uh semi-analysis inference Xbenchmark because it it hits exactly the data that people want and they do it in in a very reproducible way. to get a data to get a data point on that semi- analysis chart you have to have a um whether it's GitHub or hugging face a
>> 嗯,MLPerf 的每一次迭代——现在大概到 7.1 之类了。但很多创业公司会说:“哦不,那太麻烦了。”“而且如果我们真按所有规则来跑,成绩也不会那么好看。”所以他们就挑一个漂亮的结果放到幻灯片上去讲。我觉得这是它一直没真正流行起来的一个原因。另一个原因是,今天大家真正关心的是每瓦 token 数、或者每美元 token 数。而大多数 MLPerf 的基准并没有真正回答这个问题。就连我觉得最接近的那一个,回答的方式也不太对。所以现在大家展示的——你在 Hot Chips 上确实也看到了一些这样的幻灯片——是 SemiAnalysis 的 InferenceMAX 基准,因为它正好命中了大家想要的数据,而且做得非常可复现。要在 SemiAnalysis 那张图上拿到一个数据点,你必须有一个代码仓库——GitHub 或者 Hugging Face 都行——别人可以下载、运行,在那套硬件上跑出来,
便签笔记
37:32
repository of the code that you can load and run and and show on that hardware with this model this is is the tokens per second tokens >> is this something you run yourself or they do it for you >> they they do it yeah >> and they run it on your hardware and give you a number is that how it works [sighs] >> okay >> I don't remember the details of that but it it's it's >> the seems to be one that's catching on >> I think people people are using inference text more than MLF today. >> Yeah. >> Yeah. That ML Perf does require like a full-time team >> and not not an ins and not a small team either. So >> yeah, that's one of the issues with it.
用这个模型,得出每秒多少 token。 >> 这是你自己跑,还是他们帮你跑?>> 是他们跑的,对。 >> 他们在你的硬件上跑,然后给你一个数字,是这样吗?[叹气] >> 好吧。>> 具体细节我记不清了,但它…… >> 这个好像确实越来越流行。 >> 我觉得现在人们用>> InferenceMAX 比用 MLPerf 更多。>> 对。>> 是啊。MLPerf 确实需要一个全职团队。 >> 而且还不是个小团队。所以 >> 对,这是它的问题之一。
便签笔记
09AI 的社会影响与硬件人的责任
38:10
>> So it ended up just being too intensive to run like like the the old the old TPC benchmarks were a big effort for these database companies. So expensive and and this alternative was working pretty well. It was both expensive and it didn't really give you the number you wanted. >> It It was okay. Other than that, it was great. [laughter] >> All right. I I I'm supposed to warn the AV people to to uh we'll do a poll after this. You'll show the poll results of my question after this question. Okay. Uh so the and this is a more open-ended the silicon gold rush has had significant has significant economic societal and even geopet geopolitical implications beyond the technology itself. What do you believe are the most profound impacts of this rapid acceleration in AI hardware development and what responsibility do hardware architects like us and leaders of these companies bear in shaping that future responsibly?
>> 所以最后就变成投入太大、跑不动了。就像以前的 TPC 基准,对那些数据库公司来说也是一项很大的工程。太贵了,而这个替代方案效果挺好。它既贵,又不能给你想要的那个数字。>> 除此之外,它挺棒的。[笑声] >> 好的。我得提醒一下音视频的同事,这个问题之后我们要做个投票。你们会在这个问题之后展示我那道题的投票结果。好。那么下面这个问题更开放一些:这场硅片淘金热已经带来了重大的经济、社会甚至地缘政治影响,远远超出技术本身。你们认为,AI 硬件这种飞速加速发展带来的最深远的影响是什么?而像我们这样的硬件架构师、以及这些公司的领导者,在负责任地塑造那个未来上又该承担怎样的责任?
便签笔记
39:09
Bill, I think it's your turn. Okay. So I I wouldn't call it the silicon gold rush. I'd call it the AI gold rush because it's really Russian AI and I think it's really benefiting almost all aspects of our lives. So if you look at medicine, I think you know this is where one of its early applications was in image analysis. But now it's also being applied to diagnosis where it's helping doctors give more accurate diagnosis. And we can have personal health coaches that sort of, you know, look at us and say, "I don't think you should eat that piece of cake. You're over your calorie intake today." Um, >> that takes a uh [laughter] >> and and what if you just ignored that coach?
Bill,我想该你了。好。我不会把它叫做硅片淘金热,我会叫它 AI 淘金热,因为真正的推动力是 AI,而且我认为它几乎让我们生活的方方面面都受益。比如你看医疗领域,我觉得它最早的应用之一是影像分析。但现在它也被用到诊断上,帮助医生做出更准确的诊断。我们还可以有个人健康教练,它会看着我们说:“我觉得你不该吃那块蛋糕,你今天的热量已经超标了。” 嗯, >> 那得有点儿……[笑声] >> 那如果你干脆无视那个教练呢?
便签笔记
39:44
>> Yeah, well that's up to you. But at least he gave you the advice. You know, in education, I'm I'm, you know, a, you know, recovering educator. Um, you know, every student could have a personalized tutor that really understands, you know, what motivates them, how they learn, and and presents the material in a way that gets them through the, you know, difficulties of of grasping uh concept complex concepts. Um, in engineering, AI is already automating a lot of our tasks. I recently needed to design a piece of hardware. I basically wrote the spec and set cloud loose and uh, after fixing a couple errors in my spec, it it came out with a [laughter] a pretty pretty good design.
>> 是啊,那就随你了。但至少它给了你建议。你知道,在教育方面,我算是个“退休的”教育工作者。你知道,每个学生都可以有一个个性化的导师,它真正理解是什么在激励他们、他们怎么学习,并用合适的方式讲解材料,帮他们跨过理解复杂概念时的那些困难。在工程领域,AI 已经在自动化我们的很多工作了。我最近需要设计一块硬件。我基本上就是写好规格说明,然后把 Claude 放出去干,在修正了我规格里的几个错误之后,它 [笑声] 给出了一个相当不错的设计。
便签笔记
40:19
>> Did it fix this? >> No, it it produced exactly what I asked for, which was not the right thing. You have to learn how to write write a good spec, but it's basically enabling, you know, engineers in all sorts of fields to move up the ladder, right? They're no longer doing the, you know, you know, junior level calculations and stuff. They're deciding what needs to get built and they're, you know, managing a team of agent minions that, um, you know, actually carry out the work and a lot of business processes are are moving in the same way. So, we're all becoming much more productive. So, I think in entertainment, it's helping us produce really, you know, wonderful compositions, whether it's in, you know, movies or games or or or music. Um, and so I think there's tremendous number of benefits about AI. But on the flip side, you know, any great technology can be used for good and can be used used for evil. And so, you know, the most obvious ones with AI are deep fakes. And I think we need to move very rapidly to, you
>> 它把这个修好了吗?>> 没有,它完全照我要求的做了出来,而我要求的并不是对的东西。你得学会怎么写一份好的规格说明。但它基本上让各行各业的工程师都能往上走一层,对吧?他们不再去做那些初级的计算之类的活儿了。他们负责决定要造什么,然后管理一队智能体“小弟”,由它们实际去执行工作。很多业务流程也在朝同样的方向走。所以我们都变得更高产了。在娱乐领域,我觉得它在帮我们创作出非常棒的作品,不管是电影、游戏还是音乐。所以我认为 AI 带来的好处非常多。但另一方面,任何伟大的技术都可以用来行善,也可以用来作恶。就 AI 来说,最明显的就是深度伪造。我认为我们需要非常快地推进内容溯源和
便签笔记
41:09
know, embrace, you know, provenence and authentication. So that basically unless you see an image as being, you know, authenticated, you know, appropriately, you just assume, um, that it's a fake. Um, I think that there's, you know, danger with, um, you know, um, various nefarious uses. Cyber has gotten a lot of attention lately. Although even though people like to talk about the AI side of cyber, the vast majority, I think it's like over 80% of all successful cyber attacks are are human engineering. They're fishing. Of course, a AI makes that better. Um you can use AI for pathogen design. Um and then there's um the big you the big risk is as we make everybody more productive that the the nature of employment is going to change and and we'll need more people in certain jobs and fewer people in other jobs. So we need to find some way of of easing that transition for the people who are yeah >> dropping out of one and moving into the other.
认证。基本上,除非你看到一张图片经过了恰当的认证,否则你就该假定它是假的。我觉得还存在各种恶意用途的危险。网络安全最近受到了很多关注。虽然大家喜欢谈网络安全里 AI 的那一面,但绝大多数——我记得超过 80% 的成功的网络攻击靠的都是社会工程,是钓鱼。当然,AI 会让钓鱼更厉害。你也可以用 AI 来设计病原体。还有一个大风险是,随着我们让每个人都更高产,就业的形态会发生变化,某些岗位会需要更多人,另一些岗位需要更少人。所以我们得想办法,帮那些人平稳过渡 >> 从一个行业退出、转到另一个行业。
便签笔记
42:03
>> Yeah. Help. How do we reskill people whose jobs change? >> Norm you got a take on this? Um yeah, I think uh one of the biggest impacts is going to be in science. Uh we they've made a lot of like super impressive uh results which don't usually make the popular press. So at Google Deep Mind they folded all the proteins known to mankind >> and got a Nobel Prize. >> Got [laughter] a Nobel Okay. Yeah. Um and uh there's there's many other examples in in science. So another thing is uh Google is using it in virtually all of their applications but it's it's very subtle and in in many places you don't really notice it. So for example, one of the first applications was uh Google maps.
>> 是的。我们该怎么帮那些工作发生变化的人重新培训技能?>> Norm,你对这个有什么看法?嗯,我觉得最大的影响之一会体现在科学上。他们做出了很多超级令人惊叹的成果,只是通常上不了大众媒体。比如在 Google DeepMind,他们把人类已知的所有蛋白质都折叠出来了 >> 还拿了诺贝尔奖。>> 拿了 [笑声] 一个诺贝尔奖。好吧。是的。科学上还有很多其他例子。另外一件事是,Google 在用它,几乎用在他们所有的产品里,但用得非常不动声色,很多地方你根本注意不到。比如举个例子,最早的应用之一是 Google 地图。
便签笔记
43:02
Uh that was when Louise was running uh maps and you know the original Google maps would say you know drive you know 1,000 ft and turn right on Tennyson Street right and by integrating the street view database now it says things like you know drive a thousand feet and turn right at the Shell gas station and that es And you know that's much more accessible to people and especially at night when the street signs aren't lit. You know Sunnyale has nice lit up ones but >> Palo Alto we don't. So [laughter] so I there's a lot of subtle things like that which I think people won't notice but you know we'll just take for granted after a while.
那还是 Louise 负责地图的时候。你知道,最早的 Google 地图会说:往前开 1000 英尺,然后在 Tennyson 街右转。而通过整合街景数据库,现在它会说:往前开1000 英尺,在壳牌加油站右转。你知道,这对人们来说好用多了,尤其是在晚上路牌没有灯光的时候。桑尼维尔的路牌照明不错,但 >> 帕洛阿尔托就没有。所以 [笑声]所以有很多这样细微的地方,我觉得人们不会注意到,但你知道,我们只会把它们当作理所当然,过一段时间之后。
便签笔记
10能源是最大瓶颈
43:56
>> Okay. Let's see. Can you show the poll? Does that work? Okay, so here's my question. Uh, the biggest bottleneck to continued widespread deployment of AI chips is energy availability, semig manufacturing capacity, cooling of hot chips, memory capacity, bandwidth, cost of AI chips. Don't know. >> More people don't know. Okay. Well, I guess Oh, this is [laughter] >> dynamic. Dynamic. We we we won't uh we uh it's still changing, but clearly uh right now it's uh energy availability. Well, maybe I should just maybe I'll give it 30 seconds. [laughter] >> So, what what what you guys could answer this question yourselves and you you could disagree with the audience. What what what do you think about what's this biggest limiter of those those things?
>> 好的。我们看看。能把投票结果放出来吗?行吗?好,那我的问题是:AI 芯片持续大规模部署的最大瓶颈是——能源供应、半导体制造产能、高热芯片的散热、内存容量、带宽、AI 芯片的成本,还是不知道。>> 更多人选了“不知道”。好吧,我想……哦,这个 [笑声] >> 是动态的。动态的。我们不会……呃我们呃它还在变,但很明显现在就是能源供应。嗯,也许我该给它 30 秒时间。[笑声] >> 那么,你们几位其实可以自己回答这个问题,也可以不同意观众的看法。你们怎么看,这里面哪个才是最大的限制因素?
便签笔记
44:53
>> Well, I was just going to talk about the energy availability. >> Okay, go ahead. >> Um, so one of the issues is, you know, people are talking about gigawatt data centers or 5 gawatt data centers and the transmission lines are very hard to get permits for and communities don't like to have, you know, big power lines running over their houses and stuff. And so a a lot of the hyperscalers and other companies are moving to a model where they're doing on-site uh power generation and you can do that various ways. Some companies use uh natural gas to generate electricity but if you use uh if you put it in the right place uh you can get most of the power from wind and solar. And so that's the approach that we're trying to take >> which is car carbon-f free energy and will not uh will not uh you know add to the the the load of the on the climate >> right >> but by going off the grid and using carbon free any resources you can build data centers without uh clobbering things. Yeah, you you have an issue that
>> 嗯,我本来正想说能源供应这块。>> 好,请讲。>> 嗯,问题之一是,大家现在都在谈吉瓦级的数据中心,甚至 5 吉瓦的数据中心,而输电线路很难拿到许可,社区也不喜欢那种大电线从自家房子上头架过去之类的。所以很多超大规模厂商和其他公司都在转向一种模式,就是在现场自己发电,实现方式有好几种。有些公司用天然气发电,但如果你把它建在合适的位置,大部分电力可以来自风能和太阳能。这就是我们想走的路线 >> 也就是零碳能源,不会……不会给气候增加负担>> 对 >> 而且通过离网、使用零碳能源,你可以建数据中心而不会把事情搞砸。是啊,你会遇到一个问题:数据中心需要稳定的电力,而太阳
便签笔记
46:09
you need steady power for your data center and uh you know the sun is only out during the day and when there are no clouds and the wind only works when the wind's blowing. But u you know in general we've seen the various sources of supply being reasonably well balanced. I think it's just the natural market forces that are deploying resources to do that. But land power shell which is sort of the three things that you need um you know to build a data center tends pot tent to drive that demand. And you know, while a lot of people are are looking at natural gas and colllocating generators, I think um you know, gas turbines are sold out for like the next 5 years because of of this demand. And there are companies that are taking engines off of airplanes and converting them to to be used as as generators for uh um stranding lots of passengers somewhere.
只在白天、而且没云的时候才有,风也只有刮风的时候才管用。不过总体来说我们看到各种供应来源之间还算相当平衡。我觉得这就是自然的市场力量在配置资源。不过土地、电力、厂房外壳——差不多就是建一个数据中心需要的三样东西——往往会推动这块需求。而且你知道,虽然很多人在看天然气、在就地配置发电机组,但我觉得燃气轮机因为这波需求,未来大概 5 年的产能都已经卖光了。还有些公司在把飞机上的发动机拆下来,改装成发电机用——结果就是一堆乘客被困在某个地方。
便签笔记
46:57
[laughter] >> Yeah. But but if that's what we do, if we if we go to natural gas, we're going to have a giant carbon footprint. If if we don't >> It's often used to back some more renewable source. They'll have they'll have solar and wind, but because you know you can't count on that and we don't have the storage capacity at this point, although there are some very very good technologies that are >> and so you think uh like lithium batteries aren't going to be a >> I I think lithium batteries are too expensive per kilowatt hour to even to ride through the kind of 20 to 100 hour um thing you really need to bridge a gap when you might not have you know sunny days for a period of time. There are thermal batteries that are looking very promising and there are people starting to build data centers with those that are cheap enough per per kilowatt hour that they could fill that gap.
[笑声] >> 是啊。但如果我们真这么干,如果走天然气这条路,碳足迹会非常大。如果我们不…… >> 它通常是给可再生能源做后备的。他们会上太阳能和风电,但因为你没法完全指望它们,而目前我们又没有足够的储能,不过确实有一些非常好的技术…… >> 那你觉得像锂电池不会成为…… >> 我觉得锂电池按每千瓦时算太贵了,撑不起那种 20 到 100 小时、你真正需要用来填补空缺的时长,就是可能连着一段时间都没有晴天的情况。有一些热储能电池看起来非常有前景,已经有人开始用它们建数据中心了,每千瓦时的成本便宜到足以填上那个缺口。
便签笔记
47:45
>> So if you're doing serving uh to consumers though, >> yeah, >> most consumers I mean there's a dial cycle, right? And so you don't have to have full power at 3:00 a.m. when you're serving. So >> it has to do with with keeping use of a capital resource, right? you spent, you know, $10 billion building a data center and you'd like to have that capital resource be busy around the clock. And so if the consumers here are all sleeping, you'll sell those GPU cycles to the consumers on the other side of the world where it's >> where it's the latency is too large.
>> 不过如果你是在给消费者提供服务的话,>> 嗯,>> 大多数消费者……我是说存在昼夜周期,对吧?所以凌晨三点你并不需要满功率去提供服务。所以 >> 这跟资本资源的利用率有关,对吧?你花了 100 亿美元建了个数据中心,当然希望这笔资本资源全天候都在忙。所以如果这边的用户都在睡觉,你就把那些 GPU 算力卖给地球另一边的用户,那边正好是 >> 那边的延迟就太大了。
便签笔记
48:18
>> Yeah, it depends on the use. Yeah, [laughter] >> we like we like quick responses for our users. >> What what about uh nuclear is also carbon-f free and there's these small modular reactors. Uh I think Google's certainly made an announcement in this space. Do you guys have thoughts about nuclear as to help data centers? >> Yeah, we we're you know um you we're looking at all all technologies. We actually have a guy whose job it is to sort of keep track of energy technologies and and advise on on on data centers. Um and it looks promising but you know it's always been um expensive compared to you know if you look at the what they call they call the levelized cost per kilowatt hour um you know compared to to natural gas. I mean natural gas is is very tough to beat and even if you're worried about um you know the the carbon natural gas with carbon sequestration I think still winds up beating um nuclear on uh you know levelized cost per kilowatt hour. I'm a big fan of pumped hydro, but uh [laughter] in Illinois where I'm from,
>> 是啊,这要看用途。[笑声] >> 我们希望给用户快速的响应。>> 那核能呢?核能也是零碳的,而且现在有小型模块化反应堆。我记得谷歌在这方面肯定发过公告。你们对用核能支持数据中心有什么看法?>> 是的,我们……嗯,我们所有技术都在看。我们其实专门有个人,他的工作就是跟踪各种能源技术,并就数据中心给出建议。嗯,它看起来有前景,但你知道它一直都比较贵,如果你看所谓的平准化度电成本,也就是每千瓦时的成本,跟天然气比的话。我是说天然气非常难被打败,甚至就算你担心碳排放,带碳封存的天然气我觉得在平准化度电成本上还是能赢过核能。我个人很喜欢抽水蓄能,但是呃 [笑声] 在我老家伊利诺伊州,
便签笔记
11观众问:内存、量子、固化芯片
49:22
it's not too practical. >> Okay. Uh now I'm going to go to audience questions just about when I said I would too. Uh what are the biggest ch and the way the audience audience questions get voted up and down. So this is the most popular one. What are the biggest challenges to resolving the known supply chain barriers associated with AI hardware which I presume is the over demand compared to the supply. So do you have a take on that? The biggest or what's the biggest challen I guess what's the biggest challenges?
这不太现实。>> 好。我说过也会转到观众提问,现在就来。观众问题是靠投票上下排序的,所以这是票数最高的一个:要解决与 AI 硬件相关的已知供应链障碍,最大的挑战是什么——我猜这里说的是需求超过了供给。你们对此有什么看法?最大的挑战……我是说最大的挑战是什么?
便签笔记
49:54
I >> mean the biggest challenge is it takes a long time to build a fab. >> Yeah. Um and and I think you know right now the the place that's being felt the most painful is on the memory side and and the memory manufacturers are loving it because they can you know charge many times more for the same part if they'd had them in plentiful supply. So they maybe they're less motivated to build the fabs than they otherwise would be. >> Yeah. Um, I think it's it's it's always this thing where you're trying to project demand. Then when that demand exceeds what you projected, you know, you you've got this two or three year delay to actually build the production capacity.
我 >> 我是说最大的挑战就是建一座晶圆厂要花很长时间。>> 是啊。嗯,我觉得现在最痛的地方是在内存这边,而内存厂商乐坏了,因为同样一颗芯片他们能卖出好几倍的价钱,要是供应充足就卖不上去了。所以他们建厂的动力可能反而没那么足。>> 是啊。嗯,我觉得这事永远都是这样:你要去预测需求,然后当需求超出你的预测时,你就得面对两三年的滞后才能真正把产能建起来。
便签笔记
50:29
>> If if I had been paying more attention, I probably could have figured this out before it happened to some degree because uh memory densities DRAM scaling has stopped basically. And so since computers keep needing more and more memory, uh even ignoring the need for, you know, memory in AI, I I I think it's finally tipped the balance where for quite a while they there'll be uh it'll be a better business. It used to be one of the worst businesses, uh DRAM manufacturing. >> So yeah, basic you've you've taught me that SRAMM is already plateaued.
>> 如果我当时更留心一点,我大概在这事发生之前就能在一定程度上想明白,因为内存密度……DRAM 的微缩基本上已经停了。而计算机对内存的需求还在不断增加,就算不算 AI 对内存的需求,我觉得这次天平终于翻过来了,会有相当长一段时间它是门更好的生意。它以前可是最糟糕的生意之一,呃,我说的是 DRAM 制造。>> 是啊,基本上……你教过我 SRAM 早就到顶了。
便签笔记
51:08
>> Yeah. And uh but you you DRAMM has basically plateaued but the demand hasn't plateaued, right? And so it's going to be a long time. Yeah. And I I think it'll be interesting for the M companies because their history is over uh overbuild fabs and then the prices go down. They've got de a couple of decades of that. And so are they just going to say we're highly profitable. What's wrong with the way things are today? And just >> Well, fortunately competition hopefully will fix that and they'll we're highly profitable. we'd like a little bit more share and so they'll build their fab out and then the other guys will too and >> hopefully it will >> yeah but fabs are expensive and as you pointed out take a long time not only to build but to get the yield up and all the other things >> I'm looking forward to what you guys are going to say publicly about this next question [laughter] how significant will the impact of quantum computing be on the topics discussed this evening when the technology matures
>> 对。而且 DRAM 基本上也到顶了,但需求没有到顶,对吧?所以这会持续很久。是啊。而且我觉得对内存厂商来说会很有意思,因为他们的历史就是产能建过头,然后价格暴跌,这样折腾了好几十年。所以他们会不会干脆说:我们现在利润很高,今天这个局面有什么不好?就…… >> 好在竞争大概会解决这个问题,他们会说我们利润很高,但还想多要点份额,于是就去扩厂,然后别家也跟着扩,>> 但愿会这样 >> 是啊,但晶圆厂很贵,而且正如你说的,不只是建起来要很久,把良率做上去以及其他所有事情也都要很久>> 我很期待你们对下一个问题会公开说些什么 [笑声]等量子计算这项技术成熟之后,它对今晚讨论的这些话题会有多大影响?
便签笔记
52:04
>> I would say very very little so um I think quantum computing is >> I'm so glad you're here [laughter] is a is a great technology. >> Bill Deli said that >> not Dave Patterson. Okay, >> you get beat up by [laughter] Benson tomorrow. >> Yeah. So very little go >> and but the the reason is there are sort of two really good quantum algorithms, right? One is to use quantum computing to simulate quantum chemistry and the other is Shor's algorithm which is to use quantum computing to um basically factor the product of two large prime numbers and break a lot of modern cryptography. There's there's a third algorithm, Grover's algorithm, but for optimization, but it only gives you quadratic speed up, not exponential speed up. But all of these algorithms have the property that they take advantage of the fact that quantum computers are large computation, small data machines. In people's wildest dreams, they'd like to get a few thousand error corrected cubits. Um, and then you get from the exponential speed
>> 我会说非常非常小。嗯,我觉得量子计算 >> 我真高兴你在这儿 [笑声] 是一项很棒的技术。>> 这话是 Bill Dally 说的 >> 不是 Dave Patterson 说的。好吧,>> 你明天会被 Benson 收拾一顿 [笑声]>> 是啊。所以影响很小…… >> 原因是,真正好的量子算法大概就两个,对吧?一个是用量子计算来模拟量子化学,另一个是 Shor 算法,就是用量子计算去分解两个大质数的乘积,从而破解很多现代密码学。还有第三个算法,Grover 算法,用于优化,但它只带来平方级加速,不是指数级加速。但所有这些算法都有一个共同点:它们利用了量子计算机是“大计算、小数据”机器这一事实。人们最疯狂的梦想,也就是拿到几千个纠错量子比特。然后靠叠加带来的指数级
便签笔记
52:58
up of superp position, you get to do tremendous amounts of calculations on a few thousand bits. AI is the other way around. It's a large data, relatively small computation problem. So quantum computing is not going to have a measurable impact on either training or inference of for AI. >> I I I know who I'm going to call on the questions of the future. Yeah, it's I mean the AI re there we this we were 70 years into AI, right? And there was a lot of ideas that didn't work, right? It's been learn from data, the big data machine learning thing. That's the breakthrough and that's not like you said so >> the area where it will have an effect is not on computation I don't think but uh communication so uh the last uh touring award or maybe the one before it I guess uh what >> Nobel touring award for quantum oh yeah the quantum for training award you're right >> yeah uh u quantum uh transmission and cryptography >> but I thought quantum transmission was very low latency but not much data.
加速,你就能对几千个比特做海量的计算。而 AI 恰好相反,它是大数据、计算量相对较小的问题。所以量子计算对 AI 的训练或推理都不会有可测量的影响。>> 我知道以后关于未来的问题该点谁了。是啊,AI 这块……我们已经在 AI 上走了 70年了,对吧?中间有很多想法都没成,对吧?真正的突破是“从数据中学习”,是那套大数据机器学习,而那并不像你说的那样。>> 我觉得它会产生影响的领域不是计算,而是通信。所以上一届图灵奖,或者再上一届,我想是…… >> 诺贝尔奖、量子那个……图灵奖,哦对,是量子方向的,你说得对>> 对,量子传输和密码学 >> 但我以为量子传输延迟很低,但传的数据量不大。
便签笔记
54:04
Is that right? >> Uh okay. That we're out of our you don't I mean if you're sending uh a 300 bit long key you you don't need high bandwidth. >> Yeah. >> And then you can send all the rest of it >> encoded and then it doesn't matter. >> Your keys need to be longer than 300 bits. >> Yeah. [laughter] Uh this may be outside of our expertise but that's the next question. What are the national security implications associated with AI hardware maturation? I guess [clears throat] if you've worked in stuff like that, Bill, do you have any >> I'm not sure I understand the question.
是这样吗?>> 呃,好吧,这超出我们的……我是说如果你要发的是一把 300 位长的密钥,你并不需要很高的带宽。>> 是啊。>> 然后其余内容你都可以加密发送 >> 那就无所谓了。>> 你的密钥得比 300 位更长才行。>> 是啊。[笑声] 呃,下一个问题可能超出我们的专业范围了。AI 硬件走向成熟会带来哪些国家安全方面的影响?我想 [清嗓子] Bill,如果你做过这类工作,你有没有…… >> 我不太确定这个问题在问什么。
便签笔记
54:46
>> Okay. I think I don't know that one either. Uh this is a specific company. I don't know if you're going to What is your take on etched architecture? essentially etching an LLM on the chip or so so like Talis or something like that. >> Yeah. Well, Talis would be the extreme where they're planning to use ROM. Um >> so, uh these these architectures are less programmable and they're kind of building the model into the hardware, right? Well, my my understanding is that that's and >> people can correct me if I'm wrong, but etched the company started out with that thesis and then it became more and more programmable and now it's >> is it they try to make a product?
>> 好吧。我对这个也答不上来。呃,这个问的是一家具体的公司,不知道你们愿不愿意…… 你们怎么看 Etched 的架构?本质上就是把大模型刻蚀到芯片上,类似 Taalas 那种做法。>> 嗯。Taalas 算是极端做法,他们打算用 ROM。嗯 >> 所以这些架构可编程性更低,等于是把模型直接做进硬件里,对吧?嗯,我的理解是…… >> 我要是说错了大家可以纠正我,Etched 这家公司一开始是抱着这个思路的,后来变得越来越可编程,现在…… >> 他们是想做出一个产品吗?
便签笔记
55:30
>> Yeah. >> Yeah. Well, I I thought I read something in the popular press that Google was building some ship that had >> Would you like to reveal any secrets? [laughter] >> That there's research into everything. >> Oh, that's a great answer. But but I think it's it's interesting because if you if you take the talis approach if you say you know that's they're going to the extreme not even that we're going to take a particular model and totally optimized for that but we're going to take the set of weights and burn them into into ROM because that you know will wind up being you know maybe four times as dense as SRAMM um and way less dense than DRAM by the way.
>> 是啊。>> 嗯。我好像在大众媒体上看到过,说谷歌在做某种芯片…… >> 你想爆点料吗?[笑声] >> 什么方向我们都有研究。>> 哦,这回答真妙。不过我觉得这挺有意思的,因为如果你采用 Taalas 那种做法——他们走的是极端,不只是针对某个特定模型彻底优化,而是要把那一整套权重直接烧进 ROM 里,因为那样密度大概能做到 SRAM 的四倍左右,顺便说一句,比 DRAM 的密度还是低得多。
便签笔记
56:08
um you've got to have figured that there's some model that is not changing very fast because right now you know the frontier models are coming out with a new release every month or so. Um and even if you look at the various open source models there's a clever new idea every couple weeks that you know would completely break even trying to you know make it too closely matched to the model even if you made the weights programmable. >> Yeah I think we're in the early days still and programmability is critically important. Yeah, I mean I think this is what kind of what you said. This is is a completely different universe from CPUs with million line legacy code and millions of programmers. It's really small amount of code and uh not as many there and there's tremendous pressure to innovate and so that's allowing this that makes it easier not only for the hardware people makes easier for the software people or the algorithm people to innovate and deploy it. So it's it's a remarkable rate of change and this
嗯,那你就得先认定有某个模型不会变得太快,因为现在前沿模型差不多每个月就出一个新版本。而且就算你看各种开源模型,每隔几周就会冒出一个聪明的新想法,那会彻底毁掉你把硬件跟模型绑得太紧的做法,哪怕你把权重做成可编程的也一样。>> 是啊,我觉得我们还处在早期阶段,可编程性至关重要。是啊,我是说我觉得这你说的那种情况。这跟那种有着上百万行遗留代码、上百万程序员的 CPU 世界完全是两个宇宙。代码量真的很小,人也没那么多,而且创新的压力非常大,所以这就让事情变得更容易了——不只是对做硬件的人更容易,对做软件的人、做算法的人来说,创新和落地部署也更容易。所以这个变化速度是很惊人的,而
便签笔记
12边缘计算、模拟计算与 AI 设计芯片
57:05
idea that you could pick a model and get a chip out in time and how long is that lifetime and stuff? It's an interesting bet but you know it'll be interesting to watch how it works. Okay. Uh what is the right balance between centralized and distributed compute? What are the implications for edge compute hardware? Well maybe we could just change that. You know what do you get? You know, we all think about the data center. >> Well, you know, we we ship an awful lot of um products that go into autonomous vehicles. Every Mercedes above a certain level has an NVIDIA processor and NVIDIA software in it for its self-driving features. And and that's the example of if you can put it in the data center, you do because it's much more economical to deliver compute in a data center. But for reasons of latency or you you have to be able to operate reliably with a network partition, you may need to compute on the edge or in some cases you're acquiring lots and lots of data.
你可以选定一个模型、然后及时把芯片做出来,这个芯片的生命周期又有多长之类的想法?这是一场有意思的赌注,不过看它怎么发展会很有意思。好。呃,集中式计算和分布式计算之间怎样才算合适的平衡?这对边缘计算硬件有什么影响?嗯,也许我们可以换个说法。你能得到什么呢?你知道,我们都会想到数据中心。>> 嗯,你知道,我们出货了非常多用在自动驾驶汽车上的产品。某个级别以上的每一辆奔驰车里都有 NVIDIA 的处理器和 NVIDIA 的软件,用来实现它的自动驾驶功能。这就是一个例子:如果你能放在数据中心里做,你就放在数据中心,因为在数据中心提供算力要经济得多。但是出于延迟的原因,或者你必须在网络分区的情况下依然可靠运行,你可能就需要在边缘端计算;还有些情况是你在采集大量的数据,
便签笔记
58:00
You can't afford to send it all the way back to the data center, you have to do some reduction on the edge. But I think the general rule is if you can afford to do it in the data center. If one of those other factors doesn't limit you, you do it there. But if you have to do it on the edge, you do. and and it drives a different set of requirements. >> Yeah. And the Google Pixel phones have edge TPU, so it handles a lot of stuff that's relatively simple, like speech recognition is pretty simple these days compared to these other big models. And so, but if if there's something that's too big for it, uh then it yeah, it needs to go over the network if it's available. Do you talk to the people who do the computers for Whimo? Are is that are you familiar with that at all? Is that another part of the companies there?
你没法把所有数据都传回数据中心,就必须在边缘先做一些数据缩减。但我认为一般的原则是:如果你有条件在数据中心做,如果刚才那些因素没有限制你,你就在那儿做。但如果你不得不在边缘做,那就在边缘做,而这会带来一套完全不同的需求。>> 是的。而且 Google Pixel 手机上有 edge TPU,所以它能处理很多相对简单的任务,比如语音识别,跟那些大模型比起来,如今已经算相当简单了。所以,但如果有什么东西对它来说太大了,那就得走网络——前提是有网络。你们跟做 Waymo计算平台的人有交流吗?你对那边熟悉吗?那算是那边公司的另一个部门吗?是吗?
便签笔记
58:51
You know, they set up Google to be part of Alphabet and then Whimo is W. And so [laughter] we we we don't Yeah. talk that much. >> Okay. Um, since floating point precision achieves comparable AI results, will there be ever be room for analog rather than digital systems? >> Yeah, >> you guys must hear this all the time. >> What about analog for uh you you can do multiply so cooly in analog? Isn't that the future? >> Yeah. So, one of the things that when you actually build real systems for a living, um, you have to do is test them to see whether they're right or not. And if you have something that's analog or, you know, doesn't get the same result each time, it makes it really hard to test uh the device to see whether it's working correctly or not. I mean, if if you're uh probing wafers, you have to send uh test sequences and then you read back bits at tremendous speed and and they have to match one for one. So, a analog even if it can be more effective in some circumstances, uh you have this problem that you might
你知道,他们把 Google 设成 Alphabet 旗下的一部分,然后 Waymo 是 W 开头的。所以 [笑] 我们其实没怎么聊过。>> 好的。嗯,既然浮点精度已经能达到相当不错的 AI 效果,那模拟计算相对数字系统还会有机会吗?数字系统?>> 是啊,>> 这个问题你们肯定经常被问到。>> 模拟计算怎么样?呃,你在模拟域里做乘法可以做得那么漂亮,那不就是未来吗?>> 是的。所以,当你真的以造实际系统为职业时,有一件事你必须做,就是测试它们,看它们对不对。如果你手上的东西是模拟的,或者每次得到的结果都不完全一样,那就很难去测这个器件、判断它工作得对不对。我是说,如果你在做晶圆探针测试,你得送进去测试序列,然后以极高的速度读回比特,而且必须一位一位完全对上。所以模拟计算即便在某些场合可能更有效,呃,你还是会碰到这个问题:你未必每次都得到同样的答案。
便签笔记
60:21
not always get the same answer. >> And so, how do you know if the if the chip works or not? >> Yeah. Yeah. Yeah, but even if you could test it, I've seen few cases where analog actually has an advantage. We we continuously re-evaluate this because it it's a very seductive, you know, argument that, oh, you know, you you want to do a matrix multiply, we'll just, you know, put, you know, the activations in as um, you know, current and then you have, you know, a resistor and if you think of that as a conductance, you know, V is I * G. And you you've gotten your matrix multiply for free. And by the way, you can then sum a lot of those together also for free and um you wind up, you know, getting sort of free computation. Well, it it's not not free. There's no such thing as as a free lunch or a free multiply. Um it it costs something, but what's even worse is in analog, you don't really have a reliable way of storing a value to to store it for more, you know, for very short instant of time, you can put the value on a
>> 那你怎么知道这颗芯片到底能不能用?>> 对对对。不过就算你能测它,我也很少见到模拟真的有优势的情况。我们一直在反复重新评估这件事,因为这个说法非常诱人,你知道,就是说:哦,你要做矩阵乘法,那我们就把激活值当作电流输进去,然后你有一个电阻,把它看成电导,你知道,V 就是 I 乘 G。这样你的矩阵乘法就白得了。而且顺便一提,你还可以把很多这样的结果免费地加起来,最后你就相当于得到了免费的计算。可事实上它并不免费。天下没有免费的午餐,也没有免费的乘法。呃,它是有代价的,但更糟的是,在模拟域里你其实没有可靠的办法来存住一个值。要存久一点——你可以在很短的一瞬间把值放在电容上,但它会漏掉,
便签笔记
61:16
capacitor, but it's going to leak off, especially in modern semiconductor technologies. And um if if you need to move it any distance, it's very expensive to try to move it in analog. So, so typically what people do is they'll build a little array to do um you know in-memory compute analog do do a you know matrix multiply of some you know activations times weights and and then you sum the results and get the the value out and then you have to do an A tod conversion and if you look at the the fundamental requirements of energy to do that a TOD conversion it winds up limiting you to something you know if you wanted to be at say 8 bit precision to something on the order of of five terops per watt whereas with conventional digital technology ology.
在现代半导体工艺下尤其如此。而且如果你需要把它搬动一段距离,在模拟域里搬运的代价非常高。所以,人们通常的做法是搭一个小阵列来做,呃,你知道,存内计算的模拟运算,做一个矩阵乘法,把激活值乘上权重,然后把结果求和,得到那个数值,接着你就得做模数(A/D)转换。而如果你去看做那次模数转换在能量上的根本要求,最后它会把你限制在——你知道,假如你想要 8 比特精度——大概每瓦五个 TOPS 的量级;而用传统的数字技术,
便签笔记
61:57
We've built accelerators that are 100 terops per watt. So you're you're down by more than an order of >> when you include the conversion. >> If you if you can do if you can do it without converting, >> you can stay analog, you could potentially make an attractive device, but if you have to convert, you're going to lose. >> Yeah. And uh I haven't been downstairs to the museum part lately, but uh you know it starts out uh with some analog computers and there's a reason why [laughter] they stopped being used around World War II.
我们造出过每瓦 100 TOPS 的加速器。所以你要低一个数量级以上 >> 这是把转换算进去之后。>> 如果你能不做转换就完成,>> 你能一直待在模拟域里,那你有可能做出一个很有吸引力的器件,但如果你必须做转换,那你就输定了。>> 是的。呃,我最近没去楼下的博物馆区,不过你知道,它的开头就是一些模拟计算机,而它们在二战前后不再被使用,是有原因的 [笑]。
便签笔记
62:28
So >> uh this what you know in addition to uh you know the the excitement around software tools and talking to my younger colleagues a lot of excitement about using AI for hardware design and so this question is what will the next chip design software look like? Will they be radically different from what we have used for the last 30 years? This seems like a the set one of you both of you must have an opinion on this. >> Yeah. So I think the chip design software is going to be an LLM that you talk to and uh >> okay that's >> you're very quotable. That's >> I want this chip.
所以 >> 呃,除了大家对软件工具的兴奋之外,我跟一些更年轻的同事聊,大家对用 AI 做硬件设计也非常兴奋。所以这个问题是:下一代芯片设计软件会是什么样?它们会跟我们过去 30 年用的东西有根本性的不同吗?这看起来是你们两位都一定会有看法的问题。>> 是的。我觉得芯片设计软件会变成一个你可以跟它对话的 LLM,呃 >> 好的,这 >> 你这话很适合被引用。这>> 我想要这样一颗芯片。
便签笔记
63:06
>> So we should sell sell sell my ECAD stock tomorrow. >> Well, you know, I think they they have a lot of expertise and a lot of the pieces that are needed to plug that together. >> Okay. So we need to you need to capture that knowledge in your LLM. >> Yeah. And and the LM needs to know how to run the tools, right? If you're going to need to run those tools >> to do, you know, the the place and route and the verification. But I think you know to me I look at the where we wind up spending a lot of the human time in in designing a chip and it's in verification >> LLMs are really good at writing tests right you can say um make sure this works cover all the edge cases and it will write a whole test suite for you and that's where you know 75% of our labor goes >> yeah I I was going to say the same thing it it's that you know the results we presented at hot chips showed you know like 10% improvements and doing the design and stuff, but design verification there's it's the biggest part of the team is the DV team and that
>> 那我明天该把我的 ECAD 股票卖掉了。>> 嗯,你知道,我觉得他们有很多专业积累,也有把这件事拼起来所需要的很多零件。>> 好。所以我们需要——你需要把那些知识装进你的 LLM 里。>> 是的。而且这个 LLM 需要知道怎么去运行那些工具,对吧?你得让它去跑那些工具 >> 来做,你知道,布局布线和验证。但我觉得,对我来说,我会去看我们在设计一颗芯片时人力时间到底花在哪儿,答案是花在验证上 >> LLM 非常擅长写测试,对吧,你可以说:确保这个能工作、覆盖所有边界情况,它就会给你写出一整套测试集,而那正是我们大概 75% 的人力投入所在 >> 是的,我本来也想说同样的话,就是我们在 Hot Chips 上发表的结果显示,在做设计之类的事情上大概有 10% 的提升,但设计验证——团队里最大的一块就是 DV 团队,而他们
便签笔记
64:05
they need all the help they can. >> Yeah. And we're seeing bigger than 10%. We're seeing you know for average engineers a doubling of productivity and for really good engineers 10x >> what I was talking about was >> layouts circuit basically performance. >> Oh okay. Yeah. >> Oh you're talking about performance not human productivity. Okay. >> Yes. Yeah. So, but you you're talking about uh for all tasks or for uh for uh verification. >> Yeah, I'm trying to remember the benchmark study that was for was for a particular it was for a particular task.
是多少帮助都不嫌多的。>> 是的。而我们看到的提升比 10% 大得多。我们看到,对于普通工程师,生产力翻倍;对于非常优秀的工程师,能到 10 倍 >> 我刚才说的是>> 版图、电路,基本上是性能方面。>> 哦,好的。是的。>> 哦,你说的是性能,不是人的生产力。好的。>> 对,是的。那,不过你说的是所有任务,还是呃,还是只针对验证?>> 嗯,我在努力回想那个基准研究,那是针对某个特定的——是针对某个特定任务的。
便签笔记
64:34
Um you know the average came out to be about 3x. I mean it made engineers just way more productive. So I think that that's going to be the the design tools of the future and and there are a whole bunch of startups now and and again this is the the thing in AI the real the real gold miners are the people attacking a vertical. There are a whole bunch of startups attacking the uh EDA vertical. Uh let me do uh one more. Yeah, I'm sure. What do you think about uh a CPU GPU ecosystem to make space for compute in memory or memory bound workloads?
呃,你知道,平均下来大概是 3 倍。我是说,它让工程师的效率大大提高了。所以我认为那就会是未来的设计工具,而且现在有一大批创业公司,再说一次,这就是 AI 里的规律:真正淘到金的是那些切入某个垂直领域的人。有一大批创业公司正在攻 EDA 这个垂直领域。呃,让我再来一个问题。是啊,我相信。你们怎么看 CPU-GPU 生态给存内计算或访存受限的负载留出空间?
便签笔记
65:13
>> Well, what do you think about how about making the CPU closer to the GPU? You must have a comment on that. Well, [laughter] well, you know, our ours are pretty close, right? So, they uh they they sit on what we call our, you know, um chip to chip envy link. It's a it's actually, you know, technology internal that we call GRS because it's ground referenced simply. It's a single-ended very high-speed link. And and so, um you have to ask what do you get by making them close? And and we get a couple things. One is we want to provision that um link between the GPU and the CPU so that all of the and I forget what it is you know 1.8 8 terabyte per second of uh of LPDDR5 bandwidth that's attached to the uh to the the Vera CPU and and I I don't memorize these numbers that maybe but all of that CPU memory bandwidth can be routed to either of the two GPUs that's attached to it. So the link is provisioned so that link is is never a bottleneck. Then then the question is you know what do you need in terms of
>> 那,你觉得把 CPU 做得离 GPU 更近怎么样?你对这个肯定有话说。嗯,[笑] 嗯,你知道,我们的已经挺近的了,对吧?它们,呃,它们连在我们所谓的芯片到芯片的 NVLink 上。这其实是我们内部叫 GRS 的技术,因为它是地参考信号(ground referenced signaling)。它是一种单端的超高速链路。所以,呃,你得问:把它们放近了你能得到什么?我们得到了几样东西。一是我们希望把 GPU 和 CPU 之间那条链路的带宽配置到足够高,让所有的——我记不清具体数字了,大概是 1.8TB/s 的 LPDDR5 带宽,那是挂在 Vera CPU 上的,我并没有把这些数字都背下来,也许,但所有这些 CPU 内存带宽都可以路由到与它相连的两颗 GPU 中的任意一颗。所以这条链路的配置保证它永远不会成为瓶颈。接下来的问题是,你在延迟方面需要什么,呃,因为从
便签笔记
66:11
latency and um because fetching from that memory isn't necessarily a relatively high latency operation. Um the little bit that you add with the link winds up not being that material and it's only a single link. You're not going through a switch or anything like that. And then there's control interactions. But if you think about the, you know, it's basically when you launch a a kernel that you need to to do that and and that's, you know, there's enough latency with the other parts of that that making that link faster isn't that critical. It turns out we have many tens of little risk 5 CPUs on our GPU, but they're all there for little housekeeping operations that the user never sees. Um, you know, for the CPUs that the user programs, you know, the, you know, NVLink chip to chip is actually just about as close as you want it and is is kind of the right the right balance.
那块内存取数据本身就不算是低延迟的操作。呃,链路多加的那一点点延迟最后并不算重要,而且只是一条链路,你不需要经过交换机之类的东西。然后还有控制层面的交互。但如果你想想,你知道,基本上就是在启动一个 kernel 的时候需要那些交互,而那,你知道,其他环节本身就有足够的延迟,所以把这条链路做得更快也没那么关键。事实上,我们的 GPU 上有好几十个小的 RISC-V CPU,但它们都是用来做用户永远看不到的小型管理性操作的。呃,你知道,对于用户会去编程的那些 CPU 来说,芯片到芯片的 NVLink 差不多已经是你想要的最近距离了,也算是一个恰当的平衡。
便签笔记
66:56
How how about for TPUs? >> Yeah. So, I think long term the CPU that's associated with the accelerator is going to be doing mainly h housekeeping and like babysitting kind [laughter] of chores. I think the the real interesting thing that's happening now is agenic computing. So if you say you know write me a program it's got to compile it somewhere and you're not going to compile it with on a matrix uh multiplier systolic array right so you've got to have uh real CPUs with you know serious memory systems and stuff and serious networking in the same data center and >> right but in the data center it doesn't have to be that close because you're firing off a compile job >> right but I think that's that's where the the CPU really will come in is doing things like compiling uh >> yeah it >> doing the tool calls.
那 TPU 那边呢?>> 是的。我觉得长期来看,跟加速器搭配的那颗 CPU 主要会做一些管理性的、类似照看杂务[笑] 的活儿。我觉得现在真正有意思的事情是智能体计算(agentic computing)。所以如果你说帮我写个程序,它总得在某个地方编译,而你不会拿一个矩阵乘法的脉动阵列去编译,对吧,所以你必须有真正的 CPU,配上像样的内存系统之类的东西,以及同一个数据中心里像样的网络 >> 对,但在数据中心里它并不需要离得那么近,因为你只是发出一个编译任务 >> 对,但我觉得那正是 CPU 真正会派上用场的地方,做编译这类事情,呃>> 是的,它 >> 做工具调用。
便签笔记
13给下一代架构师的建议
67:54
>> Yeah, >> it doesn't. Yeah, but if you're going to do compiling cast, I I don't quite see why it it has to be close to do that because that's up to you. >> No, it doesn't have to be close. It's just in the same data center. >> Um what I was I I I was going to do a closing question, but we we have a suggestion as well. Uh basically about advice to younger people. I was thinking like high school students, college students. The one of the questions we have here, let's see if I can find it. Uh I'm 10 years old.
>> 是的,>> 它不必。是的,但如果你要做编译之类的事,我不太明白为什么它必须离得很近才行,因为那取决于你。>> 不,它不必离得近,只要在同一个数据中心里就行。>> 呃,我本来打算问一个收尾的问题,不过我们这儿也有一个建议。呃,基本上是关于给年轻人的建议。我原本想的是高中生、大学生。我们这里有一个问题,我看看能不能找到它。呃,我 10 岁。
便签笔记
68:28
What advice do you have to me to prepare for my future? I I I think your kids are all of our kids. I don't think any of us have Well, I have grandkids that old, but [laughter] I I What would you say to 10 year olds? Do you have any advice 10-year-olds? Then I'll ask you about older students. >> So, you know, um, study math and science. Um, I think that you need a good a good foundation and and there's no substitute for having a really strong, you know, um, background in mathematics and and and basic sciences.
你有什么建议能帮我为将来做准备?我觉得你们的孩子就是我们大家的孩子。我想我们当中大概没人有 呃,我有那么大的孙辈,不过 [笑] 我 你会对 10 岁的孩子说什么?你对 10 岁的孩子有什么建议吗?然后我再问你关于年纪更大一些的学生。>> 那,你知道,呃,学好数学和科学。呃,我觉得你需要一个良好的基础,而且有没有什么能替代扎实的数学和基础科学功底。
便签笔记
69:02
>> How about you, Norm? >> Uh, I think another important thing is communication skills are really important. So, uh, learning how to write well. I know uh llms can do it for you, but I I think there's there's going to be cases where you don't have one handy and you still need to communicate well. >> Yeah. Yeah. Think clearly do that. So, yeah. My last question is going to be given the incredible pace of innovation we're witnessing, what advice would you each of you offer to young computer architects and engineers entering the field today looking to make their mark in the computer architecture space? You guys have certainly made great marks in your careers. As you look back, uh, you know, what what let you make those marks and, you know, what you think you should offer? What advice would you give for this current generation?
>> 你呢,Norm?>> 呃,我觉得另一件重要的事是沟通能力真的很重要。所以,呃,学会把东西写好。我知道呃大模型可以帮你写,但我我觉得总会有些时候你手边没有大模型,而你还是得把话说清楚。>> 对。对。清晰地思考,做到这一点。所以,是的。我最后一个问题是,鉴于我们正在见证的创新速度如此惊人,你们每个人会给刚进入这个领域的年轻计算机架构师和工程师什么建议?领域的年轻人,希望在计算机体系结构领域留下自己的印记?你们几位在各自的职业生涯中确实留下了非常了不起的印记。回头看,呃,你们觉得是什么让你们做到这些的?你们觉得应该给出什么样的建议?你们会给现在这一代人什么建议?
便签笔记
69:53
>> Well, I think the world's very different today. I mean, I um I had great fun early in my career. I designed a machine at Bell Labs, you know, you know, drawing all the circuits with pencil on vellum and having a technician who did the layout, you know, for them. and and it actually all worked on first silicon without ever doing a simulation. Um, but you know, having great manual logic design skills, I think it no longer has any value at all. [laughter] Um, >> I I know a lot about punch cards. I but I think I I would have three pieces of advice for um an aspiring computer architect. The first is to master a vertical. The second is to be very broad in your understanding of computer technology. And then the third is to make AI your partner. And and I think you know the reason for the first is I think a lot of the value today is in understanding what to design not in having great you know manual logic design skills to realize it because the tools will help you know the AI will help you realize it making AI your your
>> 嗯,我觉得今天的世界很不一样。我是说,我呃,我职业生涯早期玩得很开心。我在贝尔实验室设计过一台机器,你知道,用铅笔在描图纸上画出所有电路,然后由一位技术员来做版图,你知道,帮他们做。而且实际上第一次流片就完全跑通了,中间从来没做过仿真。呃,但是你知道,拥有出色的手工逻辑设计能力,我觉得已经完全没有价值了。[笑声] 呃,>> 我我对打孔卡了解很多。不过我想我我会给一位有志成为计算机架构师的人三条建议。第一是精通一个垂直领域。第二是对计算机技术有非常广博的理解。然后第三是让 AI 成为你的伙伴。而且我觉得,第一条的理由在于,我认为今天的价值大多在于搞清楚该设计什么,而不是拥有出色的手工逻辑设计能力去把它实现出来,因为工具会帮你,AI 会帮你实现它,把 AI 变成你的伙伴。呃,但另一方面,我看到很多计算机架构师把自己局限住了,因为
便签笔记
70:48
partner. Um but then I see many computer architects who limit themselves because you know they'll ask the memory guy oh can we build a memory like this? And the guy will say no you can't do that. But if you have a really broad understanding of circuit design you realize you know he's not thinking enough out of the box. You can build a memory that that that does that. This machine at Bell Labs, by the way, had had a 3T DRAM um on it. >> Wow. >> Um all all the all the circuit guys told me, "Oh, that's not going to work."
他们会去问搞存储的人,哦,我们能不能造一个这样的存储器?那个人会说不行,你做不到。但是如果你对电路设计有非常广博的理解,你就会意识到,他的思路还不够跳出框框。你是可以造出一个能做到那件事的存储器的。顺便说一句,贝尔实验室的那台机器上就有一个 3T DRAM。>> 哇。>> 呃,所有所有搞电路的人都跟我说:“哦,那是行不通的。”
便签笔记
71:15
>> And and uh I went and I, you know, prototyped it up and and and convinced myself that it would work. We actually did did a test chip before we did the final chip to make sure it would. But um you you have to be broad enough in your understanding that you can do circuit design when needed and and things like that to sort of get outside the box and get outside the limits of the conventional thinking of what the toolbox you have available to you. Um and then I think you know you know people really need to figure out what is a uniquely human part of the human AI partnership um in in doing computer architecture and and you know practice at being good at what the human can do because there's no point in doing what the AI can do because it's just going to do it better.
>> 然后我我就去,你知道,把它做成了原型,说服了我自己它是可行的。我们实际上在做最终芯片之前先做了一颗测试芯片,以确保它能行。但是,嗯,你必须在理解上足够宽广,在需要的时候能做电路设计之类的事情,这样才能跳出框框,突破你手头那套工具所带来的传统思维的局限。嗯,然后我觉得,你知道,大家真的需要搞清楚,在人与AI的合作中,在做计算机体系结构这件事上,哪一部分是人类独有的,然后去练习,把人类能做好的那部分做好,因为去做AI能做的事情没有意义,因为它只会做得比你更好。
便签笔记
71:54
But but you have to uh be able to so far you have to be able to recognize when it's screwed up, right? So So do you have to know how to do it manually to be able to recognize when it's >> I can definitely recognize when it's screwing up, but maybe that's because I can do it manually. >> Yeah. Well, yeah. >> The other thing you can do is you can get one AI to check the other one. >> Okay. How about you, Norm? Advice. Advice to the next generation. >> Yeah. uh for computer architects it's got a long history uh and you know going back 70 years and or more and there are a lot of lessons learned along the way and I think you know the one of the best things to do is just learn about all those historical lessons because some of them apply today some don't apply today and Um, there's no point like reinventing something that didn't work the last seven or eight times people tried it.
但是,你必须能够——至少到目前为止——能够识别出它什么时候搞砸了,对吧?所以,你是不是必须知道怎么手动去做,才能识别出它什么时候……>> 我肯定能看出它什么时候搞砸了,但也许那正是因为我会手动做。>> 是啊,嗯,是的。>> 你还可以做的另一件事是,让一个AI去检查另一个AI。>> 好的。那你呢,Norm?有什么建议?给下一代的建议。>> 嗯,对计算机体系结构师来说,这个领域有很长的历史,你知道,可以追溯到70年甚至更久以前,一路走来积累了很多经验教训。我觉得,最值得做的事情之一,就是去了解所有这些历史上的经验教训,因为其中有些今天依然适用,有些则不适用。还有,嗯,没必要去重新发明那些前七八次别人尝试都没成功的东西。
便签笔记
72:59
[laughter] Okay, with that I think we invite Mark back up on the stage. [applause] >> Yeah, thank [applause] thank you so much Dave uh Norman Bill. What a great conversation. Let's give them one more round of applause. [applause] And and when you think about it, we are called the Computer History Museum, but our conversations really are about the past, present, and future. And so that's really what the museum is about. So, thank you all for coming. Uh, thank you for your questions. Thank you Mark and Mary Stevens for your support tonight.
[笑声] 好的,那么我想我们请Mark重新上台。[掌声] >> 是的,非常感谢[掌声]非常感谢Dave、Norman、Bill。这真是一场精彩的对话。让我们再一次为他们鼓掌。[掌声]仔细想想,我们虽然叫计算机历史博物馆,但我们的对话其实是关于过去、现在和未来的。这才是这家博物馆真正的意义所在。所以,感谢大家的到来。呃,感谢你们的提问。感谢Mark和Mary Stevens今晚对我们的支持。
便签笔记
73:35
And have a safe travel home. And thank these guys one more time. [applause]
祝大家回家路上平安。再一次感谢这几位嘉宾。[掌声]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

NVIDIA 首席科学家 Bill Dally 与 Google TPU 负责人 Norm Jouppi 在计算机历史博物馆对谈(Dave Patterson 主持),核心判断是:这轮 AI 芯片热潮不同于 80 年代 Lisp 机、90 年代超算的泡沫,因为有真实经济需求和极简的应用形态;但护城河已从芯片本身转向整机系统工程,软件护城河正被 AI 编码填平,数值精度的"低垂果实"已近摘完,而真正的瓶颈是能源与内存产能。

核心要点

  • 这轮热潮"能站住"的三个特征:真实需求、应用极简、无历史包袱。 Dally 对比 90 年代 DARPA 战略计算计划催生的超算热——没有真实市场、应用是上百万行的"落满灰的旧代码"(dusty decks),只能加速 1% 的核心,Amdahl 定律吞噬收益,Thinking Machines 等公司相继倒闭。而 Transformer 是个简单结构,加速几个原语就能全局提速,且用户愿意快速迭代。Jouppi 补充资本规模:Google 宣布明年 1050 亿美元 capex,纽约修一条百年漏水隧道却差 10 亿美元。
  • "卖铲子的"也在给"卖内存的"交钱。 两人自认是淘金潮里的 Levi's 和 Stanford(卖铲子的),无论谁淘到金都赚钱;但随即自嘲两家都在给内存厂商付大钱——Micron 上季度利润相当于此前 19 年之和,出货量持平、利润率飙升。
  • GPU 与 TPU 并未真正"趋同",寒武纪爆发后必有大灭绝。 面对"训练加速器已趋同演化"(大计算 die + 脉动阵列 + 环绕 HBM + 最快 SerDes)的质疑,Jouppi 说奇形怪状的创业公司会像寒武纪怪虫一样灭绝;Dally 认为 2 万英尺高度看相似,但细节差异巨大:NVIDIA 在数值格式(NVFP4 领先 MX 格式,同精度下可带来 2 倍优势)与稀疏支持(2015 年 NeurIPS 论文,Ampere 起硬件化)上领先。他类比 90 年代末上百家图形芯片创业公司最终只剩 NVIDIA 和 ATI 两家。
  • 各自羡慕对方什么。 Dally 羡慕 TPU 只有 Google 一个客户,能按自己想法设计;NVIDIA 占推理和训练市场 68% 份额,大客户的功能请求都得"至少装装样子"塞进硬件。他还喜欢 TPU 的 3D 环面网络——里面能看到他合著者 Brian Towles 的手笔(Jouppi 承认"我们把那本书研究得很仔细")。Jouppi 则强调 TPU 从第二代(2017)起就是白纸起步的超算设计,2018 年起就液冷,至今 8 年,没有功能层层叠加互相冲突的历史包袱。
  • 护城河是整个系统,创业公司缺的是"疼过"的经验。 Dally 说产品是全套硬件+软件+配置:Pascal 时代(约 2015)起的 DGX SuperPOD 让上万 GPU 开机即用,而 DOE 的 Titan、Summit、Sierra 超算硬件到位后通常要花 6 个月才能调通。两家的扩展网络是最大差异:Google 用 3D 环面+光路交换绕开故障,NVIDIA 用 Clos 拓扑+专有低延迟链路(scale-up)加常规以太网(scale-out)。Jouppi 引《数据中心即计算机》一书:这些教训不亲自建过超算不会懂。
  • 软件护城河正在被 AI 填平。 Jouppi 遵循导师 Hennessy 的忠告"能在编译期做的别拖到运行时",TPU 架构由 10 人设计,其中 2 人来自编译器团队。Dally 回忆 2006 年 CUDA 时代要移植上千个 Fortran 应用是巨大负担,而 2010 年与吴恩达合作做 cuDNN 时发现深度学习只有几个关键内核,"简直是一股清风"。但如今"给 Claude 一份库的规格,周末就能重造出来",两人一致认为软件护城河"正在被沙子填满,快能走过去了"。
  • 14 年每年 2 倍性能,工艺只贡献了 3 倍;数值精度红利接近尾声。 Dally 称 NVIDIA 自 2012 年 Kepler 以来每年翻倍,与 90 年代全靠工艺的微处理器时代形成对比;从 FP32 到 NVFP4,精度轴还剩"几圈",之后要靠稀疏、电路、局部性、更好的模型来提升每瓦 token 数,他估计手头的想法够撑四五代(8–10 年)"然后我就可以退休了"。Jouppi 讲了 BF16 的由来:Jeff Dean 早期用 FP32 计算后直接截断低 16 位——数值分析师听了会像"指甲刮黑板",但确实管用,让 TPU 能无痛复现 CPU 结果,验证正确后再逐步采用 FP8/FP4。
  • 芯片周期 2.5–3 年 vs 模型思路每 3 个月一变,只能"瞄准鸭子前方"。 Jouppi 说不可能为特定模型定制芯片,只能把矩阵和向量运算做好。Dally 指出 OpenAI 在 Hot Chips 发表的 Jalapeño 芯片是为一个模型设计的特例;通用供应商必须面对 DeepSeek 的 MLA、混合注意力(三层状态空间配一层全 n² 注意力)、稀疏 top-k 注意力等各种变体对内存带宽/算力配比的不同需求。Jouppi 认为 MoE 是最具颠覆性的变化,因其对互连带宽和延迟要求极高——这直接催生了最新 TPU 的 Dragonfly 拓扑。对"把模型烧进 ROM"的 Etched/Taalas 路线,Dally 指出前沿模型每月一版、开源社区每两周一个新点子,锁定模型的赌注风险巨大。
  • MLPerf 没能成为 AI 时代的 SPEC。 Patterson 观察 Hot Chips 上各家像 80 年代用 MIPS 指标互相指责"别人都在撒谎"一样,只报 tokens/秒却不说模型多大。Dally 说 MLPerf 失败的两个原因:几乎没人提交(创业公司嫌麻烦、按规则跑就不好看,宁可挑一个数据放 PPT),且没有回答业界真正关心的每瓦/每美元 token 数;反而 SemiAnalysis 的 InferenceMAX 因结果可复现(必须提供 GitHub/HF 代码库)正在流行。Jouppi 补充 MLPerf 需要一个不小的专职团队。
  • 最大瓶颈:能源与内存产能,而非芯片本身。 现场投票"能源可用性"居首。Jouppi 说吉瓦级数据中心的输电线难获许可,超大规模厂商转向场内发电,Google 尽量用风光;Dally 说燃气轮机已售罄 5 年,有公司把飞机引擎拆下来当发电机,锂电池每千瓦时太贵无法跨越 20–100 小时的间歇缺口,热储能更有希望;核电按平准化成本仍打不过天然气(甚至加碳捕集后也是)。内存方面,Jouppi 承认 DRAM 微缩基本停滞而需求未停,DRAM 从"最糟的生意之一"翻身,但内存厂商几十年过度建厂—价格崩盘的记忆可能让它们不急于扩产。

结论与值得注意的细节

  • 量子计算对 AI"影响极小"(Dally 语,Patterson 打趣"幸好是你说的不是我"):量子机是"大计算、小数据"(几千个纠错量子比特),AI 是"大数据、小计算",两者结构相反;量子的价值在通信/密钥分发,不在算力。
  • 模拟计算再次被否决:Jouppi 从测试角度——每次结果不同就无法晶圆探测验证;Dally 从能耗角度——模拟矩阵乘看似免费,但存值会漏电、搬运昂贵,加上 ADC 转换后 8 位精度约 5 TOPS/W,而数字加速器已做到 100 TOPS/W,差一个数量级以上。"博物馆楼下的模拟计算机二战前后就被淘汰了,是有原因的。"
  • 未来的 EDA 是"一个你跟它说话的 LLM"(Jouppi)。两人都指出芯片设计 75% 的人力在验证,LLM 写测试用例极强;Dally 称普通工程师生产力翻倍、顶尖工程师 10 倍。Dally 亲身经历:写好规格交给 Claude,修了几处规格错误后得到"相当不错的设计"——它精确产出了他要的东西,但那不是对的东西,"你得学会写好规格"。
  • CPU 的角色:Jouppi 认为加速器旁的 CPU 长期只做"看家/带孩子"的杂活,真正需要 CPU 的是智能体计算——编译、工具调用——但不必紧邻,同一数据中心即可;Dally 透露 NVIDIA GPU 里有几十个小 RISC-V 核做用户看不见的管理工作。
  • 给年轻人的建议:Dally 三条——精通一个垂直领域、对计算技术理解要广(别被"内存专家说做不到"限制住,他在贝尔实验室曾顶着所有电路工程师的反对做出 3T DRAM)、把 AI 当伙伴并练习人类独有的部分,同时得能看出 AI 什么时候搞砸了("也许是因为我会手工做");Jouppi 两条——学好写作与沟通,以及学习计算机架构 70 年的历史教训,"别去重新发明前七八次都没成功的东西"。给 10 岁孩子:学数学和科学。
  • Dally 最后纠正:这不该叫硅淘金潮,应叫 AI 淘金潮;真正的淘金者是攻打垂直领域的人(包括正涌入 EDA 领域的创业公司),风险则在深度伪造(应默认未认证的图像为假)、网络攻击(80% 以上仍是社会工程)、病原体设计与就业结构转变。
核心句型 · 10
1. X is working on everything they apply it to
“AI is working on everything they apply it to, like Jeff Dean said.”
用 apply … to 的定语从句强调「凡用之处皆有效」,适合描述一项技术的普适性;可仿写 The method works on every dataset we throw it at。
2. the demand for … is insatiable
“The demand for more tokens, more flops, more AI is insatiable”
insatiable 搭配 demand/appetite,连续三个 more 形成递进排比,表达需求无止境;写市场分析时可直接套用。
3. It turns out (that) …
“It turns out there was no intense economic demand.”
引出与预期相反的事实,语气中性,常用于回顾历史或实验结果;比 actually 更书面,适合叙述「后来才发现」。
4. … would come up and bite you
“Amdahl's law would come up and bite you because it was the other 99% of the code”
拟人化表达「某规律最终会反噬你」,形象且口语;可用于描述被忽视的约束条件最终显现。
5. the reason why A is that B
“The reason why Nvidia and Google have both been successful is that a bunch of these other ideas weren't as capable”
正式的因果解释句式,主语用 the reason why 引出结果,that 从句给原因;注意用 that 而非 because。
6. don't put off until X what you can do at Y
“Don't put off until runtime what you can do at compile time”
改编自谚语 Don't put off until tomorrow what you can do today,what 从句作宾语后置;可套用于任何「早做优于晚做」的原则。
7. all the low hanging fruit has been picked and we need to climb higher up the tree
“All the low hanging fruit has been picked and we need to climb higher up the tree to find those 2x fruits”
把 low hanging fruit 这一习语延展成完整比喻,形象说明容易的优化已用尽;演讲中延展习语是让听众记住观点的常用技巧。
8. you've got to aim ahead of the duck
“As Norm pointed out, you've got to aim ahead of the duck”
打猎类比,表示要预判目标的移动而不是瞄准现在;适合描述面向未来的规划与产品设计。
9. There's no such thing as a free X
“There's no such thing as a free lunch or a free multiply.”
固定句式否定「免费」的可能,末尾替换名词可制造幽默与专业感;此处把 lunch 换成 multiply 呼应技术语境。
10. there's no point (in) doing what … can do
“There's no point in doing what the AI can do because it's just going to do it better”
there's no point (in) + 动名词表示「做某事没意义」,what 从句作宾语;适合表达分工或取舍的建议。
词汇精讲 · 134 · 按出现顺序
Cambrian explosion n. phr. 0:14
寒武纪大爆发;喻指某领域短期内种类激增
emeritus /ɪˈmerɪtəs/ adj. 1:07
荣休的(原文误拼为 ameritus)
curators /kjʊˈreɪtərz/ n. 1:07
策展人
precursor /prɪˈkɜːrsər/ n. 2:43
先驱,前身
flow control n. phr. 2:43
流量控制(网络术语)
latency /ˈleɪtənsi/ n. 2:43
延迟,时延
skeptical /ˈskeptɪkl/ adj. 3:50
怀疑的
hype /haɪp/ n. 3:50
炒作,大肆宣传
prefetch /ˌpriːˈfetʃ/ v./n. 3:50
预取(提前把数据取入缓存)
prestigious /preˈstɪdʒəs/ adj. 5:40
有声望的
insatiable /ɪnˈseɪʃəbl/ adj. 6:53
无法满足的,贪得无厌的
dusty decks n. phr. 6:53
陈年老代码(源于打孔卡时代的行话)
wound up phr. v. 7:49
最终落得(某种结局);wind up 的过去式
going bust phr. 7:49
倒闭,破产
kernel /ˈkɜːrnl/ n. 7:49
核心计算例程;(此处)程序中最耗时的关键部分
primitives /ˈprɪmətɪvz/ n. 7:49
基本原语,底层基本操作
frenzy /ˈfrenzi/ n. 7:49
狂热
capex /ˈkæpeks/ n. 8:35
资本支出(capital expenditure 缩写)
prospectors /ˈprɑːspektərz/ n. 9:32
探矿者,淘金者
verticals /ˈvɜːrtɪklz/ n. 9:32
垂直行业,细分领域
strike it rich idiom 9:32
一夜暴富
through the roof idiom 10:05
高得离谱,暴涨
convergent evolution n. phr. 10:05
趋同进化
lineages /ˈlɪniɪdʒɪz/ n. 10:05
谱系,世系
systolic /sɪˈstɑːlɪk/ adj. 10:05
脉动(阵列)的;数据像心跳一样节律性流经处理单元
implosion /ɪmˈploʊʒn/ n. 11:13
内爆,急剧崩塌
appendages /əˈpendɪdʒɪz/ n. 11:13
附肢,附属物
extinction event n. phr. 11:13
灭绝事件
transcendental functions n. phr. 12:14
超越函数(指数、对数、三角函数等)
nuance /ˈnuːɑːns/ n. 12:43
细微差别
followed suit idiom 12:43
跟进,照做
sparsity /ˈspɑːrsəti/ n. 12:43
稀疏性(原文误拼 sparity)
entertain /ˌentərˈteɪn/ v. 13:34
(此处)考虑、认真对待(请求)
envious /ˈenviəs/ adj. 14:14
羡慕的
fond of phr. 14:14
喜爱
weeding out phr. v. 14:56
淘汰,剔除
survival of the fittest idiom 14:56
适者生存
clean sheet of paper idiom 15:31
从零开始,白纸设计
prescient /ˈpreʃənt/ adj. 16:59
有先见之明的(原文误拼 precient)
moat /moʊt/ n. 17:35
护城河;喻指竞争壁垒
captive /ˈkæptɪv/ adj. 17:35
自有的,专属的
bringing it up phr. v. 18:32
(系统)调通、启动起来
availability /əˌveɪləˈbɪləti/ n. 18:32
可用性(系统正常运行时间比例)
proprietary /prəˈpraɪəteri/ adj. 19:26
私有的,专有的
overhead /ˈoʊvərhed/ n. 19:26
开销(额外的性能成本)
the hard way idiom 19:53
通过吃苦头、走弯路(学到)
netlist /ˈnetlɪst/ n. 20:45
网表(芯片电路连接的描述文件)
put off phr. v. 21:40
推迟
compile time n. phr. 21:40
编译期(相对 runtime 运行期)
integral /ˈɪntɪɡrəl/ adj. 22:32
不可或缺的,构成整体的
port /pɔːrt/ v. 22:32
移植(软件到另一平台)
a breath of fresh air idiom 23:38
令人耳目一新的事物
profile /ˈproʊfaɪl/ v. 24:23
做性能剖析
fuse kernels phr. 24:23
算子融合(合并多个计算核以减少访存)
Pollyannaish /ˌpɑːliˈænɪʃ/ adj. 24:56
盲目乐观的(原文误拼 polyianish)
honed /hoʊnd/ v. 24:56
磨砺,精细打磨
turning … loose phr. 24:56
放开手让……去干
squeezed /skwiːzd/ v. 25:54
榨取(性能)
exponent /ɪkˈspoʊnənt/ n. 25:54
(浮点数的)指数位
shave bits off phr. 25:54
削减比特位
heyday /ˈheɪdeɪ/ n. 26:52
全盛期
low hanging fruit idiom 26:52
唾手可得的成果
axes /ˈæksiːz/ n. 26:52
维度,轴(axis 的复数)
truncating /ˈtrʌŋkeɪtɪŋ/ v. 28:17
截断
fingernails on a chalkboard idiom 28:17
令人极不舒服(如指甲刮黑板)
leaderboards /ˈliːdərbɔːrdz/ n. 30:04
排行榜
in volume phr. 30:04
批量地,大规模地
provisioning /prəˈvɪʒənɪŋ/ n. 31:30
资源配置
hybrid /ˈhaɪbrɪd/ adj. 32:30
混合的
aim ahead of the duck idiom 33:01
打提前量(射击移动目标要瞄准其前方)
anticipate /ænˈtɪsɪpeɪt/ v. 33:01
预判,预先考虑
disruptive /dɪsˈrʌptɪv/ adj. 33:47
颠覆性的
mixtures of experts n. phr. 33:47
混合专家模型(MoE)
metric /ˈmetrɪk/ n. 34:29
衡量指标
rallied around phr. v. 35:07
团结拥护
level playing field idiom 36:07
公平竞争环境
cut through the BS phr. 36:07
戳穿废话、直指要害(口语,BS 为粗语缩写)
cherrypick /ˈtʃeripɪk/ v. 36:40
挑选对自己有利的(数据)
reproducible /ˌriːprəˈduːsəbl/ adj. 36:40
可复现的
catching on phr. v. 37:32
流行起来
open-ended /ˌoʊpən ˈendɪd/ adj. 38:10
开放式的
geopolitical /ˌdʒiːoʊpəˈlɪtɪkl/ adj. 38:10
地缘政治的
recovering /rɪˈkʌvərɪŋ/ adj. 39:44
(幽默)「戒掉……的」,如 recovering educator 自嘲已脱离教职
move up the ladder idiom 40:19
向上晋升,转向更高层级的工作
minions /ˈmɪnjənz/ n. 40:19
喽啰,小跟班
flip side n. phr. 40:19
另一面,反面
provenance /ˈprɑːvənəns/ n. 41:09
来源,出处(原文误拼 provenence)
nefarious /nɪˈferiəs/ adj. 41:09
邪恶的,不法的
pathogen /ˈpæθədʒən/ n. 41:09
病原体
easing that transition phr. 41:09
缓和过渡
reskill /ˌriːˈskɪl/ v. 42:03
再培训技能
take for granted idiom 43:02
视为理所当然
bottleneck /ˈbɑːtlnek/ n. 43:56
瓶颈
permits /ˈpɜːrmɪts/ n. 44:53
许可证
hyperscalers /ˈhaɪpərˌskeɪlərz/ n. 44:53
超大规模云厂商
clobbering /ˈklɑːbərɪŋ/ v. 44:53
(口语)重创,搞砸
stranding /ˈstrændɪŋ/ v. 46:09
使滞留、困住
carbon footprint n. phr. 46:57
碳足迹
ride through phr. v. 46:57
撑过(供电中断等时段)
around the clock idiom 47:45
全天候
levelized cost n. phr. 48:18
平准化成本(全生命周期折算的单位成本)
carbon sequestration n. phr. 48:18
碳封存
pumped hydro n. phr. 48:18
抽水蓄能
fab /fæb/ n. 49:54
晶圆厂(fabrication plant)
tipped the balance idiom 50:29
使天平倾斜,改变局面
yield /jiːld/ n. 51:08
良率
quadratic /kwɑːˈdrætɪk/ adj. 52:04
平方的,二次的
error corrected qubits n. phr. 52:04
纠错量子比特(原文误拼 cubits)
superposition /ˌsuːpərpəˈzɪʃn/ n. 52:58
(量子)叠加态
etching /ˈetʃɪŋ/ v. 54:46
刻蚀;此处喻把模型固化进硅片
burn them into ROM phr. 55:30
把(数据)烧录进只读存储器
network partition n. phr. 57:05
网络分区(网络断裂成互不连通的部分)
probing wafers phr. 58:51
晶圆探针测试
seductive /sɪˈdʌktɪv/ adj. 60:21
诱人的
conductance /kənˈdʌktəns/ n. 60:21
电导
leak off phr. v. 61:16
(电荷)泄漏
in-memory compute n. phr. 61:16
存内计算
order of magnitude n. phr. 61:57
数量级
place and route n. phr. 63:06
布局布线(芯片物理设计步骤)
edge cases n. phr. 63:06
边界情况
single-ended /ˌsɪŋɡl ˈendɪd/ adj. 65:13
单端的(信号相对地参考,而非差分)
provision /prəˈvɪʒn/ v. 65:13
配置(资源、带宽)
material /məˈtɪriəl/ adj. 66:11
(此处)重要的,有实质影响的
housekeeping /ˈhaʊskiːpɪŋ/ n. 66:11
日常维护性工作
babysitting /ˈbeɪbisɪtɪŋ/ n. 66:56
照看;此处喻低价值的辅助工作
agentic /eɪˈdʒentɪk/ adj. 66:56
智能体式的(AI 自主执行多步任务)
no substitute for phr. 68:28
无可替代
make their mark idiom 69:02
留下印记,做出成就
vellum /ˈveləm/ n. 69:53
描图纸,牛皮纸
first silicon n. phr. 69:53
首次流片
aspiring /əˈspaɪərɪŋ/ adj. 69:53
有志于……的
out of the box idiom 70:48
跳出常规思维
screwed up phr. v. 71:54
(口语)搞砸
reinventing /ˌriːɪnˈventɪŋ/ v. 71:54
重新发明(已有之物)
理解自测 · 11 题
1. Bill Dally 认为本轮 AI 芯片热潮与以往有哪三个根本区别?

三点:一是真实且旺盛的经济需求,AI「用在什么上都管用」,对算力的需求无止境;二是应用相对简单,Transformer 只有少数关键算子,容易看清该造什么硬件;三是用户愿意快速迁移,没有「陈年老代码」的包袱。他在「这次淘金热为何不同」一节中对比了 90 年代 DARPA 资助的超算热:那时没有真实市场,应用是几十万到上百万行的旧代码,只能加速一小部分,阿姆达尔定律使整体收益有限,大量公司如 Thinking Machines 最终倒闭。

2. Norm Jouppi 用了哪两个数字来说明当前投资规模的反常?

他对比了纽约通往新泽西的百年老隧道工程 10 亿美元的资金缺口,与 Google 宣布的明年 1050 亿美元资本支出,其他公司投入也相仿。这一对比出现在开头讨论「淘金热」定义时,意在说明 AI 基础设施投入已远超传统公共基建。他随即指出结论:会有赢家和输家,所有人都在尽可能快地竞跑。Patterson 后来还补充了美光一季度利润等于此前 19 年之和的例子,说明内存环节的价格暴涨。

3. 关于 NVIDIA 14 年来性能提升的来源,Dally 给出了什么数据?

自 2012 年 Kepler 起,NVIDIA 基本做到每年 2 倍的性能提升,14 年累计约上万倍,但其中只有 3 倍来自制程工艺,其余全部来自架构设计、数值格式(从 FP32 到 NVFP4)、稀疏性支持等创新。这与 90 年代微处理器全靠制程进步形成对照。他在「精度削到头后靠什么翻倍」一节中承认数值精度这条路已接近尽头,但还有稀疏化、电路、局部性、更好的模型等维度,估计好点子还能支撑四五代。

4. 为什么 Jouppi 认为 GPU 与 TPU 的相似并非「趋同进化」的简单结果?

他借用寒武纪大爆发之后的大灭绝来解释:许多创业公司提出了各种奇特设计,但因能力不足而「灭绝」,NVIDIA 和 Google 的成功是选择压力筛选的结果,而非设计者互相靠拢。Dally 进一步补充:应用确实决定了基本形态,GEMM 需要矩阵单元、softmax 需要向量单元,所以幸存者必然都具备这些;但在数值格式、稀疏性、互连拓扑等大量细节上两家差异很大。结论是相似来自共同的应用约束和淘汰机制,而非趋同。

5. Dally 为什么说 CUDA 软件护城河「已被显著削弱」?他的推理链是什么?

他先承认软件极其关键:NVIDIA 在 CUDA 生态投入巨大,模型上线后半年内靠剖析工具和算子融合还能再翻倍性能。但他随即指出,如今只要有库的规格说明,让 Claude Code 之类的工具去重建并不困难,「跟它说一声,过个周末就行」。推理链是:护城河的价值在于积累难以复制,而 AI 编码工具把复制成本压到极低,因此多年打磨的库仍有便利优势,却不再构成壁垒。Jouppi 附和说护城河「正在被沙子填满,快能走过去了」。

6. 为什么两位嘉宾都认为硬件不能为某个特定模型专门设计?

核心是时间错配:Jouppi 指出从想法到量产需要 2.5 到 3 年,而机器学习研究者约每 3 个月冒出一个新想法;Dally 补充前沿模型每月一个新版本,开源社区每几周一个新技巧。因此像 Taalas 那样把权重烧进 ROM、或 Etched 早期把 Transformer 固化进芯片的路线,出货时很可能已过时,即便密度有 4 倍收益也得不偿失。两人的结论是硬件只能押注通用的矩阵、向量运算并把它们做好,同时「瞄准鸭子前方」预判两三年后的需求,可编程性至关重要。

7. Patterson 为何把 Hot Chips 上的「每秒 token 数」类比为 80 年代的 MIPS 指标?MLPerf 又为何未能像 SPEC 那样成功?

80 年代 RISC 论战中各家用「每秒百万指令」比拼,但怎么定义、跑什么程序全由厂商自定,指标失去意义,最终催生了厂商合作的 SPEC 基准。Hot Chips 上的「tokens/s」同样不说明模型大小和条件,重演了这一问题。Dally 给出 MLPerf 不成功的两个原因:一是提交需要一支不小的全职团队,创业公司宁可挑单项漂亮结果放上幻灯片;二是业界真正关心的是每瓦或每美元 token 数,MLPerf 未直接度量。SemiAnalysis 的 InferenceMAX 因可复现且直击这一指标而更受欢迎。

8. Dally 反驳模拟计算的论证分几层?关键数据是什么?

三层。第一,Jouppi 从制造角度指出模拟电路结果不可精确复现,晶圆测试要求逐位匹配,无法判断芯片是否合格。第二,Dally 指出「免费矩阵乘法」是幻觉:模拟值无法可靠存储(电容会漏电)也无法低成本远距传输。第三,决定性数据是模数转换的能耗下限:8 位精度下模拟存内计算约 5 TOPS/W,而 NVIDIA 数字加速器已达 100 TOPS/W,相差 20 倍。结论:除非全程不做转换,否则模拟必输。Jouppi 还补充博物馆里的模拟计算机在二战后被淘汰「是有原因的」。

9. 如果有人反驳说「Google 有 DeepMind 前沿研究者,硬件必然比 NVIDIA 更贴合模型」,Jouppi 会如何回应?

他在讨论中已直接回应了这一假设。Patterson 提出「能接触到推动前沿的人是理论上的大优势」,Jouppi 承认「某种程度上是」,但强调硬件周期 2.5 到 3 年、模型思路 3 个月一变,前沿研究者的具体建议在芯片出货时多半已过时,所以只能为通用算子设计。Dally 补充,真正能贴合模型的是 OpenAI 那种只为一个模型设计芯片的公司,而 Google 和 NVIDIA 都要支持多种模型。因此两人都会说:贴近研究者有助于方向判断,但不会转化为决定性的硬件优势。

10. 把 Dally「护城河在整个系统」的论点放到今天的 AI 芯片创业公司身上,还成立吗?请结合讲者的证据判断。

成立,且讲者给出了正反两面证据。正面:DOE 超算 Titan、Summit 硬件到位后仍需半年调通,而 DGX SuperPOD 的标准配置能让上万块 GPU 开机即用;Jouppi 也说系统教训只能靠亲手建大机器吃苦获得,创业公司缺这一课。反面:Dally 承认软件层的壁垒被 AI 编码工具削弱。综合看,硬件系统集成(互连、供电、液冷、可靠性)仍是难以速成的护城河,软件护城河则在变浅,这与他预言 AI 芯片创业潮会像 90 年代图形芯片那样收敛到两三家幸存者一致。

11. Dally 说「AI 能做的事它都会做得更好,人应专攻人类独有的部分」,Patterson 追问「不会手动做还能识别 AI 犯错吗」。这一分歧对英语学习者的 AI 使用有什么迁移意义?

Dally 的分工原则是找出人机合作中人类独有的部分并专门练习;Patterson 的追问指出识别错误的能力可能依赖于自己会做,Dally 回答「我能看出它搞砸,但也许正因为我会手动做」,并提出让一个 AI 检查另一个 AI。Jouppi 则强调学历史教训、写作能力,因为「总有手边没有大模型的时候」。迁移到语言学习:AI 可以翻译、润色,但判断译文是否准确、语气是否合适仍需学习者自己具备基础能力;Dally 关于「学会写好规格说明」的经历也说明,把需求表达清楚本身就是人类独有且需要练习的技能。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.127He won a Nobel here for AlphaFold. Then he left. - John Jumper 下一期 · NO.129 →Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY 内容仅供学习 · thesophielab.com