视频库 / NO.129
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

节目发布 2026-06-30 · Sequoia Capital
迪伦·帕特尔 肖恩·马奎尔 黄共宇
本期追问 · 点击跳到视频对应位置
23:16 AI 三年来效率翻百倍,功劳在芯片还是在模型?32:16 模型和芯片互相定制后,还能说谁比谁更好吗?43:52 专用 AI 芯片会不会只是一个局部最优的陷阱?50:46 AI 数据中心的高杠杆狂建,何时会变成泡沫?
归入 Ⅰ·05 整体能大于部分之和吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文是红杉资本合伙人肖恩·马奎尔(Sean Maguire)与索尼娅·黄(Sonya Huang)在 SemiAnalysis 办公室对其创始人兼首席分析师迪伦·帕特尔(Dylan Patel)的一次访谈。帕特尔 2020 年以博客起家,如今把公司做成了覆盖半导体供应链与 AI 基础设施的头部研究机构。谈话从他在加油站长大的童年讲起,一路谈到推理基准、软硬件协同设计、英伟达与 TPU 之争、算力紧缺与 Neocloud 的兴起。本文依据现场录音编译整理,仅删去口语枝节与寒暄,论证与细节悉数保留。

加油站里的第一个神经网络

马奎尔: 我们此刻在 SemiAnalysis 的办公室,对面是迪伦·帕特尔。我是红杉的肖恩,旁边是我的合伙人索尼娅·黄。你做成的事相当惊人。五年前,半导体在西方根本不是什么性感的行业,在东方也许是,但西方的人差不多已经把它忘了。你没有忘,而且押得很重,做出了这个领域大概最顶尖的研究公司,从极其技术的细节到供应链再到宏观图景,一直在给全世界补课。外界传言 SemiAnalysis 最近营收过了一亿美元,我不知道准不准,但不管数字是多少,你们都在狂飙。

帕特尔: 传言的准确度,取决于信息来源的准确度。

马奎尔: 还有传言说你可能要做一支风险投资基金。我在圈子里经常听到有人想跟 SemiAnalysis 沾上关系,你已经建起了一个受信任的品牌,所以你无论做什么都会奏效。这显然只是你旅程的开端,恭喜。但这一切是怎么发生的?先从背景讲起,你是怎么走到今天的?

帕特尔: 我在小生意人家庭里长大。父母开一家汽车旅馆,我们就住在旅馆里,马路对面是我们家的加油站。我常开玩笑说,我训练的第一个神经网络,就是根据走进加油站的人的长相和族裔,判断他要买哪种烟。香烟全都陈列在货架顶上,我个子太矮够不着,况且那个年纪卖烟本来也不合法,但反正我得把踏脚凳挪到正确的位置。

马奎尔: 我的第一份工作也是在法定年龄之前,这种经历很好。

帕特尔: 我可没有工资,家族生意嘛。

马奎尔: 一样。

帕特尔: 比如一位满头卷发的白人老太太走进来,我就把踏脚凳挪到骆驼牌那边。不同年龄、不同人群、不同职业、不同族裔,我都提前把凳子挪好。我说这是我训练的第一个神经网络,因为如果等他们开口告诉我,我再挪凳子再爬上去,那就慢了,不如提前就位。薄荷烟、100 毫米细支烟,这些我都门儿清。

真正的起点是我八岁生日。我的生日在五月,四月份 Xbox 360 发布了。父母问我生日想要什么,我没要生日礼物,而是说想要圣诞礼物。我们家过圣诞节,但我当时觉得他们绝不可能在圣诞节送我 Xbox 360,所以生日那天我就预先要了这份圣诞礼物。结果圣诞节真的收到了。几个月后,住在阿拉巴马的表弟要来我家过春假,他们家也开汽车旅馆。他年纪在我和哥哥之间,哥哥偏运动型,对 Xbox 不太上心,偶尔玩玩。可我很想让表弟觉得我酷,在电话里吹了好多次「我有 Xbox」。然后 Xbox 坏了。那是一个著名的硬件缺陷,叫「死亡红环」(red ring of death)。长话短说,我把机器拆开,把温度传感器短接,就修好了。在这之前我试了很多别的招,全都不管用。这就是我进入硬件世界的方式,潘多拉的盒子打开了。

到十二岁,我已经泡在各种论坛上,读得多、发帖也多。那正是 Reddit 吞掉所有其他论坛的时候,于是我成了安卓、苹果、谷歌以及硬件板块的版主,同时盯着英特尔、英伟达、AMD 这些板块,还有装机板块。所有这些论坛我都在看、在读、在发帖,其中一些我还在做版主。我看着智能手机从极简单的东西一路狂奔,在架构上很多方面比 PC 更先进;GPU 也一样,一条条评论追着看。而且我始终带着经济的视角,因为我是小生意家庭出身,总在看经济账。有一段时间,网上那些「胡子大叔」都爱 AMD 的 GPU,我自己也因为性价比买过一块 AMD,但要论技术上谁更强,我总是说,不,英伟达更好,因为他们用更小的芯片做出更高的性能和更好的能效,毛利率也更高。我那时就常讲英伟达在 GPU 领域的利润率比 AMD 高,很有意思。

黄: 那时你才十二岁。

帕特尔: 十二岁开始做版主,整个青春期和高中阶段都在做这个。

马奎尔: 除了半导体,还有别的奇怪爱好吗?

帕特尔: 打了海量的《星际争霸 2》,一度是北美天梯的宗师段位。

马奎尔: 相当认真。所以你在好几件事上都到了痴迷般的程度。

帕特尔: 痴迷是好事。

黄: 成绩怎么样?

帕特尔: 还行,大部分是 A,但我觉得无聊或者不喜欢的课就不行,比如西班牙语的分数就不怎么样。顺便说一句,我现在西班牙语说得很流利,所以这事挺傻的。

黄: 也许这就是你没拿到好成绩的原因。

帕特尔: 公平地说,我是后来才学的西班牙语。总之成绩过得去,对亚裔父母来说够用了。比大多数同学好,但没有拼命去刷全 A。

离职、丧亲与匿名博客的诞生

黄: 所以你完全是互联网教出来的学生,专业知识是这样养成的。你在什么时候决定创办 SemiAnalysis?创业以来最大的意外是什么?

帕特尔: 我上了大学,拿了几个跟半导体无关的学位,然后在一家小型量化风险公司做了两年量化。接着发生了一连串事情。一是我的奖金被人截胡了。我利用市场里的一个风险漏洞,给公司赚了好几百万美元的无风险收入,我估计超过一千万,结果别人抢了功劳。后来虽然补发了,但我跟公司之间的社会契约已经破裂。二是我的祖父母跟我们一起住在汽车旅馆里,我跟他们非常亲。祖母得了痴呆症,认不出我了,后来从楼梯上摔下来,出了悲惨的意外,去世了。这些都发生在 2020 年初。另外还有些感情上的事。几件事叠在一起,我非常低落。

然后新冠来了。哥哥说,你来跟我住吧。他在纳什维尔,我就过去了。我们当时想,封城也就几周,你在我这儿住着,过去了再回家。典型的立 flag。封城持续了很久。在哥哥家住了几个月,我不知道自己在干什么,一切按他的规矩来,他和当时的未婚妻、现在的妻子都在,我基本上得踮着脚过日子。但我已经不在乎那份工作了,于是在网上发帖比平时更勤。我一直是个重度发帖者,也一直炒股,做空新冠、做多新冠都赚了不少,半导体短缺也是那时候的事。总之,我对发帖入了迷。

后来我在网上跟人吵架,对方把我人肉了,公开了我匿名账号背后的身份。我当时吓坏了,停了三周没发帖。然后我想,我在干什么?我为什么要在乎?于是我以真名开始发。我以前也有过博客,这回做了一个正经的,叫 SemiAnalysis。二十四岁生日那天,我发了两篇文章。它当时还不是一份通讯,但反响非常大,因为不再是匿名账号,而是真名实姓,而且我在那两篇上下的功夫远超平常,不是随手发帖,是认真写博客。你们现在还能翻出来看,谈不上多好,但在当时是网上能找到的关于半导体最好的内容。我就一直写、一直写,开始接到很多咨询业务。

2020 年我还处在崩溃状态,不知道自己想要什么,于是把家当打包,开着皮卡,买了个能架在车斗上的帐篷和一张气垫床,跑遍全美的国家公园。一周里两三四天住在随便哪家汽车旅馆,把房价砍到一晚三十美元,在房间里工作。周末就读书,经常是教科书,在某个国家公园里或者徒步时听有声书,关于半导体、关于 AI,关于我在乎的一切。那六个月我把自己的知识水平拉高了一大截。整个过程我都是一个人,一直在写博客。所有人都问,迪伦你到底在干什么?

马奎尔: 那是星链之前,还是星链最早期?

帕特尔: 星链之前。所以确实很像「你到底在干什么」。之后我去拉美游历了一年,先是跟朋友,后来跟前女友。然后 2021 年底、2022、2023、2024 年,我从 2020 年年中起就一直没有固定住所,但我跑遍了全世界的行业会议,一年四十多场,不管它处在供应链的哪个环节。我看到一场就说,这个有意思,去。我参加第一场会议时就惊了,你能直接跟专家聊,他们也愿意跟你聊,你兴奋得不得了。半导体行业里全是老头子,他们很少见到对这个行业真正兴奋的年轻人,所以特别乐意讲。你只需要开口问。

一年四十场会议怎么学供应链

马奎尔: 供应链里有没有哪个环节,或者哪场会议,特别改变了你对半导体世界的看法?或者你当时觉得、现在也觉得被严重低估的?

帕特尔: 展会和学术会议的差别非常大。我玩得最开心的显然包括 NeurIPS,因为那是两万名 AI 研究者,年龄段基本跟我一样,很好玩,而且他们是顶尖研究者,你能学到很多,派对也多。另一端是日本某个随便什么化学品会议,三百个日本人,其中二十个来自 ASML,二十个来自台积电,二十个来自英特尔,只有这些人说英语,其他人全说日语。你会想,好吧,他们也挺有意思的。我有一项本事,就是不管对方什么背景、什么身份,我都能跟他建立联系,找到共同感兴趣的话题,通常是技术话题。

最有意思的会议往往是最大的那些,因为最大的事在那里发生。但真正令人兴奋的是那些小众的,比如 SPIE。有 IEEE,还有 SPIE,是另一个生态。SPIE 的会议细节深到极点。我去的每一场,尤其是 SPIE 先进光刻和 SPIE 光掩模,第一次去时,我听到的东西百分之九十都不懂。回去猛读,当然有了些背景,第二次去大概懂一半,第三次懂四分之三。到现在我去了还是觉得,我依然没有完全搞懂这里在发生什么。而 NeurIPS 你去个两三次,就能明白什么是神经符号推理、这是什么、那是什么,很快就能把整张地图画出来。可供应链里有些部分是如此艰深、如此技术化,你光是弄清楚正在发生什么就要花很长时间。

去会议有几层目的。你要理解研究本身,但发表出来的研究只是一部分,你真正关心的是这项研究如何与产业技术交叉,如何与今天已有的东西不同。没有一篇论文会告诉你今天的产线上在发生什么,你得去问人,建立联系,然后你就了解了供应链:哦,原来这家公司给那家公司供货,虽然没有任何地方公开写过;这种化学品大概这个价,一台设备大概用这么多。

马奎尔: 然后你就会听到那些恐怖故事:某种化学品短缺,把供应链的某一段整个打乱,最后发现全世界只有三家公司能生产它。

帕特尔: 我最喜欢的一个故事,就是在那场几乎没人说英语的日本会议上,一位日本人用很蹩脚的英语告诉我,他父亲上世纪八十年代在这个行业,当时全世界唯一一座生产某种化学品的工厂烧毁了,结果内存价格翻了两三倍。我说,哇,跟今天没什么两样。

马奎尔: 一点都没变。

推理市场与 InferenceX 基准

黄: 推理会成为地球上最大的市场,乃至地球之外最大的市场。同意还是不同意?

帕特尔: 显然,token 的使用会是最大的市场,token 创造的价值也会是最大的市场。我认为「token 经济学」(tokenomics),也就是 token 的使用、AI 的采用,是当下正在发生的最重要的事。推理,不管是开放模型还是封闭模型,会是全世界最大的市场之一,比石油大得多,比很多别的行业都大得多。AI 推理会占到 GDP 的好几个百分点。

黄: 你们做的 InferenceX 我认为已经是行业标准了。讲讲为什么做它、它是做什么的,以及人们对推理性能基准测试有哪些误解?

帕特尔: 先退一步说。SemiAnalysis 做的事里,很多是给机构客户做的研究,还有订阅业务,但也有很大一部分是:这个东西搞清楚会很酷,那我们就想办法搞清楚,然后公开发出来。这种事越做规模越大。我们在 GPU 基准测试、训练性能、推理性能上都这么干过。但我们最终发现,推理基准测试都是「时点式」的:你测一次,花一段时间,发布出来,又慢又艰深又过时,因为模型一直在变。我感觉每周都有新模型,不是中国模型就是别的,就今天,Mythos 5 和 Fable 发布了。软件层也一样,PyTorch、vLLM、SGLang,新驱动、新的什么东西一直在发,实际上这些库大多一周更新两次。

所以软件一直在变,性能也就一直在变。新的推理优化不断出来,被合并进去,一次又一次的突破持续把效率推高、把成本压低。这就是为什么同等质量的模型成本一年能降六十倍,太惊人了。要跟上这个节奏,就不能做时点式基准,基准必须是活的,也就是持续在最新硬件上跑最新模型。于是我们启动了这个项目,并且得到了生态的广泛支持。这只有在我们对生态有足够号召力时才可能,我们说动了 CoreWeave、Crusoe、Nebius、甲骨文、微软、亚马逊、谷歌和 OpenAI 给我们捐算力,跟 SGLang、vLLM 合作,现在还有 Radix Arc 和 InRact,也就是领导这些开源工作的私营公司。英伟达、AMD、谷歌、亚马逊也都参与了,因为我们正在加入 TPU 和 Trainium。捐给我们的硬件超过五千万美元,等 TPU 和 Trainium 上线,应该会超过一亿美元。大约十五种芯片,每天都在跑这些基准,跑的是最新的模型:月之暗面最好的模型、阿里最好的模型,大概五家中国实验室最好的开源模型,还有美国最好的开源模型,GPT-OSS、Nemotron 等等。每天自动跑,跑在专门给我们做推理基准的服务器上,扫过大量不同的配置和优化类型。所有结果公开,所有配置公开。

这样我们就得到了帕累托最优曲线。很多时候人们比较推理性能,是拿别人的一个次优曲线或次优点,去比自己的最优点。这就好比让我开保时捷去跟专业赛车手比,我当然开得慢。推理基准也是这样。所以我们做了开源的容器,覆盖曲线上每一点的最优配置,这条曲线的一端是交互性(interactivity),也就是模型回应我的速度,另一端是批大小(batch size),也就是我同时服务多少用户。现在任何人想要最优点,直接去 InferenceX 下载来跑就行,愿意的话可以每天查,甚至可以自动下载该模型的最优配置,推理性能就接近峰值。

吞吐与延迟:一条决定一切的曲线

马奎尔: 在你看来,吞吐量对交互性这条曲线是最重要的曲线吗?

帕特尔: 是的。硬件、基础设施、模型、应用层,几乎一切都是这条曲线的下游。你要的是超快、超低延迟,不在乎成本?那就把批大小压得很低,大量使用推测解码(speculative decoding)或多 token 预测(multi-token prediction)之类的技巧,可用的技术非常多。还是说你在批量处理海量文档,根本不在乎这些?那就不用那些牺牲成本效率来换取单用户速度的技术,而是把尽可能多的用户打包在一起,文档处理一整晚也无所谓。现在我们对待 AI 基础设施还是一刀切的。但随着时间推移,会分化出批处理型负载和即时响应型负载,整条曲线对用户都会有意义。我们在 Anthropic 那里已经看到了,Claude Code 的快速模式比常规模式贵得多;OpenAI 的优先队列也是同样的逻辑。

黄: 问个笨问题,成本是怎么进入这张图的?

帕特尔: 举个假想的例子。我的批大小是一百,每个用户每秒十个 token,那么这一块算力总共每秒产出一千个 token,这是曲线的一端:超慢,每秒十个 token。另一端是每秒五百个 token,但只有一个用户;或者说每秒二百五十个,一个用户。中间还有一些更接近帕累托最优的点,普通人其实想要每秒五十到一百个 token,再配上我能打包在一起的用户数。所以这条曲线就是:同一块硬件,总产出是每秒一千个 token 还是二百五十个,取决于我打包了多少用户。于是有些负载会想要那四倍的成本下降,因为同一单位硬件能做一千而不是二百五十;有些用户则愿意付四倍的价钱,因为他们不在乎价格,在乎时间,使用 token 的那个人很贵,或者这个反馈回路本身很贵。

太空算力与每瓦智能的预测

马奎尔: 让你猜一下,时间范围你自己选,十年或十五年,你觉得会有多大比例的推理算力在太空里发生?

黄: 可以是零,可以是百分之五十。肖恩会说百分之九十九。

帕特尔: 这个问题很难。我说个反共识的,至少是站在 SpaceX 对立面的看法。顺便说,我热爱 SpaceX,如果我能买股票,我一定会买它的 IPO。

黄: 不构成投资建议。

帕特尔: 不构成投资建议,我们两边都不构成。我不认为太空数据中心在未来三到五年内真的重要。但话说回来,二十年后,我认为绝大多数算力会去太空。真正的决定因素是时间范围,是在地面上建电力的成本,以及地面上到底能拿到多少电力。而我对推理会消耗多少吉瓦乃至太瓦的预期,对我个人来说是一条很疯狂的曲线。

马奎尔: 你的预测是多少吉瓦?

帕特尔: 到 2030 年,光 OpenAI 和 Anthropic 加起来就会超过一百吉瓦,再加上 Meta、谷歌等等,这是一个巨大的算力总量,都用于推理。到 2040 年左右会是太瓦级。我们将获得的生产力曲线是这样的,所以推理部署会非常巨大。如果看 2040 年,我认为新增算力中会有一半以上进入太空。但看 2030 年,我认为不到百分之一。

黄: 你认为每瓦智能(intelligence per watt)一直在提高吗?而且我们现在的每瓦智能跟人类生物大脑之间似乎还有巨大差距。你觉得我们会缩小这个差距吗?如果会,增益从哪里来?

帕特尔: 这常常取决于你在做什么。一台 TI-84 计算器在做算术这件事上的每瓦智能远超人类,而它是三十年前的东西。当然这么说有点傻。

黄: 我说的是通用智能。

帕特尔: 就通用智能而言,InferenceX 做的一件事就是同时测量所有这些硬件的功耗和成本。所以我们不只提供吞吐量对交互性的曲线,还提供成本对交互性、功耗对交互性的曲线。至于每瓦智能有没有提高:我刚才说过同等基准水平下成本降了六十倍,每瓦智能也有类似的提升,没有正好六十倍,接近四十倍,因为有些效率提升不是功耗层面的。但每瓦智能每年都有巨大改善,至少今年、去年、前年、大前年都是如此,我预计这会继续。至于跟人脑的差距,我们差着好几个数量级。幸好这无关紧要,我们可以给计算机投入大量电力,给计算机供电比给人脑供能容易得多。人脑会生病、有疾病,还挑食。

黄: 还要睡觉。

帕特尔: 没错。

硬件、系统、模型:百倍从何而来

马奎尔: 就这个主题再问一个。在我看来,无论是每瓦智能还是每美元智能,输入大致有三个层次。硬件层面的改进,硬件本身更高效;底层系统优化,内核级别的改进、矩阵乘法库之类;还有最高层的模型算法改进。我的感觉是,过去三年大部分增益来自硬件层面,一部分来自模型层面。你同意吗?你认为未来会是什么样?内核层面还有很多油水可榨吗?

帕特尔: 肖恩,我完全不同意你。

马奎尔: 太好了,这正是我问这个问题的原因。

帕特尔: 可以把它看成这三层。从这个角度,从 Hopper 到 Blackwell,也就是过去三年我们拥有的全部硬件进步,在 DeepSeek 上最优化的部署大约有三十倍的改进,这在 InferenceX 上能看到。但过去三年每瓦智能的提升远大于此,很大一部分来自模型层。三年前是 GPT-4,现在一个较小的 Qwen 模型,总参数二百七十亿、激活参数二十亿,就比它强得多。所以模型层有巨大的改进,硬件层有相当可观的改进,但关键在那个协同设计层,我认为这才是重要的。你看任何一个模型的架构,最著名的、公开可见的当然是 DeepSeek。

马奎尔: 对,DeepSeek 从协同优化、从内核级别的内存优化里拿到了巨大的效率增益。

帕特尔: 是的,当然有内核,但更深一层,是你按芯片的硬件架构来设计模型。DeepSeek V3 里所有专家(expert)的形状都是为 Hopper 优化的,V4 则是为 Blackwell 和华为的芯片优化的。有意思的是,尽管 TPU 客观上是极好的芯片,整个 DeepMind 都跑在上面,Anthropic 至少在预训练上也全部用它,但 TPU 跑 DeepSeek 很烂,而它跑另一些在英伟达上跑不好的模型却非常出色。优化已经深到这种程度:专家的形状、网络 IO 模式、集合通信怎么做、注意力机制的算术强度,所有这些都是在模型、硬件和中间的基础设施软件之间协同优化的。很难把增益拆开来算。

马奎尔: 我的理解是,过去几年中国在这方面做得比西方好得多,DeepSeek 是最早真正这么做的模型之一。

帕特尔: 我不这么认为。更准确地说,是西方不告诉别人自己在做什么。OpenAI 没有公开 GPT-4o 有多稀疏、形状多大这些东西,但 GPT-4o 的规模跟 DeepSeek V3 大致相当,略小一些,而且如果我没记错,4o 出得更早一点。

马奎尔: 所以你的看法是,这三层同时在推进,速度大致相当,而最大的增益出现在你协同优化的时候?

帕特尔: 我会说模型层的增益比软件基础设施层和硬件层都多,但每一层都有创新。真正最大的增益,也是顶尖实验室的美妙之处,在于三层一起协同优化。Anthropic 用很多种硬件,但他们不太在 TPU 上做推理,主要用 TPU 训练,推理大量放在 Trainium 和 GPU 上。GPU 更像是万金油。但他们优化了硬件、优化了模型、优化了一切,所以做得到。OpenAI 以前的模型更多为 Hopper 优化,现在更多为 Blackwell 优化。谷歌也一样,Gemini 2 是为 TPU v6e 深度优化的,Gemini 3 也是,下一代 Gemini 则是为 TPU v7 优化的。你把这些模型拿到旧硬件上跑,其实并不怎么样。

所以协同优化是最重要的,术语叫软硬件协同设计(software-hardware co-design)。这也是我日常工作里最令人兴奋的部分:你看某一层,这里有一堆创新,每一层都有一堆创新。真正的突破性创新,是当你跨越几层,把它们一起协同优化、协同设计,于是本来这里两倍、那里两倍、再一个两倍,相乘是八倍,结果却是一百倍,因为你跨三层做了优化。这就是实验室里令人兴奋的地方。英伟达这样的公司也是,他们不算在模型层做协同优化,但会从模型层的一点点一直往下游做到硅片。台积电也是,他们协同优化的不只是制造,而是从零部件、耗材、设备一路向上游,直到客户告诉他们的芯片设计。这是横跨抽象栈许多层的协同优化。

内存墙与每平方毫米一瓦

马奎尔: 这种优化里总会有某个地方成为瓶颈,落在后面,需要被拉上来,或者先打个补丁。让你预测一下,栈的任何层级都行,未来一年你最密切关注的瓶颈是什么?不一定是供应链,不一定是规模,可以是技术本身的。是内存的改进吗?

帕特尔: 内存是个容易的答案,大家都在谈,但我不从供应链角度谈,我从技术角度谈。内存的容量和带宽一直提升得很慢。NAND 单元大约二十五年前发明,DRAM 单元大约四十年前发明,单元本身没有过重大突破。NAND 就是一个很简单的门,DRAM 单元也是。管线里有些东西可能会带来巨大创新,但过去五年我们真正做的只是把 HBM 堆得更多、跑得更快。不过未来几年会有新的创新:不再把 HBM 跟芯片分开堆叠,而是把内存直接堆到芯片上,带宽会爆炸式增长。这个领域有些有意思的公司,也有些公司在尝试有意思的 PoC。我认为内存带宽是最大的瓶颈之一。

另一个是,至少过去二十年,硅的历史里有个规律:一块数据中心或桌面芯片的功耗,看一眼面积就能预测,上限大约是每平方毫米一瓦。一块一百平方毫米的芯片,功耗一般在一百瓦或略低。看英伟达最新的硅片、TPU 最新的硅片,仍然在每平方毫米一瓦这个范围。所以芯片现在到了一千四百瓦,英伟达下一代 Rubin 是两千瓦,再往后 Rubin Ultra 大概是四千瓦之类。但这其实是在增加硅的面积。令人兴奋的是,我们终于开始做一件事,而且已经在开发中:真正把远超每平方毫米一瓦的功率灌进硅片。这意味着你需要的硅更少了。当然它跑在更高的功率上,某些情况下效率更低,但硅的用量减少了。

马奎尔: 要克服散热问题。

帕特尔: 散热问题,还有电气干扰问题,各种各样的问题都会冒出来,所以这是个艰难的工程问题,所以我们一直卡在一瓦左右。但令人兴奋的是,整个世界都在试图改变这些。

供应链的另一个部分也有类似的情况。人们总说能源很难,我们有能源瓶颈。是的,但其实有些很简单的解法,稍微想想就能想到。比如美国有能力生产数以百万计的卡车柴油发动机,你可以在装配线上很轻松地把它们改成烧天然气,然后接上一台电机反向驱动,让电机发电而不是驱动车轮。这样你往一个美国能造几百万台的东西里灌天然气,就发出电来了。你会说,维护起来太麻烦了吧,一个数据中心场地上得摆几百台。其实你直接从汽修店里把人拉来修卡车发动机就行。我不想说这很简单,反正我自己修不了。

马奎尔: 你说到了一个很好的点:因为西方过去二三十年没怎么想过半导体,甚至更广义的硬件,所以我们没有多少创新,最聪明的头脑没有在想怎么改进这些东西。

帕特尔: 能去做广告,谁要去做硬件呢?

英伟达对 TPU:护城河还剩什么

黄: 我特别想问,英伟达对 TPU,你怎么看?

帕特尔: 大家都想在这个问题上二选一,但它其实是这样:看两年后,谷歌会通过自己的供应链生产一千多万颗 TPU,英伟达会生产几千万颗 GPU,谷歌一年会造出价值一千多亿美元的 TPU,英伟达会是五千亿以上,或者别的什么数,我不做具体估计。

黄: 这不是营收预测,只是思想实验。

帕特尔: 对,或者叫研究。

黄: 你受过媒体培训。

帕特尔: 当然,在为 SpaceX 的 IPO 做准备呢。你们在 SpaceX 里持仓很大?

马奎尔: 我们很幸运,是很大的投资者。

帕特尔: 那就说得通了。谷歌 TPU 对英伟达 GPU,双方都有真正站得住的论点。英伟达会说,我们有交换机,我们是通用的。TPU 会说,我们更优化,实际上更节能,我们的网络对某些网络架构更优化。双方都有这样的对攻点,我可以面不改色地论证 GPU 远胜 TPU,也可以论证 TPU 远胜 GPU。但归根结底是软硬件协同设计。按 OpenAI 模型的走向,他们用 TPU 可能会是一个糟糕的决定;按 Anthropic 和谷歌模型的走向,用 GPU 训练可能是一个糟糕的决定。

马奎尔: 根本差别在哪里?

帕特尔: 有很多。最简单的一个,矩阵乘法单元的尺寸不同,于是你做的矩阵乘法形状、你用的注意力机制、注意力机制的结构方式、专家的结构方式都不同。

马奎尔: 所以你认为 OpenAI 和 Anthropic 正在走向非常不同的模型架构?

帕特尔: 我认为他们的模型架构相当不同。OpenAI 的稀疏得多,这有好处;Anthropic 的仍然稀疏,但总体更稠密,这有另外的好处。还有很多别的,比如网络拓扑。英伟达所有芯片都接在 NVLink 交换机上,谷歌没有交换机。NVLink 只能连七十二颗 GPU,谷歌的 ICI 能以超高带宽连八千颗芯片,但因为没有交换机,你得穿过别的芯片才能到达。这里面有取舍,有正有负,而这些会影响模型架构。不必非要说哪个更好,因为你没法把它们单独拿出来衡量,它一直延伸到模型层。

马奎尔: 我记得很长一段时间我都认为英伟达的可编程性和 CUDA 是巨大的护城河。至少在过去三到六个月,这个叙事在我脑子里变了:模型公司不再在乎要不要为别的芯片写定制内核,说要用四五种芯片就用四五种。Claude 和 Codex 在做这类优化工作上其实相当不错。而且模型公司并没有一万家,是几十家的量级,每家都需要可编程性。所以「成千上万的大客户需要 CUDA 兼容」这个基本前提,似乎在变。

帕特尔: CUDA 护城河、软件护城河,确实至少部分被解开了,因为模型太擅长写代码,所有软件在这种情况下都被商品化了。但我认为有另一层,人们叫它 CUDA 护城河,实际上跟 CUDA 无关,而是这样一件事:DeepSeek、Kimi、智谱、阿里、腾讯,还有最近出了一个很棒模型的小米,所有这些公司的模型都是为 GPU 协同设计的。我想在 TPU 上跑它们,某些情况下就跑不好。谷歌只好自己建开源模型生态,于是有了 Gemma。所以这不是 CUDA 护城河,而是下游产品更多为英伟达优化,这些公司又把模型开源了,Nemotron 也是开源的。它们的使用者,比如推理 API 提供商、拿开源模型为企业业务做定制的强化学习公司,全都处在这个事实的下游:好吧,生态用英伟达,那我也得用英伟达,虽然我根本不在乎写 CUDA 内核,模型写得很好,但问题在于这个专家的维度是多少、隐藏维度是多少,因此在英伟达 GPU 上跑就是比在 TPU 上好,反过来也一样。如果谷歌真的开源出很好的模型,同样的事就会发生:人们拿来一跑,说这在英伟达 GPU 上跑得不怎么样,我干脆去租 TPU 或者买 TPU。

对小团队来说,你会想用所有的开源软件,vLLM、SGLang、PyTorch 这些。但大实验室不一定需要。OpenAI 很早就把 PyTorch 分叉了,Anthropic 和其他人也都不太依赖这些东西的开源实现,他们要么分叉了,要么自己造了,不需要依赖开源。所以对他们来说更像是:我选最好的硬件,然后为这个最好、最省钱的硬件把模型和基础设施软件从头到尾协同设计一遍。

马奎尔: 而且让 AI 帮我写所有这些软件。

Cerebras 与快模式的经济账

黄: 你怎么看 Cerebras?

帕特尔: Cerebras 是一家非常有创新力的公司。在市场的某些位置上他们真的很强,超快推理,我认为那是个大市场。SemiAnalysis 内部几乎只用快速模式。

马奎尔: 顺便说,我很欣赏你们在核算上的自律,不知道那是一次性的展示还是一贯如此,把每项任务花的钱和投资回报都算出来。分析非常棒。

帕特尔: 谢谢,我们做得挺认真的,那是我们写的「暗 GDP」(Dark GDP)那篇文章。我们还按天追踪每个人的 token 花费,谁的花费突然飙升,我就问,你干了什么?对方一解释,好,谢谢,看起来值。接着过我的日子。快速模式对高端任务显然很有价值,我能想到太多超快 token 值得付费的场景。反过来我也能看到很多场景不需要超快 token,市场不会为此付费,会去用 GPU 和 TPU。Cerebras 最大的风险在于,我认为你想用快速模式的恰恰是最好的模型,小模型你未必会用快速模式。在金融市场这个判断可能不成立,比如 Jane Street 那种高频交易,或者中频交易。但归根结底,在 Cerebras、Groq 这类基于 SRAM 的芯片上跑超大模型、超长上下文是非常困难的。那么问题来了,如果模型变得太大怎么办?如果 OpenAI 的模型不是几千亿或一两万亿参数,而是十万亿以上,我不认为它还能装进 Cerebras。再加上长上下文,如果是一百万的上下文长度,就很难自圆其说了。而到目前为止,我们看到实验室的营收和用量大头都在他们最好的模型上,即便模型涨价也是如此。有数据显示,Fable 今天才发布,就已经有大量用户转向 Fable 和 Mythos,也就是更高一档的模型,虽然它贵得多。

马奎尔: 这是按金额算的量吧?按 token 算呢?

帕特尔: 谁在乎按 token 算的量?重要的是钱。如果 F-150 的单价是五倍、销量只有一半,我才不在乎 Mini Cooper 或者丰田凯美瑞卖了二十万辆。

马奎尔: 好吧。

帕特尔: 所以美国最赚钱的市场是皮卡。我多半是在开玩笑,但道理如此。

马奎尔: 我确实认为这是你做得极好、跟几乎所有人都不一样的地方:你在技术之外极其在乎经济账。能把这两件事接起来的人非常少。

帕特尔: SemiAnalysis 内部很好玩,我们有九十个人,很大一部分是横跨整条供应链的技术专家和工程师,另一大块是以前在对冲基金的人。你会看到这样的争论:有人说这个不重要,有人说可是成本呢,工程师说不不不,这项技术最酷。这种争吵是自发的。我们很不正式,而且考虑到我曾经是论坛版主,你可以想象我有多享受。

马奎尔: 别跟猪摔跤,因为猪乐在其中。

帕特尔: 正是。

黄: 接着这个话题,在半导体领域有没有哪些话题一听就让你炸毛?比如一个人一开口说「内存是瓶颈」,你就觉得这人肯定是……

帕特尔: 那句倒是真的。真正让我火大的是有人说「AI 没有投资回报」。什么投资回报?还有否认模型进步的人,说模型没在变好,它们不会推理、不会思考,会走进死胡同然后停滞。兄弟,能力曲线一直在往右上方走。他们说,看,这个基准没提升。那是因为已经到百分之九十了,看新基准,你们把旧的饱和了,新的正在飙升。我觉得这才是问题所在。半导体确实极其复杂,我不怪别人不懂。我每天都在从别人那里学到半导体供应链的新东西,而我从十二岁做版主算起,可以说已经研究了十八年,全身心投入,这是我唯一在乎的事。即便如此,抽象栈的层次太多了。就昨天我还了解到一种年销售额一亿美元的化学品,我说,哇,不知道还有这个,它用在哪道工序。在一个几千亿美元的产业里,一亿美元的销售额算什么。

马奎尔: 但它是必需的。

帕特尔: 它是必需的,每颗芯片都要用。你会想,好吧,毕竟有一千道工序。就像有人说,你喜欢半导体是吧,把每道工序都报出来。得了吧。我觉得最好笑的是,有些人所有事实都摆在面前,结论却完全错了。

马奎尔: 我们这行也天天发生。

帕特尔: 我的态度不是为此生气,而是尽快把它做对。

十年视角与局部最小值陷阱

黄: AI 现在是全世界最重要的事,短期瓶颈太多,我们聊了很多短期的东西。有没有更长期、比如十年尺度上让你特别兴奋的事?我们谈了轨道数据中心,比如硅光子,你觉得十年尺度上是被低估还是高估?还有别的吗?

帕特尔: 十年尺度上,太空我觉得极其疯狂、极其棒,太空数据中心、小行星采矿这些,我对 SpaceX 的愿景非常兴奋。再说一次,不构成投资建议。半导体这边,市场的巨大波动和巨大的事件,往往就取决于某件事早一年还是晚一年发生。共封装光学(co-packaged optics)就是这样,大家都知道十年内会实现,争的是 2027、2028、2029 还是 2030 年,但某个时点一定会来。我觉得更有意思的是有些公司,比如你们投了 Naveen Rao 的公司吗?

马奎尔: 投了。

帕特尔: 他在试图同时在硅层、软件抽象层和模型层做创新,而且他完全明白这不是「几年内做成」的事。

马奎尔: 不是两年的时间尺度。

帕特尔: 不是几年的尺度,是长期押注。类似这样的事,比如要把模拟计算和基于能量的模型(energy-based models)之类的疯狂东西一次性全上,这很令人兴奋。多半不会成功,但令人兴奋,我真的很期待。

马奎尔: 至少肯定不会很快成功。

帕特尔: 对,我应该说,肯定不会很快成功。我相信 Naveen。有意思的是,他是我在这个行业里最早认识的人之一,2020 年,实际上是 2019 年,那时我还是匿名的。我在网上钓他,他开始回复,我就把话头转到私信,再转到电话。他是整个半导体行业里第一个跟我说上话的重要人物。

马奎尔: 这说明了他的为人。以我的经验,他总是在帮年轻一代,总是在发掘人才。

黄: 他做 Mosaic 的时候也太超前了,我记得当时听过他的路演。

黄: 你认为这个生态的终局是什么?每家实验室、每家超大规模云厂商都有自己的芯片?Trainium 看起来现在跑通了。所以终局是每家实验室、每家云厂商至少在推理上用自研芯片,训练可能找英伟达或者别人?还是别的什么?

帕特尔: 我认为每家都会去试,也有人会放弃。归根结底供应链很重要,你能引入什么技术很重要,行业越大,供应链多元化就越会发生。现在所有人的芯片长得差不多:中间一块大逻辑计算芯片,左右是 HBM,上方是网络,下方是 PCIe 和其他 IO。Trainium、TPU、英伟达的芯片结构完全一样,大多数初创公司也是,Groq 和 Cerebras 除外,他们在做奇怪的东西,但那很酷。往前走,硬件架构和模型架构会进一步分化,人们会把两者协同优化,而其中一些会落进局部最小值。就像梯度下降,大家都在往最优解走,有些人会冲进一个局部最小值,问题是你怎么跳出来,怎么挪回全局最小值。

某种程度上,英伟达总会比其他任何人的芯片更通用,至少在并行 AI 计算上是这样,因为他们有那么多关心不同东西的客户,会一直在设计上给反馈。一个真正的最小值总会比他们好,但那是不是局部最小值?TPU、Trainium、Groq、Cerebras 或者谁的设计,在这里优化得很棒,但终局其实在那边,那他们就错了。也许他们在一段时间里很好,然后就错了。这才是真正的问题。所以我认为通用 AI 算力会有一个大市场。你跟实验室的人聊,他们甚至不知道一年后自己会用什么架构,真的不知道。他们有很多研究押注,这正是令人兴奋的地方,但不知道会走向哪里。他们大体知道手头有什么硬件,试着去协同优化,但如果模型架构出现新的突破,比如把注意力机制换成别的什么,谁知道呢,最好的硬件就会变。那么,人们会在硬件上做五年的投资,全押在一个更专用的 ASIC 上吗?还是会保留一桶更通用的算力?

所以你看到谷歌以每小时十一美元一颗 GPU 的价格向 xAI 租 GPU,太疯狂了,价格极高,算力当然紧缺,但这依然很疯狂,何况他们有 TPU。这里面有问题值得问:他们为什么这么做?谷歌其实有三个不同的 TPU 设计项目:跟博通做的 TPU,跟联发科做的是另一种架构,还有第三个,我不方便透露,架构跟前两个很不一样。不是简单地找几家供应商做同一种架构,是不同的架构。所以人们意识到局部最小值会发生。因此我认为每家都会有自己的 ASIC 项目,每家都会部署几十亿、几百亿美元的自研 ASIC,谷歌一年会部署几千亿美元。但最终他们也会有不用 TPU 的负载。谷歌那些不属于 Gemini 和 DeepMind 的项目,有些主要用 GPU,不用 TPU;也有些主要用 TPU。比如药物发现,或者 Waymo,你可能不想用 TPU,我不说是哪一个。不同的架构押注,不同的 AI 路径。科学 AI 的算法模式可能跟通用智能模型不同。所以多样性会继续扩散。而且市场已经这么大了,各种利基会被切出来,公司可以守住自己的利基赚到钱,哪怕大头归了英伟达、TPU 和 Trainium。

算力紧缺、毛利率与杠杆之忧

黄: 我们聊聊数据中心建设。看图表,每算力小时的价格显示我们正处在一场疯狂的算力紧缺之中,而且供需两端都紧:长时程智能体的需求飙升,供给端所有数据中心建设都在延期。你认为可预见的未来都会是算力紧缺,还是某个时点会缓解?

帕特尔: 每个季度我们部署的算力都远多于上一季度,建成的数据中心也多于上一季度。今年即便算上延期也会有二十吉瓦,明年算上延期会超过三十吉瓦。延期当然什么都会有,硬件的东西都可能延期,这就是生活。我们会不会一辈子都算力紧缺?取决于模型会怎样。Mythos 5、Fable 5 的潜在市场(TAM)不是 Opus 的两倍那么简单,模型好太多了,能做的任务多太多了,市场大得多。可全世界的算力在过去六个月并没有翻倍。从 Opus 4.5 发布到现在大概七八个月,4.6、4.7、4.8 都是改进,但 Fable 和 Mythos 是一次巨大的阶跃,同期全世界的算力没有翻倍,也没有翻两番,而 AI 能做的有用任务的数量和价值却翻了。那接下来会怎样?显然 Anthropic 二季度已经盈利了,剔除股权激励后净利润为正,三季度可能连算上股权激励也盈利。他们已经赚钱到这个地步。他们一个 Opus token,至少是 Opus 4.8 的 token,按 API 定价的毛利率超过百分之八十。他们有很多交易会把公司整体毛利率往下拉一点,比如 Bedrock 和 Vertex 的分成,但单 token 毛利率极高。所以他们有能力以高于市价买下每一颗 GPU。他们也确实以高于市价从 SpaceX 买了 GPU,比谷歌付的价低,因为签得早。这是别的公司做不到的,比如靠风投的公司,或者毛利率不为正的公司。成本收益比是这样的:我算力不够了,每租一颗 GPU、每一颗 TPU、每一颗 Trainium,我都能立刻转手卖出 token,而且是正毛利。如果我毛利率百分之七十五,算力成本翻倍,也没事,还有百分之五十。如果是租的,开新的算力节点对他们来说也不怎么需要人手。所以我的净营业收入照样上涨,我会以我愿意付的任何价格去租 GPU。

黄: 我的问题几乎是反过来的:这场算力建设会不会在某个时点让人半夜惊醒?今天早些时候有条推文,Crusoe 公开说有一位客户要求暂停某个数据中心的建设。整个生态里所有人现在杠杆都拉得很高,就是要建、要建、要建。高杠杆加高增长,作为投资人,这让我非常紧张。

帕特尔: 等等,高杠杆加高增长意味着一小笔股权有巨大的上行空间。你又不是债权投资人,你是股权投资人。

黄: 冲。

帕特尔: 你得去上私募股权学校,只做杠杆收购。

黄: 我就是私募股权出身。

马奎尔: 她把学校忘了,做风投太久了。

黄: 我现在只看营收倍数。说正经的,你看到这样的迹象吗?你担心吗?

帕特尔: 我明白你的意思。这又回到模型的问题。如果模型正在扩大经济上有价值的工作总量,也就是你刚才提到的我们那份「暗 GDP」报告,如果模型能做的工作扩张得没有算力快,潮水就会转向。过去六个月,潮水非常明显地偏向这一边:模型能做的工作、它们能做的工作的市场,扩张得比算力快,所以价格上涨。当然完全可能突然之间模型进步停了。你跟 Anthropic 或 OpenAI 的任何人聊,也许他们喝了自家的迷魂汤,但他们基本都说,不不不,模型还在进步。当下的方法可能会在某处停滞,我不确定在哪里。但看起来我们对模型的持续快速改进是有清晰路径的。事实上模型改进得比六个月前、一年前更快,因为,我不想叫它递归自我改进,但模型在帮助写所有的基础设施代码,让下一代模型越来越早发布。这是一个准递归自我改进的循环,模型越来越好,越来越快。

但资本确实是个大问题,这就是谷歌融资的原因。他们持有天量的 SpaceX 股份,大概百分之五。

马奎尔: 我觉得还多一点,但差不多。

帕特尔: 一度有百分之十。拉里·佩奇在一百亿美元估值时投了十亿美元,拿了百分之十,后来被稀释。那是史上最伟大的投资之一,干得好,拉里。所以他们知道自己手里有一千亿美元,锁定期结束后九个月左右就能卖,加上他们所有的毛利润,可他们建了模型一算,说我们需要融资,于是就发行了。这太疯狂了,它告诉你他们认为自己需要花多少钱。Meta 也宣布要融资,股价跌了,人们不喜欢,但所有这些公司都会融资,要么发债要么发股。某个时点资金的闸门总要放慢。但现在,亚马逊每加一颗 GPU、TPU、Trainium,谁加都在赚更高的收入、赚毛利。

每吉瓦不等值:数据中心分层

马奎尔: 我铺垫一下再变成一个问题。听你讲的时候,我脑子里在转一个几乎可以解释 Crusoe 那件事的替代假说。用石油打比方:沙特每桶油的生产成本比很多国家低得多,而且油的纯度高、杂质少,炼起来更容易。我的问题是,看今天落地的每一吉瓦,比如正在上线的二十吉瓦,你看到多大程度的同质性?你可以用你觉得合适的任何指标,但比如谷歌的吉瓦是不是比大多数 Neocloud 的吉瓦值钱两倍,因为他们有光交换机,做了很多年,知道怎么做功率平滑?这可能是那个替代假说:擅长建数据中心的人应该建到极限,因为需求太大,他们又强得多;而我们可能正在看到不那么擅长的人开始挨打的早期迹象。我不知道真相,只是好奇你怎么想。

帕特尔: 这方面已经有指标了。Trainium 租给 Anthropic 和 OpenAI 的价格是每吉瓦不到一百亿美元。GPU,至少在过去六个月的疯狂之前,通常在每吉瓦一百二十到一百三十亿美元,这是 Neocloud 的租价,甚至跟亚马逊比也是,亚马逊现在卖 GPU 也是一百三十亿左右。

马奎尔: 我的理解是,亚马逊对那个价格有一定补贴,实际差距比这还大。

帕特尔: 不到一百亿,但里面有些奇怪的结构。

马奎尔: 我的理解是 Anthropic 在让 Trainium 变得好用上出了大力,写了所有的库之类。我听到的所有说法都是 Trainium 的硬件真的非常好,而且在快速变好,Anthropic 现在大量在用,希望我们能看到那个价格涨上去。

帕特尔: 他们那笔交易其实有个下限机制:如果表现不好就更便宜,差到一定程度可以取消;如果表现很好,价格就高一些。但实际上 Trainium 落在每吉瓦不到一百亿。而 GPU,SpaceX 跟谷歌那笔交易是每吉瓦两百五十亿之类的疯狂数字,也就是每兆瓦每年两千五百万美元的租金。我当时就说,这个分化太夸张了。当然,如果亚马逊今天卖 Trainium,因为算力短缺,价格大概会高于一百亿。

但这种分化在数据中心层面已经存在了。数据中心的租价,如果做托管,不含算力,只是电力和机房,一般按每千瓦每月多少美元定价。以前是六十美元,现在成交价在一百二十到一百六十之间。但不同质量的数据中心差别很大,我见过高到两百的,那是客户信用评级不太好、而数据中心相当不错的情况。也见过还在一百的,在印度甚至低到八十,因为电网不可靠,网络连接不好,数据中心本身也一般,但好歹是个数据中心。所以这种巨大的价差已经存在了。

数据中心建设的坑通常就是直接失败。很多人失败,就四个人,说我买了燃气轮机,付了定金,我要建数据中心,然后延期、延期、延期,失败。所以你得按概率、按时间、按滞后去给那些差劲的团队和不差劲的团队加权。我们的数据中心模型就是干这个的,我们追踪每一个数据中心,根据他们用的设备等等给每一个做这件事。

你提到谷歌的一点:在一个一吉瓦的数据中心里,他们其实会放一点五吉瓦的硬件。因为他们从工作负载一路到底都了如指掌,能把电力在各处调度。通常一吉瓦的算力,功耗利用率大概在百分之六十到七十,注意是功耗,不是硬件利用率,硬件总有人租。而他们让那百分之六十到七十正好等于一吉瓦,把整整一吉瓦用满。你还会看到包括谷歌在内的人跟电力公司做这样的交易:我知道这个电网能稳定供一吉瓦,但一年里除了三天之外其实能供两吉瓦,那你给我两吉瓦,到那三天就叫我关掉。他们真的会这么做。这类招数要求对负载、备用电力、现场发电机等有极致的管理,才能把两吉瓦稳定地维持住。做到这些的人能收更多钱。要么是我明明只有一吉瓦却卖了两吉瓦,因为那三天我能用电池、燃气之类顶过去;要么是我想出了在现场发电的办法,别人都没有一吉瓦而我有,所以我能很快交付。不一定是单价更高,而是我卖出了更多吉瓦。有时候卖更多吉瓦的杠杆,就是每一吉瓦以不同价格出售。

我认为在数据中心和能源层,更多是「有还是没有」,以及「延期还是不延期」,比较二元。但在算力这一层,有意思的东西多得多。一吉瓦给 Anthropic,客观上比给 OpenAI 产生的收入更多。而两家现在手里的每一吉瓦似乎都能卖掉,OpenAI 和 Anthropic 都有速率限制、token 上限之类的问题,尤其是 Codex 5.5 出来之后,好用多了。同样,如果你给 SpaceX 一吉瓦……

马奎尔: 我的猜测是,他们对硬件的使用大概比大多数人都好。我觉得人们低估了他们从星链积累的网络经验,还有从特斯拉积累的电力管理经验。

帕特尔: 像 Brett Mayo 这样的人非常厉害。

马奎尔: 相当强。所以对我来说,这可能是很多人的分析里缺失的一块,我不确定答案。

帕特尔: 还有一点:CoreWeave 建一吉瓦,他们的 GPU 算力在性能和可靠性上客观上比亚马逊、谷歌、微软都好,我们测过。但问题是,谷歌在算力上线六个月前就把它卖掉了,然后拿着签好的合同去做债务融资,靠那份信用背书去付他们已经发出的订单。而 SpaceX 是说,不不不,这个现在就在跑,直接买。有没有资产负债表来做这件事,差别很大,这也会让你每兆瓦的收入高得多。

Neocloud 为什么能赢,黄仁勋在下什么棋

黄: Neocloud 这个机会为什么会存在?五年前你问我,我会说超大规模云厂商会包揽一切。你刚才还说 CoreWeave 的性能比他们好。这个机会为什么存在,宏观层面和执行层面分别是什么?

帕特尔: 2023 年我写过一篇让亚马逊非常恨我的报告,叫《亚马逊云危机》。我讲的是亚马逊曾经是最好的云,因为他们有 Nitro 网卡,提供租户隔离,整个虚拟机管理程序跑在网卡上,所有 CPU 核心都能拿来卖;他们有自制的 SSD,买裸 NAND 自己组装,成本更低;还有自研的 Graviton CPU,把每核心成本压下来。所有这些让他们能卖更多核心,有更好的安全性、更好的网络、更好的存储,但这全是为传统 CPU 云世界准备的。在 AI 云里,很多这类东西反而伤害性能。Nitro 网卡对性能不利,现在也仍然更差,虽然迭代了几轮追上不少。很多安全方面的东西无关紧要,因为我不是在把用户按时间片切开,或者把一个插槽切给很多用户。没人在一台八卡服务器上只租一颗 GPU,没人在一个七十二卡的机柜上只租一颗,人们租整个机柜,而且租很多个机柜。也没有「我租六小时然后还回去」这回事,全是长期合同。

所以 GPU 租赁市场的机制让超大规模厂商的很多专长失去了意义,其中一些甚至有害。谷歌和亚马逊的网络性能,他们的定制网络对传统 CPU 和他们原本做的事更好,但对 AI 并不合适。另一些情况下,比如微软靠自建数据中心省钱,但他们的数据中心团队其实不怎么样,需求可预测时还行,等到要把全年预测翻倍时,他们就摔了个大跟头,只好去买一堆 Neocloud 的容量。所以一是性能,二是上市时间。在那些庞大的组织里,没人会因为把数据中心建得更快而发财。但看 Crusoe,Chase 和团队里的其他人,我本来想点名,还是算了,这些人如果更快交付算力就会发财,他们是高杠杆的股权持有人。

马奎尔: 而且他们也都是比特币出身,虽然你不该这么说。

帕特尔: 他们主管数据中心的人是从微软来的。

马奎尔: 我只是逗你。不过在一个波动极大的市场里待过,你确实会学到很多。

黄: 你觉得黄仁勋在下什么棋?

帕特尔: 黄仁勋绝对憎恨一个所有权力都握在超大规模厂商手里的世界。他往那些随便什么 AI 实验室里砸钱是有原因的,我甚至不确定其中一些砸得有没有道理,但他在砸钱,把它们抬起来,跑遍全世界跟人说你该投这家公司,因为他想创造一个多极世界。这就是他为什么喜欢中国实验室,他要一个多极世界。一个只有 OpenAI、Anthropic 和谷歌模型的世界,他就完了。

马奎尔: 对。

帕特尔: 一个只有超大规模厂商在建算力的世界,他也完了。所以他当然要把分配的枪口对准 Neocloud,给他们的集群兜底,能做的都做。因为今天卖给 Crusoe、卖给 CoreWeave、卖给谷歌和亚马逊的 GPU 对他来说都是同一个价,但五年后,Crusoe 和 CoreWeave 还活着,就意味着谷歌 TPU 更弱,亚马逊 Trainium 更弱;更多推理在封闭实验室之外发生,对他公司更有利。所以 Neocloud 生态,还有那些新实验室(Neolab),很多都有英伟达的投资,这是狂野西部,有人会失败,很多人会失败,但有些会冒出来成为真正优秀的团队。比如很奇妙的 Crusoe,一群加密货币的人转去建数据中心、做伴生气发电;比如 CoreWeave,最初是一群纽约对冲基金的人,也做过加密货币,后来建起来了。当时跟他们同期起步的很多人就没冒出来,失败了。

马奎尔: 我得说这两个团队都非常出色,值得很多赞誉。这也正是你的意思。

帕特尔: 对,我的意思是,往水里撒一堆饵,最好的鱼会自己找到路活下来。Neocloud 是这样,他也希望 Neolab 也这样。能不能有 Neolab 真正冒出来,走着瞧。但 Thinking Machines 已经有几亿美元的年化经常性收入了,这相当了不起,尽管媒体上说他们流失了很多人才,可 Tinker 一个上线不到六个月的产品做到几亿美元 ARR,很了不起。我们希望别的 Neolab 也这样。他要的就是一个多极世界。

马奎尔: 真心恭喜你的成功。最后说一句,我亲眼见过一点:听众大概能从你说话里听出你有多拼,你显然十多年来一直在玩命工作,才有了这几年的天时地利。你做成的事难以置信,而且我知道这只是开始。

帕特尔: 非常感谢。

黄: 谢谢你接受访谈。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 加油站里训练的第一个神经网络 ▶ 正在看
6:42 离职、丧亲与匿名博客的诞生 ▶ 正在看
10:05 一年四十场会议怎么学供应链 ▶ 正在看
13:26 推理市场与 InferenceX 基准 ▶ 正在看
18:07 吞吐与延迟:一条决定一切的曲线 ▶ 正在看
20:18 太空算力与每瓦智能的预测 ▶ 正在看
23:16 硬件、系统、模型:百倍从何而来 ▶ 正在看
29:11 内存墙与每平方毫米一瓦 ▶ 正在看
32:16 英伟达对 TPU:护城河还剩什么 ▶ 正在看
38:37 Cerebras 与快模式的经济账 ▶ 正在看
43:52 十年视角与局部最小值陷阱 ▶ 正在看
50:46 算力紧缺、毛利率与杠杆之忧 ▶ 正在看
56:28 每吉瓦不等值:数据中心分层 ▶ 正在看
64:13 Neocloud 为什么能赢,黄仁勋在下什么棋 ▶ 正在看
本期小问 · 档案清单
23:16 AI 三年来效率翻百倍,功劳在芯片还是在模型? ▶ 正在看
32:16 模型和芯片互相定制后,还能说谁比谁更好吗? ▶ 正在看
43:52 专用 AI 芯片会不会只是一个局部最优的陷阱? ▶ 正在看
50:46 AI 数据中心的高杠杆狂建,何时会变成泡沫? ▶ 正在看
本期讲者
迪伦·帕特尔SemiAnalysis 创始人兼首席分析师。2020 年以博客起家,将公司发展为覆盖半导体供应链与 AI 基础设施的顶尖研究机构,主导 InferenceX 持续推理基准项目。
肖恩·马奎尔红杉资本合伙人,早期投资 SpaceX,关注硬件、半导体与 AI 基础设施。
黄共宇红杉资本合伙人,专注 AI 应用与模型公司投资,私募股权出身。
01加油站里训练的第一个神经网络
0:00
I think it's really fun inside of semi analysis because we have 90 people and like a big chunk of them are technologists engineers across the whole supply chain. Um, and then a big chunk is people who are formerly at hedge funds. And you see these arguments like people are like, "Oh, well that doesn't matter." And it's like, then someone's like, "Well, but cost." And then someone the engineers like, "No, no, no, but this technology is the coolest." And you see this you see this organically like fight it out. Um, and and we're pretty informal and you know, given the fact that I was a for moderator, you can imagine what the the enjoying it. You don't wrestle with a pig because a pig enjoys it, right?
我觉得在 SemiAnalysis 内部真的很有意思,因为我们有 90 个人,其中很大一部分是技术人员、工程师,覆盖整个供应链。然后还有很大一部分是以前在对冲基金工作的人。你会看到这样的争论:有人会说:“哦,那个不重要。”然后另一个人说:“可是成本呢。”接着工程师会说:“不不不,但这个技术是最酷的。”你会看到这种争论自然而然地吵起来。而且我们挺随意的,考虑到我以前是论坛版主,你可以想象那种乐在其中的感觉。你不该跟猪摔跤,因为猪乐在其中,对吧?
便签笔记
0:33
Exactly.
没错。
便签笔记
0:50
We're here in the semi analysis office with Dylan Patel. You know, I'm Sean from Sequoia. My partner Sonia Huang. It's pretty insane what you've done. Semi semis 5 years ago were not very sexy in the west. They were sexy in the east but uh people here in the west had kind of forgotten about them. You did not forget about them though. You went very long. You created probably the premier research company in the space that's been educating the world and you know the state of the art from very technical details to supply chain you know to the bigger picture. Um there's rumors that semi analysis recently passed 100 million of revenue. I don't know how accurate those are. Whatever the numbers are you guys are crushing.
我们现在在 SemiAnalysis 的办公室,和 Dylan Patel 在一起。我是红杉资本的 Sean。这是我的合伙人 Sonya Huang。你做成的事情相当疯狂。五年前半导体在西方并不怎么性感。它们在东方很性感,但西方这边的人算是把它们给忘了。不过你没有忘。你重仓押注了它。你创立了这个领域里大概是最顶尖的研究公司,一直在向全世界普及知识,从非常技术性的细节,到供应链,再到更宏观的图景。有传言说 SemiAnalysis 最近收入突破了一亿美元。我不知道这些传言有多准确。不管数字是多少,你们都做得非常出色。
便签笔记
1:30
>> It's it's as accurate as the information is. Yeah. Cool. You know, you never you never know. Um there's also rumors that you might start a venture fund like you know I I hear all the time in the ecosystem people wanting you know affiliation with semi analysis. You you've built this trusted brand and so whatever you do it's working. This clearly like just the beginning of the journey for you. Congratulations all of that. But how did this happen? Like how did you first question is like what is the background? How did you kind of get to where you are now? Well, well, when I was a young boy in the, you know, coming out of the womb. No. So, so, okay. So, I grew up in like a small business. My parents had a motel. We lived in the motel. We laid our gas station. So, you know, uh I was selling. You know, I joke a lot of times the first neural network I trained was uh racially and and visually profiling people based on when they enter the gas station, which cigarette to uh pick. Right. Basically,
>> 这个信息有多准,它就有多准。是啊。酷。你永远不会知道。还有传言说你可能要开始做风险投资基金,我在生态圈里总是听到有人想和 SemiAnalysis 建立某种关联。你打造了一个受信任的品牌,所以不管你做什么都行得通。这显然只是你这段旅程的开始。恭喜你取得这一切。但这是怎么发生的?我的第一个问题是:你的背景是什么?你是怎么走到今天这一步的?嗯,嗯,当我还是个小男孩,你知道,刚从娘胎里出来的时候。不是啦。那,好吧。我是在一个小生意的家庭里长大的。我父母开了一家汽车旅馆。我们就住在那家汽车旅馆里。我们还开了加油站。所以,你知道,我那时候在卖东西。我经常开玩笑说,我训练的第一个神经网络就是根据人们走进加油站时的样子,对他们进行种族和外貌上的画像,判断该拿哪种香烟。对吧?基本上就是,
便签笔记
2:23
you know, the cigarettes were all extrudeed across the top and I was too short to actually like, you know, reach them and technically it wasn't legal to sell cigarettes at that age, but whatever. I I had to move the step stool over to the right area. >> I started working my first job was before it was legal, too. So, but it's good experience. >> Well, I didn't get paid, right? It's a family business. Same. >> Same. Um, but yeah, we had our motel and then across the street was our gas station. So, you know, sometimes, you know, you know, someone would walk in and so like if a old white lady with curly hair walked in, I'd move the ladder or the step stool over to where the camels are. And if you know and you know different different age, demographic, profession, you know, race, etc. I would move the step stool over and I joke this is the first neural network I trained because if I waited for them to tell me I'd have to like move it over and then I'd step up versus like just being ready. Um so you know
你知道,香烟都摆在最上面一排,而我个子太矮,够不到它们,而且严格来说那个年龄卖香烟也不合法,但无所谓啦。我得把小凳子挪到对的位置。>> 我开始工作、我的第一份工作也是在合法年龄之前。不过那是很好的经历。>> 呃,我没拿到工资,对吧?那是家族生意。一样。>> 一样。不过是的,我们有那家汽车旅馆,然后马路对面是我们的加油站。所以,你知道,有时候,有人会走进来,比如说,如果一位卷发的白人老太太走进来,我就会把梯子或者小凳子挪到放 Camel 香烟的地方。而如果,你知道,不同的年龄,
便签笔记
3:06
menthols versus, you know, 100 slims and all these things, you know, I joke that's the first neural network I trained. But I grew up in family businesses. Um lived in a motel and um it all really goes back to when I was like, you know, it was my 8th birthday. Um, my birthday's in May. Um, and it was April when the Xbox 360 was announced. Um, for my birthday, I didn't ask for the Xbox or I didn't ask for a birthday gift. My parents asked what I wanted. I asked for it for Christmas. Uh, we celebrated Christmas, but there was no way, at least at the time, I thought there was no way they would ask would give me the Xbox 360 for Christmas and so I got it for I asked for my birthday for tab for Christmas. Anyways, Christmas comes around, I get it. Um, you know, fast forward a couple months, my cousin who lives in Alabama, they also lived in a motel, was going to come over for spring break, um, for his spring break and we were going to just hang out at my house and he's in between me and my older brother in age,
薄荷味的,还有那种100mm细支之类的,你知道吧,我常开玩笑说那是我训练的第一个神经网络。但我是在家族生意里长大的。嗯,住在汽车旅馆里,这一切其实都要追溯到我大概8岁生日的时候。嗯,我生日在五月。嗯,Xbox 360是四月份发布的。嗯,生日的时候我没要Xbox,或者说我没要生日礼物。我爸妈问我想要什么,我说我想圣诞节要它。呃,我们过圣诞节,但当时我觉得他们绝对不可能圣诞节给我买Xbox 360,所以我生日的时候就说我圣诞节要那个。总之,圣诞节到了,我拿到了。嗯,快进几个月,我住在阿拉巴马的表弟,他们家也住在汽车旅馆里,春假要来我家,嗯,他放春假,我们就打算在我家一起玩,他的年龄在我和我哥之间,
便签笔记
3:53
brother's a bit more jockey. Um, so he didn't really care too much about the Xbox. He played sometimes, but he didn't really care. Um, but my cousin, you know, I wanted to think I'm him to think I was cool, right? You know, so I bragged many times on the phone. I was like, "Yeah, I got an Xbox." And then the Xbox broke. There was something there's a hardware defect called the red ring of death. Um, but long story short, I had to open it up and, you know, short the temperature sensor and it fixed it.
我哥比较偏运动型。嗯,所以他对Xbox不怎么感兴趣。他偶尔也玩,但不太在意。嗯,但我表弟嘛,你知道,我想让他觉得我很酷,对吧?所以我在电话里吹了好多次牛,说"是啊,我有台Xbox"。然后Xbox就坏了。有个硬件缺陷叫"红环死亡"。嗯,长话短说,我得把它拆开,然后把温度传感器短接一下,就修好了。
便签笔记
4:16
Um, but I there was many other tricks I tried first and none of them worked. Um, and so that's sort of how I like got into hardware. I was like open Pandora's box. By the time I was 12, I was like on these forums a lot reading uh posting a lot and this is around the time when Reddit ate all other forms and so I became a moderator of you know Android and Apple and Google as well as like hardware and was watch you know looking at Intel, Nvidia and AMD and all these other forms right? I was build a PC. All these forms I was watching, reading, posting a lot, but some of them I was moderating a lot. Um, and so, you know, smartphones, watching smartphones develop from like very simple to speed racing to being technologically more advanced than PCs, um, in many ways architecturally and and same with like, you know, all the in GPUs like just tracking and watching at that, reading every comment. Um, always having the economic tinge because I grew up in a small business. So, I was always looking at the economics, right? There was a
嗯,但在那之前我还试了很多别的招,都没用。嗯,所以我差不多就是这么入了硬件的坑。感觉像是打开了潘多拉的盒子。到我12岁的时候,我已经泡在那些论坛上大量地看、大量地发帖,而那正是Reddit吞并其他所有论坛的时期,所以我成了安卓、苹果、谷歌以及硬件板块的版主,还关注英特尔、英伟达、AMD等等各种论坛,对吧?还有装机论坛。这些论坛我都在看、在读、在发帖,其中有些我还当版主。嗯,所以你知道,智能手机,看着智能手机从非常简单发展到拼参数竞速,再到在架构上很多方面比PC还先进,嗯,GPU那边也一样,就是一直追踪、一直看,每条评论都读。嗯,因为我在小生意人家庭长大,所以总带着一点经济视角。我总是在看经济账,对吧。有段时间
便签笔记
5:06
time where all the like I'd say neck beards on the internet loved AMD GPUs and like I personally had bought an AMD GPU too because price performance but then when it came down to like what's technically better I'd always be like no no no Nvidia is better because they use a smaller chip to get you know better performance at better power efficiencies and their margins better and and so like I would always like talk about how Nvidia's margins were better than than AMD's in the GPU landscape and so it's like very fun.
网上那些我姑且叫"技术宅老哥"的人都超爱AMD的显卡,我自己也买过AMD的显卡,因为性价比高。但真要说技术上谁更好,我总会说不不不,英伟达更好,因为他们用更小的芯片就能在更好的能效下拿到更好的性能,而且他们利润率更高。所以我总在讲英伟达在GPU这块的利润率比AMD高多少,这特别有意思。
便签笔记
5:32
>> And you were 12 at the time. I started moderating when I was 12, but this is all through my teenage tween age and high school years, right? >> Do you have any other weird hobbies or was it just semis? >> I played a ton of Starcraft. At one point, I was grandmaster on the North American ladder. Starcraft 2. >> Very serious. >> So, you've gotten just obsessively good at multiple things. >> Yeah. I mean, it's it's it's obsession is good. >> How were your grades? >> Um, they were decent. Um, I would say like I had mostly A's, but they're classes that I like were thought were really boring or, you know, I just didn't enjoy. Um, like Spanish I got like not the greatest grades. Um, you know, but but it was like I speak fluent Spanish by the way, so it's really dumb.
>> 而当时你才12岁。我12岁开始当版主,但整个过程贯穿了我的青春期、少年时期和高中阶段,对吧?>> 你还有别的什么奇怪爱好吗,还是只有半导体?>> 我玩了超多《星际争霸》。有段时间我在北美天梯上是宗师段位。《星际争霸2》。>> 相当认真啊。>> 所以你在好几件事上都痴迷到很强的程度。>> 是啊。我觉得,痴迷是件好事。>> 你成绩怎么样?>> 嗯,还不错。我大部分科目都是A,但有些课我觉得特别无聊,或者就是不喜欢。嗯,比如西班牙语我成绩就不怎么样。嗯,不过顺便说一句,我现在西班牙语说得很流利,所以这挺蠢的。
便签笔记
6:15
But like it's just sort of >> maybe that's why you didn't get a good grade. >> I didn't learn Spanish till later to be fair. But yeah, so sort of my grades were fine, right? Like I mean I like they were fine enough for Asian parents. >> I was better than most of school but you know it wasn't like you know tryh hard maxing for like you know all A's. >> Okay. So you're very much a student of the internet then this is how you how you develop this expertise. At what point do you decide to start semi analysis and what's been the biggest surprise since starting the company?
但就是那种 >> 也许这就是你没考好的原因。>> 公平地说,我是后来才学的西班牙语。但总之我的成绩还行,对吧?我是说,对亚洲家长来说算过得去。>> 比学校里大多数人好,但也不是那种为了全A拼命卷的程度。>> 好的。所以你基本上是互联网自学出来的,这就是你积累这些专业知识的方式。那你是什么时候决定创办SemiAnalysis的?创业以来最让你意外的是什么?
便签笔记
02离职、丧亲与匿名博客的诞生
6:42
>> Yeah so I went to school I got a few degrees in stuff that wasn't related to semiconductors. Um was a quant for two years at a small quant risk firm. Um and then basically, you know, there's a culmination of events that happened, right? One was that my um you know, sort of like I got screwed out of a bonus. I I'd made my company many millions of revenue of risk-free revenue because I exploited like a risk, you know, thing in the market. Um you know, I think well over 10 million and they then someone else took credit for my work and all this sort of stuff. But eventually I did get rightsized. But you know, I lost a social contract with the company I was working with. Um add some you know, my grand my grandparents grew up in my house with us, right? are in the motel with us. Uh they lived with us and so you know very close with them and my grandmother got dementia and she forgot who I was and she she fell down some stairs and had like a tragic accident and passed away. So all of that happened
>> 是这样,我上了大学,拿了几个和半导体无关的学位。嗯,在一家小型量化风险公司做了两年量化。嗯,然后基本上是好几件事凑到一起,对吧。一件是我,嗯,我被坑掉了奖金。我给公司赚了好几百万的无风险收入,因为我利用了市场上的一个风险漏洞。嗯,我觉得远超一千万,然后别人把我的功劳拿走了,诸如此类。不过最后我确实被优化了。但你知道,我和当时那家公司之间的社会契约破裂了。嗯,还有就是,我祖父母和我们住在一起,对吧,就在汽车旅馆里和我们一起。呃,他们和我们同住,所以我和他们非常亲近,然后我奶奶得了痴呆症,她不认得我了,她从楼梯上摔下来,出了很惨的意外,去世了。所有这些都发生在
便签笔记
7:29
in early 2020. Um additionally there were some like you know girl things and so you know there's a few things that happened that made me like kind of very sad. Um and and and so all of those things sort of culminated. Then co happened and my brother's like dude just just come stay with me. He lived in Nashville so I came and stayed with him in Nashville. We were like, "Oh, lockdowns will be a few weeks. You can stay with me while they happen and then you can go back home and you know, whatever." Famous last words.
2020年初。嗯,另外还有一些感情上的事,所以有那么几件事让我特别难过。嗯,这些事全都堆到了一起。然后疫情来了,我哥说,"兄弟,你就过来跟我住吧。"他住在纳什维尔,所以我就去纳什维尔和他一起住。我们当时想,"哦,封控就几周吧。封控期间你可以住我这儿,然后就能回家了,随便怎样。"结果这话成了著名的flag。
便签笔记
7:54
Lockdowns lasted much longer. But, you know, living with my brother for a few months, you know, was like sort of like, okay, didn't know what I was doing. I was now at my brother's home. Um, everything was his rules. You know, sort of like, you know, him and him and his fiance at the time, now wife, you know, were like there. And so, like, I basically had to tiptoe around, but I didn't care about my job. And so, I was like posting even more than normal. I'd always been posting a lot on the internet. I'd always been trading stocks a lot, but like I made a lot of money shorting COVID and long in COVID and like all this stuff. Semiconductor shortages happened around then too. And anyways, I was like very much obsessed with posting and and things like that.
封控持续得久多了。不过,跟我哥一起住了几个月,感觉就是,好吧,我也不知道自己在干嘛。我当时住在我哥家里。嗯,一切都按他的规矩来。你知道,他和他当时的未婚妻、现在的妻子都在那儿。所以我基本上得小心翼翼地过日子,但我根本不在乎我的工作。所以我发帖发得比平时还多。我一直都在网上大量发帖,也一直在大量炒股,我在疫情期间做空又做多赚了不少钱,诸如此类。半导体缺货潮也差不多是那时候发生的。总之,我当时特别痴迷于发帖之类的事。
便签笔记
8:27
And eventually um around that time someone I got into an argument with someone on the internet and they doxed me, right? They they publicly revealed my identity for my anonymous account. And at the time I was like, "Oh no, I scared. I stopped posting for like three weeks and I was like, what am I doing? Why do I care?" So then I just started posting under I had like had like blogs and stuff as well. I made a real blog, semi- analysis, and on my 24th birthday, I posted um you know, two blogs and and then from there it just like it was not a newsletter, but I got so much traction because now instead of posting on an anonymous name, it was a real name and I put a lot more effort into those two posts than I usually did. Instead of like posting on the internet, it was like real effort into the blog. Um you can actually go back and read those if you want. They're they're not that great, but you know, they were they were good for the time. They were the best stuff you could find on the internet
然后嗯,大概在那段时间,我在网上跟人吵了一架,他们把我人肉了,对吧。他们公开了我匿名账号背后的真实身份。当时我想,"哦不,我好怕。"我停更了大概三周,然后我想,我这是在干嘛?我干嘛在乎这个?于是我就开始用真名发了,我本来也有一些博客之类的。我做了一个正经的博客,叫SemiAnalysis,在我24岁生日那天,我发了嗯,两篇博客,然后从那以后就,它当时还不是newsletter,但我获得了巨大的关注,因为这次不是用匿名账号发,而是用真名,而且那两篇我比平时投入了多得多的精力。不像随手在网上发帖,那是真的花心思写的博客。嗯,你现在还能回去读那两篇。它们没那么好,但在当时算不错了。它们是当时互联网上关于半导体
便签笔记
9:11
about semis. Um, and and I just kept posting, posting, posting. I started getting a lot of consulting business. You know, 2020, I also sort of I was again crashing out. Didn't know what I wanted to do. So, I uh packed everything up or sort of I I I took my truck, I bought a tent that fits on the back of the tent, truck, um, bought a air mattress, whatever, and would like and drove around all these national parks all around America. And so, like two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like $30 a night for a room.
你能找到的最好的东西。嗯,然后我就一直发一直发一直发。我开始接到很多咨询业务。你知道,2020年,我又一次有点崩溃,不知道自己想做什么。所以我,呃,把东西全打包,或者说我开着我的皮卡,买了个能装在车斗上的帐篷,嗯,买了个充气床垫之类的,然后就开着车跑遍了全美国的国家公园。所以一周里有两三天或三四天,我会随便找家汽车旅馆住,把房价砍到一晚30美元。
便签笔记
9:38
And I was work on something else stuff. And then the weekends I'd read books and oftentimes read textbooks um while in some random national park or hiking and listen to audiobooks um about semiconductors about AI about all the things that I cared a lot about and got way more educated over these six months where I'm just like going to every national park. Um and the whole time I was I was alone the whole time I was posting blogs. Um everyone was like DD what the are you doing >> pre-larink or the very early days of Starling?
然后我会在那儿做点别的事。周末我就读书,经常是读教科书,嗯,在某个国家公园里或者徒步的时候,听有声书,嗯,关于半导体、关于AI、关于所有我特别在乎的东西。在那六个月里我学到的东西多得多,就是不停地去每一个国家公园。嗯,全程我都是一个人,全程我都在发博客。嗯,所有人都说,兄弟你他妈到底在干嘛 >> 那是星链之前,还是星链非常早期的时候?
便签笔记
03一年四十场会议怎么学供应链
10:05
>> Pre-star link pre-tar link. Um yeah, so it was like very much like what are you doing? Um I travel around Latam again like for for a year initially with my friend and then with my ex you know for you know about a year and then I go then 22 23 24 end of 21 22 23 and 24 I'm completely I'm still completely homeless since mid2020 right um but I'm traveling around to every conference in the world. I go to 40 plus conferences a year no matter where in the supply chain it is. I'm like oh that looks interesting. I guess I'll go to that. And I'm like I went to one conference like wow this is amazing. you get to talk to the experts and they just like they they're they they're going to talk to you because and then you're so excited and in the case of semiconductors everyone's a boomer so it's like it's great to like you know they're like they don't see young people who are like excited about it so they're really happy to tell stuff and so you just have to ask on this was there like a part of the supply chain or one of
>> 星链之前,星链之前。嗯,是的,所以大家就是那种"你在干嘛啊"。嗯,我又跑去拉美转了一圈,一开始跟我朋友,大概一年,然后跟我前任,你知道,大概一年,然后到了22、23、24年,21年底、22、23、24年我完全,我从2020年中开始就一直是无家可归的状态,对吧。嗯,但我跑遍了全世界的各种会议。我一年参加40多场会议,不管是供应链的哪个环节。我就想,哦,那个看起来挺有意思,那我去吧。我去了一场会议,心想,哇,这太棒了。你能跟专家们聊天,而他们真的愿意跟你聊,然后你就特别兴奋。而且在半导体这行,大家都是老一辈,所以特别好,他们很少见到对这个这么有热情的年轻人,所以他们非常乐意跟你讲,你只需要开口问。>> 在这些
便签笔记
10:56
these conferences that you know particularly changed your view of the semi-world or that you felt then or feel now is particularly underrated. I think I think the trade shows like r and conferences range really widely. Um obviously some of the you know the ones I have the most fun at you know include NERPS. Why why is that? Because it's 20,000 AI researchers and they're generally in my distribution of age range. So it's like a lot of fun but they're also like leading AI researchers and it's a lot of fun and lot you learn a lot. Um there's also a lot of parties and then it ranges all the way to like you know there's random chemical conference in Japan where it's 300 Japanese dudes. It's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Uh, everyone else speaks only Japanese, and you're like, h, I guess they're still pretty interesting and fun. I think I think like one thing that I have like a skill set of is like I'm able to bond with anyone regardless
会议里,有没有哪个供应链环节或者哪场会议特别改变了你对半导体世界的看法,或者是你当时觉得、现在依然觉得被严重低估的?我觉得展会和会议的差异非常大。嗯,显然,我玩得最开心的那些,包括NeurIPS。>> 为什么是它?因为那里有两万名AI研究者,而且他们的年龄段大致跟我一样。所以特别好玩,但他们同时也是顶尖的AI研究者,很有意思,你能学到很多。嗯,而且派对也很多。然后另一个极端是,日本有个不知名的化学会议,全场300个日本大叔。大概有20个ASML的人、20个台积电的人、20个英特尔的人,而这些是全场仅有的会说英语的人。呃,其他人只说日语,你就想,嗯,我猜他们还是挺有意思、挺好玩的。我觉得我有一项技能,就是我能跟任何人建立联系,不管
便签笔记
11:45
of their background and like who they are. I'm able to talk to them, find something interesting to talk about. Oftentimes, it's the tech stuff, but, you know, it's it's and so I think like the most interesting conferences are oftentimes like, you know, the really big ones because that's where the biggest stuff is happening. Um but I think the niches that are really really exciting is like you know SPIE um so there's IE which is international electrical engineering something um and there's SPI which is another ecosystem.
他们的背景是什么、他们是什么样的人。我都能跟他们聊天,找到有意思的话题。通常是技术的东西,但总之,所以我觉得最有意思的会议往往是那些特别大的,因为最重磅的事情都在那儿发生。嗯,但真正让我兴奋的细分领域,是像SPIE这种,嗯,有IEEE,就是国际电气工程什么什么,嗯,还有SPIE,那是另一个生态。
便签笔记
12:11
SPIE conferences are super super deep in details. Every single one that I went to, especially like SPI advanced lithography or SPI photo mask, I went to them the first time I didn't even understand 90% of what I heard. And then I read red, red, red, I had made some context, of course, and then next time I went, I understood like half of what I went to. Third time I went, I understood like 75% of what I went to. Even now, I went and I was like, I still don't understand everything that's going on. Whereas like you go to like Nurips, you know, a couple times you can understand, okay, what's neurosymbolic reasoning?
SPIE的会议在细节上极其极其深入。我去过的每一场,尤其是SPIE先进光刻或者SPIE光掩模,第一次去的时候我听到的东西90%都听不懂。然后我读啊读啊读,当然也积累了一些背景知识,第二次去的时候,我大概能听懂一半。第三次去,我能听懂大概75%。就算是现在,我去了还是觉得,我依然不能完全搞懂在发生什么。而像NeurIPS那种,你去个几次就能明白,好,什么是神经符号推理。
便签笔记
12:41
Okay, what's this? What's that? like you can you can kind of get a mapping of what everything is pretty quickly but some parts of the supply chain are so arcane and so deep and so technical. It takes a lot of times for you to even understand what's happening in and you know on everything right um for every research paper doesn't necessarily mean you didn't you know you go to a conference for a few reasons right you understand the research you understand but like it's all the research that's being published but what you really care about is understanding how does that research intersection intersect with technology also how does that research differ from what's there today and none of these research papers tell you what's happening today but then you just ask people and you you build contacts and you learn and then you like learn about the supply chain and oh this company supplies this company even though it's not publicly stated anywhere or like you know you learn that the the this chemical is like cost about this much
技术,还有那些研究和今天的现状有什么不同,而且这些研究论文都不会告诉你现在到底在发生什么,但你就是去问人,你去建立人脉,你去学,然后你就会了解到整个供应链,哦,原来这家公司给那家公司供货,尽管这在任何公开资料里都没写过,或者你会了解到,这种化学品大概值多少钱,一台设备大概要用掉多少,然后你 >> 你会听到那种恐怖故事,比如
便签笔记
04推理市场与 InferenceX 基准
13:26
and a tool uses about this much and you >> hear you hear the horror stories of like this chemical had a shortage and it totally threw off this part of the supply chain and then it turns out there's only three companies in the world that make that chemical and it's like >> my favorite one is I learned uh a Japanese guy at that specific Japanese uh conference that I went to where no almost no one spoke English in very broken English he told me about how uh his father worked in this in in in this industry in the 1980s that the the only factory in the world that built this chemical uh burned down and that caused memory prices to like double or triple and I was like wow not too different from today >> not not not at all >> crazy >> um >> inference going to be the biggest market on earth biggest market beyond earth agree or disagree >> um I mean obviously use of tokens is going to be the biggest market um and the value that's created from tokens is going to be the biggest market but I think tokconomics sort of the use of
这种化学品短缺了,就把供应链的这一环彻底打乱了,然后你才发现全世界只有三家公司能生产这种化学品,那种感觉就是 >> 我最喜欢的一个例子是,我在那个日本的会议上认识了一个日本人,那个会几乎没人说英语,他用非常蹩脚的英语跟我讲,他父亲在八十年代就在这个行业里工作,当时全世界唯一一家生产这种化学品的工厂烧掉了,结果导致内存价格翻了一倍甚至两倍,我当时就想,哇,这跟今天也没差多少 >> 一点、一点都没差 >> 太夸张了 >> 嗯 >> 推理会成为地球上最大的市场吗,甚至是地球之外最大的市场你同意还是不同意 >> 嗯,我是说,很明显 token 的使用会成为最大的市场,而 token 所创造的价值也会是最大的市场,但我觉得 tokenomics,也就是 token 的使用、AI 的普及,才是当下正在发生的最重要的事情,
便签笔记
14:19
tokens adoption of AI sort of is the most important thing that's happening and inference whether it's open models or closed models will be like one of the biggest markets in the world much bigger than oil I think much bigger than like you know many other parts like inference of AI will be you know many percentage points of the GDP yeah right >> what you've done with inference X I think is you know industry standard maybe say a word on why you started it what it does and you know what do people misunderstand about uh performance benchmarking on inference >> yeah So, so to zoom back, right, like semi analysis, uh, we do a lot of stuff that's like, you know, a lot of it is like research for institutional clients and and our subscription versus products, but a lot of it is also like, hey, you know, this would just be cool to figure out. Let's figure out how to figure it out and just post it publicly.
而推理,不管是开源模型还是闭源模型,都会成为世界上最大的市场之一,比石油大得多,我觉得会比石油大得多,也比很多其他领域大得多,AI 推理会占到 GDP 好几个百分点,对吧 >> 你们用 InferenceX 做出来的东西我觉得已经是行业标准了,能不能说说你为什么要做它,它到底是干什么的,以及大家对推理性能基准测试有哪些误解 >> 好,那我先往回讲一点,就是SemiAnalysis,我们做很多事情,其中很大一部分是给机构客户做研究,还有我们的订阅和产品业务,但也有很大一部分是那种,嘿,这事儿要是能搞明白就太酷了,那我们就想办法搞明白,然后公开发出来。
便签笔记
15:02
And that gets, you know, more and more scale. And so we've done this with a lot of GPU benchmarking and testing and training performance and inference performance, but you know, ultimately we saw like inference benchmarking was like point in time. you know, you test it and you take some time, you release it and it's like slow and arcane and out outdated because models change all the time. Every I feel like every week there's a new model whether it's a Chinese model or you know today mythos 5, Fable dropped and new models are coming out all the time. Um on the software layer uh PyTorch, VLM, SG lang um new drivers, new new something drops, you know, in fact the update cycle for most of these libraries is twice a week.
这样就能获得越来越大的影响力。所以我们在很多 GPU 基准测试、评测、训练性能和推理性能上都这么做了,但最终我们发现,推理基准测试基本上是某个时间点的快照。你测一轮,花上一段时间,发布出来,结果又慢又晦涩,而且已经过时了,因为模型一直在变。感觉每周都有新模型出来,可能是中国的模型,也可能是今天 Mythos5、Fable 发布了,新模型一直在冒出来。嗯,在软件层面,PyTorch、vLLM、SGLang,新驱动、新的什么东西一直在发布,事实上,这些库里大多数的更新周期是一周两次。
便签笔记
15:41
So you basically have the software updating all the time and therefore performance changing. Um you know new inference optimizations are coming out and those get updated and and so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down which is why we've seen you know model cost drop for equivalent quality by like 60x a year. It's incredible. Um but to stay on top of that you can't have point in time benchmarking. You need to have benchmarks be living and breathing i.e. you know constantly running on the latest hardware on the latest models. And so we embarked on a project and we got a lot of buyin from the ecosystem. This was only possible because we had you know enough aura with some of the ecosystem where we're able to get coreweave and cruso and nebus and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us um compute and then we were able to work with SG Lang and VLM and now Radix Arc and InRact uh which are the private
所以基本上软件一直在更新,性能也就一直在变。嗯,新的推理优化不断出现,然后又被更新进去,所以我觉得这就是一次又一次接连不断的突破,持续推动效率提升、成本下降,这也是为什么我们看到,同等质量下模型成本一年下降了大概 60 倍。这太惊人了。但要跟上这个节奏,你就不能做时间点式的基准测试。你需要让基准测试是活的、会呼吸的,也就是说,持续地在最新硬件、最新模型上跑。所以我们就启动了一个项目,而且我们获得了整个生态的大力支持。这之所以能实现,是因为我们在生态里积累了足够的声望,能让 CoreWeave、Crusoe、Nebius、Oracle、微软、亚马逊、谷歌和 OpenAI 都给我们捐算力,然后我们又能和 SGLang、vLLM,还有现在的 Radix Arc和 InRact 合作,这些是主导那些开源项目的私营公司,
便签笔记
16:30
companies who are sort of leading those efforts um the open source efforts um to collaborate with us. We're able to get Nvidia and AMD and Google and Amazon now because we're adding TPUs and trrenium uh to collaborate. Now we've got all these people collaborating. We've got over $50 million of hardware uh donated to us. Um once we launch TPUs and trainum it actually should be over $und00 million of hardware. Um you know maybe about like 15 different chip types all running these benchmarks every single day on all the latest model, right? the best model from Moonshot, the best model from Alibaba, the best model from um there's about five different Chinese models, the best open source models, the best Chinese labs there. We run benchmarks on their models every day and then also the best US open source models um GPTOSS, Neotron, etc. So we're running these benchmarks every day um in an automated fashion and they run on these these servers that are dedicated to us for inference benchmarking and we
跟我们一起协作。我们也拿到了英伟达、AMD,还有谷歌和亚马逊的合作,因为我们正在加入 TPU 和 Trainium。现在这些人全都在跟我们合作。我们收到了超过 5000 万美元的硬件捐赠。嗯,等我们把 TPU 和 Trainium 上线之后,硬件价值应该会超过 1 亿美元。嗯,大概有 15 种不同的芯片型号,每天都在最新的模型上跑这些基准测试,对吧?Moonshot 最好的模型、阿里巴巴最好的模型、还有大概五个不同的中国模型里最好的、最好的开源模型、那边最好的中国实验室。我们每天都在他们的模型上跑基准测试,同时也跑美国最好的开源模型,GPT-OSS、Nemotron 等等。所以我们每天都在以自动化的方式跑这些基准测试,它们跑在那些专门给我们做推理基准测试的服务器上,我们会
便签笔记
17:21
sweep across so many different configurations and optimization types and then what it creates is and all the results are public and all the configurations are public. So now we have the paralo optimal curve because a lot of you know times when people are comparing inference performance they're like taking a suboptimal curve or point for someone else and comparing it to their optimal one. And it's like, well, yeah, I can make I can I can stick, you know, if I drove a Porsche versus like some some race car driver, obviously I'd drive it slower. The same thing with inference benchmarking. And so what we did is we created open- source uh basically containers for the optimal points across every uh point on the interactivity, i.e. how fast is it responding to me versus you know batch size, i.e. how many users am I simultaneously serving curve? And so now anyone who wants the optimal point can just go to inference X download it and run that as the optimal point and they can check every day if they want or they
横扫非常多不同的配置和优化方式,最后产出的结果,以及所有的配置,全部都是公开的。所以现在我们有了帕累托最优曲线,因为很多时候人们在比较推理性能的时候,会拿别人一个次优的曲线或者次优的点,去跟自己的最优点比。那就好比说,是啊,我要是开一辆保时捷,去跟一个专业赛车手比,那我肯定开得更慢。推理基准测试也是一样。所以我们做的事情是,我们开源了基本上就是各个最优点的容器,覆盖交互性曲线上的每一个点,也就是它对我的响应有多快,对应到 batch size,也就是我同时服务多少用户,这样一条曲线。所以现在任何人想要那个最优点,直接去 InferenceX 下载下来跑就行了,那就是最优点,他们想的话可以每天检查一次,或者
便签笔记
05吞吐与延迟:一条决定一切的曲线
18:07
can even autod download the most optimal point for that model and and their inference performance will be near peak. Um >> is that curve like the most important curve in your opinion? The throughput interactivity curve is the most important one. >> Yeah, I think I think um most things in hardware infrastructure uh model application layer everything is downstream of that curve, right? Is it is it something that needs to be super super fast, super low latency? Um, and I don't really care about the cost, so I make batch size very low and I use techniques like speculative decoding or multi-token prediction heavily and and there's so many, you know, possible techniques there. Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things. I don't use these techniques that actually are worse on cost efficiency but help you with speed for an individual user because I just want to pack a bunch of users. I don't care if the document takes all
甚至可以自动下载那个模型对应的最优点,这样他们的推理性能就会接近峰值。嗯 >> 在你看来这条曲线是不是最重要的曲线?吞吐量与交互性的曲线是最重要的那条。>> 是的,我觉得,嗯,硬件基础设施、模型、应用层,几乎所有东西都是那条曲线的下游,对吧?它是不是那种需要超级超级快、超低延迟的东西?嗯,而且我并不太在乎成本,所以我把 batch size 压得很低,然后大量使用投机解码或者多 token 预测这类技术,这里面可用的技术非常多。还是说,其实我是在批处理一大堆文档,我根本不在乎这些东西,我不会用那些其实在成本效率上更差、但能提升单个用户速度的技术,因为我只想把一堆用户打包在一起。我不在乎这份文档是不是要处理
便签笔记
18:53
night to process, right? Um, and right now the way we treat AI infrastructures, it's like one-sizefits-all. But over time, we're going to get to the point where, you know, there's stuff where you you have batch workloads or, you know, you need instant response and there's there's the whole curve that's going to matter for uh users. And so we see this with entropic, right? Cloud code fast mode cost way more than regular mode. Um, and same with open eyes priority Q thing. Um, >> sorry, dumb question. How does cost factor into the chart? So if if I let's say imaginary example, I have 100 I have a batch size of 100, >> okay?
一整晚,对吧?嗯,而现在我们对待 AI 基础设施的方式,基本上是一刀切。但随着时间推移,我们会走到那样一个阶段,就是有些场景是批处理负载,有些场景需要即时响应,整条曲线都会对用户产生影响。我们在 Anthropic 身上就看到了这一点,对吧?Claude Code 的 fast 模式比常规模式贵得多。嗯,OpenAI 的优先队列也是一样。嗯 >> 抱歉,问个傻问题。成本在这张图里是怎么体现的?比如说举个假想的例子,我有 100,我的 batch size 是 100,>> 好。
便签笔记
19:25
>> And I can do 10 tokens per second per user. So in total I'm doing a thousand tokens per second uh off of that one piece of compute. That's one side of the curve. Super slow, 10 tokens per second. Um you know, other side is I have uh uh 500 tokens per second, but I only have one user. And so maybe 250 tokens per second, one user. And then there's points on the middle that are more fraal optimal, right? the average person actually wants like 50 or 100 tokens a second and maybe you know the this the the number of users I can batch together. So the curve is okay a thousand tokens total uh per second or 250 tokens total per second depending on how many users I batch and there's a curve in the middle and so ultimately some workloads will actually want the 4x cost decrease because the same unit of hardware can do a th000 versus 250 and some users I'll pay 4x more because I don't care about the price I care about time because the person using the tokens is expensive or the feedback loop that I
>> 然后每个用户我能做到每秒 10 个 token。那总共就是每秒 1000 个 token,用的是同一块算力。这是曲线的一端。超级慢,每秒 10 个 token。嗯,另一端是我能做到每秒 500 个 token,但我只有一个用户。或者说每秒 250 个 token,一个用户。然后中间还有一些点是更接近帕累托最优的,对吧?普通人其实想要的是每秒 50 或者 100 个token,然后可能,我能打包多少用户在一起。所以这条曲线就是,要么总共每秒 1000 个 token,要么总共每秒250 个 token,取决于我打包了多少用户,中间是一条曲线,所以最终有些负载其实会想要那 4 倍的成本下降,因为同样一份硬件能做 1000 而不是 250,而有些用户我愿意多付 4 倍的钱,因为我不在乎价格,我在乎的是时间,因为用这些 token 的那个人很贵,或者说我这里的反馈循环很贵。如果让你猜的话,时间跨度你自己定,
便签笔记
06太空算力与每瓦智能的预测
20:18
have here is expense is is expensive. If you had to guess, you choose the time frame 10 years or 15 years. What percent of inference compete do you think will happen in space? >> Can be 0% 50% >> Sean 99% like >> this is a tough one. Um you choose the time frame like 10 whatever time frame and you're >> so I think I think the non- consensus or at least against SpaceX thing, you know, I love SpaceX by the way and I totally would buy the IPO if I could buy stocks. Um >> not investment. >> Not investment advice. Thank you. Thank you. not invested ice um from either um I don't think that space data centers will really matter in the next um you know 3 to 5 years um with that said I think in you know 20 years I think the vast majority of compute will be going in space um and so the real real factor there is sort of you know what's the cost it's the time frame it's the cost of building power on terrestrial land and how much power you going to be able to do on terrestrial land and I think obviously my views of where inference
10 年还是 15 年,你觉得会有百分之多少的推理算力发生在太空?>> 可以是 0%、50% >> Sean,99% 这种 >> 这题不好答。嗯,时间跨度你自己选,10 年或者别的什么,你 >> 我觉得,非共识的观点,或者至少是跟 SpaceX 唱反调的观点是,顺便说一句我很喜欢 SpaceX,如果我能买股票的话我肯定会打新。嗯 >> 这不是投资建议。>> 不是投资建议。谢谢,谢谢。我们俩都没有持仓。嗯,我不觉得太空数据中心在接下来三到五年里会真的有多重要。话虽如此,我觉得二十年后,绝大部分算力都会跑在太空里。所以真正的关键因素其实是,成本是多少、时间跨度是多久、在地面上建电力的成本有多高,以及地面上到底能做出多少电力。而且很显然,我对推理会
便签笔记
21:19
you know you know how many gigawatts or terowatts are devoted to inference is it's a crazy curve for me personally >> what's your forecast how many gigawatts or >> um yeah I think I think by you know 2030 just open anthropic we'll have over 100 gigawatts combined um and then you'll add you know meta and Google and you know so on and so on so forth it's it's a humongous amount of compute that will be dedicated to inference um and by like 2040 it'll be terowatts right um the the curve of like productivity that we're going to get and so you know inference deployments is going to be huge And so if you look at like 2040, I think like you know probably more than half of the incremental compute will be going in space. But if you look at 2030, I think it's sub 1%.
消耗多少吉瓦或者太瓦的看法,对我个人来说那是一条疯狂的曲线 >> 你的预测是多少吉瓦或者 >> 嗯,我觉得到 2030 年,光是 OpenAI 和 Anthropic 加起来就会超过 100 吉瓦,然后你再加上 Meta、谷歌等等等等,那将是一个极其庞大的算力规模,全部用于推理。嗯,到 2040 年就会是太瓦级别了,对吧,我们将获得的那种生产力曲线,所以推理部署的规模会非常巨大。所以如果你看 2040 年,我觉得新增算力里可能有超过一半会跑到太空去。但如果你看 2030 年,我觉得连 1% 都不到。
便签笔记
21:56
>> Do you think intelligence per watt has been increasing? Uh and then it seems like there's still a giant gap between where we are intelligence per watt versus like human biology and so like if we are do you think we are to close that gap? And if so, where is that game going to come from? >> Yeah, I think I think it often depends on what you're doing too, right? Like a TI84 is way more intelligence per watt in terms of doing math than us and it's like 30 years old, right? Obviously this is like a dumb dumb you know sort of >> general intelligence.
>> 你觉得每瓦智能一直在提升吗?而且看起来我们现在的每瓦智能跟人类生物体之间还有巨大的差距,那你觉得我们能不能把这个差距补上?如果能,那个突破会来自哪里?>> 是的,我觉得这往往也取决于你要做什么,对吧?比如一台 TI-84 在做数学这件事上,每瓦智能远比我们高,而且它都 30 年前的东西了,对吧?当然这是个很傻很傻的>> 通用智能。
便签笔记
22:20
>> Yeah. But general intelligence wise um so one of the things inference X does is we also measure the power and cost of all of these this hardware. And so we offer not just you know throughput versus interactivity we offer cost versus interactivity. We offer power versus interactivity. And so as far as has you know intelligence per watt been increasing? Um I mentioned you know it's been a 60x cost decrease for same benchmark level. Um we've also seen the same on on uh intelligence per watt. Um it's not been it's not been exactly 60x.
>> 对。但就通用智能而言,嗯,InferenceX 做的事情之一就是我们也会测量所有这些硬件的功耗和成本。所以我们提供的不只是吞吐量与交互性,我们还提供成本与交互性、功耗与交互性。那么关于每瓦智能是不是一直在提升,嗯,我提到过同等基准水平下成本下降了 60 倍。嗯,我们在每瓦智能上也看到了类似的情况。嗯,倒不是正好 60 倍。
便签笔记
22:49
It's been closer to like 40x. Uh some of the efficiencies are nonp power ways, but there's been a humongous improvement in in intelligence per watt on an annual basis at least so far this year, last year, year before, year before. And I expect that to continue as far as where we are from the human brain. We're we're many orders of magnitude away. Thankfully, doesn't really matter. We can devote a lot of power to computers. Much easier to power computers than human brains. Like you know, we have sickness, disease, and like food preferences. sleep.
更接近 40 倍。有些效率提升并不是来自功耗层面,但每瓦智能确实有了极其巨大的改善,至少按年来看,今年、去年、前年、大前年都是如此。我预计这会持续下去。至于我们离人脑还有多远,我们还差好几个数量级。好在这其实无所谓。我们可以给计算机投入大量电力。给计算机供电比给人脑供电容易多了。你知道,人会生病、会得病,还有饮食偏好、睡眠这些。
便签笔记
07硬件、系统、模型:百倍从何而来
23:16
>> Uh yeah, exactly. >> Let me just ask one more question on the like on the general theme in my opinion in terms of like you know intelligence per watt or intelligence per per dollar like any any of these metrics. I think there's kind of three levels of input. You can get hardware improvements that are where the hardware is more efficient. You can get lowlevel systems optimizations like kernel level you know improvements matri multiplication libraries you know things like that or you can get like highlevel like model level algorithmic improvements you know at the highest level. It see like to me it seems like in the last three years most of the gains have come from hardware level and you know and some from the model level like do you think that that is what do you agree with that do you think that's what look like in the future do like do you think there's a bunch of juice to squeeze in a say like kernel level like >> yeah Sean I completely disagree with you by the way great great that's why I'm
>> 呃,对,没错。>> 我再就这个大主题问一个问题,在我看来,就是那种每瓦智能或者每美元智能,随便哪个指标。我觉得输入端大概有三个层次。你可以获得硬件层面的改进,也就是硬件本身更高效。你可以获得底层系统优化,比如 kernel 层面的改进、矩阵乘法库之类的东西。或者你可以获得高层的、模型层面的算法改进,也就是最高层。在我看来,过去三年里大部分的提升来自硬件层面,还有一部分来自模型层面。你同不同意这个看法?你觉得未来会是什么样子?你觉得在比如 kernel 层面还有很多水可以榨吗?>> 是的,Sean,顺便说一句我完全不同意你的看法,很好很好,这正是我问这个问题的原因
便签笔记
24:21
asking this question >> um okay so I I think you know one way is to look at as these three different layers Um and in that sense like okay from hopper to blackwell which is all we've had over the last three years roughly 30x improvement on deepseeek on the most optimized deployment which is you know you can see on inference there's about a 30x improvement but you know over the last three years um we've had way more improvement intelligence per watt a lot of that coming from the model layer right if you look back three years it's GPD4 now it's like you know you know maybe like Quen one of the smaller Quen models that's like you know 27B parameters total and like 2 billion active is like way better. Um, and so you've got this huge improvement on model layer, you've got this pretty sizable improvement on hardware, but it's that co-design layer and I think that's that's what's important, right?
>> 嗯,好,我觉得一种看法是把它拆成这三个不同的层。嗯,从这个角度说,好,从 Hopper 到 Blackwell,这也是我们过去三年里就只有的这一代变化,在 DeepSeek 上大约是 30 倍的提升最优化的部署方式,你知道,在推理侧大概能看到 30 倍左右的提升,但你知道,过去这三年里,我们在每瓦智能上的提升要大得多,其中很大一部分来自模型层,对吧。你回看三年前那是 GPT-4,而现在可能就是 Qwen 里比较小的那种模型,比如总参数 27B、激活参数才 20 亿,效果却好得多。所以说,你在模型层拿到了巨大的提升,在硬件上也有相当可观的提升,但真正关键的是协同设计这一层,我觉得那才是最重要的,对吧?
便签笔记
25:05
If you look at the architecture of, you know, any of these models, but deepseek is the most famous one at least, uh, that's public and people have seen. >> Yeah, Deepseek got huge efficiency gains from like co-op optimization or kernel level optimizing memory. >> Yes, I I think it's it's it's like kernels of course, but it's actually you build the hardware architecture for the chip. So if you look at the shapes of all the experts in in DeepSeek, uh V3, they were all optimized for Hopper. And if you look at for V4, they're optimized for Blackwell and Huawei's chip. And what's interesting is despite the fact that TPUs are objectively an amazing chip, you know, and and they run all of deep mind and they do all the training uh for anthropic as well on the pre-training side at least. TPUs suck at running deepseeek, but they are really really great at running other kinds of models that don't run well on NVIDIA.
如果你去看这些模型的架构,随便哪一个都行,不过 DeepSeek 至少是最有名的,因为它是公开的,大家都见过。>> 是啊,DeepSeek 靠着协同优化、内核层面的显存优化,拿到了巨大的效率提升。>> 对,我觉得,当然有内核的因素,但实际上是你要围绕芯片去构建硬件架构。你看 DeepSeek V3 里所有专家的形状(shape),它们全都是针对 Hopper 优化过的。再看 V4,它们是针对 Blackwell 和华为的芯片优化的。有意思的是,尽管 TPU 客观上是一款非常出色的芯片,你知道,DeepMind 全部都跑在上面,Anthropic 的训练——至少是预训练那部分——也是在 TPU 上做的,但 TPU 跑DeepSeek 就很糟糕,可它跑那些在英伟达上表现不好的其他类型的模型,反而非常非常出色。
便签笔记
25:51
there is some level of such deep optimization that has been done um whether it be shapes uh network IO uh patterns you know how you do the collectives how you do um things around you know the the arithmetic intensity of the attention mechanism all these different things are co-optimized between the model and the and the and the hardware and the infrasoftware in between and it's it's hard to say you can disentangle the games >> do you think that like my understanding is that like China has done this a lot better than the west the last few years like in the deep sea was one of the first models to really like do this.
这里面存在着某种极深层次的优化,不管是形状、网络 IO 模式,还有你怎么做集合通信(collectives)、怎么处理注意力机制的算术强度(arithmetic intensity),所有这些不同的东西,都是在模型、硬件以及夹在中间的基础软件之间做协同优化的,所以很难说你能把这些收益拆解开来。 >> 你觉得,我的理解是,过去几年中国在这方面做得比西方好得多,比如 DeepSeek 就是最早真正这么做的模型之一。
便签笔记
26:27
>> I don't necessarily think so. think it's more so that the west doesn't tell people what they do right like open eye didn't tell people that you know GP40 was uh how sparse it was what the shape size was all these things but GP40 is roughly the same size slightly smaller than deepseeek v3 and 40 came out you know a little bit earlier right if I recall correctly >> so is your is your view that like all three of these things have been happening simultaneously at like roughly the same rate and the most the biggest gains are when you just co-optimize I would say I would say there's been more gains on the model layer than on that co-op than than on the sort of software infrastructure layer and the hardware layer. Um, but there's been innovations on every layer and and and really the biggest gain and the beauty of the best labs is when they co-optimize all three, you know, and and and that's what like you know when Enthropic is is, you know, even though they used many different kinds of hardware, they don't really inference
>> 我倒不一定这么认为。我觉得更多是因为西方不会把自己做的事告诉别人,对吧,比如 OpenAI就没有告诉大家 GPT-4o 有多稀疏、形状尺寸是多少,这些都没说。但 GPT-4o 的规模跟 DeepSeek V3 差不多,还略小一些,而且如果我没记错的话,4o 出来还稍微早一点。>> 那你的观点是不是说,这三件事其实是同时在以大致相同的速度推进,而最大的收益出现在你把它们协同优化的时候? 我会说,我会说模型层的收益比协同优化那一层、比软件基础设施层和硬件层都要多。不过每一层都有创新,而真正最大的收益、最漂亮的地方,在于最好的实验室能把这三层一起协同优化。你知道,就像 Anthropic,尽管他们用了很多种不同的硬件,但他们其实不怎么在 TPU 上做推理,主要是在 TPU 上训练。他们大量的推理是跑在
便签笔记
27:19
too much on TPUs. They mostly train on TPUs. um and and they inference a lot on cranium and GPUs and GPU is more a jack of all trades but they've optimized their hardware they're optimized their model they've optimized everything so they can do that whereas open AI they're you know prior models were optimized for hopper more now they're more optimized for blackwell and you know you you step forward through time these these um these labs and and and the same with Google right they've they've they've optimized you know Gemini 2 was really optimized for the TPU uh v uh v6e or tp Gemini 3 was and then Gemini uh you know the next Gemini that's coming out is really optimized for TPV7.
Trainium 和 GPU 上,GPU 更像是个万金油。但他们优化了硬件,优化了模型,什么都优化了,所以才能做到这一点。而 OpenAI 呢,他们之前的模型更多是针对Hopper 优化的,现在更多是针对 Blackwell 优化。随着时间往前推进,这些实验室都是这样,谷歌也一样,对吧,他们也做了优化——Gemini 2 是真正针对 TPU v6e 优化的,Gemini 3 也是,而接下来要发布的下一代 Gemini,则是针对 TPU v7 深度优化的。
便签笔记
27:56
Um and so sort of like a lot of these things are being co-optimized and actually when you pull that model and put it run it on the old hardware it's really not that great. Um and so I think a lot of this co-optimization is is the most important thing. It's called software hardware co-design and that's what's like really exciting about like you know sort of what what you know I think my day-to-day is like you know great you get to look at one layer there's all these innovations happening here there's all these innovations happening on every layer. The real breakthrough innovation is when you leaprog a few layers, you co-optimize and co-design them, and now all of a sudden you've you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x because you've optimized across all three layers. And so that's what's really exciting about sort of like what you see at the labs, which you see at like a company like Nvidia who's not co-optimizing on the
所以说很多东西都在被协同优化,而实际上,当你把那个模型拿出来放到老硬件上去跑,效果就真的没那么好了。所以我觉得这种协同优化是最重要的事情。它叫做软硬件协同设计,这也是特别让人兴奋的地方,就是我觉得我的日常大概就是这样——很好,你可以盯着某一层看,这里在发生各种创新,那里也在发生创新,每一层都有创新。真正突破性的创新,是当你跨越几个层级,把它们一起协同优化、协同设计的时候,突然之间,原本这里 2 倍、那里 2 倍、再那里 2 倍,本来相乘是 8 倍,结果实际上是 100 倍,因为你在三层之间都做了优化。所以这就是实验室里那些事情特别让人兴奋的地方,你在英伟达这样的公司也能看到——它本身并不算在模型层做协同优化,但也沾一点边,
便签笔记
28:41
model layer per per se, but a little bit from the model layer all the way downstream to, you know, silicon. Or you look at a company like TSMC, they're co-optimizing not just, you know, fabrication, but all the way from the components and the consumables and the tools all the way upstream to what the designs, their chips are, the customers are telling them is this co-optimization across many layers of the abstraction stack. >> There will always be bottlenecks somewhere in that optimization though that are like lagging behind and then need to get pulled forward, >> you know, >> and band-aids to toact.
从模型层一路往下游做到硅片。或者你看台积电这样的公司,他们协同优化的不只是制造工艺,而是从零部件、耗材、设备一路往上游,一直到设计、他们的芯片是什么样、客户告诉他们什么,这就是跨越抽象栈多个层级的协同优化。>> 不过在那种优化里,总会有某些地方成为瓶颈,会落在后面,然后需要被拉上来, >> 是啊, >> 还得打补丁凑合应付。
便签笔记
08内存墙与每平方毫米一瓦
29:11
If you had to predict like what are at any level of the stack, it can be literally anywhere. What are some of the bottlenecks you're most like you're kind of tracking most acutely the next year? And not necessarily in the supply chain, not in like scale, but in terms of the actual um and it can it can be in the supply chain too, but just like you know, is it memory improvements? Is it is it that like just like scaling? So memory memory is memory is an easy one that everyone's talked about, but I'm not going to talk about from a supply chain angle. I'm talking about from a technology angle, right? Memory um capacity and bandwidth have been improving very slowly. The NAN cell was invented like 25 years ago. The DM cell was invented like 40 years ago and there's been no major breakthrough in in cell like you know how what a NAND cell is. Obviously NAND is like a very simple gate or DAM cell. There there is stuff that could come down the pipeline that could be hugely innovative. But even over the last, you know, five
如果让你预测一下,在这个栈的任何一层,真的可以是任何地方——你最密切关注的、明年最关键的瓶颈会是哪些?不一定是供应链方面的,也不是规模方面的,而是实际的——当然也可以是供应链的,但我是说,是内存的改进吗?还是说就是单纯的扩规模? 内存的话,内存是个大家都在谈的、很容易想到的例子,但我不打算从供应链的角度讲,我要从技术角度来说。内存的容量和带宽一直提升得非常慢。NAND单元大概是 25 年前发明的,DRAM 单元大概是 40 年前发明的,在单元本身上一直没有重大突破,就是 NAND 单元到底是什么这件事。显然 NAND 就是个非常简单的门,DRAM 单元也是。确实有一些可能在未来出现的东西会带来巨大的革新。但即便是过去这五年,
便签笔记
30:06
years, all we've really done is make the HBM, you know, more stacks, faster, but actually there's like new innovations coming in the next few years where instead of, you know, stacking the HPM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode. Um, and so there's interesting companies in that space and interesting PC's that companies are trying to do there. I think like memory bandwidth is one of the biggest. Another one is um for the history of like silicon basically for the last two decades at least you know how many watts a chip is can be easily predicted just by looking at it for for a data center or desktop chip it it peaks up at one watt per millimeter squared and so if a chip is 100 millime squared generally the power consumption is around 100 or a little bit less um and if you look at the newest Nvidia silicon the newest TPU silicon it's still on that range of one watt per millimeter squared so you know chips are now getting to you know, 1400 watts.
我们真正做的也只是把 HBM 堆得更多层、跑得更快而已。不过接下来几年确实会有一些新的创新,不再是把 HBM 和芯片分开堆叠,而是直接把它堆在内存直接放在芯片上,这会让带宽暴涨。嗯,所以这个领域里有一些很有意思的公司,也有一些有意思的公司在尝试做这方面的事情。我觉得内存带宽是最大的瓶颈之一。另外一个是,嗯,纵观硅芯片的发展史,基本上至少过去二十年,你知道一颗芯片是多少瓦特,光看它就能轻松预测出来——对于数据中心或者桌面级芯片来说,它的峰值大概是每平方毫米一瓦,所以如果一颗芯片是100平方毫米,一般功耗就在100瓦左右或者稍微低一点,嗯,而且你去看英伟达最新的芯片、最新的TPU芯片,它仍然在每平方毫米一瓦这个范围内。所以你知道,芯片现在已经做到了1400瓦。
便签笔记
30:58
Next generation is 2,000 watts for Nvidia. Um, with Reuben and such. Uh, and and you move forward to Reuben Ultra, it's going to be like 4,000 watts or something like that. But really, there's increasing the amount of silicon. What's exciting is we're now finally doing things and and it's in development right now where you actually can pump the amount of power into the silicon uh to be way more than one watt per millimeter squared. And now that all of a sudden means you need less silicon. Obviously, it's running at higher power.
英伟达下一代是2000瓦。嗯,就是Rubin那些。呃,再往后到Rubin Ultra,大概会是4000瓦之类的。但实际上,这是在增加硅的用量。让人兴奋的是,我们现在终于在做一些事情——而且现在正在研发中——你其实可以把远超每平方毫米一瓦的功率灌进硅里。而这一下子就意味着你需要的硅更少了。显然,它是在更高的功率下运行。
便签笔记
31:25
It's less efficient in some cases, but you reduce the amount of silicon and you're able to like >> like over thermal issues, >> thermal issues. Um there's uh interference of like electrical interference issues. There's all sorts of different issues uh that crop up and that's why it's a hard engineering problem. That's why we've stuck at about one. But what's exciting is the world is trying to change these things. I think interesting like in a different part of the supply chain it's sort of like you know people people will talk about like energy is hard and you know we have energy bottlenecks and it's like yeah but there's actually like very simple solutions you know one could think of right um take the millions of diesel engines for trucks that the US has the capacity to make um you can very trivially convert them to be using for gas uh in the assembly line and then stick them up to a electrical motor like back driving it so the electrical motor generates electricity rather than the electrical motor causing the the
在某些情况下效率更低,但你减少了硅的用量,而且你能够——>> 比如克服散热问题,>> 散热问题。嗯,还有呃电气干扰之类的干扰问题。会冒出各种各样的问题,呃,所以这是个很难的工程问题。这也是为什么我们一直卡在一瓦左右。但让人兴奋的是,整个行业正在试图改变这些东西。我觉得有意思的是,在供应链的另一个环节,有点像,你知道,人们会说能源很难搞,说我们有能源瓶颈,然后就像是,是啊,但其实是有非常简单的解决方案的,你知道,你可以想到的,嗯,比如说,把美国有能力生产的数百万台卡车柴油发动机拿过来,嗯,你可以非常轻松地在装配线上把它们改成烧天然气,呃,然后把它们接到一台电动机上,反向驱动它,这样电动机就发电,而不是电动机去带动轮子转动,而是反过来做。这样
便签笔记
09英伟达对 TPU:护城河还剩什么
32:16
rotation of the wheel, for example, but doing it the opposite direction. And now you've generated electricity by pumping gas into something that us can make millions of. Um, and then, okay, well, that sounds like a pain in the ass to uh service, right? Because now you have to have hundreds of these on a data center site. Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines. Actually, it's actually pretty trivial to not I don't want to say it's trivial, I couldn't do it. Um >> I think you're making a really good point which is that like because the west wasn't really thinking about semic even hardware more broadly the last 20 30 years we didn't have like much innovation we'd have the best minds like thinking about how do you improve these >> why why would you why would you want to go work in hardware when you can uh make ads to ads >> yeah exactly >> um okay I'm dying to ask Nvidia versus TPU what are your thoughts >> um I think I think like everyone wants
你就通过往一个美国能造几百万台的东西里加气来发电了。嗯,然后,好吧,那听起来维护起来很麻烦,对吧?因为现在你在一个数据中心场地上要有几百台这种东西。嗯,其实呢,你可以直接从汽车修理店里把人挖出来,让他们到��跑着去修卡车发动机。其实这真的挺简单的——我不想说它很简单,反正我是干不了。嗯 >> 我觉得你说到了一个非常好的点,就是因为西方在过去二三十年里其实并没有认真思考半导体、甚至更广义的硬件问题,我们没有太多创新,也没有让最顶尖的头脑去思考怎么改进这些东西。>> 为什么呢?你为什么要去做硬件,明明你可以去做广告呢?>> 是啊,没错。>> 嗯,好,我特别想问一下英伟达对比TPU,你怎么看?>> 嗯,我觉得,我觉得大家都想在这两者里选一个站队,但这其实取决于
便签笔记
33:09
to pick one or the other for this, but it's really like a function of like look, you know, you look two years from now, Google's going to make 10 plus million TPUs and through their supply chain and Nvidia is going to make, you know, many more million tens of millions of GPUs and both are going to be 100 plus billion dollar, you know, well, Google's going to be 100 plus billion dollars, you know, of TPU created a year and and Nvidia will be, you know, 500 plus or, you know, whatever. I'm not making a specific estimate.
你看,两年之后,谷歌会造出一千多万块 TPU,通过他们的供应链,而英伟达会造出,你知道,多得多的、成千上万万块 GPU,两家都会是 1000多亿美元的,嗯,好吧,谷歌会是每年 1000 多亿美元的 TPU 产值,而英伟达会是,你知道,5000 亿以上,或者,你知道,随便吧。我并不是在给一个具体的估算。
便签笔记
33:34
>> This is not revenue forecast. This is just a thought experiment. >> Yeah. Or research. >> You've been media trans. >> Absolutely. you know, getting ready for the SpaceX idea. >> Um, are you guys big in SpaceX? Okay, so that makes sense. Um, >> we're very lucky to be very large investors. >> Awesome. Awesome. Um, so I would say um the the case of sort of like Google TPUs versus uh Nvidia GPUs, they both have like points that are really like in their favor, right? You know, Nvidia will be like, "Oh, well, we have switches and we're general purpose." And and TPUs will be like, "Well, we're more optimized. actually more energy efficient and our network is actually more um optimized for certain types of network architectures. And so you have like these counterpoints that both would really uh get into and you know I could with a straight face argue with you like that GPUs are way better than TPUs or TPUs are way better than GPUs but it comes down to hardware software codeesign. So actually the way OpenAI's
>> 这不是营收预测,这只是一个思想实验。>> 是的,或者说是研究。>> 你已经被媒体训练过了。>> 绝对的。你知道,为 SpaceX 那个想法做准备。>> 嗯,你们在 SpaceX 上投得很重吗?好,那就说得通了。嗯,>> 我们很幸运能成为非常大的投资人。>> 太棒了,太棒了。嗯,所以我想说,嗯,谷歌 TPU 对阵英伟达 GPU 这件事,两边都有各自非常有利的点,对吧?你知道,英伟达会说:"哦,我们有交换机,而且我们是通用的。"而 TPU 那边会说:"我们更优化,实际上能效更高,我们的网络其实针对某些类型的网络架构做了更多优化。"所以你会有这些针锋相对的论点,双方都真的会去深挖,你知道,我可以一本正经地跟你论证说 GPU 比 TPU 好得多,或者 TPU 比 GPU 好得多,但归根到底这取决于软硬件协同设计。所以实际上,按照 OpenAI 模型的发展方向,用 TPU 对他们来说可能会是个糟糕的决定。
便签笔记
34:24
models are headed, it would be a terrible decision for them to use TPUs potentially. And the way that Enthropic and Google's uh models are headed, it's actually a terrible decision potentially for them to train with GPUs. I mean, it'd be fun to what's the what's the fundamental difference there? >> There's various things, right? Like the size of the matrix multiply unit is different as a as a very simple thing. And therefore, the shape of the matrix multiply you do, the attention mechanism you use, uh the way that attention mechanism is structured, the way the experts are structured. So you think so open AI and anthropic are converging the very different model architectures. I think they're I think they have quite different model architectures. In fact, um, you know, open eyes are much more sparse, um, and that has benefits. And then anthropics are, you know, they're still sparse, but more dense in general, and and that has different benefits. And there's many other things, right? The network topology, right? Nvidia, all of
而按照 Anthropic 和谷歌模型的发展方向,用 GPU 来训练对他们来说其实也可能是个糟糕的决定。我是说,会挺有意思的——那里根本的差别是什么?>> 有很多方面,对吧?比如矩阵乘法单元的大小就不一样,这是个很简单的例子。因此,你做的矩阵乘法的形状、你用的注意力机制、这个注意力机制的结构方式、专家(experts)的组织方式都不同。所以你觉得,OpenAI 和 Anthropic 正在收敛到非常不同的模型架构上。我认为他们的模型架构相当不同。事实上,嗯,你知道,OpenAI 的稀疏得多,嗯,这有它的好处。而 Anthropic 的,你知道,他们仍然是稀疏的,但总体上更稠密一些,这又有不同的好处。还有很多其他方面,对吧?网络拓扑,对吧?英伟达,他们所有的芯片都连到交换机上,
便签笔记
35:10
their chips are connected to switches, NVLink switches. For Google, they have no switch. Um, but what they've done is they've been able to, you know, Nvidia, the NVLink can only connect 72 GPUs. for Google, their ICI can connect 8,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch. And so there's like there's trade-offs there. There's positives and negatives and that influences the model architecture. It's not necessarily that you should uh you know claim one is better than the other because at the end of the day, how do you say that this is better than that when you can't measure them in isolation because it also extends up to the model layer, right? Um >> but I remember for a long time thinking you know one the programmability of Nvidia and just CUDA as as such a big moat. It seems to me that narrative has kind of changed at least in my mind for the last three six months like model companies no longer care about if we have to write custom kernels for you
NVLink 交换机。而谷歌,他们没有交换机。嗯,但他们做到的是,你知道,英伟达的NVLink 只能连接 72 块 GPU。而谷歌,他们的 ICI 能以超高带宽连接 8000 块芯片,但你必须经过其他芯片才能到达目标,因为没有交换机。所以这里面是有权衡的。有好处也有坏处,而这会影响模型架构。并不一定说你就应该,你知道,宣称某一个比另一个好,因为归根到底,你怎么能说这个比那个好呢当你无法孤立地衡量它们,因为它还会向上延伸到模型层,对吧?嗯 >> 但我记得很长一段时间里我都觉得你知道,Nvidia 的可编程性、CUDA 本身就是一条很深的护城河。但在我看来,至少在过去三到六个月里,这套叙事似乎有点变了,比如模型公司已经不太在意——如果我们得为另一款芯片手写自定义 kernel,那就写呗。必要的话我们可以同时适配四五种芯片。
便签笔记
36:04
know this other chip so be it. We'll work with four or five chips if we have to. Um Claude and Codeex are actually quite good at doing a lot of that optimization work. And so it seems like some of the and then and then it's you know it's not like there's 10,000 model companies that are each you know each need programmability. There's on the order of tens maybe model companies and so it seems to me that like if you the fundamental premise of like tens of thousands of big customers that need CUDA compatibility like it seems that kind of thesis is is changing in the last >> Yeah. I mean I mean certainly the CUDA mode and software remote is at least partially uh disentangled because you know models are just great at coding and all software gets commoditized in that case. I do think there is some level of like open source and you know what people call the CUDA mode is not actually anything to do with CUDA but it's like the fact that DeepSeek Kimmy and and Zippui and and Alibaba and Tens all these all these companies Xiaomi had
嗯,Claude 和 Codex 其实相当擅长做大量这类优化工作。所以看起来有一些……而且,你知道,并不是说有一万家模型公司,每一家都需要可编程性。模型公司大概也就几十家这个量级,所以在我看来,如果你那个根本前提——有成千上万家大客户需要 CUDA 兼容性——这个论点似乎正在改变,在过去…… >> 是啊。我是说,我是说,CUDA 这条护城河、软件护城河,至少已经被部分瓦解了,因为你知道,模型本来就特别擅长写代码,那样一来所有软件都会被商品化。不过我确实觉得,有某种程度的开源……而且你知道,人们所说的 CUDA 护城河其实跟 CUDA 本身没什么关系,而是说 DeepSeek、Kimi、智谱、阿里巴巴、腾讯,所有这些公司——小米最近也出了一个很棒的模型——他们的模型
便签笔记
37:02
an awesome model recently their models are co-designed for GPUs and therefore if I want to run them on TPUs actually in some cases they don't run really well on TPUs now Google just has to create their own open source model ecosystem or open source models themselves so they have the Gemma models and and so you end up with like well that's not really CUDA as a moat it's that the downstream product is more optimized for Nvidia and in these cases these companies are just open sourcing them or like Neotron is just open sourcing it and then the users of it for example the open you know the inference uh API providers the RL companies that are trying to take open models and customize them for company's business use cases all these different companies are downstream of the fact that like okay well I guess I need to use Nvidia because the ecosystem uses Nvidia even though I don't partic particularly care about writing CUDA kernels because the models are great at that, but it's like the shape of like
都是针对 GPU 协同设计的,所以如果我想在 TPU 上跑它们,实际上在某些情况下它们在 TPU 上跑得并不好。那现在 Google 就只能自己去打造一套开源模型生态,或者自己做开源模型,所以他们有了 Gemma 系列模型。于是最后你会发现,这其实并不是 CUDA 作为护城河,而是下游的产品对英伟达做了更多优化。而在这些情况下,这些公司只是把它们开源出来,比如 Neotron 就是直接把它开源了,然后它的使用者,比如那些做推理 API 的服务商、那些想把开源模型拿来针对企业业务场景做定制的强化学习公司,所有这些不同的公司都处在下游,结果就是:好吧,我看来只能用英伟达,因为整个生态都在用英伟达。哪怕我并不特别在意写 CUDA kernel,因为模型在这方面已经很擅长了,但问题在于那种结构,比如
便签笔记
37:51
well this expert the the demod is this and you know the hidden dimension blah blah blah is this right and and so therefore it's better to run on Nvidia GPUs than it is on TPUs and vice versa right if Google were to actually open source really good models you know this would be the same thing right people would take their models and they'd be like oh wow these don't run that well on Nvidia GPUs um I should actually just rent TPUs or buy TPUs and do it on there for small teams you're going to want to use all the open source software like VLM MSG laying um pietorch all that stuff but the big labs they don't necessarily need to use all that right open I forked PyTorch long ago and you know anthropic and all these other people don't necessarily rely heavily on the open- source implementation of you know these things they forked things or built it on their own already and so they don't need to rely on the open source and therefore now it's more like you know I'll choose the best hardware and I'll co-design my model and
这个专家的维度是多少、隐藏层维度是多少之类的,对吧?所以跑在英伟达 GPU 上就比跑在 TPU 上更划算,反过来也一样。对吧,如果谷歌真的开源一些非常好的模型,情况也会是一样的:大家会拿着他们的模型,然后发现,哇,这些在英伟达 GPU 上跑得不太行,我其实应该去租 TPU 或者买 TPU,在上面跑。对小团队来说,你会想用所有那些开源软件,比如 vLLM、SGLang、PyTorch 这些东西。但那些大实验室,他们不一定需要用这些,对吧?OpenAI 很久以前就 fork 了 PyTorch,而且你知道,Anthropic 还有其他这些公司,也不一定重度依赖这些东西的开源实现。他们要么 fork 了,要么自己已经从头搭了一套,所以他们不需要依赖开源,因此现在更像是:我会挑最好的硬件,然后针对那个硬件把我的模型和基础设施软件从头到尾协同设计一遍,
便签笔记
10Cerebras 与快模式的经济账
38:37
infrastructure software through and through for that hardware uh that is the best and most costefficient >> and you know I'll have AI help me write all that software. >> What do you think of Cerebrus? >> I think Cerebrus is a really innovative company. Um I I think in in some spots of the market they're really really good. Um very fast inference. I think that's a big market. Uh we use fast mode almost exclusively at semi analysis. Um >> by the way I love how disciplined you've been about accounting for I don't know if that was one exhibit you did or if you do it consistently but accounting for the dollar spent and the ROI on each task.
选出最好、最具成本效益的那个。>> 而且你知道,我会让 AI 帮我写所有那些软件。》你怎么看 Cerebras?》我觉得 Cerebras 是一家非常有创新力的公司。呃,我觉得在市场的某些细分领域,他们确实做得非常非常出色。呃,推理速度非常快。我认为这是个很大的市场。呃,我们在 SemiAnalysis 几乎只用快速模式。呃,》顺便说一句,我很喜欢你在这方面的严谨——我不知道这只是你做的某一张图表,还是你一直都这么做——就是把花掉的每一块钱和每项任务的投资回报率都算清楚。
便签笔记
39:12
>> Awesome analysis. >> Yeah. Yeah. We we we uh we do it pretty diligently and so thank you. That was the dark GDP article that we wrote. Um and so and and also like track everyone's token spend by day and if someone's like spiked up I'm like what did you do? It's like okay thank you for telling me that that seems worth it. Cool. On with my day. I think fast mode is obviously worth a lot for high-end tasks, right? I could just see so many different use cases where you know super fast tokens are worth it. I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and and therefore uh the market won't pay for them and they'll use GPUs and TPUs instead. I think the big risk for Cerebrus is I mostly think the best models are the ones that you want to use fast mode on and small models you necessarily might not use fast mode on.
》很棒的分析。》是的,是的。我们我们我们呃我们做得挺认真的,所以谢谢你。那就是我们写的那篇「暗 GDP」文章。呃,还有就是,我还会按天追踪每个人的 token 消耗,如果谁突然飙上去了,我就会问:你干了啥?然后对方说清楚,我就想,行,谢谢你告诉我,这看起来是值得的。不错。继续过我的日子。我觉得 fast mode 对高端任务显然很有价值,对吧?我能想到非常多不同的使用场景,超快的 token 输出是有价值的。但我也能看到反面:有很多场景其实不需要超快的 token,因此市场不会为此付费,他们会转而用 GPU和 TPU。我觉得 Cerebras 最大的风险在于,我基本上认为,你想开 fast mode 的是那些最好的模型,而小模型你未必会想开 fast mode。
便签笔记
39:55
I could see that being wrong with you know financial markets maybe or something like that like a Jane Street high frequency trading or something like that um or medium frequency trading. Um but ultimately you know running really large models at really long context is very difficult on SRAMM based chips like Cerebras like Grock and so now it all of a sudden is like you know what happens then if like the models get too big right if open's model is not you know on the order of uh you know hundreds of billions parameters or you know low trillion parameters but it's actually 10 plus trillion parameters now all of a sudden I don't think that that will fit on cerebrus right and then if that doesn't with a long context length right if you have a million context length now that makes it really difficult to justify you know, and and as all so far we've seen the bulk of revenue and usage at the labs be on their best model. Even when the model price has gone up, we've seen that. Um there's some data that
我觉得这一点在金融市场之类的场景可能就不成立了,比如像 Jane Street 那样的高频交易之类的,或者中频交易。嗯,但归根结底,在基于 SRAM 的芯片上跑超大模型、超长上下文,是非常困难的,比如 Cerebras、比如 Groq。所以现在突然就变成,你知道,会发生什么呢,如果模型变得太大了呢?如果 OpenAI 的模型不是几千亿参数量级,或者一两万亿参数量级,而实际上是十万亿以上参数,那这下我觉得那就装不进 Cerebras 了对吧。然后如果再加上超长上下文,如果你有一百万的上下文长度,那就真的很难说得通了,你知道。而且到目前为止,我们看到各家实验室的大部分收入和使用量都集中在它们最好的模型上。哪怕模型价格上涨了,我们也还是看到这种情况。嗯,有些数据显示,尽管 Fable 今天才刚发布,就已经有非常多的人切换到了 Fable 和
便签笔记
40:44
shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos, sort of that next tier model, even though it's way more expensive. And so um >> is that and that's volume by dollars totally. But was that volume by tokens? >> Well, I guess who cares about volume by tokens? It's about the dollars. >> Fair enough. >> Right. If I don't care that there's, you know, uh, you know, I don't know, 200,000 Mini Coopers or Toyota Camry sold if if, uh, you know, I don't know, Ford50s are 5x ASP and they sell only half as much.
Mythos,就是那个更高一档的模型,尽管它贵得多。所以嗯——>> 那是按美元算的用量吗?完全是。但那是按 token 算的用量吗?>> 呃,我觉得谁在乎按 token 算的用量呢?关键是美元。>> 有道理。>> 对吧。如果我不在乎,你知道,呃,我也不知道,卖出了 20 万辆 Mini Cooper 还是丰田凯美瑞,如果,呃,我也不知道,福特50系的平均售价是5倍,但销量只有一半。
便签笔记
41:12
>> Okay. Right. >> And then and and therefore the most lucrative market is pickup trucks in America. Right. Mostly being facicious, but like >> I do think this is one of the things that you've done so well and differentiates you from almost everyone else is that you you care so much about the economics in addition to the technology. And I think very few people bridged those two thing things well. And so >> I think I think it's really fun inside of semi analysis because we have 90 people and like a big chunk of them are technologist engineers across the whole supply chain. Um and then a big chunk is people who are formerly at hedge funds and you see these arguments like people are like oh well that doesn't matter and it's like then someone's like well but cost and then someone the engineers like no no but this technology is the coolest. You see this you see this organically like fight it out. Um and and were pretty informal and you know given the fact that I was a for moderator is you can imagine what the
>> 好的,对。>> 所以说,美国最赚钱的市场就是皮卡。对吧。当然我大部分是在开玩笑,但是 >> 我确实觉得这是你做得特别好、并且几乎能把你和所有其他人区分开的地方,就是你除了关注技术之外,也非常在意经济账。我觉得很少有人能把这两件事很好地打通。所以说 >> 我觉得在 SemiAnalysis 内部真的挺有意思的,因为我们有90个人,其中很大一部分是覆盖整条供应链的技术专家和工程师。嗯,然后还有很大一部分是以前在对冲基金做过的人,你就会看到这种争论,有人说"哎这个不重要",然后又有人说"可是成本呢",然后工程师就说"不不不,这个技术才是最酷的"。你会看到这种争论自然而然地打起来。嗯,而且我们内部氛围挺随意的,你知道,考虑到我以前是个论坛版主,你可以想象那个场面有多热闹 >> 别跟猪摔跤,因为猪很享受这个过程。
便签笔记
42:01
the enjoying it >> you don't wrestle with a pig because a pig enjoys it. >> Exactly. Just on this topic before going to the next question. Are there like trigger topics in semis for you? You know like if someone's like which is like such a meme you think this person must be a like if you know if it's like oh you like memory is the bottleneck. I mean it's true but like um I think I think moreover the one that really gets me is people are like AI has no ROI >> infuriates me right like there's like what's the ROI or like denying model progress right there's these people that are like models aren't getting better they're not reasoning they can't think they're going to deadend and plateau and it's like bro the line has been up and to the right in terms of capabilities this entire time and they're like look this benchmark didn't improve that's cuz it said 90% look at the new benchmark you saturated now they're skyrocketing, right? It's like I think that's more so the issue and challenge. Like I think semis are
>> 完全正确。在进入下一个问题之前,就这个话题——在半导体领域有没有什么话题是你的雷点?你懂的,就是那种一说出来就特别 meme、你就觉得这人肯定是……比如说"哦你觉得内存是瓶颈"。我是说这确实没错,但是——嗯,我觉得更让我上头的是那些人说"AI没有 ROI" >> 这让我特别火大,对吧,就是那种"ROI 在哪",或者否认模型的进步,对吧,有些人就说模型没有变得更好、它们不会推理、不会思考、马上就要走进死胡同、要停滞了,然后我就想,兄弟,能力这条曲线一直都是往右上方走的啊,然后他们说"你看这个基准测试没提升",那是因为已经到90%了,你看新的基准测试,你以前刷爆了,现在这些新的分数正在直线上升,对吧?我觉得这才更是问题和挑战所在。我觉得半导体真的非常复杂,我不会怪别人
便签笔记
43:02
really complex and I don't fault people for um lacking like understanding of it. Like I learn stuff every day about the semiconductor supply chain from people and I've been studying it for you know arguably 18 years since I started moderating the forums when I was 12 right like you know arguably been studying it for that long but even then like and it's like live breathed and that's all I care about but there's so many layers of the abstraction stack it's like like I learned about a new chemical that does like a hundred million dollars of sales like yesterday and I'm like whoa didn't know this one existed and what process it did and it's like but it's like you know you learn about things all the It's like okay hundred billion dollar sales in a you know couple hundred billion dollar industry is whatever but like you know it's like >> but it's essential >> it's essential and it's like actually every chip requires it. It's like wow I guess there are a thousand process steps and you know it's like oh yeah you like
嗯,对它理解不够。比如我每天都还在从别人那里学到关于半导体供应链的新东西,而我研究这个已经,你知道,可以说有18年了,从我12岁开始当论坛版主算起,对吧,可以说研究了这么久,但即便如此,而且我是真的活在里面、天天呼吸这些东西,我在意的就只有这个,但这个抽象堆栈的层数实在太多了,比如说,我昨天才知道有一种新的化学品,一年销售额大概一亿美元,我当时就想,哇,我都不知道这东西存在,也不知道它是用在哪道工序的,就是这样,你知道,你总是在不断学到新东西。就像,好吧,在一个几千亿美元规模的行业里,一亿美元的销售额好像不算什么,但你知道,就是说 >> 但它是必需的 >> 它是必需的,而且实际上每颗芯片都需要它。就让人觉得,哇,我猜大概有一千道工序吧,你知道,就像别人说"哦你懂半导体啊,那你把每道工序都说一遍"。
便签笔记
11十年视角与局部最小值陷阱
43:52
semiconductors name every process step. It's like no come on. What what I think is the most funny is when people have all the facts in front of them and then they get the conclusion completely wrong. Um and that's >> that happens in our job all the time too. >> Yeah. >> Yeah. I mean, I can't I I get I I think my attitude is not to be mad that you do that. It's to do it as fast as possible. >> I think the industry because it's so it's just like AI is the most important thing in the world right now and there's so many near-term bottlenecks. We talk a lot about the near-term. Are there longer term things that you're really excited about? Like say on a 10-year time frame? We talked about orbital data centers, but like like siliconics, you think they're underrated or overrated on a 10-year time frame? Are there other things that on a 10-year time frame?
那我只能说,得了吧。我觉得最好笑的是,有些人明明手里握着所有事实,却把结论搞得完全相反。嗯,这个 >> 这种事在我们这行也天天发生。>> 是的。>> 是啊。我是说,我的态度不是因为别人这样而生气,而是要尽可能快地把事做对。>> 我觉得这个行业,因为它实在太……就是说 AI 现在是世界上最重要的事情,而且有这么多近期的瓶颈。我们聊了很多近期的事。有没有一些更长期的东西是你特别兴奋的?比如说10年的时间尺度?我们聊过轨道数据中心,但比如硅光之类的,你觉得在10年尺度上它们是被低估还是高估了?还有没有别的什么放在10年尺度上值得关注?
便签笔记
44:32
Yeah, I mean I think on SP I think space is like super crazy awesome in the 10-year time frame that I'm you know for space data centers and all these sort of mining asteroids and all these things which is you know super excited about the vision of SpaceX right um again not investment advice before you hop in um I think I think on the semiconductor side tremendous market movements and tremendous like things can happen just when like things happen one year later or sooner and so that's all like technology that like you know in terms of like co-ackage optics like well like everyone knows it's going to happen by the end of the decade the the debate is like 27 7 28 29 2030 but some point along there it's going to happen. I think the more interesting thing is like there's companies like um I did you guys invest in Navian Ral's company?
是的,我是说,在太空这块,我觉得10年尺度上太空真的超级酷,你知道,太空数据中心、小行星采矿这些东西,我对 SpaceX 的愿景真的非常兴奋,对吧,嗯,再说一次,在你冲进去之前,这不构成投资建议。嗯,我觉得在半导体这边,会有巨大的市场波动,巨大的……有些事情早一年或晚一年发生,结果就完全不同,所以那些都是技术层面的东西,比如说共封装光学,大家都知道它在这个十年结束前一定会发生,争论的只是27、28、29还是2030年,但在这段时间里的某个点它一定会发生。我觉得更有意思的是,有些公司,比如说,嗯,你们有投 Naveen Rao 的公司吗?
便签笔记
45:12
>> We did. >> Okay. Yeah. So I think like he's trying to innovate on like the silicon layer on the software abstraction layer and the model layer simultaneously and he fully understands that it's not a like a you know we're going to do this in a few years. >> It's not a two-year time frame. >> Yeah. It's not a few year time frame. It's a long-term bet. Um, and like stuff like that is like, okay, we're going to bring like potentially like analog compute with energy based models and like all this crazy all at once.
>> 我们投了。>> 好的。是的。所以我觉得,他是在同时在硅片层、软件抽象层和模型层上做创新,而且他非常清楚这不是那种"我们几年内就能搞定"的事。>> 这不是两年的时间尺度。>> 对。这不是几年就能成的时间尺度。这是一个长期押注。嗯,这类事情就是,好吧,我们要把模拟计算和基于能量的模型,还有这一堆疯狂的东西一次性全都做出来。
便签笔记
45:35
It's like that's exciting. Probably won't work, but you know, that's exciting and I I like really look forward to >> definitely won't work quickly. >> Yeah, definitely won't work quickly is what I should say. I believe in Deaveen and like, you know, I I I met him very, you know, I think he's one of the first people I met in the industry um, funnily enough, like in 2020 or 2021. Um, actually 2020. Yeah. It says something about him. I think he's someone in my experience. He's always trying to >> I baited him on the internet. I baited him on the internet. That's >> He's always trying to help the younger generation. He's trying to identify talent. And >> he was also so ahead of his time with Mosaic. I remember getting pitched.
这就很让人兴奋。大概率不会成功,但你知道,那很令人兴奋,我真的很期待>> 肯定不会很快成功。>> 对,我应该说的是,肯定不会很快成功。我相信 Naveen,而且,你知道,我认识他很,你知道,我觉得他是我在这个行业里最早认识的人之一,嗯,说来有趣,大概是2020或2021年。嗯,其实是2020年。是的。这也说明了他的为人。在我的经历里,他是那种一直在努力 >> 我在网上钓他。我在网上钓了他一把。就是 >> 他一直在努力帮助年轻一代。他一直在努力发掘人才。而且>> 他做 Mosaic 的时候也太超前于时代了。我记得当时被 pitch 过。
便签笔记
46:10
>> No, it was 2019. I was still I was still anonymous then actually. I I baited him on the internet and he started replying and then I just took it to DMs and then took it to a call and like that was the first person who's like really important that I talked to in the entire semiconductor industry. funny. >> Um, but yeah, sorry to interrupt. >> That's funny. What do you think is the end state of the ecosystem? Like do you think every lab, every hyperscaler just has its own chips? Like train seems like it's now working, right? So do you think we end up with every lab, every hyperscaler has it own chips at least for inference and then maybe for training you go to Nvidia or whoever or what do you think is the end state?
>> 不对,那是2019年。我那时候还是匿名的。我在网上钓了他一下,他就开始回复,然后我就把话题转到私信,再转到通话,那是第一个对我来说真正重要的人……是我在整个半导体行业里聊过的所有人里的。挺有意思的。>> 嗯,不过,抱歉打断你了。>> 这挺有意思的。你觉得这个生态最终会是什么样子?比如你觉得每个实验室、每个超大规模云厂商都会有自己的芯片吗?训练芯片现在看起来也跑通了,对吧?所以你觉得最后会不会变成每个实验室、每个超大规模厂商至少在推理上都有自己的芯片,然后训练可能还是去找英伟达或者别家?你觉得终局是什么样?
便签笔记
46:44
>> I think everyone will try and stop trying. I think ultimately um you know supply chains matter. what technology you can bring in matters and more and more as the industry gets bigger supply chain diversification happens. Um you know right now everyone's chip more or less looks the same. It's a big logic compute die in the center and there's some HBM on the right and left and on the top and bottom top side is networking and then the bottom side is PCIe and other IO. Um and that is the exact same structure for tranium TPU Nvidia chips. Um and most of the startups um not Grock and Fris are doing weird but that's cool you know um I think like as you step forward we're going to get more bifurcation of hardware architecture and model architecture and therefore people are going to co-optimize them and you know some of them will end up in local minimas right you know as we're you know if this is like gradation gradient descent like people are like trying to go to the most optimized solution some
>> 我觉得所有人都会去试,然后又都会放弃。我觉得说到底,供应链是很重要的,你能拿到什么样的技术也很重要,而且随着行业越来越大,供应链多元化会越来越明显。嗯,你知道,现在大家的芯片长得差不多都一样。中间是一颗很大的逻辑计算裸片,左右两边是一些 HBM,上下两侧——上面是网络,下面是 PCIe 和其他 IO。嗯,这就是Trainium、TPU、英伟达芯片完全一样的结构。嗯,大多数创业公司也是——Groq 和 Furiosa 除外,他们搞得比较奇怪,不过那也挺酷的。我觉得,往前走,我们会看到硬件架构和模型架构出现更多分化,因此大家就会去做软硬件协同优化,然后你知道,其中一些人最后会陷进局部最小值,对吧,就像我们说的如果这就像梯度下降一样,大家都想走到最优解,那有些
便签笔记
47:40
people will race to a local minima and then the question is like how do you leap how do you scoot back over to like the absolute minima and some to some extent like a general more Nvidia will always be more general purpose than anyone else's chip in general um at least on a parallel AI compute basis because they have so many customers who care about different things who will always give them feedback in the design you know the minima will always be better than them but is that minima a local minima like is is the TPU or tranium or grock or cerebras or whoever's design optimized awesomely for here but in the end state actually you got to go over here and so they're the wrong >> um and Maybe they make a great time, they're great for a little bit of time, but then they end up being wrong. It's like that's the real question. Um, and so I think I think there will be a big market for general purpose AI compute.
人就会一路冲进一个局部最小值,然后问题就变成了:你怎么跳出去,怎么挪回到全局最小值那边去。某种程度上说,英伟达在通用性上总是会比任何其他人的芯片更强,嗯,至少在并行 AI 计算这个层面上,因为他们有太多客户,这些客户关心的点各不相同,会持续给他们的设计提供反馈。你知道,那个最小值本身总会比他们做得更好,但问题是那个最小值是不是局部最小值?就是说 TPU、Trainium、Groq、Cerebras 或者随便谁的设计,在当下这个点上优化得非常棒,但终局其实你得走到另外一个地方去,那他们就是错的。>> 嗯,也许他们能风光一阵,有那么一小段时间他们非常棒,但最后被证明是错的。这才是真正的问题所在。嗯,所以我觉得通用 AI 算力会有一个非常大的市场。
便签笔记
48:27
Um, because you talk to people at labs, they don't even know what architecture they're going to be doing in a year. Like, right, like they literally don't know what architecture they're going to be doing in a year. They have bets. They have many research bets and and that's this exciting thing, but they don't know where where it's going. generally they like know what hardware they have and they're trying to co-optimize but ultimately like if a new breakthrough happens on model architecture it's like just replace the tension mechanism with something else right who knows or you know all of a sudden you know something happens the best hardware will change and therefore like are people going to make fiveyear investments on hardware solely on you know an an asich that is more specialized or are they going to do so they're going to have some bucket of more general purpose compute and so you see this with like Google's paying $11 an hour per GPU to XAI for G for GPUs, right? Like that's insane, right? It's a
嗯,因为你去跟实验室的人聊,他们连自己一年后会用什么架构都不知道。真的,他们是真的不知道一年后自己会在做什么架构。他们有很多下注,他们有很多研究方向的押注,这也正是让人兴奋的地方,但他们不知道方向会走到哪里去。一般来说,他们知道自己手上有什么硬件,也在努力做协同优化,但归根结底,如果模型架构上出现新的突破,比如说把注意力机制换成别的东西,对吧,谁知道呢,或者你知道,突然之间发生了什么事,最优的硬件就变了。所以说,大家会不会在硬件上做五年期的投入,而且全部押在一个更专用的 ASIC 上?还是说他们会留一部分更通用的算力?所以你会看到,比如谷歌为了 GPU 向 xAI 支付每张 GPU 每小时 11 美元,对吧?这太夸张了,对吧?这个价格非常高,当然算力是稀缺的,等等等等,
便签笔记
49:13
very high amount of uh obviously compute is limited and and so on and so forth, but it's like very like insane, but at the same, you know, despite the fact that they have TPUs and so there's like some questions there like why do they do that? Um Google actually has three different design programs for TPUs. They're making a TPU with Broadcom. That's a different architecture than the TPU with MediaTek. That's a different TPU than the architecture that is, you know, I won't disclose, you know, by research. Um but, you know, they're they're making different architectures.
但这真的挺离谱的。而与此同时,你知道,他们明明自己有 TPU,所以这里就有一些疑问,比如他们为什么要这么做?嗯,其实谷歌有三个不同的 TPU 设计项目。他们和博通一起做一款 TPU。那和跟联发科一起做的 TPU 是不同的架构。而那款 TPU 又不同于另一款——那个我不方便透露,是我们研究得到的。嗯,但你知道,他们确实在做不同的架构。
便签笔记
49:38
It's not just like, oh, they're making TPUs with a couple vendors. It's the same architecture. It's different architectures. And the third one is a very different architecture from the first two. And so, I think people recognize that the local minima can happen. And therefore, um, I think everyone will have their own ASIC program. I think everyone will deploy billions of dollars of their own AS6, tens of billions of dollars. In the case of Google, hundreds of billions of dollars a year of their own AS6. But ultimately, they're also going to have workloads that don't use TPUs, right?
不是那种“哦,他们只是找了几个供应商代工 TPU”,同一个架构。不是的,是不同的架构。而第三个跟前两个的架构差别非常大。所以我觉得大家是意识到局部最小值这种情况可能发生的。因此,嗯,我觉得每家都会有自己的 ASIC 项目。我觉得每家都会部署几十亿美元自己的 ASIC,甚至几百亿美元。对谷歌来说,一年几千亿美元规模的自研 ASIC。但归根结底,他们同样会有一些不用 TPU 的工作负载,对吧?
便签笔记
50:06
Some of the Google bets that are not Gemini Deepbind actually primarily use GPUs. They don't use TPUs. Um, some of them also primarily use TPUs, right? It's a bit of a broad thing, but like, you know, maybe for drug discovery or for Whimo, you might not want to use TPUs. I won't say which one it is, but like, you know, there's there's there's there's different architecture bets and different paths for AI. AI for science may have different algorithmic patterns than than general intelligence AGI models. Um, and so I think we'll see we'll see diversity continue to proliferate. Yeah. and and and because the market has gotten so big, niches will be carved out and so that's makes it possible for companies to have their niche and actually make money even if the majority of the pie goes to Nvidia and TPU and tranium.
谷歌那些不属于 Gemini/DeepMind 的押注里,有一些其实主要用的是 GPU,而不是 TPU。嗯,其中也有一些主要用 TPU,对吧?这个说法有点笼统,但比如说,做药物发现或者做 Waymo,你可能就不想用TPU。我不会说具体是哪一个,但你知道,确实存在不同的架构押注和不同的 AI 路径。AI for science 的算法模式可能跟通用智能、AGI模型很不一样。嗯,所以我觉得我们会看到多样性继续扩散。是的,而且因为市场已经变得这么大,会有很多细分市场被切出来,这就让一些公司有机会守住自己的细分领域并且真的赚到钱,哪怕蛋糕的大头被英伟达、TPU 和 Trainium 拿走。
便签笔记
12算力紧缺、毛利率与杠杆之忧
50:46
>> Yeah. >> Okay. Love that. Can we talk about the data center buildout? Like one, it seems like I mean by all accounts if you look at the charts like dollars per compute hour, we are in the middle of a crazy compute crunch. Um and it seems like it's both a demand and supply side crunch, right? demand for long agents skyrocketing, supply, all these data center buildouts are delayed. Um, do you think this we're in a compute crunch for the foreseeable future or do you think it alleviates at some point?
>> 是的。>> 好,我很喜欢这个说法。我们能聊聊数据中心建设吗?第一,看起来——我是说从各方面看,如果你看那些图表,比如每算力小时的价格,我们正处在一场疯狂的算力紧缺之中。嗯,而且看起来这既是需求端也是供给端的紧缺,对吧?长时间运行的 agent 需求暴涨,而供给端所有这些数据中心的建设都在延期。嗯,你觉得我们在可预见的未来会一直处于算力紧缺吗,还是说某个时点会缓解?
便签笔记
51:10
>> Yes, every quarter we're deploying vastly more compute than the prior quarter and there's more data centers built than the prior quarter. Um, this year there's going to be 20 gigawatts uh even accounting for the delays and next year there's going to be more than 30 gigawatts accounting for the delays. Um, of course delays happen on everything, right? Anything hardware can have a delay. That's that's just the reality of life. Are we gonna have a compute crunch for the rest of our lives? It depends on what happens with models. But like the TAM for Mythos, you know, Mythos 5, Fable 5 is not just like 2x that of Opus, right? The model is so much better and it can do so many more tasks that the Tamford is way larger than that. And yet compute in the world did not double in the last, you know, six months, right? From, you know, Opus or maybe like seven or eight months since Opus 45 launched to now. huge you know 46 47 48 were improvements but fable and methos were like a huge step function
>> 是这样,我们每个季度部署的算力都远远超过上个季度,建成的数据中心也比上个季度多。嗯,今年就算把各种延期都算进去,也会有 20 吉瓦,明年把延期算进去也会超过 30吉瓦。嗯,当然什么东西都会延期,对吧?只要是硬件就可能延期。这就是现实。我们会不会一辈子都处在算力紧缺里?这取决于模型会怎么发展。但比如说 Mythos 的 TAM,你知道,Mythos 5、Fable 5 的 TAM 可不只是 Opus 的两倍,对吧?模型强了这么多,能做的任务多了这么多,所以 TAM 的扩张远不止那个倍数。然而过去六个月里,世界上的算力并没有翻倍,对吧?从 Opus——或者说自从 Opus 4.5 发布到现在大概七八个月。当然 4.6、4.7、4.8都是改进,但 Fable 和 Mythos 是一次巨大的阶跃式提升。同一时间段内,世界的算力并没有翻倍,也没有翻两番,
便签笔记
52:02
improvement the world's compute did not double in that or or quadruple or whatever in that same time frame but the demand for useful tasks that can be done by AI the number of useful tasks and the value of them that can be done by AI has and so now the question is what happens well obviously anthropic in Q2 is profitable their net income profitable um excluding stockbased compensation um And and I think by Q3 they may even be profitable including stockbased compensation. That's like how profitable they're getting. And their margins on a on a on an Opus token, at least Opus 48 token is like north of 80% for the API price. They've got a lot of deals where their total corporate gross margins gets clawed down a little bit uh because of like how they do bedrock deals and vertex deals and things like that. But ultimately their their per token margin is so high. Well, then if you don't have the cap, they have the capability to pay ultimately every GPU they buy at above market rate. You know, they also bought
或者随便什么倍数。但 AI 能完成的有用任务的需求——AI 能完成的有用任务的数量以及它们的价值——确实增长了。所以现在的问题是接下来会怎样。嗯,很明显 Anthropic 在第二季度是盈利的,净利润为正——嗯,是在不计股权激励的情况下。嗯,而且我觉得到第三季度,他们可能连算上股权激励也能盈利。你看他们赚钱能力已经到这个程度了。而他们在一个 Opus token 上的毛利率,至少是 Opus 4.8 的 token,按 API 价格算是超过 80% 的。他们有很多交易会把公司整体毛利率往下拉一点,嗯,因为他们做 Bedrock 交易、Vertex 交易之类的方式。但归根结底,他们单 token 的利润率非常高。那么,如果你不受上限约束,他们其实有能力用高于市场价的价格去买每一块 GPU。你知道,他们也从 SpaceX 那里
便签笔记
53:02
GPUs at above market rate from SpaceX, which is below the rate of Google, but that's because they signed earlier. Um, you know, it's it's something that, you know, other companies, maybe a ventureback company or company that's not really got positive uh margins can't necessarily do, right? What is the cost benefit ratios like every GPU I rent because I'm out of compute capacity I can immediately turn around and sell tokens on it or every TPU or every tranium I can immediately sell tokens on it at a positive margin and if I'm running 75% gross margin and I double the cost of the compute it's fine I'm still running 50% gross margin and spinning up more compute nodes is not really necessarily a human requiring task for them if they're renting them and so ultimately it's like well my NOI still goes up right and and so I'm going to rent GPUs at whatever price at some level whatever price I want to pay I can pay.
以高于市场价买过 GPU,那个价格低于谷歌付的价,但那是因为他们签得更早。嗯,你知道,这是那种其他公司——比如一家风投支持的公司,或者一家利润率根本不为正的公司——做不到的事,对吧?成本收益比是什么样的?就是说,因为我算力不够,我每租一块 GPU,我马上就能在上面卖 token;每一块 TPU、每一块 Trainium,我都能马上在上面以正毛利卖 token。如果我原本跑着 75% 的毛利率,就算算力成本翻倍也没关系,我还有 50% 的毛利率,而且对他们来说,如果是租来的算力,多拉起一些计算节点也不太需要人力投入,所以说到底就是:我的净营业收入还是在往上走,对吧?所以我会以任何价格去租 GPU,在某种意义上,我愿意付多少就能付多少。
便签笔记
53:47
>> I have almost the reverse question of like at some point does this compute build out go bump at night? Earlier today I think there was a tweet like Cuso publicly said one of their customers had asked to halt construction on one of their data center buildouts. Like it seems like everybody in the ecosystem is so levered right now to like we got to build, we got to go build, we got to build. High leverage high growth to me is like makes me very very nervous as investor. Like >> wait hold on. High leverage high growth means small amount of equity has huge upside. You're not a debt investor.
>> 我的问题几乎是反过来的:到某个时候,这场算力建设潮会不会让人夜里睡不着觉?今天早些时候我看到一条推文,好像 Crusoe 公开说他们有一个客户要求暂停某个数据中心项目的建设。感觉整个生态里的每个人现在杠杆都开得很足,就是“我们必须建、必须去建、必须建”。高杠杆加高增长,作为投资人这让我非常非常紧张。比如—— >> 等等,先打住。高杠杆高增长意味着一小笔股权有巨大的上行空间。你又不是债权投资人。
便签笔记
54:14
>> You're a credit you're an equity investor, right? >> Let's go. >> Um, >> look, you got you got to go to the school of private equity. Levered buyouts only. >> I actually come from the school of private equity. >> Oh, awesome. >> She forgot the school. It's been a VC for too long. >> Yeah. >> No, I just do revenue multiples. No, but are you do you see any signs of that? Are you worried about that? >> I I I see what you mean. Right. And that sort of goes back to the model point, right? Obviously if the models expanding the total economic valuable like work sort of the dark GDP uh report that we did and the you mentioned earlier um if the work that these models can do does not expand faster than the compute capacity then that tide turns right and over the last six months that tide has been you know very much levered in this direction of um you know the models can do more work or can is exp or expanding their TAM of work they can do faster than the compute is increasing And so prices go up. It's very possible that
>> 你是股权投资人,不是信贷投资人,对吧?>> 冲啊。>> 嗯—— >> 你看,你得去上一上私募股权那一课。只做杠杆收购。>> 我其实就是私募股权出身的。>> 哦,那太棒了。>> 她把那套东西忘了,做 VC 做太久了。>> 是啊。>> 不,我现在只看收入倍数了。不过说正经的,你有看到这方面的任何迹象吗?你对此担心吗?>> 我明白你的意思。对。这某种程度上又回到了模型那个点,对吧?很显然,如果模型在扩大总的经济价值——就像我们做的那份“暗 GDP”报告里说的,你刚才也提到了——嗯,如果这些模型能做的工作扩张速度赶不上算力增长的速度,那潮水就会转向,对吧。而过去六个月里,这个天平一直是非常明显地倒向这一边的:嗯,模型能做更多工作,或者说它们能做的工作的 TAM 扩张得比算力增长更快。所以价格在涨。当然也完全有可能,模型进展突然停滞。
便签笔记
55:09
all of a sudden model progress stops. You talk to anyone at Enthropic or OpenAI, maybe they're drinking the Kool-Aid, but you talk to basically all of them, they're like, "No, no, no, no. Model progress still go up." Um, and so, you know, ultimately, you know, current methods could stall somewhere. I'm not sure where that would be. It seems like we have line of sight to model improvement, rapid model improvement. And in fact, models are improving faster than they were six months ago or a year ago because there's I wouldn't call it recursive self-improvement, but basically the engineer the models are helping write all the info and and launch the next model sooner and sooner and sooner. So you've got this like pseudo recursive self-improvement loop going and so the models are getting better and better and better faster. Um and so but ultimately, you know, capital is a big problem which is why Google raised capital. You know, they they've got an ungodly amount of SpaceX, right?
你去问 Anthropic 或者 OpenAI 的任何人,也许他们是自己给自己灌迷魂汤,但你基本上问遍所有人,他们都会说:“不不不不,模型进展还会继续往上走。”嗯,所以说到底,现有的方法可能会在某个地方卡住。我也不确定那会是在哪里。看起来我们对模型改进是有清晰视野的,而且是快速的改进。而且事实上,模型现在的进步速度比六个月前或者一年前更快了,因为——我不会把它叫作递归自我改进,但基本上就是模型在帮着写各种东西、帮着把下一代模型越来越快地推出来。所以你就有了这么一个类似伪递归自我改进的循环,模型因此变得越来越好,而且越来越快。嗯,但归根结底,你知道,资本是个大问题,这也是谷歌为什么去融资。你知道,他们手上有多到离谱的 SpaceX 股份,对吧?
便签笔记
55:56
They own like 5% of the company. >> I think a little more, but yeah. Yeah, maybe. >> I think at one point they had like 10%. >> Larry Page invested a billion dollars at a $10 billion valuation, got 10% of the company, it got diluted, like all this. But that was one of the greatest investments of all time. Good job, Larry. The guy. >> So, they know they have like a hundred billion dollars in the bank that they can sell in, you know, nine months or whatever from the lockup. >> And they have all the gross profit they do, and yet they still modeled that. and they were like we need to raise capital and so they did an offering and it's like that's insane. So that tells you how much they think they need to spend.
他们大概持有这家公司 5%。>> 我觉得还要多一点,不过嗯,也可能吧。>> 我记得有一度他们持有大概 10%。>> 拉里·佩奇当年以 100 亿美元估值投了 10 亿美元,拿到公司 10%,后来被稀释了,诸如此类。但那是史上最伟大的投资之一。干得漂亮,拉里。牛人。>> 所以他们知道自己账上有差不多一千亿美元的东西,等锁定期过了,大概九个月之后就能卖掉。>> 而且他们还有那么多毛利润,可他们还是把账算了一遍,然后觉得我们需要融资,于是就做了一次发行,这太夸张了。这说明他们认为自己需要花多少钱。
便签笔记
13每吉瓦不等值:数据中心分层
56:28
But capital is like really, you know, you know, Meta's do Meta did announce that they're going to do a raise. Stock tanked. People don't like it, but you know, that's all these companies are going to raise capital, whether it be debt or equity. At some point, money spiggots will have to, you know, slow down. But right now, every GPU that Amazon adds, they're making higher revenue or every TPU or tranium, you know, whoever anyone adds is is making is making gross profit. I do a little bit of a tea up on this to turn into a question for you. But like >> as we talk about this for me, the thing that's going my through my head that's that is almost an alternative hypothesis for like the Crusoe example. I'm going use an analogy in oil like in oil Saudi Arabia has way lower cost per barrel to produce oil than a lot of other countries. There's also like the purity of the oil. A lot of you know Saudi has generally like very low contaminants in their oil which makes refining easier all of this. The question for me is like
但资本真的是,你知道,Meta 也宣布了要融资。股价跌了。大家不喜欢这样,但你知道,这些公司都会去融资,不管是债还是股。到某个时候,钱的水龙头总要拧小一些。但现在,亚马逊每加一块 GPU,他们的收入就更高,每加一块 TPU 或 Trainium 也是一样,你知道,任何人加上去都在产生毛利润。我稍微我在这上面铺垫一下,把它变成一个问题问你。但是 >> 我们聊这个的时候,我脑子里一直在想的是对于 Crusoe 那个例子来说,几乎算是一个替代假设。我打个石油的比方,就像石油行业里沙特阿拉伯的单桶开采成本比很多国家都低得多。而且还有石油的纯度问题沙特的石油里杂质通常非常少,这让炼化更容易诸如此类。我的问题是,你看每一吉瓦落地的产能,如果说
便签笔记
57:26
when you look at for every gigawatt that's being put in the ground if call it the 20 gigawatts coming online today like how much like how much homogeneity do you see in those gigawatts? Is it something like and I don't you can tell me whatever metric you think is right but like are Google's gigawatts two times more valuable than say most Neoclouds because they have optical switches and they have like they've been doing it for a long time and like they know how to do power smoothing because I think this could be the alternative hypothesis that some of the people that are it's like the people that are good at at building data centers they they should just do it to the max because there's so much demand and there's so much better than it, but then maybe we're starting to see the early signs of the people that are like not as good at it kind of getting hit a little. So I like I don't know the reality here. I'm just curious how you think about this.
今天有 20 吉瓦上线,这些吉瓦之间的同质化程度有多高?是不是某种……你可以用你觉得合适的任何指标来说,但比如说谷歌的吉瓦是不是比大多数 Neocloud 的价值高两倍?因为他们有光交换机,而且他们干这行很久了,他们知道怎么做功率平滑。因为我觉得这可能是另一种解释:有些人就是那种特别擅长建数据中心的人,他们就应该开足马力干到底,因为需求这么大,而且他们比别人强太多;但也许我们现在开始看到一些早期迹象,就是那些不太擅长这活儿的人开始受到一点冲击。所以我不太清楚这里的真实情况,就是好奇你怎么看这件事。
便签笔记
58:18
>> So so far um there there are metrics for this, right? So uh tranium sells at sub10 billion per gawatt rental rate uh to anthropic and to open aai. GPUs at least before the craziness of the last six months usually went around 12 to$13 billion per gigawatt. So the rental rate and this is from a neocloud versus Amazon even and now when Amazon sells GPUs they'd also be 13 or so >> and my understanding of that also is that those number like Amazon subsidized that a little bit so that it's like I actually think the numbers were even like I think the disparity was even more >> it's less than 10. It's less than 10 but there's like some weird basically >> and like look I my understanding obvious like anthropic played a big role in making tranium useful in terms of you know writing all the libraries etc and and so like I >> everything I hear is that tranium's really freaking good hardware and it's getting way like way better and obviously anthropic now using it a lot so hopefully we would see that price go
>> 目前来说,其实是有指标能衡量这个的,对吧?比如 Trainium 的租赁价格是每吉瓦不到 100 亿美元,卖给 Anthropic 和 OpenAI。而 GPU,至少在过去半年这波疯狂之前,通常是每吉瓦 120 到 130亿美元。所以这个租赁价格,而且这还是 Neocloud 和亚马逊的对比,现在亚马逊自己卖GPU 的话也是 130 亿左右。 >> 而且我的理解是,那些数字里亚马逊还补贴了一点,所以其实我觉得实际数字甚至更……我觉得差距甚至更大。 >> 是低于 100 亿。是低于 100 亿,但里面有些奇怪的……基本上 >> 而且你看,我的理解显然是,Anthropic 在把 Trainium 变得好用这件事上出了很大力,比如写各种库之类的,所以我 >> 我听到的所有说法都是 Trainium是真的非常好的硬件,而且在变得越来越好,现在 Anthropic 显然也在大量使用它,所以希望我们能看到那个价格
便签笔记
59:18
up you know like per the the deal they did was actually like there was a floor mechanism them and like it if it didn't do well it would be like cheaper and then to the point where it's cancelceable and you know if it if it did really well the price is kind of higher um but effectively um less than 10 right is is where tranium shakes out at whereas GPUs I mean this the SpaceX deal again was like 25 or something crazy billion dollars per gigawatt or $25 million per megawatt right a year rental rate with Google I was like that's a crazy divergence now obviously if if if Amazon was selling tranium today it' probably be more expensive than 10 because the comput shortages, but you you do see this already in the sense of uh with data centers oftent times a rental price of a data center if you're doing collocation, right? Not compute in there, but just power. Here's the data center. Um you you price it generally on a uh dollars per kilowatt per month. And so they used to be $60 per kilowatt hour per month, and now you
往上走。他们做的那笔交易其实是有一个价格下限机制的,如果表现不好,价格就会更便宜,甚至到可以取消的程度;如果表现非常好,价格就会高一些。但实际上,Trainium 最终落在 100 亿以下这个区间,而 GPU 呢,我是说 SpaceX那笔交易又是每吉瓦 250 亿美元这种夸张的数字,也就是每兆瓦每年 2500 万美元的租赁价,跟谷歌做的。我当时就想,这个差距太夸张了。当然,如果亚马逊今天卖 Trainium,价格可能会比100 亿更贵,因为算力短缺。但你其实已经能看到这种现象了,比如数据中心方面,很多时候数据中心的租赁价格,如果你做的是托管(colocation),对吧?里面不含算力,只有电力。就是给你一个数据中心。你一般是按每千瓦每月多少美元来定价的。以前是每千瓦每月 60美元,现在你看到的成交价大概在 120 到 160 之间。但不同
便签笔记
60:12
see things transacting at anywhere from like 120 to 160. Um but different quality data centers, this actually you've I've seen data centers go as high as 200. um when the customer is not such a great credit rating and then the data center is a pretty good one. And I've seen stuff go as low as 100 still or in like India go like as low as 80 because the grid's not reliable, the internet connection's not great and it's a pretty mid data center but at least it's a data center. Um and so you you see this huge discrepancy there already.
品质的数据中心差别很大,我见过高到 200 的数据中心。那是客户的信用评级不太好,而数据中心本身相当不错的情况。我也见过低到 100 的,或者在印度那种低到 80 的,因为电网不可靠、网络连接也不太好,数据中心也就一般般,但至少它是个数据中心。所以你已经能看到这里巨大的差异了。
便签笔记
60:41
Um in the case of like data center construction, usually the pitfalls they just fail. There's a lot of people who fail, you know, claim they're g they're like they're like four guys they they're like, "Yeah, we here I bought some turbines. I put the money down for them. I'm gonna build a data center." And then they get delayed, delayed, delayed, and fail. Um, so you have to like probability, weight, time, weight, time lag, the teams that suck versus don't. Um, and and sort of, you know, our data center model does that. We kind of track every data center uh and try and do this for every single one based on, you know, equipment that they're using and all these things. One of the things you mentioned about Google is you know in a gigawatt data center they'll actually put like 1.5 gigawatts of hardware and because they have such understanding all the way from workload to u you know they're able to slosh the power around and so instead of you know constantly you know a gigawatt of compute which
至于数据中心建设方面,通常的坑就是干脆失败。有很多人失败,你知道,他们号称……就四个人,说:“对,我买了几台涡轮机,我付了定金,我要建一个数据中心。”然后就是一拖再拖,最后失败。所以你得按概率加权、按时间加权、按时间滞后来算,看哪些团队烂、哪些不烂。我们的数据中心模型就是干这个的。我们基本上追踪每一个数据中心,试着对每一个都做这样的评估,基于他们用的设备等等各种因素。你刚才提到谷歌的一点是,在一个吉瓦级的数据中心里,他们实际会塞进大概 1.5 吉瓦的硬件,因为他们从工作负载一直到底层都非常了解,所以他们能把功率来回调配,这样就不用像通常那样,一吉瓦的算力在功耗上一般只跑在 60% 到 70% 的
便签笔记
61:26
typically runs at like 60 or 70% utilization in terms of power consumption not utilization of the hardware someone's always renting it um they're now running it at like you know you know that 60 to 70% means it's at a gigawatt and you're using the full gigawatt um you see people doing deals with including Google with utilities where they're like, "Oh, well, I know this grid can sustainably take a gigawatt, but you know, except for three days of the year, you can actually do two gigawatts, so give me two gigawatts and then just tell me to turn off." And so they'll do that. And so these sorts of tricks and then you need to have supreme management of workload, backup power, all these things, um, generators on site to figure out how to actually keep it 2 gigawatts sustainably. When people do this, they're able to charge more. Whether it be I'm actually selling two gigawatts despite only having one gigawatt because those three deers deal days I'm be able to deal with via battery, gas, etc. or I figured out how
利用率——这里说的是功耗,不是硬件利用率,硬件总是有人在租的——他们现在跑到的是,那个 60% 到 70% 就相当于一吉瓦,而你实际用满了整个吉瓦。你还能看到有人做这种交易,包括谷歌和电力公司谈:“我知道这个电网可以稳定支撑一吉瓦,但除了一年中的三天之外,其实可以支撑两吉瓦,所以给我两吉瓦,需要的时候通知我关掉就行。”他们就这么干。所以就是这类技巧,然后你需要对工作负载、备用电源这些有极强的管理能力,现场配发电机,想办法真正把两吉瓦稳定跑起来。做到这些的人,就能收更高的价钱。可能是我实际卖出了两吉瓦,尽管我只有一吉瓦,因为那三天我可以靠电池、燃气之类的顶过去;也可能是我搞定了
便签笔记
62:15
to build power on site. Now I have a gigawatt where no one else does and so I'm able to do it quickly. Um it's not necessarily transacting for a higher price. It's that I'm selling more gigawatts. And sometimes there are levers where you're selling more gigawatts is where where each gigawatt is selling at a different price. Um I think it's more on the data center and energy layer. It's more about just having it versus not and then that being delayed or not. It's more binary. But on the compute side, I do think there's a lot more um interesting work there.
现场发电,于是我手上有一吉瓦而别人没有,所以我能很快交付。这不一定是以更高的单价成交,而是我卖出了更多吉瓦。有时候也有一些杠杆,让你卖出更多吉瓦时,每一吉瓦的售价是不一样的。我觉得在数据中心和能源这一层,更多是有没有的问题,以及会不会延期,比较偏二元。但在算力这一层,我确实觉得里面有更多有意思的东西。
便签笔记
62:40
Right? A gigawatt given to Enthropic is objectively worth more revenue than a gigawatt given to OpenAI. And it seems that both of them could sell every gigawatt that they have right now. Uh given rate limit problems and token max limit and all these sorts of things at OpenAI and anthropic. Uh especially since Codex 5.5 came out, it's much better. And then likewise, if you gave a gigawatt to SpaceX, you know, they turn >> my my guess, like my suspicion is that they're they probably make better use of the, you know, hardware than most people. Um, just like I think people underestimate how much networking experience they have from Starlink in particular and also how much just like power management experience they have via from Tesla.
对吧?给 Anthropic 一吉瓦,客观上带来的收入就是比给 OpenAI 一吉瓦更多。而且看起来两家现在有多少吉瓦都能卖光。考虑到 OpenAI 和 Anthropic 那边的速率限制问题、token 上限之类的种种情况。尤其是 Codex 5.5 出来之后,好用多了。同样地,如果你给 SpaceX一吉瓦,他们会…… >> 我的猜测、我的直觉是,他们大概能比大多数人更好地利用这些硬件。我觉得大家低估了他们从 Starlink 积累的网络方面的经验,也低估了他们从特斯拉那边积累的电力管理经验。
便签笔记
63:26
>> Yeah. people like Brett Mayo are like incredible like >> pretty good. >> Yeah. >> And so I I think that like for me that's actually I think probably the thing that might I don't actually know the answer but I think that might be missing from the analysis a lot of people are doing. >> I think I think it's also the fact that when Coreweave builds a gigawatt even though their GPU compute is objectively better than Amazon or Google or Microsoft's in terms of performance. We've tested the performance and reliability. Um, the problem is Google sells it six months before they have it up and they need to turn around and take that paper that they signed to get debt uh with that credit backing and then turn around so they can actually pay for the PO that they've already issued you know for the order that they've already issued. Whereas SpaceX was like no no no this is running now buy it right and it's it's a big discrepancy when you have a balance sheet to do that versus not and that also helps your revenue per
>> 对。像 Brett Mayo 这种人简直厉害得不行 >> 相当强。>> 是的。>> 所以我觉得这一点,对我来说其实是……我也不知道确切答案,但我觉得很多人在做分析时可能忽略了这一块。>> 我觉得还有一点是,当 CoreWeave 建一吉瓦的时候,尽管他们的 GPU 算力在性能上客观来说比亚马逊、谷歌或微软的更好。我们测过性能和可靠性。问题在于,谷歌是在真正把产能建起来的六个月前就把它卖掉了,然后他们得拿着签下的这份合同去凭这个信用背书借债,再回过头来才能真正支付他们已经发出的采购订单,就是他们已经下的那笔订单。而 SpaceX 是那种“不不不,这个现在就在跑,买就完了”。有没有一张能撑住这种玩法的资产负债表,差别是很大的,而且这也让你每兆瓦的收入高得多。
便签笔记
14Neocloud 为什么能赢,黄仁勋在下什么棋
64:13
megawatt like be much higher. >> Why does the Neocloud opportunity even exist? Because if you had asked me five years ago, I would have said the hyperscalers are going to own this. And you know, you mentioned just now core weight has better performance than than the hyperscalers. Like what why does this opportunity exist maybe at the macro level and then in the execution level? >> Yeah. So in 2023 I wrote a report that had uh Amazon really hate me. Um it was called Amazon cloud crisis. So I talked about how Amazon was the best cloud because they had their nitro nicks which offered like tenant isolation. all the hypervisor ran on the nick and then you could sell all the cores and they had you know custom SSDs that they made and they'd buy the raw NAND and they'd have lower cost because they'd buy the raw nand and build their own SSDs um and you know they had their custom graviton CPUs and that drove down cost for for per core and so they had all these things that enabled them to sell more cores
>> 为什么 Neocloud 这个机会会存在?因为如果你五年前问我,我会说超大规模云厂商会把这块全吃下来。而且你刚才也提到 CoreWeave 的性能比超大规模厂商更好。那这个机会为什么会存在?也许先从宏观层面讲,再讲执行层面?>> 好。2023 年我写过一篇报告,让亚马逊非常讨厌我。标题是《亚马逊云危机》。我在里面讲了亚马逊为什么是最好的云:因为他们有 Nitro 网卡,能提供租户隔离,整个hypervisor 跑在网卡上,于是所有 CPU 核心都能拿去卖;他们还有自研的 SSD,他们直接买裸 NAND,成本更低,因为买裸 NAND 自己造 SSD;还有他们自研的 Graviton CPU,把每核心的成本压了下来。所以他们有这一堆东西,让他们能卖出更多核心、有更好的安全性、不错的网络——但这些都是为传统 CPU、
便签笔记
65:01
have better security good networking for but this was all for the traditional CPU better storage for the traditional you know cloud world but in the AI cloud a lot of this stuff hurt performance right these nitro necks bad for performance, still are worse performance. Although they've caught up a lot because they've had a couple iterations to like, you know, improve them, but they're still worse for performance. Um, a lot of the security stuff doesn't matter because it's not like I'm time splicing users or splicing a socket into many users, right? It's like no one buy rents a single GPU and an 8GPU server. No one rents a single GPU in a 72GPU rack. They rent the whole rack and in fact, they rent many of the racks. And so, and then and then there's no like, oh, I rent for six hours and I give it back. it's everyone has these long-term contracts.
为传统云世界准备的更好存储。可到了 AI 云里,这里面很多东西反而拖累性能,对吧?这些 Nitro 网卡对性能不利,现在性能还是更差。虽然他们追上了不少,因为已经迭代了好几代做改进,但性能上仍然更差。而且很多安全方面的东西根本不重要,因为这里不是我把用户做时间切片、或者把一个插槽切给很多用户,对吧?没人会在一台八卡 GPU 服务器里只租一张 GPU,也没人会在 72 卡的机柜里只租一张 GPU。他们是整柜租,实际上是租很多个机柜。而且也不存在“我租六小时然后还回去”这种事,大家签的都是长期合同。
便签笔记
65:42
So, the mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away. Um, and a lot of the expertise that they did have were actually some of them were detrimental, right? Network performance for Google and Amazon. It was they had custom networks that were better for traditional CPU and for the stuff that they were doing, but actually worked for AI. Um and then in other cases it's like well you know Microsoft would save money by building their own data centers but their data center teams are actually were not actually that great and so when it came time to run you know when it was predictable building it was like fine when it came time to like actually double your forecast for the year it's like they fell on their face and they had to go get a bunch of NeoCloud capacity. I think so performance I think you know I think time to market's another one right ne you know these massive organizations no one's getting rich from building this data center faster right but you look at Crusoe for
所以 GPU 租赁市场的运作方式,让超大规模厂商的很多专长都失效了。而且他们确实拥有的一些专长,其中有些反而是有害的,对吧?比如谷歌和亚马逊的网络性能。他们有自研网络,对传统 CPU、对他们当时在做的那些事更合适,但放到 AI 上其实并不合适。另外有些情况是,比如微软想通过自建数据中心来省钱,但他们的数据中心团队其实并没有那么强,所以当需要真正上量的时候……在建设可预测的时候还行,可一旦要把当年的预测翻倍,他们就摔了个大跟头,只好去外面买一大堆 Neocloud 的产能。所以我觉得,性能是一点,上市速度是另一点,对吧?在这些庞大的组织里,没人会因为把这个数据中心建得更快而发财。但你看 Crusoe 这种例子,Chase 还有团队里其他所有人,我本来
便签笔记
66:33
example Chase and and and and all the other people at the team you know I was going to name some people at the team I'd rather not you know these all these people are getting rich if they deliver these this compute faster they're they're you know they're lever they're hyperlevered equity owners >> hey look they're also all coming from Bitcoin And and you know they you're not supposed to say that. >> Uh I mean a lot of the data center like their main data center guy came from Microsoft. >> I don't know. I'm just I'm just teasing.
想点几个团队里的人名,还是算了——这些人如果能更快把算力交付出来,是真的会发财的,他们是高杠杆的股权持有者。 >> 而且你看,他们其实也都是从比特币那边过来的。这话好像不该说。>> 呃,我是说,数据中心那边很多人,比如他们主要的数据中心负责人是从微软来的。>> 我不知道,我就是逗你玩的。
便签笔记
66:56
But it's uh you know it's like you learn you learn a lot when you're in a very high fluctuation you know market. >> What do you think was Jensen playing for each chess? >> Jensen absolutely hates a world where all the hyperscalers have all the power. There's a reason he's like blowing money on like random AI labs that like I don't even know if like it makes sense to but like you know he's blowing money and pumping them up and going to you know everyone around the world and saying you should invest in this company because he wants to create a multipolar world.
但你知道,在一个波动极大的市场里待过,是真的能学到很多东西的。>> 你觉得黄仁勋在下一盘什么棋?>> 黄仁勋绝对痛恨一个所有权力都掌握在超大规模厂商手里的世界。所以他才会到处砸钱投一些我都不确定是否说得通的 AI 实验室,你知道他就是在砸钱、把它们捧起来,然后跑遍全世界跟所有人说,你们应该投这家公司,因为他想创造一个多极的世界。
便签笔记
67:25
That's why he loves Chinese labs because he wants to create a multipolar world. A world where open anthropic and Google models are the only models is one in which he's screwed. Yep. >> Right. Um a world in which you know the hyperscalers are the only ones building compute is one he's screwed in. Yeah. >> And so, you know, of course he needs to point the allocation gun at NeoClouds, help back stop their clusters, do anything and everything because while today a GPU sold to Cruso and a GPU sold to um Coree and a GPU sold to Google and Amazon are all the same price for him, five years from now Cruso and Coree existing means Google TPU will be weaker and means Amazon tranium will be weaker and more inference being done with you know non-clos model labs is is better for firm. So I think you know the neocloud ecosystem is you know it's these people that are wild west these neolabs as well a lot of them have investments from Nvidia it's the wild west some will fail many will fail but you know some will emerge as really
这也是他喜欢中国实验室的原因,因为他想要一个多极世界。一个只有 OpenAI、Anthropic 和谷歌的模型的世界,对他来说就完蛋了。没错。>> 对。一个只有超大规模厂商在建算力的世界,对他来说也是完蛋的。是的。>> 所以当然,他得把分配大权这把枪指向 Neocloud,帮它们的集群兜底,无所不用其极。因为虽然今天卖给 Crusoe 的一张 GPU、卖给CoreWeave 的一张 GPU,跟卖给谷歌、亚马逊的 GPU,对他来说都是同一个价,但五年后,Crusoe 和 CoreWeave 的存在意味着谷歌的 TPU 会更弱、亚马逊的 Trainium 会更弱,而更多推理是由非封闭模型实验室来做的,这对他更有利。所以我觉得 Neocloud 这个生态,就是这群狂野西部式的人,还有这些新兴实验室也一样,很多都拿了英伟达的投资,就是狂野西部,有些会失败,很多都会失败,但也会有一些跑出来成为非常出色的团队,比如说,很奇妙地,Crusoe 这样一帮搞加密货币出身的人
便签笔记
68:20
great teams whether it be you know oddly cruso who's a bunch of crypto guys who then started building data centers and doing flared gas stuff or you know corweave who initially was a bunch of New York hedge fun guys >> they were also and then they were doing >> and crypto guys but then they they they like built you know there were a lot of people who didn't bubble up like them started around the same time just failed Right. And so I think you know um >> I gotta say both those teams are phenomenal. They deserve a lot of credit and it's like that's your point. But >> yeah, I mean my point is like he he you know you throw it's like you throw a bunch of like bait into the water and the best fish will figure out and survive, right? Um and and sort of the same way with the Neoclouds and and and he hopes the Neols as well. We'll see if any of the Neolabs really bubble up, but like you know Thinking Machines has a few hundred million dollars of ARR, right? That's pretty impressive even though they've had you know in the media
然后开始建数据中心、搞伴生气发电什么的,或者比如 CoreWeave,他们最初就是一帮纽约的对冲基金的人 >> 他们也是,然后他们在做 >> 还有搞加密货币的人,但后来他们就真的建起来了,你知道,其实有很多没能冒头的人,跟他们差不多同时起步,只是失败了。对吧。所以我觉得,你知道,嗯>> 我得说,这两个团队都非常出色。他们值得被高度肯定,这也正是你说的那个意思。但>> 是啊,我的意思是,他,你知道,就好比你往水里扔一大把饵,最厉害的鱼会想办法活下来,对吧?嗯,Neocloud 差不多也是这个道理,还有他也希望 Neolab 是这样。我们看看有没有哪家 Neolab 真的能冒出来,不过你知道,Thinking Machines 已经有几亿美元的 ARR 了,对吧?这挺厉害的,尽管媒体上说
便签笔记
69:05
it's like oh they've lost all this talent. It's like, well, but Tinker is doing a few hundred million of ARR. Like, that's pretty impressive for out of the gate a product that's less than six months old or whatever. Um, and and you, you know, we hope the same happens to other Neolabs and and so um you know, he wants a multipolar world. >> Truly, congratulations on the success. Thank you. Just the last thing I'll say is I've seen a little bit of this. I think the public, they can probably tell from listening to you how hard you work, but like it's clear you've just been working your ass off for more than a decade and it, you know, led to the last few years of being in the right place, right time. But like it's unbelievable what you've accomplished and I know it's just the beginning. So, >> thank you so much.
哦他们流失了这么多人才。但你看,Tinker 做到了几亿的 ARR。对一个上线还不到六个月左右的产品来说,这已经相当厉害了。嗯,而且你知道,我们希望其他 Neolab 也能这样,所以嗯,你知道,他想要一个多极的世界。>> 真心恭喜你取得的成功。谢谢。我最后再说一点,这些我多少也看在眼里。我想公众听你讲话大概就能感觉到你有多拼,但很明显你已经拼命干了十多年,这才有了过去这几年的天时地利、恰逢其时。但你所取得的成就真的令人难以置信,而我知道这才刚刚开始。所以,>> 太感谢了。
便签笔记
69:42
>> Thank you for doing this. >> Awesome.
>> 谢谢你来参加。>> 太棒了。
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

SemiAnalysis 创始人 Dylan Patel 认为 AI 效率的真正跃迁不来自硬件、系统或模型任何单一层的改进,而来自跨层的软硬件协同设计(co-design)——三个 2x 叠加不是 8x 而是 100x——并由此推演出芯片路线之争、算力紧缺、Neocloud 崛起与英伟达战略的底层逻辑。

核心要点

  • 协同设计才是 100x 的来源:Hopper→Blackwell 在 DeepSeek 最优部署上约 30x 提升,但过去三年"每瓦智能"的增益更多来自模型层(三年前是 GPT-4,如今 27B 总参数/2B 激活的小 Qwen 已远超之)。真正的突破是把专家形状、注意力算术强度、集合通信模式与芯片架构联合优化:DeepSeek V3 的专家形状为 Hopper 定制,V4 为 Blackwell 和华为芯片定制,所以 TPU 虽是好芯片却"跑 DeepSeek 很差"。他反驳"中国更擅长 co-design"——西方只是不公开,GPT-4o 与 V3 规模相近且更早。
  • GPU vs TPU 没有绝对优劣,取决于模型走向:矩阵乘单元尺寸、注意力结构、专家结构、网络拓扑(NVLink 交换机连 72 卡 vs Google ICI 无交换机直连 8000 芯片)都互相耦合。OpenAI 模型更稀疏,用 TPU 可能是灾难;Anthropic/Google 模型相对更稠密,用 GPU 训练可能是灾难。Anthropic 主要在 TPU 上预训练、在 Trainium 和 GPU 上推理;每代 Gemini 都为对应 TPU 代际定制,放回旧硬件性能就不佳。
  • CUDA 护城河已部分瓦解,但"生态形状"形成新的锁定:模型自身写内核已很强,大实验室早已 fork PyTorch,不再依赖开源栈。但 DeepSeek、Kimi、Qwen、小米等开源模型都为 GPU 协同设计,下游推理商与 RL 公司被迫用英伟达——这不是 CUDA 的护城河,而是模型形状的护城河;Google 因此必须自己做 Gemma 生态。
  • 吞吐-交互性曲线是一切的下游:同一硬件可做"100 用户×10 tok/s"或"1 用户×250 tok/s",价差 4x。目前 AI 基础设施是一刀切,未来会按工作负载分化(批处理 vs 即时响应),Claude Code fast mode 和 OpenAI 优先队列是先兆。InferenceX 每天在约 15 种芯片、超 5000 万美元(加 TPU/Trainium 后过亿)捐赠硬件上跑最新开源模型,公开帕累托最优配置,防止厂商拿别人次优点对比自己最优点。
  • 效率增速惊人但离人脑仍差数个数量级:同等质量的模型成本每年降约 60x,每瓦智能约 40x/年。不过人脑差距"无关紧要"——给计算机供电远比供养人类简单。
  • 技术瓶颈在内存与功率密度,而非只是供应链:NAND 单元 25 年、DRAM 单元 40 年无根本突破,HBM 只是堆得更多更快;下一步是把内存直接堆在芯片上让带宽爆发。二十年来芯片功率密度卡在约 1 W/mm²,英伟达从 1400W 到 Rubin 2000W、Rubin Ultra 约 4000W 靠的是加硅面积;正在研发突破 1 W/mm² 的方案,代价是散热与电气干扰。能源瓶颈也有土办法:把美国能大规模生产的柴油卡车发动机改燃气反推电机发电,靠汽修工维护。
  • 算力紧缺由模型 TAM 扩张快于算力增长驱动:今年新增约 20 GW,明年超 30 GW(已计入延误),但 Fable/Mythos 5 的可用任务范围远超 Opus 2x,而全球算力半年内并未翻倍。Anthropic Q2 已实现(剔除股权激励的)净利润,Opus 4.8 API 单 token 毛利超 80%,故它能以高于市价租每一块 GPU——毛利 75% 时算力成本翻倍仍有 50%。风险信号:若模型进步停滞、可做的经济工作不再快于算力扩张,潮水就会反转。
  • 算力质量高度不均:Trainium 租金低于 100 亿美元/GW,GPU 通常 120–130 亿,Google 从 SpaceX 租 GPU 高达约 250 亿/GW(每月每 kW 数据中心托管价从 60 美元涨到 120–160,差信用可达 200,印度约 80)。Google 在 1 GW 数据中心塞 1.5 GW 硬件靠工作负载级功率调度榨满,并与电网签"平时 2 GW、一年停 3 天"的协议。一 GW 给 Anthropic 产生的收入客观高于给 OpenAI;SpaceX 凭 Starlink 网络与 Tesla 电力经验也被低估。
  • Neocloud 存在的原因是超大规模云的传统优势在 AI 上反成累赘:Nitro 网卡、租户隔离、定制网络都为 CPU 云优化,对整机架长租的 GPU 市场无用甚至拖累性能(CoreWeave 实测性能与可靠性优于三大云);大公司没人因为建得快而暴富,Crusoe 这类高杠杆股权持有者有强动机。黄仁勋投神经云与新实验室(含中国实验室)是为造多极世界——若只有三家闭源实验室和超大规模云,英伟达就完了。
  • 终局是多元而非"每家一款 ASIC 独占":目前所有芯片结构雷同(中央逻辑 die + 四周 HBM + 上网络下 IO),未来硬件与模型架构会分叉,有人会陷入局部最优。实验室连一年后的模型架构都不知道,所以通用算力永远有市场;Google 同时有 Broadcom、联发科和第三条不同架构的 TPU 项目,且部分非 Gemini 业务主要用 GPU。

结论与值得注意的细节

  • Patel 预测 2030 年 OpenAI+Anthropic 合计超 100 GW 推理算力,2040 年达太瓦级;太空数据中心 3–5 年内无关紧要(2030 年占增量 <1%),但 2040 年超过一半的增量算力将在太空。
  • 对 Cerebras:快速推理是真市场(SemiAnalysis 内部几乎只用 fast mode),但 SRAM 芯片难以承载 10 万亿级参数、百万上下文的模型,而实验室收入和用量集中在最强模型——Fable/Mythos 发布当天就有大量用户涌向更贵的高档模型,"看美元不看 token 量"。
  • 他最反感的论调是"AI 没有 ROI"和"模型不再进步"——基准饱和后新基准又飙升;"拿着全部事实却得出完全错误的结论"最可笑,但更重要的是尽快犯错并修正。
  • 模型改进正在加速:不是严格的递归自我改进,但模型已在帮工程师写基础设施并提前发布下一代,形成伪递归循环。资本是硬约束——Google 手握约 1000 亿美元 SpaceX 持股仍去融资,Meta 宣布融资股价大跌。
  • 个人背景细节:8 岁修 Xbox 360 红环故障入门硬件,12 岁起当 Reddit 硬件版主,曾是星际争霸 2 北美天梯宗师;2020 年被人肉后放弃匿名、24 岁生日开博,此后四年"无家可归"地跑遍全球每年 40+ 场会议;SPIE 光刻会议第一次听懂不到 10%,至今仍未全懂。他最看重的技能是能与供应链任何环节的人建立联系,而"1980 年代唯一一家化学品工厂烧毁导致内存价格翻倍"的旧事至今仍在重演。
核心句型 · 9
1. It's as … as the … is.
“It's as accurate as the information is.”
用同级比较回避正面回答:准确度取决于信息来源。适合被问到不便证实的数字时,既不否认也不确认。仿写:It's as reliable as the source is.
2. Long story short, …
“Long story short, I had to open it up and short the temperature sensor and it fixed it.”
口语中跳过冗长过程直奔结果的标准开场。讲经历时用来压缩叙事,让听者知道你只给结论。
3. everything is downstream of …
“Everything is downstream of that curve”
用「上下游」表达因果依赖:某事物决定其后所有环节。适合论证某个指标或决策的根本性。仿写:Hiring is downstream of culture.
4. It's not that …, it's that …
“That's not really CUDA as a moat, it's that the downstream product is more optimized for Nvidia”
否定表面解释、给出真实原因的对比结构。用于纠正常见误解,先破后立,语气清晰有力。
5. instead of being … to X, it's actually Y
“Instead of being multiplicative to 8x, it's actually 100x”
先给按常识推算的结果,再用 actually 抛出反直觉的真实值,制造落差感。适合强调协同效应或超预期的数据。
6. I could with a straight face argue that … or …
“I could with a straight face argue with you like that GPUs are way better than TPUs or TPUs are way better than GPUs”
表示两种对立立场都有充分理由,自己能为任一方辩护。用来铺垫「答案取决于条件」的结论。
7. whether it be A or B
“Whether it be open models or closed models”
正式书面的让步结构,be 用虚拟语气原形。列举两种情况都不影响结论时使用,比 whether it is 更书面。
8. the question is like how do you …
“The question is like how do you leap how do you scoot back over to like the absolute minima”
把复杂讨论收束为一个核心问题的口语句式。先承认现状,再用 the question is 聚焦真正未解的难点。
9. You throw … into the water and the best … will survive.
“You throw a bunch of like bait into the water and the best fish will figure out and survive”
用撒饵捕鱼比喻广撒网、任其优胜劣汰的策略。适合描述投资、招聘或产品试错。
词汇精讲 · 156 · 按出现顺序
wrestle with a pig phr. 0:00
跟猪摔跤;喻与胡搅蛮缠者纠缠只会自己弄脏、对方却乐在其中
went very long phr. 0:50
重仓做多;金融用语,此处喻全力押注某领域
premier /prɪˈmɪr/ adj. 0:50
首屈一指的,最顶尖的
crushing /ˈkrʌʃɪŋ/ v. 0:50
(口语)做得极出色,大获全胜
affiliation /əˌfɪliˈeɪʃn/ n. 1:30
隶属关系,关联
profiling /ˈproʊfaɪlɪŋ/ n. 1:30
(根据外貌特征)画像、归类判断
step stool n. 2:23
踏脚小凳
jockey /ˈdʒɑːki/ adj. 3:53
(俚)运动型的,像运动员的(jock 的形容词化)
short /ʃɔːrt/ v. 3:53
使短路;此处指短接传感器
Pandora's box phr. 4:16
潘多拉的盒子;一旦打开就无法收回的麻烦或沉迷
tinge /tɪndʒ/ n. 4:16
一丝,少许(色彩、气息、倾向)
neck beards n. 5:06
(网络俚语)技术宅,不修边幅的极客
price performance n. 5:06
性价比
margins /ˈmɑːrdʒɪnz/ n. 5:06
利润率,毛利
grandmaster /ˈɡrændˌmæstər/ n. 5:32
宗师(游戏或棋类最高段位)
obsession /əbˈseʃn/ n. 5:32
痴迷,执念
tryh hard maxing phr. 6:15
(网络俚语 try-hard maxing)拼命卷到极致
quant /kwɑːnt/ n. 6:42
量化分析师
culmination /ˌkʌlmɪˈneɪʃn/ n. 6:42
顶点;多件事的汇聚、集大成
screwed out of phr. 6:42
被坑掉(应得之物)
rightsized /ˌraɪtˈsaɪzd/ v. 6:42
(委婉)被裁员,被「优化」
dementia /dɪˈmenʃə/ n. 6:42
痴呆症
Famous last words phr. 7:29
「著名遗言」;讽刺某句乐观预言很快被打脸
tiptoe around phr. 7:54
小心翼翼地行事,回避冲突
shorting /ˈʃɔːrtɪŋ/ v. 7:54
做空
doxed /dɑːkst/ v. 8:27
人肉,公开他人真实身份信息
traction /ˈtrækʃn/ n. 8:27
(产品、内容的)关注度、势头
crashing out phr. 9:11
(俚)情绪崩溃,撂挑子
boomer /ˈbuːmər/ n. 10:05
婴儿潮一代;泛指老一辈
underrated /ˌʌndərˈreɪtɪd/ adj. 10:56
被低估的
distribution /ˌdɪstrɪˈbjuːʃn/ n. 10:56
分布(统计术语,此处指年龄分布)
lithography /lɪˈθɑːɡrəfi/ n. 12:11
光刻(芯片制造核心工艺)
photo mask n. 12:11
光掩模(光刻中的图形母版)
arcane /ɑːrˈkeɪn/ adj. 12:41
晦涩的,只有少数人懂的
intersect with phr. 12:41
与……交叉、交汇
threw off phr. 13:26
打乱,使失常
broken English n. 13:26
蹩脚英语
percentage points n. 14:19
百分点
zoom back phr. 14:19
拉远视角,回到全局
point in time phr. 15:02
某一时刻的(快照式的)
outdated /ˌaʊtˈdeɪtɪd/ adj. 15:02
过时的
relentless /rɪˈlentləs/ adj. 15:41
不停歇的,无情的
living and breathing phr. 15:41
活的、持续更新的
embarked on phr. 15:41
着手,开始(一项事业)
buyin n. 15:41
(buy-in)认同与支持
aura /ˈɔːrə/ n. 15:41
气场;(网络用语)声望、面子
sweep across phr. 17:21
横扫;(测试中)遍历各种配置
suboptimal /ˌsʌbˈɑːptɪməl/ adj. 17:21
次优的
interactivity /ˌɪntərækˈtɪvəti/ n. 17:21
交互性(此处指单用户响应速度)
batch size n. 17:21
批大小(同时处理的请求数)
throughput /ˈθruːpʊt/ n. 18:07
吞吐量
downstream of phr. 18:07
处于……的下游;由……决定
latency /ˈleɪtənsi/ n. 18:07
延迟
speculative decoding n. 18:07
投机解码(用小模型预测、大模型验证以加速生成)
one-sizefits-all adj. 18:53
一刀切的,通用型的
fraal optimal adj. 19:25
(Pareto optimal 的误听)帕累托最优的
non- consensus adj. 20:18
非共识的,与主流不同的
terrestrial /təˈrestriəl/ adj. 20:18
陆地的,地球上的
humongous /hjuːˈmʌŋɡəs/ adj. 21:19
(口语)巨大无比的
incremental /ˌɪŋkrəˈmentl/ adj. 21:19
增量的,新增的
orders of magnitude phr. 22:49
数量级
juice to squeeze phr. 23:16
可榨取的余地、潜力
kernel /ˈkɜːrnl/ n. 23:16
内核;此处指 GPU 上运行的底层计算程序
experts /ˈekspɜːrts/ n. 25:05
专家(MoE 混合专家模型中的子网络)
objectively /əbˈdʒektɪvli/ adv. 25:05
客观地
collectives /kəˈlektɪvz/ n. 25:51
集合通信(多芯片间的同步操作)
arithmetic intensity n. 25:51
算术强度(每字节数据对应的计算量)
disentangle /ˌdɪsɪnˈtæŋɡl/ v. 25:51
解开,理清(纠缠的因素)
sparse /spɑːrs/ adj. 26:27
稀疏的(模型中每次仅激活部分参数)
jack of all trades phr. 27:19
万金油,样样通
leaprog /ˈliːpfrɔːɡ/ v. 27:56
(leapfrog)跳跃式越过
multiplicative /ˈmʌltɪplɪˌkeɪtɪv/ adj. 27:56
相乘的,倍增的
consumables /kənˈsuːməblz/ n. 28:41
耗材
abstraction stack n. 28:41
抽象栈(从底层硬件到高层软件的分层结构)
band-aids n. 28:41
创可贴;喻临时补救措施
acutely /əˈkjuːtli/ adv. 29:11
敏锐地,密切地
stacks /stæks/ n. 30:06
堆叠层(HBM 的芯片堆叠)
crop up phr. 31:25
冒出来,突然出现
trivially /ˈtrɪviəli/ adv. 31:25
轻而易举地
a pain in the ass phr. 32:16
(粗俗)麻烦事
thought experiment n. 33:34
思想实验
with a straight face phr. 33:34
一本正经地(说本可能引人发笑或有争议的话)
counterpoints /ˈkaʊntərpɔɪnts/ n. 33:34
对立论点,反驳点
trade-offs /ˈtreɪdɔːfs/ n. 35:10
权衡取舍
in isolation phr. 35:10
孤立地,单独地
moat /moʊt/ n. 35:10
护城河;喻企业的竞争壁垒
so be it phr. 36:04
那就这样吧,无妨
commoditized /kəˈmɑːdɪtaɪzd/ v. 36:04
被商品化(失去差异化溢价)
through and through phr. 37:51
彻头彻尾,从头到尾
lucrative /ˈluːkrətɪv/ adj. 39:12
利润丰厚的
flip side n. 39:12
反面,另一面
ASP n. 40:44
平均售价(average selling price)
facicious /fəˈsiːʃəs/ adj. 41:12
(facetious)戏谑的,不当真的
infuriates /ɪnˈfjʊrieɪts/ v. 42:01
使暴怒
plateau /plæˈtoʊ/ v. 42:01
进入平台期,停滞
saturated /ˈsætʃəreɪtɪd/ v. 42:01
饱和(基准分数触顶)
fault /fɔːlt/ v. 43:02
责怪,挑错
co-ackage optics n. 44:32
(co-packaged optics)共封装光学
analog compute n. 45:12
模拟计算
baited /ˈbeɪtɪd/ v. 45:35
(网络用语)故意挑衅引对方回应
hyperscaler /ˈhaɪpərˌskeɪlər/ n. 46:10
超大规模云厂商(亚马逊、微软、谷歌等)
bifurcation /ˌbaɪfərˈkeɪʃn/ n. 46:44
分叉,分化
local minimas n. 46:44
局部最小值(优化中并非全局最优的解)
scoot back phr. 47:40
挪回去
ASIC /ˈeɪsɪk/ n. 48:27
专用集成电路
vendors /ˈvendərz/ n. 49:38
供应商
proliferate /prəˈlɪfəreɪt/ v. 50:06
激增,扩散
carved out phr. 50:06
开辟出,切分出
by all accounts phr. 50:46
据各方所说
compute crunch n. 50:46
算力紧缺
alleviates /əˈliːvieɪts/ v. 50:46
缓解
TAM n. 51:10
可寻址市场总量(total addressable market)
step function n. 51:10
阶跃函数;喻跳跃式而非渐进的提升
stockbased compensation n. 52:02
股权激励
clawed down phr. 52:02
被拉低,被压缩
ventureback adj. 53:02
(venture-backed)风投支持的
NOI n. 53:02
净营业收入(net operating income)
go bump at night phr. 53:47
(things that go bump in the night)夜里出没的鬼怪;喻令人不安的隐患
levered /ˈlevərd/ adj. 53:47
加了杠杆的
Levered buyouts n. 54:14
杠杆收购
revenue multiples n. 54:14
收入倍数(估值方法)
drinking the Kool-Aid phr. 55:09
盲信,被自家叙事洗脑
stall /stɔːl/ v. 55:09
停滞,熄火
line of sight phr. 55:09
清晰可见的路径
ungodly /ʌnˈɡɑːdli/ adj. 55:09
(口语)多得离谱的
diluted /daɪˈluːtɪd/ v. 55:56
(股权)被稀释
lockup /ˈlɑːkʌp/ n. 55:56
锁定期(禁售期)
tanked /tæŋkt/ v. 56:28
(股价)暴跌
spiggots /ˈspɪɡəts/ n. 56:28
(spigots)水龙头;喻资金来源
tea up phr. 56:28
(tee up)铺垫,引出
contaminants /kənˈtæmɪnənts/ n. 56:28
杂质,污染物
homogeneity /ˌhoʊmədʒəˈniːəti/ n. 57:26
同质性
power smoothing n. 57:26
功率平滑
floor mechanism n. 59:18
价格下限机制
divergence /daɪˈvɜːrdʒəns/ n. 59:18
分歧,差距
collocation /ˌkoʊloʊˈkeɪʃn/ n. 59:18
(colocation)托管(只租机房与电力)
credit rating n. 60:12
信用评级
discrepancy /dɪˈskrepənsi/ n. 60:12
差异,不一致
pitfalls /ˈpɪtfɔːlz/ n. 60:41
陷阱,隐患
slosh /slɑːʃ/ v. 60:41
晃荡;此处指灵活调配(功率)
sustainably /səˈsteɪnəbli/ adv. 61:26
可持续地
balance sheet n. 63:26
资产负债表
tenant isolation n. 64:13
租户隔离
hypervisor /ˈhaɪpərˌvaɪzər/ n. 64:13
虚拟机监控程序
time splicing n. 65:01
(time slicing)时间切片
detrimental /ˌdetrɪˈmentl/ adj. 65:42
有害的
fell on their face phr. 65:42
栽了大跟头
time to market n. 65:42
上市速度
hyperlevered adj. 66:33
高杠杆的
multipolar /ˌmʌltiˈpoʊlər/ adj. 66:56
多极的
back stop v. 67:25
(backstop)兜底,托底
wild west n. 67:25
狂野西部;喻无序而机会丛生的领域
flared gas n. 68:20
伴生气(油田燃烧放空的天然气)
bubble up phr. 68:20
冒头,脱颖而出
ARR n. 69:05
年度经常性收入
out of the gate phr. 69:05
一起步就
理解自测 · 11 题
1. Dylan Patel 是怎么走上硬件研究这条路的?

直接答案是一台坏掉的 Xbox 360。第一章里他回忆八岁生日后得到的主机出现「红环死亡」故障,为了在表弟面前保持面子,他自己拆机、短接温度传感器修好了它,由此打开硬件的「潘多拉盒子」。到 12 岁他已在 Reddit 硬件、安卓、苹果等板块当版主,每天追踪英特尔、英伟达、AMD 的动向,且因家里开汽车旅馆和加油站而始终带着经济视角看技术。

2. InferenceX 为什么要做成「每天自动运行」而不是定期发布报告?

因为推理性能不是静态的。Dylan 在第四章指出,模型几乎每周都有新发布,vLLM、SGLang 等推理框架的更新频率是一周两次,新的推理优化持续涌现,因此任何「某一时刻」的测试结果发布时就已过时。他给的数据是同等质量下模型成本一年下降约 60 倍。要跟上这种节奏,基准必须「活着」:约 15 种芯片、超过 5000 万美元捐赠硬件,每天在最新开源模型上遍历配置,并公开全部结果与配置。

3. 视频中给出的 2030 年和 2040 年算力预测分别是什么?太空数据中心在其中占多少?

Dylan 预测到 2030 年,仅 OpenAI 与 Anthropic 合计推理算力就超过 100 吉瓦,加上 Meta、谷歌等会更多;到 2040 年达到太瓦级。关于太空,他的立场分两段:未来 3~5 年太空数据中心「不会真的重要」,2030 年占比不到 1%;但 20 年后大部分新增算力会上天,2040 年新增算力中可能超过一半在太空。决定因素是地面电力的建设成本与可获得的总量。

4. Dylan 用什么例子说明「协同设计」不是中国独有?

他在第七章反驳 Sean「中国在协同设计上领先」的判断时,举出 GPT-4o:其规模与 DeepSeek V3 相当且略小,发布还更早,只是 OpenAI 没有公开稀疏度和形状等细节。他的推论是:西方实验室早就在做同样的事,差别只在于是否披露。DeepSeek 的特殊之处是公开,而非首创。

5. 为什么 Dylan 认为「三层各提升 2 倍」最后能得到 100 倍而不是 8 倍?

关键在于跨层协同设计存在超线性收益。第七章里他先承认硬件(Hopper 到 Blackwell 约 30 倍)和模型层(GPT-4 到小型 Qwen)各有巨大进步,但真正的突破是把专家形状、注意力机制的算术强度、集合通信模式与芯片的矩阵乘法单元、网络拓扑一起调优,此时各层的收益不再是独立相乘,而是互相放大。他也承认这些收益事后无法拆解归因,这正是「协同」的证据。

6. 为什么 TPU 客观上是优秀的芯片,却跑不好 DeepSeek?

因为 DeepSeek 的架构是为英伟达硬件量身定制的。第七章说明 DeepSeek V3 所有专家的形状针对 Hopper 优化,V4 针对 Blackwell 和华为芯片;TPU 的矩阵乘法单元大小、无交换机的 ICI 互联拓扑与之不匹配。反过来,TPU 跑那些在英伟达上表现不佳的模型却非常好。由此引出他的结论:无法在脱离模型的情况下孤立比较芯片优劣。

7. Dylan 如何重新定义「CUDA 护城河」?

他认为传统意义上的软件护城河已被部分瓦解,因为 Claude、Codex 等模型能直接写自定义 kernel,且模型公司只有几十家。但第九章里他指出锁定效应换了形式:DeepSeek、Kimi、Qwen、小米等开源模型都针对 GPU 协同设计,下游的推理 API 服务商和 RL 定制公司只能跟着用英伟达。所以「护城河」是开源模型生态的硬件偏好,不是 CUDA 本身,谷歌用 Gemma 开源模型来对冲正是这个逻辑。

8. 「局部最小值」这个比喻在讨论专用芯片时表达了什么担忧?

担忧是专用芯片可能在当下最优,但在架构演进的终局上是错的。第十一章里 Dylan 说实验室连一年后的模型架构都不确定,若注意力机制被替换,最优硬件就会改变。TPU、Trainium、Groq、Cerebras 都可能「风光一阵然后被证明错误」。英伟达因客户多、反馈广而始终最通用,是这个不确定期的对冲工具。谷歌同时运行三个不同架构的 TPU 项目,被他视为对该风险的自觉分散。

9. 从 Anthropic 的毛利率推导,为什么它愿意以高于市场价租用 GPU?

推理链如下:第十二章给出 Anthropic 2026 年 Q2 已实现不计股权激励的盈利,Opus 4.8 的 API 单 token 毛利率超过 80%。若整体毛利率为 75%,算力成本翻倍后仍有 50% 毛利;每块新增 GPU 都能立即变成可售 token,且租来的算力扩容不需要额外人力。因此净营业收入随算力线性上升,只要有利润的实验室受速率限制约束,任何价格的算力都值得买。这也解释了它为何以高价从 SpaceX 采购。

10. 如果有人反驳「数据中心狂建就是泡沫」,Dylan 会怎样回应?

他会先把问题转化为一个可观察的条件:只要模型能完成的有用工作扩张速度快于算力增长,价格就会上涨;反之潮水转向。第十二章里他说过去六个月天平明显倒向前者,Fable/Mythos 使可寻址市场远超 Opus 的两倍而算力未翻倍。他也承认模型进展可能停滞,但指出模型正在加速下一代模型的开发,形成伪递归改进。至于杠杆,他认为对股权投资人而言高杠杆高增长意味着巨大上行,风险主要落在债权方。

11. 「每吉瓦价值不同」的论点,放到传统电力或石油行业还成立吗?

部分成立,但机制不同。视频第十三章说明数据中心已按质量分层:托管价格从每千瓦每月 60 美元涨到 120~160,信用差的客户配好机房可达 200,印度电网不稳处低至 80;谷歌能在一吉瓦机房装 1.5 吉瓦硬件并与电网签超额合约。这与 Sean 用沙特低成本高纯度石油做的类比相通:同一单位产能因质量与运营能力而价值不同。不同之处是算力层还叠加了模型毛利率差异,Dylan 明确说给 Anthropic 一吉瓦比给 OpenAI 一吉瓦收入更高,而石油下游没有这种由客户决定的价值分化。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.128CHM Live | The Silicon Gold Rush: How AI is Driving the Development of New Chips 下一期 · NO.130 →Inside Sequoia's Investment Committee | How the SpaceX & Citadel Deals Went Down | Julien Bek
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY 内容仅供学习 · thesophielab.com