视频库 / NO.127
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

He won a Nobel here for AlphaFold. Then he left. - John Jumper

节目发布 2026-06-22 · Machine Learning Street Talk
约翰·江珀 TTim Scarfe 伊曼纽尔·恩吉
本期追问 · 点击跳到视频对应位置
19:57 一个只把一类实验预测准的窄机器,凭什么算重大科学突破?34:39 蛋白质折叠的突破,靠的是把人类先验写进架构,还是靠更多数据?36:44 能算出答案却说不出理由,这样的机器让我们更懂生命了吗?36:44 预测的事实少到能写在一张卡片上,才叫真的理解吗?
归入 Ⅱ·03 能预言,就等于能解释吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2024 年,约翰·詹珀(John Jumper)因 AlphaFold 在蛋白质结构预测上的突破,与戴米斯·哈萨比斯共同获得诺贝尔化学奖;不久前他宣布离开 Google DeepMind,转投 Anthropic。本文是他在离职消息公布前接受播客 Machine Learning Street Talk 主持人蒂姆·斯卡夫(Tim Scarfe)访谈的实录,节目另穿插了非洲结构生物学家伊曼纽尔·恩吉(Emmanuel Nji)的一段自述。詹珀谈了蛋白质折叠这个半世纪难题的来龙去脉、AlphaFold 各代架构的真实归因,以及预测、控制与理解三者的分界。本文依据现场录音编译整理,仅删去口语枝节。

诺奖得主离职,与一个反命题

主持人:约翰·詹珀带领的团队做出了 AlphaFold,这个系统预测了两亿个蛋白质结构。2024 年,他获得诺贝尔化学奖。现在他要离开 DeepMind 了。AlphaFold 究竟解决了什么?还有什么没解决?它能否成为「AI 赋能科学」的范本?

詹珀:我并不喜欢人们套用「苦涩的教训」(the bitter lesson)的那种方式。事实上,AlphaFold 2 恰恰是它的反面。我们不是在告诉你一切,我们不是整个细胞的模型。你提出假设,去试,去测,十次里有九次发现自己错了。要是你十次错九次,你就是一个非常成功的机器学习研究者,你的产出高得惊人。

半世纪瓶颈:只有序列,能否算出形状

主持人:半个世纪以来,结构生物学一直卡在一个巨大的瓶颈上。DNA 容易读,蛋白质结构却不容易。蛋白质最初是一条氨基酸链,随后往往在细胞的协助下折成一个三维形状,这个形状决定它结合什么、催化什么化学反应、待在细胞的哪个位置,乃至它到底能不能工作。从机器学习的角度看,问题是:只给你序列,能不能预测折叠、预测结构?我们这个文明对世界的了解超过了此前任何文明,却一直被这一个问题困住:蛋白质是怎么折起来的?

每两年有一场大型科学实验,叫 CASP。世界各地的团队聚在一起,看能不能根据序列预测出那些刚刚做完实验、但尚未公开的蛋白质结构。几十年来进展都是渐进的,直到 2020 年,詹珀的团队拿出了远远甩开对手的结果。对许多单链目标,AlphaFold 的预测与实验结构贴得如此之近,以至于主办方宣布这个问题「基本上已经解决」。一个过去可能要专家花一年才能得到的蛋白质结构,现在几分钟就能预测出来并投入使用。

2 亿个预测结构公开,与诺奖裁定

主持人:DeepMind 本可以把成果捂在手里,但他们选择了公开:发布了一个包含两亿多个预测结构的数据库。要强调「预测」这个词,这些不是实验结果,但它们是生物学界如今大量新研究的起点。今天,AlphaFold 的用户超过三百万人,遍布一百九十多个国家。2024 年,诺贝尔委员会做出了正式裁定:化学奖一半授予大卫·贝克(David Baker)的计算蛋白质设计,另一半授予哈萨比斯和詹珀的蛋白质结构预测。AI 成了化学家的一件新工具,让他们看到此前结构生物学家做梦都想不到的分子分辨率。

詹珀:那天到十点半左右,我说,好吧,看来今年不是了。我跟妻子说了,她说:不,不,再等等。她话音未落,我的手机亮了,一个来自瑞典的电话。谢天谢地,那不是世界上最恶劣的恶作剧电话。

主持人:这一切让詹珀的下一步格外耐人寻味。几天前他宣布离开 Google,加入 Anthropic。要注意,詹珀做的不是 Claude 或 Gemini 那样的通用预测架构,而是高度结构化、为一个特定目的专门设计和打磨的系统。Anthropic 为什么对此感兴趣?我们只能猜测。

这件事也不只关乎 DeepMind 和 CASP 竞赛。世界各地的结构生物学家如今因为能用上这个数据库,可以做出新东西、做出新产品,甚至可能拯救生命。我采访了在非洲工作的结构生物学家伊曼纽尔·恩吉。

恩吉:我回去只做了一次蛋白质纯化,采集了数据,然后结合 AlphaFold,不到两三个月就拿到了结构。

主持人:也就是说,可能要几年的工作压缩到了几个月。

我们的录制在 Anthropic 的消息公布之前。前一晚我有幸和詹珀共进晚餐,算是热身,把想聊的话题都过了一遍。有一点让我印象很深:詹珀对 AlphaFold 解决了什么、没解决什么,说得异常谨慎。他不把它包装成生命的模型或治病的模型,他给它的定位窄得多,也许反而更激进:一台把某一类结构生物学测量预测得足够好、从而改变科学家下一步能做什么的机器。

蛋白质是什么:会自己装好的宜家书架

詹珀:AlphaFold 如今被称作 AI 与科学的一个里程碑,但它真正关心的是:怎么用 AI 去解决那些人做不了的、真正困难的、需要做多年实验的问题。在 AlphaFold 这里,就是蛋白质结构预测。

这是「机器学习街谈」,不是「生物学街谈」,所以我从头讲。DNA 是生命的说明书,但它到底告诉你造什么?它告诉你的众多事情之一,是怎么造蛋白质。蛋白质是细胞里的纳米机器,几千个原子,细胞的活都是它们干的。DNA 上每三个字母,指示往这个蛋白质上再加一块。蛋白质本身是一条长链,由一台同样由蛋白质和 RNA 组成的小机器一块一块串起来。你造出的是一根绳子,由二十种化学基团组成,相当于二十种字母,人们也确实用字母表来记它们。这二十种各不相同,我的博士导师能满怀深情地给你讲每一种的特别之处。绳子造好之后,它会自己组装:扭转、卷曲、折叠成紧凑而有趣的形状,有螺旋,有片层,而这个折好的形状才是真正起作用的东西。我常打的比方是:你买了一个宜家书架,打开箱子,它自己装好了。

于是生物学里有一个核心问题,存在了七十多年:我能读 DNA,现在读得非常好,你的听众里很多人可能都测过自己的基因组;可是要搞清楚哪怕一个蛋白质的结构,也极其困难。那够得上一个博士课题。典型的时间大概是一年,要按钱算,大概十万美元换一个答案。而这对生物学至关重要,因为我们想知道这些蛋白质如何工作。它们折错了,有时就是疾病;它们正常工作时,就是细胞的零件,细胞的一切功能都在其中。细胞为什么能动?是一台巨大的蛋白质机器在旋转,产生驱动细胞的力。人类基因组里大约有两万种不同的蛋白质。

科学家的办法是求助于庞大无比的机器,通常是同步辐射光源,小镇那么大,用来产生极亮的 X 射线。而即便如此,也要先做极难的实验,设法让蛋白质结晶,这一步就要花上好几年。结晶之后,再解一个数学问题(我们也许会谈到,也许不会),最后得到一张蛋白质的图像。有了这张图,常常就带来一整片理解:哦,原来人群里发现的这个 DNA 变异可能影响帕金森病,因为你看,它正好在蛋白质的这个位置,这就说得通了。

人们研究这个问题很久了。为单个蛋白质颁发的诺贝尔奖几乎数不过来,核糖体,还有很多别的。整个社会投入巨大,累计收集了大约二十万个结构,我们做 AlphaFold 时是十四万左右。每一个都仍然极其困难。我记得看博士生在快毕业时做报告,题目是「朝着结晶某某蛋白质迈进」,意思是:我要成为博士了,但这个蛋白质我大概结不出来。

我讲了半天蛋白质,还没讲我们做了什么。我们做的是,基于公开的实验数据(全部是公开数据),开发了一个新的深度学习系统,预测蛋白质结构的准确度大幅提升。典型精度达到一个原子半径以内,开始能与某些实验方法比肩。但更重要的是快得惊人:得到一个蛋白质结构只需五到十分钟,而不是一年。我应该哪天算一下这个时间比例。当然它还极易扩展。我们预测了两亿个蛋白质的结构,基本上覆盖了所有已测序物种的所有蛋白质。我们把它完全开放,科学家们用得如火如荼。

有了结构之后,生物学家怎么用

主持人:你们发布了这个数据库,地图一下子亮了。全世界的科学家都能拿这些结构去做各种下游任务。但说得具体些,蛋白质在身体里干活,我们有了结构,就可以做药物发现之类的事。中间的差距在哪里?有了结构,人们现在能做什么?

詹珀:正确的理解是,它是生物学研究的起点。看看人们做出的那些漂亮研究。最近刚出的一项,是科学家想弄明白胆固醇在体内是怎么被搬运的:到底是什么东西把胆固醇从一处运到另一处?它的突变可能怎样导致高胆固醇、心脏病?有一个非常奇特而漂亮的蛋白质把胆固醇整个包住,这个形状我们直到几个月前这篇论文发表才知道。

他们的做法很有代表性,是科学家使用 AlphaFold 最常见的方式之一:实验技术和 AlphaFold 结合。他们用冷冻电镜拍了一张非常模糊的图像。过去人们戏称冷冻电镜是「团块学」(blobology),现在好多了,但那仍然是很粗糙的图,看不出原子细节。然后他们跑 AlphaFold,发现 AlphaFold 给出的形状几乎正好嵌进那个团块里。于是既得到了验证,又得到了细节,一下子就有了一个漂亮的原子模型。接下来就可以问:这个蛋白质上的变异在哪里?会影响什么?会怎样影响它搬运胆固醇?再往后才是:怎么针对它做药?要不要去结合这个蛋白质?

从蛋白到药物:知道该拧哪颗螺丝

詹珀:在药物开发里,AlphaFold 大致有两三种用法。但我首先要说,药物开发最难的地方在于,我们对生物学如何运作了解得很不够。阻碍我们治愈比如说自闭症的,并不是「我们知道就是那一个蛋白质,只要有了它的结构病就治好了」。那是牵涉全身的复杂疾病。我们要做的,是把生物学一层层剥开、理清楚,才能知道是哪些蛋白质、它们怎么相互作用、最终怎样导致这些表型。生物学家在所有这些尺度上工作。AlphaFold 的贡献是告诉你:这个你原本都不知道重要的蛋白质,它是这样的。

举个几年前的例子。身体里有各种回收机制,把不再需要的蛋白质清除掉。人们发现在细胞发育的某个阶段,有几百个基因被关闭了,却不知道具体是哪个蛋白质在起作用。他们做了些遗传学实验,找到一个几乎从未被研究过的人类蛋白质,叫 Midnolin。把它敲低,那些蛋白质就不再被回收。他们知道的差不多就这些,另外还知道它不走标准途径。他们跑了 AlphaFold,看到一些有暗示性的片段。然后他们把 Midnolin 和大约五百个对敲低有响应的蛋白质一一配对,全部跑了一遍。这就是生物学家积累证据的方式。结果在大约百分之四十的配对里,AlphaFold 给出了同一种非常特定的模式:目标蛋白质的某一段被夹在 Midnolin 的两个部分之间,像钳子一样夹住。

然后他们去做实验:如果把 AlphaFold 说被夹住的那一段从蛋白质上去掉,会怎样?结果那个蛋白质在细胞里就不再被降解了。他们试了大约十个例子,九个完全如此,一个只是部分减弱。再回头看 AlphaFold 的预测,发现那个蛋白质有两处被夹住;把第二处也去掉,降解就彻底消失了。于是他们对一个从未想过的新蛋白质有了机制层面的理解,确切知道它在细胞分裂这个关键阶段是怎么识别底物的。

接下来的问题是:怎么把这种知识变成药?AlphaFold 2 是五年前的东西,大约一年前我们做了 AlphaFold 3,思路是:不只做蛋白质,把「蛋白质电影宇宙」全做了。我刚才说蛋白质会结合胆固醇,那是一种非蛋白质的脂类分子。同样很重要的是,它们也结合药物。药物是小分子,二三十到五十个原子,粘在蛋白质上,改变它的行为。这个问题你根本没法问 AlphaFold 2:这个药是怎么粘上去的?它会说,你得给我一个蛋白质。除非你的药本身是蛋白质,有些确实是。AlphaFold 3 把范围扩展到 PDB 里出现的整个分子世界,也就是细胞里的整个世界。现在我们可以说:这个药就粘在这里。

世界各地的人在用这些想法,也在建自己的东西。比如 Alphabet 旗下的 Isomorphic Labs,就是受 AlphaFold 突破的启发成立的,他们想真正用它来做药物设计:既然这些技术终于有了预测力,能不能用它设计一个小分子、设计一种药,让它结合上去、改变这台机器的运作方式,最终让人恢复健康。

要理解药物开发有多难,我觉得最好的比喻是一个老笑话。一家大工厂,最重要的一台机器停了。他们请来一位技师,他看了看,走到某颗螺丝或螺母跟前,拧了四分之一圈,工厂轰然重启。厂方说太好了,谢谢,请开账单。他说,一万美元。

主持人:知道该拧哪一颗。

詹珀:对,拧那一下五毛钱,知道拧哪儿是剩下的全部。我觉得这个比喻很贴切:我们既在学怎么拧,也就是怎么造药;也在学在细胞这座庞大复杂的工厂里,治病到底需要动哪里。

AlphaFold 的谦逊:只预测一个实验

主持人:可这实在太复杂了,不是吗?人体是活的,里面有一整套复杂的、适应性的、代偿性的机制在交响。这里的思路是,我们要提出一种机制性的理解,据此设计非常有效的干预。但机器学习教给我们的恰恰是相反的一课:我们关于事物如何运作的直觉基本不管用,我们需要大量数据,需要大量试错。这里会不会也是这样,像打地鼠,你动了一处,别处就补偿上来?

詹珀:从某种意义上说,非常重要的一点是 AlphaFold 的谦逊。我们说的是:我们在预测这个实验会给你什么结果。我们不是在告诉你一切,我们不是整个细胞的模型。我们是这个你一直在做、每次要花一年的实验的预测器。所以我们的有效性是有据可查的,我可以准确刻画我们复现那个实验的程度。然后人们自己想办法把这台机器用到我们没预料到的地方:发现新机制,跑几千次预测去找哪两个蛋白质会粘在一起,找出复杂系统里未知的组件。人们在不断把它推得更远。

但从某种意义上说,我们是窄的:我们预测的是一篇科学论文的结果。这类论文常常发在《自然》《科学》《细胞》这样的大刊上。我们按一下按钮就能得出《自然》级别的科学成果,但只限于一个非常窄的类别:某个特定蛋白质的结构。而生物学是一个广阔无边的宇宙,最终我们得想清楚:我们要把自己锚定在什么数据上?要预测哪些实验?并且预测得极好。机器学习的另一个故事也许就是:预测得还行,只是还行;预测得极好,就开始产生神奇的机器。语言模型、图像生成都是这样,蛋白质也是。所以我们造的不是一台包打天下的生物学机器,或者说,真要造的话,它会长得更像语言模型而不是窄预测器。但我们做的东西是真正有用的。

等变性只值 2.5 分:拆解成功的归因

主持人:我们能不能把 AlphaFold 各个版本的预测架构过一遍?第一版是 CNN,最新一版是扩散模型。第二版我们昨晚聊过,它有一个结构模块。几何深度学习被谈论得很多,我觉得人们把好处错误地归给了那些对称性,SE(3) 对称性之类。给我讲讲这个过程吧,你开头也说过,你们一开始真的想把人类的理解灌进去,后来经验告诉你们并非如此。

詹珀:有两三件事要说。首先,把 AlphaFold 3 叫「扩散模型」,我要提出异议,不是技术上的,而是主题上的。我们喜欢把东西装进盒子,喜欢用最顶层的那一个比特来解释为什么它管用。「哦,他们从 CNN 换成了……」。真实的答案是这样的:AlphaFold 1 作为一个网络,预测的是问题的一个子部分。它从生物学数据(进化相关性)出发,输出的是类似几何的数据(原子间距离),中间是一个 CNN,而且是从计算机视觉那里直接拿来的现成 CNN,别人做的。所有蛋白质特有的部分都包在机器学习外面。

AlphaFold 2 的思路是:与其借用图像识别的科学再套到蛋白质上,不如为蛋白质本身建立一门科学。「人类视觉系统正好是折叠蛋白质所需要的东西」,这句话不成立,人类过去和现在都不擅长预测蛋白质结构。那我们怎么把所有部件都造出来?

确实有一个 SE(3) 的部分。AlphaFold 2 是迭代着造出来的,分很多阶段,SE(3) 部分实际上是最早造出来的那一块。但到最后,AlphaFold 2 的主体是一个巨大的主干,我们叫它 Evoformer,是轴向注意力加上一堆别的东西,占了百分之九十以上的计算量和准确度。它产出的是一个 n 乘 n 的表示。

你从两份数据出发:一是蛋白质序列,二是所有进化上相关的蛋白质的序列。蛋白质结构变化很慢,我的蛋白质的结构在大多数情况下和酵母里的相似,有时和大肠杆菌里的也相似。你把许多相关序列抓过来,常常有几百上千条,把这些信息喂进 Evoformer。它有两路轴向注意力,相当于「我们对几何的看法」与「我们对进化的看法」之间的对话,两套表示。最后我们取几何那一路的 n 乘 n 表示,中间其实还加了一个损失,让它对原子间距离做分类预测。然后交给所谓的结构模块,最好把它理解成一个「几何化引擎」:你有关于 n 个位置的 n 平方个预测,总得有人把它们协调成一个结构。

这里用的是 SE(3) 不变的注意力(说「不变」,是因为我们每一层都把它坍缩回来)。这一块是我的想法,甚至在刚进 DeepMind 时就有了:应该把「点」放进去。我那时就在想蛋白质残基:主链有三个原子,可以给它对齐一个坐标系,它极其刚性;而所有残基彼此不同的那些部分,就挂在这个坐标系之外。如果你把它们都对齐到各自的参照系,在参照系里操作点,那就很自然。你可以拿注意力机制,让它在局部坐标系里投射出点,做变换,然后用这些点之间的距离来给注意力加偏置。我们叫它「不变点注意力」(invariant point attention,IPA)。挺好玩的。

但比这更重要的,几乎可以肯定,是「定义参照系」这件事本身。有了参照系,我们就能写下一个非常有意思的损失函数,叫「参照系对齐点误差」(frame aligned point error,FAPE)。它问的是:在第 i 个残基的参照系里,其他所有人在哪里?这是局部配准的,你得到 n 平方个误差,再取平均。我认为这才是真正关键的,是早期的突破之一。

当然,最好玩的部分是 SE(3) 不变性。但别忘了,我们一开始没有任何几何数据,只有非几何数据,几何是在中间涌现出来的。我们的初始化叫「黑洞初始化」:把所有残基堆在同一个点上,这是世界上最不物理的结构。同样重要的是,我们故意不尊重蛋白质已知的对称性。比如已知的是,残基之间是分开的,一个残基里这个原子和那个原子之间是 1.3 埃,正负 0.015。事实上在 AlphaFold 1 里做优化时,我们真的把它当成一条多关节机械臂在扭,通常三百个残基。你要做所有这些扭转,几何非常难看。三百个,不对,应该是九百个关节的机械臂,它的几何非常糟糕,意味着优化器要走很多步。所以一个重要的决定是:拆开它,把残基当成「残基气体」,这样四步、八步就能推进,而不是扭来扭去要走的那么多步。然后我们用了等变性,它有帮助。

真正令人意外的事在后面。也许是因为早期的报告,也许是因为很多人本来就在做等变性,几何深度学习那时很火,人们说:啊,他们提到了我的关键词,那一定就是它管用的原因。我记得当时有点困惑,心想:好吧,我们会非常仔细地做消融实验,论文一发,大家就明白了。论文发了,我记得第五行叫「无 IPA」。AlphaFold 2 在 GDT 尺度上比 AlphaFold 1 高了大约三十分,三十分是总账。我们做了这些消融,几乎全都很小。去掉不变性、等变性,损失大约两分,也许两分半。它有贡献,但三十分里只贡献了两分半。我以为这下可以盖棺定论了,结果根本没有。人们仍然把 AlphaFold 2 说成等变性的伟大胜利,从不谈 FAPE。他们谈「等变 Transformer」,还以为是别人发明的。其实我们在 2018 年就做了等变 Transformer,我记得是 2018 年 10 月。做了一个,试着改进了一下,发现改进那部分并不能让 AlphaFold 变好,就转向下一件事了。我们对这些东西冷酷地讲实证,但它确实很酷。所以我觉得大家是被它勾住了。

昨晚吃饭时我们也聊到这个:全局 SE(3) 对称性并不是一种很强的对称性,远不如「所有残基的置换不变性」那样大而有力。置换不变性大概才是 AlphaFold 的主对称性,我们用的是只带相对位置编码的 Transformer,而且相对位置编码是截断的。这里不像物理学,写下对称群,从对称群推出物理定律,得到标准模型。这是一个混乱的现实问题里的一种对称性,它大概并不能把问题钉死。所以它是好的,但我们不该迷恋某一个好东西,不该把它神化。

十八支二垒安打,而非一支本垒打

詹珀:我最喜欢的一条审稿意见,是 AlphaFold 2 投稿后收到的:「这是六七篇论文的想法量。」我觉得说得对。是许许多多想法加起来,才成了一个变革性的系统。用棒球来比喻,不是一两支本垒打,而是十八支二垒安打。这些中等规模的胜利叠在一起,才合成了一个变革性的系统。

消融实验里我们有时会做双重消融。我记得有一组是「无循环(recycling)加无 IPA」,同时关掉两样东西,性能就崩了。我的理解是:我们需要解决很多问题,大多数问题我们用了两种办法解决,因为两种比一种好。你把两样都敲掉,楼可能就塌了。那一组掉了大概十二到十五分,是我们最大的消融,但也只有到 AlphaFold 1 差距的一半。我记得做消融时我说:各位,我们从来没有跌破过 AlphaFold 1 的水平。

这些消融很多后来进了 AlphaFold 3。我们说,好,等变性不那么重要。另一组消融是,不给原始遗传信息、只给成对相关性,结果差了一两分,所以也许我们一直处理这些原始序列并没有那么必要。我们还做了一种可解释性的投影,把每一层投射成一个结构,做成动画,能看到 AlphaFold 的大部分容量都花在几何上优化结构。除了前几层,它更像一台几何引擎,而不是进化引擎。于是我们说:那为什么不把 Evoformer 砍到只剩几层,后面接一个简化得多的版本,叫 Pairformer?结果性能反而提高了。我们基本上是用消融实验来听机器学习在告诉我们什么。

不管你心里怎么想,做机器学习就是:你思考数据,你看数据,你提出假设。也许等变性重要?你试,你测,十次里九次发现自己错了。要是你十次错九次,你就是一个非常成功的机器学习研究者,产出高得惊人。你由此建立起局部的直觉,知道这个问题需要什么、怎么运作。在我看来,你建立的是一门局限于你这个领域的科学,一个围绕这些架构的、蛋白质结构预测的局部想法流形。我们会形成这样的戒律:不得把一维表示放在 Evoformer 附近,要用二维表示,否则性能会掉。

AlphaFold 2 做到半年或一年时,我们的成对处理里混用着轴向注意力和卷积,架构和最终版有些不同。我记得有人做了一个实验,直接删掉卷积层,不是换成注意力层,就是删掉,不加任何参数,参数严格变少,模型反而更准了,验证损失改善了。机器学习里通常不会这样,没人说「去掉参数,泛化就变好」。但卷积很可能在主动妨碍我们想学的东西,我对原因有一个假说。所有这些教训和探索,都在回答一个问题:深度学习在蛋白质上是怎么泛化的?你必须建立这种知识和专长。

架构抵得上 100 倍数据

詹珀:AlphaFold 和 AlphaFold 2 真正特别的地方,是我们把生物学假设、物理与几何假设、经验,以及多年来在这一个问题、这一个数据集上撞得头破血流所获得的手感,混在了一起。AlphaFold 1 和 AlphaFold 2 用的是完全相同的数据。我们决定只增加评估数据,训练数据一点不加。这个决定的效果非常大。AlQuraishi 实验室做过一项很好的研究,重新训练 AlphaFold 2,只用 PDB 的百分之一:不是十四五万个结构,而是一万五千个左右。他们发现,用百分之一的 PDB 训练的 AlphaFold 2 仍然比 AlphaFold 1 更准。所以可以说,我们放进 AlphaFold 2 的架构和训练思想,干干净净地抵得上一百倍的数据。

预测、控制、理解:三者的分界

主持人:昨晚我们聊到你们积累的那些默会知识,先放一放。机器学习这门事业,是为我们不理解的东西建立理解的模型,我们在造外星造物。你举过一个很好的例子:假设我在生成文本,我们可能天真地以为,边边角角上呈现出来的东西就是在干重活的东西,其实承重的可能完全是别的东西。我们的直觉大多不管用。

我知道你对「理解」这个词过敏。但对我来说,理解就是拥有一个能做出那件事的生成模型。如果你能为某个现象造一个物理模拟器,假定它不损失信息,我会说你理解了它。而机器学习是对既有样本建模,也就是说,它建模的不是东西怎样被构造出来,而是最后的结果。你昨晚说过,「生命游戏」是个很好的例子,它学的是路径,不是终点。你还讲了一件很有意思的事:AlphaFold 有循环机制,可以把东西反复送过整个结构。看起来发生的是,它一开始解决最复杂的问题,然后就是精修,再精修。这有点像生命游戏,它不是在模拟创造过程,更像在每一步学习如何精修和优化结构。

詹珀:我想先区分三件事:预测、控制、理解。预测是说,我要做一件事,我的机器会给出什么值?将来屏幕上会出现什么?控制是说,我将来要测这个东西,我要它测出来是十七。理解则很像预测,只是多了一个人在回路里。理解意味着我拥有一小组事实,用这一小组事实就能做预测,而且这些事实我能传达给另一个人,紧凑到能写在一张索引卡上。这差不多就是理解。

这些机器让我们能预测,能控制。此刻,理解还得我们自己去推导。我们现在可以在这个造物上做实验,可以去看两亿个预测结构,而不只是二十万个实验结构,以此帮助我们理解。但它不替我们完成理解这个动作。它完成的是预测,也许还有控制。

然后还有一件事,我认为是机器学习里非常重要的概念:有你编程的算法,和你得到的算法。或者说,机器学习就是代码遇上数据,产出权重。机器学习里一场持久的争论就是:多少工作是代码做的?多少工作是数据做的、最后落在权重里?我们在 AlphaFold 里看到的,从某种意义上说,是一个非常漂亮、直观的算法,一个我们在某种意义上能理解的算法:逐次几何精修。我用几句话就传达给你了,你大概脑子里已经看见了,虽然我想你没看过那些动画,它们在我们《自然》论文的补充材料里。这是人类早就想出来的算法。也许应该做梯度下降,加一个经验模型让每一步更正确,也许 AlphaFold 就该这么工作。但那不是我们编程写进去的。不过我们在某种意义上也是这么想的。我们想循环机制时,想的是:真奇怪,不管问题多难,AlphaFold 都得在 n 层之内给出答案,也许该多给它几层;可我的 GPU 内存不够了,那就把输出再送回去跑一遍,不用多占内存。但即便没有循环,我认为 AlphaFold 也在学这种迭代。我们只是放进了一个代码层面的想法、架构层面的想法,去帮助它从数据里本来就要学的那个过程。

回到刚才残基间距的例子,我们没有告诉 AlphaFold 那个距离,我们知道数据会朝它大喊:i 和 i 加 1 相隔 1.3 埃。所以谈到人类的理解,我并不喜欢人们套用「苦涩的教训」的方式。事实上 AlphaFold 2 恰恰是它的反面:我们做了一大堆专门的东西,因为我们的数据是有限的。而现在到了语言模型这里,我们发现数据仍然是有限的,互联网是有限的。所以从中该得出的结论不是「别做架构研究」,而是:对哪些东西该进代码、哪些该从数据里推导出来,保持谦逊。看看缺了什么,理解深度学习正在试图学习的那个算法,想想怎么加速它、怎么加入假设,尤其是怎么加入「通信」。我们在架构里做的最重要的事,就是修改哪些单元之间通信、怎么通信。这些都是我们把理解转化为迭代过程的方式。你要造一个精巧的几何物体,居然要迭代,这不该让任何人吃惊。

生成文本也一样。一种天真的看法是,这些模型是逐词生成器,所以完全不知道两三个词之后会发生什么。可你不可能不知道后面怎么写就写下一个词。我大多数时候开口说一句话,是知道它怎么结尾的。有时会改,但我会稍微往前想,才能完成任务。所以我们在这些模型里看到的理解、我们想要的那些结构,有时会涌现,有时不会。而我们太爱神化那些强加上去的高层想法了。

回到 AlphaFold 3。你说它是扩散模型,但我要说,它和图像扩散模型是不一样的扩散模型。首先,它有一个巨大的主干,完全不是扩散模型,只跑一次。结构大概就是在那个主干里决定的,扩散只是像当年的结构模块一样的几何化引擎:拿到一组质量相当好的约束,这些约束里已经很清楚地包含了结构,它只是把细节解出来。图像模型不同。你看那些早期的扩散模型生成图像,先出来一些彩色团块,然后才决定这些团块是什么。而且它们很明显是后来才决定的,因为你可以在过程中间停下来再跑一次,得到对同一批团块的不同解释。

AlphaFold 3 里有一件有趣的事。在 AlphaFold 2 里,通过把中间层投射出来,我们能看到它先解决什么:它先解局部细节、局部片段,再把局部片段拼起来,是聚合式的,这很自然,最容易预测的是局部结构,最难预测的是最大尺度的结构。这是 AlphaFold 2 的工作方式。而在 AlphaFold 3 里,你拿到加了大量噪声的坐标,第一件要解决的事恰恰是:比如有两个蛋白质,它们怎么结合?两个团块相对位置在哪?它们的高斯分布是什么?AlphaFold 2 最后才解决的问题,AlphaFold 3 的扩散必须最先实现。它怎么做到的?答案并不是它先想出一个取向再围绕它去搭蛋白质,因为它要的是一个正确答案,或者说一个很窄的分布。真正的答案是:前面那个大网络,加上扩散网络的第一次前向,就已经解决了整体结构;扩散后面做的,是把之前没解出来的细节实现出来,本质上是在其中采样。所以技术上它是扩散,但它离 AlphaFold 2 近得多。选扩散只是出于很具体的技术原因,说白了是关于几何的偷懒:它让配体更好处理,让一些键长和局部细节更好处理。它不是「先画团块、最后决定团块是什么」意义上的扩散。

所以人们喜欢说「这个管用是因为它是 Transformer」。可「因为它是 Transformer」解释不了为什么聊天模型在过去三四年里进步了这么多,解释不了那些研究,解释不了研究者每天在做什么。所有这些细节,都远比「是不是 Transformer、是不是扩散模型」这个我们爱谈的高层比特重要得多。而且,即便是扩散机制,也并不按图像里那种「渐进精修」的方式工作。对图像来说,也许先画团块再决定团块的含义,这个说法都未必是全部的故事;对蛋白质来说它肯定不是,因为最难的问题就是大尺度结构。

通用还是定制:智能与表征之争

主持人:这在某种意义上通向我们之前谈的「建构性复杂度」。我很想听听你对通用人工智能的总体看法。以语言模型为例,我们训练它基本上是行为克隆:有一个丰富的、适应性的生成过程,产生了所有这些语言,我们拿它们训练模型。对我来说,智能是对粗粒度表征的适应性获取。文化和语言一直在变,我们发明 unalive 这个词来绕过社交平台的过滤器,这是语言能动性的一个例子。而我们注意到,当我们对语言模型做迭代的、适应性的精修,做主动微调和适应,它们就变得有点智能了:学到新表征,会适应。某种意义上,即便它们脱离了生成路径,也能拿一个代码解,比如 AlphaEvolve,不断精修,效果非常好。你觉得我们现在造的这些造物,是不是并不具备同一类型的通用性?你对智能总体怎么看?

詹珀:表征这个问题非常重要,但远不如五年前人们所相信的那样重要,至少显式意义上不是。就像我们刚才谈的,AlphaFold 做的事情里,有些是代码强迫它做的,凡是代码强迫的它当然都做了;但它做的很多事并不是被强迫的,是因为要对数据建立一个好的预测模型,它就必须学会,必须找到好的中间表征。

从某种意义上说,机器学习里最诱人的想法永远是:我知道最终一定得有这么个东西,所以我要在代码里专门留个位置,用这个名字命名,强迫机制去实现这个「高层概念构造器」。这曾经很流行:我要让概念成为单元,我要通过这个损失在中间层强制解耦表征,它就存在这里。去试试这个想法是合理的。但我们大量看到的是,你以为智能所需要的那些东西,很多是靠拼命把下一个词预测得极好而发展出来的。它们不是因为你预测下一个词而产生,而是因为你在这件事上做得极好才产生。这些泛化的空间、表征、对概念的理解,是被数据非常缓慢地逼出来的。所有那些对数线性关系,大家最不喜欢的那种函数形式:性能随努力的指数线性上升,我们在缩放律里天天看到。我们确实得到了概念和表征,但我们还没有答案的是:怎么更便宜地得到它们?

有时可以靠编程。我们得到了类似记忆的东西:现在语言模型会给自己写笔记,再把笔记取回来;我们发现最好不断提醒智能体它在做什么,免得它在长轨迹里忘掉。所以我们训练权重,造出造物,发现缺陷,然后往往能用某种软件外壳把缺陷糊上。但这还不能直接回流到机器学习本身。你没法给它套一个外部记忆的外壳,再把它蒸馏回网络里,得到一个不再需要外壳的、记忆惊人的模型。这一步,我们还没弄明白怎么做。

主持人:很遗憾,约翰,我们得收尾了。詹珀博士,很荣幸请到你,非常感谢。

詹珀:非常愉快,谢谢。

非洲的一千名科学家

主持人:前面说过,伊曼纽尔·恩吉在非洲工作,他在培训科学家。不只是给他们 AlphaFold 的使用权限,而是教他们怎么用、怎么解读结果、怎么用这个数据库去设计实验。

恩吉:我的研究方向是疟疾和肠道细菌的药物发现,同时也参与非洲本地研究者的能力建设,用的就是 AlphaFold 这样的工具。以前非洲的科学家用不上昂贵的结构生物学设备。有了 AlphaFold,这些研究者现在能做过去不可能的复杂实验,去对付疟疾、艾滋病以及各种抗生素耐药的感染。在我自己的药物发现研究里,我用 AlphaFold 来解冷冻电镜数据的结构,也用它来梳理蛋白质的作用机制。

主持人:对他来说,AlphaFold 的冲击是前后判若两个世界。

恩吉:当时要解一个蛋白质的结构,真的非常非常难。我试了好几年,将近四五年,都没成功。那是十多年前了。有了 AlphaFold 之后,我回去只做了一次蛋白质纯化,采集数据,结合 AlphaFold,不到两三个月就拿到了结构。

主持人:现在他的目标是尽可能多地培训科学家使用这项技术,造福人类。

恩吉:今年在 Google DeepMind 和瑞典研究理事会的资助下,我们把培训规模扩大到一百人,培训质量没有下降,反而有所提高。基于这个结果,我们打算在接下来十年里每年培训一百名科学家,目标是十年内让近一千名非洲科学家能有效使用这个工具,并形成一个新兴的结构生物学实践者社区,专注于非洲高发的疾病。

主持人:这就是本期 AlphaFold 特辑。感谢约翰和伊曼纽尔。和约翰的对谈非常有意思,他很鼓舞人,因为他本身就是一个证明:尽管我们谈论各种通用基础模型,但要真正推进科学的前沿、做出尖端应用,还是需要大量的工程、领域知识和扎实的专长。我们的很多模型最终会是混合的、高度定制的。AlphaFold 就是这类可以用来推进科学的混合模型的一个存在性证明。约翰,祝你在 Anthropic 的新岗位一切顺利。感谢收看。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 诺奖得主离职,与一个反命题 ▶ 正在看
0:36 半世纪瓶颈:只有序列,能否算出形状 ▶ 正在看
3:01 2 亿个预测结构公开,与诺奖裁定 ▶ 正在看
7:36 蛋白质是什么:会自己装好的宜家书架 ▶ 正在看
12:21 有了结构之后,生物学家怎么用 ▶ 正在看
17:13 从蛋白到药物:知道该拧哪颗螺丝 ▶ 正在看
19:57 AlphaFold 的谦逊:只预测一个实验 ▶ 正在看
22:24 等变性只值 2.5 分:拆解成功的归因 ▶ 正在看
30:38 十八支二垒安打,而非一支本垒打 ▶ 正在看
34:39 架构抵得上 100 倍数据 ▶ 正在看
36:44 预测、控制、理解:三者的分界 ▶ 正在看
45:12 通用还是定制:智能与表征之争 ▶ 正在看
49:17 非洲的一千名科学家 ▶ 正在看
本期小问 · 档案清单
19:57 一个只把一类实验预测准的窄机器,凭什么算重大科学突破? ▶ 正在看
34:39 蛋白质折叠的突破,靠的是把人类先验写进架构,还是靠更多数据? ▶ 正在看
36:44 能算出答案却说不出理由,这样的机器让我们更懂生命了吗? ▶ 正在看
36:44 预测的事实少到能写在一张卡片上,才叫真的理解吗? ▶ 正在看
本期讲者
约翰·江珀Google DeepMind 前研究总监,AlphaFold 团队负责人,2024 年因蛋白质结构预测与 Demis Hassabis 分享诺贝尔化学奖一半。近期宣布离开 DeepMind 加入 Anthropic。
Tim Scarfe机器学习播客 Machine Learning Street Talk(MLST)主持人,长期与深度学习研究者讨论智能、表征与泛化等基础问题。
伊曼纽尔·恩吉结构生物学家,在非洲从事疟疾与肠道细菌的药物发现研究,同时开展本地研究者的结构生物学能力建设培训。
01诺奖得主离职,与一个反命题
0:00
I don't really love the bitter lesson as people try and apply it. In fact, AlphaFold 2 is the opposite of that. Protein folding is 1 of these holy grail type problems in biology. We predict nature level science with the press of a button in a very narrow category of nature level science of the structure of a specific protein. John Jumper led the team behind AlphaFold, the system that predicted 200,000,000 protein structures. In 2024, he won the Nobel Prize for chemistry. And now Jumper is leaving DeepMind.
我并不太喜欢人们套用“苦涩的教训”的那种方式。事实上,AlphaFold 2 恰恰相反。蛋白质折叠是生物学中那种“圣杯”级别的难题。我们只需按一个按钮,就能做出自然级别的科学预测,当然是在非常狭窄的一类自然级科学里——也就是某个特定蛋白质的结构。John Jumper 领导了 AlphaFold 背后的团队,这个系统预测了 2 亿个蛋白质结构。2024 年,他获得了诺贝尔化学奖。而现在,Jumper 要离开 DeepMind 了。
便签笔记
02半世纪瓶颈:只有序列,能否算出形状
0:36
But what did AlphaFold solve? What remains unsolved? And could AlphaFold be the template for AI for science? We are not trying to tell you everything. We are not a model of the entire cell. You try it. You measure. 9 times out of 10, you find out you're wrong. Right? If you're wrong 9 times out of 10, you're a very successful machine learner. You're incredibly productive. So for half a century, structural biology had a massive bottleneck. DNA was easy to read, but protein structures were not. A protein structure begins as a chain of amino acids and then often with help from the cell they settle into a three-dimensional shape and that shape determines what it binds, what chemistry it catalyzes, where it sits in the cell and whether it even works at all. But from a machine learning perspective, if you only have the sequence, can you predict the fold?
但 AlphaFold 到底解决了什么?还有什么仍未被解决?AlphaFold 能否成为“AI for science”的范本?我们并不想告诉你一切。我们不是整个细胞的模型。你去试,你去测。十次里有九次,你会发现自己错了。对吧?如果你十次里错九次,那你已经是个非常成功的机器学习研究者了。你的产出高得惊人。半个世纪以来,结构生物学一直有一个巨大的瓶颈。DNA 很容易读取,但蛋白质结构不行。蛋白质结构一开始是一条氨基酸链,然后往往在细胞的帮助下,折叠成一个三维形状,而这个形状决定了它结合什么、催化什么化学反应、位于细胞的什么位置,以及它究竟能不能正常工作。但从机器学习的角度看,如果你只有序列,你能预测出折叠方式吗?
便签笔记
1:42
Can you predict the structure? We've discovered more about the world than any other civilization before us. But we have been stuck on this 1 problem: how do proteins fold up? So every 2 years, there's a big scientific experiment called CASP. Essentially, teams from around the world gather to see if they can predict protein structure from sequences based on recently done but not yet publicly available experiments. So for many decades the progress was incremental until 2020 when John Jumper's team, AlphaFold, they produced a result which was significantly better than the competition.
你能预测出结构吗?我们对世界的认识,已经超过了此前任何一个文明。但我们一直卡在这一个问题上:蛋白质是怎么折叠起来的?所以每两年都会有一场大型科学实验,叫做CASP。本质上,就是全世界的团队聚在一起,看谁能根据序列预测出蛋白质结构,依据的是那些最近已完成但尚未公开的实验数据。几十年来,进展一直是渐进式的,直到 2020 年,John Jumper 的团队带着 AlphaFold,拿出了一个远远优于其他竞争者的结果。
便签笔记
2:25
For many single chain targets the predictions from AlphaFold were so close to the targets that the organizers of the event said that the problem had been essentially solved. So a protein structure that might have taken a year of specialist work can now be predicted and operationalized in minutes. Everyone, the whole team. It's been an incredible effort. Congratulations on this work. It is really outstanding. AlphaFold represents a huge leap forward that I hope will really accelerate drug discovery and help us to better understand disease.
对许多单链目标来说,AlphaFold 的预测与真实结构如此接近,以至于赛事的组织者宣布,这个问题基本上被解决了。所以,一个原本可能需要专业人员做上一年的蛋白质结构,如今几分钟内就能预测出来并投入实际使用。各位,整个团队,这真是一次了不起的努力。祝贺你们做出这项成果。真的非常出色。AlphaFold 代表着一次巨大的飞跃,我希望它能真正加速药物研发,并帮助我们更好地理解疾病。
便签笔记
032 亿个预测结构公开,与诺奖裁定
3:01
It's mind blowing. You know, these results were, for me, having worked on this problem so long after many, many stops and starts and will this ever get there, suddenly, this is a solution. We'd solve the problem. And fair play to DeepMind, so they could have kept this close to their chest, but they decided to release it. They released a database with over 200,000,000 predicted protein structures. Now it's important to emphasise the word predicted. These are not experiments but these are the basis for a lot of interesting new search work that the world of biology can now perform.
这令人震撼。你知道,对我来说,我在这个问题上做了那么久,经历了无数次的中断和重启,心想这到底能不能做成,然后突然之间,答案就摆在那里了。我们把这个问题解决了。而且要给 DeepMind 点个赞,他们本可以把成果捂在自己手里,但他们决定公开发布。他们发布了一个包含超过 2 亿个预测蛋白质结构的数据库。这里要强调一下“预测”这个词。这些不是实验结果,但它们是生物学界如今能够开展的许多有趣新研究的基础。
便签笔记
3:38
And today AlphaFold is now used by more than 3,000,000 people in over a 190 countries. So in 2024 the Nobel Prize Committee made a formal verdict. Half of the chemistry prize went to David Baker for computational protein design and the other half went to Demis Hassabis and John Jumper for protein structure prediction. So AI had become a new tool for chemists, a way of seeing molecular structure at a level of resolution that structural biologists couldn't even have dreamed of before. So it's it's absolutely wonderful to be here.
如今 AlphaFold 已经被 190 多个国家的 300 多万人使用。所以在 2024 年,诺贝尔奖委员会给出了正式的裁定。化学奖的一半授予了 David Baker,表彰他在计算蛋白质设计上的贡献,另一半授予了 Demis Hassabis 和 John Jumper,表彰他们在蛋白质结构预测上的工作。所以 AI 已经成为化学家的一种新工具,一种观察分子结构的方式,而这种分辨率是结构生物学家过去做梦都不敢想的。能来到这里真是太好了。
便签笔记
4:16
It's it's truly an extraordinary honor to tell you about this work, to tell you about the work of our team in protein structure prediction. So by about 10:30, I said, oh, well, I guess not this year. And I told my wife, and she goes, no. No. Wait. And as, like, as she's telling me to wait, my phone lights up with a phone call from Sweden. And thankfully, it was not the world's meanest prank call. And And all of this makes John's next move very interesting. So just a few days ago, he announced his departure at Google, and he's going to Anthropic. It's important to note that John Jumper was not building generic prediction architectures like Claude or like Gemini. It was extremely structured, designed and engineered for the purpose of doing a specific specific thing.
能向大家介绍这项工作、介绍我们团队在蛋白质结构预测方面的工作,真是莫大的荣幸。所以到了大概十点半,我就说,唉,看来今年是没戏了。我跟我太太这么说,她说,不。不,再等等。就在她让我再等等的时候,我的手机亮了,是一通来自瑞典的电话。谢天谢地,那不是世界上最缺德的恶作剧电话。而这一切让 John 的下一步动向变得非常有意思。就在几天前,他宣布从谷歌离职,要去 Anthropic。这里要说明的是,John Jumper 当时并不是在构建像 Claude 或 Gemini 那样的通用预测架构。它是高度结构化的,是为了完成某一件非常具体的事情而专门设计和打造的。
便签笔记
5:08
Why might that be interesting to Anthropic? We can only guess. So this isn't just about Google DeepMind and, you know, the CASP competition and and whatnot. There are now structural biologists around the world that can innovate and build new products and save lives potentially because they have access to this protein database. I spoke with Emmanuel Nji. He is a structural biologist based in Africa. I went back and did just 1 protein purification, collected the data, and we used AlphaFold. In combination, I got the structure in less than 2, 3 months.
那这为什么会让 Anthropic 感兴趣呢?我们只能猜测。所以这不只是关于 Google DeepMind、CASP 竞赛之类的事。现在全世界的结构生物学家都可以去创新、开发新产品,甚至可能拯救生命,因为他们能够用上这个蛋白质数据库。我和 Emmanuel Nji 聊过。他是一位在非洲工作的结构生物学家。我回去只做了一次蛋白质纯化,收集了数据,然后我们用了 AlphaFold。两者结合,我在不到两三个月的时间里就拿到了结构。
便签笔记
5:46
So yeah, we're talking about potentially years of work compressed into several months. That's pretty good. Agents are getting smarter every day but even the best agents get stuck without good context. And this is where Notion comes in. With the recent launch of custom agents Notion becomes the collaborative AI platform where agents and humans work side by side. And now their new developer platform is turning that into infrastructure that developers can work on. The way I think about it is it is the kind of curated materialized memory plane for all of my work.
所以是的,我们说的是原本可能要好几年的工作被压缩到了几个月。这相当不错。智能体每天都在变得更聪明,但即使是最好的智能体,没有好的上下文也会卡住。这正是 Notion 发挥作用的地方。随着最近推出自定义智能体,Notion 成了一个让智能体和人类并肩工作的协作式 AI 平台。而现在他们新的开发者平台正把这一切变成开发者可以在其上构建的基础设施。我是这么理解的:它就像是我所有工作的一个精心整理、具象化的记忆层。
便签笔记
6:25
And the great thing about Notion is it's so easy to work with programmatically. It has a CLI. It has an MCP server. It has agents built into it. Right? So that means using my phone, I can ask an agent to go and do some research or to put some information in there, I can sync things to my calendar. Without Notion, I honestly don't think I'd be able to do anything that I do on MLST. So it's really really good. So I highly recommend you give Notion's platform a go. You can sign up at notion.com/mlst and you'll be supporting the show if you do. And now let's get back to John Jumper.
Notion 最棒的一点是,用编程方式操作它非常方便。它有命令行工具。它有 MCP 服务器。它内置了智能体。对吧?所以这意味着我用手机就能让智能体去做点调研,或者把一些信息放进去,我还能把内容同步到日历。说实话,没有 Notion,我觉得我在 MLST 上做的这些事一件都做不成。所以它真的非常好。我强烈推荐你去试试 Notion 的平台。你可以在 notion.com/mlst 注册,这样也是在支持我们这个节目。现在让我们回到 John Jumper。
便签笔记
7:00
Now we filmed this before the Anthropic announcement so it was really cool to speak with John. He's such an inspirational guy. I was lucky enough to have dinner with him the night So we had a bit of a warm up conversation and, you know, we drilled in to the various topics we wanted to discuss. 1 thing that struck me is John is unusually careful about what AlphaFold does and does not solve. So he doesn't sell it as a model of life or, you know, as a model of curing disease. He sells it as something a little bit more narrow and possibly more radical actually, a machine that predicts 1 class of structural biology measurement well enough to change what scientists can do next.
我们是在 Anthropic 官宣之前录的这期节目,所以能和 John 聊天真的很棒。他是个特别能给人启发的人。我很幸运在前一天晚上和他共进了晚餐,所以我们先做了一点热身式的交流,也深入探讨了我们想讨论的各种话题。让我印象很深的一点是,John 对于 AlphaFold 解决了什么、又没有解决什么,格外谨慎。所以他不会把它包装成生命的模型,或者治愈疾病的模型。他把它说成是一个更窄一点、但可能也更激进的东西:一台能把某一类结构生物学测量预测得足够准的机器,准到足以改变科学家接下来能做什么。
便签笔记
04蛋白质是什么:会自己装好的宜家书架
7:36
So now I give you, John Jumper. I mean, AlphaFold itself is this kind of, I guess, I now say landmark in AI and science, but it's really about how do we use AI to solve problems that humans can't, that are really hard, that we go and we do years long experiments. And in the case of AlphaFold, it's this problem of protein structure prediction. This I guess it's machine learning street talk, not biology street talk. So, you know, DNA is the instruction manual for life, but what does it actually tell you what to build?
那么现在,有请 John Jumper。我是说,AlphaFold 本身可以算是 AI 与科学结合的一个里程碑,但它真正关乎的是我们如何用 AI 去解决人类解决不了的、非常困难的问题,那些需要做长达数年实验的问题。就 AlphaFold 而言,就是蛋白质结构预测这个问题。我想这里是“机器学习街谈”,不是“生物学街谈”。所以,DNA 是生命的说明书,但它究竟告诉你要造出什么呢?
便签笔记
8:09
And it tells you 1 of the many things it tells you how to build are proteins. And these are little nanomachines, couple thousand atoms in the cell that actually do the work of the cell. And so 3 letters of your DNA tell you how to add 1 extra piece to this protein. This protein is kind of a long string, and there's a tiny machine itself made of proteins and RNA that's built 1 kind of string at a time. And you you make this rope of 20 types of chemical groups. So it's kind of 20 types of letters, and people, of course, use the alphabet for these things. And each of them are quite different.
它告诉你的众多东西之一,就是怎么造蛋白质。这些是小小的纳米机器,细胞里由几千个原子组成,真正干着细胞里的活。而你 DNA 中的三个字母告诉你如何给这个蛋白质再加上一个组件。这个蛋白质有点像一条长链,而有一台本身由蛋白质和 RNA 构成的小机器,一次一段地把这条链造出来。于是你就造出了一条由 20 种化学基团组成的“绳子”。所以就像是有 20 种字母,人们当然也就用字母表来表示它们。而且每一种都很不一样。
便签笔记
8:43
Right? My my PhD supervisor could tell you lovingly about what's special of each 1. But you build this kind of rope of the protein, and then it assembles itself. It twists. It curls. It folds up into a really kind of compact and interesting shape. It has these helices, sheets, all these things, and that's actually what works. And the the analogy I always kind of like to say is it's like you have an IKEA bookshelf, and you open the box, and it builds itself. And so this really, really central problem for maybe 70 plus years in in biology is, okay.
对吧?我的博士导师能满怀热情地跟你讲每一种的特别之处。但你把蛋白质这条绳子造出来之后,它就会自己组装。它会扭转。它会卷曲。它会折叠成一个非常紧凑而有趣的形状。它有螺旋、有折叠片层,各种结构,而这才是真正起作用的东西。我一直喜欢用的类比是:就像你买了一个宜家书架,打开箱子,它自己就装好了。所以生物学里这个存在了大概 70 多年的、非常非常核心的问题就是,好吧。
便签笔记
9:17
How do I I can read DNA. In fact, can read DNA really well now. You know, you can probably many people in your in your listeners have had their DNA sequenced. But understanding the structure of even 1 protein is extraordinarily difficult. Right? That's a worthy PhD project. I would say maybe a typical kind of time frame is a year. If you wanna put money on it, maybe a $100,000 to get 1 answer. And this is really important to biology because we wanna understand how these proteins work. When they misfold, sometimes it's disease.
我该怎么……我能读 DNA。事实上,现在读 DNA 已经读得非常好了。你知道,你的听众里可能有很多人都测过自己的 DNA。但哪怕只是搞清楚一个蛋白质的结构,都极其困难。对吧?那足以撑起一个博士课题。我觉得典型的时间跨度大概是一年。如果要折算成钱,大概要花十万美元才能得到一个答案。这对生物学非常重要,因为我们想搞清楚这些蛋白质是怎么工作的。当它们错误折叠时,有时就会导致疾病。
便签笔记
9:48
Even when they work, they are the parts of the cell. They do all the parts, the things of the cell. You know, they're beautiful proteins. The reason that, you know, cells can move. Right? Or is this giant protein machine whirling around driving the force to move cells? All of the functions of the cell, basically, are in these proteins. Humans have about 20,000 different types in different locations in in your genome. And so what scientists have done is they've gone to these enormous, enormous machines, synchrotrons normally, you know, the size of small towns, in order to produce extraordinarily bright X rays.
即便它们正常工作,它们也就是细胞的零件。细胞里所有的部件、所有的活儿,都是它们在做。你知道,它们是很美的蛋白质。细胞之所以能移动,原因就在这里。对吧?不就是这台巨大的蛋白质机器在旋转,产生驱动细胞移动的力量吗?细胞的所有功能,基本上都在这些蛋白质里。人类的基因组里,不同位置上大约有两万种不同的蛋白质。所以科学家们的做法是,他们跑到那些巨大的、巨大的机器那里,通常是同步辐射光源,差不多有小镇那么大,为的是产生极其明亮的 X 射线。
便签笔记
10:25
And even that, only after they've done really, really hard experiments trying to figure out how to what's called crystallize a protein, you know, and this takes years and years. And once they do that, and then they solve another mathematical problem that maybe we'll talk about, maybe we won't, they get 1 picture of a protein. And they get kind of this progress and often this whole wealth of understanding of, oh, okay. I can understand how this DNA change that was found in the population might affect Parkinson's because, look, it's right here on this protein, and that makes so much more sense. And so people have studied this problem for a long time.
而且即便如此,也得先做完非常非常艰难的实验,想办法把蛋白质“结晶”出来,而这要花上好多好多年。等他们做到这一步,再解决另一个数学问题——这个我们可能会聊,也可能不聊——他们才拿到一张蛋白质的图像。然后他们就取得了这样的进展,往往还收获了一大堆理解,哦,原来如此。我就能明白在人群中发现的这个 DNA 变化为什么可能影响帕金森病,因为你看,它就在这个蛋白质的这个位置上,这样就说得通多了。所以人们研究这个问题已经很久了。
便签笔记
10:59
There have been almost innumerable Nobel Prizes given for individual proteins, right, the ribosome, many others. People have an incredible societal investment, collected around 200,000 of these about a 140,000 at the time we did AlphaFold. Each 1 still extraordinarily difficult. Right? Each 1 still. I remember seeing people talk about their PhD and give their 1 of their talks near the end of their PhD, progress toward crystallizing whatever protein. Right? So I did my I'm gonna be doctor, and I probably am not gonna crystallize this protein. I guess I'm I'm telling you all about proteins and nothing about what we did. But what we did was develop a new deep learning system from the publicly available experimental data, so all very public data, that was vastly more accurate at predicting protein structures.
为单个蛋白质颁发的诺贝尔奖几乎数不胜数,对吧,比如核糖体,还有很多其他的。整个社会为此投入了难以置信的资源,一共收集了大约 20 万个这样的结构,在我们做AlphaFold 的时候大约是 14 万个。而每一个都仍然极其困难。对吧?每一个都还是这样。我记得看到有人讲他们的博士研究,在博士快结束时做报告,题目是“某某蛋白质结晶的进展”。对吧?就是说我做完了博士,我就要成为博士了,但我可能压根结晶不出这个蛋白质。我好像一直在跟你讲蛋白质,却完全没讲我们做了什么。我们做的是,基于公开可获得的实验数据——全都是非常公开的数据——开发了一套新的深度学习系统,它在预测蛋白质结构上要准确得多。
便签笔记
11:48
So predicts it to something like within the radius of an atom, right, in in typical accuracy, and an accuracy that starts to rival at least some experimental methods, but more importantly than that, is, you know, extraordinarily fast. So it takes 5, 10 minutes to get the structure of a protein instead of a year. I should at some point figure out what that ratio is in terms of time. But then also, of course, it's incredibly scalable. So we've predicted the structure of 200,000,000 proteins, basically every protein from an organism whose genome has been sequenced.
在典型精度下,它能预测到一个原子半径以内,这个精度已经开始可以与至少某些实验方法相媲美了,但比这更重要的是,它极其快。得到一个蛋白质的结构只要五到十分钟,而不是一年。我什么时候真该算算这个时间比到底是多少。当然,而且它的可扩展性也非常强。所以我们预测了 2 亿个蛋白质的结构,基本上涵盖了所有已完成基因组测序的生物体的每一个蛋白质。
便签笔记
05有了结构之后,生物学家怎么用
12:21
Right? We've made this widely available, and scientists are using it like crazy. It's absolutely amazing. You have released a database of all of these proteins and the map lit up. So now scientists from all around the world, they can access these protein structures for many downstream tasks. But to bring this to life, you know, we have proteins doing things in the body and we can use these structures, and we could do things like drug discovery and and and whatnot. But what what's the gap? So so what can people do now that they have these structures?
对吧?我们把它广泛开放出来,科学家们正疯狂地在用。太不可思议了。你们发布了一个包含所有这些蛋白质的数据库,整张地图都亮了起来。所以现在全世界的科学家都能拿到这些蛋白质结构,用于各种下游任务。但要把这些讲得具体一点,蛋白质在体内做各种事情,我们可以利用这些结构,去做药物研发之类的事情。但差距在哪里呢?那么,人们有了这些结构之后,现在能做些什么?
便签笔记
12:48
I think the right way to think about this is it's a starting point for biological research. If you think about what what people do, what are some beautiful studies that people have done, you know, we we see it all the way. 1 that just came out was scientists trying to understand how cholesterol is moved about in the body. Right? What what actually is the thing that takes cholesterol and moves it from 1 place to another? How might mutations in that affect high cholesterol, heart disease, etcetera?
我觉得正确的看法是,它是生物学研究的一个起点。如果你想想人们都在做什么,有哪些漂亮的研究已经做出来了,你知道,我们一路上都能看到。刚刚发表的一项研究是,科学家们想弄清楚胆固醇在体内是怎么被运输的。对吧?到底是什么东西把胆固醇从一个地方搬到另一个地方?其中的突变又可能如何影响高胆固醇、心脏病等等?
便签笔记
13:22
There's this beautiful weird protein that kind of wraps around it in a shape that we really didn't know until a few months ago when this paper came out. And what they were able to do actually is a it's 1 of the ways in which scientists, I think, really commonly use alpha fold is they use both experimental techniques and AlphaFold. So they used an experimental technique, cryo electron microscopy, to take an incredibly blobby picture. You know, they used to call cryo EM blobology. It's gotten much better, but it's an incredibly kind of rough picture.
有一个既漂亮又奇怪的蛋白质,会以某种形状把它包裹起来,而这个形状我们一直不知道,直到几个月前这篇论文发表。他们所做到的,其实正体现了科学家们我认为非常常见的一种使用 AlphaFold 的方式:他们同时使用实验技术和 AlphaFold。所以他们用了一种实验技术,冷冻电子显微镜,拍出了一张极其模糊的团状图像。你知道,人们过去把冷冻电镜叫做“团块学”。已经好多了,但那还是一张非常粗糙的图像。
便签笔记
13:55
And they don't really know the atomic details. And then they also run AlphaFold, and they see, well, actually, AlphaFold has this shape that almost exactly fits within this kind of blob. And so you get both confirmation and more detail, and suddenly you have this beautiful atomic model where you can start to then go and say, now how where are the changes in this protein? What might it affect? How might it affect how it takes cholesterol from 1 place to another? Then you have to go figure out. Now how do I make drugs for this?
而且他们并不真正了解原子层面的细节。然后他们也跑一下 AlphaFold,结果发现,AlphaFold 给出的这个形状几乎正好能塞进那团模糊的密度里。于是你既得到了验证,又得到了更多细节,突然之间你就有了一个漂亮的原子模型,可以开始问:那么,这个蛋白里发生变化的地方到底在哪儿?它可能影响什么?它会怎样影响这个蛋白把胆固醇从一个地方运到另一个地方?接下来你就得去搞清楚。那我该怎么针对它做药?
便签笔记
14:23
Do I I bind to this protein? I think when you see it in drug development, there's actually kind of 2 or 3 ways in which it's used. I mean, the 1st I will say is that the hardest part of drug development is that we do not know how biology works very well, right? That it's not, you know the thing preventing us from curing, say, I guess, autism, right, is not that we know exactly 1 protein. If we just had its structure, then autism would be cured. It's a huge disease that involves the whole body. And so we kind of are trying to unwrap and unravel biology well enough to figure out which proteins. How do these proteins interact?
我要不要结合到这个蛋白上?我觉得在药物研发里,它大概有两三种用法。首先我想说的是,药物研发最难的地方在于,我们对生物学的运作方式了解得并不好,对吧?比如说,阻碍我们治愈——我想想,比如自闭症——的原因,并不是我们已经确切知道某一个蛋白。好像只要拿到它的结构,自闭症就能治好了。那是一种涉及全身的、极其复杂的疾病。所以我们其实是在努力一层层剥开生物学,搞清楚是哪些蛋白参与其中,这些蛋白之间又怎么相互作用。
便签笔记
15:04
How does that ultimately contribute out to these phenotypes? And so people do biology across all these length scales. And the contribution of AlphaFold is to say, this protein, for example, that you didn't even know was important. Like, there was oh, 1 study from a few years ago was on a pro was you know, there's all sorts of recycling mechanisms in the body that take proteins it doesn't need anymore and gets rid of them or doesn't want anymore. There were hundreds of genes, in fact, that people found were turned off at a certain phase in cell development, they didn't exactly know what protein was involved.
这些又是如何最终体现为这些表型的?所以人们在各种尺度上研究生物学。而 AlphaFold 的贡献在于,它能告诉你:比如这个蛋白,你以前甚至都不知道它很重要。比如说,几年前有一项研究,是关于一个……你知道的,体内有各种各样的回收机制,把不再需要或不想要的蛋白清除掉。实际上,人们发现有几百个基因在细胞发育的某个阶段被关闭了,但他们并不确切知道是哪个蛋白参与其中。
便签笔记
15:39
They did some genetics, and they found this protein that had essentially never been studied before, a human protein called Midnolin. If you knocked it down, then these proteins didn't get recycled. And that's kind of more or less what they knew. And they knew it didn't work in the standard way. And they ran AlphaFold, and they looked at it. And they saw some pieces that were suggestive. And then they ran AlphaFold together with, you know, I think it was almost all 500 proteins that were not that were kind of responsive to knocking this protein down.
他们做了一些遗传学实验,找到了一个此前基本没人研究过的蛋白,一个叫 Midnolin 的人类蛋白。如果你把它敲低,这些蛋白就不会被回收掉。他们知道的差不多就这些。他们还知道它的作用方式不走常规路径。于是他们跑了 AlphaFold,仔细看了看结果。他们看到了一些很有提示性的片段。然后他们把 AlphaFold 和大约 500 个蛋白一起跑——我记得几乎是全部那些对敲低这个蛋白有反应的蛋白。
便签笔记
16:07
Right? So change. So this is kind of how biologists develop evidence. And they found in about 40 percent of these when they ran AlphaFold this very, very specific pattern where 1 part of that protein was trapped between 2 parts of Midnolin kind of grabbing it like clamps. And they could find and then they would go and they would do experiments. Right? Because and then they would say, well, what happens if I take this bit of protein and I remove the place where AlphaFold says it's clamped by Midnolin?
对吧?也就是有变化的那些。生物学家大致就是这样一步步建立证据的。他们发现,在跑 AlphaFold 的这些蛋白里,大约百分之四十都出现了一个非常、非常特定的模式:那个蛋白的某一段被 Midnolin 的两个部分夹住,像钳子一样把它抓住。他们能找到这个模式,然后就去做实验。对吧?他们会问:如果我拿这段蛋白,把 AlphaFold 说被 Midnolin 夹住的那个位置去掉,会发生什么?
便签笔记
16:35
And suddenly that protein doesn't drop down in the cell. Right? So and they found this on maybe the 10 examples. 9 of them worked exactly this way. 1 of them only partially was reduced in how much it's knocked down. But then they looked at the AlphaFold predictions and found that AlphaFold actually put it 2 places. And so if they take out that 2nd place as well, then the degradation is completely abolished. So now they have this mechanistic understanding of this new protein they had never thought about before, and now they know exactly how it recognizes what's developed in this really important stage of cell division.
结果那个蛋白在细胞里的水平就不再下降了。对吧?他们大概在十个例子上做了验证。其中九个完全是这样的机制。还有一个,被敲低的程度只是部分减弱。但他们再去看 AlphaFold 的预测,发现 AlphaFold 其实给出了两个结合位点。所以如果把第二个位点也去掉,降解就被彻底消除了。于是现在他们对这个以前从没想到过的新蛋白有了机制层面的理解,也确切知道了它是怎么识别在细胞分裂这个非常关键阶段产生的那些东西的。
便签笔记
06从蛋白到药物:知道该拧哪颗螺丝
17:13
So the so and now the question becomes, okay. Now how do you take that knowledge and do drug development? And that's where so AlphaFold 2 was what came out now 5 years ago. What we've done more recently about a year ago is AlphaFold 3, which says, well, let's let's not just do proteins. Let's do the protein cinematic universe. And so, you know, I said proteins bind cholesterol. Right? So this is a non protein kind of fatty molecule. More well, not more importantly, but very importantly, they also bind drugs.
那么现在问题就来了,好吧。你要怎么把这些知识拿去做药物研发?这就说到——AlphaFold 2 是五年前发布的。我们最近,大约一年前做的是 AlphaFold 3,它的想法是:咱们别只做蛋白质了。咱们把整个“蛋白质电影宇宙”都做了。比如说,我前面提到蛋白会结合胆固醇。对吧?那是一种非蛋白的脂类分子。更……嗯,不能说更重要,但同样非常重要的是,它们也会结合药物。
便签笔记
17:44
Drugs are small molecules, maybe, you know, 20, 50 atoms that stick to proteins and change how they behave. And you couldn't even ask this question to AlphaFold 2. You couldn't say, how does this drug stick? It'd say, well, you better give me a protein. Only if your drug is a protein, which some are. But AlphaFold 3, we expanded it to kind of do the whole universe of things that appear in the PDB, the whole universe of cells. And now we can say, well, this is exactly where that drug sticks. And then people around the world are using these ideas, building others.
药物是小分子,可能就二三十到五十个原子,它们黏附在蛋白上,改变蛋白的行为方式。而这个问题你根本没法拿去问 AlphaFold 2。你没法问它:这个药物是怎么结合上去的?它会说:你最好给我一个蛋白。除非你的药本身就是蛋白——有些确实是。但到了 AlphaFold 3,我们把它扩展到能处理 PDB 里出现的几乎所有种类的分子,也就是细胞里的整个世界。现在我们就能说:药物结合的位置正是在这里。然后世界各地的人都在用这些思路,去构建其他的东西。
便签笔记
18:13
For example, isomorphic labs inside Alphabet, kinda developed, inspired from the AlphaFold breakthrough are trying to say, okay. Let's really use this to start to do drug design. Let's start to take these technologies that finally work, that are finally predictive about this, and now let's see if I can design a small molecule with it, or I can design a drug that binds, that changes how this machine works, and then in a way that hopefully makes someone healthy. And I think the the best kind of analogy for how you should think about drug development really now or maybe the way to think about how hard it is is this old joke. You know?
比如 Alphabet 旗下的 Isomorphic Labs,某种程度上就是受 AlphaFold 突破的启发而成立的,他们想做的是:好,我们真正用这个来开始做药物设计。我们把这些终于能用、终于具备预测能力的技术拿过来,看看能不能用它设计一个小分子,或者设计一种能够结合、能够改变这台机器运作方式的药物,并且希望以某种方式让人恢复健康。我觉得,要理解今天的药物研发,或者说要理解它有多难,最好的类比是一个老笑话。你知道吧?
便签笔记
18:51
Do you know this joke that there's this giant factory, and 1 of the most important machines in this factory has stopped working. And they call in a technician who comes. He looks at it. He goes to some screw or some some nut and turns it a quarter turn. The factory roars back to life. And they said, that's wonderful. Thank you so much. Can we have a bill? And he says, $10,000. Yeah. And they say, what Knowing what to turn. Yeah. Yeah. So there's, you know, knowing what to turn or turning this 50¢, knowing what to turn, all the rest.
你听过这个笑话吗——有一家巨大的工厂,厂里最重要的机器之一停了。他们请来一位技师。他看了看,走到某个螺丝或者某个螺母那儿,拧了四分之一圈。工厂轰隆隆地重新运转起来。他们说,太棒了。太感谢你了。能给我们开个账单吗?他说:一万美元。是啊。他们说:什么?——知道该拧哪儿。对对。所以你看,知道该拧哪儿,拧这一下值五毛钱,知道该拧哪儿,值剩下的那些。
便签笔记
19:25
And I think this is the right analogy that we are both learning how to turn this, right, how to build drugs, and also learning in this big complex factory of the cell what do we need to do to cure disease. Yes. But it's so incredibly complex, isn't it? Because the human body is alive and there is a symphony of complex adaptive compensatory mechanisms. I guess the the idea here is that we're proposing a mechanistic understanding of how this works, which means we can design interventions that are very effective.
我觉得这个类比很贴切:我们既在学怎么去“拧”,也就是怎么造药,同时也在这个庞大复杂的细胞工厂里学习,到底要做什么才能治好疾病。是的。但这实在太复杂了,不是吗?因为人体是活的,里面有一整套复杂的自适应代偿机制在协同运作。我想这里的思路是,我们提出了一种机制层面的理解,也就意味着我们可以设计出非常有效的干预手段。
便签笔记
07AlphaFold 的谦逊:只预测一个实验
19:57
But in machine learning, we've kind of learned the opposite lesson, which is that all of our intuitions about how things work don't don't really work, we need lots of data, and we need to test lots of things. Could it be a similar thing here that, you know, it's like whack a mole. You you kind of you you you do 1 thing, and then something else compensates. I think really important in a certain sense is almost the humility of AlphaFold. In that, you know, people say, you know, we are trying to predict what this experiment will give you.
但在机器学习里,我们学到的教训恰恰相反:我们对事物运作方式的所有直觉其实都不太靠谱,我们需要大量数据,需要做大量试验。这里会不会也是类似的情况——就像打地鼠一样?你动了一个地方,结果别的地方就代偿回来了。我觉得从某种意义上说,AlphaFold 的“谦逊”反而非常重要。也就是说,我们说的是:我们想预测的是这个实验会给你什么结果。
便签笔记
20:25
We are not trying to tell you everything. We are not a model of the entire cell. We are a predictor of this experiment that you did all the time and took you a year. And so in a certain sense, I think and so we have validity in that I can characterize very well how well we're we we will reproduce that experiment. And then people figure out how to take this machine and use it in other ways that we didn't expect to find out new you know, discover new mechanisms, to try thousands of alpha fold predictions, to find 2 proteins that stick together, and find this unknown component of this complex system. So people are finding ways to push this further.
我们不是想告诉你一切。我们不是整个细胞的模型。我们是对那个你天天做、要做上一年的实验的预测器。所以从某种意义上说,我们的有效性在于我能非常清楚地刻画我们复现那个实验的程度有多好。然后人们会自己想办法,把这台机器用在我们没预料到的方向上,去发现新的……你知道的,去发现新机制,去跑上千次 AlphaFold 预测,找出哪两个蛋白会结合在一起,找出这个复杂体系里某个未知的组分。所以人们在不断把它往前推。
便签笔记
21:10
But in a certain sense, we are narrow, or we predict the result of a scientific paper. We predict the result of a scientific paper that often appears in Nature and Science and Cell and these big journals. Right? We predict nature level science with the press of a button in a very narrow category of nature level science of the structure of a specific protein. But there's this enormous wide universe of biology that ultimately we're gonna have to figure out and understand what data will we pin ourselves to, what experiments will we predict, and predict really, really well such that you know, I mean, maybe the other story of machine learning is that predicting things okay is alright.
但从某种意义上说,我们是很窄的,我们预测的是一篇科研论文的结果。我们预测的是那种常常发表在《自然》《科学》《细胞》这些顶级期刊上的论文结果。对吧?我们按一下按钮就能预测出《自然》级别的科学成果,但只是在“某个特定蛋白的结构”这一个很窄的门类里。但生物学还有一个极其广阔的世界,我们最终得搞清楚:我们要把自己锚定在什么数据上,要预测什么实验,而且要预测得非常非常好。因为——机器学习的另一个故事可能是,把事情预测得还凑合,那也就还凑合。
便签笔记
21:51
Predicting things extraordinarily well starts to produce amazing machines. We see this, of course, in language models, in image generation, but also in protein. So I think this this kind of thing, we aren't building just 1 universal biology machine, or at least if we do, it will have to look a lot more like a language model than it will kind of a narrow predictor. But we are doing something truly useful. Can we talk through the predictive architectures of of the different versions of AlphaFold? So, you know, the 1st version was was a CNN.
而把事情预测得极其准确,就开始造出令人惊叹的机器了。我们在语言模型、图像生成里都看到了这一点,在蛋白质上也是。所以我觉得,我们并不是在造一台通用的生物学机器;或者说就算要造,它也得更像一个语言模型,而不是一个很窄的预测器。但我们确实在做一件真正有用的事。我们能不能过一遍 AlphaFold 各个版本的预测架构?比如第一版是 CNN。
便签笔记
08等变性只值 2.5 分:拆解成功的归因
22:24
The last version is a diffusion model. The 2nd version, we spoke about this last night, it had a structure component. And, you know, obviously, geometric deep learning is spoken about a lot. And I think people misattributed the the benefit of having these, you know, and kind of symmetries. It did these SC 3 symmetries. And just just talk me through that process because you were kind of saying at the beginning, you were really trying to imbue your human understanding of this and then kind of experience told you differently.
最新一版是扩散模型。第二版——我们昨晚聊过——里面有一个结构模块。而且你知道,几何深度学习被讨论得非常多。我觉得人们把这些东西——那些对称性——带来的好处归因错了。它用了 SE(3) 对称性。你能不能讲讲这个过程?因为你一开始说过,你当时很想把自己对这件事的人类理解注入进去,然后经验告诉你并非如此。
便签笔记
22:51
I think there's 2 or 3 things. I mean, 1st, I would almost object not technically, but kind of thematically to AlphaFold 3 as a diffusion model. We love to stick things in boxes. We love to have the highest level bit be the answer for why these things work. Yes. Right? Oh, they switched from CNN to I think the answer is really okay. AlphaFold 1 really was, as a network, it predicted a subpart of the problem. It started from kind of biological data, evolutionary correlations. It ended in kind of geometric ish data, distance between atoms. In between was a CNN.
我觉得有两三点。首先,我几乎想反对——不是从技术上,而是从叙事框架上——把 AlphaFold 3 说成“一个扩散模型”。我们太喜欢给东西贴标签、装进盒子了。我们太喜欢用最高层的那个标签来解释这些东西为什么有效。是的,对吧?“哦,他们从 CNN 换成了……”我觉得真正的答案是这样:AlphaFold 1 作为一个网络,其实只预测了问题的一小部分。它从某种生物学数据出发,也就是进化上的共变关系,最后输出的是某种几何类的数据,即原子之间的距离。中间那段是 CNN。
便签笔记
23:27
Right? It was actually an off the shelf CNN from a computer vision that someone else had done. Okay. That was a CNN. But after AlphaFold 1 and then kind of all the protein specific bits were kind of wrapped around the machine learning. And so the I would say alpha fold 2 was, let's build the science instead of building the science of image recognition and then applying it to proteins because, you know, human visual the human visual system is exactly what we needed to fold proteins is not something true.
对吧?其实就是一个现成的、别人做的计算机视觉里的 CNN。好,那部分是 CNN。但在 AlphaFold 1 之后,所有蛋白质特有的部分基本上都是包裹在机器学习外面的。我想说,AlphaFold 2 的思路是:我们来构建这门科学本身,而不是先构建图像识别的科学,然后把它套用到蛋白质上。因为你知道,人类的视觉系统恰好就是我们折叠蛋白质所需要的东西——这种说法并不成立。
便签笔记
23:57
Right? Humans were bad, are bad at at predicting protein structures. How are we going to actually build all the pieces? Now there was an s e 3 piece. In fact, AlphaFold 2 was built iteratively. There were many stages. Actually, the s e 3 piece was the 1st part of AlphaFold built. But AlphaFold 2 at the end was really this giant trunk of an architecture we called Evoformer, which is axial attention plus a bunch of other stuff. And that is 90 plus percent of the compute and the accuracy. And then but it produces this kind of n by n.
对吧?人类过去不擅长、现在也不擅长预测蛋白质结构。我们到底要怎么把所有这些部件搭建起来?当时确实有一个 SE(3) 的部分。事实上,AlphaFold 2 是迭代式地构建出来的。经历了很多个阶段。实际上,SE(3) 那部分是 AlphaFold 最先постро建成的部分。但最终的 AlphaFold 2 真正的核心,是一个巨大的架构主干,我们叫它 Evoformer,也就是轴向注意力(axial attention)加上一堆其他东西。这部分占了 90% 以上的算力,也贡献了 90% 以上的精度。然后它会产出这样一个 n×n 的东西。
便签笔记
24:34
So you start off with 2 pieces of data. You start off with the protein sequence, and then you find the sequence of every protein evolutionarily related. And protein structure changes slowly. The the structures of my proteins are in most cases similar to the structure of proteins in yeast, sometimes even out in E. Coli. So you grab many related structures. You often have hundreds or thousands. You provide this information, and we have this this specialty architecture called Evoformer, which had 2 forms of axial attention that were kind of having a conversation between what we believed about geometry and what we believed about evolution. We had these 2 representations.
所以你一开始有两部分数据。你从蛋白质序列出发,然后去找出所有进化上相关的蛋白质序列。而蛋白质结构变化得很慢。我身上这些蛋白质的结构,在大多数情况下和酵母里蛋白质的结构是相似的,有时候甚至和大肠杆菌里的都相似。所以你会抓取很多相关的结构。往往能拿到几百甚至几千个。你把这些信息提供进去,我们有一个专门的架构叫 Evoformer,它包含两种形式的轴向注意力,某种意义上是在让我们对几何的认知和我们对进化的认知之间进行一场对话。我们有这两种表示。
便签笔记
25:09
And we end up we take the geometric bit, the n by n, which we have actually as an intermediate loss said, what are the distances between these atoms and made categorical predictions. And then we hand it to what we call the structure module, and it's best thought of as a geometrization engine. Right? If you you have n squared predictions about n positions, you're somebody's gonna have to harmonize this thing. And this used a s e 3 I guess it was invariant in the sense we collapsed it on every layer.
最后我们取出几何的那一部分,也就是那个 n×n,我们实际上把它作为一个中间损失,去问:这些原子之间的距离是多少?并做出分类式的预测。然后我们把它交给我们所说的结构模块(structure module),最好把它理解成一个「几何化引擎」。对吧?如果你对 n 个位置有 n² 个预测,总得有人来把这些东西调和一致。这里用到了 SE(3)——我想应该说是不变性(invariant),因为我们在每一层都把它坍缩掉了。
便签笔记
25:39
S e 3 invariant attention. This was actually 1. This was 1 of mine that was kind of even starting at DeepMine. I'm like, oh, we should probably put points in. Or I was thinking about pro even then, protein residues. Right? So you have this backbone, which has 3 atoms. You can align a frame to it. It's extraordinarily rigid. And I know the business in the places where all these at where all these residues differ is kind of just off that. So if you align them to reference frames then and you operate in points in those reference frames, then it's natural.
SE(3) 不变注意力。这其实是其中一个……这其实是我自己的点子之一,甚至在刚到 DeepMind 的时候就开始想了。我当时想,哦,我们大概应该把「点」放进去。或者说,我那时候想的其实是蛋白质残基。对吧?你有这个主链骨架,它有三个原子。你可以给它对齐一个坐标框架(frame)。它极其刚性。而且我知道,所有这些残基彼此不同的地方,基本上就是从那个框架延伸出去的。所以如果你把它们对齐到参考框架上,然后在这些参考框架里操作「点」,那就非常自然了。
便签笔记
26:10
And you can just take attention, and you can let it project points in its local frame. You can transform it. Then you can use the distance of those points as a way in order to bias your attention. And this is, invariant point attention is what we called it. At the end, it's kind of fun. More important than that probably was this defining of frames. Actually, almost certainly. Defining of frames let us write down a really interesting loss function. So we call it frame align, point error, or FAPE. And this was kind of saying, in the reference frame of the I th residue, where is everyone else?
然后你就可以直接用注意力,让它在自己的局部框架里投射出点。你可以对它做变换。接着你可以用这些点之间的距离来给你的注意力加偏置。这就是我们所说的「不变点注意力」(invariant point attention)。最后来看,这挺有意思的。但可能比这更重要的,是这种「框架」的定义方式。实际上几乎可以肯定是这样。框架的定义让我们能写下一个非常有意思的损失函数。我们把它叫做「框架对齐点误差」,即 FAPE。它的意思大致是:在第 i 个残基的参考框架下,其他所有残基都在哪儿?
便签笔记
26:46
And it's kind of locally registered, and then you have n squared errors, and then you average them together. That, I think, was really, really important. I think that was 1 of the breakthroughs early on was this loss. But, of course, the really fun part is an SE 3 invariance. And but remember, we didn't start with any geometric data. We started only with nongeometric data. So our geometry emerged in the middle. We started with what we would call black hole initialization, where we just stick all the residues on top of each other in the world's least physical structure, it was important also that we disrespected known symmetries of a protein.
它是局部配准的,然后你会得到 n² 个误差,再把它们平均起来。我认为这一点真的非常非常重要。我觉得早期的突破之一就是这个损失函数。当然,真正有意思的部分是 SE(3) 不变性。但要记住,我们一开始并没有任何几何数据。我们一开始只有非几何的数据。所以我们的几何是在中间过程里涌现出来的。我们从所谓的「黑洞初始化」开始,就是把所有残基全都叠在同一个位置上,形成世界上最不符合物理的结构。同样重要的是,我们刻意不去遵守蛋白质已知的对称性/约束。
便签笔记
27:21
For example, the known symmetries of a protein are these residues are separated. The atoms this atom and that atom in a residue are separated by 1.3 angstroms plus or minus 0.015. And in fact, even in AlphaFold 1, when we would do the optimization, we would actually use turning kind of like a jointed robot arm to optimize these, say, typically 300 residues. And so you would do all this twisting. You would actually have a very ugly geometry. The geometry of a 300 joint or actually, no. Sorry. Wouldn't be 300.
举例来说,蛋白质已知的这些约束是:这些残基是分开的。一个残基里的这个原子和那个原子之间相距 1.3 埃,上下浮动 0.015。事实上,即使在 AlphaFold 1 里,我们做优化的时候,实际上是像操作一条带关节的机械臂那样,去转动优化这些——比如说典型情况下的 300 个残基。所以你得做各种扭转。你实际上会得到非常难看的几何结构。一条 300 关节的……其实不对。抱歉,不是 300。
便签笔记
27:51
It would be a 900 joint robot arm is really bad, and so that means that your optimizer has to take many steps. So 1 of the important things is let's just break it up. Let's just treat them as a residue gas, we called it, so that this can proceed in, say, 4 steps, 8 steps instead of the number of steps of this twisty geometry. And then we used equivariance, and it helped. But 1 of the things that was really surprising, think, is maybe because of the early talk or maybe people were working on equivariance.
应该是一条 900 关节的机械臂,这非常糟糕,也就意味着你的优化器必须走很多步。所以其中一件重要的事情就是:干脆把它拆开。我们干脆把它们当作一团「残基气体」——我们是这么叫的——这样一来整个过程就能顺利进行,比如说 4 步、8 步,而不是这种扭来扭去的几何结构所需要的步数。然后我们用了等变性(equivariance),它确实有帮助。但我觉得有一件事真的很让人意外,可能是因为早年的那些报告,也可能是因为当时很多人本来就在做等变性的研究。
便签笔记
28:19
Geometric deep learning has been very popular, and people said, ah, they mentioned my keyword. That must be the reason it worked. And I remember being a little bit confused, and I thought, okay. But we'll we'll very carefully ablate this. We did quite a few ablations for AlphaFold 2. And we'll publish the paper, everyone will realize. And we published a paper, and I remember the 5th row was called no IPA. So AlphaFold 2 was about 30 points on the GDT scale better than AlphaFold 1. Right? So so that's that's the kind of 30 points is is your thing.
几何深度学习那时候非常火,大家一看,啊,他们提到了我这个关键词,那肯定就是它管用的原因。我记得我当时还有点困惑,心想,好吧。但我们会非常仔细地对这一点做消融实验。我们为 AlphaFold 2 做了相当多的消融实验。然后我们把论文发出去,大家就都会明白了。我们把论文发了出去,我记得第 5 行叫做「no IPA」(去掉不变点注意力)。AlphaFold 2 在 GDT 指标上比 AlphaFold 1 高了大约 30 分,对吧?所以这 30 分就是你要去解释的那个东西。
便签笔记
28:51
And we did these ablations. They were all small, almost all small. Removing the invariance, the equivariance cost about 2 points. Right? And you could measure it, maybe 2 and a half. Right? So it it contributed, but it contributed 2 and a half out of 30. And I thought that would put it to bed, and it put it didn't even put it to bed at all. People still talked about AlphaFold 2 as the great victory of equivariance. They they never talk about FAPE. And they talk about, you know, an equivariant transformer that they think came from others. And, actually, we did this equivariant transformer in, like, 2018. I actually remember it was October 2018.
我们做了这些消融实验,结果差异都很小,几乎全都很小。去掉不变性、等变性,大概只损失 2 分。对吧?你是能测出来的,可能是 2 分半。对吧?所以它确实有贡献,但 30 分里它只贡献了 2 分半。我本以为这下总该把这事儿说清楚了,结果根本没有。大家还是把 AlphaFold 2 说成是等变性的伟大胜利。他们从来不谈 FAPE。他们谈的是那个等变 Transformer,还以为是别人提出来的。而实际上,我们做那个等变 Transformer 是在,大概是……2018 年。我其实还记得是 2018 年 10 月。
便签笔记
29:28
So it was like, we did 1. We tried a little bit to improve it. It didn't make AlphaFold better to try and improve that part, so we went on to the next thing. Right? We're kind of ruthlessly empirical about it, but it's a very cool thing. And so I think people really hooked on to it. And I think what it really happens is I think we we were talking about it kind of at dinner last night at this outfall dinner, but the equivariance is 1 like, global s c 3 symmetry is not a very powerful symmetry. It's not nearly as kind of big and powerful as a symmetry like, oh, all the residues are permutation invariant.
所以就是说,我们做了一版。我们也稍微试着改进过它。但改进那一部分并没有让 AlphaFold 变得更好,所以我们就转去做下一件事了。对吧?我们在这方面是相当冷酷地讲究实证的,但它确实是个很酷的东西。所以我觉得大家真的就抓着这一点不放了。我觉得实际情况是——昨晚在那个 outfall 晚宴上我们其实也聊到过——等变性,比如全局的 SE(3) 对称性,其实并不是一个很强大的对称性。它远远比不上那种真正又大又强的对称性,比如说,所有残基之间是置换不变的。
便签笔记
30:04
Right? So we do still have permutation invariance as probably the big symmetry of alpha fold. Right? We have a transformer that is relative position coded only. We clip the relative position coatings. But I think this particular symmetry, it's not like physics where you write down the symmetry group and then you derive the laws of physics from your big symmetry group and you get the standard model. This is this is asymmetry of a messy real world problem that probably doesn't pin it down so much. So I think it's good, but we shouldn't obsess about 1 good thing.
对吧?所以我们确实还保留了置换不变性,那大概才是 AlphaFold 里最重要的对称性。对吧?我们的 Transformer 只用了相对位置编码。而且我们还会对相对位置编码做截断。但我觉得这个特定的对称性,它不像物理里那样:你写下对称群,然后从这个大对称群推导出物理定律,最后得到标准模型。然后你就得到标准模型。这里的对称性是一个乱糟糟的现实世界问题里的对称性,它大概并不能把答案框死。所以我觉得它是好东西,但我们不该只盯着某一件好东西不放。
便签笔记
09十八支二垒安打,而非一支本垒打
30:38
Or you can't you don't wanna valorize things. My favorite review of of AlphaFold 2, we got the reviews back when we submit the paper. And 1 of them said, this is 6 or 7 papers worth of ideas. Right? And I think I think that was that was right. There are many, many ideas that added up to be a transformative system. And many, you know, to use a baseball analogy, it's not 1 or 2 home runs. It's, you know, 18 doubles. Right? That it it's really, you know, these midsize wins stacked together and together make a transformative system.
或者说,你不该把某一件事捧得太高。我最喜欢的一条关于 AlphaFold 2 的审稿意见——我们投稿后收到审稿意见,其中一条说:这篇文章里的想法够写六七篇论文了。对吧?我觉得这个说法是对的。有非常非常多的想法叠加在一起,才造就了一个变革性的系统。而且,用一个棒球的比喻来说,这不是靠一两支本垒打,而是靠十八支二垒安打。对吧?真的就是这些中等规模的进步叠加起来,合在一起才成就了一个变革性的系统。
便签笔记
31:11
Now we would sometimes find in our ablations, we ran a double ablation. I think it was no recycling and no IPA. We turned off 2 things, and performance cratered. Right? And I think it was kind of there are many problems we needed to solve. We solved most of them 2 ways because it was better than solving 1. And if you knock out both things, then your building maybe collapses. Or this was maybe a 12 or 15, which was our biggest ablation, which was still only half the gap to AlphaFold 1. Right? We I remember doing the ablations and saying, guys, we've never crossed alpha fold 1 performance.
不过在消融实验里我们有时候会发现,我们做过一次双重消融。我记得是同时去掉循环(recycling)和 IPA。我们把两样东西一起关掉,性能就崩了。对吧?我觉得大概是因为,我们要解决的问题很多。其中大部分问题我们都用两种方式去解决,因为这比只用一种方式更好。所以如果你把两样都敲掉,你的楼可能就塌了。要么就是大概 12 分或 15 分,那是我们最大的一次消融,但即便如此也只占到与 AlphaFold 1 差距的一半。对吧?我记得做完消融实验后我说,各位,我们从来没有掉回到 AlphaFold 1 的水平。
便签笔记
31:44
But a lot of those ablations actually went into alpha fold 3. So we said, okay. Well, equivariance isn't super important. We had another ablation, that if we take out, you know, giving the raw genetic information and give the pairwise correlations, that's 1 or 2 worse. So maybe this fact that we're processing these all the time is not so good. We looked at we made this kind of interpretability kind of projection of each layer into a structure and made movies and could see that most of AlphaFold's capacity was spent optimizing the structure geometrically.
但那些消融实验的很多结论后来都进了 AlphaFold 3。所以我们说,好吧,等变性并不是那么重要。我们还做过另一个消融:如果不给原始的基因序列信息,而只给成对的相关性统计,结果只差 1 到 2 分。所以也许我们一直在处理这些原始信息这件事,并没有那么好。我们还做了一种可解释性的分析,把每一层都投影成一个结构,做成动画来看,结果发现 AlphaFold 的容量绝大部分都花在从几何上优化结构上了。
便签笔记
32:16
It's much more a geometry engine than an evolution engine outside the 1st few layers. So we said, okay. Why don't we just cut back this Evoformer to just operate a few layers, and then we'll do a much simpler version called a Pairformer, and that improved performance. And we basically used these ablations to say this is what the machine learning is telling us. Whatever we may, you know, feel like, being a machine learner is all about, you know, you think about the data. You look at it. You come up with hypotheses.
除了最前面几层之外,它更像是一个几何引擎,而不是一个进化信息引擎。所以我们说,好吧。那我们干脆把 Evoformer 砍到只保留几层,然后做一个简单得多的版本,叫 Pairformer,结果性能反而更好了。我们基本上就是用这些消融实验来判断:这就是机器学习告诉我们的结论。不管我们自己心里怎么想,做机器学习无非就是:你思考数据。你去看数据,然后提出假设。
便签笔记
32:45
Maybe equivariance is important. You try it. You measure. 9 times out of 10, you find out you're wrong. Right? If you're wrong 9 times out of 10, you're a very successful machine learner. You're incredibly productive. And and you build this local intuition. You build this notion of what the problem needs and how it works, and you build, in my view, a kind of science local to your area, kind of a local manifold of ideas in protein structure prediction near these architectures. We would develop intuitions like, thou shalt not put a 1 d representation rather than a 2 d representation anywhere near the Evoformer or your performance will go down.
也许等变性很重要。你去试,然后你去测量。十次里有九次,你会发现自己错了。对吧?如果你十次里错九次,那你就是个非常成功的机器学习研究者了。你的产出效率高得惊人。然后你就建立起了这种局部的直觉。你会形成一种认识:这个问题需要什么、它是怎么运作的,然后你就建立起了,在我看来,一种属于你自己领域的局部科学,一种围绕这些架构的、蛋白质结构预测中的局部想法流形。我们会形成这样的直觉,比如:切不可把一维表示(而不是二维表示)放在Evoformer 附近的任何地方,否则你的性能就会下降。
便签笔记
33:26
Maybe a year, 6 months into AlphaFold 2, at some point, we had a mix of axial attention and convolutions in our pairwise processing. It was somewhat different architecture. And I remember someone did an experiment where they just deleted the convolutional layers, not like added attention layers. Deleted the convolutions, added no parameters, just strictly fewer parameters, and the model got more accurate. The validation loss improved. And that doesn't normally happen in machine learning. They don't say remove parameters and your generalization will improve.
大概是 AlphaFold 2 做了一年、或者说半年的时候,某个阶段我们在成对特征的处理里混用了轴向注意力和卷积。那时的架构和现在有些不一样。我记得有人做了个实验,他们只是把卷积层删掉了,并没有去加注意力层。删掉卷积,不加任何参数,参数量严格变少了,结果模型反而更准了。验证集损失变好了。这种事在机器学习里通常是不会发生的。没人会说,你把参数删掉,泛化能力反而会变好。
便签笔记
33:58
But convolutions were probably actively harmful to learning what we wanted to learn, and I have kind of a hypothesis of what that is. But all of these kind of lessons and explorations were about how does deep learning generalize in proteins? And you have to build that knowledge and that expertise. And what's really, really special in AlphaFold, in AlphaFold 2, and then I'll I guess I should talk about AlphaFold 3 in a minute, is that we have kind of the mix of biological hypothesis, physical geometric hypothesis, and experience, and the kind of, you know, tactile feel of years of banging our head against this particular problem and data set, actually.
但卷积可能对我们想要学习的东西是有实实在在的害处的,我大概有个猜想,能解释这是为什么。但所有这些经验教训和探索,都是围绕着一个问题:深度学习在蛋白质领域是怎么泛化的?你必须一点点积累出那些知识和专业经验。AlphaFold 真正特别的地方,AlphaFold 2 里——我想我待会儿应该也讲讲 AlphaFold 3——就在于我们把几样东西糅合在了一起:生物学假设、物理和几何上的假设、经验,还有那种,你知道的,多年来在这个特定问题和数据集上一次次碰壁磨出来的手感。
便签笔记
10架构抵得上 100 倍数据
34:39
AlphaFold 1 and AlphaFold 2 were the exact same data. We decided to just have more eval data and not bump our training data at all. And and the effect of that was really, really large. There was a a wonderful study by the Al Qureshi lab, which was retraining AlphaFold twos and train them on 1% of the PDB. So instead, you know, so instead of a 100 a 150,000 structures, it was something like 15,000. And they found AlphaFold 2 on 1% of the PDB was more accurate than AlphaFold 1. So you can really say that the the architecture and training ideas that we put into AlphaFold 2 were worth a clean 100 x in data. We were speaking last night about all of the tacit knowledge that you guys acquired during, but maybe we'll park that just for a 2nd.
AlphaFold 1 和 AlphaFold 2 用的是完全相同的数据。我们决定只增加评测数据,训练数据一点都不加。结果这么做的效果非常非常显著。AlQuraishi 实验室做过一项非常出色的研究,他们重新训练 AlphaFold 2,只用 PDB 的 1% 来训练。所以,你知道,不是十万、十五万个结构,而是大概一万五千个。他们发现,只用 1% PDB 训练的 AlphaFold 2 比 AlphaFold 1 还要准。所以你真的可以说,我们放进 AlphaFold 2 里的架构和训练思路,实打实抵得上 100 倍的数据量。昨晚我们聊到了你们在这个过程中积累的所有那些隐性知识,不过这个咱们先放一放。
便签笔记
35:25
But the the enterprise of machine learning is about building models of understanding for things that we don't understand. So, you know, we're we're building these these alien artifacts. And you gave a wonderful example of, you know, imagine I'm generating some text. And we we might think naively that just like the way it's rendered around the edges is actually the thing doing the heavy lifting. But actually the the the load bearing thing might be something completely different. So many many of our intuitions don't don't really work.
但机器学习这门事业,本质上就是为我们并不理解的东西建立起理解的模型。所以,你知道,我们其实是在造这些外星造物。你举了一个很棒的例子,就是说,想象我正在生成一段文本。我们可能会天真地以为,它在边缘呈现出来的那种方式,才是真正在承担主要工作的东西。但实际上,真正起承载作用的可能是完全不同的东西。所以我们的很多直觉其实并不成立。
便签笔记
35:50
But for me, and I know you're allergic to the word understanding. But for me, understanding is is possession of a generative model which can do the thing. So if you could create a physics simulator of some phenomenon, assuming that it's it's not lossy, I would say that you understood that thing. And we think of machine learning as modeling extant examples, which means rather than how it's constructed, how it's built, it's modeling the thing at the end. But you were describing, know, like the Game of Life is a great example of this.
不过对我来说——我知道你对「理解」这个词很过敏——但对我来说,理解就是拥有一个能够做成那件事的生成模型。所以如果你能为某个现象造出一个物理模拟器,前提是它不是有损的,那我就会说你理解了那个东西。而我们通常把机器学习看作是对现存样本的建模,也就是说,它建模的不是它如何构成、如何被造出来,而是最终结果那个东西。但你刚才描述的,你知道,生命游戏就是一个很好的例子。
便签笔记
36:18
It's kind of like it's modeling it's it's it's learning the path, not the destination. But you were describing something fascinating last night, which is that we have, for example, this recycling mechanism in AlphaFold. And you can place things through the structure many many times. And what seems to happen is at the beginning, it solves the most complex problem and then it's kind of refining, it's kind of refining. And this is a little bit like the game of life. It's not it's not like it's simulating the creation.
这有点像是在建模——它学的是路径,而不是终点。但你昨晚描述了一件很有意思的事,就是比如说,我们在 AlphaFold 里有这种循环回收(recycling)机制。你可以让数据反复多次地通过这个结构。而似乎发生的情况是,一开始它先解决最复杂的问题,然后就是在不断精修、不断优化。这有点像生命游戏。并不是说它在模拟这个创生的过程。
便签笔记
11预测、控制、理解:三者的分界
36:44
It's almost like it's, at any point, learning how to refine and optimize the structure. Okay. So we I think we should distinguish 3 things. Predict, control, understand Yes. 1st. So predict means that you say, I'm gonna do a thing. What am I gonna what will be this value of my machine? What will appear on my computer screen in the future? That is predict. Control is I want to measure this thing in the future, and I want it to come out 17. Right? That's control. Understand is a lot like predict, except there's a human in the loop.
更像是在任何一个时点,它都在学习如何精修和优化这个结构。好的,那我想我们应该区分三件事。预测、控制、理解。是的。第一,预测的意思是,你说,我要做一件事。我要……我这台机器的这个值会是多少?将来我的电脑屏幕上会出现什么?这就是预测。控制则是,我想在将来测量这个东西,而且我希望它的结果是 17。对吧?这就是控制。理解跟预测很像,只不过中间多了一个人。
便签笔记
37:16
Understand means that I have such a small collection of facts that you will predict, and you will do it with facts that I can communicate to another human in kind of this compact fix fits on an index card. That's almost understand. And so I think these machines let us predict. They let us control. We have to derive our own understanding at this moment. Right? We can experiment now on the artifact. We can look at the 200,000,000 predicted structures, not just the 200,000 experimental structures in order to help us understand.
理解意味着,我掌握了非常少的一组事实,用它们你就能做出预测,而且这些事实我可以传达给另一个人,它足够紧凑,能写在一张索引卡上。这差不多就是理解了。所以我认为,这些机器让我们能够预测。它们让我们能够控制。而此刻,理解还得靠我们自己去提炼。对吧?我们现在可以在这个产物上做实验。我们可以去看那 2 亿个预测出来的结构,而不只是那 20 万个实验测定的结构,来帮助我们理解。
便签笔记
37:53
But it doesn't do the act of understanding for us. It does the act of predict and maybe control. Now, though, then there's maybe 1 other thing. There is the algorithm, and it's really important, I think, concept to machine learning. There's the algorithm you program and the algorithm you get. Or, you know, machine learning as code meets data produces weights. And so 1 of the all 1 of the kind of lasting debates in machine learning, how much work is done by the code, How much work is done by the data that ends up in the weights?
但它并不能替我们完成“理解”这件事。它做的是预测,也许还有控制。不过呢,可能还有另外一件事。有算法这回事,我觉得这是机器学习里非常重要的一个概念。存在你编写的算法,和你最终得到的算法。或者说,机器学习就是代码遇上数据,产出权重。所以机器学习里一个长期存在的争论就是:有多少工作是代码完成的?有多少工作是由那些最终体现在权重里的数据完成的?
便签笔记
38:25
And so what I think you what we see in AlphaFold in a certain sense is a very beautifully intuitive algorithm, an algorithm we can, in some sense, understand, right, that it does successive geometric refinement. I communicated that to you in a few words. You probably almost saw it in your head even though I don't think you've seen the these videos. I mean, they're in the supplement of our Nature paper. But but that is an algorithm that humans already came up with. Maybe we should almost do, you know, gradient descent and some empirical model that makes each thing more correct.
所以我觉得,从某种意义上说,我们在 AlphaFold 里看到的是一个非常漂亮、非常直观的算法,一个我们在某种程度上能够理解的算法,对吧,它做的是逐次的几何精修。我用几句话就把它讲给你听了。你脑子里大概已经能想象出画面了,尽管我觉得你并没有看过这些视频。它们其实放在我们那篇《自然》论文的补充材料里。但那是人类早就想出来的算法。也许我们几乎应该做的就是梯度下降,加上某个能让每一步都更接近正确的经验模型。
便签笔记
38:57
Maybe that's how AlphaFold should work. But that wasn't what we programmed. But we also still thought about it in a certain sense. So we were thinking about things like recycling in terms of, you know, wow, isn't it weird that alpha fold in n layers has to give an answer for no matter how hard this problem is? And maybe we should give it some more layers. And my GPU is out of memory, so maybe I should just, you know, run it back through so I don't have to have more memory. But even without that, I think AlphaFold, even without recycling, was learning this kind of iteration.
也许 AlphaFold 就该这么工作。但那并不是我们编写出来的东西。不过我们在某种意义上还是思考过这些的。比如说,我们当时在琢磨“循环回收”(recycling)这类东西,想的是哇,这不是很奇怪吗?不管问题有多难,AlphaFold 都得在 n 层之内给出一个答案。也许我们应该多给它几层。可我的 GPU 显存不够了,那也许我就干脆把结果再送回去跑一遍,这样就不用占更多显存了。但我觉得,就算没有这个,AlphaFold 即使不做 recycling,也在学习这种迭代的方式。
便签笔记
39:29
And then we put in a kind of code idea, architectural idea to help this process that it was gonna learn from the data. Going back to the earlier thing about exactly how far residues are apart, we didn't tell AlphaFold that. We knew that the data would scream at it, that I and I plus 1 were 1.3 angstroms apart. So I think when we think about our human understanding, I think 1 of the you know, I don't really love the bitter lesson as people try and apply it. In fact, AlphaFold 2 is the opposite of that.
然后我们放进去一个代码层面的想法、架构层面的想法,来帮助它从数据中学到的这个过程。回到前面说的残基之间到底相距多远那件事,我们并没有告诉 AlphaFold 这一点。我们知道数据会冲着它大喊:第 i 个和第 i+1 个残基之间相距 1.3 埃。所以我觉得,当我们思考人类的理解时,我觉得……我其实不太喜欢人们试图套用“苦涩的教训”(bitter lesson)的方式。事实上,AlphaFold 2 恰恰是它的反面。
便签笔记
40:05
We did a whole bunch of specialty stuff because our data is not finite. And in fact, now that we've gone to language models, we found our data is still finite. The Internet is finite. So I think, you know, don't do architectural research is the wrong thing to draw from it. But have some humility about which things go into your code and which things will be derived from your data. Look at what's missing. Understand the algorithm that deep learning that the deep learning is trying to learn, how can you accelerate it, how can you add hypotheses, and where you especially you add kind of communication.
我们做了一大堆专门化的东西,因为我们的数据并不是无限的。而且事实上,现在我们转向语言模型之后,我们发现数据依然是有限的。互联网是有限的。所以我觉得,从中得出“不要做架构研究”这个结论是错的。但要保持一点谦逊,想清楚哪些东西该进你的代码,哪些东西会从你的数据中自然推导出来。看看缺了什么。理解深度学习正在试图学习的那个算法,你怎么去加速它,怎么去加入假设,以及在哪些地方尤其是你在哪些地方加入某种通信。
便签笔记
40:36
The most important thing we would do within the architecture is modify which units communicated and how. I think all of these have been kind of how we drive understanding to ultimately make an iterative process. And it should shock no 1 that if you're trying to make an intricate geometric object that you are going to iterate. Or similarly, if you think about generating text. Right? And 1 kind of naive assumption that people will make is that these are next word generators, so they have no idea what's gonna happen in 2 words ahead or 3 words ahead. But, of course, you can't think of the you can't write down the next word without I don't I don't start a sentence not knowing how it's gonna end most of the time. Right?
我们在架构中会做的最重要的事情,就是修改哪些单元之间进行通信、以及如何通信。我认为所有这些某种程度上都是我们推进理解的方式,最终形成一个迭代的过程。如果你想要造出一个精细的几何对象,那你必然要迭代,这一点不应该让任何人感到意外。或者类似地,如果你想想生成文本。对吧?人们会做的一个天真假设是,这些模型只是下一个词的生成器,所以它们根本不知道两个词或三个词之后会发生什么。但当然,你没法在不……我是说,我大多数时候不会在不知道一句话会怎么结尾的情况下就开始写这句话。对吧?
便签笔记
41:23
At some points, I change, but I think ahead a little bit in order to accomplish my task. And so the understanding that we see built into these models are kind of the structures that we want sometimes emerge and sometimes don't. And I think we valorize the high level ideas that impose, for example, an AlphaFold 3 coming back to AlphaFold 3. Right? You said it is a diffusion model. But I would argue it's a different diffusion model than an image model. Maybe well, there's some different for 1 thing, there's a huge trunk that is not in any way a diffusion model that's only run once.
有时候我会中途改变,但为了完成我的表达,我确实会往前想一点。所以我们在这些模型中看到的那种理解,其实就是我们想要的那类结构,有时候会涌现出来,有时候不会。而我认为我们会高估那些高层次的想法,比如说强加在……回到 AlphaFold 3,就以 AlphaFold 3 为例。对吧?你说它是一个扩散模型。但我会说,它是一种不同于图像模型的扩散模型。也许……嗯,有一些不同之处,首先,它有一个巨大的主干(trunk),那部分完全不是扩散模型,而且只运行一次。
便签笔记
41:59
That trunk is probably where the structure is actually determined, and the diffusion is just like the structure module was a geometrization engine that took a set of really quite good constraints that had very clear notion of the structure within those constraints and solved up the details. I think AlphaFold 3 diffusion is similar, and it's especially similar because, in fact, in images, okay, you start generating an image and you see especially these early trained diffusion models generate kind of colored blobs, and they start to decide what those colored blobs mean.
那个主干很可能才是结构真正被确定的地方,而扩散部分就像结构模块一样,是一个几何化引擎,它接收一组相当不错的约束——这些约束里已经对结构有非常清晰的概念——然后把细节补齐。我认为 AlphaFold 3 的扩散过程也是类似的,而且尤其相似,因为事实上,在图像里,好的,你开始生成一张图像,你会看到——尤其是那些早期训练的扩散模型——它们先生成一些彩色的斑块,然后才开始决定这些彩色斑块意味着什么。
便签笔记
42:29
And they pretty clearly kind of decide what those colored blobs will mean later because you could stop them in the middle of the process and run them again and get a somewhat different interpretation of those colored blobs. In AlphaFold 3, you actually have an interesting thing that if you look at AlphaFold 2, we can kind of, through this process of projecting out intermediate layers, see what it solves 1st. And it basically solves local details, local pieces. It starts to put local pieces together.
而且很明显,它们是在后面才决定这些彩色斑块将意味着什么,因为你可以在过程中途把它停下来再重新跑一遍,就会得到对这些彩色斑块略有不同的解释。在 AlphaFold 3 里,其实有个有意思的现象:如果你看 AlphaFold 2,我们可以通过把中间层投射出来的这种方法,看到它最先解决的是什么。结果基本上是,它先解决局部细节、局部片段。然后它开始把局部片段拼到一起。
便签笔记
42:52
It's agglomerative, in how it solves a structure as is kind of natural. The easiest thing to predict is your local structure. The hardest thing to predict is your largest scale structure. That's how alpha 2 works. If you look at alpha fold 3 and you take coordinates, which you've added a very large amount of noise to, well, the very 1st thing you have to solve is how, say, you have 2 proteins, how do they associate it? Where are their 2 blobs relative to each other? What are their Gaussians? So the the problem that AlphaFold 2 is solving last is the problem that AlphaFold three's diffusion has to realize 1st.
它解决结构的方式是凝聚式的(agglomerative),这也算是很自然的。最容易预测的是局部结构。最难预测的是最大尺度的结构。AlphaFold 2 就是这么工作的。如果你看 AlphaFold 3,你拿到一组坐标,而你往里面加了非常大量的噪声,那么你首先要解决的问题就是,比方说,你有两个蛋白质,它们是怎么结合到一起的?它们那两团东西相对彼此的位置在哪儿?它们的高斯分布是什么样的?所以 AlphaFold 2 最后才解决的那个问题,恰恰是 AlphaFold 3 的扩散模型必须最先搞定的问题。
便签笔记
43:28
And how does it do it? The answer is not that it comes up with an orientation and builds the protein around it because, of course, it's going for 1 correct answer or at least a very narrow distribution. The answer is really the big network before it plus the 1st pass through the diffusion network is solving the overall structure. And then the diffusion is realizing in any details it couldn't solve before, it's basically sampling among. So it is diffusion technically, but it's much closer to AlphaFold 2.
那它是怎么做到的呢?答案并不是说它先定出一个朝向、再围绕这个朝向把蛋白质搭起来,因为它当然是要得到一个正确答案,或者至少是一个非常窄的分布。真正的答案是:前面那个大网络,加上扩散网络的第一次前向传播,就已经把整体结构解决掉了。然后扩散过程再去补上之前解决不了的那些细节,本质上就是在里面做采样。所以从技术上讲它确实是扩散模型,但它其实更接近 AlphaFold 2。
便签笔记
43:57
I think there's no reason that it was kind of very specific technical reasons around kind of laziness and geometry that made diffusion a really good choice for Alpha Fold 3. It made it easier to handle ligands and handled some bond distances and local things. But it's not like diffusion in the same way as, oh, it's drawing the blobs and deciding what they mean at the end. So I think all of these are people like to think of these like to say, this works because it's a transformer. And this works because it's a transformer doesn't explain why chat models have gotten vastly better in the last 3, 4 years.
我觉得没有什么别的原因,只是有一些非常具体的技术理由,围绕着某种“偷懒”和几何结构,才让扩散成为 AlphaFold 3 一个很好的选择。它让处理配体变得更容易,也能处理一些键长和局部性质的问题。但它并不是那种意义上的扩散——哦,先画出一团团色块,最后再决定它们代表什么。所以我觉得,这些东西大家都喜欢这么想、喜欢说:这个能work是因为它是个 Transformer。但“这个能work是因为它是 Transformer”解释不了为什么聊天模型在过去三四年里变得好这么多。
便签笔记
44:35
It doesn't explain all the research. It doesn't explain what researchers do every day. All of these details are far more important than this high level bit of, is it a transformer, is it a diffusion model, that we wanna talk about. And then also even these diffusion mechanisms don't work in the way of kind of progressive refinement that makes sense for images. Right? Maybe you'll make colored blobs, and you'll decide what those colored blobs mean. Even that, I think you can argue, maybe not entirely the story, but it's definitely not the story for proteins because that's the hardest problem is the large scale structure. I mean, a sense, this is leaning towards this idea of constructive complexity that we were talking about before.
它解释不了所有的研究工作。它解释不了研究人员每天在做什么。所有这些细节,都比“它到底是不是 Transformer、是不是扩散模型”这种我们爱聊的高层次问题重要得多。还有就是,即便是这些扩散机制,其工作方式也不是那种对图像来说很合理的“逐步细化”。对吧?可能你会先生成彩色的色块,再决定这些色块是什么意思。我觉得就算是对图像,这也未必是全部的故事,但对蛋白质来说这肯定不是事实,因为最难的问题恰恰是大尺度结构。某种意义上,这其实呼应了我们之前聊过的那个“构造性复杂度”的想法。
便签笔记
12通用还是定制:智能与表征之争
45:12
And I'd love to get your your general take on on what this means for artificial general intelligence. Because with language models, for example, we train them basically with behavior cloning. So, you know, we have this this rich adaptive generative process and we we generate all of this language and we train language models on them. And for me, intelligence is the adaptive acquisition of coarse grained representations. Culture and language is changing all of the time. So we invent the word unalive to get around the filters on social media platforms and that's an example of ling linguistic agency.
我很想听听你总体上怎么看这对通用人工智能意味着什么。因为拿语言模型来说,我们基本上是用行为克隆的方式训练它们的。你知道,我们有这么一个丰富的、自适应的生成过程,我们生成了所有这些语言,然后拿它们去训练语言模型。对我来说,智能就是自适应地获取粗粒度表征的能力。文化和语言一直都在变。比如我们发明 unalive 这个词来绕过社交平台的过滤机制,这就是语言能动性的一个例子。
便签笔记
45:43
Language models, we noticed that when we do this iterative adaptive refining with active active fine tuning and adaptation, they become a bit intelligent. They they learn new representations and and they adapt. And in a way, what they're doing is even though they're ungrounded from the path, they can they can take a code solution like AlphaEvolve and they can refine it and they can refine it. And it seems to work really really well. But are we in this regime, do you think, that we're not necessarily building artifacts that have the same type of generality.
对语言模型,我们注意到,当我们通过主动的微调和适应做这种迭代式的自适应精炼时,它们会变得有点“聪明”。它们学到新的表征,并且会适应。某种程度上,它们做的事情是——尽管它们脱离了现实世界的锚定,它们可以拿一个像 AlphaEvolve 那样的代码方案,不断地精炼、再精炼。而且这看起来效果非常非常好。但你觉得,我们是不是处在这样一种状态:我们并不一定在造出具有同样这种通用性的产物?
便签笔记
46:15
I mean, what what do you think about intelligence in general? So this question of representations is very, very important and far less important than people believe 5 years ago in the explicit way. So just like we were talking about the things that AlphaFold does and the things that AlphaFold is forced to do by its code or or, you know, obviously, everything that's forced to do by its code, it does. But many things it does, it does without being forced. Because it had to learn it to make a good predictive model of the data.
我是说,你总体上怎么看待智能这件事?表征这个问题非常非常重要,但又远没有五年前大家以为的那么重要——我是指以那种显式的方式。就像我们刚才聊 AlphaFold 做的那些事,以及 AlphaFold 被它的代码逼着去做的那些事——当然,凡是被代码强制的,它都会去做。但它做的很多事情,并不是被强制的。因为它必须学会这些,才能对数据建立一个好的预测模型。
便签笔记
46:51
It had to find good intermediate representations. So in a certain sense, I think the most seductive idea in in machine learning is always there's this thing I know will have to be in the end there in the end, so I'm gonna have to have a u I'm gonna have to have a place in my code that is named that and then forces the mechanism to high level concept builder thingamajigger. Right? And that was a very popular kind of I'll make the concepts units. I shall force disentangled representations via this law sometimes on the intermediate layer.
它必须找到好的中间表征。所以某种意义上,我觉得机器学习里最有诱惑力的想法总是:有这么个东西,我知道它最后一定得在那儿,所以我得在我的代码里专门留一个位置,给它起个名字,然后强迫这个机制变成一个“高层概念构造器”之类的玩意儿。对吧?当年很流行这种做法:我要让概念成为单元,我要通过某个损失函数强制得到解耦的表征,有时候是加在中间层上。
便签笔记
47:30
This is where it will store those. And that's reasonable to go test. But what we've seen a lot of is that a lot of the things that you would imagine needed or needed for intelligence are developed by desperately trying to predict the next token really, really well. And they're not they're not developed because you predict next tokens at all. They develop because you do a really, really good job at it. And so these kind of generalized spaces, representations, understanding of concepts is forced very slowly with data.
这里就是它存放那些东西的地方。去做这样的测试是合理的。但我们大量看到的情况是,很多你会以为是智能所必需的东西,其实是在拼命地去把下一个词预测得非常非常好的过程中自然发展出来的。而且它们并不是因为你在预测下一个词才发展出来的。它们是因为你把这件事做得非常非常好,才发展出来的。所以这些泛化的空间、表征、对概念的理解,是靠数据非常缓慢地逼出来的。
便签笔记
48:05
Right? Pretty much all the kind of you know, there's a lot of log linears or my you know, everyone's least favorite functional is right. The things go up as the things go up linearly with the exponent of effort that we see all the time in our scaling laws. But we do get these concepts and representations, and what we don't really, I think, have an answer for is how do we get them cheaper? Now sometimes we can get them via programming. Right? We get memory like things. Now we have language models writing notes for itself and then retrieving those notes. So we find out it's better to keep reminding agents what they're doing so they don't forget over long trajectories.
对吧?基本上都是那种,你知道,有很多对数线性的关系,或者说大家最不喜欢的那种函数形式,对吧。东西随着投入的指数而线性增长——这在我们的 scaling law 里一直都能看到。但我们确实得到了这些概念和表征,而我们真正还没有答案的问题是:怎么才能更便宜地得到它们?有时候我们可以通过编程的方式得到。对吧?我们得到了类似记忆的东西。现在我们让语言模型给自己写笔记,然后再把笔记检索回来。所以我们发现,不断提醒智能体它正在做什么会更好,这样它在很长的轨迹上就不会忘。
便签笔记
48:46
So we we build weights. We build artifacts. We find deficiencies. We can often paper over those deficiencies in some kind of software harnesses, but then we don't yet know we don't that doesn't immediately drive back into exactly your machine learning, or you don't, like, put a harness with external memory and then distill it back into the network and have amazing memory things that no longer need this harness. That, we haven't figured out how to do. Unfortunately, John, we we have we have to wrap it.
所以我们训练出权重。我们造出各种产物。我们发现缺陷。我们常常能用某种软件脚手架把这些缺陷掩盖过去,但接下来我们还不知道——我们不知道怎么把这些东西直接反哺回机器学习本身;你没法说,搭一个带外部记忆的脚手架,然后把它蒸馏回网络里,从而得到一个不再需要脚手架、记忆能力超强的模型。这个我们还没搞明白怎么做。很遗憾,John,我们得收尾了。
便签笔记
13非洲的一千名科学家
49:17
But doctor John Jumper, it's been an an honor to have you on the show. Thank you so much for joining us today. Been tremendous fun. Thank you. So as I said earlier, Emmanuel Nji, he's based in Africa, and he's actually training scientists. Not just he's not just giving them access to AlphaFold, he's training them how to use it, how to interpret the results, and how to help scientists build experiments using the database. Yes. So my research focus on on drug discovery for malaria and enteric bacteria.
但 John Jumper 博士,很荣幸能请到你上我们的节目。非常感谢你今天来参加。非常开心,谢谢。像我前面说的,Emmanuel Nji 在非洲,他实际上在培训科学家。他不只是让他们能用上 AlphaFold,他还在教他们怎么使用、怎么解读结果,以及怎么帮助科学家利用这个数据库来设计实验。是的。我的研究聚焦于疟疾和肠道细菌的药物发现。
便签笔记
49:46
And then I'm also involved in capacity building for Africa based researchers, using tools like AlphaFold. Initially, African scientists didn't have access to expensive structural biology tools. With AlphaFold, these researchers can now do complex experiments that were not possible before and tackle diseases such as malaria, HIV, and other, antibiotic resistant infection. In my own research on on drug discovery, so I use AlphaFold in terms of solving structures of cryo EM data, and I also utilize that to to to map out the mechanisms of the proteins.
同时我也参与非洲本地研究人员的能力建设,教他们使用像 AlphaFold 这样的工具。一开始,非洲的科学家用不上那些昂贵的结构生物学工具。有了 AlphaFold,这些研究者现在可以做以前不可能做的复杂实验,去攻克疟疾、艾滋病,以及其他抗生素耐药性感染等疾病。在我自己的药物发现研究中,我用 AlphaFold 来解析冷冻电镜数据的结构,我还用它来梳理蛋白质的作用机制。
便签笔记
50:39
So for him, AlphaFold was so impactful. Like, if you think about the before and after, we're living in a different world now. At that time, to face a protein was, like I said, was really, really difficult. And so I tried several years, close to 4, 5 years, and it wasn't successful. And with AlphaFold imagine this is more than 10 years ago. With AlphaFold, I went back and did just 1, protein purification, collected the data, and we used AlphaFold. In combination, I got the structure in less than 2, 3 months.
所以对他来说,AlphaFold 的影响非常大。你想想之前和之后的对比,我们现在生活在一个完全不同的世界里。当年,要解出一个蛋白质,就像我说的,真的非常非常困难。我试了好几年,将近四五年,都没成功。而有了 AlphaFold——你想想那可是十多年前的事了。有了 AlphaFold,我回过头去只做了一次蛋白纯化,收了数据,然后用上 AlphaFold。结合起来,我在不到两三个月里就拿到了结构。
便签笔记
51:20
And now it's his goal to train as many scientists as he can how to use this technology for the betterment of humankind. This year, with funding from Google DeepMind and Swedish Research Council, we have scaled up to 100, and there's no drop in the quality of the training. In fact, it was there was an improvement. So based on this based on this, we want to train 100 scientists every year for the next 10 years. So we're targeting close to 1,000 African scientists in the next decade to be able to utilize this tool effectively, and then we want to form an emerging community of structural biology practitioners working on prevalent diseases in Africa.
现在他的目标是尽可能多地培训科学家,教他们如何用这项技术造福人类。今年,在 Google DeepMind 和瑞典研究理事会的资助下,我们把规模扩大到了一百人,而培训质量并没有下降。事实上,反而还有提升。所以基于这一点,我们希望在未来十年里每年培训一百名科学家。也就是说,我们的目标是在未来十年里让接近一千名非洲科学家能够有效地使用这个工具,然后我们希望形成一个新兴的结构生物学从业者社群,专注于非洲的高发疾病。
便签笔记
52:16
So that was the AlphaFold show. Thank you very much to John and Emmanuel. Yeah. The the conversation with John was very interesting. He's he's so inspiring because I think he is testament to the fact that even though we talk about all of these general purpose foundation models, to really advance the frontier and to build cutting edge applications in science. We need to do a lot of engineering. We need, you know, domain knowledge. We need serious expertise. And a lot of our models will actually look quite hybrid.
以上就是本期的 AlphaFold 节目。非常感谢 John 和 Emmanuel。是的,和 John 的这场对话非常有意思。他非常鼓舞人心,因为我觉得他证明了一件事:尽管我们一直在谈这些通用的基础模型,但要真正推进前沿、要在科学领域做出尖端应用,我们需要做大量的工程工作。我们需要领域知识。我们需要真正过硬的专业能力。而且我们的很多模型实际上会是相当混合式的。
便签笔记
52:44
They'll look quite customized. And I think AlphaFold is a kind of proof of existence for the types of hybrid models that we can deploy to further the field of science. John, I wish you the very best of luck in your new position, at Anthropic, and thanks for watching the show.
它们会是相当定制化的。我觉得 AlphaFold 就是一个存在性证明,说明我们可以部署什么样的混合模型来推动科学的发展。John,祝你在 Anthropic 的新岗位上一切顺利,也谢谢大家收看本期节目。
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

AlphaFold 的成功不在于某个时髦架构(等变性、扩散模型),而在于十几个中等规模的工程与科学想法叠加,且它刻意保持「窄」——只预测一类实验结果,而非模拟整个细胞;John Jumper 用这一反「苦涩教训」的路径拿下诺贝尔奖后,离开 DeepMind 加入 Anthropic。

核心要点

  • 解决的是 70 年悬而未决的瓶颈问题:DNA 测序早已廉价,但测一个蛋白质结构典型需要一年、约 10 万美元,要靠小镇大小的同步辐射装置和多年的结晶尝试;人类基因组约 2 万种蛋白,全球花了数十年也只积累了约 14 万(现 20 万)个实验结构。AlphaFold 2 把单个预测压缩到 5–10 分钟,精度达原子半径量级,2020 年 CASP 主办方宣布问题「基本解决」。
  • 开放数据库带来规模效应:DeepMind 公开发布了 2 亿个预测结构(几乎覆盖所有已测序基因组的蛋白),目前超过 300 万用户、190 多个国家在使用。非洲结构生物学家 Emmanuel Nji 原本 4–5 年未能解出的蛋白,用一次纯化 + AlphaFold 在 2–3 个月内完成;他现在每年培训 100 名非洲科学家,十年目标 1000 人。
  • AlphaFold 的「谦逊」是其有效性的关键:Jumper 反复强调它不是细胞模型,只是「预测你那个做了一年的实验会得到什么」。它预测的是一类 Nature 级论文的结果,这让其准确度可被严格刻画;科学家再把这个窄预测器用于意料之外的用途(如筛选数千对蛋白是否结合)。
  • 真实用法是与实验互证:典型案例是胆固醇转运蛋白——冷冻电镜给出模糊「blob」,AlphaFold 预测恰好嵌入其中,二者互相确认得出原子模型;Midnolin 案例中,研究者对约 500 个受影响蛋白跑 AlphaFold,约 40% 显示同一「钳夹」模式,10 个实验验证中 9 个完全符合,剩下 1 个 AlphaFold 预测了两个结合位点,同时敲除后降解完全消失。
  • 药物研发的瓶颈是「知道该拧哪颗螺丝」:Jumper 用工厂技师的笑话说明——治病难在我们不了解生物学本身(如自闭症涉及全身而非一个蛋白)。AlphaFold 3 扩展到小分子、脂质等「蛋白宇宙」,能回答药物结合在哪里,Isomorphic Labs 正据此做药物设计。
  • 等变性被严重高估:AlphaFold 2 比 AlphaFold 1 在 GDT 上高约 30 分,消融实验显示去掉 SE(3) 不变注意力(IPA)只损失约 2–2.5 分;90% 以上的算力与精度来自 Evoformer 主干(轴向注意力)。真正关键的是 FAPE 损失函数(在每个残基参考系里度量其他所有残基的误差)和「残基气体」的解耦表示,但学界仍把 AlphaFold 2 讲成等变性的胜利。审稿人评价其「值 6–7 篇论文」,Jumper 类比为「18 个二垒安打而非一两个全垒打」。
  • 冗余设计与「9 次错 1 次对」的经验主义:多数问题团队用两种方式各解一次,单一消融损失小,双消融(去掉 recycling + IPA)才崩盘。有人删掉卷积层、参数减少反而验证损失改善——卷积对该任务是有害的。这些消融直接指导了 AlphaFold 3:Evoformer 缩减为几层,换成更简单的 Pairformer,性能反而提升。
  • 架构与训练思路价值 100 倍数据:AlphaFold 1 与 2 用的是完全相同的训练数据;AlQuraishi 实验室用 1% PDB(约 1.5 万个结构)重训 AlphaFold 2,仍胜过 AlphaFold 1。
  • AlphaFold 3 的「扩散」并非图像式扩散:主干网络只跑一次并基本决定结构,扩散模块类似 AF2 的结构模块,只是「几何化引擎」。AF2 是从局部到整体的凝聚式求解,而 AF3 的扩散必须先解决最难的全局相对位置——选扩散是出于处理配体和键长的具体技术便利,而非原理性优势。「它是 Transformer」解释不了这几年聊天模型为何变强,细节远比顶层标签重要。
  • 对通用智能的看法:区分预测、控制、理解——模型给我们前两者,理解仍需人类从 2 亿预测结构中自行提炼。许多能力不是被代码强制出来的,而是「拼命把下一个 token 预测好」时自发涌现;当前用外部记忆等 harness 弥补缺陷,但尚不知如何把 harness 蒸馏回权重。

结论与值得注意的细节

  • Jumper 直言「不喜欢人们对苦涩教训的套用」,AlphaFold 2 正是其反例:数据有限(互联网也是有限的),应保持谦逊地判断哪些放进代码、哪些交给数据学,尤其是调整「哪些单元之间如何通信」。
  • 节目主持人的解读:AlphaFold 证明科学前沿 AI 仍需大量领域知识与定制化混合架构;Jumper 加入 Anthropic 意味深长——他并非做通用模型出身,Anthropic 为何需要他「只能猜测」。
  • 诺奖当晚细节:10:30 仍无消息,Jumper 以为落选,妻子让他再等等,随即来自瑞典的电话打进来。
  • 2024 年化学奖一半给 David Baker(蛋白设计),另一半由 Hassabis 与 Jumper 分享(结构预测)。
  • 访谈录制于 Anthropic 消息公布之前。
核心句型 · 10
1. It's not one or two …, it's eighteen …
“It's not 1 or 2 home runs. It's, you know, 18 doubles.”
用「不是 A,而是 B」的对照修正听者预期,且两侧同属一个比喻域(棒球)。仿写时务必让 A、B 在同一意象内,否则对照失效。
2. be worth a clean N× in …
“The architecture and training ideas that we put into AlphaFold 2 were worth a clean 100 x in data”
clean 在此表「实打实、不掺水的」,用于量化某改进折算成另一资源的等价量。适合技术汇报中把定性优势换算成硬指标。
3. It should shock no one that …
“It should shock no 1 that if you're trying to make an intricate geometric object that you are going to iterate”
以「这不该让任何人意外」引出一个作者认为显而易见的结论,语气克制而带轻讽。比 obviously 更礼貌,也更书面。
4. X is not nearly as … as Y
“Global s c 3 symmetry is not a very powerful symmetry. It's not nearly as kind of big and powerful as a symmetry like …”
not nearly as … as 表「远不及」,否定强度远高于 not as … as。用于压低对方看重的因素、抬出真正关键的因素。
5. I would almost object, not technically but thematically, to …
“I would almost object not technically, but kind of thematically to AlphaFold 3 as a diffusion model.”
「不在 X 层面,而在 Y 层面反对」是学术争论中的高阶让步式表达:先承认对方事实无误,再把分歧移到框架层面。
6. There's the … you program and the … you get.
“There's the algorithm you program and the algorithm you get.”
同一名词加两个不同定语从句形成对峙,句式极简而张力强。适合表达设计意图与实际结果之间的落差。
7. have some humility about which … and which …
“But have some humility about which things go into your code and which things will be derived from your data.”
have humility about 后接并列的两个 which 从句,把「保持谦逊」具体化为一个明确的判断题,避免空洞的劝诫。
8. the thing preventing us from … is not that …
“The thing preventing us from curing, say, I guess, autism, right, is not that we know exactly 1 protein.”
先指名障碍、再否定一个常见误解,是纠偏式论证的标准开头。后面通常接 It's that … 给出真正原因。
9. nine times out of ten, you find out you're wrong
“You try it. You measure. 9 times out of 10, you find out you're wrong.”
三个短句层层推进(试→测→错),配以 nine times out of ten 的高频估计。适合描述实证工作流,节奏感强。
10. could have kept … close to their chest, but they decided to …
“They could have kept this close to their chest, but they decided to release it”
could have done 表本可如此而未如此,与 but they decided to 构成反差,用来称许某个非必然的选择。
词汇精讲 · 190 · 按出现顺序
holy grail /ˌhoʊli ˈɡreɪl/ n. phr. 0:00
圣杯;(引申)某领域长期追求而未得的终极目标
with the press of a button phr. 0:00
按一下按钮就能……;形容极其省力、瞬间完成
bottleneck /ˈbɑːtlnek/ n. 0:36
瓶颈;制约整体进展的关键环节
amino acids /əˌmiːnoʊ ˈæsɪdz/ n. phr. 0:36
氨基酸;构成蛋白质的基本单元
catalyzes /ˈkætəlaɪzɪz/ v. 0:36
催化(化学反应);引申为促成、加速
9 times out of 10 phr. 0:36
十次里有九次;绝大多数情况下
incremental /ˌɪnkrəˈmentl/ adj. 1:42
渐进的、一点一点累加的(与突破式相对)
operationalized /ˌɑːpəˈreɪʃənəlaɪzd/ v. 2:25
投入实际使用、转化为可操作的东西
leap forward n. phr. 2:25
飞跃式的进展
mind blowing /ˈmaɪnd ˌbloʊɪŋ/ adj. 3:01
令人震撼的、颠覆认知的(口语)
stops and starts n. phr. 3:01
断断续续、屡次中断又重启的过程
fair play to phr. 3:01
(英式口语)该给某人点个赞、值得肯定
kept this close to their chest phr. 3:01
把(信息、成果)捂住不外传;源自打牌不让人看牌
verdict /ˈvɜːrdɪkt/ n. 3:38
裁定、判决;此处指诺奖委员会的正式定论
resolution /ˌrezəˈluːʃn/ n. 3:38
(成像的)分辨率;此处指能看清细节的精细程度
lights up phr. v. 4:16
(屏幕)亮起;也可指人脸放光、地图上被点亮
prank call /ˈpræŋk kɔːl/ n. phr. 4:16
恶作剧电话
departure /dɪˈpɑːrtʃər/ n. 4:16
离职、离开;比 leaving 更正式
imbue /ɪmˈbjuː/ v. 4:16
注入、灌注(某种性质或情感),常用 imbue A with B
and whatnot phr. 5:08
以及诸如此类的东西(口语,收尾用)
purification /ˌpjʊrɪfɪˈkeɪʃn/ n. 5:08
(蛋白质等的)纯化;实验中分离目标分子的步骤
curated /ˈkjʊreɪtɪd/ adj. 5:46
经过精心挑选与整理的
materialized /məˈtɪriəlaɪzd/ adj. 5:46
已具象化、已落成实体的(计算机中亦指「物化」的缓存视图)
programmatically /ˌproʊɡrəˈmætɪkli/ adv. 6:25
以编程方式(而非手动点击)
give … a go phr. 6:25
试一试(英式口语)
inspirational /ˌɪnspəˈreɪʃənl/ adj. 7:00
给人启发和激励的
drilled in to phr. v. 7:00
深入探究(某话题);drill into 亦作「钻研」
struck me phr. v. 7:00
让我印象深刻、突然意识到(strike 的过去式)
landmark /ˈlændmɑːrk/ n. / adj. 7:36
里程碑(式的);标志性成就
nanomachines /ˈnænoʊməˌʃiːnz/ n. 8:09
纳米机器;此处比喻蛋白质在细胞中的功能角色
chemical groups n. phr. 8:09
化学基团;分子中具有特定性质的原子集合
lovingly /ˈlʌvɪŋli/ adv. 8:43
充满喜爱地、饱含深情地
curls /kɜːrlz/ v. 8:43
卷曲、蜷成一团
helices /ˈhelɪsiːz/ n. 8:43
螺旋(helix 的复数);蛋白中的 α 螺旋结构
compact /kəmˈpækt/ adj. 8:43
紧凑的、密实的
worthy /ˈwɜːrði/ adj. 9:17
够格的、值得的(a worthy PhD project 足以撑起一个博士课题)
misfold /ˌmɪsˈfoʊld/ v. 9:17
(蛋白质)错误折叠;与多种神经退行性疾病相关
whirling /ˈwɜːrlɪŋ/ v. 9:48
飞速旋转的
genome /ˈdʒiːnoʊm/ n. 9:48
基因组;一个生物体全部遗传物质
synchrotrons /ˈsɪŋkrətrɑːnz/ n. 9:48
同步辐射光源;产生极强 X 射线的环形加速器装置
crystallize /ˈkrɪstəlaɪz/ v. 10:25
(使)结晶;此处指把蛋白培养成晶体以便 X 射线衍射
wealth of phr. 10:25
大量的、丰富的(a wealth of understanding)
innumerable /ɪˈnuːmərəbl/ adj. 10:59
数不胜数的、无法计数的
ribosome /ˈraɪbəsoʊm/ n. 10:59
核糖体;细胞内合成蛋白质的分子机器
societal /səˈsaɪətl/ adj. 10:59
社会层面的(比 social 更强调整个社会)
vastly /ˈvæstli/ adv. 10:59
极大地、大幅地(修饰比较级)
rival /ˈraɪvl/ v. 11:48
与……相媲美、可与……匹敌
scalable /ˈskeɪləbl/ adj. 11:48
可扩展的;规模放大后仍可行
downstream tasks n. phr. 12:21
下游任务;机器学习中指基于已有模型/数据开展的后续应用
bring this to life phr. 12:21
把(抽象的东西)讲得生动具体、使之鲜活起来
cholesterol /kəˈlestərɔːl/ n. 12:48
胆固醇
mutations /mjuːˈteɪʃnz/ n. 12:48
突变;DNA 序列的改变
wraps around phr. v. 13:22
环绕、包裹住
cryo electron microscopy n. phr. 13:22
冷冻电子显微镜(cryo-EM);在低温下成像生物大分子
blobby /ˈblɑːbi/ adj. 13:22
团块状的、轮廓模糊的
unravel /ʌnˈrævl/ v. 14:23
解开、层层拆解(谜团或纠缠之物)
phenotypes /ˈfiːnətaɪps/ n. 15:04
表型;基因在个体上表现出的可观察特征
length scales n. phr. 15:04
尺度层级;从原子到细胞到整体的不同空间量级
knocked it down phr. v. 15:39
(基因)敲低,即人为降低某基因/蛋白的表达量
suggestive /səɡˈdʒestɪv/ adj. 15:39
有提示性的、令人产生某种推测的
clamps /klæmps/ n. 16:07
夹钳、夹具;此处比喻蛋白的两个部分夹住底物
degradation /ˌdeɡrəˈdeɪʃn/ n. 16:35
(生化)降解,分子被分解清除的过程
abolished /əˈbɑːlɪʃt/ v. 16:35
彻底废除、完全消除
mechanistic /ˌmekəˈnɪstɪk/ adj. 16:35
机制层面的;解释「如何运作」而非仅描述现象
stick to phr. v. 17:44
黏附在……上;此处指分子结合
symphony /ˈsɪmfəni/ n. 19:25
交响曲;引申为多因素协调运作的整体
compensatory /kəmˈpensətɔːri/ adj. 19:25
代偿性的;系统受扰后自行补偿的
interventions /ˌɪntərˈvenʃnz/ n. 19:25
干预手段;医学上指治疗措施
whack a mole /ˈwæk ə ˈmoʊl/ n. phr. 19:57
打地鼠;喻按下一处又冒出另一处的徒劳应付
humility /hjuːˈmɪləti/ n. 19:57
谦逊、不自我夸大
characterize /ˈkærəktəraɪz/ v. 20:25
刻画、定量描述(某系统的性质)
validity /vəˈlɪdəti/ n. 20:25
有效性、成立程度(科学论断的可靠性)
pin ourselves to phr. 21:10
把自己锚定在……上;固定以某物为依据
geometric deep learning n. phr. 22:24
几何深度学习;把对称性与几何结构作为归纳偏置的一派方法
misattributed /ˌmɪsəˈtrɪbjuːtɪd/ v. 22:24
归因错误、错误地把功劳归给某处
symmetries /ˈsɪmətriz/ n. 22:24
对称性;变换下保持不变的性质
thematically /θiˈmætɪkli/ adv. 22:51
在主题/叙事框架层面上(与 technically 相对)
stick things in boxes phr. 22:51
给事物贴标签归类;简化为现成范畴
evolutionary correlations n. phr. 22:51
进化相关性/共变;同源序列中共同变化的位点携带接触信息
off the shelf adj. phr. 23:27
现成的、不需定制直接可用的
iteratively /ˈɪtərətɪvli/ adv. 23:57
迭代地、一轮一轮反复地
trunk /trʌŋk/ n. 23:57
(神经网络的)主干;承担主要计算的核心部分
axial attention n. phr. 23:57
轴向注意力;分别沿矩阵行、列方向计算注意力
yeast /jiːst/ n. 24:34
酵母;常用的真核模式生物
intermediate loss n. phr. 25:09
中间损失;对网络中间层输出施加的监督信号
categorical predictions n. phr. 25:09
分类式预测;把连续量离散成若干类别来预测
harmonize /ˈhɑːrmənaɪz/ v. 25:09
使协调一致、调和(相互冲突的信息)
invariant /ɪnˈveriənt/ adj. 25:09
不变的;输入做某种变换而输出不变
residues /ˈrezɪduːz/ n. 25:39
(蛋白质中的)残基,即链上的单个氨基酸单元
backbone /ˈbækboʊn/ n. 25:39
主链、骨架;蛋白链上连续的 N-Cα-C 原子链
rigid /ˈrɪdʒɪd/ adj. 25:39
刚性的、不易变形的
bias your attention phr. 26:10
给注意力加偏置,即人为调整注意力权重的分布
locally registered phr. 26:46
局部配准的;在局部坐标系下对齐比较
emerged /ɪˈmɜːrdʒd/ v. 26:46
涌现、自发出现(未被显式设计)
disrespected /ˌdɪsrɪˈspektɪd/ v. 26:46
(此处技术用法)刻意不遵守、不顾(已知约束)
angstroms /ˈæŋstrəmz/ n. 27:21
埃;长度单位,1 埃 = 10⁻¹⁰ 米,约为原子尺度
jointed /ˈdʒɔɪntɪd/ adj. 27:21
带关节的、由铰接段构成的
break it up phr. v. 27:51
把它拆开、分解成若干部分
equivariance /ˌiːkwɪˈveriəns/ n. 27:51
等变性;输入做某变换,输出随之作对应变换
ablate /əˈbleɪt/ v. 28:19
(实验中)移除某组件以测其贡献;名词 ablation 消融实验
put it to bed phr. 28:19
把(争议、问题)彻底了结、就此定案
ruthlessly /ˈruːθləsli/ adv. 29:28
毫不留情地、冷酷地(此处形容不惜舍弃心爱想法)
empirical /ɪmˈpɪrɪkl/ adj. 29:28
实证的、以实测数据而非理论推断为准的
hooked on to phr. v. 29:28
抓着不放、执着于(某个说法)
permutation invariant adj. phr. 29:28
置换不变的;打乱元素顺序不改变结果
pin it down phr. v. 30:04
把(答案)确定下来、限定死
obsess about phr. v. 30:04
过度纠缠于、执念于
valorize /ˈvæləraɪz/ v. 30:38
抬高其价值、过度推崇
transformative /trænsˈfɔːrmətɪv/ adj. 30:38
具有变革性的、能带来根本改变的
home runs n. phr. 30:38
本垒打;喻一击制胜的重大突破
doubles /ˈdʌblz/ n. 30:38
(棒球)二垒安打;喻中等规模但可靠的进展
cratered /ˈkreɪtərd/ v. 31:11
(性能、价格)急剧下滑、崩塌
knock out phr. v. 31:11
敲除、去掉(组件或基因)
pairwise /ˈperwaɪz/ adj. 31:44
成对的;针对每两个元素的关系
interpretability /ɪnˌtɜːrprətəˈbɪləti/ n. 31:44
可解释性;理解模型内部如何运作的研究方向
capacity /kəˈpæsəti/ n. 31:44
(模型的)容量,可用于拟合与表达的能力总量
cut back phr. v. 32:16
削减、砍到更小规模
manifold /ˈmænɪfoʊld/ n. 32:45
流形;数学上局部像欧氏空间的连续空间,此处喻「想法的连续区域」
thou shalt not phr. 32:45
「汝不可……」;仿圣经十诫的古语,戏谐地表达铁律
convolutions /ˌkɑːnvəˈluːʃnz/ n. 33:26
卷积(层);靠局部滑动窗口提取特征的神经网络结构
validation loss n. phr. 33:26
验证集损失;衡量模型在未训练数据上的表现
actively harmful phr. 33:58
实实在在有害的(actively 强调不只是无用)
tactile feel /ˈtæktl fiːl/ n. phr. 33:58
(比喻)手感;靠长期动手才有的直觉判断
banging our head against phr. 33:58
(对某难题)反复碰壁、屡战屡败地死磕
bump /bʌmp/ v. 34:39
(口语)上调、增加(数量或版本号)
tacit knowledge /ˈtæsɪt ˈnɑːlɪdʒ/ n. phr. 34:39
隐性知识;难以言传、只能在实践中习得的经验
park that phr. v. 34:39
(口语)把某话题暂时搁下,稍后再谈
the enterprise of phr. 35:25
……这项事业(enterprise 此处指集体性的长期工作)
artifacts /ˈɑːrtɪfækts/ n. 35:25
人造物、造物;此处指训练出的模型本身
heavy lifting n. phr. 35:25
最费力的核心工作;do the heavy lifting 承担主要负担
load bearing adj. 35:25
承重的;喻真正支撑整个结论的关键部分
allergic to /əˈlɜːrdʒɪk/ adj. phr. 35:50
对……过敏;口语中指极其反感某说法
generative model n. phr. 35:50
生成模型;能够产生符合数据分布的新样本的模型
lossy /ˈlɔːsi/ adj. 35:50
有损的(压缩或再现过程中丢失信息)
extant /ekˈstænt/ adj. 35:50
现存的、尚存的(正式用词)
refining /rɪˈfaɪnɪŋ/ v. 36:18
精修、逐步改进使之更精确
in the loop phr. 36:44
在流程之中、参与其中(a human in the loop 人在环中)
index card /ˈɪndeks kɑːrd/ n. phr. 37:16
索引卡;小卡片,此处喻「极其精简、可写下的信息量」
derive /dɪˈraɪv/ v. 37:16
推导出、从中提炼得出
lasting debates n. phr. 37:53
长期悬而未决的争论
weights /weɪts/ n. 37:53
(神经网络的)权重;训练所得的参数
successive /səkˈsesɪv/ adj. 38:25
接连的、逐次的
supplement /ˈsʌpləmənt/ n. 38:25
(论文的)补充材料
gradient descent n. phr. 38:25
梯度下降;沿损失下降方向逐步调参的优化方法
scream at it phr. 39:29
(比喻)数据会「大声喊给它听」,即信号强到无需人工告知
specialty stuff n. phr. 40:05
专门定制的东西(specialty 作定语表「专用的」)
finite /ˈfaɪnaɪt/ adj. 40:05
有限的;与 infinite 相对
draw from it phr. 40:05
从中得出(结论、教训)
intricate /ˈɪntrɪkət/ adj. 40:36
精细复杂的、结构繁复的
shock no 1 phr. 40:36
(it should shock no one that…)这不该让任何人感到意外
impose /ɪmˈpoʊz/ v. 41:23
强加、施加(约束或结构)
constraints /kənˈstreɪnts/ n. 41:59
约束条件;限定解的取值范围
agglomerative /əˈɡlɑːmərətɪv/ adj. 42:52
凝聚式的、由小块逐级合并成整体的
associate /əˈsoʊsieɪt/ v. 42:52
(分子)结合、缔合在一起
Gaussians /ˈɡaʊsiənz/ n. 42:52
高斯分布(此处指用高斯粗略描述分子团的空间位置)
orientation /ˌɔːriənˈteɪʃn/ n. 43:28
取向、朝向;物体在空间中的角度姿态
sampling among phr. 43:28
在若干可能之间采样、抽取其一
ligands /ˈlɪɡəndz/ n. 43:57
配体;与蛋白结合的小分子(药物多为配体)
bond distances n. phr. 43:57
键长;成键原子之间的距离
progressive refinement n. phr. 44:35
渐进式细化;由粗到细逐步完善的过程
behavior cloning n. phr. 45:12
行为克隆;通过模仿示范数据学习策略的方法
coarse grained /ˌkɔːrs ˈɡreɪnd/ adj. 45:12
粗粒度的;抽象层级较高、舍弃细节的
get around phr. v. 45:12
绕过、规避(规则或障碍)
agency /ˈeɪdʒənsi/ n. 45:12
能动性;主动作出选择并施加影响的能力
ungrounded /ʌnˈɡraʊndɪd/ adj. 45:43
缺乏现实锚定的;符号未与外部世界绑定
regime /reɪˈʒiːm/ n. 45:43
(科技文中)状态区间、体制;指参数或条件所处的范围
seductive /sɪˈdʌktɪv/ adj. 46:15
极具诱惑力的、让人忍不住相信的
thingamajigger /ˈθɪŋəməˌdʒɪɡər/ n. 46:51
(口语戏谐)那个玩意儿、不知该怎么称呼的东西
disentangled representations n. phr. 46:51
解耦表征;各维度分别对应独立语义因素的表示
desperately /ˈdespərətli/ adv. 47:30
拼命地、竭尽全力地
token /ˈtoʊkən/ n. 47:30
词元;语言模型处理的最小文本单位
log linears n. phr. 48:05
对数线性关系;一个量的对数与另一量成线性
scaling laws n. phr. 48:05
缩放定律;模型性能随算力、数据、参数量变化的经验规律
trajectories /trəˈdʒektəriz/ n. 48:05
轨迹;此处指智能体连续多步的执行过程
deficiencies /dɪˈfɪʃnsiz/ n. 48:46
缺陷、不足之处
paper over phr. v. 48:46
掩盖、粉饰(问题而非真正解决)
harnesses /ˈhɑːrnəsɪz/ n. 48:46
(此处技术义)外围脚手架、支撑框架,用于驾驭模型
distill /dɪˈstɪl/ v. 48:46
蒸馏;机器学习中指把一个系统的能力压缩进另一模型
wrap it phr. v. 48:46
收尾、结束(节目或工作)
capacity building n. phr. 49:46
能力建设;发展援助用语,指提升本地机构与人员的长期能力
enteric /enˈterɪk/ adj. 49:46
肠道的(enteric bacteria 肠道细菌)
antibiotic resistant adj. phr. 49:46
抗生素耐药的
map out phr. v. 49:46
梳理清楚、绘出(机制或计划)的全貌
betterment /ˈbetərmənt/ n. 51:20
改善、增进(正式用词,常见于 for the betterment of)
scaled up phr. v. 51:20
扩大规模
prevalent /ˈprevələnt/ adj. 51:20
流行的、高发的(疾病在某地普遍存在)
testament to /ˈtestəmənt/ n. phr. 52:16
……的明证;be a testament to 证明了某事
frontier /frʌnˈtɪr/ n. 52:16
前沿;知识或技术的最外边界
hybrid /ˈhaɪbrɪd/ adj. 52:16
混合式的;由不同类型成分组合而成
proof of existence n. phr. 52:44
存在性证明;数学中只证明「有解」而不给出构造方法
理解自测 · 11 题
1. AlphaFold 出现之前,用实验方法解出一个蛋白质结构,大致需要多少时间和金钱?AlphaFold 把这个数字改成了什么?

按 Jumper 在「蛋白质是什么」一节给出的数字,实验解析单个蛋白结构的典型周期约为一年,折算成钱大约十万美元,而且这足以撑起一个完整的博士课题。他还提到,科学家往往要先在同步辐射光源这类小镇大小的装置上做实验,并且必须先攻克「结晶」这道可能耗时数年的门槛。AlphaFold 2 把这个过程压缩到五到十分钟,精度在典型情况下达到一个原子半径以内,已可与部分实验方法相比。更关键的是可扩展性:团队最终预测了约 2 亿个蛋白结构,覆盖所有已完成基因组测序生物体的蛋白。

2. 消融实验显示,SE(3) 等变性/不变性对 AlphaFold 2 的提升贡献了多少?Jumper 认为真正被忽视的是什么?

AlphaFold 2 在 CASP 的 GDT 指标上比 AlphaFold 1 高出约 30 分,而消融实验显示,去掉不变点注意力(IPA)带来的等变/不变性只损失约 2 到 2.5 分,即不到总增益的十分之一。Jumper 特意在论文中放入名为「no IPA」的那一行,以为这能终结争论,结果社区仍把 AlphaFold 2 说成「等变性的伟大胜利」。他认为被完全忽视的是 FAPE(框架对齐点误差)这个损失函数——在每个残基的局部参考框架下衡量其他所有原子的位置误差再取平均,他称其为早期真正的突破之一。他还指出,真正强有力的对称性其实是残基的置换不变性,而非全局 SE(3)。

3. Midnolin 这项研究是怎么使用 AlphaFold 的?结果如何?

研究者先通过遗传学实验发现,有几百个基因在细胞发育某阶段被关闭,并定位到一个此前几乎无人研究的人类蛋白 Midnolin:敲低它,那些蛋白就不再被降解,且作用方式不走常规途径。他们随后把 AlphaFold 与约 500 个响应该敲低的蛋白一起批量运行,在其中约 40% 里发现同一个高度特异的模式——底物的某一段被 Midnolin 的两个部分像钳子一样夹住。接着他们做湿实验验证:删掉 AlphaFold 指出的被夹位点后,该蛋白在细胞中不再下降。十个例子中九个完全符合;剩下一个只是部分减弱,回看预测发现存在第二个结合位点,双位点都去掉后降解被彻底消除。这是把 AlphaFold 当作假设生成器、再由实验闭环验证的典型流程。

4. Jumper 如何区分「预测」「控制」「理解」这三件事?为什么他说机器目前只做到前两件?

他在访谈后半段给出明确定义:预测是说出未来某次测量会得到什么值;控制是我希望未来测到的那个值等于 17,并使之实现;理解则很像预测,但中间多了一个人——理解意味着支撑预测所需的事实少到可以压缩、可以传达给另一个人,能写在一张索引卡上。按这个定义,机器能预测、也许能控制,但「压缩到可传达」这一步仍必须由人完成,所以机器不能替我们完成理解。他随即补充了一个新可能:如今我们可以在 2 亿个预测结构上做计算实验,而不再只有 20 万个实验结构,这为人类自己提炼理解提供了新的实验对象。

5. Jumper 为什么反对把 AlphaFold 3 简单叫做「一个扩散模型」?他给出的结构性理由是什么?

他说自己的反对不在技术层面,而在叙事框架层面。理由有三层:第一,扩散模块之前有一个巨大的主干(Pairformer),只运行一次,且完全不是扩散模型;结构其实主要在那里被确定。第二,扩散在这里扮演的角色与 AlphaFold 2 的「结构模块」同构——接收一组已相当好的约束,然后把细节几何化补齐,而非像图像扩散那样先画色块、再决定色块含义。第三,AlphaFold 2 的求解顺序是凝聚式的:先局部、最后全局;而 AlphaFold 3 的扩散必须最先确定全局摆放(两个蛋白如何相对结合),顺序正好相反,说明它并不遵循图像扩散那种渐进细化逻辑。他还指出选择扩散的真实动因是工程性的:便于处理配体、键长与局部化学。

6. AlQuraishi 实验室的复现实验说明了什么?它如何支撑 Jumper 对「苦涩的教训」的批评?

该实验重新训练 AlphaFold 2,但只用 PDB 的 1%,即约一万五千个结构而非十几万个,结果这个数据严重受限的 AlphaFold 2 仍比用全量数据训练的 AlphaFold 1 更准。Jumper 由此得出量化结论:AlphaFold 2 的架构与训练思路实打实抵得上 100 倍的数据量。这直接构成对「苦涩的教训」流行解读的反驳——如果算力和数据终将压倒人类先验设计,那么架构改进不该等价于两个数量级的数据。他的论证前提是数据有限性:PDB 是有限的,转向语言模型后发现互联网也是有限的。因此他的结论不是「Sutton 全错」,而是「不要做架构研究」这个推论是错的,正确态度是判断哪些先验该写进代码、哪些交给数据。

7. 主持人提出「机器学习的教训是直觉不可靠、要靠大量数据试错,那机制式的药物干预是否也像打地鼠」。Jumper 是如何回应的?

他没有正面否认代偿机制的复杂性,而是把回应落在 AlphaFold 的自我限定上,称之为「AlphaFold 的谦逊」:他们从不宣称建模整个细胞或整个生命,只宣称预测一个具体的、你天天在做且要做一年的实验会给出什么结果。这一点带来两个后果。其一,有效性可以被精确刻画——因为对标的是明确定义的实验,误差能被度量,不需要靠对生物学的整体正确性来担保。其二,超出承诺的用途是用户自己发现的:跑上千次预测去找哪两个蛋白结合、找出复杂体系里的未知组分。换言之,他用「缩小承诺范围」来化解「整体系统不可预测」的质疑,而非声称机制理解足以驾驭活体的复杂性。

8. 「一万美元的四分之一圈」这个笑话在论证中承担什么作用?它和 AlphaFold 的定位有什么关系?

笑话里技师只拧了螺丝四分之一圈就让工厂复活,账单一万美元,其中拧那一下值五毛,剩下的都是「知道该拧哪儿」。Jumper 用它说明药物研发的价值几乎全在机制知识,而非操作动作本身。这与他对 AlphaFold 的定位直接呼应:AlphaFold 让「看清零件长什么样」变得几乎免费,但并不告诉你该动哪个零件。他明确说过,阻碍治愈自闭症这类疾病的原因不是缺某个蛋白的结构,而是我们对生物学如何运作了解得太差。所以他把细胞比作「庞大复杂的工厂」,我们既在学怎么拧(造药),也在学该拧哪儿(机制),后者才是瓶颈。

9. Jumper 说 AlphaFold 学到的算法「恰好是人类早就想出来的算法」。这个说法对「代码与数据谁做的工作更多」这一争论意味着什么?为什么它未必普遍成立?

他指出 AlphaFold 学到的是逐次几何精修——一个可以用几句话讲清、听者能在脑中想象出来的算法,可它并不是被编程进去的,而是训练中出现的(recycling 等架构手段只是加速了这个它本来就在学的过程)。这对「代码 vs 数据」的争论意味着:得到的算法可以既非人工写入、又恰好落在人类可理解的范围内,所以「学出来的」不等于「不可理解的」。但这未必普遍成立,因为蛋白折叠有强几何结构、目标近乎单峰(一个正确答案或很窄的分布),恰好与人类的空间直觉合拍。他自己也在别处承认,我们造的多是「外星造物」,很多直觉不成立;他甚至说若要造通用生物学机器,它会更像语言模型而非窄预测器——那时学得的算法未必还能压缩成索引卡。

10. 如果有人反驳说「Jumper 转投 Anthropic 这一举动本身,就证明通用模型路线赢了、专用架构那套经验没用了」,按他在访谈中的立场,他会怎么回应?

按他的表述,他大概不会承认这是路线之争的胜负判决。他反对的从来不是通用模型,而是「因为它是 Transformer / 扩散模型所以有效」这类用高层标签替代解释的做法——他明确说,这种标签解释不了聊天模型三四年里的巨大进步,也解释不了研究者每天在做什么。他对通用路线给出的判断是条件式的:若真要造通用生物学机器,它会更像语言模型;同时他又强调数据有限(互联网也是有限的),所以架构研究不该被废除。他真正指出的未解难题也落在通用侧:我们能用外部记忆等脚手架掩盖模型缺陷,却还不知道如何把脚手架蒸馏回权重。因此更贴近他立场的读法是,他带着「哪些先验该进代码、哪些交给数据」这套判断力换了战场,而非宣布这套判断力已经过时。

11. AlphaFold 数据库对非洲结构生物学的影响,能否推广为「开放数据能自动消除科研资源差距」?请结合 Emmanuel Nji 的经历评估。

他的经历确实提供了极强的正面证据:同一个蛋白他此前试了四五年未果,用 AlphaFold 后只做一次纯化、收一次数据,不到两三个月就拿到结构。但把它读成「开放数据自动消除差距」会漏掉两个必要条件。一是他仍需要实验能力:他用 AlphaFold 是去解析冷冻电镜数据、梳理蛋白机制,即预测与实验结合,而非纯计算替代。二是需要人的中介:主持人特别强调 Nji 不只是让人能用上 AlphaFold,还在教怎么使用、怎么解读结果、怎么据此设计实验;今年在 Google DeepMind 与瑞典研究理事会资助下扩到 100 人,目标是十年内近 1000 人。也就是说,公开数据把成本极高的一环变得几乎免费,但把它转化为成果仍依赖资金、设备与培训网络——这与 Jumper 反复说的「AlphaFold 是起点,不是终点」是同一个结论。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.126Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again 下一期 · NO.128 →CHM Live | The Silicon Gold Rush: How AI is Driving the Development of New Chips
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY 内容仅供学习 · thesophielab.com