视频库 / NO.010ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.010
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

ICM 2026 Public Lectures - Terence Tao

节目发布 2026-08-13 · Simons Foundation
陶哲轩 主持人 Alex
本期追问 · 点击跳到视频对应位置
1:42 一百年前的数学基础危机,能为今天的 AI 冲击提供什么启示?32:39 当机器能产出证明,数学家凭什么判断一个结果值得信任?27:05 形式化与 AI 协作,会把数学追求的「理解」带向何处?13:37 数学从个体手工走向大规模协作,问题的选择标准会怎样改写?
归入 Ⅴ·05 研究是怎样做成的? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
本文为 2026 年国际数学家大会(ICM)公开讲座的现场实录。主讲人是加州大学洛杉矶分校数学系杰出教授、James and Carol Collins 讲席教授陶哲轩(Terence Tao),2006 年菲尔兹奖与 2014 年科学突破奖得主;开场致介绍词并主持本场讲座的是 Alex,讲座中「正典化」(canonicalization)这个说法即由他提出。陶哲轩在这场讲座里没有去论证 AI 到底能不能做数学,而是先假定它能,再追问:那么数学共同体的目标与价值究竟是什么。全文依据现场录音编译整理,仅删去口语枝节、寒暄与重复,论证、例证与语气一概保留。

开场介绍

主持人: 欢迎各位来到国际数学家大会公开讲座。今晚由我来介绍我们的主讲人陶哲轩,这对我来说是极大的荣幸。陶哲轩在普林斯顿大学师从 Eli Stein 获得博士学位,此后前往加州大学洛杉矶分校,一直任教至今,现在是该校数学系的杰出教授,以及 James and Carol Collins 讲席教授。如果我试图把他所有的研究成就和荣誉列一遍,我们今晚就都耗在这儿了,你们一个字也听不到他本人说话。所以我只提一句:他在此前的一届国际数学家大会上获得了菲尔兹奖,此外还有科学突破奖,以及许许多多其他奖项。讲座结束后我们会留出提问时间,有十分钟的问答环节,地点在外面右手边的前厅。现在,请和我一起欢迎陶哲轩教授。

陶哲轩: 谢谢,非常感谢。能够再一次亲身参加国际数学家大会,我实在高兴。距离上一次已经很久了,这几天我见到了许多老朋友,也一直在吸收这里的正向能量。

我今天要讲的,大概就是眼下所有人都在谈的话题:人工智能的时代,以及它对数学意味着什么。我们似乎正生活在一个剧变的年代,而这一点我们可能并不习惯,因为数学在过去一个多世纪里一直非常稳定。我们没有遇到过真正意义上的生存性危机,可现在,一切忽然开始大幅改变。

这看起来是前所未有的。但我认为,我们正在经历的事情,其实是有历史先例的。

一百年前的危机

陶哲轩: 作为一段引子,我想先讲讲一个世纪前发生的那场基础危机。

在二十世纪初之前,数学家的工作方式是半形式化的。我们有欧几里得,我们知道什么叫证明,但我们并没有把逻辑、集合、分析的基础彻底公理化,很大程度上用的是一套朴素而非正式的基础。像「集合到底是什么」「数是什么」「函数是什么」「极限是什么」这类问题,我们有一些规则,可是从来没有把这门行当的全部规矩完整地写下来。

这套做法一直管用,直到二十世纪初。当时只有哲学家在认真思考这些问题,而正是他们指出,朴素的做法其实藏着严重的毛病。比如哲学家伯特兰·罗素指出,朴素集合论是不相容的:你可以构造出一个集合,它既包含自身又不包含自身。当然还有哥德尔的不完备性定理:不存在任何一套公理系统,能够把你在算术里能提出的所有问题统统证明或者证伪。

于是我们被迫真正地去思考自己的基础,去追问什么是证明,我们这门职业的规则到底是什么。那个过程是创伤性的,它被称作基础危机,充满动荡,争论不休。

但我们最终从中得到了非常宝贵的东西:一套做数学的标准框架。我们有了一套正统的规则,一阶逻辑,策梅洛-弗兰克尔选择公理集合论(ZFC),它是标准化的,而且我们达成了共识,认为它足以支撑几乎全部的数学。这当然不是最终结论。人们仍然在研究别的基础,把它们与 ZFC 作比较。举个例子,Lean 这个形式化证明助手所依据的并不是 ZFC,而是依赖类型论。但这没有问题:我们研究别的基础,并不是站在危机之中去研究,而是像研究任何一个科学或数学对象那样去研究它。这些基础经受了几百年的严苛检验,我们信任它们。理论上当然仍有可能,这些基础里藏着某个严重的问题,可它们看上去是管用的,而且已经管用了一个世纪。

这一次危机在价值

陶哲轩: 我要提出的论点是:那场危机过去一个世纪之后,我们正在进入一个同样动荡的时期。顺便说一句,在座如果有人喜欢在文本里挑长破折号,我要强调,这一页幻灯片上的那个长破折号是人手打出来的。

但这一次,危机不在我们的数学论证里,而在我们的数学价值和实践里。我们此前一直是在一套朴素的基础上运行的,也就是关于「数学究竟是什么」的朴素理解。而突然之间,这套朴素基础推导出了一些相当古怪的结论。我们确实需要重新审视它,把它认真地重建一遍。同样,这会是一次非常值得的经历。一旦做完,我们这门学科会健康得多,也韧性得多,我们才能够把这些新技术妥当地吸纳进来。

那么我这场演讲的主题是什么?我可以把驱动一切的那个大问题,叫作「共同体回应之问」:面对现代 AI 技术以及随之而来的种种事物,数学共同体应当如何回应?何况这些技术已经具备,或者声称具备完成数学任务的能力。

这是个大问题,而且是整个共同体的问题。没有任何人,哪怕是国际数学联盟,能够定下一套规则让大家照办。我们必须真正地辩论这个问题,就像当年必须辩论数理逻辑的基础一样。不过我确实有一些看法,可以拿来开个头。

这不是一道数学题。它不是一个猜想,不会有一个证明和唯一的答案。它是一个元数学的问题,同时也是政治问题、伦理问题、文化问题、经济问题,随你怎么归类。可是我认为,那种极其审慎、极其数学化地思考问题的心智习惯,在这里非常有用。所以我打算用一种伪数学的方式来处理它:我会借用数学的语言来分析这个并非数学的问题。既然在座各位大多是数学家,我希望这样能把我想说的点讲得更清楚。

关于能力的猜想

陶哲轩: 我们关心的是共同体回应之问:面对 AI,我们该怎么做。而在它之下还有一个子问题,过去三年的争论几乎都被它占满了,虽然它其实是另一个问题。我把它叫作「AI 能力猜想」。既然说了要用数学的语言,我就把它称为猜想。严格说,它不是单个猜想,而是一族猜想。

我把它写成一个模板,里面有若干占位符,或者说变量,如果你愿意继续用数学的说法的话。因为变量太多,这个猜想几乎可以表达任何意思,但我还是把它念出来:在不久的将来的某个时刻,某些 AI 工具,以某种成本、在某种程度的人类监督之下,能够在数学的某些领域完成某些研究级别的数学任务,具有某种非平凡的成功率,以及某种程度的正确性与质量。

我在里面塞进了太多个「某种」,所以它几乎可以是任何东西。近年来大量的争论,其实已经退化成了在各个空位里究竟应该填进哪个「某种」。而这不是我今天要讲的。我认为这反倒把我们的注意力从一些更根本的问题上引开了。所以,就把它当成一族猜想。你当然可以把「某些」填成「极少数」,让它变得非常弱;你也可以把它填得非常强。各人可以在自己脑子里玩这个填词游戏。我下面只笼统地说这个猜想的强形式、弱形式和中间形式。

知道这族猜想的答案显然是重要的。比方说,如果连它的弱形式都不成立,也就是说这些工具对研究级别的任务基本无用,那就没什么可争的了:它们只是新奇玩意儿,一小撮人可以拿来玩玩,数学基本上照旧运转。反过来,如果最强的形式成立,我们今天所做的每一件事,AI 都能做得更好更快,那么显然,我们不可能一切照旧。尤其是当我们把注意力集中在 AI 最擅长的那些事情上时:解决未解决的问题看起来正是它们的强项之一,而这也正是我们大量投入的事情之一,我们甚至为此颁发奖章。那我们就不得不认真思考我们的文化和实践了。而如果真相处在中间地带,我们要想得更细。

正如我说的,你去看媒体上的争论,去听休息室里喝咖啡时的争论,去看网上的争论,会发现由于共同体回应之问的答案高度依赖于 AI 能力猜想的答案,绝大部分争论都落在了能力上。我自己当然也参与其中。回看过去三年,我谈 AI 谈得很多,其中很大一部分是谈能力,因为它确实重要。三年前有那么一段时间,AI 的能力很差,但趋势是会好得多,我们当时必须把「现在在哪里」和「技术正在往哪里去」这两件事区分清楚。不过今天这场演讲不谈这个猜想。

如果你上网,一定见过无数的宣称、反驳与争吵,围绕着这个猜想在各个强度上的真假。我们现在有很多数据点:某个特定问题被解决了;然后又有人说,这个 AI 连两位数乘法都做不对。可我们手上的绝大部分数据,并不是在恰当的科学条件下采集的。我们看到的是被选择性披露的结果,看上去很惊人,但我们不知道输入到底是什么,不知道真实的失败率是多少。而且许多披露这些结果的公司自有其动机,会尽可能把结果呈现得漂亮,同时略去一些重要信息,比如产出这些结果究竟花了多少钱。所以眼下这一整摊支持或反对该猜想的数据点,是一团非常令人困惑的乱麻。

还有一点:人们有时会把「这个猜想是真是假」和「我们希望它是真是假」混为一谈。而我认为,这两者未必是一回事。

一份可信的数据

陶哲轩: 所以我不会谈这个猜想。我唯一想给各位指出的是,我们开始有了一些更系统、更科学的 AI 能力评估。其中我认为最有前景的,是许多人都听说过的 FirstProof 挑战。

这是一项由数学家自发组织的努力,是对今天 AI 模型状态的独立评估,不由任何一家大型 AI 公司赞助。它的做法是:数学家们捐出一批十道研究级别的问题,这些问题的解答没有发表过,问题本身也没有发表过,然后拿它们去测试完全自主运行的 AI 模型。比如今年五月,一共有四套系统提交参测,我参与的团队提交了其中一套。测试的时候,他们不只问这些问题是否被解决了,还要问解答是否被写出来、被解释到了可以发表的质量:是不是真的合理拿去投给某本期刊的论文,也许只需要一点小修改。

据我所知,单套系统的最好成绩大约是十题中的五题,而四套合起来,能有七题达到可发表的质量水平,代价却不低:每套系统在每道题上的花费,从十美元到一千美元不等。这大致就是五月时的状况。他们会继续测试新的批次,这些数字也许会变,但我认为这是我们目前手上最科学的一个数据点。

一个条件性的假设

陶哲轩: 好,这个猜想我不再谈了。我要谈的是它的补集:假设这个猜想在某种意义上是真的,那么我们该怎么回应?

于是,像数学家惯常做的那样,我要提出一个假设。我把前面那一堆「某种」全部换成「合理的」:一个合理强度的猜想成立,也就是说,在合理可期的时间内,AI 将能够以合理的成功率、质量和成本,完成数学任务中合理的一部分。我不会定义「合理」是什么意思,各位可以在心里自行设定。

这是一个假设。在数学里我们常常提出一个不知真假的假设,但我们对它的推论感兴趣。我要做的是条件分析。你尽可以去争论这个假设是真是假,那不是本场演讲的重点,我只是有条件地采用它。

一旦你以此为条件,也就是接受 AI 真的能够解决我们今天作为数学家所做的相当一部分工作,那么有一件事就变得很清楚:要回答回应之问,我们现在必须去思考我们的目标和价值。我认为这才是今天真正需要讨论的问题,我说的「今天」不只是字面上的今天,而是我们所处的这个时期。

目标与价值之问

陶哲轩: 我把它叫作「目标与价值之问」:我们这个共同体,以及数学研究本身,精确的目标、任务和价值究竟是什么?

我们当然有明确的目标。发新闻稿的时候,写论文摘要的时候,申请科研经费的时候,我们会列出「本项目的目标是如何如何」。那些是显性目标。但我们正在发现,其实还有大量隐性目标,我们对它们谈得太少了。而现在,把它们谈清楚变得非常重要。

我来解释为什么。过去我们主要关注技术性的目标:我想改进这个常数,我想证明这个定理,我想找到一个满足某某性质的定义。至于价值、目标,以及「数学的意义是什么」这类问题,我们一向觉得那是人文学科的事,也许该由数学哲学家来处理,或者科学社会学家、数学教育家。我们也许过度地偏重了技术层面,因为那是我们最擅长的,而且在某种意义上更好操作,尽管数学本身往往相当艰难。

但我要说,如果你接受这个工作假设,我们就再也没有「不必把目标讲清楚」这份奢侈了。而且我认为这是健康的。无论这个假设是真、是假,还是部分为真,我都认为我们早就该谈一谈:我们为什么做数学,这门学科真正的意义是什么。谈过之后,我们只会更强。

那么,我们为什么做研究?我们为什么要去解决问题、探索数学概念?理由其实非常多,我不认为有谁列出过一张完整的清单,下面只是其中一部分。

当然,我们要解决尚未解决的问题,这很重要,既包括纯粹的问题,也就是在数学之外没有直接应用的问题,也包括许多重要的应用问题,那是数学非常重要的一种动力。我们喜欢发展新工具、新理论、新方法。我们希望更好地理解数学。我们希望理解我们周围的世界,世界充满了模式,正如我们在汉娜·弗莱(Hannah Fry)的讲座里看到的那样,而我们想理解这个世界。

但数学同时也是一门高度社会性的学科。我们对数学感到兴奋,很大一部分原因在于我们能把它传达给共同体的其他人。我们想建设这个共同体,我们想培养下一代,想让这个共同体延续下去。我们投入大量精力去提携和教育下一代数学家,而他们将做出我们做梦都想不到的事情。我们还拥有一个积累了几千年的数学知识库,这大概是这颗星球上持续发展时间最长、最具累积性的知识体系之一,而我们在为它添砖加瓦,这是了不起的事。此外还有审美价值,我们常常低估它,但它同样重要:数学是美的,你可以创造出让人欣赏几个世纪的东西,这本身就是一种价值。这份清单留个作业给各位,你们可以再添两条。

古德哈特定律

陶哲轩: 我们有这么多目标,可我们往往每次只讲其中一两个。直到不久之前,这样做都还算可以,因为这些目标是彼此对齐的。它们略有差别,但数学很难,我们离其中许多目标都还很远:解决问题难,构建理论难,什么都难。只要你离所有目标都很远,只要它们之间大体正相关,那么你朝其中一个目标推进,也就同时在朝别的目标推进。因为这种正相关,你可以只把一个目标说出口,让它充当其余目标的代理。你不必声明「我要做目标一,其实我私下里也想要目标二」,因为你知道,往目标一推进,你离目标二也会更近。所以我们从来不需要把所有目标都讲明白。

这一套一直管用,直到你开始大力优化,直到你在目标一上变得非常擅长。到了某个点,你在目标一上再往前走一步,实际上就在远离目标二、三、四、五。这个现象在经济学里非常有名,叫古德哈特定律:当一个衡量指标变成目标,它就不再是好的衡量指标。如果你为某个特定指标过度优化,代价就是你原本也想优化的那些别的目标。

对人类来说,这倒还不算太大的问题,因为人类并不那么擅长优化。但 AI 擅长,AI 公司也擅长,而这是一个危险的组合:AI 可以脱离现实的牵制去做优化,因而格外容易落入古德哈特定律的陷阱。

于是我们正在进入这样一个时期:我们的各个目标开始指向彼此竞争的方向。我认为 AI 尤其可能让我们更接近其中某一个目标,而代价是别的目标。所以我们再也承受不起每次只盯住一个目标的做法了。而对于我们确实要盯住的那些目标,我们必须非常清楚地说明,评判它们的标准到底是什么。这就是我们眼下所处的处境。

解题这件事

陶哲轩: 我没有时间把所有目标都讲一遍,所以我只挑其中一个来讲,而它恰好是近来吸引了最多注意力的那一个:解决问题。

在外部世界看数学的时候,我认为他们其实并不真的知道我们在做什么,但在他们看来,我们似乎特别痴迷于解决未解决的问题。所以我就来谈这个侧面。我要再次强调,这只是我们诸多目标中的一个,但我会把它拆解开来,单就这一个目标说几点。我们并不是全都在拼命多解题,许多数学家更愿意去构建优美的理论,或者去教育下一代,这些都是极重要的目标。我之所以聚焦于解题,是因为在这个工作假设之下,它是当下受 AI 冲击最大的那一个。

那么这里的目标是什么?如我所说,从外面看,数学的目标就是尽可能多地解决未解决的问题。我想起自己念高中的时候,完全不知道数学家可以是一份职业。等我知道真有这样一份工作之后,我脑子里的画面是:有某个资深数学家组成的委员会,会给大家分派问题,我们就像做家庭作业一样各自去做被分派的题目,整个体系就是这么运转的。当然事实完全不是这样。但外界的印象有时候确实差不多就是如此。

那么,还是用伪数学的语言。你可以把它看成一个流网络优化问题:我有一个问题的源,一个解答的汇,中间是证明生成的流,我们要做的是把从未解决问题到解答的流量最大化。

如果你真去优化这个目标,那么早在 AI 出现之前,我们就已经发现它有毛病。假如我们把目标定为「收集尽可能多的黎曼假设的证明」,很快就会明白这不是个好目标,因为你会收到大量对该问题的、并不正确的解答。好,这有个显而易见的修补办法:先生成证明,再检验它是否正确。于是我们的流变成了两步:从未解决问题出发,生成证明,得到未经验证的解答,然后经过一道独立的证明验证工序,得到已验证的解答。

而 AI 在这条流的第一步上已经变得非常好,不是对所有问题,但对某些问题相当不错;在第二步上也做得不错。事实上第二步的进展主要不是靠 AI,虽然 AI 有帮助,而是靠一项互补的发展:证明助手语言。那本身就是一个大话题,会有别的报告专门讲,那也是近年来另一项重要进展,证明助手语言的巨大改进。

所以看得见,在某些领域、对某些类型的问题上,数学中的证明生成和证明验证都加速了。不是普遍加速,但确实加速了。而如果我们以这个工作假设为条件,它还会继续加速。

证明需要被解释

陶哲轩: 但这带来一个新问题。设想现在 AI 在生成这些证明,甚至证明是用 Lean 这类形式语言写的,并且已经验证通过。我们于是有了一份十万行的证明,验证为正确,可是没有人理解它。哪怕是输入提示词把它生成出来的那些人,恐怕也不理解它。

这已经在某些领域发生了。也许你们听说过厄多斯问题库,那是一份大约一千二百道问题的清单,有些已解决,有些未解决。就在过去半年里,证明开始被生成出来,其中一些甚至被形式化进了 Lean,而没有任何人真正通读过它们,也没有任何人能为「这是一个好的证明」背书。很多情况下,连提交证明的人自己都说:这里有一个证明,但我没有资格评判它,我不知道它对不对。然后另一个人说:我在 Lean 里检查过了,可我同样不能替它担保。

这样的局面还没有完全出现,但我们已经非常非常接近这样一个场景:某个重要结果被证明了,被验证了,而没有任何人类能够理解它、解释它。那会是一个很不受欢迎的发展。

这意味着,我们的目标至少还要再加一个阶段。仅仅生成证明是不够的,仅仅验证证明也是不够的,证明还必须被解释到足够好的程度,好到能够被数学共同体传达和理解。它们不能只是一件件人工制品,一项项看上去很了不起、却不能引出后续数学的成就。它们必须被理解。所以我们需要产出写得好的解答。

目前,AI 虽然在加速前两个阶段,但在第三个阶段仍然相当弱。它擅长阐述的某些部分,不擅长另一些部分。举例来说,在最基础的层面,拼写、语法、格式,它们做得完美,甚至完美到人们现在反而更愿意读稍微有点瑕疵的文本。

但问题不只是拼写正确之类的技术层面。AI 写出来的东西,重点常常放得非常古怪。AI 生成的数学读起来很叫人抓狂:你想找出全篇最重要、最困难的地方,那往往才是一个证明里最有意思的部分,结果你发现,AI 会花整整三页去证明某个显然成立的平凡引理,然后用三行打发掉论证中真正有意思的那一段。

另外,它们往往不交代影响来自何处。这件事本身很间接:证明来自一堆权重,权重来自某些训练数据,某个先前的结果在某个环节间接地影响了这个结果。可这些 AI 生成的文本常常完全不透露这些关联从何而来。于是这些结果读起来,与文献的联系比它们本应有的要疏离得多。

也许这会变好。如果你假定 AI 能力持续提升,它们在阐述上确实在慢慢变好一点。但这件事更难,因为验证一个证明有清晰的信号,非真即假,你可以设定一个好的评分函数并加以优化,而这正是 AI 工具擅长的;至于一个证明写得好不好,我们还没有真正好用的评判标准。也许将来会有,也许证明会变得更容易读。

自然的摩擦

陶哲轩: 但那也未必是好事。有时候一个证明可以太过顺滑。如果你读过那种把所有困难都打磨得一干二净的证明,也许有过这种体验:整个证明看上去毫不费力,可你回家拿纸笔想把它复原出来,却做不到,因为它把困难藏起来了。

证明里其实有必要保留一点我称之为「自然摩擦」的东西。当一个人去证明某件事,有些步骤是容易的,他会一路滑过去,不在上面花太多时间;有些步骤是真的难,他会停下来,花些时间,仔细组织论证,设法不做多余的功夫。读者能感觉到这一点,能从中读出这个证明的难点在哪里。而 AI 呢,容易的地方它一冲而过,困难的地方它也一冲而过,看上去毫无分别。

而且吊诡的是,人在阐述中犯下的错误,有时反而能帮到读者。这听起来矛盾,但确实如此。

我举一个我自己的例子。这是我手里的一份论文复印件,让·布尔甘(Jean Bourgain)1991 年的文章,我读研究生时读的,讲的是 Kakeya 猜想和限制性猜想。事实上,昨天报告中王虹的工作,有很大一部分正是建立在这篇论文之上。对于从没读过布尔甘论文的人来说,那是一次极具教育意义的体验。

我到今天还留着我研究生时代自己做的批注。他会写「显然可以这样做」,或者「这个陈述本质上等同于那个陈述」,而我在旁边打满了各种气急败坏的问号。你们大概看不清,但这里有一条批注写着:我恨让·布尔甘。

我硬是杀出了一条路。我很幸运,Eli Stein 给我讲解过其中一些,Tom Wolff 也讲解过,最后我明白了他在做什么。我终于摸到了他思考的方式,也终于明白,那些在我看来并不等价的东西,为什么在本质上就是等价的。我学到了太多东西。以至于几年之后,比起所有别人的论文,我更愿意读布尔甘的,因为我懂得他的头脑是怎么运作的,他总是直取要害,读起来是一种享受。但这中间确实有一个学习的过程。我想,如果他的证明先经过许多层 AI 的润色,我大概就得不到我当年得到的那份教育了。

所以这不只是把论文写好的问题,而是真正消化和理解结果的问题。比尔·瑟斯顿有一句著名的话正是这个意思:尽管表面上看去如此,我们并不是在完成某种抽象的生产定额,不是在拼定义、定理和证明的产量。我们真正在推进的,是人类对数学的理解。

最稀缺的资源

陶哲轩: 那么假设一个证明写得非常好,保留了全部的自然摩擦,任何人读了都能有所收获。这还是不够。

一个结果若要真正影响数学未来的发展,它必须被共同体接受和看重。别的数学家必须真的愿意去读它、消化它,把它放进自己的工作里。当然,让它正确、让它好读,是有帮助的,但这些只是必要条件,并不充分。这有点像做菜:你可以做出一桌漂亮的饭菜,香气扑鼻,所有食材都验证过可以安全食用,可你不能逼别人吃,你也不该逼别人吃。吃不吃是他们的选择。所以你必须去说服别人,这个结果值得一读。

在很近的过去,这都还不成其为问题,因为证明是稀缺的。某个结果一旦被证出来,这个领域里的其他人自然会放下手头一切去读那篇论文,他们有充足的动力。就像食物稀缺的时候,端上桌的东西你都会吃。可现在我们被证明淹没了,多到超过我们能审读的量,我们不得不挑挑拣拣。于是我们需要被说服,才会去读一个证明。证明的审读,如今成了这个新时代里最宝贵的瓶颈之一,一种稀缺资源。

你可以设法吸引别人来读你的工作。讲一个好故事是有用的:给出叙事,展示过程,说明你是怎么走到这个结果的,哪些路走不通,你又是怎么绕过去的。人是爱听故事的。我们过去不强调过程,只让结果自己说话。这一套一直管用,直到我们找到办法,可以不经过程而自动产出结果。

而当下的这些工具,对自己的过程极其不透明。有所谓的思维链,你可以稍微掀开引擎盖看一眼,但那并不怎么有启发性。也许这一点可以靠好的实践习惯改变。可眼下,尤其是当这些证明来自专有模型或者内部模型,而公司又有动机把某些事实当作商业机密的时候,我们就是看不到过程。这真的让我们大大降低了投入时间去了解它们的意愿,因为我们看不到那个故事。

审稿人与期刊

陶哲轩: 所以解题这件事还有第四个阶段:不只是生成证明、验证证明、解释证明,它们还必须被接受。

在我们现行的体系里,让一个结果被共同体接受的标准途径,是发表在一本有声望的期刊上。它必须被数学共同体消化和接受,而我们为此建立了一整套制度:我们有期刊。但它慢。你可以去推动它,可以把论文写得更有胃口,可你没法靠 AI 来优化这一步,因为它牵涉的是人的反应。除非你想办法把人接进 AI 里去,而这我实在不推荐。所以这是流程中慢得多的一段,也是 AI 目前基本上没有产生影响的一段。这里仍然是人对人说话,才能换来那份接受。

现在的做法是:一篇论文验证完毕,够格投给期刊,我们就把它送到一位人类审稿人手上,而他自愿贡献时间,把审稿当作一种服务。他们审稿,不只是因为结果有趣,也是因为想鼓励作者成为更好的数学家。这是提携新一代数学家的方式之一。这份工作没有荣耀可言,我们刚刚为解决问题颁出了四枚奖章,可我们没有为审稿颁出四枚奖章。但它是我们这门职业不可或缺的一环,正是它把数学家个人的成就转化为集体的理解。审稿人只是一种代理,真正的机制是我们互相阅读彼此的论文,学科就是这样往前走的。

但它慢,而我们正面临一场危机:AI 生成的论文如此之多,哪怕它们都是正确的,也可能多到审稿人读不过来。这是一个持续存在的问题,我想会有别的论坛专门讨论对策。我认为期刊也将不得不开始采用 AI 工具做初筛,各刊将不得不把风格指南和各类规则规定得精确得多,好让任何一篇论文在到达人类审稿人之前,先经过一层额外的评估。这会有争议,但我认为是必要的。不过我们绝对不该把人类审稿人从流程中拿掉。我不认为一本纯粹依靠 AI 审稿的期刊会成功,尤其是因为这类工具相当容易被钻空子。如果有本期刊宣布,这篇论文已被一百位 AI 审稿人接受,那并不是共同体的接受,那仍然不意味着我们想读它。

正典化

陶哲轩: 那么这就是我们的最终目标了吗?连这个都还不是真正的终态。

哪怕一个结果已经被共同体接受,已经发表在有声望的期刊上,所有人都认可它是正确的,这仍然不是终态。终态是这样的东西:教科书,我们在课堂上讲授的材料。这个概念的标准定义是什么?各个结果应当按什么顺序来证明?如何把发表在好刊物上的一个个孤立结果,组织成一套融贯的理论?那才是数学最终抵达的状态,才是我们试图解决的所有这些问题的终点。

给我做介绍的 Alex 为此起了一个很好的名字:正典化(canonicalization)。

所以除了消化一个证明、发表一个证明之外,你还要让它成为正典。而这是所有阶段里最慢的一个。过去十年、二十年里有非常多的结果,登在最好的期刊上,却还没有进入教科书,还没有被教给学生。我们还没有找到真正确定的、正典的、正确的方式去讲授这些东西。这个过程很慢:你得去开课,得收集反馈,得非常用力地去想,究竟该怎样把大量彼此不同的结果组织起来,纳入一条融贯的叙事线。而且它需要共识:如果一个领域里有一半数学家觉得应该这么讲,另一半觉得应该那么讲,那就谈不上正典化。

而这恰恰是 AI 基本上完全无能为力的阶段。我几乎看不到 AI 在这一段流程里有什么角色。可它是最有价值的一段。如果你想把数学的某个子领域应用到数学的别的领域,或者应用到某个实际问题上,它基本上必须先被消化成这种形态。你要让一个数学分支对工程师、物理学家、生物学家有用,他们是不会去翻《数学年刊》上最新的论文的,他们要的是教科书。

数学最有价值的那些应用,只有在抵达这个最终阶段之后才被解锁,这里面也包括 AI 自身。AI 之所以在数学上如此成功,一个很大的原因,正是几个世纪以来我们一直在建立这些正典的定义,我们有讲线性代数怎么运作、群论怎么运作的教科书,我们有大量非常成熟的理论。AI 把这一切都吸收了,它靠的正是这些东西才取得成功。所以如果我们把流程的这一段切断,长期来看,受伤的是数学本身。

那么,也许这才是最终的目标:解题真正的目标,不只是生成证明,不只是验证证明,不只是解释证明,也不只是发表证明,而是把所有这些证明组织进一个正典的状态,让它们成为定本。最后这三个阶段,我喜欢称之为「证明的消化」。这不只是把食物做出来,甚至不只是把它吃下去,你还得把它消化掉,内化到它真正成为数学身体的一部分。

消化不良

陶哲轩: 这个过程有可能还会继续延伸下去,不过我能识别出来的是这五个阶段。你们可以看到,「把目标讲清楚」这个动作本身,就能把你带到离你原本设想的流程很远的地方。

而这番拆解揭示了一件事:既然刚才那一整套是「证明的消化」,那么 AI 正在大量制造的,就是我称之为「证明的消化不良」的东西。如果你是电气工程师,也许会叫它阻抗失配。

因为 AI 在前两个阶段远远强于后三个阶段,我们已经开始看到各式各样的消化不良。我们看到未经验证的证明在堆积。我们看到已验证却没人会去读的证明。我们看到可读、却没有人愿意发表的证明。还有一件事尚未发生,但只要往下推演就会到来:很快我们还会有一大堆已发表的 AI 生成的证明,而没有人知道该如何把它们转化成像样的教科书,去教育下一代。

这些都是过剩带来的问题。粗略地说,世上的问题可以分为稀缺问题和过剩问题。以食物为例,饥荒和营养不良是食物稀缺的问题;而肥胖、缺乏运动、饮食结构糟糕,这些是食物过剩的问题。我们在证明稀缺的时代里生活了好几个世纪。但如果相信这个工作假设,我们很快就会进入证明过剩的时代,我们必须适应。在食物这件事上,我们的适应方式是变得对吃什么、怎么运动有意识得多;我们的品味提高了,我们的烹饪变好了,我们拒绝掉一些技术上能吃的东西,我们不会把盘子里的每一样都吃完。我想,在数学里我们也得做类似的事情。

莱顿宣言

陶哲轩: 一旦你对自己的目标有了更清楚的认识,你才能开始制定方针,才能真正去攻这个共同体回应之问。

作为一个起点,很多人已经听说过:去年在莱顿,一群数学家在一场关于 AI 与数学的研讨会上意识到了这些问题,发起了一项自下而上的行动,试图形成一份共识声明。他们邀集了共同体的许多成员,包括在座的许多人,共同起草了如今所谓的莱顿宣言。

这份宣言并不能解决这些问题,但这正是我们需要继续多做的那类事情。它陈述了当下的处境:AI 在做什么,数学在做什么,数学究竟是关于什么的,至少陈述到共同体中相当大一部分人能够认同的程度。它有非常多的签署人,我自己也签了。而一旦你理解了共同体的目标与价值,一些切实的建议就会自然浮现出来。如果你还没看过,我真心建议去读一读这份宣言。它经过了很多轮修改,许多人提供了反馈,现在已经是一份相当不错的文件了,里面有大量建议。

比如对个人的一条建议:始终披露 AI 工具的使用。也许将来这会变得没有必要,就像我们今天不会再声明自己用了 LaTeX。但在当下这个阶段,我认为这非常重要。尤其是有一种局面是我们无论如何要避免的:所有人都在偷偷用 AI,没有人披露,没有人分享自己的最佳做法和踩过的坑。那对于摸索出如何正确使用 AI,将是极其有害的。所以我们真的需要把负责任的披露变成常态。

本着这一点,我要说明:这些幻灯片里我用了 AI,不是用来打长破折号的,而是有几处自动补全很有用,另外你们看到的那些示意图是我用 AI 生成的。其余的文字都是我自己写的。

我们还必须让审稿人的日子好过一些。我们不能一边用 AI 生成大量一百页的证明,一边把它们全倒给审稿人。我们得让这件事容易得多,负担会转移到作者身上:作者有责任把论文写到最高的阐述标准,披露所用的全部工具,在需要的地方使用形式化。

而总体上,我们要把重心从流程的第一阶段挪开。过去我们强调的是,最要紧的是抢先证出一个定理,也就是证明生成,那才是主任务,其余都是收尾清理。现在我们把第一步自动化了,它不再是瓶颈。所以,一个已经生成、却还没有整理干净、还不具备可发表形态的证明,我们应当只把它当作一个不完整的证明来看待,尚未准备好发表。我们真的要开始重视证明的消化。最后那三个阶段,阐述、发表和正典化,我们现在也是看重的,但我们应当更明确地把它们的重要性提上来。

署名的门槛

陶哲轩: 我们还需要恰当地归属贡献。现在有一个争论:如果你按下一个按钮,得到了一个证明,你有没有资格成为那篇论文的作者?莱顿宣言对此有所论述。

我个人的建议是:如果一位作者没有能力像我现在这样站在台上讲这个结果,也没有能力就他所生成的这个结果接受提问,那么他不应当成为作者;或者至少,除非有至少一个人能够把这个结果讲清楚,否则这篇论文就还不该发表。

其余的一切

陶哲轩: 我刚才只是拆解了一个方向。我画了那张数学各种目标向四面八方分叉的图,然后我挑了其中一支,解决问题。我要说明的是,它远比「尽可能多解题」那种朴素图景复杂得多。但你可以把它拆开,可以弄清楚我们究竟在做什么,那样这一面就清晰了。而这也会反过来影响一些具体问题:期刊该如何运作,未来招聘时我们该如何评估解题这项能力。

可我们在乎的目标还有很多,而这个工作假设最终会波及它们全部。所以,我刚才对解题所做的事,别人也需要对教学去做,对指导学生、对招聘、对经费申请、对科普去做。我们这门职业的每一个方面都应当摆上台面来讨论,都需要分析。那是一场漫长而庞大的讨论,我只剩十分钟,所以这些我一个都不讲了。

而针对职业的不同部分,回应也会不同。我认为有些部分我们基本上必须限制 AI,尤其是教育和训练。过早地让学生大量使用 AI 是有害的,他们需要先发展出自己内在的数学能力,然后才能学着去使用这些工具。就像我们先教算术再教用计算器,就像我们先教人走路和跑步,再把车钥匙交给他。数学里也是一样。

另外一些地方我们应当使用 AI,但我们要掌握主动,由我们自己来决定哪些用法可以接受、哪些不行,而不是让外部的力量替我们定规矩。还有一些地方,期刊这类传统载体已经不适配这种新的工作方式了,我们确实需要创造新的流程、新的基础设施。那又是一场一小时的报告,可惜今天没有时间讲。

但最重要的是,我们真的需要公开地讨论所有这些问题,不能只是被动地接受种种发展。我们需要讨论 AI 的能力,需要讨论我们的目标与价值,需要讨论共同体的回应。这也正是莱顿宣言的一项主要建议。

我就讲到这里。非常感谢各位。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 开场:介绍陶哲轩与讲座主题 ▶ 正在看
1:42 一百年前的数学基础危机 ▶ 正在看
4:19 今日危机:价值观而非论证 ▶ 正在看
6:50 AI 能力猜想与 FirstProof 数据 ▶ 正在看
13:37 条件分析:转向目标与价值观 ▶ 正在看
18:26 古德哈特定律:目标开始冲突 ▶ 正在看
22:34 解题的五阶段流网络模型 ▶ 正在看
27:05 自然摩擦:证明需要被理解 ▶ 正在看
32:39 接受与审稿:最稀缺的瓶颈 ▶ 正在看
38:41 正典化:AI 无能为力的终点 ▶ 正在看
42:19 证明消化不良与过剩时代 ▶ 正在看
44:31 莱顿宣言与共同体应对建议 ▶ 正在看
本期小问 · 档案清单
1:42 一百年前的数学基础危机,能为今天的 AI 冲击提供什么启示? ▶ 正在看
32:39 当机器能产出证明,数学家凭什么判断一个结果值得信任? ▶ 正在看
27:05 形式化与 AI 协作,会把数学追求的「理解」带向何处? ▶ 正在看
13:37 数学从个体手工走向大规模协作,问题的选择标准会怎样改写? ▶ 正在看
本期讲者
陶哲轩UCLA 数学系杰出教授、James and Carol Collins 讲席教授,2006 年菲尔兹奖与 2014 年数学突破奖得主。研究横跨调和分析、偏微分方程、组合数论,近年积极参与 AI 与形式化证明在数学中的应用与讨论。
主持人 Alex本场 ICM 公开讲座主持人,为陶哲轩致介绍词;陶在演讲中提到「正典化」(canonicalization)一词由其命名。
01开场:介绍陶哲轩与讲座主题
0:00
Welcome everyone to the public lecture of the International Congress of Mathematicians. It's a great pleasure and honor for me to introduce our speaker uh Terrence Tao. Uh Terry received his PhD from Eli Stein at Princeton University and then uh moved to UCLA where he's been ever since where he's now the distinguished professor and the James and Carol Collins chair in mathematics. Um, if I tried to list all of his various research accomplishments and accolades, we'd be here all night and you'd never get to hear a word from him. So, I'll just mention that he won the Fields Medal uh at a previous ICM, the Breakthrough Prize, many, many other awards. Um, we will take questions at the end. Uh, there's going to be a 10-minute Q&A uh outside in the foyer to the right. And otherwise, I hope you'll join me in welcoming Professor Terry Tao.
欢迎大家来到国际数学家大会的公开讲座。我非常荣幸也非常高兴地为大家介绍我们的演讲者,陶哲轩。陶教授在普林斯顿大学师从 Eli Stein 取得博士学位,之后前往 UCLA,此后一直在那里任教,现在是那里的杰出教授,以及 James and Carol Collins 数学讲席教授。嗯,如果我想把他所有的研究成就和荣誉都列一遍,我们就得在这儿待上一整晚,而你们根本听不到他说一个字。所以我只提一下,他在此前的一届 ICM 上获得了菲尔兹奖,还有突破奖,以及许许多多其他奖项。嗯,我们会在最后回答问题。会有一个10 分钟的问答环节,就在外面右手边的门厅里。那么,现在请和我一起欢迎陶哲轩教授。
便签笔记
0:53
>> [applause] >> Thank you. >> Well, thank you very much. It's such a pleasure to be here again at a at an ICM in person. It's been quite a while and I've uh reconnected a lot of old friends and uh and just absorbing the uh the positive energy. Um so my my talk is about maybe the topic everyone is talking about these days uh the the era of AI and what it means for mathematics. Uh so you know we seem to live in an era of great change and that's something that's not uh we might not be accustomed to because mathematics has been very stable for over a century. Uh things we haven't had real big existential crises or anything but now suddenly we we we uh we seem to be uh changing quite a lot.
>> [掌声] >> 谢谢。>> 嗯,非常感谢。能再次亲身来到 ICM 现场,真是太高兴了。已经有好一阵子没这样了,我在这里重新联系上了很多老朋友,也在吸收这里的正能量。嗯,我的演讲讲的可能是这些天大家都在谈论的话题——AI 时代,以及它对数学意味着什么。嗯,你知道,我们似乎生活在一个剧变的时代,而这是我们可能不太习惯的,因为数学在过去一个多世纪里一直非常稳定。我们没有遇到过真正重大的生存危机之类的事情,但现在突然之间,我们似乎正在发生很大的变化。
便签笔记
02一百年前的数学基础危机
1:42
Um and it it seems unprecedented. Um but actually I think there are historical precedents for um what we're going through. Um so um as a prologue I'll talk about um the crisis in foundations which happened a century ago. Um so before the uh turn of the 20th century uh we mathematicians operated semiformally you know we had uklid we knew about proof but uh we didn't fully acize our um foundations of logic and and and sets and analysis and and you know um we largely use kind of informal naive foundations. Um so questions like what actually is a set or what is a number? What is a function?
而且这看起来是前所未有的。但其实我认为,我们正在经历的事情是有历史先例的。嗯,所以作为开场我要讲讲一个世纪前发生的数学基础危机。嗯,在 20 世纪之交以前,我们数学家是以半形式化的方式工作的。你知道,我们有欧几里得,我们懂得证明,但我们并没有把逻辑、集合和分析的基础完全公理化,而且,你知道,我们在很大程度上使用的是一种非正式的、朴素的基础。嗯,所以像什么是集合、什么是数、什么是函数这样的问题,
便签笔记
2:21
What is a limit? Uh we had some uh rules but we didn't really write down fully the actions of mathematics. Um and that worked uh until the early 20th century uh when suddenly these philosophers who were the only ones who were thinking about these questions pointed out that there were actually some serious problems with doing things naively. So Bertrren Russell for example was a philosopher who pointed out that naive set theory was inconsistent. um you could construct a set that both contained itself and didn't didn't contain itself. Um and of course Guro has incompleteness theorems that there's no system of axioms that could ever completely um prove or disprove all the the questions you could ask in say arithmetic. Um and so we were forced to actually really think about our foundations and what uh uh what is a proof? What are the rules of our um of our profession? Um and that was traumatic. It is called the crisis in foundations. Um and it was turbulent. It was a lot of debate and argument. Uh but
什么是极限?我们有一些规则,但并没有真正把数学的公理完整地写下来。嗯,而这一直行得通,直到 20 世纪初,突然之间,那些当时唯一在思考这些问题的哲学家指出,朴素地处理这些事情其实存在一些严重的问题。比如伯特兰·罗素,他是一位哲学家,他指出朴素集合论是不相容的。你可以构造一个集合,它既包含自身又不包含自身。当然还有哥德尔的不完备性定理:不存在这样一个公理系统,能够完全证明或否证你在比如算术中所能提出的一切问题。嗯,于是我们被迫真正去思考我们的基础,思考什么是证明?我们这个行当的规则是什么?嗯,那是令人痛苦的。它被称为数学基础危机。嗯,那是一段动荡的时期,有大量的争论和辩论。但
便签笔记
3:22
we came out of it at the end with something very valuable. We had a standard uh framework for doing mathematics. You know we we have an an orthodox set of rules first order logic zera franco choice set theory uh and it's standardized and we have a consensus that it's it's good enough to do almost all of mathematics. Um now it's not the final word. uh people still study other foundations compare them to to ZFC um you know for example um um the the lean um formal proof assistant language is not based on ZFC it's based on on dependent type theory for instance um but that's fine uh we we are studying other foundations but not from a position of crisis right from we're studying it just like any other scientific mathematical object to study um and we've tested these foundations you know strenuously over the centuries and and we trust you know that we you I mean there's still theoretically the possibility that that there is a serious problem in in these foundations but but they seem to work and they they've
最终我们走出来了,并且收获了非常宝贵的东西。我们有了一套做数学的标准框架。你知道,我们有一个一套正统的规则体系——一阶逻辑、策梅洛-弗兰克尔选择公理集合论(ZFC)——它已经标准化了,我们也有共识认为它足以完成几乎所有的数学。嗯,当然这不是最终定论。人们仍然在研究其他基础体系,把它们和ZFC 做比较。比如说,Lean 这个形式化证明助手语言就不是基于 ZFC 的,它是基于依赖类型论的。但这没问题,我们研究其他基础体系,并不是出于某种危机感,我们研究它就像研究任何其他科学或数学对象一样。而且我们已经对这些基础经过几个世纪的严格检验,我们信任它。我是说,理论上仍然存在这些基础中有严重问题的可能性,但它们看起来是work的,而且它们已经
便签笔记
03今日危机:价值观而非论证
4:19
worked for a century. So I will argue that one century after uh that crisis uh we are now entering a similarly turbulent period uh and just for people who like uh spotting m dashes in text uh this was a human generated mdash just to to emphasize that um but it's not a crisis in um in our mathematical arguments uh but it's a crisis in our mathematical values and practices um that we have we had kind of been operating on naive foundations of what mathematics is. Um and suddenly they are um um reaching rather strange conclusions. Um and we do need to to reexamine and and and rebuild these foundations properly. But again, it will be a very worthwhile experience. Uh once we do that, our our subject will be much healthier, much more resilient. Um and we will be able to incorporate all these new technologies uh um properly.
运转了一个世纪。所以我要提出的观点是,在那场危机过去一个世纪之后,我们现在正进入一个同样动荡的时期——顺便说一句,对于喜欢在文本里找破折号的人,这个破折号是人类打出来的,只是为了强调一下。但这不是我们数学论证中的危机,而是我们数学价值观和实践中的危机。我们某种程度上一直是在一套关于「数学是什么」的朴素基础上运作的。而突然之间,它们得出了相当奇怪的结论。我们确实需要重新审视,并且好好地重建这些基础。但同样,这将会是一段非常有价值的经历。一旦我们做到了,我们这个学科会健康得多、韧性强得多。而且我们将能够恰当地把所有这些新技术纳入进来。
便签笔记
5:19
All right. So uh what is the topic of my talk? Um so um I can phrase it as sort of you know so the big question if you wish that motivates everything is what I call the community response question. Okay. So how should the mathematical community respond to the advent of modern AI technologies and all the related things that that come with it? Um and you know the fact that they e they either have or they claim to have capabilities to perform mathematical tasks. Um so this is a big question and it's a question for the entire community. Um you know no one not even the IMU can just lay down some rules and we can all follow them. We have to actually uh debate this question just like how we had to debate the foundations of of mathematical logic. Um but I do have some comments uh to sort of start discussion.
好,那么我这次演讲的主题是什么呢?我可以这样表述:如果你愿意的话,驱动这一切的大问题就是我所说的「共同体应对问题」。好的。那就是:数学共同体应该如何应对现代 AI 技术的到来,以及随之而来的所有相关事物?以及它们要么已经具备、要么声称具备执行数学任务的能力这一事实。这是个大问题,而且是整个共同体的问题。没有任何人——哪怕是国际数学联盟(IMU)——能直接定下一些规则让我们所有人遵守。我们必须真正地辩论这个问题,就像当年我们必须辩论数理逻辑的基础一样。不过我确实有一些看法,可以算是抛砖引玉。
便签笔记
6:08
Um now this is not a math question. This is not this is not a conjecture where there's going to be a proof and a and a single answer. this is a meta mathematical question um and you know also a political question ethical question you know cultural economic whatever um so it's not a math question however I think the mindset of thinking very carefully and mathematically about about problems is very useful so I'm going to um tackle this question pseudoathematically I'm going to use the language of mathematics uh to try to analyze this question even though it is not a mathematical question uh and hopefully because so many of you in the audience are mathematicians. This will help uh clarify the points I'm trying to make.
这不是一个数学问题。这不是一个会有证明、会有唯一答案的猜想。这是一个元数学问题,同时也是一个政治问题、伦理问题,还有文化的、经济的,等等。所以它不是数学问题。但是我认为,那种非常审慎地、数学式地思考问题的心态是很有用的。所以我打算用「伪数学」的方式来处理这个问题,我会用数学的语言来分析这个问题,尽管它并不是一个数学问题。希望因为在座各位有那么多是数学家,这样能帮助我把想讲的要点讲清楚。
便签笔记
04AI 能力猜想与 FirstProof 数据
6:50
Okay. So, we're interested in this community response question. What should we do in response to AI? And so there's a sub question which is has dominated debate over the last three years. Um uh which but it's a slightly different question. Um and so this I will call the AI capability conjecture. So I said I'm using the language of mathematics. So I'm going to call this a conjecture. Um and actually technically it's not a single conjecture, it's a family of conjectures. Um so this conjecture I'm stating here as a template, there's some placeholders or variables if you want to keep the mathematical language. Um and so uh it it's so because there's so many variables, this conjecture could mean almost anything, but I will say it anyway. So the conjecture states that at some point in the near future some AI tools will with some expense and some level of human supervision be able to accomplish some research level mathematical tasks in some fields of mathematics with some non-trivial success rate and some level of
好,我们关心的是这个共同体应对问题:面对 AI 我们该怎么办?这里还有一个子问题,它在过去三年主导了整个讨论。但它其实是一个略有不同的问题。我把它称为「AI 能力猜想」。我说过我要用数学的语言,所以我就把它叫做一个猜想。而严格来说,它其实不是单个猜想,而是一族猜想。我在这里陈述的这个猜想是一个模板,里面有一些占位符,或者说变量——如果你想继续用数学语言的话。正因为变量这么多,这个猜想几乎可以意味着任何东西,但我还是要把它说出来。这个猜想是说:在不久的将来的某个时刻,某些 AI工具将以某种成本、在某种程度的人类监督下,能够在数学的某些领域完成某些研究级别的数学任务,达到某种非平凡的成功率,以及某种程度的正确性和质量。我在里面放了这么多「某种」,以至于这
便签笔记
7:47
correctness and quality. Now um I've put so many sums in here that that this could this could be almost anything. Um and so a lot of the debate has sort of devolved in recent years as to exactly which value of sum should you put in different slots here. Um and that's not the point of my talk today. Uh I think that's actually distracting from some some more fundamental issues. Um so just think of this as family conjectures. Um of course you can make this a very weak conjecture by making a sum just a few or or you know um or you can make it very strong. you can just sort of insert in your own personal mental mad liib you know what what you can put here. So let me just sort of vaguely talk about strong and weak and intermediate forms of this conjecture.
几乎可以是任何东西。所以近年来很多辩论其实都退化成了:到底该在各个空位上填入哪个「某种」的取值。而这不是我今天演讲的重点。我认为那实际上分散了对一些更根本问题的注意力。所以就把它当作一族猜想吧。当然,你可以把「某些」设成「寥寥几个」,让它变成一个非常弱的猜想;或者你也可以让它变得非常强。你可以在自己脑子里玩个填词游戏,看看该往里填什么。所以让我就笼统地谈谈这个猜想的强形式、弱形式和中间形式。
便签笔记
8:32
All right. So it's kind of clear that this knowing the answer to this conjecture or family conjectures is important because you know for example if even weak versions of this conjecture happen to be false but that you know that these tools were largely useless for research level tasks then there's no debate you know I mean these are novelty you know some small minority of people could play with them but you know math could basically continue business as usual. Um on the other hand if the strongest versions of this conjecture are true and every single task that mathematicians do today uh AIs can do better and faster and whatever then yeah then clearly uh we cannot continue business as usual. Um particularly if we focus on the things that the AIs do best. Uh for example solving unsolved problems seems to be one of their strengths and that seems to be one of the things that we do a lot. Um we even give out medals for these things ostensibly. Um so then we have to actually uh think about uh our culture
好。很明显,知道这个猜想或这族猜想的答案是重要的。因为举例来说,如果连这个猜想的弱版本都不成立,也就是说这些工具对研究级别的任务基本无用,那就没什么好争的了。它们只是个新奇玩意儿,少数人可以拿来玩玩,但数学基本上可以照常运转。另一方面,如果这个猜想最强的版本是真的,今天数学家做的每一件事 AI 都能做得更好更快,那么是的,那我们显然就不能照常运转了。尤其是如果我们聚焦于 AI 最擅长的事情。比如说,解决未解决的问题似乎是它们的强项之一,而这恰恰又是我们做得很多的事情之一。我们甚至表面上还为这类事情颁发奖章。所以那时我们就必须真正地思考我们的文化和实践。而如果成立的是介于
便签笔记
9:26
and practices and if it's something in the middle is true then we we have to think even more carefully. Okay. So as I said if you look at all the debates in the media and probably on at coffee uh you know at common rooms and and on the internet um you know because the answer to the community response question depends so much on the answer to the AI capability conjecture. So much of the debate is has been about capability um and um yeah so and and I have certainly um participated in this if you look at what uh I've spoken about AI a lot in the last three years and I talk a lot about the capability because it is important um and um um there was certainly a time in the like three three years ago where the capability was very poor but it was trending to be a lot better and um we really had to clarify the dist distinction between where we were and where the technology was going um but this talk is not about that conjecture. Um oh hang on I will get to that but um um now if you follow the uh internet I'm sure you have heard many
两者之间的情形,那我们就得思考得更仔细了。好,正如我所说,如果你去看媒体上的所有辩论,可能还有喝咖啡时、公共休息室里以及互联网上的辩论——因为共同体应对问题的答案在很大程度上取决于AI 能力猜想的答案,所以大量的辩论都是围绕能力展开的。而我自己当然也参与过这些讨论。你看我过去三年谈了很多关于 AI 的话题,我也大量谈论能力问题,因为这确实重要。而且在大约三年前,确实有那么一段时间,能力还非常差,但趋势是会好得多。我们真的必须澄清我们当时所处的位置和技术正走向何方之间的区别。但这次演讲不是关于那个猜想的。哦等等,我会讲到那个,不过——现在如果你关注互联网,我相信你已经听过许许多多的说法,还有反驳,以及关于
便签笔记
10:34
many claims and um um and and and counter claims and and arguments about whether this conjecture is true or false in various levels. Um and so we have a lot of data points now. This is you know a specific problem has been solved and then but then uh this this AI can't even you know multiply two-digit numbers or or whatever. Um most of the data we have is not gathered under proper scientific conditions. Um we have sort of uh selectively disclosed results which look impressive but we don't know exactly what the inputs were.
这个猜想在各个层面上是真是假的论证。所以我们现在有很多数据点了。比如某个具体问题被解决了,但接着又有人说这个 AI 连两位数乘法都不会,诸如此类。我们掌握的大部分数据都不是在恰当的科学条件下收集的。我们看到的是那种被选择性披露的结果,看起来很惊艳,但我们并不确切知道输入是什么。
便签笔记
11:05
We don't know the the proper failure rate. Um and um many of the uh uh companies that are disclosing these have their own incentives to to maybe uh uh present the results in as favorable a light as possible and not review important information like say how much cost it was to to actually produce these results. Um so it is it is a very confusing mess uh right now. Um this this whole uh all the data points for against this conjecture. Um and also another point is is that um people sometimes conflate whether a con this conjecture is true or false as to whether we want it to be true or false.
我们不知道真实的失败率。而且很多披露这些结果的公司有他们自己的动机,可能会把结果尽可能地往好看的方向呈现,而不去披露一些重要信息,比如说产生这些结果到底花了多少钱。所以现在这一切非常混乱。支持或反对这个猜想的所有这些数据点都是如此。另外一点是,人们有时会把这个猜想是真是假,和我们「希望」它是真是假混为一谈。
便签笔记
11:41
Uh and I think that is um it it that may not necessarily align either. Um so I won't be talking about this conjecture. Um the one thing I will point you to is that we do have a we are starting to have some more systematic and scientific assessments of of um of AI capability. Um so uh I think the most promising is the first proof challenge which many of you have heard about. Um so this is a grassroots effort run by mathematicians. Um and it's an independent um assessment of the state of of of AI models today. It is not sponsored by any uh major AI company. Uh what it does is that it it creates a batch it mathematicians donate a batch of of 10 research level problems that have not been published. The solutions have not been published. all their all the problems themselves and they test them against completely autonomous AI models. Um and so for example back in May uh there were four harnesses submitted. I was involved in a team that submitted one of them. Um and they tested them to try to to test these
我认为这两者也未必是一致的。所以我不会谈这个猜想。我唯一想给大家指出的是,我们确实开始有了一些更系统、更科学的 AI 能力评估。我认为最有希望的是 FirstProof 挑战赛,你们很多人应该听说过。这是一个由数学家发起的草根项目。它是对当今 AI 模型状态的一个独立评估,没有任何大型 AI 公司赞助。它的做法是:数学家们贡献出一批 10 道尚未发表的研究级别问题,解答也没有发表过,问题本身也没有发表过。然后他们用这些题去测试完全自主的 AI 模型。举例来说,今年五月有四套「harness」(系统方案)提交上来。我也参与了其中一个提交团队。他们用这些问题去测试它们。在这四套方案之间,他们不仅问
便签笔记
12:45
these problems. And between those four harnesses um they didn't just ask whether the the uh these problems were solved but whether they were written up and explained in a way that would be publication quality. like would it would actually be reasonable um to publish in a paper in a journal maybe with some slight um corrections or something. Um and uh uh I think the uh the best individual record I think was like five out of 10 for one harness but collectively they were able to sort of seven out of 10 at at publication level quality but at a non-trivial cost. Uh each harness to spend on each problem the cost ranged anywhere from 10 to1,000 US. Um, so that's roughly where we were at in May. Um, now they're going to keep testing um, these batches and maybe these stats will change, but I would say this is the most scientific uh, data point we have right now.
这些问题是否被解决了,还问解答的书写和讲解是否达到了可发表的质量:也就是说,是否真的可以合理地在期刊论文中发表出来,也许只需要一些小修改。我记得单个方案的最好成绩大概是10 题里做出 5 题;但四套合起来,能达到 10 题里有 7 题达到可发表质量水平,不过成本不低。每套方案在每道题上花费的成本从 10 美元到 1000 美元不等。这大致就是我们五月时的状况。他们还会继续测试新的题批,也许这些数据会变化,但我认为这是我们目前掌握的最科学的数据点。
便签笔记
05条件分析:转向目标与价值观
13:37
Okay, so I'm not going to talk about this conject this conjecture anymore. Um, I'm going to talk about the complement of this conjecture. Um, what um, uh, what suppose this conjecture is true in a certain in a certain sense. What then do we do about commutive response? So I will make a as you know as mathematicians do we'll make a hypothesis okay I'll make a working hypothesis where all the sums that I said before will now be replaced by reasonable okay so a reasonably strong um conjecture is true you know reasonably soon AI will be able to perform a reasonable fraction of mathematical tasks with reasonable levels of success quality and cost and I will not define what reasonable means okay so but um you can sort of put in your mind what um some idea okay Um, so this is an assumption. Okay, in math we sometimes make an assumption. We don't know whether it's true or false, but we are interested in exploring it consequences. I'm going to do a conditional analysis. You can argue whether this conjecture, this hypothesis
好,我就不再谈这个猜想了。我要谈的是这个猜想的「补集」。也就是说,假设这个猜想在某种意义上是成立的,那我们对于共同体应对该怎么办?所以我要做一个假设——就像数学家常做的那样,我们提出一个工作假设:把我之前说的所有「某种」都替换成「合理的」。也就是说,一个合理强的猜想是成立的:在合理可期的将来,AI 将能够以合理的成功率、质量和成本,完成合理比例的数学任务。我不会定义「合理」是什么意思。你可以自己心里有个大概的概念。所以这是一个假设。在数学里我们有时会做假设,我们不知道它是真是假,但我们有兴趣探索它的推论。我要做的是一个条件性分析。你可以争论这个猜想、这个假设是真是假,但那不是这次演讲的重点。我只是要有条件地
便签笔记
14:37
is true or false. That is not the point of this talk. I'm just going to assume it conditionally. So once you condition on the fact that on on the hypothesis that AI will actually solve a a reasonable fraction of uh of things that we currently do today as mathematicians uh what becomes clear is is that in order to proceed to answer the response question we now have to think about our goals and our values. Uh and so really this is I think the important question we need to discuss today. uh I mean not just literally today but but um in this current era uh so this is I call the goals and values question so what are the precise goals objectives and values of our community and of mathematical research in general now we have explicit goals you know we we when we make press releases or when we write abstracts to our papers or when we apply for funding grants we list you know this is the objective of this project or whatever those are explicit goals um but what we're finding is that actually there's
接受它。所以一旦你以「AI 真的能解决我们今天作为数学家所做的事情中相当一部分」这个假设为前提,那么很清楚的一点是:为了接着回答这个应对问题,我们现在必须思考我们的目标和价值观。所以我认为这才是我们今天需要讨论的重要问题——我不只是字面意义上的今天,而是指在当下这个时代。所以这个我称之为「目标与价值观问题」:我们这个共同体、以及整个数学研究,其确切的目标、宗旨和价值观是什么?我们当然有显性目标。当我们发新闻稿,或者写论文摘要,或者申请科研经费时,我们会列出「这个项目的目标是什么」之类的。那些是显性目标。但我们发现,其实还有很多隐性目标,我们
便签笔记
15:35
also a lot of implicit goals that we don't talk about enough um and uh it's certainly very important that we do so uh so let me explain why so in the past you know we stayed we we focused mostly on technical goals you know like um I want to improve this constant you know I I want to to to prove this theorem I want to find a definition that does x y and z okay um and questions about values and goals and you know what is what is the purpose of mathematics we thought that this was a question for the humanities you know maybe philosophers of mathematics would tackle this question or or sociologists of science or um or math educators or or you know but um you know so uh we we have focused maybe overly too much uh on the technical aspects uh because that was what we were best at um and in some ways it's sort of easier to to work with uh despite me math so being uh quite quite difficult some often um but u I'm going to argue that you know if you assume the working hypothesis we will not have the luxury anymore of not
谈得不够多。而我们确实非常有必要去谈。让我解释一下原因。在过去,我们主要关注技术性目标,比如说:我想改进这个常数,我想证明这个定理,我想找到一个能实现某某功能的定义。而关于价值观、目标,以及数学的目的是什么这类问题,我们以为那是人文学科的问题:也许数学哲学家会来处理这个问题,或者科学社会学家,或者数学教育工作者。但是,我们可能过于偏重技术层面了,因为那是我们最擅长的,而且在某些方面它反而更好把握——尽管数学本身常常相当相当困难。但我要说的是,如果你接受这个工作假设,我们将不再有那种奢侈:不去披露、不去真正说明我们的
便签笔记
16:38
disclosing not really disclosing our goals. Um and I think it's healthy. So you know regardless of whether this hypothesis is true or false or partly true or whatever um I think it is actually overdue that we do need to talk about why we do mathematics and what is the real purpose of of our subject. Um and we will we will we will be stronger uh for doing so. So why do we do research? Why do we try to solve problems and and and uh um and explore mathematical concepts and there's actually a lot of reasons um and I don't think anyone has compiled a complete list uh here are just some you know so of course we solve open unsolved problems that is important um both pure problems you know uh problems that that don't have any direct application outside of mathematics but also many important applied problems so very important motivation for mathematics we like to develop new tools new theories new techniques uh we like to understand mathematics better. We like to understand the world around us. You
目标。而且我认为这是健康的。所以,无论这个假设是真是假,还是部分为真,我认为我们其实早就该谈一谈:我们为什么做数学,我们这个学科的真正目的是什么。我们会因此变得更强大。那么,我们为什么做研究?我们为什么要去解决问题、探索数学概念?其实原因有很多,我不认为有谁整理出过一份完整的清单。这里只是其中一些。当然,我们解决开放的未解问题,这很重要——既包括纯粹的问题,也就是在数学之外没有任何直接应用的问题,也包括许多重要的应用问题。这是数学非常重要的动机。
便签笔记
17:34
know, the world is full of patterns as as we saw in in Anna Fry's Lord or um you know and we want to understand um the world. Um but we also you know math is also a very social subject. Um you know um math you know we are excited about math a large part because we can communicate to to the rest of our community and we want to build that community and uh we want to we want to train the next generation. We want to sustain this community. So, you know, we we invest a lot of effort in in lifting up and educating um the next generation of mathematicians who will do things that that we can't dream of. Um and we have this huge database of mathematical knowledge that's been building for thousands of years. We have almost the the oldest continuous continually developing cumulative knowledge base um on the planet and and you know we contribute to that. That's amazing. And this aesthetic value which we often um uh downplay but it is also important.
……世界充满了各种规律,就像我们在 Anna Fry 的报告里看到的那样,我们想要理解这个世界。但同时,数学也是一门非常社会性的学科。我们之所以对数学感到兴奋,很大程度上是因为我们能把它传达给整个社群,我们想建立这样一个社群,我们也想培养下一代。我们希望这个社群能延续下去。所以我们投入了大量精力去扶持和教育下一代数学家,他们将来会做出我们做梦都想不到的事情。我们还有一个积累了数千年的庞大数学知识库。我们几乎拥有这个星球上最古老、持续不断发展、层层累积的知识体系,而我们在为它做贡献,这太了不起了。还有审美价值,我们常常把它说得很轻,但它同样重要。
便签笔记
06古德哈特定律:目标开始冲突
18:26
you know, mathematics is beautiful. Um, and uh, you know, you can create something that that people will admire for centuries and that that also is is is a value and you can, you know, it's your homework, you can add two more u bullet points to this list. So, we have all these goals um, and we tend to only state one or two of them at a time. Um, and until recently that was kind of fine uh, because they were all aligned. Um, so we had all these goals and they're slightly different goals. Um but math was hard and we're very far from attaining many of these goals, you know, solving problems is difficult, building theories is difficult, everything was difficult. Um and so as long as you're far away from all these goals, as long as they're kind of roughly correlated with each other, you can move towards one goal and you you'd also be moving towards other goals as well. So we we so because of this positive correlation um you can just state one goal explicitly and and it can be a proxy for all the others. You don't
数学是美的。你可以创造出让人们欣赏几个世纪的东西,这本身也是一种价值。这算是你们的作业,你们可以再往这个清单上加两条。所以我们有这么多目标,而我们往往一次只说出其中一两个。直到最近,这样做其实也还好,因为这些目标是彼此一致的。我们有这些目标,它们略有不同,但数学很难,我们离其中很多目标都还很远——解决问题很难,构建理论很难,什么都很难。所以只要你离所有这些目标都还很远,只要它们大致是彼此相关的,你朝一个目标前进,同时也就在朝其他目标前进。正因为这种正相关,你可以只明确说出一个目标,它就能充当其他所有目标的代理指标。你不需要说……
便签笔记
19:20
need to state, you know, if you can just say, I'm I'm going to do goal one. I also secretly want to do goal two, goal two, but I know that by pushing towards goal one, I am getting closer to go to two two as well. So, we we didn't really need to to state all our goals. Okay. Um so this works until uh you start optimizing a lot um and you become very good at say goal one. Um and at some point uh you become uh so good at goal one that any further progress towards goal one actually moves you away from goal two, three, four and five. Um and this is a very um well-known uh phenomenon in economics is known as good target law that when a measure becomes a target, it ceases to be a good measure.
……你不需要全说出来,你可以只说:我要去做目标一。我其实私下也想做目标二,但我知道只要朝目标一推进,我离目标二也会更近。所以我们其实不需要把所有目标都摆出来。好,这套做法一直有效,直到你开始大量优化,直到你在比如目标一上变得非常擅长。到了某个时刻,你在目标一上做得太好了,以至于在目标一上再往前进一步,反而会让你离目标二、三、四、五越来越远。这在经济学里是一个非常著名的现象,叫做古德哈特定律:当一个衡量指标变成了目标,它就不再是一个好的衡量指标。
便签笔记
20:03
If you optimize too much for any specific metric, uh you can it's it comes at the expense of other goals that you actually wanted to optimize as well. Um and um yeah for humans uh when humans are doing this it's not that much of a problem because humans are not that amazing at optimization uh but AIs are um and also uh AI companies are um and that is um and that is actually a dangerous combination um that AI AIS can optimize uh untethered by by actual reality um and so uh they particularly vulnerable to good hearts law.
如果你为某个特定指标优化过头,代价就是牺牲掉你本来同样想优化的其他目标。对人类来说这倒不算太大的问题,因为人类并不那么擅长优化——但 AI 很擅长,AI 公司也很擅长,而这其实是一个危险的组合:AI 可以不受现实约束地去做优化,所以它们特别容易受古德哈特定律的影响。
便签笔记
20:44
And so we are now entering a period where all the goals that we have are now pointing in opposite in in in competing directions. Um and I think um you know AI in particular may get us closer to one of these goals but at the expense of others and so um we can no longer afford to only uh focus on one goal at a time. Uh and the goals that we do focus on we have to be really clear about what our rubric is for these goals. Um so that's the situation that we are now finding ourselves in. Okay. Now I do not have enough time to talk about all of these goals. Um so I will focus on on just one of the goals but it's it's the one that somehow has attracted the most attention in in in in recent um in in recent events which is problem solving.
所以我们现在正进入这样一个时期:我们所有的目标现在都指向彼此相反、彼此竞争的方向。我特别觉得,AI 可能会让我们更接近其中某一个目标,但代价是牺牲其他目标。所以我们再也承担不起一次只盯着一个目标了。而且对于我们确实要聚焦的那些目标,我们必须非常清楚地说明衡量它们的标准是什么。这就是我们现在所处的处境。好,我没有足够的时间把这些目标全讲一遍,所以我只聚焦其中一个——但它恰好是最近这些事件中最受关注的那一个,就是解决问题。
便签笔记
21:32
Um so you know if if in so far as as uh the outside world looks at mathematics um I think they have they don't really have any any idea what we do but it seems like but um um one of the things that it looks like we are really obsessed with is solving open problems. So um I'll discuss this prof this aspect of our profession problem solving. just by note as I'm trying as I try to emphasize this is not this just one of our goals but I will deconstruct it um and uh try to make some points about this goal alone um yeah I mean we don't you know um we not all just trying to solve as many problems as possible so yeah many mathematicians are more interested in building beautiful theories or or educating the next generation or whatever these are all very important goals but I will focus on the um uh the goal of problem solving because this is the one which is currently being impacted the most by by AI if we assume this working hypothesis.
就外部世界看数学而言,我觉得他们其实并不真正了解我们在做什么,但看上去,我们特别痴迷的一件事就是解决未解难题。所以我要讲讲我们这个职业的这一面:解决问题。不过我还是要强调,这只是我们的目标之一。我会把它拆解开来,试着单就这一个目标说几点。我是说,我们并不是都在拼命解决尽可能多的问题。很多数学家更感兴趣的是构建优美的理论,或者培养下一代,等等,这些都是非常重要的目标。但我会聚焦于解决问题这个目标,因为如果我们接受这个工作假设的话,它是目前受 AI 影响最大的那一个。
便签笔记
07解题的五阶段流网络模型
22:34
Okay. So what is uh the goal here? All right. So as I said externally it looks like you know to an outsider this is the goal of mathematics just solve as many unsolved problems as possible. um kind of like as if you know um you know when I was um when I was a student um when I was a high school student I had no idea you could be a professional mathematician and my when I learned that there was such a job I imagined that there was some committee of senior mathematicians that would somehow assign people problems and we would just all work on our assigned problems as our homework assignments and and that was somehow uh the way this the uh the system worked. um uh and it is totally not the way the system works. But um that is kind of what the impression of it is from the outside sometimes.
好,那这里的目标是什么?如我所说,从外部看,在外行看来,数学的目标就是尽可能多地解决未解问题。有点像……我还是学生的时候,我上高中的时候,我根本不知道可以把数学家当成一份职业。后来我知道有这么一份工作时,我想象着有一个由资深数学家组成的委员会,会以某种方式给大家分配问题,然后我们就都去做自己被分配到的问题,像做家庭作业一样,这个体系就是这么运转的。当然完全不是这么回事。但外界有时候的印象大概就是这样。
便签笔记
23:19
Okay. So you can think of it as as okay again in the pseudo mathematical language. I'm trying to optimize this flow network. I got I've got this source of problems and this sync of solutions and there's this um uh um flow of proof generation and we're trying to maximize the flow from um open problems to solutions. Okay. So um if you try to optimize this goal um so even before AI we realized it was a flaw if we say uh I'm going to um our goal was to collect as many proofs of reman hypothesis as possible. Uh we we figured out very quickly this was not a good goal. Um because you receive lots and lots of solutions to to your goals which problems which were not correct. Um okay so fine. Okay. So uh but we have an obvious way to fix that uh is that we we first generate these proofs and then we check them to be correct. Um so now we our flow has got two steps. You take open problems, you generate proofs, you create unverified solutions but then there's a separate proof verification process and creates verified solutions.
好,你可以这样想——再用一次那种伪数学的语言——我在优化一个流网络。我有一个问题的源,一个解答的汇,中间是证明生成的流,我们要最大化从未解问题到解答的流量。如果你试图优化这个目标——其实在 AI 出现之前我们就意识到它有缺陷了——如果我们说,我们的目标是收集尽可能多的黎曼猜想的证明,我们很快就发现这不是个好目标,因为你会收到大量并不正确的解答。好吧,但我们有个显而易见的修正办法:先生成这些证明,然后检验它们是否正确。于是我们的流程变成了两步:拿到未解问题,生成证明,得到未经验证的解答,然后有一个独立的证明验证过程,产出经过验证的解答。
便签笔记
24:18
Okay. Um and AI has become very good at the first step of this uh uh flow for some problems not for all problems but for some problems it's become reason pretty good and and decently good at the second goal as well. Um in fact well in the second case it's not so much because of AI although AI helps uh but because of the complimentary development in proof assistant languages which are a whole topic in itself but there will be other talks about that that's another important uh uh development in recent years the the massive improvement in proof assistant languages. So um so we visibly the proof generation and proof verification uh in mathematics has accelerated in certain fields and for certain types of problems not not uniformly but um uh it has accelerated and if we believe if we condition on this working hypothesis it will continue to accelerate even more um but then this leads to a new problem you know suppose we now have these AI generating these proofs and maybe we even have uh proofs
AI 在这个流程的第一步上已经变得非常擅长——对某些问题,不是所有问题,但对某些问题它已经相当不错了;在第二步上也做得挺好。事实上第二步与其说是因为 AI(虽然 AI 也有帮助),不如说是因为证明助手语言这一配套发展,那本身就是一个大话题,后面会有别的报告讲到,那也是近年来一项重要进展——证明助手语言的巨大改进。所以我们看得出来,数学中的证明生成和证明验证在某些领域、对某些类型的问题上加速了——不是普遍地,但确实加速了。而如果我们相信、如果我们以这个工作假设为前提,它还会继续加速。但这就带来了一个新问题:假设我们现在有 AI 生成这些证明,甚至我们还有一些证明……
便签笔记
25:18
that are in some formal language like lean and they're verified Right? And we have these 100,000line proofs that we've verified to be correct, but no one understands them. Um, even the humans who entered in the prompts to make these to to generate the proofs, maybe they they don't understand them either. Um, and this is already happening in some areas. So you maybe you've heard about the Erdish problems. Um, it's a list of about,200 uh problems, some solved, some some unsolved. Um, and proofs for the first time in since in the last six months. um uh we we're starting to get a backwalk you know proofs are being generated um and some are even formalized in lean um and no human has actually gone through them and we don't um there's no human who can ver who can vouch this is this is a good proof um and you know if there's many cases where even the people submitting the proof say here's a proof but I don't I'm not qualified to to evaluate it I don't know whether it's correct or not and then someone else says I've checked
……是用 Lean 这样的形式语言写的,而且已经被验证过了,对吧?我们有这些十万行的证明,验证过是正确的,但没有人理解它们。甚至连输入提示词、让 AI 生成这些证明的人,可能也不理解。这在某些领域已经发生了。也许你们听说过埃尔德什(Erdős)问题,那是一份大约一千多个问题的清单,有些已解决,有些未解决。在过去六个月里,这是第一次……我们开始出现积压:证明不断被生成出来,有些甚至在 Lean 里被形式化了,但没有人真正通读过它们,没有人能担保这是一个好的证明。而且很多情况下,连提交证明的人都会说:这是一个证明,但我没有资格评判它,我不知道它对不对。然后另一个人说:我在 Lean 里检查过……
便签笔记
26:14
it in lean I still but I can't I can't vouch for it either Um and it hasn't quite happened yet but we are very very close to a scenario in which a major result gets proved and verified and no human can understand can explain it. Um so that would be uh a very uh unwelcome development. Um so what that means is that there is at least one more stage to the uh if you want to uh so our goal must get updated. It is not enough to uh to generate proofs and it's not enough to verify the proofs. The proofs need to be explained uh well enough that they can be communicated and understood by the mathematical community. Okay, they can't just be artifacts, you know, just sort of impressive achievements that don't lead to anything further in mathematics.
……但我还是没法为它担保。这种情况还没完全发生,但我们已经非常非常接近这样一种场景:一个重大结果被证明出来、被验证过,却没有人能理解它、能解释它。那会是一个很不受欢迎的发展。这意味着这个流程至少还要再加一个阶段——我们的目标必须更新。光生成证明不够,光验证证明也不够。证明还需要被解释得足够清楚,好到能够被数学社群传达和理解。它们不能只是一些成果物,只是一些看着挺了不起、却不能给数学带来任何后续进展的东西。
便签笔记
08自然摩擦:证明需要被理解
27:05
They need to be understood. Um so we need to generate solutions that are well written. Now currently AIs while they are accelerating the first two stages of this process they are still quite weak in this third stage. Um they are good at some parts of exposition but not others. Okay so for example at the base level spelling grammar formatting they they're perfect at this almost too perfect to the point where people would prefer to read slightly imperfect text at this point. Um but uh it's not just about just technically being you know spelling correct and so forth. Um often the emphasis is very strange. Um like AI generate mathematics is very frustrating to read. Um you're you're looking for um you know what was the most important what was the most difficult part? always the the interesting part of of a proof and you'll find that the AI spends just as much time on some very trivial lema that that is you know three pages to prove something obvious and then like you know three lines to prove like the
它们必须被理解。所以我们需要产出写得好的解答。目前 AI 虽然在这个流程的前两个阶段起到了加速作用,但在第三个阶段还相当弱。它们在表述的某些方面做得好,另一些方面不行。比如在最基础的层面,拼写、语法、排版,它们做得很完美——几乎完美到让人宁愿去读稍微有点瑕疵的文字。但这不只是技术上拼写正确之类的事。它们的重点常常放得很奇怪。AI 生成的数学读起来非常让人抓狂:你想找的是最重要、最困难的部分在哪儿——那永远是一个证明中最有意思的部分——结果你会发现,AI 在某个非常平凡的引理上花了同样多的篇幅,用三页去证明一件显然的事,然后只用三行去证明……
便签笔记
28:03
really interesting part of the of the argument um also they often don't disclose uh where the influences came from you know they're um I mean it's very indirect often they they the proofs come through some some some weights which come from some training data and at some point some some previous result indirectly sort of influenced the result but they they often um uh these these AI generated texts say they often disclose no um um no sense of of where these these connections came from. So it the results often feel a lot more disconnected from the literature than they should um you know so um and maybe this will get better. So if if you assume that that AI capabilities get better you you know they are slowly getting a little bit better at exposition. It's harder uh because um verification of a proof you know this is clear signal it's true or false and and uh you can you can send you can you have a good scoring function and you can optimize it and that's what AI tools are good at whether a proof is well written
……论证中真正有意思的那部分。另外,它们常常不说明影响来自哪里。这往往很间接:证明是通过某些权重产生的,那些权重又来自某些训练数据,在某个环节,某个先前的结果间接地影响了这个结果,但这些 AI 生成的文本常常完全不披露,完全没有交代这些关联是从哪儿来的。所以这些结果读起来往往比它们本应有的样子更脱离文献。也许这一点会改善。如果你假设 AI 的能力会不断变强,它们在表述上确实在慢慢变好一点。这件事更难,因为对证明的验证有清晰的信号——真或假,你有一个好的打分函数,可以去优化它,而这正是 AI 工具擅长的;至于一个证明写得好不好……
便签笔记
29:02
or or not well written u we don't yet have really good rubrics for this but maybe we will and maybe proofs will get um uh easier to read um but that's not necessarily a good thing either um sometimes a proof can be too slick Um maybe you've experienced this if you've if you've read a a a proof that has somehow been where all the difficulty has been sanded down to nothing and like this it seems like the proof is effortless and then you try to reconstruct that at home in pen and paper and you can't do it because uh it has somehow hid the difficulty. Um it actually is important for proofs to contain a little bit of what I call natural friction. Um so when a human tries to prove something there are some steps that are easy and the human will just sail through them and not spend too much time on them. And then there are some steps that really are hard and the human will pause and and take some time and and try to organize the proof carefully and and try not to do too much work and um and the reader can sense
……写得不好,我们目前还没有真正好的评判标准,但也许将来会有,也许证明会变得更容易读。不过这也不一定是好事。有时候一个证明可能太顺滑了。如果你读过那种把所有困难都打磨得一干二净的证明,你可能有过这种体验:证明看起来毫不费力,可你回家拿纸笔想把它重构出来,却做不到,因为它不知怎么把困难藏起来了。证明里其实很重要的一点,是要包含我称之为“自然摩擦”的东西。当一个人去证明某件事时,有些步骤是容易的,人会一带而过,不在上面花太多时间;而有些步骤是真的难,人会停下来,花些时间,小心地组织证明,尽量不做多余的工作。读者能感觉到……
便签笔记
29:58
this and the reader can pick up on this and and understand where the the difficult part of a proof is. Um AI, you know, easy things that will blast through but also difficult things that also blast through and it just looks the same. Um and paradoxically actually you know humans when they um they don't do this and they make mistakes in exposition can actually help the reader in some way. It's it's paradoxical but it's true. Um so I'll give you one example of my own experience. Um so this paper this is a um my copy of a paper of Jean Ban from 1991 which um I read as a grad student. It's about the kaya conjecture and the restriction conjecture. actually uh Hong Wang's uh work uh that was in uh in presented yesterday is based a lot of it is based actually on this particular paper um and for those of you who have never read a paper Jean Bugen uh it is it is a very very educational experience um so um I still have the notes of myself from as a grad student trying to get through this paper um and he was to say you know
……读者能感觉到这一点,能捕捉到它,从而明白一个证明的难点在哪里。而 AI 呢,容易的地方它一路推过去,困难的地方它也一路推过去,看上去完全一样。而且很矛盾的是,人类不这么做、甚至在表述中犯了错的时候,某种程度上反而能帮到读者。这很吊诡,但确实如此。我举一个我自己的经历。这篇论文,这是我手上的一份 Jean Bourgain 1991 年论文的复印件,我读研究生时读过它。它讲的是掛谷猜想和限制性猜想。其实王虹昨天报告里的工作,有很大一部分正是建立在这篇论文之上的。对于从没读过 Jean Bourgain 论文的人来说,那是一次非常非常有教育意义的体验。我到现在还留着我研究生时啃这篇论文做的笔记。我得说……
便签笔记
31:02
clearly I can do this and you know this this this this statement is is essent this is essentially equal to this and I have all kinds of very frustrated uh question marks and so you probably can't see it but I have a uh annotation here I hate Jean Borgan okay um but I I fought my way through this um I was very fortunate to have Eli Stein explain some of this and Tom Wolf explain and I eventually understood what he was doing and actually it was it was I finally got the way he thought and and why certain things he that I didn't see were equivalent were basically equivalent and I learned so much um you know to the point where actually a few years later I I preferred reading Jean Bane's papers over all the others because I understood how his mind worked and like he just got straight to the point um and it was always a pleasure to read that but if there was definitely a learning process um and I think if his proof had been passed through many layers of AI improvement I may not have gotten the education I did um
非常幸运,有 Eli Stein 给我讲解了一部分,Tom Wolff 也讲了,我最终才明白他在做什么,其实是我终于摸清了他的思路,明白了为什么某些我当初看不出来是等价的东西,本质上确实是等价的,我从中学到了太多,以至于几年之后我读 JeanBourgain 的论文反而比读其他任何人的都更顺手,因为我理解了他的思维方式,他总是直接切入要点,读起来一直是种享受。但这中间确实有一个学习的过程,我在想,如果他的证明经过了层层 AI 的润色加工,我可能就得不到当年那样的教育了。所以说,重点不只是
便签笔记
32:02
yeah so um right so it's it's not just about uh writing papers well. It's about really digesting and understanding the results. Um and there's a famous quote of uh Bill Thirsten to this to this effect that said you know despite appearances we are not trying to meet some abstract production quer of definitions, themsel
把论文写好,而是真正消化并理解这些结果。关于这一点,Bill Thurston 有一句很有名的话大意是说,尽管表面看起来如此,我们并不是在完成某种抽象的生产指标,不是在生产定义、定理……
便签笔记
09接受与审稿:最稀缺的瓶颈
32:39
really well and it and and and and you and it has all the natural friction and it's it's it anyone reading it will actually um benefit from it. Um that's still not enough. Um like in in order to actually influence the future development of mathematics, it needs to be accepted and valued by the community. Um you know and other mathematicians need to actually want to read it and digest it and put it into their own work. And it may not, you know, and of course making it correct and and easy to read helps, but that's only a sufficient condition. It's not necessary. Um, you know, it's like it's like cooking. You know, you can cook a a beautiful meal and, you know, it it smells great and and and all the food has been verified safe to eat and that's, you know, but you can't force people or you should not force people to eat it. Um, >> and you know, and it's it's it's their choice. Um so there is this um you know you have to convince other people that this uh this result is is worth doing.
写得非常好,而且保留了所有自然的"摩擦",任何读它的人都能真正从中受益。但这还不够。要真正影响数学未来的发展方向,它还需要被整个共同体接受和重视。其他数学家得真的愿意去读它、消化它,把它用到自己的工作里。而这未必会发生。当然,把它写正确、写得易读会有帮助,但那只是充分条件,不是必要条件。这就好比做饭:你可以做一顿很棒的饭菜,闻起来香气扑鼻,所有食材也都验证过安全可食用,但你不能强迫别人吃,或者说你不应该强迫别人吃。这是他们的选择。所以你必须去说服别人,让他们相信这个结果值得关注。
便签笔记
33:41
Now this wasn't a problem until very recently because proofs were so scarce that um the moment some some result got proven then naturally people everyone else in the field will just drop everything read the paper and they'd be motivated to um to you know so um you know basically it's like if food was is scarce you will eat whatever is is delivered onto the table but um now we are flooded with proofs uh and more than more than we can we can review uh and we now have to pick and choose um and so we Um we need to be convinced to uh to actually read proofs. Now um proof review is now one of the most precious bottlenecks um scarce resources um in this new era. Um now you can help you can induce people to uh to to read your work. Um telling good stories helps having a narrative showing the process describing uh how you arrived at at at at a um at the result uh what didn't work and how how you went around it. Uh people love stories. Um we have in the past not emphasized process. We've just let the outcomes um speak for themselves
在最近之前,这都不是个问题,因为证明太稀缺了,只要有个结果被证出来,同领域的其他人自然会放下手头一切去读那篇论文,他们有充分的动力去读。基本上就是说,如果食物稀缺,端上桌的东西你都会吃。但现在我们被证明淹没了,多到远超我们能审阅的量,我们现在不得不挑挑拣拣,所以我们需要先被说服,才会真的去读一个证明。如今,证明的评审成了最宝贵的瓶颈、最稀缺的资源,在这个新时代里。你可以做些事情来吸引别人读你的工作。讲好故事有帮助,有叙事、展示过程、描述你是怎么一步步得到这个结果的、哪些路走不通、你又是怎么绕过去的。人们喜欢故事。过去我们不太强调过程,我们只是让结果自己说话——这在过去是行得通的,直到我们找到了办法,能在没有过程的情况下
便签笔记
34:46
which um has worked until we figured out ways to automate you know um um outcomes without process. Um and the current tools they are very very opaque about their process. Um there are things called chain of thought. You can kind of look under the hood a little bit. Um but um it doesn't it's not very insightful. Um now maybe some of this is is just uh can be changed with with good practices. But um uh but currently especially if these if these proofs come from propag models or internal models where uh companies often um incentivize to to to keep certain facts corporate secrets then um yeah u we do not see the process and we and this really makes us a lot less willing to actually invest the time to to to learn about these things because we we don't see the story.
自动产出结果。而现在的这些工具,对自己的过程非常非常不透明。有个东西叫思维链,你多少能掀开盖子看一点,但它并不怎么能给人启发。也许其中一部分可以通过好的实践来改变。但目前,尤其当这些证明来自专有模型或公司内部模型时,企业往往有动机把某些事实当作商业机密保守起来,那我们就看不到过程,这真的让我们大大降低了投入时间去了解这些东西的意愿,因为我们看不到那个故事。
便签笔记
35:35
Um so there is a fourth stage to um to uh um to solving problems. So not not just generating proofs, verifying them and explaining them. They have to be accepted. Um and uh in our system currently the way we uh we have the publication system. So they have to be you know we um the standard way to have a a result accepted by the community is have published in a reputable journal. So they have to be digested and accepted by the mathematical community. And we have a system set up for this. we we have journals um and um but it's slow um as I said you you can encourage it you know you can make your your papers more appetizing to read um but you cannot just try to optimize this by an AI because it it it it involves human response okay unless you somehow wire the human into an AI which I really do not recommend. Um so this is a much slower um part of the process and it is it is one where AI has basically not made an impact currently. Um yeah it is still humans talking to humans that require that that decent acceptance.
所以解决问题还有第四个阶段。不只是生成证明、验证证明、解释证明,它们还必须被接受。在我们现有的体制里,我们有一套出版系统。所以它们必须——一个结果要被共同体接受,标准途径就是发表在有声望的期刊上。所以它们必须被数学共同体消化和接受。我们为此建立了一套系统,我们有期刊,但这个过程很慢。就像我说的,你可以去促进它,可以让自己的论文读起来更有"食欲",但你没法靠 AI 来优化这一环,因为它涉及人的反应——除非你想办法把人接进 AI 里,而我真的非常不建议这么做。所以这是整个流程中慢得多的一部分,也是目前 AI 基本没有产生影响的一环。这依然是人与人之间的交流,才能带来那种真正的接受。
便签笔记
36:42
Now the way we do this right now is that we have journals and when a paper has been verified it's been good enough to submit to a journal we send it to a human referee and they volunteer to to their time as as a service to you know um to referee these papers not just because the result is interesting but they want to encourage the authors to become better mathematicians. It's it's it's a way of of promoting um another generation of mathematicians. Um it's not prestigious work you know I mean the people you know we uh we just gave four medals for solving problems we didn't give four medals for refereeing um but it is an essential part of our prof profession and you know it it it's how we convert individual achievements of mathematicians into collective understanding um I mean the referees are proxies but you know by but you know we read each other's papers and this is how the field progresses um now um but it's slow and we are now facing um a crisis that the there's so many AI generated papers that even if
我们现在的做法是:有期刊,当一篇论文被验证过、够格投给期刊时,我们把它送给人类审稿人,他们自愿贡献自己的时间,作为一种服务,来审这些论文——不只是因为结果有趣,还因为他们想鼓励作者成长为更好的数学家。这是在培养下一代数学家的一种方式。这不是什么光鲜的工作,我是说,我们刚刚为解决问题颁了四枚奖章,我们并没有为审稿颁四枚奖章,但它是我们这个职业不可或缺的一部分,它是我们把数学家的个人成就转化为集体理解的途径。审稿人只是一个代理,但我们会读彼此的论文,这个领域就是这样往前走的。可是它很慢,而我们现在面临一场危机:AI 生成的论文太多了,即便它们都是正确的,数量也可能多到审稿人根本读不完。这是
便签笔记
37:43
they're correct there may be too many for referees to read um now uh this is an ongoing problem I think there will be other panels discussing what to do about this um I think journals will also have to start adopting AI tools as filters um um that you know certain we'll have to be much more precise about uh style guides and uh uh rules for for various journals so that um before any paper reaches a human referee there's some additional layer of evaluation this will be controversial but I think uh it will be necessary but you would we should definitely not take the human referee out of the equation um I would I don't think a journal that purely operates on AI referees will be uh successful uh especially since these tools can be gamed quite a bit and you know and and if there's a journal that says oh this paper has been has been accepted by 100 AI referees now that that's not community acceptance. That still doesn't mean that we want to read it.
一个正在发生的问题,我想后面还会有别的圆桌讨论该怎么办。我认为期刊也将不得不开始采用 AI 工具作为筛选器,我们必须把各家期刊的格式规范和各种规则定得更精确,这样在任何论文送到人类审稿人手上之前,先经过一层额外的评估。这会有争议,但我认为这是必要的。不过我们绝对不该把人类审稿人从这个环节中剔除。我不认为一个纯靠AI 审稿运作的期刊会成功,尤其因为这些工具相当容易被钻空子。而且如果有个期刊说"这篇论文已经被 100 个 AI 审稿人接受了",那并不是共同体的接受,那依然不意味着我们就想去读它。
便签笔记
10正典化:AI 无能为力的终点
38:41
Okay. So, is that our final goal? Well, even that is not our um our final state really. Um you know, so even when a result has been accepted by the community and is published and is in a prestigious journal and people all accept as correct. Um it's still not the final state. Um it's the final states are things like textbooks, you know, the the material we we teach in classes. What is the standard definition of this concept? What is the correct order in which we prove things? How do you organize individual results published in good papers into a coherent theory? Um that is the state which we mathematics ultimately and that that is the end state of all these problems that we're trying to solve. Um so Alex who introduced me has a very nice name for this is it's it's canonicalization.
好,那这是我们的最终目标吗?其实连这也不是真正的终点状态。就算一个结果已经被共同体接受、已经发表在有声望的期刊上、大家都认可它是正确的,它仍然不是最终状态。最终状态是像教科书那样的东西,是我们在课堂上教的材料。这个概念的标准定义是什么?我们证明这些东西的正确顺序是什么?你怎么把发表在好期刊上的一个个孤立结果,组织成一套连贯的理论?那才是数学最终要达到的状态,才是我们试图解决的所有这些问题的终点。介绍我出场的 Alex 给这件事起了个很好的名字,叫做"正典化"(canonicalization)。
便签笔记
39:30
So um in addition to digesting a proof and and and publishing it, you want to make it canonical. Um and this is the slowest stage of all. Um you know there's many many results in the last 10 years 20 years that you know they're in prestigious journals but they're not in textbooks yet. They're not yet taught to students. We haven't yet had the figured out the really definitive canonical correct way to teach these things. Um and this is slow. I mean you have to teach classes. You have to get feedback and and you have to really think hard about what is the correct way to or to to to um to organize lots and lots of of of of um different results and put it in one coherent narrative. Um and it needs consensus you know I mean if it would be you know if half the mathematicians in a field think you should do things this way and the other half a different way it we don't have canonicalization. Um and this is the stage in which AI is basically completely useless. Um so um like there's there's sort of no I see
所以,除了消化一个证明、把它发表出来,你还要让它成为正典。而这是所有阶段中最慢的一个。过去十年、二十年里有非常多的结果,它们躺在有声望的期刊里,但还没进教科书,还没教给学生。我们还没有想清楚教这些东西真正确定的、正典的、正确的方式。这个过程很慢——你得去开课,得拿到反馈,还得非常认真地去想:把大量大量不同的结果组织起来、放进一个连贯叙事里的正确方式到底是什么。而且这需要共识。如果一个领域里有一半数学家认为该这么做,另一半认为该那么做,那我们就没有正典化。而这个阶段,正是 AI 基本上完全帮不上忙的地方。我几乎看不到 AI 在这一环有什么角色可扮演,但它却是最有
便签笔记
40:30
basically almost no role for AI in this part of the process but it is the most valuable part um that if you want to apply any sub field of mathematics to some other area of mathematics or some problem it pretty much has to be digested in this form you know like if you want field of math to become useful to engineers or or or physicists or or or biologists or whatever I mean they they're not going to dig through you know the most recent papers in in in the annals of mathematics or whatever they they want the textbooks um and many yeah so the the most valuable applications of of of math only get unlocked once you have reached this final stage including AI itself I mean a big reason why AI is so successful at mathematics is because for centuries we've been building these canonical definitions we we have these textbooks um of you know how does linear algebra work how does group theory work you know like we have all these very very mature theories And AI has absorbed all of these. Um, and that's what what
价值的一环。因为你要把数学的任何一个子领域应用到数学的其他领域或某个问题上,它基本上必须先被消化成这种形态。比如你想让某个数学分支对工程师、物理学家或生物学家之类的人有用,他们是不会去啃《数学年刊》上最新的论文的,他们要的是教科书。所以数学最有价值的那些应用,只有到达这个最终阶段之后才会被解锁,包括AI 本身。AI 之所以在数学上这么成功,一个很大的原因就是几个世纪以来我们一直在构建这些正典定义,我们有这些教科书,讲线性代数怎么运作、群论怎么运作,我们有这么多非常非常成熟的理论,而 AI 把这些全都吸收了。它靠的就是这些
便签笔记
41:28
it uses to to um to uh uh to for success. Um, and so if we cut off this this part of the process long term, uh, it it will hurt mathematics. Um, yeah. So, um, sorry said all that. So maybe this is the final um goal that so so maybe the real goal of problem solving is not just to generate proofs and not just to verify them, not just to explain them and not just to publish them but to actually um organize all those proofs into uh a a canonical state to make them definitive. Um so this last three stages I like to call proof digestion, you know. So it's not just sort of creating food and and not even just of eating it but you have to sort of uh digest it with internalize it to the point where it really becomes part of the body of mathematics.
才取得成功。所以从长期看,如果我们把流程中的这一环切断,会伤害到数学。嗯,抱歉,说了这么多。所以也许这才是最终目标——解决问题的真正目标不只是生成证明,不只是验证证明,不只是解释证明,也不只是发表证明,而是真正把所有这些证明组织成一种正典状态,让它们成为定论。所以最后这三个阶段,我喜欢称之为"证明的消化"。不只是做出食物,甚至不只是吃下去,你还得把它消化、内化,直到它真正成为数学这个躯体的一部分。
便签笔记
11证明消化不良与过剩时代
42:19
Okay. Now um it is possible that it goes this this this uh process even continues more but these are the five stages that I was able to identify. Um but you can see that the exercise of stating the your goals clearly um can take you quite far from what you might think the uh the the process is. And this um this little deconstruction reveals one thing which is that um what AI is doing is that it is creating lots and lots of well okay if this was the process of proof digestion what AI is producing is what I call proof indigestion. Um if you're electrical engineer you might call it impedance mismatching.
好。也有可能这个过程还能继续往下延伸,但这是我能识别出来的五个阶段。你可以看到,把自己的目标清楚地说出来,这个练习能带你走到离你原本以为的流程相当远的地方。而这个小小的拆解揭示了一件事:AI 正在制造大量的——好,如果说这是证明消化的过程,那 AI 生产出来的就是我所说的"证明消化不良"。如果你是电气工程师,你可能会叫它阻抗失配。
便签笔记
42:57
um that because AI is much better at the first two stages of this process than it is at the last three um we are you know beginning to see all kinds of indigestion. We are seeing proofs accumulate that haven't been verified. We're seeing verified proofs that no one will read. Um we are seeing proofs that are readable but no one will will publish. And um this isn't happening yet but uh if we extrapolate we will soon also get a big pileup of published proofs AI generated that no one knows how to convert into into like proper textbooks and and educate um the next generation. Um so these are all problems of of of abundance you know so you can roughly speaking all the problems in the world can be classified into problems of scarcity and problems of abundance.
因为 AI 在这个流程的前两个阶段远比后三个阶段擅长,我们已经开始看到各种各样的消化不良。我们看到未经验证的证明不断堆积,看到已经验证但没人会去读的证明,看到可读但没人愿意发表的证明。还有一种情况现在还没出现,但按趋势外推,我们很快也会积压一大堆已发表的、由 AI 生成的证明,没人知道该怎么把它们转化成像样的教科书、去教育下一代。这些都是"过剩"带来的问题。粗略地说,世界上所有的问题都可以分成稀缺的问题和过剩的问题。
便签笔记
43:43
Okay. So for food for example, you know, you have famine and malnutrition. These are problems of food scarcity, but you also have have obesity and lack of exercise and um and bad diet. You know, these these are problems of of food abundance. Um and so we've lived for centuries in an area of proof scarcity. Uh but we are going to if we believe this working hypothesis, we will very very soon be an era of proof abundance. Uh and we have to adapt. Um and in the case of food, we adapted by becoming much more conscious about what is good eating and what is not. What is good exercise, what is not. Our our taste improved, our our cuisines improved. We rejected some food that you know is technically edible, but we don't eat everything that's on our plate. Um and we have to do something similar in mathematics, I think.
比如食物,有饥荒和营养不良,这是食物稀缺的问题;但你也有肥胖、缺乏运动、饮食结构糟糕,这些是食物过剩的问题。几个世纪以来,我们一直生活在证明稀缺的时代,但如果我们相信这个工作假设,我们很快就会进入证明过剩的时代。我们必须适应。在食物这件事上,我们的适应方式是变得更有意识地去分辨什么是好的饮食、什么不是,什么是好的运动、什么不是。我们的品味提高了,烹饪也进步了。我们会拒绝一些技术上可食用的食物,我们不会把盘子里的东西全吃光。我想在数学里,我们也得做类似的事情。
便签笔记
12莱顿宣言与共同体应对建议
44:31
Um so once you have a clearer idea of what your goals are, then you can actually start making policies. then you can start actually attacking this community response question. Um so um I think um a first start of this many of you have heard about this. Uh so there was a grassroots effort starting in Leiden back last year. Many mathematicians in a workshop on AI and mathematics realized this and started um trying to make a consensus statement drawing in many members of the community including many people here to draft what's known as the lighten declaration.
所以,一旦你对自己的目标有了更清晰的认识,你就可以真正开始制定政策,就可以真正着手处理"共同体如何回应"这个问题。我想这方面的一个开端,你们很多人都听说过:去年在莱顿(Leiden)有一场自下而上的行动。很多数学家在一个关于 AI 与数学的研讨会上意识到了这些问题,开始尝试形成一份共识声明,吸纳了共同体的许多成员,包括在座的很多人,共同起草了后来被称为"莱顿宣言"的文件。
便签笔记
45:03
Um and this is this doesn't solve you know these these problems but this is the type of thing we need to do continue to do more of um so um this declaration states lays out the situation what AI is doing what math is doing what math is about uh at least to the the extent that a large fraction of the community can can agree with it and there are many many signaries of this I've signed this this declaration myself um and there are practical recommendations that become clear once you understand the goals and uh and values of your community. So I really all recommend if you haven't seen it to actually read this declaration. It is um it has been it's been gone through many iterations um and many people have supplied feedback and it's quite a good document at this point. Um and it has many many um recommendations. Um so one for example for individuals is to always disclose AI tool use. Um maybe in the future this will become unnecessary. You know we don't disclose latte use anymore for instance. Um but in our current era
这份宣言并不能解决这些问题,但这正是我们需要继续多做的那类事情。这份宣言陈述并梳理了现状:AI 在做什么、数学在做什么、数学究竟是关于什么的——至少是在共同体中很大一部分人能够认同的范围内。而且还有很多
便签笔记
46:02
it is it's really important I think um particularly I think the the the one scenario that we really want to avoid in this current era is where people are everyone is secretly using AI and no one is disclosing it and no one is sharing their best practices and their mistakes and this is going to be really harmful uh for um working out how to use AI properly. So we really need to normalize responsible disclosure of AI systems. So in that vein I have used AI to generate these text not for the M dashes but um there's a couple places where autocomplete was useful uh and the diagrams that you saw I generated for instance okay but the other text was was written by myself um we need to make life a lot easier for referees uh like the we cannot just generate lots and lots of 100page proofs by AI just dump them on referees um we need to make it much easier we the burden will shift on authors to to it will be it will be incumbent on authors to to make their their papers at the highest expositional standard um
……我觉得这真的很重要,尤其是在当下这个时代,我们最想避免的一种局面就是:所有人都在偷偷用 AI,没有人披露,也没有人分享自己的最佳实践和踩过的坑。这对于摸索如何正确使用 AI 会非常有害。所以我们真的需要把负责任地披露 AI 的使用变成常态。本着这个精神:我用 AI 生成了这些文字——不是那些破折号——有几处自动补全挺有用的,还有你们看到的那些图也是我生成的,但其他文字都是我自己写的。我们还需要让审稿人的日子好过得多。我们不能就这么用 AI 生成一大堆一百页的证明,然后一股脑扔给审稿人。我们得让这件事容易得多,负担会转移到作者身上——作者有责任把论文做到最高的表述标准……
便签笔记
47:03
disclosing all their tools using formalization as needed. Um and we need to in general just move away from the first stage of the process. Okay. So in the past we would emphasize what was really important was being the first to prove a theorem proof generation. This was the this main task and everything else was clean up. Um now we've automated that first step and that's no longer the bottleneck. Um so um a proof that has been generated but not cleaned up and and made publishable that is we should we should value that only as a incomplete proof uh not yet ready for publication and we really need to start valuing proof digestion. So the three final stages exposition publication and and cononicalization we um we do value these right now but we should be much more explicit about um about um raising their importance.
……披露自己用到的所有工具,必要时使用形式化。总的来说,我们需要把重心从这个流程的第一个阶段移开。过去我们强调的是,最重要的是第一个证明出某个定理,也就是证明的生成,这是主要任务,其他都是收尾。现在我们把第一步自动化了,它不再是瓶颈。所以一个已经生成、但还没整理清楚、还不能发表的证明,我们应该只把它当作一个不完整的、尚不能发表的证明。我们真的需要开始重视“证明的消化”。所以后面三个阶段——表述、发表和规范化——我们现在也是重视的,但我们应该更明确地把它们的重要性提上来。
便签笔记
47:51
Um and we need to uh attribute things properly. You know there's a debate you know if you push pushed a button and you got a proof do you get to be an author of the paper that came out. Um so um the lighten declaration has some something to say about this. Um my suggestion is that if an author cannot give a talk like up here about the result and they cannot take questions about the result that they generated um they should not be an author or at least um the paper is not ready to be published unless there is at least one person who can actually talk about the result properly.
我们还需要恰当地进行署名归属。有一场争论:如果你按了个按钮,得到了一个证明,你算不算最终产出那篇论文的作者?莱顿宣言(Leiden Declaration)对此有一些说法。我的建议是:如果一位作者没法像我现在这样站在台上讲这个结果,也没法回答关于这个结果的提问,那他就不应该当作者;或者至少,除非至少有一个人能真正把这个结果讲清楚,否则这篇论文还不能发表。
便签笔记
48:26
Um yeah so so I've deconstructed one you know so I had this diagram of of all the goals of mathematics and it branched out all over the place and I got I had one direction which was solving problems and my point was that it's it is far more complicated than sort of the naive uh um picture of just trying to solve as many problems as possible. Okay, but you can deconstruct it. You can understand what we are that what it is we're actually doing and that really clarifies that aspect. Okay. So, you know, and this informs like how journal should operate and and how we should um assess um solving problems for future hiring and so forth. But there's many other parts of uh many other goals that we care about um and this working hypothesis will will eventually impact all of them as well. So we need to also do so what I just did for problem solving people also need to do for teaching uh for mentoring hiring grant applications outreach every single aspect of our profession should be on the table and discussed and we we
所以我拆解了其中一个。我先前有那张图,画着数学的所有目标,向四面八方分叉,我挑了其中一条方向——解决问题。我的观点是,它远比“尽可能多地解决问题”这种天真图景复杂得多。但你可以把它拆开,理解我们实际在做的到底是什么,这样就能把这一面看清楚。这也会影响到期刊该怎么运作,以及我们未来在招聘等方面该怎么评估“解决问题”这件事。但我们在意的还有很多其他部分、其他目标,而这个工作假设最终也会影响到它们全部。所以我刚才对解决问题所做的这番分析,人们也需要对教学、对指导学生、招聘、基金申请、公众推广去做——我们这个职业的每一个环节都应该摆上桌面讨论,我们……
便签笔记
49:24
need to analyze them all uh and that's a big lengthy discussion and I only have 10 minutes so I'm not going to talk about any of those um yeah so um and the response will be different in in in for different parts of our profession I think there'll be some parts of our profession where we basically have to restrict AI I um education and training in particular. I think it is uh it it can be harmful to to to give students too much AI use early on. They do need to develop their own um innate mathematical skills before they can learn to uh uh to use these tools. You know, just like we we we teach people arithmetic before they do calculators or we we teach people to walk and run before they we we give them the keys to a car. Um same in math. Um there are other places where we should um use AI but we should take the initiative and and decide what are we set the rules on on what types of AI use are acceptable which ones are not and not let external actors define uh um the rules for us. Uh there'll be some places where
……需要把它们全都分析一遍。那是一场又长又大的讨论,而我只有十分钟,所以这些我就不讲了。而且对于我们职业的不同部分,应对方式也会不同。我认为有些部分我们基本上必须限制 AI,尤其是教育和培训。我认为太早让学生大量使用 AI 是有害的。他们需要先发展出自己内在的数学能力,然后才能学会使用这些工具。就像我们先教人算术,再让他们用计算器;先教人走路和跑步,再把车钥匙交给他们。数学也一样。另一些地方我们应该用 AI,但我们应该主动出击,自己来定规则:哪些类型的 AI 使用是可以接受的,哪些不行,而不是让外部力量替我们定规则。还有一些地方……
便签笔记
50:25
traditional um venues like journals would are just not suitable for u this modern a workflow. We do need to create new workflows, new infrastructures. Um that's another hour talk which I unfortunately do not have time to to give here. But most importantly we really need to discuss um all these questions um openly and and and and not just be passive uh um you know recipients of of developments but we we need to discuss AI capability. We need to discuss our goals and values and we need to discuss community response and this also is a major recommendation of the lighten declaration. So this is where I'll stop. Thank you very much.
……像期刊这样的传统渠道,根本不适合这种现代的工作流程。我们确实需要创造新的工作流程、新的基础设施。那又是一个小时的报告,可惜我没时间在这里讲。但最重要的是,我们真的需要公开讨论所有这些问题,不能只是被动地接受各种发展。我们需要讨论 AI 的能力,需要讨论我们的目标和价值,也需要讨论社群的应对,这也是莱顿宣言的一项主要建议。我就讲到这里,非常感谢。
便签笔记
51:06
[applause]
[掌声]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

陶哲轩在 ICM 2026 公开讲座中主张:AI 带来的不是数学论证的危机,而是数学价值观与实践的危机——数学界必须像百年前应对"基础危机"那样,明确说出自己的目标,尤其要把"解题"这一目标拆解为生成、验证、阐释、发表、经典化五个阶段,并认识到 AI 只擅长前两阶段,从而应对即将到来的"证明过剩"时代。

核心要点

  • 历史类比:百年前的基础危机是今日的先例。 20 世纪初,罗素悖论与哥德尔不完备定理迫使数学界正式公理化逻辑与集合论(一阶逻辑 + ZFC),过程动荡但结果让学科更健康。陶认为一个世纪后,数学正进入类似的动荡期,只是这次危机不在论证本身,而在于此前"朴素"的、未明言的价值观与实践。
  • 有意跳开"AI 能力猜想"的争论。 他把"AI 能否做研究级数学"表述为一族充满占位符("某些工具、某些花费、某些领域、某种成功率")的猜想,指出三年来的公开辩论几乎都在争这些占位符的取值,而且现有数据多为公司选择性披露、缺乏成本与失败率信息。他明确表示本次讲座不讨论真伪,而是条件分析:假设"合理强"的版本为真,数学界该怎么办。
  • 目前最科学的能力数据点:First Proof Challenge。 这是数学家自发组织、无 AI 公司赞助的独立评测:捐出 10 道未发表的研究级问题,让完全自主的 AI 系统作答并要求达到可发表的写作质量。2026 年 5 月一轮中 4 个系统参赛(陶参与其中一支队伍),单个系统最好成绩约 5/10,合计 7/10 达到发表级质量,但每道题成本从 10 美元到约 1000 美元不等。
  • 古德哈特定律解释了为何"隐性目标"必须显性化。 数学有多重目标(解决问题、建立理论、理解世界、建设社群、培养下一代、积累知识库、审美价值)。过去这些目标正相关,只需明说一个当作其余的代理即可;但一旦某个目标被极度优化,继续推进它会损害其他目标。人类优化能力有限,而 AI 与 AI 公司极擅长优化,因此格外容易触发古德哈特效应。
  • "解题"远非"最大化解出的问题数",而是一条五阶段流水线。 ①证明生成 → ②证明验证(Lean 等证明助手大幅进步)→ ③证明阐释(写得让社群能读懂)→ ④社群接受与发表 → ⑤经典化(进入教科书、形成标准定义与叙事)。AI 在①②已显著加速,③较弱,④几乎无影响,⑤"基本完全无用"——而⑤恰是最有价值的阶段,也是工程师、物理学家乃至 AI 自身真正消费数学的形态。
  • 已出现"验证了但无人理解"的证明。 以 Erdős 问题集(约 1200 题)为例,近半年涌现大量 AI 生成、部分已 Lean 形式化的证明,提交者称"不够资格评价",验证者称"检查过但不能担保"。陶警告我们非常接近"重大结果被证明并验证,却没有任何人能解释它"的局面。
  • 好的阐释需要"自然摩擦",AI 恰恰抹平了它。 AI 文本拼写格式完美,但重点错位——对琐碎引理写三页、对关键步骤写三行,且不交代思想来源,与文献脱节。人类作者在难点处放慢、在易处快速掠过,读者能由此感知难点所在。他以自己读 Bourgain 1991 年论文(Hong Wang 昨日报告的工作即基于此)的经历为例:笔记上写着"我恨 Bourgain",但正是这种挣扎让他学会了 Bourgain 的思维方式。
  • 从"证明稀缺"到"证明过剩",审稿成为最稀缺资源。 类比食物:稀缺时代的问题是饥荒,过剩时代的问题是肥胖与坏饮食,人类靠发展"品味"与主动拒绝来适应。数学也须如此:期刊需要 AI 过滤层与更精确的风格规范,但绝不能取消人类审稿——"被 100 个 AI 审稿人接受"不等于社群接受,且这类工具易被博弈。
  • Leiden 宣言与具体建议。 陶已签署去年源于 Leiden 研讨会的社群共识文件。建议包括:始终披露 AI 使用(避免"人人偷偷用、无人分享最佳实践"的局面;他本人披露本讲座用 AI 做了少量自动补全和图表);减轻审稿负担,把写作与形式化的责任转移给作者;未整理的 AI 证明只算"不完整证明";署名标准是——若作者不能上台讲解并回答关于该结果的提问,就不应署名或论文尚不宜发表。

结论与值得注意的细节

  • 核心结论:真正的瓶颈已从"率先证明"转移到"证明消化"(阐释、发表、经典化),数学界应明确提升后三阶段的价值;AI 造成的是"证明消化不良"(他戏称电气工程师会叫"阻抗失配")。
  • 他强调 AI 之所以擅长数学,正是因为数学界几个世纪以来建成了线性代数、群论等经典化教科书体系;若切断经典化阶段,长期将反噬包括 AI 在内的一切应用。
  • 对不同领域回应应有差异:教育与训练阶段应限制 AI,学生须先培养内在能力(如先学算术再用计算器、先学走路再开车);其他领域则应由数学界自主制定规则,而非任由外部行动者定义。
  • 他把解题只当作众多目标之一,教学、指导、招聘、经费申请、公众推广等每个方面都需要做同样的拆解分析,并需要新的工作流与基础设施来取代不适应 AI 时代的传统期刊模式——限于时间未展开。
  • 幽默细节:幻灯片中的破折号是"人类生成的",以回应网上用破折号识别 AI 文本的风气;他提到刚颁出四枚菲尔兹奖是为解题,"没有为审稿颁奖",但审稿是把个人成就转化为集体理解的关键环节。
核心句型 · 10
1. It seems unprecedented, but actually I think there are historical precedents for …
“It seems unprecedented. but actually I think there are historical precedents for what we're going through.”
先承认表面印象,再用 but actually 翻转,引入历史类比。适合议论开头,仿写:It seems X, but actually there are precedents for Y.
2. It's not a crisis in A, but a crisis in B.
“It's not a crisis in our mathematical arguments but it's a crisis in our mathematical values and practices”
否定—肯定的对比结构,用于精确界定问题的性质。仿写时 A、B 需是同一范畴内的两个不同层面。
3. Regardless of whether X is true or false or partly true, I think it is overdue that we …
“Regardless of whether this hypothesis is true or false or partly true or whatever I think it is actually overdue that we do need to talk about why we do mathematics”
先撇开争议前提,再提出无论如何都成立的主张。overdue 表示「早该做」,语气比 necessary 更强。
4. As long as A, you can do X and you'd also be doing Y as well.
“As long as you're far away from all these goals, as long as they're kind of roughly correlated with each other, you can move towards one goal and you'd also be moving towards other goals as well.”
以条件句解释某种机制为何过去有效。as long as 可重复叠加多个条件;后句用 would 表示一般推论。
5. This works until you start …, and at some point you become so good at A that any further progress actually moves you away from B.
“This works until you start optimizing a lot and you become very good at say goal one. and at some point you become so good at goal one that any further progress towards goal one actually moves you away from goal two”
描述「过度优化的临界点」的经典结构:works until + so…that。适合讲任何指标失效的场景。
6. It is not enough to A, and it is not enough to B. C needs to …
“It is not enough to generate proofs and it's not enough to verify the proofs. The proofs need to be explained well enough that they can be communicated”
递进式否定,逐层抬高标准,最后给出真正的要求。not enough to 可连续排比,形成层次感。
7. X is only a sufficient condition. It's not necessary.
“Making it correct and easy to read helps, but that's only a sufficient condition. It's not necessary.”
借用数理逻辑术语表达「有帮助但不保证」。注意此处陶的口误,逻辑上他想说的是「必要而非充分」;仿写时确认方向。
8. All the problems in the world can be classified into problems of X and problems of Y.
“Roughly speaking all the problems in the world can be classified into problems of scarcity and problems of abundance.”
用二分法建立分析框架,之后用具体类比(食物)填充。roughly speaking 缓和绝对化语气。
9. It will be incumbent on X to …
“It will be incumbent on authors to make their papers at the highest expositional standard”
正式语体表达「责任落在某方」。比 X should 更书面、更强调义务转移,适合政策建议。
10. Just like we teach people A before B, … same in C.
“Just like we teach people arithmetic before they do calculators or we teach people to walk and run before they we give them the keys to a car. same in math.”
用两个并列日常类比再以 same in 收束,是口语中强调原则普适性的高效方式。
词汇精讲 · 118 · 按出现顺序
accolades /ˈækəleɪdz/ n. 0:00
荣誉、赞誉(常用复数)
foyer /ˈfɔɪər/ n. 0:00
门厅、休息厅
existential crises phr. 0:53
生存危机;关乎存亡的危机
accustomed to phr. 0:53
习惯于
unprecedented /ʌnˈpresɪdentɪd/ adj. 1:42
前所未有的
precedents /ˈpresɪdənts/ n. 1:42
先例
prologue /ˈproʊlɔːɡ/ n. 1:42
开场白、序言
semiformally /ˌsemiˈfɔːrməli/ adv. 1:42
半形式化地
inconsistent /ˌɪnkənˈsɪstənt/ adj. 2:21
(逻辑)不相容的、矛盾的
axioms /ˈæksiəmz/ n. 2:21
公理
traumatic /trəˈmætɪk/ adj. 2:21
创伤性的、令人痛苦的
turbulent /ˈtɜːrbjələnt/ adj. 2:21
动荡的、混乱的
orthodox /ˈɔːrθədɑːks/ adj. 3:22
正统的、公认的
consensus /kənˈsensəs/ n. 3:22
共识
proof assistant n. 3:22
证明助手(形式化验证软件)
strenuously /ˈstrenjuəsli/ adv. 3:22
费力地、严格地
resilient /rɪˈzɪliənt/ adj. 4:19
有韧性的、能迅速恢复的
incorporate /ɪnˈkɔːrpəreɪt/ v. 4:19
纳入、吸收
advent /ˈædvent/ n. 5:19
(重要事物的)到来、出现
lay down some rules phr. 5:19
制定规则
conjecture /kənˈdʒektʃər/ n. 6:08
猜想(数学中未证明的命题)
tackle /ˈtækl/ v. 6:08
着手处理(难题)
placeholders /ˈpleɪsˌhoʊldərz/ n. 6:50
占位符
non-trivial /ˌnɑːnˈtrɪviəl/ adj. 6:50
非平凡的、不可忽视的
devolved /dɪˈvɑːlvd/ v. 7:47
退化、蜕变(devolve into)
distracting from phr. 7:47
分散对……的注意力
novelty /ˈnɑːvəlti/ n. 8:32
新奇玩意儿
business as usual phr. 8:32
一切照常
ostensibly /ɑːˈstensəbli/ adv. 8:32
表面上、名义上
common rooms n. 9:26
(学院、系里的)公共休息室
counter claims n. 10:34
反驳主张
selectively disclosed phr. 10:34
选择性披露的
incentives /ɪnˈsentɪvz/ n. 11:05
激励、动机
in as favorable a light as possible phr. 11:05
以尽可能有利的方式呈现
conflate /kənˈfleɪt/ v. 11:05
混为一谈
grassroots /ˈɡræsruːts/ adj. 11:41
草根的、自下而上的
harnesses /ˈhɑːrnɪsɪz/ n. 11:41
(AI)运行框架、封装系统
autonomous /ɔːˈtɑːnəməs/ adj. 11:41
自主的、无人干预的
publication quality n. 12:45
可发表的质量
complement /ˈkɑːmplɪmənt/ n. 13:37
补集;补充
working hypothesis n. 13:37
工作假设
conditional analysis n. 13:37
条件分析
condition on phr. 14:37
以……为条件
explicit /ɪkˈsplɪsɪt/ adj. 14:37
明确的、显性的
implicit /ɪmˈplɪsɪt/ adj. 15:35
隐含的、隐性的
the humanities n. 15:35
人文学科
have the luxury of phr. 15:35
有……的余裕/奢侈
overdue /ˌoʊvərˈduː/ adj. 16:38
早该做的、逾期的
compiled /kəmˈpaɪld/ v. 16:38
汇编、整理
cumulative /ˈkjuːmjəleɪtɪv/ adj. 17:34
累积的
downplay /ˌdaʊnˈpleɪ/ v. 17:34
轻描淡写、低估
aligned /əˈlaɪnd/ adj. 18:26
一致的、对齐的
proxy /ˈprɑːksi/ n. 18:26
代理、替代指标
ceases to be phr. 19:20
不再是
at the expense of phr. 20:03
以……为代价
untethered /ʌnˈteðərd/ adj. 20:03
不受束缚的、脱离的
rubric /ˈruːbrɪk/ n. 20:44
评分标准、评价准则
in so far as phr. 21:32
就……而言
deconstruct /ˌdiːkənˈstrʌkt/ v. 21:32
拆解、解构
flow network n. 23:19
流网络(图论概念)
sync /sɪŋk/ n. 23:19
此处为 sink,汇点
complimentary /ˌkɑːmplɪˈmentri/ adj. 24:18
此处意为 complementary,互补的
vouch /vaʊtʃ/ v. 25:18
担保(vouch for)
backwalk n. 25:18
此处意为 backlog,积压
unwelcome /ʌnˈwelkəm/ adj. 26:14
不受欢迎的
artifacts /ˈɑːrtɪfækts/ n. 26:14
人工制品;此处指孤立的产出物
exposition /ˌekspəˈzɪʃn/ n. 27:05
阐述、(数学)写作表达
lema /ˈlemə/ n. 27:05
即 lemma,引理
scoring function n. 28:03
评分函数
slick /slɪk/ adj. 29:02
过于流畅的、油滑的(含贬义)
sanded down phr. 29:02
打磨掉、磨平
natural friction n. 29:02
自然摩擦(陶自创术语)
sail through phr. 29:02
轻松通过
blast through phr. 29:58
一冲而过、快速带过
paradoxically /ˌpærəˈdɑːksɪkli/ adv. 29:58
自相矛盾地、反常地
annotation /ˌænəˈteɪʃn/ n. 31:02
批注
fought my way through phr. 31:02
艰难地啃完
digesting /daɪˈdʒestɪŋ/ v. 32:02
消化(知识)
production quer n. 32:02
即 production quota,生产指标
sufficient condition n. 32:39
充分条件
scarce /skers/ adj. 33:41
稀缺的
flooded with phr. 33:41
被……淹没
bottlenecks /ˈbɑːtlneks/ n. 33:41
瓶颈
induce /ɪnˈduːs/ v. 33:41
诱导、促使
opaque /oʊˈpeɪk/ adj. 34:46
不透明的
look under the hood phr. 34:46
查看内部运作机制
corporate secrets n. 34:46
商业机密
reputable /ˈrepjətəbl/ adj. 35:35
有声望的
appetizing /ˈæpɪtaɪzɪŋ/ adj. 35:35
诱人的、开胃的
referee /ˌrefəˈriː/ n./v. 36:42
审稿人;审稿
prestigious /preˈstɪdʒəs/ adj. 36:42
有威望的
gamed /ɡeɪmd/ v. 37:43
被钻空子、被操纵
take … out of the equation phr. 37:43
把……排除在考虑之外
coherent /koʊˈhɪrənt/ adj. 38:41
连贯的、条理清晰的
canonicalization /kəˌnɑːnɪkələˈzeɪʃn/ n. 38:41
正典化、标准化
definitive /dɪˈfɪnətɪv/ adj. 39:30
权威的、决定性的
dig through phr. 40:30
翻找、深挖
unlocked /ʌnˈlɑːkt/ v. 40:30
解锁、释放
internalize /ɪnˈtɜːrnəlaɪz/ v. 41:28
内化
indigestion /ˌɪndɪˈdʒestʃən/ n. 42:19
消化不良
impedance mismatching n. 42:19
阻抗失配(电路术语)
extrapolate /ɪkˈstræpəleɪt/ v. 42:57
外推、推断
pileup /ˈpaɪlʌp/ n. 42:57
堆积、积压
abundance /əˈbʌndəns/ n. 42:57
丰裕、过剩
famine /ˈfæmɪn/ n. 43:43
饥荒
malnutrition /ˌmælnuːˈtrɪʃn/ n. 43:43
营养不良
edible /ˈedəbl/ adj. 43:43
可食用的
signaries /ˈsɪɡnətɔːriz/ n. 45:03
即 signatories,签署者
iterations /ˌɪtəˈreɪʃnz/ n. 45:03
迭代、修订版本
normalize /ˈnɔːrməlaɪz/ v. 46:02
使成为常态
in that vein phr. 46:02
本着这种精神、循此思路
incumbent on phr. 46:02
是……的责任
formalization /ˌfɔːrməlaɪˈzeɪʃn/ n. 47:03
形式化
attribute /əˈtrɪbjuːt/ v. 47:51
归功于、署名归属
on the table phr. 48:26
摆上桌面(供讨论)
innate /ɪˈneɪt/ adj. 49:24
内在的、天生的
take the initiative phr. 49:24
主动采取行动
venues /ˈvenjuːz/ n. 50:25
(发表的)场所、渠道
理解自测 · 11 题
1. 陶哲轩用哪一段历史来类比当前 AI 对数学的冲击?这段历史的起因与结果分别是什么?

他用一百年前的「数学基础危机」作类比。起因是 20 世纪之前数学家以半形式化、朴素的方式使用集合、数、函数、极限等概念,罗素指出朴素集合论不相容(可构造既包含又不包含自身的集合),哥德尔又证明了不完备定理,迫使数学家重新思考「什么是证明」。过程动荡且痛苦,但结果是获得了一阶逻辑加 ZFC 这样的标准框架,被检验了一个世纪。陶在开场章节据此论证:当下的危机同样会带来更健康、更有韧性的学科。

2. FirstProof 挑战赛的评测方式和 2026 年 5 月一轮的结果是什么?

FirstProof 是数学家发起的草根独立评测,不受 AI 公司赞助。数学家捐出 10 道未发表的研究级问题,用于测试完全自主的 AI 系统,且不只看是否解出,还看写作是否达到可发表质量。5 月一轮有四套 harness 提交,陶参与了其中一个团队。单个系统最好成绩是 10 题中 5 题,四套合计有 7 题达到可发表水平,每题成本从 10 美元到 1000 美元不等。陶称这是目前「最科学的数据点」,但强调后续测试可能改变这些数字。

3. 陶把「解题」拆解成了哪五个阶段?其中哪些阶段被 AI 加速、哪些没有?

五个阶段是:证明生成、证明验证、证明解释(写作与理解)、共同体接受(发表与审稿)、正典化(进入教科书与课程)。前两个阶段已明显加速:生成靠 AI,验证主要靠 Lean 等证明助手语言的进步。第三阶段 AI 仍很弱,只擅长拼写格式而重点分配失衡;第四阶段依赖人类反应,AI「基本没有影响」;第五阶段需要教学、反馈和共识,AI「基本完全无用」。陶把后三阶段统称为「证明消化」。

4. 陶在演讲中如何披露自己对 AI 的使用?他为什么要这样做?

他说明本次演讲的文字由自己撰写,但有几处用了自动补全,图表则由 AI 生成,并调侃破折号是人类打的。这样做是为了践行莱顿宣言的第一条个人建议「始终披露 AI 工具使用」。他解释理由:当前最应避免的场景是人人暗中用 AI、无人披露、无人分享最佳实践与错误,这会阻碍整个共同体学会正确使用 AI。因此需要把负责任的披露变成常态,他以身作则。这一段位于结尾的共同体应对建议部分。

5. 陶为何认为「过去不必明说数学的多重目标」,而现在必须明说?请复述其论证链。

陶的论证分三步。第一,数学有多重目标(解题、建理论、理解世界、教育、审美等),但过去数学太难,离所有目标都很远,目标之间大致正相关,因此朝一个目标前进就等于朝所有目标前进,一个显性目标可以作为其他目标的代理。第二,古德哈特定律指出,当某个指标被过度优化,进一步的进展会以牺牲其他目标为代价。第三,人类不擅长优化所以问题不大,但 AI 和 AI 公司是极强的优化器,会使目标从正相关变成竞争关系。因此现在必须同时明确多个目标并给出清晰的评分标准。

6. 为什么陶说「过于流畅的证明」反而可能有害?他用什么亲身经历支持这一观点?

陶提出「自然摩擦」概念:人类写证明时在易处快速带过、在难处放慢并仔细组织,读者能借此感知难点所在;AI 生成的证明对易难之处一视同仁,读者失去这一信号,而过于光滑的证明会把困难「打磨掉」,读者回家自己重构时却做不出来。他的例子是研究生时期读 Jean Bourgain 1991 年关于挂谷猜想与限制猜想的论文,痛苦到批注「我讨厌 Jean Bourgain」,但在 Eli Stein 和 Tom Wolff 帮助下攻克后,理解了他的思维方式,几年后反而最偏爱读他的论文。他推断若该论文被 AI 层层润色,自己可能得不到那样的教育。

7. 陶为什么说 AI 在数学上的成功本身依赖于「正典化」?这一论点有什么反直觉之处?

陶指出,AI 之所以擅长数学,很大原因是几个世纪以来人类已把线性代数、群论等理论整理成成熟的正典定义和教科书,AI 吸收了这些材料。正典化是解题流程的终点,也是数学对工程、物理、生物等领域产生应用的必经形态——外行不会去啃顶刊论文,他们需要教科书。反直觉之处在于:AI 恰恰在这一最有价值的阶段「基本完全无用」,却又以这一阶段的产物为食。因此如果 AI 生成的证明堆积而无人正典化,长期会同时损害数学和 AI 自身。这一论点出现在「正典化」章节。

8. 「证明消化不良」这一概念是如何从流程分析中推导出来的?它有哪些具体表现?

陶先把解题拆成五个阶段,然后观察到 AI 在前两阶段(生成、验证)远强于后三阶段(解释、接受、正典化),各环节速率不匹配,他借用电路术语称之为「阻抗失配」,通俗说法就是「证明消化不良」。具体表现有四种:未验证的证明不断堆积;已验证但无人愿读的证明(如 Lean 形式化的埃尔德什问题解答);可读但无人愿意发表的证明;以及按趋势外推即将出现的、已发表却无人知道如何写进教科书的 AI 证明。他进一步把这归类为「过剩问题」,与过去几个世纪的「稀缺问题」对照。

9. 陶为什么坚持「不能把人类审稿人从环节中剔除」?如果有人反驳「AI 审稿更快更便宜」,他会如何回应?

陶的立场是期刊可以用 AI 作初筛、可以把规范定得更精确,但纯 AI 审稿的期刊不会成功。他会从三方面回应「更快更便宜」的反驳:第一,AI 审稿工具「相当容易被钻空子」,一旦成为目标就会遭遇古德哈特定律;第二,审稿的本质不只是判断对错,而是把个人成就转化为集体理解,并借此培养下一代数学家,这是人与人之间的过程;第三,「被 100 个 AI 审稿人接受」并不等于共同体接受,不意味着任何人真的想读它。速度和成本解决的是前两阶段的问题,而接受阶段的瓶颈本来就是人类注意力。

10. 陶提出的作者署名标准是什么?这一标准若应用于大型协作项目或纯形式化证明,还成立吗?

陶的标准是:如果一个「作者」不能像在讲台上那样报告结果并回答关于该结果的提问,就不该署名,或至少论文在有一个能真正谈论该结果的人之前不应发表。这一标准的核心是「有人能为结果负责并解释它」,而非「每个作者都懂全部」。放到大型协作项目上,只要至少有一人能整体解释,标准仍成立,且与陶对第三阶段「证明必须被理解」的要求一致。但对纯形式化证明(如 Lean 验证的埃尔德什问题解答),当前恰恰出现了「提交者说自己不够格评价、检查者也不能担保」的情形,按陶的标准这些结果尚不能算「可发表」,这正是他想用该标准来约束的场景。

11. 陶用「食物从稀缺到过剩」类比数学从「证明稀缺」到「证明过剩」。这个类比的解释力边界在哪里?

类比的强处在于:它把问题从「AI 好不好」转换为「我们如何建立品味」——人类应对食物过剩靠的是分辨好坏饮食、提升烹饪、拒绝技术上可食用但不想吃的东西,对应到数学就是建立评价证明价值的标准、只投入稀缺的审稿注意力给值得消化的结果。它也解释了为何「讲故事、展示过程」变得重要,因为要让论文更「开胃」。类比的边界在于:食物的好坏有相对客观的营养学标准,而陶自己承认「证明写得好不好」目前没有好的评分标准;此外食物消费是个体选择,而正典化需要整个领域的共识,个体品味无法直接汇总成集体正典。因此类比适合说明「态度转变」,但对「如何形成共识」这一最慢阶段解释力有限。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.009"What Is a Strange Loop and What is it Like To Be One?" by Douglas Hofstadter (2013) 下一期 · NO.011 →AI and human evolution | Yuval Noah Harari
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com