视频库 / NO.010ASK THE BEST MINDS THE BIG QUESTIONS一人,一实验室
视频库 / NO.010
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。
第 10 期 · 回应 Ⅳ·07「研究是怎样做成的?」

ICM 2026 Public Lectures - Terence Tao

节目发布 2026-08-13 · Simons Foundation
陶哲轩 主持人 Alex
章节 · 点击跳转视频
0:00 开场:介绍陶哲轩与讲座主题 ▶ 正在看
1:42 一百年前的数学基础危机 ▶ 正在看
4:19 今日危机:价值观而非论证 ▶ 正在看
6:50 AI 能力猜想与 FirstProof 数据 ▶ 正在看
13:37 条件分析:转向目标与价值观 ▶ 正在看
18:26 古德哈特定律:目标开始冲突 ▶ 正在看
22:34 解题的五阶段流网络模型 ▶ 正在看
27:05 自然摩擦:证明需要被理解 ▶ 正在看
32:39 接受与审稿:最稀缺的瓶颈 ▶ 正在看
38:41 正典化:AI 无能为力的终点 ▶ 正在看
42:19 证明消化不良与过剩时代 ▶ 正在看
44:31 莱顿宣言与共同体应对建议 ▶ 正在看
本期讲者
陶哲轩UCLA 数学系杰出教授、James and Carol Collins 讲席教授,2006 年菲尔兹奖与 2014 年数学突破奖得主。研究横跨调和分析、偏微分方程、组合数论,近年积极参与 AI 与形式化证明在数学中的应用与讨论。
主持人 Alex本场 ICM 公开讲座主持人,为陶哲轩致介绍词;陶在演讲中提到「正典化」(canonicalization)一词由其命名。
01开场:介绍陶哲轩与讲座主题
0:00
Welcome everyone to the public lecture of the International Congress of Mathematicians. It's a great pleasure and honor for me to introduce our speaker uh Terrence Tao. Uh Terry received his PhD from Eli Stein at Princeton University and then uh moved to UCLA where he's been ever since where he's now the distinguished professor and the James and Carol Collins chair in mathematics. Um, if I tried to list all of his various research accomplishments and accolades, we'd be here all night and you'd never get to hear a word from him. So, I'll just mention that he won the Fields Medal uh at a previous ICM, the Breakthrough Prize, many, many other awards. Um, we will take questions at the end. Uh, there's going to be a 10-minute Q&A uh outside in the foyer to the right. And otherwise, I hope you'll join me in welcoming Professor Terry Tao.
欢迎大家来到国际数学家大会的公开讲座。我非常荣幸也非常高兴地为大家介绍我们的演讲者,陶哲轩。陶教授在普林斯顿大学师从 Eli Stein 取得博士学位,之后前往 UCLA,此后一直在那里任教,现在是那里的杰出教授,以及 James and Carol Collins 数学讲席教授。嗯,如果我想把他所有的研究成就和荣誉都列一遍,我们就得在这儿待上一整晚,而你们根本听不到他说一个字。所以我只提一下,他在此前的一届 ICM 上获得了菲尔兹奖,还有突破奖,以及许许多多其他奖项。嗯,我们会在最后回答问题。会有一个10 分钟的问答环节,就在外面右手边的门厅里。那么,现在请和我一起欢迎陶哲轩教授。
便签笔记
0:53
>> [applause] >> Thank you. >> Well, thank you very much. It's such a pleasure to be here again at a at an ICM in person. It's been quite a while and I've uh reconnected a lot of old friends and uh and just absorbing the uh the positive energy. Um so my my talk is about maybe the topic everyone is talking about these days uh the the era of AI and what it means for mathematics. Uh so you know we seem to live in an era of great change and that's something that's not uh we might not be accustomed to because mathematics has been very stable for over a century. Uh things we haven't had real big existential crises or anything but now suddenly we we we uh we seem to be uh changing quite a lot.
>> [掌声] >> 谢谢。>> 嗯,非常感谢。能再次亲身来到 ICM 现场,真是太高兴了。已经有好一阵子没这样了,我在这里重新联系上了很多老朋友,也在吸收这里的正能量。嗯,我的演讲讲的可能是这些天大家都在谈论的话题——AI 时代,以及它对数学意味着什么。嗯,你知道,我们似乎生活在一个剧变的时代,而这是我们可能不太习惯的,因为数学在过去一个多世纪里一直非常稳定。我们没有遇到过真正重大的生存危机之类的事情,但现在突然之间,我们似乎正在发生很大的变化。
便签笔记
02一百年前的数学基础危机
1:42
Um and it it seems unprecedented. Um but actually I think there are historical precedents for um what we're going through. Um so um as a prologue I'll talk about um the crisis in foundations which happened a century ago. Um so before the uh turn of the 20th century uh we mathematicians operated semiformally you know we had uklid we knew about proof but uh we didn't fully acize our um foundations of logic and and and sets and analysis and and you know um we largely use kind of informal naive foundations. Um so questions like what actually is a set or what is a number? What is a function?
而且这看起来是前所未有的。但其实我认为,我们正在经历的事情是有历史先例的。嗯,所以作为开场我要讲讲一个世纪前发生的数学基础危机。嗯,在 20 世纪之交以前,我们数学家是以半形式化的方式工作的。你知道,我们有欧几里得,我们懂得证明,但我们并没有把逻辑、集合和分析的基础完全公理化,而且,你知道,我们在很大程度上使用的是一种非正式的、朴素的基础。嗯,所以像什么是集合、什么是数、什么是函数这样的问题,
便签笔记
2:21
What is a limit? Uh we had some uh rules but we didn't really write down fully the actions of mathematics. Um and that worked uh until the early 20th century uh when suddenly these philosophers who were the only ones who were thinking about these questions pointed out that there were actually some serious problems with doing things naively. So Bertrren Russell for example was a philosopher who pointed out that naive set theory was inconsistent. um you could construct a set that both contained itself and didn't didn't contain itself. Um and of course Guro has incompleteness theorems that there's no system of axioms that could ever completely um prove or disprove all the the questions you could ask in say arithmetic. Um and so we were forced to actually really think about our foundations and what uh uh what is a proof? What are the rules of our um of our profession? Um and that was traumatic. It is called the crisis in foundations. Um and it was turbulent. It was a lot of debate and argument. Uh but
什么是极限?我们有一些规则,但并没有真正把数学的公理完整地写下来。嗯,而这一直行得通,直到 20 世纪初,突然之间,那些当时唯一在思考这些问题的哲学家指出,朴素地处理这些事情其实存在一些严重的问题。比如伯特兰·罗素,他是一位哲学家,他指出朴素集合论是不相容的。你可以构造一个集合,它既包含自身又不包含自身。当然还有哥德尔的不完备性定理:不存在这样一个公理系统,能够完全证明或否证你在比如算术中所能提出的一切问题。嗯,于是我们被迫真正去思考我们的基础,思考什么是证明?我们这个行当的规则是什么?嗯,那是令人痛苦的。它被称为数学基础危机。嗯,那是一段动荡的时期,有大量的争论和辩论。但
便签笔记
3:22
we came out of it at the end with something very valuable. We had a standard uh framework for doing mathematics. You know we we have an an orthodox set of rules first order logic zera franco choice set theory uh and it's standardized and we have a consensus that it's it's good enough to do almost all of mathematics. Um now it's not the final word. uh people still study other foundations compare them to to ZFC um you know for example um um the the lean um formal proof assistant language is not based on ZFC it's based on on dependent type theory for instance um but that's fine uh we we are studying other foundations but not from a position of crisis right from we're studying it just like any other scientific mathematical object to study um and we've tested these foundations you know strenuously over the centuries and and we trust you know that we you I mean there's still theoretically the possibility that that there is a serious problem in in these foundations but but they seem to work and they they've
最终我们走出来了,并且收获了非常宝贵的东西。我们有了一套做数学的标准框架。你知道,我们有一个一套正统的规则体系——一阶逻辑、策梅洛-弗兰克尔选择公理集合论(ZFC)——它已经标准化了,我们也有共识认为它足以完成几乎所有的数学。嗯,当然这不是最终定论。人们仍然在研究其他基础体系,把它们和ZFC 做比较。比如说,Lean 这个形式化证明助手语言就不是基于 ZFC 的,它是基于依赖类型论的。但这没问题,我们研究其他基础体系,并不是出于某种危机感,我们研究它就像研究任何其他科学或数学对象一样。而且我们已经对这些基础经过几个世纪的严格检验,我们信任它。我是说,理论上仍然存在这些基础中有严重问题的可能性,但它们看起来是work的,而且它们已经
便签笔记
03今日危机:价值观而非论证
4:19
worked for a century. So I will argue that one century after uh that crisis uh we are now entering a similarly turbulent period uh and just for people who like uh spotting m dashes in text uh this was a human generated mdash just to to emphasize that um but it's not a crisis in um in our mathematical arguments uh but it's a crisis in our mathematical values and practices um that we have we had kind of been operating on naive foundations of what mathematics is. Um and suddenly they are um um reaching rather strange conclusions. Um and we do need to to reexamine and and and rebuild these foundations properly. But again, it will be a very worthwhile experience. Uh once we do that, our our subject will be much healthier, much more resilient. Um and we will be able to incorporate all these new technologies uh um properly.
运转了一个世纪。所以我要提出的观点是,在那场危机过去一个世纪之后,我们现在正进入一个同样动荡的时期——顺便说一句,对于喜欢在文本里找破折号的人,这个破折号是人类打出来的,只是为了强调一下。但这不是我们数学论证中的危机,而是我们数学价值观和实践中的危机。我们某种程度上一直是在一套关于「数学是什么」的朴素基础上运作的。而突然之间,它们得出了相当奇怪的结论。我们确实需要重新审视,并且好好地重建这些基础。但同样,这将会是一段非常有价值的经历。一旦我们做到了,我们这个学科会健康得多、韧性强得多。而且我们将能够恰当地把所有这些新技术纳入进来。
便签笔记
5:19
All right. So uh what is the topic of my talk? Um so um I can phrase it as sort of you know so the big question if you wish that motivates everything is what I call the community response question. Okay. So how should the mathematical community respond to the advent of modern AI technologies and all the related things that that come with it? Um and you know the fact that they e they either have or they claim to have capabilities to perform mathematical tasks. Um so this is a big question and it's a question for the entire community. Um you know no one not even the IMU can just lay down some rules and we can all follow them. We have to actually uh debate this question just like how we had to debate the foundations of of mathematical logic. Um but I do have some comments uh to sort of start discussion.
好,那么我这次演讲的主题是什么呢?我可以这样表述:如果你愿意的话,驱动这一切的大问题就是我所说的「共同体应对问题」。好的。那就是:数学共同体应该如何应对现代 AI 技术的到来,以及随之而来的所有相关事物?以及它们要么已经具备、要么声称具备执行数学任务的能力这一事实。这是个大问题,而且是整个共同体的问题。没有任何人——哪怕是国际数学联盟(IMU)——能直接定下一些规则让我们所有人遵守。我们必须真正地辩论这个问题,就像当年我们必须辩论数理逻辑的基础一样。不过我确实有一些看法,可以算是抛砖引玉。
便签笔记
6:08
Um now this is not a math question. This is not this is not a conjecture where there's going to be a proof and a and a single answer. this is a meta mathematical question um and you know also a political question ethical question you know cultural economic whatever um so it's not a math question however I think the mindset of thinking very carefully and mathematically about about problems is very useful so I'm going to um tackle this question pseudoathematically I'm going to use the language of mathematics uh to try to analyze this question even though it is not a mathematical question uh and hopefully because so many of you in the audience are mathematicians. This will help uh clarify the points I'm trying to make.
这不是一个数学问题。这不是一个会有证明、会有唯一答案的猜想。这是一个元数学问题,同时也是一个政治问题、伦理问题,还有文化的、经济的,等等。所以它不是数学问题。但是我认为,那种非常审慎地、数学式地思考问题的心态是很有用的。所以我打算用「伪数学」的方式来处理这个问题,我会用数学的语言来分析这个问题,尽管它并不是一个数学问题。希望因为在座各位有那么多是数学家,这样能帮助我把想讲的要点讲清楚。
便签笔记
04AI 能力猜想与 FirstProof 数据
6:50
Okay. So, we're interested in this community response question. What should we do in response to AI? And so there's a sub question which is has dominated debate over the last three years. Um uh which but it's a slightly different question. Um and so this I will call the AI capability conjecture. So I said I'm using the language of mathematics. So I'm going to call this a conjecture. Um and actually technically it's not a single conjecture, it's a family of conjectures. Um so this conjecture I'm stating here as a template, there's some placeholders or variables if you want to keep the mathematical language. Um and so uh it it's so because there's so many variables, this conjecture could mean almost anything, but I will say it anyway. So the conjecture states that at some point in the near future some AI tools will with some expense and some level of human supervision be able to accomplish some research level mathematical tasks in some fields of mathematics with some non-trivial success rate and some level of
好,我们关心的是这个共同体应对问题:面对 AI 我们该怎么办?这里还有一个子问题,它在过去三年主导了整个讨论。但它其实是一个略有不同的问题。我把它称为「AI 能力猜想」。我说过我要用数学的语言,所以我就把它叫做一个猜想。而严格来说,它其实不是单个猜想,而是一族猜想。我在这里陈述的这个猜想是一个模板,里面有一些占位符,或者说变量——如果你想继续用数学语言的话。正因为变量这么多,这个猜想几乎可以意味着任何东西,但我还是要把它说出来。这个猜想是说:在不久的将来的某个时刻,某些 AI工具将以某种成本、在某种程度的人类监督下,能够在数学的某些领域完成某些研究级别的数学任务,达到某种非平凡的成功率,以及某种程度的正确性和质量。我在里面放了这么多「某种」,以至于这
便签笔记
7:47
correctness and quality. Now um I've put so many sums in here that that this could this could be almost anything. Um and so a lot of the debate has sort of devolved in recent years as to exactly which value of sum should you put in different slots here. Um and that's not the point of my talk today. Uh I think that's actually distracting from some some more fundamental issues. Um so just think of this as family conjectures. Um of course you can make this a very weak conjecture by making a sum just a few or or you know um or you can make it very strong. you can just sort of insert in your own personal mental mad liib you know what what you can put here. So let me just sort of vaguely talk about strong and weak and intermediate forms of this conjecture.
几乎可以是任何东西。所以近年来很多辩论其实都退化成了:到底该在各个空位上填入哪个「某种」的取值。而这不是我今天演讲的重点。我认为那实际上分散了对一些更根本问题的注意力。所以就把它当作一族猜想吧。当然,你可以把「某些」设成「寥寥几个」,让它变成一个非常弱的猜想;或者你也可以让它变得非常强。你可以在自己脑子里玩个填词游戏,看看该往里填什么。所以让我就笼统地谈谈这个猜想的强形式、弱形式和中间形式。
便签笔记
8:32
All right. So it's kind of clear that this knowing the answer to this conjecture or family conjectures is important because you know for example if even weak versions of this conjecture happen to be false but that you know that these tools were largely useless for research level tasks then there's no debate you know I mean these are novelty you know some small minority of people could play with them but you know math could basically continue business as usual. Um on the other hand if the strongest versions of this conjecture are true and every single task that mathematicians do today uh AIs can do better and faster and whatever then yeah then clearly uh we cannot continue business as usual. Um particularly if we focus on the things that the AIs do best. Uh for example solving unsolved problems seems to be one of their strengths and that seems to be one of the things that we do a lot. Um we even give out medals for these things ostensibly. Um so then we have to actually uh think about uh our culture
好。很明显,知道这个猜想或这族猜想的答案是重要的。因为举例来说,如果连这个猜想的弱版本都不成立,也就是说这些工具对研究级别的任务基本无用,那就没什么好争的了。它们只是个新奇玩意儿,少数人可以拿来玩玩,但数学基本上可以照常运转。另一方面,如果这个猜想最强的版本是真的,今天数学家做的每一件事 AI 都能做得更好更快,那么是的,那我们显然就不能照常运转了。尤其是如果我们聚焦于 AI 最擅长的事情。比如说,解决未解决的问题似乎是它们的强项之一,而这恰恰又是我们做得很多的事情之一。我们甚至表面上还为这类事情颁发奖章。所以那时我们就必须真正地思考我们的文化和实践。而如果成立的是介于
便签笔记
9:26
and practices and if it's something in the middle is true then we we have to think even more carefully. Okay. So as I said if you look at all the debates in the media and probably on at coffee uh you know at common rooms and and on the internet um you know because the answer to the community response question depends so much on the answer to the AI capability conjecture. So much of the debate is has been about capability um and um yeah so and and I have certainly um participated in this if you look at what uh I've spoken about AI a lot in the last three years and I talk a lot about the capability because it is important um and um um there was certainly a time in the like three three years ago where the capability was very poor but it was trending to be a lot better and um we really had to clarify the dist distinction between where we were and where the technology was going um but this talk is not about that conjecture. Um oh hang on I will get to that but um um now if you follow the uh internet I'm sure you have heard many
两者之间的情形,那我们就得思考得更仔细了。好,正如我所说,如果你去看媒体上的所有辩论,可能还有喝咖啡时、公共休息室里以及互联网上的辩论——因为共同体应对问题的答案在很大程度上取决于AI 能力猜想的答案,所以大量的辩论都是围绕能力展开的。而我自己当然也参与过这些讨论。你看我过去三年谈了很多关于 AI 的话题,我也大量谈论能力问题,因为这确实重要。而且在大约三年前,确实有那么一段时间,能力还非常差,但趋势是会好得多。我们真的必须澄清我们当时所处的位置和技术正走向何方之间的区别。但这次演讲不是关于那个猜想的。哦等等,我会讲到那个,不过——现在如果你关注互联网,我相信你已经听过许许多多的说法,还有反驳,以及关于
便签笔记
10:34
many claims and um um and and and counter claims and and arguments about whether this conjecture is true or false in various levels. Um and so we have a lot of data points now. This is you know a specific problem has been solved and then but then uh this this AI can't even you know multiply two-digit numbers or or whatever. Um most of the data we have is not gathered under proper scientific conditions. Um we have sort of uh selectively disclosed results which look impressive but we don't know exactly what the inputs were.
这个猜想在各个层面上是真是假的论证。所以我们现在有很多数据点了。比如某个具体问题被解决了,但接着又有人说这个 AI 连两位数乘法都不会,诸如此类。我们掌握的大部分数据都不是在恰当的科学条件下收集的。我们看到的是那种被选择性披露的结果,看起来很惊艳,但我们并不确切知道输入是什么。
便签笔记
11:05
We don't know the the proper failure rate. Um and um many of the uh uh companies that are disclosing these have their own incentives to to maybe uh uh present the results in as favorable a light as possible and not review important information like say how much cost it was to to actually produce these results. Um so it is it is a very confusing mess uh right now. Um this this whole uh all the data points for against this conjecture. Um and also another point is is that um people sometimes conflate whether a con this conjecture is true or false as to whether we want it to be true or false.
我们不知道真实的失败率。而且很多披露这些结果的公司有他们自己的动机,可能会把结果尽可能地往好看的方向呈现,而不去披露一些重要信息,比如说产生这些结果到底花了多少钱。所以现在这一切非常混乱。支持或反对这个猜想的所有这些数据点都是如此。另外一点是,人们有时会把这个猜想是真是假,和我们「希望」它是真是假混为一谈。
便签笔记
11:41
Uh and I think that is um it it that may not necessarily align either. Um so I won't be talking about this conjecture. Um the one thing I will point you to is that we do have a we are starting to have some more systematic and scientific assessments of of um of AI capability. Um so uh I think the most promising is the first proof challenge which many of you have heard about. Um so this is a grassroots effort run by mathematicians. Um and it's an independent um assessment of the state of of of AI models today. It is not sponsored by any uh major AI company. Uh what it does is that it it creates a batch it mathematicians donate a batch of of 10 research level problems that have not been published. The solutions have not been published. all their all the problems themselves and they test them against completely autonomous AI models. Um and so for example back in May uh there were four harnesses submitted. I was involved in a team that submitted one of them. Um and they tested them to try to to test these
我认为这两者也未必是一致的。所以我不会谈这个猜想。我唯一想给大家指出的是,我们确实开始有了一些更系统、更科学的 AI 能力评估。我认为最有希望的是 FirstProof 挑战赛,你们很多人应该听说过。这是一个由数学家发起的草根项目。它是对当今 AI 模型状态的一个独立评估,没有任何大型 AI 公司赞助。它的做法是:数学家们贡献出一批 10 道尚未发表的研究级别问题,解答也没有发表过,问题本身也没有发表过。然后他们用这些题去测试完全自主的 AI 模型。举例来说,今年五月有四套「harness」(系统方案)提交上来。我也参与了其中一个提交团队。他们用这些问题去测试它们。在这四套方案之间,他们不仅问
便签笔记
12:45
these problems. And between those four harnesses um they didn't just ask whether the the uh these problems were solved but whether they were written up and explained in a way that would be publication quality. like would it would actually be reasonable um to publish in a paper in a journal maybe with some slight um corrections or something. Um and uh uh I think the uh the best individual record I think was like five out of 10 for one harness but collectively they were able to sort of seven out of 10 at at publication level quality but at a non-trivial cost. Uh each harness to spend on each problem the cost ranged anywhere from 10 to1,000 US. Um, so that's roughly where we were at in May. Um, now they're going to keep testing um, these batches and maybe these stats will change, but I would say this is the most scientific uh, data point we have right now.
这些问题是否被解决了,还问解答的书写和讲解是否达到了可发表的质量:也就是说,是否真的可以合理地在期刊论文中发表出来,也许只需要一些小修改。我记得单个方案的最好成绩大概是10 题里做出 5 题;但四套合起来,能达到 10 题里有 7 题达到可发表质量水平,不过成本不低。每套方案在每道题上花费的成本从 10 美元到 1000 美元不等。这大致就是我们五月时的状况。他们还会继续测试新的题批,也许这些数据会变化,但我认为这是我们目前掌握的最科学的数据点。
便签笔记
05条件分析:转向目标与价值观
13:37
Okay, so I'm not going to talk about this conject this conjecture anymore. Um, I'm going to talk about the complement of this conjecture. Um, what um, uh, what suppose this conjecture is true in a certain in a certain sense. What then do we do about commutive response? So I will make a as you know as mathematicians do we'll make a hypothesis okay I'll make a working hypothesis where all the sums that I said before will now be replaced by reasonable okay so a reasonably strong um conjecture is true you know reasonably soon AI will be able to perform a reasonable fraction of mathematical tasks with reasonable levels of success quality and cost and I will not define what reasonable means okay so but um you can sort of put in your mind what um some idea okay Um, so this is an assumption. Okay, in math we sometimes make an assumption. We don't know whether it's true or false, but we are interested in exploring it consequences. I'm going to do a conditional analysis. You can argue whether this conjecture, this hypothesis
好,我就不再谈这个猜想了。我要谈的是这个猜想的「补集」。也就是说,假设这个猜想在某种意义上是成立的,那我们对于共同体应对该怎么办?所以我要做一个假设——就像数学家常做的那样,我们提出一个工作假设:把我之前说的所有「某种」都替换成「合理的」。也就是说,一个合理强的猜想是成立的:在合理可期的将来,AI 将能够以合理的成功率、质量和成本,完成合理比例的数学任务。我不会定义「合理」是什么意思。你可以自己心里有个大概的概念。所以这是一个假设。在数学里我们有时会做假设,我们不知道它是真是假,但我们有兴趣探索它的推论。我要做的是一个条件性分析。你可以争论这个猜想、这个假设是真是假,但那不是这次演讲的重点。我只是要有条件地
便签笔记
14:37
is true or false. That is not the point of this talk. I'm just going to assume it conditionally. So once you condition on the fact that on on the hypothesis that AI will actually solve a a reasonable fraction of uh of things that we currently do today as mathematicians uh what becomes clear is is that in order to proceed to answer the response question we now have to think about our goals and our values. Uh and so really this is I think the important question we need to discuss today. uh I mean not just literally today but but um in this current era uh so this is I call the goals and values question so what are the precise goals objectives and values of our community and of mathematical research in general now we have explicit goals you know we we when we make press releases or when we write abstracts to our papers or when we apply for funding grants we list you know this is the objective of this project or whatever those are explicit goals um but what we're finding is that actually there's
接受它。所以一旦你以「AI 真的能解决我们今天作为数学家所做的事情中相当一部分」这个假设为前提,那么很清楚的一点是:为了接着回答这个应对问题,我们现在必须思考我们的目标和价值观。所以我认为这才是我们今天需要讨论的重要问题——我不只是字面意义上的今天,而是指在当下这个时代。所以这个我称之为「目标与价值观问题」:我们这个共同体、以及整个数学研究,其确切的目标、宗旨和价值观是什么?我们当然有显性目标。当我们发新闻稿,或者写论文摘要,或者申请科研经费时,我们会列出「这个项目的目标是什么」之类的。那些是显性目标。但我们发现,其实还有很多隐性目标,我们
便签笔记
15:35
also a lot of implicit goals that we don't talk about enough um and uh it's certainly very important that we do so uh so let me explain why so in the past you know we stayed we we focused mostly on technical goals you know like um I want to improve this constant you know I I want to to to prove this theorem I want to find a definition that does x y and z okay um and questions about values and goals and you know what is what is the purpose of mathematics we thought that this was a question for the humanities you know maybe philosophers of mathematics would tackle this question or or sociologists of science or um or math educators or or you know but um you know so uh we we have focused maybe overly too much uh on the technical aspects uh because that was what we were best at um and in some ways it's sort of easier to to work with uh despite me math so being uh quite quite difficult some often um but u I'm going to argue that you know if you assume the working hypothesis we will not have the luxury anymore of not
谈得不够多。而我们确实非常有必要去谈。让我解释一下原因。在过去,我们主要关注技术性目标,比如说:我想改进这个常数,我想证明这个定理,我想找到一个能实现某某功能的定义。而关于价值观、目标,以及数学的目的是什么这类问题,我们以为那是人文学科的问题:也许数学哲学家会来处理这个问题,或者科学社会学家,或者数学教育工作者。但是,我们可能过于偏重技术层面了,因为那是我们最擅长的,而且在某些方面它反而更好把握——尽管数学本身常常相当相当困难。但我要说的是,如果你接受这个工作假设,我们将不再有那种奢侈:不去披露、不去真正说明我们的
便签笔记
16:38
disclosing not really disclosing our goals. Um and I think it's healthy. So you know regardless of whether this hypothesis is true or false or partly true or whatever um I think it is actually overdue that we do need to talk about why we do mathematics and what is the real purpose of of our subject. Um and we will we will we will be stronger uh for doing so. So why do we do research? Why do we try to solve problems and and and uh um and explore mathematical concepts and there's actually a lot of reasons um and I don't think anyone has compiled a complete list uh here are just some you know so of course we solve open unsolved problems that is important um both pure problems you know uh problems that that don't have any direct application outside of mathematics but also many important applied problems so very important motivation for mathematics we like to develop new tools new theories new techniques uh we like to understand mathematics better. We like to understand the world around us. You
目标。而且我认为这是健康的。所以,无论这个假设是真是假,还是部分为真,我认为我们其实早就该谈一谈:我们为什么做数学,我们这个学科的真正目的是什么。我们会因此变得更强大。那么,我们为什么做研究?我们为什么要去解决问题、探索数学概念?其实原因有很多,我不认为有谁整理出过一份完整的清单。这里只是其中一些。当然,我们解决开放的未解问题,这很重要——既包括纯粹的问题,也就是在数学之外没有任何直接应用的问题,也包括许多重要的应用问题。这是数学非常重要的动机。
便签笔记
17:34
know, the world is full of patterns as as we saw in in Anna Fry's Lord or um you know and we want to understand um the world. Um but we also you know math is also a very social subject. Um you know um math you know we are excited about math a large part because we can communicate to to the rest of our community and we want to build that community and uh we want to we want to train the next generation. We want to sustain this community. So, you know, we we invest a lot of effort in in lifting up and educating um the next generation of mathematicians who will do things that that we can't dream of. Um and we have this huge database of mathematical knowledge that's been building for thousands of years. We have almost the the oldest continuous continually developing cumulative knowledge base um on the planet and and you know we contribute to that. That's amazing. And this aesthetic value which we often um uh downplay but it is also important.
便签笔记
06古德哈特定律:目标开始冲突
18:26
you know, mathematics is beautiful. Um, and uh, you know, you can create something that that people will admire for centuries and that that also is is is a value and you can, you know, it's your homework, you can add two more u bullet points to this list. So, we have all these goals um, and we tend to only state one or two of them at a time. Um, and until recently that was kind of fine uh, because they were all aligned. Um, so we had all these goals and they're slightly different goals. Um but math was hard and we're very far from attaining many of these goals, you know, solving problems is difficult, building theories is difficult, everything was difficult. Um and so as long as you're far away from all these goals, as long as they're kind of roughly correlated with each other, you can move towards one goal and you you'd also be moving towards other goals as well. So we we so because of this positive correlation um you can just state one goal explicitly and and it can be a proxy for all the others. You don't
便签笔记
19:20
need to state, you know, if you can just say, I'm I'm going to do goal one. I also secretly want to do goal two, goal two, but I know that by pushing towards goal one, I am getting closer to go to two two as well. So, we we didn't really need to to state all our goals. Okay. Um so this works until uh you start optimizing a lot um and you become very good at say goal one. Um and at some point uh you become uh so good at goal one that any further progress towards goal one actually moves you away from goal two, three, four and five. Um and this is a very um well-known uh phenomenon in economics is known as good target law that when a measure becomes a target, it ceases to be a good measure.
便签笔记
20:03
If you optimize too much for any specific metric, uh you can it's it comes at the expense of other goals that you actually wanted to optimize as well. Um and um yeah for humans uh when humans are doing this it's not that much of a problem because humans are not that amazing at optimization uh but AIs are um and also uh AI companies are um and that is um and that is actually a dangerous combination um that AI AIS can optimize uh untethered by by actual reality um and so uh they particularly vulnerable to good hearts law.
便签笔记
20:44
And so we are now entering a period where all the goals that we have are now pointing in opposite in in in competing directions. Um and I think um you know AI in particular may get us closer to one of these goals but at the expense of others and so um we can no longer afford to only uh focus on one goal at a time. Uh and the goals that we do focus on we have to be really clear about what our rubric is for these goals. Um so that's the situation that we are now finding ourselves in. Okay. Now I do not have enough time to talk about all of these goals. Um so I will focus on on just one of the goals but it's it's the one that somehow has attracted the most attention in in in in recent um in in recent events which is problem solving.
便签笔记
21:32
Um so you know if if in so far as as uh the outside world looks at mathematics um I think they have they don't really have any any idea what we do but it seems like but um um one of the things that it looks like we are really obsessed with is solving open problems. So um I'll discuss this prof this aspect of our profession problem solving. just by note as I'm trying as I try to emphasize this is not this just one of our goals but I will deconstruct it um and uh try to make some points about this goal alone um yeah I mean we don't you know um we not all just trying to solve as many problems as possible so yeah many mathematicians are more interested in building beautiful theories or or educating the next generation or whatever these are all very important goals but I will focus on the um uh the goal of problem solving because this is the one which is currently being impacted the most by by AI if we assume this working hypothesis.
便签笔记
07解题的五阶段流网络模型
22:34
Okay. So what is uh the goal here? All right. So as I said externally it looks like you know to an outsider this is the goal of mathematics just solve as many unsolved problems as possible. um kind of like as if you know um you know when I was um when I was a student um when I was a high school student I had no idea you could be a professional mathematician and my when I learned that there was such a job I imagined that there was some committee of senior mathematicians that would somehow assign people problems and we would just all work on our assigned problems as our homework assignments and and that was somehow uh the way this the uh the system worked. um uh and it is totally not the way the system works. But um that is kind of what the impression of it is from the outside sometimes.
便签笔记
23:19
Okay. So you can think of it as as okay again in the pseudo mathematical language. I'm trying to optimize this flow network. I got I've got this source of problems and this sync of solutions and there's this um uh um flow of proof generation and we're trying to maximize the flow from um open problems to solutions. Okay. So um if you try to optimize this goal um so even before AI we realized it was a flaw if we say uh I'm going to um our goal was to collect as many proofs of reman hypothesis as possible. Uh we we figured out very quickly this was not a good goal. Um because you receive lots and lots of solutions to to your goals which problems which were not correct. Um okay so fine. Okay. So uh but we have an obvious way to fix that uh is that we we first generate these proofs and then we check them to be correct. Um so now we our flow has got two steps. You take open problems, you generate proofs, you create unverified solutions but then there's a separate proof verification process and creates verified solutions.
便签笔记
24:18
Okay. Um and AI has become very good at the first step of this uh uh flow for some problems not for all problems but for some problems it's become reason pretty good and and decently good at the second goal as well. Um in fact well in the second case it's not so much because of AI although AI helps uh but because of the complimentary development in proof assistant languages which are a whole topic in itself but there will be other talks about that that's another important uh uh development in recent years the the massive improvement in proof assistant languages. So um so we visibly the proof generation and proof verification uh in mathematics has accelerated in certain fields and for certain types of problems not not uniformly but um uh it has accelerated and if we believe if we condition on this working hypothesis it will continue to accelerate even more um but then this leads to a new problem you know suppose we now have these AI generating these proofs and maybe we even have uh proofs
便签笔记
25:18
that are in some formal language like lean and they're verified Right? And we have these 100,000line proofs that we've verified to be correct, but no one understands them. Um, even the humans who entered in the prompts to make these to to generate the proofs, maybe they they don't understand them either. Um, and this is already happening in some areas. So you maybe you've heard about the Erdish problems. Um, it's a list of about,200 uh problems, some solved, some some unsolved. Um, and proofs for the first time in since in the last six months. um uh we we're starting to get a backwalk you know proofs are being generated um and some are even formalized in lean um and no human has actually gone through them and we don't um there's no human who can ver who can vouch this is this is a good proof um and you know if there's many cases where even the people submitting the proof say here's a proof but I don't I'm not qualified to to evaluate it I don't know whether it's correct or not and then someone else says I've checked
便签笔记
26:14
it in lean I still but I can't I can't vouch for it either Um and it hasn't quite happened yet but we are very very close to a scenario in which a major result gets proved and verified and no human can understand can explain it. Um so that would be uh a very uh unwelcome development. Um so what that means is that there is at least one more stage to the uh if you want to uh so our goal must get updated. It is not enough to uh to generate proofs and it's not enough to verify the proofs. The proofs need to be explained uh well enough that they can be communicated and understood by the mathematical community. Okay, they can't just be artifacts, you know, just sort of impressive achievements that don't lead to anything further in mathematics.
便签笔记
08自然摩擦:证明需要被理解
27:05
They need to be understood. Um so we need to generate solutions that are well written. Now currently AIs while they are accelerating the first two stages of this process they are still quite weak in this third stage. Um they are good at some parts of exposition but not others. Okay so for example at the base level spelling grammar formatting they they're perfect at this almost too perfect to the point where people would prefer to read slightly imperfect text at this point. Um but uh it's not just about just technically being you know spelling correct and so forth. Um often the emphasis is very strange. Um like AI generate mathematics is very frustrating to read. Um you're you're looking for um you know what was the most important what was the most difficult part? always the the interesting part of of a proof and you'll find that the AI spends just as much time on some very trivial lema that that is you know three pages to prove something obvious and then like you know three lines to prove like the
便签笔记
28:03
really interesting part of the of the argument um also they often don't disclose uh where the influences came from you know they're um I mean it's very indirect often they they the proofs come through some some some weights which come from some training data and at some point some some previous result indirectly sort of influenced the result but they they often um uh these these AI generated texts say they often disclose no um um no sense of of where these these connections came from. So it the results often feel a lot more disconnected from the literature than they should um you know so um and maybe this will get better. So if if you assume that that AI capabilities get better you you know they are slowly getting a little bit better at exposition. It's harder uh because um verification of a proof you know this is clear signal it's true or false and and uh you can you can send you can you have a good scoring function and you can optimize it and that's what AI tools are good at whether a proof is well written
便签笔记
29:02
or or not well written u we don't yet have really good rubrics for this but maybe we will and maybe proofs will get um uh easier to read um but that's not necessarily a good thing either um sometimes a proof can be too slick Um maybe you've experienced this if you've if you've read a a a proof that has somehow been where all the difficulty has been sanded down to nothing and like this it seems like the proof is effortless and then you try to reconstruct that at home in pen and paper and you can't do it because uh it has somehow hid the difficulty. Um it actually is important for proofs to contain a little bit of what I call natural friction. Um so when a human tries to prove something there are some steps that are easy and the human will just sail through them and not spend too much time on them. And then there are some steps that really are hard and the human will pause and and take some time and and try to organize the proof carefully and and try not to do too much work and um and the reader can sense
便签笔记
29:58
this and the reader can pick up on this and and understand where the the difficult part of a proof is. Um AI, you know, easy things that will blast through but also difficult things that also blast through and it just looks the same. Um and paradoxically actually you know humans when they um they don't do this and they make mistakes in exposition can actually help the reader in some way. It's it's paradoxical but it's true. Um so I'll give you one example of my own experience. Um so this paper this is a um my copy of a paper of Jean Ban from 1991 which um I read as a grad student. It's about the kaya conjecture and the restriction conjecture. actually uh Hong Wang's uh work uh that was in uh in presented yesterday is based a lot of it is based actually on this particular paper um and for those of you who have never read a paper Jean Bugen uh it is it is a very very educational experience um so um I still have the notes of myself from as a grad student trying to get through this paper um and he was to say you know
便签笔记
31:02
clearly I can do this and you know this this this this statement is is essent this is essentially equal to this and I have all kinds of very frustrated uh question marks and so you probably can't see it but I have a uh annotation here I hate Jean Borgan okay um but I I fought my way through this um I was very fortunate to have Eli Stein explain some of this and Tom Wolf explain and I eventually understood what he was doing and actually it was it was I finally got the way he thought and and why certain things he that I didn't see were equivalent were basically equivalent and I learned so much um you know to the point where actually a few years later I I preferred reading Jean Bane's papers over all the others because I understood how his mind worked and like he just got straight to the point um and it was always a pleasure to read that but if there was definitely a learning process um and I think if his proof had been passed through many layers of AI improvement I may not have gotten the education I did um
非常幸运,有 Eli Stein 给我讲解了一部分,Tom Wolff 也讲了,我最终才明白他在做什么,其实是我终于摸清了他的思路,明白了为什么某些我当初看不出来是等价的东西,本质上确实是等价的,我从中学到了太多,以至于几年之后我读 JeanBourgain 的论文反而比读其他任何人的都更顺手,因为我理解了他的思维方式,他总是直接切入要点,读起来一直是种享受。但这中间确实有一个学习的过程,我在想,如果他的证明经过了层层 AI 的润色加工,我可能就得不到当年那样的教育了。所以说,重点不只是
便签笔记
32:02
yeah so um right so it's it's not just about uh writing papers well. It's about really digesting and understanding the results. Um and there's a famous quote of uh Bill Thirsten to this to this effect that said you know despite appearances we are not trying to meet some abstract production quer of definitions, themsel
把论文写好,而是真正消化并理解这些结果。关于这一点,Bill Thurston 有一句很有名的话大意是说,尽管表面看起来如此,我们并不是在完成某种抽象的生产指标,不是在生产定义、定理……
便签笔记
09接受与审稿:最稀缺的瓶颈
32:39
really well and it and and and and you and it has all the natural friction and it's it's it anyone reading it will actually um benefit from it. Um that's still not enough. Um like in in order to actually influence the future development of mathematics, it needs to be accepted and valued by the community. Um you know and other mathematicians need to actually want to read it and digest it and put it into their own work. And it may not, you know, and of course making it correct and and easy to read helps, but that's only a sufficient condition. It's not necessary. Um, you know, it's like it's like cooking. You know, you can cook a a beautiful meal and, you know, it it smells great and and and all the food has been verified safe to eat and that's, you know, but you can't force people or you should not force people to eat it. Um, >> and you know, and it's it's it's their choice. Um so there is this um you know you have to convince other people that this uh this result is is worth doing.
写得非常好,而且保留了所有自然的"摩擦",任何读它的人都能真正从中受益。但这还不够。要真正影响数学未来的发展方向,它还需要被整个共同体接受和重视。其他数学家得真的愿意去读它、消化它,把它用到自己的工作里。而这未必会发生。当然,把它写正确、写得易读会有帮助,但那只是充分条件,不是必要条件。这就好比做饭:你可以做一顿很棒的饭菜,闻起来香气扑鼻,所有食材也都验证过安全可食用,但你不能强迫别人吃,或者说你不应该强迫别人吃。这是他们的选择。所以你必须去说服别人,让他们相信这个结果值得关注。
便签笔记
33:41
Now this wasn't a problem until very recently because proofs were so scarce that um the moment some some result got proven then naturally people everyone else in the field will just drop everything read the paper and they'd be motivated to um to you know so um you know basically it's like if food was is scarce you will eat whatever is is delivered onto the table but um now we are flooded with proofs uh and more than more than we can we can review uh and we now have to pick and choose um and so we Um we need to be convinced to uh to actually read proofs. Now um proof review is now one of the most precious bottlenecks um scarce resources um in this new era. Um now you can help you can induce people to uh to to read your work. Um telling good stories helps having a narrative showing the process describing uh how you arrived at at at at a um at the result uh what didn't work and how how you went around it. Uh people love stories. Um we have in the past not emphasized process. We've just let the outcomes um speak for themselves
在最近之前,这都不是个问题,因为证明太稀缺了,只要有个结果被证出来,同领域的其他人自然会放下手头一切去读那篇论文,他们有充分的动力去读。基本上就是说,如果食物稀缺,端上桌的东西你都会吃。但现在我们被证明淹没了,多到远超我们能审阅的量,我们现在不得不挑挑拣拣,所以我们需要先被说服,才会真的去读一个证明。如今,证明的评审成了最宝贵的瓶颈、最稀缺的资源,在这个新时代里。你可以做些事情来吸引别人读你的工作。讲好故事有帮助,有叙事、展示过程、描述你是怎么一步步得到这个结果的、哪些路走不通、你又是怎么绕过去的。人们喜欢故事。过去我们不太强调过程,我们只是让结果自己说话——这在过去是行得通的,直到我们找到了办法,能在没有过程的情况下
便签笔记
34:46
which um has worked until we figured out ways to automate you know um um outcomes without process. Um and the current tools they are very very opaque about their process. Um there are things called chain of thought. You can kind of look under the hood a little bit. Um but um it doesn't it's not very insightful. Um now maybe some of this is is just uh can be changed with with good practices. But um uh but currently especially if these if these proofs come from propag models or internal models where uh companies often um incentivize to to to keep certain facts corporate secrets then um yeah u we do not see the process and we and this really makes us a lot less willing to actually invest the time to to to learn about these things because we we don't see the story.
自动产出结果。而现在的这些工具,对自己的过程非常非常不透明。有个东西叫思维链,你多少能掀开盖子看一点,但它并不怎么能给人启发。也许其中一部分可以通过好的实践来改变。但目前,尤其当这些证明来自专有模型或公司内部模型时,企业往往有动机把某些事实当作商业机密保守起来,那我们就看不到过程,这真的让我们大大降低了投入时间去了解这些东西的意愿,因为我们看不到那个故事。
便签笔记
35:35
Um so there is a fourth stage to um to uh um to solving problems. So not not just generating proofs, verifying them and explaining them. They have to be accepted. Um and uh in our system currently the way we uh we have the publication system. So they have to be you know we um the standard way to have a a result accepted by the community is have published in a reputable journal. So they have to be digested and accepted by the mathematical community. And we have a system set up for this. we we have journals um and um but it's slow um as I said you you can encourage it you know you can make your your papers more appetizing to read um but you cannot just try to optimize this by an AI because it it it it involves human response okay unless you somehow wire the human into an AI which I really do not recommend. Um so this is a much slower um part of the process and it is it is one where AI has basically not made an impact currently. Um yeah it is still humans talking to humans that require that that decent acceptance.
所以解决问题还有第四个阶段。不只是生成证明、验证证明、解释证明,它们还必须被接受。在我们现有的体制里,我们有一套出版系统。所以它们必须——一个结果要被共同体接受,标准途径就是发表在有声望的期刊上。所以它们必须被数学共同体消化和接受。我们为此建立了一套系统,我们有期刊,但这个过程很慢。就像我说的,你可以去促进它,可以让自己的论文读起来更有"食欲",但你没法靠 AI 来优化这一环,因为它涉及人的反应——除非你想办法把人接进 AI 里,而我真的非常不建议这么做。所以这是整个流程中慢得多的一部分,也是目前 AI 基本没有产生影响的一环。这依然是人与人之间的交流,才能带来那种真正的接受。
便签笔记
36:42
Now the way we do this right now is that we have journals and when a paper has been verified it's been good enough to submit to a journal we send it to a human referee and they volunteer to to their time as as a service to you know um to referee these papers not just because the result is interesting but they want to encourage the authors to become better mathematicians. It's it's it's a way of of promoting um another generation of mathematicians. Um it's not prestigious work you know I mean the people you know we uh we just gave four medals for solving problems we didn't give four medals for refereeing um but it is an essential part of our prof profession and you know it it it's how we convert individual achievements of mathematicians into collective understanding um I mean the referees are proxies but you know by but you know we read each other's papers and this is how the field progresses um now um but it's slow and we are now facing um a crisis that the there's so many AI generated papers that even if
我们现在的做法是:有期刊,当一篇论文被验证过、够格投给期刊时,我们把它送给人类审稿人,他们自愿贡献自己的时间,作为一种服务,来审这些论文——不只是因为结果有趣,还因为他们想鼓励作者成长为更好的数学家。这是在培养下一代数学家的一种方式。这不是什么光鲜的工作,我是说,我们刚刚为解决问题颁了四枚奖章,我们并没有为审稿颁四枚奖章,但它是我们这个职业不可或缺的一部分,它是我们把数学家的个人成就转化为集体理解的途径。审稿人只是一个代理,但我们会读彼此的论文,这个领域就是这样往前走的。可是它很慢,而我们现在面临一场危机:AI 生成的论文太多了,即便它们都是正确的,数量也可能多到审稿人根本读不完。这是
便签笔记
37:43
they're correct there may be too many for referees to read um now uh this is an ongoing problem I think there will be other panels discussing what to do about this um I think journals will also have to start adopting AI tools as filters um um that you know certain we'll have to be much more precise about uh style guides and uh uh rules for for various journals so that um before any paper reaches a human referee there's some additional layer of evaluation this will be controversial but I think uh it will be necessary but you would we should definitely not take the human referee out of the equation um I would I don't think a journal that purely operates on AI referees will be uh successful uh especially since these tools can be gamed quite a bit and you know and and if there's a journal that says oh this paper has been has been accepted by 100 AI referees now that that's not community acceptance. That still doesn't mean that we want to read it.
一个正在发生的问题,我想后面还会有别的圆桌讨论该怎么办。我认为期刊也将不得不开始采用 AI 工具作为筛选器,我们必须把各家期刊的格式规范和各种规则定得更精确,这样在任何论文送到人类审稿人手上之前,先经过一层额外的评估。这会有争议,但我认为这是必要的。不过我们绝对不该把人类审稿人从这个环节中剔除。我不认为一个纯靠AI 审稿运作的期刊会成功,尤其因为这些工具相当容易被钻空子。而且如果有个期刊说"这篇论文已经被 100 个 AI 审稿人接受了",那并不是共同体的接受,那依然不意味着我们就想去读它。
便签笔记
10正典化:AI 无能为力的终点
38:41
Okay. So, is that our final goal? Well, even that is not our um our final state really. Um you know, so even when a result has been accepted by the community and is published and is in a prestigious journal and people all accept as correct. Um it's still not the final state. Um it's the final states are things like textbooks, you know, the the material we we teach in classes. What is the standard definition of this concept? What is the correct order in which we prove things? How do you organize individual results published in good papers into a coherent theory? Um that is the state which we mathematics ultimately and that that is the end state of all these problems that we're trying to solve. Um so Alex who introduced me has a very nice name for this is it's it's canonicalization.
好,那这是我们的最终目标吗?其实连这也不是真正的终点状态。就算一个结果已经被共同体接受、已经发表在有声望的期刊上、大家都认可它是正确的,它仍然不是最终状态。最终状态是像教科书那样的东西,是我们在课堂上教的材料。这个概念的标准定义是什么?我们证明这些东西的正确顺序是什么?你怎么把发表在好期刊上的一个个孤立结果,组织成一套连贯的理论?那才是数学最终要达到的状态,才是我们试图解决的所有这些问题的终点。介绍我出场的 Alex 给这件事起了个很好的名字,叫做"正典化"(canonicalization)。
便签笔记
39:30
So um in addition to digesting a proof and and and publishing it, you want to make it canonical. Um and this is the slowest stage of all. Um you know there's many many results in the last 10 years 20 years that you know they're in prestigious journals but they're not in textbooks yet. They're not yet taught to students. We haven't yet had the figured out the really definitive canonical correct way to teach these things. Um and this is slow. I mean you have to teach classes. You have to get feedback and and you have to really think hard about what is the correct way to or to to to um to organize lots and lots of of of of um different results and put it in one coherent narrative. Um and it needs consensus you know I mean if it would be you know if half the mathematicians in a field think you should do things this way and the other half a different way it we don't have canonicalization. Um and this is the stage in which AI is basically completely useless. Um so um like there's there's sort of no I see
所以,除了消化一个证明、把它发表出来,你还要让它成为正典。而这是所有阶段中最慢的一个。过去十年、二十年里有非常多的结果,它们躺在有声望的期刊里,但还没进教科书,还没教给学生。我们还没有想清楚教这些东西真正确定的、正典的、正确的方式。这个过程很慢——你得去开课,得拿到反馈,还得非常认真地去想:把大量大量不同的结果组织起来、放进一个连贯叙事里的正确方式到底是什么。而且这需要共识。如果一个领域里有一半数学家认为该这么做,另一半认为该那么做,那我们就没有正典化。而这个阶段,正是 AI 基本上完全帮不上忙的地方。我几乎看不到 AI 在这一环有什么角色可扮演,但它却是最有
便签笔记
40:30
basically almost no role for AI in this part of the process but it is the most valuable part um that if you want to apply any sub field of mathematics to some other area of mathematics or some problem it pretty much has to be digested in this form you know like if you want field of math to become useful to engineers or or or physicists or or or biologists or whatever I mean they they're not going to dig through you know the most recent papers in in in the annals of mathematics or whatever they they want the textbooks um and many yeah so the the most valuable applications of of of math only get unlocked once you have reached this final stage including AI itself I mean a big reason why AI is so successful at mathematics is because for centuries we've been building these canonical definitions we we have these textbooks um of you know how does linear algebra work how does group theory work you know like we have all these very very mature theories And AI has absorbed all of these. Um, and that's what what
价值的一环。因为你要把数学的任何一个子领域应用到数学的其他领域或某个问题上,它基本上必须先被消化成这种形态。比如你想让某个数学分支对工程师、物理学家或生物学家之类的人有用,他们是不会去啃《数学年刊》上最新的论文的,他们要的是教科书。所以数学最有价值的那些应用,只有到达这个最终阶段之后才会被解锁,包括AI 本身。AI 之所以在数学上这么成功,一个很大的原因就是几个世纪以来我们一直在构建这些正典定义,我们有这些教科书,讲线性代数怎么运作、群论怎么运作,我们有这么多非常非常成熟的理论,而 AI 把这些全都吸收了。它靠的就是这些
便签笔记
41:28
it uses to to um to uh uh to for success. Um, and so if we cut off this this part of the process long term, uh, it it will hurt mathematics. Um, yeah. So, um, sorry said all that. So maybe this is the final um goal that so so maybe the real goal of problem solving is not just to generate proofs and not just to verify them, not just to explain them and not just to publish them but to actually um organize all those proofs into uh a a canonical state to make them definitive. Um so this last three stages I like to call proof digestion, you know. So it's not just sort of creating food and and not even just of eating it but you have to sort of uh digest it with internalize it to the point where it really becomes part of the body of mathematics.
才取得成功。所以从长期看,如果我们把流程中的这一环切断,会伤害到数学。嗯,抱歉,说了这么多。所以也许这才是最终目标——解决问题的真正目标不只是生成证明,不只是验证证明,不只是解释证明,也不只是发表证明,而是真正把所有这些证明组织成一种正典状态,让它们成为定论。所以最后这三个阶段,我喜欢称之为"证明的消化"。不只是做出食物,甚至不只是吃下去,你还得把它消化、内化,直到它真正成为数学这个躯体的一部分。
便签笔记
11证明消化不良与过剩时代
42:19
Okay. Now um it is possible that it goes this this this uh process even continues more but these are the five stages that I was able to identify. Um but you can see that the exercise of stating the your goals clearly um can take you quite far from what you might think the uh the the process is. And this um this little deconstruction reveals one thing which is that um what AI is doing is that it is creating lots and lots of well okay if this was the process of proof digestion what AI is producing is what I call proof indigestion. Um if you're electrical engineer you might call it impedance mismatching.
好。也有可能这个过程还能继续往下延伸,但这是我能识别出来的五个阶段。你可以看到,把自己的目标清楚地说出来,这个练习能带你走到离你原本以为的流程相当远的地方。而这个小小的拆解揭示了一件事:AI 正在制造大量的——好,如果说这是证明消化的过程,那 AI 生产出来的就是我所说的"证明消化不良"。如果你是电气工程师,你可能会叫它阻抗失配。
便签笔记
42:57
um that because AI is much better at the first two stages of this process than it is at the last three um we are you know beginning to see all kinds of indigestion. We are seeing proofs accumulate that haven't been verified. We're seeing verified proofs that no one will read. Um we are seeing proofs that are readable but no one will will publish. And um this isn't happening yet but uh if we extrapolate we will soon also get a big pileup of published proofs AI generated that no one knows how to convert into into like proper textbooks and and educate um the next generation. Um so these are all problems of of of abundance you know so you can roughly speaking all the problems in the world can be classified into problems of scarcity and problems of abundance.
因为 AI 在这个流程的前两个阶段远比后三个阶段擅长,我们已经开始看到各种各样的消化不良。我们看到未经验证的证明不断堆积,看到已经验证但没人会去读的证明,看到可读但没人愿意发表的证明。还有一种情况现在还没出现,但按趋势外推,我们很快也会积压一大堆已发表的、由 AI 生成的证明,没人知道该怎么把它们转化成像样的教科书、去教育下一代。这些都是"过剩"带来的问题。粗略地说,世界上所有的问题都可以分成稀缺的问题和过剩的问题。
便签笔记
43:43
Okay. So for food for example, you know, you have famine and malnutrition. These are problems of food scarcity, but you also have have obesity and lack of exercise and um and bad diet. You know, these these are problems of of food abundance. Um and so we've lived for centuries in an area of proof scarcity. Uh but we are going to if we believe this working hypothesis, we will very very soon be an era of proof abundance. Uh and we have to adapt. Um and in the case of food, we adapted by becoming much more conscious about what is good eating and what is not. What is good exercise, what is not. Our our taste improved, our our cuisines improved. We rejected some food that you know is technically edible, but we don't eat everything that's on our plate. Um and we have to do something similar in mathematics, I think.
比如食物,有饥荒和营养不良,这是食物稀缺的问题;但你也有肥胖、缺乏运动、饮食结构糟糕,这些是食物过剩的问题。几个世纪以来,我们一直生活在证明稀缺的时代,但如果我们相信这个工作假设,我们很快就会进入证明过剩的时代。我们必须适应。在食物这件事上,我们的适应方式是变得更有意识地去分辨什么是好的饮食、什么不是,什么是好的运动、什么不是。我们的品味提高了,烹饪也进步了。我们会拒绝一些技术上可食用的食物,我们不会把盘子里的东西全吃光。我想在数学里,我们也得做类似的事情。
便签笔记
12莱顿宣言与共同体应对建议
44:31
Um so once you have a clearer idea of what your goals are, then you can actually start making policies. then you can start actually attacking this community response question. Um so um I think um a first start of this many of you have heard about this. Uh so there was a grassroots effort starting in Leiden back last year. Many mathematicians in a workshop on AI and mathematics realized this and started um trying to make a consensus statement drawing in many members of the community including many people here to draft what's known as the lighten declaration.
所以,一旦你对自己的目标有了更清晰的认识,你就可以真正开始制定政策,就可以真正着手处理"共同体如何回应"这个问题。我想这方面的一个开端,你们很多人都听说过:去年在莱顿(Leiden)有一场自下而上的行动。很多数学家在一个关于 AI 与数学的研讨会上意识到了这些问题,开始尝试形成一份共识声明,吸纳了共同体的许多成员,包括在座的很多人,共同起草了后来被称为"莱顿宣言"的文件。
便签笔记
45:03
Um and this is this doesn't solve you know these these problems but this is the type of thing we need to do continue to do more of um so um this declaration states lays out the situation what AI is doing what math is doing what math is about uh at least to the the extent that a large fraction of the community can can agree with it and there are many many signaries of this I've signed this this declaration myself um and there are practical recommendations that become clear once you understand the goals and uh and values of your community. So I really all recommend if you haven't seen it to actually read this declaration. It is um it has been it's been gone through many iterations um and many people have supplied feedback and it's quite a good document at this point. Um and it has many many um recommendations. Um so one for example for individuals is to always disclose AI tool use. Um maybe in the future this will become unnecessary. You know we don't disclose latte use anymore for instance. Um but in our current era
这份宣言并不能解决这些问题,但这正是我们需要继续多做的那类事情。这份宣言陈述并梳理了现状:AI 在做什么、数学在做什么、数学究竟是关于什么的——至少是在共同体中很大一部分人能够认同的范围内。而且还有很多
便签笔记
46:02
it is it's really important I think um particularly I think the the the one scenario that we really want to avoid in this current era is where people are everyone is secretly using AI and no one is disclosing it and no one is sharing their best practices and their mistakes and this is going to be really harmful uh for um working out how to use AI properly. So we really need to normalize responsible disclosure of AI systems. So in that vein I have used AI to generate these text not for the M dashes but um there's a couple places where autocomplete was useful uh and the diagrams that you saw I generated for instance okay but the other text was was written by myself um we need to make life a lot easier for referees uh like the we cannot just generate lots and lots of 100page proofs by AI just dump them on referees um we need to make it much easier we the burden will shift on authors to to it will be it will be incumbent on authors to to make their their papers at the highest expositional standard um
便签笔记
47:03
disclosing all their tools using formalization as needed. Um and we need to in general just move away from the first stage of the process. Okay. So in the past we would emphasize what was really important was being the first to prove a theorem proof generation. This was the this main task and everything else was clean up. Um now we've automated that first step and that's no longer the bottleneck. Um so um a proof that has been generated but not cleaned up and and made publishable that is we should we should value that only as a incomplete proof uh not yet ready for publication and we really need to start valuing proof digestion. So the three final stages exposition publication and and cononicalization we um we do value these right now but we should be much more explicit about um about um raising their importance.
便签笔记
47:51
Um and we need to uh attribute things properly. You know there's a debate you know if you push pushed a button and you got a proof do you get to be an author of the paper that came out. Um so um the lighten declaration has some something to say about this. Um my suggestion is that if an author cannot give a talk like up here about the result and they cannot take questions about the result that they generated um they should not be an author or at least um the paper is not ready to be published unless there is at least one person who can actually talk about the result properly.
便签笔记
48:26
Um yeah so so I've deconstructed one you know so I had this diagram of of all the goals of mathematics and it branched out all over the place and I got I had one direction which was solving problems and my point was that it's it is far more complicated than sort of the naive uh um picture of just trying to solve as many problems as possible. Okay, but you can deconstruct it. You can understand what we are that what it is we're actually doing and that really clarifies that aspect. Okay. So, you know, and this informs like how journal should operate and and how we should um assess um solving problems for future hiring and so forth. But there's many other parts of uh many other goals that we care about um and this working hypothesis will will eventually impact all of them as well. So we need to also do so what I just did for problem solving people also need to do for teaching uh for mentoring hiring grant applications outreach every single aspect of our profession should be on the table and discussed and we we
便签笔记
49:24
need to analyze them all uh and that's a big lengthy discussion and I only have 10 minutes so I'm not going to talk about any of those um yeah so um and the response will be different in in in for different parts of our profession I think there'll be some parts of our profession where we basically have to restrict AI I um education and training in particular. I think it is uh it it can be harmful to to to give students too much AI use early on. They do need to develop their own um innate mathematical skills before they can learn to uh uh to use these tools. You know, just like we we we teach people arithmetic before they do calculators or we we teach people to walk and run before they we we give them the keys to a car. Um same in math. Um there are other places where we should um use AI but we should take the initiative and and decide what are we set the rules on on what types of AI use are acceptable which ones are not and not let external actors define uh um the rules for us. Uh there'll be some places where
便签笔记
50:25
traditional um venues like journals would are just not suitable for u this modern a workflow. We do need to create new workflows, new infrastructures. Um that's another hour talk which I unfortunately do not have time to to give here. But most importantly we really need to discuss um all these questions um openly and and and and not just be passive uh um you know recipients of of developments but we we need to discuss AI capability. We need to discuss our goals and values and we need to discuss community response and this also is a major recommendation of the lighten declaration. So this is where I'll stop. Thank you very much.
便签笔记
51:06
[applause]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

陶哲轩在 ICM 2026 公开讲座中主张:AI 带来的不是数学论证的危机,而是数学价值观与实践的危机——数学界必须像百年前应对"基础危机"那样,明确说出自己的目标,尤其要把"解题"这一目标拆解为生成、验证、阐释、发表、经典化五个阶段,并认识到 AI 只擅长前两阶段,从而应对即将到来的"证明过剩"时代。

核心要点

  • 历史类比:百年前的基础危机是今日的先例。 20 世纪初,罗素悖论与哥德尔不完备定理迫使数学界正式公理化逻辑与集合论(一阶逻辑 + ZFC),过程动荡但结果让学科更健康。陶认为一个世纪后,数学正进入类似的动荡期,只是这次危机不在论证本身,而在于此前"朴素"的、未明言的价值观与实践。
  • 有意跳开"AI 能力猜想"的争论。 他把"AI 能否做研究级数学"表述为一族充满占位符("某些工具、某些花费、某些领域、某种成功率")的猜想,指出三年来的公开辩论几乎都在争这些占位符的取值,而且现有数据多为公司选择性披露、缺乏成本与失败率信息。他明确表示本次讲座不讨论真伪,而是条件分析:假设"合理强"的版本为真,数学界该怎么办。
  • 目前最科学的能力数据点:First Proof Challenge。 这是数学家自发组织、无 AI 公司赞助的独立评测:捐出 10 道未发表的研究级问题,让完全自主的 AI 系统作答并要求达到可发表的写作质量。2026 年 5 月一轮中 4 个系统参赛(陶参与其中一支队伍),单个系统最好成绩约 5/10,合计 7/10 达到发表级质量,但每道题成本从 10 美元到约 1000 美元不等。
  • 古德哈特定律解释了为何"隐性目标"必须显性化。 数学有多重目标(解决问题、建立理论、理解世界、建设社群、培养下一代、积累知识库、审美价值)。过去这些目标正相关,只需明说一个当作其余的代理即可;但一旦某个目标被极度优化,继续推进它会损害其他目标。人类优化能力有限,而 AI 与 AI 公司极擅长优化,因此格外容易触发古德哈特效应。
  • "解题"远非"最大化解出的问题数",而是一条五阶段流水线。 ①证明生成 → ②证明验证(Lean 等证明助手大幅进步)→ ③证明阐释(写得让社群能读懂)→ ④社群接受与发表 → ⑤经典化(进入教科书、形成标准定义与叙事)。AI 在①②已显著加速,③较弱,④几乎无影响,⑤"基本完全无用"——而⑤恰是最有价值的阶段,也是工程师、物理学家乃至 AI 自身真正消费数学的形态。
  • 已出现"验证了但无人理解"的证明。 以 Erdős 问题集(约 1200 题)为例,近半年涌现大量 AI 生成、部分已 Lean 形式化的证明,提交者称"不够资格评价",验证者称"检查过但不能担保"。陶警告我们非常接近"重大结果被证明并验证,却没有任何人能解释它"的局面。
  • 好的阐释需要"自然摩擦",AI 恰恰抹平了它。 AI 文本拼写格式完美,但重点错位——对琐碎引理写三页、对关键步骤写三行,且不交代思想来源,与文献脱节。人类作者在难点处放慢、在易处快速掠过,读者能由此感知难点所在。他以自己读 Bourgain 1991 年论文(Hong Wang 昨日报告的工作即基于此)的经历为例:笔记上写着"我恨 Bourgain",但正是这种挣扎让他学会了 Bourgain 的思维方式。
  • 从"证明稀缺"到"证明过剩",审稿成为最稀缺资源。 类比食物:稀缺时代的问题是饥荒,过剩时代的问题是肥胖与坏饮食,人类靠发展"品味"与主动拒绝来适应。数学也须如此:期刊需要 AI 过滤层与更精确的风格规范,但绝不能取消人类审稿——"被 100 个 AI 审稿人接受"不等于社群接受,且这类工具易被博弈。
  • Leiden 宣言与具体建议。 陶已签署去年源于 Leiden 研讨会的社群共识文件。建议包括:始终披露 AI 使用(避免"人人偷偷用、无人分享最佳实践"的局面;他本人披露本讲座用 AI 做了少量自动补全和图表);减轻审稿负担,把写作与形式化的责任转移给作者;未整理的 AI 证明只算"不完整证明";署名标准是——若作者不能上台讲解并回答关于该结果的提问,就不应署名或论文尚不宜发表。

结论与值得注意的细节

  • 核心结论:真正的瓶颈已从"率先证明"转移到"证明消化"(阐释、发表、经典化),数学界应明确提升后三阶段的价值;AI 造成的是"证明消化不良"(他戏称电气工程师会叫"阻抗失配")。
  • 他强调 AI 之所以擅长数学,正是因为数学界几个世纪以来建成了线性代数、群论等经典化教科书体系;若切断经典化阶段,长期将反噬包括 AI 在内的一切应用。
  • 对不同领域回应应有差异:教育与训练阶段应限制 AI,学生须先培养内在能力(如先学算术再用计算器、先学走路再开车);其他领域则应由数学界自主制定规则,而非任由外部行动者定义。
  • 他把解题只当作众多目标之一,教学、指导、招聘、经费申请、公众推广等每个方面都需要做同样的拆解分析,并需要新的工作流与基础设施来取代不适应 AI 时代的传统期刊模式——限于时间未展开。
  • 幽默细节:幻灯片中的破折号是"人类生成的",以回应网上用破折号识别 AI 文本的风气;他提到刚颁出四枚菲尔兹奖是为解题,"没有为审稿颁奖",但审稿是把个人成就转化为集体理解的关键环节。
核心句型 · 10
1. It seems unprecedented, but actually I think there are historical precedents for …
“It seems unprecedented. but actually I think there are historical precedents for what we're going through.”
先承认表面印象,再用 but actually 翻转,引入历史类比。适合议论开头,仿写:It seems X, but actually there are precedents for Y.
2. It's not a crisis in A, but a crisis in B.
“It's not a crisis in our mathematical arguments but it's a crisis in our mathematical values and practices”
否定—肯定的对比结构,用于精确界定问题的性质。仿写时 A、B 需是同一范畴内的两个不同层面。
3. Regardless of whether X is true or false or partly true, I think it is overdue that we …
“Regardless of whether this hypothesis is true or false or partly true or whatever I think it is actually overdue that we do need to talk about why we do mathematics”
先撇开争议前提,再提出无论如何都成立的主张。overdue 表示「早该做」,语气比 necessary 更强。
4. As long as A, you can do X and you'd also be doing Y as well.
“As long as you're far away from all these goals, as long as they're kind of roughly correlated with each other, you can move towards one goal and you'd also be moving towards other goals as well.”
以条件句解释某种机制为何过去有效。as long as 可重复叠加多个条件;后句用 would 表示一般推论。
5. This works until you start …, and at some point you become so good at A that any further progress actually moves you away from B.
“This works until you start optimizing a lot and you become very good at say goal one. and at some point you become so good at goal one that any further progress towards goal one actually moves you away from goal two”
描述「过度优化的临界点」的经典结构:works until + so…that。适合讲任何指标失效的场景。
6. It is not enough to A, and it is not enough to B. C needs to …
“It is not enough to generate proofs and it's not enough to verify the proofs. The proofs need to be explained well enough that they can be communicated”
递进式否定,逐层抬高标准,最后给出真正的要求。not enough to 可连续排比,形成层次感。
7. X is only a sufficient condition. It's not necessary.
“Making it correct and easy to read helps, but that's only a sufficient condition. It's not necessary.”
借用数理逻辑术语表达「有帮助但不保证」。注意此处陶的口误,逻辑上他想说的是「必要而非充分」;仿写时确认方向。
8. All the problems in the world can be classified into problems of X and problems of Y.
“Roughly speaking all the problems in the world can be classified into problems of scarcity and problems of abundance.”
用二分法建立分析框架,之后用具体类比(食物)填充。roughly speaking 缓和绝对化语气。
9. It will be incumbent on X to …
“It will be incumbent on authors to make their papers at the highest expositional standard”
正式语体表达「责任落在某方」。比 X should 更书面、更强调义务转移,适合政策建议。
10. Just like we teach people A before B, … same in C.
“Just like we teach people arithmetic before they do calculators or we teach people to walk and run before they we give them the keys to a car. same in math.”
用两个并列日常类比再以 same in 收束,是口语中强调原则普适性的高效方式。
生词精讲 · 118 · 按出现顺序
accolades /ˈækəleɪdz/ n. 0:00
荣誉、赞誉(常用复数)
foyer /ˈfɔɪər/ n. 0:00
门厅、休息厅
existential crises phr. 0:53
生存危机;关乎存亡的危机
accustomed to phr. 0:53
习惯于
unprecedented /ʌnˈpresɪdentɪd/ adj. 1:42
前所未有的
precedents /ˈpresɪdənts/ n. 1:42
先例
prologue /ˈproʊlɔːɡ/ n. 1:42
开场白、序言
semiformally /ˌsemiˈfɔːrməli/ adv. 1:42
半形式化地
inconsistent /ˌɪnkənˈsɪstənt/ adj. 2:21
(逻辑)不相容的、矛盾的
axioms /ˈæksiəmz/ n. 2:21
公理
traumatic /trəˈmætɪk/ adj. 2:21
创伤性的、令人痛苦的
turbulent /ˈtɜːrbjələnt/ adj. 2:21
动荡的、混乱的
orthodox /ˈɔːrθədɑːks/ adj. 3:22
正统的、公认的
consensus /kənˈsensəs/ n. 3:22
共识
proof assistant n. 3:22
证明助手(形式化验证软件)
strenuously /ˈstrenjuəsli/ adv. 3:22
费力地、严格地
resilient /rɪˈzɪliənt/ adj. 4:19
有韧性的、能迅速恢复的
incorporate /ɪnˈkɔːrpəreɪt/ v. 4:19
纳入、吸收
advent /ˈædvent/ n. 5:19
(重要事物的)到来、出现
lay down some rules phr. 5:19
制定规则
conjecture /kənˈdʒektʃər/ n. 6:08
猜想(数学中未证明的命题)
tackle /ˈtækl/ v. 6:08
着手处理(难题)
placeholders /ˈpleɪsˌhoʊldərz/ n. 6:50
占位符
non-trivial /ˌnɑːnˈtrɪviəl/ adj. 6:50
非平凡的、不可忽视的
devolved /dɪˈvɑːlvd/ v. 7:47
退化、蜕变(devolve into)
distracting from phr. 7:47
分散对……的注意力
novelty /ˈnɑːvəlti/ n. 8:32
新奇玩意儿
business as usual phr. 8:32
一切照常
ostensibly /ɑːˈstensəbli/ adv. 8:32
表面上、名义上
common rooms n. 9:26
(学院、系里的)公共休息室
counter claims n. 10:34
反驳主张
selectively disclosed phr. 10:34
选择性披露的
incentives /ɪnˈsentɪvz/ n. 11:05
激励、动机
in as favorable a light as possible phr. 11:05
以尽可能有利的方式呈现
conflate /kənˈfleɪt/ v. 11:05
混为一谈
grassroots /ˈɡræsruːts/ adj. 11:41
草根的、自下而上的
harnesses /ˈhɑːrnɪsɪz/ n. 11:41
(AI)运行框架、封装系统
autonomous /ɔːˈtɑːnəməs/ adj. 11:41
自主的、无人干预的
publication quality n. 12:45
可发表的质量
complement /ˈkɑːmplɪmənt/ n. 13:37
补集;补充
working hypothesis n. 13:37
工作假设
conditional analysis n. 13:37
条件分析
condition on phr. 14:37
以……为条件
explicit /ɪkˈsplɪsɪt/ adj. 14:37
明确的、显性的
implicit /ɪmˈplɪsɪt/ adj. 15:35
隐含的、隐性的
the humanities n. 15:35
人文学科
have the luxury of phr. 15:35
有……的余裕/奢侈
overdue /ˌoʊvərˈduː/ adj. 16:38
早该做的、逾期的
compiled /kəmˈpaɪld/ v. 16:38
汇编、整理
cumulative /ˈkjuːmjəleɪtɪv/ adj. 17:34
累积的
downplay /ˌdaʊnˈpleɪ/ v. 17:34
轻描淡写、低估
aligned /əˈlaɪnd/ adj. 18:26
一致的、对齐的
proxy /ˈprɑːksi/ n. 18:26
代理、替代指标
ceases to be phr. 19:20
不再是
at the expense of phr. 20:03
以……为代价
untethered /ʌnˈteðərd/ adj. 20:03
不受束缚的、脱离的
rubric /ˈruːbrɪk/ n. 20:44
评分标准、评价准则
in so far as phr. 21:32
就……而言
deconstruct /ˌdiːkənˈstrʌkt/ v. 21:32
拆解、解构
flow network n. 23:19
流网络(图论概念)
sync /sɪŋk/ n. 23:19
此处为 sink,汇点
complimentary /ˌkɑːmplɪˈmentri/ adj. 24:18
此处意为 complementary,互补的
vouch /vaʊtʃ/ v. 25:18
担保(vouch for)
backwalk n. 25:18
此处意为 backlog,积压
unwelcome /ʌnˈwelkəm/ adj. 26:14
不受欢迎的
artifacts /ˈɑːrtɪfækts/ n. 26:14
人工制品;此处指孤立的产出物
exposition /ˌekspəˈzɪʃn/ n. 27:05
阐述、(数学)写作表达
lema /ˈlemə/ n. 27:05
即 lemma,引理
scoring function n. 28:03
评分函数
slick /slɪk/ adj. 29:02
过于流畅的、油滑的(含贬义)
sanded down phr. 29:02
打磨掉、磨平
natural friction n. 29:02
自然摩擦(陶自创术语)
sail through phr. 29:02
轻松通过
blast through phr. 29:58
一冲而过、快速带过
paradoxically /ˌpærəˈdɑːksɪkli/ adv. 29:58
自相矛盾地、反常地
annotation /ˌænəˈteɪʃn/ n. 31:02
批注
fought my way through phr. 31:02
艰难地啃完
digesting /daɪˈdʒestɪŋ/ v. 32:02
消化(知识)
production quer n. 32:02
即 production quota,生产指标
sufficient condition n. 32:39
充分条件
scarce /skers/ adj. 33:41
稀缺的
flooded with phr. 33:41
被……淹没
bottlenecks /ˈbɑːtlneks/ n. 33:41
瓶颈
induce /ɪnˈduːs/ v. 33:41
诱导、促使
opaque /oʊˈpeɪk/ adj. 34:46
不透明的
look under the hood phr. 34:46
查看内部运作机制
corporate secrets n. 34:46
商业机密
reputable /ˈrepjətəbl/ adj. 35:35
有声望的
appetizing /ˈæpɪtaɪzɪŋ/ adj. 35:35
诱人的、开胃的
referee /ˌrefəˈriː/ n./v. 36:42
审稿人;审稿
prestigious /preˈstɪdʒəs/ adj. 36:42
有威望的
gamed /ɡeɪmd/ v. 37:43
被钻空子、被操纵
take … out of the equation phr. 37:43
把……排除在考虑之外
coherent /koʊˈhɪrənt/ adj. 38:41
连贯的、条理清晰的
canonicalization /kəˌnɑːnɪkələˈzeɪʃn/ n. 38:41
正典化、标准化
definitive /dɪˈfɪnətɪv/ adj. 39:30
权威的、决定性的
dig through phr. 40:30
翻找、深挖
unlocked /ʌnˈlɑːkt/ v. 40:30
解锁、释放
internalize /ɪnˈtɜːrnəlaɪz/ v. 41:28
内化
indigestion /ˌɪndɪˈdʒestʃən/ n. 42:19
消化不良
impedance mismatching n. 42:19
阻抗失配(电路术语)
extrapolate /ɪkˈstræpəleɪt/ v. 42:57
外推、推断
pileup /ˈpaɪlʌp/ n. 42:57
堆积、积压
abundance /əˈbʌndəns/ n. 42:57
丰裕、过剩
famine /ˈfæmɪn/ n. 43:43
饥荒
malnutrition /ˌmælnuːˈtrɪʃn/ n. 43:43
营养不良
edible /ˈedəbl/ adj. 43:43
可食用的
signaries /ˈsɪɡnətɔːriz/ n. 45:03
即 signatories,签署者
iterations /ˌɪtəˈreɪʃnz/ n. 45:03
迭代、修订版本
normalize /ˈnɔːrməlaɪz/ v. 46:02
使成为常态
in that vein phr. 46:02
本着这种精神、循此思路
incumbent on phr. 46:02
是……的责任
formalization /ˌfɔːrməlaɪˈzeɪʃn/ n. 47:03
形式化
attribute /əˈtrɪbjuːt/ v. 47:51
归功于、署名归属
on the table phr. 48:26
摆上桌面(供讨论)
innate /ɪˈneɪt/ adj. 49:24
内在的、天生的
take the initiative phr. 49:24
主动采取行动
venues /ˈvenjuːz/ n. 50:25
(发表的)场所、渠道
理解自测 · 11 题 · 是真懂了,还是以为自己懂
1. 陶哲轩用哪一段历史来类比当前 AI 对数学的冲击?这段历史的起因与结果分别是什么?

他用一百年前的「数学基础危机」作类比。起因是 20 世纪之前数学家以半形式化、朴素的方式使用集合、数、函数、极限等概念,罗素指出朴素集合论不相容(可构造既包含又不包含自身的集合),哥德尔又证明了不完备定理,迫使数学家重新思考「什么是证明」。过程动荡且痛苦,但结果是获得了一阶逻辑加 ZFC 这样的标准框架,被检验了一个世纪。陶在开场章节据此论证:当下的危机同样会带来更健康、更有韧性的学科。

2. FirstProof 挑战赛的评测方式和 2026 年 5 月一轮的结果是什么?

FirstProof 是数学家发起的草根独立评测,不受 AI 公司赞助。数学家捐出 10 道未发表的研究级问题,用于测试完全自主的 AI 系统,且不只看是否解出,还看写作是否达到可发表质量。5 月一轮有四套 harness 提交,陶参与了其中一个团队。单个系统最好成绩是 10 题中 5 题,四套合计有 7 题达到可发表水平,每题成本从 10 美元到 1000 美元不等。陶称这是目前「最科学的数据点」,但强调后续测试可能改变这些数字。

3. 陶把「解题」拆解成了哪五个阶段?其中哪些阶段被 AI 加速、哪些没有?

五个阶段是:证明生成、证明验证、证明解释(写作与理解)、共同体接受(发表与审稿)、正典化(进入教科书与课程)。前两个阶段已明显加速:生成靠 AI,验证主要靠 Lean 等证明助手语言的进步。第三阶段 AI 仍很弱,只擅长拼写格式而重点分配失衡;第四阶段依赖人类反应,AI「基本没有影响」;第五阶段需要教学、反馈和共识,AI「基本完全无用」。陶把后三阶段统称为「证明消化」。

4. 陶在演讲中如何披露自己对 AI 的使用?他为什么要这样做?

他说明本次演讲的文字由自己撰写,但有几处用了自动补全,图表则由 AI 生成,并调侃破折号是人类打的。这样做是为了践行莱顿宣言的第一条个人建议「始终披露 AI 工具使用」。他解释理由:当前最应避免的场景是人人暗中用 AI、无人披露、无人分享最佳实践与错误,这会阻碍整个共同体学会正确使用 AI。因此需要把负责任的披露变成常态,他以身作则。这一段位于结尾的共同体应对建议部分。

5. 陶为何认为「过去不必明说数学的多重目标」,而现在必须明说?请复述其论证链。

陶的论证分三步。第一,数学有多重目标(解题、建理论、理解世界、教育、审美等),但过去数学太难,离所有目标都很远,目标之间大致正相关,因此朝一个目标前进就等于朝所有目标前进,一个显性目标可以作为其他目标的代理。第二,古德哈特定律指出,当某个指标被过度优化,进一步的进展会以牺牲其他目标为代价。第三,人类不擅长优化所以问题不大,但 AI 和 AI 公司是极强的优化器,会使目标从正相关变成竞争关系。因此现在必须同时明确多个目标并给出清晰的评分标准。

6. 为什么陶说「过于流畅的证明」反而可能有害?他用什么亲身经历支持这一观点?

陶提出「自然摩擦」概念:人类写证明时在易处快速带过、在难处放慢并仔细组织,读者能借此感知难点所在;AI 生成的证明对易难之处一视同仁,读者失去这一信号,而过于光滑的证明会把困难「打磨掉」,读者回家自己重构时却做不出来。他的例子是研究生时期读 Jean Bourgain 1991 年关于挂谷猜想与限制猜想的论文,痛苦到批注「我讨厌 Jean Bourgain」,但在 Eli Stein 和 Tom Wolff 帮助下攻克后,理解了他的思维方式,几年后反而最偏爱读他的论文。他推断若该论文被 AI 层层润色,自己可能得不到那样的教育。

7. 陶为什么说 AI 在数学上的成功本身依赖于「正典化」?这一论点有什么反直觉之处?

陶指出,AI 之所以擅长数学,很大原因是几个世纪以来人类已把线性代数、群论等理论整理成成熟的正典定义和教科书,AI 吸收了这些材料。正典化是解题流程的终点,也是数学对工程、物理、生物等领域产生应用的必经形态——外行不会去啃顶刊论文,他们需要教科书。反直觉之处在于:AI 恰恰在这一最有价值的阶段「基本完全无用」,却又以这一阶段的产物为食。因此如果 AI 生成的证明堆积而无人正典化,长期会同时损害数学和 AI 自身。这一论点出现在「正典化」章节。

8. 「证明消化不良」这一概念是如何从流程分析中推导出来的?它有哪些具体表现?

陶先把解题拆成五个阶段,然后观察到 AI 在前两阶段(生成、验证)远强于后三阶段(解释、接受、正典化),各环节速率不匹配,他借用电路术语称之为「阻抗失配」,通俗说法就是「证明消化不良」。具体表现有四种:未验证的证明不断堆积;已验证但无人愿读的证明(如 Lean 形式化的埃尔德什问题解答);可读但无人愿意发表的证明;以及按趋势外推即将出现的、已发表却无人知道如何写进教科书的 AI 证明。他进一步把这归类为「过剩问题」,与过去几个世纪的「稀缺问题」对照。

9. 陶为什么坚持「不能把人类审稿人从环节中剔除」?如果有人反驳「AI 审稿更快更便宜」,他会如何回应?

陶的立场是期刊可以用 AI 作初筛、可以把规范定得更精确,但纯 AI 审稿的期刊不会成功。他会从三方面回应「更快更便宜」的反驳:第一,AI 审稿工具「相当容易被钻空子」,一旦成为目标就会遭遇古德哈特定律;第二,审稿的本质不只是判断对错,而是把个人成就转化为集体理解,并借此培养下一代数学家,这是人与人之间的过程;第三,「被 100 个 AI 审稿人接受」并不等于共同体接受,不意味着任何人真的想读它。速度和成本解决的是前两阶段的问题,而接受阶段的瓶颈本来就是人类注意力。

10. 陶提出的作者署名标准是什么?这一标准若应用于大型协作项目或纯形式化证明,还成立吗?

陶的标准是:如果一个「作者」不能像在讲台上那样报告结果并回答关于该结果的提问,就不该署名,或至少论文在有一个能真正谈论该结果的人之前不应发表。这一标准的核心是「有人能为结果负责并解释它」,而非「每个作者都懂全部」。放到大型协作项目上,只要至少有一人能整体解释,标准仍成立,且与陶对第三阶段「证明必须被理解」的要求一致。但对纯形式化证明(如 Lean 验证的埃尔德什问题解答),当前恰恰出现了「提交者说自己不够格评价、检查者也不能担保」的情形,按陶的标准这些结果尚不能算「可发表」,这正是他想用该标准来约束的场景。

11. 陶用「食物从稀缺到过剩」类比数学从「证明稀缺」到「证明过剩」。这个类比的解释力边界在哪里?

类比的强处在于:它把问题从「AI 好不好」转换为「我们如何建立品味」——人类应对食物过剩靠的是分辨好坏饮食、提升烹饪、拒绝技术上可食用但不想吃的东西,对应到数学就是建立评价证明价值的标准、只投入稀缺的审稿注意力给值得消化的结果。它也解释了为何「讲故事、展示过程」变得重要,因为要让论文更「开胃」。类比的边界在于:食物的好坏有相对客观的营养学标准,而陶自己承认「证明写得好不好」目前没有好的评分标准;此外食物消费是个体选择,而正典化需要整个领域的共识,个体品味无法直接汇总成集体正典。因此类比适合说明「态度转变」,但对「如何形成共识」这一最慢阶段解释力有限。

精读便签
下载便签 手机:长按图片保存
← 上一期 · NO.009"What Is a Strange Loop and What is it Like To Be One?" by Douglas Hofstadter (2013) 下一期 · NO.011 →AI and human evolution | Yuval Noah Harari
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com