视频库 / NO.021ASK THE BEST MINDS THE BIG QUESTIONS一人,一实验室
视频库 / NO.021
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。
第 21 期 · 回应 Ⅱ·04「预测就是理解吗?」

Terence Tao – How the world’s top mathematician uses AI

节目发布 2026-03-20 · Dwarkesh Patel
陶哲轩 DDwarkesh Patel
章节 · 点击跳转视频
0:00 开普勒:从柏拉图立体到三大定律 ▶ 正在看
3:55 开普勒是高温度 LLM 吗 ▶ 正在看
7:38 数据先行的科学新范式 ▶ 正在看
12:18 想法零成本,验证成瓶颈 ▶ 正在看
15:06 进步为何难以事前评分 ▶ 正在看
21:14 达尔文 vs 牛顿:说服力也是科学 ▶ 正在看
27:35 从微弱信号榨取信息 ▶ 正在看
30:12 Erdős 问题:AI 的矮墙与高墙 ▶ 正在看
34:19 AI 擅广度,人类擅深度 ▶ 正在看
40:46 1–2% 成功率与幸存者偏差 ▶ 正在看
46:12 陶的亲身体验:更宽但不更深 ▶ 正在看
52:48 Lean 证明、消融与策略语言 ▶ 正在看
63:20 素数随机模型与黎曼猜想 ▶ 正在看
69:35 狐狸式学习与机缘巧合 ▶ 正在看
76:50 AI 何时取代数学家 ▶ 正在看
本期讲者
陶哲轩UCLA 数学教授,2006 年菲尔兹奖得主,研究横跨调和分析、偏微分方程、组合数学与数论。近年积极使用 Lean 形式化证明与 AI 工具,并公开跟踪 AI 在 Erdős 问题上的进展。
Dwarkesh Patel科技播客《Dwarkesh Podcast》主持人,以对 AI 研究者、经济学家和历史学者的长篇深度访谈著称。
01开普勒:从柏拉图立体到三大定律
0:00
Today, I'm chatting with Terence Tao, who needs no introduction. Terence, I want to begin by having you retell the story of how Kepler discovered the laws of planetary motion because I think this will be a great jumping off point to talk about AI for math. I've always had an amateur interest in astronomy. I've loved stories of how the early astronomers worked out the nature of the universe. Kepler was building on the work of Copernicus, who was himself building on the work of Aristarchus. Copernicus very famously proposed the heliocentric model, that instead of the planets and the Sun going around the Earth, the Sun was at the center of the solar system and the other planets were going around the Sun. Copernicus proposed that the orbits of the planets were perfect circles. His theory fit the observations that the Greeks, the Arabs, and the Indians had worked out over centuries.
今天我要和陶哲轩聊一聊,他已经无需多做介绍了。陶哲轩,我想先请你重述一下开普勒是如何发现行星运动定律的故事,因为我觉得这会是一个很好的切入点,方便我们聊聊 AI 在数学中的应用。我一直对天文学有种业余的兴趣。我特别喜欢那些早期天文学家如何一步步弄清宇宙本质的故事。开普勒是在哥白尼的基础上工作的,而哥白尼本人又是在阿里斯塔克斯的基础上工作的。哥白尼最著名的就是提出了日心说:不是行星和太阳绕着地球转,而是太阳位于太阳系的中心,其他行星绕着太阳运行。哥白尼认为行星的轨道是完美的圆。他的理论与希腊人、阿拉伯人和印度人几个世纪以来积累的观测结果是吻合的。
便签笔记
0:57
Kepler learned about these theories in his studies, and he made this observation that the ratios of the size of the orbits that Copernicus predicted seemed to have some geometric meaning. He started proposing that if you take the orbit of the Earth and you enclose it in a cube, the outer sphere that encloses the cube almost perfectly matched the orbit of Mars, and so forth. There were six planets known at the time and five gaps between them, and there were five perfect Platonic solids: the cube, the tetrahedron, icosahedron, octahedron, and dodecahedron.
开普勒在学习中了解到这些理论,然后他注意到一个现象:哥白尼预测的那些轨道大小之间的比例,似乎带有某种几何意义。他开始提出,如果你取地球的轨道,把它外接一个立方体,那么包住这个立方体的外接球面,几乎完美地吻合火星的轨道,依此类推。当时已知有六颗行星,它们之间有五个间隔,而正好有五种完美的柏拉图立体:立方体、正四面体、正二十面体、正八面体和正十二面体。
便签笔记
1:35
So he had this theory, which he thought was absolutely beautiful, that you could inscribe these Platonic solids between the spheres of the planets. It seemed to fit, and it seemed to him that God's design of the planets was matching this mathematical perfection of the Platonic solids. He needed data to confirm this theory. At the time, there was only one really high-quality dataset in existence. Tycho Brahe, this very wealthy, eccentric Danish astronomer, had managed to convince the Danish government to fund this extremely expensive observatory.
于是他提出了这个他认为绝对优美的理论:你可以把这些柏拉图立体嵌在各行星的球面之间。这看起来很吻合,在他看来,上帝对行星的设计正好匹配了柏拉图立体这种数学上的完美。他需要数据来证实这个理论。当时世界上真正高质量的数据集只有一份。第谷·布拉赫,一位非常富有、性格古怪的丹麦天文学家,成功说服丹麦政府资助建造了这座极其昂贵的天文台。
便签笔记
2:12
In fact, it was an entire island where he had taken decades of observations of all the planets, like Mars and Jupiter, at least every night for which the weather was clear, with the naked eye. He was the last of the naked-eye astronomers. He had all this data which Kepler could use to confirm his theory. Kepler started working with Tycho, but Tycho was very jealous of the data. He only gave him little bits of it at a time. Kepler eventually just stole the data. He copied it and had to have a fight with Brahe's descendants. He did get the data, and then he worked out, to his disappointment, that his beautiful theory didn't quite work.
实际上那是一整座岛,他在那里用肉眼对所有行星进行了长达数十年的观测,像火星、木星这些,只要天气晴朗,他几乎每晚都在观测。他是最后一位肉眼观测的天文学家。他手上有开普勒验证理论所需要的全部数据。开普勒开始跟第谷合作,但第谷对这些数据非常小气,每次只给他一点点。开普勒最后干脆把数据偷走了。他把数据抄了下来,还得跟布拉赫的后人打官司。他确实拿到了数据,然后他推算发现,令他失望的是,他那个优美的理论并不太成立。
便签笔记
2:53
The data was off from his Platonic solid theory by 10% or something. He tried all kinds of fudges, moving the circles around, and it didn't quite work. But he worked on this problem for years and years, and eventually, he figured out how to use the data to work out the actual orbits of the planets. That was an incredibly clever, genius amount of data analysis. And then he worked out that the orbits were actually ellipses, not circles, which was shocking for him. So he worked out the two laws of planetary motion: the ellipses, and also that equal areas sweep out equal times. Then ten years later, after collecting a lot of data—the furthest planets like Saturn and Jupiter were the hardest for him to work out—he finally worked out this third law, that the time it takes for a planet to complete its orbit was proportional to some power of the distance to the Sun.
数据与他的柏拉图立体理论差了 10% 左右。他试了各种各样的修补办法,挪动那些圆,还是不太行。但他在这个问题上钻研了很多年,最终他弄明白了如何利用这些数据推算出行星的真实轨道。那是一次极其巧妙、堪称天才的数据分析。然后他推算出轨道其实是椭圆,而不是圆,这让他非常震惊。于是他得出了行星运动的两条定律:轨道是椭圆,以及相同时间内扫过相同面积。然后又过了十年,在收集了大量数据之后——像土星、木星这些最远的行星对他来说最难处理——他终于得出了第三定律,即一颗行星运行所需的时间行星完成一圈公转所需的时间,与它到太阳距离的某个幂次成正比。
便签笔记
02开普勒是高温度 LLM 吗
3:55
These are the three famous Kepler's laws of motion. He had no explanation for them. It was all driven by experiment, and it took Newton a century later to give a theory that explained all three laws at once. The take I want to try on you is that Kepler was a high-temperature LLM. Newton comes up with this explanation of why the three laws of planetary motion must be true. Of course, the way that Kepler discovers the laws of planetary motion, or figures out the relative orbits of the different planets, is as you say a work of genius.
这就是著名的开普勒三大定律。他对这些定律没有任何解释,全都是由观测数据推出来的,直到一个世纪之后,牛顿才提出一套理论,一举解释了这三条定律。我想跟你抛出的观点是:开普勒就是一个高温度参数的大语言模型。牛顿提出了一套解释,说明行星运动三定律为什么必然成立。当然,开普勒发现行星运动定律、推算出各行星相对轨道的方式,正如你所说,是天才之举。
便签笔记
4:29
But through his career, he's just trying random relationships. In fact, in the book in which he writes down the third law of planetary motion, it's an aside on The Harmonics of the World, which is just a book about how all these different planets have these different harmonies. And the reason there's so much famine and misery on Earth is because the Earth is mi-fa-mi, that's the note of Earth. It's all this random astrology, but in there is the cube-square law, which tells you what relationship the period has to a planet's distance from the Sun.
但纵观他的职业生涯,他其实就是在不停地试各种随机的关系式。事实上,在他写下行星运动第三定律的那本书里,那只是《世界的和谐》中的一段旁白,那本书讲的是各个行星如何有各自不同的和声。而他认为地球上之所以有这么多饥荒和苦难,是因为地球的音是 mi-fa-mi,这就是地球的音符。全是这类随意的占星学内容,但其中就藏着那个立方—平方定律,它告诉你公转周期和行星到太阳距离之间是什么关系。
便签笔记
5:00
As you were detailing, if you add that to Newton's F=ma and the equation for centripetal acceleration, you get the inverse-square law. And so Newton works that out. But the reason I think this is an interesting story is that I feel LLMs can do the kind of thing of trying random relationships for twenty years, some of which make no sense, as long as there's a verifiable data bank like Brahe's dataset. "Ok, I'm going to try out random things about musical notes, Platonic objects, or different geometries, I have this bias that there's some important thing about the geometry of these orbits." Then one thing works. As long as you can verify it, these empirical regularities can then drive actual deep scientific progress.
正如你刚才讲的,如果把它和牛顿的 F=ma 以及向心加速度的公式结合起来,就能得到平方反比定律。牛顿正是这么推出来的。但我觉得这个故事有意思的地方在于,我感觉大语言模型完全可以做这种事——花二十年去尝试各种随机的关系,其中很多毫无道理,只要有一个可以验证的数据库,比如第谷的观测数据集。“好,我来试试音符、柏拉图立体、各种几何形状之类的随机想法,我有个直觉,觉得这些轨道的几何形状里藏着某种重要的东西。”然后其中一个碰巧成立了。只要你能验证它,这些经验规律就能真正推动深层次的科学进步。
便签笔记
5:44
Traditionally, when we talk about the history of science, idea generation has always been the prestige part of science. A scientific problem comes with many steps. You have to identify a problem, and then you have to identify a good, fruitful problem to work on. Then you need to collect data, figure out a strategy to analyze the data, and make a hypothesis. At this point, you need to propose a good hypothesis, and then you need to validate. Then you need to write things up and explain. There are a dozen different components. The ones we celebrate are these eureka genius moments of idea generation. Kepler certainly had to cycle through many ideas, several of which didn't work. I bet there were many that he didn't even publish at all because they just didn't fit. That's an important part of the process, trying all kinds of random things and seeing if they worked.
传统上,我们谈科学史时,想法的产生一直是科学中最有声望的部分。一个科学问题包含很多环节。你得先识别出一个问题,然后还得判断哪个问题值得做、有成果可期。接着你要收集数据,想出分析数据的策略,然后提出假设。到这一步,你需要提出一个好的假设,然后你还要去验证。之后你得把结果写出来并加以解释。这里面有十几个不同的环节。而我们所歌颂的,是那些灵光一闪、天才般产生想法的时刻。开普勒肯定也是反复琢磨过很多想法,其中好几个都行不通。我敢说还有很多他根本没发表,因为完全对不上。这也是过程中很重要的一部分——去尝试各种五花八门的东西,看看行不行得通。
便签笔记
6:41
But as you say, it has to be matched by an equal amount of verification, otherwise it's slop. We celebrate Kepler, but we should also celebrate Brahe for his assiduous data collection, which was ten times more precise than any previous observation. That extra decimal point of accuracy was essential for Kepler to get his results. He was using Euclidean geometry and the most advanced mathematics he could use at the time to match his models with the data. All aspects had to be in play: the data, the theory, and the hypothesis generation. I'm not sure nowadays that hypothesis generation is the bottleneck anymore. Science has changed in the century since.
但正如你所说,这必须有同等分量的验证与之匹配,否则就只是垃圾内容。我们歌颂开普勒,但也应该歌颂第谷孜孜不倦的数据收集工作,他的精度比此前任何观测都高十倍。正是那多出来的一位小数精度,才让开普勒得出了他的结果。他用的是欧几里得几何和当时能用到的最先进的数学,来让模型与数据吻合。所有环节都得到位:数据、理论,还有假设的生成。我倒不确定如今假设生成还是不是瓶颈了。这一个世纪以来,科学已经变了。
便签笔记
03数据先行的科学新范式
7:38
Classically, the two big paradigms for science were theory and experiment. Then in the 20th century, numerical simulation came along, so you can do computer simulations to test theories. Finally, in the late 20th century, we had big data. We had the era of data analysis. A lot of new progress is actually driven now by analyzing massive datasets first. You collect large datasets and then draw patterns from them to deduce thoughts. This is a little bit different from how science used to work, where you make a few observations or have one out-of-the-blue idea, and then collect data to test your idea.
经典意义上,科学有两大范式:理论和实验。到了二十世纪,数值模拟出现了,你可以用计算机仿真来检验理论。最后在二十世纪后期,我们有了大数据,进入了数据分析的时代。如今很多新进展其实是先从分析海量数据集开始的。你先收集大规模数据集,再从中找出模式,进而推导出结论。这和过去科学的运作方式有点不一样——过去是你做几次观测,或者突然冒出一个想法,然后再去收集数据来检验这个想法。
便签笔记
8:17
That's the classic scientific method. Now it's almost reversed. You collect big data first, and then you try to get hypotheses from it. Kepler was maybe one of the first early data scientists, but even he didn't start with Tycho's dataset and then analyze it. He had some preconceived theories first. It seems like this is less and less the way we make progress, just because the data is so much more massive and useful. Oh, interesting. I feel like the 20th-century science that you're describing actually very well describes what happened with Kepler. He did have these ideas—1595 and '96 is where he comes up with the polygons and then the Platonic objects theory—but they were wrong.
那是经典的科学方法。现在几乎反过来了:你先收集大数据,然后再试着从中得出假设。开普勒也许算是最早的数据科学家之一,但即便是他,也不是从第谷的数据集出发再去分析的。他是先有了一些先入为主的理论。看起来这种方式在今天越来越不是我们取得进展的路径了,只因为数据变得庞大和有用得多。哦,有意思。我倒觉得你描述的二十世纪科学,恰恰很好地描述了开普勒身上发生的事。他确实先有过一些想法——1595 年和 1596 年,他提出了多边形、后来又提出柏拉图立体的理论——但那些都是错的。
便签笔记
9:06
Then a few years later, he gets Brahe's data, and it's only after twenty years of trying random things that he gets this empirical regularity. It actually feels a bit closer to Brahe's data being analogous to some massive data bank of simulations, and now that you've got the data, you can keep trying random things. If it wasn't for that, Kepler would be out there just writing books about harmonics and Platonic objects, and there would be nothing to actually verify against. The data was extremely important. The distinction I was trying to make was that traditionally, you make a hypothesis and then you test it against data. But now with machine learning, data analysis, and statistics, you can start with data and through statistics work out laws that were not present before. Kepler's third law is a little bit like this, except that instead of having the thousand data points that Brahe had, Kepler had six data points.
然后几年后他拿到了第谷的数据,正是在尝试了二十年各种随机想法之后,他才得出这个经验规律。这其实更接近于把第谷的数据类比成某个庞大的仿真数据库,而一旦你有了数据,你就可以不断去试各种随机的东西。要不是有这些数据,开普勒可能就只是在那儿写关于和声和柏拉图立体的书,根本没有东西可以拿来验证。数据极其重要。我刚才想区分的是:传统上你先提出假设,然后拿数据去检验它。但现在有了机器学习、数据分析和统计学,你可以从数据出发,通过统计手段推出此前并不存在的规律。开普勒第三定律有点像这样,只不过他手里不是第谷那上千个数据点,而只有六个数据点。
便签笔记
10:09
For every planet, he knew the length of the orbit and the distance to the Sun. There were five or six data points, and he did what we would now call regression. He fit a curve to these six data points and got a square-cube law, which was amazing. But he was quite lucky that these six data points gave him the right conclusion. That's not enough data to be really reliable. There was a later astronomer, Johann Bode, who took the same data—the distances to the planets—and inspired by Kepler, he had a prediction that the distances to the planets formed a shifted geometric progression.
对每一颗行星,他知道公转周期的长度和到太阳的距离。总共只有五六个数据点,而他做的正是我们今天所说的回归。他给这六个数据点拟合了一条曲线,得到了平方—立方定律,这非常了不起。但他其实相当幸运,这六个数据点竟然给出了正确的结论。这点数据量其实不足以真正可靠。后来有位天文学家叫约翰·波得,他拿了同样的数据——各行星的距离——受开普勒启发,他预测各行星的距离构成一个平移过的等比数列。
便签笔记
10:48
He also fit a curve, except there was one point missing. There was a big gap between Mars and Jupiter. His law predicted that there was a missing planet. It was kind of a crank theory, except when Uranus was discovered by Herschel, the distance to Uranus fit exactly this pattern. Then Ceres was discovered in the asteroid belt, and it also fit the pattern. People got really excited that Bode had discovered this amazing new law of nature. But then Neptune was discovered, and it was way off. Basically it was just a numerical fluke.
他也拟合了一条曲线,只不过中间少了一个点。火星和木星之间有一大段空隙。他的定律预言那里有一颗缺失的行星。这本来像是民科理论,可是当赫歇尔发现天王星时,天王星的距离恰好符合这个规律。后来又在小行星带发现了谷神星,它也符合这个规律。人们非常兴奋,觉得波得发现了一条了不起的自然新定律。但后来海王星被发现了,结果偏差极大。说白了那就是个数值上的巧合。
便签笔记
11:26
There were six data points. Maybe one reason why Kepler didn't highlight his third law as much as the first two laws is that instinctively, even though he didn't have modern statistics, he kind of knew that with six data points, he had to be somewhat tentative with the conclusions. To ask the question about the analogy more explicitly, does this analogy make sense if in the future we have smarter and smarter AIs? We'll have millions of them, and they can go out and hunt for all these empirical irregularities. It sounds like you don't think the bottleneck in science is finding more things that are the equivalent of the third law of planetary motion for each given field, so that later on somebody can say, "Oh, we need a way to explain this. Let's work out the math. Here's the inverse-square law of gravity."
当时只有六个数据点。也许开普勒没有像强调前两条定律那样强调第三定律,一个原因就是他凭直觉——尽管他没有现代统计学——隐约知道只靠六个数据点,下结论时必须留有余地。我把这个类比的问题问得更明确一点:如果未来我们有越来越聪明的 AI,这个类比还成立吗?我们会有几百万个 AI,它们可以出去搜寻各种经验上的反常规律。听起来你并不认为科学的瓶颈在于,为每个领域找到更多相当于行星运动第三定律的东西,好让后来有人说:“哦,我们需要一个办法来解释这个。我们来把数学推出来。这就是万有引力的平方反比定律。”
便签笔记
04想法零成本,验证成瓶颈
12:18
I think AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero. It’s an amazing thing, but it doesn't create abundance by itself. Now the bottleneck is different. We're now in a situation where suddenly people can generate thousands of theories for a given scientific problem. Now we have to verify them, evaluate them. This is something which we have to change our structures of science to actually sort this out. Traditionally, we build walls. In the past, before we had AI slop, we had amateur scientists have their own theories of the universe, many of which were of very little value.
我认为 AI 已经把产生想法的成本压到了几乎为零,这和互联网把通信成本压到几乎为零非常相似。这是件了不起的事,但它本身并不创造富足。现在瓶颈变了。我们如今的处境是,人们突然可以针对某个科学问题生成上千个理论。而现在我们得去验证它们、评估它们。这就要求我们改变科学的组织结构才能理清。传统上我们是筑墙。在有 AI 垃圾内容之前,我们有业余科学家提出自己的宇宙理论,其中很多价值极低。
便签笔记
13:07
We built these peer review publication systems to filter out and try to isolate the high signal ideas to test. But now that we can generate these possible explanations at massive scale, and some of them are good and a lot are terrible, human reviewers are already being overwhelmed. Many journals are reporting that AI-generated submissions are just flooding their submissions. It's great that we can generate all kinds of things now with AI, but it means that the rest of the aspects of science have to catch up: verification, validation, and assessing what ideas actually move the subject forward and which ones are dead ends or red herrings. That's not something we know how to do at scale.
于是我们建立了同行评审的发表制度,来过滤掉噪音,尽量挑出高信噪比的想法去检验。但现在我们可以大规模地生成各种可能的解释,其中有些不错,很多则很糟糕,人类审稿人已经不堪重负了。很多期刊都反映,AI 生成的投稿正在淹没他们的投稿系统。我们现在能用 AI 生成各种东西当然很好,但这意味着科学的其他环节必须跟上:验证、确认,以及判断哪些想法真正推动了学科前进,哪些是死胡同或误导性的线索。这可不是我们知道该如何大规模去做的事。
便签笔记
14:02
For each individual paper, we can have a debate among scientists and get to a consensus in a few years. But when we're generating a thousand of these every day, this doesn't work. There's this incredibly interesting question. If you have billions of AI scientists, not only how do you gauge which ones are real progress, but how do you... This is actually a question that human science has had to face and we've solved somehow, and I’m actually not sure how we solved this. Let's say in the 1940s, if you're at Bell Labs and there are these new technologies coming out.
对单篇论文,我们可以让科学家们辩论,几年之内达成共识。但当我们每天生成上千篇这样的东西时,这套机制就失效了。这里有个特别有意思的问题。如果你有几十亿个 AI 科学家,不仅要判断哪些是真正的进展,还要……这其实是人类科学也曾面对并以某种方式解决了的问题,而我其实不太确定我们是怎么解决的。比如说在 1940 年代,如果你在贝尔实验室,当时有各种新技术冒出来。
便签笔记
14:37
Pulse-code modulation, how do you transfer signals? How do you digitize signals? How do you transfer them over analog wires? There are all these papers about the engineering constraints and the details, and then there's one which comes up with the idea of the bit, which has implications across many different fields. You need some system which can then look at that and say, "Okay, we need to apply this to probability. We need to apply this to computer science," et cetera. In the future, the AIs are coming up with the next version of this unifying concept.
脉冲编码调制、怎么传输信号、怎么把信号数字化?怎么在模拟线路上传输?有一大堆论文讲工程约束和技术细节,然后其中一篇提出了“比特”这个概念,它的影响横跨许多不同领域。你需要某种机制能识别出它并说:“好,我们得把这个用到概率论上。我们得把这个用到计算机科学上”,等等。在未来,AI 会提出这种统一性概念的下一个版本。
便签笔记
05进步为何难以事前评分
15:06
How would you identify it among millions of papers that might actually constitute progress, but which have much less in terms of general unifying ideas? A lot of it's the test of time. Many great ideas didn't actually get a great reception at the time they were first proposed. It was only after some other scientists realized that they could take it further and apply them to their own... Deep learning itself was a niche area of AI for a long time. The idea of getting answers entirely through training on data and not through first principles reasoning was very controversial, and it just took a long time before it started bearing fruit. You mentioned the bit. There were other proposals for computer architectures than the zero-one that is universal today.
在几百万篇论文里,你要怎么把它识别出来?那些论文可能也算是进展,但在提出统一性想法方面要弱得多。这在很大程度上要靠时间的检验。很多伟大的想法在最初提出时,其实并没有得到好的反响。只有等到别的科学家意识到可以把它往前推进、应用到自己的领域……深度学习本身也曾长期是 AI 里的一个小众方向。完全通过在数据上训练来得到答案、而不是靠第一性原理推理,这个想法当时争议很大,过了很久才开始结出果实。你提到了比特。除了今天通行的 0 和 1,当时还有别的计算机架构方案。
便签笔记
15:50
I think there were trits, three-valued logic. In an alternate universe, maybe a different paradigm would have shown up. The transformer, for example, is the foundation of all modern large language models, and it was the first deep learning architecture that really was sophisticated enough to capture language. But it didn't have to be that way. There could've been some other architecture that was the first to do it and once that was adopted, it would become the standard. One reason why it's hard to assess whether a given idea is going to be fruitful is that it depends on the future.
我记得有“三进制位”,也就是三值逻辑。在另一个平行宇宙里,也许会出现完全不同的范式。比如 Transformer,它是所有现代大语言模型的基础,是第一个真正复杂到足以捕捉语言的深度学习架构。但事情本来不必如此。本可以是另一种架构率先做到这一点,而一旦那种架构被采用,它就会成为标准。评估某个想法是否会开花结果之所以困难,一个原因就是这取决于未来。
便签笔记
16:29
It depends also on the culture and society, which ones get adopted, which ones don't. The base ten numeral system in mathematics is extremely useful, much better than the Roman numeral system, for instance. But again, there's nothing special about ten. It's a system that is useful for us because everyone else uses it. We've standardized it. We've built all our computers and our number representation systems around it, so we're stuck with it now. Some people occasionally push for other systems than decimal, but there's just too much inertia. It's not something where you can look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context both in the past and the future. So it may never be something that you can just reinforcement learn the same way that you can for much more localized problems.
这也取决于文化和社会——哪些被采纳,哪些没有。数学里的十进制记数法极其好用,比如说比罗马数字好太多了。但话说回来,“十”本身并没有什么特别之处。它之所以对我们有用,是因为其他人都在用。我们已经把它标准化了,我们所有的计算机和数字表示系统都是围绕它建的,所以现在被它套牢了。偶尔有人主张改用十进制以外的系统,但惯性实在太大。所以你没法把任何一项科学成就完全孤立地拿出来,在不了解其前后脉络的情况下给它一个客观评分。所以这也许永远不会是那种你可以像处理更局部化的问题那样,直接用强化学习去搞定的事情。
便签笔记
17:33
Often in the history of science when a new theory comes up that in retrospect we realize is correct, it seems to make implications that either make no sense because they're wrong, and we realize later on why they're wrong, or they're correct but seem wildly implausible at the time. As you talked about, Aristarchus had heliocentrism in the third century BC. The ancient Athenians were like, "This can't be because if the earth is going around the sun, we should see the relative position of the stars change as we're going around the sun, and the only way that wouldn't be the case is if they're so far away that you don't notice any parallax,"
科学史上常有这种情况:一个新理论出现,事后我们才意识到它是对的,但它当时推出的一些结论要么看起来毫无道理、因为确实是错的,我们后来才明白错在哪;要么是对的,但在当时看来极其难以置信。就像你说的,阿里斯塔克斯在公元前三世纪就提出了日心说。古雅典人的反应是:“这不可能,因为如果地球绕着太阳转,我们应该看到恒星的相对位置随着地球公转而变化,唯一说得通的情况就是它们远到你根本察觉不到视差”,
便签笔记
18:13
which is actually the correct implication. But there's times when the implication is incorrect and we just need to graduate to a better level of understanding. Leibniz would chide Newton and disagree with Newton's theory of gravity on the basis that it implied action at a distance, and they didn't know the mechanism, and Newton himself was sort of stunned that inertial mass and gravitational mass were the same quantity. All these things later were resolved by Einstein. But it was still progress. So the question for a system of peer review for AI would be: even if you can falsify a theory, how would you notice that it still constitutes progress relative to the thing before?
而这其实正是正确的推论。但也有些时候,推论本身是错的,我们只是需要提升到更好的理解层次。莱布尼茨曾嘲讽牛顿,反对牛顿的引力理论,理由是它隐含超距作用,而他们说不清其中的机制;牛顿本人也对惯性质量和引力质量竟然是同一个量感到相当困惑。所有这些后来都被爱因斯坦解决了。但那仍然是进展。所以对 AI 的同行评审体系来说,问题在于:即便你能证伪一个理论,你又怎么看出它相对于之前的东西仍然构成了进步?
便签笔记
18:49
Often, the ultimately correct theory initially is worse in many ways. Copernicus's theory of the planets was less accurate than Ptolemy's theory. Geocentrism had been developed for a millennium by that point, and they had made many tweaks and increasingly complicated ad hoc fixes to make it more and more accurate. Copernicus's theory was a lot simpler but much less accurate. It was only Kepler that made it more accurate than Ptolemy's theory. Science is always a work in progress. When you only get part of the solution, it looks worse than a theory which is incorrect but somehow has been completed to the point where it kind of answers all the questions. As you say, Newton's theory had big mysteries.
往往那个最终正确的理论,一开始在很多方面反而更糟。哥白尼的行星理论精度还不如托勒密的理论。地心说到那时已经发展了一千年,他们做了许多修补和越来越复杂的临时性修正,让它越来越准确。哥白尼的理论简洁得多,但准确度差很多。直到开普勒,它才比托勒密的理论更准确。科学永远是一项未完成的工作。当你只解决了一部分问题时,它看起来还不如一个虽然错误、却已经被完善到几乎能回答所有问题的理论。就像你说的,牛顿的理论有很大的谜团。
便签笔记
19:42
They had the equivalence of mass and action at a distance, which were only resolved with a very conceptually different approach centuries afterwards. Often progress has to be made not by adding more theories, but by deleting some assumptions that you have in your mind. One reason why geocentrism held on for so long is we had this idea that objects naturally want to stay at rest. This is the Aristotelian notion of physics, and so the idea that the Earth was moving… How come we weren't all falling over? Once you have Newton's laws of motion—an object in motion remains in motion and so forth—then it makes sense.
他们面对的是质量的等价性和超距作用,这些直到几个世纪后,才被一种概念上截然不同的方法所解决。很多时候,进步不是靠增加更多理论,而是靠删掉你脑子里的某些假设。地心说之所以能延续那么久,一个原因是我们有这样一个观念:物体天然是想保持静止的。这是亚里士多德式的物理观念,所以如果说地球在运动……那我们怎么没都摔倒呢?一旦你有了牛顿运动定律——运动的物体会保持运动等等——这就说得通了。
便签笔记
20:25
Conceptually, it's a very big leap to realize that the Earth is in motion. It doesn't feel like it's in motion. The biggest advances, like Darwin's theory of evolution, is the idea that species are not static. This is not obvious because you don't see evolution in your lifetime. Well, now we actually can, but it seems permanent and static. Right now we're going through a cognitive version of the Copernican revolution, where we used to think that human intelligence is the center of the universe, and now we're seeing that there are very different types of intelligence out there with very different strengths and weaknesses.
从概念上讲,意识到地球在运动是非常大的一步跨越。因为你感觉不到它在动。最大的那些进展,比如达尔文的进化论,核心就是物种不是一成不变的这个想法。这一点并不显而易见,因为你在有生之年看不到进化发生。当然现在我们其实能看到了,但当时看起来物种就是恒久不变的。眼下我们正在经历一场认知版的哥白尼革命:我们过去认为人类智能是宇宙的中心,而现在我们看到,世界上存在着非常不同类型的智能,各有非常不同的强项和弱项。
便签笔记
06达尔文 vs 牛顿:说服力也是科学
21:14
Our assessment of which tasks require intelligence, which ones don't, has to be reordered quite a bit. Trying to fit AI into our theories of scientific progress and what is hard and what is easy, we're struggling quite a lot. We have to ask questions that we've never really had to ask before. Or maybe the philosophers had, but now we all have to deal with it. This brings up a topic I've been very curious about. You mentioned Darwin's theory of evolution. There's this book, The Clockwork Universe by Edward Dolnick, which covers a lot of this era of history we're talking about.
我们对哪些任务需要智能、哪些不需要的判断,必须做相当大的重新排序。想把 AI 塞进我们关于科学进步、关于什么难什么易的既有理论里,我们相当吃力。我们不得不去问一些以前从来不需要问的问题。或者说哲学家们问过,但现在我们所有人都得面对它了。这让我想到一个我一直很好奇的话题。你提到了达尔文的进化论。有本书叫《钟表宇宙》,作者是 Edward Dolnick,讲的正是我们说的这段历史时期。
便签笔记
21:47
He has this interesting observation in there. The Origin of Species was published in 1859. Principia Mathematica was published in 1687. So The Origin of Species comes out two centuries after Principia. Conceptually, it seems like Darwin's theory is simpler. There's a contemporaneous biologist to Darwin, Thomas Huxley, who reads The Origin of Species and he says, "How stupid not to have thought of that." Nobody ever says that about Principia, chiding themselves for not having beaten Newton to gravity. So there's a question of why did it take longer?
他在书里有个很有意思的观察。《物种起源》出版于 1859 年,而《自然哲学的数学原理》出版于 1687 年。所以《物种起源》比《原理》晚了两个世纪。从概念上看,达尔文的理论似乎更简单。跟达尔文同时代有位生物学家,托马斯·赫胥黎,他读完《物种起源》后说:“我怎么这么蠢,居然没想到这一点。”却从来没人对《原理》说过这种话,为自己没抢在牛顿之前想到万有引力而自责。所以问题是:为什么它反而更晚出现?
便签笔记
22:17
It seems like a big part of the reason is what you were saying. The evidence for natural selection is overwhelming in a certain sense, but it's cumulative and retrospective, whereas Newton can just say, "Here are my equations. Let me see the moon's orbital period and its distance, and if it lines up, then we've made progress." Lucretius actually had this idea that species adapted to their environment in the first century BC but nobody really talks about it until Darwin because Lucretius couldn't run some experiment and force people to pay attention.
我觉得很大一部分原因就是你刚才说的。从某种意义上说,自然选择的证据是压倒性的,但它是累积性、回溯性的;而牛顿可以直接说:“这就是我的方程。我来看看月球的公转周期和距离,如果对得上,那我们就取得进展了。”其实卢克莱修在公元前一世纪就有过这样的想法:物种会适应它们的环境。但直到达尔文之前几乎没人认真谈论它,因为卢克莱修没法做实验,逼着人们去关注。
便签笔记
22:48
I wonder if we'll in retrospect end up seeing much more progress in domains which have this kind of tight data loop where you can verify them quite easily, even though they're conceptually much more difficult. I think one aspect of science is that it's not just creating a new theory and validating it, but communicating it to others. Darwin was an amazing science communicator. He wrote in English, in natural language. I'm speaking like a— No Lean. I have to get out of my technical mindset. He spoke in plain English, didn't use equations, and he synthesized a lot of disparate facts. Little pieces of evolution had been worked out in the past, but he had this very compelling vision. Again, he was still missing things.
我在想,我们回头看的时候,会不会发现进展更多集中在那些有紧密数据闭环、很容易验证的领域,哪怕它们在概念上要困难得多。我认为科学的一个方面是,它不只是提出一个新理论并验证它,还要把它传达给别人。达尔文是了不起的科学传播者。他用英语写作,用自然语言写作。我这话说得像——不是用 Lean。我得跳出我的技术思维。他用平实的英语表达,不用方程,而且他把大量零散的事实综合了起来。进化论的一些零碎片段过去就有人做过,但他有一个极具说服力的整体图景。当然,他仍然有欠缺的地方。
便签笔记
23:42
He didn't know the mechanism for heredity, he didn't have DNA. But his writing style was persuasive, and that helped a lot. Newton wrote in Latin. He had invented entire new areas of mathematics just to explain what he was doing. He was also from an era where scientists were much more secretive and competitive. Academia is still competitive, but it was even worse back in Newton's day. He held back some of his best insights because he didn't want his rivals to get any advantage. He was also a somewhat unpleasant person from what I gather.
他不知道遗传的机制,他没有 DNA。但他的写作风格很有说服力,这帮了大忙。牛顿则是用拉丁文写的。而且他为了解释自己在做什么,发明了全新的数学分支。他还处在一个科学家远更保密、更爱竞争的时代。学术界现在依然有竞争,但在牛顿那个年代更严重。他会藏起自己最好的一些洞见,因为不想让对手占到便宜。据我所知,他为人也有点不太讨喜。
便签笔记
24:23
It was only a couple of decades after Newton when other scientists explained his work in much simpler terms that they became widespread. The art of exposition and making a case and creating a narrative is also a very important part of science. If you have the data, it helps, but people need to be convinced, otherwise they will not push it further or take the initial investment to learn your theory and really explore it. That's another thing which is really hard to reinforcement learn on. How can you score how persuasive you are? Well, there are entire marketing departments trying to do this. Maybe it's good that AI is not yet optimized to be persuasive. There's a social aspect to science.
直到牛顿之后几十年,其他科学家用更简单的语言解释了他的工作,这些成果才广泛传播开来。所以阐述、论证、构建叙事的这门艺术,也是科学中非常重要的一部分。有数据当然有帮助,但你得说服别人,否则他们不会把它往前推进,也不会愿意付出最初的成本去学习你的理论、真正深入探索它。这是另一件很难用强化学习去训练的事。你怎么给“你有多有说服力”打分?当然,有整整一批市场营销部门在试图做这件事。也许 AI 目前还没有被优化成很有说服力,这反倒是好事。科学是有社会属性的。
便签笔记
25:19
Even though we pride ourselves on having an objective side to it, where there's data and experiment and validation, we still have to tell stories and convince our fellow scientists. That's a soft, squishy thing. It's a combination of data and painting a narrative, and it's a narrative of gaps. Even with Darwin, as I said, there were pieces of his theory he could not explain. But he could still make a case that in the future, people would find transitional forms, that they would find the mechanism of inheritance, and they did. I don't know how you can quantify that in such a precise way that you can start doing reinforcement learning.
尽管我们以科学有客观的一面为傲——有数据、有实验、有验证——但我们仍然得讲故事,去说服同行科学家。这是件柔性的、模糊的事。它是数据与描绘叙事的结合,而且是一种关于“空缺”的叙事。就连达尔文,正如我说的,他的理论里有些部分他解释不了。但他仍然能论证说,将来人们会找到过渡形态,会找到遗传的机制——后来确实找到了。我不知道你要怎么把这种东西量化到足够精确的程度,好让你能拿它做强化学习。
便签笔记
26:06
Maybe that will be forever the human side of science. One takeaway I had from reading and watching your stuff on the cosmic distance ladder… By the way, I highly recommend people watch your series with 3Blue1Brown on the cosmic distance ladder. One takeaway was that the deductive overhang in many fields could be so much bigger than people realize. If you just had the right insight about how to study a problem, you might be surprised at how much more you could learn about the world. I wonder if you think that's a product of astronomy at the particular times in history that you're studying. Or is it just that based on the data that is incident on the Earth right now, we could actually divine a lot more than we happen to know?
也许这将永远是科学中属于人的那一面。我读了、看了你关于宇宙距离阶梯的内容后有一个收获……顺便说一句,我强烈推荐大家去看你和 3Blue1Brown 合作的宇宙距离阶梯系列。我的一个收获是:很多领域里“待推导的余量”可能远比人们意识到的要大。只要你对如何研究一个问题有了正确的洞见,你可能会惊讶于你还能从这个世界上了解到多少东西。我很好奇你是否认为,这是天文学在你研究的那些特定历史时期才有的特殊现象;还是说,仅凭当下落到地球上的这些数据,我们其实就能推断出比我们现在已知的多得多的东西?
便签笔记
26:52
Astronomy was one of the first sciences to really embrace data analysis and squeezing every last possible drop of information out of the information they had because data was the bottleneck. It still is the bottleneck. It's really hard to collect astronomical data. Astronomers are world-class in extracting all kinds of conclusions from little traces of data, almost like Sherlock. I hear that for a lot of quant hedge funds, their preferred hire is an astronomy PhD, actually. They are also very interested for other reasons in extracting signals from various random bits of data.
天文学是最早真正拥抱数据分析、把手头信息里最后一滴信息都榨出来的学科之一,因为数据是瓶颈。它现在仍然是瓶颈。收集天文数据真的很难。天文学家在从极少的数据痕迹中提取各种结论这件事上是世界一流的,几乎像福尔摩斯一样。我听说很多量化对冲基金,他们最青睐的招聘对象其实就是天文学博士。他们出于别的原因,也非常想从各种零散随机的数据里提取信号。
便签笔记
07从微弱信号榨取信息
27:35
Okay, speaking of clever ideas, one of my listeners, Shawn, solved the puzzle that Jane Street made for my audience and posted a great walkthrough on X. For context, Jane Street trained a ResNet, shuffled all 96 layers, and then challenged people to put them back in the right order using only the model's outputs and training data. You can't brute force this – there's more possible orderings than atoms in the universe. So Shawn broke the problem into two different parts. First, pair the layers into 48 different blocks. And second, put those blocks in the right order.
好,说到聪明的点子,我的一位听众 Shawn 解出了 Jane Street 为我的听众出的那道谜题,并在 X 上发布了一份很棒的解题过程。背景是:Jane Street 训练了一个 ResNet,把全部 96 层打乱,然后挑战大家仅凭模型的输出和训练数据把它们恢复成正确顺序。你没法暴力破解——可能的排列数比宇宙中的原子还多。于是 Shawn 把问题拆成了两部分。第一,把这些层配对成 48 个块;第二,把这些块排成正确的顺序。
便签笔记
28:06
For pairing, Shawn realized that in a well-trained ResNet, the product of two weight matrices in a residual block should have a distinctive negative diagonal pattern. This arises as a way for the model to keep the residual stream from growing out of control. From this insight, he was able to recover the right pairings. For ordering, Shawn noticed that the model seemed to improve if he sorted the blocks by the size of their residual contributions. Starting with that rough approximation, he combined a clever ranking heuristic with local swaps to recover the exact right order.
关于配对,Shawn 意识到在一个训练良好的 ResNet 里,残差块中两个权重矩阵的乘积应该呈现出一种显著的负对角线模式。这种模式的出现,是模型用来防止残差流失控增长的一种方式。凭这个洞见,他成功还原出了正确的配对。关于排序,Shawn 注意到,如果按各个块的残差贡献大小来排序,模型表现似乎会变好。以这个粗略近似为起点,他把一个巧妙的排序启发式和局部交换结合起来,还原出了完全正确的顺序。
便签笔记
28:36
His full walkthrough is linked in the description. Don't worry if you didn't get to this puzzle in time, though. There's still one up about backdoored LLMs that even Jane Street doesn't know how to solve. You can find it at janestreet.com/dwarkesh. Alright, back to Terence! We do under-explore how to extract extra information from various signals. Just to pick one random study, I remember reading once that people were trying to measure how often scientists actually read the papers that they cite. How do you measure this? You could try to survey different scientists, but they had a clever trick.
他的完整解题过程链接在简介里。不过,如果你没赶上这道题也别担心。还有一道关于带后门的大语言模型的题目仍然开放,连 Jane Street 自己都不知道怎么解。你可以在 janestreet.com/dwarkesh 找到它。好,我们回到 Terence!我们确实没有充分探索如何从各种信号中提取额外信息。随便举个研究的例子,我记得读到过有人想测量科学家到底有多经常真的读过他们引用的论文。你怎么测这个?你可以去调查各路科学家,但他们用了个聪明的办法。
便签笔记
29:21
Many citations have little typos, like a number is wrong or punctuation is almost wrong. They measured how often a typo got copied from one reference to the next, and they could infer whether an author was just copying and pasting a reference without actually checking it. From that, they were able to infer some measure of how much attention people were paying. So there are some clever tricks to extract… These questions you posed earlier of how we can assess whether a scientific development is fruitful, interesting, or represents real progress… Maybe there are really useful metrics or footprints of this phenomenon in data.
很多引文里有小错别字,比如某个数字写错了,或者标点稍微有点问题。他们测量了一个错别字从一篇参考文献被抄到下一篇的频率,由此就能推断某位作者是不是只是复制粘贴了引文,而根本没去核对原文。从中他们就能推断出人们到底有多用心的某种度量。所以确实有一些聪明的技巧可以提取……你前面提出的那些问题,比如我们怎么评估一项科学进展是否富有成果、是否有趣、是否代表真正的进步……也许在数据里确实存在关于这种现象的、非常有用的指标或足迹。
便签笔记
08Erdős 问题:AI 的矮墙与高墙
30:12
We can examine citations and how often something is mentioned in a conference. Maybe there's a lot of sociology of science research to be done that could actually detect these things. Maybe we should get some astronomers on the case, actually. That brings us nicely to the progress that, from the outside, it seems like AI for math is making. You had a post recently where you pointed out that over the last few months, AI programs have solved fifty out of the eleven hundred odd Erdős problems. I don’t know if it’s still correct, but as of a month ago you said that there had been a pause because the low-hanging fruit had been picked.
我们可以考察引用情况,考察某个东西在会议上被提及的频率。也许有很多科学社会学的研究值得去做,真的能够检测到这些东西。也许我们该找几位天文学家来干这活儿。这正好很自然地引到了另一个话题:从外部看,AI 在数学上似乎正在取得的进展。你最近有一个帖子指出,在过去几个月里,AI 程序解决了大约一千一百多道 Erdős 问题中的五十道。我不知道现在是否还是这样,但截至一个月前,你说进展停了一阵,因为容易摘的果子已经被摘完了。
便签笔记
30:55
First of all, I'm curious if that is still the case, that we have picked the low-hanging fruit and now we're at this plateau currently. It does seem so. Fifty-odd problems have been solved with AI assistance, which is great, but there's like six hundred to go. People are still chipping away at one or two of these right now. We're seeing a lot fewer pure AI solutions now where the AI just one-shots the problem. There was a month where that happened and that has stopped, not for lack of trying. I know of three separate attempts to get frontier model AIs to just attack every single one of the problems simultaneously. They pick out some minor observations, or maybe they find that some problem was already solved in the literature, but there hasn't been any further purely AI-powered solution yet. People are using AI a lot currently.
首先我好奇,现在是否还是这样:低垂的果实已经被摘完,我们目前处在一个平台期。看起来确实如此。五十来道问题在 AI 协助下被解决了,这很棒,但还有大概六百道没解决。现在人们还在一点一点地啃其中一两道。我们现在看到的纯 AI 解法少多了,就是 AI 一次性直接把问题拿下的那种。曾经有一个月出现过这种情况,后来就停了,并不是因为大家没在尝试。我知道至少有三次独立的尝试,让前沿模型同时去攻每一道问题。它们会挑出一些次要的观察,或者发现某个问题其实文献里早就解决过,但至今还没有再出现纯靠 AI 完成的解法。目前人们大量在用 AI。
便签笔记
31:50
Someone might use AI to generate a possible proof strategy, and then another person will use a separate AI tool to critique it, rewrite it, generate some numerical data for it, or do a literature survey. Some problems have been solved by an ongoing conversation between lots of humans and lots of AI tools. But it does seem like it was this one-off thing. Maybe one analogy for these problems is that you're in some sort of mountain range with all kinds of cliffs and walls. Maybe there's a little wall which is three feet high, and one that's six feet high, and then there's fifteen feet high, and then there are some mile-high cliffs.
可能有人用 AI 生成一个可能的证明策略,然后另一个人用另一个 AI 工具去批评它、重写它、为它生成一些数值数据,或者做文献调研。有些问题是靠很多人和很多 AI 工具之间持续对话解决的。但那一波看起来确实像是一次性的事件。对这些问题,也许有个比喻是:你身处某种山脉之中,里面各种悬崖和墙壁。也许有一堵三英尺高的小墙,有一堵六英尺高的,再有十五英尺高的,然后还有一些一英里高的悬崖。
便签笔记
32:39
You're trying to climb as many of these cliffs as possible, but it's in the dark. We don't know which ones are tall, which ones are short. So we try to light some candles and make some maps, and slowly we figure out some of them are climbable. Some of them we can identify a partial track in the wall that you can reach first. These AI tools, they're like jumping machines that can jump two meters in the air, higher than any human. Sometimes they jump in the wrong direction, and sometimes they crash, but sometimes they can reach the tops of the lowest walls that we couldn't reach before.
你想尽可能多地爬上这些崖壁,但周围是黑的。我们不知道哪些高、哪些矮。于是我们试着点上蜡烛、画点地图,慢慢弄清楚其中一些是能爬的。有些我们能识别出墙面上一段可以先够到的落脚路径。而这些 AI 工具,就像能跳两米高的跳跃机器,比任何人类都跳得高。有时候它们跳错方向,有时候直接摔了,但有时候它们能够到我们以前够不着的那些最矮的墙顶。
便签笔记
33:18
We've just set them loose in this mountain range, hopping around. There was this exciting period where they could actually find all the low ones and reach them.
我们刚刚把它们放进这片山脉里,让它们到处蹦。曾有那么一段令人兴奋的时期,它们确实能找到所有矮的、并且够到它们。
便签笔记
33:32
Maybe the next time there's a big advance in the models, they will try it again, and a few more will be breached. But it's a different style of doing mathematics. Normally we would hill climb, make little markers, and try to identify partial things. These tools either succeed or they fail. They've been really bad at creating partial progress or identifying intermediate stages that you should focus on first. Going back to this previous discussion, we don't have a way of evaluating partial progress the same way we can evaluate a one-shot success or failure of solving a problem.
也许下一次模型有大的进步时,它们会再试一轮,又会有几堵墙被攻破。但这是一种不同风格的做数学方式。通常我们会爬坡式推进,做点标记,试着识别出部分性的东西。而这些工具要么成功要么失败。它们在做出部分进展、或识别出应该优先攻克的中间阶段方面,一直非常糟糕。回到刚才的讨论:我们没有办法像评估“一次性解题成功或失败”那样去评估部分进展。
便签笔记
09AI 擅广度,人类擅深度
34:19
There's two different ways to think through what you've just said. One of them is more bearish on AI progress, and one of them is more bullish. The bearish one being, "Oh, they're only getting to a certain height of wall, which is not as high as humans are reaching." The second is that they have this powerful property that once they achieve a certain waterline, they can fill every single problem that is available at that waterline, which we simply can't do with humans. We can't make a million copies of you and give each of them a million dollars of inference compute and have you do a hundred years of subjective time research on a million different problems at the same time. But once AIs reach Terence Tao-level, they could do that. Once they reach intermediate levels, they could do the intermediate version of that. The same reason that we should be bearish now is the reason we should be especially bullish. Not even when they achieve superhuman intelligence, but just when they achieve human-level intelligence,
你刚才说的这些,有两种不同的解读方式。一种对 AI 进展更看空,一种更看多。看空的那种是:“哦,它们只能够到一定高度的墙,没有人类够到的那么高。”而第二种是:它们有一个很强的性质——一旦达到某条水位线,它们就能把处在那条水位线以下的每一个问题都覆盖掉,而这是我们靠人类根本做不到的。我们没法造出一百万个你,给每一个都配上一百万美元的推理算力,让你在一百万个不同的问题上同时各做一百年主观时间的研究。但一旦 AI 达到陶哲轩级别,它们就能做到这一点。一旦它们达到中等水平,它们就能做中等版本的这件事。所以我们现在该看空的理由,恰恰也是我们该格外看多的理由。甚至不用等它们达到超人智能,只要它们达到人类水平的智能,
便签笔记
35:15
because their human-level intelligence is qualitatively wider and more powerful than our human-level intelligence. I agree. They excel at breadth, and humans excel at depth, human experts at least. I think they're very complementary. But our current way of doing math and science is focused on depth because that's where human expertise is, because humans can't do breadth. We have to redesign the way we do science to take full advantage of this breadth capability that we now have. We should have a lot more effort in creating very broad classes of problems to work on rather than one or two really deep, important problems. We should still have the deep, important problems, and humans should still be working on them. But now we have this other way of doing science.
因为它们那种人类水平的智能,在质上比我们人类水平的智能更宽广、更强大。我同意。它们擅长广度,而人类擅长深度,至少人类专家是这样。我觉得两者非常互补。但我们现在做数学和科学的方式是以深度为中心的,因为人类的专长就在深度上,因为人类做不到广度。我们必须重新设计做科学的方式,才能充分利用我们现在拥有的这种广度能力。我们应该投入更多精力去构造非常广泛的一大类问题来研究,而不是只盯着一两个真正深刻、重要的问题。我们仍然应该有那些深刻重要的问题,人类也仍然应该去攻克它们。但现在我们多了另一种做科学的方式。
便签笔记
36:10
We can explore entirely new fields of science by first getting these broad, moderately competent AIs to map it out and make all the easy observations. And then identify certain islands of difficulty, which human experts can then come and work on. I see very much a future of very complementary science. Eventually, you would hope to get both breadth and depth and somehow get the best of both worlds. But we need practice with the breadth side. It's too new. We don't even have the paradigms to really take full advantage of it. But we will, and then science will be unrecognizable after that, I think. To this point about complementarity, programmers have noticed that they're way more productive as a result of these AI tools.
我们可以去探索全新的科学领域:先让这些广博但能力中等的 AI 把它整体勾勒出来,把所有容易的观察都做掉。然后识别出若干个「困难的孤岛」,再让人类专家进场去攻克。我非常看好一个高度互补的科学未来。最终,你会希望广度和深度兼得,以某种方式取两者之长。但我们需要在广度这一侧多加练习。它太新了,我们甚至还没有真正能充分利用它的范式。但我们会有的,到那时科学将会变得面目全非,我是这么认为的。说到互补性这一点,程序员们已经注意到,用了这些 AI 工具之后他们的生产力大幅提升。
便签笔记
37:05
I don't know if you as a mathematician feel the same way, but it does seem like one big difference between vibe coding and vibe researching is that with software, the whole point is to have some effect on the world through your work. If it leads to you better understanding a problem or coming up with some clean abstraction to embody in your code, that is instrumental to the end goal. Whereas with research, the reason we care about solving the Millennium Prize Problems is that presumably that in the process of solving them, we discover new mathematical objects or new techniques that advance our civilization's understanding of mathematics.
不知道你作为数学家是否也有同感,但氛围编程(vibe coding)和「氛围做研究」之间似乎有个很大的差别:对软件来说,整个重点是通过你的工作对世界产生某种影响。如果它让你更好地理解了一个问题,或者让你想出某个漂亮的抽象并写进代码里,那都是服务于最终目标的手段。而做研究时,我们之所以在意解决千禧年大奖难题,大概是因为在解决它们的过程中,我们会发现新的数学对象或新的技术,从而推进人类文明对数学的理解。
便签笔记
37:45
So the proof is instrumental to the intermediate work. I don't know if you agree with that dichotomy or if that in any way will explain the relative uplift we'll see in software versus research. Certainly in math, the process is often more important than the problem itself. The problem is kind of a proxy for measuring progress. I think even in software, there are different types of software tasks. If you just create a webpage that does the same thing that a thousand other webpages do, there's no skill to be learned.
所以证明本身反而是服务于中间过程的工作的。不知道你是否认同这种二分,或者这是否能解释软件领域与研究领域所获提升的差异。在数学里确实如此,过程往往比问题本身更重要。问题某种程度上只是衡量进展的一个代理指标。我觉得即使在软件里,也有不同类型的软件任务。如果你只是做一个网页,做的事情和另外一千个网页一模一样,那就没什么技能可学。
便签笔记
38:19
Well, there is still some skill maybe that the individual programmer could pick up. But for boilerplate-type code, it's something that you should definitely offload to AI. Sometimes once you make the code, you still have to maintain it. There are issues with upgrading it and making it compatible with other things. I've heard programmers report that even if an AI can create the first prototype of a tool, making it mesh with everything else and making it interact with the real world in the way they want is an ongoing process. If you don't have the skills that you pick up from writing the code, that may impact your ability to maintain it down the road.
当然,可能对那个程序员个人来说还是能学到点东西。但对于样板式的代码,这绝对是应该外包给 AI 的。有时候代码写出来之后,你还得维护它。还有升级、以及让它与其他东西兼容的问题。我听程序员说,即使 AI 能做出一个工具的第一版原型,让它和其他所有东西咬合起来、让它按你想要的方式与真实世界交互,仍然是个持续不断的过程。如果你没有通过亲手写代码积累起来的那些技能,日后维护它的能力可能就会受影响。
便签笔记
39:10
So yes, certainly mathematicians, we've used problems to build intuition and to train people to have a good idea of what's true, what to expect, what is provable, and what is difficult. Just getting the answers right away may actually inhibit that process.
所以是的,数学家当然一直在用问题来培养直觉、训练人们对什么是对的、该期待什么、什么可证、什么困难形成良好的判断。如果答案直接就给你了,反而可能会阻碍这个过程。
便签笔记
39:35
I made a distinction between theory and experiment before. In most sciences, there's an equal division between the theoretical side and the experimental side. Math has been unique in that it's almost entirely theoretical. We place a premium on trying to have coherent, clean theories of why things are true and false. We haven't done many experiments as to, if we have two different ways to solve a problem, which is more effective. We have some intuition, but we haven't done large-scale studies where we take a thousand problems and just test them. But we can do that now.
我前面区分过理论和实验。在大多数科学里,理论和实验两边是势均力敌的。数学的独特之处在于,它几乎完全是理论性的。我们非常看重构建融贯、干净的理论来解释为什么某些命题为真或为假。我们很少做这样的实验:如果有两种不同的解题方法,哪一种更有效。我们有一些直觉,但没有做过大规模研究,比如取一千个问题然后逐一测试。但现在我们可以做了。
便签笔记
40:13
I think AI-type tools will actually revolutionize the experimental side of math, where you don't care so much about individual problems and the process of solving them, but you want to gather large-scale data about what things work and what things don't. The same way that if you're a software company and you want to roll out a thousand pieces of software, you don't really want to handcraft each one and learn lessons from each. You just want to find what workflows let you scale.
我认为 AI 这类工具真正会革命性改变的是数学的实验一侧,在那里你不太在意单个问题以及解决它的过程,而是想收集大规模的数据,看看什么方法有效、什么无效。就像你是一家软件公司,想要推出一千个软件产品,你并不想每一个都手工打造、再从每一个里总结经验。你只想找到哪些工作流能让你规模化。
便签笔记
101–2% 成功率与幸存者偏差
40:46
The idea of doing mathematics at scale is at its infancy. But that's where AI is really going to revolutionize the subject. I feel like a big crux in these conversations about how good AI will be for science is, I think you said this, that they're using existing techniques and modifying them. It would be interesting to understand how much progress one can make simply from using existing techniques. If I looked at the top math journals, how many of the papers are coming up with a new technique, whatever that means, versus using existing techniques on new problems? What is the overhang? If you just applied every known technique to every open problem, would that constitute a humongous uplift in our civilization's knowledge, or would that not be that impressive and useful? This is a great question, and we don't have the data to fully answer it yet. Certainly, a lot of work that human mathematicians do… When you take a new problem, one of the first things we do is we look at all the standard things that have worked on similar problems in the past, and we try them one by one.
「规模化地做数学」这个想法还处在襁褓期。但恰恰是在这里,AI 会真正颠覆这门学科。我觉得在讨论 AI 对科学有多大用处时,一个关键分歧点是——我想这是你说过的——它们是在使用已有的技术并加以改造。很有意思的一个问题是:光靠使用已有技术能取得多大进展?如果我去看顶级数学期刊,其中有多少论文是在提出某种新技术(不管这意味着什么),又有多少是把已有技术用在新问题上?这里的「未开发空间」有多大?如果你把每一种已知技术都拿去试每一个未解问题,这会不会构成对人类文明知识的巨大提升,还是说其实并没那么惊艳、也没那么有用?这是个很好的问题,我们还没有数据能完全回答它。当然,人类数学家做的很多工作……当你拿到一个新问题,我们最先做的事情之一就是把过去在类似问题上奏效的所有标准手段都翻出来,一个一个去试。
便签笔记
41:54
Sometimes that works, and that's still worth publishing because the question was important. Sometimes they almost work, and you have to add one more wrinkle to it, and that's also interesting. But the papers that go into the top journals are usually ones where the existing methods can kind of solve 80% of the problem, but then there is this 20% which is resistant and a new technique has to be invented to fill in the gaps. It's very rare now that a problem gets solved with no reliance on past literature, where all the ideas come out of nowhere. That was more common in the past, but math is so mature now that it's just so much of a handicap to not use the literature first.
有时候真的成了,而且因为问题本身重要,这依然值得发表。有时候差一点点就成了,你得再加一个小花招,那也很有意思。但能进顶级期刊的论文,通常是那种已有方法大概能解决问题的 80%,但剩下 20% 顽固不化,必须发明一种新技术来补上缺口的。如今一个问题完全不依赖既有文献就被解决,是非常罕见的,那种所有想法都凭空冒出来的情况。过去这更常见,但数学现在太成熟了,不先用文献简直是自缚手脚。
便签笔记
42:44
AI tools are getting really good at the first part of that, just trying all the standard techniques on a problem, often making fewer mistakes in applying them than humans. They still make mistakes, but I've tested these tools on little tasks that I can do, and sometimes they pick up errors that I make. Sometimes I pick up errors that they make. It's about a tie right now. But I haven't yet seen them take the next step. When there are holes in the argument where none of the things are working, then what do you do?
AI 工具在前半部分已经做得相当好了——把所有标准技术都在一个问题上试一遍,而且在应用这些技术时犯的错往往比人还少。它们仍然会犯错,但我拿这些工具试过一些我自己能做的小任务,有时候它们能挑出我犯的错,有时候我能挑出它们犯的错。目前大概是打平。但我还没见过它们迈出下一步。当论证里出现窟窿,所有办法都不奏效时,你该怎么办?
便签笔记
43:25
They can suggest random things, but often I find that trying to chase them down to make them work, and finding they don't work, wastes more time than it saves.
它们可以随便提些建议,但我常常发现,追着这些建议去验证、最后发现行不通,浪费的时间比省下的还多。
便签笔记
43:38
I think some fraction of problems that we currently think are hard will fall from this method, especially the ones that haven't received enough attention. With the Erdős problems, almost all of the 50 problems that were solved by AIs were ones for which there was basically no literature. Erdős posed the problem once or twice. Maybe some people tried it casually and couldn't do it, but they never wrote up anything. But it turned out that there was a solution, and it was just combining this one obscure technique that not many people know about with some other result in the literature.
我认为目前被我们视为困难的问题中,有一部分会因为这种方法而被攻破,尤其是那些还没得到足够关注的问题。在厄多斯(Erdős)问题上,被 AI 解决的那 50 来个问题里,几乎全都是基本没有相关文献的。厄多斯提过一两次这个问题,也许有人随手试过、没做出来,但从没写成文章。结果发现它其实是有解的,只是把某个很少人知道的冷门技术,和文献里的另一个结果结合起来就行了。
便签笔记
44:12
That's the median level of what AI can accomplish, and that's really great. It clears out 50 of these problems. So I think you will see some isolated successes. But what we found… Some people have done large-scale sweeps of these Erdős problems. If you only focus on the success stories, the ones that get broadcast on social media, it looks amazing. All these problems that haven't been solved for decades, now they're falling. But whenever we do a systematic study, on any given problem an AI tool has a success rate of maybe 1% or 2%.
这就是 AI 能达到的中位水平,而这已经很了不起了。它一下清掉了 50 个这样的问题。所以我认为你会看到一些零星的成功。但我们发现……有些人对这些厄多斯问题做过大规模的地毯式扫荡。如果你只看那些成功案例,那些在社交媒体上被广而告之的,看起来简直惊人。所有这些几十年没被解决的问题,现在纷纷倒下。但每当我们做系统性研究时,对任意给定的一个问题,AI 工具的成功率大概只有 1% 或 2%。
便签笔记
44:46
It's just that they can buy scale, and you just pick the winners. It looks great. I think there'll be a similar thing happening with the hundreds of really prestigious, difficult math problems out there. Some AI may get lucky and actually solve them, and there will be some backdoor to solve the problem that everyone else missed. That will get a lot of publicity. But then people will try these fancy tools on their own favorite problem, and they will again experience the 1% to 2% success rate. There'll be a lot of noise amongst the signal of when they're working and when they're not.
只不过它们能靠规模取胜,你只挑赢的那些来看,当然显得很棒。我想在那几百个真正声名显赫、极其困难的数学难题上,也会发生类似的事。某个 AI 可能会走运真的解出来,会存在某个所有人都错过的「后门」来解决那个问题。那会得到大量宣传。但接着人们会拿这些花哨的工具去试自己最喜欢的问题,然后再次体验到 1% 到 2% 的成功率。在「什么时候有效、什么时候无效」这个信号里,会掺杂大量噪声。
便签笔记
45:28
It will be increasingly important to collect these really standardized datasets. There are efforts now to create a standard set of challenge problems for AIs to solve, and not just rely on the AI companies to only publish their wins and not disclose their negative results. That will maybe give more clarity as to where we're actually at. Although I think it's worth emphasizing how much progress in AI it constitutes already, to have models that are capable of applying some technique that nobody had written down as applicable to this particular problem. The progress is simultaneously amazing and disappointing. It is a very strange feeling to see these tools in action. But people also acclimatize really quickly.
收集这些真正标准化的数据集会变得越来越重要。现在已经有人在努力建立一套标准的挑战题集给 AI 去解,而不是只依赖 AI 公司——它们只公布自己的胜绩,不披露负面结果。这也许能让我们更清楚地看到我们究竟处在什么位置。不过我觉得值得强调的是,这本身已经代表了 AI 相当大的进步:模型能够把某种从没有人写下来说可以用在这个特定问题上的技术给用上。这种进展同时既令人惊叹,又令人失望。看着这些工具运作,是一种非常奇怪的感觉。但人们也适应得极快。
便签笔记
11陶的亲身体验:更宽但不更深
46:12
I remember when Google's web search came out 20 years ago. It just blew all the other searches out of the water. You're getting relevant hits on the front page, exactly what you wanted. It was amazing, and then after a few years, you just took for granted that you could Google anything. 2026-level AI would be stunning in 2021. A lot of it—face recognition, natural speech, doing college-level math problems—we just take for granted now. Speaking of 2026 AI, you made a prediction in 2023 that by 2026 it would be like a colleague in mathematics?
我还记得 20 年前谷歌网页搜索刚出来的时候。它把其他所有搜索引擎打得毫无还手之力。首页给你的就是相关结果,正是你想要的。那太惊人了,然后过几年,你就理所当然地觉得什么都能谷歌到。2026 年水准的 AI 放在 2021 年会让人瞠目结舌。其中很多东西——人脸识别、自然语音、做大学水平的数学题——我们现在都觉得理所当然了。说到 2026 年的 AI,你在 2023 年做过一个预测:到 2026 年,它会像数学上的一位同事?
便签笔记
46:53
A trustworthy co-author if used correctly. Which is looking pretty good in retrospect. Yeah, I'm pretty pleased. So let's see if you can continue this streak. You personally are 2x more productive as a result of AI. What year would you say that? Productivity, I think, is not quite a one-dimensional quantity. I'm definitely noticing that the style in which I do mathematics is changing quite a bit, and the type of things I do. For example, my papers now have a lot more code, a lot more pictures, because it's so easy to generate these things now. Some plot which would have taken me hours to do, now I can do in minutes. But in the past, I just wouldn't have put the plot in my paper in the first place. I would just talk about it in words.
一个只要用得对就值得信赖的合作者。事后看来这个预测相当不错。是的,我挺满意的。那我们看看你能不能把这个连胜延续下去。你个人因为 AI 而生产力翻倍。你觉得会是哪一年?我觉得生产力并不是一个一维的量。我明显注意到,我做数学的方式正在发生相当大的变化,做的事情类型也在变。比如说,我现在的论文里有多得多的代码、多得多的图,因为现在生成这些东西太容易了。有些图以前要花我好几个小时,现在几分钟就能做出来。但在过去,我根本就不会把那张图放进论文里,我只会用文字描述一下。
便签笔记
47:41
So it's hard to measure what 2x means. On the one hand, I think the type of papers that I would write today, if I had to do them without AI assistance, would definitely take five times longer. But I would not write my papers that way. 5x? Yeah, but these are auxiliary tasks. Things like doing a much deeper literature search or supplying a lot more numerics. They enrich the paper. The core of what I do, actually solving the most difficult part of a math problem, hasn't changed too much. I still use pen and paper for that.
所以很难衡量「2 倍」意味着什么。一方面,我今天写的这类论文,如果不用 AI 辅助,肯定要多花五倍时间。但我根本不会那样去写论文。5 倍?是的,但这些都是辅助性的工作。比如做深得多的文献检索,或者提供多得多的数值计算。它们让论文更丰富。而我工作的核心——真正去攻克一个数学问题里最难的那部分——并没有太大变化。那部分我还是用纸笔。
便签笔记
48:28
But there's lots of silly things. I use an AI agent now to reformat. Sometimes if all my parentheses are not quite the right size, I used to manually change them by hand, and now I can get an AI agent to do all that quite nicely in the background. They've really sped up lots of secondary tasks. They haven't yet sped up the core thing that I do, but it's allowed me to add more things to my papers. By the same token, if I were to write a paper I wrote in 2020 again—and not add all these extra features, but just have something of the same level of functionality—it actually hasn't saved that much time, to be honest. It's made the papers richer and broader, but not necessarily deeper. You made this distinction between artificial cleverness and artificial intelligence. I would like to better understand those concepts.
但有很多琐碎的小事。我现在会用 AI agent 来重新排版。有时候我的括号大小不太对,以前我得手动一个个改,现在我可以让 AI agent 在后台把这些活儿漂亮地干完。它们确实极大加速了很多次要任务。它们还没有加速我工作的核心部分,但它们让我能在论文里加进更多东西。同样地,如果让我重写一篇我 2020 年写的论文——不加这些额外的东西,只做出功能水平相同的成果——说实话,其实并没有省下多少时间。它让论文更丰富、更宽广,但不见得更深刻。你区分过「人工机巧」和「人工智能」。我想更好地理解这两个概念。
便签笔记
49:29
What is an example of intelligence that is not just cleverness?
有什么例子是属于智能、而不只是机巧的?
便签笔记
49:40
Intelligence is famously hard to define. It's one of these things that you know when you see it. But when I talk to someone and we're trying to collaboratively solve a math problem together, there's this conversation where neither of us knows how to solve the problem initially. One of us has some idea and it looks promising, so then we have some sort of prototype strategy. We test it, and it doesn't work, but then we modify it. There's adaptivity and continual improvement of the idea over time. Eventually, we've systematically mapped out what doesn't work and what does work, and we can see a path forward, but it's evolving with our discussion.
智能是出了名的难以定义。它属于那种你一看到就知道的东西。但当我和某个人交谈、我们试着一起协作解决一个数学问题时,会有这样一段对话:一开始我们俩都不知道怎么解这个问题。其中一人有个想法,看起来有希望,于是我们就有了某种原型策略。我们去检验它,发现行不通,然后我们再修改它。这里面有适应性,想法会随着时间不断改进。最终,我们系统性地摸清了什么行不通、什么行得通,然后我们能看到一条前进的路径,而这条路径是随着我们的讨论演化出来的。
便签笔记
50:30
This isn't quite what the AIs do. The AIs can mimic this a little bit. To go back to this analogy of these jumping robots, they can jump and fail, and jump and fail. But what they can't do is jump a little bit, reach some handhold, stay there, pull other people up, and then try to jump from there. There isn't this cumulative process which is built up interactively. It seems to be a lot more trial and error and just repetition: brute force. It scales, and it can work amazingly well in certain contexts. But this idea of building up cumulatively from partial progress is what's still not quite there yet.
这并不完全是 AI 在做的事。AI 可以稍微模仿一下这个过程。回到那些会跳跃的机器人的类比,它们可以跳、失败,再跳、再失败。但它们做不到的是:跳一小步,抓住某个着力点,停在那里,再把别人拉上来,然后从那里继续往上跳。它们身上没有那种通过互动逐步累积起来的过程。看起来更多是反复试错、不断重复:暴力穷举。它能扩展,在某些场景下效果好得惊人。但这种从局部进展一点点累积起来的能力,目前还差点意思。
便签笔记
51:23
Interesting. You're saying if Gemini 3 or Claude 4.5, whatever, solves a problem, it is not the case that its own understanding of math has progressed. No. Or even if it works on a problem without solving it, it's not that its own understanding of math has progressed. Yeah. You run a new session and it's forgotten what it just did. It has no new skills to build on related problems. Maybe what you just did is 0.001% of the training data for the next generation. So maybe eventually some of it gets absorbed.
有意思。你是说,如果 Gemini 3 或者 Claude 4.5,不管哪个,解决了一个问题,并不意味着它自己对数学的理解有了长进。对。甚至它花时间钻研一个问题却没解出来,也不意味着它自己对数学的理解有了长进。是啊。你开一个新会话,它就把刚才做过的事忘了。它没有获得任何可以用在相关问题上的新技能。也许你刚做的这些会成为下一代模型训练数据里的百分之零点零零一。所以也许最终有一部分会被吸收进去。
便签笔记
51:54
So Terence talks about the importance of decomposing particularly gnarly problems into a series of easier chunks. Even if this doesn't result in the full solution, approaching problems in this way helps you build up the intuitions and practice the techniques that you'll need to keep making progress. But models today tend to struggle with these kinds of problem-solving techniques. That's where Labelbox comes in. Labelbox helps you train models not just to get the right answer, but to think the right way.
陶哲轩谈到,把特别棘手的问题拆解成一系列更容易的小块非常重要。哪怕这样做不能得出完整的解,用这种方式去接近问题,也能帮你建立直觉、练习那些持续取得进展所需要的技巧。但今天的模型往往在这类解题技巧上表现吃力。这正是 Labelbox 的用武之地。Labelbox 帮你训练模型,不只是让它得出正确答案,而是让它用正确的方式思考。
便签笔记
52:19
They've operationalized these reasoning behaviors into rubrics, giving you the ability to evaluate every important dimension of a model's output. These rubrics go beyond simple correctness. Did the model reach for the right tools? Did it check its own work and explore alternative paths? How clear was its response? These skills are useful across domains: math, physics, finance, psychology, and more. And they're becoming increasingly important as models take on harder, open-ended problems, some of which have multiple solutions and some of which we don't even know the solutions to.
他们把这些推理行为落地成了评分量表(rubric),让你能够评估模型输出的每一个重要维度。这些量表远不只是看对错。模型有没有选用正确的工具?有没有检查自己的工作、探索别的路径?它的回答清晰度如何?这些能力在各个领域都有用:数学、物理、金融、心理学等等。而且随着模型开始处理更难、更开放的问题,这些能力正变得越来越重要,其中有些问题有多个解,有些我们自己都还不知道答案。
便签笔记
12Lean 证明、消融与策略语言
52:48
Labelbox can get you rubrics tailored to your domain, helping you systematically measure and shape how your models think. Learn more at labelbox.com/dwarkesh. One big question I have is how plausible is it that if we just keep training AIs—they get better and better at solving problems in Lean—that they will continue to solve more and more impressive problems, and then we will be surprised at how little insight we got from some Lean solution to proving the Riemann hypothesis or something. Or do you think it is a necessary condition of solving the Riemann hypothesis, even by an AI that is doing it entirely in Lean, that the constructions and definitions created in the Lean program have to advance our understanding of mathematics? Or could it just be assembly code gobbledygook?
Labelbox 可以为你的领域量身定制评分量表,帮你系统地衡量并塑造模型的思考方式。了解更多请访问 labelbox.com/dwarkesh。我有个很大的疑问:如果我们只是不断训练 AI——让它们在 Lean 里解题越来越强——它们有多大可能会持续解出越来越惊人的问题,然后我们会惊讶地发现,从某个证明黎曼猜想之类的 Lean 解答里,我们几乎没得到什么洞见?还是说你认为,哪怕是一个完全在 Lean 里工作的 AI,要解决黎曼猜想,有个必要条件,就是那个 Lean 程序里创造出的构造和定义必须推进我们对数学的理解?还是说它可能就是一堆汇编代码般的天书?
便签笔记
53:40
We don't know. Some problems have been basically solved by pure brute force. The four color theorem is a famous example. We have still not found a conceptually elegant proof of this theorem, and maybe we never will. Some problems may only be solvable by splitting into an enormous number of cases and doing brute force, uninsightful computer analysis on each case. Part of the reason we prize problems like the Riemann hypothesis is that we're pretty sure a new type of mathematics has to be created, or a new connection between two previously unconnected areas of mathematics has to be discovered to make this work. We don't even know what the shape of the solution is, but it doesn't feel like a problem that will be solved just by exhaustively checking cases.
我们不知道。有些问题基本上就是靠纯暴力解决的。四色定理就是个著名例子。我们至今还没找到这个定理在概念上优雅的证明,也许永远都找不到。有些问题也许只能靠拆分成极其大量的情形,然后对每种情形做毫无洞见的暴力计算机分析来解决。我们之所以看重像黎曼猜想这样的问题,部分原因在于我们相当确信,要做成这件事,必须创造出一种新型的数学,或者必须发现两个此前毫无关联的数学分支之间的新联系。我们甚至不知道解的形态是什么样,但它感觉不像是靠穷举检验情形就能解决的问题。
便签笔记
54:30
Or it could be false actually. Okay, there is an unlikely scenario that the hypothesis is false, and you can just compute a zero off the line, and a massive computer calculation verifies it. That would be very disappointing. I do feel that fully autonomous, one-shot approaches are not the right approach for these problems. You'll get a lot more mileage out of the interplay of humans collaborating with these tools. I can see one of these problems being solved by smart humans assisted by extremely powerful AI tools. But the exact dynamic may be very different from what we envision right now. It could be a collaboration of a type that just doesn't exist yet. There may be a way to generate a million variants of the Riemann zeta function and do AI-assisted data analysis to discover some pattern connecting them that we didn't know about before. This lets you transform the problem into a different area of mathematics. There could be all kinds of scenarios.
或者它其实也可能是错的。好吧,存在一种可能性不大的情形:这个猜想是假的,你能直接算出一个不在临界线上的零点,再用海量的计算机计算去验证。那就非常令人失望了。我确实觉得,对这类问题来说,完全自主、一步到位的做法不是正确的路子。人类与这些工具协作、相互配合,能带来大得多的收益。我能想象这类问题中的某一个,是由聪明的人类借助极其强大的 AI 工具解决的。但具体的互动方式可能跟我们现在设想的很不一样。那可能是一种目前还不存在的协作形式。也许有办法生成上百万个黎曼 zeta 函数的变体,再做 AI 辅助的数据分析,发现它们之间某种我们此前不知道的模式。这就能把问题转化到数学的另一个领域。各种可能的情形都有。
便签笔记
55:51
Suppose the AI figures it out, and latent in the Lean is some brand-new construction which, if we realized its significance, we would be able to apply in all these different situations. How would we even recognize it? Again, a very naive question, but if you come up with the equivalent of Descartes' idea that you can have a coordinate system unifying algebra and geometry, in Lean code it would just look like R→R, and it wouldn't look that significant. I'm sure there are other constructions which have this kind of property.
假设 AI 解出来了,而 Lean 代码里潜藏着某个全新的构造,如果我们意识到它的意义,就能把它应用到各种不同的场景中去。我们要怎么才能认出它来?这又是个很外行的问题,但假如你想出了相当于笛卡尔那个思想的东西——用一个坐标系把代数和几何统一起来——在 Lean 代码里它可能就长成 R→R 的样子,看上去一点都不起眼。我相信还有别的构造也具有这种性质。
便签笔记
56:26
The beauty of formalizing a proof in something like Lean is that you can take any piece of it and study it atomically. When I read a paper which solves some difficult problem, there's often a big sequence of lemmas and theorems. Ideally, the author will talk their way through what's important and what's not. But sometimes they don't reveal what steps were the important ones and which ones were just boilerplate, standard steps. You can study each lemma in isolation. Some of them I can see look fairly standard and resemble something I'm familiar with.
把证明形式化成 Lean 这类东西的美妙之处在于,你可以取出其中任何一部分,单独地去研究它。当我读一篇解决了某个难题的论文时,里面常常有一长串引理和定理。理想情况下,作者会讲清楚哪些重要、哪些不重要。但有时候他们并不点明哪些步骤是关键的,哪些只是套路化的标准步骤。而你可以孤立地研究每一条引理。其中有些我一看就觉得挺标准的,跟我熟悉的东西很像。
便签笔记
57:04
I'm pretty sure there's nothing interesting going on there. But this other lemma, that's something I haven't seen before, and I can see why having this result would really help prove the main result. You can assess whether a step is really key to your argument or not, and Lean really facilitates that. The individual steps are identified really precisely. I think in the future, there will be entire professions of mathematicians who might take a giant Lean-generated proof and do some ablation on it, trying to remove parts of it and find more elegant ways. They might get other AIs to do some reinforcement learning to make the proof more elegant, and maybe other AIs will grade whether this proof looks better or not. One thing that will change quite a bit in the near future is how we write papers. Until recently, writing papers was the most time-consuming and expensive part of the job. So you did it very rarely.
我基本可以确定那里没什么有意思的东西。但另外那条引理,是我以前没见过的,而且我能看出为什么有了这个结果会对证明主定理帮助很大。你可以判断某一步对你的论证到底关不关键,而 Lean 极大地方便了这件事。每一个单独的步骤都被非常精确地标识出来。我认为将来会出现一整个数学家职业群体,他们拿着一个巨大的、AI 生成的 Lean 证明去做消融实验,试着删掉其中一部分,寻找更优雅的方式。他们可能会让别的 AI 做一些强化学习,把证明变得更优雅,也许还有别的 AI 来评判这个证明是不是看起来更好。在不远的将来会有很大变化的一件事,是我们写论文的方式。直到最近,写论文都还是这份工作里最耗时、代价最高的部分。所以你很少去写。
便签笔记
58:07
You only wrote up your results once all the other parts of your argument were checked out, because rewriting and refactoring was just a total pain. That's become a lot easier now with modern AI tools. You don't have to have just one version of your paper. Once you have one, people can generate hundreds more. One giant messy Lean proof may not be very meaningful or understandable on its own, but other people can refactor it and do all kinds of things with it. We've seen this with the Erdős problem website.
你只有在论证的其他所有部分都核实清楚之后才会把结果写出来,因为重写和重构实在太痛苦了。而现在有了现代 AI 工具,这变得容易多了。你的论文不必只有一个版本。一旦你有了一版,别人可以生成上百版。一个巨大而杂乱的 Lean 证明本身也许意义不大、也不好懂,但别人可以去重构它、对它做各种各样的事情。我们在 Erdős 问题网站上已经看到了这种情况。
便签笔记
58:42
An AI will generate a proof, and here are 3,000 lines of code that verify the proof. Then people got other AIs to summarize the proof, and people write their own proofs. There's actually post-processing. Once you have one proof, we have a lot of tools now to deconstruct and interpret it. It's a very nascent area of mathematics, but I'm not as worried about it. Some people are concerned about what happens if the Riemann hypothesis is proven with a completely incomprehensible proof. I think once you have the artifact of a proof, we can do a lot of analysis on it.
AI 会生成一个证明,然后是三千行验证这个证明的代码。接着人们又让别的 AI 去总结这个证明,人们也自己写自己的证明。确实存在后处理这一环。一旦你有了一个证明,我们现在有很多工具去拆解它、解读它。这是数学中一个非常新生的领域,但我对此没那么担心。有些人担心,如果黎曼猜想被一个完全无法理解的证明搞定了会怎么样。我认为,一旦你手上有了证明这个实物,我们就能对它做大量分析。
便签笔记
59:20
You posted recently that it would be helpful to have a formal or semi-formal language for mathematical strategies as opposed to just mathematical proofs, which is what Lean specializes in. I would love to learn more about what that would involve or look like. We don't really know. We've been very lucky in mathematics that we have worked out the laws of logic and mathematics, but this is a fairly recent accomplishment. It was started by Euclid two millennia ago, but only in the early 20th century did we finally list out the axioms of mathematics, the standard axioms of what we call ZFC, the axioms of first-order logic, and what a proof is.
你最近发帖说,如果能有一种针对数学策略的形式化或半形式化语言,而不只是针对数学证明,那会很有帮助,而 Lean 擅长的正是后者。我很想多了解一下那会涉及什么、长成什么样。我们其实并不清楚。我们在数学上非常幸运,把逻辑和数学的法则梳理清楚了,但这其实是相当晚近的成就。它由欧几里得在两千年前开了个头,但直到二十世纪初,我们才终于把数学的公理列了出来,也就是我们称为 ZFC 的标准公理、一阶逻辑的公理,以及证明到底是什么。
便签笔记
60:00
This we've managed to automate and have a formal language for. There could be some way to assess plausibility. You have a conjecture that something is true, you test a few examples, and it works out. How does this increase your confidence that the conjecture is true? We have a few sort of mathematical ways to model this, like Bayesian probability, for example. But you often have to set certain base assumptions, and there's a lot of subjectivity still in these tasks.
这些我们已经能够自动化,并且有了形式语言。也许有某种办法可以评估「合理性」。你有一个猜想,认为某件事是真的,你试了几个例子,都对得上。这在多大程度上提升了你对这个猜想为真的信心?我们有一些数学上的方式来给这件事建模,比如贝叶斯概率。但你往往必须设定某些基础假设,这类工作里仍然有大量主观成分。
便签笔记
60:44
This is more of a wish than a plan to develop these languages, but just seeing how successful having a formal framework in place, like Lean, has made deductive proofs so much easier to automate and train AI on… The bottleneck for using AI to create strategies and make conjectures is we have to rely on human experts and the test of time to validate whether something is plausible or not. If there was some semi-formal framework where this could be done semi-automatically in a way that isn't easily hackable... It's really important with these formal proof assistants that there are no backdoors or exploits you can use to somehow get your certified proof without actually proving it, because reinforcement learning is just so good at finding these backdoors. If there's some framework that mimics how scientists talk to each other in a semi-formal way, using data and argument, but also constructing narratives... There's some subjective aspect of science that we don't know how to capture in a way that we can insert AI into it in any useful way.
这更像是一个愿望,而不是开发这类语言的计划,只是看到有一个像 Lean 这样的形式框架有多成功——它让演绎式的证明变得如此容易自动化,也容易拿来训练 AI……用 AI 来创造策略、提出猜想,瓶颈在于我们只能依赖人类专家和时间的检验来判断某个东西是否说得通。如果有某种半形式化的框架,能以一种不容易被钻空子的方式半自动地做这件事……在这些形式化的证明助手里,非常重要的一点是不能有后门或者可以利用的漏洞,让你在没有真正证明的情况下拿到认证过的证明,因为强化学习实在太擅长找到这类后门了。如果有某种框架能模拟科学家之间那种半形式化的交流方式,既用数据和论证,又构建叙事……科学里有某种主观的成分,我们还不知道该怎么把它捕捉下来,好让 AI 能以某种有用的方式介入进来。
便签笔记
62:18
This is a future problem. There are research efforts to try to create automated conjectures, and maybe there are ways to benchmark these and simulate this, but it's all very new science. Can you help me get some intuition? I have two sub-questions. One, it would be very helpful to have a specific example of what something like this would look like, the way scientists communicate that we can't formalize yet. Two, it seems almost definitionally paradoxical to say you're building up some narrative or natural language explanation and then also having something which you could have formalized. I'm sure there's some intuition behind where that overlap is, and I'd love to understand that better.
这是个未来的问题。已经有一些研究在尝试做自动化的猜想生成,也许有办法给这些做基准测试、做模拟,但这门科学还非常新。你能帮我建立一点直觉吗?我有两个子问题。第一,如果能有一个具体例子会很有帮助——科学家那种我们还没法形式化的交流方式到底长什么样。第二,说你在构建某种叙事或自然语言解释,同时又要有某种本可以被形式化的东西,这几乎从定义上就自相矛盾。我相信这两者的重叠之处背后有某种直觉,我很想更好地理解它。
便签笔记
13素数随机模型与黎曼猜想
63:20
An example of a conjecture: Gauss was interested in the prime numbers and created one of the first mathematical datasets. He just computed the first 100,000 prime numbers or so, hoping to find patterns. He did find a pattern, but maybe not the pattern he was expecting. He found a statistical pattern in the primes that if you count how many primes there are up to 100, 1,000, one million, and so forth, they get sparser and sparser, but the drop-off in the density was inversely proportional to the natural logarithm of the range of numbers. So he conjectured what we now call the prime number theorem: the number of primes up to X is X divided by the natural log of X.
举个猜想的例子:高斯对素数很感兴趣,他制作了最早的数学数据集之一。他就是把前十万个左右的素数算了出来,希望能找到规律。他确实找到了一个规律,但也许不是他原本期待的那个。他在素数中发现了一个统计规律:如果你数一数一百以内、一千以内、一百万以内等等有多少个素数,素数会越来越稀疏,但这个密度下降的幅度与数值范围的自然对数成反比。于是他提出了我们今天所说的素数定理这个猜想:不超过 X 的素数个数是 X 除以 X 的自然对数。
便签笔记
64:05
He had no way to prove this. It was data-driven. This was a conjecture. It was revolutionary for its time because it was maybe the first really important conjecture of math that was statistical in nature. Normally you're talking about a pattern, like maybe the spacing between the primes has a certain regularity. But this didn't tell you exactly how many primes there were in any given range. It just gave you an approximation that got better and better as you went further and further out.
他没有办法证明它。这是数据驱动的。这是一个猜想。它在当时是革命性的,因为它也许是数学史上第一个真正重要的、本质上带统计性的猜想。通常你谈的是某种模式,比如素数之间的间隔可能有某种规律性。但这个猜想并没有告诉你在任何给定区间里到底有多少个素数。它只是给了你一个近似值,而且你走得越远,这个近似就越准。
便签笔记
64:42
It started the field of what we call analytic number theory. It was the first in many conjectures like this, many of which got proved, which started consolidating the idea that the prime numbers didn't really have a pattern, that they behaved like random sets of numbers with a certain density. They had some patterns, like they're almost all odd. They're also not actually random, they're what's called pseudo-random. There's no random number generation involved in creating the prime numbers. But over time, it became more and more productive to think of the primes as if they were just generated by some god rolling dice all the time and creating this random set.
它开创了我们所说的解析数论这个领域。它是许多类似猜想中的第一个,其中不少后来被证明了,这些工作逐渐巩固了一个观念:素数其实并没有什么规律,它们的行为就像具有某种密度的随机数集合。它们确实有一些模式,比如几乎全都是奇数。它们其实也不是真的随机,而是所谓的伪随机。素数的产生过程里并不涉及任何随机数生成。但随着时间推移,把素数想象成某个神在不停掷骰子、造出这么一个随机集合,这种思路变得越来越有成效。
便签笔记
65:26
This allowed us to make all these other predictions. There's a still-open conjecture in number theory called the twin prime conjecture, that there should be infinitely many pairs of primes that are twins just two apart, like 11 and 13. We can't prove that, and there are good reasons why we can't prove it. But because of this statistical random model of the primes, we are absolutely convinced it's true. We know that if the primes were generated by flipping coins, we would just—by random chance like infinite monkeys at a typewriter—see twin primes appear over and over again.
这让我们能做出各种别的预测。数论里有一个至今未解的猜想,叫孪生素数猜想:应该存在无穷多对只相差 2 的孪生素数,比如 11 和 13。我们证明不了它,而且有很好的理由说明我们为什么证明不了。但由于素数的这个统计随机模型,我们绝对确信它是真的。我们知道,如果素数是靠掷硬币生成的,那么纯粹出于随机——就像无限只猴子敲打字机那样——我们会一次又一次地看到孪生素数出现。
便签笔记
65:57
We have over time developed this very accurate conceptual model of what the primes should behave like based on statistics and probability. It's mostly heuristic and non-rigorous, but extremely accurate. The few times when we actually can prove things about the primes, it has matched up with the predictions of what we call the random model of the primes. We have this conjectural concept framework for understanding the primes that everyone believes in. It's the same reason why we believe the Riemann hypothesis is true, and why we believe that cryptography based on the primes is mathematically secure.
随着时间推移,我们基于统计和概率,发展出了一个关于素数应该如何表现的、非常准确的概念模型。它基本上是启发式的、不严格的,但极其准确。在少数几次我们真的能证明素数相关结论的时候,结果都与所谓素数随机模型的预测吻合。我们对素数有了这样一套猜想性的概念框架,而且大家都相信它。这也是我们相信黎曼猜想为真的原因,是我们相信基于素数的密码学在数学上是安全的原因。
便签笔记
66:35
It's all part of this belief. In fact, one reason why we care about the Riemann hypothesis is that if the Riemann hypothesis failed, if we knew it was false, it would be a serious blow to this model. It would mean there's a secret pattern to the primes that we were not aware of. I think we would very rapidly abandon any cryptography based on the primes, because if there was one pattern that we didn't know about, there are probably more, and these patterns can lead to exploits in crypto. It would be a big shock.
这全都是这套信念的一部分。事实上,我们之所以在乎黎曼猜想,一个原因就是如果黎曼猜想不成立,如果我们知道它是假的,那将对这个模型是一次沉重打击。那意味着素数里存在某种我们此前没有察觉的隐秘模式。我想我们会非常迅速地放弃所有基于素数的密码学,因为如果有一个我们不知道的模式,那很可能还有更多,而这些模式可能导致密码系统被攻破。那会是一次很大的冲击。
便签笔记
67:07
So we really want to make sure that doesn't happen.
所以我们真的很希望那种情况别发生。
便签笔记
67:15
We've been convinced of things like the Riemann hypothesis over time. Some of it is experimental evidence, and some is that the few times we've been able to make theoretical results, they've always aligned. It is possible that the consensus is wrong and we've all just missed something very basic. There have been paradigm shifts in the past in scientific history. But we don't really have a way of measuring this, partly because we don't have enough data on how math or science develops. We have one timeline of history, and we have maybe 100 stories of turning points in history.
随着时间推移,我们对黎曼猜想这类事情逐渐深信不疑。一部分来自实验证据,一部分是因为在我们少数几次能拿到理论结果的时候,它们总是吻合的。当然,共识也有可能是错的,我们大家可能只是错过了某些非常基础的东西。科学史上过去也发生过范式转变。但我们其实没有办法去衡量这件事,部分原因是我们没有足够的数据来了解数学或科学是如何发展的。我们只有一条历史时间线,也许还有一百来个关于历史转折点的故事。
便签笔记
67:53
If we had access to a million alien civilizations, each with a different development of history and science in different orders, then maybe we'd actually have a decent shot at understanding how we measure what progress is and what is a good strategy. We could maybe start formalizing it and actually having a framework. Maybe what we need to do is start creating lots of mini-universes or simulations of AI solving very basic problems in arithmetic or whatever, but coming up with their own strategies for doing these things and having these little laboratories to test.
如果我们能接触到一百万个外星文明,每个文明的历史和科学都以不同的顺序发展,那也许我们才真的有机会去理解该怎么衡量什么算是进步、什么算是好的策略。我们也许可以开始把它形式化,真正建立起一套框架。也许我们该做的,是开始创造大量的小宇宙,或者说模拟环境,让 AI 去解决算术之类非常基础的问题,但让它们自己想出做这些事情的策略,用这些小实验室来做测试。
便签笔记
68:34
There are people who investigate what's the smallest neural network that can do 10-digit multiplication and things like that. I think we could learn a lot just from evolving small AIs on simple problems. I was super excited when Mercury reached out about sponsoring the podcast because I've been banking with them for years. I think I opened my first account with them in 2023. Something I've come to appreciate over the last few years is that Mercury is constantly updating things and adding new features. Take their newest feature, Insights.
有人在研究,能做十位数乘法的最小神经网络是什么,诸如此类。我觉得光是让小型 AI 在简单问题上演化,我们就能学到很多。当 Mercury 联系我说要赞助这档播客时,我特别兴奋,因为我用他们的银行服务已经好多年了。我记得我是 2023 年在他们那儿开的第一个账户。过去这几年我越来越欣赏的一点是,Mercury 一直在更新、不断加新功能。就拿他们最新的功能 Insights 来说。
便签笔记
69:03
Insights summarizes your money in and out, showing you your biggest transactions and calling out anything that deserves extra attention. Like maybe your revenue from a particular partner has gone down, or you've got a big uncategorized purchase that needs to be investigated. It's a super low-friction way for me to keep tabs on my business and make quick decisions. For example, I try to invest any cash that I don't need on hand to keep running the business. With Insights, with just a couple of clicks, I was able to see exactly how much money I spent in each month of 2025 and that lets me know exactly how much cash I'll need for the next year or so of operations. And then I can go invest the rest.
Insights 会汇总你的资金进出,展示你最大的几笔交易,并把任何值得额外留意的地方标出来。比如某个特定合作伙伴带来的收入下降了,或者你有一笔金额很大、还没归类的支出需要去查一查。对我来说,这是一种非常省事的方式,可以随时掌握公司状况、快速做决定。举个例子,我会尽量把手头不需要留着维持公司运转的现金拿去投资。有了 Insights,只要点几下,我就能准确看到 2025 年每个月花了多少钱,这让我很清楚地知道接下来一年左右的运营需要多少现金。剩下的我就可以拿去投资了。
便签笔记
14狐狸式学习与机缘巧合
69:35
Mercury just keeps adding new features like this. Go to mercury.com to check it out. Mercury is a fintech company, not an FDIC Insured Bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. You have to learn about new fields not only very rapidly, but deeply enough to contribute to the frontier. So in some sense, you're also one of the world's greatest autodidacts. What is your process of learning about a new subfield in math? What does that look like? We talked about depth and breadth before.
Mercury 就是这样不断增加新功能。去 mercury.com 看看吧。Mercury 是一家金融科技公司,不是 FDIC 承保的银行。银行服务由 Choice Financial Group 和 Column N.A. 提供,二者均为 FDIC 成员。你必须学习新领域,不仅要非常快,而且要深入到能够为前沿做出贡献。所以从某种意义上说,你也是世界上最厉害的自学者之一。你学习一个数学新分支的流程是怎样的?那具体是什么样子?我们之前聊过深度和广度。
便签笔记
70:12
It's not a purely human-AI distinction. Humans also, I think it was Berlin who split them into hedgehogs and foxes. The hedgehog knows one thing very well, and a fox knows a little bit about everything. I definitely think of myself as a fox. I work with hedgehogs a lot, and sometimes I can be a hedgehog if need be.
这并不纯粹是人和 AI 的区别。人也一样,我记得是伯林把人分成刺猬和狐狸。刺猬把一件事知道得很透,而狐狸对什么都懂一点。我绝对把自己看作狐狸。我经常和刺猬型的人合作,必要的时候我也能当刺猬。
便签笔记
70:40
I've always had a little bit of an obsessive streak. If there's something I read about which I feel like I have the capability to understand, but I don't understand why it works and there's some magic in it… Someone was able to use a type of mathematics I'm not familiar with and get a result I would like to prove. I can't do it myself, but they could do it by their method, and I want to find out what their trick was. It bugs me that someone else can do something I think I can do, but I can't. I've always had that obsessive, completionist streak. I've had to wean myself off computer games because if I start a game, I want to play it to completion, through all the levels.
我一直都有点强迫的倾向。如果我读到某样东西,觉得自己有能力理解它,但又不明白它为什么行得通,里面好像有某种魔法……有人用了一种我不熟悉的数学,得到了一个我也想证明的结果。我自己做不到,但他们用他们的方法做到了,我就想弄清楚他们的诀窍是什么。让我难受的是,别人能做到某件我觉得自己也能做的事,可我却做不到。我一直都有那种强迫式的、追求完成度的性格。我不得不戒掉电子游戏,因为只要我开始玩一个游戏,我就想把它通关,把所有关卡都打完。
便签笔记
71:23
That's one way I learn new fields. I collaborate with a lot of people who have taught me other types of mathematics. I just make friends with another mathematician working on another area of mathematics. I find their problems interesting, but they have to teach me some of the basic tricks, what's known, and what's not known. I learn a lot from that. I found that writing about what I've learned helps. I have a blog where I sometimes record things I've learned. In the past when I was younger, I would learn something, do this cool trick, and say, "Okay, I'm going to remember this."
这是我学习新领域的一种方式。我和很多人合作,他们教会了我别的数学分支。我就是去和另一位数学家交朋友,他研究的是另一个数学领域。我觉得他们的问题很有意思,但他们得教我一些基本的技巧,教我哪些是已知的、哪些还是未知的。我从中学到很多。我还发现,把学到的东西写下来很有帮助。我有个博客,有时会把学到的东西记下来。以前我年轻的时候,我会学到点什么,做出某个很酷的技巧,然后说:“好,我要把这个记住。”
便签笔记
72:02
Then six months later, I'd forgotten it. I remember remembering it, but I can't reconstruct my arguments. The first few times, it was so frustrating to have understood something and then lost it. I resolved I should always write down anything cool that I've learned. That's part of how this blog came about. How long does it take you to write a blog post? It's something I often do when I don't want to do other work. There's some referee report or something that feels slightly unpleasant for me to do at the time.
然后六个月之后,我就把它忘了。我记得自己曾经记得,但我重构不出当时的论证。头几次,理解了一样东西然后又把它弄丢了,那种感觉特别让人沮丧。于是我下定决心,一定要把任何我学到的很酷的东西都写下来。这也是这个博客的由来之一。你写一篇博客要花多长时间?这往往是我不想做别的工作时会做的事。比如有个审稿报告之类的,当下做起来让我觉得有点不太愉快。
便签笔记
72:35
Writing a blog feels creative and fun. It's something I do for myself. Depending on the topic, it could be a quick half an hour or several hours. Because it's something I do voluntarily, time flies when I write these things down, as opposed to doing something I have to do for administrative reasons that is just drudgery. Those are tasks, by the way, that AI is really helping with nowadays. If civilization could from first principles decide how to use Terry Tao's time, as a limited resource, what is the biggest difference? What if the veil of ignorance got to decide how to use Terry Tao's time versus what it does now? This podcast wouldn't be happening.
写博客感觉有创造性、也很有意思。那是我为自己做的事。取决于题目,可能半小时就写完,也可能要好几个小时。因为这是我自愿做的事,写这些东西的时候时间过得飞快;而做那些出于行政原因不得不做的事时,那就纯粹是苦差事。顺便说一句,那类任务现在 AI 帮了很大的忙。如果整个文明可以从第一性原理出发,把陶哲轩的时间当作一种有限资源来决定怎么使用,最大的区别会是什么?如果由无知之幕来决定怎么使用陶哲轩的时间,和现在的做法相比会如何?那这档播客就不会存在了。
便签笔记
73:29
As much as I complain about certain tasks that I don't want to do, but have to do… As you get more senior in academia, you get more and more responsibilities, more committees, and whatever. I have also found that a lot of events I reluctantly went to because I was obliged to for one reason or another… Because it's outside my comfort zone, it often results in interactions with people I wouldn't normally talk to, like you for instance. I would learn interesting things and have interesting experiences.
尽管我总抱怨某些我不想做、但又不得不做的事……在学术界你越资深,责任就越多,委员会越多,等等。但我也发现,很多我因为这样那样的原因不情愿去参加的活动……正因为它超出了我的舒适区,往往会带来和我平时不会交谈的人的互动,比如和你。我会学到有趣的东西,会有有趣的经历。
便签笔记
74:01
I would have opportunities to then network with other people that I never would have before. So I do believe a lot in serendipity. I do optimize portions of my day where I schedule very carefully. But I am willing to leave some portions just to do something that is not my usual thing. Maybe it'll be a waste of my time, but maybe I will learn something. More often than not, I get a positive experience that I wouldn't have planned for. So I believe a lot in serendipity. Maybe there's a danger in modern societies, not just with AI, that we've become really good at optimizing everything. We’re not optimizing our own optimization.
我会有机会去认识一些原本永远不会认识的人。所以我很相信机缘巧合。我确实会优化一天中的某些时段,把它们安排得非常仔细。但我也愿意留出一部分时间,去做一些不是我平常会做的事。也许那会浪费我的时间,但也许我会学到点什么。多数情况下,我都会获得一些计划之外的正面体验。所以我很相信机缘。也许现代社会存在一种危险,不只是 AI 带来的:我们变得太擅长优化一切了。我们没有去优化我们的优化本身。
便签笔记
74:59
With COVID, for example, we switched a lot to remote meetings, so everything was scheduled. We kept busy in academia. We met almost the same number of people we met in person, but everything had to be planned in advance. What we lost out on was the casual knocking on a hallway door, just meeting someone while getting a coffee. Those serendipitous interactions may not seem optimal, but they are actually really important. When I was a grad student, I would go to the library to look for a journal article.
比如疫情期间,我们大量转向线上会议,所以一切都是事先安排好的。我们在学术界还是很忙。我们见到的人几乎和线下时一样多,但一切都必须提前计划。我们失去的,是那种随意敲开走廊上某扇门、或者去倒咖啡时偶遇某人的机会。那些机缘巧合的互动看起来不够优化,但其实非常重要。我读研究生的时候,会去图书馆找某篇期刊论文。
便签笔记
75:42
You had to physically check out the journal and read the article. You could browse through and sometimes the next article was also interesting. Sometimes it wasn’t, but you could accidentally find interesting things. That has basically been lost now. If you want to access an article, you just type it into a search engine or an AI, and you get exactly what you want instantly. But you don't get the accidental things you might have found if you'd done it more inefficiently.
你得把那本期刊实体借出来,然后读那篇文章。你可以顺手翻一翻,有时候下一篇文章也很有意思。有时候没意思,但你可能会意外发现有趣的东西。现在这基本上消失了。如果你想看一篇文章,你只要在搜索引擎或者 AI 里输进去,立刻就能得到你想要的。但你不会再遇到那些如果用低效方式去找、才可能碰上的意外收获。
便签笔记
76:20
I spent a year once at the Institute for Advanced Study, which is a great place with no distractions. You're there just to do research. The first few weeks you're there, it's great. You're getting all these papers written up that you've been wanting to do for a long time. You think about problems for blocks of hours at a time. But I find if I stay there for more than several months, I run out of inspiration. I get bored. I surf the internet a lot more. You actually do need a certain level of distraction in your life.
我曾经在普林斯顿高等研究院待过一年,那是个很棒的地方,没有任何干扰。你在那儿就是纯粹做研究。刚去的头几周非常棒。你把一直想写、拖了很久的论文都写出来了。你可以一连好几个小时集中思考问题。但我发现,如果我在那儿待上几个月以上,灵感就会枯竭。我会觉得无聊,会上网上得更多。其实你的生活里确实需要一定程度的干扰。
便签笔记
15AI 何时取代数学家
76:50
It adds enough randomness and high temperature. I don't know the optimal way to schedule my life. It just seems to work. I'm very curious when you expect AIs that can actually do frontier math at least as well as the best human mathematicians. In some ways, they're already doing frontier math that is super intelligent that humans can't do, but it's a different frontier from what we're used to. You could argue that calculators were doing frontier math that humans could not accomplish, but it was number crunching. But replacing Terry Tao completely.
它增加了足够的随机性,提高了“温度”。我不知道安排人生的最优方式是什么。它似乎就是管用。我很好奇你预计 AI 什么时候能真正做出至少和最优秀的人类数学家一样好的前沿数学。在某些方面,它们已经在做人类做不到的、超级智能级别的前沿数学了,但那是和我们习惯的不一样的前沿。你也可以说,当年的计算器也在做人类无法完成的前沿数学,但那只是数字运算。但要完全取代陶哲轩。
便签笔记
77:38
I mean, what do you want me for? You'll just go on all the podcasts after. It might not be the right question to ask. I think within a decade, a lot of things that math students currently do—what we spend the bulk of our time doing and a lot of stuff we put in our papers today—can be done by AI. But we will find that that actually wasn't the most important part of what we do. A hundred years ago, a lot of mathematicians were just solving differential equations. Physicists needed some exact solution to some system, and they hired a mathematician to laboriously go through the calculus and work out the solution to this fluid equation, whatever. A lot of what a 19th-century mathematician would do, you could make a call to Mathematica, Wolfram Alpha, a computer algebra package, or now more recently to an AI, and it would just solve the problem in a few minutes. But we moved on. We worked on different types of problems after that. Once computers came along—computers used to be human. People used to laboriously
我是说,那你还要我干嘛?之后所有播客都你去上了。这可能不是个合适的问题。我觉得十年之内,数学系学生现在做的很多事——我们花掉大部分时间去做的事,以及我们今天写进论文里的很多内容——都可以由 AI 来完成。但我们会发现,那其实并不是我们工作中最重要的部分。一百年前,很多数学家就只是在解微分方程。物理学家需要某个系统的精确解,他们就雇一个数学家,一步步费力地做微积分,算出某个流体方程的解之类的。19 世纪数学家做的很多事情,你现在只要调用 Mathematica、Wolfram Alpha、某个计算机代数软件包,或者更近一点调用 AI,它几分钟就把问题解决了。但我们往前走了。之后我们去研究不同类型的问题。计算机出现之后——computer 这个词以前指的是人。人们过去要费力地编制对数表,像高斯那样手算素数,而这一切现在都外包
便签笔记
79:01
create log tables and work out primes as Gauss did, and that has all been outsourced to computers. But we moved on. In genetics, to sequence the genome of a single organism, that was an entire PhD of a geneticist, carefully separating all the chromosomes and whatever. Now you can just spend $1,000 and send it to a sequencer and get it done. But genetics is not dead as a subject. You move to a different scale. Maybe you study whole ecosystems rather than individuals. I take your point but when is most mathematical progress, or almost all mathematical progress, happening by AI? If you find out this year a Millennium Prize Problem has been solved, you would put 95% odds that an AI did it autonomously.
给计算机了。但我们往前走了。在遗传学里,测出一个生物体的基因组序列,曾经是一位遗传学家整整一个博士学位的工作量,要小心翼翼地把所有染色体分离出来等等。现在你只要花一千美元送去测序,就搞定了。但遗传学并没有因此死掉。你只是转向了另一个尺度。也许你研究的是整个生态系统,而不是单个个体。我明白你的意思,但大部分数学进展、甚至几乎全部数学进展,什么时候会由 AI 做出来?如果某一年你听说一个千禧年大奖难题被解决了,你会给出 95% 的概率认为是 AI 自主完成的。
便签笔记
79:49
Surely there will be such a year. I guess I do believe that hybrid human plus AIs will dominate mathematics for a lot longer. It will depend. It will require some additional breakthroughs beyond what we already have, so it's going to be stochastic. I think AIs currently are very good at certain things, but really terrible at others. While you can add more and more frameworks on top to reduce the error rates and make them work with each other a bit more, it feels like we don't have all the ingredients to really have a truly satisfactory replacement for all intellectual tasks. It is complementary currently. It's not a replacement. Because current level AIs will accelerate science in so many ways, hopefully new discoveries and new breakthroughs will happen more quickly.
这样的一年肯定会到来。我想我确实相信,人类加 AI 的混合模式还会在数学中主导相当长一段时间。这要看情况。它需要在我们现有的基础上再有一些额外的突破,所以这件事是随机的。我认为现在的 AI 在某些事情上非常擅长,但在另一些事情上真的很糟糕。虽然你可以在上面叠加越来越多的框架来降低错误率、让它们之间配合得更好一些,但感觉我们还没凑齐所有要素,做不出一个真正令人满意的、能替代所有智力任务的东西。目前它是互补的,不是替代。因为现有水平的 AI 会在很多方面加速科学,希望新的发现和新的突破能更快出现。
便签笔记
81:01
It's also possible that by destroying serendipity we actually inhibit certain types of progress. Anything is possible at this point. I think the world is very, very unpredictable at this point in time. What is your advice to somebody who would consider a career in math or is early in a career in math, especially in light of AI progress? How should they be thinking about their career differently, if at all, as a result of AI progress? We live in a time of change. As I said, we live in a particularly unpredictable era.
但也有可能,正因为破坏了机缘巧合,我们反而抑制了某些类型的进步。此刻什么都有可能。我觉得这个世界在当下非常、非常难以预测。对于正在考虑从事数学、或者刚开始数学生涯的人,你有什么建议?尤其是在 AI 不断进步的背景下。面对 AI 的进步,他们该如何用不同的方式来思考自己的职业道路,如果确实需要改变的话,这一切又是 AI 进步带来的?我们生活在一个变化的时代。正如我所说的,我们生活在一个特别难以预测的时期。
便签笔记
81:41
Things that we've taken for granted for centuries may not hold anymore. The way we do everything, and not just mathematics, will change.
那些我们几个世纪以来习以为常的事情,可能不再成立了。我们做每一件事的方式都会改变,不只是数学。
便签笔记
81:59
In many ways, I would prefer the much more boring, quiet era where things are much the same as they were 10 years ago, 20 years ago. But I think one just has to embrace that there's going to be a lot of change. The things that you study, some of them may become obsolete or revolutionized, but some things will be retained.
在很多方面,我其实更喜欢那种无聊、安静的年代,一切都和十年前、二十年前差不多。但我觉得人只能接受接下来会有大量的变化。你研究的东西,有些可能会过时或被彻底改写,但有些会保留下来。
便签笔记
82:26
You always have to keep an eye on opportunities for things that you wouldn't be able to do before. In math, you previously had to go through years and years of education and be a math PhD before you could contribute to the frontier of math research. But now it's quite possible at the high school level, or whatever, that you could get involved in a math project and actually make a real contribution because of all these AI tools, Lean, and everything else. There will be a lot of non-traditional opportunities to learn, so you need a very adaptable mindset.
你必须一直留意那些以前做不到、而现在有机会去做的事情。在数学里,以前你得经过很多很多年的教育、拿到数学博士学位,才能为数学研究的前沿做贡献。但现在完全有可能,在高中阶段之类的时候,你就能参与一个数学项目并真正做出实际贡献,因为有了这些 AI 工具、Lean 等等。将会有很多非传统的学习机会,所以你需要一种非常能适应变化的心态。
便签笔记
83:06
There will be room for pursuing things just for curiosity and for playing around. You still need to get your credentials. For a while it will still be important to go through traditional education and learn math and science the old-fashioned way.
也会有空间让人纯粹出于好奇心去追求一些东西,去玩一玩。你仍然需要拿到你的资历。在一段时间内,走传统教育路径、用老办法学数学和科学,仍然是重要的。
便签笔记
83:28
But you should also be open to very different ways of doing science, some of which don't exist yet. It's a scary time, but also very exciting. That's a great note to close on. Terence, thanks so much. Pleasure.
但你也应该对非常不同的做科学的方式保持开放,其中有些方式现在还不存在。这是个让人害怕的时代,但也非常令人兴奋。这是个很好的收尾。Terence,非常感谢。很荣幸。
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

陶哲轩认为 AI 已把"想法生成"的成本压到近零,科学的瓶颈已转移到验证、评估与叙事说服;当前 AI 做数学的特点是"广而不深"、能一次性攻克低垂果实却无法累积部分进展,未来数学将是人机互补、以规模化"实验数学"为新范式的形态。

核心要点

  • 开普勒的故事说明"高温随机尝试 + 可验证数据库"能驱动真正的科学进步。 开普勒先有错误的柏拉图立体理论,后偷得第谷·布拉赫精度高十倍的观测数据,经二十年反复试错才得出椭圆轨道与等面积定律,第三定律甚至埋在讨论行星"和声"的占星式著作里;牛顿一百年后才给出统一解释。陶强调应同等赞扬布拉赫的数据采集——没有那多一位小数的精度,开普勒的结果不可能成立。
  • 数据点少时"拟合"极易变成数值巧合。 开普勒第三定律只靠 6 个数据点回归得出,实属幸运;波德定律同样拟合行星距离,天王星和谷神星一度"验证"了它,海王星却完全不符,最终证明只是巧合。陶推测开普勒不太强调第三定律,正是因为直觉上知道 6 个点不足以下结论。
  • AI 已使假设生成成本趋近于零,瓶颈转向验证与筛选。 类比互联网使通信成本归零却不自动带来富足:现在一个问题可以生成上千种理论,期刊已被 AI 投稿淹没,传统"同行评审几年内达成共识"的机制在每天上千篇的规模下失效,而人类尚不知道如何规模化地判断哪些想法真正推动学科前进。
  • 科学进步难以被强化学习式地客观打分,因为价值取决于未来与社会语境。 二进制"比特"本可被三进制取代,Transformer 也可能被其他架构替代,十进制并无特殊性只是惯性使然;哥白尼理论初期精度低于托勒密,牛顿的超距作用被莱布尼茨批评,都要等后人补全才显现价值。陶还指出叙事与说服是科学不可缺的"软"部分:达尔文用通俗英语综合零散证据而广受接受,牛顿用拉丁文且刻意藏私,其理论靠后人通俗化才传播开。
  • Erdős 问题的 AI 攻坚已进入平台期,单次成功率仅约 1–2%。 约 1100 个问题中 50 余个借 AI 解决,但纯 AI 一次性解决的浪潮持续约一个月后停止;已知三次用前沿模型同时攻击全部问题的尝试均未再产出新解。被解决的几乎都是"没有文献"的问题——组合一个冷门技巧与已有结果即可。社交媒体只放大成功案例,系统性测试显示每题成功率仅 1–2%,靠规模挑出赢家。
  • AI 擅长"广度",人类擅长"深度",需重新设计科学组织方式。 陶用"黑暗中的山脉"比喻:AI 是能跳两米的机器,能一次跃上所有矮墙,却不能"抓住把手停住、拉别人上来再往上跳"——无法累积部分进展,每次新会话就遗忘。他主张用广度型 AI 先测绘新领域、做出所有简单观察,再由人类专家攻克"困难孤岛"。
  • AI 将催生"实验数学"与"规模化数学"。 数学几乎纯理论,从未做过"取一千个问题测试哪种方法更有效"的大规模实验,现在可以了。顶级期刊论文通常是已有方法解决 80%、剩下 20% 需发明新技巧;AI 已能可靠完成前 80%(陶自测与自己纠错水平"打成平手"),但在无技巧可用的缺口处只会随机建议,追查这些建议往往得不偿失。
  • 陶本人的生产力变化是"更丰富、不更深"。 2023 年他预测 2026 年 AI 会成为"用得对就可信的合著者",现已兑现。如今论文里代码、图表、文献综述、数值计算大增,若没有 AI 这类论文要多花 5 倍时间;但核心难题仍靠纸笔,若只写 2020 年水平的论文,AI 并未节省多少时间。AI 主要接管了括号格式化之类的辅助性苦活。
  • "人工聪明"与"人工智能"的区别在于能否互动式地累积改进。 真正的合作是双方都不知道答案、提出原型策略、失败后修改、逐步画出可行路径;AI 目前更像反复试错的暴力搜索,解题后自身对数学的理解并未增长。
  • 对不可理解的 Lean 证明不必过度担忧,但需要"策略的形式语言"。 四色定理至今无优雅证明;若黎曼猜想被 Lean 证出,可以逐引理消融分析、由 AI 重构成更优雅版本,Erdős 网站上已有 3000 行证明被 AI 总结再改写的先例。更大的缺口是缺少能半形式化评估"猜想是否可信"的框架(如高斯从 10 万素数数据得出素数定理、素数随机模型支撑孪生素数猜想与黎曼猜想的信念),且必须防止强化学习钻后门。

结论与值得注意的细节

  • 陶预计十年内学生和论文中大量常规工作将被 AI 承担,但"那不是最重要的部分"——正如 19 世纪数学家解微分方程、人工计算员编对数表、单个基因组测序曾是整个博士课题,如今都被外包,学科却转向更高尺度的问题。
  • 他认为人机混合模式将主导数学"相当长时间",完全自主 AI 解决千禧年难题需要额外突破,是随机性事件;若黎曼猜想被证伪且靠大规模计算找到线外零点,那将是对素数随机模型的重创,基于素数的密码学会被迅速弃用。
  • 关于"认知上的哥白尼革命":人类过去以为自身智能是宇宙中心,现在发现存在优劣势迥异的其他智能,对"哪些任务需要智能"的排序要重排。
  • 他提出用"演化小型 AI 解基础问题"的模拟宇宙来研究科学进步的规律,因为人类只有一条历史时间线、约 100 个转折故事,数据不足以形式化"什么是进步"。
  • 生活方式细节:陶自认是"狐狸"而非"刺猬",靠合作、写博客固化所学(曾多次学会又忘掉,遂决定凡有趣的都写下);强调偶然性的价值——远程会议和搜索引擎消灭了走廊偶遇和翻期刊时的意外发现,在高等研究院待超过几个月反而灵感枯竭,"你需要一定程度的干扰"。
  • 对年轻人的建议:保持适应性心态,传统学历仍重要,但借助 AI 与 Lean,高中生已可能实际参与前沿数学项目;"这是一个可怕但也非常令人兴奋的时代"。
核心句型 · 9
1. X has driven the cost of Y down to almost zero, in a very similar way to how Z did …
“AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero.”
用一个已被广泛接受的历史类比来解释新现象,drive … down to 表示「把…压低到」。适合在论证中引入类比,后接 but 转折指出类比的局限。
2. It's not that … , it's that …
“It's not something where you can look at any given scientific achievement purely in isolation and give it an objective grade”
先否定一种常见理解,再给出真正的解释。something where 引导限定从句是口语中常见的松散结构,写作时可改为 something that allows you to。
3. The same reason that we should be X is the reason we should be Y
“The same reason that we should be bearish now is the reason we should be especially bullish.”
「同一个理由导向相反结论」的反转句式,制造张力。适合在辩论中翻转对方论据。注意 bearish/bullish 来自股市术语,可替换为 pessimistic/optimistic。
4. A excels at X, and B excels at Y … they're very complementary
“They excel at breadth, and humans excel at depth, human experts at least. I think they're very complementary.”
对比两者长处并归结为互补关系,避免非此即彼。excel at 后接名词或动名词;补语 human experts at least 是口语中的事后限定,可学习这种「先说再收窄」的方式。
5. Not for lack of trying
“There was a month where that happened and that has stopped, not for lack of trying.”
固定表达,意为「并非因为没努力」。放在句末作补充,语气克制而有力,用于说明失败不是由于懈怠。
6. X is simultaneously A and B
“The progress is simultaneously amazing and disappointing.”
用 simultaneously 并置两个矛盾形容词,表达复杂、矛盾的评价,比 both … and 更强调「同时并存」。适合评价新技术、新政策时避免一边倒。
7. It's made X richer and broader, but not necessarily deeper
“It's made the papers richer and broader, but not necessarily deeper.”
三个比较级并列,前两个肯定、第三个用 not necessarily 保留。可用于评估任何工具的实际收益:承认改善,同时精确划定边界。
8. What we lost out on was …
“What we lost out on was the casual knocking on a hallway door, just meeting someone while getting a coffee.”
What 引导的强调结构把「损失」放在焦点位置;lose out on 表示「错失(本可获得的好处)」。适合在讨论效率提升的代价时使用。
9. More often than not, …
“More often than not, I get a positive experience that I wouldn't have planned for.”
「多数情况下」的地道表达,比 usually 更口语、更有分量。后接的 that I wouldn't have planned for 用虚拟语气强调「计划之外」,可仿写为 that I wouldn't have expected。
生词精讲 · 166 · 按出现顺序
jumping off point phr. 0:00
切入点、起点(引出讨论的话题)
heliocentric /ˌhiːlioʊˈsentrɪk/ adj. 0:00
日心的,以太阳为中心的
enclose /ɪnˈkloʊz/ v. 0:57
围住,包住;(几何)外接
Platonic solids /pləˈtɑːnɪk ˈsɑːlɪdz/ n. 0:57
柏拉图立体(五种正多面体)
inscribe /ɪnˈskraɪb/ v. 1:35
(几何)内接;刻写
eccentric /ɪkˈsentrɪk/ adj. 1:35
古怪的,离经叛道的
naked-eye /ˈneɪkɪd aɪ/ adj. 2:12
肉眼的(不借助仪器)
jealous of phr. 2:12
(古义)小心守护、不肯与人分享的
descendants /dɪˈsendənts/ n. 2:12
后代,后人
fudges /fʌdʒɪz/ n. 2:53
含糊的凑合、修补手段(fudge 也作动词:蒙混)
sweep out phr. v. 2:53
(几何)扫过(面积)
proportional to phr. 2:53
与……成正比
high-temperature adj. 3:55
(LLM 采样)高温度的,输出更随机多样的
aside /əˈsaɪd/ n. 4:29
旁白,题外话,附带的说明
famine /ˈfæmɪn/ n. 4:29
饥荒
centripetal acceleration /senˈtrɪpɪtl əkˌseləˈreɪʃn/ n. 5:00
向心加速度
inverse-square law n. 5:00
平方反比定律
empirical regularities /ɪmˈpɪrɪkl ˌreɡjəˈlærətiz/ n. 5:00
经验规律(从观测数据中总结出的规律性)
prestige /preˈstiːʒ/ n. 5:44
声望,威望(此处作定语:最有声望的)
fruitful /ˈfruːtfl/ adj. 5:44
富有成果的,有产出的
eureka /juˈriːkə/ n./int. 5:44
「我找到了!」;灵光一闪的顿悟时刻
slop /slɑːp/ n. 6:41
(俚)粗制滥造的垃圾内容,尤指 AI 批量生成物
assiduous /əˈsɪdʒuəs/ adj. 6:41
勤勉的,孜孜不倦的
bottleneck /ˈbɑːtlnek/ n. 6:41
瓶颈,制约环节
paradigms /ˈpærədaɪmz/ n. 7:38
范式,典范模式
out-of-the-blue adj. 7:38
突如其来的,凭空冒出的
preconceived /ˌpriːkənˈsiːvd/ adj. 8:17
先入为主的,预先形成的
analogous to /əˈnæləɡəs/ phr. 9:06
与……类似的,可类比的
regression /rɪˈɡreʃn/ n. 10:09
(统计)回归分析
geometric progression n. 10:09
等比数列
crank theory /kræŋk/ n. 10:48
民科理论,怪人提出的荒谬理论
asteroid belt /ˈæstərɔɪd belt/ n. 10:48
小行星带
fluke /fluːk/ n. 10:48
侥幸,纯属偶然的巧合
tentative /ˈtentətɪv/ adj. 11:26
试探性的,不确定的,有所保留的
abundance /əˈbʌndəns/ n. 12:18
富足,充裕
peer review n. 13:07
同行评审
red herrings /red ˈherɪŋz/ n. 13:07
转移注意力的误导线索
dead ends n. 13:07
死胡同,走不通的路
gauge /ɡeɪdʒ/ v. 14:02
判断,估量
Pulse-code modulation n. 14:37
脉冲编码调制(PCM)
the test of time phr. 15:06
时间的检验
niche /niːʃ/ adj. 15:06
小众的,细分领域的
bearing fruit phr. 15:06
结出成果,产生效果
first principles n. 15:06
第一性原理,基本原理
inertia /ɪˈnɜːrʃə/ n. 16:29
惯性;(比喻)因循守旧的阻力
in isolation phr. 16:29
孤立地,脱离上下文地
in retrospect /ˈretrəspekt/ phr. 17:33
回想起来,事后看来
implausible /ɪmˈplɔːzəbl/ adj. 17:33
难以置信的,不太可能的
parallax /ˈpærəlæks/ n. 17:33
视差
chide /tʃaɪd/ v. 18:13
责备,斥责
action at a distance n. 18:13
(物理)超距作用
falsify /ˈfɔːlsɪfaɪ/ v. 18:13
证伪
millennium /mɪˈleniəm/ n. 18:49
一千年
ad hoc /ˌæd ˈhɑːk/ adj. 18:49
临时的,为特定目的凑合的
a work in progress phr. 18:49
未完成、仍在进行中的工作
equivalence /ɪˈkwɪvələns/ n. 19:42
等价,等效
cognitive /ˈkɑːɡnətɪv/ adj. 20:25
认知的
contemporaneous /kənˌtempəˈreɪniəs/ adj. 21:47
同时代的
overwhelming /ˌoʊvərˈwelmɪŋ/ adj. 22:17
压倒性的
cumulative /ˈkjuːmjələtɪv/ adj. 22:17
累积的
lines up phr. v. 22:17
(数据)吻合,对得上
disparate /ˈdɪspərət/ adj. 22:48
迥异的,零散不相关的
compelling /kəmˈpelɪŋ/ adj. 22:48
极具说服力的,引人入胜的
heredity /həˈredəti/ n. 23:42
遗传
held back phr. v. 23:42
隐瞒,保留不说
exposition /ˌekspəˈzɪʃn/ n. 24:23
阐述,解说
pride ourselves on phr. 25:19
以……为傲
squishy /ˈskwɪʃi/ adj. 25:19
软乎乎的;(比喻)模糊、不严格的
transitional forms n. 25:19
(生物)过渡形态,过渡化石
overhang /ˈoʊvərhæŋ/ n. 26:06
悬突;(比喻)尚未被利用的余量
divine /dɪˈvaɪn/ v. 26:06
推断,凭直觉猜出
squeezing every last possible drop phr. 26:52
榨干最后一滴(比喻充分利用)
quant hedge funds n. 26:52
量化对冲基金
brute force /bruːt fɔːrs/ n. 27:35
暴力穷举
residual /rɪˈzɪdʒuəl/ adj. 28:06
残差的,剩余的
heuristic /hjuˈrɪstɪk/ n./adj. 28:06
启发式方法;启发式的
backdoored /ˈbækdɔːrd/ adj. 28:36
被植入后门的
chipping away at phr. v. 30:55
一点一点地啃、逐步削减
one-shots v. 30:55
一次性直接搞定
not for lack of trying phr. 30:55
并非因为没努力尝试
critique /krɪˈtiːk/ v. 31:50
评论,批判性评价
climbable /ˈklaɪməbl/ adj. 32:39
可攀爬的
set them loose phr. 33:18
把……放出去,任其自由行动
breached /briːtʃt/ v. 33:32
攻破,突破
hill climb v. 33:32
爬坡式渐进优化
bearish /ˈberɪʃ/ adj. 34:19
看空的,悲观的(源自股市)
bullish /ˈbʊlɪʃ/ adj. 34:19
看多的,乐观的
waterline /ˈwɔːtərlaɪn/ n. 34:19
水位线;(比喻)能力阈值
inference compute n. 34:19
推理算力
complementary /ˌkɑːmplɪˈmentri/ adj. 35:15
互补的
map it out phr. v. 36:10
勾勒轮廓,绘制地图
instrumental to /ˌɪnstrəˈmentl/ phr. 37:05
作为达成……的手段
dichotomy /daɪˈkɑːtəmi/ n. 37:45
二分法,对立的两分
proxy /ˈprɑːksi/ n. 37:45
代理指标,替代物
boilerplate /ˈbɔɪlərpleɪt/ n. 38:19
样板代码,套话
offload /ˌɔːfˈloʊd/ v. 38:19
卸载,转交给他人处理
mesh with /meʃ/ phr. v. 38:19
与……啮合、协调配合
inhibit /ɪnˈhɪbɪt/ v. 39:10
抑制,阻碍
place a premium on phr. 39:35
高度重视,看重
coherent /koʊˈhɪrənt/ adj. 39:35
融贯的,前后一致的
in its infancy phr. 40:46
处于萌芽期
crux /krʌks/ n. 40:46
关键分歧点,症结
humongous /hjuːˈmʌŋɡəs/ adj. 40:46
(口)巨大的
wrinkle /ˈrɪŋkl/ n. 41:54
(口)小花招,巧妙的小改动
handicap /ˈhændikæp/ n. 41:54
不利条件,障碍
chase them down phr. v. 43:25
追查到底,穷追不舍
obscure /əbˈskjʊr/ adj. 43:38
鲜为人知的,冷门的
sweeps /swiːps/ n. 44:12
地毯式扫荡、遍历
acclimatize /əˈklaɪmətaɪz/ v. 45:28
适应(新环境)
blew ... out of the water phr. 46:12
彻底击败,远超
streak /striːk/ n. 46:53
连胜纪录;连续一段
auxiliary /ɔːɡˈzɪliəri/ adj. 47:41
辅助的
numerics /nuːˈmerɪks/ n. 47:41
数值计算
By the same token phr. 48:28
同理,出于同样的理由
adaptivity /ˌædæpˈtɪvəti/ n. 49:40
适应性
handhold /ˈhændhoʊld/ n. 50:30
(攀岩)着力点,抓手
gnarly /ˈnɑːrli/ adj. 51:54
(俚)棘手的,极难的
operationalized /ˌɑːpəˈreɪʃənəlaɪzd/ v. 52:19
使可操作化,落实为可衡量的指标
rubrics /ˈruːbrɪks/ n. 52:19
评分量表,评价标准
gobbledygook /ˈɡɑːbldiɡʊk/ n. 52:48
天书,晦涩难懂的胡话
uninsightful /ˌʌnɪnˈsaɪtfl/ adj. 53:40
缺乏洞见的
get a lot more mileage out of phr. 54:30
从……获得更多收益
interplay /ˈɪntərpleɪ/ n. 54:30
相互作用,交互
latent /ˈleɪtnt/ adj. 55:51
潜在的,隐藏的
atomically /əˈtɑːmɪkli/ adv. 56:26
以原子方式,逐个独立地
lemmas /ˈleməz/ n. 56:26
引理
ablation /əˈbleɪʃn/ n. 57:04
消融实验(逐一移除组件以测其作用)
refactoring /ˌriːˈfæktərɪŋ/ n. 58:07
重构
nascent /ˈnæsnt/ adj. 58:42
新生的,初现的
incomprehensible /ɪnˌkɑːmprɪˈhensəbl/ adj. 58:42
无法理解的
axioms /ˈæksiəmz/ n. 59:20
公理
plausibility /ˌplɔːzəˈbɪləti/ n. 60:00
合理性,可信度
deductive /dɪˈdʌktɪv/ adj. 60:44
演绎的
hackable /ˈhækəbl/ adj. 60:44
可被钻空子的,可被攻破的
exploits /ˈeksplɔɪts/ n. 60:44
(安全)可利用的漏洞
paradoxical /ˌpærəˈdɑːksɪkl/ adj. 62:18
自相矛盾的
inversely proportional phr. 63:20
成反比
natural logarithm /ˈlɔːɡərɪðəm/ n. 63:20
自然对数
analytic number theory n. 64:42
解析数论
pseudo-random /ˌsuːdoʊ ˈrændəm/ adj. 64:42
伪随机的
twin prime conjecture n. 65:26
孪生素数猜想
non-rigorous /nɑːn ˈrɪɡərəs/ adj. 65:57
不严格的
conjectural /kənˈdʒektʃərəl/ adj. 65:57
猜想性的,推测的
a serious blow to phr. 66:35
对……的沉重打击
paradigm shifts n. 67:15
范式转移
a decent shot at phr. 67:53
有相当的机会做到
autodidacts /ˈɔːtoʊdaɪdækts/ n. 69:35
自学成才者
hedgehogs /ˈhedʒhɑːɡz/ n. 70:12
刺猬(伯林比喻:专精一事者)
obsessive streak /əbˈsesɪv striːk/ n. 70:40
强迫性的性格倾向
completionist /kəmˈpliːʃənɪst/ n./adj. 70:40
完成主义者(游戏术语:非全部完成不可)
wean myself off /wiːn/ phr. v. 70:40
逐步戒掉
reconstruct /ˌriːkənˈstrʌkt/ v. 72:02
重建,重构(论证)
referee report n. 72:02
审稿报告
drudgery /ˈdrʌdʒəri/ n. 72:35
苦差事,单调乏味的工作
veil of ignorance n. 72:35
无知之幕(罗尔斯政治哲学概念)
reluctantly /rɪˈlʌktəntli/ adv. 73:29
不情愿地
serendipity /ˌserənˈdɪpəti/ n. 74:01
机缘巧合,意外的幸运发现
More often than not phr. 74:01
多数情况下
run out of inspiration phr. 76:20
灵感枯竭
laboriously /ləˈbɔːriəsli/ adv. 77:38
费力地,辛苦地
outsourced /ˈaʊtsɔːrst/ v. 79:01
外包
stochastic /stəˈkæstɪk/ adj. 79:49
随机的,概率性的
taken for granted phr. 81:41
视为理所当然
obsolete /ˌɑːbsəˈliːt/ adj. 81:59
过时的,被淘汰的
adaptable /əˈdæptəbl/ adj. 82:26
适应力强的
credentials /krəˈdenʃlz/ n. 83:06
资历,资格证明
理解自测 · 11 题 · 是真懂了,还是以为自己懂
1. 开普勒最初想用第谷的数据验证的是什么理论?结果如何?

开普勒最初想验证的是「柏拉图立体嵌套理论」:六颗已知行星之间有五个间隔,恰好对应五种柏拉图立体(立方体、正四面体等),他认为这体现了上帝设计的数学完美。拿到第谷的数据后,他发现数据与该理论偏差约 10%,各种挪动圆轨道的修补都不奏效。经过多年分析,他反而推出了轨道是椭圆、等时间扫过等面积这两条定律,十年后又得出第三定律(周期与到太阳距离的某个幂次成正比)。这一段在访谈开头(第 1–4 段),是全片类比的基础。

2. 陶哲轩提到的波得定律是什么?它在论证中起什么作用?

波得定律是天文学家约翰·波得用行星距离数据拟合出的「平移等比数列」,它预言火星与木星之间有一颗缺失行星。天王星和谷神星的发现恰好符合该规律,人们一度以为这是新的自然定律,但海王星的距离偏差极大,最终被证明只是数值巧合。陶用它作为开普勒第三定律的反例:两者都是用五六个数据点拟合曲线,开普勒得到了正确结论,波得却得到错误结论——说明小样本回归本质上不可靠,开普勒的成功带有运气成分,这也解释了他为何对第三定律的表述相对谨慎。

3. 陶哲轩关于 AI 解决 Erdős 问题的关键数字有哪些?

陶给出了几个数字:约 1100 道 Erdős 问题中,大约 50 道在 AI 协助下被解决,还有约 600 道未解;「纯 AI 一次性解出」的情况集中在一个月内出现,此后停滞,尽管至少有三次让前沿模型对全部问题扫荡的尝试;系统性研究显示,对任一给定问题 AI 的成功率大约只有 1% 到 2%。此外,被解决的 50 道几乎全是「基本没有文献」的问题,解法多为一个冷门技术加一个已有结果的组合。这些数字出现在「Erdős 问题」和「1–2% 成功率」两个章节。

4. 陶在个人生产力上给出了什么样的评估?

陶拒绝把生产力看成一维量。他说如果用今天的写法但不用 AI,论文会多花五倍时间,但那是因为 AI 让他能加入更多代码、图表、深度文献检索和数值计算等辅助内容;而攻克问题最难部分的核心工作仍靠纸笔,没有明显加速。如果重写一篇 2020 年功能水平相同的论文,其实没省多少时间。他的结论是 AI 让论文「更丰富、更宽广,但不一定更深」。这段在「陶的亲身体验」章节(第 65–67 段)。

5. 主持人把开普勒比作「高温度 LLM」,陶哲轩接受了这个类比的哪些部分,又修正了哪些?

陶接受的部分:科研包含十几个环节,想法生成只是其中被过度歌颂的一环;开普勒确实尝试了大量随机假设,很多未发表,而可验证的数据(第谷)是成功的必要条件。他修正的部分有两点:第一,想法生成必须由同等分量的验证匹配,否则就是垃圾内容,第谷的十倍精度同样值得歌颂;第二,他区分了「先假设后检验」的经典范式和当代「先数据后假设」的范式,指出开普勒是先有先入为主的理论,且拟合第三定律时只有六个数据点,与「海量仿真数据库」的类比并不完全对应。由此他把讨论从「AI 能否生成想法」转到「谁来验证」。

6. 为什么陶认为「一项科学成就是否重要」很难用强化学习来打分?

陶给出三层理由。第一,想法的价值取决于未来:深度学习曾长期边缘化,Transformer 也不是唯一可能的架构,率先被采纳者才成为标准。第二,取决于文化与社会惯性:十进制没有数学上的特殊性,只因为大家都用而无法更换。第三,正确的理论初期常常更差:哥白尼精度不如托勒密,牛顿理论有超距作用等谜团,但它们仍是进步。因此无法把成就孤立出来给客观分数,也就无法像局部问题那样设计可靠的奖励信号。这一论证在「进步为何难以事前评分」章节(第 20–26 段)。

7. 陶提出「人工机巧」与「人工智能」的区分,他判断当前 AI 缺少的具体是什么?

陶用协作解题的场景定义智能:两人都不知道怎么解,一人提出原型策略,检验、失败、修改,想法在对话中适应性地、累积地演化,最终系统地摸清可行与不可行。当前 AI 缺少的正是这种累积性:用跳跃机器比喻,它们能跳、失败、再跳,但不能「跳一小步、抓住着力点、把别人拉上来、再从那里继续」。此外它们没有状态:新会话就遗忘,解题不转化为自身技能,只可能成为下一代训练数据的 0.001%。所以当前范式本质是可扩展的大规模试错,而非交互式累积。

8. 主持人提出「看空的理由恰恰是看多的理由」,这个反转的逻辑是什么?陶如何回应?

主持人的逻辑:看空者说 AI 只能够到矮墙;但 AI 有人类没有的性质——一旦达到某个能力水位线,就能同时覆盖该水位线下的全部问题,而我们无法复制一百万个陶哲轩各给一百万美元算力去同时研究一百万个问题。所以「只够到矮墙」的现状,同时说明「达到人类水平时」将带来质变。陶表示同意,并把它概括为 AI 擅广度、人类擅深度、二者互补;他进一步指出现行科研体制围绕深度设计,需要重新设计以利用广度,例如让中等能力 AI 先绘制新领域地图、扫清易得结论,再由人类专家攻克「困难的孤岛」。

9. 陶为什么说数学一直「几乎完全是理论性的」,AI 又如何改变这一点?

陶指出大多数科学有理论和实验两翼,但数学几乎只有理论:数学家看重融贯、干净的解释,却从未大规模实验「两种解法哪种更有效」,只有直觉,没有取一千道题逐一测试的研究。AI 使这种「实验数学」首次可行:不再在意单个问题及其解决过程,而是收集大规模数据看哪些方法有效,就像软件公司要推出一千个产品时寻找可规模化的工作流。他认为「规模化地做数学」仍在襁褓期,但那正是 AI 真正颠覆学科之处。这段在第 54–56 段。

10. 如果有人反驳说「AI 只要在 Lean 里证出黎曼猜想,人类是否理解并不重要」,陶会如何回应?

陶会部分同意、部分保留。他承认有些定理可能只能靠暴力(四色定理至今无优雅证明),甚至黎曼猜想若为假也可能被纯计算证伪,那会「非常令人失望」。但他不太担心「不可理解的证明」:形式化证明的优势正是可以原子化研究每条引理,判断哪一步关键、哪一步套路;未来会有专门对 AI 生成的巨型证明做消融、精简、优雅化的数学家,也会有 AI 做总结与重构。同时他强调数学界看重黎曼猜想是因为相信它需要新数学或新联系,因此更可能的路径是人类与强大 AI 的新型协作,而非完全自主一步到位。

11. 陶关于「机缘巧合」的论点,放到 AI 精准检索和自动化科研的未来还成立吗?

陶本人正是这样迁移的。他先举三个例子:疫情期间线上会议让见面次数不降但失去走廊偶遇;图书馆翻期刊会意外读到邻篇文章,而精准搜索消灭了这种可能;在高等研究院无干扰环境待几个月后灵感反而枯竭。他的结论是生活需要一定的随机性和「高温度」,现代社会(不只 AI)过度优化的危险在于「没有优化我们的优化本身」。随后他明确把这一点挂钩到 AI 前景:AI 可能通过消灭机缘巧合而抑制某些类型的进步。因此论点不仅成立,而且是他对 AI 加速科学这一乐观预期的一个重要保留。

精读便签
下载便签 手机:长按图片保存
← 上一期 · NO.020Can AI Prove It? Terence Tao on “Big Math” and Our Theoretical Future | The Futurology Podcast 下一期 · NO.022 →Terry Tao "How to think like a mathematician" presented by the UCLA Curtis Center
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com