Why Fei-Fei Li Is Betting on Spatial Intelligence · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Why Fei-Fei Li Is Betting on Spatial Intelligence

节目发布 2026-09-04 · a16z
李飞飞 贾斯汀·约翰逊 本·米尔登霍尔
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:这场对谈录制于 World Labs 新一代世界模型 Atlas 发布的次日,三位 World Labs 的创始人同席:李飞飞(联合创始人兼 CEO、斯坦福大学教授)、贾斯汀·约翰逊(联合创始人,Atlas 的技术负责人)、本·米尔登霍尔(联合创始人,NeRF 的作者)。话题从 Atlas 究竟做了什么讲起,一路回到公司成立之初的路线抉择,再谈到创作工具、机器人仿真与「新视角预测是否 AI 完备」这样的根本问题。本文依据现场录音编译整理,只删去口语枝节、寒暄与重复,论点、例子与语气一律保留。

昨天发布了什么

主持人: 昨天是个大日子,你们发布了一个新的前沿模型,反响非常好,而且还在持续发酵。我想这场对话不妨这样安排:先把昨天发布的东西讲清楚,然后再往回走,一步步回到历史。贾斯汀,要不你先讲讲昨天发布的是什么,为什么它重要?

约翰逊: 好。Atlas 是我们新一代的世界模型(world model)。它做三件基本的事:生成世界、重建世界、仿真世界。在这三件事里面,又有几项主要能力。

第一,它的相机条件生成(camera conditioned generation)做得非常好。你可以输入一张图,再配上一条相机轨迹,用这条轨迹去操纵模型,让它沿着你想要的任意视角生成画面。

第二,它非常擅长稀疏三维重建(sparse 3D reconstruction)。你可以输入一帧,也可以输入多帧,最多一百帧真实世界的视图,用它们把真实世界重建出来。重建的结果有两种形态:一种是在空间里穿行的视频,一种是显式的三维重建。

第三,它可以用来做仿真。这一块我们展示了两样东西,一样是那些「子弹时间」的视频,在网上引起了很多关注,另一样是机器人仿真。

主持人: 什么叫子弹时间视频?

约翰逊: 这个说法来自《黑客帝国》。第一部里有个著名的镜头,尼奥往后仰倒。

主持人: 对,想起来了。

约翰逊: 就是那个。他往后倒下去,画面是慢动作,镜头绕着他整整转一圈。他们当年拍这个镜头的办法,是架一圈相机,几百台。人在摄影棚里绿幕前倒下去,几百台相机同时从各自的角度拍,再用这几百台相机的素材合成出《黑客帝国》里那个著名的镜头。

现在用 Atlas,我们最少只要三台相机就能做同样的事。不用摄影棚、不用绿幕、不用昂贵的标定。我们就是把三台 iPhone 架在三个三脚架上,拍一段正在发生的事,比如有人投篮,有人把一颗草莓丢进一碗牛奶里。然后从这三段 iPhone 视频出发,我们可以重新取景,把时间冻住,让镜头在牛奶溅起来的那一刻飞进去,拿到那种时间凝固的惊人画面。就用这么几台相机。

新视角预测

主持人: 能不能用最简单的话说清楚 Atlas 到底做什么?输入是什么,输出是什么?

约翰逊: 可以。Atlas 最核心、最根本的一条原则是:它做的是新视角预测(new view prediction)。这是一个非常根本的原语(primitive),我们认为它对基础模型来说是全新的东西,此前没有人做过。

我们知道,大语言模型建立在下一个词元预测(next token prediction)之上;我们也见过视频模型,它们建立在下一帧预测(next frame prediction)之上。Atlas 是真正意义上的新视角预测。你给它若干张某个场景的视图,或者一段对场景的描述,这些东西进入我们所说的「空间上下文」(spatial context),它隐含地描述了我们要谈论的是哪一个世界。然后你可以把一台虚拟相机指向时空中的任意一点,Atlas 会知道,从那个时空位置看过去,这个世界应该长什么样。

主持人: 现在外面有成千上万的视频模型,个个都自称世界模型,个个都说自己能出新视角。你能不能更具体地把 Atlas 和它们区分开?

米尔登霍尔: 我觉得贾斯汀刚才说的空间上下文那一点,在这里极其关键。视频模型有很多,其中不少是靠单图输入,或者靠首帧到尾帧的插值出名的。现在我们也开始看到一些模型能做那种「全参考」,一次塞进二三十张、五十张图。

但 Atlas 的关键在于,你放进去的每一帧,对模型来说都有一个空间上被锚定的含义。它不是一张任由模型自己去解读的图片,也不是那种你只能在文字提示里跟它反复掰扯、争取它照你说的做的东西。在 Atlas 里,每一张图都带着一个与之对应的三维相机位姿(camera pose),这意味着你可以以极高的精度完成重建这件事。

比如我们有这个房间四个角落各一张照片,把它们放进模型,你会得到这个房间里所有东西的精确复现。它不会去猜另一个角落有什么,也不会去猜物体之间的关系,它就是把你给它的东西原样再现出来。

这件事也可以用在创作和想象的方向上。你拿两张照片,一张来自某次 AI 生成,一张来自真实地点,你可以把它们摆放、布置到你要的位置上,从而搭出一段带明确导演意图的穿行镜头,镜头怎么看、怎么走,完全由你放内容的确切位置来决定。这跟视频模型那种只有较高层级的文字控制、只能一遍遍重抽的老虎机式体验,我认为是非常不同的。

一个模型里的两件事

主持人: 这算是把传统视频模型放大之后的自然结果吗,还是一套新的架构?

约翰逊: 我觉得它是相当新的东西,有几个原因。

我们常讲的一点是,它在同一个模型里同时做生成和重建。就像本刚才说的,这东西能拿这个房间的几张视图,把房间里的一切按你看到的样子重建出来。而历史上,重建一直是计算机视觉里自成一体的子领域,有自己的专门任务、自己的专门模型。生成则是所有文生视频模型擅长的事,是这几年那些大扩散模型擅长的事,它们在创作类应用上很好用:我想象一个从来不存在的东西。现在有了 Atlas,我们第一次把视觉智能的这两个部分放进了同一个模型,它能在一套架构里同时做三维重建和生成。

为了做到这一点,我们改了几件事。一是必须从一开始就让它多模态。这东西原生支持文本、原生支持图像、原生支持视频,也原生支持相机位姿作为模型的输入,我认为此前没有人在预训练阶段这么做过。此外,它把三维当作一种原生模态来处理。所以这个模型从最开始就是按原生多模态设计的,据我所知没有别人这样做。

主持人: 打断一下,我对这个领域不算太熟。三维指的是深度,还是模型之类的,这具体是什么意思?

约翰逊: 我们目前采用的形式是深度图(depth map)。也就是说,一帧画面带着一台虚拟相机,告诉你它在三维空间中的位置,相机位置和相机参数是模型的原生输入。挂在这个相机位置上的,既可以是 RGB,告诉你那个位置看过去是什么样子,也可以是一张深度图,告诉你那个位置在三维空间里的结构是什么样子。所以文本、图像、视频、三维相机,这些模态是它以多模态方式一起处理的。

半个世纪的第一次

李飞飞: 我想补一句,因为贾斯汀刚才说的,还有本说的,其实特别重要,而且被严重低估了。这是我们第一次让像素生成和像素重建统一起来。

在计算机视觉的世界里,这个领域已经存在半个多世纪了。我在这个领域待了几十年,真的无法告诉你有多少篇博士论文是写重建或者新视角合成的。而且我们这个领域传统上是分轨的,你去参加一次计算机视觉会议,会看到像素生成一轨、识别一轨、三维重建又一轨。而这是一个优雅的模型,它通过锚定视角、锚定视角估计,把重建和生成这两个问题合到了一起,这股力量之大,怎么说都不为过。

空间智能是什么

主持人: 我们退一步。公司刚成立的时候,我记得你说你们要解决的是空间智能(spatial intelligence)。现在有了这个新模型。以一个外行的眼光看,我觉得它非常通用:那边有下一个词元预测,这边有新视角预测,你从一组视图里得到一个新视图。你能不能勾勒一下,这一步对空间智能这个大问题究竟意味着什么?也许可以先说说空间智能到底指什么。

李飞飞: 空间智能最终必须让我们能够做三件事:生成这个空间,在其中进行推理,以及在其中编辑和交互。

我们可以争论它是三维还是四维。归根到底,加上时间维度,它是四维的。但哪怕只是三维,这些也是一个人,或者说空间智能,必须能够胜任的基本任务。在此基础上我们才谈得上渲染、仿真、以及规划行动。而要做到这些,有一个必须先解决的根本问题:理解空间的几何、结构和物理。

我确实认为 Atlas 是一次实质性的推进,因为现在对每一帧画面,你都能生成、估计出一条关键信息,也就是视角、相机位姿。这是关于空间几何最要紧的信息。正是它带来了我们在模型下游看到的那些涌现行为,我们在博客里展示了这些。

所以在通往空间智能的路上,生成像素当然是早期的一步,你所说的那无数模型都做到了。但生成真正具备空间上下文、真正被空间锚定的像素,绝对是另一个重大台阶。这一步很难,而 Atlas 迈出去了。

我们当然可以继续往下走。还有第四个维度,时间,它会带来动态性。还有更高保真的仿真,以及对空间更精细的刻画。这些都是空间智能路线图的一部分。

为什么不直接做 Atlas

主持人: 好,我很想深挖它要去哪里。但先聊聊你们是怎么走到这一步的。World Labs 成立多久了?

李飞飞: 两年半。

主持人: 而且你们之前已经发布过模型。为什么当初不直接做 Atlas?

李飞飞: 好问题。贾斯汀的团队需要很多芯片。

约翰逊: 对,把这东西的规模拉起来,需要大量 GPU。去年我们发布了 Marble 世界模型,那是我们推出的第一个真正意义上的大型世界模型,现在的 Marble 产品就跑在它上面。Marble 很棒,它能接受图片、视频、文字提示,用这些生成三维世界。

Marble 和 Atlas 最大的差别之一,恰恰在输出模态上。Marble 的输出高度集中在高斯泼溅(Gaussian splat)这一种表示上,不管你输入什么,它输出的都是一个用高斯泼溅表示的三维世界。高斯泼溅确实很有用,它很好,容易渲染,在移动设备和 VR 设备上渲染效率高,能和游戏引擎、仿真引擎互通,优点不少。但我认为在上一代 Marble 模型里,它成了一个瓶颈。

所以做 Atlas 的时候,我们把这件事重新设计了一遍。我们意识到,这些模态需要在更早的阶段就分岔出来,让它们在模型内部以更统一的方式运转。于是在 Atlas 里,最根本的原语不再是「生成一个高斯泼溅世界」,而是我们刚才说的新视角预测。这个原语既能生成 RGB 画面,也能生成三维;需要的时候,我们照样可以用它做出漂亮的高斯泼溅世界。但当我们不需要的时候,就不必让所有输出都从高斯泼溅这个口子里挤出去。要弄清楚这几种表示各自的利弊,真的付出了很多血汗。这是一方面。

另一方面,你得一级一级爬缩放的梯子。你得先做小实验、训小模型,去建立自己的判断:什么东西行得通,什么东西能扩上去。如果你能立刻知道哪条路能扩,那你当然应该直接去做那条路。但我们创办公司的时候,世界完全是另一个样子,技术也完全是另一个样子。空间智能根本没有缩放定律。我们对它该走到哪里有很多野心,但还是花了几轮迭代,才碰到这套我们认为「就是它了」的表述方式。这一套是能扩上去的。

稠密重建有多苦

主持人: 本,你是 NeRF 的作者,做过大量三维和重建的工作。对我来说,有了多张视图就能得到一个三维的东西,这件事并不显然,但你的职业生涯基本就是在把东西变成三维。能不能讲讲这一步?

米尔登霍尔: 是的,我职业生涯里绝大部分时间都在做从图像生成三维的事。公司早期我们其实反复讨论过这个问题:三维到底该怎么来?是先合成多个视角再从中构建三维,还是直接奔着三维去?领域里对这两条路哪条会胜出、哪条能更早收获成果,一直有很多不确定。

但我当时就很笃定,原因是我看到了那种近乎蛮力的扩展的力量。我说的不是真正意义上的模型扩展,是很小很小的、婴儿级别的扩展,就是过去三年里稠密重建(dense reconstruction)上发生的那种扩展。

主持人: 为什么叫「稠密」?因为我知道后面要谈稀疏,我想确保大家明白稠密和稀疏分别是什么。

米尔登霍尔: 好。我觉得这一点在商业化那一面也是根本性的挑战之一,就是三维重建技术很难做成产品。

一般人的直觉是这样的:我拍了这个物体三张照片,或者拍了这个房间六张照片,我自己看这些照片,脑子里就能把它们拼起来,能把空缺补上,能明白这是什么。但在那种数据驱动的先验,和蛮力式的稠密重建之间,一直没有真正的调和。稠密重建更接近科学成像或者医学成像:你必须说,凡是我希望出现在重建结果里的每一样东西,我都需要至少三到四个视角。

你想想看,就在这个房间里,麦克风底下、桌子底下、每一道缝隙和夹角、植物的叶片之间,要真拿到一张覆盖到每一个角落的图像,那是一件极其枯燥、极其耗时的活,你得绕着房间一点点走。你们都见过我在各种地方跑来跑去做采集。对训练有素的人来说,也许几分钟;但如果你把一台普通手机或者采集设备交给一个第一次做这件事的普通消费者,哪怕是专业人士,大概要花一个小时。我见过有人第一次扫描一个多房间的空间,走了两个小时才拿到足够的覆盖。这就是一个极其累人、极其乏味的循环。

所以我们说稠密,真的就是字面意义上的稠密。这个房间我要拍一百张、两百张、三百张照片才能采下来。而我们想做的,是把它降到三张。

主持人: 三张。

米尔登霍尔: 我们说的是五十倍到一百倍的降幅。到了那个量级,什么样的采集可以拿来重建,这笔账整个就翻过来了。你可以回头去用你已有的影像,可以用你在网上找到的素材去搭场景,可以用随手拍的视频,把过去你根本不会当成「可重建素材」的镜头挖出来,重新变成三维。这件事我们用 Atlas 玩了很多。我拿过一堆自己以前根本跑不通的采集,丢进这个系统,第一次看到了重建结果;或者拿旧采集,把其中百分之九十五的照片扔掉,再去想象一些角度,那是传统 NeRF 或者泼溅方式永远给不了我的。

李飞飞: 网站上那些演示里有一个被低估了,就是斯坦福那个。本用三到二十五张图片,把整个斯坦福方庭重建了出来。但关键在于,我们得从空中视角展示它,而每一张输入图片,都是本站在地面上、从地面拍的。所以你看到的一切都是生成出来的,但它们遵守重建的法则。这真的很神奇。

重建就是超长上下文

约翰逊: 而这正是生成和重建必须以一种根本的方式互相配合的地方。

在本刚才说的那种经典重建里,你之所以需要那么多视图,是因为我需要多张图像,把三维空间里的这个点三角化出来,从多个视角看到它。传统方案里这是刚性要求。反过来,凡是没有被这些视图拍到的东西,凡是没有在任何一张输入视图里出现过的像素,在三维重建里都会是一个洞。因为一样东西既然没有在输入里出现过,你就得去想象它,把空缺补上。而这从根本上说是一个生成过程。

所以哪怕在这个房间里,哪怕我们放本拿着单反去拍几百张视图,这位做稠密采集的世界级专家,也照样会漏掉一些地方。他不可能拍到所有麦克风的底下、所有桌子的底下,或者所有椅子腿之间的缝隙。不管你拍多少视图,总会漏掉点什么。这就是模型里必须有生成这条机制的原因,因为你永远不可能拍全。你需要模型有生成的能力去想象:根据我看到的东西,先把能三角化的部分三角化出来,然后把那些注定没被拍到的地方补上。

米尔登霍尔: 这里还有件特别有意思的事:大语言模型早就把这个道理吃透了。头几年有过一场上下文长度的军备竞赛,一路从 12.8 万到 25.6 万到 51.2 万,再到一百万。现在人人都非常具体地明白上下文的价值,你用编程模型的时候会把上下文拉满,这是个硬问题,大家都有体感。

可是在图像和视频模型这一侧,从来没有人以同样有原则的方式去推这件事。没有人去把一段一小时的视频塞进模型,再做「大海捞针」式的检索,比如把第三十七分钟的那一帧找出来。而在重建和生成这里,事情完全同构:重建其实就是上下文非常长的生成,你往里面塞了很多东西。这才是真正能把这两件事连成一条连续谱的办法。

Atlas 让我们能做到以前用 Marble 永远做不到的事。Marble 有个根本性的卡点,老实说你塞不进几张图。而在 Atlas 里,我可以拿一次六十四张图的采集,做一整栋房子的穿行,所有东西都是被锚定的,要么是拍到过的,要么是几乎拍到过的,要么是从已有内容里稍作外推的。我把过去用两千张图做的多房间采集,压到三四十张输入,穿行的效果基本一样。这在以前完全不可想象,而它全靠这个能优雅扩展、可以往里面随便倒东西的上下文窗口。

主持人: 所以可以这样理解吗:稀疏的部分是你实际拍下来的那些照片,Atlas 这个模型把其余的视图造出来,然后你再用经典的重建方法去处理。大致是这样?

约翰逊: 某种意义上是的。Atlas 的妙处在于,不管你手上有多少输入,哪怕只有一张视图,你都可以一直把 Atlas 当成渲染引擎,用它产生你想要的任何别的东西。你可以像操纵一台虚拟相机那样在里面移动:我这里有一张图,我还想要那里一张、那里一张、那里一张。你先造出几张,然后再说,好,现在给我一段稠密的穿行。你可以分步骤地做,因为它是一个自回归模型。生成的过程中,往上下文里交互式地加什么、不加什么,由你来挑。

桌子底下的那颗足球

主持人: 让我真正想不通的是这一点。我的心智模型很简单:我有四张照片,模型得在它们之间做外推,而且外推出来的东西在重建时还必须对得上,它必须是三维的。我一直觉得扩散模型这类东西是视觉上很好看但不准确。我甚至不确定这里有没有一个问题,但房间怎么就对得上了?它凭什么是三维一致的?就靠数据多吗?

约翰逊: 部分是因为我们相信缩放假设。

主持人: 顺便我必须问一句,你们着手做这件事的时候,知道它会成吗?

约翰逊: 我相当确定。

主持人: 你确定吗?

李飞飞: 我们三个人对缩放定律都有绝对的信念,这一点我觉得是有的。但具体的架构选择和数据配比,魔鬼就藏在这些细节里。我看着贾斯汀和他的团队,从「我们真的不知道这要多久」,到「好像有点生命迹象了」,再到「哇,这个要成了」。没有人做过这件事,但我们主要是对两个假设有信念:一个是缩放定律假设,另一个是下一视角预测。

约翰逊: 我确实非常笃定它会成。我不确定的是它会成得这么好、这么快。我当时想,也许我们有机会一次做成,但一个新架构、新范式的模型,第一轮预训练就直接跑通,这种事本身就很离谱。所以我估计有一定概率,我们得再转几轮预训练,才能到我们想要的质量水平。

主持人: 那现在这套架构的缩放是不是快到头了?还需要下一次突破,还是说?

约翰逊: 不不,我们处在起点。

主持人: 真的?

约翰逊: 对。

主持人: 架构不用变?

约翰逊: 对,我们基本上处在起点,而且当下的限制主要来自算力。数据当然重要,飞飞老是强调这一点,但任何事情都有瓶颈,我认为继续把这东西扩上去的主要瓶颈其实是训练算力。

开发过程中我们训了一串模型,博客里也提到了一些,就是爬缩放梯子最开始的几轮。每一次我们把模型做得更大,每一次我们训得更久,每一次我们把它放到更多芯片上,它都会明显变好。我们在博客里展示的,当然是训过的最大最好的那个,但真正卡住它的不是规模,不是数据,也不是别的什么,而是我们定了一个发布的截止日期,所以只能倒推:在这个期限之前,我们负担得起训到什么程度。

李飞飞: 这里有个内部的小故事。贾斯汀和团队从小模型往稍大的模型训,一路有路线图。今年初夏的某一天,那个模型还不是现在 Atlas 的尺寸,比它小。本和贾斯汀把数据喂进去做视角生成。你们记得那张著名的桌子吧,NeRF 论文和后来很多论文里那张花园里的桌子。那天夜里我收到一条 Slack,我们都看到了本发的那条:我们的相机从桌子底下飞过去了。

米尔登霍尔: 还有那颗足球。

李飞飞: 对,还有那颗足球。

主持人: 那颗球是涌现出来的,还是原来照片里就有?

李飞飞: 是真实存在的。

主持人: 好。

李飞飞: 那天早上我们三个人对视了一眼,说,就是它了,我们要把这个做出来。这个决定我们五秒钟就下了。因为这个结果从来没有人见过。

创作流程要的是持久状态

主持人: 本,你能不能更具体地讲讲应用场景?World Labs 一直有很多做创意的用户,他们用它来保持二维图像、电影、三维、游戏等等的一致性。这次的模型是怎么拓展了这些场景,或者更好地服务了已有的场景?之后我想聊机器人。

米尔登霍尔: 好。其实挺有意思的,我们看到人们使用 Marble 的主要方式之一,恰恰就落在新视角预测这个场景上。

主持人: Marble 是你们上一代的产品。

米尔登霍尔: 对,是我们上一代产品。很多人拿这个产品,放一张图进去,得到一个完整的三维场景,是高斯泼溅形式的,然后从不同视角截几张图,就走了。我们看着就想,那我们直接把这些图做出来不就行了,这就是生成式 AI,挺好的。而且那个过程里有不少画质损耗,大家会说,这个泼溅其实可以更好看。那好,如果我们干脆用生成的方式去建模那些视角,同时保留那种模态的控制力呢?

所以哪怕只是视图合成这个核心能力,它作为一个学术问题已经存在很久了,但那是在「你要做一次非常稠密的采集」这个前提下的。生成式的视图合成,其实是个相当新的问题。

我们看到非常多的人在创作流水线里工作。大家的工作流是多阶段的,我不认为有哪一个人是用一个大一统的模型,哪怕是 Seedance 之类,去完成整件事的。人们会先从自己偏爱的那几个图像模型里拉出一堆故事板和情绪板,然后去不同的视频工具里,把这些图当关键帧接起来,之后再去剪辑、修改。所以我们看到 Marble 有一个很小众但非常明确的用途:给你一份「理智感」,让你的生成结果能落在某个三维一致的世界里。

我自己跟各种图像模型较过劲,想让它们给我同一个房间的不同视角,每次你一看就知道,东西的位置变了,它不稳定。哪怕就这一颗用例的种子,也说明水面之下藏着价值。人们已经和持久的三维状态打了几十年交道,在虚拟环境里模拟现实中的操作,有舞台、有道具、有各种元素,无论是为了一部电影、一档节目、一张营销图,还是搭一个游戏场景。

这种有状态、可持久的特性,对人们思考空间、随时间推进地打磨一个环境,是极其关键的。人们不会用那种转瞬即逝的方式思考:生成一个,再生成一个,扔掉,只留下我的文字提示。人们想要的是攒下一批资产,用这种方式去建构一个世界。

所以我们在做的事,是通过空间上下文这套机制,以及别的一些东西,把这个层级的控制力和精度交给用户,并且让他们能调整不同模态的输入,先从相机位姿和图像开始。往前走,我们希望给人更多控制力:对场景中元素的控制、编辑、交互等等。我认为这不仅能在我们已经看到的那些领域里打开更多用途,还会扩展到任何人们需要为一个真实空间做虚拟复现,或者做一次「预想象」的地方,比如建筑和施工。

我有一次跟一个给展会搭展位的人聊天。世界上有太多东西,你平时根本不会想到它们需要被制造出来,而这里面每一样,基本上都要经过一个相当磨人的虚拟设计阶段。在这个流程里,进到三维软件的那一段,是当下最费力、最耗人工的部分。比如你从创意总监、设计总监或者建筑师那里拿到一段口头意见、一张草图、一点很随手的东西,把它映射回三维表示里,这件事占掉了百分之九十五的工作量。你开完一个会拿到反馈,回去就是一周的修改。原因就在于我们的软件已经是几十年前的东西了,它从来没有变得像玩乐高、像捏陶、像用手直接摆弄、像用铅笔画草图那样自然。

这才是 AI 真正能为人们的工作流程释放巨大价值的地方,不管是创作类的应用,还是更偏工业、偏设计的场景。这也是真正驱动我去做各种不同口味的模型、去服务这些人的原因。

机器人:从真实到仿真

主持人: 我能理解它怎么帮到创作者,因为 Marble 已经在做了,也能理解它怎么延伸到设计和建筑。你们还收购了一家机器人公司。这个聊得比较少,尤其是放在 Atlas 的语境里,它和机器人是怎么对应上的?麻烦你们展开讲讲。

李飞飞: Atlas 其实是这块拼图的关键一环。我们收购了一家公司,原来叫 Synnex。它的核心技术是什么?现在它的核心技术是一套从真实到仿真、再从仿真回到真实的系统。

在机器人的场景里这意味着什么?假设你想训练一条机械臂,让它学会在工业环境里理线。首先你需要一大堆数据,去训练一个能完成理线动作的机器人策略。然后你要评估这个策略做得好不好。最后你把机器人部署到真实的理线环境里。

为了完成训练,这家公司,也就是现在我们的机器人团队,过去做的正是本刚才说的稠密重建:你去拍一个场景的照片,然后试着把那个环境重建出来。这个过程痛苦得要命,耗时、费力,严重拖慢了机器人从真实到仿真的速度。所以 Atlas 正是这件事的下一代技术。

而且这不只关乎理线或者别的某个任务。我们应该把视野拉开,认识到机器人现在最大的问题其实是数据。总有一天会是芯片,但眼下是数据。因为要采集机器人在真实世界里运转的数据太难了。而且你不只需要采集理线、洗碗这类具体情境的数据,还有一个非常重要的步骤叫随机化:你必须把同一个环境的条件随机化。线缆不能只按这一种方式弯折,它得能以别的方式弯;盒子可以有不同的尺寸、不同的颜色、不同的盖子;东西也可以出现在场景的不同位置。所以除了能从互联网上拿到的其他数据之外,你还得走一遍真实到仿真的流程,才能凑够这类数据。这个真实到仿真的环节,会被 Atlas 极大地改善。

这只是第一部分,也就是用现有技术去满足机器人的需求,因为我们还没有一个足够稳健、能用于机器人的前沿基础模型。但 Atlas 是一个全能模型,一个多模态模型,它接受不同种类的输入,也产生不同种类的输出。你完全可以想象,下一步是让 Atlas 吃进带动态的数据,那就真的能开始弥合动作规划和机器人与 Atlas 输出之间的鸿沟。这就是大致的路线。

仿真器即规划器

约翰逊: 我想说的是,训练一个机器人策略,跟我们此前见过的任何 AI 应用都有根本性的不同。

如果你在生成一段代码、生成一张图、生成一段视频,模型本质上是在造一件成品。而这类成品,你可以在网上或者别的什么地方大量收集到。你想生成图像,世界上有大量图像;你想生成视频,有大量视频;你想生成一个代码库,也有大量代码库可以学。

机器人策略是另一回事。它产出的不是一个静态的东西,而是一个要走进世界、做出动作、试图达成目标的策略。而世界并不总是按你预期的方式回应你,意料之外的事一定会发生。所以机器人策略从根本上是一个身处真实世界、与真实世界互动的智能体,而事情总会出状况。因此关键在于,这些策略在训练时必须被暴露在部署过程中可能出岔子的每一种情况面前。仿真对机器人之所以关键,原因就在这里。

这里有两条路。一条是经典仿真:你去用自己喜欢的物理引擎,作为人类设计者去发挥想象,把完成这个任务时可能出现的所有场景想一遍,然后写出显式的代码把它们都建模出来。这是一条路,而且在有了编程智能体之后,这条路其实也被大大加速了。

另一条路是更数据驱动的仿真。我们能不能有一个学出来的模型,它理解环境、理解世界会如何回应动作,而且有时候世界的回应可能出人意料?我们能不能用尽可能多的数据训出这样的神经仿真器,再把这些学出来的神经仿真器当成训练床,去训练机器人策略?这是 Atlas 一个非常有意思的未来方向。

但事情不会止步于此。一旦你有了这样一个学出来的仿真器,它在自己的「脑子」里其实已经懂得这个世界,懂得世界会如何回应动作。那为什么仿真器本身不能成为规划器?这正是我们围绕世界模型及其通用性所持的核心论点:有一批核心的东西是模型应当理解的,包括生成世界、仿真世界、理解世界在不同情境下呈现的样子;而理解世界将如何回应一个动作,与想象「我该采取什么动作才能让世界以某种方式回应」,这两件事高度相关。

动态性与四维

主持人: 顺便恭喜发布。反响压倒性地好,我觉得这大概是今年最重要的一次模型发布,所有人都在说好话。不过我给一位这个领域的专家发过消息问他怎么看,这位很不错的人说:非常棒,很惊艳,但还需要更多的动态性。至少在机器人这边,甚至一般而言,能有一个会动的世界似乎才是理想状态。能不能聊聊这个,以及你们愿意分享、又值得一谈的其他未来方向?

约翰逊: 动态性显然会来。其实……

李飞飞: 我们已经有婴儿版的动态性了。

约翰逊: 我们确实已经有了婴儿版的动态性。这一点我觉得大家没太注意到,我们在博客里也没重点讲。上一代 Marble 世界模型从根本上是静态的,那个模型完全处理不了任何动态,这是烙进模型架构、烙进训练方式里的,整个东西从根上就是静态的。

Marble 之后我们就知道这是个大问题,而且在 Atlas 里已经修掉了。Atlas 的架构从根本上就支持动态,Atlas 的训练数据从根本上就包含动态。如果你仔细看我们发出来的一些视频……

主持人: 确实有。

约翰逊: 你看出来了。

主持人: 对,水面的波浪。

约翰逊: 对,有些例子里水面有波浪,有些生成的航拍视角里有小汽车在跑。所以动态其实已经在这个模型里了。

主持人: 不过顺便问一句,如果你要从多个视图重建三维,动态在我看来是很麻烦的事。这两者是不是有冲突?

约翰逊: 我们这里的一个论点恰恰是:如果你要做的是纯粹的三维重建,你其实希望没有动态,你希望能在时间被冻结的前提下精确地建模场景的视图。但这正是我们之前 Marble 路线的一个问题:你可以去找完全静态的数据,可这类数据很难扩量,也很难拿到更多。

我们意识到的是,哪怕你最终想要的是静态的输出,得到它最好的办法反而是让模型见识动态。你有多少动态素材就喂多少,有多少静态素材就喂多少,然后让模型自己学会把动态的部分剥离出去。

所以在 Atlas 的预训练里,模型已经见过海量的动态。而我们针对这次发布的这个检查点所做的后训练,重心更多放在静态上,更多放在空间移动而不是时间上。但我们已经有了,我相当确定这个预训练检查点里已经潜藏了大量动态。这也是我们接下来会大幅改进的方向。

主持人: 那么本,这是不是意味着我们会有四维视频?我可以走进去逛?

米尔登霍尔: 我觉得……

主持人: 看他们脸上的笑就知道了。

可编辑性是关键

主持人: 我其实觉得,就算你们现在停下来,只是做得更大、更快、更好,也足以撑起一整个产业。它像是一个非常横向的原语。那么除了「更大更快」之外,还有什么让你们兴奋的东西?尤其是在你关注的那些应用上,也就是偏内容创作、偏三维的这一侧。

米尔登霍尔: 我非常想推进的是多模态那一面。我认为不同的控制方式在这里极其关键。给这些模型加上控制条件,把里面的东西「取」出来,这件事有多重要被严重低估了,尤其是在学术圈。老实说这……

主持人: 不好意思,我完全不懂这几个词是什么意思。

米尔登霍尔: 翻译成大白话就是可编辑性。我认为可编辑性才是关键。

我们在单图模型上看到过这种东西,今年开始在视频模型上也被解锁了一部分:多轮对话式的编辑,或者说非常符合直觉地理解「我想要这个人、这个物体、这件事发生」,然后把它们拼成一幅完整的画面,而不需要你在系统里做大量手工活。它就是能理解你的意思。前沿图像模型在编辑这件事上基本已经到位了。

但我们还没看到这种能力同样有力地传导到视频,再传导到世界模型上。我们看到的是一些相当玩具化的例子,比如在那种实时模型里输入一句话,然后冒出来一只恐龙。我想把这件事做到真正的工业强度。因为这里的诀窍在于,你得加上控制,同时不能牺牲模型的质量,否则它就只是个party trick。不会有人认真考虑放弃自己正在用的前沿视频模型,换成你这个多给几个旋钮、画质却变差的模型。

所以这真的是一场博弈:如何守住我们在当前模型输出上立下的那条高线,同时把人们会向我们提出的各种有意思的东西加进来,比如我想和场景交互、控制布局、控制画面里物体和事物的身份,或者控制时间。我觉得这是一个能打开大量有趣产品和交互设计工作的维度。你在这里加的复杂度和丰富度越多,就越有可能几乎从零开始重新设计人们在计算机里与有状态的三维世界打交道的方式。这才是最终目标,就是把搭起那样一套系统所需要的全部能力都拿到手。

主持人: 很好。你这边呢?有没有什么不只是「更大更好」的新功能是你期待的?

李飞飞: 对我来说,还是回到智能的第一性原理。当涉及空间、涉及物理空间时,智能不是坐在那里不动,只是看到什么、解释什么。它真正的内核是把「看」「体验」和「交互」这个回路闭合起来。所以顺着这个阶梯往上走,正是本刚才说的那件事。

有眼睛的是动物

主持人: 很好。我觉得这里有一个有意思的概念,叫 AI 完备(AI completeness)。你听过吗?

李飞飞: 听过。

主持人: 顺便说,我听到的 AI 完备是在大语言模型的语境里,意思大概是:你得是最聪明的那个大语言模型,才能回答最聪明的大语言模型需要回答的问题,或者说你必须先解决通用智能。

约翰逊: 不是,它其实是跟图灵完备类比来的。在经典复杂性理论里,一个任务是图灵完备的,意思是我可以把这一类里的任何问题都归约到这一个问题上。3SAT 是经典例子,你可以把任何 NP 难问题归约成 3SAT,因此你可以用 3SAT 去解任何问题。

那么 AI 完备的宽松定义就是:存在这样一个根本性的原语,它本身是一个 AI 任务,但如果我能在它最广义、最完整的意义上解决这个任务,我就解决了任何智能问题。大语言模型这边的经典例子是,下一个词元预测是 AI 完备的。我记得伊利亚讲过那个著名的例子:有一本推理小说,模型得把整本读完,而小说的最后一句是「凶手就是……」,预测下一个词元。

所以你基本上可以把任何智能任务都改写成这种形式。显然,很多人相信下一个词元预测是 AI 完备的。但我们逐渐意识到一件事,本今天早些时候也在说,就是新视角预测,我们在 Atlas 里拿到的这个原语,尤其是生成式的新视角预测,它同样是 AI 完备的。

主持人: 你可以拿一部电影,把所有帧都给它,然后凶手走出来,你要预测走出来的到底是谁。

约翰逊: 正是如此。不止如此,我还可以说,我要一个世界,里面的你正在黑板上写下黎曼猜想的证明。

李飞飞: 那我们换一个演化的视角来看。新视角预测,恰恰是演化通过让动物动起来而必须解决的问题。自然给了动物眼睛,却没有给树眼睛。为什么?因为当你移动的时候,你就会看到一个新的视角。这件事,无论你把它叫作 AI 完备还是智能完备,我们都非常强烈地相信,下一视角预测就等价于下一词元预测。

主持人: 太精彩了。那就在此祝贺各位这次出色的模型发布,我们非常期待你们接下来的模型,谢谢各位前来。

李飞飞: 谢谢。

本期讲者
李飞飞斯坦福大学计算机科学教授、斯坦福以人为本 AI 研究院联合院长,主持构建了推动深度学习兴起的 ImageNet 数据集。她是 World Labs 联合创始人,主张下一步的核心是空间智能。
贾斯汀·约翰逊World Labs 联合创始人,负责模型研发,此前在斯坦福参与计算机视觉教学与研究,长期从事图像生成与视觉表示方向的工作。本场由他讲解 Atlas 的架构与训练。
本·米尔登霍尔World Labs 联合创始人,NeRF(神经辐射场,2020)论文的第一作者,该工作重塑了新视角合成与三维重建领域。本场由他讲解稀疏重建、可编辑性与创作者工作流。
章节 · 点击跳转视频
0:00 冷开场:三台相机做出子弹时间 ▶ 正在看
1:05 Atlas 能做的三件事 ▶ 正在看
2:53 新视角预测:一个新的基础原语 ▶ 正在看
8:15 空间智能是什么,走到了哪一步 ▶ 正在看
10:27 从 Marble 到 Atlas:为何不能一步到位 ▶ 正在看
13:32 稠密重建之苦:三百张降到三张 ▶ 正在看
17:27 重建就是超长上下文下的生成 ▶ 正在看
20:49 他们凭什么相信这条路会成 ▶ 正在看
24:19 创作者工作流与三维软件之痛 ▶ 正在看
29:22 机器人的瓶颈是数据,不是芯片 ▶ 正在看
35:21 动态、可编辑性与下一步 ▶ 正在看
40:56 AI 完备:新视角对标下一个 token ▶ 正在看
本期论点
本期回应
42:08
生成式新视角预测是 AI 完备的,任何智能任务都能被框定成预测下一个视角 都能算分什么都能拿来打分优化吗?贾斯汀·约翰逊
其他论点
0:30
新视角生成让子弹时间镜头从几百台摄像机降到三台,不再需要影棚、绿幕和昂贵标定 观察李飞飞
1:26
最多一百帧真实世界视角就足以重建出可自由穿行的三维空间 观察李飞飞
7:19
把像素生成与像素重建统一进同一个模型,终结了计算机视觉半个多世纪的分支割裂 李飞飞
10:15
生成在空间中被锚定、带空间上下文的像素,比单纯生成像素是更难也更关键的跨越 李飞飞
11:54
把高斯泼溅当作唯一输出表示,会成为世界模型能力的瓶颈 贾斯汀·约翰逊
12:14
世界模型最根本的原语应该是新视角预测,而不是生成一个完整的三维世界 贾斯汀·约翰逊
18:25
视角再密集也必然有拍不到的地方,三维重建必须靠生成能力补全空洞 贾斯汀·约翰逊
19:18
重建本质上就是超长上下文下的生成,重建与生成处在同一条连续谱上 本·米尔登霍尔
22:56
空间智能这套架构远未到 scaling 尽头,当前主要瓶颈是训练算力而不是数据 贾斯汀·约翰逊
28:35
3D 软件几十年来始终没有变得像捏陶、搭乐高或铅笔画草图那样直观 观察本·米尔登霍尔
30:51
机器人领域当前最大的瓶颈是数据,而不是芯片算力 李飞飞
34:56
学出来的神经模拟器既然已理解世界如何回应动作,就应该同时充当规划器 李飞飞
37:16
即使最终只要静态输出,也该让模型大量接触动态素材、自己学会把动态分离出去 做法贾斯汀·约翰逊
42:46
新视角预测是演化在动物运动时必须解决的问题,所以会动的动物有眼睛、树木没有 李飞飞
01冷开场:三台相机做出子弹时间
0:00
On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. >> You know, LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. >> This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100 extra reduction. There's a famous shot in the first Matrix movie where Neo is like falling down. Exactly.
在通往空间智能的道路上,要生成真正具备空间上下文、真正有依据的像素,这是 Atlas 迈出的非常艰难的一步。>> 你知道,大语言模型(LLM)建立在「下一个 token 预测」之上。我们看到视频模型建立在「下一帧预测」之上。而 Atlas 真正做的是「新视角预测」。>> 这才是 AI 真正能为人们和他们的工作流程释放巨大价值的地方。我们说的是大概 50 倍、100 倍的额外成本削减。第一部《黑客帝国》里有一个著名镜头,Neo 向后倒下。没错。
便签引用
0:27
They had hundreds of cameras doing that angle on a green screen. With Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. >> No one has ever seen these results. >> When you set out to do this, did you know it was going to work? >> I was pretty sure. Each time we made the model bigger, and each time we trained it for longer, it got significantly better. >> Does that mean we're going to get 4D video? Can I go walk around? >> So, big day yesterday, you launched a new frontier model, which has got amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So, maybe Justin, do you want to talk about what was launched yesterday, why it's significant.
他们在绿幕前用了几百台摄像机来拍那个角度。有了 Atlas,我们只用三台摄像机就能做到。不需要影棚拍摄,不需要绿幕,也不需要昂贵的标定。>> 从来没有人见过这样的结果。>> 你们着手做这件事的时候,知道它一定会成功吗?>> 我当时挺有把握的。我们每次把模型做得更大,每次训练时间更长,效果都会显著变好。>> 那这是不是意味着我们会有 4D 视频?我可以到处走动吗?>> 昨天是个大日子,你们发布了一个新的前沿模型,反响非常好,而且反馈还在陆续涌来。我想,也许组织这次对话的一个好方式是,先聊聊具体发布了什么,然后我们再回到历史,一路讲回来。所以,Justin,你想先说说昨天发布了什么、为什么它意义重大吗?
便签引用
02Atlas 能做的三件事
1:05
>> Yeah, so Atlas is our new next generation world model. Um, it has three basic things. It can generate, reconstruct, and simulate the world. Um, so within that, there's a couple different major capabilities. It has really good camera condition generation. So, you can input an image together with a camera trajectory with a camera trajectory and steer the model and have it generate, you know, video frames along any any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to 100 frames, um, that are views of the real world, and use those to reconstruct the real world. And that reconstruction can take the case either of of novel of video flying through the space, or an explicit 3D reconstruction of the space.
>> 好的,Atlas 是我们的新一代世界模型。嗯,它有三个基本能力:它可以生成、重建和模拟世界。嗯,在这之中又有几项主要能力。它的相机条件生成做得非常好。所以你可以输入一张图像,加上一条相机轨迹,用相机轨迹来引导模型,让它沿着你想要的任意视角生成视频帧。它非常擅长稀疏三维重建。你可以输入一帧或多帧,最多 100 帧真实世界的视角画面,用它们来重建真实世界。而这个重建的结果既可以是穿梭于空间中的新视角视频,也可以是对该空间的显式三维重建。
便签引用
1:41
Um, then finally, it can be used for simulation. Um, and for this, we show off, um, you know, these awesome bullet time videos, which got a lot of attention online, and then also robotics simulation. >> What's a What's a bullet time video? >> A bullet time video, this comes from the Matrix. You know, there there's a famous shot in the first Matrix movie where Neo is like falling down. >> [laughter] >> Oh, yeah. >> Exactly. So, and then remember in that famous shot he's like falling down, it's in slow motion, and the camera flies all the way around. Um so, that's the way that they did that shot is they had a ring of like hundreds of cameras. So, then like he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that that famous shot in in Matrix. But now with Atlas we can do this with just as few as through three cameras. So, like no studio capture, no green screen, no expensive calibration.
嗯,最后,它还可以用于模拟。嗯,为此我们展示了,嗯,那些很棒的「子弹时间」视频,它们在网上获得了很多关注,另外还有机器人仿真。>> 什么是「子弹时间」视频?>> 「子弹时间」视频,这个说法来自《黑客帝国》。你知道,第一部《黑客帝国》里有个著名镜头,Neo向后倒下。>> [笑声] >> 哦,对。>> 没错。然后记得在那个著名镜头里,他向后倒下,画面是慢动作,镜头绕着他飞快地的反过来。嗯,所以他们当时拍那个镜头的方式,就是布了一圈差不多几百台相机。所以,然后他在摄影棚里往后倒的时候,有几百台相机从那个角度、在绿幕前拍摄,然后他们用这几百台相机做出了《黑客帝国》里那个著名的镜头。但现在有了 Atlas,我们只用少到三台相机就能做到这件事。也就是说,不需要影棚拍摄,不需要绿幕,也不需要昂贵的标定。
便签引用
2:25
We can literally stick like three cameras on three iPhones on tripods, um use these to take sort of a video of something happening, um like someone shooting a basket, someone dropping a strawberry into a bowl of milk. And then from those three like iPhone videos, we can then reframe the shot, and imagine like a like freeze time, have the camera fly in like as the milk is splashing up, and get these amazing frozen time views. Um and we can do this with just uh just a couple cameras. >> Can Can you um just maybe What is the simplest description of what Atlas does?
我们真的可以把三台相机、三部 iPhone 架在三脚架上,嗯,用它们拍一段正在发生的事情的视频,比如有人投篮,或者有人把一颗草莓丢进一碗牛奶里。然后从这三段 iPhone 视频里,我们就可以重新构图,想象一下把时间冻结,让镜头在牛奶飞溅的那一刻飞进去,得到这些非常惊艳的时间静止视角。嗯,而我们只需要几台相机就能做到。>> 你能不能,嗯,就是也许——对 Atlas 做的事情,最简单的描述是什么?
便签引用
03新视角预测:一个新的基础原语
2:53
Like what goes in and what comes out? >> Yeah, so one of the one of the really core principles of Atlas, like the most fundamental thing, is it does new view prediction. Um and this is a a really fundamental primitive that we think is super exciting, a super new primitive for for base models that no one's ever done before. Right? So, we know LLMs are built on next token prediction, we've seen video models as being built on next frame prediction. Atlas is really new view prediction. Right? That given some number of views of a scene or a description of a scene, um those go into what we call a spatial context that describes implicitly what is the world that we want to talk about, then you can point a virtual camera at at an arbitrary point in space and time, and Atlas will understand what that world is supposed to look like from that position in space and time.
就是输入是什么,输出是什么?>> 对,Atlas 最核心的原则之一,最根本的一点,就是它做的是新视角预测。嗯,这是一个非常基础的原语,我们觉得它特别令人兴奋,是一个非常新的原语,是此前没有人在基础模型上做过的。对吧?我们知道 LLM 是建立在下一个 token 预测上的,我们也看到视频模型是建立在下一帧预测上的。而 Atlas 真正做的是新视角预测。对吧?就是给定一个场景的若干个视角,或者一段对场景的描述,这些内容会进入我们所说的空间上下文,它隐式地描述了我们要讨论的是一个什么样的世界,然后你就可以把一个虚拟相机对准空间和时间中的任意一点,Atlas 就能理解从那个时空位置看过去,这个世界应该是什么样子。
便签引用
3:34
>> You know, you know, Ben, that you know, with with a with a bajillion video models out there all claiming to be world models and all claiming to have novel views, and can you maybe tease apart can more concretely how this is different from like the myriad models that have come before? >> Mhm. Yeah, I think what Justin was saying about the spatial context aspect is super important here. So there's many video models, a lot of video models actually got their claim to fame from their single dimension input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omni referencing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of like specially grounded meaning to every frame you put into it. So it's not just an image that the model's going to interpret whatever way it wants or you can kind of try to argue with it in the the text prompting and get it to do something specific. With Atlas, every image actually has an associated
>> 你知道,Ben,现在外面有成千上万的视频模型,都声称自己是世界模型,都声称能做新视角,你能不能更具体地拆解一下,这和此前那一大堆模型到底有什么不同?>> 嗯,对,我觉得 Justin 刚才讲的空间上下文这部分在这里特别重要。现在有很多视频模型,很多视频模型其实是靠单一维度的输入,或者靠首帧到尾帧的插值出的名。现在我们开始看到一些模型可以做那种全方位参考,用上二三十张、五十张图片。但 Atlas 的关键在于,它对你输入的每一帧都有一种真正在空间上被锚定的含义。所以它不只是一张图片、让模型爱怎么理解就怎么理解,或者你得在文本提示里跟它反复较劲,才能让它做出某个具体的东西。在 Atlas 里,每一张图片都有一个与之关联的三维相机位姿,这意味着你可以以极高的精度
便签引用
4:25
three-dimensional camera pose and that means that you can perform this task of reconstruction with an extremely high degree of precision, right? So if we had four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room and it's not going to guess what's in the other corner, like the relationship between things. It's just going to reproduce exactly what you give it. And you can also do that in a kind of creative or imaginative sense, too. If you take two photos from different, you know, AI generations or real-world locations, you can actually position and stage those to build these kind of intentionally directed fly-throughs that are really governed by exactly the precise place that you put the content you want and where the camera's going to look and travel, which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that
完成重建这个任务,对吧?比如我们拍这个房间的四个视角,每个角落一个,你把它们输入模型,就能得到这个房间里你所看到的一切的精确复现,它不会去猜另一个角落里有什么、东西之间是什么关系。它就是会精确复现你给它的内容。而你同样也可以用一种创意或想象的方式来做这件事。如果你拿两张来自不同 AI 生成结果或者真实地点的照片,你其实可以对它们进行定位和布置,来构建那种有意设计的穿行镜头,完全由你把想要的内容精确放在什么位置、以及相机往哪里看、怎么移动来决定。我觉得这和视频模型那种更像老虎机的效果非常不同——在视频模型里你只有比较高层的文本控制,只能一遍又一遍地
便签引用
5:12
kind of higher level of text control you get with video models. >> Is this just kind of an obvious, you know, scaled-up version of a traditional video model or is it a new architecture? >> I think it's it's a pretty new thing for a couple of different reasons. One that we talk about is it does both generation and reconstruction jointly in the same model. Like Ben was saying, this thing can take a couple of views of this room and then reconstruct everything in this room exactly as you see it. And historically, reconstruction has been its own subfield in computer vision with its own specialized task, its own specialized models. And generation is what all the text all the text to video models are really good at like what all the all the big diffusion models we've seen the last couple years.
重新生成、碰运气。>> 这只是传统视频模型一个顺理成章的放大版,还是说是一种新架构?>> 我觉得它是相当新的东西,有几个不同的原因。我们常说的一点是,它在同一个模型里联合完成生成和重建。就像 Ben 刚说的,这个东西可以拿这个房间的几个视角,然后把房间里的一切按你看到的样子精确重建出来。而在历史上,重建一直是计算机视觉里一个独立的子领域,有它自己专门的任务,自己专门的模型。而生成则是所有文生视频模型真正擅长的事,就是我们过去这几年看到的那些大型扩散模型。
便签引用
5:51
And those are great for creative applications. I want to imagine something that's never been there before. Um but now with Atlas for the first time we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture. So to do that we had to make a couple changes. Um one is we had to make it multimodal from the start. So this thing natively works on text, it works on images, it works on videos. It also works on camera poses um as a native input to the model which I don't think anyone's ever done at the pre-training phase before. Um and it uses um it uses 3D as a native modality that it works on. So this thing from the beginning was designed to be natively multimodal in a way that no one else I think >> Sorry, I just um I don't know this space super well but 3D is this like depth or models and like what is that what is that mean?
那些模型很适合创意类应用——我想象一个从来不存在的东西。嗯,但现在有了 Atlas,我们第一次把视觉智能的这两个不同部分放进了同一个模型里。所以它可以在一个架构里同时做三维重建和生成。为了做到这一点,我们做了几处改动。嗯,一是我们必须从一开始就让它是多模态的。所以这个东西原生支持文本,支持图像,支持视频。它还原生支持把相机位姿作为模型输入,我觉得这是此前没人在预训练阶段做过的。嗯,而且它把 3D 当作一种原生模态来处理。所以这个东西从一开始就被设计成原生多模态的,这种方式我觉得没有别人—— >> 抱歉,我对这个领域不是特别了解,但3D 是指深度之类的吗?这具体是什么意思?
便签引用
6:38
>> Yeah, so the formulation we used so far is depth maps. >> Okay. >> Right? So um right now you can have a when you have a frame that has a a a virtual camera telling its position in 3D space, yeah, that camera position and and camera parameters are a native input to the model. And then what attached to that camera position you can have both um like RGB telling you what is that what is that position in space look like and you can have a depth map that tells you what is the spatial structure of that position in 3D space. So then you know text, image, video, 3D cameras are these modalities that this thing all does jointly in a multimodal way.
>> 对,我们目前采用的形式是深度图。>> 好的。>> 对吧?所以,现在当你有一帧画面,它带有一个虚拟相机来说明它在三维空间中的位置,那么这个相机位置和相机参数就是模型的原生输入。然后附着在这个相机位置上的,既可以有 RGB,告诉你那个空间位置看起来是什么样,也可以有一张深度图,告诉你那个位置在三维空间中的空间结构是怎样的。所以文本、图像、视频、3D 相机,这些模态都是这个模型以多模态方式联合处理的。
便签引用
7:10
>> I want to add something cuz I think what Justin just said is actually so important and also what Ben said that it's under appreciated. It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century. Um sitting here having been in this field for decades, I cannot tell you how many PhD thesis have been written on the problem of reconstruction or novel view things synthesis and also our field traditionally uh have multiple tracks. You go to a computer vision conference, you have the pixel generation track, you have some recognition track and you have 3D reconstruction track. This is a elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and the the viewpoint estimation and that's just incredibly powerful.
>> 我想补充一点,因为我觉得 Justin 刚说的其实非常重要,Ben 说的也一样,而这一点被低估了。这是我们第一次实现像素生成和像素重建的统一。在计算机视觉领域,这个方向已经存在了半个多世纪。嗯,作为一个在这个领域待了几十年的人坐在这里,我没法告诉你有多少篇博士论文是围绕重建或新视角合成这个问题写的,而且我们这个领域传统上有多个分支。你去参加一个计算机视觉会议,会看到像素生成的分会场,识别的分会场,还有三维重建的分会场。而这是一个优雅的模型,它通过锚定视角和视角估计,把重建和生成这两个问题结合、或者说统一了起来,这非常非常强大。
便签引用
04空间智能是什么,走到了哪一步
8:15
>> Can Can you Can you maybe Well, can we take a step back and then maybe you just fill something out? So, when when you started the company, I remember you saying uh you know, you want to tackle um spatial intelligence, right? And uh you know, now we have this new model. And so, it feel I mean, like as a layperson, it feels very general to me. You've got next product prediction and this is next new view prediction. >> New view prediction, right? >> So, like you can get one view out of a set of views and you have a new view.
>> 你能不能——也许我们退一步,然后请你再补充一些?就是说,你们创立公司的时候,我记得你说过,你们想攻克空间智能,对吧?而现在我们有了这个新模型。所以,我作为一个外行,感觉它非常通用。之前有'下一个 token 预测',而这个是'新视角预测'。>> 新视角预测,对。>> 就是说,你可以从一组视角里得到一个新的视角。
便签引用
8:40
Can maybe you pencil out like like how this is a significant step to this general problem of spatial intelligence? And maybe by like starting to describe what spatial intelligence is? >> Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it. >> Yeah. >> Now, we can argue is it 3D or 4D? Ultimately, it's 4D with the time dimension, but even just 3D, these are the fundamental tasks that one has to do or spatial intelligence has to enable.
你能不能大致说一说,这对空间智能这个总体问题来说是一个多重要的进展?也许可以先从描述什么是空间智能开始?>> 嗯,空间智能最终必须让我们既能生成空间本身,又能在其中进行推理,还能在其中编辑和交互。>> 对。>> 当然,我们可以争论它到底是 3D 还是 4D。归根结底它是带时间维度的 4D,但哪怕只是 3D,这些也都是必须要做到的基本任务,或者说是空间智能必须支持的。
便签引用
9:16
And then we talk about with that, you can render, you can simulate, and you can plan actions. But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space. >> Yeah. >> And I do believe Atlas is a significant step forward because now with every single frame, you have a you can generate an estimate a important piece of information, which is the the view viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the of the space. And that can lead to all the emergent behaviors we see in our in the downstream of the model, which we showed in the blog. So So in the on the path to spatial intelligence, generating pixels is definitely a a early step, which we have seen with what you call it gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step. And that is the very hard step that Atlas has taken.
然后在此基础上,我们才谈得上渲染、模拟,以及规划行动。但要做到这些,一个必须解决的根本问题是理解空间的几何、结构和物理。>> 对。>> 我确实相信 Atlas 是一个重要的进步,因为现在对每一帧,你都能生成、估计出一条重要的信息,也就是视角、相机位姿。而这正是关于空间几何最关键的信息。这也就带来了我们在模型下游看到的所有涌现行为,我们在博客里展示过。所以在通往空间智能的路上,生成像素肯定是很早的一步,我们已经在你说的那无数个模型里看到了。但生成真正具备空间上下文、在空间上被锚定的像素,绝对是另一个重大的一步。而这正是 Atlas 迈出的那一步非常难的一步。
便签引用
05从 Marble 到 Atlas:为何不能一步到位
10:27
I we definitely have, you know, we can just keep going here, right? Like there is the fourth dimension of time, which will bring in dynamics. And there is more higher fidelity simulation and delineation of the space. So this is part of the the road map of spatial intelligence. >> Great, yeah. I mean, I definitely want to like dig into like where this is going. But first, maybe let's talk about getting here. How long has World Models been in existence? >> Two and a half Two and a half, yeah. >> And so you you've actually released models before. So what why didn't you just jump right to Atlas?
我们当然还可以继续往下走,对吧?比如还有时间这第四个维度,它会带来动态。还有更高保真度的模拟和对空间的刻画。所以这是空间智能路线图的一部分。>> 很好,是的。我确实很想深入聊聊这件事会走向哪里。但首先,也许我们先聊聊是怎么走到今天的。World Labs 成立多久了?>> 两年半,两年半,对。>> 你们其实之前也发布过模型。那为什么没有直接一步到 Atlas 呢?
便签引用
11:06
>> [laughter] >> Good question. >> It's so magic, right? Yeah. >> Justin's team needs a lot of chips. >> Yeah, we need a lot of GPUs to actually [laughter] scale this thing up. So what one of the last year we released our Marble World model, and that was the first kind of big major World model that we put out. That that powers our current Marble product. And Marble and Marble is really cool. Marble can take images, it can take videos, it can take text prompts, and use these to generate 3D worlds. But one of the biggest differences between Marble and Atlas is exactly what is that output modality.
>> [笑] >> 好问题。>> 这太神奇了,对吧?是啊。>> Justin 的团队需要很多芯片。>> 对,我们需要很多 GPU 才能真正把这个东西 [笑] 规模化。去年我们发布了 Marble 世界模型,那是我们推出的第一个真正意义上的大型世界模型。它支撑着我们现在的 Marble 产品。Marble 真的很酷。Marble 可以接受图片,可以接受视频,可以接受文本提示,用这些来生成 3D 世界。但 Marble 和 Atlas 之间最大的区别之一,正是输出模态是什么。
便签引用
11:33
So, marble was really focused on Gaussian splats as an output representation. So, whatever you're inputting, um it's going to output a 3D world represented as a as a Gaussian splat. And Gaussian splats are really useful, right? They're really nice, they're easy to render, they are they they can they can render efficiently on mobile devices, on VR devices, they can interoperate with other with a with game engines, with simulation engines. There's a lot of nice things about Gaussian splats. But, you know, I think that was kind of a bottleneck in the previous marble model. So, what we did with Atlas is redesign the thing bit. Um and we realized that we need to bifurcate these modalities earlier and actually have these things all these modalities working in a more unified way in the model. So, now with Atlas, um the the the fundamental primitive is not like a good generate a Gaussian splat world. The fundamental primitive is is as we said new view prediction. Um and that can generate RGB frames, that can
Marble 非常聚焦于把高斯泼溅(Gaussian splat)作为输出表示。所以不管你输入什么,它输出的都是一个用高斯泼溅表示的 3D 世界。高斯泼溅确实很有用,对吧?它们很棒,容易渲染,可以在移动设备、VR 设备上高效渲染,还可以和游戏引擎、仿真引擎互相打通。高斯泼溅有很多好处。但我觉得那在之前的 Marble 模型里某种程度上成了一个瓶颈。所以我们在 Atlas 上做的是把这个东西重新设计了一下。嗯,我们意识到我们需要更早地把这些模态分叉出来,让所有这些模态在模型里以一种更统一的方式运作。所以现在有了 Atlas,最根本的原语不再是'生成一个高斯泼溅世界'。最根本的原语就是我们说的新视角预测。嗯,它可以生成 RGB 帧,可以
便签引用
12:19
generate 3D, and we can use those to generate a beautiful Gaussian splat worlds when you need them. Um but, we don't need to bottleneck our our our outputs through the Gaussian splats when we don't need to. And that was a that actually took a lot of, you know, blood, sweat, and tears to understand like what are all the pros and cons of these different representations. So, that's one part of it. Um the other part is you got to like climb the scaling ladder, right? You got to like work your way up and like do smaller experiments, do smaller models like to build your conviction on what's going to work and what's going to scale. Um and there's there's you know, if you could instantly know the right thing that's going to scale, you know, you should just do that. But, when we when we started [laughter] the company, the world was a very different place. Like there's there's no scaling law of special intelligence. Right. So, like when we started the company, like the world was in a very different place, the tech was
生成 3D,而当你需要的时候,我们可以用这些去生成漂亮的高斯泼溅世界。嗯,但在不需要的时候,我们不必把输出都挤过高斯泼溅这个瓶颈。而要搞清楚这些不同表示各自的优缺点,其实真的付出了很多心血。所以,这就是这只是其中一部分。呃,另一部分是你得去爬那个规模化的阶梯,对吧?你得一步步往上走,先做一些小的实验、小的模型,来建立你对什么行得通、什么能规模化的信心。呃,你知道,如果你能一下子就知道哪条路能规模化,那你直接去做就行了。但是,我们刚创立公司的时候(笑),世界跟现在完全不一样。比如说,当时根本没有所谓空间智能的规模化定律。对吧。所以我们创立公司的时候,整个世界处在一个非常不同的阶段,技术也处在非常不同的阶段。我们对它未来的走向有很多野心,
便签引用
13:04
in a very different place. We had a lot of ambitions for where we wanted it to go, but it took a couple it took a a couple iterations for us to hit upon this formulation that we thought is actually is actually like this is the one. This is this is the one that can scale out. >> You know, Ben, you know, being the creator of Nerf and doing a lot of 3D and reconstruction, so it's not so obvious to me that like if you have multiple views that you actually end up with a 3D thing. But, like you've kind of like made a career of ending up with a 3D thing. So, maybe talk a little bit about like kind of that step.
但我们迭代了好几轮,才找到这个我们认为真正对了的方案——就是它了,就是这个,这个是能规模化的。>> 你知道,Ben,作为 NeRF 的作者,做了大量三维和重建的工作,所以对我来说并不是那么显然:如果你有多个视角,最后就真能得到一个三维的东西。但是,你的职业生涯某种程度上就是在做"最终得到一个三维的东西"这件事。所以,也许可以稍微聊聊这一步。
便签引用
06稠密重建之苦:三百张降到三张
13:32
>> Yeah. Yeah. Um yeah, I mean as you said, I've spent many, many years of my career as a mass majority of my career actually working on producing 3D things from images. Um and this is actually something we we talked about a lot early on in the company even of like is this going to be the approach that produces 3D, right? Are we going to synthesize multiple views and then build 3D out of that? Are we going to try to go direct to 3D? Like there's been a lot of um uncertainty in the field around like which of those approaches kind of will win out or will kind of like reap the best advantages earlier on. Um but I did have a lot of conviction just from seeing the kind of power of what I would almost call the brute force scaling scaling at a very, very, very small baby scale, not like real model scaling, but the scaling of dense reconstruction that we had seen happening over the past 3 years before.
>> 是的,是的。呃,是啊,就像你说的,我职业生涯中的很多很多年,其实是绝大部分时间,都在研究怎么从图像生成三维的东西。呃,这其实也是我们公司早期就讨论了很多的问题,就是:这会不会是那条能产出三维的路径?我们是要先合成多个视角,再从中构建三维?还是直接做到三维?这个领域里一直有很多不确定性,就是这几种路线到底哪种会胜出,或者哪种能更早拿到最大的好处。呃,但我确实有很强的信心,因为我看到了那种我几乎可以称之为"暴力规模化"的力量——只不过是在非常非常小的、婴儿级别的规模上,不是真正意义上的模型规模化,而是稠密重建的那种规模化,那是我们在此前三年里亲眼看到发生的事情。
便签引用
14:17
Um so, basically we put up >> What Why is dense reconstruction Yeah. dense? >> Yeah, dense. >> Dense dense is >> Because I know we're going to talk about sparse and I want to make sure that people understand what is dense and what is sparse. >> Yeah, so I think this is actually even on the kind of like business and commercial side I I think been one of the challenges of productizing uh 3D reconstruction technology like at a fundamental level, right? People kind of don't uh in in a in a casual sense like you think I took three photos of this object or I took six photos of this room. Like I I look at the photos, I can understand in my mind like how this piece together. I can kind of fill in the gaps and get it, but there's just never been really any kind of reconciliation between those like really data-driven priors and then the kind of brute force dense reconstruction, which it actually is much more akin to almost like scientific or medical imaging what we did in dense reconstruction, right? You basically
呃,所以基本上我们就搭了 >> 等一下,为什么稠密重建是……是啊。"稠密"?>> 是的,稠密。>> 稠密,稠密指的是 >> 因为我知道我们接下来会聊"稀疏",我想确保大家明白什么是稠密、什么是稀疏。>> 好,我觉得这一点其实连在商业和产品层面上,都一直是把三维重建技术产品化的根本挑战之一,对吧?人们在日常的直觉里不太能理解,你会想:我给这个物体拍了三张照片,或者给这个房间拍了六张照片。我看着这些照片,脑子里就能明白它们是怎么拼在一起的。我能自动补全那些空缺,看懂它,但在这种非常依赖数据先验的能力和暴力式的稠密重建之间,一直没有真正的调和。我们做的稠密重建,其实更接近科学成像或医学成像那一类东西,对吧?你基本上必须说:我希望出现在这个重建里的每一样东西,我都至少需要
便签引用
15:08
have to say every single thing I want to appear in this reconstruction, I need at least three or four views of it. And if you think about that, like even just in this room, right? There's like under the microphone, under the table, between every different crack and crevice and the plant leaves, right? To actually truly get a picture that covers every one of those spots, it's this like very tedious and exhaustive effort to walk around the room. I think you've all seen me running around various places like capturing them. It takes, you know, for for someone who's well-trained per se, like it can take minutes, but if you hand a casual consumer or even some like kind of professional trying to do this for the first time, an average like cell phone camera or capture device, like it's going to take them probably an hour. I've seen someone for the first time trying to scan a multi-room environment spend like two hours walking through it and get enough coverage. And that's just this like very, very exhaustive and tedious loop.
三四个视角。你想想看,就拿这个房间来说,对吧?麦克风底下、桌子底下、每一道缝隙和角落、植物的叶子中间,对吧?要真正拍到覆盖每一个这样的位置,那就是一件极其繁琐、极其穷尽的活儿,得绕着整个房间走一遍。我想你们都见过我在各种地方跑来跑去做采集。对于一个受过良好训练的人来说,可能要花几分钟;但如果你把一台普通手机或采集设备交给一个普通消费者,甚至是某个第一次做这件事的专业人士,那大概得花上一个小时。我见过有人第一次扫描一个多房间的环境,花了差不多两个小时到处走,才拿到足够的覆盖。这就是一个非常非常穷尽而繁琐的循环。
便签引用
15:59
And so yeah, when we say dense, we really mean dense. It's like this room I want >> You just need like lots of >> so many photos, right? I want like 100, 200, 300 photos of this room to to capture it. And what we're trying to do is bring that down to like three. >> Three. >> Right? Well, we're saying like 50, 100 x reduction. And then that's at that scale where it just completely like flips that calculus on its head of like what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and like kind of unearth a lot of footage in the past you would never have treated as reconstructible and go back and like bring it to life as 3D potentially. This is something we've been playing around with a lot with Atlas, right? Like taking old clips. Like I've taken a bunch of my own old captures that never worked before and then put them through the system and I kind of seen a
所以是的,我们说稠密,是真的字面意义上的稠密。就像这个房间,我要 >> 你就是需要大量的 >> 特别多的照片,对吧?我需要一百张、两百张、三百张这个房间的照片才能把它采集下来。而我们想做的,是把这个数字降到三张左右。>> 三张。>> 对吧?我们说的是 50 倍、100 倍的削减。到了那个量级,它就彻底把"你能重建什么样的采集素材"这件事的算盘整个翻过来了。你可以回头去用你已有的影像。你可以用你在互联网上找到的素材,甚至用它们构建出场景。你可以用随手拍的视频,把过去那些你根本不会当作可重建素材的影像挖出来,回头把它们变成三维、让它们活过来。这也是我们用 Atlas 玩了很多的东西,对吧?比如拿老的片段。我自己就拿了一堆以前从来没成功过的采集素材,把它们丢进系统,第一次看到了重建结果;或者拿我以前的采集素材,扔掉其中
便签引用
16:44
reconstruction for the first time or taken my old captures and thrown away 95% of the photos I took and you know, imagine angles that I never would have gotten from a traditional kind of like Nerfer Splat type reconstruction. >> One thing that's under appreciated on the website of the demos is the Stanford demo where Ben showed uh anywhere between 3 to 25 images you can reconstruct that entire Stanford quad. But the thing is we had to show it from aerial view. But every single input image is been standing on the ground taking a picture from the ground. So, everything you see are generated but according to the laws of reconstruction.
95% 的照片,然后,你知道,去想象那些用传统的NeRF 或 Splat 类重建根本拿不到的角度。>> 网站上那些 demo 里有一个被低估了,就是斯坦福那个 demo,Ben 展示了用 3 到 25 张图像就能重建整个斯坦福 quad。但关键在于,我们得从空中视角来展示它。而每一张输入图像都是他站在地面上、从地面拍的。所以你看到的一切都是生成出来的,但又符合重建的规律。
便签引用
07重建就是超长上下文下的生成
17:27
And this is really magical. >> And this is where like generation and reconstruction need to interplay in a really fundamental way to solve this problem. Because under the classic kind of reconstruction stuff that Ben was talking about, like the reason you need so many views is because I need like multiple images and I need to triangulate this point in 3D space and see it from multiple viewpoints. So, that means like that's required in the traditional version. And on the flip side, anything that wasn't captured in these views, like any pixel that was not visible in one of the input views will be a hole in a 3D reconstruction.
这真的很神奇。>> 而这正是生成与重建必须以一种非常根本的方式相互配合,才能解决的问题。因为在 Ben 刚才说的那种经典重建里,你之所以需要那么多视角,是因为我需要多张图像,需要对三维空间中的这个点做三角测量,从多个视点看到它。所以,也就是说,传统方法里这些是必需的。而反过来说,任何没有被这些视角捕捉到的东西——任何在输入视角里看不见的像素,在三维重建中都会变成一个空洞。
便签引用
17:59
Because fundamentally like if a thing wasn't visible in the input views, you know, you need to imagine it to fill in the gaps. And that's fundamentally a generative process. So, even in this room, even if we set Ben loose with a DSLR and like let him like capture like hundreds of views of this room, even the world expert on doing these dense captures is still going to miss some spots. Like he's not going to get like underneath all of the microphones or underneath all the tables or like in between all the chair legs, you're always going to miss something, no matter how many views you get. So, that that's where you need generation as another mechanism in the model. Because you're never going to get everything.
因为从根本上说,如果某个东西在输入视角里没出现过,你就得靠想象去填补这些空缺。而那本质上就是一个生成的过程。所以哪怕就在这个房间里,哪怕我们让 Ben 拿着单反去拍上几百张这个房间的照片,即便是做这种密集采集的世界级专家,也一定会漏掉一些地方。比如他不可能拍到所有麦克风底下,或者所有桌子底下,又或者椅子腿之间的缝隙,不管你拍多少个视角,总会漏掉点什么。所以这就是为什么模型里还需要生成这个机制。因为你永远不可能把一切都拍全。
便签引用
18:31
So, you need to have some generative capacity for the model to imagine, oh, based on what I'm seeing, then like first triangulate what I what I can see, but then fill in the gaps of the stuff that inevitably inevitably was not captured. >> Yeah, and there's something like super cool about this that LLMs have really understood this for a for a long time, right? There was kind of one of these like context wars like the first couple years I was like, oh, we got to 128 to 256 to like 512. We got a million, right? And everyone kind of understands now at a pretty tangible level the value of, you know, you crank your context length to high when you're using your coding models. It's a hard problem. Like everyone has a feel for that. But like no one has pushed that at all on the image and video model side in the same kind of like principled way. Like no one's out there trying to like put a like an hour-long video through and do a needle in a haystack retrieval of like a frame at the 30 37-minute mark. Whereas
所以你需要让模型具备一定的生成能力,去想象:哦,根据我看到的内容,先把能看到的部分三角化重建出来,然后再把那些不可避免没被拍到的部分补上。>> 是啊,而且有一点特别酷,就是大语言模型其实很早就理解了这件事,对吧?最初那几年有过一场所谓的“上下文长度之争”,我记得是,哦,我们做到 128K、256K,再到 512K。我们做到一百万了,对吧?现在大家都能很具体地感受到上下文的价值——你用编程模型的时候会把上下文长度拉到很高。这是个难题,大家都有切身体会。但在图像和视频模型这边,还没有人用同样有章法的方式去推进这件事。没有人真的去尝试把一个一小时长的视频塞进去,然后做“大海捞针”式的检索,比如找出第 37 分钟那一帧。而在重建和生成这件事上,你
便签引用
19:17
with reconstruction and generation, you actually have the same exact thing of like reconstruction is just like generation with a really long context and you put a lot of stuff in it. Right? Like that's the way to actually build this continuum where you kind of bridge between those two things. And like Atlas, like being able to like this is something we could never do with Marvel. Marvel had this kind of fundamental blocker of like you couldn't really jam more than honestly like a couple images in. But Atlas, I can go and I can actually take like a 64-image capture and do like a fly-through of an entire house and everything is grounded by being, you know, seen or like almost seen or like slightly extrapolated from what's not there. But you're just getting these, you know, I'm taking captures I did with 2,000 images of a multi-room house and taking it down to like 30, 40 inputs and the fly-through looks like basically the same. And this is just like totally inconceivable before and it's all enabled by building
其实面对的是完全一样的问题:重建无非就是超长上下文下的生成,你往里塞了海量的东西,对吧?这才是真正打通这两者、建立起连续谱的方式。比如说Atlas,能做到这种事——这是我们用 Marvel 时永远做不到的。Marvel 有一个根本性的瓶颈:说实话,你塞不进去超过几张图像。但用 Atlas,我可以拿一组 64 张图的采集,做出整栋房子的漫游,而且一切都是有依据的——要么是看见过的,要么是接近看见过的,要么是从已有内容里稍作外推的。你得到的就是这种效果,比如我拿以前用 2000 张图拍的一栋多房间住宅的采集,把输入压缩到 30、40 张,漫游效果基本上是一样的。这在以前完全不可想象,而这一切都得益于
便签引用
20:02
this gracefully scaling kind of context window that you can dump stuff into. >> And so the way to think about it is like the the sparseness are the pictures that you physically took and then Atlas is a model creates the rest of the views and then you use classic reconstruction techniques. Is that roughly the way to think about it or >> In some sense, yeah, yeah. I mean, that's the beauty of Atlas is like you can take however many inputs you have down to like a single view and then you can always use Atlas as this this rendering engine to produce anything else you want. Right? You can you can navigate it like a virtual camera. Yeah, exactly. You can just you can say like, "Okay, I have a picture here. I want a picture there, there, there." You can make a couple of those then you can say to dense fly-through. You can do this in sequence because it's an auto regressive model. It's up to you, right, to kind of pick and choose what you add interactively into the context as you generate.
构建了这种可以优雅扩展的上下文窗口,你可以往里面不断塞东西。>> 所以可以这么理解:稀疏的那部分是你实际拍下的照片,然后 Atlas 这个模型生成其余的视角,再用经典的重建技术来处理。大致可以这么理解吗,还是说 >> 某种意义上,是的,是的。我是说,Atlas 的妙处就在于,不管你有多少输入,哪怕只有一张图,你都可以把 Atlas 当作一个渲染引擎,生成你想要的任何其他画面,对吧?你可以像操控虚拟相机一样去导航。是的,没错。你可以直接说,“好,我这里有一张照片。我想要那边、那边、那边各来一张。”你可以先生成几张,然后再说做成密集漫游。你可以分步骤地做,因为它是一个自回归模型。完全取决于你,在生成过程中交互式地挑选往上下文里加什么。
便签引用
08他们凭什么相信这条路会成
20:49
>> I mean, the thing that I just blows my mind is Listen, I just have a very simple mental mental model. I I I I four pictures and then I've got to like have the model extrapolate between them and then it has to fit when you reconstruct. Like it's got to be 3D. Like and I always think of these diffusion models as like being visually great but not accurate. And so like And and I don't even know if there's a question here, but like how how how is I like the room fits? So like how is it that it's three 3D consistent? Is it just lots of data or >> Yeah, I mean it's a part partially it's a belief in the scaling hypothesis, right? Like, you know, >> Did you By the way, but I have to ask, when you set out to do this, did you know it was going to work?
>> 我是说,真正让我惊掉下巴的是——听着,我的心智模型很简单。我有四张照片,然后得让模型在它们之间做外推,重建的时候还必须对得上。得是三维一致的。我一直觉得这些扩散模型视觉效果很棒,但并不准确。所以说……我其实也不确定这里有没有一个明确的问题,但就是——房间怎么就对得上呢?它是怎么做到三维一致的?是单纯靠海量数据,还是 >> 是啊,我觉得部分原因是我们相信 scaling 假设,对吧?就是那种,你知道的 >> 顺便问一句,我必须问,你们一开始做这件事的时候,知道它一定能成吗?
便签引用
21:32
>> I was pretty sure. >> [laughter] >> Were you sure? >> I think three of us have total conviction about the scaling law. That that I think we do. I do think the exact architecture choices and data mixtures is where the the devils are in the details. I, you know, have watched Justin and his team going from we really don't know how long this is going to take to oh, maybe sign of life to wow, this is going to work. So it it no one what no one has done it, but I think the hypothesis, two hypothesis, one is scaling law hypothesis, the other one is next viewpoint prediction. We had conviction of these two things primarily.
>> 我当时挺有把握的。>> [笑] >> 你当时确定吗?>> 我觉得我们三个人对 scaling law 是有绝对信念的。这一点我觉得我们有。但具体的架构选择和数据配比,魔鬼确实藏在细节里。我一路看着 Justin 和他的团队,从“我们真不知道这要花多久”,到“哦,好像有点起色了”,再到“哇,这真能成”。所以没有人做过这件事,但我觉得有两个假设,一个是 scaling law 假设,另一个是下一视角预测。我们主要是对这两点抱有信念。
便签引用
22:20
>> So I think I was very convicted that it was going to work. I was not sure it was going to work at this well at this fast, right? Like I thought there's a chance that we do this. Maybe it's not clear that like the first cycle of pre-training a new model with a new architecture and a new paradigm. Like the first cycle of that working is insane. So I thought there was a chance in which we had to we might have had to do a couple more turns of that of that pre-training cycle before we got to the level of quality we we wanted.
>> 所以我觉得我非常确信它能成。我只是不确定它能做得这么好、这么快,对吧?我当时想,我们有机会做成。但说实话,用全新的架构、全新的范式做一个新模型,第一轮预训练就能跑通,这本身就很离谱。所以我当时觉得,也有可能我们得再多跑几轮预训练,才能达到我们想要的质量水平。
便签引用
22:44
>> Is there is it are we kind of like at the end of like the scaling for this architecture approach? We need another breakthrough or is there >> No, no, we're at the beginning. >> Really? >> Yeah. >> Without without changing the architecture? >> Yeah, we're basically at the beginning. I think we're basically at the beginning and we're basically limited by compute at this point. All right, like data is is important as Feifei likes to point out, but like everything has a bottleneck and I think the main bottleneck on continuing to scale this thing is actually training compute.
>> 那我们是不是已经到了这套架构路线 scaling 的尽头?是不是需要另一个突破,还是说 >> 不不,我们才刚刚开始。>> 真的吗?>> 是的。>> 在不改变架构的情况下?>> 是的,我们基本上还处在起步阶段。我认为我们基本上还在起步阶段,而且目前基本上是被算力限制住了。没错,就像李飞飞喜欢指出的那样,数据很重要,但每样东西都有瓶颈,我认为继续扩展这个东西的主要瓶颈其实是训练算力。
便签引用
23:07
Right? Like during development, we trained a sequence of models. We wrote about this in the blog post a little bit, but we trained a couple of models that like the first couple of runs of the scaling ladder. Um and each time we made the model bigger and each time we trained it for longer, each time we put it on more chips, like it got significantly better. And the model size that we like the model that we showed in the blog post is obviously the biggest and best one that we trained, but the thing that was limiting it was not the scale or the data or anything like that.
对吧?比如在研发过程中,我们训练了一系列模型。我们在博客文章里稍微写了一点,但我们训练了好几个模型,也就是扩展阶梯上最初的那几次训练。嗯,每一次我们把模型做得更大,每一次我们训练得更久,每一次我们把它放到更多芯片上,它都会显著变好。而我们展示的那个模型规模,也就是我们在博客文章里展示的那个模型,显然是我们训练过的最大最好的一个,但限制它的并不是规模或数据之类的东西。
便签引用
23:32
It was literally like we had a deadline of when we wanted to release this thing and therefore we backed up what we could afford to train in time for that deadline. >> But here's a little bit of a insider story, right? Like Justin and team are are training from the smaller and slightly bigger, you know, are train having these roadmaps. And then there was one day in summer, early summer, that it's not even this the current Atlas model size, it's a smaller model. And then Ben, Justin, Ben feeded into, you know, the viewpoint generation. And remember that famous table, the garden table for Nerf paper and many papers, that overnight I got a slack. I mean, we all saw the slack from Ben that our camera flew through under the table.
真的就是,我们有一个想要发布这个东西的截止日期,因此我们倒推出在那个截止日期之前我们能负担得起训练多大的模型。>> 不过这里有个内部小故事,对吧?就是贾斯汀和团队在从更小的、稍微大一点的模型开始训练,你知道的,他们有这些路线图。然后在夏天的某一天,初夏的时候,那甚至还不是现在 Atlas 的模型规模,而是一个更小的模型。然后本、贾斯汀,本把它接入到,你知道的,视点生成里。还记得那张著名的桌子吗,Nerf 论文和很多论文里的那张花园桌子,那天晚上我收到一条 Slack 消息。我是说,我们都看到了本发的那条 Slack,我们的相机从桌子底下穿了过去。
便签引用
09创作者工作流与三维软件之痛
24:19
>> With the soccer ball. >> Uh yes, with the soccer ball. >> ball emergent or is that in the original picture? >> It's It's It was real, right? >> Okay. >> That morning the three of us looked at each other in the eyes and say, "That's it. This is We're going to build this." Like we we made a decision within 5 seconds. How This is just a no one has ever seen this result. >> Ben, um can you talk through maybe more specifically the use cases? So so uh World Labs has historically had a lot of users that were creatives and they use it for like consistency in, you know, like whatever 2D images and for movies and for 3D and for games etc. And so maybe can you talk about how this extends use cases or cater to the existing ones and then we'll I'd like to talk about robotics actually.
>> 还带着那个足球。>> 呃是的,还带着那个足球。>> 足球是涌现出来的,还是原图里就有的?>> 那是It's 那是真实存在的,对吧?>> 好的。>> 那天早上我们三个人对视了一眼,然后说:「就是它了。这就是我们要做的东西。」我们在 5 秒钟内就做出了决定。这简直是从来没人见过这种结果。>> 本,嗯,你能不能更具体地讲讲应用场景?因为呃 World Labs 一直以来有很多用户是创意工作者,他们用它来做比如,你知道的,2D 图像里的一致性,用在电影、3D 和游戏等等上面。那么也许你能讲讲这个怎么拓展了应用场景,或者怎么服务于现有的场景,然后我其实很想聊聊机器人。
便签引用
25:04
>> Yeah, sure. Um yeah, I mean it's kind of funny actually one of the kind of main ways we even saw people using Marvel plays exactly into this new view prediction case. Like a lot of our >> Marvel being the previous sorry our Marvel our previous product. >> Um like people would take that product, put an image in, get a full 3D scene as a Gaussian splat, take a couple of screenshots of it from different points of view and leave. Right? And we're like We can just make these images [laughter] and that's generative AI cool, right? So So I think like and you know, there's a lot of degradation there. They're like, "Oh, this splat could look better." And it's like, "Okay, what if we just generatively model those viewpoints with that exact modality of control?" So I think like even that core capability of just like view synthesis, um it's sort of been this academic problem for a long time. But in the sense of "Oh, you're going to do this really dense capture."
>> 好的,当然。嗯,其实挺有意思的,我们看到人们使用 Marble 的一种主要方式,恰恰就落在这个新的视角预测场景上。就像我们很多 >> Marble 是指之前的抱歉,我们之前的产品 Marble。>> 嗯,人们会拿那个产品,放进去一张图,得到一个完整的 3D 场景,以高斯泼溅的形式呈现,然后从不同视角截几张截图就走了。对吧?我们就想,我们完全可以直接生成这些图像啊 [笑]而这就是生成式 AI,酷吧?所以所以我觉得,而且你知道,那里面有很多质量损失。他们会说,「哦,这个泼溅本可以更好看。」然后就是,「好吧,那如果我们直接用那种精确的控制方式来生成式地建模这些视点呢?」所以我觉得,即使是视角合成这个核心能力,嗯,它在学术界已经是个老问题了。但那是在「哦,你要做非常密集的采集」这个意义上。
便签引用
25:48
Like like generative view synthesis is a relatively quite a new problem. And we just see so many people who uh in this creative pipeline, right? People have a multi-stage workflow, right? I don't think there's a single person out there using one monolithic model, not even C dance or whatever for for their entire task. Uh people will have this like, you know, kind of a bunch of storyboards and mood boards of images they pull out from like their favorite collection of image models. And then they'll go to different video tools and like build those together as keyframes. Then they'll go and like clip and edit those later, right? So we were seeing this like sort of, you know, niche but very specific use case for Marvel as just providing that like sanity that you can ground your generations in some kind of 3D consistent world, right? People, you know, I don't I I fought with with various image models to ask them to like give me different viewpoints of a room.
而生成式视角合成是个相当新的问题。我们看到太多人,呃,在这个创作流程里,对吧?人们有多阶段的工作流,对吧?我不认为有任何一个人会用一个单一的巨型模型,哪怕是 C dance 还是什么,来完成他们的整个任务。呃,人们会有那种,你知道的,一堆故事板和情绪板图片,是他们从最喜欢的那些图像模型里挑出来的。然后他们会去用不同的视频工具,把那些作为关键帧拼起来。接着再去剪辑和编辑,对吧?所以我们看到 Marble 有这种,你知道的,小众但非常具体的用途,就是提供那种保障,让你的生成内容可以锚定在某种 3D一致的世界里,对吧?人们,你知道,我我曾经跟各种图像模型较劲,让它们给我一个房间的不同视角。
便签引用
26:33
And every time you can just look and see, "Oh, things kind of moved around. Like it's not stable." And like even that one seed of a use case I think kind of signals that there's this value and there's hiding under the surface there. Like there's just decades of people being used to persistent 3D state like virtually modeling what they would be doing in the real world and having, you know, a stage and props and like elements there, whether it is for a movie or a show or a marketing shot or like building out game environments.
每次你一看就能发现,「哦,东西都挪位了。它不稳定。」而即便是这么一个用例的种子,我觉得也标志着这里面是有价值的,有些东西藏在水面之下。就是说,几十年来人们已经习惯了持久的 3D 状态,在虚拟中建模他们在现实世界里会做的事情,你知道的,有舞台、有道具、有各种元素,不管是为了电影、剧集、营销镜头,还是搭建游戏环境。
便签引用
27:00
Like this this statefulness and persistence is so key in how people think about spatial reasoning and like developing an environment over time. Like people don't think in this ephemeral like generate a thing, generate a thing, like just throw it away, keep my text prompts. Like people want to build this like collection of assets and like model a world in that way. So we're we're trying to provide like again with with the spatial context mechanism and other things like we're trying to provide that level of control and precision and the ability to adjust different modalities of input starting with the post images, but you know, we want to give people more control over the elements of the things in the scenes they're looking at and editing and interaction and all that as we go forward. And I think that that it unlocks like further use cases in those areas we're already seeing, but also expanding out into kind of any place people want to create a virtual replication or or like, you know, a pre-imagination of a real-world space
这种状态性和持久性,对于人们如何思考空间推理、如何随时间发展一个环境来说,太关键了。人们不会用那种转瞬即逝的方式思考,比如生成一个东西,再生成一个东西,然后直接扔掉,只留着我的文本提示词。人们想要构建这样一个资产集合,用那种方式建模一个世界。所以我们我们试图提供,还是那句话,用空间上下文机制以及其他东西,我们试图提供那种程度的控制和精度,以及调整不同输入模态的能力,从姿态图像开始,但你知道,随着我们往前走,我们希望给人们对场景中各个元素有更多控制权,他们正在查看、编辑、交互等等的那些。而且我觉得这会解锁我们已经看到的那些领域里的更多用例,同时也拓展到任何人们想要创建一个虚拟复制品,或者说,你知道,对他们需要建造的真实空间做预先想象的地方,对吧?比如建筑和施工。我曾经跟一个
便签引用
27:51
they need to build, right? For architecture and construction. Like I talked to a guy at some point building booths for conferences, right? There's just so many things in the world you don't think about need to be fabricated and every single one of those basically goes through this like pretty painstaking virtual design phase. And of that process like the part where you go into 3D software is kind of one of the most like arduous and like labor-intensive parts right now. Like like taking feedback on a 3D design from kind of like verbal commentary or sketch or really really quick stuff you got from like a creative director or like a design director or an architect or whatever. Like mapping that back into the 3D representation is like 95% of the work, right? You can have a meeting get feedback and then you go back and do a week of revisions. And that's just because like our software is kind of decades old at this point and it's it's just never became as intuitive as, you know, playing with Legos or like
给会议搭展台的人聊过,对吧?世界上有太多你想不到的东西是需要被制造出来的,而其中每一样基本上都要经过这种相当费力的虚拟设计阶段。而在那个流程中,进入3D 软件的那个环节,是目前最艰辛、最耗人力的部分之一。比如说把对 3D 设计的反馈,那种口头评论、草图,或者你从创意总监、设计总监、建筑师之类的人那里拿到的非常非常随手的东西,把那些映射回 3D 表示里,大概就是 95% 的工作量,对吧?你可能开个会拿到反馈,然后回去做一周的修改。而这仅仅是因为我们的软件到现在基本上已经有几十年历史了,它从来没有变得像玩乐高那样直观,或者像
便签引用
28:42
pottery or doing this stuff with your hands or sketching with a pencil. And this is the real place where AI can actually unlock a ton of value for people in their process, whether it's a creative application or something more industrial or design or whatever. Uh, and that like really motivates me to kind of build different flavors of of our model to cater to those kind of people. >> Yeah, I I I can understand how it helps with the creatives cuz like Marvel did that. And also how that extends to things like designer architecture. Uh, Baidu you acquired a robotics company. And so it's >> [laughter] >> it's less less >> talked about it too. We just talked about >> know. But it's the last time we talked to me especially in the context of Atlas like how that maps to robotics. So if you wouldn't mind just penciling that out.
做陶艺、用手做东西、用铅笔画草图那样。而这正是 AI 真正能在人们的流程中为他们解锁大量价值的地方,不管是创意应用,还是更工业化的、设计类的,或者别的什么。呃,这真的很激励我去打造我们模型的不同版本,来服务这类人群。>> 是的,我我我能理解这对创意工作者有什么帮助,因为 Marble 就做到了这一点。也理解它怎么拓展到设计、建筑这类领域。呃,飞飞你们收购了一家机器人公司。所以这 >> [笑] >> 这个谈得比较少>> 也聊聊那个吧。我们刚才只聊了 >> 知道。不过这是上次我们聊的时候,尤其是放在 Atlas 的背景下,它怎么对应到机器人领域。所以如果你不介意的话,就把那个讲一讲。
便签引用
10机器人的瓶颈是数据,不是芯片
29:22
>> Yeah, actually Atlas is a a key part of the puzzle. So, um, we acquired this company that was formerly known as Synnex. And what is their key technology? Right now their key technology is a system that goes from real to sim and and then sim to real. And what does that mean in robotic situation? You want to train a robotic arm to, you know, figure out how to, um, do cabling, let's say, in a in a industrial setting. Well, you need a whole bunch of data to first train a robotic policy to do these cable cables cabling activity. And then you want to evaluate if the robotic policy is doing a good job. And then you deploy the robot into the cabling environment. In order to train, what you what this company, uh, Synnex and now our robotics team used to be doing is doing exactly what Baidu was saying, dense reconstruction. You take >> All right, so Synnex >> pictures of a situation and then and try to reconstruct that environment. It's excruciatingly painful. Takes a long time, laborious, and it really blocks
>> 好的,其实 Atlas 是这个拼图中的关键一块。所以,嗯,我们收购了这家公司,以前叫 Synnex。他们的核心技术是什么?目前他们的核心技术是一套从真实到仿真、再从仿真到真实的系统。那这在机器人场景下意味着什么?你想训练一个机械臂,让它,你知道的,搞明白怎么,嗯,在工业环境里,比如说,做布线。那么,你首先需要大量数据来训练一个机器人策略去做这些布线操作。然后你要评估这个机器人策略做得好不好。接着你把机器人部署到布线环境里去。为了训练,这家公司,呃,Synnex,也就是现在我们的机器人团队,以前做的正是飞飞刚才说的那件事:密集重建。你拍>> 好的,所以 Synnex >> 一个场景的照片,然后试着重建那个环境。这个过程痛苦到极点。花很长时间,费力,而且它真的拖慢了机器人仿真中从真实到仿真的速度,对吧?所以 Atlas 真的是
便签引用
30:35
the velocity of robotic simulation real to sim, right? So Atlas really is the next generation technology for that. And this is not just for robotics cabling or anything. We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data. Because it's so hard to collect real-world data where robots, you know, are operating and and uh in order to uh not only you need to uh collect the data of, let's say, the cabling situation or dishwashing situation, whatever. There is also a very important step called randomization. Is that you have to take the same environment and then randomize the conditions. So, the cable doesn't literally only, you know, uh bend this way. It can bend a different way or the box can have different sizes, colors, different lids, and all that. And and or in different parts of the scene. So, you have to go through a real-to-sim situation in order to get enough of that data in addition to other data you can
那方面的下一代技术。而且这不只是为了机器人布线或别的什么。我们应该跳出来看,认识到目前机器人领域最大的问题其实是数据。有一天会是芯片,但现在是数据。因为要采集机器人在真实世界中运行的数据太难了,而且呃,为了呃不仅仅需要呃采集,比如说,布线场景或者洗碗场景之类的数据。还有一个非常重要的步骤叫做随机化。就是你必须拿同一个环境,然后把各种条件随机化。这样,那根线缆就不会只按这一种方式弯折。它可以用不同方式弯折,或者箱子可以有不同尺寸、颜色、不同的盖子,诸如此类。还有或者出现在场景的不同位置。所以,你必须走一遍真实到仿真的流程,才能在你从互联网获取的其他数据之外,拿到足够多这样的数据。所以,这个真实到仿真的
便签引用
31:48
get from internet. So, this real-to-sim uh step will be um you know, really helped by by Atlas. That's just the first part of this is meeting the the the robotics needs in the current technology cuz we don't yet have a a frontier foundation model that's robust enough for for uh robotics. But Atlas is a uh omni model. It's a multi uh multi-modal model. It takes on different kinds of input and generates different kind of output. You can totally imagine the next step is Atlas um taking in uh data that's in the uh in that's dynamical.
呃步骤会被,嗯,你知道的,被 Atlas 大大助力。这只是第一部分,是用现有技术去满足机器人领域的需求,因为我们还没有一个足够稳健、能用于呃机器人的前沿基础模型。但 Atlas 是一个呃全能模型。它是一个多呃多模态模型。它接受不同种类的输入,并生成不同种类的输出。你完全可以想象下一步是 Atlas 嗯接收呃那种动态的数据。
便签引用
32:31
And that can really start to bridge the gap between, you know, action planning and uh and um a robotics and and the Atlas output. So, that's all narrow map. What are you going to say something? >> Yeah, I I was going to say there's something fundamentally different about training a robotics policy compared to really any other application in AI we've seen before. Um and that's like if you're generating a piece of code, like you're generating an image, you're generating a video, I'm the model is fundamentally creating this this artifact. And that artifact like there's a lot of examples of artifacts that you can go out on the web or somewhere and collect. Right? You want to generate images, there's a lot of images out there. You want to generate videos, there's a lot of videos out there. You want to generate a code base, there's a lot of code bases out there you can learn from.
而这真的可以开始弥合,你知道的,动作规划和呃和嗯机器人以及 Atlas 输出之间的鸿沟。所以,那都是狭义的路线图。你刚才要说什么?>> 是的,我我本来想说,训练一个机器人策略,和我们之前见过的 AI 里几乎任何其他应用相比,有一点根本性的不同。嗯,就是说如果你在生成一段代码,比如你在生成一张图像,你在生成一段视频,我是说模型本质上是在创造这个产物。而那个产物,你可以在网上或别的什么地方找到并收集到很多这类产物的样例。对吧?你想生成图像,外面有大量图像。你想生成视频,外面有大量视频。你想生成一个代码库,外面有大量代码库可以供你学习。
便签引用
33:14
>> Yeah. >> A robotics policy is something fundamentally different. It's not producing a static thing. It's instead a policy that's going to go out into the world, make actions, and like try to achieve a goal. And the world's not always going to respond the way you expect, right? Unexpected stuff is going to happen. So, a robotics policy is like fundamentally an agent that is out in the real world interacting with the real world, and stuff happens. So, you need like a critical part of that is the those those policies during training need to be exposed to every possible thing that could go wrong during a deployment. Um and that's where simulation is really key for robotics, right? So, there's the then there's two angles on that. Like one is the kind of classical simulation. You can go out You can go and like go to your favorite physics engine and like try to imagine creatively as a human designer, what are all the scenarios that might happen in this in this in this when achieving this task, and then try to write explicit
>> 是的。>> 机器人策略是根本不同的东西。它产出的不是一个静态的东西。它反而是一个要走进世界、做出动作、并试图达成目标的策略。而世界并不总是按你预期的方式回应,对吧?意料之外的事情总会发生。所以,机器人策略本质上就是一个身处真实世界、与真实世界交互的智能体,而各种事情都会发生。所以你需要,其中至关重要的一点是,那些策略在训练时需要接触到部署过程中可能出错的每一种情况。嗯,这就是仿真对机器人真正关键的地方,对吧?所以,那里有两个角度。一个是那种经典仿真。你可以去找你最喜欢的物理引擎,然后作为人类设计师去创造性地设想,在完成这个任务时可能发生的所有场景有哪些,然后试着写出显式的
便签引用
34:06
code that models them all. That That's one angle. And that's an interesting angle with coding agents. Like that actually gets supercharged, too. But there's another angle, which is try to more data-driven simulation. Right? Like maybe we can Can we have a learned model that can understand how the environment how the world is going to respond to actions? And maybe it might respond in unexpected ways sometimes. Then could we build these neural simulators that are trained on as much data as we can, then use these neural simulators, these learned neural simulators, you know, as a simulation bed to train robotic policies.
代码去把它们全都建模出来。那是一个角度。而且在有了编程智能体之后,这是个有意思的角度。它实际上也被大大增强了。但还有另一个角度,就是尝试更数据驱动的仿真。对吧?比如也许我们可以我们能不能有一个学习得到的模型,能理解环境世界将会如何对动作做出回应?而且它有时候可能会以意想不到的方式回应。那我们能不能构建这些神经仿真器,它们能用尽可能多的数据去训练,然后把这些神经模拟器,这些学出来的神经模拟器,当作一个仿真环境来训练机器人策略。
便签引用
34:37
Um and that's that's a really interesting future direction of Atlas. But but then it doesn't stop there, right? So >> [laughter] >> so um but once you have, you know, this learned simulator, like this learned simulator kind of already has in its like mental brain, like it understands the world, it understands how the world is going to respond to actions, and why doesn't the simulator itself become the planner? Right? Like the same That's that's kind of the core thesis that we've had around world models and their generality, that there's some core stuff that a model should understand around generating worlds, simulating them, understanding how they appear in different situations, and, you know, understanding how the world's going to respond to an action is highly related to imagining what kind of action I need to take to make the world respond in a particular way.
嗯,这是 Atlas 一个非常有意思的未来方向。但是它还不止于此,对吧?所以 >> [笑] >> 所以,嗯,一旦你有了这个学出来的模拟器,这个学出来的模拟器其实已经在它的「大脑」里理解了世界,它理解世界会如何对动作做出反应,那为什么模拟器本身不能变成规划器呢?对吧?这其实就是我们围绕世界模型和它们的通用性所持有的核心论点:模型应该理解一些核心的东西,比如生成世界、模拟世界、理解它们在不同情境下的呈现方式,还有,理解世界会如何对一个动作做出反应,这跟想象「我需要采取什么样的动作才能让世界以特定方式做出反应」是高度相关的。
便签引用
11动态、可编辑性与下一步
35:21
>> Yeah. What what um one piece of feedback that I got I've been By the way, congrats on the launch. It was overwhelmingly positive. I think it was probably the most significant model launch this year. And yeah, everybody said glowing things. But one person who's an expert in the space who I texted was like, "What do you think?" And great person says, "It's fantastic, it's amazing, but there needs to be more dynamics." >> [laughter] >> And so it would seem to be at least in the robotics case, but generally it was kind of ideal to actually have a world that moves. And so maybe talk a little bit about that and then any other future directions that A, you're comfortable sharing, but you think are worth talking through.
>> 是的。我收到的一条反馈是——顺便说一句,恭喜发布。反响压倒性地正面。我觉得这可能是今年最重要的一次模型发布。是的,大家都赞不绝口。但我给一位这个领域的专家发短信问「你怎么看?」,这位很厉害的人说:「它非常棒,很惊艳,但需要更多动态效果。」>> [笑] >> 所以看起来,至少在机器人这个场景里,但总体上,理想情况是有一个会动的世界。所以也许可以聊聊这个,以及你们愿意分享的、觉得值得聊一聊的其他未来方向。
便签引用
35:56
>> Yeah, I mean, like dynamics is clearly going to happen. Like actually um >> We have baby dynamics. >> We actually do have baby dynamics already. And this is something I think people didn't quite appreciate, we didn't really highlight in the blog post, but like the previous marble world model, it was like fundamentally static. Like the the model just like could not handle any dynamics at all. And that was just like baked into the model architecture, baked into the training, like the whole thing was fundamentally static.
>> 是的,我是说,动态肯定会实现。其实,嗯 >> 我们已经有「婴儿版动态」了。>> 我们确实已经有初级的动态了。我觉得这一点大家不太理解,我们在博客文章里也没怎么强调——之前的 Marble 世界模型本质上是静态的。模型根本处理不了任何动态。这是写死在模型架构里、写死在训练里的,整个东西本质上就是静态的。
便签引用
36:17
>> Yeah. >> Um we already knew that that was a big problem post marble, and we already fixed it in Atlas, right? Like the Atlas architecture is already fundamentally supports dynamics. Um the Atlas training data fundamentally has dynamics. And if you look carefully in some of the videos that we've even posted, >> It actually is. >> I saw I saw you see that. >> Yeah, [laughter] the waves, the water waves. >> Yeah, so like some of the examples there's like waves in the water, like in some of the like air generated aerial views, there's like little cars moving around. So like dynamics is actually already in this model.
>> 是的。>> 嗯,Marble 之后我们就知道这是个大问题,而且我们在 Atlas 里已经解决了,对吧?Atlas 的架构从根本上已经支持动态了。Atlas 的训练数据本质上就包含动态。如果你仔细看我们发布的一些视频,>> 确实有。>> 我看到了,我看到你注意到了。>> 是的,[笑] 那些波浪,水面的波浪。>> 对,比如有些例子里水里有波浪,还有一些生成的航拍视角里有小车在移动。所以动态其实已经在这个模型里了。
便签引用
36:45
>> But is it By the way, dynamics seems very problematic to me if you're trying to reconstruct 3D from multiple views, right? And so like are these things like at odds or >> it's actually one of our thesis here is that like, you know, if you're just going to do fundamental 3D reconstruction, you actually want to have no dynamics. Like you want to be able to model like exact views of the scene with exact frozen time. >> Right. >> Um, but then like this is actually kind of a problem with our previous Marble approach, right? Like there like you can try to find data that's fully static, but that's really hard to scale and really hard to get more of. And the thing we realized is that even in the case where I want a static output in the end, the best way to get it is actually expose the model to dynamics, right?
>> 但是——顺便问一下,如果你想从多个视角重建 3D,动态在我看来是个大麻烦,对吧?所以这两者是不是有冲突?>> 其实我们的一个论点就是,如果你只是要做纯粹的 3D 重建,你其实希望完全没有动态。你希望能建模场景在完全冻结时间下的精确视角。>> 对。>> 嗯,但这其实正是我们之前 Marble 方案的一个问题,对吧?你可以试着去找完全静态的数据,但那非常难规模化,也很难拿到更多。我们意识到的是,即使在我最终想要静态输出的情况下,得到它的最佳方式其实是让模型接触动态,对吧?
便签引用
37:22
Like expose the model to as much dynamic stuff as you got, as much static stuff as you got, and let the model figure out how to factor out the dynamic stuff. So in especially in like the This is like >> And this is actually So again, in the Atlas pre-training already, like it it saw a ton of dynamics in the pre-training already. Then the post-training that we did specific to this checkpoint in this release was focused a lot more on on static stuff, focused a lot more on spatial movement and less on temporal. But like we already have we already like I'm pretty sure this the the the pre-trained checkpoint already has a lot of latent dynamics in it.
让模型尽可能多地接触动态素材,也尽可能多地接触静态素材,让模型自己搞明白怎么把动态部分分离出去。所以尤其是在——这就像 >> 这其实——所以再说一次,在Atlas 的预训练里,它在预训练阶段就已经见过大量动态了。然后我们针对这次发布的这个 checkpoint 所做的后训练,更多聚焦在静态内容上,更多聚焦在空间移动而不是时间维度。但我们已经有了,我相当确定这个预训练的 checkpoint 里已经包含了大量潜在的动态能力。
便签引用
37:52
>> Yeah. >> Um, and this is something we're going to improve quite a lot going forward. >> So so Ben, does that mean we're going to get 4D video? You can like go walk around. >> I mean, I think >> see the smile on their face. [laughter] >> So I I I actually can see if like you just stopped now and you only did kind of bigger, you know, faster, better, you could build almost an entire industry. Like I feel it feels like a very horizontal primitive. And then and if you did nothing else, but are there other things that are not just kind of bigger, faster that you're excited about with the applications you're focused on, which tend to be kind of more on the kind of content creator 3D side.
>> 是的。>> 嗯,这也是我们接下来会大幅改进的地方。>> 那 Ben,这是不是意味着我们会有 4D 视频?你可以走来走去的那种。>> 我是说,我觉得 >> 看他们脸上的笑容。[笑] >> 所以我其实能想象,如果你们现在就停下,只是做得更大、更快、更好,你几乎就能撑起一整个产业。我觉得这感觉像是一个非常横向的基础能力。就算你不做别的——但除了「更大、更快」之外,还有哪些让你兴奋的东西,尤其是在你们关注的那些应用上,也就是更偏内容创作者、3D 这一侧的。
便签引用
38:24
>> Yeah, I'm really excited about pushing that kind of multimodal aspect. I think different modes of control is so critical here. I think like it's super under appreciated, especially in the academic community, how critical it is to add control conditioning to these models to kind of get out what's inside. I mean, honestly, this >> I'm not trying to I don't even understand what those words >> So, translate in layman's language is so editability. I think editability is is the key here. >> Yeah, so I mean, this is something we've seen in like sort of single image models and starting this year in video models is starting to be unlocked in terms of oh, like getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person and this object and this thing to happen and kind of combining those all together in like one pastiche without having to do a lot of like manual work with the system. Like it just interprets it like kind of frontier image models are kind of there, right? For in terms
>> 是的,我非常期待推进多模态这方面。我觉得不同的控制方式在这里太关键了。我觉得这一点被严重低估了,尤其在学术圈里——给这些模型加上控制条件,以便把模型内部的东西释放出来,这有多关键。老实说,这 >> 我不是想…我甚至听不懂这些词 >> 所以,用大白话翻译就是可编辑性。我觉得可编辑性才是关键。>> 对,我是说,这是我们在单图模型里看到的,今年开始在视频模型里也开始被解锁的东西——就是那种多轮交互的感觉,或者说真正凭直觉去理解「我想要这个人、这个物体、这件事发生」,然后把这些全部组合成一个整体,而不需要在系统里做大量手工操作。它就是能理解你的意思——前沿的图像模型在编辑这方面基本已经做到了,对吧?但我们还没看到这种能力
便签引用
39:14
of editing. But, we haven't seen that propagate out as strongly into video yet and then into world models, right? We've seen some really kind of toy examples of oh, I can like put in a sentence and like, you know, a dinosaur appears or something with these like sort of real-time models. Um but, I want to like turn that up to really industrial strength and make that cuz like the the trick here is you got to add control, but not compromise the quality of the model or it just becomes a a party trick, basically. Like it's like no one is going to seriously think about swapping their like cutting-edge frontier video model usage for your model if you give them extra knobs, but the quality degrades. So, I think it's really that game of like how can we maintain maintain like the high bar we've set with the outputs we're able to get in the current model and then add all kinds of interesting stuff that people will ask us for in terms of like I want to interact with the scene or control the layout or control like the
同样强地扩散到视频里,再到世界模型里,对吧?我们看到的还是一些相当玩具化的例子,比如我可以输入一句话,然后一只恐龙出现之类的,用那种实时模型。嗯,但我想把这个提升到真正的工业级强度,因为这里的诀窍是:你得加上控制,但不能牺牲模型的质量,否则它基本就只是个小把戏。就是说,没有人会认真考虑把他们在用的最前沿视频模型换成你的模型——哪怕你多给了几个旋钮,但质量下降了。所以我觉得这就是一场博弈:我们怎么保持住当前模型输出所树立的高标准,同时再加上人们会向我们要的各种有意思的东西,比如我想跟场景互动、控制布局、控制里面物体和元素的
便签引用
40:01
identity of the objects and the things that we're seeing within there or control time, right? Um and I think that's like an axis where it opens up like a ton of really interesting product and interface work. The more complexity you add there and richness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful, you know, 3D worlds in the computer. Like that's that's really the end goal here. That is getting like all the capabilities you need to build that kind of system.
身份,或者控制时间,对吧?嗯,我觉得这是一个能打开大量有趣产品和交互设计工作的维度。你在那里加入越多复杂度和丰富度,就越能让你真正去思考,几乎从零重新设计人们与计算机里那种有状态的 3D世界交互的方式。这才是最终目标。也就是获得构建那种系统所需的全部能力。
便签引用
40:28
>> Awesome. Anything to get out of that as far as new functionality that you'd be excited about that's not just bigger better? >> I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this uh closing the loop between seeing and experiencing and interaction. So, just thinking about going up that ladder is exactly what uh Ben said.
>> 太棒了。关于新功能,有没有什么你特别期待的、不只是「更大更好」的东西?>> 对我来说,我想回到智能的第一性原理。当涉及空间、物理空间时,智能不是坐在那儿不动,只是看见什么或解释什么,对吧?它真正是在「看见」「体验」和「互动」之间闭环。所以,沿着这个阶梯往上走,正是 Ben 刚才说的那件事。
便签引用
12AI 完备:新视角对标下一个 token
40:56
>> Cool. I think one interesting notion there is this notion of AI completeness. >> You heard of that before? >> Yeah, yeah, I have. Yeah. >> So, like everyone >> Yeah, I I hear about AI complete, by the way, in terms of LLMs, which is like you have to be basically, you know, like the smartest LLM to answer the question what the smartest LLM will need to answer or you have to solve general intelligence. >> No, no, it's it's basically it's it's a connection to Turing completeness, right? Like the idea being that like a task is Turing complete like in classical complexity theory if like I can take any class of any problem in this category, reduce to that one problem, right? Three SAT is a classic example, right? So, you can take any NP-hard problem and reduce it to three SAT. Therefore, therefore, you can use three SAT to solve any any problem.
>> 很好。我觉得这里有一个有意思的概念,就是「AI 完备性」。>> 你听说过吗?>> 听过,听过,是的。>> 所以,大家 >> 是的,顺便说,我听到的 AI 完备是在大语言模型的语境里,大意是你基本得是最聪明的 LLM,才能回答那个最聪明的 LLM 需要回答的问题,否则你就得解决通用智能。>> 不不,它基本上是跟图灵完备性相关联的,对吧?这个想法是:在经典复杂性理论里,一个任务是图灵完备的,如果我能把这一类里的任意问题归约到那一个问题上,对吧?3-SAT 就是经典例子,对吧?你可以把任何 NP 难问题归约到3-SAT。所以你可以用 3-SAT 去解决任何问题。
便签引用
41:33
>> Yeah, it's it's a yeah, yeah. >> So, so then like the the kind of like soft definition of AI completeness is like there's this fundamental primitive that's that's an AI task, but if I could solve this AI task in its full broadest generality >> You solve that. >> it would solve any intelligence problem. And like the classic example of LLMs is like next token prediction is AI complete because I could like, you know, there's the classic example, I think from Ilya, where like I there's a mystery novel and like the thing has to read the whole mystery novel and the final sentence of the mystery novel is like, "And the killer was Predict the [laughter] next token.
>> 对,就是这样,对对。>> 那么,AI 完备性的那种「软」定义就是:存在某个基础原语,它是一个 AI 任务,但如果我能以最广泛的通用性解决这个 AI 任务 >> 你解决了它。>> 它就能解决任何智能问题。大语言模型的经典例子就是,下一个 token 预测是 AI 完备的,因为我可以,你知道,有个经典例子,我想是 Ilya 说的:有一本推理小说,模型得读完整本推理小说,而小说的最后一句是「凶手是——」预测 [笑] 下一个 token。
便签引用
42:02
So, like you could basically like frame any kind of intelligence task in terms of that. So, clearly next token prediction is something that people believe is AI complete. >> Yeah, yeah. >> But I think that something we're kind of realizing and Ben was uh talking about this earlier today is like new view prediction, this primitive that we have in Atlas, especially generative new new view new view prediction. This is also AI complete. Right? And because I could take something like >> You could you could have the the movie and you do all of the frames of the movie and then like the killer walks out and then you predict exactly who walks out.
所以你基本上可以把任何智能任务都框定成这种形式。所以很明显,下一个 token 预测是人们认为 AI 完备的东西。>> 对,对。>> 但我觉得我们逐渐意识到的一点——Ben 今天早些时候也在聊这个——就是新视角预测,我们在 Atlas 里拥有的这个原语,尤其是生成式的新视角预测。这同样是AI 完备的。对吧?因为我可以拿一个像 >> 你可以有一部电影,你把电影的所有帧都给它,然后凶手走出来,你要精确预测走出来的是谁。
便签引用
42:30
>> [laughter] >> Exactly. Not just that, but I could say like I want to have a world where like Martine is like writing a proof of the Riemann hypothesis on the >> So okay, so to take a evolutionary view, right? That new viewpoint prediction is exactly evolution had to solve by making animals move. You you nature give animals eyes. But nature didn't give trees eye. Eyes. Why? Because when you move, you see a new viewpoint. And that is the the the whether you call it AI complete or intelligence complete. So So we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction.
>> [笑] >> 正是如此。不只是这样,我还可以说,我想要一个世界,里面 Martine 正在写黎曼猜想的证明 >> 那好,从演化的角度来看,对吧?新视角预测正是演化在让动物运动时必须解决的问题。大自然给了动物眼睛。但大自然没有给树木眼睛。为什么?因为当你移动时,你就会看到一个新视角。而这就是——不管你叫它 AI 完备还是智能完备。所以我们非常坚定地相信,下一个视角预测等价于下一个 token 预测。
便签引用
43:19
>> Amazing. Well, with that, congratulations all of you on a phenomenal model launch. We're very excited for future model launches and thanks for coming. >> Thank you.
>> 太精彩了。那么,恭喜各位带来了一次出色的模型发布。我们非常期待未来的模型发布,也谢谢你们来。>> 谢谢。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

World Labs 的李飞飞、Justin Johnson 与 Ben Mildenhall 解读其新一代世界模型 Atlas:它把"新视角预测"确立为与"下一个 token 预测"同级的基础范式,首次在同一个模型里统一了三维重建与像素生成,并把空间智能的落地路径从创作工具延伸到机器人仿真。

核心要点

  • Atlas 的基础原语是"新视角预测",而非下一帧预测。 LLM 建立在 next token prediction 上,视频模型建立在 next frame prediction 上;Atlas 则是给定若干视角或场景描述构成的"空间上下文(spatial context)",再把一台虚拟相机指向空间与时间中的任意位置,模型输出该位置应当看到的画面。团队明确把它类比为 next token prediction 的空间版本,并主张它同样具备"AI 完备性"(AI-complete)——任何智能任务原则上都能被改写成一次新视角预测。
  • 与市面上号称"世界模型"的视频模型的真正区别在于每一帧都带三维相机位姿。 普通视频模型的输入图只是让模型"自由解释"的图像,只能靠文本提示反复重掷(Ben 称之为"老虎机效应")。Atlas 的每张输入图都绑定一个三维相机 pose,因此房间四角各拍一张就能精确复现房间内容,而不是猜测另一个角落有什么;也可以把来自不同 AI 生成或真实地点的照片按指定位置"摆台",做出被精确导演的穿行镜头。
  • 首次把重建与生成统一进一个模型,这在计算机视觉里是跨越半个多世纪的分野。 李飞飞指出,视觉会议历来分成像素生成、识别、三维重建等独立赛道,无数博士论文耗在重建或新视角合成上;Atlas 以视点(viewpoint)为锚把两者统一。为此模型从预训练阶段就原生多模态:文本、图像、视频、相机位姿与三维(当前形态为深度图)同为原生输入——把相机位姿放进预训练,团队认为此前无人做过。
  • 稀疏重建把所需拍摄量压到原来的 1/50 到 1/100。 传统稠密重建要求每个想出现在结果里的细节至少被三四个视角看到,桌下、麦克风下、植物叶片缝隙都要覆盖;熟练者要几分钟,普通人扫一套多房间环境可能耗两小时,一个房间常需 100–300 张照片。Atlas 把这个数字降到三张量级。实际效果:2000 张照片拍的多房间住宅压到 30–40 张输入,穿行效果基本一致;Ben 把自己以前失败的旧素材、扔掉 95% 照片的旧拍摄重新跑通。演示中的斯坦福 Quad 用 3 到 25 张地面拍摄的照片重建,展示的却是航拍视角——所有空中画面都是生成的,但受重建规律约束。
  • 生成能力不是锦上添花,而是重建在数学上必需的补丁。 传统重建要三角化就必须多视角覆盖,任何未被任一输入视角看到的像素在结果里就是一个洞。即便让 Ben 这样的世界级专家拿单反狂拍,桌底、椅腿之间仍会漏;所以模型必须具备"先三角化能看见的,再想象补上必然漏掉的"这一生成能力。团队进一步把重建重新表述为"带超长上下文的生成"——LLM 社区早已卷过 128K→1M 的上下文竞赛,而图像/视频侧几乎无人以同等原则推进;Atlas 能吞下多达 100 帧输入,前代 Marble 连塞几张图都困难。
  • Bullet time 从"上百台相机 + 绿幕"降到三部 iPhone。 《黑客帝国》那个环绕镜头靠一圈数百台相机在绿幕前拍成;Atlas 用三脚架上的三部 iPhone 拍同一事件(投篮、草莓落进牛奶),即可重新构图、冻结时间并让镜头在牛奶飞溅的瞬间穿入——无影棚、无绿幕、无昂贵标定。
  • 从 Marble 到 Atlas 的关键转变,是拆掉"必须输出高斯泼溅"的瓶颈。 Marble 把所有输入都压成 Gaussian splat 输出,splat 本身很实用(渲染快、能跑在移动端和 VR、可与游戏和仿真引擎互通),但成了架构瓶颈。Atlas 把根原语换成新视角预测,可输出 RGB 帧、可输出三维,需要时再生成高斯泼溅世界。团队称摸清各种表示的利弊"流了不少血汗"。
  • 团队的两条核心信念是 scaling law 与新视角预测,且认为 scaling 才刚开始。 公司成立两年半,起步时"空间智能根本没有 scaling law"。Justin 表示模型每次做大、训练更久、上更多芯片,都显著变好;当前瓶颈不是数据也不是架构,而是训练算力——博客里展示的最大模型之所以停在那个规模,纯粹是发布 deadline 倒推出的可承受训练量。转折点是今夏某个夜里,在比 Atlas 更小的模型上,Ben 把 NeRF 论文那张著名的花园桌场景喂进去,相机成功从桌下带球穿过;三人当场五秒内拍板。
  • 机器人是 Atlas 最实际的下一站,因为机器人当下的头号瓶颈是数据而非芯片。 World Labs 收购了原名 Synnex 的机器人公司,其核心技术是 real-to-sim-to-real:为训练比如工业布线策略,需先把真实场景稠密重建进仿真,再做随机化(线缆弯法、箱体尺寸颜色盖子、物体位置),生成足够训练与评估数据。稠密重建这一步"痛苦、漫长、劳力密集",直接卡住 real-to-sim 的速度,Atlas 正是它的下一代替代方案。
  • 机器人策略与其他 AI 产物在本质上不同,这决定了神经仿真器的价值。 代码、图像、视频都是静态产物,网上有海量同类样本可学;策略则是一个进入真实世界行动、且世界会以意料之外方式回应的 agent,训练时必须暴露于部署中一切可能出错的情形。除了用编程 agent 强化的传统物理引擎路线,另一条是数据驱动的神经仿真器。Justin 进一步推论:一个已经理解世界如何回应动作的仿真器,天然也能反推"要让世界变成某样,我该采取什么动作"——仿真器本身就可以成为 planner。

结论与值得注意的细节

动态是公认的下一个缺口,但模型里已经有"婴儿版动态"。 有专家反馈"需要更多 dynamics"。Marble 从架构到训练数据都是彻底静态的,Atlas 已在架构和预训练数据层面原生支持动态——已发布的视频里能看到水波和空中俯瞰中移动的小车。本次发布的 checkpoint 之所以显得偏静态,是因为后训练刻意侧重空间移动而非时间维度;团队相信预训练权重里已藏有大量"潜在动态"。

一个反直觉的训练洞见:想要静态输出,反而要喂动态数据。 纯做三维重建理论上最好完全没有动态(要的是精确冻结时间的视角),Marble 的思路就是去找完全静态的数据,但这类数据难以规模化。团队的发现是——把尽可能多的动态与静态数据都喂进去,让模型自己学会把动态因素分离出来。

空间智能的完整定义包含三件事:生成空间、在其中推理、以及编辑与交互。 李飞飞强调最终是含时间维度的 4D,而要做到渲染、仿真与行动规划,必须先理解空间的几何、结构与物理;相机位姿正是关于几何最关键的那条信息。Ben 的补充是产品侧的落点:控制条件(control conditioning)与可编辑性才是关键,且加控制旋钮绝不能牺牲输出质量,否则只是"派对把戏"。

被低估的市场洞察来自用户的真实行为。 不少 Marble 用户的用法是:输入一张图 → 得到一个 splat 三维场景 → 从几个角度截图 → 走人。这直接说明"生成式新视角合成"本身就是需求。没有人用单一模型完成整条工作流,创作者是在多个图像模型、视频工具、剪辑工具之间串联的——Marble 提供的稀缺价值是"让生成结果锚定在三维一致的世界里"。这一需求延伸至任何需要先做虚拟设计的行业:建筑、施工,甚至展会展位搭建;在这些流程中,把口头意见或草图改回三维模型占了约 95% 的工作量,一次评审反馈往往意味着一周返工,因为三维软件的交互方式几十年未变,从未像捏陶土或用铅笔画草图那样直觉。

李飞飞用进化论给这套范式做了收尾论证: 自然给动物眼睛却没给树眼睛,因为只有移动才会产生新视角——新视角预测是演化本身必须求解的问题,也因此可被视作与 next token prediction 等价的"智能完备"原语。

核心句型 · 9
1. No A, no B, no C.
“No studio capture, no green screen, no expensive calibration.”
三个并列名词短语连用否定,不带动词,节奏短促。适合在对比新旧方案时一口气列出被省掉的成本,比完整句更有力。
2. as few as + 数量
“Now with Atlas we can do this with just as few as three cameras”
as few as 强调数量之少(可数名词),少量用 as little as(不可数)。用于突出门槛下降,比 only 更能制造对比感。
3. What's key with X is that …
“But what's key with Atlas is that it actually has a kind of like specially grounded meaning to every frame”
先用 what's key 把焦点抬出来,再用 that 从句交代内容。适合在一堆相似方案中点出自己的关键差异,口语书面皆宜。
4. The thing that was limiting it was not …
“The thing that was limiting it was not the scale or the data or anything like that”
强调结构:the thing that … was(not)…,先设悬念再揭示。否定形式尤其适合纠正听者的默认猜测,随后再给真正原因。
5. It took … for sb to hit upon …
“It took a couple iterations for us to hit upon this formulation”
It took + 时间/轮次 + for sb to do,客观陈述代价。hit upon 表示「摸索着找到」,暗含并非一开始就想清楚,适合讲研发过程。
6. That's where … is really key.
“That's where simulation is really key for robotics”
用 that's where 把前文铺垫的问题接到解法上,是论证中的转接枢纽。仿写时前面必须先讲清一个困难,否则 where 无所指。
7. …, no matter how many …
“You're always going to miss something, no matter how many views you get”
让步状语后置,先给结论再堵退路。用于论证「这不是努力不够,而是原理上做不到」,是本场最关键的一步推理句式。
8. It's not … It's instead …
“It's not producing a static thing. It's instead a policy that's going to go out into the world”
先否定听者脑中的旧类比,再给新类比,两句独立更有停顿感。适合在引入一个容易被误解的新概念时使用。
9. One thing that's under appreciated is …
“One thing that's under appreciated on the website of the demos is the Stanford demo”
under appreciated(也写作 underappreciated)意为被低估。此句式用于把一个不起眼的细节提到台面上,随后紧跟解释为何重要。
词汇精讲 · 150 · 按出现顺序
contextualized /kənˈtekstʃuəlaɪzd/ adj. 0:00
被置于上下文中的;此处指像素带有空间上下文
grounded /ˈɡraʊndɪd/ adj. 0:00
有依据的、被锚定的;AI 语境中指输出受真实数据约束
calibration /ˌkæləˈbreɪʃn/ n. 0:27
标定、校准;此处指测定各相机的位置与参数
set out to phr. 0:27
着手做、立志做(某事)
reception /rɪˈsepʃn/ n. 0:27
反响、接受度;a warm reception 反响热烈
trajectory /trəˈdʒektəri/ n. 1:05
轨迹、路径;此处指相机运动路线
steer /stɪr/ v. 1:05
引导、操控;steer the model 引导模型输出
sparse /spɑːrs/ adj. 1:05
稀疏的;与 dense(稠密)相对
explicit /ɪkˈsplɪsɪt/ adj. 1:05
显式的、明确写出的;反义 implicit
show off phr. v. 1:41
展示、炫示(成果)
tripods /ˈtraɪpɑːdz/ n. 2:25
三脚架
reframe /ˌriːˈfreɪm/ v. 2:25
重新构图;引申为重新界定问题
splashing /ˈsplæʃɪŋ/ v. 2:25
飞溅、溅起
primitive /ˈprɪmətɪv/ n. 2:53
原语、基本单元;此处指模型训练的基础任务
arbitrary /ˈɑːrbɪtreri/ adj. 2:53
任意的、随意指定的
implicitly /ɪmˈplɪsɪtli/ adv. 2:53
隐式地、不言明地
bajillion /bəˈdʒɪljən/ n. 3:34
多得数不清(口语夸张造词)
tease apart phr. v. 3:34
细致地区分开、剖析清楚
myriad /ˈmɪriəd/ adj./n. 3:34
无数的、大量的(书面)
claim to fame phr. 3:34
成名的资本、最为人知的一点
interpolation /ɪnˌtɜːrpəˈleɪʃn/ n. 3:34
插值;此处指首尾帧之间生成中间画面
pose /poʊz/ n. 4:25
位姿(位置与朝向);camera pose 相机位姿
replication /ˌreplɪˈkeɪʃn/ n. 4:25
复现、精确复制
fly-throughs /ˈflaɪθruːz/ n. 4:25
穿行漫游镜头(相机在场景中飞行的画面)
governed by phr. 4:25
由……支配/决定
subfield /ˈsʌbfiːld/ n. 5:12
子领域、分支学科
diffusion /dɪˈfjuːʒn/ n. 5:12
扩散;diffusion model 扩散模型
natively /ˈneɪtɪvli/ adv. 5:51
原生地;指能力内建而非外挂
modality /moʊˈdæləti/ n. 5:51
模态(文本、图像、深度等不同类型的数据)
parameters /pəˈræmɪtərz/ n. 6:38
参数;相机参数指焦距、主点等内参
unification /ˌjuːnɪfɪˈkeɪʃn/ n. 7:10
统一、合为一体
synthesis /ˈsɪnθəsɪs/ n. 7:10
合成;view synthesis 视角合成
anchoring /ˈæŋkərɪŋ/ v. 7:10
锚定;anchor on sth 以某物为基准固定住
elegant /ˈelɪɡənt/ adj. 7:10
(方案)简洁而精妙的,非仅指外表优雅
layperson /ˈleɪpɜːrsn/ n. 8:15
外行、非专业人士
tackle /ˈtækl/ v. 8:15
着手解决(难题)
pencil out phr. v. 8:40
大致勾勒、粗略算一算(美式口语)
geometry /dʒiˈɑːmətri/ n. 9:16
几何结构;此处指空间的形状关系
emergent /ɪˈmɜːrdʒənt/ adj. 9:16
涌现的;指未被专门训练却出现的能力
downstream /ˌdaʊnˈstriːm/ adj. 9:16
下游的;指基础模型之后的应用环节
gazillions /ɡəˈzɪljənz/ n. 9:16
多到数不清(口语夸张,同 bajillion)
fidelity /fɪˈdeləti/ n. 10:27
保真度、还原度
delineation /dɪˌlɪniˈeɪʃn/ n. 10:27
刻画、勾勒(细致描绘)
interoperate /ˌɪntərˈɑːpəreɪt/ v. 11:33
互操作、互相打通(不同系统之间)
bottleneck /ˈbɑːtlnek/ n./v. 11:33
瓶颈;作动词指把输出挤过某个限制
bifurcate /ˈbaɪfərkeɪt/ v. 11:33
一分为二、分叉(书面)
blood, sweat, and tears phr. 12:19
血汗与泪水,形容付出极大心血
pros and cons phr. 12:19
利与弊、优缺点
conviction /kənˈvɪkʃn/ n. 12:19
坚定的信念;build conviction 建立信心
hit upon phr. v. 13:04
偶然想到、找到(好办法)
iterations /ˌɪtəˈreɪʃnz/ n. 13:04
迭代、一轮轮的版本
synthesize /ˈsɪnθəsaɪz/ v. 13:32
合成、综合生成
win out phr. v. 13:32
最终胜出、占上风
reap /riːp/ v. 13:32
收获(成果、好处);reap the benefits
brute force phr. 13:32
蛮力的、穷举式的(不靠巧法靠堆量)
productizing /ˈprɑːdəktaɪzɪŋ/ v. 14:17
把技术做成产品、产品化
reconciliation /ˌrekənsɪliˈeɪʃn/ n. 14:17
调和、使两者相容
priors /ˈpraɪərz/ n. 14:17
先验知识;data-driven priors 数据驱动的先验
akin to phr. 14:17
类似于、近似于(书面)
crevice /ˈkrevɪs/ n. 15:08
细缝、裂隙
tedious /ˈtiːdiəs/ adj. 15:08
冗长乏味、磨人的
exhaustive /ɪɡˈzɔːstɪv/ adj. 15:08
穷尽的、无一遗漏的(勿与 exhausting 混)
per se /ˌpɜːr ˈseɪ/ adv. 15:08
本身而言(拉丁语借词)
calculus /ˈkælkjələs/ n. 15:59
权衡算计;flip the calculus 彻底改变利弊账
unearth /ʌnˈɜːrθ/ v. 15:59
发掘出、翻找出(被埋没的东西)
footage /ˈfʊtɪdʒ/ n. 15:59
影像素材、片段(不可数)
aerial /ˈeriəl/ adj. 16:44
空中的;aerial view 航拍视角
quad /kwɑːd/ n. 16:44
(大学的)方形庭院,quadrangle 的简写
interplay /ˈɪntərpleɪ/ n./v. 17:27
相互作用、交互影响
triangulate /traɪˈæŋɡjuleɪt/ v. 17:27
三角测量(由多视角定位一点)
on the flip side phr. 17:27
另一方面、反过来说
set Ben loose phr. 17:59
放手让某人去干;set sb loose 撒开手
DSLR /ˌdiː es el ˈɑːr/ n. 17:59
数码单反相机
tangible /ˈtændʒəbl/ adj. 18:31
可感知的、切实具体的
crank /kræŋk/ v. 18:31
调高、加大(crank up the volume)
needle in a haystack phr. 18:31
大海捞针;AI 里指长上下文检索测试
principled /ˈprɪnsəpld/ adj. 18:31
有原则、有章法的(做法)
continuum /kənˈtɪnjuəm/ n. 19:17
连续统、连续谱(复数 continua)
jam /dʒæm/ v. 19:17
硬塞进去;jam sth in
extrapolated /ɪkˈstræpəleɪtɪd/ v. 19:17
外推得出(由已知推未知)
inconceivable /ˌɪnkənˈsiːvəbl/ adj. 19:17
不可想象的、难以设想的
gracefully /ˈɡreɪsfəli/ adv. 20:02
平滑优雅地;工程里指性能随规模平缓变化
auto regressive phr. 20:02
自回归的(逐步生成,每步依赖此前输出)
pick and choose phr. 20:02
挑挑拣拣、自行取舍
blows my mind phr. 20:49
令我大为震撼;blow sb's mind
extrapolate /ɪkˈstræpəleɪt/ v. 20:49
外推、由已知推断未知
hypothesis /haɪˈpɑːθəsɪs/ n. 20:49
假设(复数 hypotheses)
devils are in the details phr. 21:32
魔鬼藏在细节里,难点在具体实现
sign of life phr. 21:32
有起色的迹象;研发中指初步跑通
convicted /kənˈvɪktɪd/ adj. 22:20
深信不疑的(此处非「被定罪」,属硅谷用法)
paradigm /ˈpærədaɪm/ n. 22:20
范式、根本思路框架
compute /kəmˈpjuːt/ n. 22:44
算力(AI 领域的名词化用法,不可数)
backed up phr. v. 23:32
倒推回去(从截止日反推可行规模)
degradation /ˌdeɡrəˈdeɪʃn/ n. 25:04
质量下降、劣化
monolithic /ˌmɑːnəˈlɪθɪk/ adj. 25:48
单体庞大、不可拆分的
storyboards /ˈstɔːribɔːrdz/ n. 25:48
故事板、分镜
mood boards phr. 25:48
情绪板(用图片确定视觉调性)
keyframes /ˈkiːfreɪmz/ n. 25:48
关键帧
niche /niːʃ/ adj./n. 25:48
小众的;细分领域
persistent /pərˈsɪstənt/ adj. 26:33
持久保留的;persistent state 持久状态
props /prɑːps/ n. 26:33
(影视舞台的)道具
statefulness /ˈsteɪtflnəs/ n. 27:00
状态性(系统能保留先前状态)
ephemeral /ɪˈfemərəl/ adj. 27:00
转瞬即逝的、留不下的
assets /ˈæsets/ n. 27:00
资产;影视游戏里指可复用的美术资源
fabricated /ˈfæbrɪkeɪtɪd/ v. 27:51
制造、加工制作(也可指捏造)
painstaking /ˈpeɪnzteɪkɪŋ/ adj. 27:51
费尽心力的、一丝不苟的
arduous /ˈɑːrdʒuəs/ adj. 27:51
艰辛费力的(书面)
labor-intensive adj. 27:51
劳动密集型的、耗人力的
intuitive /ɪnˈtuːɪtɪv/ adj. 27:51
符合直觉的、一上手就会的
pottery /ˈpɑːtəri/ n. 28:42
陶艺、制陶
cater to phr. v. 28:42
迎合、满足(某类人的需求)
cabling /ˈkeɪblɪŋ/ n. 29:22
布线、接线作业
excruciatingly /ɪkˈskruːʃieɪtɪŋli/ adv. 29:22
极其痛苦地;excruciatingly slow 慢得难熬
laborious /ləˈbɔːriəs/ adj. 29:22
费时费力的
velocity /vəˈlɑːsəti/ n. 29:22
速度;科技公司常指研发推进速度
zoom out phr. v. 30:35
拉远看全局、跳出细节
randomization /ˌrændəmaɪˈzeɪʃn/ n. 30:35
随机化;机器人领域指域随机化
robust /roʊˈbʌst/ adj. 31:48
稳健的、抗干扰的
dynamical /daɪˈnæmɪkl/ adj. 31:48
动态的、随时间变化的
artifact /ˈɑːrtɪfækt/ n. 32:31
制成品、产物;此处指模型产出的静态成果
deployment /dɪˈplɔɪmənt/ n. 33:14
部署、投入实际运行
supercharged /ˈsuːpərtʃɑːrdʒd/ adj. 34:06
被大幅增强的、火力全开的
generality /ˌdʒenəˈræləti/ n. 34:37
通用性、普遍适用的程度
thesis /ˈθiːsɪs/ n. 34:37
论点、核心主张(不限于论文)
overwhelmingly /ˌoʊvərˈwelmɪŋli/ adv. 35:21
压倒性地、绝大多数地
glowing /ˈɡloʊɪŋ/ adj. 35:21
溢美的、极力称赞的;a glowing review
baked into phr. 35:56
内建、写死在……之中,难以事后更改
at odds phr. 36:45
相互抵触、不一致;be at odds with
factor out phr. v. 37:22
把某因素分离出去、剔除掉
latent /ˈleɪtnt/ adj. 37:22
潜在的、尚未显现的
temporal /ˈtempərəl/ adj. 37:22
时间(维度)的;与 spatial 相对
horizontal /ˌhɔːrɪˈzɑːntl/ adj. 37:52
横向通用的;指跨行业可复用的能力
conditioning /kənˈdɪʃənɪŋ/ n. 38:24
条件控制;给生成模型附加可控输入
editability /ˌedɪtəˈbɪləti/ n. 38:24
可编辑性
layman's /ˈleɪmənz/ adj. 38:24
外行的;in layman's terms 用大白话说
propagate /ˈprɑːpəɡeɪt/ v. 39:14
扩散、传播开去
party trick phr. 39:14
聚会上的小把戏,喻中看不中用
knobs /nɑːbz/ n. 39:14
旋钮;喻可调节的控制项
cutting-edge adj. 39:14
最前沿的、尖端的
axis /ˈæksɪs/ n. 40:01
轴、维度(复数 axes)
from scratch phr. 40:01
从零开始、白手起步
closing the loop phr. 40:28
闭环,把感知与行动接成回路
completeness /kəmˈpliːtnəs/ n. 40:56
完备性;Turing completeness 图灵完备性
reduce to phr. 40:56
归约到(把一类问题转化为某个问题)
NP-hard adj. 40:56
NP 难的;复杂性理论中的难度类别
mystery novel phr. 41:33
推理小说、悬疑小说
frame /freɪm/ v. 42:02
把……表述为、框定为;frame X in terms of Y
evolutionary /ˌevəˈluːʃəneri/ adj. 42:30
演化的、进化论的
equivalent /ɪˈkwɪvələnt/ n./adj. 42:30
等价物;等同的
phenomenal /fəˈnɑːmɪnl/ adj. 43:19
非凡的、了不起的
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.192Dario Amodei — “We are near the end of the exponential” 下一期 · NO.194 →Why Investors Are Rethinking Everything for the AI Era
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]