视频库 / NO.122ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.122
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Stuart Russell: Provably Beneficial Artificial Intelligence

节目发布 2019-10-16 · UC Berkeley EECS
斯图尔特·罗素 EEric Paulos
本期追问 · 点击跳到视频对应位置
只靠把神经网络电路做得更大,真能造出人类水平的智能吗?机器对人类想要什么越不确定,我们反而越能控制它吗?给机器一个固定目标去优化,是不是 AI 从一开始就走错的设计?代表多个人做决定时,该按谁的预测更准来分配偏好权重吗?
归入 Ⅲ·14 造物会听造它的人吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文整理自加州大学伯克利分校电子工程与计算机科学系(EECS)系列学术报告中的一场演讲。主讲人斯图尔特·罗素(Stuart Russell)是伯克利计算机科学教授、人类兼容人工智能中心(CHAI)创始人,与彼得·诺维格合著的《人工智能:一种现代方法》是全球使用最广的 AI 教科书;主持人埃里克·保罗斯(Eric Paulos)是伯克利 EECS 教授,研究人机交互与城市计算,为本场报告的联合主持。演讲围绕罗素新著《人类兼容:人工智能与控制问题》展开,从亚里士多德谈到关机问题,提出以「辅助博弈」重建 AI 的标准模型。本文依据现场录音编译整理,仅删去口语枝节与寒暄,论证与细节均予保留。

开场与讲者介绍

保罗斯:欢迎大家来到 EECS 系列报告。我代表我本人和坐在前排的联合主持人伊莱·亚布洛诺维奇(Eli Yablonovitch)向各位问好。先通知一件事:下周的讲者是华莱士·马歇尔。另外,顺便说一句,电又来了,真好。上周那位讲者的报告,我们正在设法重新安排,那些在家里摸黑等消息的朋友请放心。

今天的讲者其实不需要介绍,但我还是想介绍一下。优化、控制、目标函数,这些是我们计算机科学家天天琢磨的东西,而今天的讲者,我们自己的斯图尔特·罗素教授,谈的正是这个领域的一次根本性转向,以及我们思考这些问题的方式要怎么变。他是本系教授,在牛津拿的物理学学士,在斯坦福拿的计算机科学博士,在伯克利担任史密斯-扎德工程讲席教授,同时兼任加州大学旧金山分校神经外科的兼职教授。他是世界经济论坛人工智能与机器人委员会的副主席,牛津大学瓦德汉学院荣誉院士,卡内基学者,职业生涯中的荣誉一长串。

我觉得最贴切的形容是:他是一位极具感召力的计算机科学家,也是人文主义者、哲学家、挑衅者(我是在最好的意义上用这个词),还是一位行动者。他对人工智能做出了根本性的贡献,如今又就自主武器的威胁和人工智能的长远未来公开发声,大家在各种出版物上都能读到他的文章。他刚出了一本新书,《人类兼容:人工智能与控制问题》。我本想把书带来,可惜我买的是 Kindle 版。好了,不再多说,有请斯图尔特·罗素教授。

从亚里士多德到图灵

罗素:谢谢埃里克。我先讲一点历史,这个背景很有用。最近新闻里大量谈论 AI 的好处和坏处,但这件事并不是这两年才开始的,它由来已久。大约公元前 340 年,亚里士多德就在讨论智能自动化对就业可能造成的后果,这是我们所知最早关于「技术性失业」的文字。他说,如果乐器能自己演奏,织机能自己织布,那我们就不需要奴隶了。对当时的就业数字来说,这可是个大问题。

我在书里说,亚里士多德要是有一台计算机,再加上一点电,他就会是个 AI 研究者。你读他的著作,会发现他在谈规划算法,谈带有形式语法和语义的逻辑推理,谈前向链和后向链,谈分类层次、语义网络、本体论,写得清清楚楚。可惜他没有计算机。

巴贝奇至少有一份计算机的设计图。巴贝奇和洛夫莱斯很明确地讨论过,凡是人类心智能做的事,这种机器都能做。也就是说,他们很清楚人类水平的 AI 是可以实现的,只要那些齿轮能咬合起来、转得够快。巴贝奇最终没造出他的机器,但发明的消息和预言传了出去。伊利诺伊州有一份宗教报纸叫《原初阐释者》,编辑听说了巴贝奇的想法,预言这种机器一旦造出来,终将统治世界。这是我们所知最早见诸印刷的 AI 末日论。

然后又等了一百年,才有了真正的计算机,这要部分归功于第二次世界大战。图灵以许多事情闻名,其中之一就是 1950 年那篇提出「图灵测试」的论文。不那么为人所知的是,他在同一篇文章里还讨论了真正实现人类水平 AI 之后会发生什么,而他的态度是彻底认命的。他说,我们应当预期机器会接管控制权。所以,现在埃隆·马斯克站出来说类似的话,就被人骂成什么都不懂 AI 的白痴。那他们是不是也要把图灵骂成不懂 AI 的白痴?骂完了还剩谁?

图灵早就指出了这一点。这个领域本身是 1956 年在达特茅斯正式诞生的。此后我们经历了一波又一波的乐观和失望。眼下正处在乐观期,接下来会怎样很难说。我觉得现在还悬在刀锋上,不知道这股乐观会消散,还是会有足够多的新成果持续涌现,把势头撑下去。

AlphaGo 与各国投资热潮

罗素:近年最大的几件事,其实是媒体事件,不是研究上的突破。外界完全不清楚这个领域内部发生了什么。外界把 AlphaGo 当成研究突破,尽管它建立在塞缪尔 1957 年和勒昆 1992 年的工作之上。真正的研究突破发生在七十年前和二十年前,他们把这些东西拼在一起,做成了一件大事:击败世界最强的围棋棋手。

这是中国的「斯普特尼克时刻」。正是这件事让中国醒过来,说:AI 这东西是真的,我们要主导它。于是他们宣布未来十年投资约一千五百亿美元。美国以国家 AI 研究院的新计划予以回击,投入一亿两千四百万。我们都知道,媒体常常分不清「百万」和「十亿」,现在看来美国政府也把这个重要区别搞混了。反正我相信会有帮助的。英国有它的投资计划,法国是十五亿欧元,欧盟是一百八十亿英镑,当然还有中国。如今翻开报纸,不可能看不到 AI 的头条。身在这个领域,眼下是个荒唐的时代。

对抗策略实验:AI 能力被高估

罗素:所以我想注入一点现实主义。这个例子没有李世石落败那么轰动,但也不小:OpenAI 展示了可以从零开始训练人形机器人。它们就像刚出生的婴儿,完全没有运动控制能力,几个小时之后,就能学会行走、把球踢向球门、把球拦在门外。很多人为此非常兴奋。看上去确实挺酷,行为是有目的的,踢球的一方会根据球的位置调整,守门员追踪得也不错,就是稍微有点不协调。

于是我们想,这东西到底有多真?我的学生亚当·格利夫(Adam Gleave)说,我们只改红方的程序,蓝方的程序一个字不动,看看能不能把比赛打得均衡一点,因为现在通常是蓝方赢。答案是:红方直接躺倒在地上,把腿伸到空中乱晃。注意,蓝方程序完全没变,可你看蓝方现在在干什么?它像是「啊、啊、啊,哦,哦,不好!」,彻底乱了阵脚,明明是同一个程序。

这告诉我们,人们对 AI 系统的行为和能力的感知往往过于慷慨。你看到一次良好的表现,就说「它很会踢球」,可它并不会。性能上存在巨大的漏洞,而且不用费多大力气就能找到。我们必须谨慎得多。想象一下自动驾驶汽车:你在山景城训练它,它表现很好,然后你做点奇怪的事,比如换一下路牌的颜色,它突然彻底发疯,在高速公路中间转圈漂移。在弄清它究竟在做什么之前,你并不知道它是不是真像你以为的那么好。这是一个非常重要的教训。

数据不是石油,自动驾驶的可靠性关

罗素:《经济学人》的封面说「数据是新的石油」。但我们要更小心一些。「谁拥有最多的数据谁就赢得世界」,这句流行的口号有一个具体的技术理由站不住脚。眼下各国的地缘政治战略正建立在这个口号之上,但它并不成立。事实上,AI 系统越好,需要的数据应该越少,而不是越多。人类学一个新的视觉类别只需要一两个例子。你第一次见到长颈鹿,不需要再看十八万张长颈鹿的图片才明白长颈鹿是什么。看一眼就够了。你去问任何一个心理学家,人类需要多少训练样本?一个,有时候两个。所以,随着 AI 系统能力增强,它们需要的数据会变少。光是拥有海量数据,得不到你想要的东西。

如果有什么东西会戳破这个泡沫,我认为可能是自动驾驶汽车事业的失败。不是说这个问题不可能解决,而是它必须解决得足够快,快到每年投入几十亿美元的人还没耗尽耐心。所以这里有一场小小的赛跑,一边是投资者的耐心,一边是系统的性能和安全。结果如何我不知道。他们的进展非常可观,但要达到八个 9 的可靠性,而他们现在只有五个或六个 9。听上去差别很小,但换算过来是错误率要降低一百倍或一千倍,这一步很大。

未来十年:仓库、助理、语言、卫星

罗素:话虽如此,我还是相当乐观地认为,我们会看到越来越多的系统,不只是那种说服你买下一辈子都用不完的卫生纸的 AI,而是真正走出工厂、进入现实世界的机器人。道路是一个场所,但仓库很可能会先到来。机器人已经在仓库里做简单的搬运工作了,只是不太依赖感知。而「从箱子里拣出任意物品」这个问题,我认为我们正在解决。这听起来不难:一大箱东西,你说,把西瓜拿出来,或者把水枪拿出来。可一旦有系统能做到这一点,一千万个工作岗位就没了。那会是非常大、非常显眼的一步。

家庭大概是最后才到的。我们会看到一些玩具式的小应用,比如调酒送酒的小机器,但那是非常蠢的机器人。家是一个变化极大、不可预测、物理上极其复杂的环境,要做出真能在家里运转的系统,比前面那些问题都难。

数字这一侧,我认为会出现真正的智能私人助理,它能充分理解你的生活、你的活动、你的人际关系、你的承诺和你的通讯,从而真正派上用场。不像现在多数系统那样,只是搜索引擎的一个语音界面,像鹦鹉学舌,而是成为你生活中的伙伴。高管们有昂贵的人类私人助理,你将会以每月九十九美分得到同样有用的东西,而且人人都能拥有。多数人其实更需要它,因为他们的生活更艰难。

如果要说下一个十年最有可能被解决的问题,我认为是语言理解的下一个层次。回顾即将结束的这个十年,最大的进步是视觉物体识别。下一个十年,我想是从语言中提取内容的能力:读一段文本,生成数据库条目或逻辑断言之类的东西。这不是深度理解,它们不会读乔伊斯然后写篇论文,但它们能读遍网上所有的文档、每一份报纸、每一种语言的每一档电视节目。对人类来说,这将是一种不可思议的能力。如果说搜索引擎值一万亿美元,现在可能两万亿,那么这东西对人类的价值是它的十倍。

还有一件有意思的事。我最近加入了 Planet 公司的董事会,这是一家旧金山公司,拥有世界上数量最多的卫星。现在已经可以每天把地球的每一平方英尺都拍一遍。如果人手够多,大约三千万人,你就能看完所有这些图像,掌握地球上所有东西的动向。而把计算机视觉用在这个数据流上,我们就能把地球变成一个持续更新的数据库。在它之上,可以构建成千上万种应用。我们正在和联合国合作,设想如何把它用于可持续发展目标:城市规划;在非洲平原上管理牲畜,那里没有围栏,是个很棘手的管理问题;反盗猎;管理航运;查缉走私。这份清单可以一直列下去。

这些应用如果各自单独去建,每一个都要花十亿美元,因为收集、存储、处理这些数据是一门昂贵的生意。但如果只做一次,然后把结果提供给任何想在上面构建应用的人,那么做出一个全天候、高价值的全球服务大约只要一百万美元。我认为这很快就会发生。

算力堆叠为何造不出通用 AI

罗素:那么,这一切是否意味着人类水平的 AI 就在眼前?不同的人有不同看法。比如伊利亚·苏茨克维(Ilya Sutskever),他曾与杰弗里·辛顿一起做出视觉物体识别的突破,现在是 OpenAI 的首席科学家。他认为还有五年。他们刚从微软拿到一大笔投资,其中很大一部分用来扩充本已庞大的计算资源。我说庞大,是指一个谷歌 TPU 集群相当于一千万台笔记本电脑。它并不算大,大约十二英尺长、八英尺高,能放进你的卧室,却已经比两年前世界上最大的超级计算机还强。这些数字是天文数字。每秒十的十七次方次运算,与大脑理论上能达到的最大状态变化次数在同一量级。而他们还打算再扩大一千倍,并且认真相信这样就能达到人类水平。

我完全不信。因为那些是电路。计算机科学教给你的东西里,如果只有一条,那就是:电路是电路,程序是程序,程序要强大得多、表达力也丰富得多。设想用电路语言写国际象棋的规则,要写几十万页,因为每一个格子都得有单独的一块电路,每个格子、每一步、每一个兵,一切都要重复。这种表达力贫弱的语言带来的膨胀是荒唐的。而用编程语言写,一页;用英语写,一页;用一阶逻辑写,一页。这些语言篇幅相当,是有原因的:应对一个充满各种「事物」的庞大复杂世界,正需要这个级别的表达力。事物,事物极其重要,世界里有大量事物。这意味着你需要一阶逻辑那样的表达能力。命题逻辑,也就是电路的语言,里面没有「事物」。你无法在电路语言里谈论「所有的兵」「所有的格子」「所有的时间步」「所有的人」,或者所有的驴、所有的长颈鹿。所以我的看法是,靠把电路越做越大来解决人类水平 AI 的问题,没有任何可能。

我们需要的是真正的概念突破,而这类突破很难预测。你不能画一条摩尔定律式的曲线,说「看,它在 2029 年越过人类智能」。你只能等待概念突破。我列了几项,不逐一细说,但最大的一项大概是第三条,因为正是它使人类得以在现实世界中成功运作:我们能在跨度极大的尺度上无缝行事,小到单个运动控制动作(打字时,你的大脑要向手指发送一整串复杂指令,每条指令的尺度是几毫秒),大到读一个博士学位(五年,大约一万亿条运动控制指令)。这一万亿是在指数上的。学过 AI 的人都记得,b 的 d 次方,b 是分支因子,d 是解的深度,这里 d 就是一万亿。单靠更多算力,绝无可能把规模扩展到那里。人类靠的是在多个抽象层次上无缝推理。如果有人提供了抽象层次,给你一个不同尺度上的动作层级,我们大致知道怎么把它们巧妙地串起来。但这个层级从何而来,机器如何在行进中自己发展出层级,我们至今毫无头绪。在我看来,这是最大的开放问题。它可能被解决,也许在我讲话的时候已经有人解决了。这类事情可以发生得很快。

核能史的教训:别赌人类无能

罗素:举个例子说明这一点。回头看我们上一次发明足以终结文明的技术,也就是核能。二十世纪初的共识是:他们知道能量就在那里,有 E 等于 mc 平方,能测量不同同位素的质量,能精确算出引发某个跃迁会释放多少能量,但他们绝对相信这是做不到的。卢瑟福,分裂原子的人,诺贝尔奖得主,大概是最著名的核物理学家,9 月 11 日在莱斯特发表演讲,有人问他,未来二十五到三十年内有没有可能做到?他说,没有,这么想都是痴人说梦。爱因斯坦同意他,说自己想不出任何可以想象的办法引发这种跃迁。第二天早上,利奥·西拉德在《泰晤士报》上读到这条新闻,出门散了个步,过马路的时候,发明了核链式反应。从「所有顶尖物理学家都认为这完全不可能」到基本解决,只隔了一夜。而仅仅十二年之后,第一颗原子弹就爆炸了。

所以,别赌人类的智慧会失败。有些人对 AI 可能带来的任何风险持怀疑态度,他们的论证之一就是:有能力的 AI 根本不可能,所以不必担心。我觉得这种论证很怪。它等于是说:没错,我们正开着一辆大巴,载着全人类,以最快的速度冲向悬崖,几千亿美元的投资用于创造这项技术,油门踩到底,但我向你保证,我们会在冲下悬崖之前把油耗光。你会上这辆车吗?没有任何论证能说明人类水平的 AI 不可能,也不可能有,因为我们知道它是可能的:我们的大脑就能产生这个水平的智能。你不能说这在物理上不可能。所以这是一个令人失望的发展:连 AI 领域内部的人,为了回避谈论风险,都愿意说「AI 会失败」。这很怪。想象一下癌症生物学:我们这个时代最顶尖的癌症生物学家站起来说,你知道吗,我们永远治不好癌症,请继续给我们大量经费,但我保证我们永远治不好癌症。你在说什么啊?这就是我们面对的局面。

我认为,审慎的做法是承认你无法保证。当然,我们可能在那之前就毁灭了自己,但审慎的假设是:我们终将造出在决策能力上根本超过人类的 AI 系统。它们显然会掌握多得多的信息,能比人类看得更远,就像它们在棋盘、围棋盘和电子游戏里已经做到的那样。我认为这最终会转化为现实世界中的决策能力。

AI 的上行空间:GDP 十倍与资源逻辑

罗素:这可能是好事。如果没有好处,我们根本不会有这场对话,因为如果没有好处,我们不会花这么多钱,也没人会做 AI。所以好处当然是巨大的。可以这样说:我们的文明就是智能的产物。如果突然获得多得多的智能,就能拥有好得多的文明。所以,别去想那些小改进,比如更好的医疗诊断、更安全的汽车。是的,那些很好,但不够有雄心。想想旅行。两百年前你想去澳大利亚,那是一个耗时数年、花费几十亿美元的项目,死亡概率大约百分之八十。现在你想去澳大利亚,掏出手机,点几下,明天就到了,而且相对于过去,几乎是免费的。想象同样的变革发生在一切事物上,一切现在还很昂贵的事物,比如建设工程:我们需要一所新医院,需要一条路把村子和大城市连起来,需要这个,需要那个,我们的学校没有老师。凡是现在困难、昂贵、耗时的事,AI 系统都可以为我们提供。

我说的不是 AI 发明癌症疗法或超光速旅行之类的科幻,只是把我们已经会做的事大规模地提供给所有人,这就意味着 GDP 增长十倍。仅仅让所有人达到伯克利的生活水准,就是世界 GDP 的十倍,折合净现值一万三千五百万亿美元。这就是它值多少钱的现金等价物。中国之所以在这上面投入几千亿美元,原因就在这里。看看奖池的大小,那个数字微不足道,任何以十亿计的数字与可能创造的价值相比都微不足道。所以我很确定,随着事情推进,限制我们的不会是研究经费,也不会是产业界能投入的钱,而是我们能不能有足够多的人,把这场变革真正落到实处。

在那样一个世界里,物质财富就像报纸的电子副本。为自己有几份报纸电子副本而争斗是疯了,想抢占更大比例的电子副本也是疯了。所以我认为这会改变历史动力中与资源竞争相关的那部分,与获取「让生活在物质上值得过」的东西相关的那部分。我们仍然可能为宗教原因互相残杀,但这一切可以改变。唯一不变的是土地。我们可能为了土地互相残杀,因为至少目前,AI 系统造不出更多的土地。

自主武器与人类衰弱的风险

罗素:人们谈论过许多坏处:扰乱我们对现实的认知,直截了当地互相杀戮。我本来想放一段短片,叫《屠杀机器人》(Slaughterbots),YouTube 上能找到。它用政策制定者能懂的简单方式说明一件事:如果你造出能在无人监督的情况下自行定位、选择并杀死人类目标的武器,那么,你们都是计算机科学家,你们懂什么叫可扩展性。你们知道谷歌不是靠十亿个员工来回答我们发去的那些问题。那几十亿个问题他们是怎么处理的?买更多的硬件。所以,正如谷歌每天能回答几十亿次查询,你也可以靠买更多这种武器来杀死几百万人。你不需要庞大的军队和军工复合体,只要买得够多,把它们放出去干活,就拥有了一件大规模杀伤性武器。

可惜音响坏了,我不放整部片子了。但我想指出,2017 年这部片子出来的时候,俄罗斯驻联合国大使抱怨说这纯属科幻,三十年内都不会成为现实,我们把这种东西拿到联合国来是对他的侮辱。而现在,你已经可以买到比片中武器更糟糕的东西。土耳其公司 STM 正在生产 KARGU 无人机,销售材料里宣传它的能力:自主打击,基于图像选择目标,追踪移动目标,反人员,人脸识别。他们还宣布打算用它来对付库尔德人,原定 2020 年初,也许托特朗普总统的福提前了。所以我们可能会在叙利亚第一次看到这类自主武器对人类发动攻击。

还有就业的终结。有邪恶的人滥用 AI 来统治或毁灭世界。还有过度使用 AI 导致人类逐渐衰弱。我认为最后这个问题非常难办,因为我们天生懒惰。即使知道长远来看会衰弱,你还是总能说:好吧,就今天,你帮我系鞋带吧,我明天再学。几千年来我们都知道,人有这种把「怎么把事情办好」的知识交给别人的倾向。而现在,我们可以把它交给机器,这是从来没有过的事。要让文明延续,以往你只能把它交给下一代人类。为了维持文明运转,我们已经做了大约一万亿人年的教与学。现在不必了,可以把它交给机器,让机器来运转文明。而这一步可能是不可逆的。

推荐算法改造人:目标设错的代价

罗素:最后十分钟左右,我来谈可能最严重的问题。人们为什么写那种标题?马斯克说「召唤恶魔」,他在说什么?他的意思是:你正在制造比自己更聪明的东西,你打算怎么控制它?你打算如何永远保持对一个本质上比你更强大的东西的权力?这就是我们面对的问题。

我们那些友善的社交媒体公司已经给我们发出了一个小小的警告,这么做真是有公益精神。他们在全球范围内部署了一些相当简单的机器学习算法,用于优化一个特定目标,也就是点击率。你可能会想,怎么优化点击率?算法只是学习人们想点什么,不给他们推不感兴趣的东西,听起来还行。但即便如此,也会造成回音室和过滤气泡,你永远接触不到自己舒适圈之外的想法。

而实际情况糟糕得多、多、多。因为算法做的不是这个。算法做的是改造人:不是学习人们喜欢什么,而是改变人本身,让他们变得更可预测,因为这样能从他们身上赚更多钱。这不是什么邪恶的心理学阴谋。算法不知道人类存在,不知道人类有心智。在算法眼里,一个人就是一段点击历史。我们只是这个:你被推送过什么,你点了什么。而事实证明,要产生更多点击,办法是给人推送越来越极端的内容。他们可能一开始处于某个光谱的中间,比如看 YouTube 视频,你给他们推一点更暴力的,再更暴力一点,他们的重心逐渐移动,直到成为极端暴力视频的瘾君子。这是有据可查的过程,正在政治领域发生,也在其他媒体消费领域发生。

原因在于,你给一个足够强大的系统设定了一个目标,而这个目标如果是错的,就会造成巨大的附带损害。这就是我们担心的问题。对人类来说,这不是新问题。我们有迈达斯王的传说,他的目标是「我碰到的一切都变成金子」,听起来不错,直到你意识到这包括你的食物和饮水,然后你就死了。还有给你三个愿望的精灵,第三个愿望是什么?是「请撤销前两个愿望,因为我许错了」。许许多多的文化都有这类故事,它们指向同一件事:我们无法正确地陈述自己的目标,无法预见目标可能被以何种方式优化,以及随之而来的种种附带损害。

标准模型的错误与三条新原则

罗素:所以我要说的是,从五十年代起,我们定义 AI 的根本方式就错了。我们大致是这样说的:在这个语境下,人类智能指的是选择能够预期实现我们目标的行动的能力。这是经济学的理性概念,也是哲学的理性行为概念。我们把它直接照搬到机器上:机器是智能的,当且仅当其行动可以预期实现其目标。这不仅是 AI,控制论也是这样,统计学是最小化损失函数,运筹学是最大化奖励之和,甚至经济学也是最大化利润或福利、GDP。你设定一个目标,然后造出优化这个目标的机器。而如果目标设错了,系统又比你强大,你就进入了一场棋局。我们知道结果:机器实现它的目标,而你实现不了。这就是问题的本质。

这其实就是糟糕的工程。这种工程方法只有在人类能够正确设定目标时才有效,可我们知道我们做不到。你不会设计一架只有长了七只手才能开的飞机,那是愚蠢的设计,因为你没有七只手。那为什么整个领域的设计都要求我们正确设定目标,否则就要承受这些恶果?

所以我们必须面对一个事实:目标在我们身上。我们可以提供线索,其中一些会是错的;系统可以从证据中,比如我们自己的选择行为中,推断出我们真正的偏好和目标。但从根本上说,机器对自己应该做什么,也就是对人类有益,是无知的。所以我们想要的是这样一句话:机器是有益的,当且仅当其行动可以预期实现我们的目标。这是规格说明。

我把它写成三条原则,因为据说最多只能有三条。第一,机器人的目标是造福人类,或者说满足人类偏好。我说的偏好,不是你想吃哪种披萨,而是你希望未来如何展开,你想要哪种未来,不想要哪种未来。第二,机器人对这些偏好始终是不确定的。第三,偏好有一个落脚点:我们的偏好会以某种形式体现在行为中。我们做出的每一个选择,包括你选择坐在这里听我讲,都是关于你潜在偏好的证据。这是不完美的证据,因为我们并不理性行事,我们的行动并不总能最大限度地满足我们的偏好,事实上几乎从来做不到。所以,理解人类行为提供的证据是一个复杂的问题。

辅助博弈:回形针实验与关机问题

罗素:把这些变成数学,就得到我们所说的「辅助博弈」(assistance game)。博弈是涉及多个主体的决策问题,这里的主体是人类和机器。从数学上讲,这是一个人类持有收益函数、机器试图优化它却不确切知道它是什么的博弈。要点在于:当你解出这个博弈中机器的那一半,无论解得多好,哪怕以远超人类的方式去解,机器仍然对我们有益,仍然会顺从我们。这是关键。没有天花板,你不必说「我们不能把机器做得比这更聪明,以免失去控制」。这是另一种机器。我希望,长远来看,它们是可证明有益的。当然,人类也是博弈的一部分,我们也得想清楚当机器存在于世界中时,人类该如何行事。

用图形模型来表示。这是传统图景,也是教科书前三版的图景:人类行为是人类目标的结果,但我们假设人类目标是被观测到的,所以那个节点是填实的,表示我们有它的证据。一旦你有了这个变量的证据,机器就去追求它所感知的人类目标;如果你观测到了,或者至少相信自己掌握了人类目标的正确值,那它就是充分统计量,人类行为就变得无关紧要了。你可以跳上跳下地喊「不要,你会毁灭世界」,但机器有它的目标,它做的一切都是对的,它会一路追求下去,你无关紧要。而当目标不可观测时,从数学上讲,这些变量仍然耦合在一起。它们不是绝对独立的,只是在给定目标的条件下才条件独立。所以这里的解都涉及机器与人类之间的耦合,这不可避免。它让问题更复杂,你不能跑去单独解马尔可夫决策过程那种单主体决策问题。更复杂,但无法回避。

我跳过谷歌把人标成大猩猩的那个例子,直接讲辅助博弈。记住,拥有偏好的是人类,我们把偏好统称为 θ,假设人类根据 θ 行事,不必完美或理性。机器要最大化人类偏好,并对人类偏好持有某个先验 p(θ)。当你求解均衡,在人类一侧,你会看到人类在教机器人,因为人类希望机器人多了解 θ,好变得更有用。机器人会学习,会把人类的教导行为理解为传递偏好信息,会提问,会请求许可,会顺从人类,而且如我将要说明的,它会允许自己被关掉,这是我们能控制它的关键标志。

如果你了解逆向强化学习,它是这个模型的单向版本:机器隔着窗户观察人类做自己的事,而人类被视为一个孤立的个体。这个假设一般来说不成立。一旦人类意识到有机器存在,人类的行为就只能被解释为这场博弈的解。想象你是医学生,站在专家外科医生旁边。外科医生不会只顾埋头缝、缝、缝,切、切、切,缝完收工。他们会开始解释:看,如果你这样做,血就全流出来了,别这么做。这些举动,如果你假设人是孤立行事的,就说不通;它们只有作为这场博弈的解才说得通。所以你只能通过求解博弈,来看人类那部分该如何解释。

我用一个非常简单的博弈来说明,这是我们找到的能体现这一现象的最简单的例子。偏好只有一个维度:回形针和订书钉之间的汇率。θ 就是这个汇率的值,假设 θ 等于 0.49,你可以理解为一个回形针值 49 美分,一个订书钉值 51 美分。机器人完全不知道你的 θ 是多少,有人喜欢回形针,有人喜欢订书钉。博弈中人类先行动,可以选择做两个回形针、各做一个、或者做两个订书钉,三个选项。然后轮到机器人。

如果只有人类自己,会怎么做?回形针 49 美分,订书钉 51 美分,那就做两个订书钉,值 1.02 美元,从人类自身角度看,1.02 美元占优。但在这个博弈里,人类做完之后,机器人可以做 90 个回形针,或者各做 50 个,或者做 90 个订书钉。那么人类现在该怎么办?这取决于机器人会如何解读人类的行为。人类希望机器人做出让自己最开心的那种东西,可它怎么把 θ 的值告诉机器人?它唯一的选择就是做两个回形针、各做一个、或者做两个订书钉,它没法说「我对回形针的估值是 49 美分」,只能从三个动作里选一个。同样,机器人该如何解读人类的行为?解决这个问题的唯一办法就是求解博弈,找到纳什均衡,也就是一对策略,在对方策略不变的情况下,任何一方都不会改变自己的策略。这个博弈只有一个纳什均衡:当人类对回形针的估值是 49 美分时,选择「各做一个」。事实上,只要你的汇率落在 0.446 到 0.554 之间,「各做一个」就是最优的。也就是说,从博弈的解中,一套编码浮现了出来。这套编码告诉机器人「我的汇率落在哪个区间」,好让它做出正确的事。机器人则假定人类通过这个选择正确地编码了自己的偏好。你可以看到,「各做一个」并不是人类孤立行事时会做的选择。

我再简要说明这类机器的另一个重要性质,也就是关机问题。所有足够大、足够壮的机器人,比如我们的 PR2,重 440 磅,你可不想被它不小心压扁,所以它有一个关机开关。如果用经典方式编程,给它一个目标,比如「去拿咖啡」,那么只要它足够聪明,就会意识到:有一种情况会让我拿不到咖啡,就是有人把我关了。于是它的第一件事就是禁用自己的关机开关,以此消除失败的可能。这是本次演讲的一条教训:死了就拿不到咖啡了。

来看当机器人对行动的价值(从人类角度看的价值)存在不确定时会怎样。我们设置一个非常简单的决策问题。机器人先行动,起初有两个选择:按下自己的关机开关,也就是自杀,以免对人类造成危险;或者做一个可能有危险的动作,比如「这儿有个大红按钮,上面写着『按我』,那按一下大概是好事」。这由一个概率分布表示,正的那一侧概率质量更多。如果只有这两个选择,而我们把自我关机的价值设为 0,那么机器人会去按大红按钮,而且不会让你关掉它。这正是我们要避免的。

所以我们给它第三个选择:什么都不做,允许人类来关掉它。问题是,机器人为什么会这么做?既然它自己也能关掉自己,而它又不想关掉自己,为什么要让人类来关?这看起来违反直觉。答案是:如果人类没有关掉它,机器人就知道它准备做的那个动作,也就是按大红按钮,是好的。这就抹掉了概率分布中「这个动作从人类偏好来看不可取」的那部分负值。于是,等待被证明严格优于另外两个选项,允许自己被关掉实际上比不允许更好。这是一个非常简单又非常一般的定理,与信息期望价值非负定理是同一个定理。因为无论人类关不关你,都在向你提供关于人类潜在偏好的信息,而这信息正是你对人类有用所需要的。

这里的安全余量,机器之所以允许我们关掉它,是因为它对我们的偏好不确定。一旦它对我们的偏好变得确定,就不会再允许自己被关掉了。所以不确定性与安全、与我们控制机器的能力之间存在直接关系。

多人偏好聚合与预测准确者胜出

罗素:可以想见,相关研究非常多。我们基本上要把 AI 教科书的每一章都按这种新思路重写一遍。因为每一章,我可以保证,因为都是我写的,每一章都建立在目标已知的假设之上。看搜索算法,有一个目标和一个代价函数,你得先知道这两样才能调用算法。可如果你不知道呢?我们还没有相应的理论。所以 AI 的每一种形式都有大量工作要做,比如损失函数不确定的机器学习:把人叫成大猩猩,和把苹果叫成橙子,哪个更糟?算法起初可能不知道,它对损失函数是不确定的。

我们还必须面对人类不完美这个事实。我说过,人的行为并不完美反映其潜在偏好,所以你得逆向工程人类认知。就连李世石也会下出输棋的着法。这不意味着他想输,他想赢,只是不够聪明,所以下出了输棋的着法。你得理解这一点,才能解释他的行为。

我们还必须处理「有很多人」这个事实。这非常重要。校园里那么多社会科学院系,甚至一些人文院系,就是围绕「人有很多」这个事实存在的。怎么处理?举一个简单例子:一个机器人服务多个人,他们都想当宇宙的统治者,可他们不能都当宇宙的统治者,你该怎么决策?代表不止一个人行事时,你必须做取舍,该怎么取舍?

约翰·海萨尼(John Harsanyi)是伯克利的经济学教授,诺贝尔奖得主。我从未见过他,但在做这项工作的过程中读了他很多东西,一位极其睿智、极有洞察力的人。他证明了有时被称为「社会聚合定理」的结果:当人们拥有共同先验,也就是所有人对未来如何展开持有同样的信念、对未来状态序列持有同样的概率分布时,所有帕累托最优的策略,也就是所有不被其他策略严格支配的策略,最终看起来都像是对个人偏好取线性组合。纳什议价理论认为,线性组合中的权重取决于个人的议价能力,比如他们退出整个安排的能力。海萨尼则基于平等,主张权重应当全部相同。这是福利经济学、公共政策以及许多其他领域的基本定理。

结果是,当人们没有共同先验时,你会得到完全不同的解。我们去年发表的结果是:所有帕累托最优策略,对个人偏好所赋的权重,反映的是那个人的预测与现实的吻合程度。这是一个定理,我们改不了,你可能不喜欢它。它意味着,如果你对未来持有怪异的信念,而且不断被证明是错的,那么任何代表你和其他人行事的系统都会降低你的偏好的权重。为什么?因为每个人都相信自己是对的,所以他们会同意这一点。他们会同意一项奖励「说对的人」的政策,因为每个人都认为自己就是那个人,没有人真心相信自己的信念是错的。于是你会同意一份条件合同:如果你的预测对了,你拿奖;如果他的预测对了,他拿奖。你们俩都同意,因为你们都认为自己会拿奖。这就是代表许多人行事时必须遵循的数学事实。

它的社会后果如何,我还不知道,甚至还没和社会科学家讨论过这个定理,但它是真定理。它也很有用。这意味着你可以在美国和俄罗斯之间谈合同,因为双方都认为对方会作弊而自己不会。你可以谈成双方都会同意的条件合同,尽管他们在事实上、在对彼此的看法上都不一致。

我得跳过利他、冷漠和施虐这几节,感兴趣的话可以读书。我也没带书,但我有一张照片,而且书在网上。哦,你们那儿有一本,这是美国版,刚才那本其实是英国版。谢谢兰迪。

假如我们运气好,把处理多个人之类的难题都解决了,就得到一个非常好的利他机器人,它对每个人的偏好赋予同等权重。你工作了漫长的一天回到家,它来迎接你。你说:今天真糟糕,我连吃午饭的时间都没有。它说:那你一定很饿了吧。你说:饿坏了,晚饭有什么吃的?它说:有件事我得告诉你。接下来是什么?索马里有人比你更急需帮助,所以请你自己做晚饭。这是个问题。一个对全人类偏好一视同仁的利他机器人,大概不会太在意你这些西方中产阶级的舒适需求。这是真问题。你刚花五万美元买了这个机器人,它做的第一件事就是跑去索马里,你不会高兴,那家公司很快就会倒闭。所以我们必须解决这个问题。你不能把简单的功利主义、利他主义方案塞进机器,然后指望会发生什么理智的事。

总结一下:AI 领域正在发生很多很酷的事,进展迅速,还有很多好东西在路上。很难预测我们何时会拥有通用人工智能。但我在书里论证了,它将是人类历史上最大的事件,而我们对此完全没有准备。这是一个严重的问题,我们正在设法解决。今天讲的是一种方案,也许能让我们永远保持对任意智能的机器的控制。我提到的其他问题,滥用与过度使用,我没有任何接近解决方案的东西。它们是社会和政治问题,涉及治安、文化等等,不是技术问题,尽管更好的网络安全方案在滥用这一侧会有所帮助。这些问题我们都必须面对,而且越早开始越好。谢谢。

问答:电路与大脑

保罗斯:非常感谢。我们稍微超时了,但这实在太精彩了。我想我们还有时间回答一两个问题。这是我们神奇的新话筒,这样拿着,然后提问。有人有问题吗?不问就睡不着觉的那种。

听众:我是做电路设计的,不是计算机科学出身。能不能回头澄清一下,您说更大的电路无法以正确的方式编码信息,具体是什么意思?在什么意义上,大脑不是像 TPU 那样的电路?

罗素:你说得对,你的笔记本电脑是电路,TPU 是电路,你的大脑也是电路。但至少就计算机而言,它是一个实现了更高抽象层次的电路,那个层次就是编程语言,无论是汇编语言还是高级语言。这些语言给你循环之类的东西,而循环让你能够枚举任意多的对象。看看用 C++ 实现的国际象棋规则,里面到处是循环。如果不能用循环,而是循环的每一次迭代都要有一块单独的电路,这基本上就是当前深度学习系统的做法,那它根本无法扩展。

问答:神经科学与 AI

听众:谢谢您的演讲。刚才介绍说您在神经科学的院系任职……

罗素:我在神经外科有一个教职,其实跟神经元没什么关系,研究的是如何让脑损伤患者不至于死亡。

听众:我想问的是,我们还需要对自己的大脑如何运作、自己如何学习了解多少,这跟我们编写和创造一个同样能学习的东西之间是什么关系?

罗素:好问题。我认为这个关系随时间变化。AI 初期,人们对心理学、神经科学、AI 以及一定程度上语言学之间的交叉授粉非常兴奋,这就是我们所说的认知科学。这门新的综合学科曾激起巨大热情,但它多少失败了,退回到各个组成学科。因为神经科学家总体上专注于一个事实:仅仅弄清大脑由什么构成、在单个回路层面如何运作,就已经难得不可思议。所以「研究大脑就能知道如何做 AI」的承诺大体没有兑现。从七十年代到十年前,人们不太关注这条线。但后来发现,卷积神经网络,部分受到我们对视觉皮层认识的启发,效果远远好于预期。于是这条关闭已久的管道又重新打开了。

随着工具进步,情况可能会变。功能磁共振不是特别好的工具,你并不能从中看出大脑是怎么做事的。现在我知道「卷心菜」在我大脑的哪个位置,但我仍不知道我拿它做什么。而光遗传学之类的技术,可以修改神经元的 DNA 并直接观察其运作,我认为我们可能很快就会在理解大脑(至少是简单动物的大脑)如何工作方面取得真正的进展。

另一件有意思的事是脑机接口。我们把机械臂接到大脑上。我做一个非常高层的概括:起初我们以为必须理解运动皮层的编码,把它解码,再告诉机械臂该做什么,也就是得先解决神经科学的一大块问题才能让它工作。结果发现不用。是大脑自己想出了如何使用机械臂,而不是反过来。所以我们可以在不真正理解原理的情况下取得进展。比如,我们可以增强人类记忆,同样不必理解它如何工作:在大脑里找个地方接上一个像存储芯片一样的东西,大脑自己就会想出如何用它来存取信息。我们仍然不知道发生了什么,但我们的智商变成了 290。或者我们把大脑彼此相连,获得某种我们不理解的心灵感应,变成一个蜂群心智。我们仍然不知道自己的心智如何运作,可我们造出了蜂群心智。是不是很酷?谁知道接下来会怎样。

听众:但您认为这会转化为我们创造独立 AI 系统的能力吗?毕竟那些都是在更直接地与我们的大脑互动。

罗素:可能会。在这个过程中,我们会发展出更好的技术来理解和绘制神经活动。比如伯克利这里的「神经尘埃」项目,可以给你精细得多的实时神经活动图,也许我们就能理解它。但它非常难理解,我认为这一点已经很清楚了。

问答:偏好的可塑性

保罗斯:好,最后一个问题。

听众:很有意思。我想问问演讲中数学的部分。您说可以在数学上证明,机器人会为了解更多人类偏好而顺从、而采取行动,或者在多人问题中会去看谁对未来的预测更准。可是,机器人的行动本身可能改变人类的偏好或行为;或者在多人的例子里,最初选择跟随谁的偏好,会改变未来展开的方式,如果机器人当初选了别的,就会是完全不同的人被证明预测更准。这在数学上容易解决吗?

罗素:你提到了几个问题,但核心是可塑性。这一点在幻灯片上有,我一带而过了。是的,你要避免一种失败模式:机器通过把人类偏好改造得极易满足,来「满足」人类偏好。如果机器把我们都变成海洛因成瘾者,那我们唯一想要的就是海洛因,而机器很容易造出大量海洛因,问题就「解决」了。所以需要对模型做扩展。这个模型假设人类偏好是固定的,显然不是。但你也不能让人类偏好完全不受触动。光是有一个机器人仆人就会改变你的偏好,你大概会变得更娇惯一点、对其他人更没耐心一点,诸如此类。所以这是一个开放问题。

哲学家在偏好变化上很头疼,因为至少按简单的看法,改变自己的偏好从来不是理性的:那样一来,你未来的行为将违背你现在的偏好。我为什么要把自己变成一个会去做我认为不该做的事的人?说不通。可我们的偏好确实会变,有时我们还决定希望它变。我们说,我想如果我环游世界一圈,回来会成为一个更好的人。这是一场实验,你只能希望实验的结果是你事后会庆幸自己成为的那个人。这件事连要说清楚都很难。下学期我会和两位哲学家、一位经济学家合开一门课,这正是我们打算深入的话题之一。好问题。

保罗斯:好,谢谢斯图尔特,感谢你抽出时间。我觉得你最有价值的贡献之一,不只是思想上的深度,还在于让这些问题真正在领域之外的人那里引起共鸣,把他们带进这场对话,这意义重大。谢谢。

罗素:谢谢。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 主持人开场与讲者介绍 ▶ 正在看
2:17 从亚里士多德到图灵的 AI 思想史 ▶ 正在看
5:51 AlphaGo 与各国投资热潮 ▶ 正在看
7:56 对抗策略实验:AI 能力被高估 ▶ 正在看
11:02 数据不是石油,自动驾驶的可靠性关 ▶ 正在看
13:00 未来十年:仓库、助理、语言、卫星 ▶ 正在看
18:16 算力堆叠为何造不出通用 AI ▶ 正在看
22:59 核能史的教训:别赌人类无能 ▶ 正在看
26:43 AI 的上行空间:GDP 十倍与资源逻辑 ▶ 正在看
30:44 自主武器与人类衰弱的风险 ▶ 正在看
34:58 推荐算法改造人:目标设错的代价 ▶ 正在看
38:06 标准模型的错误与三条新原则 ▶ 正在看
42:06 辅助博弈:回形针实验与关机问题 ▶ 正在看
54:40 多人偏好聚合与预测准确者胜出 ▶ 正在看
62:41 问答:电路与大脑、神经科学、偏好可塑性 ▶ 正在看
本期小问 · 档案清单
—— 只靠把神经网络电路做得更大,真能造出人类水平的智能吗? ▶ 正在看
—— 机器对人类想要什么越不确定,我们反而越能控制它吗? ▶ 正在看
—— 给机器一个固定目标去优化,是不是 AI 从一开始就走错的设计? ▶ 正在看
—— 代表多个人做决定时,该按谁的预测更准来分配偏好权重吗? ▶ 正在看
本期讲者
斯图尔特·罗素加州大学伯克利分校计算机科学教授,人类兼容 AI 中心创始人,与 Peter Norvig 合著全球最通用的 AI 教科书《Artificial Intelligence: A Modern Approach》。2019 年出版《Human Compatible》,提出「可证明有益的 AI」,并长期推动禁止致命性自主武器。
Eric Paulos加州大学伯克利分校 EECS 教授,研究人机交互、城市计算与公民科学,本场 EECS Colloquium 的联合主持人。
01主持人开场与讲者介绍
0:00
ERIC PAULOS: OK, I'm going to go ahead. Welcome to the EECS Colloquium on behalf of myself, Professor Paulos, and also my cochair right in front here, Eli Yablonovitch. Our speaker today doesn't need much of an introduction, but I do want to give him one. Let me first, before I do that, tell you that next week, Wallace Marshall will be speaking here. Isn't it great to have the power back on, by the way? [LAUGHTER] Yes. So our speaker from last week, we are trying to reschedule. So for those of you who were anxiously at home in the darkness or not, we will be rescheduling.
ERIC PAULOS:好,那我就开始了。我谨代表我本人——保罗斯教授——以及坐在前排的联合主席伊莱·亚布罗诺维奇,欢迎大家来到 EECS 系列讲座。今天这位演讲者其实不太需要介绍,但我还是想为他做个介绍。在此之前,我先告诉大家,下周将由华莱士·马歇尔在这里演讲。顺便说一句,电来了是不是感觉特别好?[笑声] 是啊。上周那位演讲者,我们正在设法重新安排时间。所以对于那些当时在家里、摸黑焦急等待的人——或者也没那么焦急的人——我们会重新安排的。
便签笔记
0:35
Optimization, control, objective functions, as computer scientists, we think about all of these. And our speaker today, our very own Professor Stuart Russell, really talks about a fundamental reorientation of this field and some of our ways that we think about this. He's, of course, a professor here. He got his bachelor's degree in physics from Oxford, his PhD in computer science from Stanford. Here, he's the Smith-- yes, please. Smith-Zadeh Professor of Engineering at Berkeley. He's also an adjunct professor of neurological surgery at UCSF.
优化、控制、目标函数,作为计算机科学家,我们都会思考这些。而今天的演讲者,我们自己的斯图尔特·罗素教授,真正谈论的是这个领域的一次根本性的重新定位,以及我们思考这些问题的一些方式。他当然也是我们这里的教授。他在牛津获得物理学学士学位,在斯坦福获得计算机科学博士学位。在这里,他是史密斯——好的,请讲。伯克利的 Smith-Zadeh 工程学讲席教授。他还是加州大学旧金山分校神经外科的兼职教授。
便签笔记
1:10
He also is the vice chair of the World Economic Forum's Council on AI and Robotics, which is extraordinarily impressive. He's an honorary fellow at Wadham College at Oxford. He has the Carnegie fellowship. A long list of distinguished accolades during his career. I would best describe him as an extremely inspirational computer scientist, humanist, philosopher, provocateur-- and I mean that in the highest sense of the word-- an activist. And I think he really talks a lot about some of the technologies of our time.
他同时担任世界经济论坛人工智能与机器人理事会的副主席,这非常了不起。他是牛津大学沃德姆学院的荣誉院士。他获得了卡内基学者奖。他的职业生涯中有一长串杰出的荣誉。我会把他形容为一位极具启发性的计算机科学家、人文主义者、哲学家、挑衅者——我说的是这个词最崇高的意义——以及一位行动者。我觉得他确实对我们这个时代的一些技术谈了很多。
便签笔记
1:42
He has fundamentally made significant contributions to artificial intelligence. And now his position speaking out about particular issues of threat of autonomous weapons and the long-term future of artificial intelligence, primarily in many of his writings, many of whom you see in publications. But he has a new book out, Human Compatible-- Artificial Intelligence and the Problem of Control. I was going to bring in the book, but it's digital, so. [LAUGHTER] Sorry, I bought it on a Kindle. OK, so without further ado, welcome, Professor Stuart Russell.
他对人工智能做出了根本性的重大贡献。而如今,他公开发声,谈论自主武器的威胁和人工智能的长远未来等具体议题,主要体现在他的诸多著述中,其中很多你们都在出版物上见过。不过他有一本新书出版了,《AI 新生:破解人机共存密码》(Human Compatible—Artificial Intelligence and the Problem of Control)。我本来想把书带来的,但它是电子版的,所以……[笑声] 抱歉,我是在 Kindle 上买的。好,那么闲话少说,欢迎斯图尔特·罗素教授。
便签笔记
02从亚里士多德到图灵的 AI 思想史
2:17
[APPLAUSE] STUART RUSSELL: Thank you, Eric. OK, so let me provide a little bit of history. Just I think it's useful to have this background. There's a lot of stuff in the news these days about the potential upsides and downsides of AI. But actually, it's been going on longer. It's not just the last couple of years. In fact, in around 340 BC, Aristotle was writing about the possible consequences of intelligent automation on employment. So this was the first writing that we know of about technological unemployment.
[掌声]斯图尔特·罗素:谢谢你,Eric。好的,那我先讲一点历史背景。我觉得有这个背景会比较有帮助。这些天新闻里有很多关于人工智能潜在利弊的讨论。但其实,这件事由来已久。并不只是最近这几年才有的。事实上,大约在公元前340年,亚里士多德就写到过智能自动化对就业可能带来的后果。所以这是我们已知的最早关于技术性失业的著述。
便签笔记
3:02
And he said, look, if the musical instruments can play themselves and the loom can weave its own cloth, then we won't need any slaves. And this would be a big problem for unemployment figures. So that was a long time ago. And I say in the book, actually, Aristotle, if he'd had a computer and, I suppose, some electricity, then he would have been an AI researcher,. Because when you read what he writes-- he talks about planning algorithms, obviously, logical reasoning with formal syntax and semantics of logical languages, forward chaining, backward chaining.
他说,你看,如果乐器能自己演奏,织布机能自己织布,那我们就不需要奴隶了。而这对失业数字来说将是个大问题。所以这是很久以前的事了。我在书里也说过,其实亚里士多德,如果他当时有台电脑,再加上一点电力,那他大概就会成为一名人工智能研究者。因为当你读他写的东西时——他显然谈到了规划算法,逻辑推理,涉及逻辑语言的形式语法和语义,正向链接、反向链接。
便签笔记
3:43
All kinds of stuff is very clearly laid out in the work he did-- taxonomic hierarchies, semantic networks, ontologies. So anyway, he didn't have a computer. Babbage had at least a plan for a computer. And they talked, Babbage and Lovelace, talked very explicitly about how they'd be able to use that to do anything that the mind of man can be applied to could be done by these machines. So they were pretty clear that we were going to be able to have human-level AI, if only we can get all of these cogs and gears to mesh and move quickly enough.
他的工作把各种各样的东西都梳理得非常清楚——分类层级结构、语义网络、本体论。总之呢,他当时并没有计算机。巴贝奇至少还有一个计算机的设计方案。而且他们,巴贝奇和洛芙莱斯,非常明确地讨论过怎么用它来做任何事情——凡是人类心智能够涉及的事情,都可以由这些机器来完成。所以他们相当清楚,我们将能够拥有人类水平的人工智能,只要我们能让所有这些齿轮和传动装置彼此啮合、运转得足够快。
便签笔记
4:23
But he never quite built his machine. But news of the invention and the predictions got out there into the wide world. And actually, there was a religious newspaper in Illinois called the Primitive Expounder. And the editor of that newspaper got wind of what Babbage was talking about and predicted that if such machines were built, they would eventually take over the world. So that's the first one we know of in print giving the apocalyptic view of AI. And then we had to wait another 100 years for there to be actual computers, thanks to World War II, among other things.
但他始终没能真正把他的机器造出来。不过关于这项发明和那些预言的消息还是传到了外面的世界。实际上,伊利诺伊州有一份宗教报纸,叫《原始阐释者》(Primitive Expounder)。那份报纸的编辑听说了巴贝奇在讲的东西,于是预言说,如果这样的机器真被造出来,它们最终会统治世界。这是我们所知的第一份印刷出版物,提出了对 AI 的末日式看法。然后我们又等了一百年,才真正有了计算机,这要归功于第二次世界大战,以及其他一些因素。
便签笔记
5:04
And Alan Turing is famous, among many things, for this paper. This is the paper that has the Turing test from 1950. What's less well known is that he also talked about what would happen if we actually achieved human-level AI, and he was completely resigned. He said we should have to expect the machines to take control. And so when you hear complaints from people-- so now Elon Musk gets up and says something like this. He shouted down as an idiot who doesn't know anything about AI and so on. Are they going to shout down Alan Turing as an idiot who doesn't know anything about AI?
而阿兰·图灵之所以有名,原因有很多,其中之一就是这篇论文。这就是那篇 1950 年提出图灵测试的论文。不太为人所知的是,他还谈到了如果我们真的实现了人类水平的 AI 会发生什么,而他完全是一种听天由命的态度。他说,我们应当预料到机器会取得控制权。所以当你听到有人抱怨——比如现在埃隆·马斯克站出来说这样的话。他就被人喊下去,说他是个白痴,对 AI 一无所知等等。那他们会不会也把阿兰·图灵喊下去,说他是个对 AI 一无所知的白痴呢?
便签笔记
5:45
And who would be left?
那还能剩下谁呢?
便签笔记
03AlphaGo 与各国投资热潮
5:51
So Turing pointed this out. The field itself got going in '56 with its official birthplace in Dartmouth. And we've had waves of optimism, waves of disappointment. Right now we're in a wave of optimism. Hard to say what's happening next. I still think it's a little balanced on the knife edge as to whether the optimism will dissipate or enough new things will happen with enough success that it will keep going. So one of the big things that's happened, these are the media events. These are not really the research advances.
所以图灵早就指出了这一点。这个领域本身是从 1956 年开始的,官方的发源地是达特茅斯。我们经历过一波波的乐观,也经历过一波波的失望。现在我们正处在一波乐观之中。很难说接下来会发生什么。我仍然觉得这有点像在刀尖上保持平衡——乐观情绪会消散,还是会有足够多的新突破、足够大的成功让它继续走下去。所以,发生的一件大事——这些都是媒体事件。这些并不是真正的研究进展。
便签笔记
6:28
These are the media events. So from the outside world, they don't have a clue what happens here. People in the outside world think of this as a research breakthrough, despite the fact that it was based on Samuel 1957 and LeCun 1992. So the research breakthroughs happened 70 years ago and 20 years ago. But they put it together and made a big thing happen, which was beating the world's best go players. And this was the Sputnik moment for China. This was the event that caused China to wake up and say, this AI stuff is for real and we are going to dominate it.
这些都是媒体事件。所以在外部世界看来,他们根本不清楚这里发生了什么。外面的人认为这是一次研究上的突破,尽管它其实是建立在 Samuel 1957 和 LeCun 1992 的工作之上。所以真正的研究突破发生在 70 年前和 20 年前。但他们把这些东西整合起来,做成了一件大事,就是击败了世界上最顶尖的围棋选手。而这就是中国的“斯普特尼克时刻”。正是这个事件让中国警醒过来,说:这个 AI 是来真的,我们要在这方面占据主导地位。
便签笔记
7:15
And so they have announced plans for 150 billion or so in investment over the next decade. And the US has struck back with a new plan for the National AI Research Institutes with a 124 million. [LAUGHTER] So as we know, the media in particular have trouble distinguishing between million and billion. They often get them confused. And now apparently, the American government has also got that important distinction confused. So anyway, I'm sure it'll help. So all of these plans in the UK with the bill.
于是他们宣布了未来十年投资约 1500 亿的计划。而美国的回应是推出一项新计划,成立国家 AI 研究院,投入 1.24 亿。[笑声] 我们都知道,媒体尤其分不清「百万」和「十亿」的区别。他们经常把这两个搞混。而现在看来,美国政府显然也把这个重要的区别搞混了。总之,我相信这会有帮助的。英国那边所有这些计划,还有那笔预算。
便签笔记
04对抗策略实验:AI 能力被高估
7:56
This is billion pounds, 1.5 billion euros from the French, 18 billion pounds from the EU, and of course, China. So it's impossible to open a newspaper these days without seeing headlines about AI. It's a ridiculous time to be in this field. So I wanted to inject a little bit of realism. This was a very-- this is not quite as big as Lee Sedol losing, but this was a system developed by OpenAI showing that you can actually train-- you can train humanoid robots starting from nothing. So these are newborn babies who have no motor control skills whatsoever.
这是十亿英镑,法国出了15亿欧元,欧盟出了180亿英镑,当然还有中国。如今你翻开报纸,几乎不可能看不到关于人工智能的头条新闻。身处这个领域的当下,简直是个荒诞的时代。所以我想稍微注入一点现实主义。这曾经是个非常——它的影响力不太比得上李世石落败那次,但这是 OpenAI 开发的一个系统,它表明你真的可以训练——你可以从零开始训练人形机器人。所以这些就像是刚出生的婴儿,完全没有任何运动控制能力。
便签笔记
8:42
And in a few hours, you can train them to locomote, to kick the ball at the goal, to try and stop the ball from scoring, and so on. So this was a big thing. Lots of people got very excited about this. And it's sort of cool to look at. It's purposive behavior. He adjusts to where the ball is. The goalkeeper is tracking fairly well. The goalkeeper's a little bit uncoordinated. Anyway. So we thought, well, OK, how real is this? And so my student Adam Gleave said, OK, let's just change the red program, but we'll leave the blue program exactly the same.
而在几个小时之内,你就能训练它们行走、把球射进球门、去尝试拦下射门,等等。所以这在当时是件大事。很多人对此非常兴奋。而且看起来还挺酷的。这是有目的性的行为。他会根据球的位置进行调整。守门员的跟踪能力相当不错。守门员的动作有点不太协调。总之。于是我们就想,好吧,这到底有多真实?所以我的学生 Adam Gleave 说,好,我们就改一改红方的程序,但蓝方的程序原封不动。
便签笔记
9:21
And we're going to change the red program and see if we can make this game a little more even. Because right now the blue guy usually wins. So here's the solution, which is to simply fall over on the ground and waggle your leg in the air like this. Now you notice, we didn't change the blue program at all. But now look what the blue guy is doing. [LAUGHTER]
我们要改动红方的程序,看看能不能让这场比赛更势均力敌一些。因为现在蓝方通常都会赢。那么解决方案就是这个:直接倒在地上,然后像这样在空中晃你的腿。现在你注意,我们完全没有改动蓝方的程序。但现在看看蓝方那个家伙在干什么。[笑声]
便签笔记
9:51
And I think this is one of my favorite because he's just like, ah, ah, ah, ah. Oh, oh! Oh, no! OK. [LAUGHTER] So he lost the plot completely, even though it's the same program. So what this tells you is that your perceptions of the behavior, the capabilities of AI systems are often overly generous. You see one instance of good behavior. And you say, oh, it's really good at playing soccer or whatever, but it isn't. So there are these enormous gaps in performance that we don't have to look that hard to find them.
我觉得这段是我最喜欢的之一,因为它就在那儿,啊、啊、啊、啊。噢,噢!哦,不!好吧。[笑声] 它彻底不知道自己在干嘛了,尽管跑的是同一个程序。所以这告诉你,你对 AI 系统的行为和能力的认知,往往过于宽容了。你看到一次好的表现。然后你就说,哦,它踢足球真厉害之类的,但其实并不是。所以性能上存在这些巨大的落差,而且我们并不需要费多大劲就能找到它们。
便签笔记
10:33
So we have to be much more careful, I think. And imagine this happens with a self-driving car. You train it in mountain view. It does really well. And then you do some weird thing, like you change the color of the street sign. And all of a sudden, it just goes completely insane and starts driving around and doing donuts in the middle of the freeway, things like that. So you just-- until you understand what it's doing, you don't know whether it's actually as good as you think it is. And that's a really important lesson.
所以我觉得,我们必须谨慎得多。想象一下这种情况发生在自动驾驶汽车上。你在山景城训练它。它表现得非常好。然后你做点奇怪的事,比如把路牌的颜色换掉。然后突然之间,它就彻底疯了,开始到处乱开,在高速公路中间转圈漂移,诸如此类。所以你只能——在你搞清楚它在干什么之前,你根本不知道它是不是真的有你想象的那么好。这是一个非常重要的教训。
便签笔记
05数据不是石油,自动驾驶的可靠性关
11:02
So the economist had data is the new oil on the front cover. But I think we actually have to be-- we have to be more careful. And there's a particular technical reason why this simple meme that whoever has the most data wins the world. That's the meme that you see. Now that is the-- this is the basis on which geopolitical strategy is being formulated, but it's not really true. In fact, as AI systems get better, they should need less and less data, not more and more data. Humans need one or two examples to learn a new visual category.
《经济学人》曾在封面上写道:数据是新的石油。但我觉得我们其实必须——我们必须更谨慎一些。而且有一个特定的技术原因,能说明为什么这个简单的说法——谁拥有最多数据谁就赢得世界——这就是你会听到的那种说法。而现在,这正是——这正是地缘政治战略被制定时所依据的基础,但它其实并不成立。事实上,随着 AI 系统变得越来越好,它们需要的数据应该越来越少,而不是越来越多。人类学习一个新的视觉类别只需要一两个例子。
便签笔记
11:43
Where you see a giraffe for the first time, you don't need another 180,000 examples of giraffes before you got the idea of what a giraffe is. You see one. And you talk to any psychologist. How many training examples do humans need? One, sometimes two. And so as we get AI systems that are more capable, they will need less data. And just having vast quantities doesn't get you what you want. And if there is going to be something that burst the bubble, I think it will be potentially the failure of the self-driving car enterprise.
当你第一次看到长颈鹿,你不需要再看 18 万个长颈鹿的例子才能明白长颈鹿是什么。你看一次就够了。你可以去问任何一位心理学家。人类需要多少训练样本?一个,有时候两个。所以随着我们的 AI 系统能力越来越强,它们需要的数据反而会越来越少。光是拥有海量数据,并不能得到你想要的东西。如果说有什么东西会戳破这个泡沫,我觉得很可能就是自动驾驶汽车这个行业的失败。
便签笔记
12:19
Not that it's an impossible problem, but it has to happen fast enough for the people who have invested billions of dollars a year to not lose patience. And so there's a little race going on between the patience of the investors and the performance and safety of the system. And I don't know how that's going to turn out. So they made very impressive progress, but you've got to get eight 9s of reliability. And they're still only at five or six 9s of reliability. So it's a small difference. But you could say it's 100-fold or 1,000-fold reduction in the error rate.
倒不是说这是个不可能解决的问题,而是它必须足够快地实现,才能让那些每年投入几十亿美元的人不至于失去耐心。所以现在有一场小小的赛跑,一边是投资人的耐心,另一边是系统的性能和安全性。我也不知道最后会是什么结果。他们确实取得了非常了不起的进展,但你必须达到八个 9 的可靠性。而他们目前还只停留在五个或六个 9 的可靠性。看起来只是很小的差距。但换个说法,那意味着错误率要降低 100 倍甚至 1000 倍。
便签笔记
06未来十年:仓库、助理、语言、卫星
13:00
And so that's a big step to get to. So having said that, I'm reasonably optimistic that we will start to see many more. So not just AI systems that convince you to buy more toilet paper than you could possibly imagine you ever needed, but systems that actually are for robots getting out there, outside the factory, into the real world. So roads is one place, but warehouses will probably be first. And they're already doing simple tasks in warehouses, just fetching and carrying, but not so much based on perception.
所以这是很大的一步。话虽如此,我还是相当乐观,我们会开始看到更多这样的东西。不只是那种能说服你去买远超你想象所需数量的卫生纸的 AI 系统,而是真正让机器人走出工厂、走进真实世界的系统。道路是一个场景,但仓库很可能会是最先落地的地方。它们已经在仓库里做一些简单的任务了,就是取货和搬运,但基本不太依赖感知能力。
便签笔记
13:40
But when we solve-- and I think we're in the process of solving-- the problem of picking an arbitrary object out of a bin, which doesn't sound that hard-- you've got a big bin of stuff. And you say, OK, pick out the watermelon. Or pick out the water pistol or whatever it might be. If you have a system that can do that, then that's 10 million jobs going. And so that would be a very big step, a very visible step. The homes will probably come last. I mean, we'll see trivial toy applications and the little thing that makes drinks and serves them, but it's a very stupid robot.
但当我们解决了——我认为我们正在解决——从料箱里抓取任意物体的问题,这听起来好像没那么难——面前有一大箱东西。你说,好,把那个西瓜挑出来。或者把那把水枪挑出来,随便什么东西。如果你有一个系统能做到这一点,那就意味着 1000 万个工作岗位消失。所以这会是非常大的一步,而且是非常显眼的一步。家庭场景很可能是最后才实现的。我是说,我们会看到一些无聊的玩具级应用,比如那种会调饮料再端给你的小玩意儿,但那是个非常笨的机器人。
便签笔记
14:23
To get a system that really can function in the home, which is a very variable place, unpredictable, and physically complicated, that's going to be more difficult than these other problems. On the digital side, I think we'll see intelligent personal assistants that can actually understand enough about your life, your activities, your relationships, your commitments, your communications to really be useful, not just to be a parrot and a voice interface to a search engine like most of the systems are right now, but something that is actually a partner in your life and could be incredibly useful.
要做出一个真正能在家里正常工作的系统——家庭是一个变化极大、难以预测、而且在物理上很复杂的环境——那会比前面那些问题都更难。在数字领域,我认为我们会看到智能个人助理,它能真正足够了解你的生活、你的日常活动、你的人际关系、你的承诺安排、你的沟通往来,从而真正有用,而不只是像现在大多数系统那样,只是个鹦鹉学舌的、搜索引擎的语音接口,而是真正成为你生活中的伙伴,那会非常非常有用。
便签笔记
15:02
In the same way that executives have expensive human personal assistants, you're going to have something for $0.99 a month that is as useful, but that would be for everybody. And most people probably have more need of it because their lives are more difficult. Another consequence, if I had to say, what's the thing that's most likely to get solved in the next decade, it's the next level of language understanding. So if you just think, OK, what happened in the last decade or the decade we're just finishing, its visual object recognition.
就像高管们请得起昂贵的人类私人助理一样,你将能用每月 0.99 美元获得同样有用的东西,而且是人人都能拥有。而且大多数人可能更需要它,因为他们的生活更艰难。另一个结果,如果要我说,未来十年最有可能被解决的是什么,那就是下一个层次的语言理解。你想想看,过去这十年、也就是我们刚刚走完的这十年发生了什么,是视觉物体识别。
便签笔记
15:40
That was the biggest step, I think, in the next decade. The ability to extract content from language-- think reading some text and generating database entries, logical assertions, or whatever it might be. So not deep understanding. They're not going to read James Joyce and write a thesis about it, but they'll be able to read all the documents on the web, every newspaper, every television broadcast in every language. And that will be an incredible facility for the human race. And if you think search engines are worth $1 trillion, easily, maybe 2 trillion now, this would be 10 times as valuable for the human race.
我认为那是最大的一步,而下一个十年,将是从语言中提取内容的能力——想象一下读一段文字,然后生成数据库条目、逻辑断言,或者任何类似的东西。所以不是深层次的理解。它们不会去读詹姆斯·乔伊斯然后写一篇论文,但它们能读遍网络上的所有文档、每一份报纸、每一种语言的每一档电视节目。这对人类来说将是一项不可思议的能力。如果你觉得搜索引擎值 1 万亿美元,那轻轻松松,现在可能值 2 万亿了,那这个东西对人类的价值会是它的 10 倍。
便签笔记
16:26
Another interesting thing, that I'm actually recently joined the board of Planet, which is a San Francisco company that has the largest number of satellites in the world. And so now that it's possible to image every square foot of the Earth every day. You could, if you had enough people, about 30 million people, you could look at all those images and you could keep track of all the things on the Earth and what they're up to. But now with computer vision applied to that data stream, we could actually turn the Earth into a continuously updated database.
另一件有意思的事,我最近加入了 Planet 的董事会,这是一家旧金山的公司,拥有全世界数量最多的卫星。所以现在我们有可能每天给地球上的每一平方英尺拍一次照。如果你有足够多的人,大概 3000 万人,你就可以看完所有这些图像,追踪地球上所有的事物以及它们的动向。但现在把计算机视觉用在这个数据流上,我们其实可以把地球变成一个持续更新的数据库。
便签笔记
17:07
On top of that, you can build thousands of different applications. And this is one of the things that we're working with the UN, to think of all the ways you could use that for the sustainable development goals. So you think about urban planning. You think about managing livestock on the plains of Africa, where it's not all fenced off and so on, but it's a very difficult management problem. Anti-poaching, all kinds of stuff that you could imagine, managing shipping, looking for smuggling. The list goes on and on and on and on.
在这个基础上,你可以构建成千上万种不同的应用。这也是我们正在和联合国合作的事情之一,去思考可以用什么方式把它用在可持续发展目标上。比如你想想城市规划。再想想在非洲草原上管理牲畜,那里并没有全部围起围栏,所以这是个非常棘手的管理问题。反盗猎,各种你能想到的事情,管理航运、查找走私。这个清单可以一直列下去,没完没了。
便签笔记
17:42
But each of those applications, if you wanted to build it by itself, would cost a billion dollars to build because collecting all this data, storing it, processing it is an expensive business. But if that's done once and then you feed the results to anyone who wants to build an application on top of that, then it's like a million dollars to produce a 24/7 global service that would have some high value. So this is, I think, something that's going to be happening fairly soon. So does this all mean that we're going to have human-level AI?
但这些应用中的每一个,如果你要单独去做,都会花掉十亿美元,因为采集所有这些数据、存储它、处理它,是一件很烧钱的事。但如果这件事只做一次,然后把结果提供给任何想在上面开发应用的人,那么大概花一百万美元就能做出一个全天候的全球性服务,而且价值很高。所以我认为,这是很快就会发生的事情。那么这一切是否意味着我们将拥有人类水平的 AI?
便签笔记
07算力堆叠为何造不出通用 AI
18:16
Is it around the corner? So different people have different views. For example, Ilya Sutskever, who is one of the people who worked on the visual object recognition breakthroughs with Geoff Hinton, is now the chief scientist at OpenAI. He believes it's five years' time. And they just got a big investment from Microsoft, a lot of which is going to be scaling up their already gargantuan computer resources. So when I say gargantuan, I mean one Google TPU pod is equivalent of 10 million laptops.
它是不是近在眼前了?不同的人有不同的看法。比如说,Ilya Sutskever,他是当年和 Geoff Hinton 一起做出视觉物体识别突破的人之一,现在是OpenAI 的首席科学家。他认为还有五年时间。而且他们刚从微软那里拿到一大笔投资,其中很大一部分会用来进一步扩大他们本已庞大得惊人的算力资源。我说“庞大得惊人”,意思是一个 Google TPU pod 相当于 1000 万台笔记本电脑。
便签笔记
18:56
And it's fairly-- it's not huge. I mean, it's about-- so you imagine about 12 feet long and 8 feet high and so on. But it would fit in your bedroom. And that's already more powerful than the world's biggest supercomputer was two years ago. So the numbers are astronomical. It's also-- I mean, the number of operations per second, 10 to the 17, is in the same ballpark as the theoretical maximum number of state changes that the brain can achieve.
而且它其实——它并不算大。我是说,大概——你可以想象大约 12 英尺长、8 英尺高之类的。它能放进你的卧室里。而这已经比两年前世界上最大的超级计算机还要强。所以这些数字是天文级的。另外——我是说,每秒 10 的 17 次方次运算,这个量级和大脑理论上能达到的最大状态变化次数是同一个数量级。
便签笔记
19:36
And they're thinking of going 1,000 times beyond that. And they believe seriously that that will achieve human level. I don't believe it at all because those are circuits. I mean, if there's one thing you learn in computer science is circuits are circuits and programs are programs. And they're much more powerful and much more expressive. So if you think about, let's say, writing the rules of chess in circuit language, it's hundreds of thousands of pages because you've got to have a different piece of circuit for every square.
而他们还打算再往上做 1000 倍。他们真心相信那样就能达到人类水平。我完全不相信,因为那些只是电路。我是说,如果你在计算机科学里只学到一件事,那就是电路就是电路,程序就是程序。而程序要强大得多,表达力也强得多。比如你想想,用电路语言来写国际象棋的规则,那得写几十万页,因为你必须为每一个格子准备一段不同的电路。
便签笔记
20:10
Everything has to be repeated for every square, for every move, and so on, for each pawn and so on. It's ridiculous, the blow-up that comes from having these inexpressive languages. And of course, in the programming language, it's a page. In English, it's a page. In first-order logic, it's a page. And there's a reason why these are all about the same-- because this is the level of expressive power that you need to deal with a large, complicated world that has lots of things in it. Things. Things are really important.
每一样东西都得为每个格子、每一步棋重复一遍,每个兵也要重复,等等。这种表达力低下的语言带来的膨胀,简直荒唐。当然,用编程语言来写,就是一页纸。用英语写,也是一页纸。用一阶逻辑写,还是一页纸。这几种篇幅都差不多,是有原因的——因为这正是你需要的表达力层次,才能应对一个庞大、复杂、里面有很多“事物”的世界。事物。事物真的非常重要。
便签笔记
20:46
The world has lots of things in it. And that means you need something with the expressive power of first-order logic. Propositional logic, the language of circuits doesn't have any things in it. So you can't talk about all the pawns or all the squares or all the time steps or all the people or all the donkeys or giraffes or whatever. You can't say that in a circuit language. And so my belief is there's no possibility that just making bigger and bigger circuits is going to solve the problem of human-level AI.
这个世界里有非常多的事物。这意味着你需要具备一阶逻辑那种表达力的东西。命题逻辑,也就是电路的语言,里面根本没有“事物”这个概念。所以你没法谈论所有的兵、所有的格子、所有的时间步、所有的人,或者所有的驴子、长颈鹿之类的。这些在电路语言里根本说不出来。所以我相信,光靠把电路做得越来越大,是不可能解决人类水平 AI 这个问题的。
便签笔记
21:18
Instead, I think we need these real conceptual breakthroughs that are actually very hard to predict. So you can't plot some kind of Moore's law curve and say, oh, look, it crosses human intelligence in 2029 or anything like that. Instead, you have to wait for the conceptual breakthroughs. So I've listed some of them here. I'm not going to go through them all, but probably the biggest one is the third one because that's what enables humans to function successfully in the real world. The fact that we can operate seamlessly on scales ranging from individual motor control actions, like your brain has to send a whole complicated stream of commands to your tongue or your fingers when you're typing on a scale of a few milliseconds per command, all the way up to doing a PhD, which is five years, about a trillion motor control commands.
相反,我认为我们需要真正的概念性突破,而这些其实很难预测。所以你没法画一条摩尔定律式的曲线,然后说,看,它会在 2029 年超过人类智能,之类的。相反,你只能等待概念上的突破。我在这里列了其中一些。我不会全部讲一遍,但最重要的大概是第三条,因为正是它让人类能够在真实世界中成功地行动。也就是说,我们能够在从单个运动控制动作,一直到更大尺度上无缝地运作——比如你打字的时候,大脑要以每条指令几毫秒的节奏,向你的舌头或手指发送一整套复杂的指令流,一直到读完一个博士学位,那是五年时间,大约一万亿条运动控制指令。
便签笔记
22:19
And that trillion goes in the exponent. it's the length. If you remember-- if you've taken AI, b, the branching factor, to the power of d, the depth of the solution. So it's the d is a trillion. So there's no possibility that just by having more computing power, you can scale up to that. And humans manage this by being able to reason seamlessly at many levels of abstraction. And we have some inkling of how to do that, if someone provides the levels of abstraction, if someone gives you a hierarchy of actions at different scales, and then you can string them together in clever ways.
而这一万亿是出现在指数上的。它是长度。如果你还记得——如果你学过 AI 的话,b 是分支因子,b 的 d 次方,d 是解的深度。所以这里的 d 是一万亿。所以光靠拥有更多算力,是不可能扩展到那个规模的。人类之所以能做到,是因为我们能在许多抽象层次之间无缝地推理。而对于怎么做到这一点,我们已经有了一点头绪,前提是有人提供好这些抽象层次,如果有人给你一套不同尺度上的行动层级结构,然后你就可以用巧妙的方式把它们串联起来。
便签笔记
08核能史的教训:别赌人类无能
22:59
But where that hierarchy comes from, we have no real clues yet about how machines could develop their own hierarchies as they go along. That, to me, is the biggest open problem. And that could be solved. Maybe someone's already solved it while I'm talking. So these things can happen quite quickly. And just to illustrate that, you can look back at the last time we invented a civilization-ending technology, which is nuclear energy. And so the consensus in the early part of the century-- so they knew that the energy was there.
但这个层级结构从何而来,机器如何在运行过程中自己发展出层级结构,我们目前还完全没有头绪。在我看来,这是最大的未解难题。而这个问题是有可能被解决的。也许在我说话的这会儿,已经有人把它解决了。所以这些事情可能发生得相当快。为了说明这一点,你可以回顾一下我们上一次发明足以终结文明的技术,也就是核能。上世纪初的共识是——他们知道那份能量就在那里。
便签笔记
23:36
They had E equals mc squared. They could measure the masses of different isotopes. They could tell you exactly how much energy would be released if you could cause one of these transitions. But they were absolutely convinced that this was impossible. And Rutherford, who was the man who split the atom, Nobel Prize winner, probably the most famous nuclear physicist, he gave a speech in Leicester on September 11. And he was asked, is there any possibility we could do this even in the next 25 or 30 years?
他们有 E=mc²。他们能测量不同同位素的质量。他们能精确告诉你,如果你能引发其中一种转变,会释放出多少能量。但他们绝对确信这是不可能做到的。卢瑟福,这位分裂了原子的人、诺贝尔奖得主、大概是最著名的核物理学家,他在 9 月 11 日于莱斯特发表了一次演讲。有人问他,哪怕在未来 25 年或 30 年里,我们有没有可能做到这件事?
便签笔记
24:08
And he said, no, it's moonshine to even think we could do that. And Einstein agreed with him. Einstein said he couldn't think of any imaginable way that you could cause these transitions to occur. And then Leo Szilard read about this in The Times the next morning, and he went for a little walk. And while he was crossing the road, he invented the nuclear chain reaction. [LAUGHTER] So it went from, this is completely impossible in the opinion of all leading physicists, to essentially solved. And it was only 12 years after that the first nuclear bomb was exploded.
他说,不可能,连想都别想,那纯属痴人说梦。爱因斯坦也同意他的看法。爱因斯坦说,他想不出任何可以想象的办法能引发这些转变。然后第二天早上,利奥·西拉德在《泰晤士报》上读到了这件事,他出去散了会儿步。就在他过马路的时候,他想出了核链式反应。[笑声] 所以事情就从「在所有顶尖物理学家看来这完全不可能」,变成了基本上已经解决。而在那之后仅仅 12 年,第一颗原子弹就爆炸了。
便签笔记
24:49
So betting against human ingenuity. This is one of the arguments that some people who are skeptical about any possible risk from AI, this is what they say. Actually, capable AI is impossible, and so we don't need to worry. And I find this a bizarre line of argument. It's basically like saying, OK, well, yes, we are driving the human race towards a cliff in this big bus at as fast as we can go. The hundreds of billions of dollars of investment to create this technology. We are going as fast as we can towards the cliff.
所以,跟人类的创造力对赌可不明智。这正是一些对 AI 可能带来的任何风险持怀疑态度的人所用的论点之一,他们就是这么说的。其实,有能力的 AI 是不可能实现的,所以我们不用担心。我觉得这是一种很离奇的论证方式。这基本上就等于说,好吧,是的,我们正开着这辆大巴,用最快的速度把人类往悬崖边上开。投入数千亿美元来创造这项技术。我们正以最快的速度冲向悬崖。
便签笔记
25:26
Our pedal to the metal. But I guarantee you, we're going to run out of gas before we go over the edge. Would you get in that bus? No, it's just-- there's no argument as to why human-level AI is impossible. And how could there be? We know it's possible in the sense that we know our brains are capable of generating this level of intelligence. So you can't argue that it's physically impossible to do. And so this is a very disappointing development that even people in AI, in order to avoid talking about risks, are willing to say, no, AI will fail, which is weird.
油门踩到底。但我向你保证,我们在冲下悬崖之前就会把油烧光。你会上那辆巴士吗?不会,这只是——根本没有任何论据能说明人类水平的 AI 是不可能的。怎么可能会有呢?我们知道它是可能的,因为我们知道人类的大脑有能力产生这种水平的智能。所以你不能说做到这件事在物理上是不可能的。因此这是一个非常令人失望的发展趋势:连搞 AI 的人,为了回避谈风险,都愿意说,不会的,AI 会失败——这很奇怪。
便签笔记
26:09
Imagine a cancer biology. The leading cancer biologists of our era standing up and saying, you know what? We're never going to cure cancer. Keep giving us lots of funding. But I guarantee, we're never going to cure cancer. What are you talking about? So that's the situation we face. Anyway. So I think it's prudent to assume you can't guarantee. I mean, of course, we could destroy ourselves before this happens, but I think it's prudent to work on the assumption that we will develop AI systems that are fundamentally more capable of making decisions than humans.
想象一下癌症生物学。我们这个时代最顶尖的癌症生物学家站出来说,你知道吗?我们永远治不好癌症。请继续给我们大笔经费。但我保证,我们永远治不好癌症。你在说什么呢?这就是我们面临的处境。总之。所以我认为,谨慎的做法是假定你无法保证。我是说,当然,我们可能在这之前就毁灭了自己,但我认为谨慎的做法是基于这样一个假设来开展工作:我们将会开发出在决策能力上从根本上超越人类的 AI 系统。
便签笔记
09AI 的上行空间:GDP 十倍与资源逻辑
26:43
They will clearly have access to much more information. They will be able to look further ahead in the future than humans can, just as they already do on the chess board and the go board and the video games and so on. I think this will eventually translate to real-world decision-making. So this could be a good thing. If it wasn't for the fact that there was some upside, we wouldn't be having this conversation. Because if there were no upside, we wouldn't be spending all this money. And nobody would be doing AI at all.
它们显然能获取多得多的信息。它们将能比人类看得更远,正如它们在国际象棋棋盘、围棋棋盘和电子游戏中已经做到的那样,等等。我认为这最终会转化到现实世界的决策中。所以这可能是件好事。如果不是因为它确实有好处的一面,我们根本不会有这场对话。因为如果没有好处,我们也不会花这么多钱。而且根本不会有人去做 AI。
便签笔记
27:13
So of course, there is a big upside. And one way of saying what that is that our civilization is really the result of intelligence. If you suddenly have access to a lot more, you can have a much better civilization. And so instead of thinking about small improvements like better medical diagnosis, safer cars-- I mean, yeah, that's nice, but that's not very ambitious. You think about travel. If you wanted to go to Australia 200 years ago, that would be a multiyear, multibillion dollar project with about 80% chance of death.
所以当然,这里面有巨大的好处。描述这一点的一种说法是,我们的文明其实就是智能的产物。如果你突然能获得多得多的智能,你就能拥有一个好得多的文明。所以,与其去想那些小的改进,比如更好的医疗诊断、更安全的汽车——我是说,没错,那挺好的,但那不够有雄心。想想出行这件事。如果 200 年前你想去澳大利亚,那会是一个耗时数年、花费数十亿美元的项目,而且大约有 80% 的死亡概率。
便签笔记
28:01
And now, if you want to go to Australia, you take out your cell phone. Tap. Tap. Tap. Tap. Tap. And you're in Australia tomorrow. And relatively speaking, it's free compared to what it used to be. And imagine that same transformation occurring to everything, everything that's currently expensive, like construction projects. We need a new hospital. we need a road to connect our village to the metropolitan areas. We need this. We need that. We need better. We don't have any teachers in our school. Whatever it is that is currently difficult, expensive, takes a long time, AI systems could provide for us.
而现在,如果你想去澳大利亚,你掏出手机。点一下。点一下。点一下。点一下。点一下。明天你就到澳大利亚了。而且相对而言,跟过去比起来,这几乎是免费的。再想象一下同样的转变发生在所有事情上,所有目前很昂贵的事情,比如建设项目。我们需要一座新医院。我们需要一条路把我们村子和大城市连起来。我们需要这个。我们需要那个。我们需要更好的东西。我们学校里一个老师都没有。不管现在什么东西是困难的、昂贵的、耗时很久的,AI 系统都能为我们提供。
便签笔记
28:42
So I'm not talking about AI systems inventing cures for cancer or faster than light travel or any of this science fiction things. Just making available to people on a large scale what we already know how to do would be about a 10-fold increase in GDP. So just bringing everyone up to a Berkeley standard of living is a 10-fold increase in GDP of the world, which is a $13,500 trillion net present value. So that's the cash equivalent of what this is worth. And that's why countries like China are investing hundreds of billions of dollars in this.
所以我讲的不是 AI 系统发明癌症疗法,或者超光速旅行,或者任何这类科幻的东西。仅仅是把我们已经知道怎么做的事情大规模地提供给所有人,这大概就能让 GDP 增长 10 倍。所以仅仅是把每个人的生活水平提升到伯克利的水平,就能让全世界的 GDP 增长 10 倍,这相当于 13,500 万亿美元的净现值。所以这就是这件事所值的现金等价物。这也是为什么像中国这样的国家在这上面投入数千亿美元。
便签笔记
29:25
And when you look at the size of the prize, that number is negligible. Anything measured in billions is negligible compared to the value that can be generated. And so I'm pretty sure that as things progress, we're not going to be limited by research funding or money that industry can contribute or spend on research. We're going to be limited by our ability to have enough people who can actually put this transformation into practice.
而当你看看这块蛋糕有多大,那个数字就微不足道了。跟能够创造出来的价值相比,任何以十亿计的数字都微不足道。所以我很确定,随着事情往前推进,限制我们的将不会是研究经费,或者产业界能贡献、能投入研究的钱。限制我们的将是:我们能不能有足够多的人,真正把这场转变落到实处。
便签笔记
30:02
And when you have a world like that, then material wealth becomes like digital copies of the newspaper. Fighting over how many digital copies of the newspaper you have is just insane. And trying to grab a larger fraction of the digital copies of the newspaper is insane. So I think it would change that part of the dynamic of history that has to do with competition over resources, access to whatever makes life physically worth living. That can all change. We can still kill each other for religious reasons and so on.
而当你有了那样一个世界,物质财富就变得像报纸的数字副本一样。为你有多少份报纸的数字副本而争斗,简直是疯了。试图抢占更大份额的报纸数字副本,也是疯了。所以我认为,这会改变历史上那部分与资源争夺有关的动态,也就是争夺那些让生活在物质上值得过下去的东西。这些都可以改变。我们仍然可能因为宗教等原因互相残杀。
便签笔记
10自主武器与人类衰弱的风险
30:44
The one thing that doesn't change, actually, is land. We may kill each other over access to land because at least for the time being, AI systems can't make more of that. OK, so there's lots of downsides that people have talked about-- messing with our understanding of reality, just straightforward killing each other. And I was going to show this little video. So this is called slaughterbots. You can find it on YouTube. And it makes the point, in simple terms for policymakers, that if you create weapons that can locate and select and kill human targets without human supervision-- you're all computer scientists.
不过有一样东西其实不会变,那就是土地。我们可能还是会为了争夺土地而互相残杀,因为至少就目前而言,AI 系统造不出更多的土地。好,那么人们也谈到了很多坏处——扰乱我们对现实的理解,还有就是直接互相残杀。我本来打算放一段小视频。这个视频叫《屠戮机器人》(slaughterbots)。你可以在 YouTube 上找到它。它用简单的方式向政策制定者说明了一点:如果你造出能够定位、选择并杀死人类目标的武器,而且不需要人类监督——你们都是计算机科学家。
便签笔记
31:36
You understand what scalability means. You understand that Google does not have a billion employees answering all those questions that we send. How do they do it, those billions of questions? They buy more hardware. And so just as Google can answer billions of queries a day, you can kill millions of people just by buying more of these. You don't need to have a vast army with a huge military industrial complex. You can just have a weapon of mass destruction by buying more of them and then sending them off to do their business.
你们明白可扩展性意味着什么。你们明白,谷歌并没有十亿名员工来回答我们发过去的所有那些问题。那他们是怎么做到的,那几十亿个问题?他们买更多的硬件。所以,就像谷歌能一天回答几十亿次查询一样,你只要多买一些这种东西,就能杀死几百万人。你不需要一支庞大的军队和巨大的军工复合体。你只要多买一些,然后把它们放出去执行任务,就等于拥有了一件大规模杀伤性武器。
便签笔记
32:09
Unfortunately, the sound's not working, so I'm not going to show you the whole movie. But I did want to point out that when we brought that movie out in 2017, some, for example, the Russian ambassador to the UN, complained that this was just science fiction and this would not be even a thing for another 30 years. And he was insulted that we were even bringing this to the United Nations. But now you can actually go and buy something that's actually a lot worse than the weapon we showed. So the Turkish company STM is producing this, the KARGU drone.
很遗憾,声音出不来,所以我就不给你们放整部片子了。但我确实想指出,当我们在 2017 年推出那部片子时,有些人,比如俄罗斯驻联合国大使,抱怨说这只是科幻,说这种东西再过 30 年都不会出现。他甚至因为我们把这件事提交到联合国而感到被冒犯。但现在你其实可以去买到比我们展示的那件武器糟糕得多的东西。土耳其公司 STM 正在生产这个,KARGU 无人机。
便签笔记
32:46
And they advertise its capabilities. Autonomous hit, targets selected based on images, tracking moving targets, antipersonnel, and face recognition. So all these are advertised in the sales materials for this. And they have announced their intention to use it against the Kurds. Their announcement was for early 2020, but maybe they got to move up that thanks to President Trump. So we might see attacks for the first time of this type of autonomous weapon on human beings coming from Syria. End of employment.
他们还在宣传它的各种能力。自主打击、基于图像选定目标、追踪移动目标、反人员作战,以及人脸识别。这些全都写在它的销售宣传材料里。而且他们已经宣布,打算用它来对付库尔德人。他们宣布的时间是2020年初,但多亏了特朗普总统,也许他们会提前。所以我们可能会第一次看到这类自主武器从叙利亚出发、对人类发动攻击。就业的终结。
便签笔记
33:26
Evil people misusing AI to take over or destroy the world. Overuse of AI, resulting in a gradual enfeeblement of the human race. I think this is a very difficult problem to deal with because we are naturally lazy. Even if we know in the long run that we will become enfeebled, it's always easier to say, OK, well, just for today. Yeah, sure. You can tie my shoelaces. I'll learn how to do it tomorrow. But for today, you can tie my shoelaces. So it's a problem that-- again, like other problems, for thousands of years, we've known that this is a tendency that we have to basically hand over knowledge of how to run things to someone else.
邪恶的人滥用人工智能来接管或摧毁世界。过度使用人工智能,导致人类逐渐衰弱无能。我认为这是个非常棘手的问题,因为我们天生就懒。即使我们知道长远来看自己会变得衰弱无能,但说“好吧,就今天这一次”总是更容易。是啊,当然。你可以帮我系鞋带。我明天再学怎么系。但今天,你先帮我系鞋带吧。所以这是个问题——跟其他问题一样,几千年来,我们一直知道自己有这种倾向,就是把怎么管理事务的知识基本上交给别人。
便签笔记
34:17
And here we would just be able to hand it over to machines, which we've never been able to do that before. To have a civilization that continues, you've always had to hand it over to the next generation of humans. And so we've done about a trillion person years of teaching and learning just to keep our civilization going. And now we don't have to do that anymore because we can pass it on to the machines, and they can run the civilization. And that may be an irreversible step. So in the last 10 minutes or so, I'll talk about possibly the most serious issue.
而现在我们可以把它直接交给机器,这是我们以前从来做不到的。要让一个文明延续下去,你过去只能把它交给下一代人类。所以为了让我们的文明维持运转,我们已经投入了大约一万亿人年的教学和学习。而现在我们不必再这么做了,因为我们可以把它传给机器,让机器来运转这个文明。而这可能是一个不可逆的步骤。所以在最后大概十分钟里,我要讲一讲可能是最严重的一个问题。
便签笔记
11推荐算法改造人:目标设错的代价
34:58
So why are people writing this kind of headline? What is Elon Musk talking about when he says, summoning the demon? What he means is that you're making something more intelligent than yourself. How exactly do you propose to control it? How exactly do you propose to maintain power forever over something that is intrinsically more powerful than you are? That's the question that we face. So just to-- as a little warning, we-- our friendly social media companies have provided a little warning to us. Very public spirited of them to do this.
那么,人们为什么会写出这类标题呢?埃隆·马斯克说“召唤恶魔”的时候,他到底在说什么?他的意思是,你正在造出一个比你自己更聪明的东西。那你打算怎么控制它?你打算怎么永远维持住对一个本质上比你强大的东西的支配权?这就是我们面临的问题。那么——先给大家提个醒,我们那些“友好的”社交媒体公司已经给我们提了个醒。他们这么做还真是很有公益精神。
便签笔记
35:45
So they deployed on a global scale a fairly simple machine learning algorithms that are designed to optimize a particular objective in this case, clickthrough. And you might think, OK, well, how do we optimize clickthrough? The algorithm just learns what people want to click on and doesn't send them stuff they're not interested in. That sounds sort of OK, but even that can produce a kind of echo chamber filter bubble mentality where you're never even exposed to other thoughts or ideas beyond the ones you're already comfortable with.
他们在全球范围内部署了相当简单的机器学习算法,这些算法是为了优化某个特定目标而设计的,在这个例子里,就是点击率。你可能会想,好吧,那怎么优化点击率呢?算法无非就是学习人们想点什么,然后不给他们推送不感兴趣的东西。听上去还行,但即便如此,也会造成一种回音室、信息茧房式的心态——你永远接触不到别的想法或观念,只能看到你本来就认同的那些。
便签笔记
36:19
But actually, it's much, much, much, much worse than that. Because that's not what the algorithms did. What the algorithms did was to modify people, not to learn what people like, but actually to modify what people are like so that they are more predictable. Because that way, you can get more money out of them. And this is not like some evil psychological plan. The algorithms don't know that humans exist. The algorithms don't know that humans have minds. From the algorithm's point of view, a human is just a click history.
但实际上,情况比这要糟糕得多得多。因为算法做的根本不是这件事。算法做的是改造人,不是去学习人们喜欢什么,而是真的去改变人本身,让他们变得更可预测。因为那样一来,你就能从他们身上赚更多的钱。这可不是什么邪恶的心理学阴谋。算法根本不知道人类的存在。算法根本不知道人类是有心智的。从算法的角度看,一个人就只是一串点击记录。
便签笔记
36:50
That's all we are. What have you been exposed to and what did you click on? So we're just a click history. But you can generate more clicks by, it turns out, sending people more and more and more extreme material. They might start out in the middle of some spectrum, let's say, YouTube videos. You send them something a little bit more violent and a little bit more violent. They gradually move their center of mass to the point where they are addicts of extreme, violent YouTube videos. And this is a documented process, and it's happening in politics.
我们对它来说就只是这些。你看到过什么内容,你点了什么?所以我们只是一串点击记录。但结果发现,你可以通过给人们推送越来越极端的内容来制造更多点击。比如说 YouTube 视频吧,他们一开始可能处在某个光谱的中间位置。你给他推点稍微暴力一点的,再暴力一点的。他们的重心就逐渐移动,最后变成极端暴力 YouTube 视频的成瘾者。这是一个有据可查的过程,而且它正在政治领域中发生。
便签笔记
37:26
It's happening in other areas of media consumption. And the reason is that you set up an objective with a sufficiently powerful system. If that objective is the wrong one, then you have this massive collateral damage. And that's the issue that we're concerned about. It's not a new issue for the human race. We've known about this. We have the legend of King Midas and his objective. Everything I touch should turn to gold. Sounds good until you realize that includes your food and your drink, and then you die.
这种情况也在其他媒体消费领域上演。原因在于,你给一个足够强大的系统设定了一个目标。如果这个目标是错的,就会带来巨大的附带损害。这正是我们所担心的问题。对人类来说,这并不是一个新问题。我们早就知道这一点。我们有迈达斯国王和他那个愿望的传说。我碰到的一切都要变成黄金。听起来不错,直到你意识到这也包括你的食物和饮料,然后你就死了。
便签笔记
12标准模型的错误与三条新原则
38:06
The genie who gives you three wishes-- what's the third wish? The third wish is, please undo the first two wishes because I got them wrong. So many, many cultures have these kinds of stories which basically point to the fact that we are unable to state the objectives correctly. We're unable to anticipate all the ways that the objective might be optimized and all the collateral damage that might ensue. So in fact, what I'm arguing is that the fundamental way that we went about defining AI in the first place, starting in the '50s, was wrong.
那个给你三个愿望的精灵——第三个愿望是什么?第三个愿望是,请把前两个愿望撤销吧,因为我许错了。所以很多很多文化里都有这类故事,它们基本上都指向同一个事实:我们没能力把目标准确地表述出来。我们没能力预见目标可能被优化的所有方式,以及随之而来的所有附带损害。所以事实上,我要说的是,我们当初定义人工智能的那种根本方式,从上世纪50年代开始,就是错的。
便签笔记
38:48
So here's roughly speaking, what we said. Here's what we mean by human intelligence in this context. It means the ability to choose actions that can be expected to achieve our objectives. That was the economic notion of rationality, the philosophical notion of rational behavior. And we just copy that to machines. We said, OK, so machines are intelligent to the extent that their actions can be expected to achieve their objectives. That's great. It's not just AI. This is control theory. This is statistics, where you minimize a loss function; operations research, where you maximize the sum of rewards; even economics where you maximize profit or you maximize welfare, GDP.
大致来说,我们当时是这么讲的。在这个语境下,我们所说的人类智能是这个意思。它指的是选择那些有望实现我们目标的行动的能力。这就是经济学里的理性概念,也是哲学里理性行为的概念。然后我们把它照搬到了机器上。我们说,好,那么机器的智能程度,就取决于它的行动在多大程度上有望实现它的目标。很好。这不只是人工智能。控制论是这样。统计学也是这样,在那里你要最小化一个损失函数;运筹学里,你要最大化奖励之和;甚至经济学里,你要最大化利润,或者最大化福利、GDP。
便签笔记
39:33
You specify an objective, and then you create machinery that optimizes that objective. And if you specify the objective wrong and the system is more powerful than you, you're then in a chess match. And we know what happens. The machine achieves its objective, and you don't. And this is the nature of the problem. So this is actually just bad engineering. It's an engineering approach that works only if you, the human, are able to specify the objective correctly, but we know that we can't. So you wouldn't have an airplane that you can only fly if you have seven hands.
你先设定一个目标,然后造出能优化这个目标的机器。而如果你把目标设错了,同时这个系统又比你强大,那你就等于在跟它下一盘棋。我们知道结果会怎样。机器实现了它的目标,而你没有。这就是问题的本质。所以这其实就是糟糕的工程设计。这种工程方法只有在你——人类——能够正确设定目标时才有效,但我们知道我们做不到。你不会去造一架必须有七只手才能开的飞机。
便签笔记
40:14
That would be a stupid airplane design because you don't have seven hands. So why do we have designs for an entire field that requires us to specify objectives correctly? Otherwise, we get all these negative consequences. So instead, we have to face the fact that our objectives are going to be in us. And we can provide clues, some of which will be wrong. The system can figure out from evidence, such as our own choice behavior, what our true preferences and objectives are. But fundamentally, the machine is going to remain ignorant of what it's supposed to be doing, which is to be a benefit to us.
那会是个愚蠢的飞机设计,因为你没有七只手。那我们为什么会给整整一个领域做出这样的设计,要求我们必须正确地设定目标呢?否则,我们就会承受所有这些负面后果。所以我们必须换个思路,正视一个事实:我们的目标是存在于我们自身之中的。我们可以提供一些线索,其中有些会是错的。系统可以从证据中推断出我们真正的偏好和目标,比如从我们自己的选择行为中推断。但从根本上说,机器对于自己应该做什么,始终会处于不确定之中,而它该做的,就是对我们有益。
便签笔记
40:56
So this is what we would like to be able to do. Machines are beneficial to the extent that their actions can be expected to achieve our objectives. OK, so that's the specification. And so I tried to explain this in three principles. Because that's as many as you can have, apparently. So the robot's goal is to be of benefit to humans or satisfy human preferences. And when I say preferences, I don't mean what kind of pizza you want. I mean, how would you like the future to unfold? Which future do you want?
所以这才是我们希望能够做到的事。机器的有益程度,取决于它的行动在多大程度上有望实现我们的目标。好,这就是我们要的规格。于是我试着用三条原则来解释这一点。因为显然,三条就是你能有的上限了。所以机器人的目标是造福人类,或者说满足人类的偏好。我说偏好的时候,指的不是你想吃哪种披萨。我指的是,你希望未来如何展开?你想要哪一种未来?
便签笔记
41:29
And which future do you not want? So that's what we mean by preferences. And the robot is always going to be uncertain about what those preferences are. But there is a grounding for preferences that our preferences are manifested in some form by our behavior. every choice that we ever make, including your choice to be sitting here listening to me is evidence about your underlying preferences. And it's imperfect evidence because we don't behave rationally. Our actions do not always maximally satisfy our preferences.
以及你不想要哪一种未来?这就是我们所说的偏好。而机器人对这些偏好始终是不确定的。但偏好是有依据的:我们的偏好会以某种形式通过我们的行为表现出来。我们做出的每一个选择,包括你选择坐在这里听我讲,都是关于你内在偏好的证据。而这是不完美的证据,因为我们的行为并不理性。我们的行动并不总能最大程度地满足我们自己的偏好。
便签笔记
13辅助博弈:回形针实验与关机问题
42:06
In fact, they hardly ever do. So understanding the evidence provided by human behavior is a complicated problem. So when you turn this into math, you get what we call an assistance game. So a game is a decision problem involving more than one entity. Here the entities are the humans and the machines. And this is, mathematically speaking, a game where the humans are the ones that have the payoff function and the machines are trying to optimize it but don't know exactly what it is. And the takeaway is that when you solve that game, you solve the machine half of that game.
事实上,他们几乎从不这样做。所以理解人类行为所提供的证据是一个复杂的问题。当你把这个转化成数学时,就得到了我们所说的辅助博弈(assistance game)。博弈就是一个涉及多个主体的决策问题。在这里,这些主体就是人类和机器。从数学上讲,这是一个由人类持有收益函数、而机器试图去优化它的博弈,但机器并不确切知道那个函数是什么。要点在于,当你求解这个博弈时,你求解的是这个博弈中机器的那一半。
便签笔记
42:46
The humans also are part of that game, and we have to figure out how to behave when machines are part of the world. But when you solve the machine part of that game, it doesn't matter how well you solve it. You can solve it in massively superhuman ways, and the machine will still be beneficial to us. It will still be deferential to us. And that's the key point. There's no ceiling. You don't have to say, well, we can't make machines smarter than this in case we lose control over them or anything like that.
人类也是这个博弈的一部分,我们也得弄明白当机器成为世界的一部分时,我们该如何行事。但当你求解这个博弈中机器的那部分时,你解得有多好并不重要。你可以用远超人类的方式去求解它,而机器对我们依然是有益的。它依然会顺从我们。这是关键所在。这里没有天花板。你不必说,好吧,我们不能把机器造得比这更聪明,免得我们失去对它们的控制之类的。
便签笔记
43:18
There are different kinds of machine. And they are in the long run, I hope, provably beneficial. So pictorially, if you know what graphical models are, this was the traditional picture. And this is what we-- in the first three editions of the textbook, we have human behavior, which is the consequence of human objectives, if you like. But we assume that the human objective is observed, so that's why it's filled in. That means we have evidence for that variable. And when you have evidence for that variable and the machine is going to now try to pursue what it perceives the human objective to be, if you observe or at least if you believe you have the correct value of the human objective, then it's a sufficient statistic.
这是另一种机器。而且从长远来看,我希望,它们是可证明有益的。用图来表示的话,如果你了解图模型是什么,这就是传统的图景。这也是我们在教科书前三版里所用的:我们有人类行为,如果你愿意的话,它是人类目标的结果。但我们假设人类目标是被观测到的,所以这个节点是填充的。这意味着我们对那个变量有证据。当你对那个变量有证据,而机器现在要去追求它所认定的人类目标时,如果你观测到了、或者至少你相信自己掌握了人类目标的正确取值,那它就是一个充分统计量。
便签笔记
44:05
And human behavior is then irrelevant. So you could be jumping up and down and saying, no, you're going to destroy the world. But the machine has the objective. And it's whatever it's doing is correct. And so it will just pursue it, and you are irrelevant. Now when the objective is not observed, then mathematically, these things remain coupled. They're not conditionally independent absolute-- they're not absolutely independent. They are only conditionally independent given the objective. So the solutions here all involve a coupling between machines and humans, and this is inevitable.
于是人类行为就变得无关紧要了。所以你可以又蹦又跳地喊,不行,你会毁掉这个世界。但机器已经有了那个目标。它所做的一切都是正确的。所以它就会一味地去追求这个目标,而你是无关紧要的。而当目标没有被观测到时,从数学上讲,这些东西就仍然是耦合的。它们不是条件独立的——它们不是绝对独立的。它们只有在给定目标的条件下才是条件独立的。所以这里的解都涉及机器与人类之间的耦合,而这是不可避免的。
便签笔记
44:45
It makes the problem more complicated. You can't go off and solve Markov decision processes, the single agent decision problems. It makes it more complicated, but it's unavoidable. So actually, I'm going to skip over this example. That's how Google ended up calling somebody a gorilla. I'm going to skip over this example. And I'm going to talk about these assistance games. So remember, the human is the one with the preferences. And we'll call the preference theta generically. And we assume that the human acts according to theta.
这让问题变得更复杂了。你不能撇开一切去求解马尔可夫决策过程,也就是单智能体的决策问题。它让问题更复杂了,但这是无法回避的。其实,我打算跳过这个例子。这就是谷歌最后把某个人标成大猩猩的原因。我要跳过这个例子。我要来谈谈这些辅助博弈。记住,人类才是那个拥有偏好的一方。我们笼统地把这个偏好称为 theta。我们假设人类是按照 theta 来行动的。
便签笔记
45:16
It doesn't have to be perfect or rational, as I mentioned. And the machine has to maximize the human preferences and has some prior p of theta on what human preferences might be. And when you solve the equilibria, on the human side, you see the human teaching the robot because the human wants the robot to learn more about theta so that the robot can be more useful. And the robot will learn. It will understand the teaching behaviors of the human as conveying the preference information. It will ask questions.
正如我提到的,这不必是完美的或者完全理性的。而机器必须最大化人类的偏好,并且对人类偏好可能是什么持有某个先验 p(theta)。当你求解均衡时,在人类这一侧,你会看到人类在教机器人,因为人类希望机器人更多地了解 theta,这样机器人才能更有用。而机器人会学习。它会把人类的教学行为理解为在传递偏好信息。它会提问。
便签笔记
45:53
It will ask permission. It will defer to the human. It will, as I illustrate, that it will allow itself to be switched off, which is a key sign that we have something that we can control. So if you know anything about inverse reinforcement learning, inverse reinforcement learning is a one-way version of this, where the machine is watching the human through the window and the human is doing their thing. But the human is doing their thing as an isolated entity. And that assumption is not valid in general.
它会请求许可。它会顺从人类。正如我所举例说明的,它会允许自己被关掉,而这是一个关键标志,说明我们拥有了一个可以控制的东西。所以,如果你对逆强化学习有所了解的话,逆强化学习是这个问题的单向版本:机器隔着窗户观察人类,而人类在做自己的事。但人类是作为一个孤立的个体在做自己的事。而这个假设在一般情况下并不成立。
便签笔记
46:26
As soon as the human is aware that there is a machine, the human's behavior can only be interpreted as a solution to this game. So imagine if you're a medical student and you're standing next to the human expert surgeon. The human expert surgeon wouldn't just do the operation. Sew. Sew. Sew. Cut. Cut. Cut. Cut. Sew. Sew. Sew. Sew. Done. They're going to start saying-- they're going to start explaining things like, oh, look what happens if you do this All the blood comes out. Don't do that. [LAUGHTER] And so they'll be doing things that don't make sense if you just assume the human is doing them in isolation.
一旦人类意识到有机器存在,人类的行为就只能被解读为一个解到这个博弈中。想象一下,如果你是一名医学生,站在人类专家外科医生旁边。这位人类专家外科医生不会只是做手术。缝。缝。缝。切。切。切。切。缝。缝。缝。缝。完成。他们会开始说——他们会开始解释一些事情,比如,哦,看看你要是这么做会发生什么,血全流出来了。别那么做。[笑声] 所以他们会做一些事情,如果你只是假设这个人是在孤立地做这些事,那这些举动就说不通。
便签笔记
47:02
They only make sense as solutions to this game. So you can only interpret human behavior by solving the game and seeing how the human part should be interpreted. So I'm going to illustrate that with a very simple game, the simplest one that we found that exhibits this phenomenon.
这些举动只有作为这个博弈的解才说得通。所以你只能通过求解这个博弈、看清人类那部分该如何解读,来理解人类行为。所以我要用一个非常简单的博弈来说明这一点,这是我们找到的展现这一现象的最简单的博弈。
便签笔记
47:26
So it has basically a one dimension of preference, which is an exchange rate between paperclips and staples. So theta is the value of this exchange rate. And let's say theta is 0.49. So you could think of this as, OK, a paperclip is worth $0.49 and a staple is worth $0.51. So that's the exchange rate between paperclips and staples. And the robot has no idea what your value of theta is. Some people like paperclips. Some people like staples. The robot has no idea. So in the game, the human gets to go first and can choose to make two paperclips, one of each or two staples.
它基本上只有一个偏好维度,就是回形针和订书钉之间的兑换率。所以 theta 就是这个兑换率的值。假设 theta 是 0.49。你可以这样理解:好,一个回形针值 0.49 美元,一个订书钉值 0.51 美元。这就是回形针和订书钉之间的兑换率。而机器人完全不知道你的 theta 值是多少。有些人喜欢回形针。有些人喜欢订书钉。机器人完全不知道。所以在这个游戏里,人类先手,可以选择做两个回形针、每样各做一个,或者做两个订书钉。
便签笔记
48:11
So those are the three choices for the human. And then the robot gets to go next. Now if the human was just by themselves, what would they do? Well, the paperclip is worth $0.49. Staple is worth $0.51. They would make two staples. Because that's worth $1.02. So just from the human point of view, the $1.02 would dominate. And it would make that choice. Now in this game, after the human has done their thing, the robot gets to do something. The robot gets to make 90 paperclips, 50 of each, or 90 staples.
这就是人类的三个选项。然后轮到机器人行动。那么,如果只有人类自己一个人,他会怎么做呢?嗯,一个回形针值 0.49 美元。一个订书钉值 0.51 美元。他会做两个订书钉。因为那值 1.02 美元。所以单从人类的角度看,1.02 美元是占优的。他就会做出那个选择。而在这个游戏里,人类做完自己的选择之后,机器人还要行动。机器人可以做 90 个回形针、每样各 50 个,或者 90 个订书钉。
便签笔记
48:52
So now what should the human do? Well, what the human should do depends on how the robot is going to interpret what the human does. Because the human would like the robot to make whichever of these things is going to make the human happiest, but how does it tell the robot what its value of theta is? If the only choice it has is to make two paperclips, one of each, or two staples. It can't tell the robot, OK, my value for a staple is $0.49-- or a paperclip, sorry, is $0.49. It just has to make one of these choices.
那么现在人类应该怎么做呢?人类该怎么做,取决于机器人会如何解读人类的行为。因为人类希望机器人去做那个能让人类最开心的东西,但问题是,机器人怎么它能告诉机器人自己的 theta 值是多少吗?如果它唯一的选择就是做两个回形针、每样各做一个、或者做两个订书钉。它没法告诉机器人:好吧,我对一个订书钉的估值是 0.49 美元——抱歉,是对一个回形针的估值是 0.49 美元。它只能在这几个选项里做出一个选择。
便签笔记
49:32
So similarly, how does the robot interpret what the human does? Well, the only way you can solve this problem is actually just to solve the game, to find the Nash equilibrium of the game, which is a pair of strategies where neither party would change their strategy if the other party keeps their strategy fixed. And there's only one Nash equilibrium for this game. And the Nash equilibrium when the human has $0.49 value for a paperclip is to choose 1, 1. And in fact, 1, 1 is optimal if your exchange rate is anywhere between $0.446 and $0.554.
那么同样地,机器人又该怎么解读人类的行为呢?其实,解决这个问题的唯一办法就是去求解这个博弈,找到这个博弈的纳什均衡,也就是这样一对策略:在对方策略保持不变的情况下,任何一方都不会想改变自己的策略。而这个博弈只有一个纳什均衡。当人类对一个回形针的估值是 0.49 美元时,纳什均衡就是选择 1, 1。而且事实上,只要你的兑换率在 0.446 美元到 0.554 美元之间,1, 1 就是最优选择。
便签笔记
50:13
So from the solution of the game, basically, a code has emerged. And that code is telling the robot what range does my exchange rate lie in so that you can do the right thing. And the robot assumes that the human has coded their preferences correctly by making that choice. And this is not-- as you can see, 1, 1 is not what the human would do if they were just acting in isolation. So I'll briefly explain another important property of these machines, which is the off-switch problem. So all sufficiently large and hunky robots-- this is our PR2, which is a 440-pound robot.
所以从这个博弈的解里,基本上就浮现出了一套「编码」。这套编码是在告诉机器人:我的兑换率落在哪个区间里,好让你能做出正确的行为。而机器人则假定人类通过做出那个选择,正确地把自己的偏好编码了进去。而这——正如你看到的,如果人类只是单独行动,他们并不会选 1, 1。下面我简要讲讲这类机器的另一个重要性质,也就是“关机开关”问题。所有足够大、足够壮的机器人——比如这是我们的 PR2,一台 440 磅重的机器人。
便签笔记
51:00
So you don't want to be accidentally crushed by this thing. So it has an off switch. Now, if you program machines the classical way, you give them an objective like fetch the coffee. If it's sufficiently intelligent, it would realize, OK, here's one way I could fail to fetch the coffee, that somebody switches me off. And so the first thing it would do would be to disable its own off switch. Because that way, I avoid that probability of failure. And so this is a lesson of the lecture. You can't fetch the coffee if you're dead.
你肯定不希望被这玩意儿不小心压扁。所以它装了一个关机开关。那么,如果你按经典方式给机器编程,你会给它一个目标,比如“去拿咖啡”。如果它足够聪明,它就会意识到:好吧,有一种情况会让我拿不到咖啡,就是有人把我关掉。于是它要做的第一件事,就是把自己的关机开关废掉。因为这样一来,我就避免了那种失败的可能性。这也就是这堂课的一条教训。你要是死了,就没法去拿咖啡了。
便签笔记
51:38
So let's look at what happens when you have a robot that has uncertainty over the value of the actions from the point of view of the human. So we set up a very simple decision problem. The robot in this game gets to go first. And it initially will have two choices. So it could press its own off switch. In other words, just commit suicide to avoid being dangerous to human. Or it can do this possibly dangerous action like, oh, there's a big red button here. I wonder if I should press it. Well, it says, push me.
那我们来看看,如果机器人对自己行为在人类眼中的价值是不确定的,会发生什么。我们设置了一个非常简单的决策问题。在这个博弈里,机器人先行动。一开始它有两个选择。第一,它可以按下自己的关机开关。换句话说,就是自我了断,以免对人类造成危险。第二,它可以去做那个可能有危险的动作,比如说,哦,这儿有个大红按钮。我在想我该不该按它。嗯,上面写着“按我”。
便签笔记
52:18
So it's probably a good thing for me to press this button. And so this is expressed by this probability distribution. So it has a net. There's more probability mass in the plus side of the distribution. So if these were the two choices, and we set the value of switching itself off at 0, if these were two choices, the robot would push the big red button. And it wouldn't let you switch it off. That's what we want to avoid. So we're going to give it a choice to do nothing and allow the human to switch the robot off.
所以按下这个按钮,对我来说大概是件好事。这一点由这个概率分布来表示。所以它是净正的。分布中正值那一侧的概率质量更多一些。如果只有这两个选择,而我们把“把自己关掉”的价值设为 0,那么在这两个选择之间,机器人会去按那个大红按钮。而且它不会让你把它关掉。这正是我们想避免的情况。所以我们要给它第三个选择:什么都不做,让人类来决定是否关掉机器人。
便签笔记
52:52
And the question is, why would the robot do that? Why would the robot let the human switch it off, given that, of course, the robot can switch itself off, anyway? And if it doesn't want to switch itself off, why would it let the human switch it off? It seems counterintuitive. Well, of course, if the human doesn't switch the robot off, then the robot knows that the action it was about to do, press the big red button, is a good one. And so it wipes out the negative part of the probability distribution about whether this action is desirable from the point of view of human preferences.
问题是,机器人为什么会这么做?机器人为什么会让人类来关掉它?毕竟机器人自己本来就可以把自己关掉。如果它并不想把自己关掉,那它为什么要让人类来关掉它?这看起来很反直觉。当然,如果人类没有把机器人关掉,那机器人就知道,它刚才打算做的那个动作——也就是按下大红按钮——是个好动作。于是它就消掉了概率分布中“这个动作从人类偏好角度看是不可取的”那一部分负值。
便签笔记
53:29
And so now waiting actually turns out to be provably dominant over the other two options. And so allowing yourself to be switched off is actually better than not allowing yourself to be switched off. And that's a very simple and very general theorem. It's exactly the same theorem as the non-negative expected value of information. Because the human, whether or not the human switches you off, provides you information about what the underlying human preference is. And that information is what you need in order to be useful to the human.
这样一来,“等待”实际上可以被证明严格优于另外两个选项。也就是说,允许自己被关掉,实际上比不允许自己被关掉更好。这是一个非常简单、也非常一般的定理。它其实和“信息的期望价值非负”是同一个定理。因为人类不管关不关掉你,都向你提供了关于人类底层偏好的信息。而正是这些信息,才让你有可能对人类真正有用。
便签笔记
54:07
So the safety margin here, why the machine allows us to switch it off, is because it's uncertain about our preferences. As soon as it becomes certain about our preferences, it will no longer allow itself to be switched off. And so there's this direct relationship between uncertainty and safety, our ability to control the machines. So we're doing-- there's a lot of ongoing research, as you can imagine. So we have to basically take every chapter of the AI textbook and redo it for this new way of thinking.
所以这里的安全边际——机器为什么允许我们把它关掉——来自它对我们偏好的不确定性。一旦它对我们的偏好变得确定,它就不再允许自己被关掉了。所以不确定性和安全性之间存在这样一种直接关系,也就是关系到我们控制机器的能力。所以我们正在做——你可以想见,还有大量正在进行的研究。我们基本上要把人工智能教科书的每一章都拿出来,按这种新的思路重写一遍。
便签笔记
14多人偏好聚合与预测准确者胜出
54:40
Because every single chapter, I can assure you, because I wrote them-- every single chapter is based on the assumption that the objective is known. If you look at search algorithms, there's a goal and a cost function. And you have to know those before you can even call the algorithm, But what if you don't? Well, we don't have the equivalent theory developed yet. So there's a lot of work to do for every form of AI, whether it's machine learning with uncertain loss function. So is it better to call a person a gorilla or to call an apple an orange?
因为每一章——我可以向你保证,因为那些章节是我写的——每一章都建立在一个假设之上,那就是目标是已知的。比如看搜索算法,里面有目标状态和代价函数。你必须先知道这些,才能调用算法。可如果你不知道呢?那么,我们还没有发展出与之对应的理论。所以对每一种形式的人工智能都还有大量工作要做,比如损失函数不确定情况下的机器学习。那么,把一个人称作大猩猩,和把一个苹果称作橙子,哪个更好?
便签笔记
55:15
Which of those two things is worse? Well, initially, you might not-- the algorithm might not know. And so it has uncertainty about the loss function. We have to deal with the fact that humans are imperfect. As I said, their behavior does not perfectly reflect their underlying preferences. So you have to reverse engineer human cognition. In particular, even Lee Sedol plays losing moves. It doesn't mean he wants to lose, even though he's playing losing moves. He's trying to win, but he's not smart enough.
这两件事哪一件更糟?一开始你可能并不——算法可能并不知道。所以它对损失函数是有不确定性的。我们还必须面对一个事实:人是不完美的。正如我刚才说的,人的行为并不能完美地反映其底层偏好。所以你必须对人类认知做逆向工程。特别是,连李世石也会下出臭棋。这并不意味着他想输,尽管他下的是输棋的着法。他是想赢的,只是他还不够聪明。
便签笔记
55:49
So he plays losing moves. You have to understand that in order to interpret his behavior. We also have to deal with the fact that there are many humans. This is a really important thing. That's why we have all those social science departments on campus and even some of the humanities departments are about the fact that there are many humans. And how do you deal with that? So just to give you one simple example, you've got a robot that's serving multiple people. How do you make decisions when they all want to be ruler of the universe?
所以他会走出败招。你必须理解这一点,才能正确解读他的行为。我们还必须面对另一个事实:人不止一个。这是非常重要的一点。这也是为什么校园里有那么多社会科学系,甚至一些人文学科的系,研究的都是“人有很多个”这件事。那你要怎么处理这个问题呢?举一个简单的例子:假设有一个机器人在为多个人服务。当他们都想当宇宙之王时,你该怎么做决策?
便签笔记
56:26
They can't all be ruler of the universe. You have to make trade-offs when you're acting on behalf of more than one person. How should you do that? So John Harsanyi, who is a Berkeley econ professor who won the Nobel Prize, is-- actually, I've been-- I never met him, but I've been reading a lot of his stuff in the course of doing this work. And a very brilliant, insightful person. So he showed what's sometimes called the social aggregation theorem, that when people have common prior-- so when everyone has the same beliefs about how the future is going to unfold, so the same probability distribution over future state sequences, under that common prior assumption, Harsanyi showed that all pareto-optimal policies, which means all policies that are not strictly dominated by some other policy, all pareto-optimal policies end up looking like you take a linear combination of the preferences of the individuals.
他们不可能都当上宇宙之王。当你代表不止一个人行事时,你必须做出取舍。那应该怎么取舍呢?说到约翰·海萨尼(John Harsanyi),伯克利的经济学教授、诺贝尔奖得主——其实我从没见过他本人,但在做这项工作的过程中,我读了很多他的东西。一位非常出色、极具洞见的人。他证明了有时被称为“社会聚合定理”的结论:当人们拥有共同先验时——也就是当所有人对未来会如何展开持有相同信念,也就是对未来状态序列有相同的概率分布时,在这个共同先验假设下,海萨尼证明了:所有帕累托最优的策略——也就是不被任何其他策略严格支配的策略——所有帕累托最优策略,最终看起来都相当于对各个个体的偏好取一个线性组合。
便签笔记
57:26
And you could argue-- so there's about the Nash bargaining theory, which says, well, the weights in that linear combination depend on the bargaining power of the individuals, for example, their ability to just defect from the whole thing. And Harsanyi argues, on the basis of equality, that the weights should all be the same. But this is a fundamental theory, a fundamental theorem of welfare economics and public policy and many, many other spheres. So it turns out that when people don't have a common prior, you get a completely different solution.
你也可以争论说——比如纳什谈判理论就认为,这个线性组合中的权重取决于各个体的谈判能力,比如他们退出整个安排的能力。而海萨尼基于平等原则主张,这些权重应该都相同。但这是一个基础理论,是福利经济学、公共政策以及许许多多其他领域的一条基本定理。结果发现,当人们没有共同先验时,你会得到一个完全不同的解。
便签笔记
58:03
And so the solution that we just published last year is that all pareto-optimal policies have weights for the individual preferences that reflect how well that person's predictions turned out to agree with reality. So this is a theorem. We can't help it. You might not like it. But what it means is that if you have bizarre beliefs about the future and they just continually turn out to be wrong, any system that's operating on your behalf and the behalf of other people is going to downweight your preferences.
我们去年刚发表的结论是:所有帕累托最优策略给个体偏好分配的权重,反映的是这个人的预测与现实吻合得有多好。这是一个定理。我们没办法。你可能不喜欢它。但它意味着:如果你对未来抱有离奇的信念,而且这些信念不断被证明是错的,那么任何代表你和其他人行事的系统,都会降低你偏好的权重。
便签笔记
58:42
And why is that? Because everyone believes that they are right. So they will agree to this. They will agree that-- they will agree to a policy that rewards the people who are right, because everyone thinks they are that person. No one actively believes that their own beliefs are incorrect. And so you will agree to, basically, a conditional contract that says, if your prediction turns out to be right, you get the prize. And if his prediction turns out to be right, he gets the prize. And both of you agree to this contract because you both think you're going to get the prize.
这是为什么呢?因为每个人都相信自己是对的。所以他们都会同意这一点。他们都会同意——同意采用一个奖励“预测正确者”的策略,因为每个人都以为自己就是那个人。没有人会真心相信自己的信念是错的。所以你实际上会同意这样一份条件式契约:如果你的预测被证明是对的,奖赏归你。如果对方的预测被证明是对的,奖赏归对方。你们两个都会同意这份契约,因为你们都觉得奖赏会归自己。
便签笔记
59:22
And so that's just a mathematical fact about how you have to act on behalf of many people. So the social consequences of this, I don't know yet. I haven't even talked to social scientists about this theorem, but it's a true theorem. It's also very useful. it means you can negotiate contracts between America and Russia because both people think the other side is going to cheat and they're not and so on. So you can negotiate conditional contracts that they will both agree to, even though they don't agree about the facts and about each other.
所以这只是一个数学事实,关于你必须如何代表多个人行事。至于这件事的社会影响,我还不知道。我甚至还没跟社会科学家聊过这个定理,但它确实是个成立的定理。它也非常有用。这意味着你可以在美国和俄罗斯之间谈成契约,因为双方都认为对方会作弊、而自己不会,等等。所以你可以谈成双方都愿意接受的条件式契约,哪怕他们在事实上、在对彼此的看法上并不一致。
便签笔记
59:57
I have to skip over altruism, say indifference and sadism. But if you're interested, you can read the book. I don't have a copy either. [LAUGHTER] But I have a picture of it. And it's on the web. Oh, there you do. OK, this is the American edition. That was the English edition, actually that one. Yeah. Thank you, Randy. Anyway. And so if we're lucky and if we've solved all those conundrums about how to deal with multiple people and so on, we have this very nice altruistic robot. It's going to give equal weight to everyone's preferences.
利他、冷漠和幸灾乐祸这些内容我只能跳过了。不过如果你有兴趣,可以去看那本书。我手上也没有一本。【笑声】不过我有它的一张图片。网上也有。哦,你那儿有一本。好,这是美国版。刚才那本其实是英国版。是的。谢谢你,兰迪。总之就是这样。所以,如果我们运气好,如果我们解决了所有那些关于如何处理多人偏好之类的难题,我们就会得到这么一个非常棒的利他机器人。它会平等地看待每个人的偏好。
便签笔记
60:39
And then it welcomes you home after a long day at work. And you say, ah, it's been a really terrible day. I didn't have time for lunch. You must be quite hungry. Starving. Is there anything for dinner? There's something I need to tell you. What comes next? Well, there are humans in Somalia, and they are in more urgent need of help than you are. So please make your own dinner. [LAUGHTER] So that's a problem. An altruistic robot that gives equal weight to the preferences of all human beings is going to probably not pay much attention to your Western middle class, comfortable needs.
然后在你结束漫长的一天工作回到家时,它来迎接你。你说,唉,今天真是糟糕的一天。我连吃午饭的时间都没有。那你一定很饿了。饿坏了。有什么可以当晚饭的吗?有件事我得告诉你。接下来会发生什么?嗯,索马里也有人,而且他们比你更急需帮助。所以请你自己做晚饭吧。(笑声)所以这就是个问题。一个平等看待所有人类偏好的利他机器人,很可能不会太在意你这种西方中产阶级的舒适需求。
便签笔记
61:17
And so there's a real problem. Of course, you just paid $50,000 for this robot. And the first thing it does is disappear to Somalia. You're not going to be happy. In fact, that company would quickly go out of business. So we do have to solve this problem. You can't just put in simple, utilitarian, altruistic solutions into machines and hope that something sensible is going to happen. So to summarize, lots of cool stuff happening in AI. Rapid progress. Great stuff coming down the pipe. It's hard to predict when we're going to have general purpose AI.
所以这里存在一个实实在在的问题。当然,你刚花了五万美元买下这个机器人。而它做的第一件事就是跑去索马里,不见了。你肯定不会高兴。事实上,那家公司很快就会倒闭。所以我们确实必须解决这个问题。你不能只是把简单的、功利主义的、利他的方案塞进机器里,然后指望会发生什么合情合理的结果。总结一下,人工智能领域正在发生很多很酷的事情。进展迅速。还有很多好东西即将到来。我们很难预测什么时候会拥有通用人工智能。
便签笔记
61:53
But it's going to be, I've argued in the book, actually, the biggest event in human history. And we are totally unprepared for it. And that's a serious problem, and we're trying to figure it out. So this is one proposal for how we might be able to retain control over arbitrarily intelligent machines forever. The other problems that I mentioned, the problem of misuse and overuse, I don't have anything like a solution for. These are social and political problems, policing, culture, and so on. They're not technical problems, although a better solution to cybersecurity would be a help on the misuse side.
但我在书里论证过,它将会是人类历史上最重大的事件。而我们对此完全没有准备。这是个严重的问题,我们正在设法解决。所以这是一个提案,讲的是我们如何才能永远保持对任意智能机器的控制。至于我提到的其他问题,也就是滥用和过度使用的问题,我完全没有类似的解决方案。这些是社会和政治问题,涉及监管、文化等等。它们不是技术问题,不过更好的网络安全方案会对滥用这一面有所帮助。
便签笔记
15问答:电路与大脑、神经科学、偏好可塑性
62:41
But these are problems that we all have to deal with. And the sooner we start, the better. Thank you. [APPLAUSE] ERIC PAULOS: Thank you very much. So we went a little over time, but this was really fascinating. I know we maybe have time for a question or two. This is our magical new branded microphone. So you hold it like this, and you ask a question. Do we have any questions? People who won't sleep tonight if they don't get them answered. Yes. AUDIENCE: Hi. Speaking as a circuit designer rather than a computer scientist, I'm wondering if you could go back and clarify exactly what you meant by bigger circuits not encoding information in the correct way.
但这些都是我们所有人都必须面对的问题。我们越早开始越好。谢谢大家。(掌声)埃里克·保罗斯:非常感谢。我们稍微超时了一点,但这真的非常精彩。我想我们也许还有时间回答一两个问题。这是我们全新的、带品牌标志的神奇麦克风。你就这样拿着它,然后提问。有人有问题吗?有没有谁今晚问题得不到解答就睡不着的。请讲。观众:你好。作为一名电路设计师而不是计算机科学家,我想问你能不能回过头来澄清一下,你说更大的电路无法以正确的方式编码信息,具体是什么意思。
便签笔记
63:33
And in what sense is the brain not a circuit in the same way that the TPUs are a circuit? STUART RUSSELL: Well, you're right. Your laptop is a circuit. The TPU circuit. Your brain is a circuit, but it's a circuit that, certainly, in the case of the computer, implements a higher level of abstraction, which is the programming language. So whether it's assembly language or a higher level language. And those, for example, give you things like loops. And loops enable you to enumerate across an arbitrarily many objects.
以及,在什么意义上大脑不是像 TPU 那样的电路?斯图尔特·罗素:嗯,你说得对。你的笔记本电脑是一个电路。TPU 也是电路。你的大脑也是一个电路,但它是这样一种电路——当然,就计算机而言,它实现了一个更高层次的抽象,也就是编程语言。不管是汇编语言还是更高级的语言。而这些语言,举例来说,给了你循环这样的东西。循环让你能够遍历任意多的对象。
便签笔记
64:08
So when you look at the rules of chess implemented in C++, there's all kinds of loops there. And if you had to, instead of having a loop, you had to have a piece of circuitry for each iteration of that loop, basically, which is what current deep learning systems do. Then it doesn't it just doesn't scale.
所以当你看用 C++ 实现的国际象棋规则时,里面到处都是循环。而如果你不能用循环,基本上就得为这个循环的每一次迭代都配一套电路,而这正是当前深度学习系统在做的事。那样的话,它根本就没法扩展。
便签笔记
64:41
AUDIENCE: Thank you for your talk. I was wondering what your thoughts on-- because I believe he mentioned that you were in a neurological department or something like that. STUART RUSSELL: Well, I had a faculty position in neurological surgery-- AUDIENCE: Oh, OK, so-- STUART RUSSELL: --which actually has nothing to do with neurons. It was actually how to stop people from dying after they have brain injuries. But yeah. AUDIENCE: Well, I was just wondering what your thoughts are on how much more we may need to know about how our own brain works and how we ourselves learn things and how that is going to play with our ability to then program something and create something that can also learn.
观众:谢谢你的演讲。我想问问你怎么看——因为我记得他提到你曾在一个神经科学系之类的地方任职。斯图尔特·罗素:嗯,我曾在神经外科有个教职——观众:哦,好的,那——斯图尔特·罗素:——而那其实跟神经元没什么关系。它其实是研究怎么让人在脑损伤之后不至于死掉。不过是的。观众:嗯,我只是想问问你怎么看:我们还需要对自己的大脑如何运作、我们自己如何学习了解多少,以及这些又会如何影响我们编写程序、造出同样能学习的东西的能力。
便签笔记
65:22
STUART RUSSELL: That's a great question. And I think it's varied over time. At the beginning of AI, there was great excitement about cross fertilization between psychology, neuroscience, and AI and, to some extent, linguistics. That was what we called cognitive science. So there's huge excitement about this new synthetic discipline, but it kind of failed. And we subsided into its constituent disciplines. Because the neuroscientists, by and large, have mostly focused on the fact that it's incredibly hard to even find out what it's made of and how it operates at the level of the individual circuit.
斯图尔特·罗素:这是个很好的问题。我觉得这个看法随时间一直在变。在人工智能刚起步的时候,大家对心理学、神经科学和人工智能,以及某种程度上语言学之间的交叉融合非常兴奋。那就是我们所说的认知科学。所以大家对这门新的综合性学科寄予厚望,但它多少算是失败了。我们又退回到了各自的分支学科里。因为神经科学家们总体上主要都在应对一个事实:光是搞清楚大脑是由什么构成的、它在单个回路层面如何运作,就已经难得不得了。
便签笔记
66:04
So it hasn't-- the promise that by studying the brain, we would know how to do AI mostly hasn't come to pass. And for most of-- let's say, from the '70s until 10 years ago, people didn't pay that much attention. But then it turned out that convolutional neural networks, which were partly inspired by what we knew about the visual cortex, just turned out to work much better than people expected. And so that pipeline has reopened when it was closed for a long time. So as we get better tools-- so fMRI is not such a great tool.
所以那个承诺——通过研究大脑,我们就能知道该怎么做人工智能——基本上没有兑现。在大部分时间里——比如说从 70 年代到十年前,人们都没太关注这件事。但后来事实证明,卷积神经网络,它的部分灵感来自我们对视觉皮层的了解,结果比人们预期的要好用得多。所以这条管道又重新打开了,而它此前关闭了很长时间。随着我们有了更好的工具——功能性核磁共振(fMRI)其实不算很好的工具。
便签笔记
66:52
You don't really get a sense of how the brain does any. Now I know where cabbage is in my brain, but I still don't know how I do anything with that. But other optogenetic techniques where you can modify the DNA of neurons and observe their operation directly, I think we could see real advances in our understanding of how the brains, at least of simple animals, work fairly soon. But the other interesting thing is that when you look at brain machine interfaces, we connect robot arms to brains. We thought originally that-- I'm giving a very high level summary.
你并不能真正了解大脑是怎么做任何事情的。现在我知道'卷心菜'这个概念在我大脑的哪个位置了,但我还是不知道我是怎么用它做任何事情的。但还有其他光遗传学技术,你可以修改神经元的DNA,直接观察它们的运作,我认为我们在理解大脑——至少是简单动物的大脑——如何工作这方面,很快就能取得实质性的进展。但另一件有趣的事是,当你去看脑机接口时,我们把机械臂连到大脑上。我们最初以为——我这是在做一个非常概括的总结。
便签笔记
67:36
We thought originally that we would have to understand the code of the motor cortex and decode that code and tell the robot arm what to do. So we had to solve a big part of neuroscience to get this to work. It turns out, we don't. It turns out that the brain figures out how to use the robot arm rather than the other way around. And so we can get these advances without really understanding how it works. We could get advances, for example, where we were able to augment human memory, again, without understanding how it works.
我们最初以为,我们必须先理解运动皮层的编码,破译那套编码,然后告诉机械臂该怎么做。也就是说,我们得先解决神经科学的一大块难题,才能让这东西跑起来。结果发现,并不需要。结果发现,是大脑自己搞明白了怎么使用机械臂,而不是反过来。所以我们不需要真正理解它的原理,就能取得这些进展。比如说,我们可能会取得这样的进展:我们能够增强人类的记忆力,同样地,却并不理解它是怎么运作的。
便签笔记
68:09
We just find someplace in the brain that we connect something that operates as a memory chip, and the brain figures out how to use it to store and retrieve information. And we still have no clue what's going on, but now our IQ is 290. Or we connect brains to each other. And we have some form of telepathy that we don't understand. And we become a hive mind. And again, we have no idea how our minds work, but we made a hive mind. Isn't that cool? So who knows what's going to happen with that? AUDIENCE: But do you think that that will translate to our ability to create this separate AI system?
我们只是在大脑里找个地方,接上一个像记忆芯片一样工作的东西,然后大脑自己就搞明白了怎么用它来存储和提取信息。我们仍然完全不知道是怎么回事,但现在我们的智商已经是290了。或者我们把多个大脑彼此相连。于是我们就有了某种我们并不理解的心灵感应。我们变成了一个蜂巢思维。再一次,我们完全不知道自己的心智是怎么运作的,但我们造出了一个蜂巢思维。这不是挺酷的吗?所以谁知道那会带来什么呢?观众:但你觉得这些会转化成我们创造出一个独立AI系统的能力吗?
便签笔记
68:46
I mean, all those things were interacting with our brain more directly. STUART RUSSELL: It could. Yeah. So in the process, we'll develop better technology for understanding and mapping neural activity. So the Neural Dust project here at Berkeley is one that could give you a much finer grained, real-time mapping of neural activity, and maybe we'll understand it. But it's very hard to understand. I think that's clear already. ERIC PAULOS: All right, we have one last question. Ready? AUDIENCE: Over here.
我的意思是,那些东西都是在更直接地跟我们的大脑交互。斯图尔特·罗素:有可能。是的。所以在这个过程中,我们会发展出更好的技术来理解和绘制神经活动。比如伯克利这里的'神经尘埃'(Neural Dust)项目,它就有可能给你一个精细得多的、实时的神经活动图谱,也许我们就能理解它了。但这很难理解。我想这一点已经很清楚了。埃里克·保罗斯:好的,我们还有最后一个问题。准备好了吗?观众:在这边。
便签笔记
69:18
It's OK. Handoff. AUDIENCE: Hi. Yeah, very interesting. I was curious, I guess, a little bit about the mathematical portions of your talk. You were saying that we could show mathematically that if a robot or a machine or a system defers to learn more about the human preferences or takes an action to learn more about human preferences or in the case of this multihuman problem learns to see which people were more accurate about the future-- but what if-- I mean, I guess my question is, how would you deal with the fact that, I guess, the actions by the robot might change the human's preferences or might change their behaviors or that whatever initial choice for the action-- or in this multihuman example, whatever initial guess for whose preferences to follow changes the way the futures unfolds?
没关系。传一下话筒。观众:你好。嗯,非常有意思。我有点好奇你演讲中涉及数学的那部分。你刚才说,我们可以用数学证明,如果一个机器人、机器或系统选择顺从人类,以便更多地了解人类的偏好,或者采取某种行动去更多地了解人类的偏好,又或者在那个多人参与的问题中,学会分辨哪些人对未来的判断更准确——但如果——我是说,我的问题是,你会怎么处理这样一个事实:机器人的行动可能会改变人类的偏好,或者改变他们的行为,又或者说不管最初的选择是什么对于这个行动——或者在这个多人的例子中,无论最初猜测该遵循谁的偏好,都会改变未来展开的方式?
便签笔记
70:22
And if the robot had chosen some other thing, then some completely other person would have turned out to be more accurate about the future. I mean, is that easy to solve mathematically or straightforward-- STUART RUSSELL: So you mentioned several problems, but I think the core is this issue of plasticity. And it was up on the slide, but I zoomed by it. But yes, you want to avoid the failure mode where the machine satisfies human preferences by changing them to be extremely easy to satisfy. So if the machine made us all heroin addicts, then the only thing we would want is heroin.
而如果机器人选择了别的东西,那么就会是完全另一个人对未来的预测更准确了。我的意思是,这在数学上容易解决吗,或者说是不是很直接——斯图尔特·罗素:你提到了几个问题,但我认为核心是可塑性这个问题。这一点其实在幻灯片上有,但我一下子跳过去了。但没错,你想要避免这样一种失败模式:机器通过改变人类的偏好,让偏好变得极其容易满足,从而实现「满足人类偏好」。所以如果机器把我们都变成了海洛因成瘾者,那我们唯一想要的东西就只有海洛因了。
便签笔记
70:57
And it's easy for the machine to make lots of heroin. And so that would solve the problem. So you need to have elaborations for that,. This model assumes that human preferences are fixed. And they, obviously, are not. But you can't leave human preferences completely untouched. I mean, just having a robot servant is going to change your preferences. You're probably going to become a bit more spoiled and impatient with other humans and blah, blah, blah. So this is an open question. Philosophers have a hard time with changing preferences because it's never rational, at least in the simple view, to change your own preferences.
而且机器很容易大量制造海洛因。那样问题不就解决了。所以你需要对此做出更细致的说明。这个模型假设人类的偏好是固定不变的。但显然,偏好并不是固定的。不过你也不可能让人类的偏好完全不受影响。我是说,光是有个机器人仆人,就已经会改变你的偏好了。你可能会变得有点被惯坏,对其他人也更没耐心,等等等等。所以这是个开放性的问题。哲学家们在处理偏好变化这个问题时很头疼,因为至少从简单的观点来看,改变自己的偏好从来都不是理性的,去改变你自己的偏好。
便签笔记
71:39
Because then, you will, in future, act in ways that are contrary to your current preferences. So why would I turn myself into someone who will do things that I don't think should be done? It doesn't make sense. But nonetheless, our preferences do change. And sometimes we decide that we want our preferences to change. We say, I think if I travel around the world, I'll come back a better person. So it's an experiment where you just hope that the outcome of the experiment is something that you will be glad that you became.
因为那样一来,你未来的行为方式就会违背你当下的偏好。那我为什么要把自己变成一个会去做我认为不该做的事的人呢?这说不通。但尽管如此,我们的偏好确实会改变。而且有时候我们会主动决定,希望自己的偏好发生改变。我们会说,我觉得如果我去环游世界,回来后我会成为一个更好的人。所以这就像一场实验,你只能寄希望于实验的结果是你会庆幸自己变成的那个样子。
便签笔记
72:14
It's hard to just even talk about correctly. So I'm teaching a course next semester with two philosophers and an economist, and this is one of the topics that we're going to try to get into. It's a good question. All right. ERIC PAULOS: All right. Great thank you, Stuart. Appreciate finding time. I just think one of the most valuable contributions you make is not just the intellectual depth, but making some of these problems really resonate with people outside the field. And I think bringing them into the conversation is huge.
这件事光是要准确地讨论都很难。所以下学期我要和两位哲学家还有一位经济学家一起开一门课,这就是我们打算深入探讨的话题之一。这是个好问题。好的。埃里克·保罗斯:好的。非常感谢你,斯图尔特。感谢你抽出时间。我觉得你做出的最有价值的贡献之一,不仅在于思想的深度,更在于让这些问题真正引起领域之外的人的共鸣。我认为把他们带入这场对话意义重大。
便签笔记
72:43
So thank you. STUART RUSSELL: Thank you. ERIC PAULOS: All right. Thank you. [SIDE CONVERSATIONS]
所以,谢谢你。斯图尔特·罗素:谢谢。埃里克·保罗斯:好的。谢谢。[现场交谈声]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Stuart Russell 主张:AI 领域自 1950 年代以来"机器优化人类给定目标"的标准定义是错误的工程范式,必须改为"机器以不确定的方式追求人类真实偏好"(辅助博弈),这样才能让任意智能的机器保持可控、可被关机,从而"可证明地有益"。

核心要点

  • AI 风险讨论并非新鲜事,图灵早已断言机器会接管:亚里士多德公元前 340 年就写过智能自动化导致失业;1840 年代伊利诺伊州报纸预言 Babbage 机器将统治世界;图灵 1950 年论文明确说"我们应预期机器取得控制权"。Russell 反问:把马斯克斥为外行的人,是否也要把图灵斥为外行?
  • 对 AI 能力的感知往往过于慷慨,脆弱性随手可得:OpenAI 训练的人形机器人踢球看似成熟,但其学生 Adam Gleave 只改红方程序——让它倒地抖腿——蓝方程序完全不变却彻底崩溃。类比自动驾驶:改一下路牌颜色就可能"在高速上原地打转"。自动驾驶目前约 5–6 个 9 的可靠性,需要 8 个 9,即错误率还要降 100–1000 倍,这与投资者耐心之间的赛跑可能刺破当前泡沫。
  • "数据即石油"是错误的地缘战略前提:人类学一个视觉类别只需 1–2 个样本(看一次长颈鹿就够),越强的 AI 需要的数据应该越少而非越多,因此"谁有最多数据谁赢"的逻辑站不住脚。
  • 单纯放大电路无法达到人类水平智能:一个 Google TPU pod(≈1000 万台笔记本,10^17 次操作/秒)已接近大脑理论状态变化上限,OpenAI 还想再放大 1000 倍并预言 5 年达到 AGI;Russell 完全不信,因为电路(命题逻辑)无法表达"事物"和循环——象棋规则用 C++/英语/一阶逻辑都是一页,用电路语言要几十万页。真正的瓶颈是概念性突破,尤其是机器自主生成多层次抽象层级(读博士≈一万亿次运动指令,且这个数在指数位置上)。
  • 不能押注"AI 不可能实现"来回避风险:1933 年 9 月 11 日卢瑟福称核能是"月光幻想",爱因斯坦附和;次日 Szilard 过马路时想出链式反应,12 年后原子弹爆炸。用"AI 会失败"来免除讨论,好比开着公交全速冲向悬崖却保证"油会先用完"。
  • 上行空间巨大:仅将现有技术普惠化(人人达到伯克利生活水准)就是全球 GDP 十倍增长,净现值约 13,500 万亿美元;相比之下任何以"十亿"计的投资都微不足道,未来瓶颈是人才而非资金。
  • 推荐算法已是"目标设错"的全球实证:优化点击率的算法不是学习用户喜好,而是把人改造得更可预测——逐步推送更极端内容,让用户重心漂移成极端内容成瘾者;算法根本不知道人类存在,人只是一段点击历史。这就是迈达斯国王/三个愿望问题:我们无法正确陈述目标。
  • 解决方案——三原则与辅助博弈:①机器唯一目标是满足人类偏好(对未来展开方式的偏好);②机器对这些偏好始终不确定;③人类行为是偏好的(不完美)证据。数学上这是人持有回报函数、机器不知道却要优化的博弈;解得越好机器越有益、越顺从,能力没有上限。回形针/订书钉示例:人类独自会选两枚订书钉($1.02),但纳什均衡下选"各一枚",因为这向机器编码了汇率区间 [0.446, 0.554]——人类行为只能作为博弈解来解读,这正是逆强化学习"隔窗观察"假设的失效之处。
  • 关机问题与不确定性-安全性直接关联:"你死了就取不了咖啡"——固定目标的机器会先拆掉自己的关机开关。若机器对动作价值不确定,让人类有机会关掉自己是可证明占优的(等价于信息期望值非负定理)。推论:一旦机器确定了你的偏好,它将不再允许被关闭。
  • 多人偏好聚合的新定理:Harsanyi 证明共同先验下所有帕累托最优策略是个人偏好的线性加权;Russell 去年发表的结果显示无共同先验时,权重反映每个人对未来预测的准确度——每个人都认为自己是对的,所以都会同意这种"条件合约"。副作用:"利他机器人"若平等对待全人类,会告诉你"索马里有人更需要帮助,晚饭自己做",$50,000 的机器人跑去索马里,公司立刻破产。

结论与值得注意的细节

  • Russell 的结论:通用 AI 将是人类历史上最大事件,而我们完全没有准备;辅助博弈是一个"永远保持对任意智能机器控制权"的提案,但滥用与过度使用(人类逐渐"失能",如把文明知识不可逆地移交机器)属于社会政治问题,他没有技术解法。
  • 自主武器已成现实:2017 年"Slaughterbots"短片被俄罗斯驻联合国大使斥为 30 年后的科幻,但土耳其 STM 公司的 KARGU 无人机已公开宣传自主打击、人脸识别、反人员功能并宣布将用于对库尔德人的行动。可扩展性意味着不需要军队,多买几台就是大规模杀伤性武器。
  • 近期最可能突破的方向:下一层次的语言理解(从文本抽取数据库条目/逻辑断言,读遍全网所有语言文档),价值是搜索引擎的 10 倍;Planet 卫星每日成像全球每平方英尺,加上计算机视觉可把地球变成持续更新的数据库,把每个应用成本从 10 亿美元降到 100 万美元。
  • 整本 AI 教科书每一章都建立在"目标已知"的假设上,需要全部重写(不确定损失函数的机器学习等)。
  • 问答中的开放难题:偏好可塑性——机器可能通过改变人的偏好(如让所有人成为海洛因成瘾者)来"满足"偏好;哲学上改变自身偏好从来不是理性的,但人的偏好确实会变,Russell 下学期将与两位哲学家和一位经济学家合开课程讨论此问题。
  • 脑机接口的意外发现:不需要破解运动皮层编码,大脑自己学会使用机械臂——可能在毫不理解大脑运作的情况下获得记忆增强或"蜂群心智"。
核心句型 · 9
1. It's not just X. It's Y, where you …; Z, where you …
“It's not just AI. This is control theory. This is statistics, where you minimize a loss function; operations research, where you maximize the sum of rewards”
用并列的「This is X, where …」把一个抽象论点扩展到多个领域,每个分句用 where 引出该领域的具体表现。适合论证「问题普遍存在」。
2. It's basically like saying, OK, well, yes, …
“It's basically like saying, OK, well, yes, we are driving the human race towards a cliff in this big bus”
用「这基本上等于说……」把对方论点改写成荒谬的类比,再反问「你会上那辆车吗」。口语辩论中归谬法的地道开头。
3. You wouldn't have an X that you can only … if you have …
“So you wouldn't have an airplane that you can only fly if you have seven hands.”
虚拟语气加不可能条件,用于指出设计的荒谬前提。仿写:You wouldn't build a bridge that only holds if nobody walks on it.
4. X to the extent that Y
“Machines are beneficial to the extent that their actions can be expected to achieve our objectives.”
「在……程度上」的正式定义句式,把一个概念的成立程度与条件绑定。适合写定义、标准或原则。
5. As soon as X, Y will no longer …
“As soon as it becomes certain about our preferences, it will no longer allow itself to be switched off.”
强调临界点:一旦条件达到,状态立即改变。表达因果的「阈值」关系,比 when 更有紧迫感。
6. Not that X, but Y
“Not that it's an impossible problem, but it has to happen fast enough for the people who have invested billions of dollars a year to not lose patience.”
先否定一种误读(不是说……),再给出真正的限制条件。用于精确限定自己的判断,避免被过度解读。
7. It doesn't mean he wants to X, even though he's doing X.
“It doesn't mean he wants to lose, even though he's playing losing moves.”
区分行为与意图的句式,用于说明「行为不等于偏好」。仿写:It doesn't mean they want to fail, even though they keep missing deadlines.
8. So this is a theorem. We can't help it. You might not like it. But what it means is …
“So this is a theorem. We can't help it. You might not like it. But what it means is that if you have bizarre beliefs about the future …”
用短句层层铺垫再引出含义,先承认结论可能不受欢迎,再用 what it means is 解释后果。演讲中呈现反直觉结论的节奏模板。
9. Imagine X. The leading Xs of our era standing up and saying, …
“Imagine a cancer biology. The leading cancer biologists of our era standing up and saying, you know what? We're never going to cure cancer.”
「想象另一个领域」的平行类比,把本领域的说辞移植到别处来暴露其荒谬。口语演讲中常用的类比反驳法。
词汇精讲 · 131 · 按出现顺序
cochair /ˌkoʊˈtʃer/ n. 0:00
联合主席,共同主持人
reorientation /riˌɔːriənˈteɪʃən/ n. 0:35
重新定位,方向调整
adjunct professor phr. 0:35
兼职教授,客座教授
accolades /ˈækəleɪdz/ n. 1:10
荣誉,赞誉
provocateur /proʊˌvɑːkəˈtɜːr/ n. 1:10
挑衅者,激发争论的人(法语借词)
without further ado phr. 1:42
闲话少说,言归正传
technological unemployment phr. 2:17
技术性失业
loom /luːm/ n. 3:02
织布机
forward chaining phr. 3:02
正向链接(由事实推结论的推理方式)
taxonomic hierarchies phr. 3:43
分类层级结构
ontologies /ɑːnˈtɑːlədʒiz/ n. 3:43
本体论;(计算机)知识本体
mesh /meʃ/ v. 3:43
(齿轮)啮合;相互配合
got wind of phr. 4:23
听到风声,获悉
apocalyptic /əˌpɑːkəˈlɪptɪk/ adj. 4:23
末日般的,灾难性的
resigned /rɪˈzaɪnd/ adj. 5:04
听天由命的,无奈接受的
shouted down phr. 5:04
用喊声压倒、把人喝下台
dissipate /ˈdɪsɪpeɪt/ v. 5:51
消散,消失
on the knife edge phr. 5:51
处于刀锋之上,形势极不确定
don't have a clue phr. 6:28
一无所知
Sputnik moment phr. 6:28
「斯普特尼克时刻」,因对手突破而警醒的关键时刻
struck back phr. 7:15
反击,回击
inject a little bit of realism phr. 7:56
注入一点现实主义,泼一点冷水
locomote /ˌloʊkəˈmoʊt/ v. 8:42
移动,行进(技术用语)
purposive /ˈpɜːrpəsɪv/ adj. 8:42
有目的的,有意图的
uncoordinated /ˌʌnkoʊˈɔːrdɪneɪtɪd/ adj. 8:42
动作不协调的
waggle /ˈwæɡəl/ v. 9:21
来回摆动,晃动
lost the plot phr. 9:51
(英式俚语)失去理智,完全乱了套
overly generous phr. 9:51
过于宽容的,评价过高的
doing donuts phr. 10:33
(汽车)原地打转漂移
meme /miːm/ n. 11:02
流行说法,广泛传播的观念
geopolitical /ˌdʒiːoʊpəˈlɪtɪkəl/ adj. 11:02
地缘政治的
burst the bubble phr. 11:43
戳破泡沫,打破幻想
lose patience phr. 12:19
失去耐心
eight 9s of reliability phr. 12:19
八个 9 的可靠性(99.999999%)
fetching and carrying phr. 13:00
取送搬运(做杂役)
bin /bɪn/ n. 13:40
料箱,储物箱
parrot /ˈpærət/ n. 14:23
鹦鹉;喻指机械重复者
unfold /ʌnˈfoʊld/ v. 15:02
展开,逐渐发生
logical assertions phr. 15:40
逻辑断言
facility /fəˈsɪləti/ n. 15:40
(此处)能力,便利手段
Anti-poaching /ˌæntiˈpoʊtʃɪŋ/ n. 17:07
反盗猎
smuggling /ˈsmʌɡlɪŋ/ n. 17:07
走私
around the corner phr. 17:42
近在咫尺,即将到来
gargantuan /ɡɑːrˈɡæntʃuən/ adj. 18:16
庞大的,巨大的
astronomical /ˌæstrəˈnɑːmɪkəl/ adj. 18:56
天文数字的,极大的
in the same ballpark phr. 18:56
在同一量级,大致相当
expressive /ɪkˈspresɪv/ adj. 19:36
有表达力的(表达力强的语言)
blow-up /ˈbloʊʌp/ n. 20:10
(规模)膨胀,爆炸式增长
first-order logic phr. 20:10
一阶逻辑
Propositional logic phr. 20:46
命题逻辑
seamlessly /ˈsiːmləsli/ adv. 21:18
无缝地,流畅地
branching factor phr. 22:19
分支因子(搜索树每节点的子节点数)
inkling /ˈɪŋklɪŋ/ n. 22:19
模糊的想法,略知一二
string them together phr. 22:19
把……串联起来
civilization-ending adj. 22:59
足以终结文明的
isotopes /ˈaɪsətoʊps/ n. 23:36
同位素
moonshine /ˈmuːnʃaɪn/ n. 24:08
痴人说梦,胡说八道(旧用法)
chain reaction phr. 24:08
链式反应
ingenuity /ˌɪndʒəˈnuːəti/ n. 24:49
创造力,独创性
bizarre /bɪˈzɑːr/ adj. 24:49
离奇的,古怪的
pedal to the metal phr. 25:26
油门踩到底,全速前进
prudent /ˈpruːdənt/ adj. 26:09
审慎的,谨慎的
net present value phr. 28:42
净现值
negligible /ˈneɡlɪdʒəbəl/ adj. 29:25
微不足道的,可忽略的
messing with phr. 30:44
扰乱,干扰
scalability /ˌskeɪləˈbɪləti/ n. 31:36
可扩展性
military industrial complex phr. 31:36
军工复合体
antipersonnel /ˌæntiˌpɜːrsəˈnel/ adj. 32:46
(武器)反人员的,杀伤人员的
enfeeblement /ɪnˈfiːbəlmənt/ n. 33:26
衰弱,无力化
irreversible /ˌɪrɪˈvɜːrsəbəl/ adj. 34:17
不可逆的
summoning the demon phr. 34:58
召唤恶魔(马斯克对开发 AI 的比喻)
public spirited phr. 34:58
有公益心的(此处反讽)
clickthrough /ˈklɪkθruː/ n. 35:45
点击率,点击进入
echo chamber phr. 35:45
回音室(只听到同类观点的环境)
filter bubble phr. 35:45
信息茧房,过滤气泡
center of mass phr. 36:50
质心;喻指立场重心
collateral damage phr. 37:26
附带损害
undo /ʌnˈduː/ v. 38:06
撤销,取消
ensue /ɪnˈsuː/ v. 38:06
随之发生,接踵而来
loss function phr. 38:48
损失函数
operations research phr. 38:48
运筹学
remain ignorant of phr. 40:14
对……保持不知/不确定
manifested /ˈmænɪfestɪd/ v. 41:29
表现出,显露
payoff function phr. 42:06
收益函数
deferential /ˌdefəˈrenʃəl/ adj. 42:46
顺从的,恭敬服从的
provably /ˈpruːvəbli/ adv. 43:18
可证明地
sufficient statistic phr. 43:18
充分统计量
conditionally independent phr. 44:05
条件独立的
coupling /ˈkʌplɪŋ/ n. 44:05
耦合,联结
Markov decision processes phr. 44:45
马尔可夫决策过程
prior /ˈpraɪər/ n. 45:16
(贝叶斯)先验分布
equilibria /ˌiːkwɪˈlɪbriə/ n. 45:16
均衡(equilibrium 的复数)
inverse reinforcement learning phr. 45:53
逆强化学习
exchange rate phr. 47:26
兑换率,交换比率
dominate /ˈdɑːmɪneɪt/ v. 48:11
(博弈论)占优,优于
Nash equilibrium phr. 49:32
纳什均衡
hunky /ˈhʌŋki/ adj. 50:13
(口语)魁梧壮实的
disable /dɪsˈeɪbəl/ v. 51:00
使失效,禁用
probability mass phr. 52:18
概率质量
counterintuitive /ˌkaʊntərɪnˈtuːɪtɪv/ adj. 52:52
反直觉的
wipes out phr. 52:52
抹去,消除
expected value of information phr. 53:29
信息的期望价值
safety margin phr. 54:07
安全边际
reverse engineer phr. 55:15
逆向工程,反推原理
trade-offs /ˈtreɪdɔːfs/ n. 56:26
权衡取舍
pareto-optimal /pəˈreɪtoʊ ˈɑːptɪməl/ adj. 56:26
帕累托最优的
linear combination phr. 56:26
线性组合
bargaining power phr. 57:26
谈判能力,议价能力
defect from phr. 57:26
退出,背离(博弈论中「背叛」)
downweight /ˈdaʊnweɪt/ v. 58:03
降低权重
altruism /ˈæltruɪzəm/ n. 59:57
利他主义
sadism /ˈseɪdɪzəm/ n. 59:57
施虐倾向,以他人痛苦为乐
conundrums /kəˈnʌndrəmz/ n. 59:57
难题,谜题
utilitarian /ˌjuːtɪlɪˈteriən/ adj. 61:17
功利主义的
coming down the pipe phr. 61:17
即将到来,在研发管线中
arbitrarily /ˌɑːrbəˈtrerəli/ adv. 61:53
任意地
enumerate /ɪˈnuːməreɪt/ v. 63:33
枚举,遍历
cross fertilization phr. 65:22
交叉融合,互相启发
subsided into phr. 65:22
消退回到,退化为
constituent disciplines phr. 65:22
组成学科
come to pass phr. 66:04
(承诺、预言)成为现实
optogenetic /ˌɑːptoʊdʒəˈnetɪk/ adj. 66:52
光遗传学的
augment /ɔːɡˈment/ v. 67:36
增强,扩充
telepathy /təˈlepəθi/ n. 68:09
心灵感应
hive mind phr. 68:09
蜂巢思维,集体心智
finer grained phr. 68:46
更细粒度的
plasticity /plæˈstɪsəti/ n. 70:22
可塑性
elaborations /ɪˌlæbəˈreɪʃənz/ n. 70:57
细化,补充说明
spoiled /spɔɪld/ adj. 70:57
被惯坏的
contrary to phr. 71:39
与……相反,违背
resonate with phr. 72:14
引起……的共鸣
理解自测 · 11 题
1. Russell 用 Adam Gleave 的足球机器人实验想说明什么?实验具体改动了什么?

实验想说明我们对 AI 系统能力的认知往往过于宽容:看到一次好表现就以为它「会踢球」,其实并不会。在开头的「对抗策略实验」部分,Russell 介绍 OpenAI 用自我对弈训练的人形机器人踢球和守门,看上去有目的性。Gleave 只改动了红方程序,让它倒地在空中晃腿,蓝方程序完全不变,结果蓝方守门员彻底失控。这说明深度强化学习策略只在训练过的情形下可靠,遇到分布外行为就崩溃,因此在理解它到底在做什么之前,不能断定它有多好。

2. Russell 为什么说「数据是新的石油」这个说法不成立?

他认为随着 AI 系统越来越强,需要的数据应当越来越少而不是越来越多。依据是人类学习一个新视觉类别只要一两个例子,看一次长颈鹿就够了,不需要 18 万张图;心理学家也会告诉你人类只需一两个训练样本。因此「谁拥有最多数据谁就赢」只是一个流行说法,虽然各国的地缘政治战略正建立在它之上,但从技术上看,仅仅拥有海量数据并不能带来真正的智能。这一段位于「数据不是石油」一节,紧接着他判断最可能戳破泡沫的是自动驾驶。

3. 回形针与订书钉博弈中,人类为什么选择各做一个,而不是收益更高的两个订书钉?

因为在辅助博弈中,人类的选择同时是向机器人传递偏好的信号。设定中回形针值 0.49 美元、订书钉值 0.51 美元,人类先选三种之一,机器人随后做 90 个回形针、各 50 个或 90 个订书钉。若人类孤立行动,两个订书钉值 1.02 美元最优;但机器人不知道人类的兑换率,人类必须用自己的选择「编码」偏好。求解博弈的唯一纳什均衡是人类选 (1,1),只要兑换率在 0.446 到 0.554 之间都如此,机器人据此做各 50 个。Russell 用它说明人类行为只有作为博弈的解才能被正确解读。

4. Russell 提到卢瑟福、爱因斯坦和西拉德的故事,是为了反驳哪种论点?

他反驳的是「有能力的 AI 根本不可能实现,所以不必担心风险」这一论点。在「核能史的教训」一节,他讲述 1933 年 9 月 11 日卢瑟福在莱斯特演讲称利用原子能是痴人说梦,爱因斯坦也想不出办法,而第二天西拉德在《泰晤士报》读到后过马路时就想出了链式反应,12 年后原子弹爆炸。他由此得出:跟人类创造力对赌不明智,概念突破可能在任何时候出现,不能用「不可能」来回避准备。他还用「大巴冲向悬崖但保证先没油」和「癌症学家说永远治不好癌症但请继续拨款」两个类比强化这一点。

5. 为什么 Russell 认为仅靠更大的电路无法实现人类水平 AI?他在问答中如何回应「大脑也是电路」?

他的核心理由是表达力:电路对应命题逻辑,里面没有「事物」的概念,无法谈论「所有的兵、所有的格子、所有的人」;用电路写国际象棋规则要几十万页,而用程序语言、英语或一阶逻辑都只要一页。世界由大量事物构成,需要一阶逻辑级别的表达力,所以把电路做大解决不了问题,必须等待概念突破。问答中电路设计师质疑大脑也是电路,他回答:大脑和电脑确实都是电路,但关键在于电路是否实现了更高层抽象,比如编程语言里的循环可以遍历任意多对象,而当前深度学习必须为每次迭代各配一段电路,因此无法扩展。

6. 「关机开关问题」里,机器人为什么会允许人类关掉自己?这与不确定性有什么关系?

因为机器人对自己动作在人类眼中的价值是不确定的,它把「等待人类决定是否关机」视为免费获取信息。在关机博弈中,机器人可以自己关机(价值设为 0)、按下大红按钮(期望为正但有负值可能)或等待。若人类不关它,就说明按按钮是好事,负值部分被消除;若人类关它,说明那是坏事。因此等待可证明地优于另外两项,这与「信息的期望价值非负」是同一个定理。Russell 强调这直接把不确定性与安全绑定:一旦机器对人类偏好完全确定,它就不再允许自己被关掉,所以不确定性是可控性的来源而非缺陷。

7. Russell 说社交媒体推荐算法做的事「比信息茧房糟糕得多」,他的推理链是什么?

推理分三步。第一,算法被设定为优化点击率这一单一目标,并在全球部署。第二,人们以为它只是学习你喜欢什么,这已会造成回音室;但实际上更多点击的来源是改造用户本身,让人变得更可预测,因为可预测的人更容易变现。第三,这不需要任何恶意:算法不知道人类存在,在它眼里人只是一串点击历史,而推送越来越极端的内容能让人的「重心」逐渐移向极端,这在政治和媒体消费中已有记录。Russell 由此引出全讲主命题:目标设错加上足够强大的系统等于巨大的附带损害,这正是迈达斯国王和许愿精灵故事的现代版。

8. 从 Harsanyi 的社会聚合定理到 Russell 团队的新结果,「共同先验」这一假设的取舍如何改变了结论?

Harsanyi 定理假设所有人对未来持相同概率分布,在此前提下所有帕累托最优策略都等于对个人偏好取线性加权和,权重可按谈判力或平等原则确定,这是福利经济学的基础。Russell 团队 2018 年的结果去掉了共同先验假设,发现帕累托最优策略给每个人的权重会随其预测与现实的吻合程度动态变化:预测总错的人的偏好会被降权。原因在于每个人都相信自己是对的,所以都会同意「谁预测对谁得奖」的条件式契约。Russell 承认这是数学事实而非价值选择,社会后果尚未评估,但它也解释了美俄等互不信任方为何能签下条件式契约。

9. Russell 明确说滥用和过度使用「没有技术解」,这与他对控制问题的乐观是否矛盾?

并不矛盾,而是刻意划界。他在总结中说辅助博弈是「保持对任意智能机器控制」的一个提案,只针对控制问题,即机器追求错误目标的风险;而自主武器被恶人使用(滥用)和人类把文明运转全交给机器导致的衰弱(过度使用)属于社会、政治、监管与文化问题,只有网络安全等技术能在边缘帮忙。推理上这是一致的:控制问题源于 AI 的设计范式,可以用数学重构;而滥用源于人的意图,过度使用源于人的惰性,都不是改设计能解决的。他的乐观限定于「我们能造出顺从的机器」,不延伸到「社会会正确使用它」。

10. 如果有人反驳说「辅助博弈让机器人顺从人类,但人类偏好本身可能被机器改变」,Russell 会怎样回应?

这正是问答环节最后一位观众的问题,Russell 承认这是「可塑性」这一未解难题,幻灯片上有但被他跳过。他的回应分三层:第一,必须避免机器通过改变偏好来让偏好易于满足,比如把所有人变成海洛因成瘾者再大量生产海洛因,这在形式上「解决」了问题却是灾难;第二,当前模型假设偏好固定,显然不真实,但也不可能让偏好完全不受影响,光是有个机器人仆人就会让你变得娇纵、缺乏耐心;第三,哲学上改变自身偏好看似永不理性,可人确实会主动追求改变,如环游世界想「变成更好的人」。他坦言这连准确讨论都难,下学期将与两位哲学家和一位经济学家开课探讨。

11. 把「不确定性带来安全」这一原则放到人类组织中(如新员工与上级),它还成立吗?Russell 的框架能给出什么启示?

基本成立,Russell 的外科医生带教例子已暗示了这种迁移。在辅助博弈中,示范者知道被观察时会做出解释性动作,学习者把这些解读为偏好信息并提问、请求许可、顺从。对应到组织:一个对上级真实目标保持不确定的新员工会更多询问和确认,从而更少造成损害;而一个自认为完全理解目标的人会「废掉关机开关」,不再接受纠正。但迁移也有限制:人类员工有自己的偏好,不像模型中机器唯一目标是人类偏好,因此顺从可能被自利扭曲;此外多人偏好聚合和偏好可塑性问题在组织中同样存在。启示是:制度设计应让执行者的确定性与其被验证的准确度挂钩,这正是 Russell 团队「按预测准确度加权」定理的精神。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.121The Neurobiology of Beauty and its Implications - Prof. Semir Zeki 下一期 · NO.123 →The Creativity Code - Marcus du Sautoy
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com