[ICAPS 2020] Polanyi vs. Planning (Planning around AI's New Romance with Tacit Knowledge) · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

[ICAPS 2020] Polanyi vs. Planning (Planning around AI's New Romance with Tacit Knowledge)

节目发布 2020-10-27 · Subbarao Kambhampati
苏巴拉奥·坎帕蒂 主主持人 Alvaro
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2020年10月,国际自动规划与调度会议(ICAPS 2020)因疫情改为线上举行,亚利桑那州立大学教授苏巴拉奥·坎巴哈帕蒂(Subbarao Kambhampati,学界通称拉奥)应邀作特邀报告,题为《波兰尼对阵规划:在AI与隐性知识的新恋情周围做规划》。拉奥曾任美国人工智能促进会(AAAI)主席,长期研究人机协作与可解释AI,也是ICAPS「费斯提沃斯」(Festivus)吐槽传统的创始人。这次报告不谈论文写作,也不谈他自己的具体成果,而是借哲学家波兰尼的「隐性知识」概念,剖析深度学习时代规划研究的处境与出路。本文依据现场录音编译整理。

主持人介绍与讲者背景

主持人: 接下来是本次会议的特邀报告。很高兴向各位介绍苏巴拉奥·坎巴哈帕蒂,在座大多数人可能都叫他拉奥。他任教于亚利桑那州立大学,领导着成果丰硕的Yochan实验室,研究人机交互、可解释AI以及许多相关方向,是一个研究面极广的团队。拉奥的成就我可以讲上几个小时:他担任过美国人工智能促进会主席,拿过无数研究奖与教学奖。我特别推荐大家去看看他的个人主页,那个主页非常有趣,他还有一个YouTube频道,也值得一看。

主持人: 说实话,在ICAPS这里,拉奥不需要太多介绍,我相信各位早就听说过他。我想强调一点:ICAPS的「费斯提沃斯」传统就是他发起的,希望今天的报告里也能看到一些那种味道。最后我想说,拉奥大概是我们这个领域里最接近摇滚明星的人,他甚至有自己的IMDb页面,这一点值得一提。在把话筒交给拉奥之前,我还要告诉大家,本场报告将全程录像,各位若参与互动,请知悉这一点。

本次演讲不谈什么

拉奥: 大家好。我其实看不到你们,这种讲法很奇怪,但我在尽力让自己进入状态。我要讲的题目相当拗口,叫「波兰尼对阵规划:在AI与隐性知识的新恋情周围做规划」。

拉奥: 先说明一点:这不是那种常规的博士生论坛导师报告。那种报告我以前讲过,2009年讲过一次,那时候在座的一些人可能还没出生,我头发也多得多,肉也少得多。那次讲的是怎么做报告、怎么写论文。2013年我在IJCAI又讲了一遍,那个版本在YouTube上能找到。如果你想听听某一个人对于写论文、做报告的看法,至少是我个人的看法,可以去看看。尤其是如果今天这场讲砸了,你至少能知道什么不该做。总之,那种导师式的报告我今天不讲。

拉奥: 看起来我也可能讲一场关于自己研究的正经学术报告。结果这也不对。正如阿尔瓦罗刚才说的,我做的是人类感知的AI系统(human-aware AI systems),这项工作投射到规划领域,也就是我视为老家的这个社区,最直接的对应就是可解释规划(explainable planning)。几天前刚开过一个很好的研讨会,你们有些人可能也在场。我目前的研究概览写在刚刚发表的一篇《AI杂志》文章里,有兴趣可以去看。另外明天上午第一场有一个可解释规划的专场,我记得其中三个报告来自我们组,那会让你们了解我们现在在为什么兴奋。希望大家也去听。

拉奥: 以上两种报告我不用准备就能讲。所以我把这两个题目连同第三个题目「波兰尼对阵规划」一起发给阿尔瓦罗,跟他说:拜托,选前两个中的一个,第三个我得从头做。他作为一个好朋友,当然选了第三个,就是那个我一张幻灯片都没有的。于是过去一周我基本上是在ICAPS里晃来晃去,琢磨该跟你们说什么。今天这场报告,既不是教你怎么写论文,也不是讲我自己那套如何发ICAPS论文的具体研究,而更像是给你们一个声明式偏置(declarative bias),让你们思考当下这个时代的规划研究该怎么做。

拉奥: 所以这场报告有点费斯提沃斯的味道。年轻一些的人可能会想:费斯提沃斯是什么玩意儿?我告诉你们,费斯提沃斯是你们的传统,是你们身份的一部分。ICAPS从2005年开始搞费斯提沃斯。我还特意找出了当年我主持第一届费斯提沃斯时穿的那件库尔塔(印度长衫),发现现在还能穿得上。买一件特别大的库尔塔就有这个好处,无论后来胖成什么样都合身。

拉奥: 我还注意到,我这边现在是下午一点,欧洲大概是晚上九点,澳大利亚是早上七点。我知道你们很多人不在工作时间,所以我想建议一个饮酒游戏:每次听到「隐性」或者「显性」这两个词,就喝一口橙汁。这样一场听下来,你摄入的维生素C足够撑一整年。

Hinton的两条路

拉奥: 从这里开始。2019年,杰夫·辛顿(Geoffrey Hinton)来凤凰城做他的图灵奖演讲,他和杨立昆(Yann LeCun)都来了。我当时就在听众席上,因为就在凤凰城,我们都跑去会议中心朝圣。这是我拍的他第一张幻灯片的照片。他开场就说:让计算机做你想要的事,有两种办法。一种叫智能设计(intelligent design),另一种叫学习。

拉奥: 这个人真的会用词,恰到好处。他要是来ICAPS的费斯提沃斯,一定是天生的好手。世界上几乎任何东西前面加上「智能」二字都是好词,可他偏偏把它放在「设计」前面,立刻就变成了贬义。熟悉「智能设计论」的人都知道,那是一帮否认进化论的人的说法,科学界没有人愿意跟它扯上关系。

拉奥: 如果你看不清那张幻灯片,上面写的是:智能设计意味着有意识地弄清楚,要完成一项任务究竟该如何操纵各种表示,然后事无巨细地告诉计算机。他还用了「令人痛苦」之类的词,那是辛顿的风格。但说白了,这就是:如果你知道怎么做一件事,你就把这件事的模型告诉计算机,然后在上面加搜索算法。这正是我们做规划的人干的事。他说,这叫智能设计。任何一个稍有自尊的科学工作者都会担心自己被归到这一类里。

拉奥: 另一条路是学习:给计算机看大量的输入和对应的期望输出,让计算机自己学会从输入到输出的映射。这就相当于让进化在计算机里自行上演,AI就这么自己长出来。他就这样漂亮地开了场。

拉奥: 学习还是被告知,这对我们这个社区是个有意思的问题。我们知道,我们是给计算机提供PDDL模型的,后面还会细说。我们把自己对世界的了解告诉计算机,然后让它去处理组合爆炸。辛顿其实只是在强化当下AI的时代精神,只不过是以一种教条的、非常辛顿式的形式。

拉奥: AI技术之所以抓住了公众的想象力,很大程度上要归功于感知智能上的惊人成就,视觉、音频处理、语音识别等等,然后又把这些进展通过手机送到了街上每一个人手里。这就是AI变得家喻户晓的主要途径。但你要注意,这些进展绝大多数属于我所谓的隐性知识任务(tacit knowledge tasks)。所谓隐性知识任务,就是我们会做、却完全不知道自己是怎么做的那类任务。你要是想给「人如何看见物体」建一套理论,去问计算机视觉的人,他们试过,彻底失败了。我们不知道自己怎么处理语音,不知道自己怎么看世界。对这类任务,智能设计的路子当然行不通:你自己都不知道任务怎么做,却去告诉计算机怎么做,那肯定是错的做法。

拉奥: 但有意思的问题在于:隐性知识任务是全部吗?是唯一重要的东西吗?我们这个社区看这个问题的角度,跟世界上其他人,比如机器人或视觉社区的人,略有不同。我们清楚地知道,有很多任务是人类拥有相关技术知识的,人知道领域模型,我们的整个出发点就在这里,整个PDDL模型的事业就在这里。想到这一层,就该说说波兰尼悖论了。

波兰尼悖论:我们知道的多于能言说的

拉奥: 迈克尔·波兰尼(Michael Polanyi)是一位匈牙利博学者,他是哲学家,也做过许多别的事。他写过一本书叫《隐性维度》(The Tacit Dimension),提出了隐性知识的概念。维基百科上的定义基本就是我刚才说的:我们知道怎么做,却不知道自己是怎么做的,对做这件事的过程没有自觉意识。

拉奥: 波兰尼写这本书时是在感叹:我们太多的注意力都放在理解显性知识任务(explicit knowledge tasks)上,而不是隐性知识任务上。他说,我们知道的多于我们能说的,多于我们能用语言表达的,那为什么我们对这部分视而不见?这就是所谓的波兰尼悖论。

拉奥: 要理解显性知识任务的思维方式在我们心里扎得有多深,想想那句老话:想精通一件事,就去教它。这句话其实只对显性知识任务成立。同样,「想真正理解一件事,就把它编成程序」也是如此。你告诉我,有哪个写过基于Transformer的语言补全模型或者基于卷积网络的视觉模型的人,因此就更明白了我们是怎么看世界、怎么补全语句的?完全没有。事实上,我们一谈起「任务」,想到的从来都是显性知识任务。所以波兰尼担心的事很合理:各位,也请看看隐性知识任务吧。

人类与AI智能发展次序恰好相反

拉奥: 这张幻灯片是几年前我在ICAPS的一次报告里放过的。看看人类表现出各类智能的顺序:小孩子来到这个世界,先表现出感知与操作的智能,然后是情绪智能,再是社交与沟通智能,最后才开始显露认知与推理的智能。人类的幼儿出生时几乎已经具备了类似动物的智能,然后才逐渐长出认知推理能力,而后者是我们与人类文明联系在一起的东西。

拉奥: 有趣的是,如果你还没想过这一点:AI系统的发展次序恰恰相反。「深蓝」把国际象棋世界冠军卡斯帕罗夫打得落花流水的时候,它还远远认不出棋盘上的一枚小棋子,因为那时的AI系统根本不做视觉。但认知推理任务,AI早就在做了。做规划的人务必记住并理解这一点。那些在AlexNet之后才进入AI、以为一切都得从视觉出发然后不知怎么就能造出一个人的人不懂,但做规划的人是懂的。

拉奥: 这正是我总拿一句话开玩笑的原因:你给我找一个信心满满地宣称「未来AI不仅会学习,还将能够推理」的AI专家,我就能告诉你,这个人是在AlexNet之后靠看报纸标题进入AI的。因为AI在做感知这类隐性知识任务之前很久,就已经在做推理了。把这一点放在心里。

拉奥: 这其实解释了AI领域发生的很多事。对于那些我们多少有点自觉理论的智能侧面,给计算机编程要容易得多。这些理论未必万无一失,但至少我们知道怎么给各种规划领域写几个PDDL算子。可你试试给视觉任务写一个算子?你什么都写不出来,因为你根本不知道我们是怎么看的。所以对显性知识任务,推理与认知智能的进展快得多,因为我们可以先把至少是部分的任务模型告诉计算机,再用组合搜索之类的手段把它自动化。而对感知与操作这类智能,我们毫无自觉,只能让机器像我们自己那样去学:从观察、数据、示范、经验中学。

拉奥: 还要记住一点:如果你作为这个社会里的一个人,只从自己的原始经验中学习,从没见过别人,从没读过任何东西,从没被人告知过任何事,那你大概不会成为一个文明人,哪怕是最低限度的文明人。因为我们高度依赖显性知识任务的可传递性。当然,走路这种事你只能自己来,没人教过你走路,说话在最初阶段也没人教你。

拉奥: 我们也应当想到,隐性知识任务的学习之所以变得可行,实际上靠的是AI之外的一些正交发展,比如万维网。我们把集体潜意识整个上传到了网上,这就成了训练系统的数据。

推理导向与学习导向两种范式

拉奥: 还有一点值得注意:如果你上一门AI导论课,再上一门机器学习课,两者联系紧密,但你会发现AI这边,至少教科书里,往往是以推断为中心的。可以把设计智能体的路子大致分为两类:推断导向与学习导向。

拉奥: 推断导向是做规划的人一直以来的做法:假定模型是现成的,也就是面对显性知识任务;表示也是现成的,没有「学习表示」这回事,我们自己造出表示,把知识写进去,再设计算法在各种构型之间搜索、推理。所以它以推断为中心。我们也会对自己许一个小小的诺言,说总有一天我们要做学习,但往往迟迟做不到,至少在很长一段时间里没人真去做。于是推断始终在前台。对于有好模型的显性知识领域,这很好用。AI的发展在其历史的大部分时间里走的是这条路,规划文献也大多如此。我希望你们相信,这是件好事,因为做一个人,本来就意味着显性知识与隐性知识两者兼做。

拉奥: 学习导向则做出相反的假设。我想它最早的版本可以追溯到包容架构(subsumption architecture)之类的东西:假定智能体没有任何先验模型,专注于学习,连最基本的模型和表示都要学。你得到的往往是一个反射型智能体,然后对自己许诺,总有一天会开始做推理、做长期决策。所以这条路子倾向于推迟推断。对于没有好模型、但有大量样例和经验生成器的隐性知识领域,这很合理。这一边的研究近年来进展显著,尤其是2013年AlexNet之后。用吴恩达的话说,人类几秒钟内能做的事,计算机现在都能做了。

拉奥: 这很了不起,计算机能做人在几秒钟内做的一切。可如果人类只做几秒钟内能做的事,我们的总统就会是现在这样。人类会规划,会做长期推理,会做各种基于显性知识的推理。对人来说,两者必须结合。

拉奥: 我已经说过,为什么恰恰是隐性知识任务上的进步抓住了公众的想象力。手机能识别你的声音、能补全你的句子,这些能力非常有用,人们喜欢用。但这不是智能或智能行为的全部。

系统一系统二与显性知识的编译

拉奥: 说到这里,有些人一定在想:拉奥在讲显性知识与隐性知识,这不就是系统一与系统二吗?卡尼曼和特沃斯基提出的那套。顺便说一句,系统一、系统二只是理论,你的大脑里没有哪块叫系统一、哪块叫系统二,那只是描述人脑能力的一个比喻。系统一负责反射式推理,系统二负责深思熟虑的推理。

拉奥: 但是系统一、系统二与显性、隐性之间有一个差别,往下讲之前我希望你们弄清楚。大多数隐性知识任务确实由系统一处理。可显性知识任务可以从系统二起步,你一开始是有意识地、深思熟虑地做,后来为了效率,它被编译进系统一,变成反射行为。有人说过,文明的进步,在于我们能不假思索地做的事越来越多。你最初学微分的时候,是在算极限,f(x+h)减f(x)再除以h,h趋于零。可要是你一直这么算,你的微分能力就远远落后于时代了。所以你把它编译下去,开始反射式地做。

拉奥: 也就是说,被编译进系统一的显性知识任务,与本来就待在系统一的隐性知识任务,虽然都在系统一,仍然不一样。这就像一段直接手写的汇编程序,与一段从高级语言编译出来的汇编程序之间的区别。后一种情况下,汇编程序出了错,你可以把错误定位回高级语言的程序里,符号调试器就是这么工作的,我们就是这样调试程序的。这就是显性知识模型带来的可解释性。

拉奥: 我忍不住要给你们看一样东西,它印出来的时候,博士生论坛里的大多数人肯定还没出生。1989年,《AI杂志》上有一场席卷AI界和规划界的大论战,关于「通用规划」(universal planning)。马特·金斯伯格(Matt Ginsberg)的文章标题叫《通用规划:一个几乎普遍糟糕的主意》。事情是这样:马塞尔·肖珀斯(Marcel Schoppers)说,为什么不从一个类似PDDL的模型出发(那时候还没有PDDL),把它编译成一个反射式策略?也就是预先解决大量问题,把解记住,差不多就是记忆化。金斯伯格写了好多页指出,组合爆炸决定了你要记住的那张表会大得吓人,所以这是个普遍糟糕的主意,有时候你就是得当场思考。

拉奥: 如今时代变了。我们不再那么担心空间与在线计算之间的权衡,事实上我们很乐意用空间换在线计算时间,你可以把策略存成几个TB,只要能瞬间回答问题就行。但这个权衡,跟你最初有没有显性知识,是两码事。往下讲之前,请把这一点想清楚。

波兰尼的复仇:AI迷恋隐性知识

拉奥: 刚才我们讲的是波兰尼悖论。但你也看得出来,AI领域实际上正在上演一场「波兰尼的复仇」。AI现在与隐性知识展开了一场全新的恋情。阿基米德发现杠杆原理之后,得意忘形地说:给我一根杠杆和一个支点,我就能撬动地球。我们现在的口气差不多:给我一大堆GPU、一个足够大的数据集、一个足够深的网络,我就能给你造出通用人工智能。不只是人类水平的AI,直接就是AGI。

拉奥: 这让我很不安,部分原因是我见过AI的两面,我认为做一个人就意味着显性知识任务与隐性知识任务两者都要做。所以我写了一篇观点文章,题为《波兰尼的复仇与AI对隐性知识的新恋情》,明年二月会登在《ACM通讯》上,我的主页上也有。那篇文章里有我刚才展示的这些铺垫,还有几个额外的论点,我先讲完它们,再从那篇文章转到它对规划的直接影响。这一点我已经零星提到,但我想更直接地谈谈这一切对ICAPS的规划研究者究竟意味着什么。

数据与教条之争:可解释性与数据偏见

拉奥: 这是你们所处世界里的一种张力:数据对阵教条。波兰尼当年担心的是,我们只研究有显性知识、有教条的问题,他希望大家也研究隐性知识。可现在我们几乎走到了另一个极端:AI基本上只做隐性知识任务。甚至流行起这样一种做法:拿一个明明有显性知识模型的问题,比如数独,把它转换成海量的样例,只为了能说一句「我做了深度学习」。

拉奥: 也就是说,别提从数据到知识了,现在有一整个家庭作坊式的产业,专门把显性知识转成数据,然后喂给某个Transformer之类的东西,再设法把你一开始就有的知识恢复出来,然后写成一篇论文。听起来离奇又荒诞,但确实有人在做,你们很多人知道,你们中有些人可能就在做。

拉奥: 那么真正有意思的问题是:当我们确实有教条、并且希望系统遵循它的时候,该怎么办?这正是规划领域一直在研究的问题:在那些人们乐于并愿意提供一定领域知识的领域,你难道要把他们推开,说「别跟我讲,我只从行为里学」?在我看来这简直愚蠢。人类智能的一个标志,就是隐性知识与显性知识之间无缝的交互。而我们的钟摆,以一种奇怪的方式,从「所有模型都是错的,有些是有用的」,摆到了「模型是什么?我们要模型干什么?直接从数据到决策就行了」。这显然值得我们认真想一想。

拉奥: 在那篇面向更广泛读者的观点文章里,我把波兰尼的复仇与当下困扰AI的一系列问题联系在一起,我在别的文章里也分别写过,可以去我的主页找。举个例子,可解释性问题。一个自己学习表示的系统,学到的东西不一定对你有意义,这一点根本不应该让人吃惊。而如果你从PDDL模型之类的东西出发,至少你放进去的东西你自己是懂的,可解释性问题就简单得多。类似地,如果你不太清楚系统学到了什么,它对对抗攻击的脆弱性也会很高,因为你无法保证它会按预期工作。

拉奥: 最后一点,也许更有意思,是对数据集偏见的脆弱性。右边这两个例子:给GPT-3输入「两个穆斯林」,不管你怎么做,它给出的许多补全都是穆斯林卷入了某种坏事,杀了人,或者被杀,或者干了别的什么。在一个文明世界里,这是完全说不通的。但这也不奇怪,因为GPT-3是从我们的集体潜意识里学出来的。我们确实会有疯狂的念头,有些念头我们不会说出口,除非你行事冲动、情绪控制和冲动控制很差。正常情况下,系统一生成的东西,系统二可以拦住不让说出来,这里有一道控制。这正是对隐性知识的迷恋可能把我们带进麻烦的一个更大层面的例子。

何时学习、何时接受人类知识的智慧

拉奥: 问题于是变成:什么时候该从样例中学,什么时候该从人那里拿知识?两个方向上你都可能自欺。规划社区,我马上会给你们看,曾经指望有一种超人类的人类,能够告诉规划器搜索树上第7539号节点该不该剪掉。这就是我们当年在所谓「混合主动规划」(mixed-initiative planning)里做的事。我们指望得太多了,我们指望的是一些根本没有自己生活的人。而另一边则说,人类没什么可教我们的,我们只从数据里学。所以,弄清楚何时该学、何时该向人要知识,需要一点智慧。

拉奥: 于是我给基督教的那段宁静祷文编了一个自己的版本,主角是这个紫色的机器人:人类啊,请赐我宁静,去接受我学不会的东西;赐我数据,去学习我学得会的东西;赐我智慧,去分辨两者的不同。这场报告的一部分目的,就是试着给你们一点这种智慧。你们有些人已经具备了,有些人则通过自己写的论文表明还不具备,我稍后会点名。粗略地说:显性知识唾手可得的地方,我们就该用它。这又会引出有趣的问题:如何把它与隐性知识衔接起来?这正是我们后面要谈的。

规划领域一贯依赖显性模型

拉奥: 现在来看规划这一边,看这一切怎么投射到规划上。ICAPS风格的、主流的自动规划,绝大部分时间关注的都是显性知识任务。我们的核心信念一直是:在很多领域,人们对任务拥有可以言说的显性知识。所以我们专注于模型描述语言,STRIPS、PDDL、SHOP、RDDL等等,让人们更容易写下他们想表达的知识,然后开发高效的通用规划器来处理这些模型。

拉奥: 显性知识在规划里无处不在,我们必须承认这一点。在ICAPS,你不能说显性知识不重要,否则我们做的大部分工作算什么?规划社区一直乐于接受人类设计者轻易就能给出的显性知识,我们甚至支持过更加知识密集的路子,比如向人索取控制信息和抽象信息。正如我说的,我们有时指望的是没别的事可做、专门给我们提供模型的超人类。PDDL这样的显性模型还算容易,可我们有时还指望人给控制信息、控制规则之类的东西。

拉奥: 上一次我为显性知识这么激动,还是在2003年的ICAPS,就是那张幻灯片,我当时在痛陈基于知识的规划:我们究竟是在比较法希姆和达纳(指TLPlan与SHOP两个系统的作者),还是在比较两个规划算法各自搭配的知识库?这就变得很有意思。我们有过这些争议,但对使用显性知识,我们从来是开放的。

拉奥: 然而最近出现了一些颇为离奇的尝试,想搭上「从像素到决策」的浪潮,哪怕是在显性知识唾手可得的领域。我不知道这是想让谁刮目相看,但这类工作确实在做,AI领域有,规划社区也有。这就是波兰尼带着他的复仇进入规划的方式。而波兰尼进入规划的通道,就是强化学习(reinforcement learning)。

强化学习把波兰尼带进规划

拉奥: 让我用两张幻灯片讲讲规划与强化学习,尤其是深度强化学习之间的联系。这几张小图是我从自己在PRL研讨会(规划与强化学习研讨会)的报告里剪来的。如果你把规划看成从模型到策略的问题,那么强化学习就是直接从经验到策略。它不考虑模型,只想从经验中学出策略。规划则是给定模型,推断出策略。

拉奥: 强化学习有一个版本叫基于模型的强化学习(model-based RL),它有一个中间步骤,就是模型。原味的做法是从纯经验中学出这个模型,然后用它做推断得到规划,规划就成为策略。无模型强化学习(model-free RL)则连这个中间步骤都不要。

拉奥: 对基于模型的强化学习,我希望你们这样想:普通的学习可以直接从数据到分类决策,但大多数学习是从数据到假设、再到分类决策。假设这个中间环节,理论上并非必需,却给了我们一个很好的入口,把关于这个世界的有意思的偏置注入进去。同理,正如可以给假设加偏置,我们也可以给学到的模型加偏置。原味的基于模型的强化学习不从人类那里拿任何东西,这一点稍后再谈。但我要说,我们应当考虑把PDDL模型当作初始化,比如把部分模型作为初始化,再通过经验加以改进。这是把专家能给你的部分知识与经验结合起来的一种远为有用的方式。

拉奥: 说到强化学习与规划,我曾经说过:规划与强化学习是AI的两个核心领域,只被一个共同的问题隔开。这套话原本是说美国和英国的:两个伟大的文明,只被一门共同的语言隔开。我认为规划与强化学习确实有很多共同之处,但我们彼此交流不多,部分原因是,虽然问题相同,但各自单独发展时关注的任务差异很大。规划的人,也就是ICAPS的人,倾向于研究显性知识的规划任务;强化学习的人则主要研究隐性知识的规划任务,倒立摆平衡、抓取、操作。深度强化学习一些最漂亮的成就正是在机器人操作上,因为那些任务我们不知道怎么写出简单的显性知识模式,于是他们借助模拟器去学。模拟器的事我们马上也会谈到。

拉奥: 波兰尼通过深度强化学习进入规划,有两条路。第一条是:既然可以直接用无模型的方法学出策略,何必还要规划这一步?无模型了,规划就完全不需要了,没有中间步骤。第二条是:既然可以从经验中学出模型,而且是用你自己造的表示,也就是整个表示学习那一套,何必还要接受显性知识?这通常导致学出无法解读的模型。

拉奥: 对隐性知识任务,这两种立场当然都合理。希望你们在喝橙汁。但对于同时涉及显性知识和隐性知识的复杂任务,这两种立场就相当离奇了,无论是从解的鲁棒性看,还是从解的可解释性以及AI智能体运行过程的可解释性看。这一点要记在心里。波兰尼的复仇就是这样找上门来的。

模拟器其实也是人给的知识

拉奥: 这里我想岔开说一件事。出于某种奇怪的原因,强化学习的人,尽管从定义上讲强化学习没有任何地方规定不能从外部接受知识,毕竟模型在学习开始前完全可以用部分知识去偏置,但随着时间推移,强化学习社区基本上决定了他们不喜欢接受知识。可他们也知道,不可能真的只从经验中学。从经验中学开车,意味着车子要掉下悬崖,然后我从你的经验中学习,而不是你,因为你已经死了。既然在真实世界里赤裸裸地获取经验对机器人的健康相当有害,很多强化学习系统实际上是在外部提供的模拟器上工作的,但对接受任何显性知识依旧嗤之以鼻。

拉奥: 这有点好笑:模拟器是人造的,我们不介意接受模拟器,却不喜欢人直接给我们知识。这差不多就是理查德·萨顿(Rich Sutton)那篇《苦涩的教训》可以被解读出来的意思。而对显性知识领域来说,模拟器比部分领域知识更容易提供,这种可能性极小。你要是让我给积木世界写个模拟器,我大概会先写一个积木世界的PDDL模型,再往上加一个动作评估器,除非我做的是远远超出我们所说的积木世界的东西。总的来说,关于模拟器这个问题,请认识到:模拟器正是人类向强化学习提供极有用知识的一种方式。这一点要记住。事实上,莱斯利(指莱斯利·凯尔布林,Leslie Kaelbling)几天后的报告里也会讲到类似的内容。

拉奥: 深度强化学习对这些扭来扭去的小虫子非常在行,这是我从OpenAI那里直接复制来的图,它们能学会运动,学得相当好,而这种东西你没法写PDDL模型。可另一边,史蒂夫·钱(Steve Chien)和喷气推进实验室的人想谈任务规划与任务调度,人们想谈钻探、怎么制定钻探计划。世界上有大量的规划任务,没有简单的、遍历性的模拟器可以让你反复试错、慢慢摸索。于是问题就成了:我们怎样把这一类隐性知识任务与那一类显性知识任务衔接起来?而在两者之间,还有像任务与运动规划(task and motion planning)这样的任务,两方面的成分都有。这才是我们真正该谈的有意思的东西。

拉奥: 遇到这类情况,自底向上一路往上走不通,自顶向下一路往下也不通。我们必须考虑如何把两个方向的长处结合起来。这就是我想在剩下几分钟里传达的。

应对之道一:加入这场浪漫(不推荐)

拉奥: 怎样在AI对隐性知识的新迷恋周围做规划?第一个办法:加入这场浪漫。你们很多人在这么做,好吧,你们中有些人在这么做。毕竟,任何显性知识任务只要稍加努力,都能转换成隐性知识任务。比如,你可以把PDDL模型拍成照片,它就变成了图像,然后用卷积网络、图神经网络、Transformer,甚至昨天刚出来的Performer去分析它,把模型恢复出来。当然,Transformer和Performer对序列数据更有用,所以你可以用PDDL模型加FF规划器生成一大堆规划轨迹,把这些轨迹交给某个序列学习算法,随便哪个你喜欢的。或者跳上「从像素到决策」的花车,给积木世界各种构型之间的转换拍照,看能不能从中恢复出某种积木世界模型,再喂给FF,然后开开心心地拍拍自己的肩膀。

拉奥: 这些全都有论文。我不想具体点名,但你们很多人大概知道是哪些。私信客气地问我,我会把引用发给你。至于发表这些论文的人,我只说一句:我不是这个方向的拥趸。因为从某种意义上说,你是在自底向上,拿一个本来有显性知识的东西,硬当作隐性知识任务去解。当然能做到,可好处在哪里?它反而带来了可解释性、鲁棒性等一大堆别的问题。

应对之道二:结合两者的具体研究方向

拉奥: 另一个办法是我推荐的,这是一张密密麻麻的幻灯片,之后我会挑几点展开,希望你们中有些人会考虑。能做的好事有很多很多,这些只是我这个小脑袋想到的。你真正想要的是结合两者,让整体大于部分之和。

拉奥: 比如,研究那些同时包含显性知识成分和隐性知识成分的任务。运动与任务规划一直是个很好的领域,我很想看到它们如何从中受益。莱斯利和她的团队做的那类工作我很喜欢,他们是真的认真在把这两者结合起来。迈克尔·利特曼(Michael Littman)前几天有一个很好的报告,也是围绕这些问题。

拉奥: 在这些方向上,可以研究用学习来降低模型获取的难度。我们一直在谈这个,改进模型获取有好办法也有坏办法,稍后我会多说几句。至少在今天,你不能再说学习没有解决,所以我们不能对这个问题视而不见,这绝对值得做。

拉奥: 研究部分指定的、不完整模型下的规划。这与模型获取是相关的问题:一旦你承认手上的模型是不完整的,你就必须谈鲁棒规划,因为它不再是「对着这个已知正确的模型求一个最优规划」的局面了。

拉奥: 研究可解释性问题,尤其是系统同时具有显性知识和隐性知识时。我很喜欢可解释规划(XAIP)这个社区,他们在可解释性与解释方面做得很好。即便在有共享词汇的情况下,解释和可解释行为方面也有大量工作要做,而没有共享词汇时就更难。已经有一些有意思的方向在探索,我稍后会讲一点,这些方向绝对值得看。

拉奥: 还有一些正在进行的工作,我认为很好,比如在学习中编译控制知识。我们过去总是用干净的原理推导启发式,但其实你可以从规划器的运行轨迹和经验中学习启发式,这容易做,也值得做。当然你会失去最优性保证和信息量保证,但说实话,生活本来就没有这些保证,所以这些东西是有用的。很多人在做,本届会议上就有这方面的论文,这是个很好的方向。

拉奥: 最后,把规划技术推广到非声明式表示的场合,尤其是那些你只有一个模拟器、没人愿意给你写领域模型的地方。已经有人开始做了。我认为这些都是值得推荐、值得研究的方向。接下来我快速讲其中三个,然后收尾。

用学习获取与精化模型、声明式偏置

拉奥: 第一个是用学习来获取和精化模型。从零开始、只凭经验学模型,不太可能得到鲁棒且可解释的模型。迈克尔·利特曼有这么一张图,我要说的是,我们应该允许在学习阶段注入声明式偏置,也就是部分模型。机器学习的人经常谈归纳偏置,但他们往往还停留在把网络拓扑当作归纳偏置的阶段,这没什么不好,但作为「告诉你我想让你学什么」的手段,实在太原始了。我应该能给你背景知识,你再通过经验去改进它。所以更好的做法是,以部分模型的形式给模型学习器提供声明式偏置,然后在学习中精化。这些不是已经做完的工作,我只是说这些方向你们可以考虑。

拉奥: 这样一来,你可以从人那里拿到一个还过得去的模型作为起点,再从经验中逐步改进。它也保留了可解释性,因为你始终没有离开人给你的那个起点。从写给人看的文本中学习模型,也是一个极为多产的方向,这是我们与当下自然语言处理各种进展对接的好途径。很久以前就有人做过这件事,但那时候的NLP技术很烂,现在好多了。你应该能读菜谱,从中得到某种部分模型,然后把它作为模型学习的起点。我们在IJCAI 2018上有这方面的工作,可以去看,也有其他人在做这类工作,很有用。

拉奥: 有一件事我不建议你做:从FF的轨迹里恢复PDDL模型。你已经有PDDL模型了,为什么还要用FF生成一大堆规划,只为了把这些轨迹再反演回模型?你当然可以说,因为这样能在ICAPS的规划与学习分会上发一篇论文。但这不是个好理由。这种事很离奇,不值得做。

不完整模型下的鲁棒规划与多模型贝叶斯

拉奥: 不完整模型下的规划在这里非常相关。一旦规划器意识到它手上的模型在任何实际意义上都不能保证正确,它就必须开始考虑鲁棒性,必须把自己对模型的无知考虑进去。这就引出了不完整模型的表示问题。我们组一直在研究的,不只是STRIPS那种完全因果的、由用户担保正确的模型,而是越来越浅、越来越不能保证正确和完整的模型,然后讨论如何对着这样的模型做鲁棒规划。所以你需要思考这类模型的表示是什么,以及在这些表示上做鲁棒规划意味着什么。

拉奥: 我们组的一些例子包括:带有「可能前提条件」和「可能效果」的领域模型,2017年在《人工智能杂志》上发过一篇;还有奖励度量不确定的规划问题,如今很流行的多样化规划(diverse plans)的许多工作,正是从这里来的,那是较早的工作之一。这些我认为都相当切题。

拉奥: 一旦开始考虑多个模型,模型不确定性这个问题就可以从贝叶斯的角度看:当你对模型不确定,就意味着你实际上有许多完整的模型,其中一个是真实模型,于是你是在做贝叶斯意义上的规划。这是做鲁棒规划的一般方法。

拉奥: 而对于我更关心的人类感知规划(human-aware planning),这个问题更加突出。机器人有一个模型M_R,它本身可能就不完整,所以实际上是一组模型。而人类心中有一个关于机器人模型的近似,机器人要去估计它,所以这个估计同样是一个模型分布。机器人就不得不做多模型规划,才能判断自己的行为是否可解释。所以从这些角度看,这个问题同样非常相关。萨拉特(指萨拉特·斯里达兰,Sarath Sreedharan)今年在XAIP研讨会上有一篇很好的论文,讲可解释性度量的贝叶斯刻画,可以看看。

XAI不只是指点式解释:共享词汇与不可解读模型

拉奥: 我还应该提一下,今天早上博士生论坛上有人问一位做序贯决策可解释性、做得很好的同学:我不明白这为什么和可解释AI有关。我不太清楚他们为什么这么问,但有些人倾向于认为,可解释AI就等于对不可解读表示的可解释机器学习。我希望你们认识到,那只是整个光谱中很小的一部分。可解释AI难,但主要是作为不可解读表示的调试工具而难。指点式解释是相当原始的。如果指一指是我们彼此交流的唯一方式,我们不会有今天的文明。人与人之间的解释对协作至关重要,但它们不是指点式的,也不是智能体单方面的独白。所以可解释规划社区在做的那类工作,在整个可解释性方向上是非常切题的,这一点请记住。

拉奥: 我再半开玩笑地问一句:拿那个对抗样本来说,一辆校车加上一点噪声,就被当下大多数深度学习视觉系统认成了鸵鸟。你能问视觉系统「告诉我,第二辆校车的哪个部分让你觉得它是鸵鸟」吗?就算能,这又能有多大用处?指点式解释对动物来说够用了,但我们是人,我们真的得超越这一步。请把这一点记在心里。

拉奥: 收尾之前最后一点:处理不同的词汇很重要。可解释规划的大部分工作是在两个PDDL模型之间做的,两者可能有不同的前提条件和效果。最终我们应当把它推广到这样的情形:机器所用模型的一部分是不可解读的,不是PDDL式的模型,但机器愿意把它的解释翻译成你能理解的语言。可解释机器学习社区有一些有意思的工作,我们组也在做一些。有一篇论文,萨拉特参与的,大意是:即便你在用黑箱模型做推理,你也用人类理解的概念来提供解释,而机器理解这些概念的方式,是在它那套不可解读的表示之上学习这些概念的指称。这是可行的,也值得做。

总结:在天启之前造出可解释的规划系统

拉奥: 最后两张幻灯片。杰夫·辛顿和理查德·萨顿都在说,智能设计很粗糙,我们应该等着学习自己发生。有人会合理地问:我们来到这个世界,从阿米巴的智能进化到动物的智能,再进化到人类的智能,一路上还发展出了一个相当漂亮的系统二,而且我们有时候甚至能让彼此听懂。既然如此,我为什么还需要任何显性知识任务?也许我可以全部自底向上地做出来?我不一定要论证这做不到。但它可能要到天启那天才做得成,而我希望AI系统在那之前就能同时使用隐性模型和显性模型。

拉奥: 想想工业界,他们希望这些系统现在就能工作,而不是说:我要让那条扭来扭去的小虫子去做NASA的任务规划,看看会怎么样。即便算力再快,这要花多久也完全不清楚。而且整个问题还在于,等它们真做到了,它们的行为能不能解释。所以我们真的希望在天启之前,早一点拥有具备规划能力的可解释AI系统。

拉奥: 总结一下。指望为隐性知识任务给出显性模型,与拒绝显性知识任务上唾手可得的部分模型,同样愚蠢。你要是去问人「能不能给我一个PDDL模型,描述你是怎么看见物体的」,那跟有人说「让我只从像素转换里弄明白积木世界是怎么回事」一样愚蠢。因为明明有人能告诉你这件事是怎么做的,你完全可以把两者结合起来。在AI当前与隐性知识的浪漫周围做规划,有许多富有成果的方向,而不必无缘无故地往任务规划里加十二克Performer或者Transformer网络。我完全支持把这些技术结合起来,但我不喜欢只为了发论文而这么做。

拉奥: 这也是我为什么要放我学生们这张幻灯片。这四位得忍受我的长篇大论,每当他们想说「哇,我们能不能就为了好玩,往里塞几个图神经网络算法」,我就会说:告诉我为什么,告诉我这到底有什么意义。为此,也为了他们至少愿意让我这样安排他们,还为了我在他们身上试过这里的很多论点,我要感谢他们。当然还有那个紫色的机器人,希望它能从我的学生、从我、从我们其他人身上学会那段祷文:赐它宁静,去接受它学不会的东西;赐它数据,去学习它学得会的东西;赐它智慧,去分辨两者的不同。

拉奥: 我就以这张总结幻灯片结束:一条路是加入这场浪漫,我不太喜欢;另一条路是结合两者,让整体大于部分之和,我认为这条路更有意思。谢谢大家。

本期讲者
苏巴拉奥·坎帕蒂亚利桑那州立大学计算机教授,Yochan 实验室负责人,研究人机协作 AI 与可解释规划。2016–2018 年任 AAAI 主席,ICAPS Festivus 传统的创始人。
主持人 AlvaroICAPS 2020 博士生论坛(Doctoral Consortium)的组织者,负责介绍特邀报告人并说明会议录制安排。
章节 · 点击跳转视频
0:00 主持人介绍与开场自嘲 ▶ 正在看
5:12 Festivus 与 Hinton 的两条路 ▶ 正在看
9:38 隐性知识任务与波兰尼悖论 ▶ 正在看
13:56 AI 发展顺序:推理先于感知 ▶ 正在看
17:07 推理中心与学习中心两种路线 ▶ 正在看
20:03 系统一二与显性知识的编译 ▶ 正在看
24:14 波兰尼的复仇:数据压倒教条 ▶ 正在看
31:29 规划社区的显性知识传统 ▶ 正在看
33:42 强化学习把波兰尼带进规划 ▶ 正在看
41:00 两条出路:加入浪漫或整合 ▶ 正在看
45:54 模型获取、不完整模型与解释 ▶ 正在看
53:12 总结:宁静祷文与整体大于部分 ▶ 正在看
本期论点
本期回应
10:30
在隐性知识任务上,手工编码规则注定失败,因为人自己也说不清这件事该怎么做 靠学习智能主要靠什么长出来?
16:16
人若只靠自身原始经验学习、从未被他人告知任何事,很可能算不上文明人 要合起来智能主要靠什么长出来?
19:55
人类会做规划和长期的显性知识推理,所以推理路线与学习路线必须结合起来 要合起来智能主要靠什么长出来?
54:27
纯自下而上的学习或许原则上能达成通用智能,但等待时间太长,满足不了产业界当下的需求 要合起来智能主要靠什么长出来?
55:18
对隐性知识任务强求显式模型,与在显性知识任务上拒绝现成的部分模型,是同样愚蠢的错误 要合起来智能主要靠什么长出来?
其他论点
15:17
人类拥有有意识理论的那部分智能更容易编程,这解释了AI各子领域进展速度的差异
22:20
从显性知识编译进系统一的技能与天生的隐性技能有本质差别,前者出错时可回到高层模型定位
26:12
当前AI把本有显性模型的问题转成海量样例,只为宣称自己用了深度学习
28:05
自学表示的系统学到的表示没理由对人有意义,可解释性因而远难于从显式模型出发
39:01
强化学习拒绝人提供的显式知识、却坦然接受人手写的模拟器,是自相矛盾的
44:16
一旦承认模型是不完整的,规划目标就必须从求最优计划转为鲁棒规划 做法
46:55
应以部分模型的形式给学习器提供陈述式偏置并在学习中精化,而不是只靠网络拓扑做归纳偏置 做法
51:15
把可解释AI等同于解释难解的机器学习表示,只覆盖了可解释性谱系中很小一部分
01主持人介绍与开场自嘲
0:00
for the greater moderation of the funnel and next we will have our invited talk and it's my pleasure to introduce to you to subara kampanpatti or josh rao probably as uh most of you know him um he works in arizona state university where he leads the highly successful group of yochan um about human robot interaction and explainable ai and more things like it's a really highly uh diverse research lab that you have there and well i could speak hours about all the achievements of rao he has been president of triple a he has won multitudes of research and teaching awards i really recommend you going to to his webpage because uh he has a really funny uh webpage and uh youtube channel which i recommend you to to check out and um in general i think that drought doesn't need uh much introduction here at icubs i'm pretty sure all of you have already heard of him i want to highlight that he was the one starting the festivus tradition at icubs so hopefully we will get some of that uh today in the talk uh
为了让整个流程更顺畅,接下来是我们的特邀报告,我很荣幸向大家介绍Subbarao Kambhampati,或者大多数人熟悉的 Rao,他任教于亚利桑那州立大学,在那里他领导着非常成功的 Yochan 研究组,研究人机交互、可解释 AI 以及更多方向,那是一个研究方向非常多元的实验室。说真的,我可以花上几个小时来讲Rao 的各种成就,他担任过 AAAI 主席,获得过无数的研究和教学奖项。我强烈建议大家去看看他的主页,因为他的主页真的很有意思,还有他的 YouTube 频道,也推荐大家去看看。总的来说,我觉得 Rao 在 ICAPS 这里其实不太需要介绍,我相信在座各位都已经听说过他了。我想特别提一下,是他在 ICAPS 开创了 Festivus 这个传统,所以希望今天的报告里也能有一点那种味道,我很期待。最后我想这样结束我的介绍:我觉得他大概是我们这个领域里
便签引用
1:27
i'm looking forward to it and uh i i would uh conclude my introduction saying that uh i think is the closest that we have uh to a rockstar in uh so he has even his own imdb page so that's something worth mentioning so um before i give the the speak to to raw i would also tell you that we will be recording this this session uh so just uh so that you know if you participate uh well we are recording this so with that uh rao i think okay so you can see my shared screen right yes okay good hello everybody i really can't see you it's a strange way to give talks but i'm doing my best to get psyched up about this um as uh so we i'm going to be talking about this reasonably long mouthful of a an idea called polany versus planning planning around ai's new romance with tacit knowledge um so let's see yeah um so actually this is not the usual normal style um dc doctoral consortium um mentoring talk so i did give one of those a while back um before probably some of you were born in 2009 i had a lot more hair and lot less fat
最接近摇滚明星的人了,他甚至有自己的 IMDb 页面,这一点很值得一提。那么在把话筒交给 Rao 之前,我还要告诉大家,我们会录制这场报告,所以如果你参与互动的话,请知悉我们正在录制。那么 Rao,交给你了。好的,你们能看到我共享的屏幕吧?能看到,好的。大家好,我其实完全看不到你们,这种演讲方式挺奇怪的,不过我会尽力让自己进入状态。那么,我今天要讲的是一个名字有点长的题目,叫做「Polanyi 对阵 Planning:在 AI 与隐性知识的新恋情中做规划」。我们来看看。其实呢,这并不是通常那种博士生论坛(DC)的指导性报告。我以前确实做过一次那样的报告,那是很久以前了,2009 年,可能你们有些人那会儿还没出生。那时候我头发多得多,也瘦得多。
便签引用
3:08
at that point of time but you know and that thing about how to give talks and write papers i wound up doing that once again um in each guy a few years later in 2013 that version is actually available on youtube so if you're interested in in one person's view of how to write papers and give talks um you know at least my view you know you can try checking it out uh especially if this stock doesn't bomb because if it bombs you know at least what not to do presumably uh so that's about uh you know something that you know said more effort mentoring talk but i'm not going to give that um now it would look as if i might be giving an actual uh research talk on my research it turns out that's also not true as you know alvaro pointed out i actually give i do work in uh human very eye systems and the closest projection of that into this community into planning of course which i call consider myself a consider my home community is explainable planning uh there's a great workshop that happened um you know a couple of days back and
就是关于怎么做报告、怎么写论文的那个内容,几年后我在 2013 年的 IJCAI 上又讲了一遍,那个版本其实在 YouTube 上能找到。所以如果你想看看某个人对怎么写论文、怎么做报告的看法,至少是我个人的看法,你可以去看一下,尤其是如果今天这个报告没搞砸的话;因为如果搞砸了,那至少能让你知道什么是不该做的。总之那算是一个更偏指导性质的报告,但我今天不打算讲那个。那么现在看起来,我也许会做一个真正意义上的、关于我研究工作的学术报告——结果这也不对。正如 Alvaro 提到的,我其实做的是人本感知 AI(human-aware AI)系统方面的工作,而它投射到这个社区、投射到规划领域最接近的部分——当然,我一直把规划当成自己的老家社区——就是可解释规划。前几天有一个很棒的相关研讨会,也许你们有些人也在场。我目前研究的整体概览
便签引用
4:14
maybe some of you were there too uh so my current overview of my research is in this ai magazine article that just came out you might look at it and if you are interested apparently we rented a session tomorrow am1 unexplainable planning a bunch of the talks there i believe three of the talks there are from our group so that would give you an idea of what we are currently excited about so i hope you will come to those sessions too okay so that's about what this talk is not about um either of these talks i could have given without much of a preparation and so i asked alvaro to these two ideas and this other idea about poland versus planning i told him look you know please say one of the other two because this one i have to make the talk and as a good friend of course he said go with the third one the one for which you don't have any slides so i spent i guess last one week essentially sort of uh pottering around icaps thinking about what i'm going to tell you so that's what i'm going to do and so it
写在刚发表的这篇 AI Magazine 文章里,你们可以看看。如果感兴趣的话,明天好像有一个AM1 的可解释规划专场,我记得那里面有三个报告来自我们组,这能让你们了解我们现在在兴奋什么,所以希望你们也来听那几场。好,这就是这个报告不会讲的内容。这两种报告我不用怎么准备就能讲,所以我跟 Alvaro 提了这两个想法,还有另一个想法,就是关于 Polanyi 对阵 Planning 的。我跟他说,拜托你选前两个中的一个吧,因为第三个我得从头做幻灯片。结果作为一个好朋友,他当然说:就选第三个吧,那个你一张幻灯片都还没有的。于是我大概花了过去一周时间基本上就是在 ICAPS 会场里晃来晃去,一边想着我要跟大家讲些什么,所以我就打算这么干,于是
便签引用
02Festivus 与 Hinton 的两条路
5:12
turns out that this talk is kind of not mentoring about how to write papers kind of a talk it's also not my specific research you know to get how you know i caps papers kind of a thing but it's more of a sort of a declarative bias to get you to think about planning research in the current age uh so that's what i'll try to do so it's sort of a bit of a festivus like and you know for those of you younger ones who think what is this festivals nonsense you know festivals is your heritage it's part of who you are you know icaps is basically uh started these festivals back in 2005 it turns out that i went and found this kurta that i wore in the original festivals when i actually ran the show and so i still fit in it i guess that's one of the advantages of buying really large kurta so you can always fit even after you get quite fat um also i noticed that i am talking at 1 pm in the afternoon here for me but it's probably like 9 o'clock in europe and and and also maybe 7 o'clock in australia nine in the night evening for europe and
结果这个报告呢,既不是那种教你怎么写论文的辅导型报告,也不是讲我自己具体的研究,你懂的,不是那种教你怎么发 ICAPS 论文之类的东西,它更像是一种声明式的偏置,想让你去思考在当下这个时代该怎么规划研究,所以这就是我要试着做的事。所以它有点像 Festivus(吐槽大会)那种感觉,你们当中比较年轻的可能会想,这个 Festivus 是什么鬼东西?要知道 Festivus 是你们的传统,是你们身份的一部分,ICAPS其实早在 2005 年就办起了这个 Festivus。说来也巧,我翻出了当年我主持那一届 Festivus 时穿的这件 kurta(印度长衫),我现在还能穿得下。我想这大概就是买超大号 kurta 的好处之一吧,这样你永远都穿得下,哪怕你后来胖了不少。另外我注意到,我这边现在是下午一点,但在欧洲大概是晚上九点,在澳大利亚大概是早上七点——欧洲是晚上九点,澳大利亚是早上七点,
便签引用
6:22
seven in the morning for australia so to encoura i you know i know that you're not in the work day time so i thought i should suggest a nice um drinking game too so here is an intriguing game for you every time you hear the word tacit are explicit take a swig of orange juice and then you would have enough c vitamin to last an entire year so that's for sure okay so let's get going now so let me start with this okay so last year in 2019 um jeff hinton of course came to phoenix to take his uh um to to provide give his um uh the turing award lecture uh he and jan lakum both came and gave the lectures here um anyway so when uh jeff started giving the talk and i was actually in the audience you know because it's in phoenix we all you know made a pilgrimage to the convention center there and we were in the audience and this is a picture i took um of you know his first slide he started by saying there are two ways to make a computer do what you want one is intelligent design and the other is learning say what you mean about
所以为了鼓励一下……我知道你们现在并不在工作时间,所以我想我应该建议一个不错的喝酒游戏,所以给你们一个有趣的游戏:每次你听到「隐性(tacit)」或者「显性(explicit)」这个词,就喝一口橙汁,这样你就能补足一整年的维生素 C,这是肯定的。好,那我们开始吧。先从这个讲起。好,去年 2019 年,Geoff Hinton 当然来到了凤凰城,来做他的图灵奖演讲,他和 YannLeCun 都来这里做了演讲。总之,当 Jeff 开始做报告的时候,我其实就坐在台下,因为是在凤凰城嘛,我们都算是朝圣一样跑到了那边的会议中心,坐在观众席里。这是我拍的一张照片,是他的第一张幻灯片。他一开场就说,让计算机做你想让它做的事有两种方法:一种是智能设计(intelligent design),另一种是学习。不得不说,Jeff 这个人真的很会用词,用得恰到好处,他要是来参加
便签引用
7:36
um jeff the man knows how to use words just so just right he would have been a natural in icaps festivals really um it turns out that you can put intelligent in front of pretty much anything in the world and it would mean a good thing and he put it in front of design and it becomes a bad thing you know those of you who know the whole intelligent design theory business which is all these evolution deniers so obviously nobody in science would like to be connected to intelligent design in case you couldn't see what was on that side this is what you were saying intelligent design involves figuring out consciously exactly how you manipulate representations to perform a task and then tell the computer in detail and you know things like excruciating etc is of course jeff doing it but really if you know how to do the task you tell the computer a model of the task and then presumably you know add search algorithms on top of it the kind of thing that we do in planning and then he says that is intelligent design
ICAPS 的 Festivus,绝对是个天生的好手。有意思的是,你几乎可以把「智能的」这个词放在世界上任何东西前面,它都会变成一件好事,可他把它放在「设计」前面,它就变成了一件坏事。你们当中了解「智能设计论」这套东西的人都知道,那帮人都是进化论否定者,所以显然科学界没有人愿意跟智能设计扯上关系。万一你看不清那边写的是什么,他说的是:智能设计就是要有意识地弄清楚你到底该如何操作表示来完成一项任务,然后事无巨细地告诉计算机,用了像「令人痛苦地(excruciating)」这样的词,这当然是 Jeff 的风格。但说真的,如果你知道怎么做这个任务,你就告诉计算机这个任务的模型,然后大概再在上面加上搜索算法,也就是我们在规划领域做的那种事情。然后他说,这就是智能设计,
便签引用
8:37
so that anybody what their salt in science would be worried about being in that area um and then the other of course is learning show the computer lots and lots of examples of inputs together with the desired outputs and then let the computer learn how to map from inputs to outputs you know it's like basically just let the evolution play on in computing and let the you know ai happen that way so that was the way he starts on the talk in a very nice way um now learning versus being told okay which is basically an interesting thing for this community we know that we actually give pdl models we'll talk about a little more later but we do give computers what we know about the world and we let them do the combinatorics um so hinter is really just reinforcing the ai zeitgeist uh if only in sort of a doctrinal farm and if only in a very hintonesque form um ai technology has managed to catch public imagination of it i mean thanks in large part to the impressive feats in perceptual intelligence the things like
所以任何在科学界有点分量的人都会担心自己身处那个领域。然后另一种当然就是学习:给计算机看大量大量的输入示例以及期望的输出,然后让计算机自己学会如何从输入映射到输出。基本上就是让进化在计算中自行演进,让 AI 以那种方式产生。这就是他开场的方式,讲得非常漂亮。那么,学习 vs. 被告知,这对我们这个社区来说是个很有意思的话题。我们知道我们其实是给出 PDDL 模型的,稍后我们会多讲一点,但我们确实是把我们对世界的了解告诉计算机,然后让它去做组合搜索。所以 Hinton 其实只是在强化 AI 界的时代思潮,只不过是以一种教条的方式,以一种非常 Hinton 式的方式。AI 技术已经成功抓住了公众的想象力,我是说,这在很大程度上要归功于在感知智能方面的惊人成就,比如视觉、音频处理等等,还有语音
便签引用
03隐性知识任务与波兰尼悖论
9:38
the vision audio processing etc um voice recognition and so on and then bringing those advances to the masses on the street who you know with their cell phones okay that wound up being like a very big way ai has become um you know very popular all over the place uh most of these advances um have in fact however been in what i would call tacit knowledge tasks um tacit knowledge tasks we'll talk in a minute is basically these tasks that we do what we have no clue how we do so if you were to make a theory of how people see objects ask computer vision people they've been trying to do that and it was an object failure object with the a uh object with the oh uh but anyway it was a failure and so it was actually we don't know how we do uh voice processing we don't know how we see the world um and so these are tacit knowledge tasks and intelligent design approach of course fails because if you don't know how to do the task and if you tell the computer how to do it it's going to be the wrong way to do it
识别之类的,然后把这些进展带给街上的普通大众,带到他们的手机里。这最终成了 AI 变得如此流行、到处都是的一个非常重要的途径。不过这些进展中的大多数其实都发生在我称之为「隐性知识任务」的领域。隐性知识任务,我们等下会讲,基本上就是那些我们会做但完全不知道自己是怎么做到的任务。所以如果你要建立一套人是怎么看见物体的理论,去问计算机视觉的人,他们一直在尝试做这件事,结果是彻底失败——用「对象」失败,呃……总之就是失败了。所以其实我们并不知道自己是怎么做的。语音处理我们不知道是怎么做的,我们不知道自己是怎么看世界的。所以这些都是隐性知识任务,而智能设计的路子当然会失败,因为如果你自己都不知道该怎么做这个任务,你还去告诉计算机怎么做,那肯定是用错误的方式在做。但接下来一个有意思的问题是:隐性知识任务真的是
便签引用
10:40
but then the interesting question is are tacit knowledge tasks really the only thing and or everything though interestingly this community looks at this problem in slightly different way from the rest of the world especially people let's say on robotics or vision community um basically we do realize that in fact there are many tasks for which there is a human um know how you know it about the domain models and so on and that's what we started from that's what the whole pdl model stuff is um so when you're thinking about it that's where we start thinking about palani's paradox apalanji is this polymath hungarian um the food root among many other things he's a philosopher and one of the things that he did was this writing this book about tacit dimension that acid dimension which is basically he started talking about tacit knowledge you know his definition from wikipedia shown there essentially as i said it's the thing that we know how to do but we don't know how we do it we don't we are not
唯一的东西吗?或者说是全部吗?有意思的是,我们这个社区看待这个问题的角度,跟世界上其他人略有不同,尤其是跟机器人或视觉社区的人相比。基本上我们确实意识到,其实有很多任务是存在人类的「诀窍」的,你知道的,关于领域模型之类的,而这正是我们的出发点,整个 PDDL模型那一套就是这么来的。所以当你想到这一点时,我们就开始想到波兰尼悖论(Polanyi's paradox)。Polanyi是一位匈牙利博学家,除了很多其他身份之外,他还是一位哲学家,他做的其中一件事就是写了《隐性维度》(The Tacit Dimension)这本书,在这本书里他开始谈论隐性知识。你们看到的这个来自维基百科的定义,本质上就像我说的,是那些我们知道怎么做但不知道自己是怎么做到的事情,我们并没有有意识地觉察到自己是如何完成这个任务的。所以 Polanyi 在
便签引用
11:40
consciously aware of how we do this task so polani actually was lamenting when he wrote this book on acid dimension he was lamenting the fact that too much of our attention is typically focused on understanding uh explicit knowledge tasks rather than the tacit knowledge ones um and so and then basically so he was arguing that really we know more than we can tell we know more than we can verbalize and so how come we are ignoring that stuff that's what he was worried about so he was this was sort of been seen as polany's paradox and to understand how into how inculcated it is in our psyche about these explicit knowledge tasks remember things like you know this this uh penman uh thing saying if you want to master something teach it okay it turns out that actually really only works for explicit knowledge tasks that's the same way if you know if you really want to understand how to do something program it you know tell me if anybody who programmed um like a transformer-based uh language completion
写《隐性维度》这本书时其实是在感叹一个事实:我们太多的注意力通常都集中在理解显性知识任务上,而不是隐性知识任务。所以他当时的论点是,我们知道的其实比我们能说出来的多,我们知道的比我们能用语言表达的多,那我们怎么能忽视那部分东西呢?这就是他所担心的。所以这后来被视为波兰尼悖论。要理解显性知识任务这个观念在我们心里扎根有多深,想想那句话,你知道的,费曼说过的那种话——如果你想真正掌握某样东西,就去教它。事实证明,这其实只对显性知识任务有效。同理,如果你真的想理解怎么做某件事,就去把它编程实现出来。你告诉我,有谁编程实现了基于 Transformer 的语言补全,或者基于卷积网络的视觉,然后就搞明白了我们人是怎么看世界的
便签引用
12:55
our canonet-based vision has figured out any more about how we see the world or how we complete the language that's not at all the case so in fact we always just thought when we talk of thoughts uh tasks we thought of explicit knowledge tasks and so it seemed like a reasonable thing that polanye was worried that you know guys look at a little bit of the statute knowledge tasks too um in fact you know this is a slide that i'm shown a couple of years back in a talk that i gave at icaps um that you know if you just look at the kinds of the types of intelligences that people show uh human kids sort of come into this world showing perceptual and manipulation intelligence then emotional intelligence then social communicative intelligence and then finally start showing cognitive and reasoning intelligence and this is sort of almost like the human kids when they come in they have sort of animal like intelligence abilities that they're able to show already and then they grow over and they start showing this
或者我们是怎么补全语言的?完全不是这么回事。所以事实上,我们一直以来在谈到「思考」任务时,想的都是显性知识任务,所以 Polanyi 当时担心的那件事看起来是很合理的,就是说各位也稍微看看隐性知识任务吧。其实,这是我几年前在一次 ICAPS 演讲中展示过的一张幻灯片:如果你去看人所展现出来的各种类型的智能,人类的孩子来到这个世界,最先展现的是感知和操作智能,然后是情感智能,然后是社交沟通智能,最后才开始展现认知和推理智能。这几乎就像是说,人类的孩子刚来到世界时,已经具备了某种类似动物的智能能力,然后他们长大,开始展现出这种认知推理能力,而我们把这种能力跟人类文明联系在一起。
便签引用
04AI 发展顺序:推理先于感知
13:56
cognitive reasoning which we sort of have connected to human civilization okay um so interestingly uh in if you haven't already thought about it um ai systems developed exactly in the opposite way right we were essentially creaming uh deep blue streaming the chess champion gary caspero way before deep blue can recognize a little chess piece on the board because vision was not something that ai systems were doing before but they were doing cognitive and reasoning um tasks for quite a bit before okay this is something that people in planning should really really remember and understand uh not the people who come into ai after alex net thinking the only thing that we need to do is start from vision and somehow build a human but people in planning know better essentially um which is exactly why i keep making fun of this you know that show me an ai expert confidently proclaiming that in future ai will not only just learn but will also be able to reason too and i'll show you someone who entered ai via newspaper headlines
有意思的是,如果你还没想过这一点的话,AI 系统的发展恰恰是按相反的顺序进行的,对吧?我们早在深蓝能认出棋盘上一个小棋子之前,就已经用深蓝把国际象棋冠军卡斯帕罗夫打得落花流水了,因为视觉并不是早期 AI 系统在做的事,但它们在相当早的时候就已经在做认知和推理任务了。这是搞规划的人真的真的应该记住并理解的一点,而不是那些在 AlexNet 之后才进入 AI、以为我们唯一要做的就是从视觉出发然后设法造出一个人的那些人。搞规划的人心里更有数。这也正是我为什么老是拿这个开玩笑:给我看一个信心满满地宣称「未来 AI 不仅会学习,还将能够推理」的 AI 专家,我就能指给你看一个是通过报纸头条、在 AlexNet 之后才进入 AI 的人。因为说到底,AI 早在开始做
便签引用
15:04
after alex net because after all ai has been doing reasoning way before it started doing tacit knowledge tasks like perception so keep that in the back of my mind um it wants up that it actually explains quite a bit of what coison in ai it's a frog easier to program computers and aspects of intelligence for which we do have some kind of conscious theories they may not be foolproof but at least we know how to write a couple of pdl operators for various you know planning domains but if try writing an operator for vision task you have nothing to write because you don't actually know how vision is done you know by us um so it turns out that for these explicit knowledge does the progress in reasoning and cognitive intelligence happened much faster because we actually first started right telling the computers at least partial models of how to do those tasks and automated it with combinatoric search and various other things we are not particularly conscious at all of perceptual and manipulation
感知这类隐性知识任务之前,就已经在做推理了。所以请把这一点记在心里。事实证明这其实能解释AI 中很多事情的成因:为那些我们确实有某种有意识理论的智能方面编程要容易得多,这些理论也许不是万无一失的,但至少我们知道怎么为各种规划领域写出几个 PDDL 算子。可你要是试着为视觉任务写一个算子,你根本无从下笔,因为你其实并不知道我们人是怎么完成视觉的。所以事实证明,对这些显性知识任务来说,推理和认知智能方面的进展要快得多,因为我们其实一开始就在告诉计算机至少是部分的任务模型,然后用组合搜索和其他各种方法把它自动化。而对于感知、操作
便签引用
16:03
and you know those sorts of intelligences and so we have to depend on making machines learn exactly the way we did which is basically learn from observation data demonstration experience and so on and so forth okay um keep in the back of your mind that if you only as a human being in this society learnt only from your raw experience never saw anybody else never read anything never got told anything you are not likely to be a civilized person even at the lowest common denominator level because we do depend a lot on the ability to transfer these explicit knowledge tasks across but things like you know how to walk etc you just have to do it yourself nobody taught you how to walk you know and nobody taught you how to speak at least in the beginning parts um and you know of course that this thing became feasible you know as i as we think about the fact that this became feasible uh this learning for these tasks became feasible because really of the orthogonal ex you know extension orthogonal developments in things other than ai
这类智能,我们完全没有有意识的觉察,所以我们只能依赖于让机器以我们自己学会的方式去学习,也就是从观察、数据、示范、经验等等中学习。好,请记在心里:如果你作为一个人在这个社会里只从你自己的原始经验中学习,从没见过别人,从没读过任何东西,从没被告知过任何事,那你很可能连最低标准意义上的文明人都算不上,因为我们非常依赖于把这些显性知识任务传递给他人的能力。但像走路之类的事情,你只能自己去做,没有人教过你怎么走路,也没有人教过你怎么说话,至少在最开始的阶段是这样。当然,你们知道这件事之所以变得可行——我们想想看它为什么变得可行——这些任务上的学习之所以变得可行,其实是因为 AI 之外的一些正交的发展,
便签引用
05推理中心与学习中心两种路线
17:07
such as web and we just completely uploaded our collective subconscious to the web and then that became the data in which you are training your systems now it's also useful to realize that actually if you take an intro to ai class as against a machine learning class both of which would be really well connected but the difference you will wind up noticing is that much of ai tends to at least the textbooks tend to be oftentimes inference focused and so i basically you can think of these two broad ideas inference versus learning focused approaches to um ai or intelligent agent design in the case of inference focus is the kind of thing planning people always did you sort of assume that models are available so you're sort of looking at explicit knowledge tasks uh there are representations are available there's no such thing as learning representations we generate representations and we write our knowledge in that representation and then we generate algorithms that will um basically be able to search
比如互联网。我们把我们集体的潜意识彻头彻尾地上传到了网上,然后那就成了你现在用来训练系统的数据。另外,认识到这一点也很有用:如果你去上一门 AI 导论课,而不是机器学习课——这两者当然是密切相关的——但你会注意到的区别是,AI 的大部分内容,至少教科书里的内容,往往是以推理为中心的。所以基本上你可以想到这两大类思路:以推理为中心 vs. 以学习为中心的 AI 或智能体设计方法。以推理为中心的,就是搞规划的人一直在做的那种事:你基本上假设模型是现成的,所以你看的是显性知识任务,表示也是现成的,不存在「学习表示」这种事,我们自己设计表示,然后把我们的知识写进那个表示里,接着我们设计算法,让它能够搜索各种
便签引用
18:05
through configurations and to reasoning with respect to that so it's sort of inference focused we do make a little promise to ourselves saying you know one of these days we're going to do learning um but oftentimes we don't really get up to that at least for quite a long time people didn't get up to that and so inference winds up being in the foreground and it's great for explicit knowledge domains with good models and of course the aid of development followed this direction for much of its history and much of the planning literature has done that too and one of the things i would try to get you to believe is that this is a great thing because being human involves doing both explicit and tacit knowledge together okay um on the learning focus side you make the opposite it's sort of i think it's like the earliest versions of this have been connected to things like subsumption architecture which is assumed that agent have no apriori math models focus on learning even the primitive models and representations
配置、并基于这些表示进行推理。所以它是以推理为中心的。我们确实会对自己做一点小小的承诺,说「总有一天我们要去做学习」,但往往我们并没有真的做到,至少在相当长的时间里人们都没做到。所以推理就成了前景中的主角,这对于有良好模型的显性知识领域来说非常棒,当然,AI 的发展在其历史的大部分时间里都沿着这个方向走,规划领域的大量文献也是这么做的。而我想让你们相信的一件事是:这是一件很棒的事,因为「做一个人」意味着同时要处理显性和隐性知识。好。而在以学习为中心的一侧,你做的正相反。我想这条路最早的版本可以联系到像包容式架构(subsumption architecture)这样的东西,它假设智能体没有任何先验的数学模型,专注于连最基本的模型和表示都要学习出来,你最终至少能得到一个
便签引用
19:01
you wind up at least getting a sort of a typically reflex agent and then sort of promise to yourself that one of these days we start doing reasoning and you know longer term um decision making etc etc so this idea typically this area trend to postpone inference um and reasonable for tacit knowledge domains with no good models but a lot of examples and experience generators and there's been significant research progress in this side more recently especially after alex net after 2013 and you know that sort of got us to a point where as i think andrew hang put it anything humans can do in a few seconds computers are able to do okay that's great amazing that computers are now able to do anything that humans can do in a few seconds but if humans only did what they can do in a few seconds you will have precedence like ours right now okay so you actually humans plan humans make long-term reasoning humans to all sorts of explicit knowledge uh based uh reasoning and so on and so they sort of both have to be
典型的反射式智能体,然后对自己承诺说,总有一天我们要开始做推理、做长期的决策等等等等。所以这个思路,这个领域往往倾向于推迟推理,对于没有好模型但有大量样例和经验生成器的隐性知识领域来说是合理的。而这一侧最近取得了显著的研究进展,尤其是在 AlexNet 之后、2013 年之后。这也把我们带到了一个点上——我想是吴恩达说过的——凡是人类在几秒钟内能做的事,计算机都能做到。好,这很棒,很了不起,计算机现在能做到任何人类在几秒钟内能做的事。但如果人类只做那些几秒钟内能做完的事,你们现在就会有像我们这样的总统了。好,所以其实人类会做规划,人类会做长期推理,人类会做各种基于显性知识的推理等等,所以对人来说这两者必须结合起来。事实上,我已经
便签引用
06系统一二与显性知识的编译
20:03
combined for humans and um so in fact and i already mentioned as to why this specifically the tacit knowledge tasks being if you know our ability to do that wound up increasing and catching public imagination already but i don't want to go into more but you know certainly people like being able to use the fact like the fact that their cell phones can recognize their voice it can they can complete their sentences etcetera these are very very useful abilities but that's not necessarily the full story of intelligence or intelligent behavior okay um so um one other thing i want to mention just before since i came very close and some of you must be thinking about this law is talking about explicit versus tacit knowledge i already know about system one and system two uh which is basically kahneman and firsky did this uh i talked about it by the way system on system two are essentially theories there is no part of your brain called this what is system one this body system two it's just a a metaphorical way of
提到了为什么隐性知识任务——我们做这类任务的能力提升起来并且抓住了公众想象力。这我不想再多讲,但人们确实很喜欢用到这样的事实,比如他们的手机能识别他们的声音,能补全他们的句子等等,这些都是非常非常有用的能力,但这未必就是智能或智能行为的全部故事。好。还有一件事我想提一下,因为我已经讲得很接近了,你们当中有些人一定在想,这家伙在讲显性 vs. 隐性知识,我早就知道系统一和系统二了,也就是 Kahneman 和Tversky 提出的那套。顺便说一句,我讲过这个,系统一和系统二本质上是理论,你脑子里并没有哪个部位叫系统一,哪个部位叫系统二,它只是一种理解大脑能力、理解人类
便签引用
21:07
thinking of the brain's abilities human brain's abilities ah and that system one tends to do reflexive reasoning and system two tends to do deliberative reasoning having said that there is somewhat of a difference between system one system two and explicit and tacit i want you to understand it as we going forward here most tacit knowledge tasks do get handled by system one okay however explicit knowledge tasks can start in system two you thought start doing things deliberately before but may get compiled into system one reflexive behavior for efficiency in fact people have said that civilization progresses by the ability to do many more things without thinking than you started with so originally when you started um you know figuring out how to do differentiation you are doing limits f of x plus h minus f of x by um you know edge limit has to be zero but then if you kept doing that you would be very much behind times in terms of the ability to do differentiation so you compiled it down and you started actually doing some of
大脑能力的隐喻性方式。系统一倾向于做反射式的推理,系统二倾向于做审慎的推理。话虽如此,系统一、系统二和显性知识、隐性知识之间还是有一些差别的,我希望你们在接下来的内容里能理解这一点。大多数隐性知识的任务确实是由系统一来处理的,好吧,但是显性知识的任务可以从系统二开始,你一开始是刻意地、有意识地去做这些事,但后来可能被编译进系统一,变成一种反射式的行为,以提高效率。事实上有人说过,文明的进步就在于人能够不假思索地做的事情越来越多,比最初的时候多得多。所以最开始,当你刚在琢磨怎么做微分的时候,你用的是极限,f(x+h) 减去f(x),再除以 h,然后让 h 趋于零。但如果你一直这么算下去,那你在做微分这件事的效率上就会远远落后于时代。所以你把它编译下来了,你开始真的能反射式地做这些事情,尽管这些显性
便签引用
22:10
this stuff reflexively even though the syst the explicit knowledge does that got compiled into system one and tacit knowledge tasks that really stay in system one are sort of in both system one there's still a difference because you know it's the difference between just an assembly program or an assembly program that got compiled from a higher level language in the later case if the assembly program fails the higher level language can be you can actually localize the failure in the higher level language program and that's how we wind up debugging uh symbolic debuggers work that way so this is basically interpretability aspects come into play when you have explicit knowledge models um i can't resist showing something that was a printed uh way before probably most of you are gone in doctoral construction for sure in 1989 you know there was ai magazine had this huge raging controversy ai area as well as planning um and so magazine and this thing about universal planning universal planning and almost
知识是被编译进系统一的。而那些本来就一直待在系统一里的隐性知识任务——虽然两者都在系统一里,但还是有区别的,因为这就好比:一个是直接写出来的汇编程序,另一个是由高级语言编译得到的汇编程序。后一种情况下,如果汇编程序出错了,你其实可以在高级语言层面定位这个错误,我们就是这样调试的,符号调试器就是这么工作的。所以基本上,当你拥有显性知识模型时,可解释性方面的东西就派上用场了。我忍不住想给大家看一个东西,它印出来的时候,在座大多数人肯定还没开始读博士,那是 1989 年。当时《AI Magazine》上有一场非常激烈的争论,涉及整个 AI 领域,也涉及规划领域,就是关于通用规划(universal planning)的——“通用规划,一个几乎通用地糟糕的主意”,这是马特·金斯伯格(Matt Ginsberg)说的。因为 Marcel Schoppers 基本上是说:
便签引用
23:13
universally bad idea that's what matt ginsberg was saying because martial shoppers basically said look why don't you want to start with something like a pdl model those days there was no period but compile it down to a reflexive policy so you essentially solve lots of problems upfront and just remember the solutions you know more or less memoize okay and matt ginsberg basically writes a whole bunch of pages pointing out that the combinatorics are such that the table in which you want to remember is going to be huge and so it's a universally bad idea and that you want to think sometimes um when you end up doing this right now the times have changed right now we no longer think so we're not so worried about space versus online computation tradeoffs in fact we are very happy to throw online computation time into space so that you can put you know your policy in terabytes if you want as long as you can just answer the question very quickly and that trade-off is a different one from whether the
你看,为什么不从类似 PDDL 模型这样的东西出发(那个年代还没有 PDDL),然后把它编译成一个反射式的策略呢?这样你相当于提前把大量问题都解出来了,然后把解记住,差不多就是做记忆化(memoize)。好,然后马特·金斯伯格写了一大堆篇幅指出:组合数学上的爆炸意味着你想用来记忆的那个表会大得惊人,所以这是一个通用地糟糕的主意,你有时候还是需要去思考的。而现在,时代变了,现在我们已经不这么想了,我们不再那么担心空间与在线计算之间的权衡了。事实上,我们非常乐意把在线计算的时间换成空间,比如你想的话,可以把策略存成好几个 TB,只要你能非常快地给出答案就行。而这种权衡,跟你最初有没有显性知识是两码事,这一点
便签引用
07波兰尼的复仇:数据压倒教条
24:14
original original you had explicit knowledge or not that is something that i want you to understand as we go forward okay so now polandi paradox is what we talked about but you know as you can see while i was talking really there's a bit of a polonius revenge that's been going on in ai right um essentially ai now has this complete new romance with tacit knowledge um this is old show the whole show hold up the story about archimedes who drunk by his you know the fact that he discovered this idea of fulcrum on the labor said give me a liver and a place to stand i can move the world okay and so it's sort of like that we are right now saying give me a begin of gpu a large enough data set and a deep enough network i will create you a gi just not even human level ai just agi altogether completely um so this sort of bothered me i you know partly because i've seen both sides of ai and i think you know being human involves doing both of these both explicit fantastic knowledge tasks and so i wrote this um in a viewpoint
我希望你们在后面的内容里能理解。好,我们刚才讲的是波兰尼悖论,但正如你们看到的,我讲的过程中其实 AI 里正在上演一场“波兰尼的复仇”,对吧。本质上,AI 现在跟隐性知识陷入了一场全新的热恋。这是个老掉牙的故事了——阿基米德的故事,他因为发现了杠杆和支点这个想法而兴奋不已,说:给我一个支点和一个立足之地,我就能撬动地球。好,现在的情况有点像这样,我们现在说的是:给我一大堆 GPU、一个足够大的数据集、一个足够深的网络,我就能给你造出一个 AGI——甚至都不只是人类水平的 AI,而是彻头彻尾的 AGI。所以这多少让我有点困扰,部分原因是我见过 AI 的两面,我觉得,做一个人本来就同时涉及这两类任务——显性知识任务和隐性知识任务。所以我写了这篇
便签引用
25:20
article um which is coming out in csm um in february uh that's also available on my webpage in that address uh sal palani's revenge and ai's new romance with tacit knowledge and what you see in that viewpoint article is some of this setup that i've shown you as well as a couple more points um that i want to show you before i go from that article to its direct impact on planning which i've been talking about to some extent but you know i want to talk more directly about how does this all matter to planning folks in i caps in particular okay so this is data versus doctrine tension um in the world that you live in poland basically was worried that you are expecting you're only working on problems for which there is explicit knowledge there is doctrine and let's just work also on passive knowledge but now we have went almost the other way ai is basically we are mostly only working on tacit knowledge tasks it's almost become fashionable to take problems with explicit knowledge models such as sudoku
观点文章,它将于二月发表在 CACM 上,在我的网页上那个地址也能看到,题目是《波兰尼的复仇与 AI 与隐性知识的新恋情》。在那篇观点文章里,你会看到我刚给大家展示的这些铺垫,还有另外几点,我想在这里也讲一下,然后我就要从这篇文章转到它对规划领域的直接影响了。前面我多少已经提到了一些,但我想更直接地讲一讲,这一切对规划领域的人——尤其是 ICAPS 的各位——到底意味着什么。好,这就是数据与“教条/成文知识”之间的张力。在波兰尼所处的那个世界里,他担心的是:你们只去做那些有显性知识、有成文知识的问题,他说我们也该去做隐性知识。而现在我们几乎走到了另一个极端,AI 基本上主要只在做隐性知识的任务,现在几乎成了一种时髦:把本来有显性知识模型的问题,比如数独,转换成海量的
便签引用
26:24
and convert them into a bazillion examples just so that you can say i did deep learning okay so it's almost like forget about going from data to knowledge there is an entire cottage industry about going from knowledge explicit knowledge to data just so that you can then uh give it to some transformer or something and then try to recover the original knowledge you started with and and essentially write a paper as strange and exotic as it sounds it's being done many of you know this and some of you are probably doing this so the interesting question of course is what do we do when we actually have doctrine that we want systems to follow you know that's the kind of things that people in planning have always been looking at you know look at domains where there is some amount of domain knowledge that people are happy and willing to give you would you just spurn them and say don't talk to me i will just learn from behavior that seems like a completely uh silly thing to do to me at any rate uh one of the
样例,就为了能说一句“我做了深度学习”。好,所以这几乎就是:别提什么从数据到知识了,现在有一整个作坊式的产业,是在把知识、把显性知识变成数据,就为了能把它喂给某个 Transformer 之类的东西,然后再试图把你一开始就有的那个知识给恢复出来,然后写成一篇论文。听起来又古怪又离奇,但确实有人在这么做,你们很多人都知道,你们中有些人可能自己就在这么做。所以有意思的问题当然是:当我们真的手里有希望系统去遵循的成文知识时,我们该怎么办?你知道,这正是做规划的人一直在关注的东西——去看那些领域,那里有一定量的领域知识,而且人们很乐意、也很愿意告诉你。你会一口回绝他们,说“别跟我讲,我只从行为里学”吗?这在我看来至少是件相当愚蠢的事。人类智能的标志之一似乎就是
便签引用
27:22
hallmarks of human intelligence seems to be a seamless interplay between tacit and explicit knowledge and in a weird way i think the pendulum for us has swung from all models are wrong some are useful to what are models why do we need models we just go from data to decision directly and that is something that we need to obviously give some thought to um so in in the in the write up actually on the um on the viewpoint which is meant for a more generalized audience i actually tried to connect the poland's revenge to a whole bunch of affliction the things that are reflecting ai right now and i've written a bunch of articles you know on this separately you can look up my web page um but you know for example this whole issue of interpretability you should not be surprised at all that systems which learn their own representations don't have to make sense to you okay and you know whereas if you started from something like a pdl model that is sort of at least you understand the things that you're putting in
隐性知识和显性知识之间无缝的相互配合。而有点奇怪的是,我觉得我们的钟摆已经从“所有模型都是错的,但有些是有用的”,摆到了“模型是什么?我们要模型干嘛?我们直接从数据到决策”。这是我们显然需要认真想一想的事情。所以在那篇写给更广泛读者的观点文章里,我其实试着把波兰尼的复仇,跟当下 AI 面临的一大堆问题联系起来。我另外还就这些写过一些文章,你们可以去我的网页上找。不过举个例子,可解释性这个问题——你完全不应该感到意外:那些自己学出表示的系统,它学到的表示没有理由对你有意义。好,而如果你是从类似 PDDL 模型这样的东西出发,那你至少还理解你放进去的是什么东西,所以那里的可解释性问题要简单得多。
便签引用
28:25
so the interpretability problem is a much simpler issue there um similarly this issue of if they do if you don't quite know what they have learned then susceptibility to adversarial attacks is quite high because you can't guarantee that things will work as expected and finally of course probably much more interestingly um susceptibility to data set by us the two things that are going on on the right hand side one is gpt3 being given um from saying two muslims and irrespective of what you do many of the completions that it comes up with is that the muslims were involved in some bad thing like they killed people they got killed they did something etc it's a completely nonsensical thing in a civilized world for us to spark about and yet it's not surprising because open ai the gpd3 essentially learned from our collective subconscious we do have crazy thoughts we do have thoughts that we don't say it aloud you know unless you know unless you're impetuous and you have very poor emotion control
同样地,还有这个问题:如果你并不太清楚它们到底学到了什么,那么它对对抗攻击的脆弱性就相当高,因为你没法保证事情会按预期运行。最后,当然,可能更有意思的是数据集偏见的问题。右边有两个例子,一个是给 GPT-3 输入“两个穆斯林”,然后不管你怎么做,它给出的很多续写都是这些穆斯林卷入了某种坏事,比如他们杀了人、他们被杀了、他们干了什么等等。这在一个文明社会里是完全说不通的事,然而这并不令人意外,因为 OpenAI 的 GPT-3 本质上是从我们的集体潜意识——我们确实会有疯狂的念头,我们确实会有一些不会说出口的念头,除非,你知道的除非你性子急躁、情绪控制、冲动控制很差,否则你是不会把这些说出来的,这个过程是这样发生的:
便签引用
29:27
impulse control you don't talk about this the way this happens is what the system one generates the system too can stop from being explicitly said by you so there is a control that goes on and we don't do this and of course it's not surprising that you wind up having these problems um you know and so this is something the bigger sense of how the fascination with tacit knowledge can get us into trouble uh so the question of course is when do you learn from examples versus when do you take knowledge from the humans it's obviously you can delude yourself in both cases planning community as i'll show you in a minute was hoping for superhuman humans who will be able to tell the planner how to decide whether node number 7539 should be removed from the search tree or not that was for a while the things that we did in things called mixed initiative planning uh we were expecting way too much uh we were expecting humans who had no life to begin with and you know on the other hand this other side is saying there is nothing humans
系统一生成的东西,系统二可以阻止你把它明确说出来,所以中间是有一层控制的,而我们不会那么做,当然,你最后会碰上这些问题也就不奇怪了,嗯,所以这就是那个更宏观地看,对隐性知识的痴迷会怎样把我们带进麻烦。所以问题当然就是,什么时候该从样例中学习,什么时候该从人那里获取知识。显然,这两种情况下你都可能自欺欺人。规划这个圈子,我等一下会展示给你们看,一直期待着有超人般的人类,能够告诉规划器怎么判断第7539号节点该不该从搜索树中剪掉。有一阵子我们做的就是这类事情,叫做混合主动式规划。嗯,我们的期待实在太高了,我们指望的是那些压根没有自己生活的人类。而另一方面,另一派则说人类没有任何东西可以告诉我们,我们就只从数据里学。所以显然,弄清楚什么时候
便签引用
30:32
can tell us we'll just try to learn from data and so clearly figuring out when is it you're supposed to learn and when is it supposed to ask for data is something that requires some wisdom and so i have my own version of the serenity prayer that christianity has except for this purple robot it says human grant me the serenity to accept the things that i cannot learn and data to learn the things i can and wisdom wisdom to know the difference and part of this talk is to try to give you at least a little bit of that wisdom some of you already possess it some of you have shown that you don't have that wisdom because some of your written papers that i'll be naming names on in a minute but it's the kind of thing that we need to figure out and you know an approximate idea is where the explicit knowledge is available easily we should use it and we try then of course that opens up interesting questions as to how to bridge that with tacit knowledge which is something that we will talk about okay let's get to now the planning part
该去学习、什么时候该去要数据,是需要一些智慧的。所以我有我自己版本的基督教里的《宁静祷文》,只不过是给这个紫色机器人的,它说:人类啊,请赐我宁静去接受我学不到的东西,赐我数据去学习我能学到的东西,还有智慧,赐我分辨二者的智慧。这次演讲的一部分,就是想至少给你们一点点这样的智慧。你们中有些人已经具备了,有些人则已经表明自己并不具备这种智慧,因为你们中有些人写的论文——我待会儿会点名——但这正是我们需要弄明白的事情,而且,一个大致的想法是:在显性知识容易获得的地方,我们就应该用它。而这样一来,当然就带出了一些有趣的问题:如何把它和隐性知识衔接起来,这个我们等下会谈。好,现在我们进入规划的部分,
便签引用
08规划社区的显性知识传统
31:29
of this and how this whole thing projects on to planning um i cap style planning uh the mainstreamish i caps trail planning automated planning has for the most part focused on explicit knowledge tasks you know our central conceit has always been that there are many domains where people have explicit verbalizable knowledge about the task and and so there so basically we focused on model specification languages whether it is strips pdl shop uh redl um etc to make easier for people to write the knowledge that they want to verbalize and then developing efficient of the shelf planners for handling these models okay so explicit knowledge in planning is like all over the place we have to admit this you know in icaps we can't say explicit knowledge doesn't matter because then what the heck are we doing you know in most of the work that we are doing so planning community has always taken the easily available explicit knowledge from human designers we in fact supported even approaches that are even more knowledge intensive
以及这一整套东西如何投射到规划上。嗯,ICAPS 风格的规划,主流的 ICAPS 路线的规划,自动规划在很大程度上一直聚焦于显性知识的任务。你知道,我们的核心自负一直是:有很多领域,人们对任务拥有可以用语言表述的显性知识,因此,所以我们基本上聚焦在模型描述语言上,不管是 STRIPS、PDDL、SHOP、RDDL 等等,为的是让人们更容易把他们想要表述的知识写下来,然后再开发高效的通用规划器来处理这些模型。好,所以显性知识在规划里到处都是,这一点我们得承认。你知道,在 ICAPS 里我们不能说显性知识不重要,否则我们做的那么多工作到底算什么呢,你知道的。所以规划这个圈子一直在使用人类设计者那里容易拿到的显性知识,我们实际上甚至还支持那些知识密集度更高的做法,比如提供控制信息和抽象信息,试图从人那里把它
便签引用
32:35
providing for example control and abstraction information trying to get it from humans we as i said expected superhuman humans who had nothing else to do other than give us models sometimes okay so there is explicit model such as pdl that is easy enough but we also expected sometimes control information control rules etc so in fact the last time around i was this riled up about knowledge and explicit knowledge was back in icaps 2003 um which that slide there where i was sort of ranting about knowledge based planning and whether we are actually comparing fahim and dana or are we comparing two planning algorithms with respect to some knowledge bases that becomes kind of interesting so we had these issues but we certainly were open to using explicit knowledge but lately there are some exotic attempts at best to jump on the pixels to decisions movement even for domains where explicit knowledge is easily available i don't know whom we are trying to impress doing that but certainly there is work of that kind that's been
拿过来。就像我说的,我们指望的是超人般的人类,他们除了给我们模型之外没有别的事可做。有时候,好,所以有像 PDDL 这样足够简单的显性模型,但我们有时候还指望控制信息、控制规则等等。事实上,上一次我对知识和显性知识这么激动,还是在 2003 年的 ICAPS,嗯,就是那边那张幻灯片,我在上面有点像是在抨击基于知识的规划,以及我们到底是在比较 Fahiem 和 Dana,还是在比较两个规划算法在某些知识库上的表现,这就变得挺有意思了。所以我们那时有这些争论,但我们当然是乐于使用显性知识的。不过最近出现了一些顶多算是标新立异的尝试,想搭上“从像素到决策”的潮流,哪怕是在显性知识唾手可得的领域也这么干。我不知道我们这么做是想给谁留下深刻印象,但确实有这类工作在出现,在整个 AI 里是这样,在规划圈子里也是这样。
便签引用
09强化学习把波兰尼带进规划
33:42
happening in ai itself as well as in the planning community in general um so this essentially is the way polany will come into planning with his revenge and the palani comes to planning through rl reinforcement learning that's the general way to think about it so let me talk about two slides worth of that connection between planning and um especially deep reinforcement learning so if you view planning as the problem of going from model to policy i'm basically looking at these little pictures that i um happily uh clipped from my commitments my stock in um you know prl workshop now if you view planning as the problem of going from model to policy reinforcement learning is really going directly from experience to policy so it doesn't think about model is just try to learn the policy from experience planning is try to infer the policy given the model um now there's a version of reinforcement learning called model based reinforcement learning that has this intermediate step which has a model
嗯,所以这基本上就是波兰尼带着他的复仇进入规划领域的方式,而波兰尼进入规划的途径是RL,强化学习,大体上可以这么理解。所以让我用两张幻灯片讲讲规划和尤其是深度强化学习之间的这种联系。所以如果你把规划看成是从模型走向策略的问题——我基本上是在看这些小图,是我很开心地从我在 PRL 研讨会上的报告里剪过来的。现在,如果你把规划看成是从模型走向策略的问题,那么强化学习其实是直接从经验走向策略。所以它不去考虑模型,只是试图从经验中学出策略。规划则是在给定模型的情况下试图推出策略。嗯,现在有一种强化学习叫做基于模型的强化学习,它多了一个中间步骤,也就是有一个模型,
便签引用
34:43
and so you can learn that model in you know vanilla rl you learn that model from just experience and then you use that model uh and to do inference to get the plans okay and then plans become the policy okay so model free order doesn't even bother with the intermediate step if you're thinking about model based rl i would like you to think in terms of just normal learning can go from data to category decisions but then most learning goes from data to hypothesis to category decisions and it is hypothesis intermediate point while it is not theoretically needed winds are providing us a nice way to inject interesting biases about the world in which we live in so similarly just as you can bias the hypothesis we can bias the models that are learned and in fact you know vanilla model rl basically doesn't take anything from humans we'll talk about that in a minute but i would actually argue that we should be considering you know pdl models as the initialization for example partial models as initialization
所以你可以学出那个模型,在那种,你知道的,最朴素的 RL 里,你就只从经验中学出那个模型,然后用那个模型嗯,去做推理,得到计划。好,然后计划就变成了策略。好,所以无模型的做法压根就不操心这个中间步骤。如果你在想基于模型的 RL,我希望你这样来理解:普通的学习可以直接从数据走到类别判定,但大多数学习是从数据到假设、再到类别判定,而正是这个作为中间环节的假设——虽然理论上并不是必需的——最终给了我们一个很好的途径,把关于我们所处世界的有趣偏置注入进去。所以同样地,就像你可以给假设加偏置一样,我们也可以给学出来的模型加偏置,而且事实上,你知道,最朴素的基于模型的 RL 基本上不从人类那里拿任何东西,这个我们等下会谈,但我其实想主张,我们应该考虑把 PDDL 模型当作初始化,比如说把部分模型当作初始化,然后再通过经验去改进它,这才是把专家能够给你的部分知识结合进来的一种
便签引用
35:44
which are then improved through experience and that's a lot more useful way to combine uh partial knowledge that experts are able to give you okay so given that now rl and planning you know i once was talking about the fact that planning an ireland two central areas of ai separated only by a common problem this is what they've said about u.s and uh uk are two great civilizations you know separated by a common language so i think planning and rl really have a lot in common but we don't talk as much because partly because even though the problem is common the tasks we focus when left alone wound up differing quite a bit planning people that is you know i kept flying people tended to look at explicit knowledge planning tasks whereas rl folks mostly tended to focus on the tacit knowledge planning tasks the card pole balancing the grasping manipulation and some of the greatest neatest you know feats of deep reinforcement learning have been in manipulation robotic manipulation because those are
有用得多的方式。好,既然如此,现在 RL 和规划,你知道,我以前说过这样一句话:规划和RL 是被同一个问题分开的 AI 两大核心领域,这就像人们说美国和英国是两个伟大的文明,被同一种语言分开一样。所以我觉得规划和 RL 其实有很多共同之处,但我们彼此交流得不够多,部分原因是,尽管问题是共同的,但当各自独立发展时,我们所聚焦的任务最后差得挺远。规划圈的人,也就是你知道,ICAPS 圈的人,倾向于研究显性知识型的规划任务,而 RL 的人大多倾向于聚焦隐性知识型的规划任务:倒立摆平衡、抓取操作。深度强化学习一些最了不起、最漂亮的成果都出现在操作、机器人操作上,因为那些正是我们不知道怎么
便签引用
36:49
things for which we don't know how to write any easy explicit knowledge schemas and so they are able to learn uh with some you know simulators you know working on simulators we'll get to the simulation a minute too um now polany comes to planning through this deep reinforcement learning these two ways one is why even bother with planning step when you can just learn the policy core model free if you go model free planning is not needed at all there's no intermediate step the second is why bother with explicit taking explicit knowledge when you can learn the model in your own representation the one that you made up this whole representation learning aspect from experience okay which is typically results in learning inscrutable models um both are certainly reasonable stances both of these are certainly reasonable stencils for tacit knowledge tasks i hope you are taking your oranges but quite exotic for complex tasks that involve both explicit and tacit knowledge uh both from the point of view of
写出任何简单显性知识框架的东西,所以他们能够借助一些,你知道的,仿真器来学,在仿真器上做,我们等下也会讲到仿真。嗯,现在波兰尼通过深度强化学习进入规划,有这么两条途径:一条是,既然你可以直接学出策略,为什么还要费劲去做规划这一步呢——核心是无模型的,如果你走无模型路线规划根本就不需要,也没有中间步骤;第二点是,既然你可以自己学,为什么还要费劲去获取显式知识呢,你完全可以用你自己的表示方式去学这个模型——就是你自己造出来的那套表示,也就是从经验中学表示的这一整套东西,而这通常会导致学出来的模型是难以解读的。嗯,这两种立场当然都有道理,这两种立场确实都说得通对隐性知识类的任务来说是合理的,我希望大家能跟上,但对于同时涉及显式和隐性知识的复杂任务来说就相当奇怪了无论是从我们给出的解的鲁棒性角度来看,还是从可解释性角度来看,你知道的
便签引用
37:50
robustness of the solutions that we come up with and the interpretability of the you know the those solutions as well as the operation of our ai agents so that's something that you want to keep in mind and so that's how the problem is revenge is coming to us uh there's a small digression i want to make rl folks for stage reason even though by definition there was nothing about rl which says you should not be taking knowledge from outside because after all the model can be uh biased with a partial you know biased from you know to begin with before the learning happens and there has over a period of time rl essentially has decided that they don't like taking knowledge but then they know that they can't really learn from experience i mean learn to drive from experience would involve that like that car falling and then when you fall i learned from your experience not you you die okay so since getting experienced in the real world raw and tooth and claw can be quite deleterious to the robot's health
这些解本身、以及我们 AI 智能体的运行方式都是如此。所以这是你要记在心里的一点,那么问题就是这样摆到我们面前的。嗯,我想稍微跑个题:强化学习的人出于某种原因——尽管按定义来说,强化学习里并没有哪一条说你不能从外部获取知识,因为毕竟在学习开始之前,模型本来就可以带有偏置、带有部分知识作为出发点但久而久之,强化学习基本上认定他们不喜欢接受外部知识;可他们又知道自己其实没法真的完全从经验中学,我是说,从经验中学开车,那就意味着那辆车翻了,而当你翻车之后我是从你的经验中学到的,不是你自己学到——你已经死了。所以,既然在这个血淋淋、弱肉强食的真实世界里获取经验会严重危害机器人的健康,很多强化学习系统实际上是在外部提供的模拟器里工作的,但他们仍然
便签引用
38:51
many rl systems actually work with externally supplied simulators but still brazil at taking any explicit knowledge it's sort of funny because simulators are made by humans we don't mind taking the simulators but we don't like humans giving us knowledge that's kind of the bitter lesson versions that you can interpret you know from rich certain essay it's highly unlikely that for explicit knowledge domains simulators are easier to provide than partial domain knowledge if you require me to provide a box world simulator i probably will write a blocks word pdl model i will then throw in some sort of an action um you know evaluator you know unless i'm really taking something way beyond what we call blocks word so in general this issue of simulator do realize that simulator is the way humans wind up providing very useful knowledge for rl and that's something to keep in mind in fact something along these lines will be talked about by leslie in her talk in a couple of days later okay so deep reinforcement learning is amazingly
拒绝接受任何显式知识。这其实挺好笑的,因为模拟器也是人写的——我们不介意接受模拟器,但我们不喜欢人给我们知识。这大概就是你可以从 Rich Sutton 那篇《苦涩的教训》里读出来的一种解读。对于显式知识的领域来说,提供模拟器比提供部分领域知识更容易,这几乎是不可能的如果你要我提供一个积木世界的模拟器,我多半会先写一个积木世界的 PDDL模型,然后再加上某种动作评估器,除非我要做的东西远远超出我们所说的积木世界。所以总的来说,关于模拟器这件事,你要意识到:模拟器恰恰是人类最终为强化学习提供非常有用知识的方式,这一点要记住。事实上,Leslie 在她的报告里也会谈到类似的内容就在几天之后。好,深度强化学习在这些扭来扭去的小虫子上表现得惊人地好,你知道的
便签引用
39:57
good with these wiggly worms you know this is the open ai um thing that i just copied from um essentially they can learn how to do locomotion uh quite well and you can't write a pdl model with that okay but on the other hand you're on this other side you have steve chen and koh and jbl want to talk about mission planning mission scheduling people apparently want to talk about drilling and how to do drilling plans it's all that large has lots and lots of planning tasks for which there are no simple simulators ergodic simulators on which you can repeatedly keep you know trying out and figure out how to do the simplistic task so the question then is how do we actually bridge these kind of tacit knowledge tasks with this kind of explicit knowledge task and in fact in between there are tasks like task and motion planning which involve aspects of both and that's really the interesting thing we should be talking about okay now whenever those kinds of things happen neither the bottom up going all the way
这是我从 OpenAI 那边直接拷过来的图,基本上它们能相当好地学会怎么移动运动,而这个你是没法写出一个 PDDL模型来描述的。但另一方面,在光谱的另一端,你有 Steve Chien 那些人,还有 JPL 的人想讨论任务规划、任务调度;还有人显然想讨论钻井,以及怎么制定钻井计划这一大类里有非常非常多的规划任务,对它们来说根本没有简单的模拟器、没有那种可以让你反复不断去试、然后摸索出怎么完成任务的遍历性模拟器。那么问题就来了:我们究竟该怎么把这类隐性知识的任务和这类显式知识的任务连接起来?而且事实上,在两者中间还有一些任务比如任务与运动规划(task and motion planning),它同时涉及两方面,这才是我们真正该讨论的有意思的东西好,那么每当出现这种情况时,不管是自底向上一路往上走,还是自顶向下一路往下走
便签引用
10两条出路:加入浪漫或整合
41:00
up not the top down coming all the way down doesn't make too much sense we have to consider ways of combining the best you know aspects of both directions and that's something that i want to convey in the few remaining minutes i have okay so how do we plan around ai's new fascination with tacit knowledge well first idea is join the romance you know many of you are doing it um okay some of you are doing it um basically after all every explicit knowledge task can be converted into a tacit knowledge one if you just try a little okay so for example you can take pdl models take pictures of it and then it will become an image you can have cnns gnns transformers or even performers that say came out yesterday not to analyze them to recover the model um in fact i mean of course transformative performance will be useful in more sequential data so for example use pdl model and ff to generate a ton of planet traces submit these traces to some sequence learning algorithm again one of your favorite ones
都不太说得通。我们必须考虑如何把两个方向各自最好的部分结合起来,这也是我想在剩下的几分钟里传达的。那么,面对 AI 对隐性知识的这股新迷恋,我们该怎么规划呢?第一个想法是加入这场浪漫——你们中很多人正在这么做,嗯,好吧,你们中有些人在这么做。毕竟,每一个显式知识任务只要你稍微努力一下,都可以被转化成一个隐性知识任务。比如说,你可以拿 PDDL 模型,给它拍张图它就变成了图像,然后你可以用 CNN、GNN、Transformer,甚至昨天刚出的 Performer 去分析它们,把模型恢复出来。嗯,当然,Transformer 和 Performer 在更偏序列的数据上会更有用,所以比如说用 PDDL 模型和 FF 生成一大堆规划轨迹,把这些轨迹丢给某个序列学习算法,还是你最喜欢的那几个;或者赶上「从像素到决策」的潮流,给积木世界状态转移拍照,看看我们能不能
便签引用
42:04
jump on the pixels to decision bandwagon take pictures of the transitions of the block world configurations and try to see if we can somehow recover some sort of blocks world model that can be fed to ff and then you know pat yourself very happily on your back um these are a base in fact there are papers for all of this i just didn't want to specifically name those papers but many of you probably know these papers if you ask me nicely on dm i'll send you citations first people who have published these papers i let me just say i'm not a huge fan of this direction because in some sense you are trying to just go bottom up take something which has explicit knowledge and they're trying to solve it as a test knowledge task of course it can be done but what is the advantage you know it seems like it actually has all these other issues such as interpretability robustness aspects coming in the other one which is the recommended one and this is actually the dense slide and then i'll you know look at some pieces of it
以某种方式恢复出某种积木世界模型,然后喂给 FF,接着你就可以美滋滋地拍拍自己的肩膀了嗯,这些其实都有对应的论文,我只是不想具体点名那些论文但你们很多人大概都知道这些论文;如果你好好地私信问我,我可以把引用发给你,先给那些发表了这些论文的人我只想说,我不太喜欢这个方向,因为在某种意义上你只是在一味地自底向上:把本来有显式知识的东西,硬当成隐性知识任务来解。当然这是做得到的,但好处在哪里呢?你知道,它看起来反而带来了各种其他问题,比如可解释性、鲁棒性方面的问题。而另一个方向,也就是我推荐的那个方向——这其实是张信息密集的幻灯片,接下来我会挑其中几点讲讲,希望你们中有些人会考虑
便签引用
43:04
which i hope some of you will consider again there are many many good things that can be done these are things that sort of occurred to me for my small brain um you really want to combine and make the sum greater than the parts okay look at for example tasks that involve both explicit and tacit knowledge components motion and task planning that's been a great area i would love to see how those will benefit um you know some of the work that i really like uh the kind of work that leslie and her folks do um in season they're actually seriously realizing that you want to combine these things and there's a nice talk that mike litman was giving michael litman was giving the other day which is also sort of focuses on these things um what while looking at those kinds of directions look at making the model acquisition task easier with learning we always talk about this there are better and worse ways of improving the model acquisition i'll tell you a little more about this but certainly these days you can't say learning is not
再说一次,可做的好事情非常非常多,这些只是以我这个小脑袋所能想到的一些嗯,你真正想要的是把两者结合,让整体大于部分之和。好,比如说,去关注那些同时涉及显式和隐性知识成分的任务,运动与任务规划就是一个很棒的领域,我很想看到它们能从中获益多少嗯,我特别喜欢的一些工作,比如 Leslie 和她的团队做的那类工作,他们是真的在认真地意识到你需要把这些东西结合起来;还有 Michael Littman 前几天做的一个很不错的报告也大致聚焦于这些问题。嗯,在关注这些方向的同时,也要关注怎么用学习让模型获取这件事变得更容易。我们总在谈这个,改进模型获取有好方法也有坏方法,我待会儿会再多说一点。但可以肯定的是,如今你不能再说学习没解决,所以我们不能对这个问题视而不见,这
便签引用
43:59
solved so we have to somehow not look at this problem so that's definitely worthwhile looking at looking at planning for partially specified and incomplete models it winds up being a related problem to actual model acquisition if you admit that the model that you're working with is incomplete then you really have to talk about robust planning because it's not a it's no longer a situation of find an optimal plan with respect to this known to be correct model that's not any longer true look at interpretability issues especially when the systems have both explicit and tacit knowledge i love xaip community they are doing very good job in terms of making interpretability and explanations um then there is sort of a shared vocabulary because it turns out even then there is a huge amount of work to do in terms of explanations and interpretable behavior but when there is no shared vocabulary that becomes even more challenging there are interesting directions that are being taken i will talk a little
绝对值得关注。再有就是针对部分指定和不完整模型的规划,它最终会变成一个与模型获取相关的问题。如果你承认你手上的模型是不完整的,那你就真的必须谈鲁棒规划了,因为这已经不再是「针对这个已知正确的模型找一个最优计划」的情形了,那个前提不再成立。再看可解释性问题,尤其是当系统同时具备显式和隐性知识的时候。我很喜欢 XAIP 这个社区,他们在可解释性和解释生成方面做得非常好。嗯,然后还有共享词汇表的问题,因为事实证明,即使有共享词汇表,在解释和可解释行为方面也还有大量的工作要做;而当没有共享词汇表时,就更具挑战性了。目前已经有一些有意思的方向在被探索,我待会儿会
便签引用
44:57
about this but certainly those are directions worth looking at and of course some of the other things that are going on which i think is great are looking at compiling control knowledge while learning we always looked at deriving heuristics with you know clean principles but really you can try to learn heuristics from traces of from experience of the planner that is something that is easy to do and that's worthwhile to do of course you won't have the optimality guarantees and informedness guarantees but really you know life isn't having those guarantees and so these are useful things and lots of people are working on this even in this conference there are papers on this that's a great direction and finally look at generalizing planning technology in cases for non-declarative representations then especially for those places where you are stuck only with a simulator people don't want to write you a domain model there are some people have already started working on this too and these are i think
稍微讲一点,但这些方向确实值得关注。当然还有一些正在进行的其他工作,我觉得也很棒,就是在学习的同时编译出控制知识。我们过去总是用干净的原理去推导启发式,但其实你可以尝试从规划器的轨迹、从它的经验中学习启发式,这是很容易做的,也很值得做。当然,你不会有最优性保证和信息性保证,但说真的,生活本来就没有这些保证,所以这些都是有用的东西,而且很多人都在做这方面的工作,就连本次会议上也有相关论文,这是个很好的方向。最后,看看怎么把规划技术推广到非陈述式表示的情形,尤其是那些你只能拿到一个模拟器的场合别人不愿意给你写领域模型。已经有人开始做这方面的工作了,我认为这些是
便签引用
11模型获取、不完整模型与解释
45:54
recommended directions they are worthwhile looking at um so next what i'm sorry next what i want to do yeah next what i want to do is kind of talk about three of them quickly and then wrap up so learning for model acquisition and refinement learning models from scratch from experience alone is unlikely to result in robust and interpretable models so this picture that michael litman had i would just basically say we should be allowing declarative bias partial models into this learning stage people in machine learning talk a lot about inductive biases but oftentimes they are stuck yet with inductive bias being the topology of the network which is a fine but very primitive way to tell you what i want you to learn um i should be able to give background knowledge which you then improve further over experience so it's much better to provide declarative bias in terms of the model learner in the form of partial models which are then refined while learning it's not these are not work that's done
值得推荐、值得关注的方向。嗯,那么接下来我想做的——不好意思,接下来我想做的,对,接下来我想做的是快速讲其中三点,然后收尾。首先是面向模型获取与精化的学习:只从经验、从零开始学模型不太可能得到鲁棒且可解释的模型。所以对于 Michael Littman 那张图,我基本上会说,我们应该允许把陈述式偏置、部分模型引入到这个学习阶段。机器学习的人经常谈归纳偏置,但很多时候他们的归纳偏置就只是网络的拓扑结构,这固然可以,但它是一种非常原始的方式来告诉你我希望你学什么。嗯,我应该能够给出背景知识,然后你在经验中把它进一步改进。所以更好的做法是,以部分模型的形式给模型学习器提供陈述式偏置,然后在学习过程中不断精化它。这些还不是已经做完的工作,我只是说这些是你们也许该关注的方向
便签引用
47:03
i'm just saying these are directions that you might want to look at um that way essentially you can start with like an okay model that you got from the human and you are improving it over time uh by from experience and you know it also allows for interpretability that way because you're still sticking to something that the human started you off on learning models from text meant for human consumption is also an extremely fruitful direction this is a great way we can interface with all the advances in nlp right now because you can take long long back people used to do this but that time nlp technology sucked now nlp technology is much better so you should be able to read recipes and then get essentially some partial models which then can then be used possibly as the beginning stage for this model learning um so we actually have some work in each guy 2018 that you might look at but there are other people doing this sort of work that's very useful one thing that i would not suggest you do is recovering pdl models from ff traces
嗯,这样一来,你基本上可以从人给你的一个「还行」的模型出发,然后随着时间通过经验不断改进它;而且你知道,这样也带来了可解释性,因为你始终守着人一开始给你的那个东西。另外,从写给人看的文本中学习模型,也是一个极其有前景的方向这是我们与当前 NLP 各种进展对接的绝佳途径,因为很久以前人们就试过这么做,但那时NLP 技术很糟糕,现在 NLP 技术好多了,所以你应该能读懂菜谱,然后基本上得到一些部分模型,这些模型接着可以作为模型学习的起点。嗯,我们在 IJCAI2018 上其实有一些相关工作,你可以看看;不过也有其他人在做这类工作,非常有用。有一件事我不建议你去做,那就是从 FF 的轨迹里恢复 PDDL 模型:既然你已经有 PDDL 模型了,那你为什么还要用 FF 生成一大堆
便签引用
48:07
if you have the pdl model why are you then using ff to generate a whole bunch of plans if only to you know invert these traces back into the model of course you can say because i will that way get a paper into the icaps planning learning track but that's not a good reason to do it so that's a very exotic thing that's not a worthwhile thing to do uh planning with incomplete models minds are being quite relevant here because once a learner realizes that the model that it has a planner realizes that the model it has is not guaranteed to be correct in any real sense it needs to start thinking about robustness it needs to start thinking about taking into account the fact that it has ignorance about the model uh and so this actually leads to things like representations for incomplete models uh you know i've been sort of in my group looking at not just strips which is sort of a fully causal and certified correct model from the user towards models that are more and more shallow and less and less
计划,就为了把这些轨迹反推回模型呢?当然你可以说,因为这样我能在ICAPS 的规划与学习专题里发一篇论文,但这不是一个好理由。所以那是一件很奇怪的事,不是一件值得做的事。嗯,然后是带不完整模型的规划,这在这里就相当相关了,因为一旦学习器意识到——或者说规划器意识到——它手上的模型在任何实际意义上都不保证是正确的它就必须开始考虑鲁棒性,必须开始把「它对模型存在无知」这一事实考虑进来。嗯,这就引出了不完整模型的表示等问题。嗯,你知道,我在我的组里一直在关注:不只是 STRIPS 这种完全因果、由用户认证为正确的模型,而是走向那些越来越浅、越来越不保证正确和完整的模型,然后讨论怎么在此之上做规划
便签引用
49:06
guaranteed to be correct and complete and then talk about how to do planning robust planning with respect to that and so you need to think about the kinds of representations for these kinds of models and you need to think about what does it mean to do robust planning over these representations now some examples from our own group include domain models with possible preconditions and effects that one one has written a paper on ai journal in 2017 and of course planning problems with uncertain reward metrics which is where actually a lot of work that's even now popular about diverse plants has come about um you know it's one of the earlier works so those are i think quite relevant things um once you start thinking about multiple models in some sense this issue of model uncertainty you can take a bayesian view of it and say that when you are uncertain about the model that means you really have many complete models one of which is the true model and you so you actually are doing a bayesian account
做鲁棒规划。所以你需要思考适合这类模型的表示形式也需要思考在这些表示之上做鲁棒规划究竟意味着什么。我们自己组里的一些例子包括带「可能前提条件和可能效果」的领域模型,这方面有人在 2017 年的 AI Journal 上写过一篇论文;当然还有奖励度量不确定的规划问题,其实现在很流行的很多关于多样化计划(diverse plans)的工作就是从这里来的。嗯,你知道,这是比较早期的工作之一,所以我认为这些都相当相关嗯,一旦你开始考虑多个模型——某种意义上就是模型不确定性这个问题——你可以采取贝叶斯的视角说:当你对模型不确定时,就意味着你其实有很多个完整模型,其中之一是真实模型。所以你实际上是在做规划的贝叶斯处理,一般来说这就是做鲁棒规划的方式
便签引用
50:00
of planning that's the way to do robust planning in general um and so it turns out that when you're thinking about human eyewear planning something that much closer to my heart it turns out that this becomes even more of an issue so you now have a robot with an mr which in itself can be incomplete so it can be actually a set of models and then the human has an approximation of the robots model which then the robot is trying to estimate so this m hat rh will again be a distribution of models and essentially the robot is stuck with doing multi-model planning to figure out whether or not its behavior is interpretable so this is quite relevant even from those directions in fact there is a nice paper uh by sharath um in xaip this year on bayesian account of interpretability measures that you might look at um actually i should probably mention that you know somebody this morning was asking somebody in the doctoral consortium who was somebody who's doing good explainability work and sequentialization making
嗯,事实证明,当你在思考人机协作规划(human-aware planning)时——那是我更加钟情的方向——事实证明这个问题就变得更加突出。现在你有一个机器人,它有自己的模型 M^R,而这个模型本身可能就是不完整的,所以它实际上可能是一组模型然后人对机器人的模型又有一个近似,而机器人还要去估计这个近似,所以这个 M^R_H 的估计同样会是一个模型上的分布。于是机器人基本上只能做多模型规划,来判断它的行为是否可解释。所以从这些角度看,这也是相当相关的。事实上,今年 XAIP 上有一篇很不错的论文是 Sarath 写的,关于可解释性度量的贝叶斯处理,你们可以看看。嗯,其实我大概应该提一下今天早上有人在博士生论坛上问一位在可解释性方面做得很好的同学——他做的是可解释性和序贯决策问题——这个人问:我不明白这跟 XAI 有什么关系。我不太
便签引用
51:05
problems this person asked i don't see why it is connected to xai i don't quite know why they asked it but some people tend to think that xai essentially means explainable machine learning with inscrutable representations and i want you to kind of realize that's just a very small part of the spectrum xci is hard but mostly as a debugging tool for inscrutable representations pointing explanations are quite primitive if we have to point is the only way we can talk to each other we would not have had the civilization we have explanations between humans typically tend to be very critical for collaboration but they are not pointing and they are not a solid lucky by the agent so the kind of work that's actually being done in xap community is very much relevant um in in the in the um you know this whole interpretability direction um so that's something that you might want to keep in mind and again as i i kind of tongue-in-cheek ask if you take this adversarial example where this school bus with some noise becomes an
清楚他们为什么这么问,但有些人倾向于认为,XAI 本质上就是指对具有难解表示的机器学习做解释而我希望你们意识到,那只是整个谱系中很小的一部分。XAI 很难,但它主要是作为难解表示的一种调试工具;而「指点式」的解释是相当原始的。如果指点是我们彼此交流的唯一方式我们不可能拥有今天这样的文明。人与人之间的解释通常对协作极其关键,但它们不是靠指点的,也不是由智能体单方面给出的。所以 XAIP社区实际在做的那类工作,在整个可解释性方向上是非常相关的。嗯,这是你们可能要记住的一点。再说一次,就像我半开玩笑问过的:如果你拿这个对抗样本,就是那辆加了噪声之后、在当前大多数深度学习视觉系统看来就变成鸵鸟的校车,你能不能问视觉
便签引用
52:07
ostrich for most of the current um you know deep learned vision systems can you ask the vision system tell me which part of this particular second school bus is making you think it is an ostrich and how possibly useful can it be in in essence right so pointing explanations is okay if you're an animal but you know we are humans and we really have to go beyond that and we want to keep that in mind one last thing before i summarize is handling differing vocabularies is important much of the work in xaip is done with essentially two pdl models maybe with differing reconditioned effects kinds of things eventually we should be able to extend this in cases where parts of the models that are being used by the machine are you know sort of not they're inscrutable they're not in the pdl style models but then the machine is willing to translate its explanation in a language that you understand there's an interesting work that's going on in both exact not explainable machine learning community and then some of which we are doing
系统:告诉我这第二张校车图里到底是哪一部分让你觉得它是鸵鸟?而这在本质上又能有多大用处呢?对吧。所以指点式的解释,如果你是动物那还行,但我们是人,我们真的必须超越那个层次,这一点要记住。在我做总结之前,最后一件事是:处理词汇表不一致很重要。XAIP里的大部分工作基本上都是用两个 PDDL 模型来做的,可能前提条件和效果有所不同之类。最终我们应该能把这个扩展到这样的情形:机器所使用的模型的某些部分,你知道,并不是……难以理解,它们不是 PDDL 风格的模型,但机器愿意把它的解释翻译成一种你能理解的语言。在可解释机器学习社区里,有一些很有意思的工作正在进行其中一些是我们组在做的,呃,有一篇论文就涉及这个,本质上就是你提供
便签引用
12总结:宁静祷文与整体大于部分
53:12
in our group uh there's a paper that character involved where essentially you provide the explanations even if you are doing reasoning with the black box model you provide the explanations in terms of concepts that the humans understand and you the machine understand these concepts by learning the denotation over whatever is this inscrutable representations that you are looking at so this is doable and worthwhile doing last but two slides so um jeff hinton and rich saturn are both saying intelligent design is crude we should just be waiting for learning to happen um my so the question of course a reasonable question can be asked that we came into this world essentially we evolved to this point essentially um to from animal intelligence to human intelligence from america intelligence to animal intelligence to human intelligence and developing a pretty nifty system too along the way and we even managed to make sense to each other sometimes right so why do i need any kind of explicit knowledge task maybe i can just
解释,即使你是在用黑箱模型做推理,你也用人类能理解的概念来给出解释而机器理解这些概念的方式,是在你所面对的那些难以解读的表示之上学习它们的指称所以这是可行的,也是值得做的。还剩两张幻灯片,呃,Geoff Hinton 和 Rich Sutton 都在说,智能设计太粗糙了,我们应该干脆等着学习自己发生。呃,我,当然有人会提出一个合理的问题:我们来到这个世界,本质上我们演化到了今天这个地步,本质上,呃,从动物智能到人类智能,从微生物智能到动物智能再到人类智能,一路上还发展出了相当巧妙的系统,我们甚至有时候还能彼此听懂对方,对吧?那我为什么还需要任何形式的显式知识任务呢?也许我可以完全靠,你知道的,自下而上把这整件事搞定
便签引用
54:23
do this whole thing from you know from bottom up um i don't necessarily want to argue that it cannot be done um but it might get done by the time of rapture and i want ai systems to be able to work with asset as well as explicit models before then so think about it industry that you want to have these systems working right now rather than saying i want the wiggly worm to start doing nasa mission planning and let's see how it's going to happen um even with the fast compute power it's just not at all clear how long that's going to take and it's this whole entire issue of whether or not when they do that their behavior will be interpretable so we would really like to have interpretable ai systems with planning capabilities a little earlier than the rapture time so summary the expecting explicit models for tacit knowledge tasks is as silly as rebuffing easily available partial models for explicit knowledge tasks if you ever ask people can you give me a pdl model for how you see the objects that's just
呃,我并不是一定要论证这做不到,呃,但等到它做成的时候可能已经是末日审判了,而我希望 AI 系统在那之前就既能用隐性知识、也能用显式模型来工作。所以想想产业界吧你希望这些系统现在就能用起来,而不是说,我要让一条扭动的蠕虫开始做 NASA 的任务规划,然后看看它会怎么发生。呃,就算有飞快的算力,也完全不清楚那要花多长时间,而且还有一个整体性的问题就是当它们真做到了的时候,它们的行为是否可解释。所以我们真的希望拥有具备规划能力的可解释 AI 系统而且要比末日审判早一点。所以总结一下:对隐性知识任务期待显式模型跟在显式知识任务上拒绝唾手可得的部分模型一样傻。如果你去问人:你能给我一个 PDDL 模型吗来说明你是怎么看见物体的?这就跟他们说,让我只从像素变化里搞清楚你是怎么玩积木世界的一样傻,完全说不通
便签引用
55:36
as silly as them saying let me figure out how you do blocks world just from pixel transitions makes no sense because there's somebody who can actually tell you how it's done and you can combine that there are many fruitful directions for planning around ai's current romance with tacit knowledge does without gratuitously adding 12 grams of performer or transfer networks to mission planning tasks i'm all for combining these technologies but i'm not a big fan of doing it just for getting the papers which brings me to why i'm showing my students slide because those are the four guys who have to deal tolerate my heroines whenever they want to say wow can we just put in a couple of graph neural network algorithms you know for fun of it you know and i'll just say tell me why tell me exactly how does this matter and that is for that as well as for at least actually letting me arrange them you know and also i kind of tried many of these arguments on them i thank them and i'll show you them and of course the purple robot is hopefully
因为明明有人能直接告诉你这是怎么做到的,而你可以把两者结合起来。围绕 AI 当下对隐性知识的迷恋在规划方面有很多富有成果的方向可以走,而不必没来由地往任务规划里塞上 12 克的 Transformer 或者 Transformer 网络我完全支持结合这些技术,但我不太喜欢仅仅为了发论文而这么做这就说到了我为什么要放这张我学生的幻灯片,因为就是这四个人得忍受我的这些说教,每当他们想说,哇,我们能不能就放几个图神经网络算法进去,你知道的,图个好玩,然后我就会说,告诉我为什么,告诉我这到底有什么意义。为此,也为了他们至少真的让我把他们排好队,你知道的,还有我也算在他们身上试过很多这些论点,我感谢他们,我把他们展示给你们看。当然,那个紫色的机器人希望能
便签引用
56:38
learning from my students me and rest of us humans on how to grant its energy to accept the thing it cannot learn and data to learn the things it can and the wisdom to know the difference that said i will end with the summary slide of join the romance which i am not particularly a fan of and combine to make this some greater than parts part that i think is more fun to do and i thank you for your attention
从我的学生、我以及我们其他人类身上学到,如何把它的能量用于接受它学不了的东西用数据去学它能学的东西,以及分辨两者的智慧。话虽如此,我最后用这张总结幻灯片结束:加入这场我并不特别买账的浪漫,然后去结合,让整体大于部分之和,我觉得这部分更有意思去做。谢谢大家的聆听。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

Rao Kambhampati 在 ICAPS 2020 的邀请报告中指出,AI 界正沉迷于"隐性知识"(tacit knowledge)任务并鄙视人类给出的显性模型,而规划社区应当拒绝盲目跟风,把易于获得的显性(部分)模型与学习结合起来,让"整体大于部分之和"。

核心要点

  • Hinton 的"智能设计 vs 学习"二分法映射到规划社区就是"被告知 vs 学习"。 2019 年 Hinton 在图灵奖演讲中把"人类想清楚如何操作表示、再详细告诉计算机"贬称为"智能设计"(借用反进化论的贬义词),把"给大量输入输出样本让机器学映射"称为学习。而规划社区恰恰是给计算机 PDDL 模型、让机器做组合搜索的那一派。
  • Polanyi 悖论:我们知道的比能说出的多。 匈牙利博学者 Michael Polanyi 在《The Tacit Dimension》中定义隐性知识为"会做但说不出怎么做"的知识(如视觉、语音、行走)。Polanyi 当年抱怨学界过度关注显性知识任务;而"想精通就去教""想理解就去编程"这类格言只对显性知识任务成立——写过 Transformer 或 CNN 的人并没有因此更懂人类如何看世界。
  • AI 的发展顺序与人类儿童恰好相反。 儿童先展现感知/操作智能,再是情感、社交,最后才是认知推理;而 AI 在 Deep Blue 击败卡斯帕罗夫时还无法识别棋盘上的棋子。原因是显性知识任务能写出(哪怕不完美的)部分模型再加搜索,而视觉任务写不出任何算子。他嘲讽道:谁若宣称"AI 未来不仅会学习还会推理",此人必是 AlexNet 之后才从新闻标题进入 AI 的。
  • 推理为中心 vs 学习为中心是两种互补的智能体设计路线。 前者假设模型和表示已给定,把学习"以后再说",适合有好模型的显性领域;后者(如 subsumption 架构)不预设模型、先学反射式行为,把推理"以后再说",适合无模型但有大量经验的隐性领域。引用 Andrew Ng:"人几秒能做的事计算机都能做"——但若人只做几秒内的事,我们就不会有今天的文明。
  • 系统一/系统二与显隐性知识不等同,差别在于可解释性。 隐性知识多由系统一处理;显性知识可从系统二起步再"编译"进系统一(如微分从极限定义到反射式计算)。区别类似手写汇编与从高级语言编译来的汇编:后者出错时可在高级语言层定位,这正是符号调试器的原理。他还回顾 1989 年 AI Magazine 上 Ginsberg 批评 Schoppers "通用规划"是"几乎普遍糟糕的主意",指出当年争的是空间-在线计算权衡,如今大家乐于用 TB 级策略表换取快速响应,但这与"起点是否有显性知识"是不同问题。
  • "Polanyi 的复仇":AI 已从"所有模型都错、有些有用"摆到"要模型干什么"。 现在的口号仿佛是"给我足够多 GPU、足够大数据集和足够深的网络,我就能造出 AGI"。甚至出现了把数独这类有显性模型的问题转成海量样本、再用深度学习"恢复"原知识的"从知识到数据的家庭作坊"。其后果是可解释性差、易受对抗攻击、数据集偏见(GPT-3 对"两个穆斯林"的补全反复关联暴力,因为它学的是人类上传到网络的"集体潜意识",缺少系统二的抑制)。
  • 规划社区曾走向另一极端:期待"超人类的人类"。 混合主动规划时代曾指望人类告诉规划器"第 7539 号节点该不该剪枝",还要求提供控制规则和抽象信息。他改写宁静祷文:"请赐我宁静接受学不到的,赐我数据学能学的,赐我智慧分辨两者。"近似准则是:显性知识唾手可得时就该用,问题在于如何与隐性知识桥接。
  • Polanyi 通过强化学习进入规划,且 RL 对"知识"的态度自相矛盾。 规划是"模型→策略",RL 是"经验→策略",基于模型的 RL 中间有个学来的模型。RL 定义上并不禁止注入先验,但社区文化拒绝人类知识,却乐于接受人类编写的模拟器。他指出:对显性领域而言,模拟器不可能比部分领域知识更容易提供——让他写积木世界模拟器,他必然先写 PDDL 模型。RL 的亮点(OpenAI 的"扭动虫"运动控制、机械臂抓取)都是写不出 PDDL 的隐性任务,而 NASA 任务调度、钻井规划等则没有可反复试错的遍历式模拟器。
  • 不推荐"加入浪漫",推荐"合并求和大于部分"。 反面例子:把 PDDL 模型截图喂 CNN/GNN/Transformer;用 PDDL+FF 生成大量轨迹再用序列学习"反推"模型;从积木像素转移中恢复模型再喂给 FF——这些论文都存在(他拒绝点名,可私信索取引用)。推荐方向:任务与运动规划(TAMP)这类显隐混合任务;以部分 PDDL 模型作为"声明式偏置"初始化模型学习(比只把网络拓扑当归纳偏置强得多);从人类可读文本(菜谱等)学习部分模型(IJCAI 2018 工作);对不完整模型做鲁棒规划(AIJ 2017 的可能前提/效果模型、不确定奖励度量下的多样规划、贝叶斯多模型规划);从规划器经验中学习启发式(放弃最优性保证换取实用性);将规划技术推广到只有模拟器的非声明式表示。
  • 可解释 AI 不等于对不可解读表示的"指向式解释"。 人类之间的解释是协作的核心,而不是"指着某块像素"。问视觉系统"校车加噪声后哪部分让你认为是鸵鸟"几乎无用。XAIP 社区目前多用两个 PDDL 模型(前提效果不同)做解释,下一步应处理词汇表不同的情形:机器用黑盒模型推理,但通过学习概念在黑盒表示上的指称,用人类理解的概念给出解释。人机协同规划中,机器人自身模型 M_R 可能是一组模型,人类对机器人模型的估计 M̂_RH 又是一个分布,机器人必须做多模型规划来判断行为是否可解释(参见 Sreedharan 在 XAIP 2020 的贝叶斯可解释性度量论文)。

结论与值得注意的细节

  • 核心结论: 为隐性知识任务索要显性模型,与拒绝显性知识任务中唾手可得的部分模型,是同等愚蠢的。前者如问人"给我一个你如何识别物体的 PDDL 模型",后者如说"让我从像素转移中自己弄懂积木世界",而明明有人能直接告诉你。
  • 对"完全自底向上也许能行"的回应: 他不否认从进化式学习最终能涌现推理,但那可能要等到"末日审判"(rapture),而工业界现在就需要能同时处理显隐性知识、且行为可解释的规划系统,而不是等"扭动虫"学会 NASA 任务规划。
  • 对研究动机的提醒: 他支持技术融合,但反对"无端往任务规划里加 12 克 Performer 或 Transformer"只为发论文;他对学生的一贯要求是"告诉我这为什么重要、具体怎么起作用"。
  • 报告定位: 这不是写论文/做报告的辅导(2009、2013 版已在 YouTube),也不是他本人 XAIP 研究的综述(见其 AI Magazine 综述及次日三场 XAIP 报告),而是给规划研究者的"声明式偏置"式思想演讲,带 ICAPS Festivus(他 2005 年发起)的戏谑风格。
  • 趣味细节: 演讲开头他建议每听到 tacit 或 explicit 就喝一口橙汁,"一年的维生素 C 就够了";配套观点文章《Polanyi's Revenge and AI's New Romance with Tacit Knowledge》发表于 2021 年 2 月 CACM,其网页可下载;他引用"规划与 RL 是被同一个问题隔开的两个 AI 核心领域",仿"美英是被同一种语言隔开的两个伟大文明";Leslie Kaelbling 数日后的报告将进一步讨论模拟器作为人类知识载体的问题。
核心句型 · 10
1. There are two ways to X: one is A and the other is B
“There are two ways to make a computer do what you want one is intelligent design and the other is learning”
经典二分开场句式,先立框架再逐一展开。适合演讲或文章开头。仿写:There are two ways to learn a language: one is immersion and the other is instruction.
2. Show me an X who …, and I'll show you someone who …
“Show me an ai expert confidently proclaiming that in future ai will not only just learn but will also be able to reason too and i'll show you someone who entered ai via newspaper headlines”
讽刺性反驳句式:用「给我看 A,我就给你看 B」揭示某类人的真实来历或动机。语气锋利,适合辩论与评论。
3. X and Y are two … separated only by a common Z
“Planning and rl two central areas of ai separated only by a common problem”
化用萧伯纳名句「被同一种语言分开的两个国家」,用悖论式表达指出本应相近却彼此隔阂的两方。引用经典再改一个词,是幽默又有学识感的写法。
4. Doing X is as silly as doing Y
“Expecting explicit models for tacit knowledge tasks is as silly as rebuffing easily available partial models for explicit knowledge tasks”
对称类比句,把两种极端并列,暗示中间道路才合理。适合总结段落。仿写时注意两边结构平行。
5. X is really just doing Y, if only in a … form
“Hinter is really just reinforcing the ai zeitgeist if only in sort of a doctrinal farm”
if only in … form 表示「哪怕只是以某种形式」,用于承认某事成立但限定其方式。学术评论中常用的克制性让步。
6. I'm all for X, but I'm not a big fan of doing it just for Y
“I'm all for combining these technologies but i'm not a big fan of doing it just for getting the papers”
先表支持再划界限的表态句式。be all for 表全力支持,not a big fan of 是温和否定。适合表达有条件的赞同。
7. I don't necessarily want to argue that it cannot be done, but …
“I don't necessarily want to argue that it cannot be done but it might get done by the time of rapture”
让步式反驳:不否认可能性,转而质疑时间或代价。学术讨论中避免绝对化又保持立场的常用手法。
8. Why (even) bother with X when you can just Y?
“Why even bother with planning step when you can just learn the policy”
反问句式,用于转述或质疑某种「省事」逻辑。bother with 表「费心去做」,此处讲者用它复述对方立场再加以反驳。
9. It turns out that …
“It turns out that for these explicit knowledge does the progress in reasoning and cognitive intelligence happened much faster”
引出事实或结论的高频口语衔接语,暗示「细想之后会发现」。演讲中可用来自然过渡到论据。
10. Grant me the serenity to accept …, and the wisdom to know the difference
“Grant me the serenity to accept the things that i cannot learn and data to learn the things i can and wisdom to know the difference”
改编《宁静祷文》的三段式结构,替换关键词即可表达「接受不可改变者、改变可改变者、分辨二者」。适合作结论的金句模板。
词汇精讲 · 100 · 按出现顺序
multitudes of /ˈmʌltɪtuːdz/ phr. 0:00
大量的、无数的(正式用法)
mouthful /ˈmaʊθfʊl/ n. 1:27
拗口冗长的词语或名称
psyched up /saɪkt ʌp/ phr. 1:27
(使自己)兴奋起来、进入状态
bomb /bɑːm/ v. 3:08
(演出、报告)彻底失败、砸锅(美式口语)
wound up /waʊnd ʌp/ phr. 3:08
最终(做了某事),wind up doing 的过去式
pottering around /ˈpɑːtərɪŋ/ phr. 4:14
闲逛、慢悠悠地做琐事
declarative bias n. 5:12
声明式偏置:以显式陈述的知识作为学习或思考的先验倾向
swig /swɪɡ/ n. 6:22
大口喝、一大口(饮料)
intriguing /ɪnˈtriːɡɪŋ/ adj. 6:22
引人入胜的、令人好奇的
pilgrimage /ˈpɪlɡrɪmɪdʒ/ n. 6:22
朝圣;此处戏指专程前往聆听大师演讲
a natural /ˈnætʃərəl/ n. 7:36
天生的好手、有天赋的人
excruciating /ɪkˈskruːʃieɪtɪŋ/ adj. 7:36
极其痛苦的;此处指「事无巨细到令人痛苦」
deniers /dɪˈnaɪərz/ n. 7:36
否认者(如 evolution deniers 进化论否定者)
zeitgeist /ˈzaɪtɡaɪst/ n. 8:37
时代精神、时代思潮(源自德语)
doctrinal /ˈdɑːktrɪnl/ adj. 8:37
教条的、教义式的
feats /fiːts/ n. 8:37
壮举、非凡成就
the masses /ˈmæsɪz/ n. 9:38
普通大众、民众
have no clue phr. 9:38
毫无头绪、完全不知道
polymath /ˈpɑːlimæθ/ n. 10:40
博学家、通才
lamenting /ləˈmentɪŋ/ v. 11:40
哀叹、惋惜(lament 的现在分词)
verbalize /ˈvɜːrbəlaɪz/ v. 11:40
用言语表达、说出来
inculcated /ɪnˈkʌlkeɪtɪd/ adj. 11:40
被反复灌输而根深蒂固的
perceptual /pərˈseptʃuəl/ adj. 12:55
感知的、知觉的
creaming /ˈkriːmɪŋ/ v. 13:56
(口语)彻底击败、打得落花流水
proclaiming /prəˈkleɪmɪŋ/ v. 13:56
宣称、公开声明
foolproof /ˈfuːlpruːf/ adj. 15:04
万无一失的、不会出错的
lowest common denominator phr. 16:03
最低共同标准(原为数学「最小公分母」)
orthogonal /ɔːrˈθɑːɡənl/ adj. 16:03
正交的;引申为互不相关、独立的
collective subconscious n. 17:07
集体潜意识(此处指人类在网上留下的全部内容)
inference /ˈɪnfərəns/ n. 17:07
推理、推断
subsumption architecture n. 18:05
包容式架构:Brooks 提出的分层反应式机器人架构
apriori /ˌeɪ praɪˈɔːraɪ/ adj. 18:05
先验的(规范拼写为 a priori)
in the foreground phr. 18:05
处于前景、成为关注重点
reflex agent n. 19:01
反射式智能体:仅依据当前感知直接做出动作
metaphorical /ˌmetəˈfɔːrɪkl/ adj. 20:03
隐喻的、比喻性的
reflexive /rɪˈfleksɪv/ adj. 21:07
反射性的、不假思索的
deliberative /dɪˈlɪbərətɪv/ adj. 21:07
审慎的、深思熟虑的
compiled into phr. 21:07
被编译成;此处比喻显性知识固化为自动反应
localize /ˈloʊkəlaɪz/ v. 22:10
定位(错误所在)
symbolic debuggers n. 22:10
符号调试器:在源码层面而非机器码层面定位错误的工具
raging controversy /ˈreɪdʒɪŋ/ phr. 22:10
激烈的争论
memoize /ˈmemoʊaɪz/ v. 23:13
记忆化:缓存计算结果以免重复计算(计算机术语)
combinatorics /ˌkɑːmbɪnəˈtɔːrɪks/ n. 23:13
组合数学;此处指组合爆炸的规模
tradeoffs /ˈtreɪdɔːfs/ n. 23:13
权衡、取舍
fulcrum /ˈfʊlkrəm/ n. 24:14
支点
doctrine /ˈdɑːktrɪn/ n. 25:20
教义、信条;此处指成文的显性知识
bazillion /bəˈzɪljən/ n. 26:24
(口语夸张)无数、极大的数量
cottage industry n. 26:24
家庭手工业;引申为小规模却泛滥的某类活动
spurn /spɜːrn/ v. 26:24
轻蔑地拒绝
exotic /ɪɡˈzɑːtɪk/ adj. 26:24
奇异的、标新立异的(此处带贬义)
hallmarks /ˈhɔːlmɑːrks/ n. 27:22
标志、特征
seamless interplay phr. 27:22
无缝的相互作用
pendulum /ˈpendʒələm/ n. 27:22
钟摆;the pendulum has swung 指风向逆转
affliction /əˈflɪkʃn/ n. 27:22
痛苦、困扰(此处指 AI 面临的问题)
susceptibility /səˌseptəˈbɪləti/ n. 28:25
易受影响性、脆弱性
adversarial attacks /ˌædvərˈseriəl/ n. 28:25
对抗攻击:以微小扰动误导模型的输入
nonsensical /nɑːnˈsensɪkl/ adj. 28:25
荒谬的、毫无意义的
impetuous /ɪmˈpetʃuəs/ adj. 28:25
冲动的、鲁莽的
impulse control n. 29:27
冲动控制
delude /dɪˈluːd/ v. 29:27
欺骗、使产生错觉(delude yourself 自欺)
mixed initiative planning n. 29:27
混合主动式规划:人与规划器共同参与决策的规划方式
serenity /səˈrenəti/ n. 30:32
宁静、平和
conceit /kənˈsiːt/ n. 31:29
(学术领域的)核心设定、基本假设;亦有「自负」义
verbalizable /ˈvɜːrbəlaɪzəbl/ adj. 31:29
可用语言表述的
knowledge intensive adj. 31:29
知识密集型的
riled up /raɪld ʌp/ phr. 32:35
激动、被激怒
ranting /ˈræntɪŋ/ v. 32:35
大声抱怨、咆哮式抨击
vanilla /vəˈnɪlə/ adj. 34:43
(技术口语)最基本的、无附加改动的
inject /ɪnˈdʒekt/ v. 34:43
注入、引入
grasping /ˈɡræspɪŋ/ n. 35:44
(机器人)抓取
inscrutable /ɪnˈskruːtəbl/ adj. 36:49
难以理解的、不可解读的
stances /ˈstænsɪz/ n. 36:49
立场、态度
digression /daɪˈɡreʃn/ n. 37:50
离题、插话
deleterious /ˌdeləˈtɪriəs/ adj. 37:50
有害的(正式用语)
tooth and claw phr. 37:50
弱肉强食、残酷无情(源自丁尼生诗句)
wiggly /ˈwɪɡli/ adj. 39:57
扭动的、蠕动的
locomotion /ˌloʊkəˈmoʊʃn/ n. 39:57
移动、运动能力
ergodic /ɜːrˈɡɑːdɪk/ adj. 39:57
遍历的:可反复采样以覆盖所有状态的
bandwagon /ˈbændwæɡən/ n. 42:04
潮流;jump on the bandwagon 赶时髦、随大流
pat yourself very happily on your back phr. 42:04
自我表扬、沾沾自喜(pat oneself on the back)
partially specified adj. 43:59
部分指定的、未完全描述的
heuristics /hjʊˈrɪstɪks/ n. 44:57
启发式(搜索中估计代价的函数)
informedness /ɪnˈfɔːrmdnəs/ n. 44:57
信息性:启发式对真实代价的贴近程度
non-declarative adj. 44:57
非陈述式的(如只有模拟器而无符号模型)
inductive biases n. 45:54
归纳偏置:学习算法在数据之外的先验假设
topology /təˈpɑːlədʒi/ n. 45:54
拓扑结构(此处指网络结构)
from scratch phr. 45:54
从零开始
fruitful /ˈfruːtfl/ adj. 47:03
富有成果的
sucked /sʌkt/ v. 47:03
(俚语)很糟糕、很差劲
invert /ɪnˈvɜːrt/ v. 48:07
反转、反推
bayesian /ˈbeɪziən/ adj. 49:06
贝叶斯的:以概率分布表示不确定性的
tongue-in-cheek /ˌtʌŋ ɪn ˈtʃiːk/ adj. 51:05
半开玩笑的、故作正经的
spectrum /ˈspektrəm/ n. 51:05
谱系、范围
ostrich /ˈɑːstrɪtʃ/ n. 52:07
鸵鸟
denotation /ˌdiːnoʊˈteɪʃn/ n. 53:12
指称、外延(概念所对应的实际对象)
nifty /ˈnɪfti/ adj. 53:12
巧妙的、好用的(口语)
crude /kruːd/ adj. 53:12
粗糙的、原始的
rapture /ˈræptʃər/ n. 54:23
(基督教)被提、末日审判;此处比喻遥遥无期
rebuffing /rɪˈbʌfɪŋ/ v. 54:23
断然拒绝
gratuitously /ɡrəˈtuːɪtəsli/ adv. 55:36
无缘无故地、多余地
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.168How to Speak Clearly & With Confidence | Matt Abrahams 下一期 · NO.170 →DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux | Lex Fridman Podcast #501
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]