视频库 / NO.027ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.027
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Prof. Judy Fan: Cognitive Tools for Making the Invisible Visible

节目发布 2025-04-08 · MIT Siegel Family Quest for Intelligence
朱迪·范 乔什·特南鲍姆
本期追问 · 点击跳到视频对应位置
3:53 人类心智的什么特性,让数轴、图表这类认知工具能被持续发明?21:14 为什么解释性的图画反而要牺牲逼真度?40:54 机器在图表理解上接近人类得分,为何仍不算真正对齐?51:35 要判断机器是否拥有心智,我们还能依据什么标准?
归入 Ⅴ·06 想法如何被讲清、被记住? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
本文根据斯坦福大学心理学助理教授范洁如(Judy Fan)在麻省理工学院脑与认知科学系系列讲座上的报告与现场问答整理而成,主持人为该系教授、计算认知科学学者约什·坦嫩鲍姆(Josh Tenenbaum)。范洁如主持斯坦福的「认知工具实验室」,研究人们如何借助绘画、图示、数据图表这类外部表征去学习、发现与沟通。报告分两部分,前半部谈徒手绘画中的视觉抽象,后半部首次公开分享实验室关于数据可视化的新工作,末尾是现场问答。以下依据现场录音编译整理,仅删去口语枝节、寒暄与重复,论点、例证与现场交锋一概保留。

主持人开场

坦嫩鲍姆: 今天非常荣幸请到范洁如做我们的系列讲座报告人。范洁如是我最喜欢的认知科学家之一,不分年龄。按任何一种客观标准衡量,她都属于年轻一代里领先的那批人。她在斯坦福做助理教授已经有几年了,我想这个估计大致不差。无论怎么看,她都是我们这个领域里正在升起并且已经在发光的明星。

坦嫩鲍姆: 她拿过不少奖,格卢什科博士论文奖、美国国家科学基金会的青年学者奖(NSF Career Award),我记得这个你也拿到了。这算是现场实时的介绍。但这些都不是我们请她来的真正理由。在我认识的研究者里,不分资历、不分领域,范洁如是最有创造力的几位之一,这话我说得很认真。

坦嫩鲍姆: 她的学术出身是神经科学,所以从很多方面看,她跟这栋楼里的人是同类,跟那些从大脑里记录信号、再用计算方法去解读它们的人聊天完全没有隔阂。这类工作她做过很多。她与我们熟悉的同行合作过,比如丹·亚明斯(Dan Yamins)就是她的导师之一;她研究过视觉,她的底子有很大一部分是视觉心理物理学、视觉神经科学和计算神经科学。

坦嫩鲍姆: 但如果你看她现在的研究,也就是今天你们会看到的一部分,以及今天根本来不及看的更大一部分,她已经走向了很多别的方向,或者说走向了一个关键方向:从知觉里那些最基础、可以用漂亮而精致的计算模型描述的过程,走向那些真正使我们成为人的认知层面,无论从生物意义上还是从文化意义上。

坦嫩鲍姆: 今天你们一定会看到这一面。这意味着她研究的是艺术表达、视觉与其他媒介里的创造性表达。她对叙事表达很感兴趣,对教育、学习和教学很感兴趣,对我们如何借助符号、数据与解释来理解世界很感兴趣。她越来越靠近那些更复杂的认知过程,那些在生物和文化两个意义上都独属于人类的过程,而它们在我们这个时代的分量也越来越重。

坦嫩鲍姆: 但她没有放弃任何一分严谨,或者说没有放弃对严谨的那种口味。我跟她打过很多次交道,我知道让她夜里睡不着的是什么:怎么用我们许多人从视觉神经科学、计算神经科学里学来并为之着迷的那种严格与精确,去抓住这些真正重要又真正棘手的东西。我不能说她已经彻底解决了这个问题,因为这大概还要花些时间,但这是个了不起的挑战。

坦嫩鲍姆: 她处理这个挑战的方式让我很受启发,我也希望我们都能从中学到东西、受到鼓舞,看看她现在走到了哪里,又要往哪里去。范洁如,请。

范: 好,哇。这真是太客气了,约什。谢谢你,也谢谢在座各位。今天到现在为止有多美好,我其实说不出来。我对这个系、这个共同体,以及各位工作中体现出来的那些价值,无论是科学上的还是别的方面的,都怀有很深的喜爱和敬意,所以能用这几十分钟跟大家讲讲我们最近做的事,对我来说是一种享受。

数轴不是自然给的

范: 我们研究的是认知工具。什么叫认知工具?我们从一个再熟悉、再简单不过的东西开始:数轴。自然界没有把数轴交给我们,是我们把它发明出来的。西班牙建筑师高迪说过,自然界里没有直线,也没有尖角。可这并没有拦住我们,我们照样把它们造了出来。几百年前,我们又把数轴延伸开去,做出了直角坐标系,而它后来被证明极其好用。

范: 那在当时是货真价实的前沿思维工具,是用来推导新数学发现的工具。笛卡尔和他同时代的人意识到,可以把代数表达式和几何曲线挂起钩来,从而解决各种数学难题,包括这个困扰了世界上千年的问题。在座有多少人熟悉所谓的「提洛问题」,也就是提洛神谕之谜?它问的是:怎么把一个正立方体的体积变成原来的两倍?用那个共同体、那个时代的数学方法去做,这非常非常困难。

范: 而这就是它的解法。要评价这项发明在过去四个世纪里的影响,怎么说都很难说过头。这项技术把一大类问题,比如你想找出同时满足两个方程的那组值,转换成了另一件事:在两条曲线之间找交点。现在设想一下,这个革命性的工具接下来遭遇了什么。它变成了我们习以为常的东西。它太好用了,好用到成了每一代人受教育时都绕不开的基础设施。

范: 地球上几乎每一份数学课程大纲,都会引入一套符号记法和图形记法的组合,用来表示和操作数学对象。而我们反复琢磨的问题是:我们是怎么走到这一步的?人的心智里究竟有什么,使得这种持续不断的创新成为可能?可以从很多学科视角切进这个问题,历史学、人类学、经济学都可以,而我认为认知科学在这里也有非常重要的贡献。

让不可见变可见

范: 我想这个故事至少要从三万到八万年前讲起。解剖学意义上的现代人开始在自己的物理环境上留下标记,包括你们看到的这些洞穴岩壁,等于是把周围的物体和表面挪作他用,让它们成为意义的载体。我们当然没有停在岩壁上。人类学习与发现的历史,同「让不可见之物变得可见」的技术史深深缠绕在一起。这里是我很喜欢的几个科学史例子。

范: 这是达尔文的雀鸟,图是鸟类学家约翰·古尔德画的,达尔文与他合作得非常紧密。只有当这些个案被并排放在一起看,形态上的差异才真正凸显出来、跳到你眼前。这是伽利略用来观察木星卫星运动的望远镜,正是那种分辨率,让他有底气去质疑关于太阳系结构的正统说法。这是卡哈尔那些广受赞誉的视网膜素描,画的是显微镜下所见,告诉我们神经系统的那一部分长什么样,各部分之间又是怎么连起来的。

范: 到了二十世纪,我们有了费曼图,以物理学家理查德·费曼命名,它让我们看见亚原子粒子如何一闪而生、一闪而灭,而这些事件我们肉眼根本无法直接观察,将来也永远不可能。

范: 我想把大家的注意力先拉回这张幻灯片上的差异。请注意,其中有些图像相当细致,如果可以这么说的话,相当忠实于我们睁开眼时视觉世界呈现的样子,比如左上角达尔文的雀鸟。另一些则要图式化得多。但所有这些例子的共同点是,它们都在利用我称之为「视觉抽象」(visual abstraction)的手段,把我们看到和知道的东西,以一种突出「什么值得注意」的格式传达出去。

范: 在此之上,我们还会把这种不断扩张的自然知识拿去做别的事,用这些学习工具去创造新东西。举个例子,正是因为我们对力学有了细致的理解,才能设计并造出高精度的计时装置。技术进步很大程度上是由这样一种能力推动的:我们能够不断把对世界的理解重新表述成有用的抽象,而这些抽象又让我们有可能按照自己的设计去重新改造物理世界。于是,生物学洞见变成了生物工程,物理理论变成了先进的物理仪器,神经科学变成了医疗器械,量子力学变成了现代电子学。

范: 这就是我从中汲取灵感的那类现象的一个取样。我们持续琢磨的问题是:我们身上的什么东西,使这一切成为可能?

缺席的两块拼图

范: 过去几年我一直用这张示意图帮自己梳理其中的关键行为现象,它也可以作为一个框架,把我今天要讲的几条工作线索放进去。这一版画的是认知心理学,也就是我的本行,传统上的做法:关注人如何加工外部世界提供的信息。

范: 加进社会认知的研究之后,这张图会变得更丰富一些:它同时考虑多个个体的行为,以及他们如何互相影响。当这些活动被用来了解世界、并把知识分享给别人时,有人主张,即便由非专业者在日常情境中进行,它们也与正规科学有着重要的相似之处。

范: 沿着这个传统往下走,并且带着「理解人类是怎么做出刚才那些非凡的发现与发明」这个目标,我想主张这张图里还缺了两样关键的东西。第一样是关于认知工具或者说认知技术的解释:那些编码了信息、并且意在影响我们心智,影响我们如何思考、思考什么的物质对象。第二样,我想说,是时候接纳科学天然的另一半了,那就是工程:人们如何利用自己对世界的理解,不管这理解来自直接经验还是来自他人,去创造出新的、有用的东西。因为如果不认真对待这幅图的工程那一半,我斗胆说一句,我们永远无法解释我们所知的这个世界为什么会变成今天这样。

范: 说得直白一点,我的团队做的研究,核心就是要把这个环闭上:一方面发展心理学理论,解释我们如何找到那些能说明世界运作方式的有用抽象;另一方面同时发展理论,解释我们如何运用这些抽象去造出新东西。

范: 今天的计划是讲两条线上已经做出来的一些工作。第一部分,我会讲讲我们自认为搞清楚了什么:人如何利用视觉抽象来传达语义知识,核心案例是徒手绘画。按常规,第二部分我该讲我们关于人如何在搭建实物时学习和协调「程序性抽象」的研究,那对应工程那一半。但今天我想跟各位试点新的。我想改讲我们正在展开的一批工作,我很想听听大家的反应,那就是数据可视化的认知基础:人如何调动多种信息模态,图形元素、文字和数字,去进行统计推理。

范: 换句话说,如何从有限的证据中,去了解那些单靠一个人直接观察很难甚至不可能了解的世界面向。关于实物装配和物理推理的工作,我很乐意在酒会上跟大家聊。

相似还是约定

范: 进入第一部分。我们该怎么开始思考「人如何用视觉抽象传达自己所知与所见」这个问题?这里有三种层层叠加的行为值得区分。

范: 第一当然是视觉知觉,也就是我们如何把原始的感觉输入转换成有语义、有意义的知觉体验;有了它,才谈得上视觉生产,即生成一组标记、在物理环境中留下有意义且可见的痕迹的能力。这两者又在视觉沟通中汇合:我们如何决定把这些图形元素怎样排布、以什么顺序排布,好对别人的心智产生某种特定的影响,无论目的是告知、教学、说服、协作,还是我们能赋予这些痕迹的任何其他用途。

范: 接下来我会走一遍三项研究。先从这个问题开始:理解某一类图画在表示什么,其知觉基础是什么?为了起步,我们选了视觉抽象最具体、最熟悉的那种形态,就是徒手画出一张看起来像世上某物的画。为什么我们能毫不费力地判断,这张幻灯片左边的素描对应的是右边那只写实渲染的鸟?

范: 对这个问题有很多种回答,其中两种占主导。第一种观点认为,我们之所以觉得图画有意义、在表示某物,根本上是因为图画就是长得像世上的对象。这幅画字面意义上看起来像那只鸟,我们就是这么知道的。第二种回答认为,图画指代对象主要是约定的事,我们是从别人那里学到哪种图配哪个对象、配哪个意思的,比如这个汉字。

范: 在更早的一条工作线里,我和我的紧密合作者丹·亚明斯,以及我的博士导师尼克·布朗发现,通用视觉算法,这里指的是由多层可学习的空间卷积堆叠而成、在自然照片上训练的神经网络,能够相当强地泛化到那些一点也不写实的稀疏素描上。这似乎暗示,图画意义与相似性的问题,也许只要建立更好的腹侧视觉通路模型就能解决,尤其是那种能准确刻画相应脑区所执行运算的模型。

范: 发现这一点的不止我们。谷歌学术里能检索到的计算机视觉论文,轻轻松松就有上千篇用某种卷积网络或别的神经网络作为主干,把素描和自然图像编码起来,服务于各种应用。所有这些结果,你都可以看作是为一个升级过的、现代版的「相似论」背书。这个洞见对我们自己也很有用,还带来了别的实际后果。

范: 在更近的一项工作里,由实验室从前的硕士生查尔斯·卢主导,与加州大学圣迭戈分校的王小龙合作,我们在这些早期发现之上继续加压测试相似论:拿一个卷积网络当主干,在上面训练一个解码器,让它把素描里的局部元素映射到照片里对应的元素上,约束条件是你可以把素描揉皱、扭曲,但不能撕出洞来。

范: 这个方法能奏效,我这里只是做个演示,算不上正式结果,它提示我们,素描的各部分与它们所要表示的真实物体的各部分之间,遵循着相当强的空间约束。很好。我们已经有了足够好用、可训练的素描理解模型,可以拿去做下游应用,效果也不错。看来可以收工回家了。当然,这远不是故事的全部。

语境决定抽象层次

范: 一个静态、确定性的视觉加工解释,无法说明我们如何生成、又如何读懂这样的图,而这类图在这栋楼里到处都是。这些团块、方框、波浪线和箭头是什么意思,取决于我们正在谈什么。所以我们下一个目标是搞清楚怎么把语境信息纳进来,从而涵盖我们实际用于沟通的、更多样的图形表征,从左边这些比较忠实的图画,一直到右边那些明显是符号性的记号。

范: 我们处理这个问题的第一篇论文问的是:人们怎么知道什么时候必须画得更忠实,什么时候可以糊弄过去、画得图式一点、抽象一点?在那项研究里,我们把两个人配成一组玩一个画画游戏。画者看到的界面大致是这样,任务是把高亮的目标物体画出来,也就是第三个,而我们操纵的是界面上其他物体是什么。在「近」条件下,干扰项全部属于同一个基本层次范畴;在「远」条件下,干扰项来自不同范畴。

范: 用这个非常简单的操纵,我们发现普通人调整描绘方式的能力有多么灵活:在「近」试次上,当他们需要让画唯一地指认某一个具体样例时,他们会画得更细致、更忠实;而在「远」试次上,范畴层面的抽象就够用了,他们画得就更稀疏。这里是那项研究里收集到的一些真实作品。

范: 我们发现,在「远」试次上画者用的笔画更少、墨水更少、时间更短,而在「把目标视觉概念的身份传达给观看者」这一根本任务上仍然维持了天花板水平的准确率,观看者在这些试次上做判断的时间也更短。

范: 为了刻画这种行为模式,我们提出了一个关于画者的计算模型,由两部分组成:一个卷积网络,把视觉输入编码进一个泛用的抽象特征空间;再加一个概率决策模块,根据语境推断该画哪一种画。因为我很想赶紧讲后面的工作,这里我只给结论。模型消融实验的结论是:视觉抽象的能力(我们把它操作化为视觉编码器模块里的网络层)与对语境的敏感性,两者对于解释人们如何在恰当的抽象层次上谈论这些物体,都是不可或缺的。

范: 在更近的工作里,我们把这个想法又推了一步,不只看当前的指称语境如何影响沟通,还看新的图形约定可能在什么条件下涌现出来:当人们记得自己与同一个人此前的互动时,他们会逐渐产生更抽象、甚至可以说带有原始符号性质的记号,而这些记号的意义更强地依赖于那段共享历史。

范: 这些进展让我很兴奋。但当然,人对世界的知识远不止「这东西叫什么」和「它长什么样」。

解释图与描绘图

范: 视觉抽象还有一个特别重要的用途,在科学里尤其如此,就是传递关于事物如何运作的机制性知识。当人们做出这一步跨越时,他们脑子里发生了什么?也就是说,越过某只鸟身上那些视觉上最显眼的特征,转而去凸显底层的物理机制,比如鸟类总体上是怎么实现飞行的。

范: 所以你可以想象,当实验室里一位很出色的前博士生霍莉·休伊(现在在 Adobe Research)同样被这个问题迷住时,我有多高兴。我们从弗兰克·凯尔和同事们非常漂亮的工作中知道,人在向他人学习时会优先看重机制性解释;从芭芭拉·特沃斯基、米琪·奇、塔尼娅·隆布罗佐等人的工作中知道,人可以通过生成解释来学习。但我们意识到,关于人怎么想「视觉解释」这件事,我们还有很多不知道的:人们认为一张说明「某物如何运作」的图里应该有些什么?这样的图跟一张只求「看起来像」的普通插图,区别又在哪里?

范: 一种可能,我把它叫作「累加假设」,是说人们基本上把视觉解释看成普通描绘的加长加料版。照这张表看,解释会具备描绘所具备的一切,用来传达视觉外观,然后再额外加上关于物理机制的信息。另一种可能,我把它叫作「可分离假设」,是说人们认为解释这类图像会挑出机制性的抽象,同时大幅弱化视觉外观。这里存在一种选择性。

范: 为了把这两种可能拆开,霍莉设计了一项研究,从两个方向下手:第一,细致刻画视觉解释的内容,并与视觉描绘作比较;第二,测量这两类图像究竟能在多大程度上帮助下游的观看者完成任务,提取他们真正需要的信息,无论那是关于外观还是关于机制的信息。

范: 与其上一堂鸟类飞行课,那当然迷人但也复杂,霍莉动手做了六台新奇的小装置,它们闭合电路、点亮灯泡的机制清晰可见。这是其中一台机器,以及被试观看的教学视频。被试其实看了两遍,也就是演示了两次,就是这样。

范: 请注意,这台机器由三类零件组成。有因果零件,它们必须转动才能把灯点亮。有非因果零件,长得很像,但并不导致灯亮。还有背景性的结构件,它们颜色鲜艳、把其他零件撑起来,非常重要,但并不直接参与点灯的那条回路。

范: 每位被试都会为一部分机器画解释图,为另一部分画描绘图。在解释试次上,我们请他们设想,自己的画会被别人用来理解这个物体是怎么运作的。在描绘试次上,我们请他们设想,自己的画会被别人用来从一排长相相似的物体里认出是哪一个。操纵就这么简单。用这个流程,我们为这六台机器各收集了大量的描绘图和解释图。

范: 用肉眼扫一遍,这些画看上去是有区别的。描绘图里背景似乎稍微多一点,解释图里箭头似乎多一些。但霍莉想把这件事做得非常系统,于是她众包了标注,把每张画里的每一笔都归入四个类别之一:三类物理零件各占一类,即背景、因果、非因果,第四类是符号的兜底类别,在这个情境下指的是箭头和运动线。然后用这些标签去比较两种条件下人们各自强调了怎样的语义信息配比。

范: 她发现,虽然两种条件下人们都会画因果零件和非因果零件,但在解释图里分配给因果零件的笔画更多。他们在描绘图里对背景的强调略多于解释图;而正如我们猜想的,他们在解释图里花在符号性表达上的墨水更多,用来表示运动和零件之间的相互作用。仅凭这些结果,就已经与累加假设的强版本不相容了,因为按后者的预测,即便加入了箭头,背景、因果零件和非因果零件之间的相对强调也该保持不变。

范: 不过,这些差异再稳定,也可能只是风格上的变化,对它们所要支持的任务并没有实际的用处上的影响。所以,为了测量这些决策的功能后果,霍莉设计了三项推断任务。第一项问的是:你多容易看出操作这台机器需要哪种动作,是拉、是转还是推?这正是一张好的机制解释图应当讲清楚的。

范: 第二项测量每张图用于物体辨认的效果,这正是描绘图该干的事。第三项是更难的视觉辨别任务:你要判断两个被高亮的零件里哪一个是因果零件,这要求你在图的各部分与真实机器的各部分之间建立细致的对应关系。这种事,你只能指望一张既讲清了机制、又保留了足够多整体外观与零件组织信息的解释图来完成。

范: 按累加假设,我们应当看到解释图在所有任务上都至少不输给描绘图;而按可分离假设,解释图在动作任务上可能更好,在物体辨认任务上反而更差。霍莉的发现更符合可分离的那一版:解释图更好地传达了机制如何运作,而描绘图更好地传达了物体的身份。

范: 有意思的是,在这个样本里,解释图在第三项更难的任务上,也就是指认因果零件,并不占优。这与下面这个想法一致:由于略去了大量背景细节,某些解释图恰恰把那些能让人把画中特定部位与机器部位对应起来的信息给抽掉了。

范: 这些研究的底线是:即便是第一次被要求画一张视觉解释图,人们对「解释图里应该有什么」也共享着某种直觉。而这可能意味着牺牲视觉上的逼真度,去凸显更抽象的机制信息。更一般地说,这项工作表明,要理解人们为什么这样画、描绘图为什么长成这样,沟通语境与沟通目标至关重要;它也提供了一套实验与分析工具,用来刻画人们传达与目标、语境相关的视觉信息时所采用的策略。

SEVA 与人机差距

范: 接下来这一节我们要问:要造出能像人一样进行视觉抽象的人工系统,需要什么?为什么问这个?因为我们根本上想要的是关于视觉沟通的、有用的科学模型。我到目前为止讲的工作,代表了我们把投入优先放在哪里,那就是开发实验范式与数据集,在更广的情境里、以更完整的方式测量和刻画这些行为。

范: 与此同时,我们也花了相当大的力气去评估:那一批不断进步的机器学习系统里,究竟有哪些成员有可能继续成为有前景的候选者,用来刻画人在这些高维任务中更细致的行为模式,具体来说,既作为人类图像理解的模型,也作为图像创作的模型。

范: 我们考虑的任务设定,灵感来自毕加索这组著名的素描。其中有些非常细致,最后几幅极其抽象,但每一幅都毫无疑问是公牛。任何一个称得上合格的、关于人类视觉抽象的科学模型,都应当能够表示这些公牛彼此之间有何不同,同时又能表示它们骨子里都很「公牛」,更准确地说,是像真实的人类观察者所看到的那样「公牛」。

范: 这项工作是一次巨大的团队努力,由库辛·穆克吉主导,他这周五就要答辩,然后会以博士后身份加入实验室;霍莉、查尔斯·卢、亚埃尔·温克勒(她今天也在场)和里奥·阿吉纳-康都有贡献。

范: 我们的前提是:要检验我们是否走在通往那种科学模型的正确道路上,一个强有力的检验就是看我们能不能造出像人一样生成和理解抽象图像的算法。素描理解就属于那种看似简单、实则对通用视觉算法构成根本挑战的问题。一来,它要求对稀疏程度的差异保持稳健,因为有些素描比另一些细致得多;二来,它要求对语义歧义有容忍度,因为素描完全可以稳定地唤起多种意义。

范: 于是我们造了一个基准,叫作 SEVA,把这些挑战明确摆上台面。我们收集了九万张手绘素描,来自大约五千五百人,涵盖一百二十八个视觉概念,且在不同的「生产预算」下完成。这里是一些例子:人们看到一张照片作为提示,然后在越来越短的时间里画出来。照片取自 THINGS 数据集,不知道大家熟不熟悉。到只剩四秒的时候,那画是真的相当潦草了。我们把这些画同时拿给人和十七种当时最先进的视觉算法看,这些算法在架构取向、策略与训练方案上覆盖面很广。

坦嫩鲍姆: 插一句澄清性的问题。人在计时开始之前有没有时间先想一想,还是说他们一看到就只有四秒?

范: 这个问题折磨了我两年。没想够。不是的,他们看到图像,然后就得开始画。所以我们并不确切知道具体是怎么回事,整个试次的总时长是受限的。

坦嫩鲍姆: 所以这里面「想」和「画」是混在一起的。

范: 想和画混在一起。我认为,要真正逼近那个「公牛现象」,我应该给他们无限的规划时间,只限制执行时间,那才是正确的做法。但这一版不是那样。所以还有另一个数据集尚未诞生,能更直接地把这一点分离出来。是的,四秒的那些画确实相当「呆」,这是这个领域的专业术语。但在那种条件下,人们画出来的就是这样。

范: 于是我们手上有了这些不同的素描,在那些条件下它们就长成那个样子。然后让人和这些视觉算法执行同一项素描分类任务,这样我们就能测出每一张素描所唤起的完整标签分布。

范: 我们首先确认的是:随着给人更多时间去思考如何画、再动手画出细致的素描,这些素描对模型和对人都变得更容易辨认;用标签分布的熵来衡量,它们的歧义也更小。而且即便猜错了,也就是最高概率的那个标签不对,它至少也更可能落在正确的语义邻域里,这一点是用语言嵌入估计的。

范: 这算是让人放心。但再往深处挖,我们发现,虽然确实有些模型在识别任务上货真价实地比别的模型更强,但模型之间的表现差异,完全被模型与人之间的差距压得看不见,人类识别中的可靠信号要大得多。这在左边的表现指标上是如此,在「对一张素描意义的相对不确定性」上也是如此。所以这说明,在素描理解上,人与模型之间还有相当可观的对齐差距需要弥合。

范: 尽管如此,在我们做这个基准的当时,我们注意到用 CLIP 训练的模型表现优于其他模型,这使它成为一个合理的起点,可以在其上探索素描生成的模型。于是我们考察了一个特别酷的素描生成算法的能力,它叫 CLIPasso,开发工作由亚埃尔·温克勒主导。我们也让 CLIPasso 生成素描,并同样操纵它的生产预算。只不过这里的单位、这里的「货币」变成了笔画数,我知道这跟人那边不一样。

范: 我们发现,在四种生产预算下,人的画和 CLIPasso 的画可辨认程度相当接近。顺带说一句,可辨认度上的这种重合完全是巧合,毕竟两边的单位根本不同。而真正耐人寻味的发现是:人和 CLIPasso 「变稀疏」的方式不一样。右边纵轴上画的,是人给「他人所画素描」和「CLIPasso 在各预算下所画素描」分别指派的标签分布之间的散度。

范: 换句话说,如果你认真对待这样一种功能主义的看法,即素描是为了传达概念而存在的,并且用它所唤起的完整标签与意义分布来刻画它的意义,那么就会发现:三十二笔的 CLIPasso 作品和三十二秒的人类素描,看上去是不一样的,风格上确实有别,但它们在功能上相当接近,也就是所传达的意义集合、所唤起的意义分布相当接近。

范: 另一方面,随着生产预算被收紧,人与 CLIPasso 之间的分歧才真正开始大幅拉开。我们希望 SEVA 这个基准能成为一份有用的资源,供那些有志于开发「类人视觉抽象」模型的人使用。

数据可视化的力量

范: 现在进入第二部分,我想它会短一些。我给大家提前透露一下我们最新的、基本尚未发表的工作,关于多模态抽象,以及它们如何支撑统计推理。让我们回到笛卡尔坐标平面。

范: 在座都是科学家,用不着我说:你对世界做观测时,从来不会有这么干净的东西。我们实际收集到的往往是这样一堆数据点。它们落在哪里就是哪里,我们要从中推断某种底层结构,推断真正产生它们的那个生成过程。这一步推断是科学推理的基本构件,而我们绝不是靠把见过的一切都背下来、再拼命动脑子做到的。

范: 我们靠的是技术。在报告开头,我给各位看了徒手绘画的一些用例,它是一种特别持久、特别多用、也特别易得的工具,用来让不可见变得可见。我认为这很了不起,值得被理解。但在现代产生的技术里,影响最大的,恐怕要数数据可视化的发明。

范: 跟望远镜、显微镜一样,图表帮我们分辨出那些无法直接看见的世界部分。但与这两种光学技术不同的是,它让你看见那些用肉眼看太大、太嘈杂、太缓慢因而看不见的模式与现象。它们在新闻里无处不在,是商界与政界循证决策的基石,在每一个科学与工程领域都不可或缺。

范: 这里是历史上最早的时间序列图之一,威廉·普莱费尔画于一七八六年,展示英格兰在一七〇〇年到一七八〇年这八十年间进出口的平衡。有一段时间进口超过出口,然后在一七五〇年代关系反转,出口大幅上扬。

范: 但关键在这里。跟一幅达尔文雀鸟的素描不同,如果你此前从没见过这类图像,你可能压根不知道自己在看什么。可一旦学会了怎么看,它就是一种超能力。海量的单个观测可以被浓缩进一张图,而这张图讲的故事,你只要看一眼就能读出来。

范: 值得关心图表的理由还不止这些。正因为它是帮助人们更新和校准自己关于复杂世界的信念的有力工具,培养读图、解图乃至制图的能力,长期以来一直是这个国家 STEM 教育的目标之一,而且这件事的分量还在变重。

范: 《纽约时报》曾报道过一则新闻,其实已经是大约一年前的事了,讲的是新冠疫情造成的数学学习损失的恢复情况,基础是我们斯坦福和哈佛的教育学同行主导的研究。看上去在不少州都有橙色箭头,意味着确实有所恢复,但离恢复到位还有很长的路。我认为,能够解释人们如何用这类图像去发现并传达重要定量洞见的成功理论,会帮助我们让人们更普遍地具备所需的定量数据素养。

六项测试的错法

范: 我要简要讲三个我们正在推进的方向。第一个问题是:理解图表需要哪些底层操作?我们采取的策略是:先找到那些至少能应付关于数据可视化的问题的机器学习系统,评估它们与人的对齐程度,然后追问这些差距的来源。当然,要起步就得先有办法测量「理解」。

范: 大概长这样。假设我们一起看这张堆叠柱状图,有人问你拉斯维加斯的花生价格。花点时间把它看进去,四处扫一扫。假设那个人给了你四个选项。如果你觉得是 A,那你是对的。那个炭灰色的小方块是我加的。但面对同一个问题,几个知名的视觉语言模型给出的答案并非如此。

范: 在实验室现任研究助理阿纳夫·维尔马主导的一次堪称大力气的基准工作中(他今年秋天就要进入这里的项目),我们在六项常用的图表推理测试上,对人和 AI 系统做了细致比较,这些测试分别来自教育学、健康、可视化、心理学和机器学习等不同共同体。

范: 六项测试都尽可能以平行的方式施测给人类被试和若干所谓的多模态 AI 系统,而进入这份基准的模型,都是曾被宣称在其他视觉落地推理任务上具备胜任力的。我们不仅记录人和模型的总分,还记录它们产生的全部错误模式,这样即便模型或人答错了,我们也能看看它们错的方式是否相似。

范: 结果如何?这六项测试的名字都挺古怪,GGR、VLAT、CALVI、HOLF、HOLF-multi,最后那个难听的名字是我们自己起的,怪我;另外还有 ChartQA 的一个子集。我们计算每个模型的表现,就是横轴上蓝色、橙色、紫色、红色那几个,再与人的表现比较,人的表现是绿色。这里的人是至少上过一门高中数学课的美国成年人。

范: 先给大家一个感觉,这些人做得如何,这是我们的参照点。这是所有模型的表现。这项研究里包含了 Blip2-FlanT5 的两个变体、三个基于 LLaVA 的模型、MatCha 这样的专用系统及其基础模型 Pix2Struct,最后还有一个闭源的商业模型 GPT-4V。在所有这些测评上,无论采用较宽松还是较严格的评分标准,我们都看到了模型与人之间实实在在的差距。

范: 如果我们只依赖机器学习文献里当下最流行的那些图表理解基准,这个差距很可能会被漏掉,这也正是我们要把 ChartQA 纳进来的原因,也就是这张幻灯片最右边那一格,那里的差距看上去小得多。第三项 CALVI 特别有意思,因为它是对抗性设计的图,纵轴的取值范围很古怪,逼你非得看得很仔细不可,所以那一项上我们也看到了比较大的差距。

范: 我们还分析了完整的错误模式,在没有谁接近天花板的时候,这种分析特别能说明问题。毕竟,全部答对只有一种方式,但错的方式有很多种,而且可以错得很稳定。我们发现,尽管 GPT-4V 看上去在接近人的水平,但包括 GPT-4V 在内,没有任何一个模型产生了类人的错误模式。这一点体现在所有的点都明显落在绿色阴影区域下方,那片区域代表人类的噪声上限。

范: 所以结论是:当前不断发展的视觉语言模型,作为「可视化理解的可能认知模型」的假设空间来说,仍然是令人兴奋、值得参数化的试验床,但仍存在系统性的行为差距,值得继续追问下去,才能把这些模型的潜力真正发挥出来。

选图看得懂受众

范: 与此同时,我们也在开发实验范式,去探测可视化理解的另一个相关侧面:为你的认识目标挑选或设计合适的图表的能力。

范: 我们是这样设定问题的。设想你对某个数据集有某个问题,我一直称之为「认识目标」(epistemic goal),比如「哪一组更好」。假设左边这位要挑一张图,去恰当地改变对方的信念;但如果对方心里的问题不同,也许就需要另一张图。直觉就是这么来的。这条线索由霍莉·休伊开启,我们在一项研究里编了几百个原则上可以用真实的、公开数据集回答的问题,这里用的是 R 语言基础包自带的数据集。

范: 这是一个听着挺吓人的例子,追踪飞机与鸟撞。问题是:在阴天飞行、并在零到五千英里高度遭遇鸟撞的飞机,其平均速度是多少?然后我们给被试一份菜单。

听众: 洁如,我能问个问题吗?

范: 当然。

听众: 五千英里是什么?

范: 高度。

听众: 你确定不是五千英尺?

范: 抱歉,五……等一下。这可能是我这边的笔误。我想这个错误没有传到题目里去。

听众: 总之,别在意。

范: 我开始在意了,但我现在要克制住这个冲动。

听众: 不过如果这就是标准,这些题里恐怕到处都是各种笔误。

范: 这是用模板生成的,题目是套模板出来的。基本上这里面有不少问题问得挺滑稽,我觉得当初本可以再打磨得平顺一些。这个意见很好。

范: 总之,我们有了这一堆问题。有些问得比另一些讲究。我们给被试一份可选图表的菜单,让他们挑一张展示给别人,好帮对方回答那个问题,不论他们自己是怎么理解那个问题的。反正呈现给人的就是那些字符串。他们可以从菜单里挑柱状图、折线图或散点图,有些图更浓缩,有些更展开。然后我们统计每一种被选中的频率,构造出一个关于图表的选择分布。

范: 这是我们刺激材料中「读取数值」这一类问题的平均结果。这是其中一类任务,另外还有一些任务要求你计算两个数值之间的某种差。这里是「读取数值」这一子集。接着我们检验了人们在挑图时可能采用的各种假设策略。

范: 黑色曲线显示,他们显然存在某种偏好。紫色曲线代表「柱状图原教旨主义者」「折线图原教旨主义者」「散点图原教旨主义者」的预测,也就是不管图上画了三个变量还是更多,一律只选柱状图之类。我们也考虑过另一种可能:人们总体上偏好更有针对性的可视化,但并不真的在意分享的是哪种图型;或者他们偏好展示更多数据的图;或者偏好那些更少掩盖数据变异的图。

范: 我们考察过的候选里最好的一个,是这样一种设想:人们其实对那些「与回答该问题相关」的图表特征是敏感的。而这个「相关性」,实际上是由另外大约一千七百名被试的表现来预测的,他们要用每一张可能的图去回答每一个问题,不管问题问得多古怪。我们真的去测了他们的表现,再根据这些表现水平构造分布,然后用这项任务的数据去生成「受众敏感」假设的预测。

范: 这是只考虑「读取数值」类题目时,那些曲线的形状。而如果看整个数据集,我们发现「受众敏感」这个假设全面表现良好。我这里展示的本质上是一个模型拟合指标,定义在预测的选择分布与真实的人类选择分布之间的散度上。它有个名字,叫詹森-香农散度,也就是双向 KL 散度的平均。

范: 我对这个结果感到兴奋,一方面因为它初步验证了一种用更开放的任务去测量可视化理解的策略;另一方面因为它表明,即便是非专家,是那些并非职业科学家的人,也能察觉图表的哪些特征使它更适合回答某些问题。说实话,作为一个靠教统计学导论谋生的人,这让我看到希望。

这些测试测的是什么

范: 旅程的最后一段,我们要再一次批判性地看待「测量」这个问题。几分钟前我给大家看了用那六项数据可视化理解测试得到的结果。我们手上现在只有这些,所以基准研究里用的就是它们。而这最后一项研究,现在已经在出版流程中,由我们实验室一位出色的博士后埃里克·布罗克班克主导,阿纳夫·维尔马也参与,我们问的是:这些测试到底在测什么?它们测这些技能的方式是最好的吗?我们能不能做得更好?

范: 这里是回答这些问题的最初几步。为了找到抓手,我们先从 GGR 和 VLAT 入手,因为它们是使用最广、最成熟的两项测试。我们把 GGR 和 VLAT 合成一份综合测试,施测给一个规模大、构成多样的美国成年人样本,其实是两个样本:一个在加州大学圣迭戈分校校园招募,另一个通过 Prolific 招募,并要求在人口统计学上具有代表性。左边显示的是,在校园样本和全美代表性样本中,我们对单个题目难度的估计相当一致,我觉得这挺让人安心。

范: 右边显示的是,在一项测试上做得好的人,往往在另一项上也做得好,这提示两项测试也许在测量某些相同或相近的东西。问题是:那些东西究竟是什么?

范: 一种可能是,这两项测试追踪的是某些图表比另一些更容易理解的程度。若是如此,那些图应当在各处都稳定地难或稳定地易,比如柱状图容易,堆叠面积图更难。但在我们看来,这里显然还有别的事情在发生。

范: 对于给定的图型,表现并不总是一致,无论是在同一项测试内部还是跨测试之间;而且题目数量根本不足以在图型与表现之间建立直接联系,因为这些测试里每一类图通常只有一道题。这也让分析变得困难。

范: 我们还深入分析了人们犯错的模式,结果发现,在这两项测试上,预测这些模式的最佳变量既不是图的类型,也不是问题的类型,比如「找最大值」「识别聚类」「刻画分布」「读取数值」。这些都是描述图表任务分类体系时常见的说法,但从这次分析看来,真正更有效地解释错误模式的,是另外一些底层因素,而现有的分类体系并没有很好地描述它们。

范: 这里我展示的是,一个简洁的四因子模型,比那种按你直觉上会用来组织「读图所需技能」的分组方式,要好多少。这张幻灯片我不打算逐项拆解,它只是想说明:这里似乎存在某种东西,而它并不能明显地对应到我们平时谈论「读图的组成技能」的方式上,也就是你在教科书或教学材料里看到的那套入门框架。

范: 即便现有测评未必是以最好的方式在测量和刻画可视化理解,我们也想把这件事当作一个「杯子半满」的行动号召,去开发更好的测量工具。所以,敬请期待。

尾声:教育与设计

范: 更一般地说,我之所以认为这项确立数据可视化的知觉与认知基础的工作如此重要,是因为它让我们有机会把学到的东西用起来,最终去帮助真实教育情境中的学习者,校准他们对一个复杂多变、而我们永远只能观察到其一部分的世界的理解。这类把基础科学与我们共同生活的更大世界连起来的整合性努力,正是这一切最终的去向。

范: 我们真正想做的,是发展心理学理论,解释人如何使用我们继承下来、并且仍在不断创新的这一整套认知技术;我们想弄明白这套工具箱为什么长成今天这样,未来的认知工具可能怎样才会更好用。

范: 从长远看,我认为理解这些工具如何运作、如何把它们做得更好,真的很重要,因为正是这些工具支撑着我们最有影响力、最具生成性的两项活动。第一是教育,这是一种制度,或许更重要的是一种期待:每一代学习者都应当站在上一代的肩膀上,看得更远。第二是设计,是那一整套活动与思维习惯,它们让人不断重新想象世界可以变得多好,然后走出去把它变成现实。

范: 讲到这里,我要感谢所有参与这些工作以及实验室其他工作的人。没有一支了不起的研究团队,没有分布在很多地方、也包括这里的合作者与同行网络,这些都不可能完成。谢谢他们,也谢谢各位的专注。如果还有时间,我很乐意回答问题,虽然可能已经没时间了。好,谢谢。

问答:画错了的画

听众: 非常感谢你的报告。我好奇的一点是,在你很多关于描绘的例子里,似乎存在一条一维的轴,一端是更细致、更忠实,另一端是更稀疏。但我想问的是那些人们「跑偏」的情况,也就是他们画出的东西其实是假的、与事实相反的。比如我要画一杯水,为了表示杯子是满的,我会把里面涂成蓝色,尽管现实里杯子装了水并不会变蓝。我觉得这跟文化和语言关系很大。我说的那种语言里,水就是被形容成蓝色的;我周围的游泳池内壁通常也漆成蓝色,如此等等。所以我想知道,你怎么看这类情况?

范: 是的。哲学里的美学领域有一大批文献,我从中获益不少,它们的出发点跟你说的很像:我们究竟怎么可能理解那些「假」的图画,比如虚构人物的画像、独角兽的画像,或者装着蓝色水的杯子,哪怕现实里根本不是那样?

范: 我们在自己的实证工作里采取的路径,是不从这个前提出发。与其把它们看成「假的」,不如说这就是人在某个目标之下真实生成的数据,而我们要做的是搞清楚那个目标是什么。至于他们为什么把杯中的水画成那样,我们感兴趣的正是你指出的那种权衡:在多大程度上忠实于视觉外观,忠实于「一只玻璃杯就摆在你面前的房间里」那种看的现象学,又在多大程度上依循一种习得的约定,也就是「水该怎么画」。

范: 我认为这些都是很合理的约束来源,能解释人们为什么做出那样的表征决策。所以,与其把某些画判为好或坏、真或假,我们觉得更有用也更有产出的想法是:在那些条件下,人们画出来的就是这些画。为什么?为什么它们是这个样子,而不是别的样子?希望这有帮助。谢谢你的问题。

问答:如何诊断误读

听众: 谢谢。我很好奇你只是很简短带过、但我觉得潜力巨大的一点,就是人工系统在图形与可视化上所犯错误的「类人程度」。某种意义上,这对人工系统来说尤其有用,因为当我做一张可视化时,我知道别人如果理解正确会得到什么,但我很难想象它可能被怎样误解。所以「这张图还可能被怎样误解」本身就是一个任务,而且我觉得会非常非常有价值。我想问的是,关于图表如何被误解,我们目前最好的模型是什么?通常它们是怎么被误解的?而且这在科学上是不是一个可解释的问题,能让我们看清内部到底发生了什么?

范: 这个问题问得非常非常好,我来正面接一下。在座有些人研究的一个现象叫「视角采择」,它很难,但在某些情境下是做得到的。想象世界从别人的视角看起来是什么样,这是可能的,也有各种范式在研究它。而且看起来,快速而准确地做到这一点的能力是会随练习而改变的,所以专长与经验在塑造这种能力上是有作用的。

范: 这一点的一种表现形式是这样,这里我完全不提视觉语言模型的结果,只勾勒问题的形状:教师,真正高水平的教师所发展出的众多专长中的一种,就是在听学生描述自己怎么想一道题时,能够诊断出所谓的「迷思概念」。学生给出的最终答案可能是错的,但事情不止于「错」。相关的问题往往不只是对还是错,而是:这个迷思概念、这个误解的性质是什么?在规范上正确的表征方式与推进步骤,同学生实际做的事情之间,差在哪里?

范: 如果可以的话,我很简短地说一下,在超大规模机器学习系统的今天,我们是怎么想这个问题的:去追问这些系统在答对或答错时,究竟依赖了哪些运算、哪些概念性的基元,用的是「机制可解释性」名下的那一整套工具。你也可以把它叫作人工神经网络的认知系统神经科学,目的就是诊断正确答案和错误答案分别从哪里来。

范: 那会非常非常好。由于时间关系我不展开细节。但这是一种可能的策略,把这些系统内部执行的、并不能立刻被解读的运算,同人们通常所说的「知识图谱」,也就是一个人心里持有或不持有的那张底层知识网络,连起来。

范: 我想这是一种设定问题的方式。有时候瓶颈就在知觉本身,另一些时候瓶颈或者说落差出在某一个推理步骤上,而目标是把这些步骤真正暴露出来,并且拥有诊断它们的工具,这样原则上你就能用它去诊断任何人、任何系统身上的迷思概念、误知觉、错误和失步。这当然非常非常难。但我认为,正是这些困难的问题,是我们朝那个方向前进时可以借用的工具。希望这些多少讲通了,我关于这个问题的想法很多。

听众: 我能追问一句吗?

范: 可以。

听众: 我觉得,你在这里其实帮了我们一把,因为你把图上该注意的那个位置标了出来。

范: 我确实这么干了。

听众: 我只是想确认自己有没有理解。你在回答这个问题时可能遇到困难的原因之一,是你不知道该注意这东西的哪一部分。

范: 对。

听众: 或者你知道该看哪一部分,但不知道拿它怎么办,不知道怎么把那个红色矩形跟纵轴上的信息对起来。又或者,即便你知道怎么对起来,你也不知道怎么把它转化成问题所要的答案。就是这一类。

范: 是的,而这些正是人类教师可能会采取的教学动作。

听众: 有人甚至试着把它拆开过,比如埃里克·舒尔茨的团队,他们会去问模型或者问人:第二列那个红色矩形有多大?这就不需要任何推理。或者问:哪一个蓝色矩形最高?你们做的是这类事情吗?

范: 是的,正是这样。我们从这类刺激材料开始,这是非常非常新的工作,由亚历克莎·塔尔塔利尼主导,与克里斯·波茨合作。我们从这些真实世界里的图和问题开始,然后意识到我们需要往下钻。所以一方面是拿这些图去问各种不同的问题,包括刚才说的那些更基础的问题;另一方面我们也意识到,得把它剥到简单得多的版本,同时仍保留其中的核心多模态整合挑战。我们现在正在用数轴刺激做实验,本质上是想理解人如何判断数轴上两个数据点之间的距离。总之通往这个问题的路径有很多条,因为关键在于这件事本身很复杂,把它分解开来正是挑战的一部分。

听众: 而且这大概也关系到怎么教这些技能,把它们拆解开来。

范: 正是。对,就是这样。所以既有测评的一面,另外,我之所以在图上放那个炭灰色的方块,是因为我心里有一个猜测:迅速看向图上恰当的那个特征,其瓶颈可能在哪里。

问答:孩子怎么读图

听众: 我能再追问吗?

范: 当然,请。

听众: 我第一眼看这张图的时候,觉得它其实是有歧义的,我不确定这些柱子是不是都从零开始、我们看到的是一根挡在另一根前面。

范: 我知道,我知道。

主持人: 就那些最便宜的……

范: 就好比 dodge 参数设成了 true。

听众: 但这不是我的问题。我的问题是:你一直关注的是成年人,是那些已经接触过大量各类媒介的成年人,问的是怎样呈现信息才能让人高效地读懂。但还有另一个问题。当我们第一次见到一种新的图型时,比如这张图可能真的是我头一次见到这样堆叠起来的,至少已经很久没见过了。

范: 那挺让人兴奋的。

听众: 头一次看到的时候,它并不那么显而易见。

范: 确实不。

听众: 而当我们看惯了某一种,再看到不一样的,反而会变得更难。比如事件相关电位图里负值朝上,每次看到都让我抓狂。

范: 对,就是这样!

听众: 所以我想问,你怎么看这个问题:这类表征该如何设计,才能让孩子容易学会?孩子一开始这些表征一个都没有。你会在最早那批要教给孩子的表征里放进什么?之后又该怎么推进?

范: 这是个好问题。其实是有一些版本的。几分钟前我还和南希聊到一件相关的事,就是人究竟是怎么「打进」这类视觉输入的。某种意义上,你是在大量神经典型的、普通的视觉与认知发展之上再往上建,才能把握到不同形状处在特定空间排布中这件事。它还依赖于识读能力本身。你需要先建立起一批概念性基元和更基础的能力。

范: 也许存在某个经验序列,可以一步步搭上去。这问题问得真好。关于课程设计这个问题我没有答案,但我认为它是个很好的问题。

听众: 看起来皮层后三分之一整个都在做视觉,那是一套极其了不起的机器,能从朝向到高度到形状到景观,抽取出各种各样丰富的视觉表征,而所有这些都是我们可以利用的空间。所以做出一张好图的精髓,似乎就是想清楚怎么设计一个可视化,去接上其中某一部分机器,从而建立映射。

范: 什么样的视觉场景总体上更容易、更流畅地被加工?不那么杂乱的场景。如果你关心的是在这类场景中做视觉搜索的问题,那么其中一部分加工过程会被征用,用来确定该看这张图像的哪个子区域。你可以设想,把最早那批图设计得更贴合这类场景。

范: 然后还有一步,是把这些场景的组成部分映射到你同样需要学习的概念上,这一步一开始会非常难,随后会越来越快、越来越容易。我觉得这个过程中发生的事情真的很迷人。

主持人: 这真的很精彩。要不我们就到这里,让需要离开的朋友先走。

范: 好的。

主持人: 其他人请留步。你还会待一会儿,我们有酒会,大家有四十五分钟左右可以继续盘问你。

范: 太好了。谢谢,谢谢,谢谢。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:04 主持人介绍:从视觉神经科学到人类认知 ▶ 正在看
3:53 数轴到费曼图:让不可见变可见的工具史 ▶ 正在看
9:41 理论框架:认知工具与工程视角的缺席 ▶ 正在看
12:45 图画何以有意义:相似论与约定论 ▶ 正在看
17:40 语境决定抽象层次:绘画游戏实验 ▶ 正在看
21:14 解释图与描绘图:机制优先于外观 ▶ 正在看
29:13 SEVA 基准:素描理解的人机差距 ▶ 正在看
36:48 数据可视化:现代最有力的认知技术 ▶ 正在看
40:54 六项图表测试:模型错得不像人 ▶ 正在看
46:22 选图实验:非专家也懂受众需要 ▶ 正在看
51:35 重审测量:现有测试到底在测什么 ▶ 正在看
57:44 问答:错误图画、误读诊断与儿童读图 ▶ 正在看
本期小问 · 档案清单
3:53 人类心智的什么特性,让数轴、图表这类认知工具能被持续发明? ▶ 正在看
21:14 为什么解释性的图画反而要牺牲逼真度? ▶ 正在看
40:54 机器在图表理解上接近人类得分,为何仍不算真正对齐? ▶ 正在看
51:35 要判断机器是否拥有心智,我们还能依据什么标准? ▶ 正在看
本期讲者
朱迪·范斯坦福大学心理学助理教授,认知工具实验室负责人,研究人类如何用绘画、图表等外部表征进行学习与交流;曾获认知科学学会 Glushko 博士论文奖与 NSF CAREER Award。
乔什·特南鲍姆MIT 脑与认知科学系教授,计算认知科学领军人物,以概率程序、直觉物理与「像人一样学习」的机器研究著称;本场报告的主持人与介绍人。
01主持人介绍:从视觉神经科学到人类认知
0:04
JOSH TENENBAUM: It's my great pleasure to introduce Judy Fan, who's our colloquium speaker. Judy is one of my favorite cognitive scientists really of any age. I think she's definitely one of the leading-- by any objective standard leading of the younger generation people who are-- have just junior faculty or I don't even know how junior you are, but she's-- been assistant professor at Stanford for a couple of years. I think that's a good estimate. And she's by any means one of the rising or glowing superstars of our field.
JOSH TENENBAUM:我非常荣幸地向大家介绍 Judy Fan,她是我们这次学术报告会的演讲人。Judy 是我最喜欢的认知科学家之一,不分年龄段。我觉得她绝对是最顶尖的——按任何客观标准来看,都是年轻一代中的领军人物之一,她们——刚刚成为初级教职,我其实也不太清楚你到底有多资历尚浅,不过她——在斯坦福当助理教授已经有几年了。我觉得这个估计差不多。无论从哪个角度看,她都是我们这个领域冉冉升起、光芒四射的明星之一。
便签笔记
0:37
She's won a number of awards-- the Glushko Dissertation Prize, the NSF Career Award. I think you won that. Real-time talk introduction. But none of that really is-- that's not the reason that we invited her. I think of Judy as one of the most creative researchers I know of any stage in any field, and I really mean that deeply. She has a background coming out of neuroscience, and in that in many ways is it-- is very much at home talking to all the sorts of people who inhabit this building, who record signals from brains and try to make sense of them computationally.
她获得过很多奖项——Glushko 博士论文奖、NSF Career Award(美国国家科学基金会职业生涯奖)。我记得你拿过那个奖。实时演讲介绍。但这些其实都不是——那并不是我们邀请她来的原因。在我认识的所有研究者中,无论处在哪个阶段、哪个领域,Judy 都是最有创造力的之一,这是我发自内心的评价。她的背景出自神经科学,从很多方面来说,她跟这栋楼里的各类人交流起来都非常自如,这些人记录大脑的信号,并试图用计算的方法理解它们。
便签笔记
1:18
She's done a lot of that. She's worked with colleagues of ours whose work we know well like Dan Yamins was one of her mentors, and she has studied vision. She's-- a lot of her background is in visual psychophysics and visual neuroscience and computational neuroscience. But if you look at her research, some of which you'll see today and a lot of which you won't even see today, she's gone in many other directions or let's just say one key direction, which is from basic-- some of the most basic processes of perception and the things that you can describe with nice, elegant computational models to aspects of cognition that really are what make us human, both biologically and also culturally.
这类工作她做了很多。她与我们熟悉的一些同行合作过,比如 Dan Yamins 就是她的导师之一,而且她研究过视觉。她——她的背景很大一部分是视觉心理物理学、视觉神经科学和计算神经科学。但如果你看看她的研究,其中一些今天你会看到,还有很多今天甚至都看不到,她已经走向了许多其他方向,或者说一个关键方向,那就是从最基础的——一些最基本的知觉过程,从那些可以用漂亮、优雅的计算模型来描述的东西,走向认知中那些真正让我们之所以为人的方面,无论是从生物学上还是从文化上。
便签笔记
2:02
And you're definitely going to see some of that here. And so that means she studied things like artistic expression or creative expression in visual and other media. She's very interested in narrative expression. She's really interested in education, learning, and teaching and how we make sense of the world symbolically and through data and explanation. And she's-- so she's really moved towards these much more complex cognitive processes that are really distinctively human both biologically and culturally and increasingly really of great import in our current era.
你在这里肯定会看到一些这样的内容。这也就意味着,她研究的是像视觉媒介和其他媒介中的艺术表达或创造性表达这样的东西。她对叙事表达非常感兴趣。她真的很关注教育、学习和教学,以及我们如何通过符号、数据和解释来理解这个世界。所以她——她真的转向了这些复杂得多的认知过程,这些过程无论在生物学上还是文化上都是人类所独有的,而且在我们当下这个时代变得越来越重要。
便签笔记
2:36
But she hasn't given up any of the rigor or I would say her taste for rigor. And I know from a lot of interactions with Judy that what keeps her up at night is this is a struggle and the challenge of how do you get at these really important, hard things with the kind of rigor and precision that many of us grew up valuing and being inspired by in, for example, visual neuroscience and computational neuroscience. And I can't say she's completely solved that problem because it might take a while, but it's a really great challenge.
但她并没有放弃任何严谨性,或者我该说,她对严谨的那份品味。而且从我和 Judy 的多次交流中我知道,让她夜不能寐的是,这是一场艰难的探索,一个挑战,那就是如何你是否能以我们许多人从小就珍视、并深受启发的那种严谨和精确,去攻克这些真正重要而困难的问题——比如说,比如视觉神经科学和计算神经科学中的那种严谨。我不能说她已经完全解决了那个问题,因为那可能还需要一段时间,但这确实是个非常了不起的挑战。
便签笔记
3:07
It's one that inspires me, the way she works on it, and I hope that it's one that we can all learn from and be inspired by and to see where she is and where she's going with this. So, Judy, please tell us. JUDY FAN: OK wow. [APPLAUSE] I wasn't totally prepared for that. That was incredibly kind, Josh. I really-- thank you and thank you all for being here. It's-- I can't actually express how wonderful today has been so far. I have a lot of affection and respect for this department and community and the kind of values, scientific and otherwise, embodied in the work that you all do here, and so it really is a treat to spend a few minutes telling you about some of the work that we've been doing.
它激励着我,她研究它的方式也激励着我,我希望这也是我们大家都能从中学习、受到启发的,去看看她现在走到了哪里,以及她将带着这个方向走向何方。那么,Judy,请你来跟我们讲讲吧。JUDY FAN:好的,哇。【掌声】我完全没料到会有这样的介绍。Josh,你实在太客气了。我真的——谢谢你,也谢谢在座各位的到来。这——我其实没法表达今天到目前为止有多美好。我对这个系、这个社群,以及你们在这里的工作中所体现出的那些价值观——无论是科学层面还是其他层面的——都怀有深深的喜爱和敬意,所以能花几分钟跟大家讲讲我们一直在做的一些工作,真的是一种享受。
便签笔记
02数轴到费曼图:让不可见变可见的工具史
3:53
So we study cognitive tools. What are those? Let's start with something as familiar and simple as the number line. I've been told not to move around as much because we're using this mic. I'm going to try to do that. Nature didn't give us the number line. We invented it. So as the Spanish architect Antoni Gaudí once put it, there are no straight lines or sharp corners in nature, but that didn't stop us. We created them anyway, and we extended the number line, of course, a few hundred years ago to create rectangular coordinates, which turn out to be super useful.
我们研究的是认知工具。那是什么呢?我们先从一个再熟悉、再简单不过的东西说起——数轴。有人叮嘱我别走来走去,因为我们在用这个麦克风。我会尽量照做。数轴不是大自然给我们的。是我们发明的。正如西班牙建筑师安东尼·高迪曾经说过的,自然界里没有直线,也没有尖锐的棱角,但这并没有阻止我们。我们还是把它们造了出来。当然,几百年前我们还把数轴延伸开去,创造出了直角坐标系,事实证明它极其有用。
便签笔记
4:29
They were genuinely cutting edge tools for thought, for deriving new mathematical discoveries. When René Descartes and his contemporaries realized that you could link up algebraic expressions with geometric curves in order to solve all kinds of mathematical puzzles including this one that has stumped the world for millennia-- how many are familiar with this particular-- the Delian problem, this riddle of the Oracle of Delos? It's the problem of how do you double the volume of a perfect cube? It's really, really difficult to do using the mathematical methods of that community and time.
它们在当时确实是最前沿的思维工具,用来推导出新的数学发现。当勒内·笛卡尔和他同时代的人意识到,可以把代数表达式和几何曲线联系起来,从而解决各种各样的数学难题时——包括这个困扰了世人几千年的难题:有多少人熟悉这个特定的——德利安问题,也就是德洛斯神谕的这个谜题?它的问题是:如何把一个正方体的体积扩大一倍?用那个群体、那个时代的数学方法来做,真的非常非常困难。
便签笔记
5:04
And here's the solution for that. And it'd be really hard to overstate the impact of this invention over the last four centuries. That technology converted all kinds of problems like these where you might want to find the set of values that satisfies two equations into essentially problems that relied on locating points of intersection between two curves. Now imagine what happened next to that revolutionary tool. It became something that we could take for granted. It became so useful that it's become basically indispensable to the way every generation is educated.
而这就是它的解法。要说这项发明在过去四个世纪里的影响力,怎么强调都不为过。这项技术把各种这样的问题——比如你想找出同时满足两个方程的那组值——本质上转化成了依靠寻找两条曲线交点位置的问题。现在想象一下,这个革命性的工具后来变成了什么样。它变成了我们习以为常的东西。它变得如此有用,以至于基本上成了每一代人接受教育的过程中不可或缺的一部分。
便签笔记
5:46
Virtually every mathematics curriculum on the planet introduces a combination of symbolic and graphical notation for representing and manipulating mathematical objects. And the question that we wrestle with is how did we get here and what is it about the human mind that makes that kind of continual innovation possible. There are lots of ways of approaching that question from many different disciplinary perspectives including history, anthropology, economics, and I think that cognitive science also has a really important contribution to make here.
地球上几乎每一套数学课程,都会引入符号记法与图形记法相结合的方式,来表示和操作数学对象。而我们苦苦思索的问题是:我们是怎么走到这一步的?人类心智中究竟是什么,让这种持续不断的创新成为可能?可以从许多不同的学科视角来切入这个问题,包括历史学、人类学、经济学,而我认为认知科学在这里也能做出非常重要的贡献。
便签笔记
6:22
And I think the story begins at least 30,000 to 80,000 years ago when anatomically modern humans began to mark up their physical environments, essentially-- including these cave walls here, iconically repurposing objects and surfaces in their surroundings into carriers of meaning. We obviously didn't stop at cave walls. The story of human learning and discovery is deeply intertwined with the story of technologies for making the invisible visible. So here are just a few examples from the history of science which I love.
我认为这个故事至少要追溯到三万到八万年前,那时解剖学意义上的现代人开始在他们的物理环境上留下标记——包括这里的这些洞穴壁——把周围的物体和表面以图像化的方式重新利用,变成意义的载体。显然我们并没有止步于洞穴壁画。人类学习与发现的历史,与那些让不可见之物变得可见的技术的历史,深深交织在一起。这里是科学史上我非常喜欢的几个例子。
便签笔记
7:01
We've got Darwin's finches and these illustrations produced by John Gould, the ornithologist that Darwin worked very closely with. Only when you see these cases side by side does the kind of morphological variation begin to really become salient and pop out. We have the telescope that Galileo used to observe the movement of the moons around Jupiter. The kind of resolution that he needed in order to question the orthodoxy when it came to how the solar system was organized. We have Ramón y Cajal's much celebrated drawings of the retina as seen under the microscope, showing us what different parts of that-- part of the nervous system were like and how they were hooked up to one another.
我们有达尔文的雀鸟,以及由约翰·古尔德绘制的这些插图——他是与达尔文合作非常密切的鸟类学家。只有当你把这些标本并排放在一起看时,那种形态上的差异才真正开始变得显眼、跳出来。我们有伽利略用来观测木星卫星运动的望远镜。要质疑当时关于太阳系构造的正统学说,他需要的正是那样的分辨率。我们有拉蒙-卡哈尔那些广受赞誉的视网膜显微镜下的绘图,它们向我们展示了神经系统的那一部分中不同区域是什么样子,以及它们彼此是如何连接的。
便签笔记
7:49
And in the 20th century, we have Feynman diagrams, named after, of course, the physicist Richard Feynman, showing us subatomic particles winking in and out of existence, events that we literally could not and will never be able to observe directly with the naked eye. And more than any other species, we leverage this understanding, this expanding understanding of the world in order to-- well, what-- maybe to go back just one beat is that I want to draw your attention to the variation on this slide.
到了20世纪,我们有了费曼图,它当然是以物理学家理查德·费曼命名的,向我们展示亚原子粒子如何倏忽生灭,那些我们真的无法、也永远不可能用肉眼直接观察到的事件。而且比任何其他物种都更甚,我们利用这种理解、这种不断扩展的对世界的理解,去——嗯,也许——让我稍微往回退一步,我想请大家注意这张幻灯片上的差异。
便签笔记
8:28
So notice that some of these images are quite detailed and faithful if you will to the way the visual world looks to us when we open our eyes. Let's take Darwin's finches at the top left, for example. Others are much more schematic, but what all of these examples share in common is they leverage what I've been calling visual abstraction to communicate what we see and know about the world in a format that highlights what is relevant to notice. And then building on that, we leverage that expanding understanding of the natural world through the use of those tools for learning in order to create new things.
注意,其中有些图像相当细致,可以说相当忠实于视觉世界本来的样子。当我们睁开眼睛时呈现给我们的样子。比如左上角达尔文的雀类。另一些则要图示化得多,但这些例子的共同点在于,它们都利用了我一直所说的视觉抽象,以一种能突出关键信息的形式,传达我们对世界的所见与所知。在此基础上,我们借助这些学习工具,运用不断扩展的对自然世界的理解,去创造新的事物。
便签笔记
9:07
So, for example, our detailed understanding of physical mechanics allowed us to design and then build high precision timekeeping devices, and a lot of this technological progress has been driven by our ability to continually reformulate our understanding of the world in terms of those useful abstractions that make it possible to re-engineer the physical world according to our design. So translating biological insights into bioengineering, physical theory into advanced physical instrumentation, neuroscience into medical devices, and quantum mechanics into modern electronics.
比如说,我们对物理力学的细致理解,让我们得以设计并制造出高精度的计时装置,而这些技术进步在很大程度上源于我们不断把对世界的理解重新表述为那些有用的抽象的能力,正是它们让我们能够按照自己的设计去改造物理世界。于是我们把生物学洞见转化为生物工程,把物理理论转化为先进的物理仪器,把神经科学转化为医疗设备,把量子力学转化为现代电子技术。
便签笔记
03理论框架:认知工具与工程视角的缺席
9:41
So this is sampling the kind of phenomena that I take a lot of inspiration from, and the question we continue to wrestle with is what about us makes all of that possible. This is a schematic that I've been using over the last few years to help me think about the key behavioral phenomena in play, and that will also serve as a framework for embedding the different lines of work that I'll be sharing with you. So this is one way of illustrating the traditional mode in cognitive psychology, my home discipline, which focuses on how people process information supplied by the external world.
以上就是我从中获得诸多灵感的那类现象的一个取样,而我们一直在思索的问题是:我们身上究竟是什么让这一切成为可能。这是我过去几年一直在用的一张示意图,它帮助我梳理其中涉及的关键行为现象,同时也将作为一个框架,用来串联我接下来要与各位分享的不同研究工作。这是说明认知心理学(我的本行学科)传统研究模式的一种方式,它关注的是人如何加工处理来自外部世界的信息。
便签笔记
10:14
Here's how that picture is enriched by the study of social cognition, which considers the behavior of multiple individuals at once and how they interact with each other. When those activities are used in the service of learning about the world and sharing that knowledge with others, they've been argued to share important similarities with formal science even when pursued by non-experts in everyday contexts. So building on that tradition and with the goal of understanding how humans made all those remarkable discoveries and inventions that I shared with you just a minute ago, I'd like to argue that there's still two critical ingredients missing from this picture.
而社会认知的研究让这幅图景变得更丰富,它同时考虑多个个体的行为以及他们之间如何互动。当这些活动被用来了解世界、并把这些知识分享给他人时,有人认为它们与正规科学有着重要的相似之处,即便进行这些活动的是非专业人士、场景也只是日常生活。所以,在这一传统的基础上,为了理解人类是如何做出我刚才和大家分享的那些非凡发现和发明的,我想提出:这幅图景里还缺了两个关键要素。第一个缺失的要素。
便签笔记
10:46
First is an account of cognitive tools or technologies, material objects that encode information intended to have an impact on our minds, how and what we think. Second, I would like to argue that the time has come to embrace science's natural complement, engineering, how people leverage their understanding of the world whether it's through direct experience or socially mediated to create new and useful things. Because without a serious consideration of the engineering half of this picture, I would venture to say that we'll never be able to explain how and why the world as we know it came to be.
首先是对认知工具或认知技术的解释——那些编码了信息、意在影响我们的心智,影响我们如何思考、思考什么的实物。其次,我想说,是时候接纳科学的天然补充面了,那就是工程,也就是人们如何利用自己对世界的理解——无论这种理解来自直接经验还是社会传递——去创造新的、有用的东西。因为如果不认真对待这幅图景中属于工程的那一半,我敢说我们永远无法解释我们所知的这个世界是如何、又为何变成今天这样。
便签笔记
11:28
So at its core, if I had to state it really bluntly, research in my group aims to close this loop to develop psychological theories that explain how we go about discovering useful abstractions that explain how the world works jointly with theories that explain how we then apply those abstractions to go make new things. And my plan today is to tell you about some of the work that we've done so far along two lines. In the first part, I'm going to share with you what I think we've figured out about how people leverage visual abstraction to communicate semantic knowledge using freehand drawing as a central case study there.
所以从根本上说,如果要我直白地讲,我们团队的研究就是想闭合这个环路,去发展这样一些心理学理论——解释我们如何去发现那些能说明世界如何运作的有用抽象,同时又发现那些能说明我们随后如何运用这些抽象的理论,去创造新的东西。我今天的计划是,向大家介绍我们目前在两条主线上所做的一些工作。在第一部分,我会和大家分享我认为我们已经弄明白的一些东西:人们如何利用视觉抽象来传达语义知识,这里以徒手绘画作为一个核心的案例研究。
便签笔记
12:06
Ordinarily, in the second part, I would then have told you about our work investigating how people learn and coordinate on procedural abstractions when building physical things aligned with the engineering segment. But today, I want to try something new with you all. So I wanted to instead share with you some of our emerging work. So I'm really curious to hear what you think exploring the cognitive foundations of data visualization in which people harness multiple information modalities, graphical elements, words, and numbers to engage in statistical reasoning.
按照惯例,在第二部分,我本来会讲我们的另一项研究:人们在搭建实体物件时,如何学习并协调程序性抽象,这部分和工程那一环节相呼应。但今天,我想和大家尝试点新东西。所以我想改为和大家分享我们一些正在成形的工作。我很想听听你们的想法:我们在探索数据可视化的认知基础,在这个过程中,人们会调用多种信息模态——图形元素、文字和数字——来进行统计推理。
便签笔记
04图画何以有意义:相似论与约定论
12:45
In other words, to learn from finite amounts of evidence about aspects of the world that might be difficult or impossible to learn through direct observation by a single individual. And I'm still happy to chat with folks about our work on physical assembly and physical reasoning probably at the reception. So to dive into part one, how do we begin to think about how people use visual abstraction to communicate what they know and what they see? Here I think it's useful to think about three behaviors that build successively on one another.
换句话说,就是从有限的证据中,去了解这个世界上那些单靠个人直接观察可能很难、甚至不可能了解的方面。靠一个人是做不到的。关于我们在实体组装和物理推理方面的工作,我也很乐意在酒会上和大家聊聊。那么进入第一部分:我们该如何开始思考人们如何用视觉抽象来传达他们所知道的和所看到的东西?我觉得这里可以从三种层层递进的行为来思考。
便签笔记
13:15
So first, of course, visual perception, the problem of how we transform raw sensory inputs into semantically meaningful perceptual experiences that in turn makes it possible to even contemplate visual production, that ability to generate a set of markings that leave a meaningful and visible trace in the physical environment. And those come together during visual communication, how we decide just how to arrange those graphical elements and in what order in order to have a particular kind of impact on other minds whether it is to inform or teach, persuade, collaborate, or any other purpose that we can put those marks to.
首先,当然是视觉知觉,也就是我们如何把原始的感官输入转化为具有语义意义的知觉体验,而这又反过来让我们有可能去构想视觉产出,也就是生成一组标记的能力,这些标记会在物理环境中留下有意义且可见的痕迹。这些能力在视觉交流中汇合到一起:我们如何决定这些图形元素该如何排布,以及以什么顺序排布,才能对他人的心智产生某种特定的影响——无论是为了告知、教学、说服、协作,还是我们能用这些标记去实现的任何其他目的。
便签笔记
13:56
So this is a-- this is going to be an overview of the three studies that I'll be walking through. First, we'll begin with the question of what is the perceptual basis for understanding what a certain class a subclass of pictures represents. So to get off the ground, we started by considering maybe the most concrete and familiar instantiation of visual abstraction, creating a drawing by hand that looks like something in the world. What makes it so easy to tell that the drawing on the left of this slide is meant to correspond to the realistic bird rendering on the right.
那么这就是——这将是我要讲的三项研究的一个总览。首先,我们从这个问题开始:理解某一类、某一子类的图画所表征的是什么,其知觉基础是什么。为了起步,我们先从视觉抽象最具体、最熟悉的体现形式入手,也就是手工画出一幅看起来像世界中某个东西的画。是什么让我们如此轻易地看出,这张幻灯片左边的这幅画对应的是右边那张写实的鸟类图像?
便签笔记
14:42
There have been a lot of different ways of responding to that question. Two have been really dominant. So the first is that-- the first view is that we essentially see drawings as meaningful and representing things because drawings simply resemble objects in the world. Like this drawing literally looks like the bird and that's how we know. The second response is that drawings denote objects primarily as a matter of convention, and we only learn which drawings go with which objects in meanings from other people, so if you take this Chinese character, for example.
对这个问题,人们提出过很多不同的回答。其中有两种一直占据主导地位。第一种观点是——第一种观点认为,我们之所以把图画看作有意义、能表征事物的东西,是因为图画本身就与世界中的物体相像。比如这幅画看起来就是那只鸟的样子,我们就是这样认出来的。第二种回应是,图画之所以能指代物体,主要是一种约定俗成,我们只是从其他人那里学到了哪些图画对应哪些物体和含义,比如说这个汉字。
便签笔记
15:23
So in an earlier line of work, my collaborators-- close collaborators Dan Yamins and my PhD advisor Nick Brown, and I discovered that general purpose vision algorithms-- so in this case neural networks that compose multiple layers of learnable spatial convolutions trained on natural photographs-- were capable of generalizing fairly strongly to even quite sparse sketches that didn't look photo-realistic per se, suggesting that the problem of pictorial meaning and resemblance might be resolved simply by building better models of the ventral stream of visual processing, especially ones that accurately capture those operations performed in those brain regions.
在早期的一项研究中,我的合作者们——密切合作者 Dan Yamins 和我的博士导师 Nick Brown,还有我,发现通用视觉算法——在这里指的是由多层可学习的空间卷积组成、在自然照片上训练的神经网络——能够相当强地泛化到即使是非常稀疏、看起来并不就照片写实性本身而言,这暗示图像意义和相似性的问题也许可以简单地通过构建更好的腹侧视觉通路加工模型来解决,尤其是那些能准确刻画这些脑区所执行运算的模型。
便签笔记
16:06
We were not the only ones to discover this. There are easily thousands of computer vision papers indexed by Google Scholar that use some kind of ConvNet or other neural network based backbone to encode sketches and natural images for a variety of applications. All of those results you can think of as vindicating an updated modern resemblance based account. It was also a really useful insight for us locally with other practical consequences. So in some more recent work led by Charles Lu, a former master's student in the lab and in collaboration with Xiaolong Wang at UCSD, we built on top of those early findings to further stress test that resemblance account, basically taking a ConvNet backbone and then training a decoder on top that could map local elements in a sketch to particular-- to corresponding elements in a photograph under the constraint that you could warp or rumple the sketch but not tear any holes in it.
发现这一点的并不只有我们。在 Google Scholar 上收录的计算机视觉论文里,轻轻松松就有上千篇使用了某种 ConvNet 或其他基于神经网络的骨干网络,来对草图和自然图像进行编码,用于各种各样的应用。所有这些结果,你都可以视为对一种更新过的、现代版基于相似性的解释的支持。对我们本地的研究来说,这也是一个非常有用的洞见,还带来了其他一些实际影响。在最近的一项由 Charles Lu 主导的工作中——他是实验室之前的一位硕士生——我们还与 UCSD 的王小龙合作,在那些早期发现的基础上,进一步对这个「相似性」解释做压力测试,基本上就是拿一个 ConvNet 主干网络,然后在上面训练一个解码器,把草图中的局部元素映射到照片中对应的元素,同时施加一个约束:你可以对草图做扭曲、褶皱,但不能在上面撕出任何洞。
便签笔记
17:07
And the success of that approach-- I'm just demoing here. It's not really a result. It's more of a demo-- the success of that approach suggests that fairly strong spatial constraints govern how the parts of sketches correspond to the parts of real objects that they're meant to represent. So that's great. We've got good enough trainable models of sketch understanding to build downstream applications that work pretty well. I guess we can pack up and go home. But, of course, that's not the whole story.
这个方法的成功——我这里只是做个演示。它其实不算是一个结果,更像是一个 demo——这个方法的成功表明,草图的各个部分与它们所要表征的真实物体的各个部分之间的对应关系,受相当强的空间约束支配。所以这很棒。我们已经有了足够好的、可训练的草图理解模型,可以用来构建效果相当不错的下游应用了。我想我们可以收拾东西回家了。但当然,这并不是故事的全部。
便签笔记
05语境决定抽象层次:绘画游戏实验
17:40
A static deterministic account of visual processing falls short of explaining how we generate and make sense of drawings like these, which you might find all around this building. What these blobs and boxes, squiggles, and arrows mean depends on what we're talking about. So our next goal was to figure out how to incorporate that information about context to begin to account for a greater variety of graphical representations that we, in fact, use to communicate, ranging from more faithful pictures like those on the left here to the more obvious symbolic tokens on the right.
一个静态的、确定性的视觉加工解释,无法说明我们是如何生成并理解这类图画的——这类图画你在这栋楼里到处都能看到。这些团块、方框、波浪线和箭头到底是什么意思,取决于我们在谈论什么。所以我们的下一个目标,就是搞清楚如何把关于语境的信息整合进来,从而开始解释更大范围的我们实际上用来交流的各种图形表征方式,从左边这些更写实的图画,到右边那些更明显的符号化标记,跨度很大。
便签笔记
18:19
Our first paper tackling that issue asked how people knew when they needed to produce a more faithful drawing and when they could get away with something more schematic or abstract. So in that study, we paired two people up to play a drawing game. The sketcher saw a display that looked something like this where their goal was to draw the highlighted target object, the third one here, and we varied what the other objects in the display were. On close trials, the distractors all belong to the same basic level category whereas on far trials the distractors were from different categories.
我们探讨这个问题的第一篇论文,问的是:人们怎么知道什么时候需要画得更写实,什么时候可以画得更简略、更抽象也没关系。所以在那项研究里,我们把两个人配成一组,玩一个画画游戏。画的人会看到一个大致长这样的画面,他的目标是把高亮的那个目标物体画出来,就是这里的第三个,而我们会改变画面中其他物体是什么。在「近」试次中,干扰项都属于同一个基本层次类别;而在「远」试次中,干扰项则来自不同的类别。
便签笔记
18:56
Using this really simple manipulation, we discovered how readily ordinary folks can adjust the way they use depiction to communicate, making more detailed and faithful drawings on those close trials when they needed their drawing to uniquely identify a particular exemplar but then sparser drawings on the far trials when they could get away with these category level abstractions. So here are some examples of actual drawings that we collected in that study. And we found that sketchers used fewer strokes on far trials, less ink, less time to produce those drawings while still achieving sealing accuracy on the fundamental task of communicating the identity of the target visual concept to the viewer, who themselves took less time on those trials to make their decisions.
通过这个非常简单的操作,我们发现普通人可以多么轻易地调整他们使用图画来交流的方式,在'近'条件的试次中,当他们需要用画来唯一地指认某个特定样例时,就画得更详细、更忠实于原物,而在'远'条件的试次中,当他们可以靠这些类别层面的抽象蒙混过关时,就画得更简略。这里是我们在那项研究中收集到的一些实际画作的例子。我们发现,在'远'试次中,作画者用的笔画更少、墨水更少、花的时间也更少,但在向观看者传达目标视觉概念身份这一基本任务上,仍然达到了同样的准确率,而观看者在这些试次中做出判断所花的时间也更少。做出他们的判断。
便签笔记
19:45
Then to capture that behavioral pattern, we proposed a computational model of the sketcher that consisted of two parts. So a ConvNet to encode visual inputs into a generically abstract feature space, and then a second probabilistic decision making module that inferred what kind of drawing to make depending on the context. I'm just going to give you the take home from that study because I'm really excited to get to other work in this talk. So the take home from our model ablation experiments was that both the capacity for visual abstraction, which you operationalize as the network layer in the visual encoder module and sensitivity to context were critical for capturing how people manage to communicate about these objects at the appropriate level of abstraction.
为了捕捉这种行为模式,我们提出了一个关于'绘图者'的计算模型,它由两部分组成。一部分是卷积神经网络(ConvNet),把视觉输入编码到一个通用的抽象特征空间;另一部分是概率决策模块,它会根据语境推断应该画出什么样的画。我这里只讲这项研究的结论,因为我很想赶紧讲到这次报告里的其他工作。我们模型消融实验的结论是:视觉抽象的能力(我们把它操作化为视觉编码模块中的网络层)和对语境的敏感性,这两者对于捕捉人们如何在恰当的抽象层次上传达这些物体,都是至关重要的。也就是在合适的抽象层级上进行沟通。
便签笔记
20:34
And then in more recent work, we pushed that idea even further to understand not only the impact of the current referential context on how people communicate but also the conditions under which new graphical conventions might emerge when memory for previous interactions with the same person leads people to produce even more abstract, maybe even proto symbolic tokens over time whose meaning depends even more strongly on that shared history. So all of that was really exciting progress to me, but, of course, people possess much richer knowledge about the world than just what things are called or what they look like.
在更近期的工作中,我们把这个想法又往前推进了一步,不仅想理解当前的指称语境如何影响人们的沟通方式,还想理解在什么条件下会涌现出新的图形约定——当人们记得自己之前与同一个人的互动时,他们会随着时间产生更加抽象、甚至近乎原型符号的记号,而这些记号的含义会更强烈地依赖于双方共享的互动历史。所有这些进展都让我非常兴奋,但当然,人们对世界的知识远比'东西叫什么名字'或者'它们长什么样'要丰富得多。
便签笔记
06解释图与描绘图:机制优先于外观
21:14
And a particularly important way that we use visual abstraction, especially in science, is to transmit mechanistic knowledge about how things work. So what is going on in people's minds when they make that move? Going beyond what is visually salient let's say about a specific bird to highlight underlying physical mechanisms, for example, how birds achieve flight in general. So you can imagine my excitement when Holly Huey, a wonderful former PhD student in the lab who's now at Adobe Research, was also fascinated by this question.
而我们使用视觉抽象的一个特别重要的方式,尤其是在科学领域,就是传递关于事物如何运作的机制性知识。那么当人们做出这种转变时,他们的头脑里究竟发生了什么?超越视觉上最显眼的东西——比如某种特定的鸟——去揭示底层的物理机制,比如鸟类总体上是如何实现飞行的。所以你可以想象我有多兴奋,Holly Huey,我们实验室一位非常出色的前博士生,现在在 Adobe Research 工作,她也对这个问题着迷。
便签笔记
21:49
And while we knew from very cool work by Frank Kyle and colleagues that people privilege mechanistic explanations when learning from others and from work by Barbara Tversky, Micki Chi, Tania Lombrozo, and others that people can learn by producing explanations, we realized there was a lot that we didn't know about how people thought about visual explanations like what do people think is supposed to go into a diagram that illustrates how something works and what makes those different from an ordinary illustration that is just intended to look like something.
我们从 Frank Kyle 和同事们非常酷的研究中知道,人们在向他人学习时更看重机制性的解释,也从 Barbara Tversky、Micki Chi、Tania Lombrozo 等人的研究中知道,人们可以通过产出解释来学习,但我们意识到,关于人们如何看待视觉解释,我们还有很多不了解的地方——比如人们认为一张用来说明某个东西如何运作的图解里应该包含什么,以及这些图跟只是想画得像某个东西的普通插图有什么不同。
便签笔记
22:25
One possibility that I'll sketch, which I'll call the cumulative hypothesis, is that people basically think of visual explanations as being like extended augmented versions of ordinary depictions. So in this table, explanations would have all the things that depictions have to communicate visual appearance and then tack on information about physical mechanism. An alternative possibility, which I'll call the dissociable hypothesis, is that people think of explanations as being images that pick out mechanistic abstractions while greatly de-emphasizing visual appearance.
我要勾勒的一种可能性,我称之为累加假说,是说人们基本上把视觉解释看作普通描绘的扩展增强版本。所以在这张表里,解释会具备描绘所拥有的一切用来传达视觉外观的东西,然后再附加上关于物理机制的信息。另一种可能性,我称之为可分离假说,是说人们认为解释是那种挑出机制性抽象、同时大大弱化视觉外观的图像。
便签笔记
23:01
There's a kind of selectivity there. So to tease apart those possibilities, Holly designed a study to probe this question in two ways, first, by characterizing the content of visual explanations in detail and comparing them to visual depictions, and, second, to measure how well either of those kinds of images actually help downstream viewers perform the task-- extract information that they really needed whether it's about appearance or about mechanism. So rather than a lesson on bird flight, which is fascinating but complicated, Holly constructed six novel contraptions with a clearly observable mechanism for closing a circuit and turning on a light.
那里存在某种选择性。所以为了区分这两种可能性,Holly 设计了一项研究,从两个方面来探究这个问题:第一,详细刻画视觉解释的内容,并把它们和视觉描绘作比较;第二,衡量这两类图像各自在多大程度上真正帮助下游的观看者完成任务——提取他们真正需要的信息,无论是关于外观的还是关于机制的。所以,比起讲鸟类飞行这一课——虽然很迷人但太复杂——Holly 制作了六个新奇的装置,它们都有一个清晰可观察的机制,用来闭合电路并点亮一盏灯。
便签笔记
23:43
So here's an example of one of those machines and the instructional video that participants watched.
这是其中一台机器的例子,以及参与者观看的教学视频。
便签笔记
23:58
Participants in the study actually saw that twice-- those who demonstrated twice, but there you go. So notice that this machine consists of three different kinds of parts. There are causal parts that need to rotate to turn the light on. There are non-causal parts that look very similar but don't actually cause the light to turn on. And then background structural elements that are colorful and hoist up the other parts are really important but don't directly participate in the light activation circuit.
研究中的参与者其实看了两遍——演示了两遍,就是这样。注意,这台机器由三种不同的部件组成。有因果部件,它们需要转动才能把灯点亮。有非因果部件,它们看起来很像,但实际上并不会让灯亮起来。还有背景性的结构元件,它们色彩鲜艳、支撑起其他部件,非常重要,但并不直接参与点灯的电路。
便签笔记
24:27
Every participant produced explanations of some machines and depictions of other ones. On explanation trials, they were asked to imagine that their drawing would be used by someone else to understand how the object worked. And then on the depiction trials, they were asked to imagine that their drawing would be used by someone else to identify which object it was out of a lineup of similar looking objects. That's manipulation. Really simple. Using that procedure, we collected a large number of depictions and explanations of each of these six machines.
每位参与者都为一些机器画解释图,为另一些机器画描绘图。在解释试次中,我们请他们设想自己的画会被别人用来理解这个物体是如何运作的。而在描绘试次中,我们请他们设想自己的画会被别人用来从一排外观相似的物体中辨认出是哪一个。这就是我们的操纵变量。非常简单。用这个流程,我们收集到了这六台机器各自大量的描绘图和解释图。
便签笔记
25:00
Eyeballing them, the drawings seemed to look different. For example, maybe there's a little bit more background in the depictions. Maybe there's some more arrows in the explanations. But Holly wanted to be really systematic about this, so she crowdsourced tags, assigning every single stroke in every drawing to one of four categories, so one for each of the three kinds of physical parts, background, causal and non-causal, and then her fourth category that was a catch all for symbols. So in this particular context, we're talking about arrows and motion lines and then used those tags to compare what kind-- what balance of semantic information people emphasize in these two conditions.
粗看之下,这些画似乎确实不一样。比如,描绘图里的背景也许稍微多一点。解释图里也许箭头更多一些。但 Holly 想做得非常系统,所以她众包了标注,把每一幅画中的每一个笔画都归到四类之一:三类物理部件各一类——背景、因果和非因果——然后她的第四类是一个用来兜底的符号类别。所以在这个具体语境下,我们说的是箭头和运动线,然后用这些标注来比较人们在这两种条件下强调的是哪种——什么样的语义信息配比。
便签笔记
25:46
What she found was that while people drew causal and non-causal parts in both conditions, they allocated more strokes to the causal part and explanations than the non-causal ones. They also emphasize the background a bit more in depictions than explanations and as we suspected spent more of their ink on symbolic displays of motion and parts interacting in explanations than depictions. So already these results are incompatible with a strong version of the cumulative hypothesis, which would have predicted the relative emphasis on background.
她发现,虽然人们在两种条件下都会画因果部件和非因果部件,但在解释图中,他们分配给因果部件的笔画比非因果部件更多。他们在描绘图中对背景的强调也比解释图稍多一些,而且正如我们所猜想的,他们在解释图中把更多的笔墨花在了表示运动和部件互动的符号表现上,多于描绘图。所以这些结果已经与累加假说的强版本不相容了,因为该假说会预测,对背景的相对强调,
便签笔记
26:26
Causal and non-causal parts would be maintained even in the presence of errors. But however reliable, those differences might just amount to stylistic variation that doesn't have any impact on how useful they are for the tasks that they are intended to support. So to measure those functional consequences of those decisions, Holly designed three different inference tasks. The first asked how easily you could tell what kind of action would be needed to operate the machine, so here pull, rotate, or push, what you might expect a good mechanistic explanation to make clear.
以及因果与非因果部件的相对强调,即便在出现误差的情况下也应保持不变。但无论这些差异多么可靠,它们也可能只是风格上的变化,对这些图在其本应支持的任务中有多大用处并没有任何影响。所以,为了衡量这些决策带来的功能性后果,Holly 设计了三种不同的推断任务。第一种考察的是,你能多容易地看出操作这台机器需要哪种动作——是拉、转还是推,这正是你会期待一个好的机制性解释讲清楚的事。
便签笔记
27:06
The second was to measure how well each drawing could be used for object identification, exactly what depictions are supposed to do. And then the third was a more challenging visual discrimination task where you had to determine which of two highlighted parts was the causal one, requiring you to establish a detailed mapping between the parts of the drawing to parts of the actual machine, something that you really only expect from a mechanistic explanation that also preserved enough information about the overall appearance and organization of the parts of the machine.
第二种是衡量每幅画在物体辨认上的效果如何,而这正是描绘图本该做到的事。第三种则是一个更有挑战性的视觉辨别任务,你得判断两个被高亮的部件中哪一个是因果部件,这要求你在画中的部件与实际机器的部件之间建立起细致的对应关系,而这是你只有从一个同时保留了足够多整体外观信息、以及机器各部件组织方式信息的机制性解释中才会指望得到的。
便签笔记
27:43
So under the cumulative hypothesis, what we should see is that explanations are at least as good as depictions for all tasks but under the dissociable hypothesis that they might be better for the action task but worse for the object task. And what Holly found was more consistent with the dissociable account where explanations better communicated how the mechanism worked. But depictions were better for communicating object identity. Interestingly, explanations were not better in this sample for conveying the identity of the causal part in that challenging third task consistent with the idea that by leaving out a lot of the background details and some of those explanations that they may have abstracted away the very information that would make it easier to link up specific parts of the drawing with parts of the machine.
所以在累积假说下,我们应该看到的是:在所有任务上,解释图至少和写实描绘一样好;但在可分离假说下,它们在动作任务上可能更好,而在物体任务上更差。而 Holly 发现的结果更符合可分离的解释:解释图能更好地传达机制是如何运作的。但写实描绘在传达物体身份方面更胜一筹。有意思的是,在这个样本里,解释图在第三个有挑战性的任务中并没有更好地传达因果部件的身份,这与下面这个想法是一致的:由于省略了大量背景细节,有些解释图可能恰恰把那些本可以让人更容易把画面中的特定部分和机器的部件对应起来的信息给抽象掉了。
便签笔记
28:33
And the bottom line from those studies is that people share intuitions about what is supposed to go into a visual explanation even if this is the first time they've been asked to generate one. And it can mean sacrificing visual fidelity to emphasize more abstract, mechanistic information. And more generally this work shows how important communicative context and goals are for understanding why people draw the way they do, why depictions look the way they do, and provided some experimental and analytic tools for characterizing the strategies people use to communicate visual information that is goal and context relevant.
这些研究的核心结论是:人们对于一张视觉解释图里应该包含什么,有着共通的直觉,哪怕这是他们第一次被要求画这样一张图。而这可能意味着牺牲视觉的逼真度,来突出更抽象的、机制性的信息。更宽泛地说,这项工作显示了交流情境和目标对于理解人们为什么会那样画,以及描绘为什么会呈现出那种样子,是多么重要;同时也提供了一些实验和分析工具,用来刻画人们在传达与目标和情境相关的视觉信息时所采用的策略。
便签笔记
07SEVA 基准:素描理解的人机差距
29:13
And then in this next section, we'll be asking what would it take to develop artificial systems that are capable of human-like visual abstraction. And here's why. We fundamentally want useful scientific models of visual communication, and the work I've presented so far is representative of where we've prioritized making investments, namely the development of experimental paradigms and data sets to measure and characterize those behaviors in fuller ways across a broader range of settings. Meanwhile, we've also been investing considerable energy in evaluating how well any member of the steadily advancing cohort of machine learning systems might continue to be relevant and promising candidates for capturing more detailed patterns of human behavior in these high dimensional tasks, specifically as models of human image understanding as well as models of image creation.
接下来这一部分,我们要问的是:要开发出具备类人视觉抽象能力的人工系统,需要什么条件。原因是这样的。我们根本上想要的是关于视觉交流的、有用的科学模型;而我到目前为止介绍的工作,代表了我们优先投入的方向,也就是开发实验范式和数据集,以便在更广泛的情境下更全面地测量和刻画这些行为。与此同时,我们也投入了相当多的精力去评估:在不断进步的机器学习系统队列中,哪些成员可能持续具有相关性、有希望成为在这些高维任务中捕捉更细致人类行为模式的候选模型,具体来说就是作为人类图像理解的模型,以及图像生成的模型。
便签笔记
30:11
So the task setting that we consider takes inspiration from this famous series of drawings by Pablo Picasso. Some of these are very detailed. The last few are very, very abstract yet all unmistakably bulls. Any scientific model of human visual abstraction worth its salt ought to be able to represent the ways in which these bowls are all different from one another and somehow at the same time all bullish to their core or rather as bullish as they actually look to real human observers. So in this work, that was a huge team effort led by Kushin Mukherjee, who is defending this Friday before joining the lab as a postdoc, with contributions again from Holly as well as Charles Lu, Yael Winkler, who's here actually, and Rio Aguina-Kang.
我们考虑的任务设置,灵感来自毕加索这一组著名的系列素描。其中有些非常细致。最后几幅极其抽象,但无一例外都明白无误地是公牛。任何称得上合格的人类视觉抽象科学模型,都应当能够表征这些公牛彼此之间有何不同,同时又在某种意义上骨子里都很“公牛”,或者说,其“公牛程度”要与真实人类观察者感受到的一致。这项工作是一次巨大的团队努力,由 Kushin Mukherjee 领衔——他本周五就要答辩,之后会以博士后身份加入实验室——同样也有 Holly 的贡献,还有 Charles Lu、Yael Winkler(她其实今天就在现场),以及 Rio Aguina-Kang。
便签笔记
31:05
We began from the premise that a strong test of whether we're on the right track towards those scientific models is that we'll be able to build algorithms that can generate and understand abstract images the way that people do. Sketch understanding is one of these deceptively simple yet poses a fundamental challenge for general purpose vision algorithms because, for one, it requires robustness to variation in sparsity like some sketches are more detailed than others you could say. And, two, because sketches demand tolerance for semantic ambiguity because sketches can reliably evoke multiple meanings.
我们的出发点是这样一个前提:要检验我们是否走在通往那些科学模型的正确道路上,一个强有力的检验就是看我们能否构建出像人一样生成和理解抽象图像的算法。素描理解就是这样一类看似简单、实则对通用视觉算法构成根本挑战的问题,首先,它要求对稀疏程度的变化具有鲁棒性——可以说有些素描比另一些更细致。其次,素描要求对语义歧义有容忍度,因为素描可以稳定地唤起多种含义。
便签笔记
31:53
So we created a benchmark which we call SEVA to pose those challenges explicitly. So we collected 90,000 hand-drawn sketches made by about 5,500 people of 128 visual concepts under varying production budgets. So here are some examples of what it looks like when people had to create sketches cued by a photo. These are photos taken from the THINGS data set, if you're familiar with that, in less and less time. So by the time we get to the four seconds, they're real, real sketchy. We then took those drawings and then showed them to both people and 17 different then state-of-the-art vision algorithms representing a broad array of different kind of architectural commitments and strategies and training protocols.
所以我们创建了一个基准测试,我们称之为 SEVA,来明确地提出这些挑战。我们收集了约 5500 人手绘的 9 万张素描,涵盖 128 个视觉概念,并设置了不同的绘制预算。这里有一些例子,展示了人们在看到一张照片提示后画出的素描是什么样子。这些照片取自 THINGS 数据集,如果你熟悉的话;给的时间越来越短。所以到了四秒这一档时,画得是真的非常、非常潦草。然后我们把这些画拿去,同时给人和 17 种当时最先进的视觉算法看,这些算法代表了各种各样不同的架构取向、策略和训练方案。
便签笔记
32:40
Yes. TENENBAUM: Quick question of clarification. Do people get to think about it before you start the timer, or do they see it and then they have four seconds? JUDY FAN: That's something that's been haunting me for two years. Not enough. I think they-- no no, no. And I think-- so what we're studying-- TENENBAUM: That's finding both thinking and drawing. JUDY FAN: Thinking and drawing. So they see the image, and then they have to go OK. So this is-- so we don't know exactly how-- we-- there is this restriction on the overall trial duration.
请讲。TENENBAUM:一个澄清性的小问题。被试在你开始计时前有时间先想一想吗,还是他们一看到图就只有四秒?JUDY FAN:这个问题困扰了我两年。时间是不够的。我觉得他们——不不不。我觉得——所以我们研究的是—— TENENBAUM:那是把思考和作画一起算进去了。JUDY FAN:思考加作画。他们看到图像,然后就得马上开画。所以这个——我们并不确切知道具体是怎样的——我们——这里是对整个试次总时长的限制。
便签笔记
33:08
I think that to really get at the bowl phenomenon, I would want to give them unlimited planning time and then just limited execution time, and that would have been the way to do it. This is not that. So there's another data set that has yet to be born that isolates that more directly. Yes. So the four-second drawings are quite derpy. That's the technical term of art for that, but they are what people did in this-- in those settings. So we had all of these different sketches. They look the way they look under those conditions.
我觉得要真正抓住那个“公牛”现象,我会想给他们无限的规划时间,然后只限制执行时间,那才是应该采取的做法。但这个实验不是那样的。所以还有另一个尚未诞生的数据集,能更直接地把那一点分离出来。请讲。所以四秒的那些画相当“抽象派”。这是学术上的专业说法啦,但那确实就是人们在那种设置下画出来的东西。所以我们就有了所有这些不同的素描。在那些条件下,它们看起来就是那个样子。
便签笔记
33:41
Then both people and those vision algorithms perform the same sketch categorization task, allowing us to measure the full distribution of labels evoked by every individual sketch. The first thing we established was that as you give people more time to think and make a detailed-- think about how to make and then go and make a detailed sketch, those sketches get more recognizable to models and people. They're less ambiguous as measured by the entropy of the label distribution. And even when the guess is wrong, not the top one, it's more likely to at least be in the right semantic neighborhood estimated by language embeddings.
然后人和那些视觉算法执行同样的素描分类任务,这让我们能测量出每一张素描所唤起的完整标签分布。我们确立的第一件事是:当你给人更多时间去思考、去构思如何画并且真的画出一张细致的素描时,这些素描对模型和对人来说都变得更容易辨认。以标签分布的熵来衡量,它们的歧义性更低。而且即使猜错了、最高概率的那个不对,它也更可能至少落在正确的语义邻域里——这是用语言嵌入估计的。
便签笔记
34:22
So that's reassuring, but then we dug in a little bit deeper and found that while some models do honest to goodness perform better the recognition task in other models, the variation across models in terms of performance is totally dwarfed by the gap between models and people, the reliable signal in human recognition, in terms of both performance on the left and also relative uncertainty about the meaning of a sketch. So that suggests there's still a sizable human model gap in alignment to close here in terms of-- for sketch understanding.
这令人放心;但接着我们深入挖了一下,发现虽然确实有些模型在识别任务上比其他模型表现得更好,模型之间的表现差异却被模型与人之间的差距彻底盖过了——也就是人类识别中的可靠信号,无论是左边的表现指标,还是对一张素描含义的相对不确定性,都是如此。这说明在素描理解方面,人与模型之间仍有相当大的对齐差距有待缩小。
便签笔记
34:59
Nevertheless, at the time we conducted this benchmark, we noticed that the clip trained models were outperforming the others, making it a reasonable candidate to begin to explore generative models of sketch production built on top. So we explored the capabilities of a particularly cool sketch generation algorithm named CLIPasso whose development was led by Yael Winkler. We asked CLIPasso to generate some sketches as well and also manipulated its production budget. And in this case, the unit now, the currency, is the number of strokes, which is different, which I know.
尽管如此,在我们做这个基准测试的当时,我们注意到用 CLIP 训练的模型表现优于其他模型,这让它成为一个合理的选择,可以在其之上开始探索素描生成的生成式模型。于是我们探索了一个特别酷的素描生成算法的能力,它叫 CLIPasso,其研发由 Yael Winkler 主导。我们也让 CLIPasso 生成了一些素描,并同样操纵了它的绘制预算。而在这种情况下,单位、也就是那个“货币”,变成了笔画数量,这不一样,我知道。
便签笔记
35:36
And what we found is that even while human drawings and CLIPasso's drawings were similarly recognizable across the four production budgets-- by the way, this overlap in recognition-- in recognizability is totally coincidental given how the units are totally different. But we found that human-- the really intriguing discovery here is that humans and CLIPasso sparsify their drawings differently. So what I'm plotting on the y-axis on the right hand side here is the divergence between the label distributions assigned by people to sketches produced by other people and sketches produced by CLIPasso under those different budgets.
我们发现,虽然在四种绘制预算下,人类的画和 CLIPasso 的画可辨认度相近——顺便说一句,可辨认度上的这种重合完全是巧合,毕竟单位完全不同。但我们发现,人类——这里真正耐人寻味的发现是,人类和 CLIPasso 在稀疏化自己的画作时方式并不相同。我在右边这张图的 y 轴上画的,是人们分配给他人所作素描的标签分布,与分配给CLIPasso 在不同预算下所作素描的标签分布之间的散度。
便签笔记
36:15
So in other words, if you take a-- if you take seriously a functional view of sketches as being for communicating concepts and characterize their meaning in terms of the full distribution of the labels and meanings that it evokes, then even though the 32 stroke CLIPasso drawings and the 32 second human sketches look different when you look at them-- stylistically they are different-- they are also quite functionally similar in terms of the set of meanings that they convey in the distribution of meanings they evoke.
换句话说,如果你认真对待素描的功能性视角——即素描是用来传达概念的——并且用它所唤起的标签和含义的完整分布来刻画其意义,那么即便 32 笔的 CLIPasso画作和 32 秒的人类素描在你看上去时并不一样——它们在风格上确实不同——它们在所传达的意义集合、所唤起的意义分布方面,功能上却相当相似。
便签笔记
08数据可视化:现代最有力的认知技术
36:48
On the other hand, as you tighten the production budget, that's where you really start to see much larger divergences between people and CLIPasso. And we're hoping that this SEVA benchmark will be a useful resource for others who are interested in developing models of human-like visual abstraction. Now in the second and I think shorter part, I'll give you a sneak preview of our newest and mostly unpublished work on multi-modal abstractions and how they are used to support statistical reasoning. So back to that Cartesian coordinate plane.
另一方面,当你把绘制预算收得更紧时,你才真正开始看到人和 CLIPasso 之间大得多的分歧。我们希望这个 SEVA 基准能成为一个有用的资源,供那些有兴趣开发类人视觉抽象模型的研究者使用。现在进入第二部分,我想它会短一些;我会给大家提前剧透一下我们最新的、基本尚未发表的关于多模态抽象以及它们如何支持统计推理的工作。回到那个笛卡尔坐标平面。
便签笔记
37:36
We're a room full of scientists. You don't need me to tell when you're making observations about the world it's never this clean. Instead of perfect lines, we might actually collect something like this, a collection of data points. They land where they land and from which we try to infer some underlying structure with the actual generator-- the actual generative process that gave rise to them. That inferential move is a fundamental building block of scientific reasoning, and we sure don't do it just by memorizing everything we've ever seen and thinking real hard.
在座都是科学家。不用我说你们也知道,当你在对世界做观测时,数据从来不会这么干净。得到的不是完美的直线,我们实际收集到的可能是这样的东西:一堆数据点。它们落在哪儿就是哪儿,我们要从中推断出某种潜在结构,推断出真正产生它们的那个生成过程。这一推断动作是科学推理的一块基本构件,而我们当然不是靠把见过的一切都背下来、再拼命苦想来完成它的。
便签笔记
38:11
We use technologies. At the beginning of my talk, I showed you these examples of the use of drawings produced by hand as a particularly enduring and versatile and accessible tool for making the invisible visible. I think that's remarkable and worth understanding. But perhaps one of the most impactful technologies to have been developed in the modern era was the invention of data visualization. Like the telescope and microscope, plots help to resolve parts of the world that you can't see directly.
我们使用技术工具。在演讲开头,我给大家展示了这些手绘图的例子,它是一种特别经久不衰、用途广泛且门槛很低的工具,能让不可见之物变得可见。我觉得这很了不起,值得去理解。但现代最具影响力的技术之一,也许要算数据可视化的发明。就像望远镜和显微镜一样,图表帮助我们看清世界中那些无法直接看见的部分。
便签笔记
38:46
But unlike either of those optical technologies, it allows you to see patterns and phenomena that might be too large, too noisy, too slow to see with our own eyes. They're ubiquitous in the news, the cornerstone of evidence-based decision making and business and government. And they're indispensable in every field of science and engineering. What I'm showing you here is one of the first time series plots drawn ever by William Playfair in 1786 to show the balance of imports and exports from England over an 80-year period from 1700-1780.
但与这两种光学技术都不同的是,它让你能看到那些可能过于宏大、过于嘈杂、过于缓慢而无法用肉眼看见的模式和现象。它们在新闻里无处不在,是商业和政府中循证决策的基石。而且在每一个科学和工程领域里都不可或缺。我给大家展示的这幅,是有史以来最早的时间序列图之一,由 William Playfair 在 1786 年绘制,用来展示1700 到 1780 这 80 年间英格兰进出口的平衡状况。
便签笔记
39:19
For a while, imports exceeded exports, but then the relationship flipped in the 1750s as exports really took off. Here's the thing. Unlike a drawing of one of Darwin's finches, if you haven't seen one of these kinds of images before, it may not be obvious what you're looking at. But once you learn how, it's a kind of superpower. So many individual observations can be distilled into a single graphic that tells a story that you can read just by looking. And that's not even all the reasons to care about plots because they're such a powerful tool for helping people update and calibrate their beliefs about a complicated world, developing the skills to read and interpret and even make graphs has long been a goal of STEM education in this country, something that's becoming even more important over time.
有一段时间进口超过出口,但到 1750 年代关系反转了,出口开始大幅腾飞。问题在于。跟一幅达尔文雀的素描不同,如果你以前没见过这类图像,你可能看不出自己在看什么。但一旦你学会了,它就是一种超能力。如此多的单个观测可以被浓缩进一张图里,讲出一个你只要看一眼就能读懂的故事。而这还不是我们该重视图表的全部理由,因为它们还是帮助人们更新和校准的强大工具他们对这个复杂世界的看法,培养读图、解读图表甚至制作图表的能力,长期以来一直是本国 STEM 教育的目标之一而且随着时间推移,这件事正变得越来越重要。
便签笔记
40:15
The New York Times broke a story-- this is actually from about a year ago now-- about recovery from COVID-related learning loss in mathematics based on work led by some of our education colleagues at Stanford and Harvard, and it looks like across a bunch of different states, the orange arrows are there, which means that there's been some recovery. But there's still a long way to go. And I think that successful theories that explain how people use these kinds of images, discover, and communicate important quantitative insights will help us equip people with the kind of quantitative data literacy skills they need more generally.
《纽约时报》曾报道过一则新闻——其实是大约一年前的事了——讲的是新冠疫情导致的学习损失在数学方面的恢复情况,这项研究由我们斯坦福和哈佛的一些教育界同事牵头,看起来在很多不同的州,都出现了橙色箭头,也就是说确实有了一些恢复。但要走的路还很长。我认为,能够成功解释人们如何使用这类图像、如何发现并传达重要定量洞见的理论,将有助于我们让人们具备他们更普遍需要的那种定量数据素养技能。
便签笔记
09六项图表测试:模型错得不像人
40:54
So I'm going to highlight briefly three directions that we're pursuing in this vein. Our first question asks about the underlying operations that are needed to understand plots. So the strategy that we've been taking is to obtain machine learning systems that can handle questions about data visualizations at all, assess alignment with people, and then interrogate the source of any of those gaps-- any gaps there might be. What we need, of course, to get off the ground is some way of measuring understanding.
所以我要简要介绍一下我们在这个方向上正在推进的三条研究路线。我们的第一个问题是:理解图表所需要的底层操作到底是什么。我们采取的策略是:先找到那些完全有能力处理数据可视化问题的机器学习系统,评估它们与人类的一致性,然后再深挖这些差距的来源——如果存在差距的话。当然,要起步我们首先需要某种衡量“理解”的方法。
便签笔记
41:28
So here's what that might look like. Say we're looking together at this stacked bar plot and someone asks you about the cost of peanuts in Las Vegas. Take a moment to take it all in. Scan around. Suppose that person then gave you four options to choose from. I added the little charcoal thing. Now if you thought it was A, you would be right. But that's not what some prominent visual language models say given this very same question. So in a Herculean benchmarking effort led by Arnav Verma, a current RA in the lab who's actually headed to the ECS program right here in the fall, we've conducted careful comparisons between humans and AI systems on six commonly used tests of graph-based reasoning sourced from across the education, health, visualization, psychology, machine learning communities.
那大概会是这个样子。假设我们一起看着这张堆叠条形图,有人问你在拉斯维加斯花生的价格是多少。花点时间把整张图看进去。四处扫一扫。假设这个人接着给了你四个选项让你选。那个小炭灰色的标记是我加上去的。如果你认为答案是 A,那你就答对了。但面对同样一个问题,一些知名的视觉语言模型给出的答案却不是这个。在实验室现任研究助理 Arnav Verma 主导的一项艰巨的基准测试工作中——他今年秋天其实就要进入这里的 ECS 项目了——我们在六个常用的图表推理测试上,对人类和 AI 系统做了细致的对比,这些测试来自教育、健康、可视化、心理学、机器学习等各个领域。
便签笔记
42:35
All six of these tests were administered in as parallel a manner as possible to both human participants and several of these so-called multi-modal AI systems and that had been-- the ones that made it into this benchmark had been claimed displaying competence on other kinds of visually grounded reasoning tasks. We then recorded not only the overall score achieved by humans and these models but the full set of error patterns they produced, which allowed us to assess even when a model or a person got a question wrong to see if they're getting things wrong in similar ways.
这六个测试都以尽可能平行的方式施测于人类被试和几个所谓的多模态 AI 系统,而那些能进入这个基准的模型,此前都被宣称在其他各类基于视觉的推理任务上表现出了相当的能力。然后我们不仅记录了人类和这些模型的总体得分,还记录了他们产生的全部错误模式,这样一来,即便某个模型或某个人答错了题,我们也能看看他们是不是以相似的方式答错的。
便签笔记
43:18
What did we find? Here I'm going to show you for each of the six tests we included, they have funny names like GGR, VLAT, CALVI, HOLF, HOLF-multi-- actually we made those-- the bad name there is my fault-- and also Chart-QA, a subset from Chart QA. We recorded how well-- we computed how well each of the models-- so these are the ones that are showing up on the xticks in blue, orange, purple, and red-- how well they did compared to how well humans did. That will show up in green. And these were US adults who had taken at least one high school math class.
我们发现了什么?接下来我会针对我们纳入的六项测试逐一展示,它们的名字都挺有意思,比如 GGR、VLAT、CALVI、HOLF,HOLF-multi——其实这些是我们自己起的——那个不太好的名字得怪我——还有 Chart-QA,取自 Chart QA 的一个子集。我们记录了——我们计算了每个模型的表现如何——也就是横轴刻度上显示的这些,蓝色、橙色、紫色和红色的——把它们的表现和人类的表现做对比。人类的成绩会用绿色显示。这些人是美国成年人,至少上过一门高中数学课。
便签笔记
43:59
First, I want to give you a sense of how well these people did. That's our reference point here. This is how well all the models did. So we have in this study two variants of Blip2-FlanT5, three variants of lava-based models, specialized systems such as matcha and its base model picks destruct, and finally a closed proprietary model, GPT 4 V. Across all these assessments, we did see a meaningful gap between models and humans both under more lenient and a strict-- a more strict grading protocol. This is a gap that we might have missed had we relied exclusively on the chart understanding benchmarks that are currently most popular in the machine learning literature, so this is chart-- this is the reason why we included Chart-QA, which is the rightmost facet on this slide where the gap seems a lot smaller.
首先,我想让大家先了解一下这些人做得怎么样。这就是我们这里的参照点。这是所有模型的表现。这项研究里我们有两个 Blip2-FlanT5 的变体,三个基于 LLaVA 的模型变体,还有像 MatCha 这样的专用系统及其基础模型 Pix2Struct,最后是一个闭源的专有模型 GPT-4V。在所有这些评估中,我们确实看到模型与人类之间存在明显差距,无论是在较宽松的评分标准下,还是在更严格的——更严格的评分协议下。如果我们只依赖目前现有的图表理解基准测试,这个差距可能就被我们忽略了在机器学习文献里最流行,所以这是图表——这就是我们纳入 Chart-QA 的原因,它是这张幻灯片最右边的那一栏,那里的差距看起来小得多。
便签笔记
44:53
CALVI, which is the third one, is really interesting because these are adversarially designed plots that have funny y-axis limits that require you to really attend closely. So that's one where we also see somewhat larger gaps. We also analyze their full error patterns so-- which can be really telling when no one is quite at ceiling. Again, there's only one way to get all the questions right, but there are lots of ways that you can be wrong and wrong reliably. So we found that even though GPT 4 V might look to be approaching human level performance, none of these models, GPT 4 V included, generated human-like error patterns.
第三个是 CALVI,它非常有意思,因为这些是对抗性设计的图表,y 轴的范围设得很古怪,需要你非常仔细地去留意。所以在这一项上,我们同样看到了相对更大的差距。我们还分析了它们完整的错误模式——当没有谁的表现接近上限时,这一点会非常说明问题。再说一次,把所有题目都做对只有一种方式,但出错的方式有很多,而且可以错得很有规律。所以我们发现,尽管 GPT-4V 看起来可能已经接近人类水平的表现,但这些模型,包括 GPT-4V 在内,都没有产生类人的错误模式。
便签笔记
45:33
So this is shown by all the dots here falling well below the green shaded area which represents the human noise ceiling. So the upshot is that while currently developing, VLMs remain exciting and promising testbeds for developing and parameterizing the hypothesis space of possible cognitive models of visualization understanding, there still are these systematic behavioral gaps that are worth interrogating further to realize those models' full potential. In parallel, we've also been developing experimental paradigms to probe a related facet of visualization understanding, the ability to select-- design select-- the appropriate plot to address your epistemic goal.
这一点体现在这里所有的点都远低于绿色阴影区域,而那个区域代表的是人类的噪声上限。所以结论是,虽然目前在发展中,VLM 仍然是令人兴奋且有前景的试验平台,可以用来开发并参数化可视化理解的可能认知模型的假设空间,但仍然存在这些系统性的行为差距,值得进一步深究,才能充分发挥这些模型的潜力。与此同时,我们也一直在开发实验范式,来探究可视化理解的一个相关侧面,也就是选择的能力——设计——选择合适的图表来达成你的认知目标。
便签笔记
10选图实验:非专家也懂受众需要
46:22
So I'm going to talk to you about that next. The way we set up the problem is to imagine that there's some question that you have about a data set. Some what I've been calling epistemic goal that a person is trying to satisfy like, for example, like which group is better. Let's say the left agent is trying to pick the plot to help shift the person's beliefs appropriately, but if they had a different question in mind, maybe they might need a different plot. That's the intuition. So this is a line of work that was launched by Holly Huey, and we formulated in a study hundreds of different questions that could in principle be answered by using real, publicly available data sets, in this case, your data sets that ship with base R.
接下来我就要跟大家聊聊这个。我们设定这个问题的方式是,设想你对某个数据集有某个疑问。也就是我一直在说的认知目标,某个人想要满足的目标,比如说,比如哪一组更好。假设左边这个人想挑一张图,来帮助恰当地改变对方的看法,但如果他们心里想的是另一个问题,那可能就需要另一张图。这就是那个直觉。这条研究路线是由 Holly Huey 发起的,我们在一项研究中设计了数百个不同的问题,这些问题原则上都可以用真实的、公开可得的数据集来回答,在这个例子里,就是 base R 自带的那些数据集。
便签笔记
47:10
So here's an example of a scary one, tracking airplanes and bird strikes. What is the average speed of aircrafts flying in overcast skies that encounter bird strikes at 0 to 5K miles? We then presented participants with a menu. AUDIENCE: Judy, can I ask a question? JUDY FAN: Sure. AUDIENCE: What is 5K miles? JUDY FAN: Altitude. Yeah. AUDIENCE: You sure it's not 5,000 feet? JUDY FAN: Sorry, five-- hold on. Hold on.
这里有一个挺吓人的例子,追踪飞机与鸟击事件。在阴天飞行、在 0 到 5K 英里处遭遇鸟击的飞机,平均速度是多少?然后我们给参与者呈现了一个菜单。观众:Judy,我能问个问题吗?JUDY FAN:当然。观众:5K 英里是什么意思?JUDY FAN:高度。对。观众:你确定不是 5,000 英尺吗?JUDY FAN:抱歉,五——等一下。等等。
便签笔记
47:37
That might be a-- that might be a typo on my part. I don't think the question inherited that. AUDIENCE: [INAUDIBLE] Anyways, don't worry about it. JUDY FAN: I'm not-- I'm-- I'm-- I'm starting to, but I'm going to resist that urge right now. No, no I think that-- yeah. So-- so-- AUDIENCE: But these-- a lot if this is the standard that they're full of all sorts of typos. JUDY FAN: This is a templated-- there's a templated thing. I think there are-- basically a lot of these questions are funny, and I think they could have been smoothed over a bit more.
那可能是——那可能是我这边的一个笔误。我觉得那个问题本身应该没有沿用这个错误。观众:[听不清]总之,别在意了。JUDY FAN:我没有——我——我——我开始在意了,但我现在要克制这种冲动。不不,我觉得——是的。所以——所以——观众:但这些——如果这就是标准的话,那里面到处都是各种各样的笔误。JUDY FAN:这是套模板生成的——有一个模板化的东西。我觉得——基本上这些问题里有很多都挺搞笑的,我想它们本可以再打磨得更顺一些。
便签笔记
48:05
So this is like-- the-- yeah. Yeah, yeah, yeah. No, no, it's good. It's good. So we had all these questions. Some of them were better put than others, but we presented participants with a menu of possible graphs that they might show someone else in order to help them answer that question. However, they interpreted that question. Whatever it is, those are the strings that we're showing to people. And they could choose from this menu either a bar plot, line plot, or scatter plot. Some of them were more distilled.
所以这就像是——那个——对,就是这样。对对对。不不,挺好的。很好。所以我们有了这一系列问题。有些问题问得比另一些更清楚,但我们给参与者提供了一个图表菜单,里面是他们可能会展示给别人看的各种图,用来帮助对方回答那个问题。来帮助他们回答那个问题。不管他们是怎么理解那个问题的。不管怎样,这些就是我们展示给大家看的文字表述。他们可以从这个菜单里选择柱状图、折线图或散点图。有些图更加提炼概括。
便签笔记
48:37
Some of them were more disaggregated. And then we measured how often they picked each one in order to construct a choice distribution over plots. So here's what it looked like on average for the retrieve value questions in our stimulus set. This is one kind of task, and there are other ones that'll ask you to compute some kind of difference between two values. This is the retrieve value subset of items. We then tested various hypothetical strategies that people might have used to pick the plots that they did.
有些图则更加分散、细化。然后我们统计他们每一种被选中的频率,从而构建出一个关于图表的选择分布。这就是我们刺激材料集中「取值类」问题的平均结果。这是其中一类任务,还有其他一些任务会要求你计算两个数值之间的某种差异。这是「读取数值」这一类任务的子集。然后我们测试了各种假设的策略,看看人们在挑选自己所选的那些图表时可能用的是什么方法。
便签笔记
49:09
Clearly they are showing some kind of bias here, the black curve. The purple curves here represent the predictions of bar plot, line plot, or scatter plot purists. They just always went for the bar plot but just didn't care whether they had three variables plotted or more. We also considered the possibility that people might prefer more targeted visualizations overall but not really care about what type of graph they were sharing. Maybe they might prefer ones that show more of the data, maybe ones that hide less of the variation in the data.
很明显,他们在这里表现出了某种偏好,也就是这条黑色曲线。这里的紫色曲线代表的是「柱状图派」「折线图派」或「散点图派」这些死忠者的预测结果。他们就是一律选柱状图,完全不在乎图里画的是三个变量还是更多。我们还考虑了另一种可能:人们整体上或许更偏好针对性更强的可视化,但并不太在意自己分享的是哪种类型的图。也许他们更喜欢展示更多数据的图,也许是那些更少掩盖数据变异性的图。
便签笔记
49:47
The best candidate we found that we considered was the proposal that people are maybe actually sensitive to the features of those plots that were relevant for answering the question and, in fact, actually predicted the performance of about 1,700 other participants who tried answering every one of those questions, however weirdly put, when paired with every possible plot. So we went and actually measured performance, and then we constructed the distribution based on those performance levels and then use that task in order to-- use the data from that task in order to generate predictions of the audience sensitive hypothesis.
我们找到的、在所考虑的方案中最好的候选解释是:人们可能实际上对这些特征是敏感的那些与回答问题相关的图表,事实上,还真的预测了大约1700名其他参与者的表现,他们尝试回答那些问题中的每一个,无论问题措辞多么奇怪,并且与每一种可能的图表配对。所以我们真的去测量了表现,然后根据这些表现水平构建了分布,然后用那个任务——用那个任务的数据来生成对受众的预测敏感的假设。
便签笔记
50:25
And this is where we-- this is how the shape of those curves look when we're only considering the true value items. But if you look at the data set as a whole, we found that this audience sensitive proposal did well across the board. So what I'm showing here is essentially a model fit measure defined over the divergence between the predicted and the actual human choice distributions. Has a name, the Jensen-Shannon divergence, the average of KL divergence but both ways. So I'm excited about that result as an initial validation of a strategy for measuring visualization understanding using more open ended tasks and also because it suggests that even non-experts you could say, people who might not be professional scientists, are sensitive to those features of plots that make some suitable for answering some questions than others and honestly as someone who teaches intro stats for a living gives me hope.
这就是我们——这就是当我们只考虑真实值项目时,那些曲线的形状。但如果你把整个数据集作为一个整体来看,我们发现这种对受众敏感的方案在各方面表现都不错。我这里展示的本质上是一个模型拟合指标,它定义在预测的选择分布与实际人类选择分布之间的差异上。它有个名字,叫 Jensen-Shannon 散度,也就是双向 KL 散度的平均值。所以我对这个结果很兴奋,它初步验证了一种用更开放式的任务来测量可视化理解能力的策略,同时也因为它表明,即便是非专业人士——你可以说,那些可能不是职业科学家的人——也能敏锐地察觉到图表中的那些特征,正是这些特征让某些图表更适合回答某些问题。说实话,作为一个靠教统计学导论为生的人,这给了我希望。
便签笔记
11重审测量:现有测试到底在测什么
51:35
In the final final leg of our journey, we're going to take another critical look at this problem of measurement. So a few minutes ago, I showed you some results when using these six tests of data visualization understanding. These are the ones that we have today, which is why we use them in our benchmark study. Our question in this last study in work that's actually now in press and led by this extraordinary postdoc in our lab, Erik Brockbank, also with Arnav Verma, we're asking what are these tests actually measuring, and are they measuring the skills in the best possible way?
在我们旅程的最后一段,我们要再次以批判的眼光审视这个测量问题。几分钟前,我给大家展示了使用这六项数据可视化理解测试得到的一些结果。这些是我们目前拥有的测试,所以我们在基准研究中使用了它们。在这项最后的研究中——这项工作现已被接收发表,由我们实验室这位非常出色的博士后 Erik Brockbank 主导,Arnav Verma 也参与其中——我们要问的是:这些测试实际上在测量什么?它们是以最佳的方式在测量这些技能吗?
便签笔记
52:17
And can we do better? So here are some initial steps towards answering those questions. To get some traction, we started with these two GGR and VLAT, these being some of the most widely used and established tests. We gave GGR and VLAT as a composite test to a large and diverse sample of US adults-- actually two samples, one that was recruited on the UCSD campus and another that was recruited over prolific under the constraint that it had to be a demographically representative sample. And what I'm showing on the left is that we get a pretty convergent estimates of the difficulty of individual items in both the college campus sample and the US representative sample, which I think is reassuring.
我们能做得更好吗?下面是回答这些问题的一些初步步骤。为了找到切入点,我们从 GGR 和 VLAT 这两项测试入手,它们是使用最广泛、最成熟的测试之一。我们把 GGR 和 VLAT 合成一个测试,发给了一个规模大且多样化的美国成年人样本——实际上是两个样本,一个是在加州大学圣地亚哥分校校园里招募的,另一个是通过 Prolific 招募的,条件是必须是一个在人口统计学上具有代表性的样本。左边我展示的是,在大学校园样本和美国代表性样本中,我们对单个题目难度的估计相当一致,我觉得这一点让人安心。
便签笔记
53:05
On the right, what I'm showing is that people who did well on one test often did well on the other, suggesting that maybe the two tests are measuring some of the same things or similar things. And the question is like what? What are those things? One possibility is that those two tests track how much easier some plots are to understand than others. If so, those plots should be reliably hard or easy across the board like maybe bar plots or easy but stacked area charts are more challenging. It seemed to us that clearly there's a lot more going on here.
右边我展示的是,在一项测试上表现好的人往往在另一项上也表现好,这说明这两项测试可能在测量一些相同或相似的东西。问题是,具体是什么呢?那些东西到底是什么?一种可能是,这两项测试反映的是某些图表比另一些图表更容易理解的程度。如果是这样,那这些图表在各方面应该稳定地表现为难或易,比如条形图可能容易,而堆叠面积图更有挑战性。但在我们看来,这里显然还有更多因素在起作用。
便签笔记
53:44
Performance wasn't consistent for a given kind-- it wasn't always consistent for a given kind of plot within or across tests, and there weren't even enough items to be able to establish a direct link between the type of graph and performance because actually in each of these tests, there's generally one instance of each class of graph. So that also made it challenging. We also dug into the patterns of mistakes that people made and found that the best way to predict those patterns on these two tests wasn't the kind of plot or even the type of question.
对于某一类图表,表现并不一致——无论是在同一项测试内部还是跨测试之间,都不总是一致,而且题目数量甚至都不够,无法在图表类型和表现之间建立直接联系,因为实际上在这些测试中的每一项里,每一类图表通常只有一个实例。所以这也带来了挑战。我们还深入分析了人们犯错的模式,发现在这两项测试上预测这些模式的最佳方式既不是图表的类型,也不是问题的类型。
便签笔记
54:21
Maybe find the max or identify clusters or characterize distribution or retrieve value. These are all common ways of describing the ontology of tasks involving plots, but it really seemed from this analysis that other underlying factors that aren't well described by the ontology are really accounting for error patterns much more efficiently. So here I'm showing you how much better a parsimonious four factor model does than one that uses the groupings that you might think of to-- use to organize the set of skills needed to understand any graph.
比如找最大值、识别聚类、描述分布特征,或者读取数值。这些都是描述涉及图表的任务分类体系的常见方式,但从这项分析来看,似乎是一些无法被这套分类体系很好描述的其他潜在因素,更有效地解释了错误模式。所以我在这里展示的是,一个简约的四因子模型比使用你可能想到的那些分组方式的模型好多少——那些分组方式就是用来组织理解任何图表所需的技能集合的。
便签笔记
55:01
So I'm not going to unpack everything on this slide, but this is just meant to illustrate that there seems to be something that is going on here that isn't obviously mappable to the ways that we talk about what the component skills are that you might actually see in textbooks or in instructional materials when it comes to how to break into data visualizations. And even if existing assessments might not be testing and characterizing visualization understanding in the best possible way, we're really trying to take that as a glass half full call to action to develop improved measures.
我不打算把这张幻灯片上的所有内容都展开讲,但这只是想说明,这里似乎有些东西在起作用,而它并不能明显地对应到我们谈论组成技能的那些方式——就是你在教科书或教学材料里能看到的、关于如何入门数据可视化的那些说法。即便现有的评估工具可能并没有以最佳方式测试和刻画可视化理解能力,我们还是真心想把这看作一个「杯子半满」式的行动号召,去开发更好的测量工具。
便签笔记
55:37
So stay tuned for those. More generally, the reason I think this work establishing the perceptual and cognitive foundations of data visualization is so important is because it will give us a chance to use what we learn to eventually help people, learners in real educational settings, calibrate their understanding of a complicated and changing world that we can ever only observe a part of. And these kinds of integrative efforts if you will that connect fundamental science to that wider world that we all inhabit exemplify where we're going with all of this.
所以敬请期待。更广泛地说,我认为这项确立数据可视化的知觉和认知基础的工作之所以如此重要,是因为它将让我们有机会运用所学,最终帮助人们——真实教育环境中的学习者——校准他们对一个复杂多变、而我们永远只能观察到其中一部分的世界的理解。可以说,这类把基础科学与我们共同生活的更广阔世界连接起来的整合性努力,正体现了我们所有这些工作的方向。
便签笔记
56:22
We really want to develop psychological theories that explain how people use the suite of cognitive technologies that we've inherited and continue to innovate on. We want to understand why that toolkit looks the way it does, what future cognitive tools might work even better. In the long run, I think that understanding how these tools work and how to make them better really matters because it's like these tools that are at the heart of two of our most impactful and generative activities. First, education, which is the institution and maybe more importantly the expectation that every generation of human learners should be able to stand on the shoulders of the last and see further and design the suite of activities and habits of mind that help people continually reimagine how the world could be better, and then go out and make it true.
我们真正想做的是发展心理学理论,来解释人们如何使用我们继承下来、并持续加以创新的这一整套认知技术。我们想理解这套工具为什么是现在这个样子,以及未来的认知工具怎样才可能更好用。从长远来看,我认为理解这些工具如何运作、如何让它们变得更好,真的很重要,因为正是这些工具处在我们两项最具影响力、最富创造力的活动的核心。首先是教育,它是一种制度,或许更重要的是一种期待——期待每一代人类学习者都能站在前人的肩膀上看得更远,并设计出那一整套帮助人们不断重新想象世界如何可以更美好的活动和思维习惯,然后走出去把它变成现实。
便签笔记
57:16
So with that, I want to thank all of the folks who've been involved in this work and other lines of work in the lab. It would not have been possible without an amazing research team and network of collaborators and colleagues in many different places including here. So I want to thank all of them, all of you, for your attention, and I'm happy to take questions if we have time, which we might not. Great. Thank you. [APPLAUSE]
那么最后,我想感谢所有参与这项工作以及实验室其他研究方向的同仁。如果没有一支出色的研究团队,以及分布在许多不同地方(包括这里)的合作者和同事组成的网络,这一切都不可能实现。所以我想感谢他们所有人,也感谢在座各位的聆听,如果还有时间的话,我很乐意回答问题——不过可能没时间了。很好。谢谢。[掌声]
便签笔记
12问答:错误图画、误读诊断与儿童读图
57:44
I see hand, hand. Hand, hand, hand AUDIENCE: Thank you so much for the talk. One thing I'm curious about is so in a lot of your examples specifically relating to depictions, there seems to be this one-dimensional axis between things that are more detailed and faithful or versus more sparse. But I'm curious about cases where people watching diverge and produce things that are actually false or counterfactual. So I'm thinking if I'm drawing a glass of water and I want to indicate that it's full, I'll shade it in blue even though in real life, glasses aren't blue when they have water.
我看到有人举手,还有一位。还有、还有、还有 观众:非常感谢您的演讲。我很好奇的一点是,在您很多与「描绘」相关的例子里,似乎存在这样一个一维的轴:一端是更细致、更忠实的表达,另一端是更简略的表达。但我很好奇那些人们的做法出现分歧、并且画出实际上是错误或与事实不符的东西的情况。比如我在想,如果我画一杯水,我想表示它是满的,我会把它涂成蓝色,尽管现实生活中杯子装了水并不是蓝色的。
便签笔记
58:17
And I feel like that depends a lot on culture and language and stuff like that. I speak a language where water is described as blue. I have been around pools which are usually painted blue on the inside and so on. And that could depend. So I'm just wondering how you thought about some of those cases? JUDY FAN: Yeah. Yeah. There's a large literature in philosophy in the area of aesthetics that I take inspiration from that begins from a similar premise, which is how is it possible that we understand pictures that are false like pictures of fictional individuals or unicorns or glasses with blue water even if it doesn't look that way.
我觉得这在很大程度上取决于文化、语言之类的因素。我说的那门语言里,水就是被描述成蓝色的。我也常去那种内壁通常漆成蓝色的游泳池,等等。所以这是有可能受影响的。所以我想问的是,您是怎么看待这类情况的?朱迪·范:是的。没错。哲学中的美学领域有大量文献给了我启发,它们的出发点也很类似,那就是:我们究竟怎么可能理解那些「假」的图画——比如虚构人物、独角兽,或者装着蓝色水的杯子的图画,哪怕现实并不是那个样子。
便签笔记
58:55
The tech that we've been taking in our own empirical work is to not begin from that premise, which is that rather than thinking of them as false, this is the actual data that people are generating under some objective. And we just-- we're trying to figure out what that is. And so the reasons why they might render the water depicted in that depicted glass is for some reason-- and we're interested in those stakes that you identified-- to what degree it captures the-- faithfully the visual appearance, the phenomenology of looking at a glass of water in a room where the glass is in front of you as opposed to a kind of acquired convention for how you depict water, for example.
我们在自己的实证研究中采取的做法是不从这个前提出发,也就是说,与其把它们当作「错误的」,不如把它们看成人们在某个目标之下真实生成的数据。而我们要做的,就是设法弄清楚那个目标是什么。所以他们之所以会那样画出杯中的水,是有某种原因的——我们感兴趣的正是你提到的那些权衡:它在多大程度上忠实地捕捉了视觉外观,也就是你站在房间里看着面前那杯水时的现象学体验;还是说它体现的是一种习得的、关于如何描绘水的约定俗成。
便签笔记
59:40
I think those are very reasonable sources of constraints on why people make those representational decisions. So I think those are the kind of questions that just rather than thinking of some drawings as good or bad, false or true, we just find it a lot more useful and productive to think of these are the drawings people make under those conditions. Why? Why do they look that way, not another? If that helps. Thanks for your question. AUDIENCE: Sure. Thank you very much. I am curious about-- so I thought something you only very briefly touched on is super potentially interesting is the human likeness of errors that artificial systems make in graphics and visualization.
我认为这些都是非常合理的约束来源,能解释人们为什么做出那样的表征决策。所以我觉得,与其把某些画归为好或坏、真或假,我们发现更有用、更有成效的思路是:这就是人们在那些条件下会画出的东西。为什么?为什么它们看起来是这样,而不是别的样子?希望这回答了你的问题。谢谢你的提问。观众:好的。非常感谢。我很好奇——我觉得您只是非常简略地提到的一点其实潜在地非常有意思,就是人工系统在图形和可视化方面犯的错误与人类错误的相似性。
便签笔记
60:24
And I guess in a way actually that seems especially useful for an artificial system because if I am producing a visualization, I know what people will understand it to be if they understand it correctly, but I don't know-- I might have a lot of difficulty imagining how it would be misunderstood. And so that actually seems like a task which was literally how else might this be misunderstood. I maybe understood would be it's very, very valuable thing, and I want to ask what's-- what is our state-of-the-art model of that and how graphs are misunderstood.
我想在某种意义上,这对人工系统来说其实特别有用,因为如果我要做一个可视化,我知道人们在正确理解它的时候会理解成什么,但我不知道——我可能很难想象它会怎样被误解。所以这实际上像是一个任务,字面意义上就是:这还可能被怎样误解?我想这会是非常非常有价值的事情,我想问的是,目前关于这个问题、关于图表如何被误解,最先进的模型是什么?
便签笔记
61:05
Usually they're misunderstood. And is it an interpretable state-of-the-art problem from a scientific-- JUDY FAN: Oh, gosh. AUDIENCE: Give us insights into the workings of the actual inner workings.
通常它们都是被误解的。以及从科学角度看,这是不是一个可解释的前沿问题—— 朱迪·范:哎呀。观众:能不能让我们了解其内部的实际运作机制。
便签笔记
61:20
JUDY FAN: It's really, really good question. So I'm going to engage with it, which is so there's a phenomenon which some of us in the room study which is perspective taking, which is hard but doable in some contexts. It can be possible to imagine what the world must look like from the viewpoint of someone else and their various paradigms for studying that. It also seems like the capacity for doing that very quickly and accurately can change with practice. And so there's a role for expertise and experience to play in shaping your ability to do that.
朱迪·范:这真是个非常非常好的问题。我来回应一下:在座有些人研究的一个现象叫「视角采择」,它很难,但在某些情境下是可以做到的。我们有可能想象出世界从别人的视角看会是什么样子,也有各种研究这一点的范式。而且看起来,快速而准确地做到这一点的能力,是可以随着练习而改变的。所以专业素养和经验会在塑造你这种能力的过程中发挥作用。
便签笔记
61:57
One way in which that manifests-- I'm not going to refer to the VLM results at all here, but just to sketch the shape of the problem, it's teachers, the kind of expertise-- one of many kinds of expertise that teachers-- really skilled teachers develop is the ability to diagnose-- notice it is called misconceptions-- when they listen to learners describe how they're thinking about a problem. So it may be that the final response or answer that a student gives to a math problem is wrong, but it's not only that it's wrong.
它的一种体现方式是——这里我完全不打算提视觉语言模型的结果,只是勾勒一下这个问题的形态——就是教师。教师所具备的众多专业能力中的一种,真正优秀的教师会发展出的能力,就是在听学生描述他们如何思考一个问题时,诊断出——通常被称为「错误概念」的东西。所以,学生对一道数学题给出的最终回答或答案可能是错的,但问题不只是它错了。
便签笔记
62:30
The relevant question isn't always that it's wrong or right but rather what is the nature of the misconception or misperception or what is the gap between the normatively correct way of representing the problem and proceeding through the different steps and whatever the student did. So one way in which I think-- maybe I will just very, very briefly, if it's OK, speak to how we've been thinking about this question in the modern era of extremely large machine learning systems is to interrogate the kind of operations and conceptual primitives so to speak that these systems rely on when they get an answer right or wrong, using the suite of tools that go under the banner of mechanistic interpretability.
真正相关的问题并不总是它对还是错,而是这个误解或错误认知的本质是什么,或者说,从规范上正确地表征这个问题并一步步推进的方式,和学生实际做的事情之间,差距在哪里。所以我觉得有一种方式——也许我可以非常非常简短地说一下,如果可以的话,谈谈在当今这个超大规模机器学习系统的时代,我们是如何思考这个问题的:就是去审视这些系统在给出正确或错误答案时所依赖的那些操作,或者说概念上的基本单元,用的是统称为机制可解释性(mechanistic interpretability)的那一整套工具。
便签笔记
63:17
But you could-- we could call it cognitive systems neuroscience for artificial neural networks in order to diagnose where right answers and wrong answers come from, and it'd be really, really nice. I'm not going to go into those in detail because of the-- But that's a kind of like strategy you might take to connect what is the operations that are being performed in these systems that do not immediately lend themselves to those kind of interpretations at the same time and then connect those to what has been called knowledge graphs-- the underlying knowledge graph that a person might hold or not hold.
但你可以——我们可以把它叫做针对人工神经网络的认知系统神经科学,用来诊断正确答案和错误答案是从哪里来的,那会非常非常好。因为时间关系我就不详细展开了——但这大概就是你可能采取的一种策略,去把这些系统里正在执行的操作联系起来,而这些操作本身并不能直接对应到那类解释上,然后再把它们跟所谓的知识图谱联系起来——也就是一个人可能持有、也可能不持有的那个底层知识图谱。(承上)
便签笔记
63:55
So that's like a-- like a way of setting up the problem I think, and sometimes the issue, the bottleneck might be perceptual as such. And other times, it might be like a bottleneck or the gap might have to do with a reasoning step, and the goal is to really expose what those steps are, have tools for diagnosing them so that in principle you could use them to diagnose those misconceptions or misperceptions, mistakes, missteps in any body or any system. That's very, very, very hard, of course. But these challenging problems are I think like a tool that we can use in-- towards that end.
所以我觉得这算是一种设定问题的方式,而有时候问题、那个瓶颈,可能本身就是感知层面的。还有些时候,瓶颈或差距可能出在某个推理步骤上,而目标就是真正把这些步骤揭示出来,有工具去诊断它们,这样原则上你就能用它们来诊断那些错误概念或错误认知、错误、失误——无论是在什么主体或什么系统里。当然,这非常非常非常难。但我认为这些有挑战性的问题就像是我们可以用来朝那个方向努力的一种工具。
便签笔记
64:40
I hope some of that made sense. I-- a lot of thoughts about. AUDIENCE: Can I ask a follow up? JUDY FAN: Yeah. Yeah. AUDIENCE: So I think it makes it-- so in particular, you helped us out here by highlighting-- JUDY FAN: I know. AUDIENCE: That red, the right part of the graph to pay attention to. JUDY FAN: I did do that. AUDIENCE: And one of the reasons why you might have trouble answering this question-- I'm just trying to see if I understand-- JUDY FAN: Yeah. AUDIENCE: Is that you might not know which part of the thing to pay attention to to answer the question.
希望这些多少说清楚了一些。我——关于这个我有很多想法。观众:我能追问一个问题吗?JUDY FAN:可以。嗯。观众:所以我觉得这让它——特别是,你在这里帮了我们一把,因为你突出标记了——JUDY FAN:我知道。观众:那个红色的部分,图上值得关注的那一块。JUDY FAN:我确实那么做了。观众:而你可能很难回答这个问题的原因之一——我只是想确认我是不是理解对了——JUDY FAN:嗯。观众:就是你可能不知道该关注这东西的哪一部分才能回答那个问题。
便签笔记
65:06
JUDY FAN: Yeah. AUDIENCE: Or but-- or you might know what part-- but you might not know what to do with it how to take that red rectangle and relate it to the information on the y-axis. Or you might not know even if you know how to do that how to turn that into the ends of the question. So those are the kind of that's that-- JUDY FAN: Well, yeah. Those are the kind of teaching moves that a human educator might actually take. AUDIENCE: You could. JUDY FAN: Yeah. AUDIENCE: People have even tried to break this down like some of Erik Schultz's group, for example, where they say you could ask the model or people how big is the rectangle, red rectangle on the second column.
JUDY FAN:对。观众:或者说——要么你知道该看哪一部分——但你不知道拿它怎么办,怎么把那个红色矩形和 y 轴上的信息对应起来。到 y 轴上的信息。或者就算你知道怎么做,你也可能不知道怎么把它转化成问题的答案。所以这些就是那种——就是那种——JUDY FAN:嗯,是的。这些正是人类教育者可能会采取的教学动作。观众:你可以这么做。JUDY FAN:对。观众:已经有人试着把这个拆解开了,比如 Erik Schultz 团队的一些工作,他们会说,你可以问模型或者问人:第二列上那个红色矩形有多大。
便签笔记
65:45
JUDY FAN: Yeah. AUDIENCE: It doesn't require any reasoning but just-- JUDY FAN: Yeah. Yeah. Yeah. AUDIENCE: Or something or which blue rectangle is the highest. JUDY FAN: Yeah. AUDIENCE: For example, is that the kind of thing you're doing? JUDY FAN: Yeah. Yeah. Yeah. So we started with stimuli like these. So this is a very, very new work that's being led by-- yes, exactly. So this is being led by Alexa Tartaglini in collaboration with Chris Potts. And the-- we started with stimuli like these, like the real wild plots and questions, and then realized we wanted to drill down.
JUDY FAN:对。观众:这不需要任何推理,只是——JUDY FAN:对。嗯。是的。观众:或者类似的,比如哪个蓝色矩形最高。JUDY FAN:对。观众:比如说,你做的是这类事情吗?JUDY FAN:是的。嗯。对。我们一开始用的就是这样的刺激材料。这是一项非常非常新的工作,主导者是——对,没错。这项工作由 Alexa Tartaglini 主导,和 Chris Potts 合作。我们一开始用的是这样的刺激材料,就是真实的、原生的图表和问题,然后意识到我们需要往下深挖。
便签笔记
66:13
So there's both taking these plots and asking a variety of different questions including these more basic ones so to speak. Also we've realized that really stripping these down to much simpler versions of it that still contain the core multi-modal integration challenge. Even we're currently running these studies on number line stimuli essentially to understand how you might judge the distance between two data points on the number line. For example, but they're like all these different routes to it because this is like-- this-- the point is that this is complicated, and so decomposing it is part of the challenge.
所以一方面是拿这些图表来问各种不同的问题,包括刚才说的那些更基础的问题。我们也意识到,要把它们精简到简单得多的版本,但仍然保留核心的多模态整合难点。挑战。我们现在甚至在用数轴类的刺激材料做这些研究,本质上是想理解你会怎么判断两个数据点之间的距离。在数轴上。比如说,但其实有各种各样不同的路径可以做到这一点,因为这个——这个——关键在于这件事本身很复杂,所以把它拆解开来本身就是挑战的一部分。
便签笔记
66:57
AUDIENCE: [INAUDIBLE] --if there could be like-- JUDY FAN: Yeah. Yeah. AUDIENCE: [INAUDIBLE] how you teach these skills. JUDY FAN: Yeah. AUDIENCE: Break them down into-- JUDY FAN: Exactly. Yeah. That's right. That's right. Yeah, right. So there's assessment and then also the reason why this black shaded charcoal rectangle is there is because there's some guess that I had about what might be a bottleneck to rapidly looking at the appropriate feature of the plot. AUDIENCE: Could I follow up? JUDY FAN: Yeah, sure.
观众:[听不清] ——如果可以像—— JUDY FAN:对,对。观众:[听不清] 你是怎么教这些技能的。JUDY FAN:对。观众:把它们拆解成—— JUDY FAN:没错。对。是这样的。是这样的。对,没错。所以既有评估的部分,另外之所以这里有一个黑色炭灰色的矩形,是因为我当时有一些猜想,猜测究竟是什么造成了瓶颈,让人没法迅速看到图中合适的那个特征。观众:我可以追问一下吗?JUDY FAN:可以,当然。
便签笔记
67:26
OK, this is-- OK. AUDIENCE: When I first looked at this plot, I thought it was actually ambiguous, and I wasn't sure whether the height of the bars all started at 0 and we're seeing one of them in front of another. JUDY FAN: I know. I know. PRESENTER: With the cheapest ones-- JUDY FAN: It's like dodge equals true. Yeah. AUDIENCE: But that's not my question. My question is this you've been focusing on for adults who've been exposed to lots of different kinds of media. What are what are the best ways to present information so that people can get it and get it efficiently so forth?
好的,这个——好的。观众:我第一次看这张图的时候,其实觉得它有歧义,我不确定这些柱子是不是都从 0 开始,然后我们看到的是其中一根挡在另一根前面。JUDY FAN:我懂。我懂。主持人:用最便宜的那些—— JUDY FAN:就像 dodge equals true 一样。对。观众:但这不是我的问题。我的问题是,你一直关注的是那些接触过大量不同媒介的成年人。呈现信息的最佳方式是什么,才能让人看懂、并且高效地看懂等等?
便签笔记
67:58
But there's another question, which is I think for when we first see a new kind of graph this might actually be the first time I ever saw one that stacked up, at least it's been a while since-- JUDY FAN: That's exciting. AUDIENCE: [INAUDIBLE] --stacked up like this. JUDY FAN: They look like-- AUDIENCE: It's not so obvious when you first see it. JUDY FAN: Yeah, it's not. AUDIENCE: As we get very used to seeing them, seeing things that are different can become really harder. So, for example, event related potentials where negative goes up drives me crazy-- JUDY FAN: Right!
但还有另一个问题,就是我觉得当我们第一次看到一种新类型的图表时,这可能真的是我第一次见到这样堆叠起来的图,至少已经很久没有—— JUDY FAN:这挺让人兴奋的。观众:[听不清] ——像这样堆叠起来。JUDY FAN:它们看起来像—— 观众:第一次看到的时候并不是那么显而易见。JUDY FAN:对,确实不明显。观众:等我们非常习惯看某一种图之后,再看到不一样的图反而会变得更难。比如说,事件相关电位图里负值朝上,这让我抓狂—— JUDY FAN:对啊!
便签笔记
68:23
AUDIENCE: Every time I see it. JUDY FAN: Yeah. Right. AUDIENCE: So I wonder what do you think about the question of how these kinds of representations should be structured so that they're easy for kids to learn. JUDY FAN: Yeah. Oh. AUDIENCE: You start out with none of these representations. What would you put into the earliest ones that you want kids to learn about? How would you progress from there? JUDY FAN: Yeah. That's a good question. Yes. And actually there are versions-- I feel like there's a few minutes ago.
观众:每次看到都这样。JUDY FAN:对。没错。观众:所以我想知道,你怎么看这个问题:这类表征应该如何设计,才能让孩子容易学会。JUDY FAN:嗯。哦。观众:假设一开始这些表征你一个都没有。你会把哪些内容放进最早的那一批、想让孩子先学会的表征里?然后你会怎么一步步往后推进?JUDY FAN:嗯。这是个好问题。是的。其实是有一些版本的——我觉得就在几分钟前,
便签笔记
68:51
Nancy and I were talking about something kind of related when it comes to how people break into this class of visual inputs at all. There's a sense in which you're building on top of a lot of neurotypical ordinary visual and cognitive development in order to grasp the presence of different shapes in certain spatial arrangements. There's-- it relies on literacy as such being able to-- You have-- there are a bunch of primitive-- there are a bunch of conceptual primitives and more basic competencies that you might need to build up first.
Nancy 和我聊到过有点相关的话题,就是人们究竟是怎么最初进入这一类视觉输入的。从某种意义上说,你是建立在大量神经典型的、普通的视觉和认知发展之上,才能把握不同形状在特定空间排布中的存在。这——它依赖于识字能力本身,能够—— 你有—— 有一堆原始的——有一堆概念上的原语,还有一些更基础的能力,是你可能需要先建立起来的。
便签笔记
69:27
It might be that there is a sequence of experiences that build from the-- that's a really good question. I don't have answers to that curriculum design question, but I think it's a great question. And it feels like there are thoughts that I've-- yeah. AUDIENCE: It seems like the whole back third of the cortex does vision, and it's a spectacular set of machinery that extracts all these different rich kinds of visual representations from orientations to heights to shapes to landscapes, all of which are possible spaces we can use.
有可能存在这样一个经验序列,从——这真是个好问题。课程设计这个问题我没有答案,但我觉得这是个很棒的问题。而且感觉我确实有一些想法——嗯。观众:看起来大脑皮层后三分之一整个都在处理视觉,这是一套非常了不起的机制,它能提取出各种丰富的视觉表征,从朝向到高度到形状再到景观,所有这些都是我们可以利用的可能空间。
便签笔记
70:03
And so it seems like the essence of making a good graph is figuring out how to make a visualization that taps into some of that machinery to build a map. JUDY FAN: What makes visual scenes in general easier and more fluent to process? Things that aren't particularly cluttered. If you're if you're interested in the kind of problem of visual search over those scenes, there are aspects of those processes that are co-opted in order to identify the appropriate sub-region of this image to look at. You might imagine designing those initial-- those initial graphs to cohere more with those kinds of scenes.
所以看起来,做出一张好图的精髓,就在于想办法做出一种可视化,去调用其中一部分机制来构建映射。JUDY FAN:是什么让视觉场景整体上更容易、更流畅地被加工呢?那些不特别杂乱的东西。如果你感兴趣的是在这些场景中进行视觉搜索这类问题,这些过程中有一些方面会被调用,用来确定这张图像里该看哪个合适的子区域。你可以设想,把那些最初的——那些最初的图设计得更贴合这类场景。
便签笔记
70:39
And then there's the step of mapping the components of these scenes to concepts that you also need to learn about and that can be really hard initially and then become faster and faster and easier and easier over time. And I think it's really fascinating what is happening as that takes place. Yes. And also yes. But then also-- what should I do? PRESENTER: And one-- this-- why don't we-- this is really cool. Why don't we wrap up and allow people to who need to leave to go? JUDY FAN: Yeah. Yeah. PRESENTER: But other people, please linger here.
然后还有一步,是把这些场景的组成部分映射到一些概念上,而这些概念你同样需要去学习,一开始这可能非常困难,但随着时间推移会越来越快、越来越容易。我觉得在这个过程中发生的事情真的很迷人。是的。也确实如此。但接下来——我该做什么呢?主持人:还有一点——这个——我们不如——这真的很精彩。我们不如就到这里收尾,让需要离开的人先走?JUDY FAN:好的。好的。主持人:不过其他人,请留下来待一会儿。
便签笔记
71:16
You'll be here for a little while. We have a reception. We have 45 minutes or so for people to interrogate you. JUDY FAN: Amazing. Thank you. Thank you. Thank you. [APPLAUSE]
你会在这儿待一段时间。我们有一个招待会。大家有大概 45 分钟的时间可以来向你提问。JUDY FAN:太好了。谢谢。谢谢。谢谢。[掌声]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

斯坦福认知科学家 Judy Fan 以"认知工具"为核心,用手绘草图和数据可视化两条线索论证:人类通过"视觉抽象"把不可见的知识变为可见,这种能力受交流目标与语境驱动,而当前 AI 视觉模型在草图理解和图表推理上与人类仍有系统性差距。

核心要点

  • 认知工具是人类发明而非自然赋予的,且能改写整个思维方式。 数轴、笛卡尔坐标系把代数方程与几何曲线连接起来,解决了困扰千年的"倍立方"(Delian)问题,把解方程组转化为找曲线交点,四百年后已成为全球数学课程的标配。
  • "让不可见变可见"贯穿科学史,且抽象程度各异。 古尔德画的达尔文雀、伽利略望远镜、卡哈尔的视网膜绘图、费曼图——从写实到高度示意,共同点是用视觉抽象凸显"值得注意的相关信息"。Fan 主张认知科学需补上两块缺口:一是认知工具/技术的理论,二是"工程"(用理解去创造新事物)作为科学的补集。
  • 草图理解可部分由"相似性"解释,但不够。 在自然照片上训练的 ConvNet 能较好泛化到稀疏草图(后续工作还能把草图局部与照片局部做保拓扑映射),但静态确定性视觉模型无法解释白板上的方块箭头——含义取决于语境。
  • 人们会根据交流语境自动调节绘画抽象层级。 绘画游戏实验中,当干扰项属同一基本类别(close)时,画者用更多笔画、墨水和时间画出细节;干扰项来自不同类别(far)时则画得更简略,且识别准确率仍在天花板,观看者决策也更快。模型消融显示:视觉编码层的抽象能力和对语境的敏感度二者缺一不可。
  • "视觉解释"与"视觉描绘"是可分离的,而非累加关系。 Holly Huey 用 6 台有开关电路的自制机械做实验:解释图把更多笔画分给因果部件、更多墨水给箭头/运动线,背景更少;功能测试显示解释图更利于判断操作方式(拉/转/推),描绘图更利于识别物体身份,但解释图在"指认因果部件"任务上并无优势——因为抽象掉背景后反而丢失了图与实物的对应线索。
  • SEVA 基准显示模型与人类在草图理解上存在巨大鸿沟。 约 5,500 人在不同时间预算(最短 4 秒)下画了 128 个概念共 9 万幅草图,17 个视觉模型与人类做同一分类任务。时间越长草图越可识别、标签熵越低;但模型间差异远小于模型—人类差距,包括对草图歧义性的不确定度估计。
  • 人与 CLIPasso 的"稀疏化"策略不同。 在高预算(32 笔 vs 32 秒)下两者唤起的语义标签分布相近,但随着预算收紧,人类草图与模型草图的功能性差异急剧扩大。
  • 视觉语言模型在图表推理上未达人类水平,且错误模式不像人。 Arnav Verma 用 GGR、VLAT、CALVI、HOLF、HOLF-multi、Chart-QA 六套测试比较人类(至少上过一门高中数学的美国成人)与 BLIP2、LLaVA 系、MatCha、GPT-4V:普遍存在差距,在对抗性设计(如异常 y 轴范围)的 CALVI 上差距更大,在 ML 社区常用的 Chart-QA 上差距最小——说明只用现有 ML 基准会低估问题;所有模型的错误模式均落在人类噪声上限之下。
  • 非专家也能根据问题选择合适的图表。 用 R 自带数据集生成数百个问题,让被试从柱状/折线/散点、聚合/分解程度不同的图中选择。最能预测选择分布的假设是"受众敏感":人们偏好那些能让另外约 1,700 名被试实际答对问题的图表特征(用 Jensen-Shannon 散度评估拟合)。
  • 现有图表理解测试测的东西并不清楚。 Erik Brockbank 把 GGR+VLAT 给校园样本和美国代表性样本,题目难度估计一致、两测试成绩相关;但错误模式既不能由图表类型也不能由任务类型(求最大值、找聚类、取值等)解释,一个简约的四因子模型拟合远优于教科书式技能分类,提示需要开发更好的测量工具。

结论与值得注意的细节

  • Fan 的总目标是"闭环":既解释人如何发现有用抽象,又解释如何用抽象去造新东西;长期落点是教育与数据素养——她引用《纽约时报》关于新冠后数学学习损失恢复的报道说明其紧迫性。
  • Tenenbaum 提问点出 SEVA 的方法学局限:4 秒预算混合了思考与作画时间,Fan 承认理想设计应是"无限规划时间 + 有限执行时间",这已困扰她两年。
  • 关于"画蓝色的水"这类"不真实"的图,Fan 的立场是不把图分为对/错,而是问"人在什么目标下为什么这样画"。
  • 对 AI 错误的可解释性,Fan 提出用机制可解释性(她称之为"人工神经网络的认知系统神经科学")诊断模型出错的步骤——是感知瓶颈还是推理瓶颈,并类比优秀教师诊断学生"错误概念"的能力。
  • 关于儿童如何最先学习图表,她坦言没有课程设计答案,但认为好图表应借用视觉系统已有的场景处理与视觉搜索机制,再逐步学会把图形元素映射到概念。
  • 演讲中被现场指出刺激材料存在模板生成的措辞问题(如"5K miles"应为高度),Fan 承认部分问题措辞欠佳但并未影响实验逻辑。
  • 谈到的物理装配/程序抽象研究本次略过,仅在招待会上交流。
核心句型 · 9
1. It'd be really hard to overstate the impact of …
“It'd be really hard to overstate the impact of this invention over the last four centuries”
用否定式反衬强调「怎么说都不为过」。适合学术或演讲中评价重大影响,比 very important 更有分量。可仿写:It's hard to overstate how much X changed Y.
2. Only when … does/did … begin to …
“Only when you see these cases side by side does the kind of morphological variation begin to really become salient”
only + 状语从句置于句首,主句倒装。用于强调某条件是关键前提。仿写:Only when the data is plotted does the trend become visible.
3. I would venture to say that …
“I would venture to say that we'll never be able to explain how and why the world as we know it came to be”
委婉但坚定地提出有争议的断言,学术演讲中常用来降低攻击性。类似表达:I'd go so far as to say…
4. … falls short of explaining …
“A static deterministic account of visual processing falls short of explaining how we generate and make sense of drawings like these”
fall short of + 动名词,指某理论/方案「不足以」达到某目标。用于批评前人理论时语气客观。仿写:This model falls short of capturing context effects.
5. get away with (something more schematic)
“When they could get away with something more schematic or abstract”
本义「侥幸逃脱」,此处引申为「用较低成本的方案也能过关」。口语化,适合描述策略性偷懒。仿写:You can get away with a rough sketch here.
6. … is totally dwarfed by …
“The variation across models in terms of performance is totally dwarfed by the gap between models and people”
dwarf 作动词表「使相形见绌」,用来对比两个量级差距。仿写:The cost is dwarfed by the potential savings.
7. there's only one way to be right, but lots of ways to be wrong
“There's only one way to get all the questions right, but there are lots of ways that you can be wrong and wrong reliably”
对比句式,用于论证「错误模式比正确率更有诊断价值」。结构简单有力,可迁移到任何评估语境。
8. take that as a glass half full call to action
“We're really trying to take that as a glass half full call to action to develop improved measures”
take X as Y 把某结果重新框定为积极信号;glass half full 作定语修饰 call to action。适合把负面发现转为下一步动力。
9. rather than thinking of X as …, we find it more useful to think of X as …
“Rather than thinking of some drawings as good or bad, false or true, we just find it a lot more useful and productive to think of these are the drawings people make under those conditions”
用于回应质疑时重构问题框架,不否定对方而是提出替代视角。仿写:Rather than treating errors as noise, we find it more useful to treat them as data.
词汇精讲 · 125 · 按出现顺序
colloquium /kəˈloʊkwiəm/ n. 0:04
学术报告会、学术讨论会(大学系所定期举办)
inhabit /ɪnˈhæbɪt/ v. 0:37
居住于、栖居于(此处指在这栋楼工作的人)
at home phr. 0:37
自在的、得心应手的(be at home doing/with sth)
psychophysics /ˌsaɪkoʊˈfɪzɪks/ n. 1:18
心理物理学(研究物理刺激与主观感受关系)
of great import phr. 2:02
非常重要(import 作名词表「重要性」,正式用法)
rigor /ˈrɪɡər/ n. 2:36
严谨、严密
keeps her up at night phr. 2:36
让她夜不能寐、萦绕心头的难题
embodied /ɪmˈbɑːdid/ v. 3:07
体现、具体化(embody 的过去分词)
rectangular coordinates n. phr. 3:53
直角坐标系
cutting edge adj. phr. 4:29
最前沿的、尖端的
stumped /stʌmpt/ v. 4:29
难倒、使困惑(stump sb)
millennia /mɪˈlɛniə/ n. 4:29
数千年(millennium 的复数)
overstate /ˌoʊvərˈsteɪt/ v. 5:04
夸大;hard to overstate = 怎么强调都不为过
indispensable /ˌɪndɪˈspɛnsəbəl/ adj. 5:04
不可或缺的
wrestle with phr. v. 5:46
努力应对、苦苦思索(难题)
anatomically modern humans n. phr. 6:22
解剖学意义上的现代人
repurposing /riːˈpɜːrpəsɪŋ/ v. 6:22
改作他用、重新利用
intertwined /ˌɪntərˈtwaɪnd/ adj. 6:22
交织在一起的
ornithologist /ˌɔːrnɪˈθɑːlədʒɪst/ n. 7:01
鸟类学家
morphological /ˌmɔːrfəˈlɑːdʒɪkəl/ adj. 7:01
形态学的
salient /ˈseɪliənt/ adj. 7:01
显著的、突出的
orthodoxy /ˈɔːrθədɑːksi/ n. 7:01
正统学说、正统观念
subatomic /ˌsʌbəˈtɑːmɪk/ adj. 7:49
亚原子的
winking in and out of existence phr. 7:49
倏忽生灭、瞬间出现又消失
leverage /ˈlɛvərɪdʒ/ v. 7:49
利用、借助(本文高频词)
schematic /skiːˈmætɪk/ adj. 8:28
示意性的、图式化的
reformulate /riːˈfɔːrmjəleɪt/ v. 9:07
重新表述、重新构想
instrumentation /ˌɪnstrəmɛnˈteɪʃən/ n. 9:07
仪器设备(总称)
in the service of phr. 10:14
为……服务、用于……目的
socially mediated adj. phr. 10:46
经社会传递的、通过他人中介的
I would venture to say phr. 10:46
我斗胆说、我敢说(委婉而有力的断言)
close this loop phr. 11:28
闭合环路、形成闭环
harness /ˈhɑːrnəs/ v. 12:06
利用、驾驭(资源、能力)
finite /ˈfaɪnaɪt/ adj. 12:45
有限的
contemplate /ˈkɑːntəmpleɪt/ v. 13:15
考虑、设想
get off the ground phr. 13:56
起步、开始运作
instantiation /ɪnˌstænʃiˈeɪʃən/ n. 13:56
实例化、具体体现
denote /dɪˈnoʊt/ v. 14:42
指代、表示
as a matter of convention phr. 14:42
作为约定俗成的结果
ventral stream n. phr. 15:23
腹侧视觉通路(负责物体识别的脑区通路)
per se /ˌpɜːr ˈseɪ/ adv. 15:23
本身、就其本身而言
vindicating /ˈvɪndɪkeɪtɪŋ/ v. 16:06
证明……正确、为……正名
stress test v. phr. 16:06
压力测试、严格检验
rumple /ˈrʌmpəl/ v. 16:06
弄皱、揉皱
pack up and go home phr. 17:07
收工回家(反讽:问题已解决)
falls short of phr. 17:40
达不到、不足以
squiggles /ˈskwɪɡəlz/ n. 17:40
波浪线、弯弯曲曲的线条
get away with phr. v. 18:19
侥幸做到、用……蒙混过关
distractors /dɪˈstræktərz/ n. 18:19
干扰项(实验术语)
exemplar /ɪɡˈzɛmplɑːr/ n. 18:56
样例、典型实例
ablation /əˈbleɪʃən/ n. 19:45
消融(实验):移除模型某部件以检验其作用
operationalize /ˌɑːpəˈreɪʃənəlaɪz/ v. 19:45
操作化(把抽象概念转为可测量变量)
referential /ˌrɛfəˈrɛnʃəl/ adj. 20:34
指称的
proto symbolic adj. phr. 20:34
原型符号的、准符号的
mechanistic /ˌmɛkəˈnɪstɪk/ adj. 21:14
机制性的、关于运作机制的
privilege /ˈprɪvəlɪdʒ/ v. 21:49
优先看重、给予特殊地位
tack on phr. v. 22:25
附加、加上
dissociable /dɪˈsoʊʃiəbəl/ adj. 22:25
可分离的
de-emphasizing /ˌdiːˈɛmfəsaɪzɪŋ/ v. 22:25
弱化、降低对……的强调
tease apart phr. v. 23:01
区分开、分离(相互纠缠的因素)
contraptions /kənˈtræpʃənz/ n. 23:01
新奇装置、古怪机械
hoist up phr. v. 23:58
举起、支撑起
lineup /ˈlaɪnʌp/ n. 24:27
一排(供辨认的)候选对象
Eyeballing /ˈaɪbɔːlɪŋ/ v. 25:00
目测、粗略地看
catch all n./adj. 25:00
兜底类别、包罗一切的
amount to phr. v. 26:26
等于、相当于
fidelity /fɪˈdɛləti/ n. 28:33
逼真度、保真度
cohort /ˈkoʊhɔːrt/ n. 29:13
一批、同期群
worth its salt idiom 30:11
称职的、名副其实的
unmistakably /ˌʌnmɪˈsteɪkəbli/ adv. 30:11
明白无误地
deceptively simple adj. phr. 31:05
看似简单(实则不然)
robustness /roʊˈbʌstnəs/ n. 31:05
鲁棒性、稳健性
evoke /ɪˈvoʊk/ v. 31:05
唤起、引发
cued by phr. 31:53
由……提示、以……为线索
derpy /ˈdɜːrpi/ adj. 33:08
(俚语)傻乎乎的、粗糙笨拙的
term of art n. phr. 33:08
专业术语(此处为自嘲)
entropy /ˈɛntrəpi/ n. 33:41
熵(此处衡量分布的不确定性)
honest to goodness idiom 34:22
确确实实、货真价实
dwarfed /dwɔːrft/ v. 34:22
使相形见绌
currency /ˈkɜːrənsi/ n. 34:59
(此处比喻)计量单位、通用尺度
sparsify /ˈspɑːrsɪfaɪ/ v. 35:36
稀疏化、简化
divergence /daɪˈvɜːrdʒəns/ n. 35:36
散度、差异(信息论中衡量分布差距)
sneak preview n. phr. 36:48
抢先预览、提前剧透
inferential /ˌɪnfəˈrɛnʃəl/ adj. 37:36
推断的
enduring /ɪnˈdʊrɪŋ/ adj. 38:11
经久不衰的
ubiquitous /juːˈbɪkwɪtəs/ adj. 38:46
无处不在的
cornerstone /ˈkɔːrnərstoʊn/ n. 38:46
基石
distilled /dɪˈstɪld/ v. 39:19
提炼、浓缩
calibrate /ˈkælɪbreɪt/ v. 39:19
校准
broke a story phr. 40:15
率先报道一则新闻
in this vein phr. 40:54
沿着这个思路、在这方面
interrogate /ɪnˈtɛrəɡeɪt/ v. 40:54
深入追问、审视
Herculean /ˌhɜːrkjəˈliːən/ adj. 41:28
极其艰巨的(源自大力神赫拉克勒斯)
lenient /ˈliːniənt/ adj. 43:59
宽松的、宽容的
proprietary /prəˈpraɪətɛri/ adj. 43:59
专有的、闭源的
adversarially /ˌædvərˈsɛriəli/ adv. 44:53
对抗性地(故意设计来考验)
at ceiling phr. 44:53
达到上限、接近满分
upshot /ˈʌpʃɑːt/ n. 45:33
结果、要点
testbeds /ˈtɛstbɛdz/ n. 45:33
试验平台
epistemic /ˌɛpɪˈstiːmɪk/ adj. 45:33
认知的、与知识有关的
distilled /dɪˈstɪld/ adj. 48:05
提炼过的、汇总的(图表)
disaggregated /dɪsˈæɡrɪɡeɪtɪd/ adj. 48:37
分解的、细分的
purists /ˈpjʊrɪsts/ n. 49:09
纯粹主义者、死忠派
across the board phr. 50:25
全面地、各方面都
get some traction phr. 52:17
取得进展、找到抓手
convergent /kənˈvɜːrdʒənt/ adj. 52:17
趋同的、一致的
ontology /ɑːnˈtɑːlədʒi/ n. 54:21
本体、分类体系
parsimonious /ˌpɑːrsɪˈmoʊniəs/ adj. 54:21
简约的(模型参数少而解释力强)
unpack /ʌnˈpæk/ v. 55:01
详细展开、逐一解释
glass half full idiom 55:01
乐观看待(杯子半满)
stay tuned phr. 55:37
敬请期待
stand on the shoulders of idiom 56:22
站在……的肩膀上
counterfactual /ˌkaʊntərˈfæktʃuəl/ adj. 57:44
反事实的
phenomenology /fɪˌnɑːmɪˈnɑːlədʒi/ n. 58:55
现象学;主观体验的呈现方式
perspective taking n. phr. 61:20
视角采择(设想他人视角的能力)
misconceptions /ˌmɪskənˈsɛpʃənz/ n. 61:57
错误概念、误解
normatively /ˈnɔːrmətɪvli/ adv. 62:30
规范上、按标准应当地
conceptual primitives n. phr. 62:30
概念原语、最基本的概念单元
mechanistic interpretability n. phr. 62:30
机制可解释性(分析神经网络内部计算的方法)
lend themselves to phr. 63:17
适合于、易于被……
event related potentials n. phr. 67:58
事件相关电位(脑电研究指标)
neurotypical /ˌnʊroʊˈtɪpɪkəl/ adj. 68:51
神经典型的(非神经发育异常的)
co-opted /koʊˈɑːptɪd/ v. 70:03
被征用、被挪作他用
cluttered /ˈklʌtərd/ adj. 70:03
杂乱的
linger /ˈlɪŋɡər/ v. 70:39
逗留、留下
理解自测 · 11 题
1. Fan 用哪些科学史例子说明「让不可见变可见」的工具?各自揭示了什么?

她举了四个例子:古尔德为达尔文绘制的雀鸟图,把形态差异并置后才显现;伽利略的望远镜,提供了质疑太阳系正统模型所需的分辨率;卡哈尔的视网膜显微绘图,展示神经系统各部分及其连接;费曼图,把永远无法肉眼观察的亚原子事件画出来。这些出现在开场的工具史部分(第 10–12 段),共同点是图像本身参与了发现,而非事后装饰;且它们从写实到高度图式化构成一个连续谱。

2. 在「近/远」绘画游戏实验中,操纵变量是什么,观察到了哪些结果?

操纵变量是干扰项与目标的类别关系:近条件下干扰项与目标同属基本层次类别,远条件下来自不同类别。结果是远条件下作画者笔画更少、用墨更少、用时更短,但观看者识别准确率不变,且判断也更快。这出现在第 28–29 段。结论是普通人能根据语境需要动态调整抽象层次,并且这种调整是高效的;后续模型消融表明视觉抽象能力与语境敏感性两者缺一不可。

3. SEVA 基准的规模和设计要点是什么?

SEVA 收集了约 5500 人绘制的 9 万张素描,覆盖 128 个视觉概念,提示图片来自 THINGS 数据集,并设置了不同的绘制时限(最短 4 秒)。随后让人类和 17 个当时最先进的视觉模型做同样的分类任务,以获得每张素描唤起的完整标签分布(第 48 段)。设计上明确针对素描理解的两大挑战:稀疏度变化和语义歧义。Tenenbaum 在问答中指出时限混合了思考与作画,Fan 承认理想设计应分离规划与执行时间。

4. Playfair 1786 年的时间序列图展示了什么内容?Fan 用它说明图表与素描的什么差异?

该图展示英格兰 1700–1780 年间进出口平衡:早期进口超过出口,1750 年代反转,出口大幅增长(第 59–60 段)。Fan 借此指出图表与达尔文雀鸟图的关键差异:素描靠相似性一看即懂,而图表若没见过就不知道在看什么,必须学习;但一旦学会就像「超能力」,能把大量观测浓缩为一眼可读的故事。这解释了为什么数据素养成为 STEM 教育目标,也引出后文对读图能力测量的研究。

5. 累加假说与可分离假说的区别是什么?数据如何裁决?

累加假说认为视觉解释是描绘的增强版,保留全部外观信息再附加机制信息;可分离假说认为解释有选择地突出机制、弱化外观(第 34 段)。逐笔标注显示解释图中因果部件笔画多于非因果部件、背景更少、箭头等符号更多,已与累加假说的强版本矛盾。功能测试进一步显示解释图更利于判断操作动作,描绘图更利于辨认物体,而解释图在需要图与实物细致对应的第三任务上并无优势(第 40–43 段)。综合支持可分离假说。

6. 为什么 Fan 认为只看总分不足以判断模型与人类对齐?她的证据是什么?

因为「答对只有一种方式,答错有很多种」,错误模式比正确率更能揭示系统内部是否用相似方式加工(第 67 段)。证据是:在六项图表测试上,GPT-4V 总分看似接近人类,但所有模型的错误模式与人类错误模式的相关度都远低于人类噪声上限(人与人之间的一致性),即没有一个模型「错得像人」(第 67–68 段)。她还指出若只用 ML 界流行的 Chart-QA,人机差距会显得很小,而 CALVI 等对抗性设计的测试暴露出更大差距。

7. 人类与 CLIPasso 在素描稀疏化上的差异说明了什么?

在宽松预算下(32 秒人类 vs 32 笔 CLIPasso),两者风格不同但唤起的标签分布相似,即功能上等价;但预算越紧,两者标签分布的散度越大(第 54–56 段)。这说明在极简条件下,人类决定「保留什么、舍弃什么」的策略与基于 CLIP 语义空间优化的策略不同。推理链是:若素描意义定义为它唤起的完整分布,那么差异只在紧预算下出现,意味着人类抽象的核心机制正是在信息极度受限时才显露,这是未来建模最值得切入的地方。

8. 「受众敏感」假说是如何被验证的?它为何让 Fan「充满希望」?

假说认为人们选图时考虑的是该图对回答特定问题的实际效用。验证方法是另找约 1700 名参与者,用每一种候选图回答每一个问题,实测各图的回答表现,据此构建预测的选择分布,再用 Jensen-Shannon 散度与真实选择分布比较(第 75–76 段)。该假说在整个数据集上都优于「柱状图死忠派」「偏好展示更多数据」等简单策略。Fan 感到希望是因为这说明非专业人士也能感知图表特征与问题的适配性,作为统计导论教师,这意味着数据素养有可教的基础。

9. Fan 对现有图表素养测试提出了哪些方法论批评?

第一,两项主流测试 GGR 和 VLAT 成绩相关,但不清楚它们共同测的是什么;第二,每种图表类型通常只有一道题,无法在图表类型与表现之间建立统计关联;第三,错误模式的最佳预测因素既不是图表类型也不是任务本体(找最大值、识别聚类等),而是一个探索性四因子模型,说明教科书式的技能分类可能不对应真实认知成分(第 79–82 段)。她把这些解读为「杯子半满」的行动号召,即开发更好的测量工具,而非否定该领域。

10. 若有人反驳说「把水涂成蓝色」证明绘画本质是文化约定而非相似,Fan 会如何回应?

她在问答中已给出回应(第 87–89 段):不从「这幅画是错的」这一前提出发,而把它视为人们在某个目标下生成的真实数据,任务是推断那个目标是什么——是忠实于看水杯时的现象学体验,还是习得的描绘约定。她承认美学哲学中关于虚构图像的文献是灵感来源,但方法上更倾向于问「为什么在这些条件下画成这样而非别样」。这与她整体立场一致:相似性与约定不是非此即彼,而是共同约束表征决策的来源,需要通过实验分离其权重。

11. Fan 关于「用机制可解释性诊断错误来源」的设想,放到教育场景是否成立?有哪些限制?

她的设想是:优秀教师能从学生表述中诊断错误概念的性质,而不只判断对错;类似地,可用机制可解释性工具审视模型答对/答错时依赖的操作与概念原语,把瓶颈定位到知觉层或推理层,最终得到对人和机器通用的诊断工具(第 93–96 段)。在教育场景下,这一思路与教育研究中的「错误概念诊断」传统吻合,有成立基础。限制在于:她自己承认「非常非常难」;模型错误模式目前并不像人,因此模型内部机制未必能直接映射到学生的知识图谱;且问答中提到读图涉及「看哪里、如何映射到坐标轴、如何转为答案」多个步骤,需先把任务拆解到数轴距离判断这类最简形式,才能逐步建立诊断工具。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.026MIT Godel Escher Bach Lecture 1 下一期 · NO.028 →The full-length interview with Yuval Noah Harari | The Economist
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com