视频库 / NO.097
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Atlas of AI: Kate Crawford in conversation with Judy Wajcman

节目发布 2021-06-28 · The Alan Turing Institute
凯特·克劳福德 朱迪·瓦克曼
本期追问 · 点击跳到视频对应位置
28:44 AI 系统显得理性客观,我们凭什么就信它的结果?35:20 为什么说靠调整数据追求公平,没有抓住真正的问题?14:50 算法管理追踪到工人的肌肉与韧带,这还是效率问题吗?9:40 一本讲人工智能的书,为什么要从一座锂矿写起?
归入 Ⅵ·01 技术是中立的吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2021年春,凯特·克劳福德(Kate Crawford)的新著《AI地图集:人工智能的权力、政治与地球代价》(Atlas of AI)在英国上市数日后,艾伦·图灵研究所线上举办了一场对谈。克劳福德是微软研究院高级首席研究员、南加州大学安纳伯格学院研究教授,AI Now研究所联合创始人。与她对谈的朱迪·瓦伊克曼(Judy Wajcman)是伦敦政治经济学院社会学教授,主持图灵研究所「数据科学与AI领域的女性」项目,长期从事科学与技术研究(STS)及技术女性主义研究。本文依据现场录音编译整理,仅删去口语枝节与寒暄,论证与细节悉数保留。

STS视角与「权力登记册」

瓦伊克曼:今天有几百人在线,这本身就说明凯特作为研究人工智能社会与政治影响的学者,声望有多高。我知道在座大多数人还没来得及读这本书,它在英国真的是这两天才刚上市。所以我想先用几分钟讲讲,我为什么如此欢迎这本书。老实说,我是一路点头读完的,因为它极其丰富地示范了如何用科学与技术研究(STS)的视角去审视一项新技术,而且带着恰到好处的马克思主义色彩。对我来说,这正是思考这些问题的理想透镜。

第一,我喜欢这本书每一章都是历史性的。你把每一项技术发展放回它诞生的历史语境,展示这些起源如何嵌入技术本身,无论是人脸识别还是情绪识别。你在全书中反复强调,要理解这些系统如何运作,必须先理解它们从哪里来。这一点我完全赞同。历史视角还提醒我们,早先几轮技术革命时,我们已经有过许多同样的讨论。我是在二十世纪八十年代微电子革命时期进入这个领域的,我清楚记得当时关于技术影响的争论,无非是一种简单的二元对立:乌托邦还是反乌托邦。STS这个领域,很大程度上就是为了回应那些争论背后的技术决定论而兴起的,它强调技术从来不是自主的,而始终受社会、政治与经济力量的塑造。凯特,你说得非常清楚,人的选择嵌在设计与架构之中,用今天的话说,嵌在技术的物质性里,不管是硬件还是软件,不管是兰登·温纳(Langdon Winner)笔下的那些桥,还是AI系统。

第二,你选择用地图集(atlas)或者说地图作为透镜,这让我很感兴趣,也完全把我说服了。我得多琢磨琢磨,因为我向来对时间比对空间更在行。但我确实认为这是一个绝妙的透镜,它让你能把「基础设施」的概念延伸到行星尺度,涵盖哈特与奈格里(Hardt and Negri)所说的信息资本主义中抽象与采掘的双重运作,包括自然资源的代价及其劳动过程。所以,秉承马克思主义揭露商品拜物教的优良传统,你揭示了AI生产的物质条件,揭示它的整个生命轨迹:一个AI系统从矿物与能源开始,经过一条全球劳动供应链,到企业和国家如何将它投入运作,再到它对我们生活的影响。和所有技术一样,AI一点也不「人工」,它是不折不扣的人造物,如果允许我这么说的话,是「男人造的」(man-made)。

第三,我非常欣赏你的这个论点,我认为它是全书的核心:把AI称为「智能」,就是在援引一种极其狭窄的智能定义,一种把生活经验的复杂性压平的定义。AI系统按照数学最优化的规范逻辑运作,横跨教育、医疗、福利、信贷、刑事司法等诸多领域,结果就是只有可测量的东西才算数,这是一种以数字治理的方式。你的关键论断是:这些AI系统看起来理性而客观,而这种表象恰恰为它们的使用提供了合法性,可它们归根到底是为现存的支配性利益服务而设计的。在这个意义上你说,AI是一份权力登记册(registry of power)。

这本书做得极其出色的地方,是它并不主张这些趋势完全是新的,而是指出,行为痕迹数据的大规模积累从根本上强化并放大了这些趋势。所以我想先问你:你是什么时候意识到这些巨变的?你从一开始就在场,并且看出这是我们眼下面临的最关键的议题之一。你怎么知道现在恰好是推出这本书的完美时机?我们这些在这个领域做研究的人,一直盼着有人把这些东西整合起来,而你写出了这本漂亮的书,用如此易读、清晰、有说服力的方式把它们贯通在一起。给我们讲讲你的历程吧。

写作缘起与AI的物质生命周期

克劳福德:朱迪,我得说这是我听过最慷慨、最让人受宠若惊的开场。谢谢你,尤其因为我一直是你的作品的忠实读者,从很久以前就是。九十年代我书架上最显眼的书之一就是《技术的社会塑造》,之后的每一本我都读过,包括《时间的压迫》(Pressed for Time),我在《AI地图集》里也写到了它。所以今天能和你对谈,对我意义重大。我也要感谢艾伦·图灵研究所主办这场活动。

你问到时机,我怎么会知道2021年是合适的年份。我们刚经历了一场大流行病,看到你刚才精彩概括的那些逻辑全都在加剧。我很想说这是我算准的,但实话是,这本书写了五年多,我当年看到的那些主题会这么快变得这么无处不在,是我无论如何想象不到的。对我来说,真正的机缘是2012年加入微软研究院。那正是卷积神经网络登上中心舞台、深度学习加速的时候,各家工业研究实验室开始以前所未有的方式投资机器学习。当时已经很清楚,从计算机视觉到自然语言处理,一切都将被彻底改变。但与此同时,也存在非常真实的问题。

我们可以把这些问题当作一个生命周期来看。AI系统需要海量计算,而且计算密集程度还在不断上升;我们的消费级AI设备,无论手机还是笔记本电脑,消耗着大量矿物,从稀土到钴再到锂。驱动AI的数据必须由人来标注,这些人在加纳、菲律宾等地的远程工作平台上劳作。最后,这些系统被部署到高度泰勒化的工作环境里,算法管理系统每天都在作用于工人身上。所以你能看到,AI是一个完整而有形的系统,它对世界的影响远比把它想成抽象的计算或者无形的算法要深刻和物质得多。这本书想做的一部分工作,就是把这些系统安放到世界各地实际部署它们的具体地点和机构之中。

从锂矿看AI作为采掘业

瓦伊克曼:说得好。所以你在书的开头讲述自己去内华达州参观一座锂矿的旅程,再合适不过了。在那里,你可以看到AI在最字面意义上就是一种采掘业。我得说,我们两个都是澳大利亚人,都明白采矿在政治中有多重要,这还是说得含蓄的。你从矿开始写,我觉得非常好。我想请你讲讲那段经历,但也想问问你怎么看这样一个现象:在新冠危机期间,我们一直在安慰自己说碳足迹在下降,可你在书里特别强调,我们天天开的这些视频会议,耗费着惊人的能源。

克劳福德:完全正确。有意思的是,很多人问我,为什么一本讲人工智能的书要从一座锂矿开篇。原因恰恰在于,要把人工智能落到实处,要真正理解这些物质影响,我们必须去它被制造出来的地方,我说的「制造」是最完整意义上的制造。对我而言,这意味着去美国最后一座仍在运营的锂矿,它在内华达州一个叫银峰(Silver Peak)的地方。你能看到巨大的锂盐卤水池,泛着虹彩般的绿色,在太阳下晾晒,产出人们所说的「灰色黄金」,也就是锂离子电池的原料。这些锂离子电池当然需求极旺,iPhone里有它,特斯拉汽车和电动车里也有它。但我们面临非常严重的供应问题。我相信你也看到了,拜登政府最近刚发布一份危机文件,说必须确保驱动当前信息资本主义和行星计算的诸多矿物的供应链安全。所以此刻正是时候去思考:制造那些我们每天理所当然指望它们运转的AI系统,究竟需要多少这样的物质?这些矿物的可用性,以及它们背后的地缘政治,其实都正处在一个临界点上。

幽灵劳动与伪自动化

瓦伊克曼:好,下一章,我会尽量多讲几章。第三章讲的是制造AI所涉及的各种形式的劳动,从矿工到内容审核员,到亚马逊仓库工人,再到硅谷的工程师。这一章我也很喜欢,我知道我每一章都会这么说。我觉得它特别令人耳目一新,因为过去几年我花了大量时间参加各种关于「工作的未来」的讨论。上周我刚参加了一场,和麻省理工的大卫·奥托(David Autor)一起谈技术与工作的未来。经济学家们经常告诉我,机器人很快就会包揽一切,一份工作都不会剩下。奥托是例外,他不这么说。而你说得很对,他和我也这么认为:那场争论在某种程度上偏离了要点。更重要的是去看当下的工作体验,看这些技术如何助长日益加强的监控和算法管理系统。这个我稍后会问。但我想先请你简单谈谈「幽灵劳动」(ghost work)的重要性,就是格雷(Mary Gray)和苏里(Siddharth Suri)所说的那种,在我看来这也是一个很值得强调的点。

克劳福德:很高兴你提到玛丽·格雷和西德·苏里的工作。《幽灵劳动》是一本极其重要的书,它揭示了在远程工作平台上做工是什么滋味。这些工作实质上是数字计件工,报酬极低,属于低于贫困线的工资,而工作本身可能极其单调,也极其令人紧张。从莉莉·伊拉尼(Lilly Irani)到阿斯特拉·泰勒(Astra Taylor),许多学者都研究过这个。泰勒用了「伪自动化」(fauxtomation)这个词,指的是我们以为是自动化,实际上却是人在撑着这些系统。这种现象已经浮现有一段时间了。就AI而言,朱迪,你在数字助手方面的研究中肯定也看到了同样的模式:整条供应链上,这套系统有多大一部分其实是靠人做着相当琐碎的任务来伪造出来的。我在书里写到一个叫x.ai的系统,他们真的让人假扮数字助手,每天工作十四小时,工作条件相当痛苦,只为营造AI运转无缝的印象。这也是你的研究早就指出的问题,如今它正大行其道。

从巴贝奇到贝索斯:算法管理

瓦伊克曼:是的,关于硅谷的社会想象以及它们所扮演的角色,我们可以谈很久。但我想进入这一章的一些细节。你用一座亚马逊仓库里的大型打卡钟的画面开启这一章,我知道你去过那里。你借这个画面追溯了用自动化控制时间的漫长历史,经由E. P. 汤普森、布雷弗曼、泰勒主义的科学管理、福特制。这些都是我的老本行,可以说是我一辈子研究的材料。我想请你多讲一点算法管理系统在劳动力市场两端的应用:一端是平台工作,另一端是我一直在研究的知识工作,以及它正在如何被改变。

克劳福德:我喜欢你提问的方式,因为我们确实必须把算法管理放在两端来思考:一端是它如何被引入低工资的、传统上所谓的蓝领工作,另一端一直延伸到高薪的白领岗位。有意思的是,这些系统的谱系可以追溯得很远。查尔斯·巴贝奇(Charles Babbage)最出名的是差分机,但他也写了大量社会理论。他对未来的一个设想是,将来会有高效的系统监视工人,追踪他们每一分钟在做什么,以确保他们被最有效率地使用。这是一个相当骇人的愿景,而它确实在今天的一些系统中变成了现实。

在亚马逊的运营中心里待过一段时间,对我来说是绝对恐怖的经历。这些地方被当作把自动化(既包括算法式的,也包括机器人)与人结合起来的样板工作场所。但它们同时也是另一种东西的样板:在那种条件下工作所承受的非同寻常的心理与身体压力。我相信你也看到了,就在上个月,杰夫·贝索斯说,针对这些工作环境中工伤和压力的增加,他们的应对措施之一是引入一套新的算法管理系统,把追踪细化到工人正在使用的韧带和肌肉的层面,以求产生新的效率。在我看来,这就是巴贝奇的愿景活生生地实现了,相当噩梦般。

远程办公与情绪识别监控

瓦伊克曼:是啊,这也太像泰勒主义和他的秒表了。那另一端呢,知识工作者那一端?我同意你的看法,我们真正需要关注的变化,是机器学习工具使得追踪、助推和评估的颗粒度变得空前精细。在向居家办公转变的背景下,我自己对此非常担忧。这难道不会刺激对知识工作者的监控通过这些工具不断加强吗?最糟糕的情形,你在情绪识别软件那一章里描述得很好:不只是身体动作,连情绪状态都要被监控。你是否也感觉到,当前的新冠危机会成为加强对知识工作者监控的一个推力?

克劳福德:这已经发生了。过去十八个月,我们看到相关服务大幅激增:通过摄像头追踪员工,统计他们发了多少封邮件、开了多少次会,然后算出一个效率分数,拿去和其他远程工作者比较。提供这类服务的大量初创公司和企业,业务量都出现了猛增。类似地,还有像「四棵小树」(4 Little Trees)这样的公司,提供所谓的情绪「检测」,我们说「检测」时是要打引号的。它针对的是同样在家学习、努力跟上课程的年轻学生,用摄像头捕捉他们脸上的微表情,看他们是否在专心听讲,并试图由此推断出某种内在情绪状态。

为了写这本书,我花了很多时间研究这个问题,也专门写了一章。这里的假设,即你可以看着一个人的脸就知道其内心状态,问题极大。这种想造出一台「AI测谎仪」的欲望,在我看来从根子上就是坏的:支撑它的理念和理论是坏的,试图实现它的系统也是坏的。所以,这绝对是我认为迫切需要监管的领域之一,因为在大流行期间,我们看到这些系统越来越深地渗入日常工作的工具中。

训练数据的考古学

瓦伊克曼:是的,监管的问题我们稍后再回来谈,我确实想问你这个。但我想赶紧进入你关于数据和分类的两章,这两章非常精彩,也正是你自己的研究真正开创性的地方。作为《时间的压迫》的作者,我很清楚时间不够用,所以我真想赶紧谈到这里。先从数据开始吧。你谈到大规模的数据采集正在驱动AI的成功。我们今天的听众背景非常多样,你能不能讲讲,训练数据集到底是怎么运作的?也许还可以讲讲做标注工作的那支计件工大军所遇到的问题。

克劳福德:简单地说,监督式和非监督式机器学习系统是用大量数据构建的,这些数据用来表征世界,我们称之为训练数据。在某种意义上,训练数据成了系统实际如何运作的「基准真相」(ground truth)。对我来说,最有启发也最有意思的事情之一,就是真正去研究训练数据如何运作,也就是把训练数据集打开来看,不只看里面的数据,还要看它被分类和排序的逻辑。让我大为震惊的是,这种「挖掘」竟如此罕见。对绝大多数训练数据集来说,它们只是被当作聚合的基础设施来用:拿来就用,别想太多,用它来构建机器学习模型。

但这里恰恰藏着一个非常实在的问题。一旦你真正开始细看这些系统如何运作,一旦你开始对训练数据做这种考古,你会发现各种各样的问题。不只是人被归入二元性别,或者五分法的种族类别,这些分类是彻底坏掉的,而且大家也都知道它们坏掉了;还有根据外貌对人的品格做道德判断。今天的系统里,这种近乎「颅相学冲动」的东西,问题数不胜数。而正是在训练数据这个层面,你可以开始看到那些分类,以及它们从哪里来。

「越多越好」:数据意识形态的来源

瓦伊克曼:是的,完全同意。你在这一章里谈到,数据如今被当作一种自然资源来开采和提取,这让我印象很深。我喜欢「开采和提取」这个说法,从字面意义上的采矿,一路到数据被以同样方式对待。你还谈到,正如玛丽昂·福尔卡德(Marion Fourcade)和基兰·希利(Kieran Healy)所说的,存在一种收集更多数据的道德律令:因为我们能收集,数据就在那儿等着被收集,所以不停地收集就成了一种道德上的义务。你还指出这与另一个观念相关:数据越多,人就越可知,于是我们每个人都被造出了一个「数据分身」,或者用卢克·斯塔克(Luke Stark)的话说,一个「可缩放的主体」。你能不能解释一下你的意思?

克劳福德:有意思的是,这个表述酝酿了很久。「越多越好」这种意识形态,可以一直追溯到二十世纪后期最早的那批AI实验室。媒介史学者李肇兴(Xiaochang Li)在她关于IBM连续语音识别实验室的研究中,对此有精彩的记录。罗伯特·默瑟(Robert Mercer),没错,就是后来资助特朗普竞选以及许多别的事情的那个默瑟,当年还是一名研究科学家的时候说过,更多的数据永远是更好的数据。我们在那个历史时刻看到的,正是AI从此前占主导地位的专家系统和符号逻辑路径,转向概率方法和暴力计算路径。

从那时到现在,这种意识形态不断发展:我们应当尽可能多地采集数据,哪怕你不知道它有什么用,哪怕持有这些数据实际上带来各种责任风险,人们仍然相信它自身就蕴含着某种价值,即便你暂时看不出来。我认为我们现在必须认识到,这实际上是一条非常成问题的路径。它造成了数据的过度采集,造成了人们对数据集被合并所深感忧虑的局面,我和杰森·舒尔茨(Jason Schultz)在一篇论文里把这称为「预测性隐私伤害」。所以,数据被如此设想和理解的方式,实际上已经把我们带到了计算史上一个非常成问题的节点。

瓦伊克曼:是的。我还觉得有意思的是,这种数据积累正在如何影响我们自身的主体性,这是另一个场合的话题了:自我数据的积累里,某种程度上内置了自我优化。你就此写过很好的文章,我很喜欢你那篇论文,我现在记不清标题了,但我一直拿它来教学,它从体重秤讲起,一路谈信息的积累如何影响你对自己身体的感知,以及你在世界上的运作方式。

数据集的来源、弃用与来生

瓦伊克曼:不过我们该往下走了,我看到已经有很多提问。我想谈谈隐私问题,这方面的问题太多了。人们在谈「数据分身」时最担心的一点,当然是我们从未被征询过关于自己数据的使用、流通和提取。我记得莫罗佐夫(Evgeny Morozov)几年前曾呼吁我们应该为自己的数据获得报酬,诸如此类的讨论。不过随着这个领域日趋成熟,也许更有意思的是你谈到的另一个问题:缺乏记录数据来源的标准做法。而考虑到许多数据集是私有的,这个问题就格外突出。你说你研究过数百个数据集,那些应该都是公开的吧。对于大多数数据集其实并不在公共领域,因此像你这样想去检视它们的人根本接触不到,我们能做些什么?

克劳福德:这是一个非常实在的问题。说实话,我至今记得几年前和蒂姆尼特·格布鲁(Timnit Gebru)的早期谈话,她当时是我们微软研究院的博士后。我们都惊讶于一个事实:对于一个训练数据集从哪里来,根本不存在任何标准化的信息。你看,一块硬件会有一份数据表(datasheet),告诉你这颗半导体只能在什么温度、什么条件下使用,可数据的使用却没有任何对应的东西。这就是《数据集的数据表》(Datasheets for Datasets)那篇论文的缘起,它试图阐明:要理解一个数据集究竟为何而设计,我们需要哪些标准、信息、历史和来源,更重要的是,它什么时候就不该再用了。

这是另一件我们谈得远远不够的事:这些数据集何时应该被弃用,以及当它们被撤下之后,仍在学术种子站(Academic Torrents)之类的地方继续拥有「来生」时,会发生什么。就在过去几年里我们看到,一些训练数据集被发现有问题,创建者说好,我们撤掉,可它们继续流传,当然也继续为许多生产级系统提供养料。这是问题之一。

你指出的第二个问题是,对于大型科技公司内部训练数据集的规模和构建方式,我们根本一无所知。这些东西被当作高度机密的专有信息保存,作为研究者,我们看不到。这就引出一个大问题:我们该如何理解这些系统被用来分类、用来解读的方式,理解它们如何喂进广告技术、保险、警务等领域的逻辑。这是这个领域一个重大的研究障碍,我们必须直面它。

分类即政治:ImageNet案例

瓦伊克曼:这正好引出你的下一章,关于分类的那一章,我认为写得极其出色。让我问你几个相关的问题。那一章的核心论点是,分类的过程,在这里就是标注的分类体系,本质上是政治性的。因此,针对统计偏差的狭隘技术方案,以及为了「更公平」而调整数据的努力,都没有抓住要点。我一直对那些关于公平性(FAccT)的会议很感兴趣,对所有那些通过重塑数据来让它变得公平、平等的技术尝试也很感兴趣。我想和你谈谈这有多难,以及用那种方式来表述问题本身有什么问题。你在分类这一章里非常出色地运用了杰夫·鲍克(Geoffrey Bowker)和苏珊·利·斯塔尔(Susan Leigh Star)的研究,即分类作为秩序系统,在塑造我们如何看世界方面拥有巨大的认识论权力。读这一章时,我不禁想起早期的女性主义研究,比如桑德拉·哈丁(Sandra Harding)等许多人关于科学知识角色的研究。我们早年做过大量所谓「科学知识社会学」的工作,思考科学知识如何建构生物学上的性别二元,在其中女性被「天然地」定义为劣于男性。当你谈到AI系统如何假定,或者说如何构建、如何「烘焙」进去,把性别、种族和性取向当作天然的、固定的生物学范畴时,我想到的就是这些。你能谈谈这个核心问题吗?也许可以通过描述你和特雷弗·帕格伦(Trevor Paglen)合作的ImageNet项目来谈。

克劳福德:这个问题太丰富了,有太多层面。就我个人而言,我最早发表关于大规模数据系统偏差问题的研究,大概是在2010年。当时的观点是,偏差是一个只需收集更多数据就能解决的问题:如果这里有问题,那就再多弄点数据,问题就解决了。但过去十多年我们看到的恰恰相反:对于极大规模的数据系统,偏差实例比比皆是。苹果的信用评估算法一贯给女性更低的额度;语音识别系统识别不了女性的声音;人脸识别系统识别不了肤色较深的人。如今例子数不胜数。

我的思考发生了一个变化:我开始把偏差看作某种「巨型动物」,也就是那些大的错误、系统显眼的失效点。但我认为我们需要再往下走一层,进入构建的逻辑,我认为那才是真正更深刻地塑造系统如何看世界的东西。在很多情况下,我们不会看到那种壮观灾难式的失效点,取而代之的是一种细密的方式,人在其中被理解、被解读、被估值。你在那么多系统里都能看到这一点,从招聘到刑事司法,人被指派某种属性,可能是性别、种族,也可能是一个风险评分、一个信用评分。这些观念从哪里来,它们如何被烘焙进系统,是我们眼下的核心问题。

过去五年一个非同寻常的现象,是对公平、问责与透明议题的关注,比如FAccT会议。我认为这些是必要的干预,但并不充分。因为在许多论文里,这些被当作纯粹的技术问题,需要的是技术修补,而没有去思考:当我们摄入来自过去的数据时,我们也在吸收那些结构性不平等、偏见和观看世界的方式,而这些本身就是深有问题的。这正是我们和艺术家特雷弗·帕格伦做ImageNet项目时所看到的。

真正开始审视ImageNet这样重要的东西,是一段非同寻常的经历。我们开始研究它时,它已经上线近十年,是图像识别领域的一座巨像,可是从来没有人在具体类别的层面研究过它里面发生了什么,尤其是「人」这个类别。我们在那里发现了大量赤裸裸的种族主义和厌女的词条,还有一些完全说不通的词条,也就是没有视觉对应物的词,比如把一个人描述为「债务人」「朋友」或者「熟人」。这些观念根本不属于我们所理解的视觉名词。所以我们做的一件事,就是打开这些关于知识如何被制造的认识论问题。我认为对这种深层次的分类,需要有远为审慎的批判性介入,不仅在技术领域如此,而且从根本上说这是一个社会技术问题,这意味着需要不同类型的专业知识围坐在同一张桌子旁。

招聘算法、多样性与设计的边界

瓦伊克曼:这些问题很棘手,不是吗?一旦你承认它们是结构性问题,就很难修补了。比如亚马逊那个自动化招聘工具,它不选女性。你在书里谈到,那在某种意义上是个难以修补的问题。也许你可以展开谈谈,因为它涉及语言中内嵌的性别用法,涉及各种各样的东西,在那个应用上修修补补是解决不了的。那这会把你引向哪里?这和我们在图灵研究所的性别项目特别相关。这类例子不断冒出来,得到大量曝光,比如女性化的语音助手,然后就有讨论说该怎么办。我知道苹果内部讨论过所谓「中性声音」,要不要发明一种听不出是男是女的声音,接着又有口音的问题。所有这些都很难,对吧?就拿招聘来说,一旦认识到这些关联有多深,接下来该怎么办?

克劳福德:这把我们带回到那个问题:什么可以用技术手段解决,什么必须通过监管来解决。这非常应景,因为就在过去几个月,知名招聘公司HireVue,它曾在视频面试中使用情绪识别,高调宣布他们听取了批评,将把这一功能从系统中剥离。但他们仍在使用语音情绪检测之类的东西,也就是听你的声调,然后试图判断它是否说明了你的品格或者你是否适合被雇用。所以在招聘这个层面,我认为我们有很多问题要问:谁被看重,谁是理想主体或者理想雇员?不过朱迪,我很好奇你的项目怎么看这个问题:性别本身在这些系统中被折射的方式,你们如何应对?你认为设计能在技术层面有所贡献,还是说这更需要从政策和监管的角度来思考?

瓦伊克曼:我认为这些都得做。我们得对这个领域缺乏多样性的问题有所作为。看看顶尖的机器学习会议,女性贡献者的数量少得可怜,整个从业群体严重失衡,主要是年轻男性。我并不认为这是解决方案,但我几十年来一直强烈主张,在某种意义上,人只能从自身经验出发去设计。如果喂进设计和系统的经验范围极其狭窄,无论硬件还是软件,你得到的想象力和创造力就非常有限,这对我们所有人都没有好处。所以我认为这一点非常重要。

我们还希望在组织文化方面做更多研究,因为在我看来,我们得走进组织内部,看看那里发生了什么,为什么只有某些类型的员工在那里感到自在,为什么那里存在一种「寒冷的气候」。我认为在这些方面有所作为,会真正推动更广泛的讨论,既关于监管,也关于一种更有见识的公民身份,也就是关于这些问题的一场更广泛、更好的公共辩论,尤其是关于我们何时想用这些系统,何时不想用。把这些工具作为招聘的一部分来使用,是有区别的:对人力资源管理者或者法官来说,拥有更多数据当然是好事,拥有更多知识当然是好事。但正如弗吉尼亚·尤班克斯(Virginia Eubanks)在她的书里精彩地讨论的那样,人的自由裁量在哪里?判断在哪里?我们如何把这些结合起来,并决定何时想用这些系统,何时不用?不过还是请你自己展开谈谈这些。

创新观的失衡:技术 vs 政策

克劳福德:这正是我喜欢和你谈这些话题的原因,因为你刚才说的里面有太多可以深挖的东西。比如,我们把「创造性创新」这个观念寄托在何处。过去十五到二十年,硅谷讲述了一个关于技术创新的故事,不惜一切代价地颂扬它:快速行动,打破常规,这是我们做事的最佳方式,我们创造技术创新。但我们没有做的,而且我认为这是一个真正的损失,是去思考政策、监管、法律层面的创新,以及所有那些能让我们考量这些系统社会影响的伦理框架和方式。这些空间从未获得过同等程度的意图、兴趣、投资或者社会关注。而我们现在看到的,就是这种局面结出的果实:我们在很大程度上过度投资了一种高度技术化的关于「什么算数」的愿景。

你刚才谈到,更多数据和更多信息是不一样的。什么构成信息,以及正如你所说的,良好的判断力,这些在这里都极为重要。这又把我们带回到系统如何被设计的历史。我总是想到洛兰·达斯顿(Lorraine Daston)和彼得·盖里森(Peter Galison)关于客观性的研究,他们考察了向「机械客观性」的转变,也就是人们决定通过工具就能获得对世界更真实的描述的那个时刻。我们确实经历了这样一个时刻:机器学习被视为一种万能工具,可以应用于任何事情,从福利领域,我们在尤班克斯的书中看到了它如何让我们失望,一直到刑事司法领域,我们可以想到ProPublica的报道,想到The Markup的调查,它们都在揭示这些系统如何在那里让我们失望。

所以我们面临一个抉择:我们打算怎样使用接下来的十年?我们真正需要创新和创造力的地方在哪里?我的建议是,我们应当把思考放在政策上,放在公共辩论上,放在监管上。而且我是乐观的,因为我开始在这些领域看到一些非常积极的信号。比如在欧盟,我们看到了第一份综合性的AI监管草案。它也许还处在早期,还不完美,但在思考如何更好地监管这些系统方面,是朝正确方向迈出的一步。

伦理准则不够,要谈权力

瓦伊克曼:我想多问问伦理的问题。你在书里有一句很好的话,大意是说,此前太多的关注放在了伦理上,而对权力的关注不够。我认为这是对的,我们应当少谈些伦理,多谈些权力。你提到伦理准则正在激增,我记不清有几百份了,反正数量庞大。我想知道你怎么看这件事。你是否设想在某个时点会出现某种全球性的协议?这些准则是怎么运作的?还有,既然你显然认为我们太关注伦理而不够关注权力,是不是太多精力都流向了伦理准则这个方向?聊天区里也有不少问题,我看了一下,都是很棒的问题。所以请你多讲讲。

克劳福德:谢谢这些提问,我们稍后一定会谈到。让我从两个方面来回答。其一,在哲学上,伦理与权力的问题显然总是交织在一起的,我们有几十年乃至几个世纪的哲学可以援引,来说明这一点。但我在书中具体指的是,伦理准则被用作一种打发监管的手段,一种说「我们自己搞得定」的方式:我们会有一套伦理准则,因此不需要任何法律或者监管护栏来约束我们如何创建这些系统、这些系统可能伤害谁。坦率地说,这不够。当然,就我们如何设计和部署系统而言,阐明各种伦理原则是重要的,但还有更深层的东西:我们如何真正确保这些原则是可问责的?有什么办法能保证,一旦出了问题,同样的事不会再次发生,并且会有人真正为此负责?

正因如此,我认为监管远为必要,对权力的分析也远为必要。这个系统是否给强者赋予了更多权力?这是我们对自己构建的每一个系统都应当追问的问题。还有,这些系统以何种方式与现存的权力形式相呼应,无论是资本、警务,还是军队,不一而足。在我看来,这比那些关于安全、关于「确保不造成伤害」的高层次伦理声明有用得多,因为我们一次又一次看到,这些准则并没有在现实中兑现。

参与式监管与拒绝的政治

瓦伊克曼:有意思。硅谷现在特别流行搞用户群体、用户体验小组,仿佛那就是答案:拉几个用户进来测试一下,可那是一种非常狭窄的使用情境。而你非常强调,应当被征询、应当参与设计的,是那些受这些技术影响的人。我们这个圈子里多年来有不少关于参与式设计的运动。你能谈谈这个吗?你认为这是整件事中重要的一环吗?

克劳福德:绝对是。我和许多学者,包括安德鲁·塞尔布斯特(Andrew Selbst)等人,一直在研究的一个东西就是影响评估:如何建立这样一种影响评估机制,让社区成员对政府是否部署一套人脸识别系统、一套AI招聘系统,或者刑事司法系统内部的一套风险评估系统拥有发言权。在很多很多环节,你其实都可以引入公共辩论和公共决策,还可以拥有一种「拒绝的政治」,也就是能够说:不,我们不想在这里用这套系统,我们根本不想它被应用。这只是我们可以着手构想的参与式监管的众多机制之一。你也知道,在许多情况下,法律是在科技公司的影响下制定的,有些甚至有一部分是科技公司代笔的。当我们看到科技行业与它所服务的人群之间存在如此非同寻常的权力不对称时,我认为这种做法根本不可能把我们带到该去的地方。所以我认为,我们需要更激进地重新思考监管如何形成,以及社区如何对监管的实施拥有发言权。

瓦伊克曼:这很有意思,因为我上次在斯坦福参加的会议叫「为关怀编程」(Coding for Care),讨论的正是老年人照护中的哪些环节可以交给技术。这类讨论总是围绕着「关怀」,甚至就是这样被提出来的,而且它的表述方式是让个人尽可能长时间地独自待在家里,然后去设想能够运转的系统。正是在这样的情境中,你真的需要退后一步想想:你到底想用技术做什么,什么是不合适的,我们从哪里获得指导原则来思考这些?你对照护这个问题有什么看法吗?

克劳福德:照护劳动是一个极其重要的问题,因为它在大流行期间又一次处于核心位置。当我们处于隔离状态时,讨论的方向一直是「自动化怎样才能把我们从这种处境中解救出来」,而不是去思考「互助」是否是一个好得多的框架。对我来说,这里的问题是:为什么我们总是把技术放在中心?为什么我们总假定技术要么是万灵药,要么是问题本身,而不是去问:我们想生活在怎样的世界里?有哪些方式可以增进公平与正义?然后再问技术如何服务于这一愿景,而不是由技术来驱动这一愿景。AI辩论在很多方面都令人沮丧,因为它总是假定AI真的可以应用于任何事情,总是把它当作第一选择,而不是回头看看漫长的历史:我们究竟是如何实现人们所呼吁的那类社会变革的?答案几乎从来不只是一个技术修补。所以我认为这意味着我们要寻找不同类型的工作,而这也正是你的研究如此重要的原因。

业界反应、封面艺术与AI作为交叉学科

克劳福德:我注意到聊天区里有几个问题问到这本书的封面艺术,它此刻就在我身后。这也让我想到,艺术家在这类讨论中有多重要,文化产业在讲述关于技术如何运作、如何失败的另类故事方面有多重要。回答聊天区的提问:这幅非凡的图像出自弗拉丹·约勒(Vladan Joler)之手,我曾和他合作过一个叫「一个AI系统的解剖」(Anatomy of an AI System)的项目,我们为一台亚马逊Echo音箱绘制了一幅巨型地图。这幅封面图其实是他为书中绘制的众多设计之一,它们呈现的是这样一些观念:从人的头颅所联系的智能与颅相学的观念,到我们如何表征世界,再到大规模计算系统的行星代价。我真的认为他的作品非常了不起,谢谢这个提问。

瓦伊克曼:聊天区里有一大片精彩的问题,我肯定回答不完。我看着时间,只剩两个快问快答。第一个我忍不住要问:业界对这本书的反应如何?我看到你在计算机历史博物馆和许多不同场合都做过演讲,科技圈的反应是怎样的?

克劳福德:这很有意思。我写这本书的方式,可以说像地质地层,几乎是按一个完整的计算堆栈写的:从地球开始,经过劳动、数据、分类、国家,一直到外太空。在这个意义上,科技行业内的不同群体,根据他们处理的是问题的哪一部分,会有不同的疑问。对于从事大规模数据分析和预测的群体来说,分类与偏差的问题此刻极为紧迫,人们在寻找不同的应对方式,因为我们知道,过去几年这些问题并没有消失,我们确实需要新的思路。

而在行业的另一些部分,我觉得这一点非常鼓舞人心,出现了一组新的对话,讨论如何降低机器学习方法的能源密集度。这还很新,但极其令人振奋。我们看到那里有大量的浪费可以改进,大量的计算周期可以压缩,还有一些算法技术能够大幅降低能耗。我们需要这些,因为我们正处在气候危机之中。作为一个行业,我认为正在滋生一种关切感和责任感。所以在这个意义上,我的经历相当积极,人们提出的都是很好的问题。但与此同时,我们也确实看到人脸识别、预测性警务之类的工具在不断扩张。有很多公司和行业不会觉得这本书对他们的信念有什么启发。这也正是我认为现在进行这些公共辩论如此重要的原因:这些工具正被以各种方式使用,而其中许多方式,我想我们大多数人都会深感忧虑。

瓦伊克曼:最后一个问题。我想说,像你这样的女性走在关于AI偏差的辩论最前沿,实在太好了。我想问你,为什么会这样?我刚开始研究性别与技术时,这个领域里女性肯定不多。而现在非常显眼的是,许多正在出版的优秀著作和正在进行的研究,都由女性和女性主义者引领。你觉得这是为什么?

克劳福德:这确实有意思。研究这些议题的人,比起只研究算法的人,构成了一个多样得多的联盟。这本身就说明了一些事:那些曾经被系统边缘化、排斥、误认的人,正是对这些问题最为警觉的人,他们会不断追问,如何在这类问题更深地扎根于我们的社会之前把它们改善。

我还喜欢回头看人工智能最早的年代,二十世纪五六十年代。即便在那时,你也能看到玛格丽特·米德(Margaret Mead)和格雷戈里·贝特森(Gregory Bateson)同台,实际设计系统的人和思考其社会影响的人类学家坐在一起。我们曾一度失去了这种局面,我认为我们所看到的是一种对技术的过度优先。这是我们现在必须纠正的,因为这些系统已经不再只是实验室里的设计,或者本质上属于理论性的干预,它们正以极为不同的方式影响着全世界数十亿人。所以我认为,我们必须开始把AI重新构想为一门交叉学科(interdiscipline),一门必须深深扎根于那些每天受这些工具影响的社群之中的学科。这无疑是这个领域未来十年迫切需要的。

瓦伊克曼:太好了。剩下的就是祝贺你写出这本书。这是一本了不起的书,我诚心推荐每个人都去读。它制作精美,非常易读,条理清晰,是一项出色的工作。我手边有一杯香槟,本来我们该一起去喝一杯的,很遗憾没法当面举杯,那就让我们各自举起一小杯。祝贺你,再一次,我对这本书的推荐不能更高了。干杯。

克劳福德:能和你,和今天到场的每一位一起庆祝,对我意义重大。这是一个非同寻常的群体。我真心希望我们有朝一日能当面再做一次。

瓦伊克曼:我也希望如此,但愿疫苗给力。干杯,再次感谢。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:19 开场:STS 视角与「权力登记册」 ▶ 正在看
7:15 写作缘起与 AI 的物质生命周期 ▶ 正在看
9:40 从锂矿看 AI 作为采掘业 ▶ 正在看
12:18 幽灵劳动与伪自动化 ▶ 正在看
14:50 从巴贝奇到贝索斯:算法管理 ▶ 正在看
17:15 远程办公与情绪识别监控 ▶ 正在看
19:52 训练数据的考古学 ▶ 正在看
22:34 「越多越好」:数据意识形态的来源 ▶ 正在看
26:20 数据集的来源、弃用与来生 ▶ 正在看
28:44 分类即政治:ImageNet 案例 ▶ 正在看
35:20 招聘算法、多样性与设计的边界 ▶ 正在看
39:06 创新观的失衡:技术 vs 政策 ▶ 正在看
41:36 伦理准则不够,要谈权力 ▶ 正在看
45:23 参与式监管与拒绝的政治 ▶ 正在看
48:58 业界反应、封面艺术与 AI 作为交叉学科 ▶ 正在看
本期小问 · 档案清单
28:44 AI 系统显得理性客观,我们凭什么就信它的结果? ▶ 正在看
35:20 为什么说靠调整数据追求公平,没有抓住真正的问题? ▶ 正在看
14:50 算法管理追踪到工人的肌肉与韧带,这还是效率问题吗? ▶ 正在看
9:40 一本讲人工智能的书,为什么要从一座锂矿写起? ▶ 正在看
本期讲者
凯特·克劳福德澳大利亚学者,微软研究院高级首席研究员、南加州大学安纳伯格学院研究教授,AI Now 研究所联合创始人。著有《Atlas of AI》(2021),与 Vladan Joler 合作的《Anatomy of an AI System》被 MoMA 收藏。
朱迪·瓦克曼伦敦政治经济学院社会学教授,艾伦·图灵研究所「数据科学与 AI 领域的女性」项目负责人。科学与技术研究(STS)及技术女性主义的代表人物,著有《The Social Shaping of Technology》《Pressed for Time》。
01开场:STS 视角与「权力登记册」
0:19
hundreds of people are here today which is wonderful and attest to how well known kate is um as a scholar of the social and political implications of artificial intelligence kate is a senior principal researcher at microsoft and recently became a research professor at the university of california's annenberg school my name is judy wiseman and i lead the women in data science and ai project at the alan turing institute and we're part of the public policy group our project highlights the lack of diversity in the ai industry but also crucially makes links between this and the gender race and other social biases that kate's book so squarely addresses we only have an hour with kate today um i wish we had many hours and i hope we'll get a chance to do this in person on another occasion um but everyone is very welcome to participate through the chat function um please post your questions in the q a and i'll keep an eye on them and i'll try and integrate as many as i can as we go along today today's event will be subtitled if you can't see the subtitles hit the cc button on the bottom toolbar of your screen okay um i know not many of you will have had a chance to read kate's book because it's only literally just been launched
今天有好几百人来参加,这真是太棒了,也说明凯特作为研究人工智能社会与政治影响的学者有多广为人知。凯特是微软的高级首席研究员,最近还成为了加州大学安纳伯格传播学院的研究教授。我叫朱迪·瓦克曼,我在艾伦·图灵研究所负责“数据科学与人工智能领域的女性”项目,我们隶属于公共政策组。我们的项目关注人工智能行业缺乏多样性的问题,但更关键的是,把这个问题与性别、种族以及其他社会偏见联系起来,而这正是凯特这本书正面处理的议题。今天我们和凯特只有一个小时,我真希望能有好几个小时,也希望以后有机会能当面做这场对谈。欢迎大家通过聊天功能参与,请把问题发到问答区,我会留意,并尽量在过程中把更多问题融进来。今天的活动有字幕,如果看不到字幕,请点击屏幕底部工具栏的 CC 按钮。好,我知道你们当中没几个人有机会读过凯特这本书,因为它真的是刚刚才出版
便签笔记
1:45
it occurred to me you know that i think it's just literally come out in the uk in the last couple of days actually um so i thought i would um start by just very quickly telling you why i i so welcome the book i have to say that i nodded all the way through you know as i do um because it's such a rich illustration really of taking a science and technology studies perspective to look at a new new developments in technology and it's done with a suitably kind of marxist inflection so you know for me it's the kind of perfect um lens uh with which to think about these things okay so very quickly for a start what i love about the book is that it's historical in every chapter you like you locate technological developments in the historical context of their emergence and show how these origins are embedded in the technologies themselves whether it's facial recognition or emotion recognition that you rightly stress throughout the book that we need to understand where these systems come from to understand how they function so i'm totally sort of with you on that i mean history as well taking historical approach really reminds us that we've had a lot of these discussions uh in earlier with earlier technological revolutions and you know i
我突然想到,这本书好像就是最近这几天才在英国上市的。所以我想先非常简短地讲讲,为什么我这么喜欢这本书。我得说,我从头到尾一直在点头,我一贯如此,因为它非常丰富地示范了如何用科学与技术研究(STS)的视角去审视技术上的新发展,而且带着恰到好处的马克思主义色彩。所以对我来说,这是思考这些问题的绝佳视角。好,先快速说一点:我特别喜欢这本书的地方在于它是历史性的。你在每一章里都把技术发展放回它出现的历史脉络里,展示这些源头是如何嵌入技术本身的,不管是人脸识别还是情绪识别。你在全书中都非常正确地强调:我们必须理解这些系统从哪里来,才能理解它们是怎么运作的。所以我完全认同这一点。而且采取历史的视角也提醒我们,很多这样的讨论在更早的技术革命中我们就已经有过了
便签笔记
3:10
i came into this work during the microelectronic revolution of the 1980s and i remember well then the debates about the impact of technology um being of this simple dichotomous debate about whether they're um you know utopian or dystopian we had a lot of that discussion and in a way the field of science and techno technology studies was really trying to address the technological determinism that sort of was underlying uh those debates to very much emphasize that technology isn't autonomous but is always shaped by social political and economic forces and as you make so clear kate human choices are embedded in the very design or architecture i mean as we say nowadays the materiality of technology whether it's hardware or software you know langdon winners bridges or ai systems secondly um i i was very intrigued and absolutely um persuaded by you adopting an atlas or map as your lens i have to think about that a lot because i'm much better on time than on space you know but i actually think it's it's it's terrific um lens because it enables you to extend the notion of infrastructure to include the planetary infrastructure the dual operation of abstraction and extraction in information capitalism as hart
我是在上世纪八十年代微电子革命的时候进入这个领域的,我清楚地记得当时关于技术影响的争论,就是那种简单的二元对立:到底是乌托邦还是反乌托邦。我们那时讨论了很多这类问题。某种意义上,科学与技术研究这个领域当时正是想去回应那些争论背后的技术决定论,强调技术并不是自主的,它总是被社会、政治和经济力量所塑造。凯特,你说得非常清楚:人的选择就嵌在设计或者架构里,也就是我们今天说的技术的物质性,无论是硬件还是软件,无论是兰登·温纳笔下的那些桥,还是人工智能系统。其次,我非常好奇,也完全被你说服了——你用“图集”或者说地图作为切入视角。我得好好想想这一点,因为我对时间比对空间在行得多。但我真的觉得这是个极好的视角,因为它让你能把“基础设施”的概念扩展到行星尺度的基础设施,扩展到信息资本主义中抽象与榨取的双重运作,就像哈特和
便签笔记
4:34
negri uh put it that includes the cost of natural resources and its labor process so in the fine marxist tradition of exposing commodity fetishism you expose the material conditions of the production of ai its whole life trajectory the birth life and death of an ai system from minerals and energy by a global supply chain of labor to the way corporations on the state operationalize it and its impact on our lives like all technologies ai is not artificial it's very much man-made if i can say so um thirdly i very much like the way you argue that call ai intelligent and i think this is crucial to the whole book really is to invoke a very narrow definition of intelligence one that flattens the complexity of lived experience ai systems operate according to the normative logic of mathematical optimization across a variety of domains such as education health welfare credit criminal justice and the result is that only what can be measured counts its governance through numbers and your key point really is that
奈格里所说的那样,这里面包含了自然资源的成本和它的劳动过程。所以,秉承揭露商品拜物教的优良马克思主义传统,你揭示了人工智能生产的物质条件,它的整个生命轨迹——一个人工智能系统的诞生、生存与死亡:从矿物和能源,经由全球劳动供应链,到企业和国家如何把它运作起来,以及它对我们生活的影响。跟所有技术一样,人工智能并不“人工”,它非常实在地是人造的,如果我可以这么说的话。第三,我非常喜欢你的这个论证:把人工智能称作“智能”,我觉得这是全书的核心,其实是在援引一种非常狭隘的智能定义,一种把鲜活经验的复杂性压平的定义。人工智能系统按照数学优化的规范性逻辑在各个领域运作,比如教育、健康、福利、信贷、刑事司法,结果就是只有可测量的东西才算数,这是一种通过数字进行的治理。而你的关键观点其实是
便签笔记
5:52
while these ai systems appear as rational and objective and in a way that legitimates their use that's the crucial legitimation of their use they're ultimately designed to serve existing dominant interests in this sense you say ai is a registry of power what the book does superbly i mean superbly is i'm completely with you on all these arguments so you know um is not to argue that these trends are kind of uh completely new but rather that the massive accumulation of behavioral trace data has fundamentally intensified and amplified these trends so i'd like to begin just by asking you when you became aware of these massive changes because you know you were there right at the beginning and you sussed out that this was one of the most crucial um issues facing us now and you know how did you know that this was like totally the perfect moment to bring out this book it's like we've been wanting someone to bring this stuff together a lot of us who bring research in this area to think of thinking about it and you you've written this beautiful book that kind of absolutely brings it together in an accessible such an accessible clear and convincing way so tell us about your journey well judy i have to say that is the most generous and overwhelmingly uh just wonderful introduction
虽然这些人工智能系统看起来是理性、客观的,这种表象在某种程度上让它们的使用获得了正当性——这正是关键的正当化机制——但它们归根结底是被设计来服务既有的支配性利益的。从这个意义上说,你说人工智能是一份权力的登记册。这本书做得特别出色的地方在于,我完全认同你所有这些论证,它并不是主张这些趋势完全是新的,而是说行为痕迹数据的大规模积累从根本上强化和放大了这些趋势。所以我想先问你:你是什么时候意识到这些巨大变化的?因为你从一开始就在场,而且你看穿了这是我们当下面临的最关键的议题之一。你怎么知道现在正是推出这本书的完美时机?我们很多在这个领域做研究、思考这些问题的人,一直都盼着有人能把这些东西整合起来,而你写出了这本漂亮的书,用如此易懂、清晰、有说服力的方式把它们串到了一起。跟我们讲讲你的历程吧。朱迪,我得说,这真是我听过最慷慨、最让人受宠若惊的精彩介绍
便签笔记
02写作缘起与 AI 的物质生命周期
7:15
thank you i mean particularly because i've been such a fan of your work for so long since way back when you've when you first probably think social shaping of technology was one of the books that was like prominent on my bookshelf in the 90s and of course everything since including press for time which i write about in atlas of ai so it means a great deal to me to be here in conversation with you today um i'd also like to say a big thank you to the alan turing institute for hosting us but you know you ask this question of you know of timing how can i know that this would be the right year 2021 we've come through a pandemic we've seen the intensification of all of these logics that you've beautifully articulated you know i wish i could say i did the truth is that this book took a long time to write over five years um and there's no way that i could have imagined that the themes that i saw then would have become so present um really as quickly as they have for me it was really the great fortune of joining microsoft research in 2012 which is around the same time as convolutional neural nets were taking central stage deep learning was accelerating and certainly we were seeing industrial research labs really invest in machine
谢谢你。特别是因为我一直是你作品的忠实读者,很久以前就是了,《技术的社会塑造》大概是九十年代摆在我书架上最显眼的书之一,当然还有之后的所有作品,包括《为时间所迫》(Pressed for Time),我在《人工智能地图集》里也写到了它。所以今天能和你对谈,对我意义重大。我也要特别感谢艾伦·图灵研究所主办这次活动。不过你问到时机的问题——我怎么会知道 2021 年是对的年份呢?我们经历了一场大流行,看到了你刚才精彩阐述的那些逻辑的全面强化。我真希望我能说这是我算准的,但事实是这本书写了很久,超过五年。我当时看到的那些主题,我完全没法想象它们会这么快就变得如此切近。对我来说,很幸运的是我在 2012 年加入了微软研究院,那差不多正是卷积神经网络登上中心舞台的时候,深度学习在加速,我们确实看到工业界的研究实验室以一种全新的方式投入到机器
便签笔记
8:23
learning in a new way and it was really clear at that point that everything from computer vision to natural language processing was going to be radically changed but at the same time there were very real problems and we can think about these problems as a life cycle we can think about the ways in which ai systems require so much computation in fact they're becoming increasingly computationally intensive and so many of our consumer ai devices be it phones or laptops use enormous amounts of these minerals that everything from rare earth to cobalt to lithium and then of course the data that drives ai must be labeled by people on remote work platforms in places like ghana and the philippines and then finally they get applied in these highly tailorized work environments where algorithmic management systems are brought to bear on workers every day so you can see this full tangible system of ai having a much more profound and material impact in the world than simply thinking about it as abstract computation or immaterial algorithms so part of really what i wanted to do in this book was to situate these systems in sort of the specific locations and institutions around the world that are actually deploying them very
学习中。当时就很清楚,从计算机视觉到自然语言处理,一切都将被彻底改变。但与此同时也存在非常真实的问题,我们可以把这些问题看成一个生命周期:人工智能系统需要巨量的算力,事实上它们的算力消耗还在不断加大;我们那么多的消费级人工智能设备,无论是手机还是笔记本电脑,都要用掉大量矿物,从稀土到钴到锂;然后当然,驱动人工智能的数据必须由人来标注,这些人在加纳、菲律宾这样的地方的远程众包平台上工作;最后,这些系统被应用到高度泰勒化的工作环境中,算法管理系统每天都作用在工人身上。所以你能看到人工智能这一整套实实在在的系统,它对世界的影响比我们把它仅仅当作抽象计算或非物质算法要深刻和物质得多。所以我在这本书里想做的一部分工作,就是把这些系统放回到世界各地那些真正在部署它们的具体地点和机构里
便签笔记
03从锂矿看 AI 作为采掘业
9:40
nice and so i think it's it's most appropriate that you start the book by telling us about your sort of journey and describe your journey um to visit a lithium mine in um nevada where you can say you can see ai as an extractive industry in its most literal sense i mean i must say as we're both australians we both understand how important mining is politics you know that it's to put it mildly it was very nice you start with the mining but you know i i wanted to ask you um to tell us about that but also you know what you think about the fact that in a way during this covert crisis we've been kind of reassuring ourselves that actually our carbon footprint is going down and yet you very much stress that all this kind of zooming that we're doing like that is taking up an enormous amount of energy exactly right i mean you know it's interesting for me to one of the things that people have asked it's like why did you open a book about artificial intelligence at a lithium mine and it was precisely because you know in order to ground artificial intelligence in order to truly understand these material impacts we have to go to the places where it's made and i mean made in the fullest sense so that for me meant going to the last functioning
很好。所以我觉得非常恰当的是,你在书的开头讲述了你的旅程,描述了你去内华达州一座锂矿的经历,在那里你可以看到人工智能在最字面意义上是一个采掘业。我得说,我们俩都是澳大利亚人,我们都明白采矿在政治上有多重要,说得客气一点。所以你从采矿讲起,我觉得非常好。我想请你讲讲那段经历,同时也想听听你怎么看:某种程度上在这场新冠危机期间,我们一直在自我安慰说我们的碳足迹在下降,但你非常强调,我们现在这么多的视频会议其实消耗了巨量的能源。完全正确。有意思的是,人们常问我:你为什么要用一座锂矿来开启一本关于人工智能的书?正是因为要让人工智能落地,要真正理解这些物质影响,我们就必须去到它被制造出来的地方,而且是最完整意义上的“制造”。对我来说,这意味着要去美国最后一座还在运转的
便签笔记
10:57
lithium mine in the u.s which is in nevada in a place called silver peak where you can see these gigantic lithium-brine pools and some iridescent green color sort of drying in the sun producing uh essentially the the stuff the gray gold as it's referred to that becomes lithium-ion batteries and those lithium-ion batteries are of course in high demand you know we have them in iphones but also in you know tesla cars and evs but we have a very serious supply problem in fact i'm sure as you saw the biden administration just recently released a crisis document saying we have to essentially secure the supply chain for so many of the minerals that drive current information capitalism and planetary computation so it's very timely right now to think about just how much of this sort of material is required to make the sorts of ai systems that we just expect to work every day are actually poised on this brink in terms of how do we think about the usability of these minerals but also their geopolitics yeah okay now the next chapter i'm going to try and get through as much as i can i mean chapter three is on uh the many different forms of labor that are involved in making ai um from minors to content moderators to amazon warehouse workers to engineers in
锂矿,它在内华达州一个叫银峰(Silver Peak)的地方。在那里你能看到巨大的锂盐卤水池,泛着某种虹彩般的绿色,在阳光下慢慢蒸干,产出的实际上就是人们所说的“灰色黄金”,最终变成锂离子电池。而这些锂离子电池当然需求极大,我们的 iPhone 里有,特斯拉和电动车里也有。但我们的供应问题非常严重,我相信你也看到了,拜登政府最近刚发布了一份危机文件,说我们必须确保这些矿物的供应链安全,而正是这些矿物在驱动当下的信息资本主义和行星尺度的计算。所以现在正是时候去思考:我们每天理所当然地指望它能用的那些人工智能系统,究竟需要多少这样的物质,而这些矿物在可用性上、在地缘政治上其实都处在某种危险的临界点上。好,接下来一章——我尽量多讲一些——第三章讲的是制造人工智能所涉及的各种形式的劳动,从矿工到内容审核员,到亚马逊仓库工人,再到硅
便签笔记
04幽灵劳动与伪自动化
12:18
silicon valley and i love this chapter as well i'm gonna say this with every chapter i mean i found it really refreshing because you know i've spent a lot of my time in the last few years on these panels about the future of work i was on one just actually last week um with the guy from mit you know david otter talking about you know technology in the future of work and and you know so often economists not him he's the exception actually you know tell me that robots will soon do everything and there won't be any kind of jobs left at all and as you quite rightly say and he and i say you know what's what's you know that that debate is off the point in a way and what's more important is to look at the experience of the work now and how these technologies are kind of facilitating increased surveillance and algorithmic management systems and i want to come on to that in a minute but i wonder if you might say something quickly about the importance of ghost work um as gray and and and sorry caller because it seems to me that's an important thing to to highlight as well i'm so glad you mentioned the work of mary gray and sid siri and and and certainly ghost work i think is an incredibly important book at illuminating what it's like to you know be an a person working on a
谷的工程师。我也非常喜欢这一章,我每一章大概都会这么说。我觉得它特别令人耳目一新,因为过去几年我花了很多时间参加各种关于“工作的未来”的圆桌讨论,上周我还刚参加了一场,跟麻省理工的那位大卫·奥托一起,谈技术与工作的未来。经常有经济学家——他不算,他是例外——跟我说机器人很快就会做所有事,根本不会再有工作留下来。而正如你说得非常对的,我和他也这么说,那种争论其实是抓错了重点。更重要的是去看当下的劳动体验,去看这些技术是如何助长了更严密的监控和算法管理系统的。我等下想聊这个,但我想先请你简短说说“幽灵劳动”(ghost work)的重要性,就是格雷和苏里说的那个,因为在我看来这也是很值得强调的一点。我很高兴你提到了玛丽·格雷和西德·苏里的工作。《幽灵劳动》这本书我认为极其重要,它照亮了一个在远程众包平台上工作的人
便签笔记
13:32
remote work platform i mean these are jobs which are really digital piece work you're being paid very little and these are sort of you know sub poverty level wages for doing work that can be extremely tedious um and extremely stressful so certainly you know we've seen you know many scholars from lilliarani to uh i'm thinking here also of the wonderful work of astra taylor she uses the term photomation the the way in which we assume automation but in actual fact it's people propping up these systems and this is something that has been emerging for some time and certainly when it comes to ai um in your work on digital assistance you know i know you saw these these same patterns emerging and how much of this system is actually faked by using people all along the supply chain actually doing these quite mundane tasks and in the book i write about a system called xai where they actually had people pretending to be digital assistants um doing you know 14 hour days you know that sort of it's really quite painful work conditions but again to give the impression of seamless ai so you know this is something that again i think your work has been pointing to for some time but as well and truly in the ascendant now yes and we could um you know talk about the kind of socio-imaginary of silicon
究竟过着什么样的生活。这些工作本质上就是数字计件工,报酬极低,属于贫困线以下的工资水平,而工作内容可能极其枯燥,也极其让人紧绷。当然,我们看到从莉莉·伊拉尼到许多学者都在研究这个,我这里还想到阿斯特拉·泰勒的精彩工作,她用了“伪自动化”(fauxtomation)这个词,指的是我们以为是自动化,但实际上是人在支撑着这些系统。这个现象已经出现有一段时间了。当然,说到人工智能,在你关于数字助理的研究里,我知道你也看到了同样的模式:这套系统有多少其实是靠人假装出来的,供应链上一路都有人在做这些相当琐碎的任务。我在书里写到一个叫 x.ai 的系统,他们真的让人去假扮数字助理,一天工作 14 个小时,工作条件相当痛苦,而这一切都是为了营造出无缝人工智能的印象。所以我觉得这也是你的研究早就指出过的东西,而现在它是真真切切地大行其道了。是的,我们其实可以聊很久硅
便签笔记
05从巴贝奇到贝索斯:算法管理
14:50
valley and those issues for a long time and the role these things play but i want to um move on you know to a bit of detail about this chapter i mean you begin the chapter um with the image of a large time clock in an in in an amazon warehouse that i know you visited and you use that image to trace the long history of the use of automation to control time um you know through authors like ep thompson and braverman and through taylorism scientific management forties and you know this is all my my material you know the stuff of my life if you like um i wondered if you could tell us about um a little bit more about the use of algorithmic management systems at both ends of the labor market on the one hand in terms of kind of platform work and at the other end which is what i've been studying um you know about knowledge work and how that's being um transformed right and i love the way you've asked this question because of course we know we have to think about algorithmic management both in terms of the way that it's been introduced in low wage work and what's traditionally understood as blue collar jobs all the way through to highly paid white collar employment and and this is interesting the degree to which you know you can certainly
谷的那套社会想象,以及这些东西扮演的角色,但我想继续往下讲这一章的一些细节。你在这一章开头用了一个画面:亚马逊仓库里的一个巨大计时钟,我知道你去参观过。你用这个画面追溯了用自动化来控制时间的漫长历史,经由 E.P. 汤普森、布雷弗曼这些作者,经由泰勒主义、科学管理。这些都是我的老本行,可以说是我的人生素材了。我想请你再多讲一点算法管理系统在劳动力市场两端的使用:一端是平台劳动,另一端是我一直在研究的知识工作,以及它是如何被改造的。对,我很喜欢你这样提问,因为我们当然必须同时从两方面来看算法管理:它如何被引入低薪工作、也就是传统上所说的蓝领岗位,以及它如何一路延伸到高薪的白领雇佣。有意思的是,你确实可以把
便签笔记
16:04
trace these systems back i mean charles babbage of course who is you know known best for the difference engine used to also write a lot of social theory and you know one of his visions for the future was that you know we would have systems that would be highly efficient that would be watching workers and tracking every minute of what they do in order to make sure they're being used most efficiently and it's it's a somewhat horrific vision but it's certainly come to pass in some of the systems that we're seeing today spending time inside an amazon fulfillment center was for me just you know absolutely horrifying i mean these are sort of used as the exemplars of workplaces that combine automation both the algorithmic variety and robots with people yeah but they're also exemplars of something else of precisely the kind of extraordinary psychological and physical pressure of working under those conditions and i'm sure you saw just last month jeff bezos said that one of their responses to increased injury and stress in these work environments is to induce introduce a new algorithmic management system that will be tracking workers down to the level of the ligaments and muscles that they're using to try and produce new efficiencies now
这些系统一路追溯回去。查尔斯·巴贝奇当然最为人所知的是差分机,但他也写了大量社会理论。他对未来的设想之一就是:我们会拥有高效率的系统,它们会盯着工人,追踪他们所做的每一分钟,以确保他们被最高效地使用。这是个有点可怕的设想,但在我们今天看到的一些系统里,它确实已经成真了。对我来说,待在亚马逊配送中心里的那段时间简直触目惊心。这类地方常被当作典范,展示如何把自动化——既包括算法层面的,也包括机器人——和人结合起来。是的,但它们也是另一种东西的典范:正是那种在这样的条件下工作所承受的巨大心理和身体压力。我相信你也看到了,就在上个月,杰夫·贝索斯说他们应对这类工作环境中工伤和压力上升的办法之一,是引入一套新的算法管理系统,把对工人的追踪细化到他们正在使用的韧带和肌肉,以此挖掘新的效率。
便签笔记
06远程办公与情绪识别监控
17:15
to me this is this is babbage's vision brought to life and it's quite nightmarish yeah yeah yeah well it's so like taylorism as well isn't it and the stopwatch and um and and what about the other end in terms of kind of knowledge workers because i i'm with you on [Music] you know the the the fact that we need to really kind of focus on you know the the change in a way is the increased granular granularity of kind of tracking and nudging and assessment that the machine um tools machine learning tools facilitate and you know i'm i'm kind of myself very worried about this in the context of the shift to working from home right but given that won't that act as a spur to increasing kind of surveillance of knowledge workers through these tools and i mean the the worst thing you describe very well in your chapter on emotion detection software would be to that extent so as you say it would be bodily movements but emotional states and all of those things do you have a sense as well that this current kind of moment of the covet crisis will be a spur to increasing that kind of surveillance of knowledge workers it's already happened so certainly in the last in the last 18 months we've seen an enormous uptick in services that
对我来说,这就是巴贝奇的设想成真了,而且相当噩梦。是啊是啊,这也太像泰勒主义了,就像那只秒表。那另一端呢,知识工作者那边?因为我很认同你的看法,我们真正需要关注的变化,其实是机器学习工具所带来的那种更加细颗粒度的追踪、助推和评估。而我自己在“居家办公”转向的背景下对此相当担忧。既然大家都在家办公,这难道不会成为通过这些工具加强对知识工作者监控的推手吗?而且你在关于情绪识别软件那一章里描述得非常好,最糟糕的情况会到那种程度:不只是身体动作,还有情绪状态之类的一切。你是不是也觉得,当下这场新冠危机会推动对知识工作者的这类监控进一步加强?这已经发生了。在过去 18 个月里,我们确实看到这类服务大幅增加:
便签笔记
18:37
offer to essentially track workers through cameras through uh tracking how many times they send emails uh sit in meetings and to come up with efficiency scores that they use to compare to other remote workers and certainly the the many startups and companies that offer this have seen an enormous increase in their services similarly we've seen other companies like four little trees offer to do essentially emotion detection as it's called we call detection in scare quotes here of of young students who are working again studying at home trying to engage with classwork and using cameras to detect micro expressions in their faces to see if they're paying attention and to try and sort of if you will in infer a type of internal emotional state now certainly something that i spent quite a lot of time studying for the purposes of this book and have a chapter on it is looking at just how problematic this assumption is that you can look at somebody's face and know their internal state this desire to create an ai polygraph if you will i think is fundamentally broken at its very tap root in terms of the ideas and theories that inform it and in terms of the systems that seek to implement it
它们提供的服务本质上就是通过摄像头追踪员工,追踪他们发了多少封邮件、开了多少会,然后算出一个效率分数,用来和其他远程工作者做比较。提供这类服务的许多初创公司和企业,业务量都有巨大增长。类似地,我们也看到像 Four Little Trees 这样的公司,提供所谓的情绪“检测”——这里的“检测”要打上引号——对象是同样在家学习、努力投入课业的学生,用摄像头捕捉他们脸上的微表情,判断他们有没有在专心,并试图借此推断某种内在的情绪状态。为了这本书,我在这个问题上花了相当多时间研究,也写了一章,就是去看这个假设有多成问题:你能通过看一个人的脸就知道他的内在状态。这种想造出一台“人工智能测谎仪”的欲望,我认为从最根本的地方就是坏掉的——无论是支撑它的那些观念和理论,还是那些试图实现它的系统。
便签笔记
07训练数据的考古学
19:52
so absolutely this is one of the domains where i think we urgently need regulation because certainly under the pandemic we've seen these systems move ever deeper into the tools of everyday work yes i mean maybe we can come back to the regulation because i do want to ask you about that but i want to quickly move on to your terrific chapters on um data and classification because here's really where your own work has been absolutely sort of path breaking so i'm i'm i'm cognizant of the time as the author pressed the time um i really want to get on to this but let me just um start with data actually um you look at them you know you talk about the massive harvesting of data that's really driving the success of ai can you tell us and um we've got a very varied audience um here of so many people you know how do training data sets actually work and perhaps tell us um about the problems of labeling encountered by the army of piecemeal workers who do this work well you know put very briefly sort of supervised and unsupervised machine learning systems are built by using large amounts of data to essentially represent the world we call this training data and training data in the sense becomes
所以这绝对是我认为迫切需要监管的领域之一,因为在疫情期间,我们看到这些系统越来越深地进入日常工作的工具之中。是的,监管的话题我们也许可以待会儿再回来说,我确实想问你这个,但我想先快点讲到你关于数据和分类的那两章精彩内容,因为你自己的研究在这方面绝对是开创性的。我很在意时间,就像那本《为时间所迫》说的一样,但我真的很想聊到这里。先从数据说起吧。你谈到了大规模的数据采集,正是它在驱动人工智能的成功。今天我们的听众背景非常多元,能不能请你讲讲训练数据集究竟是怎么运作的?也许再说说那支做数据标注的计件工大军所遇到的标注问题。非常简要地说,有监督和无监督的机器学习系统,都是通过使用大量数据来表征世界而构建起来的,我们把这些数据叫做训练数据,而训练数据在某种意义上就成了
便签笔记
21:10
sort of the ground truth of how systems will actually function so for me one of the things that's been so revealing and interesting is to really study to really look at how training data works to open up if you will the training data sets and see not just the data within them but the logic in terms of how it's classified and ordered and and here i think was just such a revelation to me was to to understand that this sort of excavation as we call it is so rare that that ultimately for you know for so many training data sets they're simply used as aggregate infrastructures you know applied don't think too much about it um and use it to create models for machine learning but here in lies a very real problem because once you actually start looking closely at how these systems work once you start conducting these sort of archaeologies of training data you find all sorts of problems not just in terms of the ways in which people are being classified into you know binary gender or you know five part categorizations of race you know things that are profoundly broken and understood as such but also making moral judgments of people's characters based on their appearance i mean there are so many problems with this this type of um almost sort of will to
系统如何运作的“基准真值”。所以对我来说,特别有启发也特别有意思的一件事,就是真正去研究训练数据是怎么运作的,去把训练数据集打开来看,不只看里面的数据,还要看它背后的逻辑——数据是怎么被分类和排序的。这里让我大为震惊的是,我意识到这种我们所说的“考古挖掘”非常罕见。对那么多训练数据集来说,它们只是被当作聚合的基础设施来使用,拿来就用,不怎么细想,然后用它来训练机器学习模型。但问题就出在这里,因为一旦你真的开始仔细看这些系统是怎么运作的,一旦你开始对训练数据做这种考古学式的研究,你就会发现各种各样的问题:不只是人被分类进二元性别、或者五分法的种族类别这类从根本上就站不住脚、而且大家也都知道站不住脚的做法,还包括根据外貌对一个人的品格做道德判断。今天的系统里存在着太多这类问题,几乎可以说是一种
便签笔记
08「越多越好」:数据意识形态的来源
22:34
phrenology that we see in systems today that certainly it's it's at the level of training data where you can start to see those classifications and where they come from okay yes absolutely um you you talk about in this chapter and i was very struck by this that data is now treated as a natural resource to be mined and extracted and i like the mind and extracted um you know from the mining from the from the literal mining to the to the data being treated in this way and you talk about how there's a kind of moral imperative i mean as as marion futada and karen healey put it to collect more data because we can you know it's there for the collecting so it's kind of a morally imperative to keep doing this and how this is related to the idea that the more data we have the more knowable the person that we have now um a kind of data double or as luke stark calls it a scalable um subject being created of all of us i wonder if you could explain what you mean by that elaborate on that for us well i mean i think that's that's a you know interestingly a formulation that that's been a long time coming so this sort of ideology that more is better is something that you can see really going back to sort of the late 20th century in sort of the very early
颅相学的冲动。而正是在训练数据这个层面上,你才能开始看到那些分类,以及它们从哪里来。对,完全同意。你在这一章里谈到一点,让我印象很深:数据现在被当作一种可供开采和榨取的自然资源。我喜欢“开采和榨取”这个说法,从字面意义上的采矿,一路到数据被这样对待。你还谈到存在一种道德律令,就像马里恩·富卡德和基兰·希利说的那样,要收集更多数据,因为我们能收集——它就在那儿等着被收集,所以不停地收集几乎成了一种道德义务。这又跟另一个想法相关:我们拥有的数据越多,那个人就越可知,于是就产生了某种“数据分身”,或者像卢克·斯塔克说的“可缩放的主体”,我们每个人都被这样制造出来。能不能请你解释一下你说的是什么意思,展开讲讲?我觉得有意思的是,这个说法其实酝酿了很久。“越多越好”的这种意识形态,可以一直追溯到二十世纪晚期最早的那些
便签笔记
23:55
ai labs and this is something that the media historian chao cheng li has documented beautifully in her work about the ibm continuous speech recognition lab where robert mercer this is yes the same robert mercy you know who funded the trump campaign and you know many other things besides um back when he was actually a research scientist said that more data is always better data and indeed so what we see at that time in history is a shift away from the sorts of expert systems approaches and symbolic logic that were previously sort of dominant in ai towards sort of probabilistic methods and brute force approaches so certainly from that time to now you've seen this development of this ideology that we should always capture as much data as possible even if you don't know what it's for or you it's actually presenting forms of liability to be holding this data that it was going to have this value this kind of this this value in its form even if you couldn't quite see it yet and i think what we have to see now is that this is actually a really problematic approach and it's produced this kind of over collection of data it's produced a situation where again people are profoundly concerned about the way in which data sets are being combined to produce
人工智能实验室。媒体史学家 Xiaochang Li 在她关于 IBM 连续语音识别实验室的研究中对此做了非常漂亮的记录。在那里,罗伯特·默瑟——对,就是那个后来资助特朗普竞选、还做了很多别的事情的罗伯特·默瑟——当年他还是研究科学家的时候就说过:更多的数据永远是更好的数据。确实,我们在历史的那个时刻看到的是一次转向:从之前在人工智能中占主导地位的专家系统方法和符号逻辑,转向概率方法和暴力计算的路径。所以从那时到现在,你看到这种意识形态不断发展:我们应该总是尽可能多地采集数据,哪怕你不知道它有什么用,哪怕持有这些数据其实带来了各种责任风险;人们相信它总会有价值,某种内在于其形态之中的价值,即使你现在还看不出来。而我认为我们现在必须看清,这其实是一种非常成问题的做法,它导致了数据的过度采集,也造成了这样一种局面:人们深切担忧数据集被相互组合,从而产生
便签笔记
25:10
what we've called in the past predictive privacy harms in a paper with jason shorts so certainly i think here this this way in which data has been figured and understood has actually led us to a very problematic point in computational history yes and i also think that you know what's interesting is how this kind of accumulation of data is affecting our own subjectivity i mean i know this is a topic for another time that you know the the self-optimization that somehow built into you know the accumulation of data of the self and you've written very nicely sort of about that and i love your paper actually i i can't remember it now but i do use it for teaching where you start with um weight scales and go through talking about you know the accumulation of information and how that actually affects your own sense of embodiment and and how you're functioning in the world um but maybe we should move on i i can see there's lots of questions as well so i did um want to get on to privacy issues of which there's sort of so many and of course one of the things that people um are very concerned about when they talk about data doubles is of course that we're never sort of consulted about the use of our data and how it circulates and how it's extracted and i
我们过去和杰森·舒尔茨在一篇论文里所称的“预测性隐私伤害”。所以我确实认为,数据被这样构想和理解的方式,把我们带到了计算史上一个非常成问题的位置。是的,我还觉得有意思的是,这种数据的积累是如何影响我们自身的主体性的。我知道这是另一个话题了,就是那种内嵌在自我数据积累之中的自我优化。你对此写得非常好,我很喜欢你那篇论文,我现在想不起标题了,但我上课会用它。你在文章开头讲体重秤,然后一路谈到信息的积累,以及这如何影响你自己的具身感受、影响你在世界中的运作方式。不过也许我们该往下走了,我看到还有很多问题。我确实想聊聊隐私问题,这方面的问题实在太多了。当然,人们谈到“数据分身”时特别担心的一点是,我们的数据被怎么使用、怎么流通、怎么被榨取,从来没有人征询过我们的意见。我
便签笔记
09数据集的来源、弃用与来生
26:20
remember morizov some years ago talking it trying to um campaign for us being paid for our data and kind of all of those discussions but perhaps um as we're maturing what might be more interesting is how you talk about the problem that there's a lack of standard practices to note where data comes from and how that's particularly a problem um given that many of the data given the private ownership of many data sets you know that that when there's public data sets i mean you say you've looked at hundreds of data sets and they must be in the public realm i mean what could we do about this business you know about the fact that most of these data sets are actually not in the public realm and therefore not accessible to people like you examining them i mean this is a very real problem and and honestly i mean i can still remember sort of early conversations a few years ago with timnit gabriel who at that time uh was a postdoc with us at msr and we were marveling at the fact that there really is no sort of standard way in which we have any information about where a training data set came from so you know the way in which you would have a data sheet for a piece of hardware that says you know you can only use a semiconductor at these
记得莫罗佐夫几年前还试图倡导我们应该为自己的数据获得报酬之类的讨论。但随着我们对这个问题的认识逐渐成熟,也许更有意思的是你谈到的另一个问题:我们缺乏标准做法来记录数据从哪里来。考虑到很多数据集是私有的,这个问题就尤其严重。你说你研究过成百上千个数据集,那些必然是公开领域里的数据集。对于大多数数据集其实并不在公开领域、因此像你这样的人无法接触和检视,我们能做些什么呢?这确实是个非常现实的问题。老实说,我还记得几年前和蒂姆尼特·格布鲁的早期对话,她当时是我们微软研究院的博士后。我们当时就很惊讶:竟然完全没有一种标准方式,能让我们了解一个训练数据集从何而来。你想想,一件硬件会有数据表,会写明这块半导体只能在某某
便签笔记
27:32
temperatures you know in these sorts of conditions there was nothing equivalent for how we use data and that was the genesis of a paper that's called data sheets for data sets which is attempting to articulate what are the sorts of criteria information history and provenance that we would need to understand what a data set is actually designed for and more importantly when it should no longer be used and that is another thing that we really don't talk about enough is you know when these data sets should be deprecated and then what happens when they continue to have these afterlives in places like academic torrents long after they've been taken down and we've seen that happening just in the last couple of years where training data sets were seen to be problematic and the creators said yes we'll remove them but they continue to circulate and of course continue to inform many production level systems so that's problem number one but problem number two that you point to is the fact that we simply don't know how large or what's happening in terms of the construction of training data sets inside sort of the large technology companies these are things which are kept as very secret proprietary information
温度、某某条件下使用,可数据的使用却完全没有对应的东西。这就是那篇叫《数据集的数据表》(Datasheets for Datasets)的论文的缘起,它试图明确:我们需要哪些标准、信息、历史和来源,才能理解一个数据集究竟是为什么目的设计的,以及更重要的是,它什么时候就不该再被使用了。这是另一件我们谈得远远不够的事:这些数据集什么时候该被弃用?而当它们被下架之后,却在“学术种子”之类的地方继续拥有“来生”,又该怎么办?就在过去这几年里我们看到过这种情况:一些训练数据集被认为有问题,创建者说好,我们把它撤下来,但它们仍在继续流通,当然也继续影响着许多生产级的系统。这是第一个问题。而你指出的第二个问题是,我们根本不知道大型科技公司内部的训练数据集有多大、是怎么构建的,这些都被当作高度机密的专有信息
便签笔记
10分类即政治:ImageNet 案例
28:44
um and there are things that again as researchers we do not get to see so that raises a big question in terms of how might we understand the way that these systems are used to classify to interpret how they feed into the logics of ad tech of insurance of policing and so forth so that is a major i think research blocker in the field and something that we have to contend with right well that's a beautiful segue to your next chapter which i i think is is a stunner um your one-on-classification um so let me just ask you a bit about that um you know in that chapter your core argument is that the process of classification in this case labeling taxonomies is inherently political and that's why narrow technical solutions to to statistical bias and the quest to skew data to make it more fair miss the point i mean i've been very intrigued by all these conferences on you know the fat comfort fairness in you know and all of the technical attempts um to reshape the data to somehow make it fair and equal and i i want to talk with you about how difficult that is what the problems are in terms of thinking you know articulating the problem in those terms and in that in this um chapter on classification you make terrific use of jeff balker and
而这些东西,作为研究者我们同样看不到。所以这就带来一个大问题:我们要如何理解这些系统被用来分类、用来解读的方式,以及它们如何嵌入广告技术、保险、警务等等的运作逻辑。所以我认为这是这个领域里一个重大的研究障碍,是我们必须应对的问题。好,这正好非常自然地引到了你的下一章,我觉得那一章简直精彩绝伦,就是关于分类的那一章。让我就此问你几个问题。在那一章里,你的核心论证是:分类的过程,在这里就是标注分类体系的过程,本质上是政治性的,所以那些针对统计偏差的狭隘技术解决方案,以及那种通过调整数据让它更“公平”的努力,都没抓到要点。我一直很关注那些关于公平性(FAT)的会议,以及各种试图重塑数据、让它变得公平平等的技术尝试。我想跟你聊聊这有多难,以及用那样的方式来界定问题会带来什么样的问题。在这一章关于分类的讨论里,你非常出色地运用了杰弗里·鲍克和
便签笔记
30:06
sue lee styles work on classifications as ordering systems that have huge epistemic power in shaping how we see the world and you know reading the chapter i was really reminded of early feminist work um like sandra harding and many other people on the role of scientific knowledge you know we did a lot of work early on on you know what we used to call the sociology of scientific knowledge and we were thinking then about how scientific knowledge actually constructs biological gender binaries in which women are defined by nature as inferior to men now so i was thinking about that as you were talking about how ai systems you know assume or one might even say building baking construct gender race and sexuality as natural fixed biological categories so i wondered if you could tell us a bit about that um which is kind of core and perhaps um by describing your imagenet project with trevor and paglin such a rich question and that there are so many parts to it i mean to speak personally i i first started publishing research on the question of bias in large-scale data systems back in i think it was 2010 and certainly at that point the view was that bias was a problem that could simply be addressed by collecting more data so you know if we have an issue here
苏珊·利·斯塔的研究,他们把分类看作一种排序系统,在塑造我们如何看待世界方面具有巨大的认识论权力。读这一章的时候,我真的想起了早期的女性主义研究,比如桑德拉·哈丁和其他很多人关于科学知识之角色的工作。我们早年做了很多所谓“科学知识社会学”的研究,当时我们思考的是科学知识如何建构出生理性别的二元对立,在这套建构里女性被自然地定义为劣于男性。所以当你谈到人工智能系统如何假定、甚至可以说如何建构、“烘焙”出性别、种族和性向这些自然的、固定的生物学范畴时,我就想到了那些。所以能不能请你讲讲这一点,这是很核心的部分,也许可以从你和特雷弗·帕格伦合作的 ImageNet 项目说起。这个问题太丰富了,包含了很多层面。就我个人而言,我第一次发表关于大规模数据系统中偏差问题的研究,大概是在 2010 年。当时的普遍看法是,偏差这个问题只要收集更多数据就能解决。所以如果这里有问题,
便签笔记
31:33
let's just get more data and we'll solve it but of course what we've seen over the last 10 plus years is actually the opposite is the case for extremely large scale data systems we have many instances of bias we could think of the way in which apple's credit worthiness algorithms give consistently lower credit to women we could think about the voice recognition systems that fail to recognize women we could think about the facial recognition systems that fail to recognize people with darker skin tones i mean it's there there are countless examples now but certainly one of the things the ways in which my thinking has changed is to see bias almost as sort of the mega fauna sort of you know the the large errors these these failure points of systems but instead i think we need to actually go a step below into the logics of construction which i think are actually far more profound in shaping the way systems actually see the world and in many cases we won't see a failure point in the same way as a spectacular disaster rather it becomes this kind of fine-grained way in which people are understood interpreted and valued and you can see this in in so many systems you know from hiring to criminal justice the way that people are being assigned to you know whether it be
那就多弄点数据,问题就解决了。但当然,过去这十多年我们看到的恰恰相反:对于超大规模的数据系统,偏差的实例比比皆是。我们可以想想苹果的信用评估算法一贯给女性更低的额度;可以想想语音识别系统识别不出女性的声音;可以想想人脸识别系统识别不出深肤色的人。这样的例子现在数不胜数。但我的思考发生的一个变化是,我开始把偏差几乎看作是“巨型动物”,也就是那些显眼的大错误、系统的失效点;相反,我认为我们真正需要往下走一层,进入构建的底层逻辑,我觉得那才是更深刻地塑造系统如何看待世界的东西。而且在很多情况下,我们不会看到某个像壮观灾难那样的失效点,它反而会变成一种细颗粒度的方式,决定人如何被理解、被解读、被赋予价值。你可以在非常多的系统里看到这一点,从招聘到刑事司法,人们被归入的那些类别,无论是
便签笔记
32:51
uh again gender race but it could also be a risk score it could be a credit score where do these ideas come from and how do they become baked into systems is a core problem for us now certainly one of the extraordinary things we've seen in the last five years has been the focus on fairness accountability and transparency issues things like the fact conference these i think are necessary interventions but they're not sufficient because they always in many cases we see many papers factor these as purely technical problems that require technical fixes without thinking about the way in which by ingesting the data from the past we are absorbing those forms of structural inequality bias and ways of seeing that are actually profoundly problematic and that's certainly something that we saw working on the imagenet project with the artist trevor paglin and and certainly here it was it was extraordinary to to really again start to look at something as important as imagenet you know when we began studying it you know it had been it had been online for almost 10 years it had complete it was a colossus of image recognition um but certainly it really hadn't been studied at the level of what was happening in particular categories particularly the people
呃,还是性别、种族,但也可能是风险评分,也可能是信用评分。这些观念从哪里来、又是怎么被固化进系统里的,这是我们当下的一个核心问题。当然,过去五年里我们看到的一件了不起的事,就是大家开始关注公平性、问责和透明度的问题,比如 FAccT 这样的会议。我认为这些是必要的干预,但还不够,因为很多情况下我们看到不少论文把这些当成纯粹的技术问题,认为只要技术上修补一下就行,而没有去想:当我们把过去的数据吃进系统时,其实也在吸收那些结构性的不平等、偏见和看世界的方式,而这些恰恰是极其成问题的。我们和艺术家 Trevor Paglen 一起做 ImageNet 项目时,肯定也看到了这一点。能真正去重新审视像 ImageNet 这么重要的东西,确实是件了不起的事。我们开始研究它的时候,它已经上线快十年了,是图像识别领域的庞然大物,但在具体类别层面到底发生了什么,其实一直没有被认真研究过,尤其是「人物」这个类别。
便签笔记
34:04
category where we found so many just deeply racist misogynistic terms but also you know terms that were just made no sense terms that didn't have a visual analog like describing somebody as a debtor or a friend or an acquaintance you know these ideas that are not part of how we'd understand a sort of a visual noun so certainly one of the things we did was to start to open up these epistemic questions around how knowledge is being made and how classifications at that deep level need to be far more i think critically engaged with both in the technical fields but i think this is fundamentally a socio-technical question which means you need different sorts of expertise around the table [Music] i mean they're difficult problems aren't they because i mean once you recognize that they're structural problems then they're kind of very hard to fix i mean if i if if we think about and maybe you could elaborate a bit more on um you know amazon's heart automated hiring tool that that was not selecting women you sort of talk about how that that's a difficult problem to fix in a way well maybe you could talk about that because actually it's to do with embedded gender use of language it's to do with all
在那里我们发现了大量赤裸裸的种族歧视、厌女的词汇,还有一些根本说不通的词——那些没有视觉对应物的词,比如把某个人描述成「债务人」「朋友」或者「熟人」,这些概念并不属于我们理解的那种「视觉名词」。所以我们做的一件事,就是打开这些认识论层面的问题:知识是怎么被生产出来的,深层的分类体系为什么需要被更具批判性地对待——在技术领域内是这样,但我认为这本质上是一个社会技术问题,也就是说,你需要不同领域的专家坐到同一张桌子前。〔音乐〕我是说,这些都是很棘手的问题,对吧?因为一旦你意识到它们是结构性问题,就非常难解决了。比如说,也许你可以再多讲讲亚马逊那个自动招聘工具,它不选女性;你提到过那是个很难修复的问题。也许你可以谈谈这个,因为它其实牵涉到语言中内嵌的性别,牵涉到各种
便签笔记
11招聘算法、多样性与设计的边界
35:20
kinds of things that actually fiddling around with that app we're not going to fix and where does it sort of lead you i mean this is particularly relevant to our gender project at turin you know how does one kind of deal with things like that that you can you know that these these examples come up and they get a lot of publicity you know like the you know the female voices um and and then there's a discussion about what should we um i i mean i know at apple they had a discussion about kind of neutral void you know will we try and invent a voice um that doesn't that isn't identifiable as vascular or feminine feminine then there's an issue about accents i mean all of these things are sort of very difficult aren't they i mean if we take something like the um recruitment as a as a as an example what do you think about that where does one go once one recognizes how how deep these kind of connections are well i mean this this would bring us back to our questions around what can be addressed technically and what has to be addressed through regulation and of course this is very timely because if we think about just in the last couple of months you know one of the well-known hiring companies higher view which was using emotion recognition in its video
各种东西,光在那个应用上修修补补是解决不了的。那这会把你引向哪里呢?我是说,这跟我们图灵研究所的性别项目特别相关。人们该怎么应对这类事情?这些例子层出不穷,也引来很多关注,比如女性声音的问题,然后大家就讨论我们该怎么办。我知道苹果内部曾讨论过一种中性的声音,说我们能不能造出一种听不出是男性还是女性的声音;接着又有口音的问题。所有这些事情都非常难办,对吧?就拿招聘来说,你怎么看?当人们意识到这些关联有多深之后,该往哪儿走?嗯,这就把我们带回到那个问题:哪些可以用技术手段解决,哪些必须通过监管来解决。这当然非常应景,因为就在最近几个月,有一家知名招聘公司 HireVue,之前在视频
便签笔记
36:35
interviews uh made a sort of big announcement that yes they were listening to the criticism and they were going to devolve that from their systems but they still use things like basically voice emotion detection listening to your vocal tone and again trying to decide you know whether that says something about your character or your employability so you know again at the level of hiring i think we have lots of questions around who is valued you know who is the idea of the ideal subject or the ideal employee but i'm curious in in your project judy you know how do you see this in terms of how do you deal with this in terms of the way that gender itself is also inflected in these systems do you see this as something that design can contribute to in terms of technical designs or is this something where we have to think in more kind of policy in regulatory terms i mean i i think we have to do all of these things you know we have to do something about the you know the lack of diversity in the field that there are so few that you know when you look at the leading machine learning conferences that the number of women contributing there is kind of tiny you know that you've got a very skewed workforce that's mainly a young male workforce and i do think that is by no means a solution but i
面试里使用情绪识别,他们高调宣布说,是的,他们听取了批评,会把那部分从系统里去掉。但他们仍然在用类似语音情绪检测的东西,听你的语调,然后据此判断这是否说明了你的性格或者你的可雇佣性。所以在招聘这个层面,我觉得有很多问题:谁被认为是有价值的?所谓理想的对象、理想的员工是什么样子?不过 Judy,我很好奇你们的项目,你怎么看性别本身也被编织进这些系统里的方式?你觉得这是设计、是技术设计能够贡献力量的地方,还是我们必须更多从政策和监管的角度来思考?我觉得这些事我们都得做。比如这个领域缺乏多样性的问题——看看那些顶尖的机器学习会议,女性投稿者的数量少得可怜,劳动力构成非常失衡,主要是年轻男性。我确实认为这绝不是解决方案,但我
便签笔记
37:50
i've strongly argued as you know for decades that you can only um you know design from your experience in a sense and if there's a very narrow range of experience kind of feeding into design and systems again whether it's hardware or software you get a very kind of limited use of the imagination creativity it's not good for any of us so i think that's very important i mean i also you know we're hoping to do some more work on organizational culture because it seems to me that partly we need to get into organizations and and see what's going on and why only certain kinds of work workers feel comfortable in those organizations why there's a kind of chilly climate in there and i think that you know that doing something about those things would really contribute to a kind of wider debate both about kind of um you know about regulation and also a more informed kind of citizenship which i you know i wanted you to talk about but you know a better broader kind of public debate about these things and very much in terms of you know when do we want to use these systems and when don't we and the difference between using these tools as part of recruitment right that human resource management or a judge of course it's good to have more data right it's of course it's good to have
我几十年来一直强烈主张:某种意义上,你只能从自己的经验出发去做设计。如果输入到设计和系统里的经验范围非常狭窄——不管是硬件还是软件——你得到的想象力和创造力的运用就会非常有限,这对我们所有人都不好。所以我认为这一点很重要。另外,我们也希望在组织文化方面再做些研究,因为在我看来,我们多少需要走进这些组织内部,看看到底发生了什么,为什么只有某几类员工在那里觉得自在,为什么那里的氛围是冷冰冰的。我觉得,在这些方面做点事情,真的能为更广泛的讨论做出贡献——既包括关于监管的讨论,也包括一种更知情的公民意识,我本来就想请你谈谈这个——总之是一场更好、更广泛的公共讨论。而且很大程度上是要讨论:我们什么时候想用这些系统、什么时候不想用;以及把这些工具用在招聘、人力资源管理里,和用在法官身上,是不一样的。当然,数据多一点是好事,
便签笔记
12创新观的失衡:技术 vs 政策
39:06
more knowledge but as i think virginia eubank so well talks about in her book where's the human discretion where's the judgment you know how do we um you know combine these things and make decisions about when we want to use these systems and when we don't but you know please um please elaborate yourself on these things yes really and this is why i love talking with you about these topics because i think there's so much in that in terms of thinking about how and where we vest the idea of creative innovation and certainly for the last 15 20 years silicon valley has a really told a story about technical innovation that that has been lionized at all costs move faster break things you know this is how we do things best we we create technical innovation but what we haven't done and i think what's been a real loss here is thinking about the sorts of innovation around policy around regulation around law around all of the sorts of sort of ethical frameworks and ways in which we could consider the social implications of these systems these are not spaces that have had anything like the same amount of intent interest investment or or sort of social focus and and we can see certainly the fruits of that now when we've i think in many ways over
知识多一点也是好事,但正如我觉得 Virginia Eubanks 在她书里讲得非常好的:人的裁量权在哪里?判断力在哪里?我们该怎么把这些结合起来,去决定什么时候要用这些系统、什么时候不用。不过请你自己多讲讲这些吧。是的,真的,这也是为什么我特别喜欢和你聊这些话题,因为里面有太多值得挖的东西,比如我们把「创造性创新」这个想法寄托在何处、以何种方式寄托。当然,过去十五二十年里,硅谷讲了一个关于技术创新的故事,把它捧上了神坛,不惜一切代价——「快速行动,打破陈规」,说这就是我们把事情做到最好的方式,我们创造技术创新。但我们没有做的、我觉得真正的损失在于:我们没有去思考政策上的创新、监管上的创新、法律上的创新,以及那些能让我们考量这些系统之社会影响的伦理框架和路径。这些领域从来没有获得过同等的关注、投入或社会重视。而现在我们当然能看到它的后果,因为我们在很多方面过度
便签笔记
40:22
capitalize on a highly technical vision of you know what counts and you know you talk about it it's important to have more data versus more information you know these are different things and you know what constitutes information and as you say good judgment is is something that is so important here and this of course takes us back to to the history of of how systems get designed and i always think here of laney daston and peter galson's work on objectivity um where they look sort of to the shift in mechanical objectivity that moment when it was decided that through tools we would have you know a more truthful account of the world we've certainly had a a moment where machine learning has been seen to be this sort of all-purpose tool that can be applied to anything from welfare as we saw in the case of virginia eubanks book and we see how that fails us all the way through to criminal justice and we could think of pro-public as work and we could think of the work of of the markup looking at how these systems again are failing us there so we we have a decision to make which is how are we going to use the next 10 years and and where do we really need innovation and creativity and i would suggest that that is where we need to be thinking is around policy
押注在一种高度技术化的视野上——押注在「什么才算数」上。你刚才也说到,重要的是拥有更多数据,还是更多信息,这是两回事。什么构成信息?正如你所说,良好的判断力在这里极其重要。这当然又把我们带回到系统是怎么被设计出来的这段历史,我在这里总会想到 Lorraine Daston 和 Peter Galison 关于「客观性」的研究,他们考察了向「机械客观性」的转变——那个人们认定通过工具就能获得对世界更真实描述的时刻。我们现在肯定也处在这样一个时刻:机器学习被看作一种万能工具,可以用在任何地方,从福利(就像 Virginia Eubanks 书里写的那样,我们看到它如何辜负了我们),一直到刑事司法(可以想想 ProPublica 的报道),还有 The Markup 的工作,同样揭示了这些系统在那里如何让我们失望。所以我们面临一个选择:接下来的十年我们要怎么用?我们真正需要创新和创造力的地方在哪里?我想我们该思考的方向是政策,
便签笔记
13伦理准则不够,要谈权力
41:36
is around public debate and is around regulation and and i am optimistic because i'm starting to see certainly some very positive signs in those areas um and as we've seen of course in the eu we've seen the first omnibus draft regulations for ai and there to make perhaps still early days imperfect but a step in the right direction in thinking about how do we regulate these systems better but i wanted to ask you more about ethics actually because you have a very nice line of course you know in the book about you know too much focus in a way having been on ethics and not enough about power i mean i think this is why yes why you know that we should be focusing less on ethics and more on power and you mentioned that there's a proliferation of ethical codes right there is you know i can't remember how many hundreds but there are sort of loads of ethical codes and i just wondered what you thought about that i mean is it you know are you thinking about at some point there being some kind of global agreement or you know how do they function and also i guess whether you think which you clearly do with the sentence about you know we focus too much on ethics and and not enough about power whether you know too much energy
是公共讨论,是监管。我是乐观的,因为我确实开始在这些领域看到一些非常积极的信号。当然,在欧盟我们已经看到了第一份关于人工智能的综合性法规草案,也许还处在早期、并不完美,但在「如何更好地监管这些系统」这个方向上迈出了一步。不过我其实想再多问问你关于伦理的事,因为你书里有一句很精彩的话,大意是我们过多地聚焦在伦理上,而对权力关注不够。我想这正是你的意思:我们应该少谈一点伦理,多谈一点权力。你还提到伦理准则大量涌现——我记不清有几百份了,反正各种伦理准则一大堆。我就想知道你怎么看这件事?你有没有设想过某个时候会出现某种全球性的共识?或者说它们到底是怎么运作的?另外我猜,既然你写下了「我们太关注伦理、太不关注权力」这样的话,你显然是这么想的——你是否觉得太多精力
便签笔记
42:51
is going in the direction of ethics codes i mean i wonder if you could elaborate your thoughts on that for us there's quite a few questions in the chat i'm just sort of looking here about um oh amazing questions yes so i'd love to tell us a bit more about that yes and thank you for these questions we'll certainly get to them in in a moment um allow me to just address that in in two ways one is to say philosophically obviously you know issues of ethics and power are always intertwined um and we have you know decades in fact centuries of philosophy to point to exactly how that works but in the book what i'm referring to specifically is the way in which ethics codes have been used as a way to brush off regulation as a way to say we've got this we'll have an ethics code therefore we don't need any laws or sort of regulatory guard rails around how we're creating these systems or who they might be harming and and quite frankly this isn't enough and and certainly while i think it is important to articulate forms of ethical principles in terms of how we design systems and deploy them there is something much deeper which is how do we actually make sure that these principles are accountable what are the ways in which we can ensure
都投到了伦理准则这个方向上?不知道你能不能就此为我们展开讲讲。聊天区里还有不少问题,我正在这儿看……哦,问题都很精彩。是的,我很想听你多讲讲这个。好的,也谢谢大家提的这些问题,我们等一下肯定会讲到。请允许我从两个方面来回应。第一,从哲学上讲,伦理与权力的问题当然始终是交织在一起的,我们有几十年、其实是几百年的哲学传统可以说明这一点是如何运作的。但在书里,我具体指的是:伦理准则被用来搪塞监管的那种方式——「我们已经管好了,我们有伦理准则,所以不需要任何法律,也不需要围绕我们如何构建这些系统、系统可能伤害谁来设置监管护栏」。坦白说,这不够。当然,我认为在设计和部署系统时阐明某种伦理原则是重要的,但还有更深层的东西:我们如何真正确保这些原则是可问责的?我们有哪些方式可以确保
便签笔记
44:06
that in fact if something does go wrong that this won't happen again and there will actually be an account so this is where i think regulation is is far more necessary and certainly accounts of power i does this system actually give more power to the powerful a system you know a question we should always ask of the systems we build and also what are the ways in which these systems actually speak to existing forms of power be it capital policing the military you name it that to me is is a much more useful lens than quite high-level ethical statements around safety and you know making sure that you know we do no harm because again and again we've seen how again these codes have actually not borne out with the reality yeah well i mean it's interesting because there's such a kind of fashion in silicon valley for kind of user groups ux groups as if that's as if that's the answer you know that you get some users in and and test something out and it's such a kind of narrow context of use whereas you're kind of very much stressing that actually it's people who are affected by these technologies who should be the ones who are consulted and you know participate in the design and we've had a lot of movements in our in our circles for years about um
万一真的出了问题,它不会再次发生,并且真的会有人给出交代?所以我认为这正是监管更加必要的地方。关于权力的分析也是如此:这个系统是不是让本就有权力的人更有权力?这是我们对自己构建的系统应该始终追问的问题。还有,这些系统以什么方式与既有的权力形式发生关系——资本、警务、军队等等。在我看来,这比那些高层次的伦理宣言(关于安全、关于「不造成伤害」)要有用得多,因为我们一次又一次地看到,这些准则并没有在现实中兑现。是啊,这很有意思,因为硅谷特别流行用户小组、用户体验小组,好像那就是答案:找几个用户来测试一下。可那是个非常狭窄的使用情境。而你其实非常强调,真正应该被咨询、应该参与设计的,是那些受这些技术影响的人。在我们这个圈子里,多年来一直有很多关于
便签笔记
14参与式监管与拒绝的政治
45:23
participatory design could you say something about that do you think that's an important kind of element in all of this oh absolutely and and certainly you know one of the things that i've worked on as have many other scholars including people like andrew selpst is this idea of impact assessments you know how would you actually have impact assessments where members of the community can have a say in whether a government will deploy a facial recognition system or you know a hiring ai system or a you know risk assessment system inside the criminal justice system for example these are many many places at which you could actually have public debate and public decision making and also the ability to have a politics of refusal to be able to say no we don't want this system here we actually don't want this applied and and that is just one of many mechanisms that we could start to think of around participatory regulation so as you know in many cases laws have been you know informed and in some cases even partly written by technology companies and i think that's simply not going to get us where we need to go when we're starting to see you know such extraordinary power asymmetries between the tech sector and you know the populations that they
参与式设计的运动。你能谈谈这个吗?你觉得这在整件事里是个重要的元素吗?绝对是。我和其他很多学者——包括 Andrew Selbst 这样的人——做过的一件事,就是「影响评估」这个想法:你要怎样真正建立起影响评估机制,让社区成员能够对政府是否部署人脸识别系统、招聘 AI 系统,或者刑事司法系统内部的风险评估系统等等有发言权?有非常非常多的环节,其实是可以进行公共讨论和公共决策的,也可以有一种「拒绝的政治」——能够说不,我们不要这个系统,我们不希望它被用在这里。而这只是我们可以设想的参与式监管机制中的一种。你也知道,很多时候法律是受科技公司影响而制定的,有些甚至部分是由它们起草的。我认为,当我们开始看到科技行业与它们所服务的人群之间存在如此惊人的权力不对称时,这种做法根本无法把我们带到该去的地方。
便签笔记
46:32
serve so instead i think we need to think much more radically around how regulations are formed and the ways in which communities can have a say in the way that they're applied now that's really interesting because you know the last conference i went to in stanford was called um what was it coding for care and it was a you know a discussion exactly about what elements of care for the elderly it's always you know the discussions are always care you know but it's even posed as that right um and it's posed in terms of keeping individuals at home alone as long as possible and then thinking about systems that will operate and it's exactly sort of context like that where you really do want to step back and think well what do you want to use the technology for in a way and what what isn't appropriate and where do we get guidelines to think that through you know do you have anything yeah do you have any views about the care issue i mean i mean care work is such an important question because again it's been so central during the pandemic is to think about you know what it is when we're isolated you know and and again the way in which the debates have said how can automation get us out of this rather than thinking how might mutual
所以我认为,我们需要更激进地重新思考法规是如何形成的,以及社区如何能对法规的适用方式拥有发言权。这真的很有意思,因为我最近去斯坦福参加的那场会议叫……叫什么来着,「为照护编码」,讨论的正是对老年人照护的哪些环节可以怎样处理——讨论总是围绕照护,甚至议题本身就是这么设定的。而且它的框架是:让个体尽可能长时间地独自待在家里,然后再去想有哪些系统可以运转起来。恰恰是在这样的情境里,你真的很想退后一步想想:我们到底想用这项技术做什么?什么是不合适的?我们又该从哪里获得指导原则来把这些想清楚?你有什么看法吗?关于照护这个问题你有什么想法?照护工作是个特别重要的问题,因为在疫情期间它太核心了——想想当我们被隔离时那意味着什么。而且这些讨论的方式往往是:自动化怎样才能让我们摆脱这种处境?而不是去想,互助
便签笔记
47:44
aid be a much better framework and the question here for me is why do we always put technology at the center yeah why do we assume that technology is either going to be the you know the panacea or the problem rather than saying why don't we actually think about this problem around what kind of world do we want to live in what are the ways in which we increase equity and justice and then how technology might serve that vision rather than driving it i mean there are so many ways in which you know the ai debate is frustrating because it keeps assuming that ai really can be applied to everything and anything and it's always the first go-to approach rather than looking at the long histories of you know how do we actually produce the types of social change that people are calling for well it's very rarely just a tech fix so i think that means that we're looking for different sorts of work and again this is where your work has been so important um and i noticed that there's several questions in the chat and i'm going to have to answer this um asking about the artwork of the book and the cover which is just behind me right now um but it also makes me think of how important artists are in these sorts of discussions as well and and again the culture industries for telling different sorts of stories
是不是一个好得多的框架。对我来说,这里的问题是:我们为什么总是把技术放在中心?为什么我们总假定技术要么是万灵药、要么是问题所在,而不是说:我们为什么不真正围绕「我们想生活在一个什么样的世界里」「我们有哪些方式可以增进公平与正义」来思考问题,然后再看技术如何服务于这个愿景,而不是由技术来驱动它?关于人工智能的讨论有太多让人沮丧的地方,因为它总是假定 AI 真的可以应用于一切、任何事情,而且永远是第一选择,而不是去看那些漫长的历史:我们究竟是怎么产生人们所呼吁的那种社会变革的?而那很少只是一次技术修补。所以我觉得这意味着我们要寻找不同类型的工作,这也正是你的工作如此重要的地方。我注意到聊天区里有好几个问题,我得回答一下:有人问这本书的封面艺术,就在我身后。这也让我想到,艺术家在这类讨论中有多重要,还有文化产业在讲述不同的故事方面有多重要——
便签笔记
15业界反应、封面艺术与 AI 作为交叉学科
48:58
about how technology works and fails so to answer your question in the chat about uh this extraordinary image it's by vladan jola who i also collaborated with on a project called anatomy of an ai system where we came up with a gigantic map of a single amazon echo this cover image is actually one of many designs he has in the book illustrating ideas around the way in which again you can see from sort of the human head ideas of intelligence and phrenology to how we represent the world to the planetary costs of large-scale computational systems so i really think his work is marvelous so thank you for that question um in fact there's a field of amazing questions do you think i can get to all of them i absolutely can't i've i i've just got two quickies um yeah i'm looking at the clock can i just ask you two quickly i mean one is like i can't help but ask you what the industry response has been like to the book because because i see you you know you give it a talk at the computer history museum and to lots of different audiences what's what's been the response of the of the tech community in a way well it's interesting because i think certainly the way in which i wrote this book is you know if you think about it
关于技术如何运作、又如何失效的故事。那么来回答聊天区里关于这张非凡图像的问题:它出自 Vladan Joler 之手,我也和他合作过一个叫《一个 AI 系统的解剖》的项目,我们为一台亚马逊 Echo 画了一张巨大的地图。这张封面图其实只是他为这本书做的众多设计之一,用图像去阐释各种观念——从人的头颅形象里关于智能和颅相学的想法,到我们如何表征世界,再到大规模计算系统的行星级代价。所以我真的觉得他的作品非常了不起,谢谢这个提问。其实还有一大堆很棒的问题,你觉得我能全部回答完吗?肯定不行。我这里有两个小问题,我看了下时间,能快速问你两个吗?一个是我忍不住想问:业界对这本书的反应如何?因为我看到你在计算机历史博物馆做过讲座,也面对过很多不同的听众,科技圈的反应是怎样的?这挺有意思的,因为我写这本书的方式,如果你想一想的话,
便签笔记
50:14
it's it's almost like geological strata it's it's written almost as a sort of a full computational stack you know we start with the earth we look at labor data classification the state all the way up to outer space and so in that sense i think for different different groups within the tech sector they have different sorts of questions depending on what part of the problem they're addressing so for groups that are sort of working around large-scale data analysis and prediction issues around classification and bias are extremely timely right now and people are seeking different ways of contending with this because we know that certainly over the last few years these problems are not going away so we do need different ways of thinking then again in other parts of the industry and i find this really heartening there's a new set of conversations about how to reduce the energy intensiveness of machine learning approaches that is still very new but it's extremely exciting we're seeing that there is a lot of waste there that can actually be improved upon a lot of cycles that can actually be compressed and algorithmic techniques that will just be far less energy intensive we're going to need that you know we're at a time of climate crisis and so as a sector i think there's a growing
它几乎像地质地层,几乎是按照一整个计算堆栈来写的:我们从地球开始,看劳动、数据、分类、国家,一路上到外太空。所以从这个意义上说,科技行业里不同的群体会有不同的问题,取决于他们面对的是问题的哪一部分。对于那些做大规模数据分析和预测的群体来说,分类和偏见的问题现在非常紧迫,人们在寻找不同的应对方式,因为我们知道,过去几年里这些问题并没有消失。所以我们确实需要新的思考方式。而在行业的另外一些地方——这让我觉得很受鼓舞——出现了一批新的对话,讨论如何降低机器学习方法的能耗强度。这还非常新,但极其令人兴奋。我们看到其中有大量浪费其实是可以改进的,有大量计算周期是可以压缩的,还有一些算法技术能耗会低得多。我们会需要这些,毕竟我们正处在气候危机的时代。所以作为一个行业,我觉得有一种日益增长的
便签笔记
51:25
sense of concern and i think of responsibility so you know in in that sense i i've had really positive experiences and i think people are asking really good questions but at the same time you know we are seeing you know a real sort of shift towards you know increasing the tools of uh facial recognition predictive policing you know there are there are lots of companies and sectors that won't find this book so edifying to their beliefs and certainly this is why i think it's important that we have these public debates now because these tools are being used in so many ways that i think many of us would find deeply concerning i wanted to end by asking you well to say that it's it's just so wonderful that women like you are at the forefront of these debates about bias in ai you know it's and and i wonder why you think that is because of course when i started working on gender and technology certainly there were not a lot of women in the field and now it's very striking actually that many of the wonderful books that are coming out and the research that's being done is very much being led by women and feminists at the moment so i wanted you know why do you think that is you've got a yeah i mean it's it's certainly interesting
忧虑感,我认为还有一种责任感。所以从这个意义上说,我的经历其实相当正面,我觉得人们提出的问题都很好。但与此同时,我们也确实看到一种明显的转向:人脸识别、预测性警务这类工具在不断被强化。有很多公司和领域不会觉得这本书对他们的信念有什么启迪。当然,这正是我认为我们现在必须进行这些公共讨论的原因,因为这些工具正在以许多我们中很多人会深感忧虑的方式被使用。最后我想说,像你这样的女性走在关于 AI 偏见这些讨论的最前沿,真是太好了。我很好奇你觉得这是为什么?因为我当年开始做性别与技术研究的时候,这个领域确实没有多少女性;而现在很显著的是,许多出色的新书、许多正在进行的研究,都是由女性和女性主义者在引领。所以我想问,你觉得这是为什么呢?是的,这确实很有意思,
便签笔记
52:39
um that that we're seeing a far more sort of diverse coalition of people working on these issues i think then we have of people who are just working on algorithms and that should tell us something in terms of who are the people who have experienced marginalization ostracization misrecognition by systems other people who are going to be very alert to these questions and will be asking again how do we improve these kinds of issues before they become more ingrained in our society but also one of the things i like to do is to go back and look at the very early years of artificial intelligence in the 1950s and 60s and even back then you know you would have margaret mead i know on on the same panel as you know jeffrey bateson you'd have people who are actually designing systems um sitting there with anthropologists who are thinking about their social implications we lost that for a while there and i think what we've seen is this kind of over prioritization of the technical and it's something that we really need to correct for now because these are no longer you know systems that are just being designed in labs or that are essentially you know theoretical interventions these are systems that are affecting
我们看到,从事这些议题的人群构成,比单纯做算法的人群要多元得多,这本身就说明了一些问题:那些经历过被边缘化、被排斥、被系统错误识别的人,往往会对这些问题格外警觉,会不断追问我们要如何在这些问题更深地嵌入社会之前把它们改善。另外,我喜欢做的一件事,是回过头去看人工智能最早期的那些年,五十、六十年代。即便在那时,你会看到玛格丽特·米德和贝特森同台,会看到真正在设计系统的人和思考其社会影响的人类学家坐在一起。有一阵子我们把这个丢掉了,我认为我们看到的是对技术层面的过度优先,而这是我们现在真的需要纠正的,因为这些已经不再只是实验室里设计出来的系统,也不再只是本质上属于理论性的干预,这些系统正在影响
便签笔记
53:49
billions of people around the world and in very different ways so what that means is that i think we have to start reconceptualizing ai as an interdiscipline and as one that has to be very much grounded in the communities who are being affected by the these tools every day so so that is certainly something that i think is is urgently needed for the next decade in the space fantastic well look it it's um it just leaves me to congratulate you on the book it is a fantastic book i really recommend everyone read it it's beautifully produced it's very easy to read clearly set out i mean it is a kind of wonderful job i've got a champagne glass yes i'm going we were going to have it i'm very sorry that we've had drinks because i'm sad about that too we would go and have a drink um but i let us have a let us have a little glass um and congratulations a wonderful job thank you so much and i i couldn't recommend the book more highly cheers it means so much to celebrate this with you and with everybody who's joined today it's it's an extraordinary group um and i really hope that we can do this in person sometime i hope so too vaccines willing cheers thanks again
全世界数十亿人,而且影响方式各不相同。这意味着,我认为我们必须开始把 AI 重新构想为一门交叉学科,一门必须深深扎根于那些每天被这些工具影响的社群之中的学科。所以这肯定是这个领域未来十年迫切需要的东西。太好了。那么,我只剩下祝贺你这本书了。这是一本非常棒的书,我真心推荐所有人都去读,装帧精美,读起来很轻松,条理清晰,真的做得非常出色。我这儿有一杯香槟,是的,我们本来打算……很遗憾我们没能一起喝一杯,我也为此难过,本来可以出去喝一杯的。不过还是让我们各自举一小杯吧。祝贺你,干得漂亮。非常感谢你,我怎么推荐这本书都不为过。干杯!能和你、和今天所有到场的人一起庆祝,对我意义重大,这是一群了不起的人。我真的希望我们某天能当面这样做。我也希望,等疫苗允许的时候。干杯,再次感谢。
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Kate Crawford 与 Judy Wajcman 围绕新书《Atlas of AI》展开对谈,核心论点是:AI 并非抽象、非物质的智能,而是一套从矿产、能源、劳动、数据到国家权力的星球级抽取性产业,其"智能"和"客观"外衣掩盖了它作为"权力登记簿"服务于既有支配利益的本质,因此应对之道不是伦理准则而是规制与权力分析。

核心要点

  • AI 是字面意义上的抽取性产业。 Crawford 以内华达州 Silver Peak——美国最后一座运营中的锂矿——作为全书开篇,指出锂离子电池(手机、特斯拉、AI 设备)依赖稀土、钴、锂等矿产;拜登政府刚发布的供应链危机文件印证了"星球级计算"的地缘政治脆弱性。疫情期间人们以为碳足迹下降,但大规模远程视频通话本身耗能巨大。
  • "自动化"背后是被隐藏的人工劳动(fauxtomation)。 引用 Mary Gray 与 Siddharth Suri 的《Ghost Work》:加纳、菲律宾等地的远程平台工人以低于贫困线的计件工资标注训练数据;书中提到的 x.ai 公司让真人每天工作 14 小时假扮"数字助理",以制造无缝 AI 的假象。
  • 算法管理是 Babbage 与泰勒主义的当代复活,且横跨蓝领与白领。 Crawford 亲访亚马逊履单中心,感受到"极其恐怖"的身心压力;贝索斯上月宣布应对工伤的方案竟是引入追踪工人韧带与肌肉使用的新算法系统。疫情 18 个月来,远程办公监控服务激增(摄像头、邮件计数、会议时长→"效率评分"),Four Little Trees 等公司甚至对居家学生做面部微表情"情绪检测"。
  • 情绪识别在理论根基上就是坏的,亟需规制。 Crawford 用整章论证"看脸推断内心状态"的假设——即所谓"AI 测谎仪"——从根本上站不住脚。HireView 虽宣布撤除视频面试中的情绪识别,但仍保留通过声调判断"品格与可雇性"的语音情绪检测。
  • "数据越多越好"是一种可追溯至 IBM 语音实验室的意识形态。 据媒介史学者 Xiaochang Li 的研究,当时的研究员 Robert Mercer(后来资助特朗普竞选的那位)提出"更多数据总是更好的数据",标志着 AI 从专家系统/符号逻辑转向概率与暴力计算。其后果是过度采集、数据集合并造成的"预测性隐私伤害",以及把数据当作待开采的"自然资源"。
  • 训练数据集缺乏来源标准与"退役"机制。 Crawford 与 Timnit Gebru 合作提出"Datasheets for Datasets",类比硬件规格书,要求记录数据集的来源、用途与何时应停用。已被撤回的问题数据集仍在 Academic Torrents 等处流传并持续喂养生产系统;大公司的专有数据集则完全不对研究者开放,构成该领域的重大研究阻碍。
  • 偏见只是"巨型动物"级的表层故障,真正问题在分类的建构逻辑。 她 2010 年就发表偏见研究,当时主流认为"多收集数据就能解决",十年实践证明恰恰相反(Apple 信用算法给女性更低额度、语音识别听不懂女性、人脸识别对深肤色失效)。与艺术家 Trevor Paglen 剖析运行近十年的 ImageNet 时发现,"人"类别下充斥种族主义、厌女词汇以及"债务人""熟人"等根本无视觉对应物的标签——这是"颅相学式意志"的复现。因此 FAccT 等技术性公平修补是必要但远不充分的,因为吞食历史数据就是在吸收结构性不平等。
  • 伦理准则太多、权力分析太少。 数以百计的 AI 伦理守则被业界用作挡箭牌——"我们有伦理准则所以不需要法律"。Crawford 主张追问"这个系统是否让强者更强、它如何服务于资本/警务/军队",并通过规制确保问责。她对欧盟首份 AI 综合法规草案持谨慎乐观态度。
  • 创新应转向政策、法律与公共辩论,并赋予社区"拒绝的政治"。 硅谷 15–20 年来将"快速行动、打破常规"式技术创新奉为唯一价值,却极少投资规制与社会影响评估。她与 Andrew Selbst 等推动的影响评估机制,让社区能对政府部署人脸识别、招聘 AI、刑事风险评估系统投票甚至说"不";现行法律往往由科技公司参与甚至代笔,权力不对称必须被打破。
  • 为什么是女性与女性主义者引领这场辩论。 Wajcman 指出主流机器学习会议中女性占比极低,而当前批判性研究却多由女性主导;Crawford 回应:曾被系统边缘化、误识别的人最能察觉问题。1950–60 年代 AI 早期,Margaret Mead 等人类学家曾与系统设计者同席,此后技术被过度优先化;如今 AI 影响数十亿人,必须重构为扎根受影响社区的"跨学科"。

结论与值得注意的细节

  • 全书结构如"地质地层"或"完整计算堆栈":从地球(矿产)→劳动→数据→分类→国家→外太空,业界不同群体据此对号入座提出不同问题。
  • 业界回应两极:数据/预测领域对分类与偏见问题反应积极;关于降低机器学习能耗的新讨论令 Crawford 振奋,认为在气候危机下大量计算周期可压缩;但人脸识别、预测性警务等行业"不会觉得这本书令人愉快"。
  • 关于养老"照护"技术,Crawford 质问:为何总把技术置于中心、总把 AI 当作第一方案?应先问"我们想要怎样的世界、如何增进公平正义",再让技术服务于该愿景,而非以互助等社会框架为代价。
  • 封面及书中插图出自 Vladan Joler,二人曾合作《Anatomy of an AI System》——一幅描绘单个 Amazon Echo 全生命周期的巨型地图;Crawford 强调艺术家与文化产业在讲述技术"如何运作与失败"上的不可替代性。
  • 本书写作历时五年以上,Crawford 2012 年加入微软研究院恰逢卷积神经网络与深度学习崛起,坦言无法预见这些主题会如此迅速地成为现实。
核心句型 · 9
1. in order to X, we have to go to the places where Y
“In order to ground artificial intelligence in order to truly understand these material impacts we have to go to the places where it's made”
用「目的—手段」结构为一个反常规的做法(用锂矿开篇)辩护。适合解释研究方法或叙事选择:先亮目的,再说明为何必须走这条路。
2. X, and I mean X in the fullest sense
“We have to go to the places where it's made and i mean made in the fullest sense”
追加一句限定,把一个普通词升格为强调性概念。口语演讲中用来提示听众「这个词我用的是扩展含义」。
3. not just A, but B
“See not just the data within them but the logic in terms of how it's classified and ordered”
递进结构,B 是更深一层的重点。学术表达中常用来从表面现象推进到底层机制。
4. these are necessary interventions but they're not sufficient
“These i think are necessary interventions but they're not sufficient because…”
「必要但不充分」是批评性评价的经典框架:先肯定对方价值,再指出其局限,避免全盘否定。后面务必接 because 说明为何不充分。
5. why do we always X rather than Y?
“Why do we always put technology at the center… rather than saying why don't we actually think about this problem around what kind of world do we want to live in”
反问句加 rather than 对照,用于挑战一种默认思路并提出替代框架。适合议论文的转折段。
6. this is X brought to life
“To me this is babbage's vision brought to life and it's quite nightmarish”
用「某人的设想成真」把历史与当下连接,过去分词 brought to life 作后置定语。可替换为 realized / made real。
7. rather than A, B would be a much better framework
“How can automation get us out of this rather than thinking how might mutual aid be a much better framework”
提出替代性框架时的礼貌说法,比直接说 A is wrong 更具建设性。
8. there is something much deeper, which is how…
“There is something much deeper which is how do we actually make sure that these principles are accountable”
先承认表层做法有价值,再用 something much deeper 引出真正的问题。是从「伦理」推进到「权力/问责」的转场句式。
9. to put it mildly
“We both understand how important mining is politics you know that it's to put it mildly”
轻描淡写式反讽,暗示实际情况比说出来的更严重。插入语位置灵活。
词汇精讲 · 119 · 按出现顺序
attest to phr. 0:19
证明、证实(某事)
squarely /ˈskwerli/ adv. 0:19
直接地、正面地(处理问题)
inflection /ɪnˈflekʃən/ n. 1:45
色彩、倾向;(语言学)屈折变化。此处指理论倾向
dichotomous /daɪˈkɑːtəməs/ adj. 3:10
二分的、二元对立的
technological determinism n. 3:10
技术决定论
materiality /məˌtɪriˈæləti/ n. 3:10
物质性
commodity fetishism n. 4:34
商品拜物教(马克思主义术语)
operationalize /ˌɑːpəˈreɪʃənəlaɪz/ v. 4:34
使可操作、将…付诸实施
normative /ˈnɔːrmətɪv/ adj. 4:34
规范性的
legitimate /ləˈdʒɪtəmeɪt/ v. 5:52
使合法化、使正当化(注意此处作动词)
registry /ˈredʒɪstri/ n. 5:52
登记册、注册处
sussed out phr. 5:52
(英式口语)弄清楚、看穿
convolutional neural nets n. 7:15
卷积神经网络
computationally intensive phr. 8:23
计算密集的
tailorized /ˈteɪlərаɪzd/ adj. 8:23
泰勒化的(按科学管理原则分解、监控的)
brought to bear on phr. 8:23
施加于、作用于
situate /ˈsɪtʃueɪt/ v. 8:23
将…置于(具体语境中)
extractive industry n. 9:40
采掘业
to put it mildly phr. 9:40
说得客气一点
brine /braɪn/ n. 10:57
盐水、卤水
iridescent /ˌɪrɪˈdesnt/ adj. 10:57
泛着虹彩的
poised on this brink phr. 10:57
处于临界边缘
geopolitics /ˌdʒiːoʊˈpɑːlətɪks/ n. 10:57
地缘政治
off the point phr. 12:18
离题、没抓住要点
piece work n. 13:32
计件工作
tedious /ˈtiːdiəs/ adj. 13:32
冗长乏味的
propping up phr. 13:32
支撑、撑住
mundane /mʌnˈdeɪn/ adj. 13:32
平凡琐碎的
in the ascendant phr. 13:32
处于上升期、日益得势
socio-imaginary n. 14:50
社会想象(社会集体对自身的想象方式)
scientific management n. 14:50
科学管理(泰勒主义)
come to pass phr. 16:04
(文学语)发生、成真
fulfillment center n. 16:04
(电商)配送中心
exemplars /ɪɡˈzemplɑːrz/ n. 16:04
典范、范例
ligaments /ˈlɪɡəmənts/ n. 16:04
韧带
nightmarish /ˈnaɪtmerɪʃ/ adj. 17:15
噩梦般的
granularity /ˌɡrænjəˈlærəti/ n. 17:15
颗粒度、精细程度
nudging /ˈnʌdʒɪŋ/ n. 17:15
助推(行为经济学:温和引导)
spur /spɜːr/ n. 17:15
刺激、推动因素
uptick /ˈʌptɪk/ n. 17:15
上升、增长
scare quotes n. 18:37
表示存疑或反讽的引号
micro expressions n. 18:37
微表情
polygraph /ˈpɑːliɡræf/ n. 18:37
测谎仪
tap root n. 18:37
主根;引申为根本、根源
cognizant of /ˈkɑːɡnɪzənt/ adj. 19:52
意识到、知晓的
path breaking adj. 19:52
开创性的
harvesting /ˈhɑːrvɪstɪŋ/ n. 19:52
(数据)采集、收割
ground truth n. 21:10
基准真值(机器学习术语)
excavation /ˌekskəˈveɪʃən/ n. 21:10
发掘、挖掘
aggregate /ˈæɡrɪɡət/ adj. 21:10
聚合的、总体的
phrenology /frəˈnɑːlədʒi/ n. 22:34
颅相学(伪科学)
moral imperative n. 22:34
道德律令
knowable /ˈnoʊəbl/ adj. 22:34
可知的
expert systems n. 23:55
专家系统(早期符号 AI)
brute force n. 23:55
暴力计算、穷举法
liability /ˌlaɪəˈbɪləti/ n. 23:55
责任、负债;隐患
subjectivity /ˌsʌbdʒekˈtɪvəti/ n. 25:10
主体性
embodiment /ɪmˈbɑːdimənt/ n. 25:10
具身性、身体感
marveling at phr. 26:20
对…感到惊讶
provenance /ˈprɑːvənəns/ n. 27:32
来源、出处
deprecated /ˈdeprəkeɪtɪd/ adj. 27:32
(技术)被弃用的
afterlives /ˈæftərlaɪvz/ n. 27:32
来生;引申为(事物下架后的)残留生命
proprietary /prəˈpraɪəteri/ adj. 27:32
专有的、私有的
segue /ˈseɡweɪ/ n. 28:44
(话题)自然过渡
taxonomies /tækˈsɑːnəmiz/ n. 28:44
分类法、分类体系
skew /skjuː/ v. 28:44
使偏斜、调整(数据)
epistemic /ˌepɪˈstiːmɪk/ adj. 30:06
认识论的
baking /ˈbeɪkɪŋ/ v. 30:06
(bake in)固化、内置于
credit worthiness n. 31:33
信用资质
mega fauna n. 31:33
巨型动物;比喻显眼的大问题
fine-grained adj. 31:33
细颗粒度的
necessary but not sufficient phr. 32:51
必要但不充分(此处拆开出现)
ingesting /ɪnˈdʒestɪŋ/ v. 32:51
摄入、吸收(数据)
colossus /kəˈlɑːsəs/ n. 32:51
庞然大物、巨人
misogynistic /mɪˌsɑːdʒəˈnɪstɪk/ adj. 34:04
厌女的
visual analog n. 34:04
视觉对应物
socio-technical adj. 34:04
社会技术的
fiddling around with phr. 35:20
摆弄、修修补补
devolve /dɪˈvɑːlv/ v. 36:35
移交、剥离(此处指从系统中去除)
employability /ɪmˌplɔɪəˈbɪləti/ n. 36:35
可雇佣性
inflected /ɪnˈflektɪd/ adj. 36:35
带有…色彩的、受…影响的
skewed /skjuːd/ adj. 36:35
失衡的、有偏的
chilly climate n. 37:50
冷淡氛围(指少数群体在组织中的不友好环境)
human discretion n. 39:06
人的裁量权
vest /vest/ v. 39:06
赋予、寄托(于)
lionized /ˈlaɪənaɪzd/ v. 39:06
被奉为名流、被吹捧
capitalize on phr. 40:22
利用、押注于
mechanical objectivity n. 40:22
机械客观性(科学史概念)
all-purpose adj. 40:22
万能的、通用的
omnibus /ˈɑːmnɪbəs/ adj. 41:36
综合性的、一揽子的
proliferation /prəˌlɪfəˈreɪʃən/ n. 41:36
激增、扩散
intertwined /ˌɪntərˈtwaɪnd/ adj. 42:51
交织的
brush off phr. 42:51
搪塞、打发掉
guard rails n. 42:51
护栏;引申为防护性规则
accountable /əˈkaʊntəbl/ adj. 42:51
可问责的
borne out phr. 44:06
被证实、得到印证
you name it phr. 44:06
凡是你能想到的、等等
impact assessments n. 45:23
影响评估
politics of refusal n. 45:23
拒绝的政治
power asymmetries n. 45:23
权力不对称
mutual aid n. 46:32
互助
panacea /ˌpænəˈsiːə/ n. 47:44
万灵药
equity /ˈekwəti/ n. 47:44
公平(实质公平,区别于 equality)
go-to adj. 47:44
首选的、惯常求助的
tech fix n. 47:44
技术修补、技术万能解
geological strata n. 50:14
地质地层
contending with phr. 50:14
应对、对付
heartening /ˈhɑːrtnɪŋ/ adj. 50:14
令人鼓舞的
energy intensiveness n. 50:14
能耗强度
predictive policing n. 51:25
预测性警务
edifying /ˈedɪfaɪɪŋ/ adj. 51:25
有启迪的、有教益的
at the forefront of phr. 51:25
处于…最前沿
coalition /ˌkoʊəˈlɪʃən/ n. 52:39
联盟、联合体
ostracization /ˌɑːstrəsaɪˈzeɪʃən/ n. 52:39
排斥、放逐
misrecognition /ˌmɪsrekəɡˈnɪʃən/ n. 52:39
错误识别、误认
ingrained /ɪnˈɡreɪnd/ adj. 52:39
根深蒂固的
reconceptualizing /ˌriːkənˈseptʃuəlaɪzɪŋ/ v. 53:49
重新构想、重新概念化
interdiscipline /ˌɪntərˈdɪsəplɪn/ n. 53:49
交叉学科
vaccines willing phr. 53:49
仿「God willing」的戏仿:如果疫苗允许的话
理解自测 · 11 题
1. Crawford 为什么用内华达州的锂矿来开启一本关于 AI 的书?

因为她主张要「让 AI 落地」,必须去到它被制造出来的地方,而且是最完整意义上的「制造」。锂矿(Silver Peak)是美国唯一在产锂矿,产出的锂进入 iPhone、特斯拉等设备的锂离子电池,是 AI 硬件的物质起点。视频第 7–8 段她说明这一选择:AI 不是「非物质算法」,而是消耗稀土、钴、锂的采掘性产业,并提到拜登政府刚发布的关键矿产供应链危机文件,说明这些矿物在可用性和地缘政治上都处在临界点。

2. 什么是「伪自动化」(fauxtomation)?视频中举了什么例子?

这是 Astra Taylor 的术语,指被宣传为自动化、实际上靠人力在背后支撑的系统。视频第 10 段 Crawford 举了 x.ai 的例子:该公司让真人假扮数字助理,每天工作 14 小时,只为营造「无缝 AI」的印象。她同时提到 Mary Gray 与 Siddharth Suri 的《Ghost Work》,指出这类远程平台上的数字计件工拿的是贫困线以下的报酬,工作内容极其枯燥且高压。这一段的作用是把 AI 的「智能」外表与它依赖的隐形劳动并置。

3. 《Datasheets for Datasets》这篇论文的灵感来自什么?它想解决什么问题?

灵感来自硬件的「数据表」:一块半导体会标明只能在什么温度、什么条件下使用,而训练数据却没有任何对应的说明文档。第 20–21 段 Crawford 回忆与 Timnit Gebru 的早期对话,两人惊讶于业界完全没有记录训练数据来源的标准方式。论文试图明确数据集的来源、历史、适用目的,以及更关键的「何时不该再用」。她进一步指出,被下架的数据集仍在 Academic Torrents 等处流传并训练生产系统,弃用机制的缺失是治理空白。

4. Crawford 说她对「偏差」的看法发生了什么变化?为什么这个变化重要?

她原本(2010 年左右)和业界一样认为偏差是个可以靠「收集更多数据」解决的问题,但十多年的证据显示恰恰相反:超大规模系统偏差实例比比皆是(Apple Card 给女性更低额度、语音识别听不懂女性、人脸识别认不出深肤色)。第 24 段她把这些显眼错误比作「巨型动物」,主张要再往下走一层,进入「构建逻辑」。这一转向重要在于:真正的伤害往往不是壮观的失效,而是细颗粒度地决定人如何被理解和赋值,因此单纯的技术修补无法触及。

5. 为什么 Crawford 认为 FAccT 之类的公平性研究「必要但不充分」?

必要,是因为它们让偏差问题获得了关注和技术工具;不充分,是因为很多论文把公平当作纯技术问题,以为调整数据或算法就能修复。第 25 段她指出,当系统摄入过去的数据时,也吸收了历史中的结构性不平等和看世界的方式,这不是重新加权能消除的。第 22–23 段主持人也铺垫了这一点:分类体系本身具有认识论权力,是政治性的。所以她主张这是「社会技术问题」,需要不同领域的专家同席,以及监管而非仅靠技术。

6. 从巴贝奇到贝索斯,Crawford 想说明算法管理的什么特点?

她想说明「用监控榨取效率」的设想有近两百年的连续性,而 AI 只是让它变得空前细致。第 12 段:巴贝奇除差分机外还写过大量社会理论,设想用系统追踪工人每一分钟以确保被最高效地使用;泰勒主义的秒表是其延续。贝索斯 2021 年宣布用新算法系统追踪工人的韧带与肌肉来应对工伤,Crawford 称之为「巴贝奇的设想成真,而且相当噩梦」。要点是:变化不在于有没有监控,而在于颗粒度——这也是主持人在第 13 段所强调的。

7. Crawford 说「我们过多聚焦伦理、过少关注权力」,她具体反对的是什么?

她并非否定伦理本身——她在第 33 段明确说伦理与权力在哲学上始终交织。她反对的是伦理准则被当作「搪塞监管」的工具:企业说「我们有伦理准则,所以不需要法律和护栏」。这不够,因为准则缺乏问责机制,无法保证出事后有人负责、不再重演。第 34 段她给出替代的「权力透镜」:这个系统是否让强者更强?它如何与资本、警务、军队等既有权力接合?她认为这比「不造成伤害」的高层宣言更可检验,因为准则一再未在现实中兑现。

8. 为什么 Crawford 认为「AI 讨论总把技术放在中心」是个问题?她提出了怎样的替代顺序?

第 36–37 段她以老年照护为例:议题被设定为「让老人独自在家更久,再找系统填补」,自动化被当作首选方案。她指出这种讨论假定技术要么是万灵药、要么是问题本身,却跳过了更基本的问题。她的替代顺序是:先问「我们想生活在什么样的世界」「如何增进公平与正义」,再看技术如何服务于这一愿景,而不是由技术驱动议程。她补充历史论据:社会变革很少只是一次技术修补,互助等框架可能比自动化更合适。

9. 如果有人反驳说「多样性只是政治正确,与系统质量无关」,Wajcman 和 Crawford 会如何回应?

Wajcman 在第 29 段的回应是认识论层面的:设计只能来自设计者的经验,若输入的经验范围狭窄(年轻男性为主的团队),想象力与创造力的运用就会受限,「这对我们所有人都不好」。Crawford 在第 41 段补充了经验证据:研究这些问题的群体比做算法的群体多元得多,因为经历过被边缘化、被系统误认的人对这些问题格外警觉。她还引用 1950 年代梅西会议上 Margaret Mead 与工程师同席的历史,说明「技术优先」是后来的偏离而非常态。两人都把多样性视为系统质量问题,而非道德点缀。

10. Crawford 的「生命周期」框架如果应用到今天的大语言模型上,哪些环节会更突出?

框架的四个环节——算力/矿物、数据标注、数据来源、部署——在大模型时代都被放大。她在第 39 段已预见「降低机器学习能耗」会成为关键议题,大模型训练与推理的电力与水消耗正是这一点的延续。数据标注环节对应 RLHF 中的人工反馈与内容审核外包,与她描述的加纳、菲律宾众包劳动同构。数据来源问题(第 20–21 段)转化为网络爬取语料的版权与「弃用」争议。她的核心论点——把系统放回具体地点与机构去看——在大模型上依然成立,甚至更迫切,因为她第 42 段强调这些系统已影响数十亿人。

11. 「参与式监管」与硅谷流行的用户测试(UX 研究)有何本质区别?这一区分放到政府采购 AI 的情境中意味着什么?

第 34–35 段主持人指出用户小组是「狭窄的使用情境」,找几个用户测试产品,而 Crawford 强调应被咨询的是「受技术影响的人」,包括那些不是用户却被系统分类、评分的人。她提出算法影响评估和「拒绝的政治」:社区有权对政府部署人脸识别、招聘 AI、刑事风险评分说「不」。放到政府采购情境中,这意味着决策不应只在供应商与采购官员之间完成,而要有公共讨论环节,且法律不应由科技公司起草——她指出当前科技行业与被服务人群之间存在惊人的权力不对称,这种方式无法把我们带到该去的地方。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.096Is Reality a Controlled Hallucination? - with Anil Seth 下一期 · NO.098 →Slavoj Žižek - The Reality of the Virtual (2004)
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY 内容仅供学习 · thesophielab.com