'Weapons of Math Destruction' interview with Cathy O'Neil | Vlog · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

'Weapons of Math Destruction' interview with Cathy O'Neil | Vlog

节目发布 2022-01-28 · LSE Department of Mathematics
凯茜·奥尼尔 主主持人
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2016 年秋,《算法霸权》(Weapons of Math Destruction)出版不久,作者凯茜·奥尼尔应邀到访伦敦政治经济学院数学系,与该系访谈者进行了这场对谈。奥尼尔是数学家出身,先后任教于哥伦比亚大学巴纳德学院,任职于对冲基金德劭(D. E. Shaw),再转入数据科学行业,后参与「占领华尔街」运动,并以博客「数学宝贝」(Math Babe)为人所知。谈话从她亲历金融危机的经过说起,一路谈到算法如何在暗处伤人、监管为何失灵、专家为何失信,最后落在占领运动的余绪和写博客的心得。本文依据现场录音编译整理。

从数学系到对冲基金

主持人:今天我们请到了凯茜·奥尼尔。欢迎来到伦敦政经,凯茜。能否先介绍一下你自己和你的经历?

奥尼尔:好的。我受的是数学训练,在哥伦比亚大学的巴纳德学院当过几年教授,后来决定走进「真实世界」,2007 年进了金融业。刚进去,一切就开始崩塌了。所以那段时间待在金融业,很震撼,也很有意思。我在一家叫德劭的对冲基金干了几年,算是坐在前排,看清了什么地方出了问题,尤其是金融危机本身。

奥尼尔:我现在回头看,这场危机的底部埋着一个数学谎言,就是给抵押贷款支持证券(mortgage-backed securities)打分的那套风险模型。它把 AAA 的最高评级发给了这些证券。作为数学家,我为此感到羞耻,我想在座各位也会同意:数学本来应该让信息变得清楚,而不是把信息搅浑。所以我后来转去做了几年风险管理,想在风险模型上出点力。我在风险价值(Value at Risk)模型上干了几年,也做过信用违约互换(credit default swaps)。这段经历我自己认为是失败的。

奥尼尔:我对风险管理这一行也逐渐幻灭,最后彻底离开金融业,把手洗干净,开了「数学宝贝」这个博客。占领运动一开始我就加入了,到今天我还有一个占领小组,每周日碰头讨论。但我总得有个谋生的工作,于是进了数据科学,正好赶上这波热潮刚起。我本来是希望参与一件不会对社会造成破坏的事,就进了网络广告这个行当。可没多久我就发现,同样的错误、同样的问题,在这里又被制造出来了。我开始琢磨一套理论:什么东西会出问题,出的问题会跟金融危机平行或类似。等我有了这套关于「破坏性算法」的理论,我一下子就到处都能看见它们。又过了不到一年,我觉得必须敲响警钟了。

辞职写书

奥尼尔:于是我辞了职。我是靠写一本书来学习怎么写书的。我先写了一本讲数据科学的技术书,叫《数据科学实战》(Doing Data Science),找了代理人,把书稿卖出去,然后才真正开始写这本《算法霸权》。

主持人:也就是说,你写第一本书的时候,心里已经另有打算了?

奥尼尔:对。我当时的想法是这样的:第一,弄明白出版业是怎么运作的;第二,学会怎么写书;第三,这本书本身可以充当一份资历,人们可以说,「写过那本数据科学教科书的女人,现在写了一本书来批评数据科学」。因为我知道,你要拆掉一样东西,最好先弄懂它是怎么搭起来的。

主持人:这一条我最喜欢。

奥尼尔:顺便说一句,我不觉得这三步有哪一步真的照计划实现了,但当时的计划就是这样。事实证明,写那样一本书,和写《算法霸权》这样一本书,是两种完全不同的过程。

什么是「数学杀伤性武器」

主持人:你想出了一个很妙的书名,「数学杀伤性武器」。能说说这个名字是怎么来的,它指的是什么?

奥尼尔:我得坦白,这个名字是我的朋友亚伦·艾布拉姆斯(Aaron Abrams)替我想的,他也是数学家。我告诉他我的原名,工作书名叫「大数据的暗面」,他就想出了「数学杀伤性武器」。我喜欢它,是因为我跟他讨论过一件事。我最早碰到的「数学杀伤性武器」之一,是针对教师的「增值模型」(value-added model)。那是一套统计上漏洞很多的教师评估体系,给教师打零到一百分,打出来的分数几乎是随机的。我有一位朋友在布鲁克林当高中校长,她想把公式要来给我看,得到的答复是:「这是数学,你看不懂的。」我把这件事讲给亚伦听,说:他们在把数学当武器使。在数学家之间,这是不能接受的,这是把数学武器化。他就是在那时候想出了这个书名。

奥尼尔:笼统地说,一件数学杀伤性武器,就是一种算法,它以数学的名义,被当作武器,用来对付它所瞄准的人。这些算法几乎全都是给个人打分的评分系统。这一点很重要,它区分了两类算法。金融算法本质上是在预测市场,市场对大多数人来说是抽象概念;而数据科学算法是在预测人的行为。所以当我想到这层平行关系时,我的感受是:你可以毁掉一个市场,那很糟糕,但人们会注意到;你毁掉的是人,那显然更糟糕,可你未必注意得到。因为受害的是一个个的个人,他们失去的是自己的机会,比如他们想得到的那份工作,想申请的那张信用卡。这里不会有一场飞机失事,所有人都在新闻里看到。只有一个个孤零零的人,在自己的办公室里、自己的家里,被拒绝。

奥尼尔:所以,我说的是那些应用广泛、影响重大的算法。我想聚焦在真正关乎人们切身利益的事情上,同时这些算法又是保密的。只要同时满足这两条,你马上就会意识到:如果它是保密的,又在替很多人做重大决定,那么它几乎从定义上就是不受问责的。

大学排名的连锁反应

主持人:比如大学排名,或者银行体系?

奥尼尔:大学排名我确实谈了不少,大概在第三章,讲的是《美国新闻与世界报道》(U.S. News & World Report)的大学排名。我之所以谈它,是因为它太完美地展示了一件事:一个算法的副作用,它的外部性,说到底就是它的设计选择,能造成多么夸张的后果。你看这个排名这么多年来形成的长期反馈回路:各个大学的院长和行政部门费尽心思,只为了在排名上好看,可他们做出的这些选择,极少有哪一项真正改善了教育。所以我认为它大概算得上一件数学杀伤性武器。

奥尼尔:不过它有个耐人寻味的地方。赋予它权力的,是我们,是我们这些大学教育的消费者。而大多数数学杀伤性武器(再说一次,重要、保密、有破坏性)的权力,来自使用它的人。它们背后通常有一种权力关系:公司对雇员有权力,公司对求职者有权力,司法系统对被告有权力。可大学排名这个例子不一样,是我们太信它了,才把权力交给了它。所以我觉得,这一点我们应该克服,也许我们真的会克服。

激励与意外后果

主持人:这就引到我想问的一个问题:激励与意外后果。即便算法并不保密,比如用学生评教来评价教师,或者按某几项指标给大学排名,人人都知道指标是什么、算法怎么算,它照样可能鼓励一些与排名初衷背道而驰的行为。拿学生问卷来评价大学讲师,在这里以及英国大多数大学都很受重视,可它把「学生满意度」和「教学效果」混为一谈了,这两件事未必是一回事。

奥尼尔:完全正确。

主持人:这样一来,它可能鼓励一些不利于学生学习的行为:不去把学生推得足够远,把目标定得太容易。人会钻空子。这类情况虽然没有保密这一条,算不算某种意义上数学杀伤性武器的子类?

奥尼尔:不算,因为按定义它们不保密,所以不是数学杀伤性武器。但它是有破坏性的。我并不是说,凡是不属于数学杀伤性武器的东西就都没问题。数学杀伤性武器之所以在本质上更成问题,是因为我们不理解它,也就更难反驳它。而且再说一遍,我绝不认为学生评教是判断一位教师好坏的好办法。我自己就有体会:同一门微积分课,秋季开和春季开,来的学生完全是两类人。秋季来的是求知欲旺盛、真心热爱数学的大一新生;春季来的是快毕业的大四学生,这门课挂了四次,不修完就毕不了业。两种态度截然不同。所以我会第一个站出来说,用这种东西做高风险决策是不行的。但好消息是,作为教授,你自己可以把这番道理讲出来,而且教授们其实挺会讲这番道理。所以我认为,这类东西的权力是有上限的。

算法的好处与反建制的误读

主持人:你大部分时间都在集中谈算法结果的负面。我猜你也看到算法有正面和负面两方面,只是人们大多已经理解了它潜在的好处?

奥尼尔:是的,市场部门替我把那部分工作做了。而且通常是夸大其词。话说回来,有很多算法我每天都在用,也真心欣赏。我不是要说我们应该停止使用算法。

主持人:你是说,有些话必须有人说,而别人不说。

奥尼尔:对,我就是屋子里那个愿意站起来说「这样不行,这有问题」的人。不过我自己内心也有矛盾。我坦白一件事:我的书在布赖特巴特(Breitbart)上得到了正面评价。那绝不是我的目标。我猜这本书在某些方面确实有点反建制,但我自己当然不是这么看的。我想要的,是让我们的数据科学里真有科学。我希望你能拿出证据,证明这个数据科学算法确实管用,而不是光告诉我它管用,然后指望我盲目相信。

奥尼尔:所以我的意思有两层:第一,这些东西并没有像你们说的那样管用;第二,请拿出证据。可你去看布赖特巴特那篇评论,它写的是「凯茜·奥尼尔说算法不管用」,然后就停在那儿了,不走第二步。因为他们的目标是瓦解人们的信任,而那不是我的目标。我的目标是让你的信任建立在真相之上、建立在科学证据之上。这才是我想去的地方。我觉得这是一个非常基本的要求:你要根据这些分数开除人,那就先给我看证据,证明它有效。你会惊讶地发现,在大数据算法的世界里,问责少得可怜。

问责与对专家的不信任

主持人:我想深入谈谈「信任」这个概念。我猜,我们让算法而不是人来做决定,原因之一是:一位银行经理对客户说「抱歉,电脑说不行」,比说「这是我做的决定,我为此负责」要容易得多。这里面是不是有一层意思,专家把自己的责任推给了算法?这是否又利用了这样一种心理,即人们被引导去相信,既然是一套确定的程序,那它就一定是对的?

奥尼尔:绝对如此,你说到点子上了。我越想越觉得,我该去找一位研究官僚制度史的学者聊聊,因为我总是绕回到这一点:作为一种官僚机制,算法的设计简直不能再完美了。它正好填了那个位置,让办事员可以说「别怪我」,然后往肩膀后面一指,指向一个谁也解释不了的东西,就像墙上挂着的一块牌子:「您的分数不够」。这种隔着一臂距离推卸责任的便利,正是算法如此流行、并将继续流行的原因之一。再加上可扩展性。可扩展性当然是关键。你看,有一家公司做人格测试,很多很多家大公司都用它来筛掉他们不喜欢的求职申请。这很快就变成一套铁板一块的系统:那套评分系统不喜欢你,你就找不到工作。所以,可扩展、高效率,但最要紧的是:不是我的问题,不是我的责任,别找我问责。

主持人:反过来看,人们本来就对专家越来越不信任了。我说不好这对数学杀伤性武器是好是坏。比如英国脱欧投票之前,政治人物迈克尔·戈夫(Michael Gove)说过,这个国家已经「听够了专家的话」,他就是这样把经济学家们的意见一笔勾销的。你觉得人们有没有可能起来挑战数学杀伤性武器,去问:你这个决定是根据什么做的?有什么证据表明它是对的?有希望吗?

奥尼尔:这是一场危机。我不太确定你问的是什么,你是问人们会不会起来反对这本书,还是反对某个具体的算法?

主持人:两者都问。人们会挑战算法吗?

奥尼尔:是这样的,我自己对此也很矛盾。我希望会,或者说,我希望他们会去要证据,而且我希望他们知道证据长什么样。我显然是在为「更多证据」而呼吁。可至少在我的国家,情况已经到了危机的地步,你都不敢肯定人们还认得出什么是证据,因为他们把假的事实(我应该直接说,就是谎言)当成证据来相信。这就成了一个相当根本的问题。我其实在考虑再写一本书,谈「什么是证据」。因为在某种意义上,我觉得,特朗普当选之后,我原本是在说:我们现在在这里,我希望我们在科学严谨性上往上走到那里;可实际上我们根本不在我以为的位置,我们在更低的地方。那我就想,也许我们得先往上挪到那个位置。所以这事很棘手。我当然不想写一本书,在科学素养这件事上给「无政府状态」再添一把火。我希望人们成为更科学的科学消费者,理想的状态是,作为消费者,主动要求科学证据。但这很难。

展望:监测与审计

主持人:那么往前看,谈到证据,未来有没有可能让算法通过机器学习之类的手段,以理智的方式变得自适应、能响应?我想你在书里以及我看过的访谈里都说过,事情并不是这样运作的。因为算法如果说在学习,也只是从历史数据里学。拿银行贷款来说,一个算法,如果从一开始就不向某个被认为违约率高的街区放贷,它就永远不会观察到那里的人其实会还钱,也就永远不会知道向他们放贷是可以的。所以我们是不是被卡住了?即便算法自称能自适应,它也得不到正确的反馈。

奥尼尔:对。问题有很多种,这是其中之一。但我认为其中很多都可以解决,办法是对算法进行一种新型的监测和审计,这是我们至今没有做过的。所以我是有希望的,其实我非常乐观。我们拿招聘算法来想一想。现在白领岗位的简历筛选算法到处都是:你简历里有什么关键词,你有什么经历。它们比关键词搜索多走了几步,但说到底就是用来过滤申请的算法。可以说,如果它们是根据「过去谁被录用了」这样的历史数据训练出来的,那么这些算法很可能已经学会了带有性别歧视的做法。就当这是一个思想实验。但这其实很容易处理:你给它装一个监测器,看看在按人的标准判定为合格的女性和男性当中,各有多少比例通过了这套系统。如果发现有偏差,你就可以调整。这不是什么解决不了的问题。

主持人:你是说,人们没有在看这些东西?

奥尼尔:我认为他们只是在信任它们。我的意思是,我并不要求完美。我要求的是:我们能不能复核一下,确保那些再明显不过的偏差没有一路漏过去、被不断复制?我不认为我们现在在检查,但我们可以开始检查。

主持人:这让人意外,考虑到我们在其他方面要做多少层保护。

奥尼尔:对,这正是有意思的地方。而且我认为这不是巧合:算法受到的问责比人的流程还少。我们把算法当成客观的数学对象来拥抱,于是人们很容易说:「我们把这个有缺陷的人为流程拿掉,换成算法。」结果换上去之后,事情变得更糟,而不是更好。

主持人:而且变得无法挑战。

奥尼尔:对。这是它的复杂性和不透明带来的部分问题。但我的理论是,你并不需要把代码公开,或者做类似的事。你需要的是对这套系统建立持续的监测,检查它的输出,看总体统计:多少女性,多少男性,多少白人,看看它的运作方式是不是我们认为合理的。

主持人:这正是我们对人做决定时的做法。比如大学招生,大学会监测录取的学生来自哪些类型的学校,把这些数字管住。

奥尼尔:正是如此。所以我要说的其实就是:我们没有像要求人的流程那样,去要求电脑流程负责。这很可笑。

主持人:那怎样才能把这类系统管起来?

奥尼尔:说实话,我谈的大多数东西,因为我坚持只谈重要的事,所以通常本来就是受监管的。法律是有的,招聘、信贷、保险,大多数领域都有反歧视法。所以我要求的,其实只是执法。而现在的问题是,没有哪个监管机构具备监测算法的技术能力,它们甚至不知道该怎么提出自己需要什么。我觉得,哪怕监管者只是说一句「你们必须做这项检查,把结果交给我们看」,那也是一大步。就像大学不会告诉你具体怎么做检查,但你还是得做。这会是一种非常容易的做法。

危机中的幸存者心态

主持人:你刚才提到,你进入金融业的时机特别有意思。我们常听到有人说,金融危机应该怪数学家,因为他们弄出了那些极其复杂的金融产品定价方法。这公平吗?

奥尼尔:我认为,去责怪那些假装抵押贷款支持证券很安全的人,是绝对公平的。他们是应用数学家,在标普、穆迪、惠誉工作。而且还有整整一个行业的人,知道这件事却什么都不说。所以每当我听到有人说这是「黑天鹅事件」,说它太超乎寻常,谁也预测不到,我想说:很多人都预测到了。

主持人:这就有意思了。至少我当时得到的普遍印象是,出问题的是那些极其复杂的金融产品,那些衍生品,除了一小群数学家谁也不真正懂它。问题在于不懂数学的人以为自己懂了,其实没懂。可你说的是,情况比这还糟。

奥尼尔:穆迪内部当时有电子邮件传来传去,他们心里清楚。他们说的大意是:「我们连一头牛都能评成 AAA。」所以他们知道,圈子里的人都知道。不该有任何篡改历史的说法。事实上,我在纽约洛克菲勒中心的彩虹厅参加过一次活动,是德劭为员工办的,请的是拉里·萨默斯(Larry Summers)、艾伦·格林斯潘(Alan Greenspan)和罗伯特·鲁宾(Robert Rubin)。萨默斯当时在德劭工作,我算他的下属,他们三位是来看他的。我坐在前排,听他们三个人谈那些糟糕的抵押贷款支持证券和证券化产品。他们不知道具体会发生什么,但他们知道会出坏事。而这是在危机正式开始之前。

主持人:这个行业现在还是这样运作吗?

奥尼尔:我已经不在那里工作了。但我可以告诉你,我在同事身上看到的态度是:「我先尽量多赚钱,然后搬到犹他州去住,带着我的枪。」一种生存主义者的心态,古怪、末世的调子。我这么说吧:反正不是「我们应该去警告公众」。

占领华尔街

主持人:你之前简单提到了占领华尔街。能说说它是怎么开始的,现在又是什么情况?

奥尼尔:我是从新闻里听说占领运动的。不对,我收回这句话。我有个在金融业工作的朋友,每天上班都要穿过祖科蒂公园(Zuccotti Park)。占领运动第八天的时候,他在我的「数学宝贝」上写了一篇客座博文,基本上是在拿那些嬉皮士开玩笑,我当时也觉得挺好笑。不久之后,我去了现场,赶的是早上九点股市开盘时的游行。我本想找人聊聊,可每次下去,那些人对金融一无所知,我有点看不上他们,心里想:你们连这个都不懂。

奥尼尔:后来我又多想了想,意识到:其实谁都不懂。金融体系太庞大了,比它本该做的事大得太多。我总是想起保罗·沃尔克(Paul Volcker)的那句话:唯一真正的金融创新是自动取款机。那他们在那边到底在干什么?我在那里干了四年,也答不上这个问题,金融业的大部分东西我自己也解释不了。于是到某个时候我想通了:你其实不需要了解一个黑箱内部的每个零件,就能知道这个黑箱出了故障。这后来差不多成了我的座右铭。说来有点奇妙,我的想法是,无论这个黑箱是金融体系,还是某个具体的算法,我们都可以通过看输入和输出来分析它。

主持人:没错。

奥尼尔:所以,就因为人家说不清那些细节而给人打低分、看不起人,是不对的。他们清楚地知道输出是什么:没有工作,背着学生贷款。这个体系对整整一代人都不管用。我开始真正同情这种看法,这就是我加入的原因。

奥尼尔:至于我们现在在做什么,答案是很分散。我想最恰当的说法是,当年的占领者,几乎没有人还留在「占领」这个名字之下了,因为还在聚会的小组已经很少,我的小组也许是最后一个。但他们都在一个进步派活动人士的网络里,我们交流、沟通、组织,很多人在左翼阵营里非常活跃。

主持人:你在写书。

奥尼尔:我在写书,这是我搞活动的方式。我主持周日的会议。还有一个在哥伦比亚大学的联络人,一位经济学家,六年来一直帮我们在哥大订会议室,我打算和他合写一篇论文,讨论怎样以各种形式把工人组织起来。总之,这是一个团体,一个社群,这是最贴切的说法。它未必还叫「占领」了。

写博客的心得

主持人:最后一个问题。我们这里刚开了一个数学博客,而你有一个非常成功的数学博客。能给我们一些建议吗?

奥尼尔:我写过几篇关于写博客的博文,算是「元博客」。有一篇你们可以看看,叫「新年,开个博客」,是讲新年决心的。我把最重要的几条建议快速说一遍。第一条,坚持,每天都写。我自己现在做不到了,但你们应该做到。第二条,享受它。第三条,写你感兴趣的东西,别担心跑题,因为读者其实就喜欢这样,博客本来就是干这个的。再一条,一篇文章永远只讲一个观点,讲好它,然后停下来。有更多要说的,留到第二天。因为如果你有一批好读者,你会收到针对第一个观点的评论,你的第二个观点会因此变得更好。而且文章写太长,人们就不读了。

主持人:那一开始怎么吸引读者呢?你就是开始写,然后人们自己就找到这个页面了?

奥尼尔:我早期运气好。我写过一篇很火的博文,讲我为什么讨厌数学竞赛,这个话题居然有争议。我太讨厌数学竞赛了,讨厌到不能相信这居然有争议。因为我认为数学竞赛把女性和温和的男性淘汰得太厉害了,它让人以为数学是一项计时活动,可数学根本不是这样,这把我在数学里的乐趣全毁了。总之,长话短说,我写了这篇文章,很多人被激怒了,于是我得到了巨大的流量。那大概是我的第二十篇博文。

主持人:正如你所说,有些人拒绝了它。

奥尼尔:不过,别去做标题党。因为我真心认为,我写过的最受欢迎的一篇博文,几乎是一篇哲学性的文章,讲的是张量积(tensor products):讲我如何通过困惑、通过接受自己对张量积的困惑,学会了怎样过自己的生活。我的意思是,那样也很有趣。真的,什么都能写。有一段时间我在「数学宝贝」上开过一个性生活咨询专栏,我还挺想念它的。我一直请大家来问我关于性生活的问题,可收到的大多是「我怎样才能成为一名数据科学家」。

主持人:非常感谢你,凯茜。

奥尼尔:谢谢你们邀请我。

本期讲者
凯茜·奥尼尔数学家、数据科学家,哈佛数学博士,曾任巴纳德学院教授,2007 年加入对冲基金德劭,后转入数据科学。2016 年出版《数学杀伤性武器》,批评不透明算法对个人的伤害,并创办算法审计公司 ORCAA。
主持人伦敦政经学院(LSE)数学系的访谈者,运营一个数学博客,在访谈末尾向奥尼尔请教写博客的经验。
章节 · 点击跳转视频
0:05 从数学教授到金融危机亲历者 ▶ 正在看
2:53 「数学杀伤性武器」的命名与定义 ▶ 正在看
6:44 排名与评教:权力来自谁的信任 ▶ 正在看
10:30 要的是证据,不是瓦解信任 ▶ 正在看
11:45 算法是官僚推责的完美工具 ▶ 正在看
14:26 假事实时代,先弄清什么是证据 ▶ 正在看
15:42 不看代码,监测输出就能审计 ▶ 正在看
19:37 法律早已存在,缺的是执法能力 ▶ 正在看
20:59 金融危机:知道却不说的人 ▶ 正在看
23:35 占领华尔街与「黑箱失灵」口头禅 ▶ 正在看
26:26 写博客的四条建议 ▶ 正在看
本期论点
本期回应
5:50
一个算法只要同时保密且为大批人做重要决定,就从设计上不可问责 这套做法本来就要它技术带来的伤害,源头在哪里?
12:24
作为一种官僚机制,算法几乎是设计得最完美的推责工具 这套做法本来就要它技术带来的伤害,源头在哪里?
13:06
同一套评分系统被大量大公司采用,就会形成铁板一块的系统:它不喜欢你,你就找不到工作 这套做法本来就要它技术带来的伤害,源头在哪里?
17:45
基于历史录用数据训练的简历筛选算法,很可能已经学会了带性别歧视的做法 数据里的旧世界算法的偏见从哪里来?
18:34
算法承担的问责远少于人类流程,因为人们把算法当成客观的数学对象来接纳 这套做法本来就要它技术带来的伤害,源头在哪里?
其他论点
0:48
2008 金融危机的底层是一个数学谎言:风险模型给抵押贷款支持证券打出了 3A 评级
1:02
数学的职责是把信息说清楚,而不是把信息搞模糊
4:48
金融算法预测的是抽象的市场,数据科学算法预测的却是具体的人的行为 观察
6:56
算法的力量通常来自使用者与被使用者的权力关系,大学排名的力量却来自消费者过度的信任
9:26
学生评教不该用于高风险决策,因为不同学期的学生构成会让结果截然不同 做法
10:57
数据科学算法应当拿出有效性的证据,而不是要求人们盲目相信 做法
15:14
在推动科学严谨性之前,得先解决「什么算证据」这个更基础的问题
20:05
算法歧视的领域大多已有反歧视法覆盖,真正缺的不是立法而是执法
21:11
2008 金融危机不是无人能预料的黑天鹅事件,业内很多人早就预见到了
24:41
不必弄懂黑箱内部的全部构造,也能判断这个黑箱出了问题
01从数学教授到金融危机亲历者
0:05
well today we have cathy o'neal with us welcome to lse kathy thank you um can you start by telling us a bit about yourself and your background sure i'm a mathematician by training i was a professor at barnard college at columbia university for a couple years um when i decided to join the real world and uh i went into finance in 2007. um and then immediately everything started falling apart so that was a very you know sort of shocking but interesting time to be in finance i spent a couple years at a hedge fund called the esha sort of front row view on what was going wrong and in the financial crisis in particular like the way i look at it now is you know that there was sort of a mathematical lie at the bottom of it which was this risk model for mortgage-backed securities that gave triple a's aaa ratings to these mortgage-backed securities and you know it made me embarrassed as a mathematician i'm sure you guys would agree like what mathematics is supposed to clarify not obfuscate information so i actually went into risk for a couple years to try to help with the risk models which i considered a failure i worked with a value-added risk model for a couple years i worked on credit default swaps
今天我们请到了凯茜·奥尼尔(Cathy O'Neil),欢迎来到伦敦政经,凯茜。谢谢。嗯,能不能先请你介绍一下你自己和你的背景?当然。我受的是数学训练,在哥伦比亚大学的巴纳德学院当过几年教授,后来我决定进入真实世界,2007 年进了金融业。然后一切马上就开始分崩离析了,所以那真是一段——你知道——挺让人震惊但也很有意思的时期待在金融圈。我在一家叫德劭(D.E. Shaw)的对冲基金待了几年,算是在前排看着到底哪里出了问题,尤其是金融危机。我现在回头看,觉得这件事的底层其实有一个数学上的谎言,就是那个用于抵押贷款支持证券的风险模型,它给这些抵押贷款支持证券打出了 3A 的评级。你知道,这让我作为一个数学家觉得很难堪,我相信你们也会同意——数学本来是要把信息说清楚,而不是把它搞得更浑浊。所以我后来做了几年风险,想帮着改进风险模型,但我认为那是失败的。我做了两年风险价值模型,也做过信用违约互换。
便签引用
1:31
um became a disillusioned in risk as well and then left i'll refinance all together washed my hands of it started my blog math babe um joined occupy when when it started and i still have an occupy group we need every sunday to talk about stuff but in the meantime i needed a day job so i went into data science right when it was the beginning of the wave of the hype um hoping to be part of something that i didn't find socially destructive and i went into this the world of online advertising in it but pretty soon i discovered that the same kinds of mistakes and problems were being made and created and i started to develop a theory of like what could go wrong that would be parallel or analogous to the financial crisis and as soon as i had that theory of like destructive algorithms like i started to see them everywhere and within a year after that i was like i gotta i gotta sound the alarm so i quit my job and learned how to write a book by writing a book a technical book about data science called doing data science getting an agent selling my my manuscript and then actually writing the book so you wrote the first book the view with already know when you had another plan yeah yeah actually the way i looked at
嗯,后来我对风险这块也幻灭了,就干脆离开了,把整个金融业都放下,洗手不干了。我开了自己的博客 Math Babe,占领华尔街运动一开始我就加入了,到现在我还有一个占领小组,每周日聚在一起讨论各种事情。但与此同时我需要一份糊口的工作,所以我进了数据科学这一行,正好赶上这波炒作刚起来的时候。我本来希望自己参与的是一件不那么具有社会破坏性的事,结果我进的是网络广告这个世界。但很快我就发现,同样那些错误、同样那些问题又在被制造出来。我开始形成一套想法:什么样的事情会出错、会跟金融危机形成某种平行或者类比。而一旦我有了这套关于破坏性算法的理论,我就开始到处都看得见它们。之后不到一年,我就想:我得站出来敲警钟。所以我辞了职,用写一本书的方式学会了怎么写书——那是一本关于数据科学的技术书,叫《Doing Data Science》,我找到了经纪人、把书稿卖出去,然后真的把书写了出来。所以你写第一本书的时候,其实心里已经有另一个计划了?对,对,其实我当时的想法是
便签引用
02「数学杀伤性武器」的命名与定义
2:53
it was first of all i'll learn how publishing works second of all learn how to write a book and third of all it'll be actually a credential like people will be able to say the woman who wrote the book on data science now comes up with a book complaining about data science right because like i know that like when you're gonna try to tear down something you should be you should understand how it was built yeah right that was my favorite right and by the way like i don't think any of those steps actually happened but that was my plan it turns out writing a book like that is a very different process than writing a book like weapons of math destruction yeah you can imagine you came by this nice term uh weapons of mass destruction catch your name uh someone who could say something about how do you pick up the name and say a little bit about what that is about right well i should confess that my friend aaron abrams who's also a mathematician actually came up with a name for me um i told him the original name i the working title was the dark side of big data and he he came up with the name weapons of mass direction i like it because i had discussed with him this idea that like you know example one of the first weapon weapons of mass destruction i came across was the value-added teacher
首先,我可以搞懂出版是怎么运作的;其次,学会怎么写一本书;第三,这本身就是一种资历——大家可以说:那个写了数据科学教科书的女人,现在写了一本书来批评数据科学。因为我知道,当你要拆掉一样东西的时候,你应该先明白它是怎么被造起来的。对,这就是我最喜欢的一点。顺便说一句,我觉得这几步其实一步都没真的实现,但那确实是我的计划。事实证明,写那样一本书和写《数学杀伤性武器》这样一本书,完全是两种不同的过程。可以想象。你想出了一个很妙的说法——weapons of math destruction,跟大规模杀伤性武器(weapons of mass destruction)谐音。能不能说说这个名字是怎么来的,以及它到底指什么?好,我得坦白,其实是我的朋友亚伦·艾布拉姆斯(Aaron Abrams)——他也是数学家——帮我想出来的。我告诉他我原来的名字,工作标题叫《大数据的阴暗面》,然后他就想出了 weapons of math destruction 这个名字。我喜欢它,是因为我之前跟他讨论过这么一件事:我最早遇到的数学杀伤性武器之一,就是教师的增值
便签引用
4:10
value-added model for teachers which was this very statistically flawed assessment for teachers that would score them from zero to 100 almost randomly and when my friend who was a principal in a high school in brooklyn asked for the formula because she wanted to show it to me she was told it's math you won't understand it and i related the story to my friend aaron and i was like they're weaponizing mathematics like between mathematicians that's not okay it's a weaponization and he was that's when he came up with the title but generally speaking a weapon about destruction is an algorithm that i think of as is used as a weapon in the name of mathematics against the people that it is targeting so almost all the algorithms are actually scoring systems on individuals so that's a that's an important distinction between financial algorithms which are essentially algorithms trying to predict markets abstract concepts to most people versus data science algorithms which are algorithms that try to predict human behavior and that's why when i thought about that parallel like you can destroy a market it's a bad thing but people notice when you destroy people it's obviously a very bad thing but you might not notice because individuals lose out on their options
评估模型。那是一套统计上非常有缺陷的教师评估方法,从 0 到 100 给老师打分,而且几乎是随机的。我有个朋友在布鲁克林一所高中当校长,她想把公式拿给我看,去要公式的时候,人家告诉她:这是数学,你不会懂的。我把这件事讲给亚伦听,我说他们这是在把数学武器化——在数学家之间,这是不能接受的,这就是一种武器化。就是那时候他想出了这个书名。但一般来说,所谓数学杀伤性武器,指的是一种算法,我认为它被当作武器使用,打着数学的旗号,去对付它所针对的那些人。这些算法几乎全都是对个人的打分系统,这是一个很重要的区分:金融算法本质上是想预测市场,对多数人来说那是抽象概念;而数据科学的算法是想预测人的行为。所以当我在想那个类比的时候——你可以摧毁一个市场,那是坏事,但人们会注意到;当你摧毁的是人,那显然是非常坏的事,但人们可能不会注意到,因为受损的是一个个个体,他们失去的是自己的机会——
便签引用
5:25
for like whatever the job they wanted to get the credit card they wanted to you know use and they don't there's no sort of there's no airplane crash where everyone it's on the news yeah right yeah it's just individuals isolated in their offices in their homes being rejected yeah um so yeah so it's algorithms that are widespread and important so i wanted to focus on things that really matter to people and they're secret and as soon as you have those two you're already like oh if it's secret and it's making important decisions about a lot of people it's not accountable almost by construction things like the university banking system well i do actually talk quite a bit in the in the third chapter or so about u.s news and world report college ranking yeah the reason i talk about that is because it is such a perfect example of how um overblown the the side effects the sort of externalities of or the the design choices really of an algorithm can have like all the long term feedback loops that you see with the years college u.s news and world report in college rankings like you see you know all these college genes and administrations going out of their way to to look good with respect to the ranking but none of very few of those
比如他们想要的那份工作、他们想用的那张信用卡,而这件事不会有那种飞机失事、全都上新闻的时刻。对,就是一个个孤立的人,在自己的办公室里、在自己家里被拒绝。嗯,所以这些是覆盖面广、又很重要的算法。所以我想聚焦在那些真正影响人们生活的东西上,而且它们是保密的。一旦这两条同时成立,你就已经会想:如果它是保密的,又在对大量的人做重要决定,那它几乎从构造上就是不可问责的。像大学排名系统那样?我在第三章左右确实花了不少篇幅讲《美国新闻与世界报道》的大学排名。我讲它的原因是,它太完美地示范了一个算法的副作用——那些外部性,其实更准确地说是设计选择——可以被放大到什么程度。你能看到多年下来的种种反馈循环:所有这些大学的院系、行政部门都在千方百计地把自己在排名上弄得好看,但他们做出的那些选择里,真正能改善教育的很少,
便签引用
03排名与评教:权力来自谁的信任
6:44
choices that they make actually improve education right so i would argue that that probably is a weapon of mass destruction um but the one of the interesting things about it is that we are the ones we consumers of college are the ones that are giving it its power whereas most of the weapons of mass destruction which are again important secret and destructive um most of those are given power by the people using them like they're usually there's a power arrangement usually it's like company has power over a worker or company has power over a job applicant or justice system has power over defendant there's usually a power structure that example is like we imbue it with power by just trusting it too much and so i feel like that like we should get over that and maybe we will and that that brings me to the question about um sort of incentives and unintended consequences so even in situation where the algorithm is not secret something like um teacher evaluation if you don't bite students rv or or college ranking by certain metrics so everybody knows what the metric is and how the algorithm works but nonetheless it can encourage behavior that's not conducive to what the pacific purpose of the ranking is in first place when you think of
少之又少。所以我会说它大概算得上是一种数学杀伤性武器。不过它有个有意思的地方:赋予它权力的是我们自己,是我们这些大学的消费者。而绝大多数数学杀伤性武器——同样是重要的、保密的、破坏性的——它们的权力来自使用它们的人。通常那里存在一种权力关系:公司对员工有权力,公司对求职者有权力,司法系统对被告有权力,通常有一个权力结构。但排名这个例子是我们因为太信任它,反而赋予了它权力。所以我觉得我们应该走出这一点,也许我们真的会。这就带出我想问的关于激励和意外后果的问题。就算算法并不保密——比如教师评估,或者按某些指标做大学排名——大家都知道指标是什么、算法怎么运作,但它照样会鼓励一些无助于这个排名本来目的的行为。比如你想想
便签引用
8:02
evaluation of individual college lecturers by student surveys for example something that's very important here as in most other bridge universities so it conflates student satisfaction with effectiveness of teaching and the two things are not necessarily the same absolutely and so you know that can arguably encourage behavior that's not conducive to student learning by not pushing the students far enough by making the goals too easy people game things yeah yeah yeah now is that is does that sort of do it don't get the secrecy there but does that fall into the subcategory of wmds to some extent well no because by definition they're not secret right right so that wouldn't be a wmb but it would be destructive i mean i'm not saying that all things that are not wfds are okay okay um one of the things that sort of is inherently sort of more problematic about wmds is that because we don't understand them it's harder to argue against them and i'm not saying again i'm not saying that the student evaluations is a good way of deciding whether teacher is good because i have my own experience with teaching a calculus class in the fall versus the spring and just the different type of student yeah that shows up for those two different classes first are like oh
用学生问卷来评价某位大学讲师,这在这里和在大多数英国大学都非常重要。它把学生的满意度和教学的有效性混为一谈,而这两件事不一定是一回事。完全同意。所以你可以说,它会鼓励一些不利于学生学习的行为:不把学生逼得够紧,把要求定得太容易。人是会钻空子的。对对对。那么,这里虽然没有保密性,但它算不算某种意义上的数学杀伤性武器的子类?不算,因为按定义它们并不保密,所以那不是 WMD,但它确实是有破坏性的。我不是说凡不是 WMD 的东西就都没问题。WMD 之所以本质上更麻烦,是因为我们不理解它,所以更难去反驳它。我再说一次,我并不是说学生评教是判断一个老师好坏的好办法,因为我自己就有经验:秋季学期教微积分和春季学期教微积分,出现在这两个班里的学生类型完全不同。第一种是那种
便签引用
9:14
eager freshmen who really love math second is like you know last second semester seniors who need who failed this class four times and needed to graduate like there's very different attitudes yeah um so i'd be the first person to say that's not okay for high-stakes decisions but the the good news is like you can make that argument yourself yeah as a professor and actually professors are pretty good at making that argument so there's a there's i think there's just like a there's a limit to how much power that will have i see yeah so you you most of the time you're just focusing very much on the negative side of the use of outcomes yeah so presumably i'm guessing you see that there are positive and negative sides but people may understand largely the positive potential benefits well yes the marketing teams do that work for me yeah i mean usually over over over loan um benefits um so you know [Music] having said that like there's plenty of algorithms that i use daily and really appreciate i'm i'm not trying to and i'm also not even trying to say we should stop using algorithms you're saying there's stuff that you think needs to be said because people aren't other people right i am i am that person in the room who's like willing to stand up and say
充满干劲、真心喜欢数学的大一新生;第二种是那种最后一学期的大四学生,这门课已经挂了四次,必须过了才能毕业——态度非常不一样。所以要说这不适合用来做高利害的决定,我会是第一个站出来说的人。但好消息是,这个论证你自己就能提出来。对,作为教授你可以,而且教授们其实很擅长提出这种论证。所以我觉得这件事的权力是有上限的。我明白了。所以你大部分时候都非常聚焦在算法使用的负面。对。那我猜你也看到它有正反两面,只是人们对正面的潜在好处已经了解得够多了。是的,市场营销团队替我把那部分工作做了。通常还是过度吹嘘的好处。嗯,说到这里,其实我每天也用很多算法,也很感激它们。我并不是要——我甚至不是在说我们应该停止使用算法。你是说有些话你觉得必须有人说,因为别人不说。是的,我就是房间里那个愿意站起来说
便签引用
04要的是证据,不是瓦解信任
10:30
this is not okay and it's problematic i mean like i can't i i'm conflicted myself you know like i'll tell you what this is a confession my book was reviewed positively in breitbart that is not that wasn't my goal um but you know i guess it's kind of anti-establishment in some ways well that's not of course not how i see it right what i see is like i would like our data science to have science in it yeah i would like you to give me evidence that this data science algorithm is working rather than just tell me it works and expect me to blindly believe it so so my point would be twofold i'd say these aren't working the way you say give me evidence but the breitbart review if you look at it says hey kathy o'neil says that algorithms don't work and they don't go to that second step right because their goal is to undermine people's trust but that's not my goal my my goal is to actually have your trust founded on truth on scientific evidence and that that's where i'm trying to get and i think i think it's like a very basic demand right show me the evidence that this works if you're going to fire people yeah but based on these scores you'd be surprised how little accountability there is in the world of big data algorithms
这不行、这有问题的人。我自己其实也很矛盾。我跟你说一件让我别扭的事:我的书在《布莱巴特》上得到了正面书评,那可不是我的目标。但我猜它在某种意义上算是反建制的吧。当然那不是我看待它的方式。我想要的是:让我们的数据科学里真的有科学。我希望你拿证据给我看,证明这个数据科学算法是有效的,而不是只告诉我它有效、然后指望我盲目相信。所以我的论点是两层的:我会说这些东西并不像你说的那样有效,拿证据来。但《布莱巴特》的书评你去看,写的是:嘿,凯茜·奥尼尔说算法不管用。他们不会走到第二步,因为他们的目的是瓦解人们的信任。但那不是我的目的,我的目的恰恰是让你的信任建立在真相上、建立在科学证据上,这才是我想去的地方。而且我觉得这是一个非常基本的要求:如果你要拿这些分数去开除人,那就把证明它有效的证据给我看。你会很惊讶,大数据算法的世界里问责有多么稀少。
便签引用
05算法是官僚推责的完美工具
11:45
can i explore this idea of trust so one of i would imagine one of the reasons why we have algorithms for making decisions rather than humans is because it's much easier for say a bank manager to say to a client sorry the computer says no than it is for them to say i've made a decision and i'm accountable for it yes so is there an application in a sense of experts uh abdicating their responsible algorithms and is that taking advantage of of the fact that people are sort of led to believe that oh because it's a deterministic procedure it must be correct yeah absolutely i mean you nailed it right so the more i think about it the more i want to talk to a historian of bureaucracy because i keep coming back to this that it it's as a bureaucratic mechanism it couldn't have been perfect more perfectly designed right algorithms um fill that role where the bureaucrat can say don't blame me you know point over their shoulder to something that no one can explain like a sign on the wall yeah you have not scored high enough and so that arm's length responsibility avoidance is part and parcel to why algorithms have caught on so much and will continue that that plus scalability of course scalability is is key you have this one company
我能不能追问一下信任这件事?我猜我们之所以用算法而不是人来做决定,一个原因是:对银行经理来说,跟客户说抱歉电脑说不行,要比说这是我做的决定、我为此负责容易得多。是的。所以某种意义上,是不是存在专家把责任推给算法的现象?而这是不是在利用人们的一种错觉——因为它是一套确定性的程序,所以它一定是对的?完全正确,你说到点子上了。我越想这件事,越想找一个研究官僚制历史的人聊聊,因为我老是回到这一点:作为一种官僚机制,算法的设计简直不能更完美了。算法正好填上了那个角色:官僚可以说别怪我,然后指指自己肩后某个谁也解释不清的东西,就像墙上贴的一张告示。你的分数不够高。这种保持距离、回避责任的机制,正是算法为什么这么流行、而且还会继续流行的重要原因之一。再加上可扩展性,当然,可扩展性是关键。有这么一家公司
便签引用
13:06
building personality tests that are used in many many very very large companies to filter out um you know employee applications they don't they don't like the looks of that very quickly becomes a sort of monolithic system that if that that scoring system doesn't like you you can't get a job yeah you know so scalability efficiency but but very importantly not my problem not my responsibility i'm not accountable on the flip side of that it seems to me there's in some sense a growing distrust of experts in the first place anyway and yeah this i mean this may work against wmd i don't know you know for example we had the concept of politician michael gold say before the brexit vote here the countries had heard enough from experts dismissive this is how he dismissed the opinion of economics experts um is there any do you see any possibility that people will sort of as aware rise up against wmd by by sort of challenging them by saying well you know what basis do you make that decision yeah what's the evidence for that being right yeah and so on is there hope it's a crisis i mean i'm not exactly sure what your question is are you asking me do will people rise up against the book or is it up against a
做人格测试,被很多很多超大型公司拿来筛掉那些他们看不上的求职申请,这很快就变成一种铁板一块的系统:如果那套打分系统不喜欢你,你就找不到工作。所以,可扩展性、效率,但非常重要的是——不关我的事,不是我的责任,我不用担责。反过来看,在我看来,社会上本来就有一种对专家越来越不信任的趋势。这也许对 WMD 是件好事,我不知道。比如英国脱欧公投前,政客迈克尔·戈夫说这个国家已经听够了专家的话,他就是用这种轻蔑的方式打发掉经济学专家的意见。你觉得有没有可能,人们会随着意识提高而起来反抗 WMD,去挑战它们,问:你凭什么依据做这个决定?有什么证据说明它是对的?等等。有希望吗?这是一场危机。我不太确定你的问题是什么,你是问人们会不会起来反抗这本书,还是反抗某个
便签引用
06假事实时代,先弄清什么是证据
14:26
given algorithm both people challenge the algorithms yeah i mean i mean look here's the thing and it's it's i'm conflicted about it but like i hope so slash i hope they ask for evidence and i hope they understand what evidence looks like because obviously i am i'm trying to argue for more evidence right but i'm getting to the point of crisis in at least in my country where you're not even sure that people recognize evidence because they're they're believing false facts yes um but which aren't facts i should just say falsehoods yes um as evidence so it becomes a pretty existential problem so i mean i'm actually thinking of writing another book about the what is evidence because i feel in some sense i guess with the trump presidency i feel like i was saying we're here i want us to go here you know in scientific rigor but actually we were here yeah you know and i'm like maybe we should move up here you know so it is tricky it's tricky i don't i of course didn't want to write a book that adds to the sense of anarchy when it comes to scientific literacy i want people to become
具体的算法?后者,人们去挑战算法。是这样,我对这件事很矛盾,但我希望如此——或者说,我希望他们会要求证据,也希望他们知道证据长什么样。因为很明显我就是在为更多的证据辩护。但至少在我的国家,我们正走到一个危机点:你甚至不确定人们还认不认得出什么是证据,因为他们相信虚假的事实。是的。但那些不是事实,我应该直接说,是谎言。是的,被当成了证据。所以这变成了一个相当根本性的问题。我其实在考虑再写一本书,讲什么是证据。因为某种意义上,在特朗普当选之后,我感觉我原来是在说:我们现在在这儿,我希望我们能走到那儿,在科学的严谨性上。但其实我们是在这儿(更低的位置)。所以我想,也许我们该先往上挪到这里。所以这很棘手,真的很棘手。我当然不想写一本在科学素养上火上浇油、加剧无政府状态的书,我希望人们成为
便签引用
07不看代码,监测输出就能审计
15:42
more scientific consumers of science you know consumers like demand scientific evidence ideally um but yeah it's hard yep looking ahead then how how do we um is there any way in which talk about evidence and is there any way in which algorithms can in the future be made adaptive and responsive through machine learning or something like this in in a sensible way i mean i think you've made the point i think in your book and some of the interviews i've had with you that doesn't quite work that way because we're learning from historical data if we're learning at all in the development of algorithms [Music] take for example bank loans an algorithm is not going to know that it's okay to lend money to people in a neighborhood that was perceived to be one with a high default rate if it doesn't actually lend in the first place yeah and observe that the money gets paid by so are we sort of stuck you know even even when algorithms can claim to be adaptive because they're not getting the right feedback right right there's lots of different problems that's one of them um and i think a lot of them can be solved
更有科学素养的科学消费者,作为消费者去要求科学证据,这是理想状态。但是的,这很难。是的。那往前看,我们该怎么办?说到证据,未来有没有可能让算法通过机器学习之类的方式变得自适应、能够回应现实,而且是以一种合理的方式?我记得你在书里、也在我看过的一些访谈里提到,这其实没那么容易,因为如果我们真的在学习,那也是在从历史数据里学。比如银行贷款,如果一个算法从一开始就不向某个被认为违约率高的社区放贷,它就永远不会知道其实可以贷给那里的人,也不会观察到这些钱确实被还上了。所以我们是不是卡住了?就算算法声称自己是自适应的,它也得不到正确的反馈。对对,问题有很多种,这只是其中之一。我觉得其中很多问题是可以解决的,
便签引用
17:00
with a new kind of monitoring and auditing of algorithms that we just haven't done so i i do have hope i have i'm very hopeful actually um in the sense that like you know let's let's think about let's think about hiring algorithms right now we have a proliferation of white-collar job resume algorithms like what are the key words on your resume you know what's your experience they're like a few steps beyond keyword searches but you know they're they're algorithms that sort of filter applications and i think it's it's safe to say that they you know if they're based on historical data for who got hired in the past then they put these algorithms have probably learned sort of sexist practices let's just say as a thought experiment right but you know that is actually pretty easy to address what you do is you put a monitor on it which says well how many of the qualified you know using human being qualified women versus men actually get through this system and if you see that there's a bias you can adjust i mean it's like not unsolvable right so you're saying people aren't looking at these things i think they are just trusting them yeah i am actually like guess what i'm saying is like i'm i'm not asking for perfection i'm asking for like can we
办法是建立一种我们还没做过的、新的算法监测和审计机制。所以我是有希望的,其实我非常乐观。比如说,我们想想招聘算法。现在白领岗位的简历筛选算法泛滥:你简历上有哪些关键词、你的经历是什么,它们比关键词搜索先进了那么几步,但基本上就是用来过滤申请的算法。可以说,如果它们是基于过去谁被录用的历史数据训练的,那么这些算法很可能学到了某种带性别歧视的做法——就当作一个思想实验来说。但你知道吗,这其实相当容易处理:你在上面加一个监测器,看看在符合条件的人里——用人的判断来界定的合格者——女性和男性各有多少比例通过了这个系统。如果你看到有偏差,你就可以调整。这并不是无解的,对吧。所以你的意思是,现在人们根本没在看这些东西,只是在信任它们?对,我要说的是,我并不是在要求完美,我是在问:我们能不能
便签引用
18:21
double check that really obvious biases aren't getting through and just being propagated and i don't think we are checking but we can start checking surprisingly given the level of protection we have to do in terms of yeah well yeah and that's what's what's interesting and by the way not i don't think a coincidence that algorithms are have less accountability than human processes like we have embraced algorithms as like these objective like mathematical objects and it's really easy for people to say let's take this flawed human process and replace with the algorithm and then and then it becomes worse rather than better yeah yeah and it becomes unchallengeable and uh yes and that's that's part of the problem with the complexity and opacity of it but my theory is that you don't need to make the code transparent or something that you need to establish a sort of ongoing monitor of the system checking the outputs for aggregate statistics like how many women how many men like how many white people you know yeah um to see is this functioning in a way that we find reasonable i mean it's what we do when we make human decisions for university entries for example you know university will monitor how many of these students are from what types of
复查一下,别让那些非常明显的偏见就这么通过、就这么被延续下去。我不认为我们在检查,但我们可以开始检查。考虑到我们在其他方面要做的合规工作量,这挺让人意外的。是啊,而且有意思的是——顺便说一句我不认为这是巧合——算法受到的问责居然比人的流程还少。我们把算法当成客观的、数学的对象来拥抱,人们很容易说:我们把这个有缺陷的人工流程换成算法吧,结果反而变糟了,而不是变好。对,而且变得无法被挑战。是的,这就是复杂性和不透明带来的部分问题。但我的看法是,你并不需要把代码公开之类的,你需要的是建立一种对系统的持续监测,检查输出的总量统计:有多少女性、多少男性、多少白人,看看它的运作方式是不是我们认为合理的。这其实就是我们在人做决定时会做的事,比如大学录取,学校会监测有多少学生来自什么类型的
便签引用
08法律早已存在,缺的是执法能力
19:37
schools exactly and you know keep that under check so really all i'm saying is we don't hold computer processes accountable the way we hold human processes like and that's silly yeah indeed so how do you import systems out of the well the truth is most of the things i talk about because i insist on talking about things that are important right so usually they're regulated actually you know there are there's laws about it there's anti-discrimination laws about most of them like hiring credit insurance so really what i'm asking for is enforcement and the problem right now is that there's no regulator that has the technological capacity to monitor algorithms and they don't even know how to ask for what they need right um even i mean i feel like even if they just said you guys have to do this check and show us the results that would be a big step yeah because university doesn't tell you how to how to do your checks but you still have to do them right yeah yeah so that's that would be a sort of very easy approach you mentioned earlier um that you joined the finance sector a particularly interesting time yeah mathematicians we often hear have to be blamed for the financial crisis because they had these really complicated ways of pricing financial products
学校。没错,然后把这件事管住。所以我真正想说的只是:我们没有像问责人的流程那样去问责计算机的流程,这很荒谬。的确如此。那怎么把这套东西落到制度里?说实话,我谈的大多数事情——因为我坚持只谈重要的事——通常本来就是受监管的,其实是有法律的,反歧视法几乎覆盖了它们中的大部分:招聘、信贷、保险。所以我要求的其实是执法。而现在的问题是,没有哪个监管机构具备监测算法的技术能力,他们甚至不知道该怎么去要自己需要的东西。我觉得哪怕他们只是说:你们必须做这项检查,并把结果给我们看——那都会是很大的一步。对,就像大学不会告诉你该怎么做检查,但你还是必须做。对,是的。所以那会是一个相当容易的入手办法。你刚才提到你是在一个特别有意思的时间点进入金融业的。是的。我们常听到说数学家该为金融危机负责,因为他们搞出了那些极其复杂的金融产品定价方法,
便签引用
09金融危机:知道却不说的人
20:59
is that fair i think it's absolutely fair to blame the people that were pretending that mortgage-backed securities were safe right they were applications yes they were working at s p moody's and fitch right and then there was an entire industry of people that knew that but didn't say anything so when i hear people say that it was like a black swan event or whatever it was like way out of ordinary for many like nobody could have predicted this like a lot of people predicted it that's interesting because i mean the common at least impression i got at the time was that what went wrong was the these very complicated financial products the derivatives that nobody really understood except a very small group of mathematicians and the problem was that non-mathematicians thought they understood them and didn't but actually you're saying it was it was worse than that i mean moody's um emails were while going around like so like they knew really yeah they knew they said we would we would rate a cow i think something like we would give a triple a rating to a cow so they knew it and people on the inside knew it right there should be no like revisionist history going on in fact i was at a rainbow room um event with larry summers
这公平吗?我觉得,去责怪那些假装抵押贷款支持证券是安全的人,是绝对公平的。他们是在做应用。是的,他们在标普、穆迪和惠誉工作。而且还有一整个行业的人心里清楚,却什么都不说。所以当我听到有人说这是黑天鹅事件、说这完全超出常规、没有人能预料到——其实很多人都预料到了。这很有意思,因为至少我当时的普遍印象是,出问题的是那些极其复杂的金融产品、那些衍生品,除了极少数数学家没人真正懂,而问题在于非数学家以为自己懂、其实不懂。但你现在是说,情况比这更糟。穆迪的内部邮件当时都在流传,所以他们是真的知道?是的,他们知道,他们说过类似我们连一头牛都能给评级——我记得是说我们会给一头牛 3A 评级。所以他们知道,圈内人也知道。所以不该有什么历史修正主义。事实上我参加过一场在彩虹厅(Rainbow Room)办的活动,有拉里·萨默斯、
便签引用
22:16
alan greenspan and robert rubin who visited larry summers i worked for at the esha like those three guys were at the rainbow room in new york city and rockefeller's center for the de shaw people so i was in the front row and then three of them were talking about all the bad mortgage back securities and securitized products that was that were nobody knew it was going to happen but something bad was going to happen and they knew it and this was before the official beginning of the crisis and is it is this still the way that the sector operates i don't work there anymore but i'll tell you what like the attitude that i witnessed amongst my colleagues was you know well i'm just going to get as much money as i can yeah and then go to in you know live in utah with my gun you know it was survivalist weird kind of strange apocalyptic way it was very much let me put it this way it wasn't like we should warn the public right yeah okay uh that's all right you talked a little bit about earlier about occupying wall street yeah say something about how that started and uh what's happening with that now yeah i mean i heard about occupy through the news actually let me take that back my friend who worked in finance walked through
艾伦·格林斯潘和罗伯特·鲁宾,他们是来看拉里·萨默斯的——我当时在德劭工作。这三个人在纽约洛克菲勒中心的彩虹厅,为德劭的人办的活动,我坐在第一排。他们三个人在谈那些糟糕的抵押贷款支持证券和证券化产品,说没人知道会发生什么,但肯定要出坏事。他们知道,而且这是在危机正式开始之前。这个行业现在还是这样运作的吗?我已经不在那儿工作了,但我可以告诉你,我当时在同事中看到的态度是:反正我就尽可能多捞点钱,然后跑到犹他州去,带着我的枪过日子。那是一种带着末世感的、生存主义式的奇怪心态。我这么说吧,那绝对不是我们应该警告公众这种心态。对,好的。你刚才稍微提到了占领华尔街,能说说它是怎么开始的、现在又怎么样了吗?我其实是从新闻上听说占领运动的——不对,我收回这句话。我有个在金融业工作的朋友,每天走路穿过
便签引用
10占领华尔街与「黑箱失灵」口头禅
23:35
uh zuccotti park every day and he's he actually blogged for me on math babe guest blog like day eight of occupy and he you know was basically making fun of the hippies blog posts and i thought they were pretty funny too and i visited soon after that for the 9 a.m uh march um when the market opened and um and i was interested in talking to people but every time i went down there like the people didn't know anything about finance and i was a little bit i spurned them a little bit i was thinking to myself like you guys don't even understand this and i thought about it some more and i realized like actually nobody understands us like the financial system is so massive so over sized compared to what it it should be actually doing um you know keep coming back to the allen uh to the uh what's his name volcker paul volcker's the statement that like the only real financial innovation was the atm you know like what are they doing over there and having worked there for four years i couldn't really answer that question nor could i really explain most of finance myself so at some point i just thought to myself well i mean you don't actually need to know everything about the innards of a black box to know that black box is failing
祖科蒂公园。他还在我的 Math Babe 博客上写过客座文章,写的是占领第八天。他基本上是在拿那些嬉皮士开玩笑,我当时也觉得那些帖子挺好笑的。之后不久我就去看了,去参加早上九点开市时的游行。我本来很想跟人聊,但每次我下去,发现那些人对金融一无所知,我当时还有点看不上他们,心里想:你们连这个都不懂。后来我又想了想,意识到其实没有人真的懂——金融系统实在太庞大了,相对于它本该发挥的作用而言过度膨胀。我老是回到艾伦——不对,是保罗·沃尔克的那句话:唯一真正的金融创新是自动取款机。你知道,那他们在那儿到底在干嘛?我在那儿工作了四年,其实答不上来这个问题,我自己也没法真正解释金融里的大部分东西。所以到某个时候我就想通了:你其实不需要了解一个黑箱内部的所有细节,也能知道这个黑箱正在失灵。
便签引用
24:57
and actually that's become kind of my mantra right yeah in a kind of weird way like my feeling is like we can analyze this black box which is finance or given black box which is an algorithm by looking at inputs and outputs exactly yeah you don't you know so scoring or spurning somebody for saying they don't understand those details but they know the output for them which was for them like no jobs loss of student debt yeah um it wasn't working for an entire generation of people so i started really having sympathy for that view and that's why i joined and as for what we're doing now the answer is like it's very disparate it i i should i think i think the best way of characterizing it is that the people who were occupiers then are almost not almost none of them are actually in occupy as as such anymore because very few groups still meet mine might be the last one but they're all in a network of progressive activists and we we talk we we communicate we organize many people are very active um you know on the left so you're writing books i'm writing books that's my form of activism i run the meetings on sunday um i you know i'm actually going to write a paper with the the guy who's been in my uh my contact at columbia who's getting um who's getting our room he's an economist
其实这后来差不多成了我的口头禅。对,某种奇怪的方式上,我的感觉是:我们可以分析这个黑箱——不管是金融,还是某个具体的算法——通过看它的输入和输出。没错。所以你不该因为别人说不出那些细节就打分或者鄙视他们,因为他们知道输出对自己意味着什么:对他们来说就是没有工作、背着学生贷款。对,这个系统对整整一代人都不管用了。所以我开始真心同情这种看法,这也是我加入的原因。至于我们现在在做什么,答案是很分散。我觉得最好的描述方式是:当年那些占领者,现在几乎没有一个还在占领运动本身里面了,因为还在聚会的小组非常少,我这个可能是最后一个。但他们都在一个进步派活动者的网络里,我们互相交流、一起组织,很多人都非常活跃,在左翼这边。所以你在写书。我在写书,那是我参与行动的方式。我主持周日的聚会,我还打算跟哥伦比亚那边一直帮我们的联系人合写一篇论文,他是个经济学家,
便签引用
11写博客的四条建议
26:26
who's been getting our room for the last six years at columbia to meet so we're talking about you know how to organize workers in various forms um anyway the point is that it's it's a group it's a community it's the best way of saying it it's not necessarily occupy anymore right that's the last question yeah so we started this math blog you have a very successful math blog yeah so you can finish by giving me some tips yeah i think i totally i have i have a few blogs about blog blogging first of all metabolics yeah yeah there's one i think that you should look at which is called um like new year's start a blog it's a new year's resolution all right um and i have some tips i think the most important i'll just reel off a few of the most important tips number one is be consistent just write every day which i don't do anymore but you should um second is have fun with it third is like write what you're interested in don't worry about it if it's off topic because people actually like that that's what blogs are for third is never try to make more more than like one point yeah just one point make it make it well stop if you have more to say save the next day because what you'll get if you have a good readership is you'll get comments on that first point and your second
过去六年一直在哥大帮我们借教室开会。所以我们在讨论怎么以各种形式组织劳动者。总之,重点是它是一个团体、一个社群,这是最贴切的说法,不一定还叫占领运动了。好,最后一个问题。你开了一个数学博客,而且非常成功。对。那你最后能不能给我一些建议?好,我完全可以。我写过几篇关于写博客的博客。有一篇你应该看看,叫《新年,开个博客吧》,是关于新年计划的。我有一些建议,我挑最重要的几条快速讲一下。第一,要有稳定的节奏,每天都写——这一点我现在自己做不到了,但你应该做到。第二,要写得开心。第三,写你自己感兴趣的东西,跑题也别担心,因为读者其实喜欢那样,博客就是干这个的。还有一条,永远不要在一篇里讲超过一个观点,就讲一个,把它讲好,然后停下。如果你还有更多要说的,留到第二天。因为如果你有不错的读者群,你会收到关于第一个观点的评论,你的第二个
便签引用
27:44
point will be better and also if you make it too long people won't read it so how do you start attracting readings in the first place you just started writing and how people will come across the page yeah i was lucky early on i wrote a very popular blog post about how i hate math competitions which seems to be controversial i hate them so much that i can't even believe that's controversial um because i think math competitions are like weed out women and and nice men way too much um they give you the impression that math is a time activity which it isn't at all and it loses the fun for me all together anyway long story short i wrote this piece and a bunch of people were very provoked by it so i got an enormous amount of traffic um that might have been my you know 20th blog post okay as you were saying right some people rejected it well you know don't be clickbaity right but because i actually think the most popular blog post i ever wrote was almost a philosophical piece about tensor products you know about how i learned how to live my life by um being confused and accepting my own confusion about tensor products right okay yeah so i mean i'm just saying like that's that's fun too right it's right literally whatever i mean for a while i
观点就会更好。而且写太长了大家也不会读。那一开始怎么吸引读者呢?就是直接开始写,然后人们自然会看到这一页?我早期运气不错。我写过一篇很火的博客,讲我为什么讨厌数学竞赛——这居然是有争议的。我讨厌它讨厌到都不敢相信这还能有争议。因为我觉得数学竞赛太容易把女生、还有性格温和的男生排挤出去了,它还会让人误以为数学是一种比速度的活动,而数学完全不是这样,对我来说竞赛把乐趣全毁了。总之长话短说,我写了这篇东西,很多人被它激怒了,所以我得到了巨大的流量,那大概是我第 20 篇博客吧。好的,就像你说的,有些人不认同。你知道,不要为了骗点击而写。因为我其实觉得,我写过最受欢迎的一篇博客几乎是一篇哲学随笔,讲张量积——讲我怎么通过困惑、通过接受自己对张量积的困惑,学会了怎么过我的生活。对,所以我想说的是,那也很有意思,真的什么都可以写。我有一段时间在 Math Babe 上
便签引用
29:05
had a sex advice column on that babe which i kind of missed i was always asking people to ask me questions about their sex life but most of it was about how do i become a data scientist no well i think uh so thank you very much cassie thanks for having me very very good
开过一个性爱建议专栏,我还挺怀念的。我一直请大家问我关于他们性生活的问题,结果收到的大多数问题是:我怎么才能成为数据科学家。不会吧。好,非常感谢你,凯茜。谢谢你们请我来。非常好。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

数学家凯茜·奥尼尔(Cathy O'Neil)结合亲历 2008 年金融危机与在线广告数据科学的经验,提出"数学杀伤性武器"(WMD)概念:那些重要、隐秘且具破坏性的个人评分算法,正以数学之名逃避问责,而解药不是公开代码,而是像审计人类流程一样持续监控算法的输出。

核心要点

  • 金融危机的根源是一个"数学谎言"。 奥尼尔 2007 年离开巴纳德学院进入对冲基金 D.E. Shaw,亲眼看到给抵押贷款支持证券打出 AAA 评级的风险模型如何掩盖而非澄清信息。她之后转做风险管理(VaR 模型、信用违约互换)试图补救,同样失望,最终彻底离开金融业。
  • 数据科学在重蹈金融业覆辙。 她进入在线广告行业本想做"无社会破坏性"的工作,却发现同类错误正在复制。她由此发展出与金融危机平行的"破坏性算法"理论,一旦有了这个框架便处处看到实例,一年内决定辞职写书。为此她先写了技术书《Doing Data Science》,目的是先学会出版流程、建立"写过数据科学教科书的人"的资格再来批评它。
  • WMD 的定义是三要素:重要、隐秘、破坏性。 书名由数学家朋友 Aaron Abrams 提出,灵感来自"教师增值模型":该模型几乎随机地给教师打 0 到 100 分,而布鲁克林一位高中校长索要公式时被告知"这是数学,你不会懂"。奥尼尔称这是"把数学武器化"。她强调金融算法预测的是抽象市场,数据科学算法预测的是人的行为,被毁掉的是个体,且没有"飞机失事"式的公共事件,受害者只是在各自的家中和办公室里被默默拒绝。
  • 权力结构决定 WMD 的杀伤力。 多数 WMD 的权力来自使用者与被评分者之间的不对等关系:公司对求职者、司法系统对被告。《美国新闻与世界报道》大学排名是例外:它不隐秘,权力来自消费者自己的过度信任,却引发大学管理层围绕排名指标而非教育质量做决策的长期反馈循环。
  • 不隐秘的坏指标不算 WMD,但仍有害,关键区别在于能否反驳。 主持人举学生评教为例:它把满意度与教学效果混为一谈,鼓励降低要求。奥尼尔以自己教微积分的经历回应,秋季是热爱数学的新生,春季是挂科四次急着毕业的高年级生,同一教师得分天差地别。但因为规则公开,教授们能有效地为自己辩护,其破坏力有上限。
  • 算法是完美的官僚避责机制。 银行经理说"电脑说不行"比说"我做了决定并负责"容易得多。算法让官僚可以指着一个无人能解释的东西说"别怪我",这种"一臂之遥的责任规避"加上可扩展性,是算法流行的核心原因。一家做人格测试的公司为众多大公司筛选求职者,一旦它不喜欢你,你就找不到工作。
  • 她要的不是推翻算法,而是"让数据科学里有科学"。 她坦承自己的书被 Breitbart 正面评价,但对方只取"算法不管用"这一半,省略了"拿出证据来"这后半句。她的诉求是信任建立在科学证据上,而不是瓦解信任。她甚至考虑再写一本关于"什么是证据"的书,因为特朗普时代让她意识到公众对证据的辨识能力比她预想的还要低。
  • 反馈缺失让"自适应"成为幻觉,但监控可以补救。 贷款算法若从不向被认为高违约率的社区放贷,就永远学不到那里的人其实会还钱。招聘算法基于历史录用数据训练,很可能学会了性别歧视。但解法并不难:在系统上加一个监控器,统计有多少合格女性与合格男性通过,发现偏差就调整。她要的不是完美,而是"至少检查明显偏见没有被自动传播"。
  • 问责要求的是执法,不是新法律。 她讨论的领域(招聘、信贷、保险)大多已有反歧视法,问题是没有任何监管机构具备监控算法的技术能力,甚至不知道该要求什么。仅仅规定"你们必须做这项检查并提交结果"就是一大步,就像大学必须自查生源构成一样。
  • 数学家该为危机负责吗?该负责的是装作证券安全的人。 她驳斥"黑天鹅"说法:穆迪内部邮件称"我们会给一头牛评 AAA",业内很多人早有预见。她曾在洛克菲勒中心的 Rainbow Room 现场听萨默斯、格林斯潘和鲁宾在危机正式爆发前讨论坏账证券。同事们的态度是"尽量多捞钱然后带着枪去犹他州",而不是"我们应该警告公众"。

结论与值得注意的细节

  • "不必了解黑箱内部就能判断它失灵了"是她的座右铭。 加入占领华尔街时她起初轻视抗议者不懂金融,后来意识到没人真正懂这个体量过大的系统,包括工作了四年的她自己。看输入和输出就够了:对整整一代人而言,输出是失业和学生债务。这个思路直接延伸为她审计算法的方法论。
  • 占领运动的余波。 她的占领小组仍每周日在哥伦比亚大学聚会,可能是最后一个还在活动的小组。原来的占领者已散入进步活动家网络,她正与一位为他们订了六年会议室的经济学家合写关于组织工人的论文。写书就是她的行动主义。
  • 博客心得(Mathbabe)。 每天写、写自己感兴趣的东西、每篇只讲一个观点。她的流量突破来自第 20 篇左右一篇痛批数学竞赛的文章:竞赛过度淘汰女性和"好人",还制造数学是计时活动的错觉。但最受欢迎的一篇是关于张量积的哲学随笔,讲她如何学会接受自己的困惑。她还曾开过性咨询专栏,收到的问题却大多是"怎样成为数据科学家"。
  • 主持人引用英国脱欧前迈克尔·戈夫的"人们受够了专家",问反专家情绪是否会转化为对算法的挑战。奥尼尔的回答是矛盾的"希望如此":她希望人们索要证据,但担心他们已经分不清什么是证据。
核心句型 · 10
1. X is supposed to A, not B
“Mathematics is supposed to clarify not obfuscate information”
用「本该……而不是……」表达对某事物职能的期待与失望的落差。not 后接与 A 相反的动词,形成对比。适合批评某物偏离了本来目的。
2. learn how to X by X-ing
“Learned how to write a book by writing a book”
「靠做这件事本身来学会做这件事」,同一个动词重复出现,强调边做边学。可仿写:learned how to code by coding。
3. when you're gonna try to tear down something, you should understand how it was built
“When you're gonna try to tear down something you should be you should understand how it was built”
先建立资格再批评的方法论句式。tear down 与 built 形成拆与建的对照,适合表达「批评前先弄懂对象」。
4. as soon as you have A and B, it's already C almost by construction
“As soon as you have those two you're already like oh if it's secret and it's making important decisions. it's not accountable almost by construction”
表示两个条件一旦同时成立,结论便自然成立。by construction 是数学口吻,意为「从定义/构造上就如此」,用于强调结论的必然性。
5. I'd be the first person to say (that) …
“I'd be the first person to say that's not okay for high-stakes decisions”
先主动承认对方观点,再转向自己的重点,避免被误解为一味辩护。常与 but 连用引出真正想说的话。
6. I'm not asking for X, I'm asking for Y
“I'm not asking for perfection i'm asking for like can we double check”
「我要的不是……而是……」,用否定前半句压低对方对自己诉求的预期,再给出一个更合理、更容易接受的要求。谈判和说服中很常用。
7. you'd be surprised how little/much … there is
“You'd be surprised how little accountability there is in the world of big data algorithms”
用「你会惊讶于……有多少/多么少」引出反直觉事实。how 后接 little/much/few/many 加名词,语气比直接陈述更有冲击力。
8. the more I think about it, the more I want to …
“The more i think about it the more i want to talk to a historian of bureaucracy”
「越……越……」的经典双重比较级结构,两个 the more 各带一个从句,用来表达随思考深入而增强的念头。
9. you don't need to know everything about X to know that X is failing
“You don't actually need to know everything about the innards of a black box to know that black box is failing”
「不必了解……的全部,也能知道……」,用 don't need to … to … 表达判断的门槛。适合反驳「你不懂就没资格批评」。
10. long story short, …
“Anyway long story short i wrote this piece and a bunch of people were very provoked by it”
口语中跳过冗长过程直接讲结果的过渡语,相当于「长话短说」。放在句首,后接结论。
词汇精讲 · 102 · 按出现顺序
by training phr. 0:05
受……专业训练出身(a mathematician by training 科班出身的数学家)
hedge fund /ˈhedʒ fʌnd/ n. 0:05
对冲基金
front row view phr. 0:05
前排视角,喻指近距离目睹
mortgage-backed securities n. 0:05
抵押贷款支持证券(MBS)
obfuscate /ˈɑːbfəskeɪt/ v. 0:05
使模糊、混淆(与 clarify 相对)
credit default swaps n. 0:05
信用违约互换(CDS),一种转移违约风险的衍生品
disillusioned /ˌdɪsɪˈluːʒənd/ adj. 1:31
幻灭的、不再抱幻想的
washed my hands of it phr. 1:31
洗手不干、与之彻底撇清关系
hype /haɪp/ n. 1:31
炒作、大肆宣传
socially destructive phr. 1:31
具有社会破坏性的
analogous /əˈnæləɡəs/ adj. 1:31
类似的、可类比的(analogous to)
sound the alarm phr. 1:31
敲响警钟、发出警告
manuscript /ˈmænjəskrɪpt/ n. 1:31
书稿、手稿
credential /krəˈdenʃl/ n. 2:53
资历、资格证明
tear down phr. 2:53
拆毁;驳倒、推翻
working title n. 2:53
暂定书名、工作标题
confess /kənˈfes/ v. 2:53
坦白、承认
flawed /flɔːd/ adj. 4:10
有缺陷的
weaponizing /ˈwepənaɪzɪŋ/ v. 4:10
武器化(weaponize,把某物当作武器使用)
in the name of phr. 4:10
以……的名义、打着……的旗号
distinction /dɪˈstɪŋkʃn/ n. 4:10
区别、区分
lose out on phr. 4:10
错失、在……上吃亏
isolated /ˈaɪsəleɪtɪd/ adj. 5:25
孤立的、分散的
accountable /əˈkaʊntəbl/ adj. 5:25
可问责的、须负责的
by construction phr. 5:25
从构造上说、由定义本身决定(数学用语)
overblown /ˌoʊvərˈbloʊn/ adj. 5:25
被夸大的、过度膨胀的
externalities /ˌekstɜːrˈnælətiz/ n. 5:25
外部性(经济学:行为对第三方的副作用)
feedback loops n. 5:25
反馈循环
going out of their way phr. 5:25
不遗余力、特意去做
imbue /ɪmˈbjuː/ v. 6:44
灌注、赋予(imbue A with B)
power arrangement n. 6:44
权力安排、权力关系
unintended consequences n. 6:44
意外后果、非预期后果
conducive /kənˈduːsɪv/ adj. 6:44
有助于……的(conducive to)
conflates /kənˈfleɪts/ v. 8:02
混为一谈(conflate A with B)
arguably /ˈɑːrɡjuəbli/ adv. 8:02
可以说、按理说
game /ɡeɪm/ v. 8:02
钻空子、操纵规则(game the system)
inherently /ɪnˈhɪrəntli/ adv. 8:02
本质上、固有地
eager /ˈiːɡər/ adj. 9:14
热切的、干劲十足的
high-stakes /ˌhaɪ ˈsteɪks/ adj. 9:14
高利害的、后果重大的
presumably /prɪˈzuːməbli/ adv. 9:14
想必、大概
conflicted /kənˈflɪktɪd/ adj. 10:30
内心矛盾的
anti-establishment /ˌænti ɪˈstæblɪʃmənt/ adj. 10:30
反建制的
undermine /ˌʌndərˈmaɪn/ v. 10:30
削弱、暗中破坏
twofold /ˈtuːfoʊld/ adj. 10:30
两方面的、双重的
founded on phr. 10:30
建立在……之上
abdicating /ˈæbdɪkeɪtɪŋ/ v. 11:45
放弃(责任);退位(abdicate)
deterministic /dɪˌtɜːrmɪˈnɪstɪk/ adj. 11:45
确定性的(输入确定则输出确定)
nailed it phr. 11:45
一语中的、说得完全对
bureaucracy /bjʊˈrɑːkrəsi/ n. 11:45
官僚体制
part and parcel phr. 11:45
不可分割的一部分(part and parcel of)
arm's length phr. 11:45
保持距离的(此处指把责任推到一臂之外)
caught on phr. 11:45
流行起来、被广泛接受(catch on)
scalability /ˌskeɪləˈbɪləti/ n. 11:45
可扩展性、规模化能力
filter out phr. 13:06
过滤掉、筛除
monolithic /ˌmɑːnəˈlɪθɪk/ adj. 13:06
铁板一块的、庞大单一的
on the flip side phr. 13:06
反过来看、另一方面
dismissive /dɪsˈmɪsɪv/ adj. 13:06
轻蔑的、不屑一顾的
rise up against phr. 13:06
起来反抗
falsehoods /ˈfɔːlshʊdz/ n. 14:26
谎言、虚假陈述
existential /ˌeɡzɪˈstenʃl/ adj. 14:26
关乎存亡的、根本性的
rigor /ˈrɪɡər/ n. 14:26
严谨(scientific rigor 科学严谨性)
anarchy /ˈænərki/ n. 14:26
无政府状态、混乱
literacy /ˈlɪtərəsi/ n. 14:26
素养(scientific literacy 科学素养)
adaptive /əˈdæptɪv/ adj. 15:42
自适应的
default rate n. 15:42
违约率
auditing /ˈɔːdɪtɪŋ/ n. 17:00
审计
proliferation /prəˌlɪfəˈreɪʃn/ n. 17:00
激增、泛滥
thought experiment n. 17:00
思想实验
bias /ˈbaɪəs/ n. 17:00
偏差、偏见
propagated /ˈprɑːpəɡeɪtɪd/ v. 18:21
传播、延续下去(propagate)
opacity /oʊˈpæsəti/ n. 18:21
不透明性
unchallengeable /ʌnˈtʃælɪndʒəbl/ adj. 18:21
无法质疑的
aggregate /ˈæɡrɪɡət/ adj. 18:21
汇总的、总计的(aggregate statistics)
enforcement /ɪnˈfɔːrsmənt/ n. 19:37
执法、强制执行
regulator /ˈreɡjuleɪtər/ n. 19:37
监管机构
technological capacity n. 19:37
技术能力
keep that under check phr. 19:37
把……控制住、管住
black swan event n. 20:59
黑天鹅事件(极罕见、影响巨大、事后才被解释)
derivatives /dɪˈrɪvətɪvz/ n. 20:59
衍生品
revisionist /rɪˈvɪʒənɪst/ adj. 20:59
修正主义的(revisionist history 篡改历史叙述)
securitized /sɪˈkjʊrɪtaɪzd/ adj. 22:16
证券化的
witnessed /ˈwɪtnəst/ v. 22:16
目睹、亲眼见到
survivalist /sərˈvaɪvəlɪst/ adj. 22:16
生存主义的(囤积物资、准备末日)
apocalyptic /əˌpɑːkəˈlɪptɪk/ adj. 22:16
末日般的
make fun of phr. 23:35
取笑
spurned /spɜːrnd/ v. 23:35
鄙弃、不屑一顾(spurn)
innards /ˈɪnərdz/ n. 23:35
内部构造、内脏(口语)
mantra /ˈmæntrə/ n. 24:57
口头禅、信条
sympathy /ˈsɪmpəθi/ n. 24:57
认同、同情(have sympathy for a view 认同某种看法)
disparate /ˈdɪspərət/ adj. 24:57
分散的、各不相同的
as such phr. 24:57
就其本身而言、严格意义上
reel off phr. 26:26
一口气说出、连珠炮般列举
off topic phr. 26:26
跑题的
readership /ˈriːdərʃɪp/ n. 26:26
读者群
controversial /ˌkɑːntrəˈvɜːrʃl/ adj. 27:44
有争议的
weed out phr. 27:44
淘汰、剔除
long story short phr. 27:44
长话短说
provoked /prəˈvoʊkt/ v. 27:44
激怒、激起反应(provoke)
traffic /ˈtræfɪk/ n. 27:44
网站流量
clickbaity /ˈklɪkbeɪti/ adj. 27:44
标题党的、骗点击的
tensor products n. 27:44
张量积(抽象代数概念)
column /ˈkɑːləm/ n. 29:05
专栏
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.158Tim Wu discusses The Master Switch - Stanford Center for Internet and Society 下一期 · NO.160 →The AI Awakening: Implications for the Economy [Erik Brynjolfsson]
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]