视频库 / NO.119ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.119
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

Fairness, part 1 - Moritz Hardt - MLSS 2020, Tübingen

节目发布 2020-06-30 · virtual mlss2020
莫里茨·哈特 主持人
本期追问 · 点击跳到视频对应位置
从训练数据里删掉种族、性别等敏感属性,就能让算法公平吗?各组错误率相等与各组分数校准,为什么在现实中不能同时满足?预测被告是否会缺席庭审,真的比提供托儿和交通补贴更值得做吗?把某个公平指标当成优化目标,会不会反而诱导出更大的伤害?
归入 Ⅲ·10 公平可以被计算吗? →
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:本文整理自二〇二〇年图宾根机器学习暑期学校(MLSS 2020)线上课程,主讲人为加州大学伯克利分校电气工程与计算机科学系助理教授莫里茨·哈特(Moritz Hardt)。哈特于普林斯顿大学取得博士学位,曾任职于 IBM 研究院与谷歌,是「机器学习中的公平、问责与透明」研讨会(FAT/ML)的联合创办人,亦是教材《公平与机器学习》的合著者。本讲为其「公平性」系列教程的上半部分,聚焦分类问题中的统计公平准则及其局限。全文依据现场录音编译整理,仅删去口语枝节,问答部分保留了聊天区听众的提问。

开场与讲者介绍

主持人:大家好,欢迎回到二〇二〇年机器学习暑期学校。今天很荣幸为大家介绍下一位讲者莫里茨·哈特。哈特是加州大学伯克利分校电气工程与计算机科学系的助理教授,研究方向是算法与机器学习,尤其关注其可靠性、有效性与社会影响。他在普林斯顿大学获得计算机科学博士学位,之后曾在 IBM 研究院与谷歌任职。他是「机器学习中的公平、问责与透明」研讨会的联合创办人,也是教材《公平与机器学习》的合著者。他获得过美国国家科学基金会的职业奖、斯隆研究奖,以及 ICML 2018 与 ICLR 2017 的最佳论文奖。话不多说,有请哈特。

哈特:非常感谢介绍,也感谢这次演讲的机会。我先共享屏幕。大家能看到了吧?好。

感谢各位来听这场报告。我想给大家做一个关于公平与机器学习的小型教程,讨论在后果性决策(consequential decision making)场景中使用机器学习的局限与机会。内容分为两部分,今天是第一部分,第二部分在周四。

这是一个很难切入的话题,甚至很难说该从哪里开始。我觉得有一点必须先承认:我们生活在一个不平等、压迫与歧视无处不在的世界里,当下的政治事件就足以证明形势之严峻、行动之紧迫。而当我们用机器学习去把这个不平等世界里的种种流程形式化、规模化、加速化时,我们就有延续既有不公的危险。所以,在一个极不平等、充斥着既有不公的世界里用机器学习解决问题,必须格外小心。

不过我也想说,这里面还存在一个稍显脆弱的机会:借此重新审视决策本身,重新审视我们在各个领域是怎么做机器学习的,进而改革现有流程,而且有可能是往好的方向改。我希望机器学习也能成为这场改革的一部分。

从某种意义上说,我的教程甚至不是正确的起点。你们应该先去读那些指出人工智能与自动决策潜在问题的重要著作。近年来这方面的学术成果非常丰富,我举五本书为例:鲁哈·本杰明(Ruha Benjamin)、梅雷迪思·布鲁萨德(Meredith Broussard)、弗吉尼娅·尤班克斯(Virginia Eubanks)、萨菲娅·诺布尔(Safiya Noble)和凯茜·奥尼尔(Cathy O'Neil)各有一本。这些书对用人工智能与计算机技术去解决有社会影响的问题的危险,都做了极其重要的记述。要想了解危险究竟在哪里,应该从这些书读起。你们还应该去看乔伊·布奥拉姆维尼(Joy Buolamwini)、凯特·克劳福德(Kate Crawford)、蒂姆尼特·格布鲁(Timnit Gebru)、拉塔尼亚·斯威尼(Latanya Sweeney)、梅雷迪思·惠特克(Meredith Whittaker)等人的教程与主题演讲,他们在这一领域做了大量工作,对这项技术的危险给出了准确而重要的论述。这才是这个话题的自然起点。

我引一句本杰明的话,我觉得非常精彩。她提出了一个概念,叫「新吉姆法典」(the New Jim Code):指的是那些反映并复制既有不平等的新技术,却被宣传和理解为比先前那个时代的歧视性制度更客观、更进步。这句话点出了一件事:机器学习与人工智能会给人一种比旧式官僚流程更中立、更进步的印象,而旧式流程构成的是赤裸裸的歧视;可到头来,它只不过是把那些不平等原样复制了一遍。她对这种危险的论述极有说服力。

范围界定:后果性决策中的歧视

哈特:记住这些之后,我来交代本教程的焦点:我主要讨论后果性决策场景中的歧视。也就是说,我关心的决策者通常要做一个二元决定,比如接受或拒绝,而这个决定对当事人有实际后果:招聘、大学录取、申请信贷、刑事司法等等。你们应该以这类场景为参照。

我得说明,这一下就排除掉了世界上许多其他形式的不正义甚至不公平。不是所有的不正义都是决策场景中的歧视。所以这已经是一层范围限制。

第二个限定是,我主要从美国的视角来谈,尤其是我举的例子和法律背景。我知道在座听众的地域分布非常多样,有很多欧洲听众,也有来自世界各地的人。我的例子不一定适用于你们所在社会的法律框架,这是一种局限。但各国的法律状况本来就是碎片化的,把歧视问题当作一个全球性问题来谈并没有意义,它在不同国家呈现出截然不同的形态。

最后,我当然会讲用来处理这些问题的形式模型与框架,因为那是我的专长所在。但这绝不是要削弱或转移刚才提到的那些学术研究,它们依然成立,依然值得牢记。而且至少在我看来,这明确地为非技术性干预留出了空间,也就是不依赖预测或机器学习来解决问题的那些办法。在这个教程里,我们会遇到一些场景,你们当中许多人会被迫追问:我们究竟该不该一开始就用预测或机器学习来解决这个问题?

这就是本教程的焦点。我看到有一些问题,如果现在就有问题,我可以先答一个。好像暂时没有,那我继续。

什么算歧视:不正当的区分依据

哈特:有些人接触这个话题时会说:歧视不就是机器学习的全部意义吗?这也是这个话题早期引来的一种反应。机器学习不就是画一条决策边界,把一些点标为接受、另一些点标为拒绝吗?这必然排除一些人、接纳另一些人。既然如此,指责机器学习构成或参与歧视性做法,岂不是有点过分?

所以我要说明,本教程追究的是不正当的区分依据(unjustified basis for differentiation),而不是任何形式的区分,不是任何形式的接受与拒绝。

什么叫不正当的区分依据?一种情况是实际上不相关(practical irrelevance)。比如性取向与雇佣决定:我们不认为性取向与工作有实际关联,那它凭什么参与招聘决策?还有一种更复杂的情况,可以叫道德上不相关(moral irrelevance)。比如残障状况:为残障员工提供便利可能给雇主带来额外成本,但我们作为一个社会已经决定,这是雇主必须承担的成本;在招聘决策中,我们认定一个人有残障在道德上是不相关的,哪怕纯从统计上看它可能影响其工作表现。同样,怀孕的可能性也不应影响招聘决定,哪怕它可能减少一个人未来的工作时数。这些就是道德上不相关的例子,也正是歧视这个概念想要捕捉的东西。

还要理解一点:歧视不是一个泛用的概念,它是特定领域的。它不适用于所有领域,它针对的是那些关乎重要机会、影响人们生活的领域,而不是机器学习的任何一个无关紧要的应用。它也是特定群体的:它关乎那些具有社会显著性的类别,这些类别在历史上曾被用作不正当的、系统性的不利对待的依据。种族、性别、残障状况,这些在过去曾是歧视的靶子,至今仍是歧视这个概念的核心。所以我们关心的不是随便哪种群体身份。

这就是法律登场的地方。我只简单交代一下法律做了什么、没做什么,作为一个框架。在美国法律里,有一些受规制的领域,也就是存在某种反歧视立法的领域,也有一些不受规制。信贷是重要的一个,有《平等信贷机会法》;教育有一九六四年的《民权法案》;此外还有就业、住房、公共设施。这些都是私营部门的领域,不包括政府领域,那里的法律可能更复杂。

对机器学习来说有一点很重要:我们的很多应用,至少从经济动机上看,机器学习目前最主要的应用可能就是广告。而这些受规制的领域延伸到营销。如果你为信用卡打广告,或者为招聘打广告,那么相关规制同样适用于广告。我再次略去各类政府法律,它们有不同的规则。

美国法律承认的受保护类别包括种族、肤色、性别、宗教、国籍、公民身份等等。这些类别来自不同时期的不同立法。并不存在某一部统一连贯的反歧视法规在某个时刻诞生。这些是几十年的行动争取来的,往往是对政治运动和事件的回应,比如我们现在称作二十世纪六十年代民权运动的那段历史。这些类别至今仍然重要、相关,而且处于变动之中。

就在几周前,最高法院有一个重大判决,基本上是确认了我上一张幻灯片上的内容:针对性别的反歧视法保护延伸到性别认同和性取向。这在奥巴马政府时期被认定成立,在本届政府下曾有反复,这次判决重新确认了它继续适用于这些领域。这个例子说明,这些东西一直在演变,一直有争议,一直是政治讨论的来源。

美国反歧视法的两种原则

哈特:在美国,主要有两种法律原则。我简要说明,这样你们听到这些术语时大致知道是什么意思。要证明某人歧视了你,有两条路。

一条是差别对待(disparate treatment)。它要捕捉的是诸如有意考虑群体身份的情况:我确实查看了你在某个受保护群体中的成员身份,并据此做决定。它也试图捕捉有意的歧视,哪怕没有明确考虑群体身份,但我怀有歧视的意图和目的。可以说,这条原则追求的是程序公平:你给人分类的程序本身,表面上看不构成歧视性做法。

另一条原则叫差别影响(disparate impact)。它要捕捉更间接的歧视形式,那些可以避免的、不正当的、可能是间接的损害。它的目标更偏向分配正义,或者说尽量缩小结果上的差异。差别影响原则允许你这样论证:即便决策者的程序看起来没问题,但结果上出现了巨大的差距,你可以要求雇主为其决策造成的这种差距负责。

人们已经注意到两者之间的张力。有时候,为了避免差别影响,也就是避免不同群体在结果上出现差距,你可能需要明确把群体身份考虑进去,而这恰恰与差别对待原则相冲突。这一点也是这个领域的形式化工作一直在纠结的。稍后我们会看到,有些公平准则明确引用了群体身份,你可以追问:在差别对待原则之下,这算不算问题?它们会不会构成差别对待?

我不是法律学者,这里面显然还有很多东西,我只想让你们有个大致的把握。有几点关于法律的告诫值得记住。美国的反歧视法并不反映某一个连贯的道德理论,它也不是那个道德理论的操作化。这些立法是对社会运动、对民权运动的回应,每一部都是历经时日、经过多番争斗、在不同政党之间达成某种妥协才诞生的。它比一个自上而下的歧视概念要碎片化得多。

尤其要注意,法律没有给我们一个可以直接形式化或操作化的公平定义。作为计算机科学家,我们如果想用形式模型处理这个话题,不能指望翻开法律说「这就是定义,我们只要用正确的方式把它形式化就行」。很遗憾,事情不是这样运作的。

「无意识即公平」为何失败

哈特:那么关于这个话题,我想让你们理解的第一件事,也是我认为这个社群最先学到的一课是:不存在「无意识即公平」(fairness through unawareness)这回事。你不能指望忽略或删除敏感属性就能消除这个担忧。把所有看起来像敏感属性的东西,也就是受保护群体身份的编码,从数据里清除掉,然后祈祷一切顺利,这个想法会错得很惨。

我提到的那几本书里有大量例子,我只举一个。二〇一六年彭博社有一篇报道,讲亚马逊当日达服务的覆盖范围。亚马逊向一批美国城市推出了当日送达,以波士顿为例,你会看到恰好有一个街区被排除在外。蓝色阴影区是当日达覆盖区,灰色区没有覆盖,而那个被排除的街区是罗克斯伯里(Roxbury),一个以黑人或非裔美国人为主的街区。

人们看了这些地图就说,等一下,这太令人警惕了,因为它看起来和旧时的「红线」地图非常相似。这是一张费城的红线图:当年的贷款员手里有秘密地图,把某些区域划掉,标成红色的「危险」区,在那些街区不做生意、不放贷。《平等信贷机会法》的一部分目的就是废除这类做法,让你不能再这么干。人们看到当日达覆盖图,看出了这种相似。

这就是「无意识即公平」失败的一个例子。亚马逊几乎肯定没有看种族,那篇报道也是这么说的,报道甚至说也许他们应该看。我不在场,不知道内情,但亚马逊也许只是在预测各邮编区或各街区的购买量。购买量与社会经济地位相关,而在美国,社会经济地位又与种族相关。所以他们也许只是在预测购买量,一个标准的机器学习问题:看历史数据,跑一遍机器学习,然后画地图。这个例子说明,不考虑这些问题就直接用机器学习,会一路导致各种问题,最终撞上这种让人觉得不公平的结果。

所以,「我们的数据里没考虑那个」这种说法从来不是相关的辩护。这一点必须说得非常清楚,因为至少到二〇一六年,公司还经常拿这个理由说事:我们的数据没考虑那个,我们的模型是中立的,等等。

那么,如果忽略问题行不通,我们该做什么,或者说人们尝试过什么?这是我第一部分的提纲。我会讲一种标准视角,或者说我称之为较窄的视角:分类中的公平准则。我们会讲标准的决策理论设定与监督学习,讲人们在分类的语境下讨论过哪些公平准则,看看它们之间的关系与各自的局限。周四我会转向一种更宽的视角,主要探讨两条研究路线:一是决策的因果模型能对公平说些什么,二是社会技术系统的动态模型可能如何带来对这个问题的不同看法。

主持人:聊天区有一个提问,来自一位听众:这个议题是否应该和那个很成问题的「我不看肤色」的说法联系起来?

哈特:是的,我认为这肯定是相关的。所谓「色盲」政策的想法,确实和「无意识即公平」这一套是一路的:似乎只要你不看某人的种族或其受保护群体的身份,问题就自动消失了。所以是的,两者肯定相关。

这也是个好时机,可以多问一些关于框架的问题,因为我马上要离开背景部分,进入技术内容了。如果有人对这些还有问题,现在正合适。我给大家一点时间打字。

主持人:如果有人有问题,也可以在 Zoom 里举手,然后直接提问。

哈特:好像目前没有别的问题,那我就进入第一部分,统计公平准则。

统计公平的学术渊源与形式框架

哈特:关于公平与分类的形式化工作是从哪里来的?计算机科学绝不是最早研究这个问题的。事实上,先驱性工作来自教育测试界,也就是关心标准化考试的那批人。一位值得注意的学者是安妮·克利里(Anne Cleary),她早在一九六八年就写了一篇论文,讨论教育测试中不同的公平准则,即一场考试公平意味着什么。这和机器学习非常相似,因为归根结底你是在给人打分,试图预测某人的表现,而考试里存在群体差异,这个问题最早就是在那里浮现的。

经济学也有早期工作。加里·贝克尔(Gary Becker)一九五七年的博士论文提出了「偏好型歧视」(taste-based discrimination)的概念;菲尔普斯(Phelps)与阿罗(Arrow)在七十年代对「统计型歧视」(statistical discrimination)做了重要工作。用偏好型歧视和统计型歧视来思考歧视,至今仍是经济学家概念化这些问题的主要方式。

计算机科学基本是二〇一〇年之后才加入的,或许二〇〇八、二〇〇九年有一些早期工作,然后从二〇一六年开始研究爆炸式增长。二〇一六年是个转折点,计算机科学拥抱了公平问题,它成了我们各大会议的主要议题,人们开始广泛地做这方面的工作。

你可能会问,为什么是现在?为什么当年的工作没有解决这些问题?我认为今天算法决策的紧迫性、规模与影响范围都不一样了,算法决策这个想法走得比我们当初设想的远得多。此刻,作为机器学习研究者与工程师,我们不得不做出这些关于「某件事应当如何做」的规范性判断,而在很多情况下我们并没有做好准备。这就是为什么它现在是个紧迫的问题。

你也可以问,为什么要在机器学习的语境下研究它,而不是决策理论或经济学?那些领域也研究这个问题,也应该继续研究。但机器学习推动了算法决策的大量落地,由此催生了新的技术问题。更重要的是,这是想法落地的地方,是这些问题真正冒出来的地方。结果就是我们被迫直面它们,再也不能视而不见。

让我先回顾标准的形式化预测与决策设定,这只是标准决策理论。数据由协变量 X 描述,把它看作一个随机变量。X 可以是多维的,可以是一组特征。有一个结果变量 Y,通常是二元的,我的例子大多是二元的,它有时也叫目标变量。目标一般是给定 X 预测 Y,也就是给 X 分配一个与 Y 一致的标签。X 与 Y 是同一概率空间中的随机变量。

通常我们用监督学习产生一个评分函数,它是 X 的函数,于是成为一个新的随机变量,我记作大写的 R。评分函数把这些协变量概括成一个实值分数。然后我们按阈值规则 D 做二元决定:分数高于某个阈值就是 1,否则就是 0。可以把 1 解释为接受、0 解释为拒绝。这一切应该都很熟悉。这些随机变量在同一个概率空间里,定义在一个总体上,在我们关心的场景中总体通常由个人构成。

评分函数从哪里来?这个暑期学校里很多人在讲这个。它可以来自数据的某个参数模型,比如似然比检验,决策理论讲的就是这个:如果你有数据的参数形式,就能这样得到评分函数。也可以考虑非参数的分数,比如贝叶斯最优分数(Bayes optimal score),也就是给定数据后目标变量的条件期望,至少在平方损失下它是贝叶斯最优的。这是一个重要的评分函数,我之后会时常提到。最常见的情况是,评分函数是用监督学习从有标签的数据中学出来的:用深度学习或者任何技术,在有标签的训练数据上最小化某个损失函数。

不过对本次报告而言,评分函数怎么得来的不是重点,重点是你拿它做什么,也就是决策那部分。所以我们主要谈决策,不谈训练过程,忘掉优化和正则化,甚至忘掉泛化,因为我们谈的是总体层面。我写这些随机变量,就是因为我在总体层面讨论一切,不谈有限样本问题,不谈总体与有限样本之间的差异。

谈到决策,就有标准的混淆表:结果可以是 1 或 0,决定也可以是 1 或 0。两者都是 0,叫真阴性;都是 1,叫真阳性,这是你做对了的情况。错误的情况有两种:结果是 0 而你判为 1,是假阳性;结果是 1 而你判为 0,是假阴性。

有了这些定义就可以谈比率。真阳性率是在结果为 1 的条件下决定为 1 的概率;假阳性率是在结果为 0 的条件下决定为 1 的概率;真阴性率是在结果为 0 的条件下决定为 0 的概率;假阴性率是在结果为 1 的条件下决定为 0 的概率。这就是混淆表上的四种可能,以及由此派生的四个自然的统计量。可以把它们理解为按行归一化,所以有时叫「行统计量」,因为它们谈的是这张表的行。这只是标准图景。

那么统计公平准则做的是什么?到目前为止我们只讲了标准决策理论,根本没谈公平。公平是怎么进来的?人们把公平,或者至少把群体差异,引入这个图景的最简单方式,是再引入一个随机变量 A,编码在某个受保护类别中的成员身份。A 可以是种族的某种编码,或者性别、残障状况等等。先不管怎么编码,也先不管这种编码本身、以及一开始就收集这种数据本身可能带来的伤害。你引入这个随机变量 A,它可以与 X 相关,不必独立于 X,甚至可以是 X 的一部分,是你已有特征中的一个,只是另外给它一个字母。

有了群体身份之后,就可以谈涉及群体身份的各种统计量,进而谈群体差异,也就是这些统计量在群体之间如何不同。这正是人们所做的。这个想法至少可以追溯到六十年代克利里的工作,她当年做的正是这件事,与她当年提出的看群体统计差异是等价的。

我们会回顾三个常见准则,我逐一带大家过一遍。几乎所有已有工作都以某种方式与这三者之一相关。理解了这三个准则,文献里看到的东西大体都能在概念上归到其中一个,未必是形式上等价,但至少在概念上相关。

标准一:独立性与接受率均等

哈特:第一个也是我认为最常见的准则,是接受率均等。你可以直接要求,对任意两个群体 a 和 b,做出正向决定的比率相等。也就是说,接受率在所有群体中相同:如果一个群体有百分之八十的接受率,其他每个群体也必须是百分之八十。

可以把它推广为要求 D 与 A 统计独立,有时就叫独立性准则(independence)。你可以自己验证,如果决定独立于 A,那么这些条件接受概率必然全都相等,心算一下就能得出。你也可以把它施加在评分函数上,要求分数独立于 A,那么对分数做阈值处理,也能保证正向决定率相等。这些说法都紧密相关。要求分数独立于 A,我觉得是一种很自然的表述,它蕴含其余所有说法。当然,人们还提出了这个准则的各种变体、松弛与等价表述。

我们来对这个准则做一点压力测试:设想一些情形,其中满足这个准则并不能排除不公平的做法,决策者可以在满足准则的同时,藏起在直觉上明显不公的决定。这个准则还有别的问题,但我喜欢举的例子是这样的:你可以在一个群体里做出高质量、有依据的决定,也许你对这个群体经验丰富,而对其他群体没有;于是在另一个群体里你做的是糟糕、任意、随机的决定,但你调整得恰好让正向决定率相匹配。你可以说,在多数群体里我照常行事,经验丰富;在某个少数群体里我就抛硬币随机挑人。

我要说,这种情况可能自然而然地发生:如果你在某个群体上数据更少或者质量更差,模型在那个群体上基本失灵。举个例子,有一个叫弗雷明汉风险评分(Framingham risk score)的冠心病风险评估工具,如果我没记错,它是在二十世纪初的一个白人男性队列上建立的,后来用于其他患者,表现差得多。

在一个群体里做糟糕的决定、在另一个群体里做好的决定,这当然令人担忧。这里的道德直觉是:你不应该能拿一个群体里的真阳性去匹配另一个群体里的假阳性。我不该能说「这个群体里我有很多真阳性,为了补上,我在另一个群体里造出很多假阳性」。因为如果你随机接受,你大概率没有挑到最好的那批人,结果是给那个群体建立了一份糟糕的记录。所以我认为这类准则有这个问题。

尽管如此,这个准则非常流行,有大量工作围绕它展开。我特别想提一下里奇·泽梅尔(Rich Zemel)等人关于公平表示(fair representation)的工作。大致来说,目标是用对抗学习等深度学习技巧,训练一种独立于群体身份的数据表示,同时尽可能保留原始数据的信息。也就是从某个特征表示 X 出发,得到另一个表示 Z,让它尽可能代表 X,同时独立于群体身份。人们把这看作一种数据去偏,这个概念有一些问题,我们稍后会回到,但这方面的工作仍在持续。

我想让你们记住的是:当你看到一篇论文说「公平表示学习」之类的东西时,问问自己,它真正达成的公平准则是什么?底层的公平准则不会因为用了深度学习就变得更好。堆砌了大量机器和技巧去达成这个准则,并不意味着它就更公平,或者比独立性多了什么。掀开引擎盖,它仍然是你刚才看到的那个独立性准则。

主持人:有一个给你的问题:对二元决定而言,给不同群体设不同的接受阈值,是不是就足以满足这个准则?

哈特:这是个好问题。的确,如果你只想要独立性,有个更容易的办法:设置群体特定的阈值,把它们调到让接受率相同就行。这正好说到我的点子上:既然有这么直接的办法达成这个准则,那你指望从更复杂的办法里得到什么?应该尽量明确、精确地说清楚这一点。我认为有些论文在这方面做得比另一些好,而且我欣赏泽梅尔的工作,所以这不是针对那篇开创性的论文,而是对这一整条研究线的一般性看法。

主持人:还有一个问题:接受率均等听起来像是「结果平等」(equality of outcome)?

哈特:在某种「结果平等」的意义上,我觉得你说得对。但「结果平等」可以有不同含义,而我已经把「结果变量」这个词用在 Y 上了,所以我不太想用这个说法。不过我想你心里想的大概是对的。

标准二:错误率均等与 ROC 曲线

哈特:认识到用假阳性换真阳性这个问题令人担忧、似乎违背某种道德直觉之后,人们提出了另一组准则:错误率均等。它明确不允许你在真阳性与假阳性之间做交换,要求所有群体有相同的假阳性率和相同的假阴性率。

同样可以推广,而且推广得很漂亮:要求决定 D 在给定结果变量 Y 的条件下独立于 A。不再是简单的独立性,而是条件独立性。对评分函数也一样,要求 R 在给定 Y 的条件下独立于 A。如果你要求的是这个,其余的都随之而来:如果分数在给定 Y 的条件下独立于 A,那么用阈值规则得到的决定也满足这一点,假阳性率和真阳性率的均等也随之成立。这可能是最简洁、最温和的表述方式。

关于错误率均等(error rate parity)有一点要认识到,这让它有点棘手:它是一个事后准则。我解释一下。在决策时刻,也就是必须说是或否的时候,决策者并不知道谁是正例、谁是负例,所以无法在决策时评估这些错误率,它们对决策者是未知的。事后回看,有人可以在每个群体里收集一批最终结果为正的人和一批最终结果为负的人,看看他们当初是怎么被分类的。凭借后见之明,你可以检验这些错误率,但在决策时你做不到。

于是,在这种事后审计的场景里,群体之间的错误率差异常常让人觉得不公。这里又有强烈的道德直觉:等一下,如果一个群体里那些最终结果良好的正例,被判为负例的比率远高于另一个群体,这看上去就不公平。如果这个准则被违反,就好像某个群体承担了不成比例的不确定性负担,成了统计误差的首当其冲者。换一种说法:你希望误分类带来的伤害或不利在所有群体中相同。

可以用 ROC 曲线把它可视化。ROC 曲线是讨论阈值规则的简单办法:假设评分函数取值在 0 到 1 之间,对每个可能的阈值,画出真阳性率和假阳性率,通常得到一条曲线。你可以为每个群体各画一条,表示你的阈值规则在该群体上达到的真阳性率与假阳性率;滑动阈值,就描出这些曲线,一个群体一条。错误率均等告诉你的是:如果你的分数满足这个准则,它的 ROC 曲线必须位于所有这些曲线之下。所有曲线交集下方的区域,就是在满足这个准则的前提下可以实现的真阳性率与假阳性率的权衡。

主持人:聊天区有一位听众提问:对于招聘这类场景,似乎即便事后也无法评估这些反事实结果,那么这个准则在这类场景中还有意义吗?

哈特:问得很好。他指出,有时候一个人没被接受,我们就永远看不到其结果,没有任何记录。比如你没进哈佛,我们永远不知道你在哈佛的绩点会是多少。这一点很重要,答案是:是的,这是个必须处理的问题。以贷款为例,人们会这样估计错误率:你被这家银行拒了,但从另一家银行拿到了贷款并且还清了,那你大概也会还清第一家银行的贷款,这就给了我们估计的依据。但一般来说,这些东西就像潜在结果,你只能观察到其中一个,观察不到两个,所以通常很难估计。好问题。

回到 ROC 曲线相交的这张图,你会看到一件事:如果你强制执行这个准则,可能意味着在某个群体里你做得没有本可以做到的那么好。也就是说,你实际上让某个群体的决定变差了。我认为这是对这个准则的一个合理担忧:仅仅为了让错误率相等,你在某个群体里做得更差,而这本身可能构成对该群体的伤害,因为你明明可以做得更好,这种伤害是不正当的。这就是关于错误率的故事。

标准三:校准与充分性

哈特:回到混淆表,正如我之前说的,你也可以谈谈交换条件之后会发生什么。不谈行准则,而谈列准则:均等形如「给定决定为某值、群体为某值,Y 等于某值的概率」的表达式。也就是把 Y 和 D 对调,现在 Y 出现在事件里,D 出现在条件里,之前是反过来的。这些叫列统计量,它们在统计上同样有意义。两个标准的列统计量是错误遗漏率(false omission rate)和错误发现率(false discovery rate),取决于你选取 Y 和 D 的哪种取值。这两个量在统计学里也很重要,你同样可以问,把它们均等化会得到什么公平准则。从语法上说这样做没有任何问题。

但人们实际上做的是一件密切相关、我认为更自然、也更受欢迎的事,那就是校准(calibration)。我来解释校准是什么。一个分数 R 是校准的,如果给定分数值 r 时观察到正结果的概率恰好等于 r。这意味着你可以把分数当作概率来用,虽然它可能并不真是概率,在个体层面肯定不是,但平均而言你的分数表现得像概率。如果你看到某人分数是 0.8,你就知道,在分数为 0.8 的人当中,正结果率是 0.8。如果我看到你的心脏病风险分数是 0.9,我就知道,在分数为 0.9 的人群里平均来看,患心脏病的概率是百分之九十。

然后你可以要求分类器不仅在整个总体上校准,而且在每个群体内校准。这是一个更强的条件:多加一层对群体的条件,要求分数在每个群体内都表现得像概率。即便你只看某一个群体,该群体内的分数仍然名副其实。

想一想就能证明,按群体校准可以从一个条件独立陈述推出:结果变量 Y 在给定分数的条件下独立于群体身份。换一种说法:就预测结果这个目的而言,你的分数已经包含了你需要知道的关于群体的一切。换句话说,如果我唯一的目标是预测 Y,那么在观察到分数 R 之后,我没有兴趣再去询问你的群体身份。这是一个很自然的保证:知道了你的分数,就不必问你的群体身份。看到 0.7,它在所有群体里意思相同,我不必说「请再告诉我你的群体,如果是这个群体我加 0.1,如果是那个群体我减 0.2」。不需要做这种心理体操,分数在所有群体里意思一样。

因此,校准是一个事前的保证。决策者看到分数值 r,在决策时就知道正结果的频率是多少:分数 0.8 意味着,在每个群体内,得到 0.8 的人平均有百分之八十的心力衰竭率。你不需要询问群体身份。

但别对校准过于兴奋。它不是关于个体的保证。玛丽拿到 0.8,不意味着玛丽个人,在她精确的特征条件下,有百分之八十的心力衰竭概率。它只意味着,在得到 0.8 分的人当中平均而言,正结果的频率是百分之八十。这就是校准。

还有一件事要注意:按群体校准,以及一般意义上的校准,往往是无约束学习自然产生的结果。它不一定是对学习的一个约束,而常常是机器学习本来就给你的东西。有一个定理说,在某些条件下,偏离按群体校准的程度,也就是如果你放松这个条件之后偏离了多少,被你的分数相对于贝叶斯风险、或者说相对于贝叶斯最优分数的差距所上界。比较你学到的评分函数与贝叶斯最优分数,两者性能的差距给出了按群体校准被违反程度的上界。你机器学习做得越好,越接近贝叶斯最优分数,按群体校准就满足得越好。这来自我和刘丽迪雅(Lydia Liu)、马克斯·西姆乔维茨(Max Simchowitz)几年前的工作。换句话说,看到无约束监督学习近似地产生校准,你不必惊讶。

这里有一张图:在 UCI 成人数据集上直接跑逻辑回归,不做任何调参,就会看到这样的校准图。按分数十分位从一到十,看正结果的比率,男性和女性基本都是对角线。完美校准就是一条对角线,这已经非常接近了。

主持人:聊天区又有一个问题:分数校准与道德相关性的问题怎么联系起来?我们可能校准得同样好,但某个群体的人在现实中还是可能处境更差,怎么调整分数来处理这一点?

哈特:这是个很棒的问题,我很喜欢你的思路,你已经把道德相关性这个概念内化了。我可以用一个例子说明校准的问题,而且我认为这同样是最优学习本身的问题,因为最优的机器学习总是校准的,贝叶斯最优分数总是校准的。

假设我想预测职场生产力,并把生产力定义为你未来十年的工作小时数。假设我机器学习做得极好,给出的正是精确的贝叶斯最优分数:我知道在你的特征条件下未来工作小时数的确切期望。问题在于,一个有残障的人,或者未来可能怀孕的人,分数会更低,仅仅因为分数正确地识别了未来工作时数的损失。可这正是我们认定在招聘中道德上不相关的东西。所以这是个很好的例子,说明按群体校准解决不了道德相关性的问题,最优预测一般也解决不了。最优预测会利用它能利用的一切:只要某条信息能提升预测精度,它就会纳入,哪怕我们作为一个社会认定它不相关。

子群体漏洞与不可能性定理

哈特:关于这些群体公平定义还有一点要说,它不限于校准,而是对所有准则都成立:如果你在两个群体之间保证了某个公平准则,这可能导致在其中一个群体内部出现违反,而且往往是更触目的违反。这催生了关于子群体公平的工作,有两篇论文开启了这条线:一篇是卡恩斯(Kearns)、尼尔(Neel)、罗斯(Roth)与吴(Wu),另一篇是赫伯特-约翰逊(Hébert-Johnson)、金(Kim)、莱因戈尔德(Reingold)与罗斯布卢姆(Rothblum)。

卡恩斯等人有一个虚构的示例:有蓝、绿两个群体,你在每个群体里接受相同比例的个体,用这些圆圈表示。然而,在一个群体里你只接受男性候选人,在另一个群体里只接受女性候选人。这样一来,只要看这些子群体,就会发现独立性准则被严重违反。你可以为所有准则构造类似的例子。

我认为这是一个真实的担忧。宾夕法尼亚大学卡恩斯团队把它叫作「公平选区操纵」(fairness gerrymandering),我喜欢这个说法。如果你强制施加某个约束,机器学习通常会以最偷懒的方式去满足它,所以你不应指望在子群体内部会有什么聪明的事发生。均等化某个统计量之后,子群体里很可能有恼人甚至令人警觉的事情在发生。

我们回顾一下进展到哪里了。还有半小时,时间充裕。我们看了三个准则:第一个是独立性,第二个是条件独立性,R 在给定 Y 的条件下独立于 A,第三个是另一种条件独立性,Y 在给定 R 的条件下独立于 A。第一个蕴含接受率均等,第二个蕴含错误率均等,第三个蕴含按群体校准。我喜欢这样表述,因为用条件独立性来理解它们非常容易。我想你们到这个阶段应该已经听过图模型或因果的课,理解条件独立性,这也是记住它们的好办法。

把多个公平准则摆出来之后,人们自然会问:能不能全都要?能不能构造一个神奇的预测器,同时满足这一切,得到最好的世界?人们很快发现,答案是不能。形式上可以这样陈述:这三个准则中的任意两个在一般情况下互斥。除非在退化的条件下,你不可能全部拥有。

我只给出一个形式陈述,略去其他可以证明的陈述,这一个是错误率均等对校准。定理如下:假设基础率(base rates)不相等,也就是两个群体的正结果比率不同,一个群体的正结果率高于另一个;再假设你的决策规则不完美,错误率非零,至少有一个假阳性和一个假阴性。这是一般情况,因为我们生活在一个不平等的世界里,群体的基础率往往不同,而我们通常也没有完美的预测器,预测器会犯错。在这些一般性的假设下,如果满足按群体校准,那么错误率均等必然失败。特别地,如果你只做无约束的机器学习,得到了按群体校准,那就意味着你默认就违反了错误率均等。强制其中一个,另一个就变成假的。

最初的工作来自乔尔德乔娃(Chouldechova),以及克莱因伯格(Kleinberg)、穆莱纳森(Mullainathan)与拉加万(Raghavan),他们证明了类似的陈述,不完全是这个。后来有大量跟进工作研究放松这些准则之后会怎样。放松之后的情况不那么容易说清,对某些松弛版本存在一些非平凡的权衡是可能的。

COMPAS 之争与两重告诫

哈特:但我想转而谈的是,这些权衡如何主导了围绕一个重要问题的学术讨论,那就是美国的再犯预测。二〇一六年五月左右,ProPublica 发表了一篇里程碑式的报道,题为《机器偏见》,指控全国各地使用的一款预测未来罪犯的软件对黑人存在偏见。它引发了大量讨论。

COMPAS 之争的核心,至少投射到学术圈、投射到计算机科学界所讨论的部分,是这样的:有一个叫 COMPAS 的风险评分,美国许多司法辖区用它来评估被告获释后再次犯罪的可能性。当你必须决定是羁押还是释放一名待审被告时,它用来估计该被告获释后再犯的风险。法官可以部分依据这个分数决定羁押,它不是法官看的唯一东西,但可以用。

ProPublica 考察了那些最终没有再次犯罪的被告,注意到在这些人当中,黑人被告被评为高风险的比例远高于白人被告。换个说法,这个分数的假阳性率对黑人被告显得远高于白人被告:黑人被告更可能被贴上高风险的标签,尽管他们并没有再犯。

COMPAS 的开发商 Northpointe 回应说:没错,但我们的分数是按群体校准的,而黑人被告在美国的再犯率更高,所以这是不可避免的。他们的意思是,我们只是在做本职工作,把分数校准好。如果分数不校准,法官就得做心理体操:被告是白人就加三分,是黑人就减一分之类。为了避免这种事,分数必须按群体校准,而这不可避免地导致那个差距。

不管好坏,这大体上就是那场讨论的内容。学者们被这个想法吸引:这个复杂的问题似乎归结为某种统计上的权衡,可以用数学工具和严格论证去处理。

我想对 COMPAS 之争加两重告诫。第一重是,错误率均等和校准在这个场景中都不能排除不公平的做法,所以争论它们已经有点像在打稻草人,因为两者都排除不了你真正该担心的事。特别是,刑事司法中什么是公平,不是靠符合其中某一条准则就能解决的,我们不能把正义或刑事正义的问题归约为其中一条准则。我想给你们一个构造性的例子来说明这一点,说明这两条性质都不是公平证书,不是公平的证明。

先看错误率均等在这种场景里为什么有问题。考虑两个群体,蓝色群体和橙色群体。阈值规则是:分数高于 0.5 的一律羁押,低于 0.5 的一律释放。分数写在这些小人下面,蓝色群体从 0.1 到 0.7,橙色群体从 0.2 到 0.9。再假设这些分数就是这些人真实的再犯概率,也就是给定个体协变量后再犯的条件期望,姑且当作真实概率。

基于这些真实概率,可以算出羁押率和假阳性率:蓝色群体羁押率百分之三十八,橙色群体百分之六十一。橙色群体被羁押的人多得多,所以独立性被违反了。橙色群体的假阳性率也高得多,所以错误率均等也不满足。

你可能会问,怎么修?假设你是一个警察部门,被指控在一个群体里的羁押率远高于另一个群体,你能做什么?这个例子的要点是:达成错误率均等有非常坏的办法,达成独立性也有非常坏的办法。比如,你可以对橙色群体更激进地执法,逮捕更多低风险的人,也就是那些再犯风险很低的人,然后把他们释放。这样你只是在给橙色群体里将被释放的被告数量注水。这么做之后,羁押率可以降到百分之四十二,假阳性率降到百分之二十六,两个群体看起来就相近了。

主持人:这里有一个问题,问答区里的:对于一个具体问题,我们该如何选择正确的准则?有什么建议吗?

哈特:好问题,我稍后会回到这个问题。先让我把这个例子讲完,因为正讲到一半。等我讲到公平准则的用途时,请提醒我。

所以,这很令人警惕。它说明,如果你激励某人去满足错误率均等,他们可以用偷偷摸摸的方式达成,甚至用伤害更大、引入额外伤害的方式达成。作为激励、作为要追求的目标,这可能是一个非常成问题的准则。

校准也有一个值得指出的问题。再次假设这些是真实的再犯概率,假设我们知道它们。现实中通常不是这样,但为了这个例子姑且假设。顺便说一句,这两个例子都来自科比特-戴维斯(Corbett-Davies)、皮尔逊(Pierson)、费勒(Feller)、戈尔(Goel)与胡克(Huq)二〇一七年的论文,是他们慷慨提供给我的。

还是羁押所有高于 0.5 的人,要求满足校准,但我想释放这个群体里的所有人,就想给这个群体放一马。校准允许我做的一件事是:选出一个子集,让它们的分数平均到某个特定值。这里我挑出分数为 0.1、0.1、0.2、0.2、0.2 和 0.6、0.6、0.6、0.7、0.7 的这些人,它们平均下来是 0.4。校准允许这样的平均操作:把所有这些人的分数都替换成它们的平均值 0.4。这样得到的分数依然是校准的,而现在没有人被羁押,所有人都低于 0.5 的羁押阈值。所以在校准之下也可能发生这种偷偷摸摸的操作,而这个准则检测不出来。

未出庭案例:该不该做预测

哈特:现在说关于刑事司法这个例子的第二重、也是更重要的告诫,然后再谈这些公平准则的适用范围,也就是刚才那个问题。

人们在刑事司法的语境下长篇大论地讨论这些公平准则,讨论所有这些权衡,讨论从不同准则中能学到什么。但还有一个问题:我们一开始就把自己绑在预测视角、统计视角上,是不是错过了问题的大部分?这个问题是不是比统计视角所暗示的复杂得多?学术辩论围绕这些准则之间的张力展开,人们也正确地指出警务中存在反馈回路,会扭曲测量:你只在你执法的地方测得犯罪,这就不利于那些执法更密集的群体,在美国主要是黑人街区,你在那里收集到更多犯罪记录,所以你一开始的数据就很糟糕。但这也不是我的主要观点。

我的主要观点是:有时候问题不在于我们怎么预测、用什么数据、选什么目标函数,而在于我们一开始就选择把这个问题表述为一个预测问题。有一个例子我觉得非常有说服力,就是未出庭(failure to appear)问题:预测被告会不会到庭。这在美国刑事司法系统里同样是个重要问题。一种做法与再犯预测类似:你有一个被告,给他定了三个月后的开庭日期,你希望他到时出现,又不确定他会不会来,于是你可以做机器学习来预测未出庭。很多人在做这件事,很多人在推动把它作为全国刑事司法改革的一部分落地,用预测未出庭的风险评分取代人类法官。想法是:分数高就继续关押,分数低就释放。

从这个时点到开庭可能相隔三个月甚至更久。如果你先被关三个月,第一,把人关在监狱里通常很昂贵;第二,这对当事人是毁灭性的。他们会丢掉工作,失去社会联系,生活被彻底打乱。即便他们原本处境相当安全稳妥,等到出狱、洗清指控时,很可能也回不去了,要挣扎着重新站起来。审前羁押是极具破坏性的。

我觉得有说服力的替代方案是认识到:人们未能出庭,往往是因为缺少托儿服务或交通工具,因为工作时间冲突,雇主不放人,或者仅仅是因为开庭次数太多。几年前的一届「公平、问责与透明」会议上,有一位曾经身为被告的黑人讲述,他在最终被免除指控之前经历了二十二次开庭。不是一次,是二十二次。如果你在第十五次没有出现,就会被记录为未出庭,哪怕其余各次你都到了。

认识到这一点,更合理的做法也许是给人们发放托儿券或交通券,强制雇主放人出庭,同时努力减少开庭次数,让这份负担更容易承受。你可以采取这些措施去缓解问题,而它们没有一项和预测有关。这个替代方案实际上已经成为哈里斯县(Harris County)诉讼和解协议的一部分。哈里斯县是得克萨斯州的一个县,有人对该县使用这类风险评分提起诉讼,和解协议要求哈里斯县在法院提供免费托儿服务,建立双向沟通系统等等,落实这类结构性干预,让贫困被告不至于一开始就处于劣势、被这个系统进一步惩罚。

在这种情况下,我很难为一开始就使用预测辩护,因为落实这些别的措施显得更直接、更重要。我个人一直觉得这条论证很有说服力,作为一个搞机器学习的人,面对这些替代方案,我很难继续坚持预测那一套。

统计视角的局限与总结

哈特:这里我要回到刚才那个问题:这些统计准则究竟窄在哪里,什么让它们受限?统计公平准则把数据生成分布当作既定的,只用 X、Y、R、A 这些随机变量的联合统计量,也就是观测性的联合统计,你也可以往同一个总体里加入任何其他随机变量。结果是,它们能告诉你的,被限制在总体的联合统计量所能告诉你的范围之内。特别是,你不能改变总体,不能考虑反事实假设,不能干预产生这个联合分布的世界。从某种意义上说,它非常被动:把世界的现状照单全收,于是失去了一整套干预手段。

周四我要做的,就是探讨这个问题:如果我们承认统计视角可能太窄,排除了太多干预,把焦点放在了不一定针对正确问题的地方,那么我们该如何把显著的社会事实与社会语境纳入考量?如何让我们的模型更能感知那些真正重要的更宽泛的问题?这是周四第二部分的主题。

现在总结一下,然后再回答一些问题。第一件要记住的事:「无意识即公平」是失败的。

第二件事更微妙一些。我仍然觉得这些公平准则有意思,因为它们能触发不同的道德直觉,我认为这是一个被低估的机制。人们常常想要一个公平的定义,然后失望地发现这些东西不是公平的定义。它们确实不是公平的定义,绝对不能证明某件事是公平的。但它们确实会触发道德直觉,这只是一个经验陈述。ProPublica 对错误率均等不成立感到愤怒,这就是一种道德直觉,而且很难反驳,因为它捕捉到了我们关于决策应当如何做出的某种道德理解。

因此,这些准则仍然可以用来揭示关于决策的各种规范性问题,帮我们更好地把握在某个场景里,决策的做法究竟哪里让我们不安。它们会引出各种权衡与张力,我认为这是一场重要的对话。这场对话不是要一劳永逸地裁定什么公平、什么不公平,它是一个更微妙的机制,帮助我们讨论自己对决策应当如何做出的理解与期待。

但正如我说的,公平准则本身不能作为公平的证明。我认为至少在我描述的这个统计设定之内,不可能有公平的定义。我确信,在这个统计设定之内不存在令人满意的公平定义。公平准则本身也不是好的目标函数。正如我在错误率均等与警务的例子中展示的,你不希望把这些准则变成激励,不希望激励人们去追求它们。把它们当作衡量现实世界的量规、测量工具,是一回事;要求人们去优化它们,是另一回事。

问答:干预、多分类、内容审核、身份

主持人:有几个问题,问题还在不断进来。第一个来自一位听众,他说:你对预测是否过于悲观了?在有些场景里我们不知道什么是正确的干预,而对给定 X 时 Y 的条件期望做出好的估计,实际上可能帮助我们找到干预。比如在未出庭的例子里,可能会发现带着幼儿的女性更容易未出庭,从而凸显出托儿服务是更好的干预。

哈特:这是很好的观点,谢谢提问,也谢谢你来听报告。人们是怎么发现缺乏托儿服务是未出庭的重要因素的?他们大概也在某处用了统计。我绝不是在对统计的使用做笼统的指控,那确实走得太远、太悲观了。但他们提问的方式很可能非常不同:我们对未出庭的原因做因果推断吧。的确,这个条件期望可能以某种方式融入那种理解,或者某种形式的统计至少能给出好的答案。也有可能,仅仅是把预测做得更好,你就能通过观察模型找出一些显著因素,从而帮助理解。这些我都同意。

但我认为这是一种不同的问题框架,我们首先要在「这是我们要追求的问题框架、这是我们做统计的目的」上达成一致。只要我们说目的就是把这些风险评分装上去,用来给审前羁押分类,我们就可能无意中放弃了问题的这些不同框架。不管怎样,我认为都会有好的技术工作的空间。提出这些替代干预,不是说这些事情里就没有任何有趣的研究问题,它只是对问题的一次重新框定。

主持人:下一个问题:在多分类与回归任务中,二元公平准则会增加哪一层复杂性?

哈特:有一点要说明:这些条件独立或独立的陈述,可以直接推广到任意随机变量。离散的肯定可以,稍微费点力气也能处理连续变量。你可以有任何离散的群体身份变量,任何取多个标签的离散结果变量等等。所以这些条件独立陈述已经适用于多分类与回归任务,这也是我喜欢用这种方式表述的部分原因。好问题。

主持人:另一位听众问:在预测是必需的场景里,比如内容审核,你对公平性要求有什么看法?最近的研究显示,许多模型对非裔美国人英语内容的标记率更高。有没有什么公平准则能在这类预测必要的场景中缓解偏见?

哈特:很好的问题。有些场景里,仅仅因为人力资源不足,我们就需要求助于某种形式的算法审核或内容策展。姑且接受这个前提,虽然这本身可能是另一场辩论。是的,毒性分类、情感分类、辱骂分类等等,都存在针对比如非裔美国人说话者的严重偏见,因为训练数据里他们使用的词语或短语与标签之间存在相关性,导致他们更容易被标记。这是个大问题。我认为这会违反其中许多准则,错误率均等肯定被违反了。这又是一个例子:这里触发的道德直觉是有效的,我们觉得这非常令人担忧,而且确实如此。

但我不认为这些准则在这种情况下能给你缓解问题的办法。像情感分类或毒性分类这样的任务,我认为本身就是很成问题的预测任务,因为目标变量非常模糊。我知道我们是在「必须这么做」的前提下讨论,但从某种意义上说,我也想重新审视这个前提,因为我不认为这是表述这些问题的正确方式,它会撞上这些问题。这些准则本身并不会告诉你该怎么做,它们只会把这标记为一个问题,不会告诉你接下来往哪里走。在这些情况下该怎么办,目前有很多工作正在进行。

主持人:另一位听众问:校准是否应该考虑个体特征,以防止在群体校准之下群体内部的那种随机化?

哈特:正是如此。这就是我举的橙色群体和蓝色群体的例子,做那种平均操作会得到不希望看到的结果。给定 X 时 Y 的条件期望,即便在具体特征层面也满足校准:对每一个 X 都成立的这个条件,是一个强得多的条件,基本上就是最优预测的条件,是一个强得多的设定。有一些论文尝试逼近这种更强的条件,这与子群体问题有关。我想刚才提问的那位听众本人就有这方面的工作,斯坦福的欧姆·雷因戈尔德(Omer Reingold)、盖伊·罗斯布卢姆(Guy Rothblum)、迈克尔·金(Michael Kim)也有论文,尝试对更丰富的子群体实现校准,这会逼近那种更强的保证。

主持人:还有一位听众问:之前有人提到接受率与结果平等的关系,能否澄清一下这种关系?

哈特:我想请之前提问的那位说明他所说的「平等」是什么意思,等我们在这一点上达成共识,我回答起来会更有把握。

主持人:还有一位听众问:你觉得由谁来做这项工作重要吗?我们都属于某些群体,有些人比别人更有特权,我们该如何看待这一点?

哈特:我先回答第二句。第一个问题我一开始没完全看懂,是问我是否认为,从身份的角度看,由谁来研究这些问题是相关的?是的,这很重要。我认为在这个问题上,我作为一个享有特权的白人男性,视角是非常有限的。我一直是从一个很有限的视角出发研究形式模型。我无法讲述被歧视和被压迫的亲身经历,正因如此,让那些能讲述亲身经历的声音居于中心,是绝对重要的。报告开头我列了一批做这件事的学者,我认为那才是你们该开始的地方:对问题的记述,对经历的记述,这是绝对不可或缺的。

所以,这个教程不应该是你们在这个话题上的唯一来源。它只是对这一领域形式化与技术性工作的一种视角,你们绝不该止步于此。我认为正确的方式是,我们都从多种不同的来源、从多元的来源学习,其中一些更直接地讲述亲身经历。这是个重要的问题。我刚意识到这是一条私信提问,我却公开回答了,希望这没有违背你的本意。

主持人:提问者说本来就是打算公开的。刚才有位听众问过结果平等的问题,他现在在嘉宾席里,也许可以直接说。

听众:你好。我之前问的是「结果平等」的澄清。我知道时间快到了,如果你们想结束,我可以晚点再问。

哈特:没关系,我们给它一分钟。

听众:其实我并不知道「结果平等」到底是什么意思,我不是搞政治或伦理学的,它只是听起来很重要,所以我才请人澄清。

哈特:我觉得可以这样理解:如果把「结果」理解为你的决定,而你只是在比较这个决定上的接受率,那么这就是某种意义上的结果平等。但我不认为这个准则捕捉到了道德哲学家所说的「结果」,那大概是一场我们之后该另找时间进行的更长的对话。

听众:明白。有没有什么资源讨论这两者的对应关系?

哈特:我不认为有人尝试过。这是个好问题,也是我和合著者索隆·巴罗卡斯(Solon Barocas)经常碰到的:从这些准则到道德哲学之间,没有清晰的映射,反过来也没有。不是说对每一条准则你都能指出「它追踪的是道德哲学里的这一块」。我们曾经尝试过这样的综合,一度抱有很高的期望,但没走多远。可能有很多人能给你更好的回答,但我一时想不出哪篇论文讨论了这种精确的对应。

听众:非常感谢,这个回答本身就有帮助,至少说明了这是个大问题。

哈特:是的,机器学习中的公平研究与哲学之间的对应关系,在我看来非常不清楚。如果你在哲学系有同事,这是个很值得和他们聊的问题。

结束语与 MLSS 花絮

主持人:非常感谢莫里茨精彩的报告,我想大家今天都从你这里学到了很多。有一个信息通知:半小时后我们有一场与莫里茨的圆桌讨论,可以把一些讨论延续到圆桌上。休息之前,还有几段来自组织方的短视频,我来放一下。

视频发言人:在我看来,机器学习完全建立在统计依赖之上:我们看到某些量与另一些量相关,然后用一个去预测另一个。通常机器学习不会追问这些依赖从何而来,而我认为它们来自底层的因果关系。如果我们关心的是当世界上某些东西发生变化时该怎么做机器学习,也就是在一个问题上训练、在另一个问题上测试,这种情况在动物和人类身上时刻发生,那我们就必须更细致地看待事物。

视频发言人:这个问题很难。我在研究上不算系统,不会在一月一日醒来就决定明年要做什么,然后按部就班地执行。对我来说,一个很大的灵感来源是和业界的人交流,了解那些我甚至不知道存在的实际问题。这比读文献高效得多。我倾向于做自己喜欢的东西,而不是因为某个方向热门、大家都在做就去做。我在很大程度上是被做研究的乐趣驱动的,被那些不管别的、我就是觉得有意思的东西驱动。我本来不想用「热情」这个词,因为作为意大利人,热情太多有点老套,但这是真的。我认为人应该被好奇心和热情驱动,那是一种非常好、非常健康的驱动力。

视频发言人:最好的体验大概是讲者与学员之间的亲近氛围。对我个人而言,五年前我第一次参加 MLSS 是作为学员,如今是作为讲者。最有趣的其实是那件 T 恤:当你走出这个圈子,经常有人问你为什么参加。

主持人:非常感谢大家,本节到此结束,二十分钟后我们在与莫里茨·哈特的圆桌讨论上再见。谢谢,再见。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:02 开场与讲者介绍 ▶ 正在看
2:14 范围界定:后果性决策中的歧视 ▶ 正在看
8:44 什么算歧视:不正当的区分依据 ▶ 正在看
11:04 美国反歧视法的两种原则 ▶ 正在看
18:11 「无意识即公平」为何失败 ▶ 正在看
24:02 统计公平的学术渊源与形式框架 ▶ 正在看
33:22 标准一:独立性与接受率均等 ▶ 正在看
42:04 标准二:错误率均等与 ROC 曲线 ▶ 正在看
50:09 标准三:校准与充分性 ▶ 正在看
58:03 子群体漏洞与不可能性定理 ▶ 正在看
63:52 COMPAS 之争与两重告诫 ▶ 正在看
73:10 未出庭案例:该不该做预测 ▶ 正在看
78:29 统计视角的局限与总结 ▶ 正在看
82:57 问答:干预、多分类、内容审核、身份 ▶ 正在看
94:29 结束语与 MLSS 花絮 ▶ 正在看
本期小问 · 档案清单
—— 从训练数据里删掉种族、性别等敏感属性,就能让算法公平吗? ▶ 正在看
—— 各组错误率相等与各组分数校准,为什么在现实中不能同时满足? ▶ 正在看
—— 预测被告是否会缺席庭审,真的比提供托儿和交通补贴更值得做吗? ▶ 正在看
—— 把某个公平指标当成优化目标,会不会反而诱导出更大的伤害? ▶ 正在看
本期讲者
莫里茨·哈特时任加州大学伯克利分校电气工程与计算机科学系助理教授,普林斯顿博士,曾任职 IBM 研究院与谷歌;FAT/ML 研讨会联合创始人,教科书《Fairness and Machine Learning》合著者,Equalized Odds 公平标准提出者之一。
主持人MLSS 2020 图宾根暑期学校组织方成员,负责介绍讲者与转达聊天区提问。
01开场与讲者介绍
0:02
[Music] okay hello everyone welcome back to the machine learning summer school 2020 so today this is a clip it's my pleasure to introduce our next speaker Mike Hart simorgh Hart is an assistant professor in the Department of Electrical Engineering and computer science at the University of California Berkeley so he investigated algorithm and machine learning with focuses on reliabilities validities and societal impact after obtaining a PhD in computer science from Princeton University's hee-ho position at IBM research cooker wizard in Google burn he also a co-founder of the workshop on fairness accountability and transparency in machine learnings and co-author of the textbook fairness and machine learnings he has received an NSF Career Award as a Sloan fellowship and best paper award at ICML 2018 and iclear 2017 so without further ado more it
[音乐] 好的,大家好,欢迎回到 2020 年机器学习暑期学校,今天这一节,我很荣幸能够介绍我们的下一位讲者,Mike Hardt——Moritz Hardt 是加州大学伯克利分校电气工程与计算机科学系的助理教授,他研究算法与机器学习,重点关注可靠性、有效性以及社会影响。在普林斯顿大学取得计算机科学博士学位后,他曾任职于 IBM 研究院,也曾在谷歌工作。他还是机器学习中的公平、问责与透明度研讨会(FAT/ML)的联合创始人,以及教科书《公平与机器学习》的合著者。他曾获得 NSF 职业奖(CAREER Award)、斯隆奖学金,以及 ICML 2018 和 ICLR 2017 的最佳论文奖。那么废话不多说,交给你
便签笔记
1:41
yeah thank you so much for introducing me I really appreciate the opportunity to speak here and I will share my screen now let me see
是的,非常感谢你的介绍,我真的很感激能有机会在这里演讲,我现在来分享我的屏幕让我看看
便签笔记
02范围界定:后果性决策中的歧视
2:14
here we go are you seeing my screen now excellent alright so thank you so much for joining me to this talk I want to give you a little bit of a tutorial on fairness and machine learning and some of the limitations and opportunities of using machine learning in the context of consequential decision making I will split that across two parts the first one is today the other one is coming up on Thursday and I'll tell you what I'll have to plan for for both days so it's it's a tough topic to engage with it's a difficult topic and it's hard to even say where to start and I think it helps to acknowledge that we live in a world of pervasive inequality and oppression and discrimination and current political events certainly are testament to how great that situation is and how urgent it is to to act and as we use machine learning to formalize scale accelerate processes in this world of inequality we run danger of perpetuating existing patterns of injustice okay so we have to be careful in the way that we use
好了,你们现在能看到我的屏幕了吗?很好,那么非常感谢大家来听这场报告,我想给大家做一个关于公平与机器学习的小教程,讲讲在重大决策场景中使用机器学习的一些局限和机会。我会把内容分成两部分,第一部分是今天,另一部分在周四,我会告诉大家两天分别的安排。所以这是一个很难切入的话题,它是一个困难的话题,甚至很难说该从哪里开始,我觉得有帮助的做法是承认:我们生活在一个普遍存在不平等、压迫和歧视的世界里,而当前的政治事件无疑印证了这种状况有多严重、采取行动有多紧迫。当我们用机器学习去把这个不平等世界中的流程形式化、规模化、加速化时,我们就有可能延续既有的不公正模式。好的,所以在一个非常不平等、存在大量既有不公正的世界里,我们必须
便签笔记
3:34
machine learning for problem solving in a world that is very unequal and has a lot of existing non-justice but I would say that there's also a somewhat fragile opportunity to revisit decision making and revisit how we do machine learning in various domains and reform existing processes possibly for the better so I'm hoping that machine learning can also be a tool that is part of the reform here and I would like to say that you know in some sense my tutorial is not even the right place to start I would say you should start by reading the important scholarship that points out the you know the the potential issues with using AI and automated decisions there's been tremendous scholarship just to give you five books that appeared in recent years by ruja Benjamin Meredith Broussard Virginia Ewbank Sofia Noble and Kathy O'Neill and all of these books give you a tremendously important accounts of the dangers of using AI and and computer technology for solving problems that have societal impact and
谨慎地使用机器学习来解决问题。但我也想说,这里同样存在一个多少有些脆弱的机会——重新审视决策方式,重新审视我们在各个领域如何做机器学习,并改革现有的流程,或许能让它变得更好。所以我希望机器学习也能成为这场改革中的一种工具,而且我想说,从某种意义上讲,我的这个教程甚至都不是合适的起点。我认为你们应该先去读那些重要的学术著作,它们指出了使用人工智能和自动化决策可能带来的问题。这方面已经有大量出色的研究,这里只举最近几年出版的五本书为例,作者分别是 Ruha Benjamin、MeredithBroussard、Virginia Eubanks、Safiya Noble 和 Cathy O'Neil,所有这些书都为你提供了极其重要的论述,讲述在解决具有社会影响的问题时使用人工智能和计算机技术的危险,
便签笔记
4:40
these are the books that you should start with to get a sense of what the what the dangers are and there's a lot of important scholarship you should also check out tutorials and keynotes given by joy bull and we Nika Crawford Timna gabru Latanya Sweeney Meredith Whittaker and others that have tremendous and done tremendous work in this area and have given the very accurate and important accounts of the dangers of using this technology and so this is in some sense the natural place to start and my tutorial though and let me give you one quote that I find fascinating from guru Benjamin spoke so she introduces something she calls the new gym code and she calls that the employment of new technologies that reflect and reproduce existing inequities but that are promoted and perceived as more objective and progressive than the previous systems of discrimination discriminatory systems of a previous area so this points out you know that machine learning can have an AI can have sort of this effect of being perceived as more
这些是你们应该首先阅读的书,以便对其中的危险有一个认识。还有很多重要的学术工作,你们也应该去看看 Joy Buolamwini、Kate Crawford、TimnitGebru、Latanya Sweeney、Meredith Whittaker 等人所做的教程和主旨演讲,他们在这一领域做了大量卓越的工作,并对使用这项技术的危险给出了非常准确而重要的论述,所以从某种意义上说,这就是这是我教程里自然的起点,我想先给大家分享一段我觉得非常有意思的引述,来自 RuhaBenjamin。她提出了一个概念,叫做"新吉姆代码"(the New Jim Code),指的是那些反映并再生产既有不平等的新技术的应用,但这些技术却被推广、被认为比以前那些歧视性的系统更客观、更进步比以前那个时代的系统更好。所以这就指出了,你知道,机器学习、人工智能可能会产生这样一种效果:被认为更
便签笔记
5:44
neutral and being perceived as more progressive than the old Burke Roddick processes that constitute a blatant discrimination but in the end it just reproduces these inequities and she gives a powerful persuasive account of that danger so with that said on keeping that in mind the focus for this tutorial is for not primarily on discrimination and consequential decision making settings so I'll be talking about you know decision-makers that you know have to make typically a binary decision something like accept or reject in a setting that is of has consequences for the individual hiring you know college admissions applications in criminal justice and so forth that's the kind of setting you should think about and I should say that this excludes right off the bat many other forms of injustice or even unfairness that exists in the world and not everything that is a you know form of injustice is you know discrimination in a decision-making setting okay so this is already limiting the scope the second
中立、更进步,比起那些构成公然歧视的旧官僚流程,但最终它只是再生产了这些不平等。她对这种危险给出了非常有力、有说服力的论述。所以说到这里,带着这个想法,本教程的重点并不主要放在歧视上,而是放在后果性决策的场景上。所以我要讲的是,你知道,那些决策者,他们通常需要做出一个二元决策比如在接受或拒绝这样的场景中,会对个人产生实际后果,比如招聘、大学录取申请、刑事司法等等,这就是你应该想到的那类场景。而且我要说明的是,这一开头就排除了世界上存在的许多其他形式的不公,甚至是不公平,并不是每一种不公都属于决策场景中的歧视。好,所以这已经限定了讨论范围。第二
便签笔记
6:55
you know caveat is that I'll be talking primarily about or from a us-centric perspective especially so far as my examples in the legal backdrop go and I appreciate that this audience is geographically very diverse we have a lot of strong European audience here and people from all around the world and so my examples do not necessarily apply to the societal legal framework in your communities and this is is a limitation but the the the legal situations are fragmented and that it doesn't make sense to talk about this problem of discrimination sort of globally it takes on very different forms in different countries and so finally of course I will be talking about formal models and frameworks to address these questions that's because that's where my expertise is but that is not in any way meant to dis dissenter or distract from the scholarship that I just mentioned that still stands and is important to keep in mind it also decidedly leaves room at least in my opinion for non technical interventions so things that are not you
个需要说明的是,我主要会从以美国为中心的视角来讲,尤其是在我举的例子和法律背景方面。我也理解在座的听众来自世界各地,我们这里有很多来自欧洲的听众,也有来自世界各地的人,所以我的例子未必适用于你们所在的社会和法律框架。社群,这确实是一个局限,但法律状况是碎片化的,所以笼统地在全球层面上讨论歧视这个问题其实没什么意义,它在不同国家会呈现出非常不同的形式,所以最后,我当然会讲到用来处理这些问题的形式化模型和框架,那是因为那正是我的专长所在,但这绝不意味着要贬低或转移大家对我刚才提到的那些研究成果的注意力,那些研究依然成立,也值得我们始终牢记。至少在我看来,这也明确地留出了空间给非技术性的干预手段,也就是那些不使用预测或机器学习来解决问题的方法。所以在本次教程的过程中,我们
便签笔记
8:08
know using prediction or machine learning to solve a problem so we will in the course of this tutorial encounter scenarios where many of you will be challenged to you know ask the question should we be using prediction or machine learning in the first place to solve this problem okay so that's that's the focus for this tutorial and I see there are some questions so maybe if you have questions already at this point I can take one of them if you will
会遇到一些场景,其中很多人会被促使去思考这样一个问题:我们究竟是否应该一开始就用预测或机器学习来解决这个问题。好,这就是本次教程的重点,我看到有一些问题了,所以如果你们现在已经有问题了,我可以先回答其中一个,如果你们愿意的话
便签笔记
03什么算歧视:不正当的区分依据
8:44
let's see okay maybe I'm there are no questions here okay so some of you might come to this and say well isn't discrimination and this was some of the early reactions to this topic isn't discrimination the whole point of machine learning you know isn't it the whole point to draw some decision boundary and you know label some of the points accept and some of the points reject and that's going to exclude some people and include others isn't that the whole point of machine learning so wouldn't it be a little bit rich to accuse machine learning of a sort of constituting or participating in discriminatory practices and so the concern that we'll trace in this tutorial is with unjustified with an unjustified basis for differentiation not just any form of you know distinguishing or between individuals not just any form of you know accepting or rejecting individuals in a certain setting okay and so unjustified basis for differentiation could be something like practical irrelevance like in the context of sexual orientation and
我看看,好吧,这里好像没有问题了,好的。你们当中有些人可能会说,歧视不就是机器学习的全部意义所在吗?这也是这个话题早期引发的一些反应之一歧视不就是机器学习的全部意义吗?它的全部意义不就是要划出某条决策边界,然后把一些点标为接受,把一些点标为拒绝,这样必然会排除一些人、纳入另一些人,这不就是机器学习的意义所在吗?所以,指责机器学习构成或参与了歧视性做法,这难道不是有点太过分了吗?所以我们在本教程中要梳理的关切,是那种没有正当理由的区分依据,而不是任何形式的、对个体之间的区分不只是在某个场景中接受或拒绝某些个体的任何形式,好吧,所以不正当的区分依据可能是像"实际不相关"这样的东西,比如在性取向和雇佣决定的背景下,我们认为这对这份工作并没有实际相关性,那为什么它应该成为招聘决定的一部分呢
便签笔记
9:52
employment decisions we don't think this is practically relevant for the job so why should it be part of the hiring decision but there's also more complicated notion perhaps of moral irrelevance where you might have something like disability status and you know that accommodating somebody's disability might incur additional cost to the employer but we decided as a society that this is a cost that the the employer has to absorb and we considered morally irrelevant that somebody has a disability when it comes to hiring decisions even if on purely statistical grounds that might affect their their work okay similarly the possibility of pregnancy should not have an influence on hiring decisions even if it you know might offset the number of hours you can work in the future okay so these are examples of moral irrelevance and this is sort of thing that discrimination is trying to or is trying to capture and it's important to understand that discrimination is not a general concept it's domain-specific okay so there it
我们认为这对这份工作并没有实际相关性,那为什么它应该成为招聘决定的一部分呢?但还有一个也许更复杂的概念,叫做"道德上不相关",比如残疾状况这样的情况,你知道,为某人的残疾提供便利可能会让雇主承担额外成本,但我们作为一个社会已经决定,这个成本是雇主必须自己吸收的,而且我们认为在招聘决定中,某人是否有残疾在道德上是不相关的决定中,即便纯粹从统计角度看,这可能会影响他们的工作表现,好吧。同样地,怀孕的可能性也不应该影响招聘决定,即便这可能会减少你未来能工作的小时数,好吧。所以这些就是"道德上不相关"的例子,这也是歧视这个概念试图捕捉的东西。而且很重要的一点是要理解,歧视并不是一个笼统的概念,它是有领域特定性的,好吧。所以它并不适用于每一个领域,它试图讨论的是
便签笔记
04美国反歧视法的两种原则
11:04
doesn't apply to to every every domain is trying to sort of talk about the names domains that are concerned with important opportunities that affect people's lives not just any you know it doesn't apply to any fertilis example of use of machine learning so it implies to to these consequential domains and it's also group specific it's about socially salient categories that have served as the basis of unjustified and systematically adverse treatment in the past so race gender disability status these are things that have been used a two-disc or have been discriminated against in the past and and continue to be a focus of what discrimination means so it's not just any group membership not just any group status that we're concerned with and so this is where the law comes in and I'm going to give you just like a very brief you know a sense of what the law does and doesn't do just as a framing in the United States law they're regulated domains domains that have some kind of anti-discrimination
那些涉及重要机会、影响人们生活的领域,而不是任意的什么领域,你知道,它并不适用于任何随便举出的机器学习使用案例。所以它适用于这些有重大影响的领域,而且它也是群体特定的,它关注的是那些具有社会显著性的类别,这些类别在过去一直是不正当的、系统性的不利对待的依据。所以种族、性别、残疾状况,这些都是过去曾被用来歧视,或者说曾遭受歧视的因素,并且它们仍然是歧视这个概念所关注的焦点。所以并不是任意的群体成员身份,也不是任意的群体状态是我们所关心的。所以这就是法律介入的地方,我要给你们讲的只是非常简要的,你知道,让你们大致了解一下法律做了什么、没做什么,只是作为一个框架。在美国法律中,有受监管的领域,也就是那些有某种反歧视立法的领域,还有一些没有。信贷
便签笔记
12:11
legislation and others that don't credit and is an important one and there's that Equal Credit Opportunity Act there is education with the Civil Rights Act of 1964 employment housing on public accommodation and these are all private sector domains these leave out government domains where the law might be more complicated and one important thing for machine learning because a lot of the applications that we do are you know perhaps the main application of machine learning now is is advertising at least win comes financially financial incentives and these regulated domains extend to marketing so if you're advertising for a credit card or you're advertising for a hiring ad then these are regulations extend to the advertisement okay and again I leave out all sorts of government laws that also have different rules and they're legally recognized protected classes in the United States race color sex religion nationality citizenship and so on and these came about in in in different legislations a
是很重要的一个,有《平等信贷机会法》;还有教育,有1964年的《民权法案》;就业、住房、公共设施,这些都是私营部门领域,它们没有涵盖政府领域,那里的法律可能更复杂。对机器学习来说有一件重要的事,因为我们做的很多应用,你知道,也许机器学习现在的主要应用就是广告,至少从财务激励的角度来说是这样。而这些受监管的领域延伸到了营销,所以如果你在为信用卡做广告,或者你在投放招聘广告,那么这些法规就会延伸到那个广告上,好吧。同样,我这里略去了各种政府法律,它们也有不同的规则。在美国,有法律认可的受保护类别:种族、肤色、性别、宗教、国籍、公民身份等等。而这些是在不同的立法中形成的,
便签笔记
13:24
different you know at different times so it's not like there was just one you sort of coherent anti-discrimination statute or that came about at some point these things were a hard fought over time through decades of activism and were often responses to political movements and events like what we now call the civil rights movement of the 1960s and they continue to be important and relevant and in flux these kind of categories okay so just recently a couple of weeks ago there was a big Supreme Court decision that said you know basically we affirm that what I just had on the previous slide the protection of or the anti-discrimination law for sex extends to gender and extends to sexual orientation which was something that the Obama administration you know held to be true and was sort of in flux under the current administration and this decision kind of showed that or reaffirmed that it continues to be it continues to extend to these domains and protects you know gender okay and so this is an important example to show you
你知道,是在不同的时间点形成的。所以并不是说曾经有一部统一连贯的反歧视法规,或者说它是在某个时间点一下子出现的。这些东西是经过几十年的行动主义抗争,一点点争取来的,而且往往是对政治运动和事件的回应,比如我们现在所说的1960年代的民权运动,它们至今仍然重要还有相关性,而且这些类别一直在变动。好,就在最近,几周前有一项重要的最高法院判决,它说基本上就是,我们确认了我上一张幻灯片上写的内容,也就是对性别的保护,或者说反歧视法中关于 sex(生理性别)的规定延伸到 gender(社会性别),也延伸到性取向。这一点在奥巴马政府时期是被认定为成立的,而在本届政府下有点摇摆不定,这项判决算是表明了,或者说重申了它继续成立,继续延伸到这些领域,保护所谓的 gender。好,所以这是一个很重要的例子,让你看到这些事情一直在演变,一直存在争议,而且
便签笔记
14:43
that things continue to evolve and are continue to be contentious and you know continue to be a source of political discussion okay so the United States there are two legal doctrines primarily and I'll state them briefly so that you know when you hear these terms what that roughly means there are two ways that you can try to show that somebody discriminated against you one of them is disparate treatment and disparate treatment tries to capture things like purposeful consideration of group membership if I actually look you're you know your group membership status in one of these protected groups and make a decision based on that and it's tries to capture intentional discrimination with even if the group membership is not explicitly considered but I have an intention a purpose of discrimination that I'm pursuing and then there is an end you could say that the goal here is to capture procedural fairness to capture that you know the procedure by which you sort people it doesn't on the face of it sort of
一直是政治讨论的来源。好,那么在美国主要有两种法律原则(doctrine),我简单说一下,这样你以后听到这些术语时大致知道是什么意思。有两种方式可以用来证明某人歧视了你,其中一种叫差别对待(disparate treatment),差别对待试图涵盖的是这样一些情况:有意地考虑群体身份,比如我确实去查看了你在这些受保护群体中的身份归属,并据此做出决定。它试图涵盖的是故意歧视,即使群体身份没有被明确地纳入考量,但我心里有一个歧视的意图、一个歧视的目的。然后还有另一种,你可以说这里的目标是捕捉程序公平,也就是确保你对人进行区分所用的程序,表面上看不构成歧视性做法。还有另一种原则叫差别影响(disparate impact),它试图涵盖的是
便签笔记
15:51
constitute a discriminatory practice there's another doctrine called disparate impact which tries to capture sort of more indirect forms of discrimination and tries to gather things that are sort of avoidable or unjustified harm possibly indirect and so the goal here is more distributive justice or to minimize differences in outcomes so disparate impact would allow you to argue even if sort of procedurally this looks okay what the decision-maker is doing somehow down the road there's a big disparity and you're you're trying to hold the the employer responsible for on the this disparity that they cost in their decision-making and people have recognized some tension between the two so you know sometimes avoiding disparate impact so avoiding that you have some kind of disparity between groups down the road you might want to take group membership explicitly into account and that is something that it poses sort of a tension between these two doctrines and that has certainly also been something that the formal work
那些更间接的歧视形式,试图囊括那些本可避免的或者说不正当的伤害,可能是间接造成的。所以这里的目标更偏向分配正义,或者说尽量减少结果上的差异。因此差别影响允许你这样主张:即使程序上看起来没问题,决策者的做法在后续某个环节还是造成了很大的差距,而你是想让雇主为他们在决策中造成的这种差距负责。人们也意识到这两者之间存在张力。比如说,有时候为了避免差别影响,也就是为了避免在后续环节出现群体之间的某种差距,你可能就需要把群体身份明确地纳入考量,而这恰恰在这两种原则之间造成了某种张力,这一点当然也是这个领域的形式化研究一直在挣扎的问题。所以稍后我们会看到,我们要讲的一些公平性
便签笔记
17:03
in this area has struggled with so later we'll see that the some of the fairness criteria that we'll go through make explicit reference to a group membership and as such you can ask if that would be a problem for under the disparate treatment doctrine so would they constitute this pair of treatment okay this just to give you a sense I'm you know not a legal scholar there's obviously a lot more that's going on here I just want you to have sort of working grasp what this looks like and importantly there are some caveats about the law that I think are worth keeping them in mind there it's not like the anti-discrimination law in the United States reflects one coherent moral theory and it's just sort of an opera operationalization of that moral theory that's not what it is right so these legislations were responses to to activism to civil rights movements each of them came about over time and was the bait you know involved you know many fights and with some kind of agreement between different different
标准会明确地引用群体身份,因此你就可以问,这在差别对待原则下会不会是个问题,也就是它们是否构成差别对待。好,这只是给你一个大致的感觉,我不是法律学者,这里显然还有很多更复杂的东西,我只是想让你对它大概是什么样子有一个可用的把握。而且很重要的是,关于法律有一些需要注意的地方,我认为值得记在心里。美国的反歧视法并不是说体现了某一套连贯的道德理论,它也不只是那套道德理论的一种操作化实现,事实并非如此,对吧。这些立法是对社会运动、对民权运动的回应,它们各自是随着时间逐步出现的,其间经历了很多博弈,并且是在不同政党之间达成某种妥协的产物,对吧。所以这并不是……它比
便签笔记
05「无意识即公平」为何失败
18:11
political parties right and so this is not you know it's more fragmented than just one sort of top-down idea of discrimination that takes on different shapes okay and in particular the law does not give us a fairness definition that we could readily formalize or operationalize so as computer scientists if we're trying to work with formal models on this topic we can't just turn to the law and be like okay here's the definition and we just need to sort of formalize this and the right way so this is not how that's going to how that's going to work unfortunately okay so then what is the first thing that kind of comes you know I want you to understand about this topic and I think the first thing that we the community sort of learned is that there's no such thing as fairness through awareness an awareness okay no such thing as ignoring or removing sensitive attributes hoping that this will somehow you know get rid of the concern okay so the idea that if you just remove all the all the things that
某种自上而下的、单一的歧视概念要碎片化得多,只是在不同场合呈现出不同形态。好,特别要说的是,法律并没有给我们一个可以直接形式化或操作化的公平定义。所以作为计算机科学家,如果我们想在这个话题上使用形式化模型,我们不能直接求助于法律,然后说:好,定义在这儿,我们只要以正确的方式把它形式化就行了。很遗憾,事情不会是这样运作的。好,那么关于这个话题,第一件我希望你理解的事情是什么呢?我认为我们这个社群学到的第一件事就是,不存在所谓的“通过无意识实现公平”(fairness through unawareness)。好,也就是说,指望忽略或删除敏感属性,就能以某种方式消除这个问题,这是行不通的。好,那种想法就是:只要你把所有那些
便签笔记
19:24
look like a sensitive attribute meaning like an encoding of a protected group remove that from the data scrub that from your data and then hope for the best okay that goes terribly wrong and there's a you know the the books I mentioned contain numerous examples and I'll just give you one to give you a sense of this here was an article from Lumbergh from 2016 about Amazon's same day delivery coverage and so basically what happened is that Amazon rolled out same-day shipping to a bunch of US cities and if you look at Boston for example you will see that there is a single neighborhood that is left out okay so the blue shaded area is same-day coverage the gray area is no coverage and this neighborhood turns out to be Roxbury which is a predominantly black or african-american neighborhood in Boston and that was the one neighborhood that was excluded okay and so people you know looked at these maps and they were like wait a minute this looks very alarming because it looks quite similar actually to these
看起来像敏感属性的东西——也就是对受保护群体的某种编码——从数据里删掉,把它从你的数据里清洗干净,然后祈祷一切顺利。好,这会出大问题。我提到的那几本书里有大量这样的例子,我这里只举一个,让你有个直观感受。这是彭博社 2016 年的一篇报道,讲的是亚马逊当日送达的覆盖范围。基本情况是,亚马逊在美国一批城市推出了当日送达服务,如果你看波士顿的例子,你会发现有一个街区被排除在外。好,蓝色阴影区域是当日送达覆盖区,灰色区域是不覆盖。而这个街区正好是罗克斯伯里(Roxbury),它是波士顿一个以黑人或非裔美国人为主的街区,而它就是唯一被排除在外的街区。好,于是人们看了这些地图,就说:等一下,这看起来非常令人警觉,因为它其实跟以前那些“红线”(redlining)地图很像。这是一张
便签笔记
20:33
old redlining maps here's one of Philadelphia we're lending officers had secret maps where they kind of you know had sort of scratched out certain areas who call it them red as hazardous here where you wouldn't do business you wouldn't give out loans in those neighborhoods okay and part of the equal affair of Fair Credit Opportunity Act is it was to to get rid of practices such as this that you could no longer do this so people looked at these same day coverage maps and and saw those assemblance and so this is an example of like you know the failure of fairness to unawareness Amazon almost certainly didn't look at race and that's also what the article said in fact the article said maybe they should look at it but maybe Amazon I don't know I wasn't there but perhaps Amazon was just trying to predict the number of purchases that were made that would be made and in these zip codes or in these neighborhoods and that correlates with socioeconomic status which in the United States correlates with race and so maybe
费城的地图,当时放贷人员有一些秘密地图,他们会把某些区域划掉,用红色标注为“高风险”,意思是在那些街区你不做生意、不发放贷款。好,《平等信贷机会法》的部分目的就是要消除这类做法,让你不能再这么干。所以人们看这些当日送达覆盖地图时,看出了那种相似性。这就是一个例子,说明“通过无意识实现公平”是行不通的。亚马逊几乎可以肯定并没有看种族这个变量,报道里也是这么说的,事实上那篇报道还说他们也许应该去看一看。但也许亚马逊——我不知道,我又不在现场——但也许亚马逊只是想预测这些邮编区或这些街区里会产生多少订单量,而这跟社会经济地位相关,在美国社会经济地位又跟种族相关。所以也许他们只是在预测购买量,也就是一个
便签笔记
21:36
they were just predicting purchases so a standard machine learning problem purchase prediction you look at past data you do your machine learning thing and then you draw the map okay and so this is an example of just using machine learning without concern for these problems we'll just kind of you know make you lead to you know some issues down the road and you will run into these kind of outcomes that strike people as as unfair okay and so things like we don't consider that in our data as never relevant and it bears being very clear about this because as of like I think at least 2016 companies would certainly often make this argument that we don't consider that in our data you know our models are neutral and so on okay so what should we do instead or what have people try to do instead if if ignoring the problem doesn't work what else can we do okay and so here's my overview for what I'm going to tell you about in part one I'll talk about sort of a standard perspective or what I call maybe a narrower perspective and talk
标准的机器学习问题:购买预测。你看历史数据,做你那套机器学习流程,然后画出地图。好,所以这就是一个例子:在不关心这些问题的情况下直接使用机器学习,就会在后续带来一些问题,你就会遇到这类让人觉得不公平的结果。好,所以像“我们的数据里没有考虑这个变量”这种说法,从来都不构成免责理由。这一点值得说清楚,因为我觉得至少到 2016 年,企业确实还经常拿这个来辩解:我们的数据里没有考虑那个,我们的模型是中立的,等等。好,那么我们应该怎么做呢,或者说如果忽略问题行不通,人们尝试过什么别的做法?我们还能做什么?好,接下来是我要讲的内容的总览在第一部分,我会讲一个比较标准的视角,或者说我称之为相对狭义的视角,来谈
便签笔记
22:48
about fairness criteria and classification so we'll talk about the standard decision theory set up and supervised learning and we'll talk about what kind of fairness criteria our people have discussed in the context of classification and we'll see some of the relationships and limitations that they have and on Thursday a works sort of toward a broader perspective and entertain and particularly two avenues of research one is causal models of decision making and what that can say about fairness and the other our direction is dynamic models of socio technical systems and how that might lead to a different view of this problem there's a new message there was a question from young could would should one relate this issue to the very problematic I don't see color statement yeah so I think that is definitely related so this idea of colorblind policies is definitely related to sort of decide your fairness through unawareness that somehow if you don't look at you know you know somebody's race or status in some protected group
公平性标准和分类问题,所以我们会讲标准的决策论设定和监督学习,还会讲人们在分类的语境下讨论过哪些公平性标准,我们会看到其中一些相互关系和它们的局限性;而周四那一讲会走向更宽的视角,具体会展开两条研究路线,一条是决策的因果模型,以及它对公平性能说些什么,另一条方向是社会技术系统的动态模型,以及这可能如何带来对这个问题的不同看法。有一条新消息,有个来自 Young 的问题:能不能把这个议题和那个很成问题的说法"我看不见肤色"联系起来?是的,我认为这肯定是相关的,所谓"色盲"政策的想法,肯定和那种通过"无意识"来实现公平的思路有关,就是说如果你不去看某个人的种族,或者他在某个受保护群体中的身份,
便签笔记
06统计公平的学术渊源与形式框架
24:02
that's somewhere that makes the problem go away so yes I think that is certainly related so this is a good point maybe to ask more questions about the frame because I'm sort of leaving the part of giving a little bit of background and I'm jumping into some of the technical material so if anyone has more questions about about this net would be a good time let's give you a moment to maybe type in a question before I move on okay so it seems like there is currently no other question so let me jump into the first part part oh yes so I think I'm just in case anyone has a question you can raise your hand in ensue and then you can directly ask the question yeah but but we don't have yeah never no worse let me let me move ahead and go into part one about sort of statistical fairness criteria okay so where did formal work on fairness and classification come from and computer science is actually not at all the first to study this and in fact there was pioneering work in the educational testing communities of people that are
问题就好像自然消失了。所以是的,我认为这当然是相关的。这可能是个好时机,大家可以多问一些关于这个框架的问题,因为我差不多要讲完背景介绍的部分了,接下来我要进入一些技术内容了,所以如果还有人对这部分有更多问题,现在会是个好时机。我给大家一点时间打字提问,然后我再往下讲。好,看起来目前没有其他问题了,那我就进入第一部分吧——哦对,我想说的是,万一有人有问题,你也可以举手,然后直接口头提问。嗯,不过我们现在没有……好吧,没关系,那我就往下讲,进入第一部分,关于统计意义上的公平性标准。好,那么关于公平性和分类的形式化研究是从哪里来的呢?计算机科学其实完全不是最早研究这个的,事实上教育测评领域早就有开创性的工作,那些人关心的是标准化考试之类的问题,
便签笔记
25:31
concerned with you know standardized testing and things like this and a notable scholar here is n Cleary who in 1968 already wrote a paper on different fairness criteria for educational testing so what it means for a test to be fair and that's very similar to two sort of machine learning because you know you're scoring people in the end or some outcome and and you're trying to basically predict somebody's performance and so there are group differences in testing and that's that's where they were this first surfaced there's also early work in economics Gary Becker rode his thesis and 1957 introduced a notion of taste based discrimination Phelps and arrow and the 70s important work on statistical discrimination and those ways of thinking about discrimination in terms of taste based discrimination and statistical discrimination are still very much the way that economists conceptualize some of these questions and so this is some of the early work computer scientists computer science join mostly post 2010
其中一位很有影响的学者是 N. Cleary,她在 1968 年就已经写了一篇论文,讨论教育测评中不同的公平性标准,也就是一场考试怎样才算公平。这跟机器学习非常相似,因为你最终也是在给人打分,或者预测某种结果,本质上你是在预测某人的表现。考试中存在群体差异,这就是这个问题最早浮现的地方。经济学里也有早期工作,Gary Becker 在 1957 年的博士论文中提出了基于偏好(taste-based)的歧视概念,Phelps 和 Arrow 在70 年代做了关于统计性歧视的重要工作。这些从基于偏好的歧视和统计性歧视来理解歧视的方式,至今仍然是经济学家概念化这些问题的主要方式。所以这些是早期的工作。计算机科学家大多是 2010 年之后才加入的,可能 2008 到 2009 年有一些早期工作,
便签笔记
26:45
some early work you know in 2008-2009 maybe and then an explosive increase in work since 2016 so in 2016 a tipping point where computer science I think embraced this problem of fairness it became like a major topic at our conferences and people started working on this very broadly so you might ask why today and you know why didn't this work back then already solved these questions I think today we're seeing a different urgency scale and reach of algorithmic decisions and maybe they're going the the idea of making algorithmic decisions has gone much further than we initially thought and so right now we're confronted as machine learning researchers and engineers were confronted with making these decisions normative judgments about how certain things are supposed to be done and we're not that well prepared to them to make these judgments in a lot of cases so that's why this is kind of an urgent problem right now and you could ask why study this in the context of machine learning why machine learning and not
然后从 2016 年起工作量爆炸式增长。2016 年是个转折点,我认为计算机科学在那时接纳了公平性这个问题,它变成了我们会议上的一个主要话题,人们开始非常广泛地研究它。所以你可能会问为什么是今天?为什么当年那些工作没有把这些问题解决掉?我认为今天我们面对的是不同的紧迫性、规模和算法决策的影响范围,也许算法决策这个想法已经走得比我们当初设想的远得多。所以现在,作为机器学习研究者和工程师,我们被迫要去做这些决定,去做关于某些事情应该怎么做的规范性判断,而在很多情况下我们并没有做好做这些判断的准备。这就是为什么这在当下是个很紧迫的问题。你也可以问,为什么要在机器学习的语境下研究这个?为什么是机器学习,而不是
便签笔记
27:58
decision theory or economics and so on and these fields also study that and it should continue to do so but machine learning feels a lot of the adoption of algorithmic decision making and as such motivates of new technical problems but importantly you know it this is sort of where the where rubber meets the road right this is where kind of these problems come up and as a result we have to we're forced to to grapple with them ok we can no longer ignore them so let me remind you of just a standard formal prediction and decision-making setting this is just standard decision theory so you have data described by covariates X think of that as a random variable X could be multi-dimensional doesn't to be a single thing could be like an array of features if you want and there's an outcome variable Y often called often binary and most of my examples going to be binary and that sometimes called a target variable and the goal is generally to predict Y given X okay to assign label to X that agrees with y
决策论或经济学等等?这些领域也在研究,而且应该继续研究。但机器学习承载了算法决策的大量应用落地,因此催生了新的技术问题。但更重要的是,这就是所谓"实际落地、真刀真枪"的地方,这些问题就是在这里冒出来的,结果就是我们不得不去正面应对它们,好吧,我们再也无法忽视它们了。那么让我先帮大家回顾一下标准的形式化预测与决策设定,这就是标准的决策论。你有由协变量 X 描述的数据,把它想成一个随机变量,X 可以是多维的,不一定是单一的东西,如果你愿意,它可以是一组特征。然后有一个结果变量 Y,通常是二值的,我大部分例子都会是二值的,它有时也被称为目标变量。目标通常是在给定 X 的条件下预测 Y,好吧,就是给 X 分配一个与 y 一致的标签。所以 x 和 y 都是同一空间上的随机变量。通常
便签笔记
29:02
okay so x and y are both random variables in the same space and you know typically we use supervised machine learning to produce a score function R that is a function of X that becomes a new random variable I call it capital R the score function to kind of summarize these these covariates into a single real valued score and then we make binary decisions according to the threshold rule D which just says 1 if the score is above some threshold and 0 otherwise okay so this is your threshold in the score on the end and if somebody's above the threshold they get accepted you know can interpret one as acceptance C or as rejection and the lowest threshold you get rejected okay so this should strike you as very familiar and these are random variables in the same probability space and define a population over typically individuals in in the settings that will be concerned with so where do score functions come from this is the Munim board many others in this workshop in summer school are talking about you
我们用监督式机器学习来得到一个打分函数 R,它是 X 的函数,于是它就成了一个新的随机变量,我把它叫做大写的 R,这个打分函数把这些协变量归纳成一个实数值的分数。然后我们根据阈值规则 D 来做二元决策,也就是分数高于某个阈值时输出 1,否则输出 0。好,这就是你在分数上设的阈值,如果某人在阈值之上,他就被接受了。你可以把 1 解释为接受,把 0 解释为拒绝,低于阈值就被拒绝。好,这些应该会让你觉得非常熟悉。这些都是同一概率空间上的随机变量,并定义了一个总体,在我们要讨论的场景里这个总体通常是由个体组成的。那么打分函数是从哪来的呢?这也是这个研讨班和暑期学校里 Munim 以及其他很多人在讲的内容。打分函数可以来自数据的某种参数化模型,它
便签笔记
30:11
know score functions could come from some parametric model of the data it could be like a likelihood ratio test that's sort of where the decision theory look the decision theory is about so if you have a parametric form for the data you can you can come up with your score function that way you could consider nonparametric scores such as the base optimal score which is just the expectation of the target variable given the data that is what's called the Bayes optimal score or at least Bayes optimal on squared loss and that's an important score function that all sometimes make reference to and most commonly you know these core functions learn from label data using unsupervised learning using supervised learning so use any of these techniques people have deep learning whatever and you come up with a score function that minimizes some loss function on a set of label data label training data for the purpose excuse me for the purpose of this talk how you get the score function is and not the focus of this talk so the focus
可以是类似似然比检验的东西,这大致就是决策论关心的内容。所以如果你对数据有一个参数化形式,你就可以那样构造出打分函数。你也可以考虑非参数的分数,比如贝叶斯最优分数,它就是给定数据时目标变量的期望,这就是所谓的贝叶斯最优分数,至少是平方损失下的贝叶斯最优。这是一个重要的打分函数,我后面有时会提到它。而最常见的情况是,这些打分函数是用监督学习从带标签的数据中学出来的——抱歉,是监督学习。你可以用人们有的任何技术,深度学习什么的,得到一个打分函数,使它在一组带标签的训练数据上最小化某个损失函数。就本次讲座而言——不好意思——就本次讲座而言,你怎么得到这个打分函数并不是重点,本次讲座的重点是你拿它来做什么,也就是
便签笔记
31:15
of this talk is what you do with it the decision part okay so we'll mostly talk about decisions and not about the training process so we'll forget about optimization and regularization we'll even forgot about forget about generalization because we'll talk about the population level okay so I'm writing these random variables because I'm talking about everything at the population level I'm not talking about the you know finite sample issues I'm not talking about the difference between population and finite sample okay so I'm talking about decisions at the population level and so when comes to decision-making there's the standard confusion table here where you contrast the outcome that could either be 1 or 0 with a decision that could also be 0 1 and if both are 0 you call that a true negative at both are 1 you call that a true positive these are some of your good cases where you did it right and default the wrong cases could be a false positive where the outcome is 0 but you declared it to be 1 so that's a false
决策的部分。好,所以我们主要谈决策,而不谈训练过程,所以我们不去管优化和正则化,我们甚至不谈泛化,因为我们讨论的是总体层面。好,我之所以写成这些随机变量,是因为我讲的一切都是在总体层面上,我不讨论有限样本的问题,我也不讨论总体与有限样本之间的差别。好,所以我谈的是在总体层面上的决策。那么说到决策,这里有个标准的混淆矩阵,你把结果(可以是 1 或 0)和决策(也可以是 0 或 1)对照起来看。如果两者都是 0,就叫做真负例;两者都是 1,就叫做真正例,这些是你做对了的好情况。而做错的情况可能是假正例,也就是结果是 0 但你判成了 1,那就是假
便签笔记
32:14
positive and it could be that the outcomes actually 1 but you declared to be 0 that's a false negative you said negative when it was actually a positive instance ok oops ok this is a copy and paste error this disparate impact was a copy and paste thing forget about what's what I want to say is that there's something called the truth see I'm gonna this is so confusing I'm just gonna erase that and go back to the presentation so and once you have to find these you can talk about rates you can talk about the rate of making positives true positives that's the probability that the decision is 1 conditioned on the outcome being 1 you can talk about the false positive rate that's the condition probability of making one decision given that the outcome is you similarly the true negative rate is the probability of negative decision given a negative outcome and the false negative rate is the probability of a negative decision given a positive outcome so these are the four possibilities here on this confusion table and these are our
正例;也可能结果其实是 1 但你判成了 0,那就是假负例,也就是实际是正例你却说成了负例。好,哎呀,这里是个复制粘贴的错误,这个"差别性影响"是复制粘贴带过来的,别管它。我想说的是,还有一个叫做……唉,太乱了,我就直接把它擦掉,回到幻灯片吧。好,一旦定义了这些,你就可以谈各种"率"了,你可以谈做出正例判断的率,也就是真正例率,它是在结果为 1 的条件下决策为 1 的概率。你也可以谈假正例率,也就是在结果为 0 的条件下做出正例决策的条件概率。类似地,真负例率是在结果为负的条件下做出负例决策的概率,而假负例率是在结果为正的条件下做出负例决策的概率。所以这就是这个混淆矩阵上的四种可能情况,也是从这张表里导出的四个很自然的统计量。你
便签笔记
07标准一:独立性与接受率均等
33:22
four natural ways of you know natural statistics derive from this table you could think of these as kind of normalizing by the rows so these are sometimes called the row wise statistics because they talk about the rows of this confusion table so in your in your [Music] numerator if you will you're talking about rows up this table okay so that's just that's the standard picture and so what do statistical fairness criteria do so far we haven't talked about the switch to standard decision theory we haven't talked about fairness at all how does fairness come in the simplest way that people have sort of tried to work fairness into the picture or at least group differences into the picture is to introduce an additional random variable a that encodes membership status in a protected class so a could be some encoding a race some encoding a gender or disability status and so on never mind how you encode that and that there may be harm to such encodings in the first place and collecting such data
可以把它们理解成按行做归一化,所以它们有时被称为按行的统计量,因为它们讲的是这张混淆矩阵的行。所以在分母(呃,分子)里,你说的是这张表的行。好,那这就是标准的图景。那么统计意义上的公平性标准是做什么的呢?到目前为止我们还没谈公平——我们讲的都是标准决策论,完全还没谈公平。公平是怎么进来的呢?人们把公平,或者至少把群体差异,纳入这个框架的最简单方式,就是引入一个额外的随机变量 A,它编码了一个人在受保护类别中的身份。所以 A 可以是某种种族的编码、某种性别的编码,或者残障状况等等。先不管你怎么编码它,也不管这种编码本身以及收集这类数据本身可能带来伤害。总之你引入这个随机变量 A,它可以
便签笔记
34:29
in the first place but you introduce this random variable a it could be correlated with X it doesn't have to be independent of X it could be part of X even it could be part of the features that you already have you just give it another letter okay you just introduce another letter for it and then what you can do once you you have this group membership you can talk about different statistics different statistical quantities involving your membership and then you can talk about group differences you talk about how these statistical quantities differ by group okay and that's exactly what what people have done here so the idea dates back to you know like I said that at least the 1960s with a work of and Clery who already did exactly this so it's it's you know it's equivalent to to what you proposed back then to look at these group difference statistical group group differences and we'll review three common criteria and I'll walk you through them and almost everything that's been done is related
和 X 相关,不必与 X 独立,它甚至可以是 X 的一部分,可以就是你已有特征里的一个,你只是给它另起了一个字母而已。好,你只是为它引入了另一个字母。然后一旦你有了这个群体身份,你就可以谈涉及这个身份的各种统计量、各种统计指标,接着你就可以谈群体差异了,也就是这些统计量在不同群体之间有何不同。好,这正是人们在这里做的事。这个想法可以追溯到,像我说的,至少 1960 年代 Cleary 的工作,她当时做的正是这件事。所以这跟当年提出的做法是等价的,都是去看这些群体之间的统计差异。我们会回顾三个常见的标准,我会带大家逐一过一遍,几乎所有已有的工作都以某种方式和这三个标准之一相关。好,所以如果你理解了
便签笔记
35:29
in one way or the other to one of these three criteria okay so if you understand these criteria typically if you see something in the literature you can morally relate it to one of these three maybe not necessarily formally but it'll be at least conceptually related to one on the first one that comes to mind and this is I'd say the most common one is equalizing the acceptance rate okay so you can simply say that for any two groups a and B we're going to require that the rate of making positive calls in that group is is or is is equal in all groups okay so the rate of positive acceptance or acceptance decisions should be the same in all group okay if one group has an eighty percent acceptance chance every other group must also have an eighty percent acceptance rate okay so the acceptance rate is equal in all groups you can generalize this by requiring the D is statistically independent of a I an independent of Al sometimes just call that the independence criterion so you can convince yourself that if the decision
这些标准,通常你在文献里看到什么东西,都能大致把它归到这三个之一,也许在形式上不完全一致,但至少在概念上是相关的。第一个,也是最容易想到的,我认为也是最常见的一个,就是让接受率相等。好,你可以简单地要求,对于任意两个群体 A 和 B,在该群体中做出正例判断的比例是相同的,也就是在所有群体中都相等。好,所以接受决策的比例在所有群体中应该一样。好,如果一个群体的接受概率是百分之八十,那么其他每个群体也必须有百分之八十的接受率。好,所以接受率在所有群体中相等。你可以把它推广为要求 D 在统计上与 A 独立,与 A 独立,有时就把这叫做"独立性"标准。你可以自己说服自己:如果决策
便签笔记
36:35
is independent of a then these conditional probabilities will all work out to be the same okay so you can just do the math and you're in your head probably if de the decision and the attribute a are independent then all these conditional probabilities acceptance probabilities have to be the same and you get this okay you can also apply this to the score function and ask that the scores independent of a and then if you threshold the score that'll also ensure equal positive rates so these things are all closely related and you know you can ask for if you ask for the score being independent of a I find that a very natural way of stating it and that implies all these other things and again there are all sorts of variance relaxations equivalent formulations that people have proposed of this criterion so let's start with stress testing this criterion a little bit so let's try to think about situations where enforcing this criteria asking for it does not rule out unfair practices so where a decision-maker
与 A 独立,那么这些条件概率就都会是一样的。好,你在脑子里算一算就行了,如果决策 D 和属性 A 独立,那么所有这些条件概率、也就是接受概率,都必须是一样的,于是你就得到了这个。好,你也可以把它用在打分函数上,要求分数与 A 独立,那么你对分数做阈值处理后,同样能保证各组的正例率相等。所以这些东西都是紧密相关的。你知道,如果要求分数与 A 独立,我觉得这是一种非常自然的表述方式,而且它蕴含了其他所有这些性质。同样,人们还提出了各种各样的变体、放松版本和等价的表述形式,都是围绕这个标准的。那我们先来给这个标准做点压力测试,我们来想想在哪些情况下,强制满足这个标准、要求它成立,并不能排除不公平的做法。也就是说,决策者可以在满足这个标准的同时,
便签笔记
37:39
could hide blatantly unfair and some intuitive sense decisions you know while satisfying this criterion and the example I like to give and there are other issues with it but the example I like to give is that you could make good and informed decisions in one group maybe you have a lot of experience with one group and not the others and so you make poor arbitrary random decisions in another group but you do it in such a way that you match the equal that the positive rate okay so you could say in this majority group you're doing what your obvious always been doing and you have a lot of experience for what you're doing and in some other representative group you just toss a coin and select random folks and you know this is not you know I'd say this can happen naturally on its own if you have less data or poor data on one group and so your model is basically not working well on that group let me give you an example there's something called the Framingham risk score for coronary heart disease
掩盖某种直觉上明显不公平的决策。我喜欢举的例子——当然这个标准还有其他问题——我喜欢举的例子是:你可以在一个群体中做出很好、很有依据的决策,也许你对某个群体有大量经验,而对其他群体没有,于是你在另一个群体中做出糟糕的、任意的、随机的决策,但你做的方式恰好让正例率相等。好,所以你可以说,在这个多数群体里,你照着你一贯的做法来,你对自己在做什么有大量的经验,而在另一个群体里,你就抛硬币,随机挑一些人。而且你知道,这种情况其实完全可能自然而然地发生:如果你在某个群体上数据更少或数据质量更差,你的模型在那个群体上就基本不管用。我给你举个例子,有一个叫 Framingham 冠心病风险评分的东西,
便签笔记
38:40
that was created on a cohort of white men I think at the beginning of the 20th century if I recall correctly and then it was used for other patients and its performance was much worse for other patients okay and so this idea of making poor decisions in in one group and in good decisions in another is of course concerning and the moral intuition here is that you shouldn't be able to match true positives in one group with false positives in another right I shouldn't be able to say oh in this group I have lots of true positives and I'm making up for it by having lots of false positives in the other group because you know if you accept you know random instances presumably you're not selecting the best instances that you could selecting the best people I could and as a result you sort of setting up a poor track record for that group okay so I think this is an issue with that kind of criterion I should say nonetheless this criterion is very popular and there's a lot of work on it there was in particular one thing
如果我没记错,它是在 20 世纪初基于一个白人男性队列建立的,后来被用在其他病人身上,而它在其他病人身上的表现要差得多。好,所以这种在一个群体里做出糟糕决策、在另一个群体里做出良好决策的情况,当然是令人担忧的。这里的道德直觉是,你不应该被允许用一个群体里的真正例去抵消另一个群体里的假正例,对吧,我不该能够说,哦,在这个群体里我有很多真阳性,而我通过在另一个群体里制造很多假阳性来弥补这一点在另一个群体里,因为你知道,如果你接受随机的样本,那你大概就没有在挑选最好的你本可以挑选的样本——挑选我能找到的最优秀的人——结果你就等于给那个群体制造了一份很差的历史记录所以我认为这类标准是存在问题的。不过我得说,这个标准非常流行,围绕它有大量的工作。我特别想提一下的是 Rich Zemel
便签笔记
39:43
I want to mention is work by rich Seema and others on fair representation we're roughly speaking the goal is to use you know deep learning tricks such as adversarial learning and other things to train a representation of the data that is independent of group membership while representing the original data as much as possible okay so you start from some kind of feature representation X and you try to get another feature representation called as Z that's that's you know as representative as possible of X while being independent of the of the of group membership and so people sort of think of this as kind of data D biasing and there are some issues with that with that notion that we'll return to in a moment but there continues to be a lot of work for it what I want you to sort of keep in mind is you know when you see one of these papers that says Oh fair representation learning or something you know ask yourself what is actually the fairness criterion that's being achieved okay and the underlying
等人关于公平表示(fair representation)的工作,粗略地说,目标是利用深度学习的技巧,比如对抗学习之类的方法来训练一种与群体归属无关的数据表示,同时又尽可能多地保留原始数据的信息好,所以你从某种特征表示 X 出发,试图得到另一种特征表示,叫做 Z,它尽可能地能代表 X,同时又独立于群体归属。人们某种程度上把这看作是数据去偏(de-biasing)。这个概念存在一些问题我们待会儿会回来讨论,但这个方向仍然有大量的工作在进行。我希望你们记住的是,当你看到这类论文说“哦,公平表示学习”之类的时候,你要问自己,实际上被满足的公平性标准到底是什么。而底层的公平性标准并不会因为
便签笔记
40:52
fairness criterion doesn't get better by using deep learning okay so just because there's a lot of machinery and a lot of parata's in achieving this criterion that doesn't necessarily mean there's somehow it's it's more fair or there's something more to it than just you know independence okay so under the hood it's still the same independence criterion that that you saw okay so there might be a question so the question for Martin is for binary decisions is having different acceptance thresholds for different groups enough where this criterion yes so this is a good question in particular an easier way to achieve independence if you just want independence you could just accept the you know you could have group specific thresholds and set them up in such a way that the acceptance rate is the same right and so this is kinda gets to my point like if there is such a straightforward way to achieve this criterion what are you hoping to get out of the more sophisticated way and maybe try to be specific or precise about what
用了深度学习就变得更好。所以,仅仅因为在实现这个标准的过程中用了大量的机器和大量的参数,那并不必然意味着它就更公平,或者说除了独立性之外还有别的什么内涵好,所以在底层,它仍然是你们刚才看到的同一个独立性准则。好的,可能会有问题,所以这个问题是问 Martin 的:对于二元决策来说,给不同群体设置不同的接受阈值是否就够了?这个准则——是的,这是个好问题。特别是,如果你只是想要独立性,有一个更简单的办法可以实现它,你完全可以直接接受……你知道,你可以设置群体特定的阈值,并且这样设定,使得接受率相同,对吧?所以这其实正好呼应了我的观点:如果有这么一个直截了当的方法就能满足这个准则,那你希望从更复杂的做法里得到什么?也许可以试着把这一点说得具体或精确一些。我觉得,有些论文在这方面做得比
便签笔记
08标准二:错误率均等与 ROC 曲线
42:04
this this is and I think you know some of these papers do a better job at this than others and I like Rich's work so this is not specific to the initial work here more of a general point about the this line of work there's a another question equalizing acceptance rates seems like equality of outcome yeah under some notion of equality of outcome I think you're right but we might you know you know that I think equality of outcome could mean different things and I don't want to necessarily use that terminology because I already use outcome variable for Y so I don't necessarily want to phrase it that way but I think you probably have to write the right thing in mind so that seems it seems good so okay let me move on so recognizing this issue of trading off false positives with true positives which seems concerning it seems to violate some moral intuition people proposed another set of criteria which equalizes the error rates okay so it specifically does not allow you to trade true and false positives and it will
另一些好。我很喜欢 Rich 的工作,所以这并不是针对这里最初那篇工作,更多是关于这一整条研究路线的一个总体评论。还有另一个问题:让接受率相等看起来像是结果平等。是的,在某种结果平等的概念下,我认为你说得对,但我们可能……你知道,我觉得结果平等可以有不同的含义,而且我不太想用这个术语,因为我已经把 Y 称作结果变量了,所以我不太想那样表述。但我想你脑子里想的大概是对的,所以这看起来……看起来不错。好,那我继续讲。所以,意识到这种在假阳性和真阳性之间权衡的问题——这看起来令人担忧,似乎违背了某些人的道德直觉——人们提出了另一套准则,让错误率相等。好的,它明确地不允许你去权衡真阳性和假阳性,它会要求所有群体有相同的假阳性率,并且所有群体有
便签笔记
43:19
require that all groups have the same false positive rate and all group have the same false negative rate okay so you know these these rates have to be these two rates have to be the same in all groups okay and again you can generalize this which is kind of nice you can require a deed the decision to be independent of a given the outcome variable why okay so instead of having just independence you now have a conditional independence and that's one way to generalize this it also makes sense for score functions again you can require R to be independent of a given Y and again if that's what you require all these other things follows if your score for instance independent if a given Y you get that for the decision if you use the threshold rule and you get that for the false and true positive rates as well okay so this is maybe the most simple and gentle way of stating this and so one thing to realize about error rate parity which makes it a little bit tricky is that it's a post hoc criterion
同样的假阴性率,好的,所以你知道,这些这些率必须——这两个率必须在所有群体中都相同,好的。同样地,你可以把它推广,这一点还挺不错的:你可以要求 D,也就是决策,在给定结果变量 Y 的条件下独立于 A,好的。所以不再只是独立性,现在变成了条件独立性,这是推广它的一种方式。这对分数函数同样成立,你可以要求 R 在给定 Y 的条件下独立于 A。同样地,如果你要求的是这个,其他所有东西都会随之成立。比如说,如果你的分数在给定 Y 的条件下独立于 A,你就能在使用阈值规则时把这一点带到决策上,而且假阳性率和真阳性率也同样满足。好的,所以这大概是陈述这件事最简单、最温和的方式。关于错误率平等(error rate parity)有一点需要意识到,这让它有点棘手:它是一个事后(post hoc)标准。让我解释一下这是什么意思。在做决策的时候,当你需要
便签笔记
44:22
so let me explain what that means so at decision time when you need to actually say yes or no the decision maker doesn't know who is a positive and who's a negative instance okay so they cannot evaluate at decision time these kind of you know error rate error rates they're not they're unknown to them okay so in hindsight somebody can collect a group of positive instances you know people that have turned out to have a positive outcome and if you feel that have turned out to have a negative outcome and you can do that in each group and you can see you can look at how they were classified so after the fact with the benefit of hindsight you can you know said test these error rates but you can do that ahead of time at decision time and so group differences you know or error rate differences are sort of you know in this kind of audit setting where you after the fact you look at the decisions often strike people as unfair so again there's some strong moral intuition here that well wait a minute if all these you know
真的说“是”或“否”时,决策者并不知道谁是正样本、谁是负样本,好的,所以他们在决策时无法评估这类错误率,这些率对他们来说是未知的,好的。所以事后来看,有人可以收集一组正样本,也就是那些最终有正面结果的人;以及那些最终是负面结果的人。你可以在每个群体里都这么做,然后你就能看到,你可以去看他们当初是怎么被分类的。所以在事情发生之后、有了事后信息,你就可以去检验这些错误率,但你没法提前、在决策的时候做这件事。所以群体之间的差异,或者说错误率差异,在这种事后审计的场景里——你在事后回看这些决策——常常让人觉得不公平。所以这里同样有一种很强的道德直觉:等一下,如果所有这些正样本,也就是那些实际上有
便签笔记
45:28
positive instances people with a with a good outcome ended up or were actually classified negatively you know at a much higher rate than another group that's you might seem unfair okay so it seems like one group if this if this criterion is violated it seems like one group but they're a disproportional burden of uncertainty so it would be a sort of the front of these you know statistical you know errors and more than other groups okay and so in some sense maybe a different way of stating this is that you want the harm or the disadvantage from miss classification you want that to be the same in all groups you can visualize this using the ROC curve the ROC curve is a simple way to talk about threshold rules where you plot for every possible threshold from between 0 and 1 let's say your school function is between zero one you plot the true positive rate you plot the false positive rate that typically gives you you know a curve you can plot these curves for each group or a curve that looks like this you can plot these for
好结果的人,最终却被判为负面,而且比例远高于另一个群体,那你可能会觉得这不公平,好的。所以如果这个标准被违反了,看起来就像是有一个群体承担了不成比例的不确定性负担,也就是说他们会更多地承受这些统计上的错误,比其他群体承受得更多,好的。所以从某种意义上说,换一种表述方式就是:你希望误分类带来的伤害或不利,在所有群体中都是一样的。你可以用 ROC 曲线把这一点可视化。ROC 曲线是讨论阈值规则的一种简单方式:对于每一个可能的阈值——假设你的分数函数取值在 0 到 1 之间——你画出真阳性率、画出假阳性率,这通常会给你一条曲线。你可以为每个群体画出这些曲线,或者一条长这样的曲线。你可以为每个群体画一条,所以这就是你的阈值规则在每个群体中
便签笔记
46:37
each group so this is the true and false positive rates that your threshold rule achieves in each group as you slide the threshold you trace out these curves one for each group and what you know error rate parody tells you or what this criterion will tell you is that the RC curve of your score if it satisfies this criteria and must be must be under all under all the curves okay so the area that's under the intersection of all these curves that are these are the trade-offs between true and false positive rate that you can realize while satisfying this criterion okay so this is the thing that you can achieve and so one interesting point about this is that if you look at this you can see if there's a question David is asking for something like hiring it seems like you would not have the ability to evaluate these counterfactual even in hindsight does this criteria not make sense in such thing so that's a great question David's pointing out that sometimes you know we simply if you don't get accepted
达到的真阳性率和假阳性率。当你滑动阈值时,你就描出了这些曲线,每个群体一条。而错误率平等告诉你的,或者说这个标准告诉你的是:如果你的分数满足这个标准,它的 ROC 曲线必须在所有曲线之下,好的。所以位于所有这些曲线交集之下的那块区域,这些就是你在满足这个标准的前提下能实现的真阳性率和假阳性率之间的权衡,好的。所以这就是你能达到的范围。关于这一点有一个有趣的地方:如果你看这张图,你就能看出来,如果——有个问题,David 问:对于像招聘这样的场景,似乎你连事后也没有能力去评估这些反事实,那这个标准在这种情况下是不是就没有意义了?这是个很好的问题。David 指出的是,有时候我们根本就——如果你没有被录取,我们就根本看不到
便签笔记
47:48
you know we simply do not see you know your outcome ever we just have no record of it let's say if you don't get into Harvard we never know what your GPA at Harvard would have been okay and so that's that's an important point and the answer is yes so if people do this for lending or yes this is an issue that you need to address so if you do this for lending for example people will try to estimate these air rates by saying well you were turned down by this bank but you got alone with this other bank and you paid back your loan with that other Bank and so you probably would have also paid it back with that first Bank that towards you turned you down and so that gives us a sense of you know how we can estimate these things but yeah and in general you know these things are you know like potential outcomes you only observe one of them but not both and so in general these are can be hard to estimate is a good question okay so returning to this image of these intersecting oracy curves one thing that
你的结果,我们完全没有相关记录。比如说,如果你没被哈佛录取,我们永远不会知道你在哈佛的 GPA 会是多少,好的。所以这是很重要的一点,答案是:是的。如果人们在信贷场景里这么做,这确实是一个你需要处理的问题。比如说在信贷里,人们会试着这样估计这些错误率:他们会说,你被这家银行拒了,但你在另一家银行拿到了贷款,而且你在那家银行按时还清了贷款,所以你很可能在第一家拒绝你的银行那里也会还清。这就给了我们一点感觉,知道怎么去估计这些量。但确实,总体来说这些东西就像潜在结果(potential outcomes)一样,你只能观察到其中一个,看不到两个。所以一般来说这些可能很难估计。这是个好问题。好的,回到这张相交的 ROC 曲线图,你会看到的一点是:
便签笔记
48:56
you see is that you know if you enforce this criterion that might imply that in some group you're not doing as well as you could okay so it means that somehow you're actually making the decisions worse for some group okay and this is I think a valid concern with this criterion that it could mean you're you're doing more poorly in some group just because you want to equalize the error rates and then that could constitute you know harm against that group okay because you could you could be doing better and so that is somehow unjustified that you're you're harming them in sort of doing worse than you could okay so that's that's the story about error rates and again like I said before returning to this this confusion cable table you could also talk about what happens when you swap you know the conditioning so instead of talking about these row wise criteria you talk about column wise criteria so you could equalize expressions of the form probability that y equals something given that the decision is something and
如果你强制执行这个标准,那可能意味着在某个群体里你做得不如你本可以做到的那么好,好的。也就是说,你其实是在让某个群体的决策变差了,好的。我认为这是对这个标准一个成立的质疑:它可能意味着你在某个群体里表现更差,只是为了把错误率拉平,而这可能构成对那个群体的伤害,好的,因为你本可以做得更好,所以这在某种程度上是没有正当理由的——你在伤害他们,做得比你本可以做到的更差,好的。所以这就是关于错误率的故事。再一次,就像我之前说的,回到这张混淆矩阵,你还可以讨论把条件反过来会发生什么。也就是说,不再讨论这些按行的标准,而是讨论按列的标准。所以你可以去拉平这样形式的表达式:在决策是某个值、群体是某个值的条件下,
便签笔记
09标准三:校准与充分性
50:09
that group is something okay so this is you're swapping Y and D now Y appears in the argument and D appears in the condition before it was the other way around so you could swap this conditioning these are called column-wise rates and they're also statistically meaningful specifically the two standard column-wise rates are false emission rate and false discovery rate depending on which setting of Y and D you choose here and these are two important things that could study in statistics and you could again say okay what fairness criteria do you get when you when you you know equalize those okay so there's syntactically there's nothing wrong here you could do that as well but there's something closely related that people do instead which is you know I'd say a bit more natural and has sort of gained more traction and that's the idea of calibration so let me tell you what calibration is so a score R is calibrated if the probability that you see a positive outcome given score value little R is equal to little R okay
Y 等于某个值的概率,好的。所以这里你把 Y 和 D 调换了,现在 Y 出现在参数里,D 出现在条件里,之前是反过来的。所以你可以调换这个条件,这些被称为按列的率,它们同样具有统计意义。具体来说,两个标准的按列率是错误遗漏率(false omission rate)和错误发现率(false discoveryrate),取决于你在这里选择 Y 和 D 的哪种取值。这是统计学里两个重要的量,你同样可以问:如果把这些拉平,会得到什么样的公平性标准?好的,所以从形式上看这没什么问题,你完全可以这么做。但有一个与之密切相关、人们实际上更常用的东西,我觉得它更自然一些,也获得了更多的关注,那就是校准(calibration)的概念。让我讲讲什么是校准。一个分数 R 是校准的,如果在分数值为小 r 的条件下看到正面结果的概率恰好等于小 r,好的。这意味着你可以把
便签笔记
51:16
so what that means is you can pretend your score as a probability although it might not actually be one it's certainly not at an individual level but on average your score you know behaves like a probability so if you see somebody with score value 0.8 you know that they have a positive outcome rate of 0.8 okay so you know if I know you have if I see you're at your your heart disease risk score is 0.9 then I know you have our 90% chance of heart disease among people on average over people with a score of 0.9 okay so and then you could ask that your classifier is not just calibrated on the whole population but calibrated within each group okay so you can require this stronger condition that you satisfy calibration in each group okay so you throw in an extra conditioning on the group and you want that your score behaves like a probability within each group so even if you restrict your attention to a specific group you know that within that group do your score says or you know means what it claims to
你的分数当作一个概率来看待,尽管它实际上可能并不是——在个体层面上肯定不是——但平均而言,你的分数表现得像一个概率。所以如果你看到某个人的分数值是 0.8,你就知道他们的正面结果发生率是 0.8,好的。所以你知道,如果我看到你的心脏病风险分数是 0.9,那我就知道你有 90% 的心脏病概率——是在所有分数为0.9 的人中平均而言,好的。然后你可以进一步要求,你的分类器不只是在整个人群上校准,而是在每个群体内部都校准,好的。所以你可以要求这个更强的条件:在每个群体中都满足校准,好的。于是你在条件里额外加上群体这一项,你希望你的分数在每个群体内部都表现得像一个概率。所以即便你只把注意力限制在某个特定群体上,你也知道在那个群体内部,你的分数说的就是它字面上
便签笔记
52:28
say okay and so if you think about it you can again show that this of calibration by group follows from a conditional independence statement and that statement is why the outcome variable is independent of group membership conditional on the score it's a different way of saying that is for the purpose of predicting the outcome your score contains everything you need to know about the group okay so in other words if I'm my only goal is to predict why the outcome then after observing the score are I have no interest in also soliciting your group membership and so that's not a very natural guarantee that you know if I see if I know your score I don't need to ask what your group membership is so if I see you're 0.7 that means the same thing in all groups I don't have to ask you know hey also please tell me your group membership and if your group membership is this I'm going to add point one if your group membership is that I'm going to subtract point two thank you I don't need to do those mental gymnastics I know that the
声称的意思,好的。而且如果你想一想,你同样可以证明按群体的校准可以由一个条件独立性陈述推出,那个陈述就是:结果变量 Y 在给定分数的条件下独立于群体成员身份。换一种说法就是:为了预测结果,你的分数已经包含了关于群体你需要知道的一切,好的。换句话说,如果我唯一的目标是预测结果 Y,那么在观察到分数 R 之后,我就没有兴趣再去打听你的群体成员身份了。这是一个非常自然的保证:如果我知道你的分数,我就不需要再问你属于哪个群体。所以如果我看到你是 0.7,这在所有群体中意思都一样,我不需要说,嘿,还麻烦你告诉我你属于哪个群体,如果你属于这个群体我就加零点一,如果你属于那个群体我就减零点二,谢谢。我不需要做这些心算体操,我知道这个
便签笔记
53:35
score means the same thing in all groups and as a result calibration is an a priori guarantee okay so the decision-maker sees the score value R and at decision time knows that based on this what the frequency of a positive outcome is okay so score point eight means an 80% rate of heart failure on average over people who receive score point eight within each group okay so you don't need to solicit group membership but let's not get carried away by calibration it's not meant to be you know sort of a guarantee about individuals okay so just because Mary gets a score of point eight doesn't mean that Mary individually like conditional on her exact features has a heart failure a chance of 80% that's not what it means it just means on average over people that received the score point eight the frequency of positive outcomes is eighty percent okay so that's calibration what another thing that you should be aware of is that group calibration and calibration in general often follows from unconstrained learning okay
分数在所有群体中含义相同。因此校准是一个事前(a priori)的保证,好的。所以决策者看到分数值 R,在决策的时候就知道,基于这个值,正面结果的频率是多少,好的。所以分数零点八意味着在每个群体内部,得到零点八分的人平均有 80% 的心衰发生率,好的。所以你不需要去打听群体成员身份。但我们也别对校准过度乐观,它并不是要给出关于个体的保证,好的。所以,仅仅因为 Mary 得了零点八分,并不意味着 Mary 个人——在她确切特征的条件下——有 80% 的心衰概率,它的意思不是这个。它只是说,在所有得到零点八分的人中平均而言,正面结果的频率是 80%,好的。这就是校准。另一件你应该知道的事情是:群体校准,乃至一般意义上的校准,常常是无约束学习自然带来的结果,好的。
便签笔记
54:44
so it's not necessarily a constraint on I'm Christine learning it's it's often what machine learning gives you in the first place okay so they're you know as a theorem that says under some conditions the deviation from satisfying group calibration so how far off you are if you want to relax that is upper bounded by how good your score is relative to the Bayes risk or relative to the base optimal score okay so if you compare your score function you'll earn score function to the Bayes optimal score the difference in performance will give you an upper bound on the violation of group calibration so the better you you better you are at machine learning the closer you are to the base optimal score the better the more your group calibration is satisfied okay there's a from work with Lydia Liu and Maxim Cove it's from a couple years ago and so in other words you shouldn't be surprised to see calibration follow approximately from unconstrained supervised learning and here's a picture if you just run
所以它不一定是对机器学习的一种约束,它往往本来就是机器学习会给你的东西,好的。有一个定理说,在某些条件下,偏离群体校准的程度——也就是如果你允许放松,你偏离了多远——是有上界的,这个上界取决于你的分数相对于贝叶斯风险、或者说相对于贝叶斯最优分数有多好,好的。所以如果你把你学到的分数函数和贝叶斯最优分数比较,性能上的差距就给出了群体校准违反程度的一个上界。所以你的机器学习做得越好,你越接近贝叶斯最优分数,你的群体校准就满足得越好,好的。这来自我和 Lydia Liu 以及 Max 几年前的一项工作。所以换句话说,你不应该对“校准近似地由无约束的监督学习自然得到”这件事感到意外。这是一张图,如果你就用逻辑回归在 UCI Adult 数据集上开箱即用地跑一下,完全不做调参
便签笔记
55:46
logistic regression on the UCI adult data set out-of-the-box just no tuning or whatever you will see this calibration plot okay so if you look at the score deciles from one to ten and you look at the rate of positive outcomes you will see that these are basically both diagonal lines for male and female which means that if it were perfectly calibrated you'd see a diagonal line and this is pretty close to it there was another question from the chat the question is how does the score calibration linked to problems of moral relevance we could be equally calibrated but people of group could end up being worse in reality how do we just scores for this so this is a great question that's exactly and I like the way you think you already internalize this idea of moral relevance so one example that can show you an issue with calibration and I think it's it's also you know just an issue with optimal learning right so optimal machine learning is always always calibrated right the Bayes optimal score is always
或者随便怎么跑,你就会看到这张校准图。好,如果你看从一到十的分数十分位,再看正例结果的发生率,你会发现男性和女性基本上都是对角线,这意味着如果它是完美校准的,你就会看到一条对角线,而这个已经相当接近了。聊天区还有一个问题,问题是:分数校准和道德相关性的问题之间有什么联系?我们可能在各组上都同样校准,但某个群体的人最终在现实中处境更糟。我们该怎么针对这一点调整分数?这是个很好的问题,正是如此,我很喜欢你的思路,你已经内化了道德相关性这个概念。有一个例子可以向你展示校准存在的问题,而且我觉得这其实也是最优学习本身的问题,对吧。最优的机器学习永远都是校准的,对吧,贝叶斯最优分数总是校准的。所以这里就有一个问题,一个最优分数可能存在的问题。比如说我想预测
便签笔记
56:53
calibrated so here's an issue potentially an issue with an optimal score let's say I want to predict productivity in the workplace and now define productivity as the number of hours you work in the next ten years and I'm so good at machine learning that I in fact give you the exact optimal base base score okay so I know the exact expectation of the number of hours you work conditional on your features and so the issue with that is that somebody who has a disability or somebody who might become pregnant in the future might have a lower score simply because you know your score correctly identified this loss of work hours in the future but this is something we deem morally irrelevant to hiring right so again this is a good example of where group calibration doesn't solve this problem of moral relevance right generally optimal prediction does not okay so optimal prediction will use whatever it can if you know a piece of information gives you boost and predictive accuracy it'll it'll include that even if as a society
职场生产力,现在把生产力定义为你未来十年工作的小时数,而我的机器学习水平高到我实际上能给你精确的最优贝叶斯分数。好,也就是说我知道在给定你的特征下,你工作小时数的精确期望。那么问题就在于,一个有残疾的人,或者一个未来可能怀孕的人,分数可能会更低,仅仅因为你的分数正确地识别出了未来工作小时数的这种损失。但这是我们认为在招聘中道德上不相关的东西,对吧。所以这又是一个很好的例子,说明群体校准并不能解决道德相关性的问题,对吧。一般来说最优预测做不到这一点。好,最优预测会利用一切能用的东西,如果某条信息能提升你的预测准确率,它就会把它用上,哪怕作为一个社会
便签笔记
10子群体漏洞与不可能性定理
58:03
we deem it irrelevant okay so that's a great question another thing I should say but group about all these group fairness definitions and this is not specific to calibration but it you know it holds for all of them if you ensure fairness criteria between two groups that can lead to violations and often more striking violations within one of these groups okay and so this motive motivated work on ensuring group fairness and there were two papers that started this line of work one by Kearns Neil Roth and whoo and the other by Herbert Johnson Kim Rangel and Roth learn and so one of the illustrative examples from Kurds at all is the following made-up made-up situation where you have two groups blue and green and you accept you know the same fraction of individuals in each of these two groups on indicated by these circles and however in one group you only accept male candidates and the other group you only accept female candidates and so if you now look within these subgroups you get a terrible
我们认为它是不相关的。好,这是个很好的问题。关于所有这些群体公平性定义,我还应该说一件事,这不是校准特有的,而是对所有定义都成立:如果你在两个群体之间保证公平性标准,这可能会导致违反,而且往往是在其中某个群体内部出现更触目惊心的违反。好,这就催生了关于保证子群体公平性的工作,有两篇论文开启了这条研究线,一篇是 Kearns、Neel、Roth 和Wu 的,另一篇是 Hébert-Johnson、Kim、Reingold 和 Rothblum 的。Kearns 等人给出的一个说明性例子是这样一个虚构的情形:你有两个群体,蓝色和绿色,你在这两个群体中接受了同样比例的个体,就是这些圆圈表示的。然而,在其中一个群体里你只接受男性候选人,而在另一个群体里你只接受女性候选人。所以如果你现在往这些子群体里看,你就会发现独立性标准被严重违反了。好,你也可以对所有其他
便签笔记
59:07
violation of you know the independence criteria okay and you could imagine examples for all the criteria as well and I'd say this this is a real concern you know the the UPenn group around occurrence calls this fairness gerrymandering and I like that expression I think this is a real concern if you enforce some constraint usually machine learning will satisfy this constraint in the laziest possible way and so you should not expect to get anything clever you know within the subgroups if you equalize some kind of statistic okay you probably will see something annoying or something alarming going on in the subgroups okay so let's recap what where we are and you know have another half an hour that's plenty of time for for add so that's recap we saw three criteria and the first of them is you know the independence criteria the second one is this conditional independence of our giving are as independent of a conditional on Y and the third one is you know another conditional independence where Y is
标准想象出类似的例子。我要说这是一个真实的隐患,宾大 Kearns 那个团队把这叫做“公平性选区划分(fairness gerrymandering)”,我很喜欢这个说法。我认为这是个真实的隐患:如果你强加某个约束,机器学习通常会以最偷懒的方式满足这个约束,所以你不应该指望在子群体里得到什么聪明的结果。如果你把某种统计量拉平了,好,你多半会看到子群体里发生一些让人不舒服、甚至令人警觉的事情。好,我们来回顾一下讲到哪儿了,我们还有半个小时,时间很充裕。那么回顾一下,我们看了三个标准,第一个是独立性标准;第二个是这个条件独立性,即在给定 Y 的条件下 R 与 A 独立;第三个是另一种条件独立性,即在给定 R 的条件下 Y 与 A 独立。第一个意味着相等的
便签笔记
60:26
independent of a conditional R and you know the first one implies equal acceptance rate second one equal error rates the third one implies calibration Burgh group so this is the way I like to state them just because this makes it really easy to you know I appreciate them in terms of these conditional independencies and I think you've all at this point program you've had some lectures on graphical models or causality so you understand conditional independence at this points probably this will be also a good way for you to memorize these and so one question that people will ask you know once you lay them out once you lay out multiple fairness criteria you can ask him we have them all so is there a way to maybe you know get the best of all worlds and create some amazing predictor that just satisfies these all and people quickly realize the answer's no so there is an if formally I could state this as any two of these three criteria are mutually exclusive in general so except under degenerate conditions you cannot have
接受率,第二个意味着相等的错误率,第三个意味着分组校准。这就是我喜欢的表述方式,因为这样特别容易用这些条件独立性来理解它们。我想在座各位到这个阶段应该都上过一些关于图模型或因果推断的课,所以你们大概已经理解条件独立性了,这对你们记住这些标准也是个好办法。那么一旦你把多个公平性标准摆出来,人们会问一个问题:我们能不能全都要?有没有办法两全其美,造出某个既满足这个又满足那个的神奇预测器?人们很快就意识到答案是否定的。形式化地说,可以表述为:这三个标准中任意两个一般来说是互相排斥的。所以除非在退化的情形下,你不可能同时拥有它们全部。我只给你们讲一个形式化的
便签笔记
61:28
them all and I'm gonna just give you one formal statement and and exclude the other statements that you could prove here table which is error rate parity versus calibration the theorem goes as follows if you assume unequal base rates so you assume that you know that you know two groups have different different rate of positive outcomes so your groups are unequal in terms of outcomes so currently one of the groups has a high rate of positive outcomes then another and if you assume that your decision rule is imperfect excuse me so you assume that your your decision has nonzero error rates so it makes at least one false positive and makes at least one false negative call and this is the general situation in general because we live in a world of inequality groups tend to have different base rates and also you know typically we don't have a perfect predictor so our particulars make errors and under these assumptions which are the general case if you satisfy calibration by group then this implies that error parity fails
命题,其他你们本可以在这里证明的命题就略过了,这个命题是关于错误率均等与校准的。定理是这样的:如果你假设基础率不相等,也就是假设两个群体有不同的正例结果发生率,也就是说你的群体在结果上是不平等的,其中一个群体的正例结果率比另一个高;再如果你假设你的决策规则是不完美的,也就是说你的决策有非零的错误率,它至少会做出一个假阳性判断和至少一个假阴性判断——这就是一般情况,因为我们生活在一个不平等的世界里,各群体往往有不同的基础率,而且通常我们也没有完美的预测器,所以我们的预测器会犯错——在这些假设下,也就是一般情况下,如果你满足分组校准,那么这就意味着错误率均等
便签笔记
62:45
error rate parity fails in particular if you just do unconstrained machine learning and you get a calibration break group that also means that you by default have a violation of air a parity okay so enforcing one of them makes the other become untrue and there's been the initial works here whereby children each OVA and Kleinberg milena town and Raghavan who proved similar statements not exactly this and then there's been a lot of follow work on what if you relax them and it's not so not so easy to say what happens in some of these cases when you relax it and there's possibility death for some of these realizations some trade-offs are is some non-trivial trade-offs are possible but what I want to go into instead is how these these trade-offs you know dominated sort of an academic discourse around an important problem namely that of recidivism prediction and and in the United States and there was a sort of landmark or landmark article by ProPublica called machine bias that alleged that there was
不成立。特别地,如果你只是做无约束的机器学习,然后得到了分组校准,那也就意味着你默认就违反了错误率均等。好,所以强制满足其中一个就会让另一个不成立。这方面最初的工作是 Chouldechova,以及 Kleinberg、Mullainathan 和Raghavan,他们证明了类似的命题,不完全是这一个。之后又有很多后续工作研究如果你放松这些条件会怎样,在其中一些情形下放松之后会发生什么并不那么容易说清楚,有些情况下确实存在可能性,某些非平凡的权衡是可以做到的。但我更想讲的是,这些权衡是如何主导了围绕一个重要问题的学术讨论的,那就是美国的再犯预测问题。ProPublica 有一篇算是里程碑式的文章,叫做《机器偏见(Machine Bias)》,声称有一款在全国各地用来预测未来罪犯的软件,
便签笔记
11COMPAS 之争与两重告诫
63:52
a software used across the country to predict future criminals and inspires against blacks and this was you know written in 2016 I think came out of May 2016 and it sparked a lot of conversation around this topic and the essence of the compass debate at least serve if you projected to academic circles what academics talked about in computer science at least there is a risk or you know called Kemp comm pass in many jurisdictions in the United States use the score to evaluate the likelihood that a defendant will go on to commit another crime if there were to be released okay so if you have to decide whether to detain or release a defendant that goes up for trial this is used to estimate the risk of you know that defendant committing another crime if they were released okay and judges may detain the defendant in part based on this score this is not the only thing I look at but they can use the score and so probably looked at defendants that ultimately did not receive a so ultimately did not commit another crime
对黑人存在偏见。这篇文章是 2016 年写的,我记得是 2016 年 5 月发表的,它引发了围绕这个话题的大量讨论。COMPAS之争的核心,至少投射到学术圈之后,至少在计算机科学界大家谈论的是:有一个风险分数,叫做 COMPAS,美国许多司法辖区用这个分数来评估一名被告如果被释放,将来再次犯罪的可能性。好,所以如果你要决定对一名等待审判的被告是羁押还是释放,这个分数就用来估计该被告如果被释放再次犯罪的风险。好,法官可以部分基于这个分数来决定羁押被告,这不是他们看的唯一因素,但他们可以使用这个分数。于是 ProPublica 考察了那些最终没有再次犯罪的被告,他们注意到,在这些
便签笔记
64:59
and they noticed that of these you know defendants that ultimately did not commit another crime black defendants were much more likely to be labeled high risk by the score than white defendants okay so another way it's saying this is that the false positive rate of the score appears to be much higher for black defendants than for white defendants okay so black defendants are much more likely to be labeled high risk even though they do not go on to commit another crime okay and so compounds the you know that the maker of comm-pass called Northpoint he said yes that's true but our scores are calibrated by a group and black defendants have a higher rate of recidivism in the United States and hence this is unavoidable okay so they're saying we're just doing our job here we're calibrating the score that's what we need to do if the score wasn't calibrated the judge would have to do some mental gymnastics around maybe I have to add three points if the defendant is white and and subtract one of the defendants black and so on and to
最终没有再次犯罪的被告中,黑人被告被这个分数标记为高风险的可能性远高于白人被告。好,换一种说法就是,这个分数的假阳性率对黑人被告来说似乎远高于白人被告。好,也就是说黑人被告更有可能被标记为高风险,尽管他们并没有再次犯罪。好,于是 COMPAS 的开发方,叫做 Northpointe,他们说:是的,这是真的,但我们的分数是分组校准的,而在美国黑人被告的再犯率更高,因此这是不可避免的。好,他们的意思是我们只是在做我们该做的事,我们在校准这个分数,这是我们必须做的。如果分数没有校准,法官就得在脑子里做一堆体操,比如“如果被告是白人我是不是得加三分,如果被告是黑人得减一分”等等。为了
便签笔记
66:10
avoid that you know these scores have to be calibrated by group and that inevitably implies that you get this this disparity okay so this is what kind of you know this this discourse you know or a lot of what this discourse was about for better or worse and you know academics were intrigued by this idea that this complex problem seemed to have boiled down to some statistical trade-off that you could you know approach with mathematical tools and rigor okay so let me add a first word of caution to the contest debate and another word of caution to it the first word of caution is that you know both air raid parody and calibration do not rule out unfair practices in the setting and so it's it's already a bit of a straw man to argue about because neither of these rule out things you should actually worry about okay and so in particular what's fair in criminal justice is not settled by a field to one of these criteria so we can reduce the question of justice and/or criminal justice to one of these criteria and I
避免这种情况,这些分数必须分组校准,而这就不可避免地意味着会出现这种差异。好,这大致就是这场讨论的样子,或者说这场讨论的很大一部分内容,不管是好是坏。学者们对这个想法很感兴趣:这个复杂的问题似乎归结成了某种统计上的权衡,可以用数学工具和严谨的方法去处理。好,那么让我对 COMPAS 之争提出第一点告诫,以及另一点告诫。第一点告诫是,错误率均等和校准都不能排除这个场景下的不公平做法。所以争论这个本身就有点像稻草人,因为这两个都不能排除你真正应该担心的事情。好,特别是,刑事司法中什么是公平的,并不能由这些标准中的某一个来决定。所以我们不能把正义、把刑事司法的问题化约成这些标准之一。我
便签笔记
67:25
want to give you an example of construction if you will to to illustrate this point particularly these two two properties are not meant to be fairness certificates they're not proofs of fairness so let me give you an example of why error rate parity is an issue in this kind of setting so consider these two groups and let's say we have two groups the black and the blue group and the orange group and our a threshold rule is that we detain everyone above a score 0.5 and we release everyone below score 0.5 and our scores are you know written here under these stick figures so they go from point 1 to point 7 the blue groovin from point 2 to point 9 and they in the in the in the orange group and let's further assume that these scores are actually the probability of recidivism for these you know these individuals so let's assume that these scores are just the conditional expectation of recidivism given the given the individual's covariance okay so these are the true probabilities if you will and based on
想给你们举一个构造性的例子来说明这一点,特别是这两个性质并不是公平性证书,它们不是公平性的证明。那么我先给你们举个例子,说明为什么错误率均等在这种场景下是有问题的。考虑这两个群体,假设我们有两个群体,蓝色群体和橙色群体,我们的阈值规则是:分数高于 0.5 的人全部羁押,低于 0.5 的人全部释放。分数就写在这些火柴人下面,蓝色群体的分数从 0.1 到 0.7,橙色群体的从 0.2 到 0.9。再进一步假设这些分数其实就是这些个体的再犯概率,也就是假设这些分数就是在给定个体协变量条件下再犯的条件期望。好,所以这些是真实的概率,可以这么说。基于这些真实概率,你可以算出羁押率,也可以算出
便签笔记
68:40
these true probabilities you can compute the detention rate and you can compute the false positive rate and you'll see that the blue group has a detention rate of 38 percent and the orange group has a detention rate of 61 percent so you detain substantially more people and the orange group so that means independence is violated these scores I certainly don't satisfy independence and also the false positive rate is much higher in the orange group so they don't satisfy error rate parity either so you know the false positive rate in the orange group is much higher than in the blue group so you might say how do we fix this all right so if you're if you're let's say a police department and you're you're accused of like having this this you know detention rate that's much higher and one group than the other what could you do about it and the point of this example is that there are really bad ways of achieving error rate parity and really bad ways of achieving you know independence for example one thing you could do you could
假阳性率。你会看到蓝色群体的羁押率是 38%,而橙色群体的羁押率是 61%。所以你在橙色群体里羁押了明显更多的人,这就意味着独立性被违反了,这些分数肯定不满足独立性。而且假阳性率在橙色群体里也高得多,所以它们也不满足错误率均等。橙色群体的假阳性率远高于蓝色群体。那你可能会说我们怎么修复这个问题?好,假如你是一个警察部门,你被指责说你在一个群体上的羁押率比另一个群体高得多,你能做点什么呢?这个例子的要点在于,实现错误率均等有一些非常糟糕的办法,实现独立性也有一些非常糟糕的办法。举个例子,你可以做的一件事是,你可以
便签笔记
69:48
please the orange group more aggressively and you could say we're going to arrest more low-risk individuals in the orange group so more individuals that are have a very low risk of recidivism as a result we're you know we're going to arrest them what we're going to release them and so we're simply padding the number of you know defendants in the orange group that we're gonna release and if we do that you know we simply arrest more people in that group if we do that we can actually make it such that the tension rate drops to 42% and the false positive rate drops to 26% and now these things look like they're similar so there's a question here that I wanna get ad says there's a Q&A but I don't quite see how do I get to the Q&A maybe somebody can just send this in Chad I don't nothing happens the question to you so how should we choose the right criteria for a specific problem any suggestion yeah how should you choose the the right criterion for a specific problem so this is a good question and I'm gonna return
更严厉地对橙色群体执法,你可以说我们要逮捕橙色群体里更多的低风险个体,也就是更多再犯风险非常低的人,结果就是我们把他们抓起来,然后再把他们放掉。这样我们只是在给橙色群体里我们会释放的被告人数注水。如果我们这么做,就是在那个群体里多抓一些人,这样做之后我们实际上可以让羁押率降到 42%,假阳性率降到 26%,现在这些数字看起来就很接近了。这里有个问题我想回答一下,有人说有个 Q&A,但我不太知道怎么进到 Q&A,也许有人可以直接发到聊天区。我点了没反应。那个问题是:我们该如何为一个具体问题选择正确的标准?有什么建议吗?对,你该如何为一个具体问题选择正确的标准。这是个好问题,我过一会儿会回到这个问题。让我先把这个例子讲完,
便签笔记
71:02
to this in a moment let me first wrap up this example because we're in the middle of something but remind me to to get to that or ask that question again when I'm after this example when I'm discussing so the purpose of fairness criteria yes great so okay so this is alarming right so what this is saying is if you if you incentivize you know someone to satisfy error rate parity they could achieve this criterion in sneaky ways in ways that are even much more harmful or introduce additional harm okay so this is not as an incentive as an objective to pursue this can be a very problematic criteria there's also an issue with calibration that's worth pointing out again assume that these are the true probabilities of reoffending suppose we know them this is typically not true but for the purpose of this example let's suppose we do it and let's say we'd seen everyone above 0.5 and by the way I should say both of these examples are from a paper by Sam Corbett Davies Pearson fellow girl and Huck from 2017
因为我们正讲到一半,但请提醒我回到那个问题,或者等我讲到这个例子之后、讲公平性标准的目的时再问一次。好,太好了。好,这很令人警觉,对吧。这说明的是,如果你去激励某人满足错误率均等,他们可能会以耍花招的方式达成这个标准,而且是以危害更大、或者引入额外伤害的方式。好,所以这个作为一种激励、作为一个要追求的目标,可能是个非常有问题的标准。校准也有一个值得指出的问题。再次假设这些是真实的再犯概率,假设我们知道它们——通常这并不成立,但为了这个例子我们就假设知道。假设我们羁押所有 0.5 以上的人。顺便说一句,我应该说这两个例子都来自 Sam Corbett-Davies、Pierson、Feller、Goel 和 Huq 2017 年的一篇论文,
便签笔记
72:07
so they kindly provided these examples to me so let's say we detain everyone above above 0.5 and let's say you know we have to satisfy calibration but I want to release everyone in this group I just want to give this group a break and just release everyone in this group one thing I can do with calibration is that I can take a subset of the instances that average to particular score value okay so here I'm picking out a subset of the instances point one point one point two point two point two and point six point six point six point seven point seven I can take all these instances they average two point seven and so one thing calibration allows me to do is this kind of like averaging operation where I take all their scores and I assign them the average value of their scores okay so I replace all of these scores by point four point four point four and these these scores are again calibrated okay so this this again satisfies calibration and now no one is detained everyone is above below the you know detention
他们很慷慨地把这些例子提供给了我。那么假设我们羁押所有 0.5 以上的人,再假设我们必须满足校准,但我想把这个群体里的所有人都释放,我就是想给这个群体网开一面,把这个群体里的人全都放了。在校准下我能做的一件事是,我可以取一个平均分数正好等于某个特定分值的实例子集。好,这里我挑出一个实例子集:0.1、0.1、0.2、0.2、0.2,以及 0.6、0.6、0.6、0.7、0.7,我可以把这些实例全都取出来,它们的平均值二点七。校准允许我做的一件事就是这种平均操作:我把他们所有人的分数拿过来,然后给他们都赋上这些分数的平均值。好,所以我把这些分数全部替换成零点四、零点四、零点四,而这些分数依然是校准过的。好,所以这同样满足校准性,而现在没有人被羁押了,所有人都在你们知道的那个 0.5 的羁押阈值之下。所以你看,校准这里也可能
便签笔记
12未出庭案例:该不该做预测
73:10
threshold of 0.5 so again it could be sneaky things happening with calibration that are not detected by this criterion so let me you know make this second and a sort of more important point about this criminal-justice example and then talk about the scope of this fairness criteria which was this question so this is all you know this discussion happened people discussed these fairness criteria at length in the context of criminal justice and they discover you know this curtain discussed all these trade-offs and and what you can learn from these different criteria but there's also this question of you know are we missing out substantial parts of the problem by committing ourselves to this kind of prediction perspective this kind of statistical perspective in the first place is this not a much more complex problem than what the statistical perspective would suggest okay and so the scholarly debate was around the tension between these fairness criteria so I'm rightfully pointed out that there
藏着一些暗箱操作,而这个标准根本检测不出来。那么,让我就这个刑事司法的例子讲第二点,也是更重要的一点,然后再来谈谈这些公平性标准的适用范围,这也正是刚才那个问题。所以,这一整套讨论确实发生过,人们围绕刑事司法的场景详尽地讨论过这些公平性标准,他们发现了——刚才讲到的——所有这些取舍关系,以及你能从这些不同标准中学到什么。但同时还有一个问题:我们是不是一开始就把自己限定在这种预测的视角、这种统计的视角里,从而错过了问题中相当重要的部分?这难道不是一个远比统计视角所暗示的要复杂得多的问题吗?好,所以学术界的争论是围绕这些公平性标准之间的张力展开的。有人很正确地指出,警务中存在这种反馈回路,
便签笔记
74:12
are these you know feedback loops and policing and that bias your measurements you only measure crime where you police and so that disadvantages groups where you police more in the United States that's predominantly like black neighborhoods and and that's where you collect more crimes and as well as you start from terrible data to begin with but that's also not my main point my name point is when is the issue not how we predict and what data we use and what you know objective function we we choose but the problem is that we choose to solve this problem or attack this problem formulated as a prediction problem in the first place there one example I find very persuasive and this is the issue of failure to appear in predicting failure to be in court as an important issue again in the United States criminal justice system one approach is similar to recidivism you could try to predict failure to peon courts so if you have a defendant and you set a court appointment three months from now and you don't you know you want
它会让你的测量产生偏差——你只在你派警力的地方才测得到犯罪,所以这就不利于那些你派驻警力更多的群体。在美国,那主要是黑人社区,那也正是你采集到更多犯罪记录的地方;再加上你一开始拿到的数据就非常糟糕。但这也不是我的主要观点。我的主要观点是,问题不在于我们怎么预测、用什么数据、选什么目标函数,问题在于我们一开始就选择把这个问题当作一个预测问题来表述、来处理。有一个例子我觉得非常有说服力,就是「未出庭」这个问题——把预测未出庭当成一件重要的事,同样是在美国的刑事司法系统里。一种做法跟累犯预测类似:你可以试着去预测被告是否会不出庭。假设你有一名被告,你给他定了三个月后的开庭时间,而你,你们知道的,希望他能按时出现在
便签笔记
75:15
them to show up for their court appointment and you know you wondering if they will and so one thing you could try to do is again do some sort of machine learning and predict failure to be in court and there are lots of people doing that and lots of people pushing for these things to be implemented as part of the criminal justice reform in the country and to replace you know human judges by these you know risk scores that fail to a predict failure to be in court and so the idea is if your score is high you are kept in jail and if your score is low you're released and because there could be three months or more between your this event and the court appointment if you were jailed for three months first of all that's usually expensive to put people in jail second it is devastating to the individual they lose their job they lose their social ties it's absolutely disruptive and devastating and even if they were in a fairly safe good position to begin with by the time to get out of jail and they're cleared up they're charged they
法庭上,你会想知道他到底会不会来。于是你可能会尝试的一件事,就是又一次搞某种机器学习,去预测未出庭。确实有很多人在做这个,也有很多人在推动把这些东西作为全国刑事司法改革的一部分落地实施,用这些能预测未出庭的风险分数去取代人类法官。这个想法就是:如果你的分数高,你就被继续关押;如果分数低,你就被释放。而由于从这个环节到开庭之间可能隔着三个月甚至更久,如果你被关了三个月——首先,把人关在监狱里通常成本很高;其次,这对个人是毁灭性的,他们会丢掉工作,失去自己的社会关系,这是绝对具有破坏性和毁灭性的。哪怕他们一开始处在相当安全、不错的位置上,等他们出狱、指控被撤销的时候,他们多半已经回不去了,会非常吃力地想重新站稳。所以说,
便签笔记
76:17
probably are no longer they're struggling to get back so this is you know jailing somebody a pre trial is very disruptive and so the alternative that I find you know compelling is to recognize that people fail to pee in court due to lack of childcare or transportation due to their work schedules they might have an employer that doesn't let them go or simply due to too many court appointments so there was you know at one of the you know recent conferences on fairness accountability transparency couple years ago there was somebody who had been in the position of a defendant a black defendant who had 22 court appointments before he was finally cleared of of the charge and so you had to make 22 court appointments it's not just one that's 22 court appointments before you're cleared and if you fail to appear on a court appointment number 15 then you are you know it's recorded as failure to appear even if you make all the other appointments so there the issue here is that you know if you recognize this alternative it might
审前羁押一个人是极具破坏性的。因此,我觉得有说服力的另一条路,是要认识到人们不出庭是因为缺乏托儿服务或交通条件,是因为他们的工作安排——他们的雇主可能不让他们请假,或者干脆就是因为出庭次数太多了。在最近一次关于公平、问责与透明的会议上——几年前——有一位曾经身为被告的黑人被告,他在最终被撤销指控之前一共要出庭 22 次。所以你得出庭 22 次,不是一次,是 22 次开庭,你才能被洗清。而如果你在第 15 次开庭时没有出现,那就会被记录为「未出庭」,哪怕其他所有次你都到场了。所以这里的问题在于,如果你认识到还有这种替代方案,那么更合理的做法也许是给人们提供
便签笔记
77:27
make more sense to give people you know childcare vouchers or transportation vouchers to make sure compel employers to let people show up for court appointments and to also work on reducing the court appointments so that it becomes easier to manage this burden okay and you could implement steps to mitigate these issues none of them have anything to do with prediction and this alternative is actually part of the Harris County lawsuit settlement so Harris County is a county in Texas that includes the city of Dallas and there was a lawsuit and against you know the county that's using these risk scores and the settlement requires Harris County to provide free child care courthouses develop a two-way communication system and so on and so to implement some of these these structural interventions so that poor defendants are not you know right off the bat disadvantaged and further punished by this this system okay and so in this case I find it very hard to like make a case for the prediction to for using prediction in the first
托儿券或交通补贴券,去确保、去要求雇主允许员工出庭,同时也去努力减少出庭次数,让这个负担更容易应付。好,你是可以采取这些措施来缓解这些问题的,而它们没有一个跟预测有关系。而且这个替代方案实际上已经写进了哈里斯县诉讼的和解协议里。哈里斯县是德州的一个县,包含达拉斯市,当时有人起诉了这个使用风险分数的县,和解协议要求哈里斯县在法院提供免费托儿服务、建立双向沟通系统等等,也就是落实一些这类结构性的干预措施,让贫穷的被告不至于从一开始就处于劣势、并被这套系统进一步惩罚。好,所以在这个案例里,我觉得非常难为「一开始就使用预测」这件事
便签笔记
13统计视角的局限与总结
78:29
place when it seems so much more immediate and and important to implement these other steps so my personal I always find this this line of argument very persuasive and I find it hard as a machine-learning person to continue to insist on this prediction thing in in light of these alternatives okay so this is where I'm going to return to the question that was asked right so what is what exactly is narrower around these around these statistical criteria what makes them limited right and so these statistical fairness criteria I take the data generating distribution as a given and the work with nothing but the joint statistics the observational joint sticks statistics of these random beer was ax yra and you can throw in other random variables in the same population whatever you want and so as a result whatever they can tell you is confined to what the joint statistics of the population can tell you but in particular you cannot change the population you cannot entertain service pathetical and you cannot intervene on
辩护,因为实施那些其他措施显得更直接、更重要得多。所以就我个人而言,我一直觉得这条论证线非常有说服力,作为一个搞机器学习的人,我觉得在有这些替代方案的情况下,很难再坚持非要用预测不可。好,那么现在我要回到刚才被问到的那个问题:这些统计标准到底窄在哪里?是什么让它们变得有局限?好,这些统计公平性标准把数据生成分布当作既定的,它们只处理联合统计量——也就是 A、X、Y、R 这些随机变量的观测联合统计量。你还可以往里加其他同一总体中的随机变量,随便你加。结果就是,它们能告诉你的东西完全被限制在这个总体的联合统计量所能告诉你的范围内。特别是,你不能改变这个总体,不能考虑反事实的假设情形,也不能对
便签笔记
79:34
the world that produced this joint distribution you in some sense it's very reactive you take everything as given about the world as it is and you lose a whole range of interventions and so what I'm going to do on Thursday is you know going to this question of yes if we recognize that a statistical perspective is perhaps too narrow and rules out too many interventions and just puts the focus in a place where it's not necessarily addressing the right question how do we take salient social facts and and social context into account how do we go about making our models more aware of you know the broader questions that are important here so this will be subject part two and Thursday and I'm gonna wrap up and and then take some more questions so first thing I want you to remember is fairness to run awareness fails and so the second thing is is a bit of a more nuanced question or nuanced takeaway I still find these fairness criteria interesting as and so far as they trigger you know different moral
产生这个联合分布的世界进行干预。某种意义上它非常被动:你把世界当下的样子全盘当作既定,于是你就失去了一整批可能的干预手段。所以我周四要做的,就是去谈这个问题:如果我们承认统计视角也许太窄了,排除了太多干预可能,而且把关注点放在了一个不一定在回答正确问题的地方,那我们该如何把重要的社会事实和社会情境纳入考虑?我们该怎么让我们的模型更能意识到那些更宏观、更重要的问题?这就是第二部分、周四的主题。我现在要收尾了,然后再回答一些问题。所以,第一件我希望你们记住的事是:「无意识即公平」是行不通的。第二件事则要更微妙一些。我依然觉得这些公平性标准很有意思,因为它们能激发不同的道德
便签笔记
80:44
intuitions and I think that's sort of an underrated mechanism right and so I think people often want this kind of fairness definition and then they're disappointed that these things aren't fairness definitions they are decidedly not fairness definitions they are definitely not certifying that something is fair but they do trigger and that's just an empirical statement they trigger these moral intuitions right for pública was really outraged that this error rate parody failed I mean this is a moral intuition and it's hard to argue with that moral intuition because it captures something about our moral understanding of how decisions should be made and so as a result these criteria can still be applied to surface different normative questions about decision making and sort of get a better handle on what bugs us about the way decisions are made in a certain setting and they can lead to these different trade-offs and tensions and I think that's an important conversation that is not about let's
直觉。我认为这是一个被低估的机制。我觉得人们常常想要那种「公平性定义」,然后发现这些东西不是公平性定义就很失望——它们确确实实不是公平性定义,它们绝对不是在认证某个东西是公平的。但它们确实能激发——这只是一个经验性的陈述——它们确实能激发这些道德直觉。ProPublica 当时对错误率均等不成立这件事感到非常愤慨,这就是一种道德直觉,而你很难去反驳这种道德直觉,因为它捕捉到了我们对「决策应当如何做出」的道德理解中的某种东西。所以,这些标准仍然可以被用来揭示关于决策的不同规范性问题,让我们更好地把握在某个特定情境下,究竟是什么让我们对决策方式感到不安,它们也能引出这些不同的取舍与张力。我认为这是一场重要的对话,它不是要「一劳永逸地定下什么是公平、什么是不公平」,它不是这么回事,
便签笔记
81:42
settle this once and fall what's fair what's not fair that's not what it is it's a more nuanced you know mechanism that's at play here it helps us discuss you know our our understanding and our expectations of how decisions ought to be made okay but like I said soon as the firm's criteria on their own cannot be a proof of fairness and I think you know at least within the set up the statistical setup that I described I think there cannot be a fairness definition I'm convinced that there is not satisfactory fairness definition within this statistical setup and nor are fairness criteria on their own a good objective function okay so as I showed in this example about error rate parity and policing you don't want to make these criteria incentives okay so you don't want to incentivize people to to to pursue these criteria they can be you know if you use them as a gauge as a measurement device on the world as it is that's one thing but if you ask people to optimize for them that's another thing so yeah there are there
而是一种更微妙的机制在起作用,它帮助我们讨论我们对「决策应当如何做出」的理解和期待。好,但就像我说的,这些公平性标准本身不能作为公平性的证明。而且我认为,至少在我描述的这套统计设定里,我认为不可能存在一个公平性定义——我确信在这个统计设定内不存在令人满意的公平性定义。同样地,这些公平性标准本身也不是好的目标函数。好,正如我在那个关于错误率均等与警务的例子里展示的,你不会想把这些标准变成激励。好,你不会想去激励人们去追求这些标准。如果你把它们当作一个量表、一个测量当下世界的工具,那是一回事;但如果你要求人们去优化它们,那就是另一回事了。好,有一堆问题,我看一下——
便签笔记
14问答:干预、多分类、内容审核、身份
82:57
are a bunch of questions let me see so two questions or three okay questions keep coming in so that's good so let me start with girls question hi girl thanks for joining the talk by the way say aren't you too pessimistic about prediction taking away focus from better interventions I'm thinking of settings in which we don't know how what the right interventions are and a good estimate of the conditional expectation of Y given X may actually help us identify them for example in the failure to appear example it may become salient that women with young children are more likely to fail to appear highlighting childcare as a better intervention yeah so this is this is a good point so you know and thanks asking a question that's a great question so I think you know the when people like how do people discover that you know lack of childcare is an important factor in failing to appear in court right they probably used you know statistics somewhere right I mean I'm not making a sweeping charge against the
两个问题,或者三个。好,问题还在不断进来,那很好。让我从 Girl 的问题开始,嗨 Girl,谢谢你来听这场演讲。问题是:你对「预测会把注意力从更好的干预上引开」是不是太悲观了?我想的是那些我们并不知道正确干预是什么的场景,此时对 Y 在给定 X 下的条件期望做一个好的估计,也许恰恰能帮我们把它们识别出来。比如在未出庭这个例子里,可能会凸显出有幼儿的女性更容易不出庭,从而突出托儿服务是一个更好的干预。是的,这是个很好的观点,谢谢你提这个问题,很棒的问题。所以我想说,人们当初是怎么发现缺乏托儿服务是导致未出庭的一个重要因素的?他们大概确实用了某种统计方法,对吧。我并不是在全盘指责统计学的使用,完全不是,我
便签笔记
84:05
use of statistics not by any means I you know I think that's that would in fact be go too far and be too pessimistic but they probably posed the question very differently they were asking let's do let's say causal inference on what the causes are of failing to PN court and it's true that this conditional expectation might somehow feed into that understanding and might give you you know or some form of statistics at least might give you a good answer to that question and it's also possible that's simply doing better at prediction you know might you know you might be able to look at the model and figure out some salient factors and it helps you understand that so I agree with that but I think that's a bit of a different framing and we first need to agree that this is the framing of the question that we're going to pursue and that this is the purpose of why we're doing statistics and as long as we say the purpose is to simply install these risk scores as a way to sort our pretrial detention we might accidentally forego
觉得那样其实就走得太远、太悲观了。但他们提出问题的方式很可能很不一样:他们问的是,比如说,对「未出庭的原因是什么」做因果推断。确实,这个条件期望可能会以某种方式融入那种理解,可能会给你——或者说至少某种形式的统计方法可能会给你——这个问题的一个好答案。而且也有可能,单纯把预测做得更好,你也许就能去看看模型、找出一些显著的因素,从而帮助你理解这件事。所以这一点我同意。但我认为那是一个有点不同的问题框架,我们首先得就「这是我们要追求的问题表述方式、这是我们做统计的目的」达成一致。而只要我们说目的就是把这些风险分数装上去、用来给审前羁押排序,我们就可能在无意中放弃
便签笔记
85:09
or you know forego the possibility of these different framings of the question but I think you know there will be room for a good technical work either way so I'm not you know these other interventions it's not meant to be like oh there's there's absolutely no interesting research question going on with these things it's just a reframing of the problem next questions by America which level of complication is added to the binary fairness criteria and face of multi-class and regression tasks so one thing I should say is when it comes to these conditional independencies or independent statements that are stated these directly generalized to arbitrary random variables okay so arbitrary it's certainly discrete and you can work with a little bit of effort you can also get continuous variables but certainly discrete random variables you could have any discrete group membership variable any discrete outcome variable that takes on multiple labels and so on and as a result these and conditional independent statements
这些不同的问题框架的可能性。不过我认为,不管走哪条路,都会有很好的技术工作可做。所以我不是说——那些其他干预措施——我不是要表达「哦,这些东西上完全没有有意思的研究问题」,这只是对问题的一次重新框定。下一个问题来自 America:在多分类和回归任务面前,二元的公平性标准会增加多少复杂度?有一件事我应该说明:就这些条件独立性或者所陈述的独立性命题而言,它们可以直接推广到任意随机变量。好,任意——至少离散变量肯定可以,稍微费点力气,连续变量也可以处理,但离散随机变量肯定没问题:你可以有任意的离散群体归属变量、任意取多个标签的离散结果变量等等。因此这些条件独立性命题其实已经适用于多分类和回归任务了,这也是我
便签笔记
86:13
are they already applied to multi class and regression tasks and that's part of where I like to state it that way so that's a great question thanks Indira asks what are your thoughts about fairness requirements and settings where prediction is required say content moderation recent research shows that many models flag content with African American English at higher rates are there any fairness criteria that can mitigate bias ease in these settings where prediction is a necessary and provide okay oh is it necessary even okay so I got I was able to parse that okay so that's a great question so um there are scenarios where simply for a lack of human resources we you know need to resort to some form of algorithmic moderation or content curation let's accept that premise I mean that's maybe different debate yeah but let's accept that premise and yes so I think you know a lot of things like toxicity classification sentiment classification you know abuse classification and so on suffers from terrible biases against you
喜欢这样陈述它们的原因之一。所以这是个很好的问题,谢谢。Indira 问:在那些必须做预测的场景下,你对公平性要求怎么看?比如内容审核。最近的研究显示,很多模型对含有非裔美国人英语的内容标记率更高。在这些预测是必要的场景里,有没有什么公平性标准能缓解这类偏见?好,哦,还有「它真的必要吗」。好,我理解这个问题了。好,这是个很好的问题。嗯,确实有一些场景,仅仅因为缺乏人力资源,我们不得不诉诸某种算法化的审核或内容筛选。我们先接受这个前提——这也许是另一场辩论了——但我们先接受这个前提。是的,我认为很多东西,比如毒性分类、情感分类、辱骂分类等等,都存在针对——比方说——非裔美国人说话者的严重偏见,
便签笔记
87:27
know let's say african-american speakers due to correlations and and the training data between you know you know words you that they use or you know phrases that they use that would would be scored you know and in a way that makes them more likely to be flagged and this is a big problem so yeah I think this would violate many of these criteria certainly you know error a parody for example is violated and it's again one of these things where look I think the moral intuition that's triggered here is valid right we find this like very concerning and it and it is right and so I don't think the in this case give you a way to mitigate these these issues right so I think something like a sentiment classification or toxicity classification I think these are very problematic prediction tasks to begin with just because the target variable is very ambiguous and I think it's you know I know we're working off of the premise that this is what you have to do but in some sense I would also like to revisit
这是因为训练数据中存在相关性:他们使用的某些词或某些表达会被打上某种分数,使他们更容易被标记出来。这是一个很大的问题。所以是的,我认为这会违反其中很多标准,比如错误率均等肯定会被违反。而这又是这类情况之一:我认为这里被激发出来的道德直觉是正当的,我们确实觉得这非常令人担忧,这种感觉是对的。但我不认为在这种情况下,它们能给你缓解这些问题的办法。所以我认为,像情感分类或者毒性分类这类任务,我觉得它们从一开始就是非常有问题的预测任务,仅仅因为目标变量本身就非常含混。我知道我们是在「这是你必须做的事」这个前提下讨论的,但某种意义上我也想重新审视
便签笔记
88:35
that premise because I don't necessarily think that's the right way to phrase these problems and it runs into these issues and I don't think these criteria on their own just tell you what to do I think they would just flag this as an issue that would not tell you where to go from there and that's you know currently a lot of them working and they'll be going on on what should you do in these cases another question from Aaron shouldn't calibration account for individual features to prevent the randomization in subject to group calibration inside a group exactly so yeah so this is the example of I had with a orange or the blue group where you do this kind of averaging operation and you get some undesirable results the conditional expectation of Y given X would satisfy calibration even at the at the specific feature level so for every for every X this is the condition that you write that's a much stronger condition that's basically the condition of optimal prediction so that's that's a much much stronger setting and there are
这个前提,因为我并不认为那是表述这些问题的正确方式,它会撞上这些问题。而且我不认为这些标准本身就能告诉你该怎么做,我认为它们只会把这个标记为一个但这并不能告诉你接下来该往哪走,而这正是目前很多人在做的工作,关于在这些情况下你到底该怎么办。另一个来自 Aaron 的问题:校准是不是应该考虑个体特征,以防止在满足群体校准的前提下在群体内部做随机化?完全正确。是的,这就是我之前举的橙色群体和蓝色群体的例子,你做那种平均操作,结果得到一些不理想的后果。给定 X 的 Y 的条件期望会在具体特征层面上满足校准,也就是对每一个 X 都成立,这是你可以写出的条件,那是一个强得多的条件,基本上就是最优预测的条件,所以那是一个强得多得多的设定。确实有一些论文试图逼近这种更强的条件,
便签笔记
89:38
papers you know that try to you know approach this kind of stronger condition relates to the subgroup a subgroup problem so I think gal who asked the question earlier has to work on that and and on Marengo a guy rust loom Michael came at Stanford have papers on trying to achieve calibration for richer subgroups which would approach this kind of stronger guarantee George is asking someone earlier touched upon the relationship between acceptance rates and equality of outcome can you clarify that relationship please okay so I will ask the person who asked that earlier to specify what they mean of by quality and and then we can I say I feel more confident asked and answering that once we're on the same page about what you mean they're young is asking do you feel it relevant who works do you feel it we belong ok let me start with a second sentence we belong to groups and some of us are more privileged than others how should we consider this ok I would love to also understand your first question
这和子群体问题有关。我想之前提问的 Gal 就在做这方面的工作,另外斯坦福的 Omer Reingold、Guy Rothblum、Michael Kim 也有论文,试图为更丰富的子群体实现校准,这就会逼近这种更强的保证。George 问:之前有人提到了接受率和结果平等之间的关系,能否请你澄清一下这个关系?好的,那我想请之前提问的那位说明一下他们所说的“平等”指的是什么,然后我们再来谈,我觉得等我们对你的意思达成一致之后,我回答起来会更有把握。接下来有人问:你觉得谁来做这些研究重要吗?我们所属的群体……好的,我先从第二句开始:我们都属于某些群体,我们中有些人比其他人更有特权,我们该如何看待这一点?好的。我也很想弄清楚你的第一个问题,我不太能读懂,也许你可以把第一个问题再说一遍。哦,我是不是
便签笔记
90:55
which I can't quite parse so maybe you can repeat the first question oh do I think it is relevant who as a matter of identity who works on these things yeah it is important I think it's one of these questions where my own perspective as a privileged white man is very limited and I've worked on formal models from the perspective from a very limited perspective right and I cannot tell you know the I cannot speak of lived experience of discrimination and and an oppression and that's why it is absolutely important to Center the voices that can tell you about the lived experience here right and that's that's you know at the beginning of the talk I gave you a number of scholars that do that and and I think that's where you should start I think that's the the count of the problem the account of the experience and that is absolutely essential so yes I you know I think I you know this is this tutorial should not be your only source on this topic it's it's it's one perspective on the formal and technical work in this area
认为“谁来做这些研究”从身份的角度看很重要?是的,这很重要。我觉得这是那类问题之一:我作为一个享有特权的白人男性,我自己的视角是非常有限的,我一直是从一个非常有限的视角在做形式化模型的工作,对吧。我没法讲述那种被歧视和被压迫的亲身经历,正因如此,把那些能讲述这种亲身经历的声音置于中心位置是绝对重要的,对吧。这也是为什么在演讲一开始,我给了你们一批做这方面工作的学者的名字。我认为那才是你应该开始的地方,我认为那是对问题的呈现、对经历的呈现,这是绝对不可或缺的。所以是的,我认为这个教程不应该是你在这个主题上唯一的信息来源,它只是关于这个领域形式化和技术性工作的一个视角,但你无论如何都不应该
便签笔记
91:59
but by all means you should not leave it at that and there's I think the the way this should work is that we all learn from a number of different sources and from a diverse set of sources some of them speaking more directly to the lived experiences and yeah I think this is an important question I just realized now that this was a private question I answered this in public I hope I didn't that was not against your your intent ok so these are all the questions I can see I'm also out of time so it was intended as public ok good summary um there was the earlier question by George and you mentioned so just was now among the panelists maybe he could yeah hello um I asked about the clarification for a cult Eve outcome I know it's really close in time so we can end this now if you want I can ask it later are you boy don't we give it a minute okay sure well I don't actually know what quality um sort of means I'm not like a political person or ethical or ethics was it just sounded important somehow
止步于此。我觉得应有的方式是:我们都从多种不同的来源、从多元化的来源去学习,其中一些来源更直接地讲述亲身经历。是的,我觉得这是一个重要的问题。我刚刚才意识到这是一个私下提的问题,而我公开回答了,希望我没有违背你的本意。好的,这就是我能看到的所有问题了,我的时间也用完了。哦,原来这个问题本来就是打算公开提的,好的,那就好。嗯,之前 George 提过一个问题,你也提到了,他刚刚加入了嘉宾席,也许他可以——是的,你好,嗯,我问的是关于结果平等(equality of outcome)的澄清。我知道时间真的很紧了,如果你想的话我们现在就可以结束,我可以之后再问。要不我们就再给它一分钟吧?好的,当然。其实我并不清楚这里的“平等”指的是什么。嗯,某种程度上意思是我不是那种搞政治的人,或者说伦理的——伦理这个词只是听起来比较重要,所以我才要求你澄清一下
便签笔记
93:13
this is why I asked for the clarification yeah yeah I think you can think of it as the as some equality of you know outcome if you think of the outcome as your decision and you're just comparing acceptance rate on that decision I in sort of if you serve I don't think this criterion captures what a moral philosopher would call it call it outcome and I think that's probably a longer conversation that we should have later on but okay cool do you have enough top you had to know then e-resources for which that might discuss those two things I don't think anybody has attempted to so this is a good question right so what this is something's my co-author Solan Broca's and I often run into there's no clear mapping from these criteria to moral philosophy and and back so it's not like you know for each criteria can say this is the piece of moral philosophy that it's tracking and I think that still we tempted some such at some point we had high hopes of doing such a synthesis and we didn't get very far so I don't I
澄清,对对,我觉得你可以把它理解成某种结果上的平等,如果你把结果理解为你的决策,然后你只是在比较该决策的通过率的话我觉得如果你去深究,这个标准并没有抓住道德哲学家所说的"结果"的含义,我觉得这可能需要更长的一次讨论,我们以后再聊,不过好的好,你有足够的资源可以推荐吗?就是那种会讨论这两者关系的资料我觉得还没有人尝试过做这件事,所以这是个好问题。这正是我和我的合作者 Solon Barocas经常遇到的问题——从这些标准到道德哲学之间没有清晰的对应关系,反过来也一样。也就是说,你没办法对每一条标准都说"这对应的是道德哲学里的这一块"。我觉得我们曾经也尝试过,某个阶段我们还挺有雄心想做这样一个综合梳理,但没走多远。所以我不知道,可能有很多人能给你更好的答案,但我一时
便签笔记
15结束语与 MLSS 花絮
94:29
don't know if so there's probably many people that could give you a better answer but I don't know from the top of my head of the right paper to look for that exact correspondence or than some such discussion okay well thank you very much that perspective that does help even identify that it is a big problem yeah so the correspondence between work and fairness in Michigan and philosophy I think is is very unclear to me I think there's no you know that's a good question if you have colleagues in the philosophy department that's a good question to chat with them about okay so [Music] thanks a lot more for a wonderful lecture I learned like you I think everyone learned also learned a lot from you today's just a bit of information so we have a router table discussion with more in half an hour so I guess we can extend some of the discussion to the router tables and we have before we go for a break we have short videos right from the organizer so let me show the videos from my point of view and machinery is all based on
想不出有哪篇论文正好讨论这种对应关系,或者类似的讨论。好的,非常感谢,这个视角确实有帮助,至少让我意识到这是个大问题对,所以机器学习中的公平性与哲学之间的对应关系,我觉得对我来说是很不清晰的。我觉得这确实是个好问题,如果你在哲学系有同事,这是个很值得跟他们聊的问题。好的,那么[音乐]非常感谢你带来这么精彩的讲座,我学到了很多,我想大家今天也都从你这儿学到了很多。说几句事务性的信息:半小时后我们还有一场圆桌讨论,Moritz 也会参加,所以有些讨论我们可以延续到圆桌上。在休息之前,我们还有主办方提供的几段短视频,那我来放一下这些视频在我看来,整套机制都建立在统计依赖关系之上,我们观察到某些量与其他
便签笔记
96:09
statistical dependencies so we see that certain quantities are related to other quantities and then we predict one thing from the other one now usually machine and it doesn't asked where these dependencies come from and I think they come from underlying causal relationships and if we're interested in how to do machine learning when certain things in the world change so when we are trained have the training data in one problem but we are being tested on another problem which is happening all the time in animals and humans then turn into that we have to look at things in every little bit more detail I think this is too complicated now oh so this is this is very hard okay I tend to be I'm not really systematic in my research in the sense that I don't wake up on January 1st that we're gonna join that what I want to do next year and then I just just carry out my agenda so I think a big source of inspiration for me is talking to people that you know work in industries Coco with are the practical problems that I'm not even
量相关,然后我们用一个来预测另一个。而通常机器学习并不追问这些依赖关系从何而来。我认为它们来自底层的因果关系。如果我们关心的是:当世界上某些东西发生变化时该怎么做机器学习——也就是我们在一个问题上有训练数据,却在另一个问题上被测试,而这在动物和人类身上时时刻刻都在发生——那我们就必须把事情看得更细致一些。我觉得这个太复杂了,哦,这个,这个很难。好的,我这个人其实不太系统在做研究这件事上并不系统,意思是我不会在1月1日醒来就决定明年要做什么然后照着自己的日程表执行。我觉得对我来说一个很大的灵感来源是跟人聊天,那些在业界工作的人,聊他们遇到的、我完全没意识到的实际问题,这比读文献
便签笔记
97:23
aware of more than reading it's much more efficient talking to people and you know understanding what other important things I tend to work on things I like I'm sort of an arrest and I believe that as opposed to working on stuff that everybody works on because it's hot so I am dreaming a lot driven a lot by the pleasure in doing research and something that I really find interesting irrespective to everything else I didn't want to use the word passion because being Italian there's a lot of passion it's a little bit cliche but it is true I mean I think people should be driven by curiosity and passion because that's a very good driving force a very healthy crinum force the best experience is probably the close atmosphere between lecturers and participants and for me personally it's a first opportunity to have attended emesis five years ago and now a dear speaker but the funniest is actually the t-shirt so when you are outside the community you will often get asked why you attended for achieving and
更有效率——跟人聊天,去了解还有哪些重要的事情。我倾向于做我喜欢的东西我算是有点特立独行吧,我相信这一点,而不是因为某个方向很热大家都在做就去做。所以我在做研究这件事上,很大程度上是被乐趣驱动的,被那些我真正觉得有意思的东西驱动,跟其他一切无关。我本来不想用"激情"这个词,因为我是意大利人,激情太多了,有点老套但这是真的。我是说,我觉得人应该被好奇心和热情驱动,因为那是非常好的驱动力,非常健康的驱动力。最好的体验大概是讲者和参与者之间那种亲近的氛围。对我个人而言,五年前我第一次有机会参加 MLSS,现在成了讲者。不过最有意思的其实是那件 T 恤,当你在这个圈子之外时,人们经常会
便签笔记
98:46
is breeds pretty close to me okay thank you very much everyone so that would be the end of this session and so you see you again in 20 minute also for the router table discussion with more it heart okay thank you very much and bye bye
问你为什么去参加什么什么,这对我来说挺亲切的。好的,非常感谢大家,这场就到这里,20分钟后再见,还有和 Moritz 的圆桌讨论。好的,非常感谢,再见
便签笔记
99:32
you
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

Moritz Hardt 系统梳理了机器学习分类中的三大统计公平性准则(独立性/等接受率、错误率均等、按组校准),证明它们在一般条件下互斥,并通过 COMPAS 再犯预测与"未出庭"案例论证:这些准则只能触发道德直觉、充当审计工具,既不是公平性的定义或证明,更不应作为优化目标;更根本的问题是——是否该用预测来解决这个问题。

核心要点

  • "无知即公平"(fairness through unawareness) 彻底失败:删除敏感属性无法消除歧视。典型案例是 2016 年彭博报道的 Amazon 当日达覆盖图——波士顿唯一被排除的社区是黑人聚居的 Roxbury,与历史上的"红线"(redlining) 拒贷地图惊人相似。Amazon 几乎肯定没看种族,只是预测购买量,但购买量与社会经济地位相关,后者在美国又与种族相关。"我们的数据里没有这个变量"不是辩护理由。
  • 歧视的概念是领域特定且群体特定的:关注的是"无正当理由的区分",包括实际无关(如性取向与雇佣)和道德无关(如残疾、怀孕可能性——即使统计上影响工时,社会仍决定雇主必须承担该成本)。它只适用于信贷、教育、就业、住房等涉及重大机会的领域,且延伸到相关广告投放;受保护类别(种族、性别、宗教等)是几十年民权运动斗争的产物,且仍在变动——讲座前几周美国最高法院刚裁定性别保护延伸至性别认同与性取向。
  • 法律不提供可形式化的公平定义:美国有"差别对待"(disparate treatment,关注程序公平、意图) 和"差别影响"(disparate impact,关注结果分布) 两种学说,且二者有张力——避免差别影响可能需要显式考虑群体成员身份,这又可能构成差别对待。反歧视法是零散的政治妥协,不是某个统一道德理论的操作化,计算机科学家不能指望从法律里直接抄出定义。
  • 三大准则可统一表达为(条件)独立性陈述:引入群体变量 A、得分 R、决策 D、结果 Y。(1) 独立性 R ⊥ A → 等接受率;(2) 分离性 R ⊥ A | Y → 等假阳性率和假阴性率(错误率均等);(3) 充分性 Y ⊥ A | R → 按组校准。这种写法天然推广到多分类和回归任务。几乎所有文献中的公平准则在概念上都能归入这三类,用深度学习做"公平表示学习"也不会让底层准则更公平——本质仍是独立性。
  • 每个准则都能被"钻空子":等接受率允许在多数群体做出精准决策、在少数群体随机抛硬币(Framingham 心脏病风险评分建于白人男性队列,对其他人群效果差就是自然发生的版本);错误率均等是事后(post hoc)准则,决策时不可知,且强制均等可能故意压低某个群体本可达到的表现(ROC 曲线交集之下);校准允许把一组人的分数取平均替换(如把 0.1~0.7 的人全改成 0.4),仍满足校准却让所有人都低于拘留阈值。
  • 不可能性定理:只要两组基准率不同且分类器不完美(至少有一个假阳性和一个假阴性),按组校准就必然导致错误率均等失败——而这是现实的一般情况。更关键的是,无约束的标准监督学习本身就近似给出按组校准(偏差上界由与贝叶斯最优得分的差距决定;UCI Adult 数据上直接跑逻辑回归,男女校准曲线都接近对角线),所以默认情况下错误率均等就已被违反。
  • COMPAS 之争是围绕稻草人的争论:ProPublica 指出未再犯的黑人被告被标高风险的比例远高于白人(假阳性率不均),Northpointe 回应说分数按组校准、黑人再犯基准率更高故不可避免。两者都对,但两个准则都不排除不公平做法:警方可以通过对某群体更激进地逮捕低风险者(例中拘留率 61%→42%、假阳性率→26%),"漂亮地"满足错误率均等,同时造成更大伤害。把准则当激励目标会诱发这类操纵。
  • 更根本的问题是"是否应该用预测":以预测"未出庭"为例,被告不出庭的主因是缺乏托儿、交通、雇主不放行、出庭次数过多(有黑人被告在洗清指控前被要求出庭 22 次)。审前羁押数月会让人失业、断绝社会联系。德州 Harris County 诉讼和解方案要求提供免费法院托儿、双向通讯系统等结构性干预——这些都与预测无关。统计准则把数据生成分布视为既定,只能在观测联合分布内工作,无法设想对世界的干预,因而排除了整个干预空间。

结论与值得注意的细节

  • 主讲人的总体立场:这些准则不是公平性定义,也不能作为公平性证明或目标函数;在他所述的统计框架内,他确信不存在令人满意的公平性定义。但它们仍有价值——作为"测量仪器"触发道德直觉(ProPublica 的愤怒就是一种真实的道德直觉),帮助人们表面化关于决策应如何做出的规范性问题。
  • 公平准则与道德哲学之间没有清晰对应关系:他与合著者 Solon Barocas 曾试图做这种综合但没能走远。
  • "公平选区划分"(fairness gerrymandering,Kearns 等人):在两个群体之间强制某准则,机器学习会以最懒惰的方式满足它,往往在子群体(如群体内按性别)造成更严重的违反。
  • 校准不是个体层面的保证:得分 0.8 只表示得分为 0.8 的人群平均阳性率为 80%,不是某个人的概率。贝叶斯最优预测也解决不了道德无关问题——预测未来工时时,残疾或可能怀孕的人会被"正确地"打低分。
  • 错误率在真实场景中常常不可观测(被哈佛拒绝的人永远没有哈佛 GPA);信贷领域用"被 A 银行拒绝但在 B 银行还清贷款"来近似估计。
  • 主讲人明确承认自身作为特权白人男性视角的局限,强调此教程只是形式化工作的一个侧面,应优先阅读 Ruha Benjamin("新吉姆密码")、Meredith Broussard、Virginia Eubanks、Safiya Noble、Cathy O'Neil 等学者关于亲身经历与技术危害的著作。
  • 本讲仅覆盖"狭义"统计视角;第二部分(周四)将转向因果模型和社会技术系统的动态模型,探讨如何把社会事实与干预纳入模型。
核心句型 · 9
1. It helps to acknowledge that …
“I think it helps to acknowledge that we live in a world of pervasive inequality”
用「承认……是有帮助的」引入一个沉重但必要的前提,语气克制、不说教。适合在报告开头交代背景假设。
2. This is not in any way meant to … / not meant to be …
“That is not in any way meant to dissenter or distract from the scholarship that I just mentioned”
预防性澄清:先否认一种可能的误读,再说明真实意图。学术演讲中界定立场的常用结构。
3. Isn't … the whole point of …? So wouldn't it be a little bit rich to …?
“Isn't discrimination the whole point of machine learning … wouldn't it be a little bit rich to accuse machine learning of …”
用两个反问先替对方把反驳说足,再逐条回应。呈现对立观点时比直接否定更有说服力。
4. X does not get better by using Y
“The underlying fairness criterion doesn't get better by using deep learning”
简洁地拆穿「手段复杂≠结果更好」的错觉。可套用:The argument doesn't get stronger by adding more data.
5. Let me add a word of caution to …
“Let me add a first word of caution to the contest debate and another word of caution to it”
引出保留意见或告诫的礼貌说法,比 but 更正式。可用 a word of caution / a note of caution。
6. There are really bad ways of achieving X
“There are really bad ways of achieving error rate parity and really bad ways of achieving independence”
点出「指标达标不等于问题解决」。适合批评 KPI 式思维:满足条件的方式本身可能有害。
7. If you use them as A, that's one thing; but if you ask people to B, that's another thing
“If you use them as a gauge … that's one thing but if you ask people to optimize for them that's another thing”
「一回事……另一回事」的对比结构,用于区分同一工具的两种用法(度量 vs 目标)。
8. I find it hard to make a case for X in light of Y
“I find it hard as a machine-learning person to continue to insist on this prediction thing in light of these alternatives”
表达「考虑到 Y,很难为 X 辩护」,含蓄地放弃立场。make a case for 与 in light of 是学术写作高频搭配。
9. This is a good question, and I'm gonna return to this in a moment
“So this is a good question and I'm gonna return to this in a moment let me first wrap up this example”
演讲中处理打断的模板:肯定问题、承诺回应、先收尾当前内容。wrap up 表示「收尾」。
词汇精讲 · 118 · 按出现顺序
without further ado phr. 0:02
闲话少说,废话不多说(介绍讲者时的固定用语)
consequential /ˌkɑːnsəˈkwenʃəl/ adj. 2:14
有重大后果的,影响深远的
pervasive /pərˈveɪsɪv/ adj. 2:14
无处不在的,普遍存在的
testament /ˈtestəmənt/ n. 2:14
证明,明证(be testament to 是……的证据)
perpetuating /pərˈpetʃueɪtɪŋ/ v. 2:14
使延续,使永久化(perpetuate,多含贬义)
revisit /ˌriːˈvɪzɪt/ v. 3:34
重新审视,重新考虑
scholarship /ˈskɑːlərʃɪp/ n. 3:34
学术研究成果(此处非「奖学金」)
inequities /ɪnˈekwətiz/ n. 4:40
不公正,不平等(inequity,强调公正性而非数量差异)
blatant /ˈbleɪtənt/ adj. 5:44
公然的,明目张胆的
right off the bat phr. 5:44
一开始就,立刻
caveat /ˈkæviæt/ n. 6:55
警告,附加说明,但书
backdrop /ˈbækdrɑːp/ n. 6:55
背景(法律/历史背景)
fragmented /ˈfræɡmentɪd/ adj. 6:55
碎片化的,零散的
a little bit rich phr. 8:44
(口语)有点过分、说不过去(指责缺乏立场)
unjustified /ʌnˈdʒʌstɪfaɪd/ adj. 8:44
无正当理由的
accommodating /əˈkɑːmədeɪtɪŋ/ v. 9:52
为……提供便利,照顾(残障)需求
incur /ɪnˈkɜːr/ v. 9:52
招致,产生(费用)
absorb /əbˈzɔːrb/ v. 9:52
承担,消化(成本)
salient /ˈseɪliənt/ adj. 11:04
显著的,突出的(socially salient 社会显著的)
adverse /ˈædvɜːrs/ adj. 11:04
不利的,有害的
public accommodation n. 12:11
公共设施(美国法律术语,指对公众开放的商业场所)
statute /ˈstætʃuːt/ n. 13:24
成文法,法规
in flux phr. 13:24
处于变动之中
contentious /kənˈtenʃəs/ adj. 14:43
有争议的
doctrines /ˈdɑːktrɪnz/ n. 14:43
(法律)原则,学说
disparate treatment n. 14:43
差别对待(美国反歧视法术语,指故意歧视)
disparate impact n. 15:51
差别影响(中立程序造成群体间结果差异)
distributive justice n. 15:51
分配正义
down the road phr. 15:51
日后,在后续环节
working grasp n. 17:03
可用的基本理解
operationalize /ˌɑːpəˈreɪʃənəlaɪz/ v. 18:11
操作化,使可实际执行
scrub /skrʌb/ v. 19:24
清除,擦掉(数据)
predominantly /prɪˈdɑːmɪnəntli/ adv. 19:24
主要地,占多数地
redlining /ˈredlaɪnɪŋ/ n. 20:33
红线歧视(拒绝向特定社区放贷的做法)
hazardous /ˈhæzərdəs/ adj. 20:33
危险的,高风险的
socioeconomic status n. 20:33
社会经济地位
strike people as phr. 21:36
给人以……印象,让人觉得
entertain /ˌentərˈteɪn/ v. 22:48
考虑,抱有(想法、假设)
avenues /ˈævənuːz/ n. 22:48
途径,路线(研究方向)
taste based discrimination n. 25:31
偏好型歧视(经济学术语,因个人偏见而歧视)
tipping point n. 26:45
转折点,临界点
normative /ˈnɔːrmətɪv/ adj. 26:45
规范性的(关于「应该如何」)
where rubber meets the road phr. 27:58
真正接受检验之处,实际落地的关键环节
grapple with phr. 27:58
努力应对,设法解决
covariates /koʊˈveriəts/ n. 27:58
协变量(统计学术语)
likelihood ratio test n. 30:11
似然比检验
nonparametric /ˌnɑːnpærəˈmetrɪk/ adj. 30:11
非参数的
regularization /ˌreɡjələrəˈzeɪʃən/ n. 31:15
正则化
confusion table n. 31:15
混淆矩阵
row wise adj. 33:22
按行的
stress testing n. 36:35
压力测试(检验标准在极端情况下是否失效)
arbitrary /ˈɑːrbətreri/ adj. 37:39
任意的,武断的
cohort /ˈkoʊhɔːrt/ n. 38:40
队列(流行病学研究中的一组人)
track record n. 38:40
历史记录,过往表现
adversarial learning n. 39:43
对抗学习
under the hood phr. 40:52
在底层,在内部机制上
trading off phr. 42:04
权衡,以……换……
post hoc /ˌpoʊst ˈhɑːk/ adj. 43:19
事后的(拉丁语)
in hindsight phr. 44:22
事后看来
audit /ˈɔːdɪt/ n. 44:22
审计,审查
disproportional /ˌdɪsprəˈpɔːrʃənəl/ adj. 45:28
不成比例的
counterfactual /ˌkaʊntərˈfæktʃuəl/ adj./n. 46:37
反事实的;反事实情形
potential outcomes n. 47:48
潜在结果(因果推断框架术语)
false discovery rate n. 50:09
假发现率
gained more traction phr. 50:09
获得更多关注/认可
calibration /ˌkælɪˈbreɪʃən/ n. 50:09
校准(预测概率与实际频率一致)
soliciting /səˈlɪsɪtɪŋ/ v. 52:28
征询,索取(信息)
mental gymnastics n. 52:28
心算体操,复杂的脑力周折
a priori /ˌeɪ praɪˈɔːraɪ/ adj. 53:35
先验的,事前的
get carried away phr. 53:35
过于兴奋,失去分寸
upper bounded adj. 54:44
有上界的
out-of-the-box adj. 55:46
开箱即用的,不加调整的
deciles /ˈdesaɪlz/ n. 55:46
十分位数
deem /diːm/ v. 56:53
认为,视为
striking /ˈstraɪkɪŋ/ adj. 58:03
显著的,触目的
gerrymandering /ˈdʒerimændərɪŋ/ n. 59:07
选区划分操纵;此处引申为通过划分群体来操纵公平指标
laziest possible way phr. 59:07
最偷懒的方式
mutually exclusive adj. 60:26
互相排斥的
degenerate /dɪˈdʒenərət/ adj. 61:28
退化的(数学上的特殊平凡情形)
base rates n. 61:28
基础率(某群体中正例的比例)
recidivism /rɪˈsɪdɪvɪzəm/ n. 62:45
再犯,累犯
jurisdictions /ˌdʒʊrɪsˈdɪkʃənz/ n. 63:52
司法管辖区
defendant /dɪˈfendənt/ n. 63:52
被告
detain /dɪˈteɪn/ v. 63:52
羁押,拘留
boiled down to phr. 66:10
归结为,简化为
straw man n. 66:10
稻草人(歪曲对方立场以便攻击的论证谬误)
stick figures n. 67:25
火柴人简笔画
padding /ˈpædɪŋ/ v. 69:48
注水,虚增(数量)
incentivize /ɪnˈsentɪvaɪz/ v. 71:02
激励,以利益诱导
sneaky /ˈsniːki/ adj. 71:02
偷偷摸摸的,暗箱的
give this group a break phr. 72:07
对这个群体网开一面
feedback loops n. 74:12
反馈回路
failure to appear n. 74:12
未出庭(法律术语)
devastating /ˈdevəsteɪtɪŋ/ adj. 75:15
毁灭性的
disruptive /dɪsˈrʌptɪv/ adj. 75:15
破坏性的,造成严重扰乱的
compelling /kəmˈpelɪŋ/ adj. 76:17
有说服力的,令人信服的
cleared of the charge phr. 76:17
被洗清指控
vouchers /ˈvaʊtʃərz/ n. 77:27
代金券,补贴券
settlement /ˈsetlmənt/ n. 77:27
(诉讼)和解协议
mitigate /ˈmɪtɪɡeɪt/ v. 77:27
缓解,减轻
make a case for phr. 78:29
为……辩护,论证……的合理性
in light of phr. 78:29
鉴于,考虑到
confined to phr. 78:29
局限于
reactive /riˈæktɪv/ adj. 79:34
被动应对的(与 proactive 相对)
underrated /ˌʌndərˈreɪtɪd/ adj. 80:44
被低估的
outraged /ˈaʊtreɪdʒd/ adj. 80:44
愤慨的
get a better handle on phr. 80:44
更好地把握、理解
settle this once and for all phr. 81:42
一劳永逸地解决(原文口误为 once and fall)
gauge /ɡeɪdʒ/ n. 81:42
量表,测量工具
sweeping /ˈswiːpɪŋ/ adj. 84:05
一概而论的,全盘的(sweeping charge 全盘指责)
forego /fɔːrˈɡoʊ/ v. 84:05
放弃(= forgo)
content moderation n. 86:13
内容审核
resort to phr. 86:13
诉诸,不得已采用
ambiguous /æmˈbɪɡjuəs/ adj. 87:27
含混的,模棱两可的
lived experience n. 90:55
亲身经历(社会科学术语,强调第一人称经验)
by all means phr. 91:59
务必,无论如何
synthesis /ˈsɪnθəsɪs/ n. 93:13
综合,融合梳理
from the top of my head phr. 94:29
一时想起来的,不假思索地
理解自测 · 11 题
1. 讲者开场为本次教程设了哪三个范围限定?

三个限定分别是:一、只讨论「后果性决策」场景中的歧视,即招聘、录取、信贷、刑事司法这类需要做接受/拒绝二元决策、且对个人有实际后果的领域,排除了世界上其他形式的不公;二、以美国的法律与案例为背景,讲者承认这对来自欧洲等地的听众未必适用,但法律状况本就碎片化,无法全球统一讨论;三、只讲形式化模型与框架,因为这是他的专长,但他强调这不是要贬低 Benjamin、Eubanks、Noble 等批判性学者的工作,也明确为「非技术干预」留出空间。这些限定出现在第 2–7 段的开场部分。

2. 亚马逊当日送达的例子说明了什么,讲者对它的机制作了怎样的推测?

这个例子说明「无意识即公平」(fairness through unawareness)是失败的。2016 年彭博社报道,亚马逊在波士顿推出当日送达时,唯独排除了以黑人为主的罗克斯伯里街区,地图与 1930 年代的「红线」歧视地图高度相似。讲者推测亚马逊几乎肯定没有使用种族变量,而只是做了标准的「订单量预测」;但订单量与社会经济地位相关,后者在美国又与种族相关,于是代理变量重建了被删除的信息。结论是「我们的数据里没有这个变量」从不构成免责理由。该例在第 17–19 段。

3. 讲者归纳的三大统计公平标准分别是什么?用条件独立如何表述?

三大标准是:一、独立性(independence),R ⊥ A,即打分或决策与受保护属性无关,推出各组接受率相等;二、分离性或错误率均等(error rate parity / equalized odds),R ⊥ A | Y,即在给定真实结果下分数与群体无关,推出各组假阳性率、假阴性率相等;三、充分性或分组校准(calibration by group),Y ⊥ A | R,即知道分数后群体身份对预测结果不再提供额外信息。讲者在第 53 段说,几乎文献中所有变体在概念上都能归入这三者之一,用条件独立表述还便于推广到多分类和回归。

4. 什么是「公平性选区划分」(fairness gerrymandering)?它揭示了群体公平标准的什么问题?

这是宾大 Kearns 团队提出的说法,指在两个群体之间满足公平标准,但在子群体内部严重违反。讲者的例子:蓝组和绿组接受率相同,但蓝组只录取男性、绿组只录取女性,按性别×群体的交叉子群体看,独立性被彻底破坏(第 51–52 段)。它揭示的问题是:机器学习会以「最偷懒的方式」满足任何强加的约束,因此拉平某个总体统计量后,不应期待子群体内会自动合理,反而应预期出现令人不安的现象。这推动了 Kearns 等和 Hébert-Johnson 等关于子群体公平与多重校准的研究。

5. 为什么讲者说错误率均等是「事后」标准,而校准是「事前」保证?这一差异有什么实际含义?

错误率均等以真实结果 Y 为条件,而决策时决策者并不知道谁最终是正例、谁是负例,所以只能在事后收集各组中结果已知的人,回看他们当初如何被分类来审计(第 39 段)。校准则以分数 R 为条件,决策者看到分数 0.8 时就已知道该分数对应约 80% 的正例频率,无需再询问群体身份(第 46–47 段)。实际含义有两层:一是错误率均等更适合作为事后审计工具,校准更适合作为决策时的操作依据;二是在被拒者结果不可观测的场景(如录取、信贷)中,错误率均等的估计本身就很困难,需要借助潜在结果框架做近似(第 41–42 段)。

6. 不可能性定理的前提是什么?讲者为什么说这些前提在现实中「几乎总是成立」?

定理陈述:若两组基础率不同(正例比例不等),且分类器不完美(至少有一个假阳性和一个假阴性),则分组校准成立必然推出错误率均等不成立(第 54 段)。讲者说这两个前提是「一般情况」,理由是:我们生活在一个不平等的世界里,各群体在结果上往往确实存在差异;而现实中也几乎没有完美预测器,任何模型都会犯错。再结合第 48 段的定理——无约束学习越接近贝叶斯最优就越接近校准——就能推出:只要你「正常地」做机器学习,就默认得到校准,也就默认违反错误率均等。COMPAS 之争在数学上正是这一权衡的具体化。

7. 讲者用两组火柴人的构造例子想证明什么?「糟糕的修法」具体是怎样的?

例子想证明:错误率均等和校准都不是「公平性证书」,把它们当作激励目标会诱导出更有害的达标方式。设定是蓝组、橙组,阈值 0.5 羁押,且分数就是真实再犯概率;初始时橙组羁押率 61%、蓝组 38%,假阳性率也悬殊。「糟糕的修法」是对橙组更激进执法,多逮捕再犯风险很低的人然后释放,从而给分母「注水」,让羁押率降到 42%、假阳性率降到 26%,指标看似接近了,但实际伤害更大(第 59–62 段)。对校准的类似操作是把一组分数替换为它们的均值,仍满足校准却让所有人落到阈值之下(第 63 段)。两例均来自 Corbett-Davies 等 2017 年论文。

8. 在「未出庭」问题上,讲者为什么认为不该把它表述为预测问题?他给出的替代方案是什么?

讲者的主要论点不是数据有偏或目标函数选错,而是「一开始就把它当作预测问题」这一框架本身有问题(第 65 段)。理由:审前羁押数月对个人是毁灭性的,会失业、断绝社会关系;而人们不出庭的真实原因多是缺托儿、缺交通、雇主不放人、出庭次数过多——会议上有被告在撤诉前需出庭 22 次(第 66–67 段)。替代方案是结构性干预:提供托儿券、交通补贴、强制雇主放人、减少出庭次数;哈里斯县诉讼和解已要求法院提供免费托儿和双向沟通系统,这些没有一项涉及预测(第 68 段)。讲者说在这些更直接的措施面前,他很难为使用预测辩护。

9. 提问者 Gal 提出「好的条件期望估计可能恰恰帮助发现正确的干预」,讲者如何回应?你认为这一回应充分吗?

讲者承认这是很好的观点,并明确否认自己在全盘反对使用统计——人们发现托儿是关键因素,很可能就用了统计方法(第 73–74 段)。但他区分了两种「问题框架」:一种是对「未出庭的原因是什么」做因果推断,条件期望可以为此提供线索;另一种是把风险分数直接装进系统用来给审前羁押排序。分歧不在是否用统计,而在做统计的目的和问法,必须先就框架达成一致。这一回应大体充分,因为它把争论从「预测 vs 干预」转移到「预测服务于什么目标」;但也留下一个未完全解决的问题——现实中同一模型往往既用于理解也用于排序,两种框架未必能干净地分开。

10. 讲者说这些公平标准「确定不是公平定义」,却又说它们「有意思」。这两个说法如何调和?

调和点在于标准的功能定位:它们不是证明公平的证书,也不能当目标函数,但作为「测量当下世界的量表」,它们能经验性地激发道德直觉,浮现出关于决策方式的规范性问题(第 70–72 段)。最好的例证是 ProPublica 对 COMPAS 假阳性率差异的愤慨——那是一种真实的道德直觉,捕捉到了人们对「决策应当如何做出」的理解,很难简单反驳。因此标准的价值在于帮助我们讨论并厘清「究竟是什么让我们对某种决策方式感到不安」,引出不同的取舍与张力,而不是一劳永逸地裁定什么是公平。讲者同时断言:在这套统计设定内,不存在令人满意的公平定义。

11. 把讲者「统计标准只使用观测联合分布、无法考虑干预」的批评迁移到内容审核场景,他的立场会是什么?这一立场在必须做预测的场景下还成立吗?

讲者在回答 Indira 时已经给出了立场(第 76–78 段):即便接受「缺乏人力、必须用算法审核」的前提,毒性、情感分类对非裔美国人英语的高标记率也会违反错误率均等等标准,这里被激发的道德直觉是正当的;但标准本身只能「标记问题」,不能告诉你接下来怎么办。他进一步质疑前提:毒性分类的目标变量本身高度含混,是「一开始就有问题的预测任务」。迁移到「统计视角只处理观测分布」的批评,他会主张:与其在既定数据分布内优化某个公平指标,不如追问标签是怎么产生的、审核规则能否改变、能否引入人工复核等干预手段。在预测确实不可避免的场景下,这一立场仍成立,但它给出的更多是「重新框定问题」的方向,而非可直接执行的算法方案——这也正是讲者坦承的领域现状。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.118Amartya Sen: The Idea of Justice 下一期 · NO.120 →The Ethical Algorithm | Michael Kearns & Aaron Roth | Talks at Google
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com