视频库 / NO.088ASK THE BEST MINDS THE BIG QUESTIONS
视频库 / NO.088
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

The Information: A History, a Theory, a Flood | James Gleick | Talks at Google

节目发布 2011-03-24
本期追问 · 点击跳到视频对应位置
8:34 媒介如何改变我们所知的世界?39:00 听来的知识可靠吗?27:19 知道得越多,就离真相越近吗?20:50 我们造的东西,何时开始管我们?
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
编者按:2011年春,作家詹姆斯·格雷克(James Gleick)携新作《信息简史》(The Information: A History, a Theory, a Flood)做客 Google 总部的「Talks at Google」系列。格雷克曾任《纽约时报》科学记者,著有《混沌》与费曼传记《天才》,是当代最重要的科学与技术史作者之一。这次活动没有采用讲座形式,而是由作者抛出问题,与在场的 Google 工程师围绕「搜索还是过滤」「权威与众包」「回音室与意外之喜」「付费墙与信息洪流」等话题展开了一场即兴对话。本文依据现场录音编译整理,仅删去口语枝节与重复,问答内容均以讲者第一人称连贯成文。

从《信息简史》说起

谢谢大家来听我讲。我以前来过这里,那是因为我参与了关于 Google 大规模复制版权图书那桩诉讼的一些谈判。那次经历挺有意思,这次能回来我很高兴,也有点意外,居然有这么多人愿意买我这本古雅装帧的纸质书。

Google 在我的书里只在结尾处短暂露面。动笔时,我本来设想会在这里做大量采访,让 Google 成为全书结尾的一大主角。结果没有这么写。书里关于维基百科的内容反而比关于 Google 的多得多。某种意义上,我觉得这是因为写 Google 已经没有必要了:你们无处不在,早就是人们生活的一部分,用不着我再来介绍。

这本书的前提,用我平时不太用的说法来讲就是:通往此处的那条路,通往 Google 的那条路,起点在二十世纪中叶的一个确切时刻。那是 1948 年,克劳德·香农(Claude Shannon)在贝尔实验室发表了一对技术论文,刊登在《贝尔系统技术期刊》上,几乎一出手就是成熟形态,很快就被人称为「信息论」(information theory)。后来它以《通信的数学理论》(The Mathematical Theory of Communication)为名出版成书。你们有谁手头有这本书吗?只有寥寥几位。我原本以为在这样一个地方,它应该算是一种案头必备。我认为不管你在公司做哪一块业务,这本书里都有大量与你相关的东西。

我第一次听说它,是在写第一本书时从研究混沌的科学家那里听来的。虽然过去很多年了,我还记得当时的诧异:居然有「信息论」这种东西,一门关于信息的数学理论。在我看来,信息这个模糊而无定形的东西,本来不像是能用数学去分析的对象。而如今不必我多说,我们生活的这个世界里,信息早已不只是指令、新闻、流言,甚至不只是文字。它是一个庞大的范畴,包括音乐、图像,以及一切能被数字化、存进电脑、再用 Google 搜到的东西。

我相信这些事物之间存在着一种联系。香农的理论,以及此后数学家和紧随其后的计算机科学家做的全部工作,构成了我们这个世界的底层结构。它不只是一种技术基础,还在一种更真实、更宽泛的意义上支撑着一切。我在尽量避免说「在哲学意义上」。

有了这个前提,书的结构也就随之而来:它要回到人类历史的开端,因为信息时代并不只是我们谈了五十年的那件事。顺带一提,确实是五十年,据《牛津英语词典》记载,「信息时代」这个说法最早是 1960 年开始流传的。可我们现在明白了,整部人类史都是信息时代。印刷机是一种信息技术,电报、电话也是,而在有文字记载的历史开端,字母表就是。这就是我这本书的基本想法。组织这个故事有点难:它从开头讲起,也就是说,它从中间,从 1948 年讲起,然后回溯,再往前推进。

我不想在这里做一场讲座,更想来一次对话。说实话,我总觉得应该是我在采访你们才对。所以让我先提一个问题。

读者买的是被删掉的新闻

我职业生涯是在《纽约时报》起步的,那是一家伟大的信息企业,至今仍是。时报有位睿智的老前辈常说:读者花钱买的不是我们登出来的全部新闻(「所有适合刊印的新闻」),读者花钱买的是我们删掉的那些新闻。

如果你是上一代的典型《纽约时报》读者,时报很可能是你唯一的信息来源。我记得祖父在地铁上读它,要用一种特定的方法折叠才能在地铁上读,不然就会缠成一大团纸。曾经有段时间,这种折叠时报的方法是学校里教的。但它是有限的,众所周知,它由纸做成,就那么大。时报编辑的自负在于:只要你读完整份报纸,你就掌握了全部新闻。

不仅如此,头版是高度结构化的,你们可能知道,现在依然如此。每天都有大量正式的讨论,决定哪篇报道应该领衔(放在右上角),哪篇是第二重要的(放在左上角,作为副头条),哪些报道配得上折线以上的标题,哪些放在折线以下,标题该占两栏还是三栏还是更多。这里运用的判断力被称为「新闻判断」(news judgment),时报的编辑们以拥有这种品质而自豪。他们知道,比如说,一个人被谋杀并不具有固定的重要性。有些谋杀案完全不重要,根本不值一提,事实上几乎所有谋杀案都是如此。但假如死者恰好是斯卡斯代尔一位医生,还是一本著名减肥书的作者,而被告恰好是被他抛弃的情人,一位斯卡斯代尔的中年贵妇,那这桩谋杀案就非常重要。时报从不为做出这样的判断而道歉。

以上都是为了引出一个问题。在我看来,Google 作为一个赫赫有名的搜索引擎,和我描述的这个多少有点神话色彩的《纽约时报》一样,本质上同样建立在过滤之上。你们提供的服务,不只是从互联网上抓取信息呈现给读者,而是把用户不想找的信息滤掉,正如时报当年过滤掉那些新闻一样。今天的互联网用户经常抱怨自己被信息轰炸。一方面,人人都为不必再依赖那些陈旧的、权威的、反民主的来源而欢欣鼓舞。

昨晚我在伯克利做了一场演讲,内容和今晚不同,但讲到某处我用了「民主」这个词,在互联网上的民主事物和权威事物之间设置了一种对立,后排一位女士立刻喊了一声:「Democracy Now!」这大概就是伯克利吧。

我刚才说到哪儿了。我们这些互联网上的新闻消费者是矛盾的,而且公开表达这种矛盾。一方面,人们很高兴摆脱了《纽约时报》这类权威的专制,摆脱了那些资产阶级的、亲商的、反天主教的,随便你想给时报编辑安上什么偏见;人们从数百万博客和其他网站、其他来源那里获得了力量,这些来源算不算「新闻」,取决于我们想不想做势利眼。可同样是这些人,转头就抱怨:我怎么知道什么是真的?什么是假的?我怎么找到对我有意义的东西?

我知道 Google 有一个庞大的,我不确定该不该叫「新闻业务」,总之有 Google 新闻。我的理解是,它要么完全自动化、算法化、众包化,要么至少大部分如此,没有多少人类编辑的干预。我认为眼下正在进行一场伟大的实验:智能的众包(如果你们这里是这么称呼的),或者说由算法来决定什么对用户真正重要的做法,能否胜过老派的人工判断。老派的人工判断包括《纽约时报》,也包括图书出版商。人们曾经根据出版社的名字买书。我的出版社是 Pantheon,在座有些年纪够大的人会把这个名字和某种如今差不多已经过时的写作风格联系在一起。我想现在没多少人再根据出版社愿意选择出版什么来买书了。一份诗歌刊物的编辑是一个权威来源,是反民主的。但每一位博主当然也以自己的方式反民主:他或她在做个人判断,你可以选择信任或不信任。

这就是我的问题。我很想听听在座的各位,无论赞同还是反对:Google 的首要功能是过滤,而不是搜索?

权威与策展:众包能胜过编辑吗

现场有人指出,我提到的其实是两个在 Google 内部有专门名称的话题,一个是「权威」(authority),另一个是「策展」(curation),也就是编辑,那个决定什么值得读、什么不值得读的选择者。策展人挑选什么应当保留、什么有价值、什么没价值,而且往往是你依赖的单一或少数几个来源。在人人都能当出版者、至少能当推荐者的新时代,我们如何判定谁是好的策展人,谁是坏的?人人手里都有麦克风了,我们该指望哪些策展人,该看哪些标志?

这正是问题所在,我不知道有什么好的解决办法。所以我才指望你们来告诉我。

不过我可以先直截了当地回答一点:就新闻而言,我认为众包策展的质量很糟糕。我觉得 Google 新闻毫无用处。它就在那儿,我也把它加了书签,但我仍然需要《纽约时报》。时报今天正式启动了拖延已久的实验,开始强制读者付费。没人知道结果会怎样,我在祈祷他们能撑下去。至于报业今天的凄凉处境,我不必在此复述。

当然,这不是 Google 的主业。我和世界上所有人一样,是 Google 的重度用户,我觉得这个搜索引擎,不用我奉承你们,你们能占据主导地位是有充分理由的:它真的管用。我所知道的其他任何引擎都做不到这么好。我知道你们遇到过困难,也在做调整,这很好。看人们想出多少办法来玩弄这套系统,是很有意思的事。

真正令人担忧的是,金钱如今扮演了多大的角色。我不知道在一个完全没有广告的世界里,Google 会是什么样,说实话我不知道它会不会不一样。我没有和这里的任何人讨论过,但我能感觉到你们花了大量精力把广告结果和非广告结果分隔开,我觉得这没问题。然而,有一种复杂的动力在起作用,我猜你们不可能不知道:人人都想被看见,人人都在争夺注意力。注意力如今成了有限的资源。于是,即便撇开广告的买卖不谈,人们也会,我不想简单地说「玩弄系统」或「欺骗搜索引擎」,人们会有意调整内容,以便排得更靠前。

与此相似,《纽约时报》如今也在某种程度上根据实时得到的读者偏好数据调整版面。这和我开头描述的情形完全不同。那时的编辑不知道、也不在乎读者想读什么。那些编辑身上有一种巨大的、我觉得颇为可爱的傲慢:他们坐在一间屋子里开会,说这件事很重要,因为我多年的从业经验告诉我它重要;虽然昨天华盛顿某个机构发生的这件事看上去很无聊,也上不了电视,但我知道它埋下了一颗种子,几周、几个月甚至几年后会发芽,所以我们要把它放在头版。他们经常错,经常带着偏见,其实永远带着偏见。可我相信,他们那样做,在某种意义上产出的东西比现在更有价值。现在他们被每分每秒的反馈推着走,随时知道谁在读什么。他们知道哪篇文章被转发得最多;如果他们把转发最多的文章提到页面上方,这里面有好的一面,因为有时读者确实比编辑懂得多;也有坏的一面,因为有时候那不过是查理·辛(Charlie Sheen)的八卦,而我不想要那种东西。

现场有人反驳说,大量依赖小报作风和八卦的出版物早在互联网之前就存在,《纽约时报》虽远不及别家,但也参与其中。这说明迎合大众的做法本来就是这些报纸获得庞大发行量的重要原因,何况它们当年还是极少数的新闻出口之一,纽约人口那么多,优质新闻来源又那么少,它们注定会有一定的受欢迎程度,于是可以有一点傲慢。但它们也不可能完全不在意什么标题能抓住眼球,不管那对读者的生活有没有真正的价值。互联网或许把这一点放大到了极致,但这并非网络独有。

我完全同意,这一点说得很好。我确实在给这场对话加码,把《纽约时报》描述得带了几分神话色彩。我说的那个《纽约时报》多少是个虚构的生物。每有一个人喜欢从时报获取新闻,就有好几倍的人更喜欢《纽约邮报》《每日新闻》或《国民问询报》。你可以说,那些就是他们选择的策展人,这没什么不好。你说得对,这和今天的情形极为相似:人们可以从一份菜单上做选择。

说到底,我真正相信的是什么?这也是为什么无论在书里还是在这里,我都无法为这个困境给出一个干脆的解决方案,因为最后我总是落到一些老生常谈上:我们是独立的个体,我们要做各自的选择,这意味着选择我们想追随的博主。有些博主在我关心的一些领域里,确实走在《纽约时报》前头。所以我再说一遍,那个《纽约时报》只是一个便于讨论的神话。

但是,就 Google 的主搜索引擎来说,或者就 Google 新闻的首页来说,如果你们不根据对每一位用户的了解来调整结果(当然,这在某种程度上正日益成为一种选项),你们如何做到对所有人都合适?有人回答说,这是一个非常困难的问题,而他们正在解答的途中。

编辑即算法:相似 vs 公民所需

接下来发言的是一位出身新闻世家的听众。这位听众先亮明立场:自己读《纽约时报》,希望它不要消失,也愿意为此付钱。我在此谢过。

这位听众从做编辑的家人那里学到的是:编辑决定的,是在所有采集来的新闻中,什么是这个社区需要知道的。你读的每一份报纸都有一套特定的信念体系。编辑本身就是一种算法,一种关于什么重要的算法。而 Google 在 1997 或 1998 年之所以对这位听众变得重要,是因为当时其他搜索引擎的结果开始变得「不自然」,商业部门和「新闻」部门之间的墙被拆掉了(这里的「新闻」是对算法搜索的一种宽泛比喻)。Google 把广告挡在搜索结果之外,这种张力至今仍在:一个需要资金、靠广告驱动的企业,如何同时给用户想要的东西,而用户想要的未必是广告主想要的?

所以问题在于:Google 更像一个非常好用的卡片目录,帮你找到与你告诉它的东西相似的东西;编辑则是在告诉你什么重要。算法上的差别就在于如何决定什么重要。对 Google 来说,重要的是「和你给我的东西相像的东西」;对编辑来说,重要的是「你需要知道什么,才能在一个民主社会里健康地生存」。这就是难点。

至于我问题的后半部分,众包的问题在于:众包非常擅长抹平个体的巨大高峰和巨大低谷。一位编辑的表现会有惊人的高峰和惊人的低谷,希望高峰和平均水平多于低谷,而其他来源,博客、其他报纸,都能帮助削平这个单一信号的低谷。可另一方面,它们削掉低谷的同时也削掉了高峰,你可能再也得不到某位编辑在某个领域里的卓越洞见。这也是报纸相对于 Google 的一个问题。最后一点,报纸并不覆盖人类活动的全部领域,只覆盖亮点。你在《纽约时报》上读不到深入的 DNA 科学、深入的物理学、深入的计算机科学,因为那些内容不具备普遍适用性。而我们的社会分工越来越细,所以你需要其他来源,期刊、搜索引擎、学术搜索引擎,把你引向那些领域里重要的东西。这就是选择哪个来源、它用什么算法、你指望它有多大广度之间的差别。归根结底,需要的是一种混合模式。

猜测意图:搜索比「相似」复杂得多

我觉得这番话非常有意思,我想抓住其中一点深究。关于 Google 主搜索引擎「返回和你要求的东西相似的东西」这个说法,在我看来,事情要复杂得多,也许更像报纸或博物馆策展人做的事,超出了刚才所承认的程度。

假如我输入「饮水机」,你们并不知道我要什么。你们心里很清楚,我在提问方面多少有点笨拙。我其实认为,如何使用 Google,如何组织问题以得到想要的东西,是人们必须培养的一项生活技能。而你们大概希望这项技能变得不那么必要,对吧?你们希望人们只管自由联想就能得到想要的东西。但这是不可能的,因为我在任何时刻想要的、想找的东西,可能非常具体:我可能有个医疗问题,可能想买件东西,也可能只是突然对某件事好奇。我知道我说的都是你们早就知道的事。可「和你输入的东西相似」到底是什么意思?你们并不只是在镜像用户的搜索词,「饮水机」「赛马」,或者「怎样做某某」,我猜这一定是 Google 搜索最常见的开头之一,虽然我从不这么开头。你们同样需要一些线索。

我想说的这一点,和刚才那个关于《纽约时报》不是物理学期刊的说法有关。时报是试图报道物理的,我在那儿工作时就写物理,那恰好是我的专业领域。但没错,我写的时候心里得装着一个虚构的时报读者,此人对物理有某种程度的了解,而这个程度是完全随意设定的。这意味着我得删掉一些东西,也意味着我对某些读者来说解释过度了。可当有人在搜索框里输入什么的时候,Google 不正面临着一模一样的问题吗?有人回应说,我描述的正是整个公司面临的挑战。我说不可能知道用户想要什么,可搜索团队的每个人显然都在为此努力。

回音室与意外之喜

刚才那位听众补充说,他想讲的其实是搜索领域一个叫「同质相吸」(homophily)的概念。推荐系统的一大问题(Google 的算法极为复杂,他自己也不完全理解)在于:Google 是通用目的的,而《纽约时报》面向的是受过教育的普通公民,六年级阅读水平以上,求的是一般性的认知;Google 可能服务于多种目的,而新闻理应有明确的受众,所以算法和报纸编辑是相似的。他所指的是「同质相吸」会导致博客圈里所谓的「回音室」(echo chamber):人们只订阅和自己观点一致的博客,永远没有异议,没有建设性的反面意见,最终在某种意义上走向愚昧。这是推荐系统和搜索引擎面临的挑战之一,很难产生意外之喜(serendipity)。而编辑理应对整个领域都有了解,并且有意教育读者去认识对社区重要的全部事件。

这又是一个真正关键的问题。我要说,它对那个神话中的《纽约时报》和对 Google 同样重要。在神话般的旧日时光里,你会把整份报纸读完,当然不是所有人都这样,但有些人是。有些人只是翻页而已,整个广告模式就建立在这上面。于是,即便你一辈子从未对印度发生的任何事情感兴趣,第五版上可能有一小块关于印度的报道,而因为你正无聊,没有别的东西争夺你的注意力,又因为你信任《纽约时报》,你就会去读它。这就是意外之喜。

在现代的网络版《纽约时报》上,这在某种程度上消失了,因为没人按顺序读报纸。正如刚才所说,人们倾向于去找自己已经感兴趣的东西。想读日本的消息,就可能花大量时间读,而对某个本可以在不同技术架构下偶然接触到的意外话题,一个字也看不见。

那么,这是否意味着,总体上,像我这样如今从网上而不是印刷品获取新闻的人,意外之喜的体验变少了?我不确定,真的不确定。我试着通过观察自己的一天来判断。我确实在随意地、偶然地看见许多以前不会看见的东西,我还能沿着路径往下走,这在印刷报纸上是做不到的:看到一个让我心动的东西,然后真的从这里走到那里。我不确定我们失去了意外之喜,也不知道这一点如何适用于 Google。

Madonna 测试:多样性与个性化

一位并非搜索团队的听众接着补充,说自己的话不代表官方立场,但 Google 搜索确实考虑多重信号。人们往往以为搜索就是找那个唯一正确的答案,实际上考虑的因素很多:来源是否权威、是否受欢迎、是否来自过去一直可靠的地方。这位听众十年前有一项检验搜索引擎的早期测试:输入「Madonna」,如果只得到那位歌手,就是糟糕的搜索引擎;如果同一页结果里既有歌手,也有基督教的圣母,还有黑圣母雕像,那就说明引擎在认真猜测我可能在找什么,并让我可以从那里继续深入。所以把搜索看作「找到正确答案」是危险的。看看 Google 结果页的变化,你得到的不只是文本页面,还有图片区块,还可以按新闻、按博客来筛选结果,这是一个非常积极的方向。

这引出了一个非常有意思的问题。假设一千个人输入「Madonna」这个词,其中一部分人找的是歌手,对基督教话题毫无兴趣;另一部分人从没听说过那位歌手,我猜这部分人少一些。但你们并不知道谁是谁。或者你们其实知道?你们可以把系统设计成知道的样子,记录用户的历史和以往的搜索,从而对此人找的是歌手还是耶稣之母,掌握比随机猜测更好的信息。做不到这一点的话,你们就呈现十条结果,这十条未必包含那一千个输入「Madonna」的人各自想要的一切。我没有问过任何人,但我推测结果一直在受到用户点击行为的影响。可是,除非用户在用你们的工具栏或你们的操作系统,否则追踪用户总有一个限度。我不知道,我也不是在打听。

而这一切又绕回到意外之喜的问题上:你们并不希望结果完美无缺。对一个心里想着流行歌星而输入「Madonna」的用户来说,被迫面对 Madonna 的其他几种含义,也许反而更有价值。我不知道,我开始把自己绕进去了。

实时对话、流感预测与真相仲裁

又有听众指出,搜索所获得的信号至少有一部分本质上是某个概念或某个网站的受欢迎程度,这确实可能让结果向某个方向倾斜。但就 Google 和《纽约时报》编辑的角色对比而言,一个区别是:可以把一次搜索看成一场对话。我搜「Madonna」,得到的结果也许不是我预期或想要的,于是我说,好吧,那我搜「Madonna」加上别的词。这几乎是用户和搜索引擎之间的一场实时对话。编辑得到的反馈则来自读者来信,那是事后的,等信寄到,编辑早已转向下一件事了。所以你们可以追踪用户的下一次搜索,再下一次搜索,把信息拼在一起。点击也是一种信号:用户点了某个结果,这个信号回到你们那里,说明排序做得还算合理。

这有它的局限,正如你们所说。但同样明显的是,用算法判断来源是否可靠也有限度。你们有人提到这也是方程的一部分。维基百科这类众包的信息事业也是如此:它们有一套逐步逼近真相、至少是某种准确性的过程。然而,认为奥巴马不是在美国出生的人,数量仍然大得惊人,而且网上有一股可观的势力。

我在书里提到过一个众包做得好的活例子,你们肯定都熟悉,因为它经常被引用:Google 能根据人们搜什么、在哪里搜,比疾病控制中心更快、更准确地预测流感疫情的传播。另一方面,如果 Google 被骗局、时尚风潮或错误信息裹挟,我不认为这就意味着 Google 失职。也许这就是你们选择的模式:你们本来就不该充当真相的仲裁者。同样,我相信新闻机构也越来越不愿意做真相的仲裁者,而更愿意做中立的记录者,记录人们怎么想、人们在争论什么。无论哪种方式,都是福祸参半。

对此,有人回应说,这里面掺杂着好几种「权威」的概念,需要拆开来看。但先说一点:如果 Google 是人类欲望的一面镜子,那么你不喜欢镜中的影像,某种意义上是我们的问题,某种意义上又不是。Google 确实希望不只做大众口味的仲裁者,而是尽量接近权威性的真相。这里被抛来抛去的「权威」至少有三种:一是待读内容的作者本人的权威;二是策展人的权威,他挑选你应该读的东西,哪怕你不知道自己应该读,比如印度,或者因为太新而你从未接触过的概念;三是对你所提问题的回答的权威,也就是搜索引擎的权威。在这位听众看来,一个接到了明确请求的搜索引擎,未必有义务提供意外之喜。你搜「Madonna」,我们却给你看一些你不知道的、与 Madonna 相关的乐队,这固然是意外之喜,但显得有点牵强,毕竟你问的是 Madonna。恰当的做法是给你多种不同的话题,告诉你这是一个宽泛的查询,请你再做一次细化,告诉我们你找的是哪一个 Madonna。他们更关心的权威问题是:用什么标志,电子的、语义的或其他的,来判定网上谁对用户的问题给出了最好的回答。这正是人们所说的靠博客来玩弄系统之类的事情,尽管他们不认为博客真像人们说的那样常常浮到结果顶端。

于是他们反过来问我:写作研究时,我是怎么判断什么是权威来源、什么是值得信任的来源的?不是编辑,而是内容的作者本人,我看的是什么?

我如何判断可信来源

我是从报社记者做起的,所以我走的是几条老式的路子。老式技巧之一是:你出门去采访人,然后采访更多的人。不幸的是你有截稿期,时间总会用完,所以报纸上满是错误。报纸也有经验丰富的文字编辑,能抓住其中一些错误。对我来说,时间尺度后来变长了,因为我写的是杂志文章或者书,有更多的时间。

如今,你们肯定知道,各种各样的机构都有绝对的规定,禁止把维基百科当作事实来源。如果你是报社记者,出了错,对编辑说「我是从维基百科上查到的」,你就完蛋了。同样,如果你是大学生,交的论文引了维基百科,我很确定在大多数学校,规则是维基百科不是可信来源。然而同样真实的是,维基百科极其可靠,也极其有价值。我毫不羞于承认,写这本书时我花了大量时间在维基百科上查东西。我花在 Google 图书上的时间甚至更多,哪怕当时我正参与对它的诉讼。你们愿意的话,可以叫我伪君子。

书里的东西也并非全都可信。信息洪流(洪流,我书名副标题的最后一个词)的伟大之处在于,你可以成为自己的策展人。你可以自己挑选,也必须自己运用判断力。这是一种挑战,当然也是一种责任。

付费墙与信息洪流下的变现难题

一位自称可能代表着「垂死的老一代」(至少在 Google 内部算是)的听众说,自己非常愿意为编辑的专业能力付费,而且相信很多人也愿意以各种方式付费。所以《纽约时报》确实必须改变商业模式:有些人只愿意用注意力来付费,另一些人愿意用钱来付费,但他们大概不会再订纸质报纸了。这位听众订了一辈子的报纸,几年前才停掉,曾是时报的付费高级会员,后来时报把钱退了回来,原因不明。这位听众的意思是,这个国家有一些编辑专业能力的瑰宝,时报是其中之一,希望它不要放弃。他们拥有的东西一定有办法变现,也许他们还不知道那是什么,但希望他们继续寻找。

我希望你是对的,但我不完全确定。这里有一个技术问题,我看不到理想的解决方案,甚至不确定有没有好的方案:东西一旦上网,就太容易传播了。旧世界里,信息提供者能赚钱,很大程度上靠的是不便:纸张,图书馆。《纽约时报》希望《赫芬顿邮报》做它的扩音器和读者来源,但不希望《赫芬顿邮报》反过来抢走它的饭碗,提供足够多的内容,让人们不必再付钱给时报。我的情况也一样。你们不必付十美元(有补贴的价格,让我心跳加速),你们现在就可以在网上免费读到这本书相当大的一部分。我不知道具体多少。我希望那种阅读体验不太令人满意,希望你们时不时会被一页读不到的内容打断。但毫无疑问,有些人在旧世界里为了读某个话题的十页内容会把整本书买下来,如今不必了。时报会拿到我的钱,不管他们竖起什么样的防火墙,我看他们也拿到了你的钱。但他们无法阻止人们传播那些文章的全部或部分,而一旦有人这么做,我不确定剩下的还够不够。我不知道。

这位听众回应了两点。第一,有些人会读我这本书的十页,而他们以前一页都不会读,因为在线读那十页太容易了;其中又有一部分人会因此去买整本书,而且不止花十美元。这位听众自己也是作者,作品被大量盗版,整本书都能下载,但并不为此耿耿于怀,书照样有人买。总之,这位听众比我乐观。第二,这位听众每天多次访问《纽约时报》网站,把整个首页看一遍,倒是有点希望它更像纸质报纸的头版,因为那是一种有意思的新闻选择。这位听众不愿意放弃这样一个想法:我们有信任的人。而且不只是信任,还有审美上的认同。有些出版物,他们读它只是因为觉得会有意思,未必相信里面的大部分内容,但知道这些人有一种有趣的视角。没错,信息容易传播,没错,一切都只是比特;可另一方面,那些只管转发比特的人,并不在乎提供与原创者相同的体验。顺便说一句,在这里真正吃亏的是音乐人,因为那真是逐比特的复制。音乐人、电影人等等,那是更难的问题。

我倒想回应你刚才说的另一件事:是的,我也希望《纽约时报》网站的头版更像纸质头版。我前面描述过纸质头版是如何按照一种理应富有意义的方式结构起来的。可另一方面,我一天要看十次网络版,我并不指望它每次都一样,不指望我回来时它还是老样子。所以这里有一个动态信息问题,没有人有好的解决办法。光知道我已经来过,甚至知道我读过什么,都还不够。有人提醒说,至少订阅周日版的读者,可以下载一种叫「复刻版」(replica)的版本,就是头版的原样,虽然记得去下载有点麻烦,但如果你渴望那种东西,它是有的。

熵、麦克斯韦妖与气候变化的政治化

有听众把话题转向了物理。我前面谈到香农,谈到信息也许是一个更宽泛的概念,也许有某种哲学含义。这位听众更关心它的物理含义:香农定义了一个量,叫「熵」(entropy),结果它恰恰是热力学里研究的那个东西。这位听众从未得到过一个好的解释,问我有没有什么可说的。

好极了。此刻你从我这儿得不到一个好的解释,因为它实在太难了。但书里有,有一整章讲熵,讲麦克斯韦妖(Maxwell's demon)。在这个场合展开太多了,对我这颗小脑袋来说也太多了。我只讲一个笑话。顺便说,你说的完全正确:香农推导出一条消息的概率数学时,发现这套数学与热力学里已有的东西完全吻合,于是他开始把「熵」这个词几乎当作「信息」的同义词来用。贝尔实验室流传着一个都市传说:是约翰·冯·诺伊曼(John von Neumann)建议香农用「熵」这个词的,理由是,这样一来谁也不知道他在说什么。

最后一个问题来自另一个角度。我们谈了很多人们如何获取信息、这如何改变、有什么问题、有什么机遇。这位听众想问的是:既然情况已经变了,人们对信息的需求也许更大了,可他们得到的都是经过各自过滤器筛选的信息,那么这个世界将会、应该、可以如何更有效地与不断演变的媒体环境合作,围绕那些政治敏感、数据密集的议题展开最有效的对话?这位听众给了一个具体的例子:气候变化。这里面有开放与透明的问题,牵涉大量数字活动和电子邮件;有政治问题,人们各取所需;还有科学问题,严肃的科学家不习惯这样的世界。我怎么看今天科学与公众的沟通,以及围绕它的政策?

我想我可以这样总结:情况一团糟,而我们正处在一个我认为可能是调整期的阶段。我希望这是一个调整期。我不确定这些问题有什么简单的答案。

气候变化在某种意义上是一个典型的例子。二十年前,我还是《纽约时报》的科学记者,我记得自己写过气候变化拼图里几个具体的小块。当时我根本没想到这件事会有半点争议。一切都相当直截了当,没有被政治化。如今它已被政治化到如此激烈的程度,以至于很多人会谈论你「信不信」气候变化,就像信不信上帝一样,这在我看来相当古怪。它本该是一种技术性的知识,一种非政治的知识。我们都知道为什么不是这样:因为金钱对我们太多信息渠道的影响。就气候变化而言,我认为这一点非常清楚。有一些行业,石油和天然气行业,花大钱资助伪智库,在我看来是在花钱雇研究者撒谎。还有另一些人,并无恶意,也不自觉,只是被这些来源散播的错误信息扭曲了看法。至于福克斯新闻,我就不展开了。

我磕磕绊绊想说的是:在我们这个复杂的世界里,信息来源没有真正的纯洁可言,我们这些信息消费者面临的挑战只会越来越大,必须时刻警惕,保持怀疑。这同样适用于关于药品的信息,适用于那些更明显的政治议题,也以更微妙的方式适用于那些不那么政治化、与金钱关系不那么直接的事情。

到目前为止,我认为这家公司做得相当不错,而人们应该继续对它保持警觉。你们也是一家营利企业。尽量不作恶吧。谢谢大家。

排版 + 横图 + 来源,粘贴即成稿
章节 · 点击跳转视频
0:00 开场:作者与《信息简史》 ▶ 正在看
1:25 香农 1948:通往 Google 之路的起点 ▶ 正在看
5:49 《纽约时报》:读者买的是被删掉的新闻 ▶ 正在看
8:34 核心命题:Google 是过滤器而非搜索 ▶ 正在看
12:43 权威与策展:众包能胜过编辑吗 ▶ 正在看
20:50 编辑即算法:相似 vs 公民所需 ▶ 正在看
24:30 猜测意图:搜索比「相似」复杂得多 ▶ 正在看
27:19 回音室与意外之喜 ▶ 正在看
30:55 Madonna 测试:多样性与个性化 ▶ 正在看
35:08 实时对话、流感预测与真相仲裁 ▶ 正在看
39:00 作者如何判断可信来源 ▶ 正在看
41:47 付费墙与信息洪流下的变现难题 ▶ 正在看
47:01 熵、麦克斯韦妖与气候变化的政治化 ▶ 正在看
本期小问 · 档案清单
8:34 媒介如何改变我们所知的世界? ▶ 正在看
39:00 听来的知识可靠吗? ▶ 正在看
27:19 知道得越多,就离真相越近吗? ▶ 正在看
20:50 我们造的东西,何时开始管我们? ▶ 正在看
01开场:作者与《信息简史》
0:00
uh James click is here today to speak about his new book the information a history a theory a flood it's been called a re revelatory Chronicle that shows us how information has become the modern era's defining quality uh in this work he refers to the alphabet as a founding technology of information uh the telephone fax machine calculator and ultimately the computer are only the latest Innovations on that uh devised for saving manipulating and communicating knowledge unafraid of grappling with the sometimes extremely dense scientific theories behind uh behind his work this well- researched yet very approachable book is sure to be an instant success much like his previous bestsellers uh chaos which helped popularize popularized the chaos theory among other ideas in the mid90s and genius a biography of legendary physicist Richard feineman he's a leading chronicler of Science and modern technology and we're very happy to have him at Google today so please welcome today's author James [Applause] pleas thank you Jesse thank you all for coming thank you for having me here um I've been I've been here before because I was involved with uh some of the negotiations over the lawsuit against Google over the copying of of all those copyrighted books and I had fun then and
今天詹姆斯·格雷克(James Gleick)来到这里,谈他的新书《信息简史:一段历史、一种理论、一场洪流》。这本书被誉为一部富有启示性的编年史,向我们展示信息是如何成为现代社会的决定性特质的。在这部作品里,他把字母表称为信息的奠基性技术,而电话、传真机、计算器乃至最终的计算机,都不过是为储存、处理和传播知识而发明的最新创新。他毫不畏惧地去啃自己作品背后那些有时极其艰深的科学理论,这本考据扎实却又非常平易近人的书,肯定会像他此前的畅销书一样一炮而红——比如《混沌》,它在九十年代中期让混沌理论以及其他一些思想广为人知;还有《天才》,传奇物理学家理查德·费曼的传记。他是科学与现代技术领域首屈一指的记录者,我们非常高兴今天能请他来 Google,所以请大家欢迎今天的作者詹姆斯。(掌声)谢谢你,Jesse,谢谢大家过来,谢谢你们请我来。呃,我以前来过这里,因为我参与过一些谈判,就是关于那起针对 Google 复制所有那些有版权书籍的诉讼。当时我还挺开心的,
便签笔记
02香农 1948:通往 Google 之路的起点
1:25
I'm I'm glad to be back and I I'm really glad I'm a little surprised to see so many people people buying um my book in this quaint format um Google makes a brief appearance at the end of my book when I began writing it I actually imagined that I would do a lot of reporting here and and that Google would be sort of a big feature of the end of the book and it didn't work out that way uh Wikipedia there's much more about Wikipedia in the book than there is about Google and in a way I think it's because it became unnecessary um you are so ubiquitous you're so much a part of people's lives I don't need to tell you the premise of the book is to put it here in a way I don't usually put it that the road that leads here the road that leads to Google begins at a particular point in the middle of the 20th century in 1948 with the publication almost full-blown of what was quickly referred to as information Theory by Claude Shannon at Bell Labs um in the form of a pair of technical articles in the in the Bell systems technical Journal B yeah Bell systems technical journal and then it was
我很高兴能再回来。我真的很高兴,也有点惊讶,看到这么多人在买我这本书,而且是这种古雅的形式。Google 在我书的结尾处短暂地出现了一下。刚开始动笔的时候,我其实设想自己会在这里做大量的采访报道,Google 会是全书结尾的重头戏,但结果并不是那样。呃,维基百科——书里关于维基百科的内容比关于 Google 的多得多。某种程度上我觉得,那是因为写 Google 变得没有必要了:你们太无处不在了,已经深深嵌进了人们的生活。我其实不需要跟你们讲这本书的前提是什么——不过我还是换一种我平时不太用的说法讲一下:通向这里的那条路,通向 Google 的那条路,起点是二十世纪中叶的某个特定时刻,1948 年,一种几乎一出场就已成熟、很快被称为「信息论」的东西问世了,作者是贝尔实验室的克劳德·香农,形式是《贝尔系统技术期刊》上的两篇技术论文。对,《贝尔系统技术期刊》,后来它又
便签笔记
2:59
republished it's a book the mathematical theory of communication and do you have it do have you do any of you own it it's a few people uh I would think at a place like this it should be a sort of I think that there's a lot in it that is relevant to all of you whatever part of the business you happen to be in I first heard about it um from chaos scientists when I was working on my first book uh and I remember even though it was quite a while ago being somewhat taken aback by the just the idea that there was such a thing as information Theory as a mathematical theory of something that I thought of as being not a natural subject for mathematical analysis this vague amorphous thing called information now as I hardly need tell you we live in a world where information is not just instructions news gossip it's not even just words it's in a big thing that includes music and images and anything that you can digitize and store on a computer and search via Google um there's a connection between these things I believe that Shannon's
重新出版了——成了一本书,《通信的数学理论》。你们有这本书吗?你们当中有人拥有它吗?有几个人。呃,我原以为在这样一个地方,这本书应该算是……我觉得书里有很多内容跟在座各位都有关,不管你们做的是这行里的哪一块。我第一次听说它,是在写第一本书、接触混沌学者的时候。我记得——虽然那已经是相当久以前的事了——我当时很有点吃惊,仅仅是「竟然存在信息论这么个东西」这个想法本身:一门关于某种东西的数学理论,而我原本觉得那东西根本不是做数学分析的天然对象,那个模糊的、没有定形的、叫做「信息」的东西。当然,现在我几乎不用跟你们说,我们生活在这样一个世界里:信息不只是指令、新闻、八卦,甚至不只是文字,它是个很大的范畴,包括音乐、图像,以及任何你能数字化、存进计算机、再通过 Google 搜索到的东西。我相信这些事情之间是有联系的:香农的
便签笔记
4:27
Theory Shannon's and then all of the work that followed it by mathematicians and soon after computer scientists lie underneath the structure of our world not just as a technical Foundation but also in a more in a more genuine way in a in a broader way I'm I'm trying to avoid using the word in a philosophical way with this premise Then I then the structure of my book is also it also involves going back in time to the beginning of human history with the view that the information age isn't just the thing that we've been talking about for 50 years and it is 50 years by the way according to the Oxford English Dictionary the information age um is an expression that first started getting tossed around in 1960 um but all of human history is an Information Age we know that now the printing press is and information technology the telegraph the telephone and at the beginning of recorded history the alphabet so that's the the basic idea of my book it was a little bit difficult to organize the story starts at the beginning that is
理论,香农的理论,以及随后由数学家、很快又由计算机科学家做出的所有工作,都躺在我们这个世界的结构底下——不只是作为技术基础,还以一种更真切的方式、更宽泛的方式存在。我在努力避免说「以一种哲学的方式」。有了这个前提,我这本书的结构也就包含了回到过去、回到人类历史开端的部分,其视角是:信息时代并不只是我们过去五十年一直在谈论的那个东西。顺便说一句,确实是五十年——根据《牛津英语词典》,「信息时代」这个说法是从 1960 年才开始被人挂在嘴边的。但整个人类史都是一个信息时代。我们现在知道,印刷机是一种信息技术,电报、电话也是,而在有文字记载的历史之初,还有字母表。所以这就是我这本书的基本想法。它有点难组织,故事从开头讲起——也就是说,
便签笔记
03《纽约时报》:读者买的是被删掉的新闻
5:49
the story starts in the middle in 1948 and then goes back and then goes forward um I don't want to give a lecture about it here I would rather have a sort of conversation because I feel I should be still sort of feel I should be interviewing you so but so let me let me start with a a sort of question um I used to work at the New York Times that was where I began my career and the New York Times was a great information Enterprise and still is something that uh some wise old head at the Times used to say was that what the reader is paying for is not all of the news we put in all the News That's fit to print what the reader is paying for is all the news that we leave out if you were a typical New York Times reader in a previous generation the New York Times might have been your only source of information I mean I remember my grandfather reading it on the subway and there was a certain way you had to fold it in order to read it on the subway otherwise you just end up in a big tangle of paper and and in there was
故事其实是从中间开始的,从 1948 年讲起,然后往回走,再往前走。我不太想在这儿做一场讲座,我更愿意来一场对话,因为我总还是觉得应该由我来采访你们才对。所以,那我先抛一个问题吧。我以前在《纽约时报》工作,那是我职业生涯开始的地方。《纽约时报》是一家了不起的信息企业,现在依然是。时报有些老练的智者以前常说:读者掏钱买的并不是我们放进去的所有新闻——「一切适合刊印的新闻」——读者掏钱买的,是我们拿掉的所有新闻。如果你是上一代典型的《纽约时报》读者,这份报纸可能就是你唯一的信息来源。我记得我祖父在地铁上读它,你得用某种特定的方式把它折起来才能在地铁上看,否则你手里就是一团乱纸。曾经有那么一段时间,
便签笔记
7:06
a time when that way of folding the New York Times was taught in school but it was finite famously it was made of paper and it was only so big and the the conceit of the editors of the times was that if you read the whole thing you were up to date on the news not only that the front page was as as you may know it still is highly structured uh a huge amount of discussion formal discussion took place every day about what news article should lead the paper that is be placed in the upper right corner and what was the second most important article which should be the off lead in the upper left and and which stories were Worthy of um headlines above the fold or below the fold and whether headlines should be two columns or three columns or more and um this judgment the the exercise of judgment that was involved here was called news judgment and editors at the times were very proud of possessing this quality of news judgment and knowing for example that the murder of a human being um was not a fixed quality that that some murders of human beings could be completely
学校里还会教怎么这样折《纽约时报》。但它是有限的,众所周知它是纸做的,就那么大。而时报编辑们的自负之处在于:如果你把整份报纸读完,你就跟上了所有新闻。不仅如此,头版——你们可能知道,现在依然如此——是高度结构化的。每天都要花大量时间进行正式的讨论:哪条新闻应该做头条,也就是放在右上角;第二重要的是哪一条,也就是放在左上角的次头条;哪些报道配得上放在折线以上或折线以下的标题;标题应该跨两栏、三栏还是更多。而这里所涉及的这种判断力的运用,被称为「新闻判断力」。时报的编辑们非常以自己具备这种新闻判断力为荣,比如懂得:一个人被谋杀,其分量并不是固定的,有些命案可能完全
便签笔记
04核心命题:Google 是过滤器而非搜索
8:34
unimportant and not even worth mentioning virtually all of them in fact but if the human being happened to be um a doctor living in Scarsdale who was the famous author of a a diet book and the uh accused murderer happened to be his spurned lover a Scarsdale matron um well that was was a really important murder and there and the times did not apologize for making these judgments this is all leading up to a question it seems to me that Google which is famously a search engine just as much as the new the possibly mythical New York Times that I'm describing is also um based upon filtering that um the purpose you serve is not just to pull information from the internet and present it to the reader it's to filter out the information that the user is not looking for just as the New York Times was filtering out all of that news that users of the internet internet today so often complain is bombarding them on the one hand everybody is thrilled that they no longer have to rely on stock authoritative anti-democratic I gave a talk um last
无关紧要,甚至不值一提——事实上几乎所有命案都是如此。但如果被害人恰好是住在斯卡斯代尔的一位医生,是一本著名减肥书的作者,而被控的凶手恰好是被他抛弃的情人、一位斯卡斯代尔的贵妇——那么,这就是一桩非常重要的谋杀案了。而时报对做出这些判断毫无歉意。这些都是在铺垫一个问题。在我看来,Google——众所周知是一个搜索引擎——就跟我刚才描述的那个(可能已成神话的)《纽约时报》一样,同样是建立在过滤之上的。你们所起的作用不只是把信息从互联网上抓出来呈现给读者,而是把用户并不想要的信息滤掉,就像《纽约时报》当年滤掉的那些新闻一样——而今天互联网用户老是抱怨自己被这些东西狂轰滥炸。一方面,每个人都为不必再依赖那些老套的、权威的、反民主的东西而兴奋。我昨天晚上在伯克利做了一场演讲,
便签笔记
10:06
night in Berkeley uh that was not like what I'm saying tonight but at some point I used the word Democratic in a there was I was setting up some sort of dichotomy between Democratic things on the internet and authoritative things and a woman yelled out from near the back democracy Now is that is that Berkeley or um what was I saying we we consumers of news on the internet have ambivalence Express ambivalence on the one hand they are delighted to be freed from the tyranny of of of authorities like the New York Times with their um bgea or business oriented or anti catholic whatever Prejudice you have you want to assign to the editors of the New York Times um biases and um empowered by the millions of blogs and other websites and other sources of what we may or may not call news depending on whether we want to be snobs but the same people often complain then how do I know what's true how do I know what's false how do I find what's meaningful for me now I know that Google
内容跟我今晚讲的不一样,但讲到某个地方我用了「民主的」这个词,当时我是在互联网上民主的那类东西和权威性的那类东西之间设了某种二分法,结果后排附近有个女士喊了一句「Democracy Now!」——这算是伯克利特色吗?呃,我刚说到哪儿了。我们这些在互联网上消费新闻的人是矛盾的,会表现出一种矛盾心理:一方面,他们很高兴能从《纽约时报》这类权威的专制中解放出来——不管你想安到时报编辑头上的偏见是什么,是资产阶级的、还是商业导向的、还是反天主教的,随便哪种偏见;同时又被数以百万计的博客和其他网站、其他我们可能称之为、也可能不称之为新闻的来源赋予了力量——这取决于我们想不想摆出一副势利的姿态。但同样是这些人,又常常抱怨:那我怎么知道什么是真的?怎么知道什么是假的?我怎么找到对我有意义的东西?我知道 Google
便签笔记
11:29
has has a a big I don't know if you call it a news Operation but there's Google news and um I take it that it is um either entirely automated algorithmic crowdsourced or at least mostly that there's not much inter intervention from Human editors um I think it's there's a great experiment underway um whether intelligent crowdsourcing if that's what you call it here or algorithmic approaches to deciding what really matters to the user can be better than the old-fashioned human counterparts which include the New York Times or a book publisher people used to buy books based on uh the name of the publisher my publisher is Pantheon some people are old enough to associate that name with a certain kind of writing that's more or less obsolete these days I don't think many people buy books based on what the the publisher is is willing to choose to
有一块挺大的——我不知道该不该叫它新闻业务——但有 Google News。我理解它要么完全是自动化的、算法的、众包的,要么至少大部分是这样,人工编辑的介入不多。我觉得,这里正在进行一场了不起的实验:智能众包(如果你们在这儿是这么叫的话),或者用算法的方式来决定什么对用户真正重要,究竟能不能胜过老派的人类同行——那些人类同行包括《纽约时报》,也包括图书出版商。人们过去会根据出版社的名字来买书。我的出版社是万神殿(Pantheon),有些人年纪够大,会把这个名字跟某一类写作联系起来,而那类写作如今差不多已经过时了。我不觉得现在还有很多人会根据出版社愿意选择
便签笔记
05权威与策展:众包能胜过编辑吗
12:43
publish um um an editor of a poetry Journal is an authoritative Source anti-democratic but every blogger of course is is anti-democratic in in his or her own way is making personal judgments that you can choose to trust or not trust so um so that's the that's my question I'd love to hear from anybody who might agree or disagree that Google's primary function is filtering and not searching uh it's an excellent question I'm sure everybody's got their own answer I'd like to just start with maybe a back and forth rep on this uh because you bring up two subjects that we have specific names for them you mentioned one of them which is Authority and the other which I think is Loosely referred here as curation which is to say the editor the the Chooser of what is important to read and what is not what's the word again uh like a a curator someone who is is choosing selecting what ought to be kept what is valuable and what's not valuable and and it's a a single Source or a few sources you may go to uh to depend on on what it is you need to know uh as as I I'd like to hear first your thought on uh the quality of what your perceived quality is of the of crowdsource crowd source curation and uh in lie of a new era where everybody has
出版什么来买书。呃,一份诗歌期刊的编辑是一个权威性的来源,是反民主的;但当然,每个博主也都以他或她自己的方式是反民主的,都在做个人化的判断,你可以选择信任或不信任。所以,这就是我的问题。我很想听听有没有人同意或不同意这个说法:Google 的首要功能是过滤,而不是搜索。这是个很棒的问题。我相信每个人都有自己的答案。我想先跟你来回聊一下,因为你提到了两个我们有专门叫法的主题。你提到了其中一个,就是权威性;另一个在这儿我们大致叫做「策展」,也就是编辑,那个挑选什么值得读、什么不值得读的人——那个词怎么说来着?呃,就像策展人,一个去挑选、决定什么该被留下、什么有价值什么没价值的人。而且它是单一来源,或者少数几个来源,你会根据自己需要知道什么而去找它们。我想先听听你的想法:你觉得众包式策展的质量如何?在这样一个人人都有
便签笔记
14:06
access to be a publisher or at least a recommender of lists how do we decide who is a good curator and a bad curator everybody has a microphone now and so who who what curators should we be looking to and what flag should we be looking for yeah well this is the problem there's no good solution to this problem that I that I know of that's why I was hoping you would you would tell me I mean I think so first of all I'll give you a straightforward answer to what what I think is the quality of the curation in regard to news I don't I think it's crap I mean I I don't find Google News at all useful uh it's there it's I've got it bookmarked I still need the the New York Times the New York Times today is beginning its long delayed experiment in forcing people to pay um nobody knows how that's going to turn out I'm praying that they can make a go of it meanwhile well I don't need to rehearse the entire sad state of the newspaper industry today that's not Google's main function I'm like everybody else in the world a huge Google customer and I find the search engine to be I don't need to flatter you I mean it's you're dominant for a good reason it really works um nobody else that I know does as well I know you've had difficulties and I know
机会成为出版者、或者至少成为榜单推荐者的新时代里,我们该怎么判断谁是好的策展人、谁是坏的策展人?现在人人都有一个话筒,那我们该去看哪些策展人?该留意哪些信号?是啊,这正是问题所在。据我所知,这个问题没有什么好的解决办法,所以我才希望你能告诉我。我想,我先直接回答一下:我觉得在新闻方面,策展的质量怎么样?我觉得很烂。我是说,我完全不觉得 Google News 有用。它就在那儿,我也把它加了书签,但我还是需要《纽约时报》。《纽约时报》今天开始了它拖延已久的实验,要让人们付费。没人知道结果会怎样,我祈祷他们能成功。同时,呃,我也不用在这儿把今天报业的整个惨状再复述一遍,那不是 Google 的主要功能。我跟世上其他所有人一样,是 Google 的重度用户,我觉得这个搜索引擎——我不需要恭维你们——你们占据主导地位是有充分理由的,它是真的好用。据我所知没有别人做得这么好。我知道你们遇到过困难,我也知道
便签笔记
15:29
you're making adjustments and that's great um it's fascinating to see how many ways people come up with to try to game the system um it's genuinely worrisome how much of a role money plays now I don't know what Google would look like I don't I honestly don't know whether it would look different in a world where there just was no advertising involved anywhere in the picture I know that uh I mean I perceive without having discussed it with anybody here that a lot of attention is being paid to keeping a separation between advertising results and non-advertising results and I think that's fine um nevertheless there's a dynamic at work that's complicated that I know you're you you've got to be a I mean I infer that you've got to be aware of it where um everybody wants to be seen everybody is fighting for attention attension is the finite quantity now and so even apart from buying or selling ad advertising people try to I don't want to just say game the system or trick the search engine people tailor their content in order to rank more highly analogously the New York Times to some
你们在做调整,那很好。呃,看到人们想出这么多花样来试图操纵这个系统,真的很有意思。金钱如今扮演了多重的角色,这一点真让人担忧。我不知道 Google 会是什么样子——我真的不知道,在一个整个图景里任何地方都没有广告的世界里,它会不会不一样。我知道——我是说,虽然没跟这里的任何人讨论过,但我能感觉到——你们花了很多心思在广告结果和非广告结果之间保持隔离,我觉得那很好。不过尽管如此,还是有一种复杂的动力在起作用,我知道你们肯定——我是说我推测你们肯定意识到了:每个人都想被看见,每个人都在争夺注意力,注意力如今才是那个有限的量。所以哪怕撇开买卖广告不谈,人们也会去——我不想直接说成操纵系统或者欺骗搜索引擎——人们会去调整自己的内容,以便排名更靠前。类比来看,《纽约时报》在某种
便签笔记
16:59
ENT rearranges its offerings based on information it's getting in real time about what people like to read That's entirely different from what I described at the beginning where editors didn't know and didn't care what people wanted to read there was a tremendous and I think Charming arrogance in involved in the editors who would sit around in a room and have a meeting and say this is important because my years of experience in the business tell me it's important and because even though this this event that took place in uh some agency in Washington yesterday looks pretty boring and it's not going to be on TV I know that it's planting a seed that is going to emerge in a few weeks or a few months or even a few years and so we're going to put this on the front page um often they were wrong often they were prejudiced always they were prejudiced um I believe that in some way they were producing more valuable results that way than they are if they get pushed around now by the minute to- minute feedback they get about who wants to read what you know they know what's the most emailed article if they take the most emailed article and then move that up on the page there are some things about that that are good sometimes the users
程度上也会根据它实时获得的、关于人们喜欢读什么的信息,来重新安排自己的内容编排。这跟我一开始描述的完全不同——那时候编辑们不知道、也不在乎人们想读什么。那些围坐在一个房间里开会的编辑身上,有一种巨大的、我觉得也挺可爱的傲慢:他们会说,这条重要,因为我在这行干了这么多年,我的经验告诉我它重要;因为,尽管昨天发生在华盛顿某个机构里的这件事看起来挺无聊,也不会上电视,但我知道它埋下了一颗种子,几周后、几个月后、甚至几年后会冒出来,所以我们要把它放到头版。他们常常是错的,常常是带偏见的,而且永远是带偏见的。但我相信,从某种意义上说,他们那样做产出的结果,比他们如今被那种关于谁想读什么的分分钟反馈推着走要更有价值。你知道,他们知道哪篇文章被邮件转发得最多,如果他们把转发最多的那篇挪到页面上方——这里面有些方面是好的,有时候用户
便签笔记
18:27
do know more than they do and there's some things about it that are bad because sometimes it's just Charlie Sheen and I don't want that I think the sheer existence of uh massive Publications that rely heavily on tabloids gossip and the fact that the New York Times is a purveyor of that certainly not nearly as much as some others but but they they participate in that game as well uh the fact that they've existed long before the internet shows that popularism to some extent is a big part of what gave these newspapers such large circulation besides the fact that they were one of a very very few outlets for any news right that New York has such a large population they had so few quality news sources anyway they were just bound to get some level of popularity and therefore could be a little arrogant but without paying some attention to what a head what was an attention grabbing headline regardless of whether or not that provided real value for the user's life I I think the internet has maybe uh hyperextended that concept but I don't think it's Unique to to the web and I absolutely AG I absolutely agree with you you're making a very good point and I'm I'm definitely loading the conversation here by by describing the New York Times in this slightly mythical
确实比他们懂得多,但这里面也有一些不好的地方,因为有时候大家关心的就只是查理·辛那种事,而我不希望是那样。我觉得,光是存在那么多严重依赖小报八卦的大型刊物这件事,再加上《纽约时报》本身也在兜售这类东西——当然远不像另外一些媒体那么严重,但它们也确实参与了这个游戏——而且它们早在互联网出现之前就存在了,这说明迎合大众口味在某种程度上正是这些报纸能拥有那么大发行量的重要原因之一。除此之外还有一点,当时能提供新闻的渠道非常非常少,对吧?纽约人口那么多,优质新闻来源却那么少,它们本来就注定会获得一定程度的受欢迎度,因此也就可以稍微傲慢一点,但前提是仍然得多少留意一下什么样的标题能抓住眼球,而不管那对读者的生活是否真的有价值。我觉得互联网也许把这个概念极度放大了,但我不认为这是网络独有的。我完全同意你的看法,你这一点说得非常好。而我在这儿确实是在给这场对话预设立场,因为我把《纽约时报》描述得有点带着神话色彩
便签笔记
19:38
way the New York Times I'm talking about is a little bit of a fictional creature and it is you know for everybody who liked to get their news from The New York Times there were many times more people who preferred the New York Post or the daily news or the national Inquirer or or and those were well you would say those were the curators that that they that they chose and that's fine and and and you're right that is very much analogous to what's happening today where people can choose from among from a menu of choices that's fine um it's really what I believe in the end and it's why I can't neither in the book nor here can I come up with a really crisp way of saying what I think is the solution to this dilemma because I always end up with some platitudes you know we are individuals and and we're going to make individual choices and that means choosing the bloggers that we want to follow and the New York Times is just going to be you know there are there are some bloggers I follow who who are definitely ahead of the New York Times on some of
我说的那个《纽约时报》,多少有点像是个虚构出来的生物。要知道,在那些喜欢从《纽约时报》获取新闻的人之外,还有好几倍的人更愿意看《纽约邮报》、《每日新闻》或者《国家询问报》之类的,而你会说,那些就是他们自己挑选的内容筛选者,这没什么问题。而且你说得对,这跟今天发生的事情非常相似——人们可以从一份菜单式的选项里挑选,这也没问题。这其实正是我最终所相信的,也正因如此,无论是在书里还是在这儿,我都拿不出一个特别干脆利落的说法来讲清楚我认为这个两难困境的解法是什么,因为我最后总是落到一些老生常谈上:我们都是独立的个体,我们会做出各自的选择,这意味着去挑选我们想追随的博主,而《纽约时报》就只是……你知道,我关注的有些博主,在我感兴趣的某些领域上,确实是走在《纽约时报》前面的
便签笔记
06编辑即算法:相似 vs 公民所需
20:50
the areas that I'm interested in that's why I say again that that's a mythical convenience um but but for the main Google search engine let's say or for the Google News homepage if you aren't adjusting the results based on your knowledge of each individual user which of course is to some extent increasingly an option how do you how do you become all things to all people I think that's a a very difficult question that we're actually in the midst of answering but I'll let question step up before so hi um and to put my bias out front um I read the New York Times um and I hope that it doesn't go away um and I'll pay for it um to keep it not going away but and they thank you for that yeah and I come I also come from a family of uh journalist so um that maybe part of the reason but to your original question about Google News versus you know an editor um what I was always taught from family members who are editors was that an editor decides what's important for this community to know of the news that I've called and
在我感兴趣的那些领域上,所以我才要再说一遍,那只是一个神话式的方便说法。不过,就拿谷歌的主搜索引擎来说,或者谷歌新闻的首页来说,如果你不根据你对每一个具体用户的了解去调整结果——当然,这在某种程度上正越来越成为一个可选项——那你要怎么做到对所有人来说都是他们想要的那个东西呢?我觉得这是个非常难的问题,而我们其实正在回答它的过程中。不过我先让提问的人上来。你好。先把我的立场亮出来:我看《纽约时报》,我希望它不要消失,我也愿意花钱订阅,好让它不消失。——他们要谢谢你。是啊。而且我也出身于一个新闻从业者的家庭,所以这可能也是部分原因。但回到你最初那个关于谷歌新闻和编辑之间对比的问题,家里做编辑的人一直教我的是:编辑决定的是,在我筛选和收集到的新闻当中,哪些是这个社群需要知道的重要内容
便签笔记
22:02
and gathered um and so any particular paper you read has a particular belief system the editor is an algorithm if you will about what's important um the the thing that in my personal opinion not speaking as a Google representative that made Google important for me back in like 97 or 98 whenever I started using it was that other searchings were starting to have inorganic results where the wall between the business side and the journal journalistic side was split journalism being a loose reference to algorithmic search um so Google kept the ads out of the search results and that's still I think attention today because how do you keep an Enterprise going that needs money that driven by advertising and giving the user something they want which is not necessarily what the advertiser wants so there's that tension so I think the question is um Google is more like a card catalog that's very easy to use that helps you find things that are like the things that you tell it um whereas an editor is telling you what's important and so the algorithmic difference there is um how you decide what's important so for Google what's important is what's like what you gave
以及收集到的。所以你读的任何一份特定的报纸,都带着一套特定的信念体系,编辑本身就是一种关于“什么重要”的算法,如果你愿意这么说的话。以我个人的看法——不是以谷歌代表的身份说话——让谷歌在1997还是1998年、反正就是我开始用它的那会儿对我变得重要的原因是,别的搜索引擎当时开始出现非自然的(付费)结果,商业那一侧和新闻那一侧之间的那堵墙被打破了,这里的“新闻”是宽泛地指算法搜索。而谷歌把广告挡在了搜索结果之外,我觉得这在今天仍然是一种张力,因为你要怎么维持一家需要钱、靠广告驱动的企业运转,同时又给用户他们想要的东西,而那未必是广告主想要的东西,这里面就存在这种张力。所以我觉得问题在于:谷歌更像是一个非常好用的卡片目录,帮你找到跟你告诉它的东西相似的东西;而编辑则是在告诉你什么是重要的。所以两者的算法差异就在于,你如何判定什么是重要的。对谷歌来说,重要的就是跟你输入的内容相似的东西
便签笔记
23:10
me um as opposed to what do you need to know to exist in a democracy healthfully um so that's the trick and the problem with crowdsourcing the second half of your question is I think that crowdsourcing is very good for evening out the stupendous highs and the stupendous lows of individuals um so one editor will have stupendous highs and stupendous lows in their performance um hopefully more highs and average level than lows but other sources will help mitigate the lows so bloggers other newspapers are all ways to even out one variable signal that editor um but on the other hand not only did they take away the lows they take away the highs so you may not get the Brilliance of one editor's knowledge of a sphere which is also a problem with newspapers versus Google last piece is that um a newspaper isn't covering all of the domains that human endeavor happens in only at the highlights you don't get deep DNA science you don't get deep physics deep computer science in the New York Times because it's not generally applicable to all those communities but our society is specializing more and more so you need other sources like journals search engines scholarly search engines um that drive you towards what's important and those areas so that's part of the
而不是“你在一个民主社会里健康地生活需要知道什么”。所以这就是难点所在。至于你问题的后半部分,众包的问题在于,我认为众包非常擅长把个体那种极高和极低的表现拉平。一个编辑的表现会有非常高的时候,也有非常低的时候——但愿高点和平均水平多于低点——而其他来源可以帮忙削平那些低点,所以博主、其他报纸,都是把编辑这一个变量信号拉平的方式。但另一方面,它们不只是削掉了低点,也削掉了高点,所以你可能就享受不到某一个编辑在某个领域里那种卓越的洞见,这也是报纸相对于谷歌的一个问题。最后一点是,报纸并不会覆盖人类活动的所有领域,只覆盖亮点。你在《纽约时报》上看不到深入的DNA科学,看不到深入的物理学、深入的计算机科学,因为那些内容对所有这些群体来说并不具有普遍适用性。可我们的社会正变得越来越专业化,所以你需要别的来源,比如期刊、搜索引擎、学术搜索引擎,把你引向那些领域里重要的东西。所以这就构成了一部分
便签笔记
07猜测意图:搜索比「相似」复杂得多
24:30
difference between which source you choose what algorithm they use and what breadth you expect from it so I think you need a mixed model is the end of the day statement yeah I okay I think that was really interesting and and uh I'd like to home in on one of the many pieces of what you said and you might need to come back to the microphone because I want to ask I want to ask you more well you'll you judge uh the part about um Google's main search engine giving you back something that's like what you asked for is that how you put it it's it seems to me it's it's a lot more complicated than that and maybe it's a little more like what a newspaper does or a museum curator does than than you're admitting I mean if I ask for whatever if I if I type in drinking fountains you don't know I'm probably as you know very well somewhat inept at framing the question I actually think um a life skill that uh people have to develop is how to how to use Google how to how to frame the question to get what you want and and you presumably would like that to be less of a necessary skill right I mean you want people to just free associate and get what they want but that's impossible because you're what I might want what I might be looking for at any given moment it it could be something very specific
差别——你选择哪个来源、它们用什么算法、你期待从中得到多大的覆盖面。所以我觉得,最终的结论是你需要一种混合模式。是啊,我觉得这真的很有意思。我想聚焦在你刚才讲的众多内容中的一点上,你可能得再回到麦克风前,因为我还想再问你一些——好不好由你判断——就是关于谷歌主搜索引擎会返回给你“跟你所要求的相似的东西”这一点,你是这么说的吧?在我看来,这件事要复杂得多,而且也许比你承认的更接近报纸或者博物馆策展人所做的事。我是说,如果我搜索——随便什么——如果我输入“饮水机”,你并不知道……而且你也很清楚,我在把问题表述清楚这方面有点笨拙。我其实觉得,人们必须培养的一项生活技能,就是怎么使用谷歌,怎么把问题表述出来才能得到你想要的东西。而你们大概是希望这项技能变得不那么必要吧?我是说,你们希望人们只要自由联想就能得到他们想要的东西。但那是不可能的,因为我在任何一个特定时刻想要的、正在找的东西,可能非常具体
便签笔记
26:01
you know I might have a medical problem or I might have something I want to buy or I might just suddenly be curious about something I I know I'm I'm only saying what you already know um but but what do you mean it's like what you put in it's not you're not just mirroring um in some way the user's search phrase drinking fountains or horse races or um how do I do something I'm sure that must be a very common beginning of of Google searches even though it's none of mine I never start a search that way um you too need to get some Clues I mean part of what I'm saying is related to your final point about the New York Times not being a journal of physics the New York Times tries to cover physics I I wrote about physics when I was working there it happened to be my special area but you're right I had to write about it from from with the knowledge that there was a mythical New York Times reader who had a certain level of knowledge about physics and that was completely arbitrary you know um and that meant I had to leave stuff out and it meant I
你知道,我可能有个医疗问题,或者我有什么东西想买,或者我可能就是突然对某件事好奇起来。我知道,我说的都是你早就知道的事。可是,你说“它跟你输入的东西相似”是什么意思呢?你并不只是以某种方式把用户的搜索词——“饮水机”、“赛马”,或者“我该怎么做某件事”——原样映射回去吧。我相信“我该怎么……”肯定是谷歌搜索非常常见的开头方式,尽管我从来不这么搜,我从不那样开始一次搜索。你们也需要一些线索。我想说的一部分,其实跟你最后讲的那点有关,就是《纽约时报》不是一本物理学期刊。《纽约时报》确实会试着报道物理学,我在那儿工作的时候就写过物理,那恰好是我的专门领域。但你说得对,我必须在这样一种认知下去写:存在一个神话般的《纽约时报》读者,他对物理有某种程度的了解,而那个程度完全是任意设定的。这意味着我必须省略掉一些东西,也意味着我
便签笔记
08回音室与意外之喜
27:19
had to overex exlain some things for some people but doesn't Google have exactly that problem when somebody types something into into that search field you're describing the challenge that the entire company faces you said it's impossible to to know what we mean and yet uh I I think certainly anybody who works on the search engine themselves and please feel free to come back up and respond I just didn't want to leave a dead mic I'll have a quick response oh sure so yeah what I was basically alluding to was uh the concept in at least in the search sphere I've heard it called hopo which is that uh one of the main problems with recommender systems and definitely Google has a very complex algorithm that I certainly don't understand completely and um the the problem is given the general purpos of Google as opposed to the New York Times which is aimed at a general educated citizen right sixth grade level or above um so um looking for General awareness like Google may serve many purposes and presumably news should be targeted so the algorithms were similar to an editor of a paper um but the the thing I was alluding to basically was this concept of homophily that you you end up with an echo chamber as it's called in the
对某些人来说不得不把有些东西解释得过于详细。可是当有人在那个搜索框里输入内容的时候,谷歌难道不是正好面临着同样的问题吗?你描述的是整家公司都要面对的挑战。你说过要知道我们究竟是什么意思是不可能的,不过我觉得,任何真正做搜索引擎的人——请随时再上来回应——我只是不想让麦克风空着。我可以简短回应一下。哦当然。是这样,我刚才基本上是在暗指一个概念,至少在搜索这个领域里我听到过一个叫“同质性”的说法,就是推荐系统的主要问题之一——谷歌当然有一套非常复杂的算法,我肯定没法完全理解——问题在于,谷歌是通用型的,而《纽约时报》是面向受过一般教育的公民的,对吧,六年级阅读水平以上。所以,寻找一般性的认知的话,谷歌可以服务于很多目的,而新闻大概应该是有针对性的,所以那些算法就更像一份报纸的编辑。但我刚才主要暗指的其实就是这个“同质性”的概念,也就是你会陷入一个所谓的“回音室”,在
便签笔记
28:30
blogosphere where people only subscribe to the blogs of the people who agree with them and there's never any dissent or constructive Counterpoint um so you end up with basically idiocy in some sense right um and so that's one of the challenges with recommendations and search engines it's very hard to get Serendipity whereas an editor presumably is knowledgeable across the whole span and is interested in educating about the whole span of events that are important to the community they audience okay now this is another this is a really key issue and and I would argue that it's just as important again to the mythical New York Times as it is to Google the question of serendipity um in the olden days the mythical olden days you would read the whole paper and so I mean if you were so inclined of course not everybody did but some people did some people just turned the page the whole advertising model depended on that so even if you weren't particularly interested you never cared in your entire life about anything that happened in India there might have been something on page five this big that happened in India and because you were otherwise bored and you didn't have anything else fighting for
博客圈里,人们只订阅那些跟自己观点一致的人的博客,从来没有异议,也没有建设性的反面意见,于是你在某种意义上就落入了愚蠢。所以这是推荐系统和搜索引擎的挑战之一——很难制造出“意外的惊喜”;而编辑大概是在整个领域范围内都有见识的,并且有兴趣去就那些对社群、对他们的受众重要的整个事件范围进行教育。好,这又是另一个……这是个真正关键的问题,而且我要说,“意外发现”这个问题对那个神话中的《纽约时报》来说,跟对谷歌一样重要。在过去,在那个神话般的旧时代,你会把整份报纸读一遍——我是说,如果你有这个习惯的话,当然不是每个人都这样,但有些人是,有些人就是一页一页翻过去,整个广告模式都依赖于这一点。所以哪怕你并不特别感兴趣,哪怕你这辈子从来没关心过印度发生的任何事,第五版上可能就有那么一小块讲印度的东西,而因为你反正也无聊,也没有别的东西在争夺
便签笔记
29:42
your attention and because you trusted the New York Times you would read that and that was Serendipity to some extent that is lost with the modern online New York Times because nobody's reading the paper in order people tend as you just said to look for what they're already interested in if they want to read about what's happening in Japan they might you know spend a lot of time doing that and never see a word about some serendipitous thing that they would in with the technology constructed differently have been exposed to does that mean overall in general that people who are getting their news now as I am online and not so much from printed sources have a less serendipitous experience I'm not sure I'm really not I try to I try to figure it out by looking at my own day I mean I certainly am seeing things accidentally willy-nilly that I wouldn't have seen before and I'm able to follow paths you know the way you you couldn't follow a path in a printed newspaper and see something that strikes my fancy and actually go from
你的注意力,又因为你信任《纽约时报》,你就会去读它,那在某种程度上就是意外的收获。而这一点在今天的《纽约时报》网络版上多少丢失了,因为没人再按顺序读报纸了。就像你刚才说的,人们倾向于去找他们本来就感兴趣的东西,如果他们想读日本发生了什么,可能就会花很多时间在那上面,而永远看不到关于某件意外之事的只言片语——那件事如果技术架构设计得不一样,他们本来是会接触到的。那这是不是意味着,总体而言,像我这样现在主要从网上而不是从印刷品获取新闻的人,意外发现的体验变少了?我不确定,我真的不确定。我试着通过观察自己的一天来搞清楚这件事。我确实会偶然地、七零八落地看到一些以前不会看到的东西,而且我能顺着线索往下追——你在印刷报纸上是没法那样顺藤摸瓜的——看到某个正合我意的东西,然后真的从
便签笔记
09Madonna 测试:多样性与个性化
30:55
here to there I'm not sure that we've lost Serendipity and I and I don't know how that applies to Google but um yeah next question um yeah I just sort of wanted to follow up but I'm not on the search team so don't take anything I say as authoritative but certainly in Google search there are multiple signals that are considered you know people tend to think of it kind of as like well what's the one right answer but there are lots of things that are considered you know are the sources authoritative are the sources popular of the sources come from a a place that's been reliable in the past and you know one of my uh early tests this is decade ago but for whether a search engineer was doing a good job or not was I'd type in Madonna as a search and if all I got was the singer then that was a bad search engine say again if all I got was was pardon I didn't hear what you type in oh Madonna okay and if all I got was the singer then that was a bad search but if I got you know the singer plus the Madonna in plus you know the Christian Madonna or the black Madonna statue all in that page of results then I knew that the search engine was doing me a good job of doing a good job of trying to find you know what I might be
这里跳到那里。我不确定我们真的失去了意外发现,我也不知道这一点该怎么套用到谷歌身上。不过,好,下一个问题。嗯,我只是想接着说几句,但我不在搜索团队,所以别把我说的话当成权威说法。不过在谷歌搜索里,确实会考虑多种信号。你知道,人们往往会把它想成“唯一正确的答案是什么”,但其实要考虑的东西非常多:这些来源权威吗?这些来源受欢迎吗?这些来源是不是来自一个过去一直很可靠的地方?我早年——这是十年前了——用来判断一个搜索工程师做得好不好的一个测试是:我会输入“麦当娜”作为搜索词,如果返回给我的全都是那个歌手,那这就是个糟糕的搜索引擎。你再说一遍?如果返回的全是……不好意思,我没听清你输入的是什么。哦,麦当娜。好的。如果返回给我的全是那个歌手,那这次搜索就很糟;但如果我得到的是歌手,加上那个叫麦当娜的(其他事物),再加上基督教里的圣母像、黑色圣母雕像,全都出现在那一页结果里,那我就知道这个搜索引擎在努力找出我可能
便签笔记
32:16
looking for and would allow me to to continue on from that point and so I think there's a danger in looking it search as as finding the right answer and I think you know if you look at how the Google results page has changed you know you're getting um not only text pages but a section of images um you can filter by whether you want news by whether you want blogs that that return those results um or not and and I think that's that's a really positive direction well I'm you're raising a you're raising a really interesting issue so so if a thousand people type in the word Madonna some number of them are looking for the singer and couldn't care less about the Christian issues and some number of them have never even heard of the singer probably a smaller number I'm just guessing um but you don't know or do you that is you could construct you could set things up so that you did know so that so that you kept track of the user's history and previous searches and so forth and you might have um better than random information about whether they were looking for the the singer or for the the mother of Jesus um failing that you present 10 results the 10 results may or may not include everything that the thousand different people who have typed the word Madonna in want I presume without having asked
在找的东西方面做得不错,并且能让我从那个点继续往下探索。所以我觉得,把搜索看成是“找到那个正确答案”是有危险的。而且你看谷歌结果页的变化就知道,你现在得到的不只是文字页面,还有一块图片区域,你可以按是否要新闻、是否要博客来筛选返回的结果,我觉得这是一个非常积极的方向。嗯,你提出了一个非常有意思的问题。假设有一千个人输入“麦当娜”这个词,其中一部分人是在找那位歌手,根本不在乎基督教方面的东西;另一部分人可能压根没听说过那位歌手,人数大概更少,我只是猜测。但你并不知道谁是谁——或者你其实知道?也就是说,你是可以做到的,你可以把系统设置成让你知道,比如记录用户的历史记录和过往搜索等等,那样你就能得到比随机猜测更好的信息,来判断他们要找的是那位歌手还是耶稣的母亲。而在做不到这一点的情况下,你就呈现十条结果,这十条结果未必涵盖了那一千个输入“麦当娜”的人各自想要的所有东西,我猜是这样,虽然我没问过
便签笔记
33:44
anybody that that the results are constantly being influenced by some knowledge about what people click on but there's got to be a limit to how far you can follow the user unless they're using one of your toolbars or one of your operating systems I don't know I'm not asking um and it all twists back to to these other questions of serendipity because you don't want it to be perfect maybe it's more valuable for a user who types in Madonna thinking about the pop star to be confronted with some of the other variations of Madonna I don't know I'm beginning to get tangled up I think that's a very interesting point in terms of um at least part of the signal that we get from search point of view is in essence the popularity of a concept or popularity of a website so um that could certainly skew things in a certain way but I'm also thinking back to your comparison between like the role of Google and the role of like an editor for New York Times for instance I think one difference that I see is like in a way I can look at a search that someone conducts as almost like a conversation where I'm looking for let's say Madonna and I get back a set of results that maybe that's not what I expected or what
……任何人,搜索结果都在不断受到‘人们点击了什么’这类信息的影响,但你能追踪用户到什么程度总归是有限的,除非他们用了你们的工具栏或者你们的某个操作系统——我不知道,我也不是在追问。而这一切又绕回到关于‘意外收获’的那些问题上,因为你并不希望它太完美,也许对一个输入‘Madonna’、心里想的是那位流行歌星的用户来说,看到‘Madonna’的其他几种含义反而更有价值——我不知道,我自己都开始绕晕了。 我觉得这一点非常有意思。从搜索的角度看,我们得到的信号里至少有一部分,本质上就是某个概念的流行度、或者某个网站的流行度,所以那确实可能让结果朝某个方向偏。不过我也在想你刚才做的那个类比,比如把 Google 的角色和《纽约时报》编辑的角色相比。我觉得我看到的一个区别是,某种程度上我可以把一个人做的搜索几乎看成一场对话:比如我在找‘Madonna’,然后拿回一组结果,也许那不是我预期的、或者不是我
便签笔记
10实时对话、流感预测与真相仲裁
35:08
I wanted then I basically go okay well maybe I want to search for Madonna and something else MH um in a way that's almost like a real- Time conversation that's going on between the the the user and the search engine first it's like an editor you get the feedback through maybe letters from your readers but that's like after the fact and by the time like you get the letters like the feedback is kind of like you moved on to the next thing already so maybe that's a little bit of a difference right so you can you can keep track of what the what the user's next search is and then the search after that and put and put the information together presum some also like what you said about like the The Click is a signal right so if the user is clicking on something that's a signal that that gets back to us and go okay well we maybe doing a reasonable job in of rning the results yeah that's limits as you say but uh there are obviously also limit to the extent to which you're able to judge algorithmically whether sources are reliable or not I mean some of you have suggested that that's part of the equation too but um and it is with other crowdsourced information Enterprises like Wikipedia there is some process of zeroing in on what one
想要的,那我基本上就会说,好吧,那我可能要搜‘Madonna’加上别的什么词。嗯,某种意义上这几乎就是用户和搜索引擎之间实时进行的一场对话。而如果你是编辑,你得到的反馈可能来自读者来信,但那是事后的了,等你收到信的时候,那个反馈——你其实已经在做下一件事了。所以这可能算是一点区别,对吧?你可以追踪用户的下一次搜索,再下一次搜索,把这些信息拼在一起。 大概也包括你说的那个——点击是一种信号,对吧?如果用户点了某个东西,那就是一个反馈给我们的信号,我们就知道,好,也许我们在结果排序上做得还算合理。 是的,正如你说的,这是有局限的。但显然,你们用算法判断某个来源是否可靠,这方面的能力也是有限的。我是说,你们中有人提到过那也是这套体系的一部分。而在其他众包式的信息机构里,比如维基百科,确实存在某种逐步逼近的过程,逼近人们所
便签笔记
36:22
would consider the truth at least some kind of accuracy and yet um you know the number of people who think that Barack Obama was not born in the United States remains staggeringly large and and there is a significant web presence and um I think a live example of good crowd sourcing that I that I mentioned in the book is is um one I'm sure you're all familiar with because it's so often cited as Google's um ability to make accurate predictions about the spread of flu EP epidemics fast than the Centers for Disease Control based on what people are searching for and where um on the other hand if Google is swept Along by hoaxes or fads or misinformation I don't think that necessarily means that Google is not doing its job it may be that that's the model that you've chosen that you're not supposed to be arbit Arbiters of Truth to to some extent again news organizations um I believe are less and less inclined to be Arbiters of Truth and more um neutral reporters of what people think and what people argue about it's either way it's a mixed blessing yeah I
认为的真相,至少是某种程度的准确性。可即便如此,你知道,认为奥巴马不是在美国出生的人数量仍然大得惊人,而且在网上有相当可观的存在感。 我在书里提到过一个很好的众包正面例子,我相信你们都很熟悉,因为它被引用得太多了,就是 Google 能够根据人们搜索什么、在哪里搜索,比疾控中心更快地对流感疫情的传播做出准确预测。 另一方面,如果 Google 被恶作剧、跟风潮流或者错误信息裹挟着走,我并不认为这就一定意味着 Google 没做好本职工作。也许那正是你们选择的模式——你们不该在某种程度上充当真相的裁判者。同样地,新闻机构,我认为,也越来越不愿意做真相的裁判者,而更多是中立地报道人们怎么想、人们在争论什么。不管怎样,这都是喜忧参半的事。 是啊,我
便签笔记
37:50
think there's a a some some concepts of authority that I'd like to tease out and then see if we can isolate parts of this argument but I will say one thing which is uh if Google is a mirror of the desires of humanity uh it's it's in some sense our problem and in some sense not that you don't like what that reflection is looks like right and so we do want to be Arbiters of not just popularism but some semblance of a closer to authoritative truths but there's so many concepts of authority being tossed about here one is the authority of the actual author of The whatever whatever is to be read and the other is the authority of a curator to choose things that you should read whether or not you should whether or not you even knew you should be reading about them topics you didn't think about like India or or uh Concepts that you hadn't been exposed to because they're new and then there's the authority of the answer to the question you asked which is a search engines Authority in my opinion it's not necessarily the job of a search engine whom you've made a request to to be serendipitous if if if you if you search for Madonna and we start showing you things about maybe bands related to Madonna that you didn't
觉得‘权威’这个概念里有几层意思,我想把它们拆开,看看能不能把这场讨论的几个部分区分清楚。不过我先说一点:如果 Google 是人类欲望的一面镜子,那么某种意义上这是我们的问题,某种意义上又不是——如果你不喜欢镜子里照出来的样子的话。所以我们确实希望不只是做大众口味的裁判,而是做某种更接近权威真相的裁判。 但这里被抛来抛去的‘权威’概念实在太多了。一种是内容本身的作者的权威性;另一种是策展人的权威性,由他来挑选你该读什么——不管你该不该读,甚至不管你有没有意识到自己该读这些,比如你从没想过的话题,像印度,或者你从未接触过的、因为它是全新的概念;还有一种是‘对你所提问题的答案’的权威性,在我看来那才是搜索引擎的权威。 我认为,搜索引擎在你向它提出请求之后,未必有义务去制造意外惊喜。如果你搜‘Madonna’,我们却开始给你看一些跟 Madonna 有关的、你原本不
便签笔记
11作者如何判断可信来源
39:00
know about which is serendipitous that seems a bit of a stretch since you did ask us about Madonna it seems only right of us to give you maybe a variety of of different topics saying this is a general query could you explain which Madonna you were looking for with another refinement or by giving by giving us more information but that sense of authority I think the problem of authority that we are more concerned about is what flags do we use uh electronic or semantic or otherwise to to determine who has given the best answer on the web to the question the user's answering asking and these are the things you talk about people gaming through blogs or what have I don't think blogs really float up as much as people claim they do to the top of our results sometimes they do but you know just from your sty of information how do you determine when you did your research how do you how do you determine what is authoritative Source what's a source you should trust not an editor but the writer of the content itself what what are you looking at oh me who me please as a source of authority who I trust yeah well the I began as a newspaper reporter and so that's one there are various old-fashioned um paths to an answer to that question
知道的乐队,这算是意外惊喜——但这有点牵强,因为你问的就是 Madonna。我们似乎理应给你一些不同方向的选项,说这是个宽泛的查询,你能不能通过进一步限定、或者提供更多信息,说明你找的是哪个 Madonna。 但说到权威感,我觉得我们更关心的权威问题是:我们该用什么标志——电子的、语义的或者别的——来判断,网上谁对用户所问的问题给出了最好的答案。这也是你谈到的,人们通过博客等等来做手脚。我并不认为博客真的像人们说的那样经常冒到我们结果的顶部,有时候确实会,但你知道…… 从你对信息的研究来说,你做研究的时候是怎么判断的?你怎么确定什么是权威来源、什么是你应该信任的来源——不是编辑,而是内容的作者本人?你会看哪些东西? 哦,问我吗?我吗? 是的,请说,作为你信任的权威来源。 好吧,我最早是做报纸记者的,所以那算一条路。要回答这个问题,有各种老派的路子,
便签笔记
40:14
but one of the old-fashioned techniques is you go out you interview people and then you interview more people and um unfortunately you have a deadline and so you run out of time and newspapers are full of errors um newspapers also have copy editors with experience who can catch some of these errors then uh for me the time scales get longer because now I'm working on a magazine article or a book and I have more time but um you know these days there as as I'm sure you know lots of lots of different sorts of Institutions have absolute rules against trusting Wikipedia as a factual Source if you're a newspaper reporter and you make a mistake and you tell your editor well I got that from Wikipedia you're you're toast likewise if if you're um a college student and you turn in a paper I'm pretty sure at most places the rule is Wikipedia is not a trustworthy Source nevertheless it's also true that Wikipedia is incredibly reliable and Incredibly valuable I spent a lot of time I'm not ashamed to admit looking things up in w pedia when I was working on my book um I spent even more time using Google Books even while I was participating in the lawsuit against it call me a hypocrite if you want um everything in a book is not trustworthy and the great thing about the flood of
但其中一种老派的做法就是:你出去采访一些人,然后再采访更多人。不幸的是你有截稿期,所以时间总会用完,于是报纸上满是错误。当然报纸也有经验丰富的文字编辑,能抓出其中一些错误。 对我来说,后来时间尺度变长了,因为我开始写杂志文章或者写书,时间更充裕。但如今——我相信你们也知道——有各种各样的机构立了绝对的规矩:不许拿维基百科当事实来源。如果你是报纸记者,出了错,你跟编辑说‘我是从维基百科上看到的’,那你就完了。同样,如果你是大学生,交上一篇论文,我很确定大多数学校的规矩都是:维基百科不是可信来源。 尽管如此,维基百科同样极其可靠、极其有价值。我不怕承认,写书的时候我花了大量时间在维基百科上查东西。我花在 Google 图书上的时间甚至更多——尽管当时我还是那场针对它的诉讼的参与方之一,你要说我伪君子也行。 书里的东西也不是样样都可信。而这场信息洪流——‘洪流’是我书副标题的最后一个词——的好处在于,你可以自己当自己的策展人,你可以自行取舍,你必须运用判断力。这既是
便签笔记
12付费墙与信息洪流下的变现难题
41:47
information the flood being the last word of my subtitle is that you get to be your own curator you get to pick and choose and you have you have to exercise judgment it's um it's a challenge and it's a responsibility certainly so um I may horror of Horrors represent you know like a dying older generation you know at least at Google but um I am very much willing to pay for editorial expertise and I think a lot of other people are in various ways right so basically the New York Times does have to change its business model you know some people will only pay with their attention whereas other people will pay you know with money they probably won't pay for the paper subscription anymore I subscribed all my life until a few years ago and then I stopped you know I was a times premium subscriber and then they refunded my money I'm not quite sure why they did that um but so you know I guess my point is that there are certain you know just jewels of editorial expertise um in this country the times is one of them um and my gosh I hope they don't give up um I'm sure there's a way for them to monetize what they have um they may not know what it is yet and I hope they keep looking uh I hope you're right I'm not completely sure and there's a you know there's a technical issue that that I
挑战,当然也是一种责任。 所以,说来可怕,我大概代表了一个正在消亡的老一代人,至少在 Google 这里是这样。但我非常愿意为编辑方面的专业能力付费,而且我觉得还有很多人也愿意,以各种方式。所以基本上,《纽约时报》确实得改变它的商业模式——有些人只愿意用注意力来付费,另一些人则愿意用真金白银来付费。他们大概不会再订阅纸质报纸了,我订了一辈子,几年前停了。你知道,我曾是《时报》的高级订户,后来他们把钱退给我了,我也不太清楚为什么。 总之,我想说的是,这个国家有一些堪称编辑专业能力的瑰宝,《时报》就是其中之一。天哪,我真希望他们别放弃。我相信他们一定有办法把手里的东西变现,也许他们还没找到那个办法,但我希望他们继续找。 我希望你是对的,我不太确定。而且有一个技术上的问题,我
便签笔记
43:10
don't see an ideal solution for or I'm not sure there's even a good solution and that is once it's online it's so easily transmittable a lot of the ability of people to make money as inter information providers in the old world had to do with inconvenience paper you know libraries um The New York Times wants the Huffington Post to be an amplifier and a source of readers but it doesn't want the Huff Huffington Post to to then eat its lunch by being by providing enough of the material that people don't need to then you know pay the times for it it's the same it's the same for me you can read you don't have to pay $10 subsidized um be still my heart you don't have to you don't have to pay $10 you can read a substantial amount of the of the book free right now online I don't know exactly how much I hope it's not a very satisfying reading experience I hope that every so often you're interrupted with a a page that you can't read but but but there's no question that some number of people who in the old world would have bought the
看不到理想的解决办法,甚至不确定有没有一个还不错的办法:那就是一旦东西上了网,它就太容易被传播了。在旧世界里,信息提供者能赚到钱,很大程度上靠的是‘不方便’——纸张、图书馆之类的。 《纽约时报》希望《赫芬顿邮报》成为一个放大器、一个读者来源,但它不希望《赫芬顿邮报》反过来抢走它的饭碗,也就是提供了足够多的内容,以至于人们不再需要为此向《时报》付费。 对我来说也是一样。你可以读——你不用付那 10 块钱——我的心啊——你不用付 10 块钱,现在就可以在网上免费读到这本书相当大的一部分。我不知道具体是多少,我希望那不是一种很舒服的阅读体验,我希望每隔一阵子你就会被一个读不了的页面打断。但毫无疑问,在旧世界里本来会为了读关于某件事的 10 页内容而买下这本
便签笔记
44:26
book just to read 10 pages about something will no longer have to um the times can they they've got my they'll get my money whatever kind of firewall they'll put up and I see they have yours but but they will not be able to prevent people from transmitting all or part of some of those articles and once they do that I'm not sure if there's enough left I I don't I don't know so two things uh one is I think there are people who will read 10 pages of your book who wouldn't have read any of it before because it's so easy to read those 10 pages online and of those there are a fraction who will then go out and buy the whole thing and not for 10 bucks either um actually I'm an author too um and my my work is widely pirated you know you can download my entire books but you know I don't I don't begrudge that and and actually people are still buying the books so you know I think that um I'm basically seem to be more optimistic than you are about this the other thing is I visit the New York Times website multiple times a day and I look at the entire front page and I kind of wish it was more like the front page of the physical paper um because you know I thought that was an interesting selection of news um so you know I'm I'm just I'm not willing to
书的那些人,现在有一部分不必再买了。 《时报》嘛,他们已经拿到、也还会拿到我的钱,不管他们立什么样的付费墙,我看他们也拿到了你的钱。但他们没法阻止人们把那些文章全部或部分地转发出去,而一旦人们这么做了,我不确定还剩下多少东西。我不知道。 那我说两点。第一,我认为会有人读你书里的 10 页,而这些人在以前根本不会读任何一页,因为在网上读那 10 页太容易了;而在这些人里,会有一部分人后来去把整本书买下来,而且花的还不止 10 块钱。其实我自己也是作者,我的作品被盗版得很厉害,你可以下载到我全部的书。但你知道,我并不因此心怀怨恨,而且实际上人们还是在买书。所以我觉得,在这件事上我好像基本上比你乐观。 第二点是,我一天要访问《纽约时报》网站好多次,我会把整个首页都看一遍,而且我还挺希望它更像纸质报纸的头版,因为我觉得那是一种很有意思的新闻取舍。所以你知道,我就是不愿意
便签笔记
45:48
give up on this idea that you know we have people we trust and it isn't just trust by the way it's also people whose Aesthetics we agree with you know there are other Publications that I read simply because you know I I this will be interesting you know I may not even believe a lot of what's in here but I know that you know these guys have an interesting slant and I'd like to read it um yes it's easy to transmit the information yes it's all just bits but on the other hand you know the people who are just transmitting the bits don't care about providing the same experience that the people who originate them do um by the way the people who really get screwed here of course are musicians because they you know it really is a bit forbit copy um you know musicians and filmmakers and so forth that's a harder problem you know you can tell us your thoughts on that if you like yeah no actually I want to tell you my thoughts on another thing you just said which is yes um I too wish that the online front page of the New York Times was more like the physical front page and I I described how that's structured in a what's supposed to be a meaningful way on the other hand I consult the online page 10 times a day
放弃这个想法:我们有一些值得信任的人。而且顺带说一句,这不只是信任的问题,也关乎审美上的合拍——有些刊物我读它,纯粹是因为我觉得这会很有意思,我甚至可能并不相信里面很多内容,但我知道这些人的视角很有意思,我愿意读。 是的,信息很容易传播;是的,一切都只是比特。但另一方面,那些只是搬运比特的人,并不在乎提供和原创者一样的体验。 顺便说一句,在这件事上真正吃亏的当然是音乐人,因为那真的是逐比特的复制。音乐人、电影人之类的,那是个更难的问题。你要是愿意,也可以谈谈你对这个的看法。 嗯,其实我想说的是你刚才说的另一件事。是的,我也希望《纽约时报》的网络首页更像纸质头版,我在书里描述过那种版面是如何以一种本应富有含义的方式组织起来的。但另一方面,我一天要看十次网络版首页,
便签笔记
13熵、麦克斯韦妖与气候变化的政治化
47:01
and I don't expect it to be the same every time I don't expect it to be the same when I come back so there's a there's a dynamic information problem that nobody has a good solution to it's not enough for them to know that I've already been there or even what I've read comments on seeing it online like the front page okay one comment on seeing it like the front page um all subscribers at least I get the Sunday paper um they have version you can download called replica which is the front page even though it's kind of inconvenient to remember to go get the replica but if you're craving that it's there okay sorry you were you were next uh yeah so um I apologize I'm going to change topic a little bit you were talking before about Shannon uh and about the possibility that information is maybe this broader concept that maybe has some philosophical implication or you know I'm in particular interested in the physical implication you know so Shannon defines this quantity entropy and it turns out to be exactly the same thing that's studied in thermodynamics yeah and I've never sort of gotten a good explanation of what that's all about and I wonder if you could sort of oh great have anything to say about that yes you're not going to get a good explanation from me right this second because it's really hard but it is in
我并不指望它每次都一样,我不指望我回去的时候它还是老样子。所以这里有一个动态信息的问题,没人有好办法解决——光知道我来过、甚至知道我读过什么,还不够。 关于把网页做得像纸质头版这件事有个评论吗?好,一个关于像头版那样呈现的评论。 所有订户——至少我订的是周日版——都有一个可以下载的版本叫 replica(复刻版),就是原样的头版,虽然要记得去把那个复刻版调出来有点麻烦,但如果你很想要,它是有的。 好的,抱歉,下一位是你。 好的,我先道个歉,我要稍微换个话题。你之前谈到香农,谈到信息也许是一个更宽泛的概念,可能带有某种哲学意涵。你知道,我特别感兴趣的是它的物理意涵——香农定义了‘熵’这个量,而它恰好和热力学里研究的东西完全一样。 是的。 我从来没听到过一个像样的解释来说明这到底是怎么回事,我想问你能不能——哦太好了——对此说点什么? 是的,此时此刻你不会从我这儿得到一个像样的解释,因为这实在太难了。不过它在
便签笔记
48:16
the book there's a whole chapter about entropy in the book um and Maxwell's demon and it's just uh it's too much for this it's too much for my little brain in this in this context I'll I'll just tell this joke that um that um what you said is exactly correct by the way that when Shannon worked out the mathematics of um probabilities of a message they it turned out that the mathematics absolutely mirrored uh what already existed in thermodynamics and he started using the word entropy practically as a synonym for for information and uh what they there was the the urban legend that went around at Bell Labs was that John Von noyon had had advised Shannon to use the word entropy because then no one would know what he meant very nice we have time for uh just one more question thank you so much for your for your talk um I have a question from a different angle from what we've been talking about we've been talking a lot about how people get information how it's changed the problems with it maybe the opportunities around that but what
书里,书里有一整章讲熵,还有麦克斯韦妖。这个话题在这种场合太庞大了,对我这颗小脑袋来说也太庞大了。 我只讲个笑话吧。顺便说一句,你说的完全正确:当香农推导出消息概率的那套数学时,结果发现那套数学和热力学中已有的东西完全吻合,于是他开始把‘熵’这个词几乎当作‘信息’的同义词来用。而贝尔实验室里流传的都市传说是,冯·诺伊曼建议香农用‘熵’这个词,因为这样就没人知道他在说什么了。 很妙。我们还有时间再问最后一个问题。 非常感谢你的演讲。我想从一个和刚才不太一样的角度提个问题。我们谈了很多人们如何获取信息、它发生了什么变化、其中的问题,也许还有随之而来的机会。但我
便签笔记
49:33
I'd like your thoughts on is given that things have changed and we are there is a demand for perhaps more information but people are getting these filtered information that's their own filter and all these things that we talked about um how do you think the world will go should go could go to um more effectively uh con um work with the evolving media world to uh have an most effective dialogue around politically sensitive data Rich issues and let me throw out there a specific issue that you can focus on as an example which would be change so there's a lot of issues in terms of open and transparency this has to do with a lot of digital activity and emails and so forth there's a lot of political issues people can find what they want there's science issues silid scientists aren't used to this world what do you what do you sort of think about where we are today in terms
想听听你的看法:既然事情已经变了,人们对更多信息也许有需求,但他们拿到的却是被过滤过的信息——他们自己的过滤器,以及我们谈到的这一切——那你觉得,世界会怎样、应该怎样、可以怎样更有效地与这个不断演变的媒体环境打交道,从而围绕政治上敏感、数据密集的议题展开最有效的对话? 我给你抛一个具体的议题作为例子,你可以聚焦在上面,那就是气候变化。围绕开放和透明有很多议题,牵涉到大量的数字活动、邮件之类的东西;有很多政治议题,人们总能找到他们想要的东西;还有科学层面的议题,科学家们并不习惯这样的世界。你怎么看我们今天在
便签笔记
50:46
of the scientific communication to the world and the policy around it yeah I I guess I I can sum this up by saying things are kind of a and uh we're in what I think may be a period of adjustment I hope it's a period of adjustment I'm not sure what are the easy answers to to these questions climate change is you know such a canonical example in a way 20 years ago when I was a science reporter at the New York Times I remember I'm I wrote a few things about particular small pieces of the climate change puzzle and it never it didn't occur to me at that time that there was going to be the least bit of controversy about it it was all pretty straightforward it was not politicized it has become so intensely politicized there are so many people who even to talk about whether you believe in climate change you know like believing in God strikes me as kind of weird it should have been a technical kind of knowledge a non-political kind of knowledge we all know why that's not the case it's because of the influence of money on so much on so many of our channels of
科学向公众传播、以及相关政策方面所处的位置? 是啊,我大概可以这样总结:情况有点糟,而且我们正处在一个我认为可能是调整期的阶段——我希望那是个调整期。我不确定这些问题有什么容易的答案。 气候变化在某种意义上真是个太典型的例子。二十年前,我在《纽约时报》做科学记者的时候,我记得我写过几篇关于气候变化拼图中某些具体小片段的报道,当时我压根没想到这件事会有丝毫的争议性,一切都相当直截了当,它并没有被政治化。而如今它被政治化到了极其严重的程度,有太多人甚至会谈论‘你是否相信气候变化’——就像相信上帝一样,我觉得这挺怪的。它本该是一种技术性的知识,一种非政治的知识。 我们都知道为什么不是这样:因为金钱对我们如此之多的信息渠道的
便签笔记
52:09
information in the case of climate change I think it's very clear um there are Industries the oil and gas industry that spend a lot of money on um pseudo think tanks they pay researchers in my opinion to lie there are other people who unconsciously and without evil intent find their views distorted by uh misinformation that that comes from these sources and you know don't I won't even get started on Fox news but um but I'm I guess what I'm stumbling around trying to say is that there is no no Real Purity in our complicated world in in um sources of information and and there's just ever more of a challenge for us consumers of information to watch out and be skeptical and the same thing applies to information about pharmaceuticals and and and to more obviously political subjects and it applies in more subtle ways to things that are even less political or connected to money um this company so far I think has been doing a pretty good job and people should remain nervous about it you're a profit-making Enterprise too um try to
影响。在气候变化这件事上,我觉得非常明显:有些行业——石油和天然气行业——在所谓的‘伪智库’上砸了很多钱,在我看来他们是花钱雇研究者去撒谎。还有另一些人,他们并无恶意,也并非有意为之,却在这些来源散播的错误信息影响下形成了扭曲的看法。至于福克斯新闻,我就不开这个头了。 我磕磕绊绊想说的大概是:在我们这个复杂的世界里,信息来源没有真正的纯净可言,而对我们这些信息消费者来说,保持警觉和怀疑只会变得越来越有挑战。这同样适用于关于药品的信息,适用于更明显带政治色彩的话题,也以更微妙的方式适用于那些政治性更弱、或与金钱关联更少的事情。 到目前为止,我觉得这家公司做得相当不错,而人们也应该继续对它保持一点不安。你们也是一家营利性企业。嗯,努力
便签笔记
53:35
do no evil thank you all [Applause]
不作恶吧。谢谢大家。[掌声]
便签笔记
视频总结 · 一句话概括与核心要点

一句话概括

詹姆斯·格雷克以《信息简史》为引子,提出"通往 Google 的道路始于 1948 年香农的信息论",并借用"神话化的《纽约时报》"这一参照物,与 Google 员工就搜索引擎的本质是过滤而非搜索、算法策展与人工编辑的优劣、权威与偶遇性(serendipity)、付费墙与信息洪流下的个人判断责任等问题展开了一场对谈。

核心要点

  • 信息论是现代世界的底层结构,而非仅是技术基础。 格雷克认为通往 Google 的起点是 1948 年克劳德·香农在《贝尔系统技术期刊》上发表的两篇论文(后结集为《通信的数学理论》)。他最初在写《混沌》时从混沌科学家那里听说信息论,惊讶于"信息"这种模糊概念竟能被数学化;如今信息早已不只是新闻和文字,而是一切可数字化、可存储、可搜索的东西。
  • "信息时代"远不止 50 年,全部人类史都是信息时代。 据《牛津英语词典》,"信息时代"一词 1960 年才开始流行,但印刷术、电报、电话乃至字母表本身都是信息技术。因此书的结构是"从中间(1948)开始,回溯到起点,再向前推进"。书中关于 Wikipedia 的篇幅反而多于 Google,因为 Google 已无处不在到无需专门介绍。
  • Google 的核心功能是过滤,而不是搜索。 格雷克借《纽约时报》的旧格言"读者付费买的是我们没登的新闻"来类比:报纸头版的严格结构(右上头条、左上次头条、折线上下、栏宽)体现了编辑引以为傲的"新闻判断"——绝大多数凶杀案不值一提,但斯卡斯代尔减肥医生被情人谋杀就是大新闻。Google 同样是在为用户滤掉不需要的信息。
  • 用户对算法策展的态度是矛盾的。 人们一方面庆幸摆脱了"权威、反民主"的编辑机构(伯克利听众还为此高喊"Democracy Now"),另一方面又抱怨"我怎么知道什么是真的、什么对我有意义"。格雷克直言 Google News 的策展质量"是垃圾",他仍离不开《纽约时报》;但对 Google 搜索本身高度肯定——"你们占据主导是有原因的"。
  • 金钱与注意力争夺正在扭曲信息渠道。 注意力已成为有限资源,人们为排名而"裁剪"内容而非单纯作弊。同样地,《纽约时报》如今也根据"被转发最多的文章"实时调整版面——格雷克认为旧时编辑"不知道也不在乎读者想读什么"的傲慢反而产生了更有价值的结果,因为实时反馈得到的常常只是查理·辛这类八卦。Google 员工反驳:在互联网之前,小报式的民粹就已是报纸大发行量的基础,网络只是放大了它。
  • 编辑与算法各有取舍:编辑是"一个算法",众包抹平了高峰与低谷。 一位 Google 员工提出:编辑决定"这个社区需要知道什么",Google 则更像一个好用的卡片目录,返回"与你输入相似的东西";众包能削平单个编辑的失误,但也削掉了其洞见的高峰;报纸不覆盖深度物理、DNA 科学等专业领域,故社会需要"混合模型"。格雷克回应:Google 并非简单镜像查询词,用户实际上必须学会"如何提问"这项生活技能,而 Google 也和报纸一样面临"预设读者知识水平"的问题。
  • 同质化(homophily)与回音室是推荐系统的根本难题,偶遇性的得失尚无定论。 旧报纸让人在无聊时顺手读到印度的一则小新闻,这是被动的偶遇;在线阅读则各取所需。格雷克以"Madonna"查询为例讨论:一千个用户中大多找歌手,少数找圣母,好的搜索引擎应同时呈现多义结果;Google 员工认为搜索是一场实时对话(改写查询、点击即信号),这是编辑靠读者来信无法企及的反馈速度。但格雷克坦言,自己在网上跟随链接的漫游未必比翻报纸更少偶遇,"我真的不确定"。
  • 权威有多重含义,Google 并不打算成为"真理仲裁者"。 Google 员工区分了三种权威:内容作者的权威、策展人的权威、搜索引擎给出答案的权威,并称"如果 Google 是人类欲望的镜子,那不喜欢镜中影像某种程度上是我们自己的问题"。格雷克举例:Google 流感趋势能比 CDC 更快预测疫情,是众包的成功;但"奥巴马非美国出生"的信念仍有庞大网络存在——被骗局与谣言裹挟,不一定意味着 Google 失职,而是它选择了不做仲裁者的模式。
  • 信息洪流的代价是每个人必须自己做策展人。 记者和大学生都被明令禁止引用 Wikipedia,但格雷克坦承写书时大量使用 Wikipedia 和 Google Books(甚至在参与对 Google 的版权诉讼期间);书里的内容也并非全部可信。"洪流"的好处是你可以自己挑选,代价是"这是一种挑战,也是一种责任"。
  • 付费墙与在线传播的矛盾没有理想解。 旧世界里信息提供者的收入很大程度上依赖"不便利"(纸张、图书馆);《纽约时报》希望 Huffington Post 做放大器而非吃掉它的午餐;格雷克自己的书也能在线免费读到相当篇幅。一位自称作品被广泛盗版的 Google 员工则更乐观:免费读 10 页的人中会有一部分去买整本,且愿为编辑专业能力付费的人依然存在。

结论与值得注意的细节

  • 格雷克对"如何在信息洪流中辨别真伪"没有给出干脆的答案,最终只能落到"我们是个体,要自己选择信任哪些博主和媒体"这类他自嘲为"陈词滥调"的说法;他承认自己描述的《纽约时报》是"略带神话色彩的虚构生物"。
  • 关于香农熵与热力学熵:一位听众问两者为何数学形式完全一致,格雷克表示书中有一整章讲熵与麦克斯韦妖,但当场只讲了一个贝尔实验室的都市传说——冯·诺伊曼建议香农用"熵"这个词,因为"这样就没人知道你在说什么"。
  • 谈到气候变化时,格雷克回忆 20 年前作为《纽约时报》科学记者报道此题时完全没料到会有争议;如今它被强烈政治化,原因是石油天然气行业资助的"伪智库"和被他称为"拿钱撒谎"的研究者,连"你是否相信气候变化"这种问法本身(像信不信上帝)都显得怪异。他的结论是:信息渠道没有纯净可言,消费者只能保持怀疑,这同样适用于制药信息乃至更隐蔽的领域。
  • 一个实用细节:《纽约时报》订户可下载名为 Replica 的版本,还原纸质版头版排布,回应了双方都"怀念纸质头版结构"的感叹。
  • 结语颇具张力:格雷克说 Google"目前做得还不错,但人们应保持警惕——你们也是营利企业,努力'不作恶'吧"。
  • 时代背景值得注意:本次对谈发生于 2011 年前后(《纽约时报》付费墙刚上线、Google Books 诉讼、"奥巴马出生地"阴谋论盛行),当时讨论的回音室、注意力经济、算法与人工策展之争在今天仍然完全适用。

这期还没有生成核心句型(制作精读 PDF 时会一并生成)。

词汇精讲 · 110 · 按出现顺序
revelatory /ˈrevələtɔːri/ adj. 0:00
揭示性的,有启发性的
Chronicle /ˈkrɑːnɪkl/ n. 0:00
编年史,纪事
grappling with phr. 0:00
努力应对,与……搏斗(难题)
dense /dens/ adj. 0:00
(内容)艰深的,晦涩的
quaint /kweɪnt/ adj. 1:25
古雅的,老派而有趣的(此处指纸质书)
ubiquitous /juːˈbɪkwɪtəs/ adj. 1:25
无处不在的
premise /ˈpremɪs/ n. 1:25
前提,基本假设
full-blown /ˌfʊl ˈbloʊn/ adj. 1:25
完全成熟的,全面展开的
taken aback phr. 2:59
吃惊,愕然
amorphous /əˈmɔːrfəs/ adj. 2:59
无定形的,模糊的
lie underneath phr. 4:27
潜藏于……之下,构成……的底层
tossed around phr. 4:27
(说法)被随意使用、到处流传
Enterprise /ˈentərpraɪz/ n. 5:49
企业;事业
tangle /ˈtæŋɡl/ n. 5:49
乱糟糟的一团
finite /ˈfaɪnaɪt/ adj. 7:06
有限的
conceit /kənˈsiːt/ n. 7:06
自负;(此处)一厢情愿的设定
above the fold phr. 7:06
报纸折线以上(最显眼位置);引申指网页首屏
news judgment n. 7:06
新闻判断力(新闻业术语)
spurned /spɜːrnd/ adj. 8:34
被拒绝的,被抛弃的
matron /ˈmeɪtrən/ n. 8:34
(略带贬义)中年贵妇
mythical /ˈmɪθɪkl/ adj. 8:34
神话般的,虚构的
bombarding /bɑːmˈbɑːrdɪŋ/ v. 8:34
狂轰滥炸,连续不断地灌输
dichotomy /daɪˈkɑːtəmi/ n. 10:06
二分法,对立
ambivalence /æmˈbɪvələns/ n. 10:06
矛盾心理
tyranny /ˈtɪrəni/ n. 10:06
暴政,专横
snobs /snɑːbz/ n. 10:06
势利眼,自命清高者
crowdsourced /ˈkraʊdsɔːrst/ adj. 11:29
众包的
obsolete /ˌɑːbsəˈliːt/ adj. 11:29
过时的
curation /kjʊˈreɪʃn/ n. 12:43
策展;内容筛选与编排
in lie of phr. 12:43
(应为 in light of)鉴于,考虑到
crap /kræp/ n. 14:06
(口语)垃圾,废物
make a go of it phr. 14:06
把事情做成,获得成功
rehearse /rɪˈhɜːrs/ v. 14:06
(此处)复述,一一列举
flatter /ˈflætər/ v. 14:06
奉承
game the system phr. 15:29
钻系统空子
tailor /ˈteɪlər/ v. 15:29
量身定制,调整以适应
analogously /əˈnæləɡəsli/ adv. 15:29
类似地
arrogance /ˈærəɡəns/ n. 16:59
傲慢
planting a seed phr. 16:59
埋下种子,为日后发展铺垫
pushed around phr. 16:59
被摆布,被牵着走
tabloids /ˈtæblɔɪdz/ n. 18:27
小报(以八卦煽情为主)
purveyor /pərˈveɪər/ n. 18:27
供应者,兜售者
circulation /ˌsɜːrkjəˈleɪʃn/ n. 18:27
(报刊)发行量
hyperextended /ˌhaɪpərɪkˈstendɪd/ v. 18:27
过度延伸(原为医学术语「过伸」)
loading the conversation phr. 18:27
有倾向性地引导谈话
crisp /krɪsp/ adj. 19:38
(表述)干脆利落的
platitudes /ˈplætɪtuːdz/ n. 19:38
陈词滥调
all things to all people phr. 20:50
面面俱到,讨好所有人
in the midst of phr. 20:50
正处于……之中
inorganic /ˌɪnɔːrˈɡænɪk/ adj. 22:02
非自然的(此处指付费干预的搜索结果)
stupendous /stuːˈpendəs/ adj. 23:10
惊人的,极大的
evening out phr. 23:10
抹平,使均匀
mitigate /ˈmɪtɪɡeɪt/ v. 23:10
缓解,减轻
home in on phr. 24:30
聚焦于,锁定
inept /ɪˈnept/ adj. 24:30
笨拙的,不擅长的
free associate phr. 24:30
自由联想
arbitrary /ˈɑːrbətreri/ adj. 26:01
任意的,武断的
alluding to phr. 27:19
暗指,间接提及
recommender systems n. 27:19
推荐系统
homophily /hoʊˈmɑːfəli/ n. 27:19
同质相吸(社会网络术语)
echo chamber n. 27:19
回音室(只听到同类观点的环境)
blogosphere /ˈblɑːɡəsfɪr/ n. 28:30
博客圈
dissent /dɪˈsent/ n. 28:30
异议
Counterpoint /ˈkaʊntərpɔɪnt/ n. 28:30
对立观点;(音乐)对位
Serendipity /ˌserənˈdɪpəti/ n. 28:30
意外发现,机缘巧合
so inclined phr. 28:30
有此倾向的
willy-nilly /ˌwɪli ˈnɪli/ adv. 29:42
随意地,杂乱无章地
strikes my fancy phr. 29:42
合我心意,引起我兴趣
authoritative /əˈθɔːrəteɪtɪv/ adj. 30:55
权威的
couldn't care less phr. 32:16
毫不在乎
failing that phr. 32:16
若做不到那一点,退而求其次
skew /skjuː/ v. 33:44
使偏斜,扭曲
tangled up phr. 33:44
绕进去了,陷入混乱
after the fact phr. 35:08
事后
zeroing in on phr. 35:08
逐步聚焦、收敛到
staggeringly /ˈstæɡərɪŋli/ adv. 36:22
惊人地
epidemics /ˌepɪˈdemɪks/ n. 36:22
流行病,疫情
hoaxes /ˈhoʊksɪz/ n. 36:22
骗局,恶作剧谣言
Arbiters /ˈɑːrbɪtərz/ n. 36:22
仲裁者,裁判
mixed blessing n. 36:22
喜忧参半的事
tease out phr. 37:50
梳理出,抽丝剥茧地厘清
semblance /ˈsembləns/ n. 37:50
表象,近似
a bit of a stretch phr. 39:00
有点牵强
refinement /rɪˈfaɪnmənt/ n. 39:00
(查询的)细化,精炼
float up phr. 39:00
浮上来(此处指排名上升)
you're toast phr. 40:14
(口语)你完蛋了
copy editors n. 40:14
文字编辑,校对编辑
hypocrite /ˈhɪpəkrɪt/ n. 40:14
伪君子
horror of Horrors phr. 41:47
(戏谑)天哪,最糟糕的是
monetize /ˈmɑːnətaɪz/ v. 41:47
变现,货币化
jewels /ˈdʒuːəlz/ n. 41:47
珍宝,瑰宝
transmittable /trænsˈmɪtəbl/ adj. 43:10
可传播的
eat its lunch phr. 43:10
抢走某人的饭碗,夺走其市场
be still my heart phr. 43:10
(戏谑感叹)我的心啊,哎呀
firewall /ˈfaɪərwɔːl/ n. 44:26
(此处指)付费墙
pirated /ˈpaɪrətɪd/ adj. 44:26
被盗版的
begrudge /bɪˈɡrʌdʒ/ v. 44:26
吝惜,对……心怀不满
Aesthetics /esˈθetɪks/ n. 45:48
审美,美学趣味
slant /slænt/ n. 45:48
(报道的)倾向,视角
get screwed phr. 45:48
(俚语)被坑,吃大亏
craving /ˈkreɪvɪŋ/ v. 47:01
渴望
thermodynamics /ˌθɜːrmoʊdaɪˈnæmɪks/ n. 47:01
热力学
entropy /ˈentrəpi/ n. 48:16
urban legend n. 48:16
都市传说
data Rich adj. 49:33
数据密集的
canonical /kəˈnɑːnɪkl/ adj. 50:46
典型的,经典的
politicized /pəˈlɪtɪsaɪzd/ adj. 50:46
被政治化的
pseudo think tanks n. 52:09
伪智库
stumbling around phr. 52:09
磕磕绊绊地(表达)
Purity /ˈpjʊrəti/ n. 52:09
纯净,纯粹

这期还没有生成自测题(制作精读 PDF 时会一并生成)。

精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.087The Debate Over “Understanding” in AI’s Large Language Models 下一期 · NO.089 →Cailin O'Connor: The Misinformation Age
订阅苏菲周报 每周一封:本周入库的精读、一个值得带走的问题、一条苏菲按。免费,随时退订。
免费 · 每周一封 · 一键退订
苏菲拉底 THE SOPHIE LAB · ASK THE BEST MINDS THE BIG QUESTIONS 内容仅供学习 · thesophielab.com