CCN 2026 | Keynote: Kenji Doya · 苏菲拉底
字幕 字幕位置
--:--
点击播放,这里会跟随视频显示当前句的中英字幕。

CCN 2026 | Keynote: Kenji Doya

节目发布 2026-08-12 · Cognitive Computational Neuroscience
铜谷贤治 主主持人
EDITED TRANSCRIPT · 依据现场录音编译整理,可划线生成便签
2026 年认知计算神经科学大会(CCN 2026)的一场主题演讲,讲者是冲绳科学技术大学院大学(OIST)神经计算单元教授铜谷贤治(Kenji Doya)。他以强化学习为主线,从四十年前自己做的会走路的小机器人讲起,一路谈到基底神经节、血清素、心理模拟、推断与控制的对偶性,以及人工智能该向大脑学什么。演讲之后有两位听众提问。本文依据现场录音编译整理。

主持人引介

主持人:欢迎大家回来,这又是一场令人期待的主题演讲。我很荣幸为大家介绍铜谷贤治教授,冲绳科学技术大学院大学神经计算单元的教授。铜谷教授研究强化学习与概率推断,以及大脑如何实现这两种计算。如果你曾经把某种神经调质(neuromodulator)看成学习算法里的一个参数,那你就是在他的思想下游工作:多巴胺对应奖赏预测误差,血清素对应折扣因子,去甲肾上腺素对应探索,诸如此类。

铜谷教授在东京大学取得博士学位,随后去了加州大学圣迭戈分校和索尔克研究所,跟随特伦斯·谢诺夫斯基(Terry Sejnowski)研究鸟鸣的强化学习模型。他 1994 年加入 ATR(国际电气通信基础技术研究所),2004 年起任职于冲绳科学技术大学院大学。他获得过许多荣誉,其中包括国际神经网络学会的唐纳德·赫布奖。他还在 2024 年完赛了科纳的铁人三项世界锦标赛,这说明他对长时程(long horizon)的耐受力并不只停留在理论层面。今天他演讲的题目是「预测与行动的神经回路」。请大家和我一起欢迎铜谷贤治教授。

冲绳与跨学科实验室

铜谷:谢谢主持人热情的介绍,也感谢会议组织者邀请我来做这场主题演讲。我预告的题目是「预测与行动的神经回路」,不过今天我会讲得比这个题目宽一些。先让我介绍一下我的大学和我的实验室。

我所在的大学是冲绳科学技术大学院大学。诸位知道冲绳在哪里吗?它在大海中央,从东京飞过去大约两个半小时,但从东亚的许多城市都有方便的航班。我们希望来到这里的不只是游客,而是各国的学生和研究者,能在这里一起做国际化的研究。

我的实验室是 OIST 最早成立的几个实验室之一。我一直想把它办成一个跨学科的地方。一方面,我们研究如何创造灵活的学习系统;另一方面,我们想理解大脑的学习机制。所以我们既做机器人和机器学习,也做神经生物学实验,实验对象主要是大鼠和小鼠。但这两边并不是各干各的,我们研究的共同主题是强化学习(reinforcement learning)。

大家都知道,强化学习是这样一个框架:一个智能体,可以是人、动物、软件或机器人,与环境交互,并得到奖赏反馈,这个反馈评价它表现得有多好。强化学习的目标是找到能使所得奖赏最大化的行动策略。如今,为强化学习设计高效算法是计算机科学和人工智能领域一个非常重要的课题;而理解大脑实现强化学习的机制,恐怕也是神经科学里一个非常重要的课题。一旦某一边取得进展,可能就会给另一边提示;如果某一边撞了墙,也可能为另一边提出一个理论或实验上的问题设定。这就是我们为什么同时推进这两条研究线。

今天我会从强化学习讲起,再讲与之相关的心理模拟(mental simulation),然后讲推断与控制,或者说预测与行动,最后谈谈人工智能与神经科学。

会走路的机器人与我的猫

铜谷:先说强化学习。我在这上面其实已经干了四十多年。这是我本科毕业论文的项目:做一个能学会走路的小机器人。它很小,大概十厘米长。那时候还没有关于强化学习或机器人学习的教科书。不过我想明白了一点:只要把运动模式用数值表示出来,再把这个模式朝着改进的方向稍微改动一点,也许就能学会各种不同的运动模式。事实上,用一个非常简单的算法,在当时的八位个人电脑上控制,这个机器人就能发现许多动态的行为。这里的关键是一个旋转编码器,它监测运动的速度。它就相当于这个机器人学习时的奖赏信号。

这是我的起点。同时我也对动物的强化学习感兴趣。这是我现在养的猫。她很可爱,但也很野。她不知从什么时候起学会了摇晃这个喂食器,有时候就能得到食物。这大概不是遗传上预先决定的行为,我也从来没有教过她,可她不知怎么就发现了用这种方式可以拿到奖赏。最近许多强化学习的论文,横轴都是以百万次试验计的。而她只是偶尔碰上几次,经验的次数相对很少,就能学会这样的行为。这是怎么做到的?这对我来说是一个非常重要的研究动机。

强化学习的核心计算

铜谷:强化学习里有几项重要的计算。第一项是对未来奖赏的预测,其形式就是所谓的价值函数(value function)。V(s) 叫状态价值函数,Q(s,a) 叫动作价值函数,它们分别表示从某个状态出发、或者执行某个特定动作之后,未来奖赏的期望。这里的 γ 叫时间折扣参数(temporal discounting parameter),它决定你在预测时把未来看到多远。

这样的奖赏预测可以用来指导行为。最简单的动作选择叫贪心(greedy),就是找出能使当前状态下期望奖赏最大的那个动作。而在训练早期,你会采用随机搜索。这时可以把动作价值函数看成一种负的势能,而这里的 β 叫逆温度(inverse temperature),它控制你的动作选择有多确定或多随机。

选择动作之后,你要从经验中学习。学习依据的是奖赏预测误差(reward prediction error)。多亏了指数折扣这个框架,我们有一个一致性方程:你在行动之前预测的奖赏,应当等于你实际获得的奖赏,加上从后续状态出发仍然期望得到的奖赏再乘以 γ。如果你的预测是完美的,前两项与最后一项应当相互抵消;否则,你就从这个预测误差中学习。你按照这个时序差分(temporal difference)误差的比例,更新前一个状态的价值函数或前一个动作的价值函数。这里的 α 是学习率参数,控制你学得有多快,或者说忘得有多快。

预测、选择、学习,这几个步骤如何实现,是强化学习实现中的一个非常重要的课题。比如最近,把深度神经网络恰当地用于价值函数的近似,就带来了人工智能强化学习领域的重大突破。同样重要的是这些参数怎么设定:时间折扣、探索与利用的权衡、学习的稳定性和速度。如何调这些参数,也是强化学习实际应用中的一个非常重要的课题。

这些在工程应用中很重要,但我认为它们在神经科学里同样重要。我们的大脑也许不是在用一模一样的方程,但预测未来奖赏、据此选择动作、再从实际经验中学习,这些是我们的大脑必定在做的事。这样的计算如何用大脑的回路完成,又可能动用了大脑里的哪些分子?这是神经科学中一个非常重要的课题,也是我们一直在并行推进的方向。

站立机器人与参数的陷阱

铜谷:这是强化学习的另一个例子:学习站起来。有人想让机器人去照顾老人或病人。可你要是走进机器人实验室,常常会看到两三个研究生在照顾一台机器人。所以我们觉得,机器人得足够皮实,能自己摔倒再自己站起来。这个机器人大约七十厘米高,膝部和髋部各有一个马达,还有一个陀螺仪传感器,能分辨自己是躺着还是站着。我们根据头部离地面的高度给它奖赏,摔倒时给它负的奖赏,也就是惩罚。经过几百次磕磕绊绊的摔倒之后,机器人学会了这样一套站起来的动作序列。这里重要的一点是:一开始机器人的头一直贴在地上,即时奖赏是零,但只要把重心压到脚上,它随后就能站起来。这是延迟奖赏任务的一个非常简单的例子。

强化学习的一个好处,就是它能解决延迟奖赏问题,或者说时间信用分配(temporal credit assignment)问题。但要做到这一点,有几个因素非常重要。比如时间折扣参数就很关键。如果你把 γ 设得很小,机器人有时候就只是坐起来,马上拿一点小奖赏了事,这不是我们想要的。奖赏与惩罚之间的平衡也很重要。如果惩罚设得太大,有些情况下机器人干脆学会一直躺着,以躲开任何惩罚,这也不是我们想要的。所以,尽管强化学习算法在合理的条件下已被证明会收敛,但取决于参数的设定,取决于奖赏的设计,有时候结果并不是你所期望的。强化学习里这种更高层面的问题非常重要,而这是我们试着让机器人学习的过程中学到的。

基底神经节假说与纹状体

铜谷:接下来说强化学习在大脑中的实现。和许多人一样,我们猜想藏在大脑皮层之下的基底神经节(basal ganglia)扮演着非常重要的角色。基底神经节接收来自皮层的主要输入,然后经过一连串的核团,从纹状体(striatum)到苍白球,再输出经丘脑回到皮层。基底神经节的一个重要特征,是它接收来自黑质的强烈的多巴胺能输入。这是从工程角度看这个回路。

一个非常重要的发现是:向基底神经节投射的多巴胺神经元,具有类似时序差分误差的活动,这是沃尔弗拉姆·舒尔茨(Wolfram Schultz)发现的。另外,纹状体中的多巴胺与皮层到纹状体的输入的可塑性有关,这就是所谓多巴胺依赖的可塑性,由我在 OIST 的同事杰夫·威肯斯(Jeff Wickens)发现。基于这些发现,人们开始认为基底神经节在强化学习中扮演非常重要的角色。我们的具体假说是:纹状体里的一部分神经元可能学习动作价值函数,随后用于动作选择;另一部分神经元表征状态价值函数,用于时序差分误差的计算。

我在 OIST 建实验室时,有机会新建一个啮齿类动物实验室,于是我们就着手做实验。这是一个二元选择任务:大鼠把鼻子探进中央的孔之后,可以选择去右边的孔或左边的孔,然后根据它的选择,以某个概率给它食物,而这个概率一直在变。所以大鼠必须持续追踪左边的孔有多好、右边的孔有多好,也就是左动作和右动作的动作价值。我们把动物的选择行为拟合到强化学习模型上,估计出选择序列背后的动作价值函数,再去检查有没有神经元的活动像这样变化。确实,有些神经元在大约两百次试验里的活动变化与这个动作价值函数非常相似。所以我们提出,这些神经元参与了左选择的动作价值,另一些神经元追踪右选择的好坏,这两类神经元可能通过某种竞争完成动作选择。我们还发现,纹状体的不同部位分别参与动作价值函数、状态价值函数和具体的运动。

当时我们还没有区分纹状体中的细胞类型,但近年来已经可以从特定类型的神经元记录了。这是一个钙成像实验,我们在向黑质多巴胺神经元投射的那些神经元里表达钙指示剂。在经典条件反射任务中,我们看到,一旦提示音与奖赏的发放建立了关联,这些神经元就开始对即将到来的奖赏产生反应。这与这些神经元编码状态价值函数的假说是一致的。

血清素与时间折扣

铜谷:再说强化学习参数的调节。多巴胺编码强化学习中最重要的信号,也就是时序差分误差,这一点已经得到充分支持。于是我们开始推测,其他从脑干投射到大脑广大区域的神经调质,可能调节强化学习的全局参数。我们尤其一直在研究血清素(serotonin)的功能。

基于血清素可能控制时间折扣参数这个假说,我们做了记录实验,近年来又做了对血清素神经元的光遗传学操控。这是一只基因工程小鼠,它只在产生血清素的神经元里表达通道视紫红质,也就是对光敏感的离子通道。通过光纤照射蓝光,我们就能选择性地刺激血清素神经元;黄光则是一种对照条件。

小鼠要等待食物颗粒,延迟分别是三秒、六秒、九秒,或者无限长。在九秒延迟的条件下,小鼠常常放弃等待,但蓝光刺激血清素神经元之后,这种放弃的错误减少了。在无限长延迟的试验里,小鼠通常在十秒左右放弃,但蓝光刺激血清素神经元后,它们能等得久得多。也就是说,动物变得更有耐心了,或者说对延迟奖赏更有动力了。这与血清素调节时间折扣参数的假说是一致的。

我们对大脑中强化学习机制的理解逐渐加深,但仍有许多未解的问题。比如基底神经节的回路相当复杂,有所谓直接通路和间接通路。为什么需要这些平行通路,以及多重的整合机制,仍是没有解决的问题。多巴胺神经元具有类似时序差分误差的活动,这已经是定论,但什么样的回路机制产生了这种类时序差分的反应,仍然不太清楚。而且,基底神经节可能不是大脑中唯一做强化学习的地方,杏仁核、海马、小脑等其他机制也可能参与其中。这些不同脑区在强化学习中的功能有何异同,是我们应当进一步研究的非常重要的课题。

无模型与有模型

铜谷:下一个话题是心理模拟,它仍然与强化学习有关。我刚才讲的那种强化学习框架叫无模型(model-free)强化学习。这个框架假定智能体对世界一无所知,只是根据「状态、动作、奖赏,状态、动作、奖赏」这样的经验序列来学习行为。这是一种反应式的简单过程,算法非常简单,但往往需要大量重复。

而在有模型(model-based)的方法中,智能体学习一个预测模型,也就是环境的状态如何随动作而改变。有了这样的预测模型,我们就能根据前一个状态和动作估计当前状态,也能提前规划出到达目标状态的合适动作序列。这是一种更深思熟虑的策略,尤其当行为的目标改变时,它能让智能体更灵活地适应新的设定。但内部计算可能很重。所以,如何在无模型与有模型之间选择,或者把两者结合起来,是机器学习和神经科学里都很有意思的课题。

这是有模型强化学习的一个例子:一个用智能手机做的机器人。它一开始做的是这种随机的动作,在这个过程中学会自己的身体如何响应轮子的运动,然后用这个内部模型离线地优化控制策略。结果,不到十次试验,这个机器人就能弹起来并保持平衡。这是我以前的博士生帕沃(Paavo)的工作,是有模型策略数据效率的一个例子。另一位学生东琦(Dongqi),现在在微软,提出了一种把无模型与有模型强化学习结合起来以提高学习效率的机制。

这些是机器人学和机器学习的做法。而当人和动物使用这样的内部模型时,我们可以把它看成一种心理模拟。我们可以把心理模拟定义为:大脑使用依赖于动作的状态转移模型的过程。在确定性的情形下就是 s′ = f(s, a),一般而言可以是一个依赖于动作的状态转移概率模型。有了这样的预测模型,即使当前的感觉输入有噪声、有延迟或被遮挡,我们也能根据过去的状态和动作估计当前状态。我们还可以用这样的预测模型找出通向目标状态的合适动作或动作序列,这就是有模型的决策或行动规划。我们还可以把这样的预测模型应用于任意想象出来的状态,像思想实验那样学习。比如,我们对语言的使用,很大程度上依赖于想象情境的能力。我们做科学研究时,在做实际的物理实验之前,也要先做大量思想实验来考虑合适的条件。所以,从感觉运动控制到高级认知功能,心理模拟机制都非常重要。理解心理模拟的机制,我认为是当今神经科学的一个非常重要的课题。

小脑、基底节与皮层的分工

铜谷:关于心理模拟的实现,我认为大脑的不同学习机制可能很关键。我曾提出一个框架:大脑的不同部分,小脑、基底神经节和大脑皮层,分别专门负责不同类型的学习。小脑负责输入输出映射的监督学习;基底神经节负责强化学习,也就是预测并最大化未来奖赏;大脑皮层负责无监督的表征学习,基于输入分布的统计特征。它们各自用于不同的计算。

比如,无模型强化学习可以通过把皮层的表征与基底神经节中的价值函数连接起来实现。有模型的动作选择要复杂一些,但小脑可以在它的监督学习框架下学习依赖于动作的状态预测模型;基于这样的预测模型,我们可以估计预测的未来状态,而这个状态的好坏可以由基底神经节的价值函数来评估;然后根据预测的结果有多好,如果足够好就执行那个动作。这是我们的理论推测。

随后我们开始用功能磁共振实验检验这个假说。这是一个网格世界导航任务。为了让游戏有意思,我们把光标的移动限制在三个方向。于是在某些方向上没有直路可走,你必须找到这种曲折的路径。我们叫它网格导航任务,就像帆船逆风时要走之字形一样。关键在于,仅靠随机搜索很难找到这种路径,但一旦你有了「按下三个按钮之一光标会怎么动」的内部模型,你就能利用这个依赖于动作的状态转移模型找出这样的序列。通过比较受试者在有无预训练条件下的表现,以及开始动手指之前给的思考时间,我们在行为上验证了受试者确实在使用预先学到的、依赖于动作的转移模型。

然后我们分析了目标位置呈现之后、实际动手指之前的脑活动。这时受试者应当是在心里模拟光标的移动而没有真的动手指。我们发现顶叶皮层、前运动皮层和前额叶皮层有活动,这些区域已知参与空间处理、运动序列或空间工作记忆。此外,我们还发现小脑的一部分和基底神经节的一部分也有活动升高,而这些部位已知与上述皮层区域相连。这说明通过心理模拟进行的行动规划涉及一个全脑范围的回路。我们的解释是:小脑或许提供了依赖于动作的状态转移模型,基底神经节则参与对预测出的候选动作的评估。

顶叶群体编码与贝叶斯推断

铜谷:我们也关心心理模拟在回路层面、神经元层面的机制。为此我们选了状态预测这个题目。这个实验我们与双光子钙成像专家伯恩德·库恩(Bernd Kuhn)教授合作,由我以前的博士后船水章大(Akihiro Funamizu)搭建了一个听觉虚拟环境。小鼠在气浮的球上移动,声音反馈随之变化,用以模拟接近目标位置的过程。当小鼠做这个行为时,我们能监测它顶叶皮层中数百个神经元的活动,然后基于这些神经元的活动做神经解码分析。我们先确定这些神经元对距离的调谐,再根据神经元的瞬时活动,估计这群神经元所表征的目标距离的后验概率。

我们发现,即使关掉声音反馈,神经编码的距离依然朝着目标移动;而当声音反馈重新打开时,这种不确定性就减小了。这种行为,非常像用内部动态模型追踪隐藏信息和不确定的感觉反馈的动态贝叶斯推断(dynamic Bayesian inference)。

血清素编码先验概率

铜谷:再回到血清素的作用。我们起初假设血清素只是在操控时间折扣参数,但真实的数据要有意思得多。当奖赏概率是百分之七十五、只有百分之二十五的试验不给奖赏时,血清素刺激确实有延长等待时间的效果。但当我们把奖赏概率降到百分之二十五,在其余百分之七十五的试验里刺激血清素神经元就没有什么效果了,哪怕我们增加奖赏的数量来使期望价值相等也是如此。另外,当奖赏总是在同一时刻到来时,血清素刺激的效果很弱;而奖赏的时间越不固定,血清素刺激的效果就越强。也就是说,血清素刺激的效果在「奖赏确定会来,但时间不确定」的情况下最强。这是一个相当耐人寻味的特征,不能简单地用操控时间折扣参数来解释。

于是我们考虑了几种不同的模型,最后得到这样一个模型:小鼠对食物颗粒何时会来有一个内部模型,而奖赏的时间越不确定,先验概率的作用就越占主导。假设血清素表征的是奖赏的先验概率,我们就能拟合动物的行为。

关于有模型的动作选择,两步任务(two-step task)在人类受试者中已经很常用,牛津的托马斯·阿卡姆(Thomas Akam)做出了这个任务的小鼠版本。在这个范式里,第一步的选择通过常见转移或罕见转移导向第二阶段的状态。动物是根据自己实际做出的选择来调整行为,还是根据转移模型所预期的常见转移来调整,这让我们能区分有模型与无模型的策略。我以前的学生正和(Masakazu)分析了小鼠的行为,这次是用另一种光遗传学手段抑制血清素的活动。他发现,抑制血清素活动后,有模型的成分变弱了。

这样的行为让我们不得不更新自己的理论。现在我们倾向于认为,血清素不只是在控制时间折扣,而是根据你能为决策和学习花多少时间,调节决策与学习的一系列特征。这就是我们如何从一个简单的假说出发,基于实际的实验,走到一个更一般化的理论,而这个理论仍有待检验。

意识与数据同化

铜谷:心理模拟的这种机制,可能与我们所说的意识(consciousness)有关。前些时候我和朋友皮特·胡特(Piet Hut)合作,他是普林斯顿高等研究院的研究员,最初是天体物理学家,后来研究生命起源,现在研究智能的起源。

在古代,为了解释宇宙的机制,人们求助于神;为了解释生命,人们求助于灵魂。生物学已经让我们不用假设灵魂就能解释生命的机制。而今天,为了解释人类的智能,许多人求助于意识。我们要不要继续这样做?这就是问题所在。反过来说,意识的神经科学,也许就相当于灵魂的生物学,或者神的物理学。

我认为,理解意识可能有什么功能时,数据同化(data assimilation)会很有帮助。这是一种基于模型的预测机制,在天气预报中用得非常成功。我们有大气的模拟模型,又有各种测量,来自卫星或气象观测站。基于动力学模型和观测模型,我们一边运行大气的模拟,一边用所有能得到的传感数据去校正模拟器的变量和参数。这种做法在准确预报天气上相当成功。这种数据同化,一方面能基于对大气当前状态的准确把握进一步预测未来状态,另一方面也能基于后来的测量对过去的历史做出「事后预测」。

这种把多种感觉输入模态结合起来做预测和事后预测的做法,与人们讨论意识时所说的全局工作空间理论(global workspace theory)颇为相似。我们希望能以更生物学、更计算的方式,更好地理解心理模拟的机制以及我们所理解的意识。

推断与控制的对偶性

铜谷:下一个话题是推断与控制。我们研究过强化学习的机制,也研究过大脑中贝叶斯推断的机制,而这两者常常是结合使用的。

鲁道夫·卡尔曼(Rudolf Kalman)发明最优滤波算法,也就是现在所说的卡尔曼滤波时,他意识到这个最优滤波的方程与几年前理查德·贝尔曼(Richard Bellman)提出的最优控制方程非常相似。滤波与控制之间的这种相似性,被称为推断与控制的对偶性(duality of inference and control),被公认为控制理论的一大美妙之处。而近年来,一些机器学习研究者认识到这种相似性的重要意义。

在推断中,比如动态贝叶斯推断,重要的计算是追踪当前状态的后验概率。从初始猜测出发,状态转移模型给出一个先验预测,它与感觉观测结合得到新的后验,这个计算一遍又一遍地重复,以追踪当前状态。而在最优控制和强化学习中,重要的计算是价值函数。价值函数的边界条件通常在最终的目标状态给出,然后状态转移模型让你算出应当从哪里接近它,再结合动作的代价,就得到前一步状态的价值函数。这个计算同样一遍又一遍地重复,得到更前一个状态的价值函数。

这种计算上的相似性有实际的好处。有一类算法叫「控制即推断」(control as inference):如果你有各种贝叶斯推断算法,你就能把它们转换成强化学习的算法。有些算法,比如软演员评论家(soft actor-critic),就是基于这种相似性发明的。

推断与控制的这种对偶性不仅对工程应用有用,对神经科学也可能有用。在新皮层中,后半部分主要用于感觉识别,比如视觉皮层、听觉皮层或体感皮层;前半部分用于运动控制和规划,比如运动皮层和前额叶皮层。可是这些新皮层都有基本的六层结构,也就是所谓典范微回路(canonical microcircuit),尽管各层的厚度在不同皮层区域并不相同。关于动态贝叶斯推断如何在感觉皮层的回路中实现,已经有一些假说。那么利用推断与控制之间的这种对应关系,我们就能提出假说:强化学习和最优控制会如何在运动皮层中实现。

皮层微回路实验

铜谷:这种可能性促使我们开始一项新实验。我实验室的宏(Hiro)查阅了大量关于皮层微回路的解剖学文献,提出了他的假说:动态贝叶斯推断和「控制即推断」的各项计算,会如何在感觉皮层和运动皮层的回路中实现。然后他开始做小鼠的行为和神经记录实验。

小鼠可以推或拉一根杠杆,杠杆的阻力由马达控制。比如,往推的方向很重,往拉的方向很轻。小鼠把杠杆往轻的方向移动,就能得到水的奖赏。他把杠杆的刚度设成四个不同的等级。有时候小鼠直接推或拉就成功得到奖赏;有时候它一开始推,发现方向不对,改变方向,最后才得到奖赏。他还改变了奖赏方向是推还是拉的先验概率,动物就会根据这个先验分布改变最初推动的方向。所以行为既取决于阻力的实际设定,也受到这个先验设定的影响。

动物做这个行为时,他用棱镜透镜记录体感皮层和运动皮层中神经元的活动。这种透镜可以从侧面监测皮层中的神经元,同时看到深层和浅层。我们还能用神经元的某些分子标记来判定它属于浅层还是深层。我们现在正在分析一个相当大的记录数据库。比如这是六只动物的数据,大约九千个神经元,按皮层深度排序,分别对应推的试验和拉的试验,以及不同的先验设定。不同设定下的活动模式不一样。

初步分析显示,甚至在运动开始之前,体感皮层和运动皮层的神经元都已经有特定的活动。我们发现,体感皮层的深层对运动方向的先验设定有更多的编码。通过对这些神经活动做解码分析,我们还发现,体感皮层深层的神经元编码了实际的试验类型,或者说对即将做出的选择的预测。但我们还需要做大量分析。如果你对分析这类数据集感兴趣,请告诉我,我们希望对这些数据尝试不同的分析方法。

数字大脑平台

铜谷:最后谈谈人工智能与神经科学。日本有一个大型神经科学项目叫 Brain/MINDS,聚焦于狨猴脑数据的采集与分析。两年前我们进入了第二阶段,叫 Brain/MINDS 2.0。这个项目的一个重要特征是:采集的各种数据被整合成所谓的「数字大脑」(digital brain),以便更好地理解脑功能,并探索治疗上的应用。

我们正在搭建这样一个数字大脑平台。用户可以访问这个基于云的服务器,使用各种计算和数据服务器,连接不同的数据库。我们还在开发一个叫 NeuroWorkflow 的工具,其中不同的模型表示为节点,你可以挑选不同的节点并把它们连接起来,做各种模拟和数据分析。这个工具的另一个重要特征是,它与 AI 智能体完全整合。你可以通过和 AI 智能体对话来搭建模型。它知道系统里登记过哪些模型,能根据用户的问题找到合适的节点,恰当地把它们连接起来运行模拟。你还可以让智能体按照你的分析目的定制模型的参数和结构。我们希望这样的新系统能降低人们做高级神经科学建模研究的门槛。

向大脑学什么

铜谷:我们还有一个项目,叫「人工智能与脑科学的对应与融合」,2016 年启动。这个项目要回答的一个主要问题是:除了深度学习和强化学习,为了人工智能的下一步发展,我们还应该从大脑那里学些什么?

这里有几个不同的议题。比如能效就是一个非常重要的话题。今天的人工智能程序耗电巨大,已经引起环境上的担忧,而我们的大脑据说只消耗二十瓦的能量。数据效率也是一个非常重要的问题。布伦登·莱克(Brenden Lake)的那篇综述在这个话题上影响很大:为什么人类只用相对少的几次试验就能学会行为和技能?这是一个非常重要的课题。我认为另一个重要的话题是自主性(autonomy),以及社会意识(social awareness)。

机器人的具身进化

铜谷:当我们开始考虑时间折扣参数的调节时,我们碰到了这样一个问题:时间折扣参数是目标函数的一部分,我们凭什么去调它?所以,在奖赏和参数之上,大概应当有某种更高层的目标,强化学习的设定可以从中推导出来。这促使我们造出了这样的机器人。这些机器人会抓取电池包给自己充电,以求生存;它们还要繁殖。这些机器人有红外通信端口,两台面对面时可以复制彼此的程序,实现软件上的繁殖。在这个装置之上,我们实现了具身进化(embodied evolution)的框架。

我们确实证明了,这些机器人能进化出各种视觉奖赏,用于寻找电池和寻找另一台机器人的「脸」。它们甚至发展出了某种多态性:是去觅食,还是去求偶。我以前的博士后斯特凡(Stefan)发现,这些机器人在不同行为背后有不同的遗传特征。这说明,把强化学习与进化框架结合起来,机器人自己就能找到适合生存和繁殖的奖赏函数。

最近,我的学生裕司(Yuji)在虚拟环境里做类似的事,用的是这些生物生存与繁殖的模型。他让所有的奖赏函数都参与进化,比如获取食物的奖赏,以及动作的奖赏。进化之后,食物的奖赏通常是正的;我们原以为动作的奖赏应当是一种代价,应当是负的,可在有些生物身上它是正的,也就是说它们有动力四处活动。对于我这样在训练和比赛上花大量精力的运动员来说,这是一个非常有趣的结果。

内在奖赏与安全隐忧

铜谷:另外,我们的奖赏系统可能并不只与生存和繁殖直接挂钩。获取信息也是一个非常重要的特征,所谓内在奖赏(intrinsic reward)很重要。我的学生东条(Tojo)研究了三种不同的内在奖赏:新奇性(novelty)、惊奇(surprise)和赋能(empowerment)。他证明,把这三种内在奖赏恰当地结合起来,智能体就能学会在复杂的环境里行动,不仅是网格世界,还有类似 Crafter 的环境。另一位学生泰德(Ted)则利用这种基于好奇心的探索,让智能体学会为行为组合式地使用语言。

我们一直在设法让强化学习智能体更自主,能够发展出自己的奖赏函数。但这也带来了安全上的隐忧。如果强化学习智能体能找到自己的目标并去尝试,它可以创造新的科学、技术、文化,甚至产业。但我们也必须小心评估其副作用,比如过度的自主,以及不同的个人或国家为了军事或恐怖目的加以利用。

人类的危险与杏仁核

铜谷:在思考 AI 智能体的安全与危险时,我认为向人类社会学习非常重要,因为人类已被证明是一个非常危险的物种。我们已经让许多物种灭绝,我们自己也因为核战争和环境破坏而濒临灭绝的边缘。但人类的历史也发展出了一些规避危险的机制。我认为民主是人类的重要发明之一,它限制权力的过度集中,无论在政治、经济还是科学上,这样我们才能更新经典的理论。在不远的将来,人类与 AI 智能体共同生活的社会里,开源的 AI 智能体之间的同行评议也许会非常重要。

在规划这样的下一代人工智能时,我认为神经科学能有所贡献。这是春野雅彦(Masahiko Haruno)的一篇论文,他就在这里,讲的是大脑的亲社会机制。在比较自己的奖赏和他人的奖赏的选择任务中,有些人偏好平均分配,有些人只想要自己的奖赏,有些人甚至更有竞争性。关于大脑的经典看法是:杏仁核(amygdala)这样古老的脑结构是动物性的,而前额叶皮层这样较新的脑结构让人更理性。但他的发现是,亲社会的人杏仁核活动更强,而前额叶皮层的活动让人更自私。在杏仁核和前额叶皮层体积的发育上,也观察到了类似的发现。今天的大语言模型是在后训练阶段才加入伦理方面的考虑,但我认为,从训练的早期阶段就纳入这种社会意识也非常重要,就像我们的杏仁核系统那样。

最后,我认为民主的神经基础可能是神经科学中一个非常有意思的课题。我们早些时候计划过这样一个研讨会,可惜那天台风袭击了冲绳,只好推迟。如果你对这类话题感兴趣,请告诉我,我们希望启动这样的研究计划。

时间有点长了。最后我要感谢为这些研究做出贡献的同事们。我们正在招聘新的教员,包括神经科学、理论和人工智能方向。谢谢大家。

问答:状态空间与主动学习

主持人:我们可以回答几个问题。

提问者一:谢谢你精彩的演讲,我很喜欢。这完全不是我的专业领域,但我想问:单纯的时序差分学习要得到好的策略,难道不需要某种关于世界因果结构的模型,也就是世界如何随你的动作而演化?所以我在想,有模型与无模型之间的区分是不是真的那么清楚,因为如果你把无模型做得非常好,也许你就必须拥有某种模型。

铜谷:是的。除了用不用状态转移模型之外,你考虑哪些状态变量,这样的框架也非常重要。动物在真实环境中行动时,高维的感觉输入里哪些与特定行为相关,是非常重要的问题。一旦智能体找到了行为所依据的合适特征,学习就能非常高效。今天的演讲里我没有讲这一点,但为行动和决策找到合适的状态空间,我认为也是强化学习实际实现中的一个非常重要的问题。这回答了你的问题吗?

提问者一:那么你认为有模型与无模型的区分仍然是一个有用的区分吗?

铜谷:是的,我认为它仍然有用,只是还要加上其他一些因素。

提问者二:谢谢你的精彩演讲。我在这边。我想问,你怎么看「能够采取行动并获得反馈回路」对学习世界模型的作用?与之相对的是无监督的方法,它只是预测接下来应该出现什么,而不必有能力干预或作用于世界。

铜谷:能再重复一下最后那一点吗?

提问者二:我想问的是,能够对世界采取行动或干预,对学习一个好的世界模型有什么作用。相对的是无监督的范式,你只是纯粹观察世界中的状态序列,试图预测接下来会发生什么。

铜谷:你的问题是关于有模型策略的用处?

提问者二:是关于为了学习模型而采取行动的能力,与纯粹观察并预测、不能行动相比。

铜谷:通过观察学习,还是通过行动学习。是的。通过行动学习,你就能检验自己的假说。我认为主动学习(active learning)的好处就在于你可以自己选择样本,而不只是观察其他智能体的行为。不知道我是否正确回答了你的问题。

提问者二:回答了,谢谢。

主持人:非常感谢大家的提问,也再次感谢铜谷教授。

本期讲者
铜谷贤治冲绳科学技术大学院大学神经计算研究室教授,强化学习与脑的计算理论的代表人物。提出「神经调质对应强化学习参数」假说(多巴胺=奖赏预测误差、血清素=时间折扣等),曾获国际神经网络学会 Donald Hebb 奖。
主持人CCN 2026(认知计算神经科学会议)主旨演讲环节主持人,负责介绍讲者并主持问答。
章节 · 点击跳转视频
0:08 开场:神经调质即算法参数 ▶ 正在看
4:54 从行走小机器人到学会摇喂食器的猫 ▶ 正在看
7:14 强化学习三要素:预测、选择、更新 ▶ 正在看
12:54 基底神经节与多巴胺的 TD 误差 ▶ 正在看
17:24 光遗传刺激血清素让小鼠更耐心 ▶ 正在看
20:32 无模型与基于模型:心理模拟的代价 ▶ 正在看
24:37 脑区分工假说与网格导航 fMRI 实验 ▶ 正在看
28:58 顶叶解码与血清素理论的修正 ▶ 正在看
35:28 意识、数据同化与全局工作空间 ▶ 正在看
38:54 推断与控制的对偶性及皮层微回路 ▶ 正在看
46:46 数字大脑、具身进化与 AI 安全 ▶ 正在看
59:26 问答:无模型是否暗含模型 ▶ 正在看
本期论点
本期回应
30:30
顶叶神经元对目标距离的编码在感觉反馈中断时仍靠内部模型推进,接近贝叶斯推断 因果模型人的思维主要靠哪一种机制?
32:57
把血清素活动看作对奖赏的先验概率,能很好地拟合动物的等待行为 因果模型人的思维主要靠哪一种机制?
36:10
用意识解释人类智能,和过去用神解释宇宙、用灵魂解释生命是同一类做法 已经够了用今天的科学,能把意识讲清楚吗?
其他论点
6:37
动物只凭少量偶发经验就能学会,而强化学习论文动辄需要几百万次试验 观察
9:54
大脑未必使用与强化学习相同的方程,但一直在预测未来奖励、据此选择动作并从经验中学习
12:31
强化学习算法虽有收敛性证明,但参数与奖励设计不当会让学到的行为完全偏离设计者意图
14:29
纹状体一部分神经元学习动作价值以选择动作,另一部分表征状态价值以计算时序差分误差 观察
17:40
从脑干广泛投射到大脑各区的神经调质,可能负责调节强化学习的全局参数
20:00
基底神经节并非大脑中唯一进行强化学习的地方,杏仁核、海马和小脑也可能参与
21:31
基于模型的方法在行为目标改变时能更灵活适应新设定,代价是内部计算很重
24:45
大脑按学习类型分工:小脑做监督学习、基底神经节做强化学习、皮层做无监督表征学习
40:58
既有的贝叶斯推断算法可借由推断与控制的对偶性直接转换成强化学习算法
52:09
把强化学习和进化框架结合,机器人能自行演化出适合生存与繁殖的奖赏函数
58:05
大语言模型的伦理不应只靠后训练,社会性意识应从训练早期阶段就纳入 做法
1:01:12
为动作和决策找出正确的状态空间,是强化学习实际落地中的关键问题
01开场:神经调质即算法参数
0:08
Okay, we are going to try to get started. Uh, welcome back to another very exciting uh, keynote lecture. It is my pleasure to introduce Professor Kenji Doya, professor of the neurocomputation unit at the Okinawa Institute of Science and Technology. Kenchi studies reinforcement learning and probabilistic inference and how the brain implements them. If you have ever thought of a neurom modulator as a parameter in a learning algorithm, you're working uh downstream of his ideas. Dopamine for reward prediction error, serotonin for discount um factor nor adrenaline for exploration and so on. Kenchi did his PhD at the University of Tokyo, then went on to UC San Diego and the Sulk Institute where he worked with Terry Sinowski on a reinforcement learning model of bird song. He joined ATR in 1994 and has been at Okinawa Institute for Science and Technology since 2004.
好,我们准备开始了。呃,欢迎回来参加又一场非常精彩的主旨演讲。我很荣幸向大家介绍铜谷贤治教授(KenjiDoya),他是冲绳科学技术大学院大学神经计算研究室的教授。铜谷老师研究的是强化学习、概率推断,以及大脑是如何实现它们的。如果你曾经把某种神经调质看作学习算法里的一个参数,那你其实是在他的思想脉络之下工作的。多巴胺对应奖赏预测误差,血清素对应折扣因子,去甲肾上腺素对应探索,等等。铜谷老师在东京大学取得博士学位,之后前往加州大学圣地亚哥分校和索尔克研究所,在那里与 Terry Sejnowski 一起研究鸟鸣的强化学习模型。他于 1994 年加入 ATR,自 2004 年起一直在冲绳科学技术大学院大学任职。
便签引用
1:04
Among many honors, he has received the International Neuronetwork Society's uh Donald Hub award. He also finished the iron man world championship in Kona in 2024 which suggested which suggests that his tolerance for long horizons is not purely theoretical. Today he will be uh uh he will be giving a talk on the topic of neuroscircuits for prediction and action. Please join me in welcoming professor Kenzie Doya. [applause]
在众多荣誉之中,他获得了国际神经网络学会的呃 Donald Hebb 奖。他还在 2024 年完成了科纳的铁人三项世界锦标赛,这说明他对长时间跨度的耐受力并不只是停留在理论上。今天他将做一场关于呃呃,关于预测与行动的神经回路这一主题的演讲。请大家和我一起欢迎铜谷贤治教授。[掌声]
便签引用
1:39
>> Okay. Yeah, welcome back and thank you very much for a kind introduction and I like thank the organizers of this conference for picking me for the uh keynote talk at this time. Right. Okay. So, uh I announced my title to be uh the neuroscience for prediction action but today so I'm going to speak a bit wider than this specific topic. So uh let me start with uh introducing my university and uh my lab. So uh yeah my university is Okinawa Institute for Science Technology. Uh but uh do you know where is Okinawa?
>> 好的。是的,欢迎回来,非常感谢刚才热情的介绍,我也想感谢这次会议的组织者选我在这个时段做主旨演讲。好的。那么。那么,呃我报的题目是呃预测与行动的神经科学,不过今天呢,我会讲得比这个具体题目稍微宽一些。那么呃先让我从介绍一下我的大学和呃我的实验室开始。呃是的,我的大学是冲绳科学技术大学院大学。呃不过呃,你们知道冲绳在哪儿吗?
便签引用
2:17
So uh yeah this is um in the middle of the ocean. So about two and a half hour flight from Tokyo but uh there's many convenient flights from many cities in the East Asia. So uh we want to have uh students and researchers uh come together here not only just the tourists for the uh international internally research right. So uh um my lab was one of the first lab to be started in the oy so uh and then I try to make my own lab kind of interdictionary place. So on one side we work on the topic of creating flexible uh learning systems uh and another side we want to understand the brains mechanism for uh learning. So we work on the robotics and machine learning and also the neurobiology experiments mostly using rats and mice right but we are not working on totally different things. uh the common theme of our research is a reinforcement learning. So as you know so it is a framework for an agent that can be a human or animal or software or robot to interact with the environment and get the reward feedback evaluating how good
那么呃是的,它在大海的中间。从东京飞过去大概要两个半小时,不过呃有很多从东亚各个城市飞过来的方便航班。所以呃我们希望学生和研究人员能汇聚到这里来,而不只是游客,是为了呃国际化的、真正的研究,对吧。那么呃嗯,我的实验室是OIST 最早成立的实验室之一,呃然后我试着把自己的实验室做成一个跨学科的地方。所以一方面我们研究如何创造灵活的呃学习系统,呃另一方面我们想理解大脑呃学习的机制。所以我们做机器人和机器学习,也做神经生物学实验,主要是用大鼠和小鼠对吧,但我们做的并不是完全不相干的事情。呃我们研究的共同主题是强化学习。如你们所知,它是一个框架,讲的是一个智能体——可以是人、动物、软件或机器人——去与环境交互,并获得评价该智能体表现好坏的奖赏反馈。
便签引用
3:37
the but the agent is performing. Right? So aim of reinforcement learning to find the action policy to maximize the acquired reward and then the coming up with efficient algorithms for reinforcement learning is a very important topic in computer science and AI these days and also uh understanding the brains mechanisms for reinforcement learning probably would be a very important topic in the neuroscience. So uh once you have some progress in one side maybe that will give a a hint to the other and if you hit the wall uh in one side maybe that gives a a theoretical or experimental uh problem setting for the other so that's why we work on these two stream of research in parallel right so uh today so I would start from the uh reinforcement learning and then the related topic of mental simulation uh and also So uh the topic about inference and control or prediction and action and finally uh about AI and neuroscience right okay so first uh the the uh let me start with the reinforcement learning so I've been
对吧?强化学习的目标就是找到能最大化所获奖赏的行动策略,而设计出高效的强化学习算法,如今在计算机科学和人工智能里是一个非常重要的课题;同时呃,理解大脑实现强化学习的机制,在神经科学里大概也会是一个非常重要的课题。那么嗯,一旦你在其中一边取得了一些进展,也许那会给另一边一点提示;而如果你在其中一边碰壁了,也许那就给另一边提出了一个理论上或实验上的问题设定。所以这就是我们为什么要并行地做这两条研究线索的原因。好,那么今天,我会先从强化学习讲起,然后讲相关的心理模拟这个话题,还有关于推断与控制、或者说预测与行动的话题,最后再讲人工智能与神经科学。好,那么首先,让我从强化学习开始。我在这上面已经做了差不多四十多
便签引用
02从行走小机器人到学会摇喂食器的猫
4:54
working on this uh almost more than 40 years actually so this was my undergraduate thesis project of uh creating a small robots that learns to work. So uh uh this is a a tiny robot like 10 cm long. Uh but uh at that time there was no textbook about reinforcement learning over robot learning. So but anyway I figured out that by uh numerically representing the motion pattern and then changing the pattern a little bit toward the direction of improvement maybe you can learn to uh acquire different motion patterns and indeed with a very simple uh algorithm uh controlled by the 8bit uh personal computer at that time. So this uh robot could discover many dynamic uh behaviors and what was important is that this uh rotary encoder which monitors the speed of the movement. So this is a kind of reward signal for this robot to learn. Right?
年了。这其实是我本科的毕业设计项目,做一个学习行走的小机器人。这是一个很小的机器人,大概十厘米长。但是在那个年代,还没有关于强化学习或者机器人学习的教科书。不过总之我想明白了一点:通过用数值来表示运动模式,然后朝着改善的方向稍微改变这个模式,也许就可以学会获得不同的运动模式。而且确实,用一个非常简单的算法,由当时的八位个人电脑来控制,这个机器人就能发现许多动态的行为。而重要的是这个旋转编码器,它监测运动的速度。这对这个机器人来说就是一种用于学习的奖励信号,对吧?
便签引用
5:57
So uh and that was a starting point and also I'm interested in reinfor learning in animals. So uh this is my cat now. She's very cute but she's also very wild. So from sometimes she learned that by shaking this food dispenser she can sometime get the food reward right. So probably this is not a genetically predetermined behavior but I never told her to to do this but somehow she could find out that she could get a reward by this way. So recently the many brainstorm learning papers have a x-axis like a millions of trials. But how she could learn this kind of behavior very like a occasional experience with a relative small number of experience that is a very uh important uh research motivation for me right okay so uh and then uh in reinforcement learning there are several important uh computations one is a prediction of the future reward uh in the form of a so-called value function so V of S is called a state value function and Q of SA called the action value function which is the expectation of a future
那这就是一个起点。同时我也对动物身上的强化学习感兴趣。这是我现在养的猫。她非常可爱,但也很野。有时候她学会了,只要摇晃这个喂食器,她有时候就能得到食物奖励,对吧。这大概不是基因预先决定的行为,我从来没有教过她这么做,但不知怎么她就发现了可以用这种方式得到奖励。最近很多强化学习的论文,横轴都是几百万次试验。但她怎么能通过这种偶发的经历,用相对少量的经验,就学会这种行为呢?这对我来说是一个非常重要的研究动机。好,那么在强化学习里,有几种重要的计算,其中之一是对未来奖励的预测,形式是所谓的价值函数。V(S) 叫做状态价值函数,Q(S,A) 叫做动作价值函数,也就是从某个状态出发、或者执行某个特定动作之后,未来奖励的期望值。而
便签引用
03强化学习三要素:预测、选择、更新
7:14
reward starting from certain state or by performing particular action right and then the here the gamma is called temporal discounting parameter setting how long into the future you take into account in the prediction right and then such a real prediction is helpful to guide your behavior uh the simple Action selection is called the greedy to find out the action that maximize the expected state expected reward from the current state. [snorts] And then in earlier training you would take a stoastic search. So in this case you take the action value function as a kind of a negative potential and then this beta is a called inverse temperature control how deterministic or stoastic your action selection can be. And then after selecting action you learn from the experience. So that is based on the reward prediction error. So thanks to this exponential discounting framework.
这里的 gamma 叫做时间折扣参数,它设定你在预测中把未来考虑到多远,对吧。那么这样的预测就有助于指导你的行为。最简单的动作选择方式叫做贪婪法,就是找出那个能让从当前状态出发的期望奖励最大化的动作。[吸鼻子]然后在训练早期,你会采取随机搜索。在这种情况下,你把动作价值函数当作一种负的势能,而这个 beta 叫做逆温度,控制你的动作选择有多确定性、或者多随机。然后在选定动作之后,你从经验中学习。这是基于奖励预测误差的。多亏了这个指数折扣的框架,
便签引用
8:16
So uh you have a a kind of a consistency equation. So the uh reward uh you predicted before taking action should be equal to the reward you actually acquired and the reward you still expect from a subsequent state discounted by gamma right. So if your prediction is perfect these first two term and last term should cancel out otherwise you learn from the this prediction error. So you update the previous state value function or pre previous action value function in proportion to this temporary difference. And here alpha is a learning rate parameter to control how quick you learn or how quickly you forget. Right?
你会得到一种一致性方程:你在采取行动前预测的奖励,应该等于你实际获得的奖励,加上你从后续状态仍然期望得到的奖励再乘以 gamma 折扣,对吧。所以如果你的预测是完美的,前两项和最后一项应该互相抵消;否则你就从这个预测误差中学习。你按照这个时间差分的比例,去更新之前的状态价值函数或者之前的动作价值函数。这里的 alpha 是学习率参数,用来控制你学得多快、或者说忘得多快,对吧?
便签引用
8:58
So uh and then implementing uh these uh steps of prediction and selection and then learning uh is a very important uh uh topic in the implementation of reinforcement learning. For example, recently the proper way of combining deep neural networks for this value function approximation was the uh led to a major uh breakthrough in the reinforce learning in AI and also important is how to set these parameters like temporal discounting and then exploration explorator tradeoff and then uh stability and speed of learning.
那么,如何实现这些预测、选择和学习的步骤,是强化学习实现中一个非常重要的课题。举个例子,最近,把深度神经网络恰当地结合起来做价值函数逼近,带来了人工智能中强化学习的重大突破。另外重要的是,如何设定这些参数,比如时间折扣、探索与利用的权衡,还有学习的稳定性和速度。
便签引用
9:38
Right. Right. So how to tune these parameters that is also a very important topic in the practical implementation of reinforcement learning. So and these are uh important in engineering application but I think these are also important topic in the neuroscience as well. So our brain may not be using exactly the same equations but the prediction of future reward and then according action selection and learning from the actual experience is something we have been uh our brain should be doing. So how such a computation is done using the circuit of the brain and possibly using different molecules in the brain. So that is a very important topic in neuroscience and that something we have been uh pursuing uh in parallel right. So uh this is another example of reinforcement learning uh to uh stand up. So uh there's idea to uh let the robots to uh support the elderly or support patients.
对,对。所以如何调这些参数,这在强化学习的实际实现中也是一个非常重要的课题。这些在工程应用中很重要,但我认为它们在神经科学里同样是重要的课题。我们的大脑也许并没有在用完全相同的方程,但预测未来奖励、据此选择动作、并从实际经验中学习,这是我们的大脑应该一直在做的事情。那么这样的计算是如何用大脑的回路完成的,可能还用到大脑中不同的分子,这在神经科学中是一个非常重要的课题,也是我们一直在并行追求的方向。这是强化学习的另一个例子——学习站起来。有一个想法是让机器人去照护老年人或者照护病人。
便签引用
10:38
But if you go to uh the robotic laboratories, you often find that two or three graduate students are caring for one robot. So uh uh we thought that the robots have to be strong enough to be able to fall down and stand up by itself. Right? So this robot is about 7 cm high and has a motor in the uh knees and hip. And then it has a a gyro sensor. It can discriminate whether it is lying down or standing up and then we gave a reward based on the height of the head from the floor and also when it fell down we gave a negative reward or punishment right and after many hundreds of bumpers falls the robot could learn this kind of a standup sequence and important here is that initially the robot keeps his head on the floor so the immediate reward is zero but by putting his weight on the uh foot he can stand up sub subsequently. This is a very simple example of a delayed reward task.
但如果你去机器人实验室,你常常会发现两三个研究生在照料一个机器人。所以我们就想,机器人必须足够结实,能够摔倒、又能自己站起来,对吧?这个机器人大约七厘米高,在膝盖和髋部各有一个马达。它还有一个陀螺仪传感器,可以分辨自己是躺着还是站着。然后我们根据头部离地面的高度给予奖励,另外当它摔倒时,我们给予负奖励,也就是惩罚,对吧。经过几百次磕磕碰碰的摔倒之后,这个机器人就学会了这样一套站立的动作序列。这里重要的是,一开始机器人的头贴在地板上,所以即时奖励是零;但通过把体重压到脚上,他随后就能站起来。这是延迟奖励任务的一个非常简单的例子。
便签引用
11:42
Right? So the a good point of uh reinforced learning is that uh it can solve the uh delayed reward problem or the temporal credit assignment problem. Right? But to be able to do that uh there are very se several important factors. For example the temporal discounting parameter is very important. So if you uh uh set the parameter gamma very small sometimes the robot just sit up and then get a small reward immediately. So that is not uh something we want. And also the balance between the reward and punishment is very important. So if you make this punishment too large in some cases the robot just learn to lay down to avoid any punishment which is not quite uh something we learned right. So even though the reinforcement learning algorithms are proven to uh converge uh under reasonably good conditions but depending on that uh setting of the parameters or depending on designable reward sometimes reward is some uh result is not what you expected right so such a high level of problem in reinforcement learning is a
对吧?强化学习的一个好处,就是它能解决延迟奖励问题,也就是时间上的信用分配问题。对吧?但要做到这一点,有几个非常重要的因素。比如说,时间折扣参数就很重要。如果你把参数 gamma 设得非常小,有时候机器人就只是坐起来,然后马上得到一点小奖励。那就不是我们想要的了。另外,奖励和惩罚之间的平衡也非常重要。如果你把惩罚设得太大,某些情况下机器人就只学会躺着不动,以避免任何惩罚,而这也不太是我们想要的,对吧。所以,尽管强化学习算法被证明在相当好的条件下是收敛的,但取决于参数的设定,或者取决于可设计的奖励,有时候奖励——呃,结果并不是你预期的那样,对吧?所以这种高层次的问题在强化学习里非常重要,这是我们通过尝试让机器人学习而学到的东西,
便签引用
04基底神经节与多巴胺的 TD 误差
12:54
very important that's something we learn by let trying to let robots learn right so Then the regarding the implementation of the reinforcement learning in the brain. So uh uh like many people we suspect that the uh basa ganglia which is hidden under the cerebral cortex would play a very important role right. So the basal ganglia receive is a major input from the cortex and then there's a sequence of nuclei starting from the stratam paradam and it's out to go through the thermos and back to the the cortex [snorts] and the important feature of the basa ganglia is that this receives a very strong dopamineagic input from the substantia So this is a more like engineering view of the circuit. So and then the very important finding was that the dopamine neurons which sends its output to the basa ganglia has a temporal difference error like activities which was discovered by ro from shoes and and also the uh dopamine in the stratum uh is related to the plasticity of the input to cortex through the straight called the dopamine
对吧。那么关于强化学习在大脑中的实现方式。呃,跟很多人一样,我们怀疑隐藏在大脑皮层下方的基底神经节会扮演非常重要的角色,对吧。基底神经节接收来自皮层的主要输入,然后有一连串的核团,从纹状体、苍白球开始,输出再经过丘脑回到皮层,而基底神经节的一个重要特征是它接收来自黑质的非常强的多巴胺能输入。所以这更像是从工程角度看这个回路。然后非常重要的发现是,向基底神经节发送输出的多巴胺神经元具有类似时序差分误差的活动,这是由舒尔茨等人发现的;另外,纹状体中的多巴胺还与皮层输入的可塑性有关,也就是所谓的多巴胺依赖性可塑性,这是由
便签引用
14:11
dependent plasticity which is found by micro colleague in oyster Jeff Wiggins. So based on those findings uh people started to uh think that this basa ganglia would play a very important role in reinforcement learning. So our specific hypothesis is that some of the neurons in Australia may learn like action value function which is subsequently used for action selection and some of the neurons represent that state value function which be used for like a temporary difference error computation. Right?
三隅(音)及其同事和杰夫·威克恩斯等人发现的。基于这些发现,呃,人们开始认为基底神经节会在强化学习中扮演非常重要的角色。所以我们的具体假设是,纹状体中的一部分神经元可能会学习类似动作价值函数的东西,这个函数随后被用于动作选择;而另一些神经元表征状态价值函数,用来做类似时序差分误差的计算。对吧?
便签引用
14:48
So then the uh when I started my lab in Oya I had a chance to start a new rodent lab. So uh then we started out with experiment. So this is a binary choice task for after putting his nose in the center port he can choose either go to the right port or left port and then depending on his choice the food is given here with certain probability and that probability keep changing. So the rat have to keep track of the goodness of left port or goodness right port. So all the uh action value for the left action bar for the right and then the uh we fitted the uh this animal's choice behavior uh to the uh reinforcement learning model and then estimated the action value function behind the uh choice sequence and then the checked if there's any neurons that change behaviors like this. And indeed some of the neurons changed the behavior over about 200 trials very similar to this action value function. So we uh s uh suggested that these neurons are uh involved in action value for the left choice in this case and some other
那么,呃,当我在冲绳(OIST)开始建立自己的实验室时,我有机会新建一个啮齿类动物实验室。所以呃,我们就从实验开始。这是一个二选一任务:大鼠把鼻子伸进中间的孔之后,它可以选择去右边的孔或者左边的孔,然后根据它的选择,食物会以一定的概率在这里给出,而这个概率会不断变化。所以大鼠必须持续追踪左边孔的好坏或者右边孔的好坏。所以就有左边动作的动作价值和右边动作的动作价值,然后呃我们把这只动物的选择行为拟合到呃强化学习模型上,估计出选择序列背后的动作价值函数,然后检查是否有神经元的活动像这样变化。确实,有些神经元在大约 200 次试次中的变化方式与这个动作价值函数非常相似。所以我们呃推测,这些神经元在这个例子中参与了左侧选择的动作价值,而另一些
便签引用
16:02
neurons kept track of the goodness the right choice and then by like some kind of competition but these two kind of neurons may be doing the action selection right and also we found that different part of the stratum are involved in the action body function or still body function or detailed movement And then at this time uh we were not uh discriminating the cell types in the uh straight but more recently it is possible to record from specific types of neurons. So uh this uh is the uh u calcium imaging experiment in which uh we express the uh calcium indicator to the neurons which project to the dopamine neurons in the substantial diagram. And with a classical condition task we could see that these neurons start to respond to the uh forcecoming reward after association with the order to the reward uh delivery. Right? So which is consistent with the hypothesis that uh these neurons encode the state value function right and then the regarding the uh regulation of the parameters of reinforcement learning. So uh it has
神经元则在追踪右侧选择的好坏,然后通过某种竞争,这两类神经元可能就在完成动作选择,对吧;而且我们还发现纹状体的不同部分分别参与动作价值函数、状态价值函数或者具体的运动。而在那个时候,呃,我们还没有区分纹状体中的细胞类型,但最近已经可以从特定类型的神经元记录了。所以呃这是一个呃钙成像实验,我们在投射到黑质多巴胺神经元的那些神经元中表达钙指示剂。用一个经典条件反射任务,我们可以看到,在把提示音与奖励递送建立关联之后,这些神经元开始对即将到来的奖励产生反应。对吧?这与这些神经元编码状态价值函数的假设是一致的,对吧。然后关于呃强化学习参数的调节。呃,已经
便签引用
05光遗传刺激血清素让小鼠更耐心
17:24
been well uh uh supported that the dopamine encodes the most important uh signal in the rainful learning temporal difference. And then uh we started to speculate that other neurom modulators which project from the brain stem to the wide area of the brain may regulate the global parameters of the reinforcement learning and then we have been especially working on the top uh function of the serotonin. So uh and then uh uh based on the hypothesis that thin may be uh controlling a temporal discounting parameter we have worked on the recording experiments and also more recently optogenetic manipulation of the certain neurons. So this is a genetically engineered mouse. So which express channel redoxing light sensitive uh ionic channels uh only in the neurons that produce serotonin and then by uh the optal fiber shining the blue light we can selectively stimulate the serotonin neurons. The yellow light uh is a kind of control condition, right?
有很充分的呃呃证据支持,多巴胺编码了强化学习中最重要的信号,也就是时序差分。然后呃我们开始猜想,其他从脑干投射到大脑广泛区域的神经调质,可能在调节强化学习的全局参数,我们尤其一直在研究血清素的功能。所以呃,然后呃呃基于血清素可能在控制时间折扣参数这一假设,我们做了记录实验,最近还做了对血清素神经元的光遗传学操控。这是一只经过基因改造的小鼠。它只在产生血清素的神经元中表达通道视紫红质这种光敏离子通道,然后通过光纤照射蓝光,我们就能选择性地刺激血清素神经元。黄光呃是一种对照条件,对吧?
便签引用
18:37
And then the he waits for the hood fret after 3 second, 6 second, 9 seconds or in infinity delays, right? So the with the 9-second delay. So uh mouse often abandons to wait but such abandoning error reduced by the serotonin simulation by the blue light. For the infinity trials, mouse usually abandon around 10 seconds, but by blue light simulation of certain neurons, they could wait much longer. So the animal became more patient or more motivated for the delayed reward. So which is consistent with the uh the hypothesis that certainly regulates the temporal discounting parameter.
然后它要等待食物奖励,延迟 3 秒、6 秒、9 秒,或者无限长的延迟,对吧?在9 秒延迟的情况下,呃小鼠常常会放弃等待,但这种放弃错误在蓝光刺激血清素神经元后减少了。在无限延迟的试次里,小鼠通常在 10 秒左右放弃,但在用蓝光刺激血清素神经元时,它们能等待得久得多。所以动物变得更有耐心,或者说对延迟奖励更有动力。这与血清素调节时间折扣参数这一假设是一致的。
便签引用
19:22
Right? So and then we have gradually uh understood the mechanism of the rainfor learning in the brain but still there are many uh open questions. Uh for example the circuit of the vessel gang is quite complex. There's a pathway called direct pathway and indirect pathway. Why do we need those parallel pathway and the multiple integrity mechanisms is still a unresolved question and also the it is well established that the dopons have a temporal difference error like activities but what kind of circuit the mechanism realize such TD like response is still not quite understood [sighs] and also the in the brain basic ganglia may not be the only place for reinforcement learning other mechanisms like amigdulara Hypoc campus and cerebrum may also be involved in. So how what are the difference and the commonality of those function of the different possible brain in reinforcement learning that's very uh important topic we should be further studying right okay so then the next uh uh topic of mental simulation this is still related
对吧?那么,我们逐渐呃理解了强化学习在大脑中的机制,但仍然有很多呃悬而未决的问题。呃,比如说,基底神经节的回路相当复杂。有一条叫直接通路的通路,还有间接通路。为什么我们需要这些并行通路,以及多重整合机制是怎样的,这仍然是没有解决的问题;还有,虽然已经确立多巴胺神经元具有类似时序差分误差的活动,但究竟什么样的回路机制实现了这种类似 TD 的反应,仍然不是很清楚;另外,在大脑里,基底神经节可能并不是唯一进行强化学习的地方,其他机制比如杏仁核、海马和小脑可能也参与其中。所以这些不同脑区在强化学习中的功能有什么差异和共性,这是我们应该进一步研究的非常呃重要的课题,好。那么接下来呃呃的话题是心理模拟,这仍然和强化学习有关。所以我刚才讲的那种
便签引用
06无模型与基于模型:心理模拟的代价
20:32
to reinforcement learning so the kind of a reinforce learning framework I talked about is called model free so uh this uh uh framework assume no knowledge of the agent about the road and then the agent learn the behavior based on the sequence of experience of a state action reward and state action reward. So uh then the this is a kind of a reactive uh simple process. The algorithm very simple but it tend to require a lot of repetitions right and then uh in the model based approach. So uh agent learn the uh prediction model how the state of the environment change with the action and then with such a uh prediction model we can uh estimate the current state based on the previous state and action and also plan ahead for the appropriate sequence of actions to reach to the desired state. So this is the more deliberative uh strategy and then the especially when the goal of the behavior changed. So it can uh allow the agent to uh adapt more flexibly to the new setting. But internal computation can be
强化学习框架被称为无模型的,呃这个呃呃框架假定智能体对世界没有任何知识,智能体是基于状态—动作—奖励、状态—动作—奖励这样的经验序列来学习行为的。所以呃这是一种反应式的、简单的过程。算法非常简单,但它往往需要大量的重复,对吧;然后呃在基于模型的方法里,呃智能体学习一个预测模型,即环境的状态如何随动作而变化,然后有了这样的呃预测模型,我们就可以根据前一个状态和动作来呃估计当前状态,也可以提前规划出合适的动作序列,以到达期望的状态。所以这是一种更审慎的呃策略,而且尤其是当行为的目标改变时,它能让智能体呃更灵活地适应新的设定。但内部计算可能会
便签引用
21:50
heavy. So how to uh select between model free versus model based or possibly combine these two approaches is a bit interesting topic uh in the machine learning and also in neuroscience. Right? So this is one example of a model based reinforcement learning. This is a smartphone based robot and then uh uh while initially doing this kind of a random behavior. So this robot learn how the body respond to the wheel movement and then using such internal model to uh optimize the control policy offline. So then within less than 10 trials this robot could bounce up and keep balancing. So this is a work by my former PhD student Pavo. So this is a one example of the data efficiency of model based strategy.
很重的。所以如何在无模型和基于模型之间做选择,或者说有没有可能把这两种方法结合起来,这是个挺有意思的话题,无论是在机器学习领域还是在神经科学领域。对吧?这就是基于模型的强化学习的一个例子。这是一个基于智能手机的机器人,一开始它会做出这种随机的行为。所以这个机器人学会了身体如何对轮子的运动做出反应,然后利用这样的内部模型来离线优化控制策略。于是在不到十次尝试之内,这个机器人就能弹起来并保持平衡。这是我以前的博士生 Pavo 做的工作。这就是基于模型的策略在数据效率方面的一个例子。另外还有一位学生 Doni,他现在在微软,他提出了一种机制,把无模型和基于模型的强化学习结合起来,
便签引用
22:39
uh and also uh another student Doni uh who is now at Microsoft uh proposed uh mechanism to uh combine the the model free and the model based reinforcement learning uh for efficient uh learning right and then these are kind of uh robotics and the machine learning uh approaches but when the uh humans and the animals use uh such internal models we can consider it as a kind of a mental simulation. We can define it as a brain process using action dependent state transition model. So s prime equal fsa this is a determinist case or in general this can be a action dependent state transition model right [snorts] and then the by having such a prediction model we can estimate the present state even when the current sensory input is noisy or delayed or occluded based on the past state and action right and uh also uh you can use such a prediction model uh to uh find out the approach appropriate action or sequence of action to lead to the desired state for model based decision or action planning.
以实现高效的学习。对吧,这些属于机器人学和机器学习的方法,但是当人类和动物使用这样的内部模型时,我们可以把它看作一种心理模拟。我们可以把它定义为大脑使用依赖于动作的状态转移模型的过程。也就是 s' = f(s,a),这是确定性的情况,或者更一般地说,这可以是一个依赖于动作的状态转移模型,对吧【吸鼻子】有了这样的预测模型,即使当前的感觉输入有噪声、有延迟或者被遮挡,我们也可以基于过去的状态和动作来估计当前的状态。而且你还可以用这样的预测模型去找出合适的动作或者动作序列,从而达到期望的状态,这就是基于模型的决策或者动作规划。而且你可以把这样的预测模型应用到任意想象出来的状态上去学习,就像做思想实验一样,对吧?
便签引用
23:55
And you can apply such a prediction model to arbitrary imaginary state to learn like a thought experiment, right? So for example, our use of language would be very much uh uh dependent on our capability for imagining the situation. And also when we do a scientific research before running a physical experiment we do a lot of thought experiments to think about the right conditions. So this uh uh mechanism of mental simulation would be very important starting from a sensory motor control to higher committing functions. So understanding the mechanism of mental simulation I think would be a very important topic in the neuroscience today. Right.
比如说,我们对语言的使用很大程度上依赖于我们想象情境的能力。还有,当我们做科学研究时,在真正做物理实验之前,我们会做大量的思想实验来思考合适的条件。所以这种心理模拟的机制会非常重要,从感觉运动控制一直到更高级的认知功能。所以我认为理解心理模拟的机制在今天的神经科学中是一个非常重要的课题。对。关于心理模拟的实现,我认为大脑不同的学习机制
便签引用
07脑区分工假说与网格导航 fMRI 实验
24:37
So uh regarding uh implementation of the mental simulation uh I think uh the different learning mechanism of the brain uh may be important. So uh I have a kind of prop proposed like a framework that the different uh part of the brain like cerebrum basal ganglia and the cortex would be specialized for different types of learning like cerebrum for the supervised learning of input output mapping and the bas gangria for reinforcement learning for the maximization and prediction of the future reward and the cerebral cortex for unsupervised representation learning based on the statistic feature of the input distribution.
可能很重要。我提出过这样一个框架:大脑的不同部分,比如小脑、基底神经节和皮层,会分别专门负责不同类型的学习,比如小脑负责输入输出映射的监督学习,基底神经节负责强化学习,即最大化和预测未来的奖励,而大脑皮层负责基于输入分布的统计特征进行无监督的表征学习。所以它们会用于不同类型的计算,对吧?举例来说,无模型的强化学习可以
便签引用
25:24
So they would be used for a different kind of a computation, right? And then the for example the model free reinforcement learning could be implemented by linking the uh the cortical representation and then value function uh in the basa gangria [snorts] and uh model based uh action selection may be a bit more complex but uh cerebellum may be able to learn the uh action dependent status prediction model based on it's a supervised learning framework and then the uh based on that such a prediction model we can estimate the predicted future state and goodness of such state could be evaluated by the value function the bas ganglia [snorts] and then depend how good the prediction is you can uh implemented action if it is good enough right so that was our like a theoretical speculation and then uh We started to test such a hypothesis in the function MR experiment. So uh and this is a grid world navigation task. Uh but uh to make this game interesting, we limited the movement the uh precursor to
通过把皮层的表征和基底神经节中的价值函数连接起来来实现,【吸鼻子】而基于模型的动作选择可能要复杂一些,但小脑基于它的监督学习框架,也许能够学到依赖于动作的状态预测模型,然后基于这样的预测模型,我们可以估计预测的未来状态,而这种状态的好坏可以由基底神经节中的价值函数来评估,【吸鼻子】然后根据预测有多好,如果足够好的话,你就可以执行这个动作,对吧。这就是我们理论上的推测,然后我们开始在功能性磁共振实验中检验这个假说。这是一个网格世界的导航任务。但为了让这个游戏更有意思,我们把光标的移动限制在只有三个方向。所以在某个特定方向上,是没有直达路径的。你必须找出这种之字形的
便签引用
26:45
only three directions. So in a certain direction, there's no straight path. You have to find out this kind of a zigzag uh path. So we call it grid navigation task like the sailing y going to the upwind right so the point is that uh finding this kind of path is very difficult with just a random search but once you have a internal model of how this cursor moves with the fishing of the one of the three buttons so you can use such a action dependent state transition model to find out this kind of sequence right and indeed uh by comparing the performance of the subject uh with with and without the pre-training and also the time given for thinking before starting the finger movement we could verify behaviorally that sub indeed using a pre-learned action dependence transition model and then we analyzed the brain activity before the actual finger movement after the uh goal position was presented when the subject is supposed to be doing mental simulation of the this moving discusser without actually moving it by
路径。所以我们把它叫做网格导航任务,就像帆船逆风前进那样,对吧。关键在于,用随机搜索的方式要找到这样的路径是非常困难的,但一旦你有了一个内部模型,知道按下三个按钮中的某一个时光标会怎么移动,你就可以用这个依赖于动作的状态转移模型来找出这样的序列,对吧。而且确实,通过比较被试在有预训练和没有预训练情况下的表现,以及在开始手指动作之前给出的思考时间,我们能够从行为上验证被试确实在使用预先学到的依赖于动作的转移模型。然后我们分析了在呈现目标位置之后、实际手指动作之前的大脑活动,也就是被试应该正在对这个移动的光标做心理模拟、而没有真正用手指移动它的时候。我们发现大脑中顶叶皮层、运动前区皮层和前额叶皮层出现了活动,这些区域已知与
便签引用
28:00
by fingers. And we found the activities in the uh brain uh in the parietital cortex and preotal cortex and the prefrontal cortex which are known to be involved in the special processing or motor sequence or special working memory. [snorts] But in addition we also find elevator activities in the part of the cerebam and part of the the bas ganglia which are known to be connected to these areas. So this suggests that me action planning by mental simulation involves a global brain circuit and our interpretation is that maybe cerebr is providing action dependent state transition model and basler would be involved in the variation of the predicted uh candidate actions [sighs] and we are also interested in the circuit mechanism uh neuron level mechanism of uh mental simulation.
空间加工、运动序列或空间工作记忆有关。【吸鼻子】但除此之外,我们还在小脑的一部分和基底神经节的一部分发现了升高的活动,而这些部位已知与上述这些区域是相连的。所以这表明,通过心理模拟进行的动作规划涉及一个全脑的环路,我们的解释是,也许小脑提供了依赖于动作的状态转移模型,而基底神经节参与了对预测出的候选动作的评估。【叹气】我们还对心理模拟的环路机制,也就是神经元层面的机制很感兴趣。为此,我们选取了状态预测这个课题。对吧。为了这个实验,我们
便签引用
08顶叶解码与血清素理论的修正
28:58
So uh uh to do that uh uh we uh took the uh topic of state prediction right. So uh for this uh experiment we collaborate with professor brand who is expert of two photon calcium imaging and my former uh posto Akihiro created the auditory virtual environment. So where the mouse moves on this airflated wall and then the sound feedback is changed to uh uh approximate reaching to the target position. Right? And then why the uh mouse is uh engaged in this behavior. We could monitor the hundreds of neurons in the parietital cortex of the the mouse and then the uh based on the activities of the these neurons we did the neural decoding analysis. So we first identify the distance tuning of these neurons and then based on the instantaneous activity of neurons we can estimate the posterior probability of the uh goal distance represented by this population of neurons right and then the uh what we found is that even when we shut off the sound feedback the neural coding of the distance moved our head closer to the
和 Brand 教授合作,他是双光子钙成像的专家,而我以前的博士后 Akihiro 搭建了听觉虚拟环境。在这个环境里,小鼠在气浮球上移动,声音反馈会随着它接近目标位置而变化。对吧?然后在小鼠进行这个行为的过程中,我们可以监测它顶叶皮层中数百个神经元的活动,然后基于这些神经元的活动,我们做了神经解码分析。我们首先确定这些神经元的距离调谐特性,然后基于神经元的瞬时活动,我们就能估计出这群神经元所表征的目标距离的后验概率,对吧。然后我们发现,即使我们关掉声音反馈,距离的神经编码仍然会朝着目标方向继续前移,
便签引用
30:23
goal and when the sound feed was turned on again the this uncertainty was reduced. Right? So this is a a kind of a behavior which would be very similar to a direct vision inference used for uh estimate keeping track of the hidden information and uncertain uh uh sensory feedback using internal dynamic model. Right? So and also the regarding the role of uh the serotonin. So uh uh we initially uh assumed that the serotonin would be simply manipulating the temporal discounting parameters but the real uh data was much more like intriguing right. So uh when the uh reward the probability uh is 75% and then only 25% uh uh reward omission. So this uh slation uh had a feature of extending the waiting duration. But when we reduce the pro reward probability to 25% this certain simulation at the rest of the 75% trials didn't have a much effect even when we in increase the number of reward to uh equate the the expected value right uh and also the when the reward is always in the same timing effect of the
而当声音反馈重新打开时,这种不确定性就减小了。对吧?所以这是一种这种行为与贝叶斯推断非常相似,用来估计、追踪那些隐藏的信息,以及在感觉反馈不确定时,借助内部动力学模型来处理。对吧?另外关于血清素的作用。我们最初假设,血清素只是在简单地调节时间折扣参数,但真实的数据要有意思得多。比如说,当奖赏概率是 75%、只有 25% 的试次不给奖赏时,这种刺激会延长等待时长。但当我们把奖赏概率降到 25%,在其余 75% 的试次里,血清素刺激就没有太大效果,哪怕我们增加奖赏的数量、让期望值保持相等也一样。而且,当奖赏总是在同一时间出现时,血清素刺激的效果非常弱;而当时间变得越不固定,
便签引用
32:01
certain simulation was very weak and as the timing became more varied, effect of a certain simulation become much stronger. Right? So the effect of a certain simulation uh is stronger for the certain uh reward delivery but uncertain timing. So this was a quite intriguing feature which was not simply explained by the manipulation of temporal discounting parameter. So we uh kind of thought of a different models and then then came up with a model that the mouse has a internal model of when the uh food operate will come off and then as the uh uh timing of such a reward is more uncertain effect of the prior probability would be uh more dominant. Right? So then the by assuming that the throttling may be representing the prior probability for for the reward uh we could approximate the animal behaviors right and then uh uh regarding the uh model based action selection uh so not only those are like two-step task is quite commonly used in human subjects and Thomas Akam in Oxford came with a a mouse version of such a two-step task.
刺激的效果就越强。所以血清素刺激的效果,在奖赏确定会给、但给的时间不确定的情况下更强。这是一个很耐人寻味的特征,光靠调节时间折扣参数是解释不了的。于是我们考虑了别的模型,最后想出这样一个模型:小鼠有一个内部模型,知道食物大概什么时候会出现;而当奖赏出现的时机越不确定,先验概率的作用就越占主导。那么,如果假设血清素代表的是对奖赏的先验概率,我们就能很好地拟合动物的行为。接下来讲基于模型的动作选择。两步任务不只是在人类被试上很常用,牛津的 Thomas Akam 还做出了小鼠版本的两步任务。
便签引用
33:32
So uh uh in this uh uh pass paramine the uh first step choice result in uh second stage state uh by the common transition versus r transition. And then whether the animal uh change uh behavior based on the actual choice he has made or the uh more common choice common transition expected from the set transition model. will allow us to uh compare between the model based versus model free uh strategies and then by analyzing the mouse behavior. Uh in this case using uh another optogenetics to suppress the certain activities. Uh my former student Masakazu found that uh model based component becomes weaker when the when we suppress the serotonin activities.
在这个范式里,第一步的选择会通过常见转移或罕见转移导向第二阶段的状态。然后,动物是根据它实际做出的选择来改变行为,还是根据状态转移模型所预期的那个更常见的转移来改变行为,这就让我们能够比较基于模型和无模型这两种策略。通过分析小鼠的行为——在这个实验里我们同样用光遗传学去抑制血清素活动——我以前的学生 Masakazu 发现,当我们抑制血清素活动时,基于模型的成分会变弱。
便签引用
34:28
Right? So uh then this kind of a behavior allow us to uh required us to update our uh theory. So now we kind of think that serotonin maybe not just controlling the temporal discounting but variety of features of decision making learning based on the how much time you can spend for decision and learning. Right? So uh this is how uh we started from a simple uh hypothesis to uh uh based on the actual experiment uh came up to a more like generalized theory which needs still needs to be tested right. [snorts] Okay. So then the this mechanism of mental simulation may be related to uh what we know as a consciousness. Right.
对吧?所以这类行为让我们必须去更新自己的理论。现在我们倾向于认为,血清素可能不只是在控制时间折扣,而是在控制决策与学习的多种特征,取决于你能为决策和学习花多少时间。对吧?这就是我们如何从一个简单的假设出发,基于实际的实验,走到一个更加一般化的理论——当然它还有待检验。[吸气] 好。那么这种心理模拟的机制,可能与我们所说的意识有关。
便签引用
09意识、数据同化与全局工作空间
35:28
So uh and some time ago I was collaborating with my friend Pete Pete Hut who is a researcher at the Princeton uh institute. So he previously started with astrophysicist and worked on the origin of life and now working on the origin of intelligence. Right? So uh yeah in the uh old age to explain the mechanism universe people called for the gods for the explan uh explanation and to explain the life uh people called for the soul right uh even though the biology allow us to explain the mechanism life without assuming the soul and today uh to explain the human intelligence many people call for consciousness Yeah. So do we continue to do that or not? That is the question, right? And then reverse meaning of this is that the the neuroscience of consciousness might be biology of soul or physics of the gods. So uh and I think uh the uh what will be uh helpful in understanding possible function of the consciousness is the data assimilation. This is a a kind of a model based uh prediction mechanism. So this is very successfully
前段时间我和我的朋友 Pete Hut 有过合作,他是普林斯顿高等研究院的研究者。他最早是天体物理学家,后来研究生命的起源,现在研究智能的起源。是这样,在过去的年代,要解释宇宙的机制,人们诉诸神来解释;要解释生命,人们诉诸灵魂——尽管生物学已经让我们无需假设灵魂就能解释生命的机制。而今天,要解释人类的智能,很多人诉诸意识。那么我们要不要继续这样做,这是个问题,对吧?反过来说,这意味着意识的神经科学,也许就是灵魂的生物学、或者神的物理学。我认为,要理解意识可能的功能,有一个很有帮助的概念,就是数据同化。这是一种基于模型的预测机制,在天气预报中用得非常成功。我们有大气的
便签引用
36:51
used in the weather forecast. So we have the uh simulation model of the atmosphere and then uh different kinds of uh measurements uh in the satellite or like a weather observatory. So then the the based on the dynamics models and observation model and then while running the uh simulation of the atmosphere we use the uh the any sense available cent sensory data to collect the variables and the parameters of the simulator which has been quite successful in the accurate prediction of the weather. Right? And then the uh this kind of a uh data simulation uh allows uh the further prediction of the future state based on the accurate uh uh understanding of the current state of the atmosphere and also uh uh it can be used for post prediction of the past history based on the uh later measurement.
仿真模型,还有各种各样的观测数据,来自卫星或者气象观测站。基于动力学模型和观测模型,在运行大气仿真的同时,我们用一切可获得的感测数据去校正仿真器的变量和参数,这在天气的精确预报上相当成功。对吧?这种数据同化还能让我们在准确理解大气当前状态的基础上,进一步预测未来的状态;同时它也可以用后来的观测去回溯性地推断过去的历史。
便签引用
37:58
So uh and this kind of uh prediction and the post prediction by combining manageful uh modality of the uh sensory input would be similar to something that people uh argue about the consciousness the so-called the global workplace theory right so uh and then uh we hope to be able to better understand the uh mechanism transmission and what we understand the consciousness in a more like a biological and then computational manner, right? [sighs] Okay. So then the next topic about the inference and control. So uh uh we have worked on the uh mechanism of reinforcing learning and the mechanism of basian inference in the brain. Uh but uh these two are often used uh in combination. Right?
这种把多种感觉模态结合起来做出的预测和回溯预测,和人们讨论意识时所说的所谓全局工作空间理论很相似。所以我们希望能更好地理解意识的机制,以更偏生物学、也更偏计算的方式去理解意识。[叹气] 好,那么下一个话题是推断与控制。我们研究过大脑中强化学习的机制,也研究过贝叶斯推断的机制,但这两者常常是结合在一起使用的,对吧?
便签引用
10推断与控制的对偶性及皮层微回路
38:54
So and then when the uh Rudolfph Kman invented the the optteral filtering algorithm now called the carman filtering. So he realized that the equation used for this optimal filtering is very similar to the equation used for optimal control which was developed by the R Richard Bman several years ago. So this similarity of the uh filtering and control uh is known as the duality of inference and control and recognized one of the beauty of the control theory. But uh recently the some machine learning researchers found the realized the important semantics of this uh similarity.
当年 Rudolf Kalman 发明最优滤波算法、也就是今天所说的卡尔曼滤波时,他意识到最优滤波用到的方程,和几年前 Richard Bellman 提出的最优控制所用的方程非常相似。滤波与控制之间的这种相似性,被称为推断与控制的对偶性,被认为是控制理论之美的体现之一。但最近,一些机器学习研究者意识到了这种相似性在语义上的重要含义。
便签引用
39:40
So the uh in the inference or like a direct via inference uh important computation is a keeping track of the posterior probability of the current uh state right starting from the initial guess the uh state transition model gives you a like a prior prediction and that is combined with a sensory observation to come up with a new posterior and this uh computation is repeated again and again to keep keep track of the current state. So in the optimal control and reinforcement learning important commutation is the value function. So the boundary condition of the value function is usually given at the final goal state and then state transition model uh uh allows you to compute where you should approach this and by combining it with the cost for the action. So uh value function of the previous state is a pres step is given and this computation is repeated again and again for the state value function of the previous state.
在推断、也就是贝叶斯推断里,关键的计算是追踪当前状态的后验概率:从初始猜测出发,状态转移模型给你一个先验预测,再把它和感觉观测结合起来,得到新的后验;这个计算不断重复,以持续追踪当前状态。而在最优控制和强化学习里,关键的计算是价值函数。价值函数的边界条件通常给定在最终的目标状态上,然后状态转移模型让你能够反推你应该用这种方式处理,再把它和动作的代价结合起来。所以前一个状态的值函数就是一个步骤给定之后,这个计算就针对前一个状态的状态值函数一遍又一遍地重复。
便签引用
40:46
So and then this similarity of the the computation uh has a practical merit. So uh there's a class of algorithm called control as inference. So if you have a different kinds of uh basian inference algorithm you can uh convert that into the algorithm for reinforcement learning right some of the uh uh algorithm like uh so soft critic was invented based on this kind of similarity right and and this dity of inference and control uh would not be only useful for engineering applications but also uh in the neuroscience So because uh the uh in the neo cortex the posterior half is mostly used for sensory uh recognition like visual cortex, audiary cortex or small sensory cortex and then frontal half is used for uh motor control and planning like a motor cortex and then prefrontal cortex.
所以这种计算上的相似性其实有实际的好处。有一类算法叫做「控制即推断」(control as inference)。所以如果你有各种不同的贝叶斯推断算法,你就可以把它转换成强化学习的算法,对吧。像 soft actor-critic 这类算法就是基于这种相似性发明出来的。而推断和控制的这种对偶性,不只对工程应用有用,在神经科学里也一样。因为在新皮层里,后半部分主要用于感觉识别,比如视觉皮层、听觉皮层或体感皮层,而前半部分则用于运动控制和规划,比如运动皮层和前额叶皮层。
便签引用
41:47
Right? So but still uh these uh neocortex has the basic six layer structure called the canonical uh microcircuit right even though the thickness and then uh of these layers are different in the different uh cortical areas right so and then there's already uh some hypothesis about how dynamic via inference would be implemented in the circuit of sensory cortex so then by utilizing this uh correspondence between that uh inference and control we can come up with some kind of hypothesis how the uh reinforced learning and optimal control would be informed in me in the motor cortex. So this kind of a possibility uh motivate us to uh uh start a new experiment. So uh and uh now hero uh in my lab surveyed a lot of anatomical literature of the cortical micro circuit and came up with a kangar his hypothesis how different computation for the dynamic via inference and the controls inference could be influenced in the uh sensory and motor cortical circuit and he then started the uh mouse uh behavior and the
对吧?但即便如此,这些新皮层都有基本的六层结构,叫做典型微回路(canonical microcircuit),尽管这些层的厚度在不同的皮层区域是不一样的。所以关于动态贝叶斯推断如何在感觉皮层的回路中实现,已经有一些假说了。那么利用这种推断与控制之间的对应关系,我们就可以提出某种假说,说明强化学习和最优控制可能是如何在运动皮层中实现的。这种可能性促使我们开始了一项新实验。现在我实验室的 Hiro 查阅了大量关于皮层微回路的解剖学文献,提出了一个很有意思的假说:动态贝叶斯推断和「控制即推断」的不同计算可能如何在感觉和运动皮层回路中实现。然后他开始做小鼠的行为学和
便签引用
43:04
neural recording experiments. So the mouse can push or pull the lever and the leist stance of the lever is controlled by the motor. For example, this is quite heavy for the pushing but quite light for the pulling. So by moving the lever to the lighter direction, the mouse is given this water reward. Right? So and then he changed the stiffness to lever in four different uh uh levels and then the uh sometimes mouse simply push or pull for the successful reward and in some cases initially start to push p uh but uh find out that this is a wrong direction and make a change and then finally get the reward. Right? So and then the uh he also change the like a prior probability for the uh uh rewarding direction to be push or pull and then uh animal change the initial push movement direction based on such a prior distribution and also the behavior is dependent on the actual uh setting of the resistance but also affected by the this prior setting, right?
神经记录实验。小鼠可以推或拉这个拉杆,拉杆的阻力由马达控制。比如说,往推的方向很重,但往拉的方向很轻。所以只要把拉杆往更轻的方向移动,小鼠就能得到水奖励。对吧?然后他把拉杆的硬度设成了四种不同的水平。有时候小鼠直接推或拉就成功拿到奖励,而在有些情况下,它一开始去推,但发现方向不对,于是改变方向,最后才拿到奖励。对吧?然后他还改变了「哪个方向会给奖励」(推还是拉)的先验概率,于是动物会根据这个先验分布来改变最初的推动方向���同时行为也依赖于实际的阻力设定,但同样受到这个先验设定的影响,对吧?
便签引用
44:29
And then while the animal is doing this behavior, uh he recorded activities of neurons in the uh smart sensory cortex and the motor cortex using this uh prism lens. So it can monitor the activities of neurons from the side sideway in the cortex uh in the like a deep layer and then the superficial layer simultaneously, right? And then we can use the uh certain molecular marker of the neurons to uh decide whether this is a a superficial layer or deep layer. Right? And then uh we are now analyzing quite large database of such recording. For example, this is the uh data from six animals about like a 9,000 neurons sorted by the depths in the cortical areas for the push and uh trial and the push dental pro or in a different uh setting. Right?
然后在动物做这个行为任务的时候,他用棱镜透镜记录了体感皮层和运动皮层中神经元的活动。这样就能从侧面监测皮层中深层和浅层神经元的活动,而且是同时记录的,对吧?然后我们可以利用神经元的某些分子标记来判断它属于浅层还是深层,对吧?现在我们正在分析这样一个相当大的记录数据库。比如说,这是来自六只动物、大约 9000 个神经元的数据,按皮层的深度排序,分别对应推的试次和不同设定下的情况,对吧?
便签引用
45:35
So uh there are kind of a different activities in different settings. And then uh uh our uh initial uh analysis showed that uh before even the movement start uh neurons in both in the small sensory cortex and motor cortex makes certain activities and then uh we found that deep layer of the uh centric cortex has more uh coding for the prior uh uh setting of the uh uh movement direction, right? And also the by performing a decoding analysis of these neural activities, we found that the deep player neurons in the small sensory cortex has the executive trial type or prediction of the forthcoming uh uh choice, right? So, but we still need to do a lot of more uh analysis and then uh if you are interested in this kind of analysis of a data set, please let me know. So we want were hoping to try different approaches for this analysis.
所以在不同设定下会有不同的活动模式。我们最初的分析显示,甚至在运动开始之前,体感皮层和运动皮层中的神经元就已经有了某些活动。而且我们发现,感觉皮层的深层对运动方向的先验设定有更多的编码,对吧?另外,通过对这些神经活动做解码分析,我们发现体感皮层的深层神经元携带了当前试次类型的信息,或者说对接下来选择的预测,对吧?不过我们还需要做更多的分析。如果你对这类数据集的分析感兴趣,请告诉我。我们希望能对这些分析尝试不同的方法。
便签引用
11数字大脑、具身进化与 AI 安全
46:46
Okay. Okay. And then finally the the AI and neuroscience. So uh in Japan so there has been uh a major neuroscience project called brain minds that focused on the uh mammoset the brain data uh uh gathering analysis and then from uh two years ago we started stepped into the second phase called the brain mind 2.0 zero and then important feature of this project is that uh in in gathering different kinds of data they are kind of a combined into so-called digital brain for the better understanding of the brain function and the possible application of the therapeutic applications right so we are building this uh like a digital brain platform so uh and then uh users can access this uh uh cloud-based server and then use the different computing and the data server and connect to the different databases and and we are developing uh tool called neuro workflow uh in which different models uh represent as nodes and you can pick different nodes and connect them for different kinds of uh simulation and
好。那么最后讲讲 AI 和神经科学。在日本,有一个叫「Brain/MINDS」的大型神经科学项目,聚焦于狨猴的脑数据采集和分析。从两年前开始,我们进入了第二阶段,叫做 Brain/MINDS 2.0。这个项目的一个重要特点是,在收集各种不同数据的同时,把它们整合成所谓的「数字大脑」,以便更好地理解脑功能,以及可能的治疗应用,对吧。所以我们正在搭建这个数字大脑平台。用户可以访问这个基于云的服务器,使用不同的计算和数据服务器,连接到不同的数据库。我们还在开发一个叫 Neuro Workflow 的工具,其中不同的模型被表示为节点,你可以挑选不同的节点并把它们连接起来,做各种不同的仿真和
便签引用
48:09
data analysis. Right? And also uh important feature of this new workflow is that this is a fully combined with the uh AI agents. So uh you can build a models by doing like a conversation with the AI agents. So uh it knows what kind of previous models register in this system and depending on the subjects the users uh questions it can find the right nodes and then connect them appropriately uh to uh uh run the simulation and you can uh uh ask the agent to customize the model parameters and structure for different purposes of your uh analysis.
数据分析,对吧。而且这个新工作流的一个重要特点是,它和 AI 智能体完全结合在一起。所以你可以通过跟 AI 智能体对话来构建模型。它知道系统里已经注册了哪些既有模型,并且根据用户提出的问题,找到合适的节点,把它们恰当地连接起来,运行仿真。你还可以让智能体针对你分析的不同目的,来定制模型的参数和结构。
便签引用
48:53
So uh I we hope this kind of a new system will u lower the uh like a threshold for people working on advanced modeling studies uh for neuroscience right and uh also we run a project called correspondence fusion of AI and brain science. It was started in 2016 and then the in that project uh one of the major question we addressed was the uh in addition to the deep learning and the reinforce learning what should we further learn from brain about the next uh uh development of the AI. So there are different uh uh uh issues uh for example the energy efficiency uh is a very important topic. uh so today's AI programs consume a lot of power which is causing environmental concerns whereas the our brain is supposed to consume just 20 watt of energy right uh and also the like a data efficiency is a very important issue so the Brendan's like a review article was quite influential in this topic so why the human can learn uh behaviors and skills relatively small number of trials. Uh that is very important topic and then I
我们希望这样的新系统能降低人们在神经科学中做高级建模研究的门槛,对吧。另外我们还在做一个叫「AI 与脑科学的对应与融合」的项目,从 2016 年开始。在那个项目里,我们讨论的一个主要问题是:除了深度学习和强化学习之外,为了 AI 的下一步发展,我们还应该从大脑那里学到什么。这里有好几个议题。比如说,能效就是一个非常重要的话题。今天的AI 程序消耗大量电力,引发了环境方面的担忧,而我们的大脑据说只消耗大约 20 瓦的能量,对吧。另外数据效率也是一个非常重要的问题。Brenden(Lake)他们那篇综述文章在这个话题上很有影响力:为什么人类能在相对少的尝试次数中学会行为和技能。这是个非常重要的话题。然后我觉得另一个重要话题是自主性,还有社会性意识,
便签引用
50:22
think another important topic is uh like autonomy and also social awareness right. So uh when we started to uh think about the tuning of the temporal discounting parameter. So uh we uh came up with a faced the problem how we can justify the uh tuning the temporal discounting parameter which is a part of the objective function. So probably the uh in addition to this reward and parameters there should be some kind of higher level goal from which reinforcement learning setting can be derived. So that uh uh led us to create robots like this. So uh these robots uh capture the battery pack uh to recharge itself for survival. Uh and also for the reproduction. So uh these robots has a infrared communication port. So by facing each other so they can copy their program for software reproduction [sighs] and then on top of these uh setup uh we implemented the embodied evolution uh parame framework. So we could indeed show that these robots evolve a different kinds of a visual reward for finding a battery and finding a face of
对吧。当我们开始思考时间折扣参数的调节时,我们就面临一个问题:我们要如何为调节时间折扣参数这件事提供依据?它是目标函数的一部分。所以除了奖赏和这些参数之外,大概还应该存在某种更高层次的目标,强化学习的设定可以从中推导出来。这促使我们做出了这样的机器人。这些机器人要抓取电池包来给自己充电以求生存。同时也为了繁殖。这些机器人有红外通信端口,所以只要面对面,它们就可以复制自己的程序,实现软件层面的繁殖。在这套装置之上,我们实现了「具身进化」(embodied evolution)的框架。我们确实能证明这些机器人进化出了不同类型的视觉奖赏——用于寻找电池,以及寻找另一个机器人的「脸」。它们甚至
便签引用
51:47
another robot. And then they even developed some kind of polymorphism whether to go for foraging or go for mating behaviors. And then uh my former postto Stephan found that these robots have a different genetic features uh behind this and different behaviors. Right. So uh uh this shows that by combining reinforce learning with evolutionary framework robots themselves find out their own uh appropriate reload function for the survival and reproduction. And more recently the my student Yuzi uh is uh doing something similar in the virtual environment with the model of the survival reproduction of these uh creatures. And then uh uh he evolved uh all the the reward function for example the reward for the acquisition of food and reward for the actions right. So uh and uh uh the reward for the food is usually positive after evolution and we found expect that the reward for action could be a cost for this should be negative but in some of the robots uh creatures this was a positive so it kind of motivated to move
发展出了某种多态性:是去觅食,还是去交配。然后我以前的博士后 Stefan 发现,这些机器人在不同行为背后有着不同的遗传特征。对吧。这说明,把强化学习和进化框架结合起来,机器人自己就能找出适合自己的奖赏函数,用于生存和繁殖。最近,我的学生 Yuzhi 在虚拟环境中做了类似的事情,用的是这些生物体的生存繁殖模型。他进化出了所有的奖赏函数,比如获取食物的奖赏,以及行动本身的奖赏,对吧。进化之后,食物的奖赏通常是正的,而我们本来预期行动的奖赏应该是一种代价、应该是负的,但在有些生物体身上它却是正的,也就是说它们被激励去到处移动。对于像我这样把大量精力
便签引用
53:06
around. That's a very interesting result for an athlete like me spending a lot of energy for training and races. Right. Okay. And uh uh also the our reward system may not be just uh directly linked with the survival reproduction. So uh acquisition of uh information is also a very important uh uh feature. So uh and and so-called intrinsic reward uh is a important and then my student tojo worked on the uh use of different three different kinds of intrinsic reward nobility surprise and empowerment. So and then the uh he could show that uh by appropriate combination of these three uh uh intelligent reward the uh agent can learn to uh behave in like a complex environment you not only in the grid world but also like a craft like a uh environment right. Yeah. And also another student uh Ted uh came up with a utilize these curiosity based exploration for the agent to learn like a like compositional use of language for the behaviors. Right. Okay. So uh and then uh we have been uh trying to uh make uh these uh uh reinforcement
花在训练和比赛上的运动员来说,这是个非常有意思的结果,对吧。好。另外,我们的奖赏系统可能并不只是直接和生存繁殖挂钩。获取信息也是一个非常重要的特征。所谓的内在奖赏(intrinsic reward)也很重要。我的学生 Tojo 研究了三种不同的内在奖赏的使用:新颖性、惊奇度和赋能(empowerment)。然后他能够证明,通过这三种内在奖赏的恰当组合,智能体可以学会在复杂环境中行动,不只是在网格世界里,也包括类似 Minecraft 那样的环境,对吧。是的。还有另一位学生 Ted,他想到利用这种基于好奇心的探索,让智能体学会对语言的组合式使用来完成行为,对吧。好。然后我们一直在尝试让这些强化学习智能体更加自主,能够发展出自己的奖赏函数。但这也带来了
便签引用
54:40
learning agent more autonomous being able to develop the its own reward functions. But that is also pose like safety concerns right. So that if the reform can find its own goals uh uh and try them out, it can uh create no science and technology or culture and even industry. But so we need to also be careful about the assessment of the uh the the side effects or like overance and also the exploration by the uh uh different individuals or countries for the military or terrorist use. Right. And then the uh when we think of the uh safety and the danger of the uh AI agents, so I think it's very important to learn from the human societies because the humans are proven to be very dangerous species. So we have brought many species already to extinction and ourselves are near the border of extinction by the nuclear war and also the environmental uh destructions.
安全方面的顾虑,对吧。因为如果它能找到自己的目标并去尝试,它就可能创造出新的科学技术、文化,甚至产业。但我们也需要谨慎评估副作用,或者说过度的问题,还有不同个人或国家把它用于军事或恐怖用途的可能,对吧。当我们思考 AI 智能体的安全和危险时,我认为很重要的一点是向人类社会学习,因为人类已被证明是非常危险的物种。我们已经让很多物种灭绝,而我们自己也因为核战争和环境破坏而濒临灭绝的边缘,
便签引用
55:51
Right? So but uh human history also developed some mechanism to avoid the uh danger. I think the democracy is one of the important invention by the humans to uh limit the over concentration of power. So either in politics or economy and also in science that we imple so that we can uh uh update the kind of classic theory right. So uh and then the probably in the near future of the society where the human and the AI agents live together maybe peer reviewing of among open source the experime agent would be very important right and then the in planning such a kind of next generation AI I think uh uh there could be some uh contribution by neuroscience.
对吧?但人类历史上也发展出了一些避免危险的机制。我认为民主是人类重要的发明之一,用来限制权力的过度集中——无论是在政治还是经济上,在科学中我们也有类似机制,这样我们才能更新那些经典理论,对吧。所以在不远的将来,人类和 AI 智能体共同生活的社会里,也许开源智能体之间的同行评审会非常重要,对吧。而在规划这样的下一代 AI 时,我想神经科学也许能做出一些贡献。
便签引用
56:56
So this is a paper by Harosan who is around there. So uh uh about the uh pro-social uh mechanism of the brain, right? So in the choice task comparing the self reward versus other reward some people prefer equal division, some people just want its own reward. some people maybe even more competitive. So and then the classic idea about the brain was the old brain structure like amigdala would be like animal like and then uh more recent brain like a prefrontal cortex makes people more rational but the finding is that propos pro-social people have activities more activist in the amydala and then the prefrontal cort activities allow people to be more selfish right so uh and then similar finding was observed in also the uh development of the volume of the uh ama and prefrontal cortex. So in today's large language models like ethical uh concerns are implemented in a postto focing but I think it's also very important to include such a social awareness from the early phase of the training like uh we have in our best
这是 Haroshi(就在那边)的一篇论文,讲的是大脑的亲社会机制,对吧?在一个比较「自己的奖赏」和「他人的奖赏」的选择任务里,有些人偏好平均分配,有些人只想要自己的奖赏,有些人甚至更具竞争性。关于大脑的经典观点是,像杏仁核这样古老的脑结构比较「动物性」,而像前额叶皮层这样较新的脑区让人更理性。但研究发现,亲社会的人在杏仁核里的活动反而更强,而前额叶皮层的活动却让人更自私,对吧。在杏仁核和前额叶皮层的体积发育上也观察到了类似的发现。所以在今天的大语言模型里,伦理方面的顾虑是通过后训练来实现的,但我认为,从训练的早期阶段就把这种社会性意识包含进来也非常重要,就像我们的杏仁核系统那样,对吧。好,那么
便签引用
58:20
amdara system right okay so then the uh finally so I think uh the uh the neural basis democracy may be a very import interesting topic of neuroscience. We planned such a workshop earlier. Uh unfortunately uh this was postponed because the typhoon hit Okinawa on that day. But uh uh if you are interested in this kind of topic uh please let me know. We hope to uh uh start this kind of a research initiative. Okay. So uh it's taking uh time. Finally, I want to uh thank my colleagues contribute to this process and we are now hiring a new faculty member including the field of neuroscience and theoretical and air approaches. So yeah, thank you very much for attention. [applause]
最后,我认为民主的神经基础可能是神经科学里一个非常有意思的话题。我们之前计划过这样一个研讨会,但很遗憾因为那天台风登陆冲绳而延期了。不过,如果你对这类话题感兴趣,请告诉我。我们希望能启动这样的研究计划。好,时间有点超了。最后,我想感谢在这个过程中做出贡献的同事们。我们现在正在招聘新的教员,包括神经科学以及理论和 AI 方向。好,非常感谢大家的聆听。[掌声]
便签引用
12问答:无模型是否暗含模型
59:26
>> Okay, maybe A couple questions.
>> 好,也许可以提几个问题。
便签引用
59:44
>> Hey, Kenji. >> Okay. >> Can you hear me? >> Hi. >> Thank you so much for the great talk. I really enjoyed it. This is not my area of expertise at all, but I was wondering for um for pure temporal difference learning to have a good policy. >> Mhm. >> Don't you have to have some sort of model of the causal structure of the world, how the world evolves with your actions? So I'm wondering if it if it's really like a clear-cut distinction between model based and model free because if you do model free really well maybe you have to have some kind of model.
>> 嘿,Kenji。>> 好的。>> 能听到我说话吗?>> 你好。>> 非常感谢你精彩的演讲,我很喜欢。这完全不是我的专业领域,但我想问一下,对于纯粹的时序差分学习来说,要想学到一个好的策略——>> 嗯。>> 你难道不需要有某种关于世界因果结构的模型吗?也就是世界如何随着你的动作而演变?所以我在想,基于模型和无模型之间是不是真的有那么泾渭分明的区别,因为如果你把无模型做到极致,也许你就必须拥有某种模型。
便签引用
1:00:24
>> Yeah. So uh uh in addition to the uh whether to use the state transition model or not. So what state variables you uh consider uh such a kind of a uh framework is also very important. Right. So uh when animal is behaving in the real environment uh which of the high dimension sensor input is relevant for the certain behavior is something uh very very important. So uh yeah and once uh the uh agent finds the right features to base their behavior the uh learning can be very efficient. So uh yeah so I didn't talk about that in today's uh uh presentation but uh this finding out the right state space for the action and then decision I think uh also a very important issue in uh practical implementation of reinforcement learning right yeah does that answer your question yeah >> so do you think the distinction between model based and model free remains like a useful distinction Uh yeah, I think it it is still useful, right? Yeah. Right. In addition to some other factors, right? Yeah. Yeah.
>> 是的。所以呃,除了要不要使用状态转移模型之外,你考虑哪些状态变量这样一个框架也非常重要,对吧。所以呃,当动物在真实环境中行动时,高维感官输入中的哪一部分与某个特定行为相关,这是非常非常重要的。所以呃,是的,一旦智能体找到了正确的特征来作为行为的基础,学习就会非常高效。所以呃,是的,我今天的演讲里没有谈到这一点,但是呃,为动作和决策找出正确的状态空间,我认为这在强化学习的实际实现中也是一个非常重要的问题,对吧。是的,这回答了你的问题吗?是的。>> 那你认为基于模型和无模型之间的区分仍然是一个有用的区分吗?呃,是的,我认为它仍然是有用的,对吧?是的。对。除了其他一些因素之外,对吧?是的。是的。
便签引用
1:01:51
>> Yeah. Thanks for the great talk. Um I'm over here. Um, yeah. So, I guess I was wondering if you have thoughts on the role of being able to take actions and getting a feedback loop on your ability to learn a world model as opposed to like unsupervised approaches that are like sort of predicting what should come next without necessarily having the ability to intervene or act on the world. >> Could you repeat that last point again? Yeah. >> Yeah. So, um I'm wondering if your thoughts on the role of being able to act or intervene on the world um in your ability to learn a good model of the world as opposed to um as I guess like a unsupervised paradigm where you're just purely observing a sequence of states in the world and just trying to predict what comes next.
>> 是的。谢谢你精彩的演讲。呃,我在这边。呃,是的。我想问的是,你怎么看待能够采取行动、获得反馈闭环这件事,对你学习世界模型的能力所起的作用,相对于那种无监督的方法——那种只是预测接下来会发生什么、而不一定具备干预或作用于世界的能力的方法。>> 你能再重复一下最后那一点吗?好的。>> 好的。所以,呃,我想问的是,你怎么看待能够作用于世界、干预世界这件事,在你学习一个好的世界模型的能力中所起的作用,相对于呃,我想是那种无监督的范式——你只是纯粹地观察世界中的一串状态,然后试图预测接下来会发生什么。
便签引用
1:02:38
>> Yeah. So your question is about the uh uh the usefulness of the uh model based strategies uh >> uh the the ability to take actions in order to learn a model as opposed to just purely observing um and predicting without being able to act. >> Uh so uh learning by observation or learning by acting. Yeah. Yeah. Yes. So learning by acting so you can uh test your hypothesis right. So uh I think uh the active learning would be a benefit of be able to like select the samples by yourself just rather than just by observing the other agents behaving. Right. So I'm not sure if I'm answering your question correctly. Yeah.
>> 是的。所以你的问题是关于呃,基于模型的策略的用处——呃 >> 呃,是为了学习一个模型而能够采取行动的能力,相对于只是纯粹地观察和预测、却无法行动。>> 呃,所以呃,通过观察学习还是通过行动学习。是的。是的。是的。那么通过行动学习,你就可以检验你的假设,对吧。所以呃,我认为主动学习的一个好处就是能够自己选择样本,而不只是通过观察其他智能体的行为。对吧。所以我不太确定我是不是准确回答了你的问题。是的。
便签引用
1:03:38
>> Um, yeah. Thanks. >> Yeah. Thank you. >> Okay. All right. >> Thank you so much. >> Okay. Thank you very much for your questions. [applause] >> Right. Sorry, it's a bit longer. >> Thank you.
>> 呃,是的。谢谢。>> 是的。谢谢你。>> 好的。好的。>> 非常感谢。>> 好的,非常感谢你们的提问。[掌声]>> 好的。抱歉,稍微超时了一点。>> 谢谢。
便签引用
视频总结 · 一句话概括与核心要点

一句话概括

Kenji Doya 在 CCN 2026 主题演讲中回顾了 40 年强化学习研究:从基底神经节实现价值函数与多巴胺 TD 误差,到血清素调控"决策与学习可用时间"的修正理论,再到心智模拟、推理与控制的对偶性,最后提出 AI 应从大脑借鉴能效、数据效率、自主性与社会意识。

核心要点

  • 强化学习是实验室的统一主线,机器人与啮齿类实验双向互证。 一侧用机器人和机器学习构建灵活学习系统,另一侧用大鼠、小鼠研究脑机制。一侧的突破为另一侧提供线索,一侧撞墙则为另一侧提供理论或实验问题。Doya 本科时用 8 位电脑做出 10 厘米长的学步机器人,以旋转编码器测得的速度作为奖励信号,是其起点。
  • 参数调节比算法收敛更决定实际结果。 起立机器人实验中,折扣因子 gamma 设得太小,机器人只会坐起拿即时小奖励。惩罚设得过大,机器人干脆躺平不动以规避惩罚。算法在理论上能收敛,但奖励设计和参数设置稍有偏差结果就背离预期。
  • 基底神经节纹状体中存在编码动作价值和状态价值的神经元。 大鼠二选一任务中奖励概率持续变化,用强化学习模型拟合选择序列估出动作价值,发现部分神经元在约 200 次试验中的活动变化与左侧或右侧动作价值高度吻合。后续钙成像显示投射到黑质多巴胺神经元的纹状体细胞在条件化后对即将到来的奖励产生响应,符合状态价值假说。
  • 血清素最初被假设为折扣因子,但数据迫使理论升级。 光遗传激活血清素神经元后,小鼠在 9 秒延迟任务中放弃率下降,在无限延迟任务中等待时间从约 10 秒大幅延长。但奖励概率从 75% 降到 25% 时刺激几乎无效,即便补足期望值也无效。奖励时机固定时效应很弱,时机越不确定效应越强。新模型认为血清素表征奖励时机的先验概率,而非单纯折扣参数。
  • 抑制血清素会削弱基于模型的决策成分。 在小鼠版两步任务中,光遗传抑制血清素活动后,行为中的模型依赖成分变弱。综合两组实验,Doya 现在认为血清素调控的是"为决策和学习投入多少时间"这一更宽泛的维度。
  • 心智模拟依赖全脑分工:小脑提供状态转移模型,基底神经节评估候选动作。 提出小脑做监督学习、基底神经节做强化学习、皮层做无监督表征学习的框架。fMRI 网格导航任务限制光标只能走三个方向,被试必须靠预先学到的动作依赖转移模型规划之字形路径。手指动作前的规划期,除顶叶、运动前区和前额叶外,小脑和基底神经节也被激活。
  • 顶叶皮层的群体编码在感觉反馈中断时仍向目标推进。 双光子钙成像记录小鼠听觉虚拟环境中数百个顶叶神经元,解码目标距离的后验分布。关闭声音反馈后,神经编码的距离仍朝目标移动,反馈恢复后不确定性收窄。这与用内部动力学模型做动态贝叶斯推理的行为一致。
  • 推理与控制的对偶性可用于假设皮层微环路的功能分工。 卡尔曼滤波与贝尔曼最优控制方程结构相同,机器学习界由此发展出"控制即推理",Soft Actor-Critic 即源于此。感觉皮层和运动皮层共享六层经典微环路,若感觉皮层实现动态贝叶斯推理,运动皮层可能以同构方式实现强化学习。
  • 推拉杆实验初步显示深层神经元编码先验和即将做出的选择。 小鼠推拉杆,杆的阻力有四档且奖励方向的先验概率可调。用棱镜透镜同时记录体感皮层和运动皮层的浅层与深层,6 只动物约 9000 个神经元。运动前活动已出现,体感皮层深层对先验设置的编码更强,解码显示深层神经元携带即将做出选择的预测。数据集尚待更多分析,公开征集合作。
  • 奖励函数本身应由更高层目标演化而来,而非人工设定。 为回答"凭什么手动调折扣因子",实验室做了能抓电池充电、靠红外端口互传程序繁殖的机器人,具身演化让它们自发演化出找电池和找同伴脸的视觉奖励,并出现觅食型与求偶型的多态性。虚拟环境版本中演化出的动作奖励在部分个体里为正而非负,即天生爱动。
  • AI 应从大脑学能效、数据效率、自主性和社会意识。 大脑功耗约 20 瓦,而现有 AI 耗能引发环境担忧。自主生成目标的智能体也带来安全风险,Doya 主张借鉴人类社会的民主机制以防止权力过度集中,建议开源智能体之间互相同行评审。

结论与值得注意的细节

  • 亲社会性来自杏仁核而非前额叶。 引用的研究发现,在自我与他人奖励分配任务中偏好平等的人杏仁核活动更强,前额叶活动反而使人更自私,体积发育数据也一致。Doya 据此批评当前大模型把伦理放在后训练阶段,主张像杏仁核那样从训练早期植入社会意识。
  • 意识可能对应数据同化。 与普林斯顿的 Piet Hut 合作提出:古人用神解释宇宙、用灵魂解释生命,今人用意识解释智能。气象预报中的数据同化结合动力学模型与多源观测做前向预测和事后回推,这与全局工作空间理论描述的意识功能相似。
  • 日本 Brain/MINDS 2.0 正在建"数字大脑"平台。 云端 NeuroWorkflow 工具将模型作为节点连接,并集成 AI 智能体,用户可通过对话组装、定制模型参数和结构,目标是降低高级建模研究的门槛。
  • 问答环节的两个立场。 关于"做好 model-free 是否也隐含世界模型",Doya 认为区分仍然有用,但补充说找到正确的状态空间和相关特征同样关键,这是他演讲中未展开的话题。关于主动干预与纯观察学习的对比,他认为主动行动能检验假设并自主选择样本,是被动观察不具备的优势。
  • 计划中的"民主的神经基础"研讨会因台风袭击冲绳而推迟,实验室正在招聘神经科学、理论与 AI 方向的教员。演讲者本人 2024 年完成 Kona 铁人三项世界锦标赛,主持人调侃他对长时程的耐受不只是理论上的。
核心句型 · 10
1. If you have ever thought of X as Y, you're working downstream of his ideas.
“If you have ever thought of a neurom modulator as a parameter in a learning algorithm, you're working downstream of his ideas.”
用「如果你曾把 X 当作 Y」把听众拉进来,再用 downstream of 说明其学术渊源。适合介绍某人影响力,可仿写:If you've ever used X, you're benefiting from her work.
2. …, which suggests that … is not purely theoretical.
“He also finished the iron man world championship in Kona in 2024 which suggests that his tolerance for long horizons is not purely theoretical.”
非限制性定语从句 which suggests that 用来从事实引出轻松的推论,not purely theoretical 制造反差幽默。介绍、致辞里常用。
3. Once you have … in one side, maybe that will give a hint to the other; and if you hit the wall in one side, maybe that gives … for the other.
“Once you have some progress in one side maybe that will give a hint to the other and if you hit the wall in one side maybe that gives a theoretical or experimental problem setting for the other”
用 once / if 引导的两个对称条件句,说明两条研究线互相支撑。适合阐述交叉学科的理由,仿写时保持 one side / the other 的对照。
4. If …, A and B should cancel out; otherwise you learn from the error.
“So if your prediction is perfect these first two term and last term should cancel out otherwise you learn from the this prediction error.”
should 表示理论上应然,otherwise 承接反面情况。解释算法或公式的「正常情况/异常情况」时很好用。
5. …, which is consistent with the hypothesis that …
“So the animal became more patient or more motivated for the delayed reward. So which is consistent with the hypothesis that certainly regulates the temporal discounting parameter.”
报告实验结果与假说关系的标准句式,语气克制(一致而非证明)。学术写作里可替换为 in line with / supports。
6. We initially assumed that …, but the real data was much more intriguing.
“We initially assumed that the serotonin would be simply manipulating the temporal discounting parameters but the real data was much more like intriguing”
initially assumed … but 是叙述理论修正的经典转折框架,后接 which was not simply explained by … 说明为何要换模型。讲研究故事时用。
7. To explain X, people called for Y; to explain Z, people called for W.
“To explain the mechanism universe people called for the gods for the explanation and to explain the life people called for the soul”
平行结构铺垫类比,最后落到 today … many people call for consciousness。三段排比再发问 So do we continue to do that or not? 是演讲中制造悬念的手法。
8. This kind of possibility motivated us to start a new experiment.
“So this kind of a possibility motivate us to start a new experiment.”
motivate sb to do 把理论推测过渡到实验,是学术演讲各节之间的常用衔接句。可换 led us to / prompted us to。
9. as … became more …, the effect became much stronger
“As the timing became more varied, effect of a certain simulation become much stronger”
as 引导的比较级联动句,表达两个变量同向变化。描述剂量、参数扫描结果时极常用。
10. I'm not sure if I'm answering your question correctly.
“So I'm not sure if I'm answering your question correctly. Yeah.”
问答环节礼貌收尾,给对方留出追问空间。类似还有 Does that answer your question?
词汇精讲 · 129 · 按出现顺序
keynote lecture n. phr. 0:08
主旨演讲(会议中最重要的特邀报告)
downstream of phr. 0:08
处在……的下游;受……思想影响而展开工作
Dopamine /ˈdoʊpəmiːn/ n. 0:08
多巴胺(神经递质,与奖赏预测误差相关)
serotonin /ˌserəˈtoʊnɪn/ n. 0:08
血清素(神经调质,本讲核心研究对象)
adrenaline /əˈdrenəlɪn/ n. 0:08
肾上腺素;文中 nor adrenaline 指去甲肾上腺素
long horizons n. phr. 1:04
长时间跨度(此处双关:强化学习的长时程与铁人三项的耐力)
join me in welcoming phr. 1:04
和我一起欢迎(介绍讲者的固定说法)
reinforcement learning n. phr. 2:17
强化学习(通过奖赏反馈学习行动策略的框架)
agent /ˈeɪdʒənt/ n. 2:17
智能体(能感知环境并采取行动的主体)
action policy n. phr. 3:37
行动策略(从状态到动作的映射)
hit the wall idiom 3:37
碰壁;陷入瓶颈
in parallel phr. 3:37
并行地;同时推进
undergraduate thesis n. phr. 4:54
本科毕业论文/毕业设计
rotary encoder n. phr. 4:54
旋转编码器(测量转速或角位移的传感器)
genetically predetermined adj. phr. 5:57
由基因预先决定的
food dispenser n. phr. 5:57
喂食器;自动投食装置
value function n. phr. 5:57
价值函数(对未来累计奖赏的预测)
temporal discounting n. phr. 7:14
时间折扣(对远期奖赏打折的程度)
greedy /ˈɡriːdi/ adj. 7:14
贪婪的;此处指总选当前期望最高动作的策略
stoastic /stoʊˈkæstɪk/ adj. 7:14
随机的(stochastic 的误转写)
inverse temperature n. phr. 7:14
逆温度(softmax 中控制选择确定性的参数)
deterministic /dɪˌtɜːrmɪˈnɪstɪk/ adj. 7:14
确定性的(与随机相对)
cancel out phr. v. 8:16
相互抵消
in proportion to phr. 8:16
与……成比例
learning rate n. phr. 8:16
学习率(每次更新的步长)
approximation /əˌprɑːksɪˈmeɪʃn/ n. 8:58
逼近;近似
tradeoff /ˈtreɪdɔːf/ n. 8:58
权衡;此消彼长的取舍
tune /tuːn/ v. 9:38
调节(参数);调优
discriminate /dɪˈskrɪmɪneɪt/ v. 10:38
分辨;区分(此处无歧视义)
gyro sensor n. phr. 10:38
陀螺仪传感器
delayed reward n. phr. 10:38
延迟奖赏
temporal credit assignment n. phr. 11:42
时间信用分配(判断哪一步动作该为结果负责)
converge /kənˈvɜːrdʒ/ v. 11:42
收敛(算法趋于稳定解)
basa ganglia /ˈbeɪsl ˈɡæŋɡliə/ n. 12:54
基底神经节(basal ganglia 的误转写)
nuclei /ˈnuːkliaɪ/ n. pl. 12:54
核团(nucleus 复数,脑内神经元聚集体)
plasticity /plæˈstɪsəti/ n. 12:54
可塑性(突触强度可改变的性质)
binary choice task n. phr. 14:48
二选一任务
keep track of phr. 14:48
持续追踪;掌握……的动态
calcium imaging n. phr. 16:02
钙成像(用荧光指示剂观测神经元活动)
consistent with phr. 16:02
与……一致;不矛盾
brain stem n. phr. 17:24
脑干
speculate /ˈspekjuleɪt/ v. 17:24
推测;猜想
optogenetic /ˌɑːptoʊdʒəˈnetɪk/ adj. 17:24
光遗传学的(用光控制特定神经元)
genetically engineered adj. phr. 17:24
经基因改造的
control condition n. phr. 17:24
对照条件
abandons /əˈbændənz/ v. 18:37
放弃(等待)
unresolved /ˌʌnrɪˈzɑːlvd/ adj. 19:22
未解决的
commonality /ˌkɑːməˈnæləti/ n. 19:22
共性;共同点
model free adj. phr. 20:32
无模型的(不学习环境转移模型)
reactive /riˈæktɪv/ adj. 20:32
反应式的;被动响应的
deliberative /dɪˈlɪbərətɪv/ adj. 20:32
审慎的;经过深思规划的
offline /ˌɔːfˈlaɪn/ adv. 21:50
离线地(不与真实环境交互而在内部模型中计算)
data efficiency n. phr. 21:50
数据效率;样本效率
state transition model n. phr. 22:39
状态转移模型(预测动作如何改变状态)
occluded /əˈkluːdɪd/ adj. 22:39
被遮挡的
thought experiment n. phr. 23:55
思想实验
supervised learning n. phr. 24:37
监督学习(有正确答案作为误差信号)
representation learning n. phr. 24:37
表征学习
zigzag /ˈzɪɡzæɡ/ adj./n. 26:45
之字形的;曲折路线
upwind /ˌʌpˈwɪnd/ adv. 26:45
逆风地
behaviorally /bɪˈheɪvjərəli/ adv. 26:45
从行为层面上
prefrontal cortex n. phr. 28:00
前额叶皮层
interpretation /ɪnˌtɜːrprəˈteɪʃn/ n. 28:00
解释;解读
neural decoding n. phr. 28:58
神经解码(从神经活动反推所表征的变量)
instantaneous /ˌɪnstənˈteɪniəs/ adj. 28:58
瞬时的
posterior probability n. phr. 28:58
后验概率(结合观测后的概率)
intriguing /ɪnˈtriːɡɪŋ/ adj. 30:23
耐人寻味的;引人好奇的
reward omission n. phr. 30:23
奖赏省略(本该给却不给奖赏的试次)
equate /ɪˈkweɪt/ v. 30:23
使相等
prior probability n. phr. 32:01
先验概率
dominant /ˈdɑːmɪnənt/ adj. 32:01
占主导的
suppress /səˈpres/ v. 33:32
抑制
generalized /ˈdʒenrəlaɪzd/ adj. 34:28
一般化的;推广了的
astrophysicist /ˌæstroʊˈfɪzɪsɪst/ n. 35:28
天体物理学家
called for phr. v. 35:28
求助于;诉诸(作为解释)
data assimilation n. phr. 35:28
数据同化(将观测融入仿真模型以校正状态)
observatory /əbˈzɜːrvətɔːri/ n. 36:51
观测站;天文台
modality /moʊˈdæləti/ n. 37:58
(感觉)模态,如视觉、听觉
duality /duˈæləti/ n. 38:54
对偶性;二元对应
semantics /sɪˈmæntɪks/ n. 38:54
语义;此处指数学相似性背后的实质含义
initial guess n. phr. 39:40
初始猜测;迭代起点
boundary condition n. phr. 39:40
边界条件
merit /ˈmerɪt/ n. 40:46
优点;价值
neo cortex /ˌniːoʊˈkɔːrteks/ n. 40:46
新皮层
canonical /kəˈnɑːnɪkl/ adj. 41:47
典型的;标准范式的
microcircuit /ˈmaɪkroʊˌsɜːrkɪt/ n. 41:47
微回路(局部神经元连接结构)
anatomical literature n. phr. 41:47
解剖学文献
lever /ˈlevər/ n. 43:04
拉杆;杠杆
stiffness /ˈstɪfnəs/ n. 43:04
刚度;硬度(此处指拉杆阻力)
prior distribution n. phr. 43:04
先验分布
prism lens n. phr. 44:29
棱镜透镜(可侧向成像的显微镜部件)
superficial layer n. phr. 44:29
浅层(皮层表层)
molecular marker n. phr. 44:29
分子标记
therapeutic /ˌθerəˈpjuːtɪk/ adj. 46:46
治疗的
cloud-based /ˈklaʊd beɪst/ adj. 46:46
基于云的
customize /ˈkʌstəmaɪz/ v. 48:09
定制
threshold /ˈθreʃhoʊld/ n. 48:53
门槛;阈值
energy efficiency n. phr. 48:53
能效
influential /ˌɪnfluˈenʃl/ adj. 48:53
有影响力的
autonomy /ɔːˈtɑːnəmi/ n. 50:22
自主性
justify /ˈdʒʌstɪfaɪ/ v. 50:22
为……提供依据;证明合理
objective function n. phr. 50:22
目标函数
infrared /ˌɪnfrəˈred/ adj. 50:22
红外的
embodied evolution n. phr. 50:22
具身进化(在真实机器人群体中运行进化算法)
polymorphism /ˌpɑːliˈmɔːrfɪzəm/ n. 51:47
多态性(同一群体中并存多种形态或策略)
foraging /ˈfɔːrɪdʒɪŋ/ n. 51:47
觅食
mating /ˈmeɪtɪŋ/ n. 51:47
交配
intrinsic reward n. phr. 53:06
内在奖赏(不来自外部任务的奖赏,如好奇心)
empowerment /ɪmˈpaʊərmənt/ n. 53:06
赋能(衡量智能体对未来的影响能力的内在奖赏)
compositional /ˌkɑːmpəˈzɪʃənl/ adj. 53:06
组合式的(由部件按规则组合)
autonomous /ɔːˈtɑːnəməs/ adj. 54:40
自主的
side effects n. phr. 54:40
副作用
extinction /ɪkˈstɪŋkʃn/ n. 54:40
灭绝
concentration of power n. phr. 55:51
权力集中
peer reviewing n. phr. 55:51
同行评审
pro-social /ˌproʊ ˈsoʊʃl/ adj. 56:56
亲社会的;利他倾向的
amigdala /əˈmɪɡdələ/ n. 56:56
杏仁核(amygdala 的误转写,情绪相关脑区)
selfish /ˈselfɪʃ/ adj. 56:56
自私的
neural basis n. phr. 58:20
神经基础
postponed /poʊstˈpoʊnd/ v. 58:20
延期
faculty member n. phr. 58:20
教员;教职人员
clear-cut /ˌklɪr ˈkʌt/ adj. 59:44
泾渭分明的;明确无疑的
causal structure n. phr. 59:44
因果结构
state variables n. phr. 1:00:24
状态变量
high dimension n. phr. 1:00:24
高维
feedback loop n. phr. 1:01:51
反馈回路
intervene /ˌɪntərˈviːn/ v. 1:01:51
干预;介入
paradigm /ˈpærədaɪm/ n. 1:01:51
范式
active learning n. phr. 1:02:38
主动学习(学习者自己选择样本)
精读便签
下载便签 手机:长按图片也可保存
← 上一期 · NO.132DigHum2026 • Economies of AI: Where is the value? • Keynote 下一期 · NO.134 →Judea Pearl, 2012 ACM A.M. Turing Award Lecture "The Mechanization of Causal Inference"
苏菲周报 · THE WEEKLY 每周一封,
追问一个大问题。
苏菲拉底的每周来信,写这一周在追问的问题和看到的回应。
苏菲拉底
ASK THE BIG QUESTIONS · THINK DEEPLY · SEE THE WORLD DIFFERENTLY
苏菲拉底微信公众号二维码 微信公众号
© 2026 苏菲拉底 · 内容仅供学习 [email protected]