多模态认知的层次化基于能量的模型
文章背景与核心概要
本文介绍了一种名为 IM-LEPP(集成多模态潜空间基于能量的预测加工,Integrated Multimodal Latent Energy-based Predictive Processing)的全新层次化、基于能量的多模态认知框架,它将视觉与语言有机地结合在一起。IM-LEPP 将生成式神经网络视为认知动力学的有效理论(类似于热力学中的统计力学),不再单纯局限于模拟底层神经回路,而是将认知建模为通过学习所得的能量地形流动的潜状态。
该模型在架构上采用了基于前颞叶(anterior temporal lobe)模特的“核心-辐射”(Hub-and-Spoke)层次结构,使视觉物体、场景以及语言单元的预测编码管线汇聚到一个共享的非模态核心中。IM-LEPP 不仅能够对注意力动力学、非注意盲视、内克尔立方体双稳态等认知现象给出机械式的解释,还能复现或推导诸如惊讶度理论(surprisal theory)、N400/P600 ERP 成分以及花园路径重分析等心理语言学标记。与现代 Transformer 大语言模型相比,它为数据高效的语言习得提供了一种具有可证伪性的替代方案。
摘要 (Abstract)
我们提出了 IM-LEPP(集成多模态潜空间基于能量的预测加工),这是一个用于多模态认知的层次化、基于能量的模型,它将先前提出的单模态模型(LEPP)扩展到了视觉与语言的集成。
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language.
遵循生成式神经网络是认知动力学有效理论的观点——类似于统计力学与热力学的关系,IM-LEPP 将认知建模为通过学习所得的能量地形流动的潜状态,而不是对神经回路的直接解释。
Following the view that generative neural networks are effective theories of cognitive dynamics, analogous to how statistical mechanics relates to thermodynamics, IM-LEPP models cognition as latent states flowing through learned energy landscapes rather than as an account of neural circuitry.
架构框架 (Architectural Framework)
- 核心-辐射层次结构: 建立在 Lambon Ralph 等人的控制语义认知框架基础之上,其中视觉物体、场景和语言单元的预测编码管线汇聚到一个以颞前叶为模特的共享非模态核心中。
- Hub-and-Spoke Hierarchy: Grounded in the controlled semantic cognition framework of Lambon Ralph et al., where predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub modeled on the anterior temporal lobe.
- 上下文条件化: 每个管线的预测都受当前核心状态的条件化制约(而非被其覆盖)。这既保持了管线自身的特性,又让每个预测都反映了完整的多模态上下文。
- Contextual Conditioning: Each pipeline's prediction is conditioned by (rather than overwritten by) the current hub state. This preserves pipeline-specific identity while letting every prediction reflect the full multimodal context.
核心贡献与发现 (Key Contributions & Findings)
- 认知现象: 为注意力现象(包括非注意盲视和内克尔立方体双稳态)提供了机制性的解释。
- Cognitive Phenomena: Provides a mechanistic account of attentional phenomena, including inattentional blindness and Necker-cube bistability.
- 心理语言学: 恢复或激发了独立确立的研究发现,包括惊讶度理论、N400/P600 ERP 成分以及花园路径重分析。
- Psycholinguistics: Recovers or motivates independently established findings, including surprisal theory, the N400/P600 ERP components, and garden-path reanalysis.
- 对比反差: 针对下一个词预测中的轨迹敏感性,提供了与 Transformer 语言模型的可证伪对比,并探讨了相对于大语言模型的数据高效语言习得。
- Comparative Contrast: Offers a falsifiable contrast with transformer language models regarding trajectory-sensitivity in next-word prediction, alongside discussions on data-efficient language acquisition relative to LLMs.
- 子系统与理论: 勾勒出了语义/情景记忆子系统,将该模型置于预测加工、自由能原理、JEPA(联合嵌入预测架构)以及层次化时间记忆的背景之下,并提出了具体的实验预测。
- Subsystems & Theory: Outlines a semantic/episodic memory subsystem, situates the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory, and proposes concrete experimental predictions.
关键词 (Key Words)
- 预测加工 (Predictive Processing)
- 预测编码 (Predictive Coding)
- 基于能量的模型 (Energy-Based Models)
- 扩散模型 (Diffusion Models)
- 有效理论 (Effective Theory)
- 计算神经科学 (Computational Neuroscience)
- Predictive Processing
- Predictive Coding
- Energy-Based Models
- Diffusion Models
- Effective Theory
- Computational Neuroscience
全文与访问链接 (Full-Text & Access Links)
- PDF 版本: 查看 PDF
- 许可证: 知识共享署名 4.0 (
查看许可证)
- PDF Version: View PDF
- License: Creative Commons Attribution 4.0 (
view license)
文献计量与研究工具 (Bibliographic & Research Tools)
- 引用: NASA ADS | Google Scholar | Semantic Scholar
- 交互式工具: alphaXiv | CatalyzeX | Hugging Face
- Citations: NASA ADS | Google Scholar | Semantic Scholar
- Interactive Tools: alphaXiv | CatalyzeX | Hugging Face