跳转至

文章背景与核心概要

现代人工智能模型通过高-dimensional 的神经表示来处理复杂任务,但这些表示的内部机制由于其复杂性而难以解释。为了揭示深度网络背后的通用规律,本文研究了视觉、音频和语言等多种模态下现代AI模型的内部表示,发现分类任务在表征空间中会产生一种通用的几何结构。类内变异性并非随机分布,而是紧密围绕真实类别及其竞争“对手”类别的中心点(centroids)进行组织。

通过开发解析平均场理论,研究人员证明了高精度分类依赖于一组稀疏的中心点坐标,并且该理论能够跨架构和模态准确预测分类准确率。这一框架不仅为深度网络如何进行分类提供了一个简洁的解释,还建立了几何表征与稀疏特征提取方法(如稀疏自编码器)之间的理论联系,表明深度网络中的分类受制于嵌入在高维表征空间中的稀疏、对齐中心点的结构。


Sparse Prototype Code Underlies Classification and Prediction Across Modalities

Authors: Yehonatan Avidan, Daniel D. Lee, Haim Sompolinsky
Date: 16 August 2026
Subject: Machine Learning (cs.LG)
DOI: 10.48550/arXiv.2608.15632

Authors: Yehonatan Avidan, Daniel D. Lee, Haim Sompolinsky
Date: 16 August 2026
Subject: Machine Learning (cs.LG)
DOI: 10.48550/arXiv.2608.15632


Summary

本论文通过分析现代AI模型的高维神经表示,研究了其内部机制。作者证明了跨视觉、音频和语言模态的分类任务共享一种通用的几何结构。他们发现,类内变异性并非随机的;相反,它是围绕真实类别及其竞争“对手”类别的中心点构建的。通过开发解析平均场理论,研究人员表明,准确的分类依赖于这组中心点坐标中的稀疏子集。该框架为深度网络如何执行分类提供了一个简洁的解释,并建立了与稀疏特征提取方法(如稀疏自编码器)之间的理论联系。

This paper investigates the internal mechanisms of modern AI models by analyzing their high-dimensional neural representations. The authors demonstrate that classification tasks across vision, audio, and language modalities share a universal geometric structure. They find that within-class variability is not random; rather, it is structured around the centroids of the true class and its competing "rival" classes. By developing an analytical mean-field theory, the researchers show that accurate classification relies on a sparse set of these centroid coordinates. This framework provides a parsimonious explanation for how deep networks perform classification and establishes a theoretical link to sparse-feature extraction methods like sparse autoencoders.


Abstract

神经表示已成为研究现代AI模型内部机制的核心工具,然而其复杂的高维结构使得它们难以解释。我们表明,分类任务催生了一种通用的表征几何结构,该结构在视觉、音频和语言处理的最先进模型中是共享的。

核心结构在于,类内变异性在表征空间中并非随机的。相反,其与分类器相关的分量与该类别自身的中心点以及竞争对手类别的中心点具有强大且结构化的相关性。基于这一观察,我们推导出了一个解析平均场理论,该理论主要受真实类别和竞争对手类别中心点坐标沿线变异性的支配,同时辅以对真实类别半径的全局重正化,以补偿真实表征的非高斯统计特性。

该理论准确地预测了跨架构和模态的分类准确率。相关的几何量随模型规模系统性地改善,反映了观察到的准确率提升。该理论的一个显著特征是其稀疏性:准确预测只需要与真实类别及其最强对手相关的一小组中心点坐标——这把我们的框架与稀疏自编码器等稀疏特征提取方法连接了起来。总之,这些结果为神经表示提供了一个简洁的预测理论,并表明深度网络中的分类受制于嵌入在完整高维表征空间中的稀疏、中心点对齐的结构。

Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. We show that classification tasks give rise to a universal representational geometry, shared across state-of-the-art models in vision, audio, and language processing.

The key structure is that within-class variability is not random in representation space. Instead, its classifier-relevant component has strong and structured correlations with the class's own centroid and with the centroids of its competing classes. Building on this observation, we derive an analytical mean-field theory governed mainly by the variability along true-class and rival-class centroid coordinates, together with a global renormalization of the class radius that compensates for the non-Gaussian statistics of real representations.

The theory accurately predicts classification accuracy across architectures and modalities. The relevant geometric quantities improve systematically with model scale, mirroring the observed gains in accuracy. A striking feature of the theory is its sparsity: accurate prediction requires only a small set of centroid coordinates associated with the true class and its strongest rivals—connecting our framework to sparse-feature extraction approaches such as sparse autoencoders. Together, these results provide a parsimonious predictive theory of neural representations and suggest that classification in deep networks is governed by a sparse, centroid-aligned structure embedded within the full high-dimensional representation space.


Accessing the Paper

访问论文: * 查看 PDF * HTML 页面(实验性) * TeX 源码


Metadata

元数据: * 评注: 33页,13张图,14个表 * 一级学科: 机器学习 (cs.LG) * 二级学科: 无序系统与神经网络 (cond-mat.dis-nn);人工智能 (cs.AI);机器学习 (stat.ML)

  • Comments: 33 pages, 13 figures, 14 tables
  • Primary Subject: Machine Learning (cs.LG)
  • Secondary Subjects: Disordered Systems and Neural Networks (cond-mat.dis-nn); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)