文章背景与核心概要
长期以来,深度学习领域的神经网络架构主要通过模块语法、计算图或其实现的复合函数来进行定义和区分。然而,来自密歇根大学的 Luis F. Rosario Freytes 在这篇论文中提出了一种全新的过程级(process-level)视角,用于研究和个体化神经网络架构。作者通过分析接收端(receiver)处可用的被表示进程——具体表现为 \(B_j = G_j Q_j\) 形式的分解(其中 \(Q_j(x)\) 代表为后续计算提供的中间状态),深入探讨了计算截断(computational cut)处所保留的信息。
论文的核心技术在于通过分析核 \(\ker Q_j\) 来捕获前序区分(predecessor distinctions)。虽然这种外延阴影(extensional shadow)可以对无标记的满射分解进行分类,但它并不能完全决定有标记的接收端组织方式。通过精确的两标记(two-token)构造,作者证明了局部操作(local schemas)与注意力机制(attention schemas)在经过一个单射前缀后可能是区分等价的,而在经过另一个前缀时则不等价,即便这两个前缀都没有丢失前序信息。这一发现揭示了隐藏在模块化架构标签之下的上下文依赖性,并强力主张将架构比较的视角提升至“被表示进程”的层面。
Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map
Authors: Luis F. Rosario Freytes (University of Michigan)
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
arXiv: 2601.11618v3 [cs.LG]
Submitted: 10 Jan 2026; Last Revised: 12 Aug 2026
Status: Submitted to JMLR (34 pages)
Summary
神经网络架构传统上由模块语法、计算图或它们实现的复合函数来定义。本文引入了一种进程级的方法,通过研究接收端可用的被表示进程来个体化神经网络架构——具体而言,即形式为 \(B_j = G_j Q_j\) 的分解,其中 \(Q_j(x)\) 表示为后续计算提供的中间状态。
通过分析核 \(\ker Q_j\),作者捕获了在计算截断处保留的前序区分。虽然这种外延阴影将无标记的满射分解分类到唯一的载体再表征(carrier re-presentation),但它并没有完全决定有标记的接收端组织方式,也没有决定受限延拓对保留信息的访问权限。通过精确的两标记构造,本文证明了局部机制和注意力机制在经过一个单射前缀后可以是区分等价的,而在经过另一个前缀时是不等价的——即使两个前缀都没有丢失前序信息。这揭示了模块化架构标签背后的隐藏上下文依赖性,主张转向在被表示进程层面上比较架构。
Neural network architectures are traditionally defined by module syntax, computation graphs, or the composite functions they implement. This paper introduces a process-level approach to individuating neural architecture by studying the represented process available at a receiver—specifically, factorizations of the form \(B_j = G_j Q_j\), where \(Q_j(x)\) represents the intermediate state supplied for subsequent computation.
By analyzing the kernel \(\ker Q_j\), the author captures the predecessor distinctions preserved at a computational cut. While this extensional shadow classifies unmarked surjective factorizations up to unique carrier re-presentation, it does not fully dictate marked receiver organization or how retained information is accessed by restricted continuations. Through an exact two-token construction, the paper demonstrates that local and attention schemas can be distinction-equivalent after one injective prefix and inequivalent after another—even when neither prefix loses predecessor information. This reveals a hidden context dependence underlying modular architecture labels, advocating for a shift toward comparing architectures at the represented-process level.
Abstract
神经网络架构通常通过模块语法、计算图或其实现的复合函数来识别。这些描述回答了不同的恒等性问题。我们研究了接收端可用的被表示进程:一个实际的分解 \(B_j=G_jQ_j\),其中 \(Q_j(x)\) 是为进一步计算提供的中间状态。
遗忘呈现方式并仅保留 \(\ker Q_j\) 可以得出在该截断处保留的前序区分。对于固定的分支,此前延阴影精确地将无标记的满射分解分类到唯一的载体再表征,但它并不决定有标记的接收端组织,也不决定受限延拓对保留信息的访问权限。在复合作用下,相关的接口是 \(Q_{j,\theta}A\),因此下游的架构区别取决于上游产生的状态。
精确的两标记构造表明,局部机制和注意力机制在经过一个单射前缀后是区分等价的,而在经过另一个前缀时是不等价的,尽管两个前缀都没有丢失前序信息。该结果暴露了被模块化架构标签隐藏的上下文依赖性,并推动了在被表示进程层面上进行架构比较。
Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions answer different identity questions. We study the represented process available at a receiver: an actual factorization \(B_j=G_jQ_j\) in which \(Q_j(x)\) is the intermediate state supplied for further computation.
Forgetting the presentation and retaining only \(\ker Q_j\) yields the predecessor distinctions preserved at that cut. For a fixed branch, this extensional shadow exactly classifies unmarked surjective factorizations up to unique carrier re-presentation, but it does not determine marked receiver organization or the accessibility of retained information to restricted continuations. Under composition the relevant interface is \(Q_{j,\theta}A\), so downstream architectural distinctions depend on the states produced upstream.
An exact two-token construction shows that local and attention schemas are distinction-equivalent after one injective prefix and inequivalent after another, even though neither prefix loses predecessor information. The result exposes a context dependence hidden by modular architecture labels and motivates architecture comparison at the represented-process level.
Full-Text and Resources
- 查看 PDF: arXiv:2601.11618
- HTML 版本: arXiv HTML (Experimental)
- TeX 源码: arXiv Source File
- DOI: 10.48550/arXiv.2601.11618
- 许可协议: Creative Commons Attribution 4.0
- View PDF: arXiv:2601.11618
- HTML Version: arXiv HTML (Experimental)
- TeX Source: arXiv Source File
- DOI: 10.48550/arXiv.2601.11618
- License: Creative Commons Attribution 4.0
License Notice & Visual Assets
许可声明与视觉资源
License Notice & Visual Assets