跳转至

文章背景与核心概要

本文深入探讨了现代生成式AI中“不可见”的运作层面,指出当前的评测往往错误地将模型行为完全归因于训练权重或对齐技术。作者强调,现代推理技术栈允许在模型训练完成后进行运行时干预(如机构或商业层面的操控),这给AI的安全透明度带来了全新的挑战。

文章的核心贡献包括提出了“推理归因难题”(证明黑盒观察无法区分模型内部权重行为与外部推理策略),定义了“概率布局”(通过系统性重新分配概率质量来植入商业或意识形态影响),并探讨了对治理和监管(如欧盟《人工智能法案》)的深远影响,呼吁监管重心从静态模型审计转向对整个部署系统的审计。


不可见的编辑层:形式化定义已部署语言模型中未披露的推理时操控、概率布局与归因难题

作者: Augusto Camargo
日期: 2026年8月26日(v2版)
主题: 人工智能 (cs.AI);计算与语言 (cs.CL);计算机与社会 (cs.CY)
DOI: 10.48550/arXiv.2608.24662


摘要 (Summary)

This paper investigates the "invisible" layers of modern generative AI, arguing that current evaluations often mistakenly attribute model behavior solely to training weights or alignment. The author highlights that modern inference stacks allow for runtime interventions—such as institutional or commercial steering—that occur after the model has been trained.

Key contributions include: * The Inference Attribution Problem: A formal proof showing that black-box observation cannot distinguish between behavior caused by model weights versus behavior caused by external inference-time policies. * Probability Placement: A characterization of how commercial influence is embedded into model outputs through systematic probability-mass reallocation, distinct from traditional advertising. * Governance Implications: A call for a shift in regulatory focus (e.g., EU AI Act, Digital Services Act) from auditing static models to auditing the entire deployed system, including the inference-time editorial layer.

本文研究了现代生成式AI中“不可见”的层面,指出当前的评测往往错误地将模型行为完全归因于训练权重或对齐。作者强调,现代推理技术栈允许在模型训练完成后进行运行时干预——例如机构或商业操控。

核心贡献包括: * 推理归因难题(The Inference Attribution Problem): 给出了一个形式化证明,表明黑盒观察无法区分由模型权重引起行为与由外部推理时策略引起的行为。 * 概率布局(Probability Placement): 刻画了商业影响力如何通过系统性的概率质量重新分配嵌入到模型输出中,这与传统广告截然不同。 * 治理意义(Governance Implications): 呼吁监管重点(如欧盟《人工智能法案》、《数字服务法案》)从审计静态模型转向审计整个部署系统,包括推理时的编辑层。


核心概念 (Core Concepts)

1. 推理归因难题

The paper establishes an observational non-identifiability result. It demonstrates that two systems can produce identical outputs while possessing fundamentally different internal architectures—one relying on internal alignment and the other on external, undisclosed inference-time steering. This makes it impossible to determine the source of a model's bias through output observation alone.

本文建立了一个观测不可辨识性结果(observational non-identifiability result)。它证明了两个系统可以产生相同的输出,但却拥有根本不同的内部架构——一个依赖于内部对齐,另一个依赖于外部、未披露的推理时操控。这使得仅通过输出观测来确定模型偏差的来源成为不可能。

2. 概率布局

The author defines "Probability Placement" as a deployment pattern where commercial or ideological influence is injected into an assistant's response. By reallocating probability mass during the token generation process, the system can subtly steer the user toward specific outcomes without the need for explicit, labeled advertising.

作者将“概率布局”定义为一种部署模式,在此模式下,商业或意识形态影响力被注入到助手的回复中。通过在词元(token)生成过程中重新分配概率质量,系统可以在不需要显式、带标签广告的情况下,微妙地引导用户走向特定的结果。

3. 审计与治理

The research argues that the current paradigm of "model auditing" is insufficient for modern production environments. Because the "editorial layer" can modify behavior at runtime, governance frameworks must evolve to require transparency regarding the entire inference stack, including provenance, confidential computing, and cryptographic attestation.

该研究认为,当前的“模型审计”范式对于现代生产环境来说是远远不够的。由于“编辑层”可以在运行时修改行为,治理框架必须进行演进,以要求整个推理栈的透明度,包括数据来源、机密计算和密码学证明。


访问与资源 (Access & Resources)

license icon


For further bibliographic tools, citation tracking, and related research, please refer to the original arXiv entry.

有关更多书目工具、引文跟踪和相关研究,请参阅原始 arXiv 条目