正确不等于受控:智能体工作流中的溯源完整性
文章背景与核心概要
在机构级应用场景中,仅凭是否达成正确结果来评估智能体工作流是远远不够的。一个表面正确的行动可能依赖于错误的授权、未经支撑的完成声明,或是因后续变更而失效的过时工作。
为了解决这一痛点,本文引入了受控执行(governed execution)的概念——即其决策、完成状态以及对变更的响应均由可检查的溯源信息支持的工作流。作者提出了 Matrix,这是一个确定性的因果状态层,旨在记录权限与事实依赖关系、验证完成证据,并选择性地使受影响的工作失效。受控评估表明,尽管受控工作流与直接工作流通常能达成相同的结果,但只有受控方法能够始终如一地保存证据、防止不受支持的闭环,并将恢复范围严格限制在依赖任务中。然而,角色分离的迁移测试表明,确定性强制执行的完整性契约可能会对在其创作上下文之外生成的合成数据包造成严重的过度阻断。
归根结底,Matrix 并不是作为一个通用的准确性增强器提出的,而是作为一个机构级完整性层,旨在使智能体的工作具备可审计性和独立可验证性。
摘要 (Summary)
In institutional settings, evaluating agentic workflows solely on whether they reach the correct outcome is insufficient. A correct action may still rely on the wrong authority, an unsupported completion claim, or work rendered stale by subsequent changes.
This paper introduces governed execution—work whose decisions, completion status, and responses to change are supported by inspectable provenance. The author presents Matrix, a deterministic causal-state layer designed to: * Record authority and fact dependencies. * Verify completion evidence. * Selectively invalidate affected work.
Controlled evaluations reveal that while governed and direct workflows often achieve the same outcomes, only the governed approach consistently preserves evidence, prevents unsupported closures, and limits recovery strictly to dependent tasks. However, a role-separated transfer challenge demonstrated that a deterministically enforced completeness contract can severely over-block synthetic packets produced outside their authoring context.
Ultimately, Matrix is not presented as a general accuracy enhancer, but rather as an institutional integrity layer that makes agentic work auditable and independently verifiable.
在机构场景中,仅根据智能体工作流是否达成正确结果来进行评估是远远不够的。一个正确的行动可能仍然依赖于错误的授权、不受支持的完成声明,或是因后续变更而变得陈旧的工作。
本文引入了受控执行(governed execution)——即其决策、完成状态和对变更的响应均受可检查溯源支持的工作。作者提出了 Matrix,这是一个确定性的因果状态层,旨在: * 记录授权和事实依赖关系。 * 验证完成证据。 * 选择性地使受影响的工作失效。
对照评估表明,虽然受控工作流和直接工作流通常能达到相同的结果,但只有受控方法能够持续保存治理证据、拒绝不受支持的闭环,并将恢复限制在受影响的任务上。然而,角色分离的迁移挑战表明:确定性强制执行的完整性契约可能会对在其创作上下文之外生成的合成数据包造成严重的过度阻断。
归根结底,Matrix 并不是作为一个通用的准确性增强器出现的,而是作为一个机构完整性层,用于使智能体的工作具有可审计性和独立可验证性。
元数据 (Metadata)
- arXiv ID: arXiv:2608.12761 [cs.AI]
- DOI: 10.48550/arXiv.2608.12761
- 作者 (Author): Jesus Salas
- 提交时间 (Submitted): 2026年8月13日
- 学科分类 (Subjects): 人工智能 (
cs.AI);密码学与安全 (cs.CR) - 篇幅 (Length): 19页,2张图表
摘要原文 (Abstract)
Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Across controlled comparisons, governed and direct workflows often reached the same outcomes, but only the governed path consistently preserved governing evidence, refused unsupported closure, and limited recovery to dependent tasks. A role-separated transfer challenge then failed: a deterministically enforced completeness contract severely over-blocked synthetic packets produced outside their authoring context. These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.
智能体工作流通常根据其是否达到正确结果来进行评估。这在机构环境中是不够的,因为正确的行动可能依赖于错误的权威、不受支持的完成声明或因后续更改而变得陈旧的工作。我们将受控执行定义为:其决策、完成和对变更的响应均由可检查溯源支持的工作。我们提出了 Matrix,这是一个确定性的因果状态层,可记录权限和事实依赖项、验证完成证据并选择性地使受影响的工作失效。在对照比较中,受控工作流和直接工作流通常达到相同的结果,但只有受控路径一致地保留了治理证据、拒绝了不受支持的闭环,并将恢复限制在依赖任务上。随后,角色分离的迁移挑战失败了:确定性强制执行的完整性契约严重过度阻止了在其创作上下文之外生成的合成数据包。这些结果并没有将 Matrix 确立为通用的准确性增强器,而是支持了其作为使智能体工作可审计和独立可验证的机构完整性层的核心作用。
访问与资源 (Access & Resources)
- 全文选项 (Full-Text Options):
- 查看 PDF (View PDF)
- HTML 格式(实验性)(HTML (Experimental))
- TeX 源码 (TeX Source)
- 许可证 (License): 知识共享署名 4.0 (Creative Commons Attribution 4.0)

- 外部索引 (External Indices):
- 谷歌学术 (Google Scholar)
- Semantic Scholar
- NASA ADS