文章背景与核心概要
尽管机器智能在符号世界中取得了巨大成功,但在物理世界中却面临着结构性的停滞——即“冷启动死锁”(cold-start deadlock):物理人工智能(Physical AI)的发展需要数据才能运作,然而生成有意义的数据又必须依赖已部署的智能系统。本文作者来自北京筑思科技(Persagy Science and Technology Co.),他们提出了一种以“人工物理世界”(如建筑物、工业设施和基础设施)为核心的基础框架。由于这些设计产物在物理实例化之前就已被有意地构成并记录在可读的档案中,因此其规范是通过颁布确立的,而不是从数据中平均得出的。
该论文做出了四项核心贡献:首先,从四世界本体论中推导出了构成先验框架的合法性准则,并通过适配方向(direction of fit)进行检验;其次,确立了包含语法、概念、知识和实例在内的四层下界(layering lower bound);再次,在五个工业领域中注册了部署声明,并建立了包含32个类别的失效模式词汇表;最后,通过五项可证伪的预测来支撑该框架。该研究将大语言模型(LLMs)精确定位为“档案的阅读者,而非档案本身”,为解决物理AI的死锁问题提供了崭新的理论路径。
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World
Authors: Jiang Jiang (1), Yifu Sun (1), Qi Shen (1)
(1) Persagy Science and Technology Co., Beijing, China
Submitted: 15 August 2026
arXiv: 2608.15147 [cs.AI]
DOI: 10.48550/arXiv.2608.15147
📋 Summary
While machine intelligence has largely conquered the symbolic world, it continues to face a structural stall within the physical world—a "cold-start deadlock" where physical AI requires data to function, yet requires a deployed intelligence to generate meaningful data.
This paper introduces a foundational framework centered around the artificial physical world (e.g., buildings, industrial facilities, and infrastructure). Because these designed artifacts are intentionally constituted and documented with readable archives prior to their physical instantiation, norms are promulgated rather than averaged from data. The authors establish a legitimacy theory, layer lower bounds, industrial failure-mode vocabularies, and falsifiable predictions to resolve the physical AI deadlock, positioning Large Language Models (LLMs) as readers of the archive rather than the archive itself.
📖 Abstract
机器智能已经征服了符号世界,但在物理世界中却陷入了停滞。这种停滞是结构性的:物理人工智能(Physical AI)面临着冷启动死锁——没有数据就没有智能,没有已部署的智能就没有数据。我们的论点是:这种死锁是真实存在的,但分布并不均匀,而这个例外有一个名字:人工物理世界。建筑物、工业设施和基础设施都是经过刻意构建和记录的:设计好的产物附带了可读的档案,这些档案先于并构成了它们的实体;在这里,规范是在实例出现之前就已颁布的,而不是从实例中平均得出的。本文有四项贡献:
- 从四世界本体论中,我们推导出了构成先验框架的合法性准则:当且仅当对象域是被刻意构成且留下了可读档案时,先验提取才是合法的;该准则可通过适配方向进行检验——偏离构成规范是世界中的违规行为,而不是对模型的修正。
- 我们确立了一个分层下界:任何此类框架至少包含四层——语法、概念、知识、实例——因为四个构建目标配对成了互不相容的载体。
- 我们在五个工业领域注册了部署声明,并建立了一个32类别的失效模式词汇表。
- 我们将该框架押注于五个可证伪的预测上,其中核心预测可在公共工程记录上进行检验:如果它失败了,该框架也将失效。
半形式化论证支持了这些主张(附录 A):无档案世界中规则覆盖率的戈德式边界(Gold-type boundary)、封闭概念层上失效约简的可判定性结果,以及证书锚定演算的边界定理。大语言模型在这里找到了一个受尊崇的位置——作为档案的阅读者,而不是档案本身。这是三部姊妹篇作品中的第一部;其余姊妹篇将探讨故意留白的问题。
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions.
- From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model.
- We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers.
- We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary.
- We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails.
Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
📊 Paper Metadata & Metrics
论文元数据与指标:
| 类别 (Category) | 详情 (Details) |
|---|---|
| 主要学科 (Primary Subject) | 人工智能 (cs.AI) |
| 次要学科 (Secondary Subjects) | 多智能体系统 (cs.MA) |
| ACM 类别 (ACM Classes) | I.2.4; I.2.0 |
| 篇幅 (Length) | 55 页,3 个图表,75 篇参考文献 |
| 许可协议 (License) | CC BY-NC-ND 4.0 ![]() |
Category Details Primary Subject Artificial Intelligence ( cs.AI)Secondary Subjects Multiagent Systems ( cs.MA)ACM Classes I.2.4; I.2.0 Length 55 pages, 3 figures, 75 references License CC BY-NC-ND 4.0
🔗 Full-Text & Access Links
全文与访问链接:
- 查看 PDF (View PDF)
- HTML 版本 - 实验性 (HTML Version (Experimental))
- TeX 源码 (TeX Source)
- 外部档案 (External Archives): NASA ADS | 谷歌学术 (Google Scholar) | 语义学者 (Semantic Scholar)
- View PDF
- HTML Version (Experimental)
- TeX Source
- External Archives: NASA ADS | Google Scholar | Semantic Scholar
🗂️ Submission History
提交历史:
- [v1] 2026年8月15日 星期六 09:50:19 UTC (74 KB)
- [v1] Sat, 15 Aug 2026 09:50:19 UTC (74 KB)
