文章背景与核心概要
几十年来,机器验证(Machine Verification)一直是一项成本高昂且局限于特定领域的冷门工作。本文展示了生成式AI如何根本性地颠覆这一现状,将机器验证转化为一种必不可少的、高生产力的工具。作者引入了名为“Salt方法”的工作流,在该工作流中,AI代理在严格的证明内核(Proof Kernel)监督下运行。通过将数学断言视为经内核检查的制品(Artifacts),单个研究人员能够指挥一组AI代理开发整个系统——从应用代码、经验证的编译器一直到RISC-V处理器——期间没有人工编写RTL代码,也没有进行人工证明审查。该项目成功实现了芯片流片(Tape-out),并保持了经验证的零错误记录。
这项突破性研究表明,在AI运行的速度下,机器验证不仅经济划算,而且是提升生产力的关键所在——它是不可腐蚀的裁判,使单个人能够安全地大规模指挥自主机器工作。在短短五 weeks 内,研究人员利用消费级AI订阅服务,引导AI代理团队完成了从软件应用到物理硅片的完整设计闭环。
具备权威性的AI:从应用层到硅皮
作者: Jason Hickey
日期: 2026年8月21日
标识符: arXiv:2608.21356
主题: 软件工程 (cs.SE);人工智能 (cs.AI);硬件架构 (cs.AR);计算机科学中的逻辑 (cs.LO)
摘要 (Summary)
For decades, machine verification has been a costly, niche endeavor. This paper demonstrates that generative AI fundamentally inverts this dynamic, turning machine verification into an essential, high-productivity tool. The author introduces the "Salt method," a workflow where AI agents operate under the strict supervision of a proof kernel. By treating mathematical claims as kernel-checked artifacts, a single researcher was able to direct a fleet of AI agents to develop a system—from application code through a verified compiler to a RISC-V processor—without human-written RTL or manual proof review. The project successfully achieved a silicon tape-out with a verified, error-free record.
几十年来,机器验证一直是一项成本高昂、仅限于少数特例领域的工作。本文证明了生成式AI彻底颠覆了这种动态关系,将机器验证转变为一项必不可少的、高生产力的工具。作者引入了 “Salt方法” 这一工作流,其中AI代理在证明内核的严格监督下运行。通过将数学断言视为经过内核检查的制品,单名研究人员得以指挥一组AI代理开发出一个系统——从应用代码、经验证的编译器一直到RISC-V处理器——全程无需人工编写RTL(寄存器传输级)代码或进行人工证明审查。该项目成功实现了硅片流片(Tape-out),并拥有经过验证的零错误记录。
抽象 (Abstract)
For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relationship: at AI speed, machine verification is not only economical but essential to productivity — it is the incorruptible referee that lets one person safely direct autonomous machine work at scale.
在过去的六十年里,机器验证一直是主要的成本开销,只有针对极特殊的制品时才负担得起。在此,我们报告称生成式AI颠覆了这种关系:在AI的速度下,机器验证不仅经济,而且对生产力至关重要——它是不可收买的裁判,使一个人能够安全地大规模指导自主机器的工作。
In five weeks, one researcher on consumer AI subscriptions directed a small fleet of AI agents from application code, through a verified compiler and executive, to a RISC-V processor taped out on a community silicon shuttle; no proof passed through human review, and no RTL was written by a human. The working discipline — the Salt method — rests on a proof kernel no hallucinated proof can pass: mathematical claims travel between agents as kernel-checked artifacts, and human attention is reserved for statements, designs, and rulings.
在五周内,一名拥有消费级AI订阅的研究人员指挥一小支AI代理舰队,从应用代码出发,经由经过验证的编译器和执行程序,最终设计出在社区硅晶 shuttle 上流片的 RISC-V 处理器;期间没有任何证明经过人工审查,也没有由人类编写的 RTL。这一工作准则——即 Salt 方法——依赖于一个任何产生幻觉的证明都无法通过的证明内核:数学断言在代理之间作为经内核检查的制品传递,而人类的注意力则保留用于陈述、设计和裁决。
Verification is stated link by link, from the Lean 4 kernel to SAT-checked equivalence at the silicon boundary. We publish the complete accounting: theorem provenance, a pre-registered token meter, floor-bounded human time, and an error ledger whose catch numbering runs to #256 — a monotone counter over the mathematics campaign's append-only flags ledger, maintained 2026-07-07 to 2026-07-20 (one number, #79, was never assigned; later catches are recorded un-numbered) — against zero incorrect proofs reaching the record.
验证是环环相扣地进行的,从 Lean 4 内核一直到硅片边界处经 SAT 检查的等价性。我们公布了完整的核算数据:定理来源、预注册的 Token 计量器、有下限约束的人类工时,以及一个错误分类账本(其捕获编号一直到 #256——这是数学攻坚战中仅追加标志 ledger 的单调计数器,维护于 2026-07-07 至 2026-07-20 期间;其中一个编号 #79 从未被分配;后续的捕获则以无编号形式记录)——最终实现了零错误证明进入记录的目标。
获取论文 (Accessing the Paper)
附加元数据 (Additional Metadata)
- 评论: 17页,6张图表
- 提交历史: [v1] 2026年8月21日 星期五 17:59:16 UTC