继承式智能体记忆预算验证中的计划指针与记录指令形式
文章背景与核心概要
本文探讨了在记忆继承和预算验证约束下,不同的记录指令形式(如纯ID、长度匹配标准或两者的组合)如何影响语言模型的检索与决策行为。作者通过横跨12项注册研究、总计14,760次测试实验,评估了特定的提示词修改如何引导智能体访问归档的源记录。
研究结果表明,不同的提示词设计(如计划指针、标准说明及附加ID等)对模型的检索行为有着显著且复杂的影响。该研究的所有完整数据集、冻结软件包、代码和手稿均已公开存档于Zenodo,为理解和优化大语言模型智能体的记忆检索机制提供了宝贵的经验数据。
计划指针与记录指令形式在继承式智能体记忆预算验证中的应用
Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
摘要
一个继承了六条单行记忆的智能体在行动前最多只能提取一条归档的源记录;写入存储库中的指令可以引导这一选择:指向该记录的指针、识别该记录的标准,或者两者兼具。通过对单一工具谱系进行的12项注册研究(14,760次尝试),我们测量了在各种形式下请求的流向。在六个直连提供商模型上,长度匹配的标准比纯ID高出 +35.0 个百分点 [+31.2, +38.8](研究D);但在九个通过OpenRouter服务的模型面板上,这种对比未能通过其注册的优越性规则(研究E)。追加ID取消了三个Claude模型上的标准(Opus 5: 从 40/40 降至 0/40;研究F-x);六个字节匹配的修改使得每个精确字符串都产生其自身的效果(研究G),而以每单元80次运行进行的反向运行使得三十个复制对比中有十五个落在容差范围内,十五个未解决,没有一个超出范围(研究J中,确认行在Opus 5上提升了 +96.0 个百分点,两个信用点的预算恢复了所有三个模型上的目标;在五个标准字符串中,后缀的取消效应对Opus 5的五个措辞中的四个以及Fable 5.1的所有五个措辞都成立(研究H2);在第二个存储库中,所有模型都遵循了标准(研究H1)。延续到决策阶段,标准将选择推向当前记录(Opus 5上 +100.0 个百分点),但在Fable 5.1上则背离了它(研究I)。一个单字符计划指针的效果(+78.0 个百分点;研究B,在其第一个存储库报告修正之后)在前瞻性注册的反向运行中得出了相同的结论(研究B')。所有结果都是精确编辑在固定面板上具有注册区间的描述性效应,并不包含机制声明。
An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both. Across twelve registered studies on one instrument lineage (14,760 attempts) we measured where the request goes under each form. On six direct-provider models a length-matched criterion exceeded a bare id by +35.0 points [+31.2, +38.8] (Study D); the contrast failed its registered superiority rule on a nine-model OpenRouter-served panel (Study E). Appending the id cancelled the criterion on three Claude models (Opus 5: 40/40 to 0/40; Study F-x); six byte-matched edits gave each exact string its own effect (Study G), and a re-run at eighty runs per cell left fifteen of thirty replication contrasts within the margin, fifteen unresolved and none beyond (Study G'). A ratification line (+96.0 points on Opus 5) and a budget of two credits restored the target on all three (Study J); across five criterion strings the suffix's cancellation held for four of the five wordings on Opus 5 and all five wordings on Fable 5.1 (Study H2); in a second store every model followed the criterion (Study H1). Continued into a decision, the criterion moved the choice toward the current record (+100.0 points, Opus 5) and away from it on Fable 5.1 (Study I). A one-character plan pointer's effect (+78.0 points; Study B, after a correction of its first repository report) returned the same verdict under a prospectively registered re-run (+81.7 points; Study B'). All results are descriptive effects of exact edits on fixed panels with registered intervals and no mechanism claim.
文档元数据
Document Metadata
Document Metadata
| 字段 | 详情 |
|---|---|
| 标题 | 继承式智能体记忆预算验证中的计划指针与记录指令形式 |
| 作者 | Kazuki Nakayashiki |
| 提交于 | 2026年9月3日 |
| 主要主题 | 信息检索 (cs.IR) |
| 次要主题 | 人工智能 (cs.AI);计算与语言 (cs.CL) |
| arXiv 标识符 | arXiv:2609.03450 |
| DOI | 10.48550/arXiv.2609.03450 |
| 存档 DOI | 10.5281/zenodo.22267221 |
Field Details Title Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory Author Kazuki Nakayashiki Submitted On September 3, 2026 Primary Subject Information Retrieval ( cs.IR)Secondary Subjects Artificial Intelligence ( cs.AI); Computation and Language (cs.CL)arXiv Identifier arXiv:2609.03450 DOI 10.48550/arXiv.2609.03450 Archive DOI 10.5281/zenodo.22267221
访问与全文链接
Access and Full-Text Links
Access and Full-Text Links
- PDF 版本: 查看 PDF
- 实验性 HTML: arXiv HTML 视图
- TeX 源码: 下载源码
- 许可证: 知识共享署名 4.0
- PDF Version: View PDF
- Experimental HTML: arXiv HTML View
- TeX Source: Download Source
- License: Creative Commons Attribution 4.0
提交历史
Submission History
Submission History
- [v1] - 2026年9月3日 星期四,07:05:38 UTC (84 KB)
- [v1] - Thu, 3 Sep 2026, 07:05:38 UTC (84 KB)