文章背景与核心概要
本文探讨了一种先进的AI幻觉检测与缓解框架,该框架将受HOPE启发的嵌套学习架构与连续记忆系统(CMS)及语义相似度缓存相结合。作者在一个包含310个提示词(其中包含217个认知不确定性查询和93个虚构诱导压力测试)的混合基准上进行了评估,实现了一个通过开放发言权协议(OFP)进行编排的三阶段流水线。
核心研究结果揭示了多维可靠性指标的微妙变化:性能方面,聚合的总幻觉分数在端到端上提升了其可达成范围的6.1%,其中高达97.7%的提升是在第一轮审查阶段实现的;在维度差异上,分数的提升实际上是一维的——83.5%的总提升仅由“显式语境化”驱动,而对于追踪无根据内容最为关键的“事实声明密度”指标则保持平稳;在可持续性方面,语义缓存成功处理了47.7%的模型调用,优化了计算开销。
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
Authors: Diego Gosmar, Deborah A. Dahl
Identifiers: arXiv:2605.29055 [cs.AI] | DOI: 10.48550/arXiv.2605.29055
Submission History: Submitted May 27, 2026; Revised August 12, 2026 (v2)
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
Authors: Diego Gosmar, Deborah A. Dahl
Identifiers: arXiv:2605.29055 [cs.AI] | DOI: 10.48550/arXiv.2605.29055
Submission History: Submitted May 27, 2026; Revised August 12, 2026 (v2)
Summary
本文研究了一个先进的框架,该框架通过将受HOPE启发的嵌套学习架构与连续记忆系统(CMS)和语义相似度缓存相结合,来检测和缓解AI幻觉。通过包含310个提示词(包含217个认知不确定性查询和93个虚构诱导压力测试)的混合基准进行评估,作者实现了通过开放发言权协议(OFP)编排的三阶段流水线。
关键发现在多维可靠性指标上展现出了微妙的特性: * 性能增益: 聚合的总幻觉分数在端到端上提升了其可达成范围的6.1%,其中惊人的97.7%的增益是在最初的审查阶段实现的。 * 维度差异: 分数的提升实际上是一维的——83.5%的总增益完全由显式语境化驱动,而事实声明密度(追踪无根据内容最关键的指标)则保持不变。 * 可观测性: 可观测性指标充当了一个注释通道,而非直接的响应属性,它在审查阶段跃升了147%,随后在通道停止传播时回落。 * 人工与跨模型验证: 三位人类标注员对93个压力提示词进行的独立标注显示,最终答案中有10.8%(95%置信区间为5.9–18.7)仍将虚构项目呈现为真实内容(\(\alpha = 0.586\))。跨系列评判模型(Llama 3.1、Gemma 4和Qwen 3)重新对930个输出进行了评分,并与人类评估高度一致(\(\rho = -0.772\) vs 原评估者的 \(-0.477\))。 * 可持续性: 语义缓存成功服务了47.7%的所有模型调用,优化了计算开销。
Summary
This paper investigates an advanced framework for detecting and mitigating AI hallucinations by coupling a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching. Evaluated on a hybrid benchmark of 310 prompts (including 217 epistemic-uncertainty queries and 93 fabrication-induction stress tests), the authors implement a three-stage pipeline orchestrated via the Open Floor Protocol (OFP).
Key findings reveal nuances in multi-dimensional reliability metrics: * Performance Gains: The aggregated Total Hallucination Score improves end-to-end by 6.1% of its attainable range, with an overwhelming 97.7% of this gain achieved during the very first review stage. * Dimensional Discrepancy: The score improvements are effectively one-dimensional—83.5% of the total gain is driven solely by Explicit Contextualization, while Factual Claim Density (the metric most critical for tracking unsupported content) remains flat. * Observability: The observability indicator acts as an annotation channel rather than a direct response property, jumping 147% at the review stage before falling back when the channel ceases propagation. * Human and Cross-Model Verification: Independent labeling by three human annotators on the 93 stress prompts revealed that 10.8% (95% CI 5.9–18.7) of final answers still present invented items as real (\(\alpha = 0.586\)). Cross-family judges (Llama 3.1, Gemma 4, and Qwen 3) re-scored 930 outputs and aligned strongly with human evaluations (\(\rho = -0.772\) vs. \(-0.477\) for the original evaluator). * Sustainability: Semantic caching successfully served 47.7% of all model calls, optimizing computational overhead.
Metadata & Reference Information
元数据与参考信息
- 主题: 人工智能(
cs.AI);多智能体系统(cs.MA) - 篇幅: 33页,9张图表
- 全文链接: 查看 PDF | HTML 版本 | TeX 源码
- 许可证: 知识共享署名 4.0
Metadata & Reference Information
- Subjects: Artificial Intelligence (
cs.AI); Multiagent Systems (cs.MA)- Length: 33 pages, 9 figures
- Full-Text Links: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0
Associated Tools & Resources
关联工具与资源
- 文献计量工具: Explorer, Connected Papers, Litmaps, scite Smart Citations
- 代码与数据: alphaXiv, CatalyzeX Code Finder, DagsHub, Gotit.pub, Hugging Face, ScienceCast
- 演示与推荐系统: Replicate, Hugging Face Spaces, TXYZ.AI, Influence Flower, CORE Recommender
Associated Tools & Resources
- Bibliographic Tools: Explorer, Connected Papers, Litmaps, scite Smart Citations
- Code & Data: alphaXiv, CatalyzeX Code Finder, DagsHub, Gotit.pub, Hugging Face, ScienceCast
- Demos & Recommenders: Replicate, Hugging Face Spaces, TXYZ.AI, Influence Flower, CORE Recommender