文章背景与核心概要
评估检索增强生成(RAG)系统面临着一个根本性挑战:标准的边缘指标往往无法捕捉各个组件之间如何交互,也无法追踪错误如何在评估流水线中传播。为了解决这一问题,本文作者引入了 RAT(一种统一的贝叶斯评估框架),它根据流水线底层的的信息流,对检索成功率、拒答行为(abstention behavior)以及答案正确性进行联合建模。
该研究的主要贡献包括:通过条件分解区分了任务成功率与生成器在给定检索结果下的合规行为;通过对 27 种 RAG 配置的实验,揭示了在传统边缘指标下看似相同的系统之间存在显著的行为差异;进行了标注分配分析,证明了从信息论角度出发,检索成功标注在估计策略依从性时比任务成功标注提供的信息量大得多;此外,该模型还扩展支持将“LLM作为裁判(LLM-as-a-judge)”的标注整合为经过校准的噪声观测,使从业者能够在统一的概率框架内,将有限的人工判断与低成本的自动化评估高效结合。
The RAT: A Unified Bayesian Model for RAG Evaluation
The RAT: A Unified Bayesian Model for RAG Evaluation
arXiv: [2608.24753 [cs.CL]]
DOI: 10.48550/arXiv.2608.24753
Submitted on: August 25, 2026
Primary Subject: Computation and Language (cs.CL)
Secondary Subjects: Artificial Intelligence (cs.AI)
arXiv: [2608.24753 [cs.CL]]
DOI: 10.48550/arXiv.2608.24753
Submitted on: August 25, 2026
Primary Subject: Computation and Language (cs.CL)
Secondary Subjects: Artificial Intelligence (cs.AI)
Authors
- Pius von Däniken
- Felix Matthias Saaro
- Mark Cieliebak
- Jan Deriu
Authors
- Pius von Däniken
- Felix Matthias Saaro
- Mark Cieliebak
- Jan Deriu
Executive Summary
Evaluating Retrieval-Augmented Generation (RAG) systems presents a fundamental challenge: standard marginal metrics often fail to capture how individual components interact or how errors propagate through the evaluation pipeline.
To address this, the authors introduce The RAT, a unified Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness according to the pipeline's underlying information flow.
Key Contributions:
- Conditional Decomposition: Distinguishes between task success (whether the user received a correct answer stemming from generator success) and whether the generator behaved appropriately given the retrieval outcome. When applied across 27 RAG configurations (spanning 3 datasets, 3 retrievers, and 3 generators), the model reveals significant behavioral differences between systems that otherwise appear identical under marginal metrics.
- Annotation Allocation Analysis: Demonstrates that retrieval-success annotations are substantially more informative than task-success annotations when estimating policy adherence, supported by an information-theoretic explanation for this asymmetry.
- Noisy Observation Calibration: Extends the model to seamlessly incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to effectively combine limited human judgments with cheaper automated assessments within a unified probabilistic framework.
Executive Summary
Evaluating Retrieval-Augmented Generation (RAG) systems presents a fundamental challenge: standard marginal metrics often fail to capture how individual components interact or how errors propagate through the evaluation pipeline.
To address this, the authors introduce The RAT, a unified Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness according to the pipeline's underlying information flow.
Key Contributions:
- Conditional Decomposition: Distinguishes between task success (whether the user received a correct answer stemming from generator success) and whether the generator behaved appropriately given the retrieval outcome. When applied across 27 RAG configurations (spanning 3 datasets, 3 retrievers, and 3 generators), the model reveals significant behavioral differences between systems that otherwise appear identical under marginal metrics.
- Annotation Allocation Analysis: Demonstrates that retrieval-success annotations are substantially more informative than task-success annotations when estimating policy adherence, supported by an information-theoretic explanation for this asymmetry.
- Noisy Observation Calibration: Extends the model to seamlessly incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to effectively combine limited human judgments with cheaper automated assessments within a unified probabilistic framework.
Links & Resources
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution 4.0

Links & Resources
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution 4.0
Citation Tools & References
- BibTeX Citation: Available directly on arXiv
- External Indices:
- Google Scholar
- Semantic Scholar
- NASA ADS
Citation Tools & References
- BibTeX Citation: Available directly on arXiv
- External Indices:
- Google Scholar
- Semantic Scholar
- NASA ADS