文章背景与核心概要
随着大型语言模型(LLM)被越来越多地部署为信息检索(IR)评估中的相关性评估员,理解框架(Framing)对判断可靠性的影响至关重要。本文引入了“角色条件设定”(Persona Conditioning)作为一种诊断机制,用于探测LLM评估员的敏感性。
该研究利用多样化的面向任务的角色(来自PersonaHub和NVIDIA Nemotron-Personas-USA),实例化了五个独特的评估员角色——分别专注于意图解释、领域专业知识、对比判断、证据验证和全局搜索质量评估——并辅以标准的UMBRELA基线。通过在TREC DL20和RAG24上的六个LLM主干模型进行评估,研究表明:LLM评估员的敏感性呈现出结构化而非均匀分布的特征,通常表现为调整评估的严格度或解释,而不会引发大规模的相关性反转;高容量模型能够保持系统排名的致性,而较小的模型则表现出更大的角色诱导不稳定性;敏感性高度集中在特定的系统类型中,特别是DL20上的神经排序器/重排器以及RAG24上的RAG管道。
Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation
As large language models (LLMs) are increasingly deployed as relevance assessors in information retrieval (IR) evaluation, understanding how framing impacts judgment reliability is critical. This paper introduces persona conditioning as a diagnostic mechanism to probe LLM assessor sensitivity.
Using diverse task-oriented personas (from PersonaHub and NVIDIA Nemotron-Personas-USA), the authors instantiate five unique assessor roles—focusing on intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment—alongside a standard UMBRELA baseline. Evaluating across six LLM backbones on TREC DL20 and RAG24, the study demonstrates that: * LLM assessor sensitivity is structured rather than uniform, typically adjusting assessment strictness or interpretation rather than triggering wide-scale relevance reversals. * High-capacity models maintain system-ranking agreement, whereas smaller models exhibit greater persona-induced instability. * Sensitivity is heavily concentrated in specific system types, particularly neural rankers/rerankers on DL20 and RAG pipelines on RAG24.
📌 Summary
📌 Summary
arXiv: 2608.10385 [cs.IR]
Accepted at: CIKM 2026
Submitted: August 11, 2026
Authors: Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini
arXiv: 2608.10385 [cs.IR]
Accepted at: CIKM 2026
Submitted: August 11, 2026
Authors: Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini
📋 Metadata
📋 Metadata
| Field | Details |
|---|---|
| Primary Subject | Information Retrieval (cs.IR) |
| Secondary Subjects | Artificial Intelligence (cs.AI) |
| DOI | 10.48550/arXiv.2608.10385 |
| Full-Text Links | View PDF | HTML Version | TeX Source |
| License | Creative Commons Attribution 4.0 |
Field Details Primary Subject Information Retrieval ( cs.IR)Secondary Subjects Artificial Intelligence ( cs.AI)DOI 10.48550/arXiv.2608.10385 Full-Text Links View PDF | HTML Version | TeX Source License Creative Commons Attribution 4.0
🛠️ Additional Resources & Features
🛠️ Additional Resources & Features
- Code & Data Integration: Explore code and implementation findings via Hugging Face, CatalyzeX Code Finder, and DagsHub.
- Bibliographic Tools: Connect via Google Scholar, Semantic Scholar, NASA ADS, and Connected Papers.
- Code & Data Integration: Explore code and implementation findings via Hugging Face, CatalyzeX Code Finder, and DagsHub.
- Bibliographic Tools: Connect via Google Scholar, Semantic Scholar, NASA ADS, and Connected Papers.