跳转至

基于认知动机的隐喻解释多维评估框架

文章背景与核心概要

当前的隐喻解释评估方法通常依赖于整体质量评分,这无法揭示解释质量的结构组成,也无法展现人类判断在哪些地方达成共识或产生分歧。为了解决这一局限性,作者引入了一个具有认知动机的框架,将隐喻解释的质量细分为六个理论基础扎实的维度。

通过包含 11,200 个评分的密集标注研究,研究人员发现:1. 隐喻解释质量本质上是多维的;2. 标注者的分歧具有系统性而非随机性;3. 这六个维度自然地融合为一个共享聚类,并伴随两个独立的判断轴。一项探索性可行性研究进一步证明,标准的自动评估流程能够恢复部分底层结构——成功预测最具区分度的维度,同时将其错误直接与人类分歧相关联。这些发现表明,与整体评分相比,多维评估能够提供更丰富的诊断见解,并且针对开放式生成任务的自动评估器应当根据其保持人类判断结构完整性的能力来进行评估。


Summary

Current methods for evaluating metaphor explanations typically rely on holistic quality ratings, which fail to reveal how explanation quality is structured or where human judgments converge and diverge. To address this limitation, the authors introduce a cognitively motivated framework that breaks down metaphor explanation quality into six theoretically grounded dimensions.

Through a dense annotation study involving 11,200 ratings, the researchers discovered that: 1. Explanation quality is inherently multidimensional. 2. Annotator disagreement is systematic rather than random. 3. The six dimensions naturally collapse into a single shared cluster alongside two independent axes of judgment.

An exploratory feasibility study further demonstrates that standard automatic evaluation pipelines can recover parts of this underlying structure—successfully predicting the most discriminative dimensions while correlating their errors directly with human disagreement. These findings suggest that multidimensional evaluation offers richer diagnostic insights than holistic scores, and that automatic evaluators for open-ended generation tasks should be assessed based on how well they preserve the structural integrity of human judgment.

Currents methods for evaluating metaphor explanations typically rely on holistic quality ratings, which fail to reveal how explanation quality is structured or where human judgments converge and diverge. To address this limitation, the authors introduce a cognitively motivated framework that breaks down metaphor explanation quality into six theoretically grounded dimensions.

Through a dense annotation study involving 11,200 ratings, the researchers discovered that: 1. Explanation quality is inherently multidimensional. 2. Annotator disagreement is systematic rather than random. 3. The six dimensions naturally collapse into a single shared cluster alongside two independent axes of judgment.

An exploratory feasibility study further demonstrates that standard automatic evaluation pipelines can recover parts of this underlying structure—successfully predicting the most discriminative dimensions while correlating their errors directly with human disagreement. These findings suggest that multidimensional evaluation offers richer diagnostic insights than holistic scores, and that automatic evaluators for open-ended generation tasks should be assessed based on how well they preserve the structural integrity of human judgment.


Document Metadata

Field Details
arXiv ID arXiv:2608.15828 [cs.CL]
Subjects Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Authors Ana Naveriani, Jakob Suchan, Stefano Zoia, Mehul Bhatt, Antonio Lieto, Gian Luca Pozzato
Submission Date August 16, 2026
Comments Preprint of paper accepted at INLG 2026
DOI 10.48550/arXiv.2608.15828
Field Details
arXiv ID arXiv:2608.15828 [cs.CL]
Subjects Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Authors Ana Naveriani, Jakob Suchan, Stefano Zoia, Mehul Bhatt, Antonio Lieto, Gian Luca Pozzato
Submission Date August 16, 2026
Comments Preprint of paper accepted at INLG 2026
DOI 10.48550/arXiv.2608.15828