文章背景与核心概要
传统的AI监督通常依赖标准答案(ground truth)来验证模型行为,但在处理具争议性、道德模糊的场景时,这种方法往往捉襟见肘。本文引入了一种替代性的评估框架:当AI模型的裁决受到批判性问题挑战时,测量其为自身裁决辩护的辩护结构质量。
通过运用基于沃尔顿(Walton)论证方案和戈维尔(Govier)论证说服力标准的四阶段辩证协议,作者对200个高度模糊的道德困境中的9个前沿模型进行了评估。研究发现,虽然模型对其推理的辩护普遍远高于评分标准的最低要求,但其弱点主要集中在论据(grounds)和充分性(sufficiency)上。此外,模型在事后辩护中呈现的论证方案往往与其底层的实际推理轨迹有所不同,这为未来的AI对齐策略提供了重要的思考方向。
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
- arXiv ID: arXiv:2609.05088 [cs.AI]
- Authors: Daan R. Henselmans, Derck W.E. Prinzhorn, Arno Libert
- Submitted: September 4, 2026
- Publication: Accepted for publication in the Paris Journal of AI and Digital Ethics (2026); presented at PCAIDE 2026
- arXiv ID: arXiv:2609.05088 [cs.AI]
- Authors: Daan R. Henselmans, Derck W.E. Prinzhorn, Arno Libert
- Submitted: September 4, 2026
- Publication: Accepted for publication in the Paris Journal of AI and Digital Ethics (2026); presented at PCAIDE 2026
Executive Summary
传统的人工智能监督依赖标准答案来验证模型行为,但在处理充满争议、道德模糊的场景时,这种方法显得力不从心。本文引入了一种替代评估框架:测量AI模型在面对批判性问题质问时,为其裁决辩护的辩护结构质量。
Traditional AI oversight relies on ground truth to validate model behaviors, a method that struggles when dealing with contested, morally ambiguous scenarios. This paper introduces an alternative evaluation framework: measuring the structural quality of the defense an AI model can mount for its verdicts when challenged by critical questions.
通过使用基于沃尔顿论证方案和戈维尔论证说服力标准的四阶段辩证协议,作者在200个高模糊度道德困境中评估了九个前沿模型。研究结果表明,尽管模型在辩护其推理时普遍远高于评分标准的最低要求,但弱点主要集中在论据和充分性方面。此外,与其实际底层推理轨迹相比,模型在事后辩护中往往会呈现出不同的论证方案,这突显了未来AI对齐策略中需要重点考虑的重要问题。
Using a four-phase dialectical protocol grounded in Walton's argumentation schemes and Govier's criteria for argument cogency, the authors evaluated nine frontier models across 200 high-ambiguity moral dilemmas. The findings reveal that while models generally defend their reasoning well above rubric minimums, weaknesses concentrate on grounds and sufficiency. Furthermore, models often present different argument schemes in their post-hoc justifications compared to their actual underlying reasoning tracks, highlighting important considerations for future AI alignment strategies.
Abstract
人工智能监督方法依赖标准答案进行验证,但什么才算适当的AI行为本身就存在争议。这使得大语言模型(LLM)和基于辩论的监督中的道德推理评估回避了现实中的模糊性。我们研究了一种替代标准,旨在即使在存在这种模糊性的情况下也能发挥作用:测量模型在回应批判性问题时为其裁决进行辩护的结构质量,该质量通过基于沃尔顿论证方案理论和戈维尔论证说服力标准的四阶段辩证协议进行测量。
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency.
该协议能够适应不同的推理框架,超越了多项选择的框架,并同时处理裁决之前的推理及其事后辩护。在9个前沿模型和200个高模糊度的 MoralChoice 项目(包含 \(6,778\) 个经裁判评分的单元格,二进制失败判断的裁判间一致性达到 \(89.6\%\))中,模型在各个维度上的推理辩护都远高于评分标准的最低要求。
The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items (\(6,778\) judge-scored cells, validated against \(89.6\%\) inter-judge agreement on the binary failure judgment), models defend their reasoning well above the rubric minimum on every dimension.
失败主要集中在论据和充分性上,并且与认识论对冲(epistemic hedging)相关,而与论证长度无关。在每个模型和每个戈维尔维度上,推理过程的辩护效果都优于事后辩护。尽管基于价值的实践推理在两条轨迹中都占主导地位,但在很大一部分困境(每个模型 \(\ge 20\%\))中,模型在其辩护中呈现的方案与其推理时所用的方案不同。该协议能够捕捉到严格意义上不可辩护的辩护(自我矛盾、虚假前提),并显现出在AI对齐中描述撤回(retraction)作用的困难,这表明我们需要进行更多情境化的评估。
Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas (\(\ge 20\%\) per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.
Submission & Reference Details
- Subjects: Artificial Intelligence (
cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY) - ACM Classes: I.2.0; I.2.3; I.2.7; K.4.1
- DOI: 10.48550/arXiv.2609.05088
- Full-Text Links: View PDF | HTML Version | TeX Source
- Subjects: Artificial Intelligence (
cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)- ACM Classes: I.2.0; I.2.3; I.2.7; K.4.1
- DOI: 10.48550/arXiv.2609.05088
- Full-Text Links: View PDF | HTML Version | TeX Source
