文章背景与核心概要
针对穆斯林群体的网络仇恨言论往往表现为隐蔽的、具有文化代码的多语言表达,这使得传统的AI内容审核工具难以察觉。尽管标准模型可以实现较高的分类准确率,但它们通常作为“黑盒”运行,存在偏见、过度审查或审核不足的重大风险——特别是当脱离必要的社会文化背景时。
为了克服这些挑战,本文提出了一种训练期可解释性框架,旨在将模型的推理过程直接与人类标注的理由(rationales)进行对齐,从而同步提升分类性能与可解释性。该研究在英文数据集 HateXplain 和印地英语(Hinglish)数据集 BullySent 上进行了评估,证明了基于梯度和注意力的正则化技术不仅能提高F值,还能增强模型解释的合理性与忠实度,为构建多语言、具文化感知能力的内容审核系统提供了有前景且可扩展的路径。
训练期可解释性对于多语言仇恨言论检测:将模型推理与人类理由对齐
arXiv ID: arXiv:2608.26125 [cs.CL]
作者: Muhammad Deedahwar Mazhar Qureshi, Sannaan Khan, Muhammad Atif Qureshi, Wael Rashwan
会议: 已被 NeurIPS Workshops 2025 接受
提交时间: 2026年6月22日
📌 执行摘要
Online hate targeting Muslim communities frequently manifests in subtle, culturally coded, and multilingual expressions that evade conventional AI moderation tools. Although standard models can achieve high classification accuracy, they often operate as "black boxes," carrying significant risks of bias, over-censorship, or under-moderation—especially when detached from essential sociocultural contexts.
To overcome these challenges, this paper proposes a training-time explainability framework designed to align model reasoning directly with human-annotated rationales, thereby simultaneously boosting classification performance and interpretability.
针对穆斯林群体的网络仇恨经常表现为微妙的、带有文化密码的多语言表达,这些表达往往能规避常规的AI审核工具。尽管标准模型可以实现高分类准确率,但它们通常作为“黑盒”运行,携带着偏见、过度审查或审核不足的巨大风险——特别是当脱离了核心的社会文化背景时。为了克服这些挑战,本文提出了一种训练期可解释性框架,旨在将模型推理直接与人工标注的理由对齐,从而同时提升分类性能和可解释性。
🔍 核心亮点与方法论
- 核心创新: 引入了一种训练期正则化方法,利用基于梯度和注意力的方法引导模型进行类似人类的推理。
- 评估数据集: 在 HateXplain(英语)和 BullySent(印地英语)上进行了测试,有效地捕捉了反穆斯林仇恨在不同语言中的流行程度和细微差别。
- 可解释性技术: 使用多种归因工具评估了解释的质量、准确性以及跨方法的一致性:
- LIME
- 积分梯度 (Integrated Gradients)
- 梯度 \(\times\) 输入 (Grad \(\times\) Input)
- 注意力机制 (Attention mechanisms)
- Core Innovation: Introduces a training-time regularization approach using gradient- and attention-based methods to guide models toward human-like reasoning.
- Evaluation Datasets: Tested on HateXplain (English) and BullySent (Hinglish), effectively capturing the prevalence and nuances of anti-Muslim hate across different languages.
- Interpretability Techniques: Evaluates explanation quality, accuracy, and cross-method agreement using several attribution tools:
- LIME
- Integrated Gradients
- Grad \(\times\) Input
- Attention mechanisms
📈 主要发现
- 基于梯度和注意力的正则化技术显著提高了模型的 F值(F-scores)。
- 该框架增强了模型解释的合理性(plausibility)和忠实度(faithfulness)。
- 模型成功捕捉到了识别隐性反穆斯林仇恨言论所需的特定文化语言线索。
- 为通往真正多语言、具文化感知的内容审核系统提供了一条有前景且可扩展的路径。
- Regularization techniques based on gradients and attention significantly improve model F-scores.
- The framework enhances both plausibility and faithfulness in model explanations.
- Models successfully capture culturally specific linguistic cues necessary for identifying implicit anti-Muslim hate speech.
- Offers a promising, scalable path toward truly multilingual, culturally aware content moderation systems.
🔗 链接与资源
- 全文访问:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 数字对象唯一标识符 (DOI): 10.48550/arXiv.2608.26125
- 许可协议: 知识共享署名 4.0 国际版
- Full-Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- Digital Object Identifier (DOI): 10.48550/arXiv.2608.26125
- License: Creative Commons Attribution 4.0 International
📜 摘要
针对穆斯林群体的网络仇恨通常以带有文化密码的多语言形式出现,从而规避了传统的AI审核。此类系统尽管准确,但仍然不透明,并存在偏见、过度审查或审核不足的风险,特别是当脱离社会文化背景时。我们提出了一种训练期可解释性框架,将模型推理与人类标注的理由对齐,从而改善分类性能和可解释性。我们的方法在 HateXplain(英语)和 BullySent(印地英语)上进行了评估,反映了这两种语言中反穆斯林仇恨的普遍存在。使用 LIME、积分梯度、Grad X Input 和注意力机制,我们评估了准确性、解释质量以及跨方法的一致性。结果表明,基于梯度和注意力的正则化提高了 F 值,增强了合理性和忠实度,并捕获了用于检测隐性反穆斯林仇恨的特定文化线索,为多语言、具文化感知的内容审核提供了一条路径。
Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a training-time explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.