文章背景与核心概要
随着大语言模型(LLM)在创意写作领域的应用日益广泛,如何客观、准确地评估其生成内容的创造力成为了一个关键挑战。本文由 Alessandro Tutone 等研究人员撰写,深入探讨了当前自动化评估方法在衡量大语言模型创造力方面的有效性。
研究团队通过收集人类对来自 WritingPrompts 数据集的人类创作与 AI 生成短篇故事在 11 个创造力维度上的评估,将其与自动化客观指标以及“LLM-as-a-Judge”(大模型作为裁判)框架的评分进行了系统性对比。结果表明,现有的自动化评估方法与人类主观判断之间存在严重的错位:传统自动化指标与人类感知几乎呈零相关;而基于大模型的裁判则表现出对 AI 生成内容的系统性偏见,更青睐其表面文采,却忽视了人类创造力中至关重要的不可预测性和深度。该研究凸显了将多维度、主观性的创造力简化为计算指标的内在局限性。
大语言模型创造力自动评估的局限性
作者: Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
日期: 2026年8月24日(v2版本:2026年8月27日)
DOI: 10.48550/arXiv.2608.23705
学科: 计算与语言 (cs.CL);人工智能 (cs.AI);计算机与社会 (cs.CY)
摘要 (Summary)
这项研究调查了当前用于评估大语言模型(LLM)创造力的自动化方法的有效性。通过对比人类对短篇故事(来自 WritingPrompts 数据集)的判断与自动化指标及“LLM-as-a-Judge”框架的结果,作者发现了一个关键的错位。研究表明,自动化指标与人类感知表现出接近零的相关性,而基于 LLM 的裁判对 AI 生成的内容表现出系统性偏见,偏爱风格上的润饰,而非人类创造力特征所具有的不可预测性和深度。这些发现强调了量化创造力这一主观且多维度表达的固有难度。
This research investigates the efficacy of current automated methods for evaluating creativity in Large Language Models (LLMs). By comparing human judgments of short stories (from the WritingPrompts dataset) against automated metrics and "LLM-as-a-Judge" frameworks, the authors identify a critical misalignment. The study reveals that automated metrics show near-zero correlation with human perception, while LLM-based judges exhibit a systematic bias toward AI-generated content, favoring stylistic polish over the unpredictability and depth characteristic of human creativity. The findings underscore the inherent difficulty in quantifying the subjective, multidimensional nature of creative expression.
摘要正文 (Abstract)
大语言模型(LLM)生成文本的能力日益增强,在需要创造力的领域甚至能够挑战人类表现,然而评估 LLM 生成内容中的创造力仍然是一项重大挑战。
在此,我们探讨了当前的自动评估方法是否能够可靠地捕捉人类对创造力的判断。我们收集了人类对来自 WritingPrompts 数据集的人类和 AI 生成的短篇故事在 11 个创造力维度上的评估,并将这些判断与自动客观指标及 LLM-as-a-Judge 评估进行了对比。
我们的实验揭示了自动评估与人类评估之间存在实质性的错位。特别是,基于 LLM 的裁判对 AI 生成的故事表现出系统性的偏好,一致地偏爱其风格特征,而忽视了人类创作文本的不可预测性及其他特质。此外,相关性分析表明,广泛使用的自动指标在人类生成和 AI 生成的故事中,与人类判断的对齐度都接近于零,这表明它们未能捕捉到创造力的重要维度。这些发现凸显了当前创意文本自动评估方法中的根本局限性,并强调了将创造力的多维度和主观本质简化为计算指标的难度。
Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge.
Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations.
Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.