使用多重编码器实现计算高效的自动化创造力评估
文章背景与核心概要
自动化创造力评估(Automated Creativity Assessment, ACA)长期以来面临着两难境地:传统的评估方法要么依赖于计算资源消耗巨大的大型语言模型(LLM),要么采用缺乏实际准确性的简单方法。为了解决这一痛点,本文作者引入了一种基于多重编码器(Poly-Encoders)的新颖方法,成功在保持高评估准确性的同时实现了极高的计算效率。
研究团队利用包含约 18,000 条人工评分问答响应的《科学创造性思维测试》(Scientific Creative Thinking Test)公共数据集,对基于小型预训练 BERT 家族编码器的 Poly-Encoder 进行了微调。实验结果表明,该方法与人类评分员之间的皮尔逊相关系数最高可达 \(r = 0.74\)(95% 置信区间为 \([0.73, 0.75]\)),其性能完全可以媲美资源密集型的大型语言模型。这项研究弥合了高性能与低计算开销之间的鸿沟,使得在消费级硬件上大规模实施自动化创造力评估成为可能,尤其在教育场景中具有广阔的应用前景。
摘要 (Summary)
本文介绍了一种利用多重编码器(Poly-Encoders)进行自动化创造力评估的新颖方法,在保持高评估准确性的同时兼顾了计算效率。传统上,自动化创造力评估依赖于资源消耗巨大的大型语言模型(LLM),或者缺乏实际准确性的简单方法。
This paper introduces a novel approach to automated creativity assessment utilizing Poly-Encoders, bridging the gap between high evaluation accuracy and computational efficiency. Traditionally, automated creativity assessment relied on resource-intensive Large Language Models (LLMs) or simplistic methods lacking practical accuracy.
作者在一个包含约 18,000 条人类评分问答响应的《科学创造性思维测试》(Scientific Creative Thinking Test)公共数据集上,对 Poly-Encoder 进行了微调(利用了小型预训练的 BERT 家族编码器)。所提出的方法取得了与大型语言模型相媲美的性能,与人类评分员的皮尔逊相关系数高达 \(r = 0.74\)(95% 置信区间 \([0.73, 0.75]\))。这种高效率显著降低了计算需求,使得在消费级硬件上进行可扩展的自动化创造力评估成为可行,特别是在教育语境中。
The authors fine-tuned a Poly-Encoder—leveraging small pre-trained BERT-family encoders—on a public dataset containing approximately 18,000 human-rated question responses from the Scientific Creative Thinking Test. The proposed method achieved performance comparable to heavy LLMs, registering Pearson correlations of up to \(r = 0.74\) (95% CI \([0.73, 0.75]\)) with human raters. This efficiency significantly lowers computational demands, making scalable, automated creativity assessment viable on consumer-grade hardware, particularly within educational contexts.
论文元数据 (Paper Metadata)
- arXiv ID: arXiv:2608.26165 [cs.CL]
- 作者: Sam Grouchnikov, Phillip Gregory, Jiho Noh
- 提交时间: 2026年7月13日
- 录用会议: AIED 2026 (DOI: 10.1007/978-3-032-29755-6_32)
- 主要学科: 计算与语言 (
cs.CL) - 次要学科: 人工智能 (
cs.AI)
- arXiv ID: arXiv:2608.26165 [cs.CL]
- Authors: Sam Grouchnikov, Phillip Gregory, Jiho Noh
- Submitted: July 13, 2026
- Accepted Venue: AIED 2026 (DOI: 10.1007/978-3-032-29755-6_32)
- Primary Subject: Computation and Language (
cs.CL)- Secondary Subject: Artificial Intelligence (
cs.AI)
摘要原文 (Abstract)
Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally efficient and accurate automated creativity assessment. We fine-tuned a Poly-Encoder on a public dataset from the Scientific Creative Thinking Test, comprised of approximately 18,000 human-rated question responses. Our method leverages small pre-trained BERT encoders, achieving performance comparable to fine-tuned Large Language Models while significantly reducing computational demands. Experiments with the BERT-family models and poly-code counts achieved Pearson correlations of up to \(r = 0.74\), 95% CI \([0.73, 0.75]\) with human raters, matching the performance of resource intensive LLMs. This study bridges the gap between high performance and computational efficiency, potentially enabling widespread implementation of automated creativity assessment on accessible consumer-grade hardware. With some limitations, our findings suggest that Poly-Encoders are a promising alternative to LLMs for practical, scalable creativity assessment in various contexts, especially educational.
访问链接与资源 (Access Links & Resources)
- Full-Text: View PDF | HTML (Experimental) | TeX Source
- External Indices:
- Google Scholar
- Semantic Scholar
- NASA ADS