跳转至

文章背景与核心概要

本文挑战了学术界将大语言模型(LLM)的输出视为单一“AI语言”的普遍做法。作者提出,LLM具有独特且模型专属的语言特征,这与人类的个体语言特征(idiolects)相媲美。通过对比2024年由六个模型组成的语料库与2026年新生成的语料库,该研究表明,尽管存在明显的代际风格转变,但每个模型仍保持着独特且可识别的语言特征。

作者认为,采用个体语言特征(idiolectal)的框架对于推进法医语言学、文本检测以及基于使用的语言演变研究至关重要。该研究为理解AI生成文本的多样性提供了更坚实的理论基础,并对自然语言处理和计算语言学的未来研究方向产生了深远影响。


超越“AI语言”:大模型输出具有类个体语言特征(Idiolect)的研究

作者: Karolina Rudnicka, Thomas Stephan Juzek
日期: 2026年8月6日
学科: 计算与语言 (cs.CL);人工智能 (cs.AI)
标识符: arXiv:2608.06589


摘要 (Summary)

This paper challenges the common academic practice of treating Large Language Model (LLM) output as a monolithic "AI language." Instead, the authors propose that LLMs possess distinct, model-specific linguistic signatures comparable to human idiolects. By comparing a 2024 corpus of six models with a newly generated 2026 corpus, the study demonstrates that while there is a clear generational shift in style, each model maintains a unique, identifiable linguistic profile. The authors argue that adopting an idiolectal framework is essential for advancing research in forensic linguistics, text detection, and usage-based language evolution.

本文挑战了学术界将大语言模型(LLM)的输出视为单一“AI语言”的普遍做法。相反,作者提出LLM具有独特、模型专属的语言签名,这与人类的个体语言特征(idiolects)具有可比性。通过对比2024年包含六个模型的语料库与2026年新生成的语料库,该研究表明,尽管在风格上存在明显的代际演变,但每个模型都保持着独特且可识别的语言画像。作者认为,采用个体语言特征框架对于推进法医语言学、文本检测以及基于使用的语言演变研究至关重要。


核心发现 (Key Findings)

  • Idiolectal Signatures: Despite the collective "AI language" label, individual models exhibit unique stylistic traits.
  • Generational Evolution: A comparative analysis between 2024 and 2026 cohorts reveals significant shifts in stylistic patterns over time.
  • Quantifiable Variation: The study highlights dramatic differences in linguistic habits; for example, contraction frequencies within the 2026 cohort ranged from 1,200 to over 30,000 per million words.
  • Methodological Impact: The authors suggest that treating LLM output as idiolectal provides a more robust framework for:
    • LLM-generated text detection.
    • Forensic linguistics.
    • Research on language variation and change.
    • Usage-based linguistic approaches.
  • 个体语言特征(Idiolectal Signatures): 尽管有“AI语言”这一统称,但各个模型表现出独特的文体特征。
  • 代际演变: 对2024年和2026年模型群组的对比分析表明,随着时间的推移,文体模式发生了显著变化。
  • 可量化的差异: 该研究强调了语言习惯上的巨大差异;例如,在2026年的模型群组中,缩写(contractions)的使用频率从每百万字1,200次到超过30,000次不等。
  • 方法论影响: 作者建议,将LLM输出视为个体语言特征能够为以下领域提供更稳健的框架:
    • LLM生成文本的检测。
    • 法医语言学。
    • 语言变异与演变研究。
    • 基于使用的语言学方法。

出版详情 (Publication Details)

  • Format: 33 pages, 6 figures, 6 tables.
  • Context: Submitted as a chapter to the post-workshop volume "Corpus Linguistics 2040" (Digital Linguistics series).
  • Classification: MSC 68T50; ACM I.2.7.
  • 格式: 33页,6张图表,6张表格。
  • 背景: 作为章节提交至会后论文集《语料库语言学 2040》(数字语言学系列)。
  • 分类: MSC 68T50; ACM I.2.7.

访问与资源 (Access & Resources)