文章背景与核心概要
本文挑战了学术界将大语言模型(LLM)的输出视为单一“AI语言”的普遍做法。作者提出,LLM具有独特且模型专属的语言特征,这与人类的个体语言特征(idiolects)相媲美。通过对比2024年由六个模型组成的语料库与2026年新生成的语料库,该研究表明,尽管存在明显的代际风格转变,但每个模型仍保持着独特且可识别的语言特征。
作者认为,采用个体语言特征(idiolectal)的框架对于推进法医语言学、文本检测以及基于使用的语言演变研究至关重要。该研究为理解AI生成文本的多样性提供了更坚实的理论基础,并对自然语言处理和计算语言学的未来研究方向产生了深远影响。
超越“AI语言”:大模型输出具有类个体语言特征(Idiolect)的研究
作者: Karolina Rudnicka, Thomas Stephan Juzek
日期: 2026年8月6日
学科: 计算与语言 (cs.CL);人工智能 (cs.AI)
标识符: arXiv:2608.06589
摘要 (Summary)
This paper challenges the common academic practice of treating Large Language Model (LLM) output as a monolithic "AI language." Instead, the authors propose that LLMs possess distinct, model-specific linguistic signatures comparable to human idiolects. By comparing a 2024 corpus of six models with a newly generated 2026 corpus, the study demonstrates that while there is a clear generational shift in style, each model maintains a unique, identifiable linguistic profile. The authors argue that adopting an idiolectal framework is essential for advancing research in forensic linguistics, text detection, and usage-based language evolution.
本文挑战了学术界将大语言模型(LLM)的输出视为单一“AI语言”的普遍做法。相反,作者提出LLM具有独特、模型专属的语言签名,这与人类的个体语言特征(idiolects)具有可比性。通过对比2024年包含六个模型的语料库与2026年新生成的语料库,该研究表明,尽管在风格上存在明显的代际演变,但每个模型都保持着独特且可识别的语言画像。作者认为,采用个体语言特征框架对于推进法医语言学、文本检测以及基于使用的语言演变研究至关重要。
核心发现 (Key Findings)
- Idiolectal Signatures: Despite the collective "AI language" label, individual models exhibit unique stylistic traits.
- Generational Evolution: A comparative analysis between 2024 and 2026 cohorts reveals significant shifts in stylistic patterns over time.
- Quantifiable Variation: The study highlights dramatic differences in linguistic habits; for example, contraction frequencies within the 2026 cohort ranged from 1,200 to over 30,000 per million words.
- Methodological Impact: The authors suggest that treating LLM output as idiolectal provides a more robust framework for:
- LLM-generated text detection.
- Forensic linguistics.
- Research on language variation and change.
- Usage-based linguistic approaches.
- 个体语言特征(Idiolectal Signatures): 尽管有“AI语言”这一统称,但各个模型表现出独特的文体特征。
- 代际演变: 对2024年和2026年模型群组的对比分析表明,随着时间的推移,文体模式发生了显著变化。
- 可量化的差异: 该研究强调了语言习惯上的巨大差异;例如,在2026年的模型群组中,缩写(contractions)的使用频率从每百万字1,200次到超过30,000次不等。
- 方法论影响: 作者建议,将LLM输出视为个体语言特征能够为以下领域提供更稳健的框架:
- LLM生成文本的检测。
- 法医语言学。
- 语言变异与演变研究。
- 基于使用的语言学方法。
出版详情 (Publication Details)
- Format: 33 pages, 6 figures, 6 tables.
- Context: Submitted as a chapter to the post-workshop volume "Corpus Linguistics 2040" (Digital Linguistics series).
- Classification: MSC 68T50; ACM I.2.7.
- 格式: 33页,6张图表,6张表格。
- 背景: 作为章节提交至会后论文集《语料库语言学 2040》(数字语言学系列)。
- 分类: MSC 68T50; ACM I.2.7.
访问与资源 (Access & Resources)
- View Full PDF
- DOI: https://doi.org/10.48550/arXiv.2608.06589
- Citations: NASA ADS | Google Scholar | Semantic Scholar
- 查看完整 PDF
- DOI: https://doi.org/10.48550/arXiv.2608.06589
- 引用: NASA ADS | Google Scholar | Semantic Scholar