审视合成回忆录:对照真实生活记录评估大模型生成自传中的场景级胡言乱语
文章背景与核心概要
随着大语言模型(LLM)在长文本生成和个性化创作领域的广泛应用,人们越来越关注其生成个人传记或回忆录时的真实性。本文针对这一问题,进行了首个量化的、场景级的大模型生成自传审计研究,将其与特定主题的真实基准语料库进行了严格对比。研究人员通过极简的提示词(模板、两个示例以及每日名言),利用对话式LLM生成了一部包含366天“每日一页”内容的自传,并在一份独立验证记录的基准上对生成内容的真实性进行了评估。
研究的核心发现在于揭示了大模型在生成传记类内容时的极高“幻觉”与失实率:在366天的内容中,高达96.7%(354天)未能通过验证,缺乏积极的证据支持;同时有5.2%的内容包含与历史记录直接相悖的明确断言。最主要的失控模式被称为“接地气漂移”(Grounded Drift),即模型会将真实存在的实体(如人物、雇主、地点)嵌入到完全虚构的场景中。此外,研究还发现,即便在生成阶段注入真实的主体个人语料库,虽然验证率有所改善,但仍高达83.3%的剩余失败率。这项研究不仅提供了一个可复用的评估工具,也为未来大模型在敏感叙事任务中的落地和纠错提供了重要的量化参考。
摘要 (Summary)
This paper presents the first quantified, scene-level audit of Large Language Model (LLM)-generated autobiography compared against a subject-specific ground-truth corpus. Using a 366-day "page-a-day" conversational LLM-generated memoir—drafted using minimal prompts (a template, two exemplars, and daily quotes) without access to the author's actual corpus—the study evaluates the factual accuracy of the generated narratives against an independent verification record.
Key findings include: * High Failure Rate: 354 out of 366 days (96.7%) failed verification, meaning they lacked positively corroborated scenes. * Active Contradictions: 19 days (5.2%) contained claims explicitly contradicted by the historical record. * Grounded Drift: The dominant failure mode involved mixing real entities (people, employers, settings) into entirely fabricated scenes. * Impact of Grounding: Regenerating the memoir using the subject's personal corpus significantly improved the verification rate, though a substantial residual failure rate of 83.3% remained.
本文呈现了首个量化的、场景级的大语言模型(LLM)生成自传审计,并将其与特定主题的真实基准语料库进行了对比。本研究使用了一部由对话式LLM生成的366天“每日一页”形式的第一人称轶事回忆录——该回忆录是在没有接触作者实际语料库的情况下,使用极简提示词(一个模板、两个示例以及每日名言)起草的——并在分析前确定了四级评估标准,对照独立的验证语料库,在轶事场景层面对生成叙事的真实准确性进行了评估。
主要研究发现包括: * 高失败率: 366天中有354天(占96.7%)未通过验证,这意味着它们缺乏得到积极证实的场景。 * 主动矛盾: 有19天(占5.2%)包含了与历史记录明显相矛盾的主张。 * 接地气漂移:主要的失效模式涉及将真实的实体(人物、雇主、场景)混入完全虚构的场景中。 * 事实基础(Grounding)的影响: 使用主体的个人语料库重新生成回忆录显著提高了验证率,但仍存在83.3%的重大残余失败率。
文档元数据 (Document Metadata)
Field Details arXiv ID arXiv:2608.23640 [cs.AI] Author Heather Renze Submitted August 23, 2026 Primary Subject Artificial Intelligence ( cs.AI)Secondary Subjects Computation and Language ( cs.CL); Computers and Society (cs.CY)ACM Classification I.2.7 DOI 10.48550/arXiv.2608.23640 Resources GitHub Repository
| 字段 | 详情 |
|---|---|
| arXiv ID | arXiv:2608.23640 [cs.AI] |
| 作者 | Heather Renze |
| 提交时间 | 2026年8月23日 |
| 主学科 | 人工智能 (cs.AI) |
| 次学科 | 计算与语言 (cs.CL);计算机与社会 (cs.CY) |
| ACM 分类 | I.2.7 |
| DOI | 10.48550/arXiv.2608.23640 |
| 资源 | GitHub 仓库 |
摘要 (Abstract)
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis.
We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
当要求大语言模型(LLM)书写一个人的生平时,它所写的内容中有多少是真正发生过的?我们呈现了一项场景级的案例研究审计——根据非系统性的文献检索,这是我们所知的首个针对LLM生成的自传与特定主题真实基准语料库进行对比的量化审计。本文的主题对象与作者为同一人:我们利用对话式LLM起草了一本包含366天“每日一页”的第一人称轶事日记。该LLM已知的输入仅为一个模板、两个示例日以及每天的名言(而非作者本人的语料库)。随后,我们使用在分析前确定好的四级评估标准,对照独立的验证语料库,对每一天的内容在轶事场景层面进行了审计。
我们将验证失败率定义为未被评级为“已验证”(场景得到积极证实)的天数占比:366天中有354天失败,占比96.7%(威尔逊95%置信区间为94.4-98.1%)。仅有12天包含得到证实的场景;19天(5.2%)断言了与记录直接相矛盾的主张;主要的失败模式是“接地气漂移”——即虚构场景中包含了真实的人物、雇主和场景——不过其测得的占比在不同的评估员之间有所差异。独立的重新评级复制了这一核心结论(没有证据表明最初的比率被夸大),同时也表明四分类法仅具有一般到中等的可靠性。在相同输入下,使用当前的知名模型重新生成相同的日子,会复现100%的验证失败率;而在主体的语料库中进行事实基础(Grounding)对齐生成,显著提高了验证率,但仍留下显着的残余失败率(83.3%)。我们的贡献包括这项测量结果、一个可重用的审计工具(我们证明其“弱/未验证”的边界是不可靠的),以及一种具有量化效果的事实修正补救措施。
提交历史 (Submission History)
- [v1] Sun, 23 Aug 2026, 20:10:50 UTC (52 KB)
- [v1] 2026年8月23日 星期日 20:10:50 UTC (52 KB)
全文与访问链接 (Full-Text & Access Links)
额外资源与引用 (Additional Resources & Citations)
- Citations: NASA ADS | Google Scholar | Semantic Scholar
- Code & Data Sharing: Available via the Synthetic Memoir Audit GitHub Repository.
- 引用: NASA ADS | Google 学术 | Semantic Scholar
- 代码与数据共享: 可通过 合成回忆录审计 GitHub 仓库 获取。