拦截袋鼠:基于人工词库、主动探查以及利用大语言模型作为信息提供者与假设提出者的实验性星际语言学
文章背景与核心概要
星际语言学(即研究与那些以不同于人类框架对现实进行分类的智能体进行交流的学科)自 1960 年汉斯·弗赖登塔尔(Hans Freudenthal)提出《Lincos》以来,一直是一项纯粹的推测性事业。本文将星际语言学推向了一门实验科学的新高度。
通过将配置了刻意不兼容的人工词库的大语言模型(LLM)作为信息提供者,作者们将奎因(Quine)的“翻译不确定性”问题(具体研究了“袋鼠效应”,即某个词被错误且无声地附着在错误指代物上的现象)进行了操作化。在 400 多次模拟和实际运行中,作者展示了一套严谨的协议:该协议能成功拦截诱饵、优雅地处理信息提供者带来的噪声,并通过“生成-测试”的 LLM 循环恢复复杂的脚本外术语。
📌 摘要 / Summary
Astrolinguistics—the study of communication with minds that categorize reality differently from human frameworks—has remained a purely speculative endeavor since the introduction of Hans Freudenthal’s Lincos in 1960. This paper transitions astrolinguistics into an experimental science.
Using large language models (LLMs) configured with deliberately incompatible, constructed lexicons as informants, the authors operationalize Quine’s problem of the "indeterminacy of translation" (specifically studying the "kangaroo effect," where a word is silently and incorrectly attached to the wrong referent). Across more than 400 simulated and live runs, the authors demonstrate a rigorous protocol that successfully intercepts decoys, handles informant noise gracefully, and recovers complex out-of-script terms via a generate-and-test LLM loop.
星际语言学——即研究与那些以不同于人类框架对现实进行分类的智能体进行交流的学科——自 1960 年汉斯·弗赖登塔尔(Hans Freudenthal)的《Lincos》问世以来,一直是一项纯粹推测性的事业。本文将星际语言学转变为一门实验科学。
通过利用配置了刻意不兼容、人工构建词库的大语言模型(LLM)作为信息提供者(informants),作者将奎因的“翻译不确定性”问题(特别是研究“袋鼠效应”,即一个词在不知不觉中被错误地附着在错误的指代物上)进行了操作化。在 400 多次模拟和实测运行中,作者展示了一套严谨的协议,能够成功拦截诱饵、优雅地处理信息提供者带来的噪声,并通过生成-测试的 LLM 循环恢复复杂的脚本外术语。
🛠️ 方法论与实验设计 / Methodology & Experimental Design
- The Informants: Two language models equipped with deliberately incompatible category systems:
- Informant A: Encodes shape, color, and motion.
- Informant B: Fuses color with motion, encodes parity, and completely lacks shape-based categorization.
- The Orchestrator: A fully scripted intermediary that translates between the two differing category systems.
- The Protocol: Combines cross-situational elimination, pre-registered predictive probes, active scene selection, a strict recovery round, and a quarantine phase.
- 信息提供者(The Informants): 两个配备了刻意不兼容分类系统的大语言模型:
- 信息提供者 A: 编码形状、颜色和运动。
- 信息提供者 B: 将颜色与运动融合,编码奇偶性,并且完全缺乏基于形状的分类。
- 协调器(The Orchestrator): 一个完全脚本化的中间件,用于在两个不同的分类系统之间进行翻译。
- 协议(The Protocol): 结合了跨情境消除、预注册预测探针、主动场景选择、严格的恢复轮次以及隔离阶段。
📊 关键发现 / Key Findings
- Accuracy & Coverage: Under tested conditions, the full protocol produced zero undetected mistranslations and exceeded a passive baseline’s coverage (\(d = 0.62\)).
- Defeating "Kangaroo Traps": Injected semantic decoys defeated naive ostension and pure statistical learning in 100% of runs. However, the full protocol intercepted every decoy. When discriminating evidence was ontologically unavailable, the protocol correctly declared Quinean equivalence classes rather than guessing.
- Robustness to Noise: Under informant noise, performance degraded gracefully:
- Up to 2% per-word noise: Zero persistent kangaroos.
- At 10% noise: The protocol predominantly abstained rather than generated errors.
- Out-of-Script Generalization: When encountering terms outside the scripted hypothesis space (such as a history-dependent relational term and an XOR contextual homonym), a generate-and-test loop—where an LLM proposed rules and the script verified them—successfully scaled coverage (\(0\% \to 18\% \to 72\% \to 100\%\)) while maintaining zero undetected mistranslations.
- 准确度与覆盖率: 在测试条件下,完整协议实现了零未检测到的错误翻译,并且超越了被动基线的覆盖率(\(d = 0.62\))。
- 击破“袋鼠陷阱”: 注入的语义诱饵在 100% 的运行中击败了朴素的指称法(ostension)和纯统计学习。然而,完整的协议拦截了每一个诱饵。当本体论上无法获得区分性证据时,协议会正确声明奎因等价类(Quinean equivalence classes)而不是盲目猜测。
- 对噪声的鲁棒性: 在信息提供者存在噪声的情况下,性能呈现出平稳的退化:
- 每词噪声高达 2% 时: 零持续存在的“袋鼠”错误。
- 在 10% 噪声时: 协议主要选择弃权,而不是产生错误。
- 脚本外泛化: 当遇到超出脚本化假设空间的术语(例如历史依赖的关系词和 XOR 上下文同音词)时,“生成-测试”循环——即由 LLM 提出规则并由脚本进行验证——成功扩大了覆盖率(\(0\% \to 18\% \to 72\% \to 100\%\)),同时保持了零未检测到的错误翻译。
💡 结论 / Conclusion
The study concludes that under the tested conditions, correctness is a property of the protocol, while coverage is a property of the instruments.
该研究得出结论:在所测试的条件下,正确性是协议的属性,而覆盖率则是仪器的属性。
🔗 全文与访问链接 / Full-Text & Access Links
- View PDF: arXiv:2608.19124 PDF
- HTML Version: arXiv HTML (Experimental)
- TeX Source: arXiv Source
- External Bibliographic Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS
- 查看 PDF: arXiv:2608.19124 PDF
- HTML 版本: arXiv HTML (Experimental)
- TeX 源码: arXiv Source
- 外部文献检索工具:
- Google Scholar
- Semantic Scholar
- NASA ADS