arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拦截“袋鼠”:基于人工词汇、主动探测及大型语言模型作为信息提供者与假设提出者的实验天体语言学

Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers

Francesco Cordella, Mauro Cappelli

arXiv 2608.19124首次发表:更新:

发表机构

ENEA(意大利国家新技术、能源与经济可持续发展局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将天体语言学从推测变为实验,通过含不兼容人工词汇的语言模型、脚本协调器及结合多策略的协议,成功拦截翻译不确定性的“袋鼠效应”,提升覆盖度与正确性,还能恢复假设空间外的词语。

AI 中文摘要

天体语言学——即与对现实的分类方式不同于人类的心智进行交流——自Freudenthal于1960年提出Lincos以来一直处于纯粹的推测状态,本文将其变为实验性研究。两个具有刻意不兼容人工词汇的语言模型被用作拥有完整真实信息的信息提供者:其中一个编码形状、颜色和运动,另一个将颜色与运动融合,编码奇偶性且缺少形状编码;同时,一个完全脚本化的协调器负责在两个分类系统之间进行翻译。核心失效模式是“袋鼠效应”:即词语无声地关联到错误指称物——这是Quine的翻译不确定性的可操作化形式。在400多次模拟和实际运行中,一种结合了跨情境消除、预注册预测探测、主动场景选择、更严格的恢复轮次及隔离的协议,在测试条件下未产生未被检测到的误译,且覆盖度超过被动基准(d=0.62)。注入的“袋鼠陷阱”在100%的运行中击败了朴素指示法和纯统计学习,而完整协议拦截了所有诱饵,且在区分性证据在本体论上不可用的情况下,会声明Quine等价类而非猜测。在信息提供者存在噪声时,协议表现出平滑退化:每词噪声率达2%时,无“袋鼠”持续存在;在10%时,协议主要弃权(不执行)而非出错。最后,脚本化假设空间之外的词语(一个依赖历史的关系术语和一个XOR语境同音异义词)通过生成-测试循环被恢复:其中大型语言模型提出规则,脚本验证规则;覆盖度随提出者能力提升而扩展(0%→18%→72%→100%),同时全程未出现未被检测到的误译。在测试条件下,正确性是协议的属性,覆盖度是工具的属性。

英文摘要

Astrolinguistics -- communication with minds that categorize reality differently from ours -- has been purely speculative since Freudenthal's Lincos (1960). We make it experimental. Two language models with deliberately incompatible constructed lexicons (one encoding shape, color, and motion; the other fusing color with motion, encoding parity, and lacking shape) serve as informants with complete ground truth, while a fully scripted orchestrator translates between the two category systems. The central failure mode is the kangaroo effect: the silent attachment of a word to the wrong referent -- Quine's indeterminacy of translation, operationalized. Across 400+ simulated and live runs, a protocol combining cross-situational elimination, pre-registered predictive probes, active scene selection, a stricter recovery round, and quarantine produced no undetected mistranslations under the tested conditions and exceeded a passive baseline's coverage (d = 0.62). Injected kangaroo traps defeated naive ostension and pure statistical learning in 100% of runs, while the full protocol intercepted every decoy and, where discriminating evidence is ontologically unavailable, declared Quinean equivalence classes instead of guessing. Under informant noise it degrades gracefully: zero kangaroos persist up to 2% per-word noise; at 10% the protocol predominantly abstains rather than errs. Finally, words outside the scripted hypothesis space (a history-dependent relational term and an XOR contextual homonym) are recovered by a generate-and-test loop in which an LLM proposes rules and the script verifies them: coverage scales with proposer capability (0% -> 18% -> 72% -> 100%) while undetected mistranslations stayed at zero throughout. In the tested conditions, correctness is a property of the protocol; coverage is a property of the instruments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑