发表机构
Middlebury College(米德尔伯里学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究测试LLMs在亲属关系推理中对熟悉词汇与非ce谓词的呈现差异,发现具体准确性显著高于任意准确性,且差距可通过推理预算和提示干预缩小,表明模型能力受呈现方式影响。
AI 中文摘要
我们测试了大型语言模型在关系以熟悉词汇或明确定义的非ce谓词表达时,是否同样能解决形式上匹配的亲属关系问题。在500对配对图中,局部Qwen3.8-27B的具体准确性超过任意准确性35.6个百分点,Gemma 4 26B-A4B为26.6,Gemma 4 31B为12.0,Qwen3.8-Max为5.4。所有四个配对差距在统计上均显著。推理预算和提示语言干预可以大幅减少这种差异,表明它是可修改的而非固定缺陷。最简结论是行为层面的:在这些任务上,模型表现出的关系能力并非对呈现方式无差异。显式定义提供了形式关系,但并未使非ce谓词像嵌入学习语言关联中的熟悉词汇那样可用。
英文摘要
We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.