发表机构
Max-Planck-Institute for Biological Cybernetics; University Hospital Tübingen(马克斯·普朗克生物控制论研究所; 图宾根大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型的溯因推理能力,引入Elenchos评估框架,通过给定形式系统及变异对应物,让智能体判断变异并推断规则修改,发现模型存在检测-归因分离问题,相互作用变异下性能降,推理时间收益递减。
AI 中文摘要
大语言模型在模式识别和文本生成方面表现出色,但其溯因推理能力仍 poorly understood。本文引入了Elenchos(以苏格拉底式交叉检验方法命名),这是一个生成式评估框架,将溯因推理作为结构反问题来衡量。给定一个参考形式系统和一个可能变异的对应物,智能体必须确定是否发生了变异并推断出导致行为差异的规则修改。评估前沿和中层大语言模型发现了一致的检测-归因分离:模型常能识别系统已改变,但难以识别导致差异的潜在变异。在相互作用变异下性能大幅下降,模型常只能找回部分潜在变异。初步证据还表明增加推理时间收益递减,更大推理预算下改善有限,不过这一发现需进一步验证。
英文摘要
Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for the resulting behavioral differences. Evaluating frontier and mid-tier LLMs reveals a consistent detection-attribution dissociation: models often recognize that a system has been altered but struggle to identify the latent mutations causing the observed discrepancies. Performance degrades substantially under interacting mutations, where models frequently recover only a subset of the underlying mutations. Preliminary evidence also suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets, though this finding requires further validation.