arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39341cs.AIcs.CLcs.LG

理解即无套利:有界荷兰赌作为语言模型的定义与训练目标

Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models

Daniel Dragonevskiy

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出以无套利定义语言模型理解,通过有界荷兰赌量化逻辑不一致性,并引入Arbitr训练框架显著降低模型的可利用性,同时揭示缩放错觉,表明一致性是知识必要但非充分条件。

中文摘要 AI 辅助

语言模型仅仅是在预测词元,还是真正理解其输出内容?我们通过无套利视角使这一问题变得可量化。若一个计算能力受限的交易者无法通过针对模型在逻辑相关主张上的概率下注而获得保证收益(即“荷兰赌”),则该模型在一定程度上理解了相应词汇。我们建立了三个理论结果:第一,由于完全的逻辑一致性在计算上不可行,理解本质上是分级的,而非绝对的。第二,我们证明标准下一词预测的精确最优解在不同问题格式下固有地不一致;缺陷在于训练目标,而非架构。第三,我们表明不确定性沿推理链可预测地累积,使无根据的过度自信本身成为套利机会。为解决此问题,我们提出Arbitr训练框架,其中对抗性交易者因逻辑不一致而惩罚模型,并配以校准锚点以防止无信息坍缩。在Qwen2.5和Phi-3.5模型上的五项预注册实验中,我们证明标准模型在不同表述下高度可利用。Arbitr将这种可利用性降低数个数量级,且不牺牲任务准确率,该效果成功迁移至未见逻辑模式和新模型家族。关键的是,我们揭示了一种缩放错觉:在7B参数规模下,接近零的测量不一致性常伴随极端且无根据的自信。我们得出结论:尽管Arbitr强制严格的逻辑一致性,但一致性是知识的必要条件,而非充分条件。

英文摘要

Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one

补充信息

↑