arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39496cs.LGcs.CL

当正确答案缺失时:Jev 中算术依赖的拒绝瓶颈

When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev

Jike Zhong, Ming Li, Yuxiang Lai

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示 Jev 模型在算术决策中因依赖数值答案而存在拒绝瓶颈,正确拒绝率仅 7%,通过简单决策阈值可提升至 79% 且不损失答案存在准确率。

中文摘要 AI 辅助

诸如 Jev 之类的类型化决策模型通过直接从预定义选项中选择,为决策工作流中的生成式大语言模型提供了一种高效替代方案。当候选集不包含有效答案时,TypeSafe 建议包含一个“其他”或“以上皆非”选项以支持拒绝。然而,在本报告中,我们识别出一个算术依赖的拒绝瓶颈:Jev 在可用时能可靠地选择正确的数值答案,但当正确答案缺失时,尽管存在明确的拒绝选项,它却经常接受不正确的替代项。在配对的算术问题上,答案存在时的准确率达到 99%,而正确拒绝率降至 7%。此外,这一差距在数值量级、运算深度、上下文表述和拒绝标签上持续存在,并扩展到时间计算和容量取整等场景。然而,原生布尔验证在相同的答案缺失算术案例上实现了 99% 的精确匹配准确率,表明即使模型成功验证了候选正确性,分类式拒绝也可能失败。最后,我们表明,在独立开发问题上选择的简单决策阈值将算术拒绝准确率从 7% 提升至 79%,同时保持 97% 的答案存在准确率,从而在不进行重新训练或额外推理的情况下大幅缓解了该失败。

英文摘要

Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.

发表机构

  • University of Southern California(南加州大学)
  • University of Florida(佛罗里达大学)
  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑