arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通用决策模型:超越Jev的基准测试与洞见

General Decision Models: Benchmarking and Insights Beyond Jev

Feiyu Duan, Jiayu Lin, Jia Wang, Jun Xiang, Jialiang Wu, Xinnong Zhang, Hanqi Yan, Siyuan Wang, Zhongyu Wei

arXiv 2610.03935首次发表:更新:

发表机构

Shanghai Innovation Institute; Fudan University; Tongji University; Chinese University of Hong Kong(上海创新研究院; 复旦大学; 同济大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出JEVal双语基准(11,257实例)评估通用决策模型,发现其在证据充分时高效但置信度估计过乐观,长程交互中可靠性下降,并推出InnerJev-4B/27B模型,通过自蒸馏实现与Jev相当性能且推理仅需0.1秒。

AI 中文摘要

通用决策模型(如Jev)最近作为大型语言模型(LLM)在结构化判断和选择任务上的高效替代方案而出现。但这些模型究竟能可靠地处理哪些类型的决策?当个体决策被组合成更大的系统时,它们的行为又会如何变化?为了研究这些问题,我们引入了JEVal,一个包含10个应用领域、36个数据集、共11,257个实例的双语基准,并评估了涵盖通用决策模型和生成式LLM的25种模型配置。我们的结果表明:(1)当决策可以从现有证据中解决时,通用决策模型最具竞争力,但在需要专家知识或可靠的置信度估计时则表现较弱:它们通常能识别最可能的结果,却会大幅高估其概率。(2)在涉及长时程、多步交互的更动态和现实系统中,快速局部决策的优势被系统层面的可靠性失败所抵消。在τ-bench上,更快的局部决策缩短了中位回合时间,但随着决策错误在长轨迹中累积,任务成功率降低。(3)在大规模社会模拟中,决策模型在个体响应预测上接近强生成式LLM,且推理成本显著更低,但在用户画像方面仍较弱,并表现出更大的总体估计误差和系统性偏差。最后,我们提出了InnerJev-4B和InnerJev-27B,它们通过推理到读出自蒸馏(Reasoning-to-Readout Self-Distillation)将开放权重LLM自身的推理内化为单次前向传播的首词决策,其中InnerJev-27B在JEVal上与Jev表现相当,同时回答典型查询仅需约0.1秒。

英文摘要

General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark comprising 11,257 instances from 36 datasets across 10 application domains, and evaluate 25 model configurations spanning general decision models and generative LLMs. Our results show that (1) general decision models are most competitive when decisions can be resolved from available evidence, but weaken when they require specialist knowledge or faithful uncertainty estimation: they can often identify the most likely outcome while substantially overstating its probability. (2) In more dynamic and realistic systems involving long-horizon, multi-step interactions, the advantages of fast local decision making are offset by reliability failures at the system level. on $τ$-bench, faster local decisions reduce median episode time but lower task success as decision errors accumulate over long trajectories. (3) In large-scale social simulation, decision models approach strong generative LLMs on individual response prediction at substantially lower inference cost, yet remain weaker in user profiling and exhibit larger aggregate estimation errors and systematic bias. Finally, we propose InnerJev-4B and InnerJev-27B, which internalize an open-weight LLM's own reasoning into a single-pass first-token decision through Reasoning-to-Readout Self-Distillation, with InnerJev-27B performing on par with Jev on JEVal while answering a typical query in about 0.1 s.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑