arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

系统切换:快速决策模型何时应停下来思考?

System Switch: When Should a Fast Decision Model Stop and Think?

Gian Luca Bailo

arXiv 2610.09683首次发表:更新:

发表机构

Independent Researcher, Recco (Genoa), Italy(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究快速决策模型何时应暂停并交由推理模型处理,在Doom环境中测试门控切换策略,发现延迟低置信度决策可提升表现,但闭环中无变体到达出口,且推理模型会误判门状态。

AI 中文摘要

双过程智能体将快速策略与缓慢的深思熟虑模型配对。在实时环境中,缓慢模型通常持续运行;在回合制智能体和机器人规划器中,它会在不确定性或检测到失败等事件时被调用。我们研究了一个快速学习的行动者,它做出每一个决策,并且仅在门控开启时将控制权交给推理视觉语言模型,而游戏则持续运行。我们使用闭环Doom和新的开放“System One”类型化决策模型,通过一个通用的HTTP接口提供服务。在900个保留问题上,(i)从0.15B到9B参数规模的零样本决策模型在错误中选择收集物品的频率是随机概率的1.6至1.8倍,且不受选项顺序影响,尽管顺序会改变某些模型的准确率;(ii)准确率、校准度和敏感性(置信度区分正确与错误答案的能力)是截然不同的:相似准确率的模型在AUROC上差异很大,且最敏感模型的置信度追踪的是它失败的情境类型,而非哪些答案错误;(iii)离线情况下,将最不自信的30%决策延迟给推理模型所获得的收益,与行动者的AUROC成正比(秩相关0.87);在保留游戏上选择行动者和比率时,使用doomLaya的选项顺序收益为+0.13 [0.08, 0.18],使用打乱选项收益为+0.08 [0.02, 0.14],其中推理贡献了大约一半;(iv)在闭环中(33个游戏,三个随机种子),没有变体到达出口。承诺执行计划,无论是推理者的还是固定探索规则的,会打开更多门,并使静止不动的行动者开始行动;使用规则时,智能体更频繁死亡。当被告知某些门需要钥匙时,推理者将普通门误认为锁门,而状态无法区分这两者;没有该知识时,它又回到收集行为。我们发布了代码、提示、数据和日志。

英文摘要

Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models' accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor's AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya's option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner's or a fixed explore rule's, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.

Comments14 pages, 1 figure, 4 tables. Code, prompts and data: https://github.com/dexmac221/doomgemma (branch system-switch)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑