跳出思维定式:语言模型能否有选择地依赖外部指导?
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
- Southern University of Science and Technology(南方科技大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Box$^2$-Bench基准,研究语言模型能否有选择地依赖外部指导,通过反事实监督微调和强化学习提升模型鲁棒性,揭示了选择性依赖是智能体可靠性的关键维度。
AI中文摘要:
智能体框架通常通过人类设计的工作流程来提升语言模型的能力,但随着模型能力的增强,不可靠的指导可能越来越限制其执行。我们将从有用指导中获益同时覆盖不可靠指导的能力称为“跳出思维定式”。我们引入了Box$^2$-Bench基准,该基准在保持模型和任务不变的同时,改变工作流程的可靠性,以隔离模型如何调节其对指导的依赖。在Box$^2$-Bench上,前沿模型通常能从可靠指导中获益,但在指导具有误导性或变得不可靠时仍易受影响。为了测试这种能力是否可学习,我们使用不良工作流程训练两个开放权重模型,保留良好工作流程用于评估。我们探索了两种互补的训练策略:反事实监督微调提高了鲁棒性,而基于结果的强化学习可以将平衡转向更多地使用有帮助的工作流程。我们进一步发现,这种行为扩展到工作流程之外的其他形式的外部信息,改善了同伴纠正和对损坏记忆的鲁棒性。总之,我们的结果确定了对外部信息的选择性依赖是智能体可靠性中一个未被任务性能单独捕获的维度。
英文摘要:
Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.