arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05224cs.CV

先做首要之事:教导基于大语言模型的智能体在次要需求前优先满足必须需求

First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

  • School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Tianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao, Gongshen Liu, Zhuosheng Zhang, Cheng Yang

AI总结:

该研究针对现有基于大语言模型的智能体无法按优先级处理用户需求的问题,提出FTF-rl方法,显著提升了智能体的任务成功率及逻辑、数学推理泛化能力。

AI中文摘要:

多模态大语言模型(MLLM)的最新进展极大激发了人们对其作为自主智能体完成现实任务潜力的热情。然而,要求智能体满足用户复杂结构化需求的场景仍未得到充分探索。本研究考察了三种不同需求场景下的推理任务:(i)必须需求唯一确定唯一可行解;(ii)多个答案满足必须需求,并通过次要需求进行优先级排序;(iii)没有候选解满足必须需求,此时智能体应弃权(不执行)。我们在3649个精心构建的问题上评估了最先进的MLLM,这些问题反映了现实服务场景,包括电子商务、预订、基于地图的场景或网约车场景。我们的评估显示,现有MLLM在所有场景中都表现出严重失败,它们经常误解任务需求、违反必须需求并产生无效解。为解决这一关键差距,我们提出了First Things First强化学习(FTF-rl),该方法明确针对多优先级用户需求的推理进行优化。实验结果表明,与强大的基线相比,我们的方法大幅提高了任务成功率。此外,FTF-rl在流行的逻辑和数学推理任务上表现出普遍有效性,包括LogicVista、MathVision和InfoQA。我们的发现表明,增强需求感知推理能力为提高MLLM智能体的泛化能力提供了一条简单而有效的途径。代码和数据集可在此https URL获取。

英文摘要:

Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' complex, structured requirements remain largely underexplored. In this work, we examine reasoning tasks under three distinct requirement scenarios: (i) Must-have requirements uniquely determine a unique feasible solution; (ii) Multiple answers satisfy the must-have requirements and are prioritized via the nice-to-have requirements; and (iii) No candidate solution satisfies the must-have requirements, in which case the agent should abstain from generating a response. We evaluate state-of-the-art MLLMs on 3,649 carefully constructed problems that reflect realistic service scenarios, including e-commerce, booking, and map-based or ride-hailing. Our evaluation reveals that existing MLLMs exhibit catastrophic failures in all scenarios. They frequently misinterpret task requirements, violate must-have requirements, and produce invalid solutions. To address this critical gap, we propose First Things First Reinforcement Learning FTF-rl that explicitly optimizes reasoning over multi-priority user requirements. Experimental results show that our method substantially improves the task success rate compared to strong baselines. Moreover, FTF-rl yields general effectiveness on popular logical and mathematical reasoning tasks, including LogicVista, MathVision, and InfoQA. Our findings suggest that enhancing requirement-aware reasoning capability provides a simple yet effective pathway to improve generalization of MLLM agents. Code and dataset are available at https://github.com/claire62/FTF-RL.

补充信息

↑