AI 中文总结
研究针对任务分解中因直接用检索指标作奖励易致奖励作弊及削弱域外泛化能力的问题,提出偏好引导反事实任务分解框架PCTD,通过反事实和偏好奖励改进,构建基准MTDTool,实验证明其在多方面超越现有方法。
AI 中文摘要
任务分解旨在将模糊指令转化为可执行的原子子任务,以指导高精度工具检索。但直接采用工具检索指标(如召回率或NDCG)作为任务分解奖励,易在基于强化学习的方法中引发奖励作弊。模型倾向通过重复分解等策略最大化检索匹配,这削弱了在涉及未见工具的域外场景中的泛化能力。为解决此问题,我们提出PCTD,一种偏好引导反事实任务分解框架。它通过反事实奖励量化分解对检索排名的边际因果增益,切断虚假关联源头,还引入偏好奖励对逻辑连贯性和原子性进行细粒度结构监督,鼓励模型生成高质量分解。此外,我们构建了MTDTool,专为移动多轮交互设计的任务分解基准。大量实验表明,PCTD减轻了重复分解,在检索、分解质量和域外泛化方面超越了现有最优方法。
英文摘要
Task decomposition aims to transform ambiguous instructions into executable atomic subtasks, thereby guiding high-precision tool retrieval. However, our analysis reveals that directly adopting tool retrieval metrics, i.e., Recall or NDCG, as rewards for task decomposition can easily induce reward hacking in reinforcement learning-based methods. Specifically, models tend to maximize retrieval matching through strategies such as repetitive decomposition. This spurious correlation between the shallow features of decomposition results and retrieval metric impairs generalization in Out-of-Domain (OOD) scenarios involving unseen tools. To address this issue, we propose PCTD, a Preference-guided Counterfactual Task Decomposition framework. PCTD quantifies the marginal causal gain of decomposition on retrieval ranking through a counterfactual reward, thereby cutting off spurious correlations at their source. Meanwhile, it introduces a preference reward to impose fine-grained structural supervision on logical coherence and atomicity, encouraging the model to generate high-quality decompositions. In addition, we construct MTDTool, the task decomposition benchmark specifically designed for mobile multi-turn interactions. Extensive experiments demonstrate that PCTD alleviates repetitive decomposition and surpasses SOTA methods in retrieval, decomposition quality, and OOD generalization.