发表机构
HONOR(荣耀)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究智能体工具检索问题,提出MagicSelector联合优化框架,含反事实任务分解、渐进重排和动态Top-K策略,经实验验证,该框架在工具检索准确性、OOD泛化能力和整体令牌效率方面显著优于现有方法。
AI 中文摘要
我们提出了MagicSelector,这是一个集成反事实任务分解、渐进重排和动态Top-K的联合优化框架,旨在解决智能体中工具检索的基本挑战。MagicSelector能够将模糊用户指令转化为可执行原子子任务并指导高精度工具检索。其通过三个关键贡献实现这些能力:偏好引导的反事实任务分解机制、基于自蒸馏硬负样本挖掘的渐进工具重排方法、双语义边界感知动态Top-K策略。在我们构建的MTDTool基准测试上评估,MagicSelector性能良好,显著优于现有方法。
英文摘要
We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.