arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30717cs.IR

RecToolBench:模糊用户意图下的推荐专用工具编排基准

RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent

  • Jilin University(吉林大学)
  • Squirrel AI Learning(松鼠AI学习)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiao Chen, Yicheng Zhao, Yingying Wu, Zhendong Chu, Changyi Ma, Qingsong Wen, Xuan Song

AI总结:

针对现有推荐基准忽视模糊意图和复杂工具编排的问题,提出RecToolBench基准,包含超1200个任务、13个MCP服务器和32个工具,并采用合成-模糊化-评判流水线构建,实验揭示LLM在语义参数接地和证据整合上的瓶颈。

AI中文摘要:

近年来,智能体推荐系统正从被动过滤引擎转变为遵循指令的智能体,利用外部工具来解析用户意图。然而,现有基准通常假设明确的用户意图、简化的工具环境或孤立的函数调用,导致推荐场景下现实工具编排的研究仍不充分。为弥补这一空白,我们提出了RecToolBench,一个基于模型上下文协议(MCP)的基准,用于在模糊用户指令下评估使用工具的推荐智能体。RecToolBench包含超过1,200个可执行任务,涵盖三个推荐领域、13个MCP服务器和32个工具,涉及单工具调用、并行工具调用、顺序工具链和混合工具编排。我们通过可扩展的“合成-模糊化-评判”流水线构建RecToolBench,生成可执行的模糊推荐任务,并采用基于规则的执行检查和基于量规的LLM评估来评估智能体轨迹。在代表性LLM上的实验表明,语法上有效的工具调用并不能保证成功的推荐。模型在语义参数接地、多步证据整合和基于证据的最终推荐方面存在困难,尤其是在编排复杂度增加时。我们的结果将模糊用户意图下的工具编排确定为智能体推荐系统的主要瓶颈。我们的数据和代码可在该https URL获取。

英文摘要:

Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, existing benchmarks often assume explicit user intent, simplified tool environments, or isolated function calls, leaving realistic tool orchestration for recommendation underexplored. To bridge this gap, we propose RecToolBench, a Model Context Protocol (MCP)-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. RecToolBench contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel tool calls, sequential tool chains, and hybrid tool orchestration. We construct RecToolBench with a scalable synthesize--fuzzify--judge pipeline that generates executable fuzzy recommendation tasks, and evaluates agent trajectories using rule-based execution checks and rubric-based LLM evaluation. Experiments on representative LLMs show that syntactically valid tool calls do not guarantee successful recommendations. Models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases. Our results identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems. Our data and code are available at https://github.com/ShawnChenn/RecToolBench.

补充信息

↑