arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10042cs.LGcs.AI

UserToolBench:用于工具使用类大语言模型个性化决策的用户画像隐藏型基准测试集

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Xuexiong Yin, Zechuan Chen, Yongsen Zheng, Yuxiang Zhang, Jingyuan Yang, Bin Wang, Yubin Wang, Keze Wang

首次发表
浏览论文内容

中文总结 AI 辅助

UserToolBench是针对工具使用类大语言模型的基准测试集,可评估模型从交互历史推断用户偏好等个性化决策能力,实验发现当前模型在多工具协调等方面仍存瓶颈。

中文摘要 AI 辅助

工具使用类大语言模型越来越多地被要求代表用户执行操作,但现有基准测试通常聚焦于用户画像召回、风格模仿、通用工具使用或响应层面的个性化。我们推出UserToolBench,这是一个针对工具使用类大语言模型个性化决策的基准测试集。UserToolBench测试模型是否能从交互历史中推断潜在用户偏好、识别何时需要澄清,并在信息不完整的情况下生成符合用户偏好的工具调用轨迹。该基准测试集由隐私脱敏的真实交互轨迹构建,结合了结构化角色画像、公开API风格的工具生态系统以及长周期多轮交互轨迹,包含10个用户画像、36套工具集、1065轮交互、170种独特工具,以及涵盖信息缺失、单工具和多工具场景的评估导向任务类型。对强大工具使用类大语言模型的实验表明,当前模型在个性化委托方面仍存在困难,多工具协调、缺失约束推断和长周期行为一致性仍是主要瓶颈。这些结果表明,个性化评估应超越输出是否听起来具有用户针对性,转而关注大语言模型是否为其所代表的用户做出了正确决策。

英文摘要

Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.

发表机构

  • Sun Yat-Sen University(中山大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑