arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RideWay:面向使用工具的语言智能体的高效任务完成基准

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

Qingnuan Han, Boli Fang, Mingzhi Hou, Claire Liu

arXiv 2609.17985首次发表:更新:

发表机构

Didi Global(滴滴全球)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RideWay提出以效率为中心的网约车智能体基准及成功门控指标,量化多余轮次与工具调用的惩罚,实现任务成功之外的交互效率评估。

AI 中文摘要

AI智能体通常根据其是否完成任务来评估。在交互式服务场景中,即使智能体成功完成任务,也可能因反复提问、执行冗余搜索或进行可避免的修改而使用户感到沮丧。我们引入了RideWay,一个在有状态工具调用环境中针对网约车智能体的以效率为中心的基准,以及效率效用(Efficiency Utility),一种成功门控的指标,该指标根据任务特定的参考工作量,对成功轨迹中多余的工具调用和面向用户的轮次进行折扣。人类成对偏好校准了相对惩罚,反映了聚合的服务工作流权衡:额外的对话通常会产生明显的摩擦,而额外的工具使用有时可以验证约束或保留用户意图。在58个任务和24个模型中,拟合的多余轮次惩罚约为多余工具调用惩罚的两倍。在任务不相交的保留偏好上,效率效用总体准确率达到78.7%:当轨迹在轮次上不同时准确率为90.6%,但当轨迹仅在工具调用上不同时准确率仅为随机水平——这是人类标注者最不一致的轴。因此,RideWay使交互效率在任务成功之外变得可衡量,同时揭示了基于计数的工具使用评估的边界。

英文摘要

AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired preferences calibrate the relative penalties, reflecting an aggregate service-workflow trade-off: extra dialogue often creates visible friction, whereas extra tool use can sometimes verify constraints or preserve user intent. Across 58 tasks and 24 models, the fitted penalty for excess turns is about twice that for excess tool calls. On task-disjoint held-out preferences, Efficiency Utility achieves 78.7% accuracy overall: 90.6% when trajectories differ in turns, but chance-level accuracy when they differ solely in tool calls - the axis on which human annotators agree least. RideWay therefore makes interaction efficiency measurable alongside task success, while exposing the boundary of count-based tool-use evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑