arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估开放权重电商智能体:基于环境锚定的验证方法

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification

Nimit Shah, Haitz Sáez de Ocáriz Borde

arXiv 2609.16093首次发表:更新:

发表机构

AION(AION)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建确定性电商环境,通过环境锚定验证评估开放权重智能体,揭示成功指标掩盖的多种行为缺陷。

AI 中文摘要

一次购物对话通往同一购物车有多条路径,而任务成功率将所有路径简化为单一分数。我们构建了一个确定且可复现的电商环境,该环境预先设定每次试验的顾客与轨迹参数,包括人设、难度、目标购物车以及商品揭示时间表。一个模拟消费者在该环境中,在待评估模型的辅助下尝试购买目标购物车。环境引导模拟器的行为,并记录每次助手动作及当时的环境状态。试验结束后,这些记录使评估者能够依据保留的证据评估对话的各个部分。例如,仅当顾客已提及某目标商品时,评估者才会因搜索未能展示该商品而予以惩罚。我们进一步利用这些证据,根据助手动作与预期工具调用集合的比较,对工具调用施加不同的惩罚。我们的环境还与模拟器双向交互,读取其输出以在模拟器判定顾客已过于沮丧时终止试验,并实时注入指令,指定何时探索、延迟购买某商品或回顾之前的交流。这种交互创造了开放且可验证的模拟。在八个参数规模从20B到35B的开放权重智能体上,每个智能体进行160次试验并记录44项指标,所得能力画像区分了行动不足、过度购买、不支持的属性及搜索不佳等情况,而这些在最终成功中均被掩盖。

英文摘要

A shopping conversation has many routes to the same cart, and a task-success rate reduces all of them to one score. We build a deterministic and reproducible e-commerce environment that precommits each trial's customer and trajectory parameters, including the persona, difficulty, target cart, and an item reveal schedule. A simulated consumer attempts to buy a target cart from the environment with assistance from the evaluated model. The environment guides the simulator's actions and records every assistant action alongside the environment state at that point. After the trial, these records allow the evaluator to assess individual parts of the conversation against the retained evidence. For example, the evaluator penalizes a search for failing to surface a target product only when the customer has already mentioned that product. We further use this evidence to apply different penalties to tool calls depending on how the assistant's actions compare with an expected tool-call set. Our environment also interacts with the simulator bidirectionally, reading its output to stop the trial when the simulator determines that the customer has become too frustrated and injecting directives in real time that specify when to explore, defer buying an item, or recall a previous exchange. This interaction creates an open-ended and verifiable simulation. Across eight open-weight agents from 20B to 35B parameters, with 160 trials per agent and 44 metrics, the resulting capability profiles distinguish under-action, over-purchase, unsupported product attributes, and poor search, all of which terminal success obscures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑