发表机构
University of Illinois at Chicago; Springbrand; Northwestern University; Rutgers University; Arizona State University; University of California San Diego; Carnegie Mellon University; MBZUAI; Microsoft AI(伊利诺伊大学芝加哥分校; 斯普林布兰德公司; 西北大学; 罗格斯大学; 亚利桑那州立大学; 加州大学圣迭戈分校; 卡内基梅隆大学; Mohamed bin Zayed 人工智能大学; 微软人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出智能体商业世界(ACWorld)环境,通过氛围式商业协议(VCP)实现可审计可验证的氛围式商业智能体评估,构建含两类轨道的基准并验证模型性能,分析得出过程级证据对评估智能体的必要性。
AI 中文摘要
在氛围式编程中,人们用自然语言描述软件并将实现任务委托给AI智能体。类似地,氛围式商业允许人们用自然语言表达买卖目标,并将相应任务委托给智能体。然而,商业活动需要相互独立控制的买方智能体和卖方智能体在共享市场中交互,同时保留各自的私人目标和不同的权限。我们提出智能体商业世界(Agentic Commerce World, ACWorld),这是一种用于评估此类智能体在持续交易中表现的环境。ACWorld通过其氛围式商业协议(Vibe Commerce Protocol, VCP)在更新共享交易状态前验证智能体的动作,并记录由此产生的交互,使智能体行为可审计、评估可复现。ACWorld基准包含200个任务的能力覆盖轨道和60个任务的大型目录轨道,该轨道可搜索785022个可交易列表。在10个模型中,平均得分分别为65.9%至85.6%和56.1%至91.4%。我们的分析表明,过程级证据是必要的:仅最终状态会遗漏评估错误,不完整的轨迹仍保留有用的过程信号,大型目录任务会暴露各阶段的瓶颈。
英文摘要
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.