arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StoreBench:用于评估和训练自主运营智能体的直播电商环境

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

Daksh Raghuvanshi, Ved Vedere, Yifan Wang

arXiv 2610.10942首次发表:更新:

AI 中文总结

研究推出StoreBench直播电商环境,评估7个前沿LLM在电商运营任务的表现,发现人类专家优于所有模型,Qwen3.5-27B经GRPO训练后性能显著提升,暂不公开完整环境以保基准纯洁性。

AI 中文摘要

强化学习环境如今是提升大语言模型(LLM)后训练能力的核心手段,但多数智能体基准仍为静态:仅当智能体行动时环境才变化,奖励为终端判定,且合格线设定随意。我们推出StoreBench,这是一个直播电商环境,智能体可在生产级电商后端运营一家中型在线服装店铺,用于测试不确定性下的长周期规划与经济判断能力。客户全天候下单,供应商会调整价格或出现故障,市场冲击会部分或无预警出现。智能体通过人类运营者使用的29种商家工具执行操作,且受窗口化运营预算限制,使得模拟时间与所采取的行动挂钩,因此模型延迟不会影响模拟时间。合格阈值是针对脚本锚定策略校准的,奖励机制可抵御一系列奖励作弊行为,且给定行动序列时每个回合的回放完全一致。我们在11个时长30至45天的场景及一个完整模拟年度中,在匹配推理算力的三个世界种子下评估了7个前沿LLM。没有模型平均表现能达到脚本智能分类策略的水平:表现最佳的DeepSeek-V4-Pro在任务-种子单元上的通过率为49%,而启发式策略的通过率为97%。使用相同工具和预算的人类专家得分超过所有模型(平均综合得分0.708,对比模型的0.700)。在Claude Code框架下的完整模拟年度中,多数模型表现出显著性能提升。在一次GRPO后训练运行中,仅在5个不相交任务上训练的Qwen3.5-27B,在保留的评估任务上的平均综合得分从0.136提升至0.373。我们发布了5个训练拆分示例任务、10条样本轨迹以及评分和验证工具;完整环境和评估套件暂不公开,以保持基准的纯洁性。

英文摘要

Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.

Comments22 pages, 4 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑