arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

电商基准测试集:评估大语言模型智能体在长周期自主商业运营中的表现

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu

arXiv 2608.30730首次发表:更新:

发表机构

Qwen Team, Alibaba Group; Department of Computer Science and Engineering, HKUST; Taobao & Tmall Group, Alibaba Group(通义千问团队,阿里巴巴集团; 香港科技大学计算机科学与工程系; 阿里巴巴集团淘宝天猫集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出首个开源电商长周期自主商业运营基准E-Commerce Bench,评估18个前沿LLM智能体,发现无单一模型占优,GPT-5.6 Sol盈利最高,Qwen3.8-Max-Preview为开源模型最优。

AI 中文摘要

长周期智能体任务不同于多交互回合的短任务串联,其不断演化的动态环境与长程依赖关系要求大语言模型(LLM)在数千步中持续探索、从经验中学习并调整策略。我们推出E-Commerce Bench,这是首个开源基准测试集,将多轮对手谈判与动态事件整合进为期一年的商业运营中。在365天的周期内,LLM智能体同时运营多家线上店铺,调研市场、与供应商谈判以采购库存、优化销售策略、完成订单、处理退货并管理现金流,以最大化年末总资产。为构建真实的商家侧运营环境,产品与供应商数据来自真实电商平台,而包含促销、自然灾害与供应链冲击的全年日程不断重塑需求。为保证可复现性,市场双方均为确定性设定:客户购买与退货遵循固定需求模型,谈判内核决定供应商定价、让步与决策,LLM仅用于将这些决策表述出来。我们在七个维度(含年末资产)评估了18个前沿模型,发现没有单一模型占据主导地位:GPT-5.6 Sol盈利最高,将10万初始本金增长至1431425,但在欺诈规避上排名18个模型中的第16位,运营效率落后于Fable5;在开源模型中,Qwen3.8-Max-Preview以416252领先,较GLM 5.2(高)高38%,且在整个周期中学习能力最强,通过重复订单逐步压低价格。我们的代码可在此URL获取。

英文摘要

Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrently runs multiple online stores, researching the market, negotiating with suppliers to source inventory, optimizing sales strategies, fulfilling orders, handling returns, and managing cash flow to maximize its end-of-year total assets. To construct a realistic merchant-side operating environment, the product and supplier data are derived from a real e-commerce platform, while a year-long calendar of promotions, natural disasters, and supply-chain shocks continually reshapes demand. For reproducibility, both sides of the market are deterministic: customer purchases and returns follow a fixed demand model, while a negotiation kernel determines supplier pricing, concessions, and decisions, with an LLM used only to verbalize them. We evaluate 18 frontier models across seven dimensions, including year-end assets, and find that no single model dominates. GPT-5.6 Sol earns the most, growing the 100,000 opening stake into 1,431,425, yet it ranks 16th of 18 on fraud avoidance and trails Fable5 in operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and achieves the strongest learning over the horizon, progressively bargaining down prices across repeated orders. Our code is available at https://github.com/QwenLM/E-CommerceBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑