AI 中文总结
本研究推出MerchantBench电商运营基准测试,通过365天订单级模拟评估大语言模型智能体的中长期一致性,发现其与人类参与者存在显著差距。
AI 中文摘要
大语言模型智能体正日益被评估为自主工具使用者,但大多数基准测试聚焦于有明确成功标准的有限任务。实际部署往往需要中长期一致性,即跨较长时间范围保持有目的的行为,同时根据累积证据调整决策。评估该能力需要一个持久环境,其中行动会约束未来选择、反馈以不同延迟到达,且不一致行为会产生可测量的累积效应。卖家侧电商提供了合适的评估场景,其涉及产品采购、商品上架与定价控制、现金流管理及混合延迟反馈适应等反复且相互依赖的决策。我们推出MerchantBench,这是一个基于98843条真实电商产品记录的365天订单级模拟,配备26个智能体交互工具。MerchantBench将即时可观测的上游供应商事件与延迟的下游订单结果相结合,要求智能体跟踪单个订单生命周期并回顾早期决策。我们在两个智能体框架下对8个大语言模型进行了48次运行评估,每次运行覆盖365个模拟天。结果显示,即使是最新的大语言模型与人类参与者之间也存在显著差距,最佳大语言模型配置仅达到人类参与者平均最终净资产的27.3%。
英文摘要
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.