arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

商业竞技场:在现实市场中对大语言模型智能体进行基准测试

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing

arXiv 2608.08621首次发表:更新:

发表机构

Alibaba Group; Yale University(阿里巴巴集团; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员推出Business Arena测试平台,评估15个前沿LLM智能体在跨境商店运营中的表现,发现其平均最终净资产差9倍,最优模型仍逊于人类策略,为商业智能体评估提供了现实测试基准。

AI 中文摘要

经营企业是一项具有挑战性的智能工作,经营者必须从部分信号中推断机会、在不确定性下投入资金、适应不断变化市场中的延迟结果,且需在合法交易前满足监管义务。前沿大语言模型(LLM)智能体越来越能完成复杂工作流程,但现有智能体基准很少评估其与商业相关的能力。我们推出Business Arena,这是一个受控环境,AI智能体可在其中运营一家跨境商店,长期从供应商处采购并向买家销售。我们将该竞技场建立在真实的供应链数据和权威来源校准的市场条件之上。延迟且相互关联的后果使得单个商业决策难以评判,但其综合结果可通过利润衡量。由于仅利润无法解释智能体成败的原因,我们将智能体与人类设计的策略进行比较,以估计可用机会,使用技能水平指标揭示潜在的优缺点,并将已实现的损益追溯到产生该结果的行动。我们通过机制消融法确定,良好的结果反映的是真正的商业智能,而非疏忽或模拟器特有的捷径。我们评估了15个前沿模型,发现其平均最终净资产存在9倍差异。即使是最优模型也落后于人类设计的策略,表明商业运营对LLM智能体而言仍具挑战性。技能水平分析揭示了不同运营风格,包括聚焦利润率的优质卖家、高周转率批发商和客户服务专家;而行动层面归因则识别出创造或破坏价值的采购、定价及补货决策。总体而言,Business Arena为评估端到端商业智能体迈出了构建现实且可信测试平台的第一步。

英文摘要

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑