APEX-Accounting:评估前沿模型会计任务处理能力的基准
APEX-Accounting
浏览论文内容
中文总结 AI 辅助
该研究推出会计基准APEX-Accounting,评估前沿模型的会计任务处理能力,发现模型表现存在差距,还观察到辛普森悖论,可应要求对模型开展排行榜评估。
中文摘要 AI 辅助
我们推出由Mercor与Ramp合作构建的基准APEX-Accounting,用于评估前沿模型是否能完成会计师的实际工作,任务包括账目核对、应计费用、过账交易和生成报告。私有评估集包含160个任务,分为10个领域,每个领域包含会计系统以及电子表格、PDF和其他文件。所有任务均由会计和簿记专家编写并解决,他们还制定了评分标准。在9个前沿模型中,Claude-Fable-5(Max)以56.4%的Mean Criteria@3领先,Muse-Spark-1.1(xHigh)以52.6%紧随其后。没有模型的Pass^8得分超过2.6%(GPT-5.6-Sol (Max+Pro)),最高的Pass@8得分为21.5%(Muse-Spark-1.1(xHigh))。我们将token预算从1美元提高到50美元时,观察到辛普森悖论的一种情况:随着token预算增加,得分上升,但在给定的受预算约束的测试框架内,模型在花费更多token的任务上得分更低。由于APEX-Accounting是封闭基准,可应要求对任何前沿模型运行排行榜评估。
英文摘要
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.
发表机构
- Mercor
- Ramp
机构由 AI 辅助整理,请以论文原文为准。