发表机构
Ant International(蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对支付领域大语言模型评估,提出BENCHCOMPASS基准,通过证据包构建场景任务并引入攻击变体,揭示模型不同失败模式,且基准未饱和,最佳模型得分89.6%和81.7%。
AI 中文摘要
支付操作是关键的金融基础设施,但大语言模型在该领域的价值仍不明确,因为支付规则变化迅速、证据分散,且决策依赖于交易状态、参与者角色、地区和支付通道。现有基准无法区分失败是源于缺失支付规则知识、对提供证据的使用不当,还是在输入不完善的情况下表现脆弱。我们提出了BENCHCOMPASS,一个支付领域基准,其构建流程从类型化证据包生成基于场景的任务,应用基于LLM的质量检查,创建任务输入攻击变体,并保留最终项目准入权给领域专家。该发布包含一个经专家评审的Pro基准,涵盖支付知识、基于上下文的场景推理和受攻击的开放性鲁棒性,以及一个低保证级别的Normal池,用于检查和未来策划。在16个模型变体中,BENCHCOMPASS显示出定性不同的失败模式:参数化支付知识缺失、对提供的规则推理不完整,以及未能拒绝看似合理但无效的工作流。该基准尚未饱和:最佳前沿模型在开放上下文推理上达到89.6%,在受攻击输入下达到81.7%,而一个代表性的32B开放权重模型分别达到69.8%和42.6%。基准数据和代码可在该https URL获取。
英文摘要
Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evidence is fragmented, and decisions depend on transaction state, participant role, region, and payment rail. Existing benchmarks do not isolate whether failures come from missing payment-rule knowledge, poor use of supplied evidence, or brittleness under imperfect harness inputs. We introduce BENCHCOMPASS, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts. The release contains an expert-reviewed Pro benchmark covering payment knowledge, context-grounded scenario reasoning, and Attacked Open robustness, plus a lower-assurance Normal pool for inspection and future curation. Across 16 model variants, BENCHCOMPASS shows qualitatively different failure modes: missing parametric payment knowledge, incomplete reasoning over supplied rules, and failure to reject plausible but invalid workflows. The benchmark remains unsaturated: the best frontier model reaches 89.6% on Open Context-Grounded Reasoning and 81.7% under attacked inputs, while a representative 32B open-weight model reaches 69.8% and 42.6%. Benchmark data and code are available at https://github.com/ant-intl/BenchCompass.
Comments19 pages, 5 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026