arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

支付宝-PIBench:用于编码代理的现实支付集成基准测试

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu

arXiv 2607.14573首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在评估编码代理在支付宝支付集成任务中的表现,引入支付宝-PIBench基准测试,含九个产品项目和18个任务实例,分不同场景。通过评估六个编码代理模型得出评分通过率,发现有技能条件下RPR提升,为诊断模型能力和评估支付集成指导提供可控环境。

AI 中文摘要

支付集成是一项要求苛刻的存储库级软件任务,代理必须选择合适产品、实现协调的客户端-服务器流程、验证支付结果并保持交易与业务状态的一致性。我们引入了支付宝-PIBench,这是一个用于评估编码代理在现实支付宝支付集成方面表现的基准测试。它包含九个特定产品项目和18个任务实例,分为基本功能完成和高级风险感知强化场景。特定场景的评分标准支持确定性的静态、单元、集成和端到端检查,并辅以大语言模型辅助的语义需求评估。我们评估了六个编码代理模型并报告了评分通过率(RPR)。在具备技能的条件下,平均RPR在68.58%至91.37%之间。相对于无技能条件,获得支付宝支付集成技能平均使平均RPR提高了10.31个百分点,不同模型、产品和场景的提升有所不同。方法级结果区分了源级完成、可执行支付行为和支付领域需求。支付宝-PIBench为诊断模型能力和评估支付集成中的结构化指导提供了一个可控环境。

英文摘要

Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑