AI 中文总结
针对自主智能体微支付的买方决策问题,提出协议无关的402Pilot决策层,结合PA-DCT策略在402Pilot-Bench基准中验证其自适应采购的有效性,为可编程支付补充买方决策能力。
AI 中文摘要
x402等可编程支付协议支持按请求进行微支付,但未明确自主智能体在钱包资金有限时应选择购买哪项付费服务。我们将这一买方侧问题定义为智能体原生支付决策:在钱包压力下的上下文提供商选择、仅基于已完成支付的反馈,以及动态变化的市场条件。我们提出402Pilot,这是一种协议无关的买方侧决策层,部署于自主智能体与支付执行之间,用于实现选择付费提供商的采购策略。我们用PA-DCT实例化该决策层,这是一种感知支付的折扣上下文汤普森采样策略,可在钱包压力下调整采购决策,同时从支付后反馈中学习。为评估买方侧支付策略,我们引入402Pilot-Bench,这是一个冻结回放基准,涵盖823项任务、5种异构提供商流水线和3种市场机制,每种机制均基于30对随机种子进行评估。PA-DCT在非先知策略中实现了最强的固定钱包自适应权衡:它在保持有竞争力服务质量的同时,仅消耗钱包的39%至43%,并随市场条件变化重新分配支出。在价格冲击场景下,它取得了最佳非先知PA-gap/T值;在质量、投资回报率(ROI)和PA-gap/T的9种场景-指标组合中,它还取得了最佳均值和最坏情况排名。与学习基线及组件消融的比较进一步验证了所提出决策策略的有效性与设计。这些结果表明,可编程支付必须辅以买方侧决策能力,该能力需能学习服务价值并相应调整采购决策。
英文摘要
Programmable-payment protocols such as x402 enable per-request micropayments, but they do not determine which payable service an autonomous agent should buy under a finite wallet. We formulate this buyer-side problem as agent-native payment decision-making: contextual provider selection under wallet pressure, chosen-only paid feedback, and changing market conditions. We propose 402Pilot, a protocol-agnostic buyer-side decision layer between autonomous agents and payment execution that implements purchasing policies for selecting among payable providers. We instantiate it with PA-DCT, a payment-aware discounted contextual Thompson-sampling policy that adapts purchasing decisions under wallet pressure while learning from post-payment feedback. To evaluate buyer-side payment policies, we introduce 402Pilot-Bench, a frozen-replay benchmark spanning 823 tasks, five heterogeneous provider pipelines, and three market regimes, each evaluated over 30 paired seeds. PA-DCT achieves the strongest fixed-wallet adaptive trade-off among non-oracle policies: it maintains competitive service quality while spending only 39 to 43 percent of the wallet and reallocates spending as market conditions change. It attains the best non-oracle PA-gap/T under the price shock and the best mean and worst-case ranks across the nine scenario-metric combinations of quality, ROI, and PA-gap/T. Comparisons with learning baselines and component ablations further support the effectiveness and design of the proposed decision policy. These results suggest that programmable payment must be complemented by buyer-side decision-making capable of learning service value and adapting purchasing decisions accordingly.