FinalityBench:延迟和冲突财务终局性下智能体决策的效果级基准
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
浏览论文内容
中文总结 AI 辅助
该研究提出FinalityBench基准,用于测试智能体在延迟和冲突财务终局性下的决策,发现特定门控策略准确率达85.4%,语言模型可自动发现该策略。
中文摘要 AI 辅助
商户的支付处理器、分类账、ERP(企业资源规划)和银行信息流会因消息延迟、重复、丢失和乱序而更新,导致数分钟内这四个系统对同一订单持有相互矛盾的认知。解决该异常的智能体必须决定是否发货、重新提交捕获请求、退款或等待,且知晓其中部分操作无法撤销。我们提出FinalityBench,一个用于该决策的可执行基准。它维护一个隐藏的规范事件日志,并从单独存在故障的交付流中推导每个系统的视图,因此分歧源于指定的故障语义而非人为设置。评分基于已执行的货币效果:一个回合的得分取决于商户的最终经济状况,该状况相对于在待处理捕获请求解决时告知的特权参考值计算。包含321个任务的语料库包括45个成对任务(共90个任务):这些任务在决策瞬间四个系统的视图完全相同,权威探测均返回未知,且最终正确处置方式不同。这种快照不可区分性在每个评估种子下都经过检查,而非假设;不声称所有交互轨迹都等价。在对9种程序化策略的14445个已评分回合中,按单任务准确率和成对损失的排名有7处不同:“首次信号即发货”策略以65.7%的准确率位列第二,但在成对损失中最差,因为它无法区分成对任务的两个成员。对不可逆操作设置权威终局性探测的运行时门控策略达到85.4%的准确率,且与所有轮询策略不同,它不会因pass^5损失任何内容;其剩余损失几乎完全属于直接对终局性信息定价的原型。语言模型在分层子集上达到与手写门控策略相同的准确率,但损失约两倍的资金,且无需被告知即可发现终局性门控策略。
英文摘要
A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated, dropped and reordered, so for minutes at a time the four hold contradictory beliefs about the same order. An agent resolving the exception must decide whether to ship goods, re-submit a capture, refund or wait, knowing some of those cannot be undone. We present FinalityBench, an executable benchmark for that decision. It keeps a hidden canonical event log and derives each system's view from a separately faulted delivery stream, so disagreement follows from specified fault semantics rather than being authored. Grading is on executed monetary effects: an episode is scored by the merchant's terminal economic position, relative to a privileged reference told when the pending capture resolves. The corpus of 321 tasks includes 45 twin pairs (90 tasks): tasks whose four system views are identical at the decision instant, whose authoritative probes both return unknown, and whose eventual correct dispositions differ. That snapshot indistinguishability is checked under every evaluation seed rather than assumed; equivalence over all interaction traces is not claimed. Over 14,445 graded episodes from nine programmatic policies, ranking by single-task accuracy and by paired loss disagree in 7 places: a ship-on-first-sign policy is second-best by accuracy at 65.7% and worst in the suite by paired loss, because it cannot tell the two members apart. A runtime gating irreversible actions on an authoritative finality probe reaches 85.4% and, unlike every polling policy, loses nothing to pass^5; its residual loss is almost entirely one archetype, which prices finality information directly. Language models reach the same exact rate as the hand-written gate on a stratified subset, lose about twice as much money, and discover the finality-gating strategy without being told it.