AI 中文总结
研究金融决策中智能体行为稳定性,引入DFAH-Bench基准,通过多渠道衡量。发现仅结果一致性不能完全反映稳定性,前沿模型存在决策与工具路径一致性差距,识别出三种行为模式,并开源相关代码和数据。
AI 中文摘要
标准评估基准衡量工具使用智能体的决策结果,而非其每次决策过程是否相同。我们引入DFAH-Bench,这是一个重放基准,通过工具调用轨迹、证据接触和决策集中度三个渠道衡量金融智能体决策中可观测的行为不稳定性,且无需访问隐藏推理文本。在涵盖10个模型和3个金融任务的8127个重放情节中,我们发现仅结果一致性是不完整的稳定性信号。前沿模型决策一致性可达95%,但工具路径一致性仅77%,仅基于结果评估会完全忽略这18个百分点的差距。在决策一致性高的前沿模型案例组中,超55%存在显著轨迹差异。我们识别出三种行为模式:模式匹配器、稳定执行者和轨迹发散者。随附存储库中发布了基准代码、度量脚本、重放日志等。
英文摘要
A financial agent can repeat a decision while changing the work behind it. DFAH-Bench operationalizes the Determinism--Faithfulness Assurance Harness (DFAH), pairing decision agreement with tool-path agreement on the same qualified replays, then extends that qualification principle to evidence, authorization, execution and task outcomes. Retrospective and prospective replay analyses expose process variation behind stable decisions. Across 570 eligible prospective episodes, decision agreement is 94.2-95.1%, while agreement on ordered tools, arguments and results is 45.0-51.5%; one stratum falls one group below its prespecified coverage minimum. A separate capture diagnostic shows that systematic omissions can preserve perfect replay agreement. Using the $τ$-Knowledge banking environment, we retain 1,080 scheduled episodes and 1,033 known native outcomes across separate cohorts with open-weight and frontier generators. Missing outcomes prevented the planned tests, so comparisons are descriptive. On the primary schedule, structural checks alone yield more successes than either gate-and-recovery bundle. The typed-choice bundle has lower mean episode cost than the generative bundle on complete task pairs, but produces fewer successes under every assignment of unknown outcomes. Input limits and recovery behavior materially shape these results. Fixed-state probes reveal higher decision agreement alongside lower agreement with constructed policy labels, and separately expose sensitivity to retained generator rationale in a selected authorization case. Together, the findings connect replay observability to evidence, authorization, completion and cost: evidence sufficiency needs direct assessment alongside repeatability.
Comments25 pages, 8 figures. Expanded version with interactive banking experiments, fixed-state gate probes, and cost analysis. Code and public artifacts: https://github.com/ibm-client-engineering/output-drift-financial-llms