概率契约:跨LLM接口的准确性、一致性与决策
Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces
浏览论文内容
中文总结 AI 辅助
提出概率契约基准,评估四种LLM接口配置在1000个世界上的准确性、一致性和决策损失,发现接口间分歧显著,平均化不保证决策改进。
中文摘要 AI 辅助
用于决策的概率应在等效请求中指向同一事件。我们引入了概率契约,这是一个连接精确有限世界后验、验证事件转换和故障感知决策评估的基准。在1,000个世界上评估了四种模型-接口配置。它们的评估在准确性、一致性和决策损失方面存在差异:Kev的总体典型后验误差低于Jev,但补集和粗化残差更大,准确性排序因层而异。在延迟成本为0.10时,Jev的Event和Choice接口在32.8%的有效对上诱导不同的二元动作。事后分析发现,分歧仅证明了各配置中平均二元对误差的11-52%。一个基本的动作区域特征解释了何时平均会相对于随机选择一个接口改变决策损失。尽管平均不能恶化该基线的预期Brier分数,但其决策效果取决于成本和交叉边界;观察到的同基线惩罚发生在已经比始终弃权(不执行)更差的配置中。该基准使这些区别可测量,而不将一致性视为准确性或将分数改进视为决策保证。
英文摘要
Equivalent probability requests can lead to different decisions even when both reports are valid. We introduce probability contracts, a benchmark that connects exact finite-world posteriors, validated event alignment, interface coherence, and failure-aware decision evaluation. Across four model-interface configurations on 1,000 worlds, Kev has lower aggregate canonical posterior error than Jev but larger complement and coarsening residuals; accuracy ordering varies by stratum. Jev's Event and Choice interfaces change the binary action on 32.8% of valid pairs at defer cost 0.10. Post-hoc analyses show that disagreement certifies only 11-52% of mean binary pair error and does not consistently outperform confidence for selection. An action-region characterization and a standard scoring-rule identity explain averaging's expected Brier guarantee relative to random interface selection, but not a decision-loss guarantee at each cost. The loss contrast takes both signs on a 99-cost grid for every configuration; small penalties where both policies beat deferral have pointwise intervals containing zero. Secondary checks specified before collection include a separate 400-root cohort, where Event/Choice effects remain configuration-dependent. A joint surface-order and answer-ID intervention shifts posttrained probabilities. Both bounded reasoning arms yield no valid probability reports, leaving their probability accuracy undefined. Probability contracts make these distinctions measurable by evaluating event semantics, posterior error, coverage, and decision cost together.