反事实工具排序:在效用、成本与特权约束下
Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints
浏览论文内容
中文总结 AI 辅助
本研究提出一种可证伪的反事实工具评估方法,通过实验区分直接回归与双重稳健估计在不同环境下的性能,并揭示模型在固定提示下无法弃权(不执行)的缺陷。
中文摘要 AI 辅助
反事实工具评估必须区分权威性、历史支持以及比较实际估计的内容。我们通过十一个可执行的企业级启发工具、精确倾向性日志和真实的本地模型上下文协议传输来研究这些区别。一项初始的45次运行合成研究被保留,随后受到30次已实现回报控制运行和15项针对1,930个独立发布的伯克利函数调用排行榜(BFCL)任务的实验的挑战。在线性设置中,全回报直接回归逆转了最初有利的双重稳健(DR)评估结果:直接回归的平均绝对误差为0.0139,而DR为0.0272。在偏移环境下,DR保持优势,误差为0.0227对比0.0948。在函数名组不相交的BFCL衍生划分上,直接和DR选择器分别获得81.85%和79.83%的平衡准确率。两个固定的本地Qwen2.5模型在相同的200个保留任务上评估,暴露了在固定提示下严重无法弃权(不执行)的问题。我们进一步刻画了在支持缺失下的策略差异:两个策略共享的不支持动作相互抵消,使得当两个绝对值均不可识别时,可以点识别增量变化。一个保留分歧的回退方案在所有五次支持缺口运行中实现了这一性质,但保守的采样界限不能证明部署改进。贡献在于一种可证伪的评估方法和独立的公共证据,而非新的DR估计器、官方BFCL排行榜分数或生产代理安全声明。
英文摘要
Counterfactual tool evaluation must distinguish authority, historical support, and what a comparison actually estimates. We study these distinctions with eleven executable enterprise-inspired tools, exact-propensity logs, and real local Model Context Protocol transport. An initial 45-run synthetic study is retained, then challenged by 30 realized-return control runs and 15 experiments on 1,930 independently released Berkeley Function Calling Leaderboard (BFCL) tasks. Full-return direct regression reverses an initially favorable doubly robust (DR) evaluation result in the linear setting: mean absolute errors are 0.0139 for direct regression and 0.0272 for DR. Under a shifted environment, DR retains an advantage, with errors 0.0227 versus 0.0948. On function-name-group-disjoint BFCL-derived splits, direct and DR selectors obtain balanced accuracies of 81.85% and 79.83%. Two pinned local Qwen2.5 models are evaluated on the same 200 held-out tasks, exposing a strong failure to abstain under the fixed prompt. We further characterize policy differences under missing support: unsupported actions shared by two policies cancel, allowing point identification of an incremental change when neither absolute value is identifiable. A disagreement-preserving fallback achieves this property in all five support-gap runs, but conservative sampling bounds do not certify deployment improvement. The contribution is a falsifiable evaluation method and independent public evidence, not a new DR estimator, official BFCL leaderboard score, or production-agent safety claim.
发表机构
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。