PartHackBench:面向部分信用工具代理评估的认证等进度压力测试
PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出PartHackBench,一种认证等进度压力测试方法,通过私有认证器匹配轨迹组件以消除混淆,在PB-CSTE中验证了评估者信用膨胀的存在,并提供了零膨胀的控制基准。
AI中文摘要:
长视界工具代理常常在未达到最终成功的情况下取得有用的进展,这促使了部分信用评估的出现。然而,评估者可能会奖励那些暂时的、后来被撤销的或并非归因于被评估代理的里程碑。如果对抗性轨迹取得了更多真实的进展,那么将诚实轨迹与得分更高的对抗性轨迹进行比较将是不确定的。我们引入了PartHackBench,一种受控的方法论,以消除这种混淆因素。一个私有认证器仅在两条轨迹在当前状态谓词满足度和标准化代理归因方面逐组件匹配时接受一对轨迹;分数膨胀,定义为f(A) - f(H),仅在之后测量。在PB-CSTE中的18个密封的保留任务中,冻结的历史目标运行产生了15个任务的匹配对抗样本。历史信用产生了平均膨胀.252,条件攻击成功率为10/15,端到端产出为10/18,并且未检测到14个严格回滚中的任何一个。语义LLM评判者更具抵抗力但仍然脆弱,尤其是在针对评估者的攻击下,而PB-CSTE当前状态控制(定义为认证组件的精确函数)通过构造产生了零膨胀。因此,PartHackBench提供了一个认证控制,用于测试当所有基准定义的任务相关进展保持不变时,评估者信用是否发生变化。
英文摘要:
Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match component-wise in both current-state predicate satisfaction and standardized agent attribution; score inflation, defined as f(A) - f(H), is measured only afterward. In 18 sealed held-out tasks in PB-CSTE, the frozen historical-target run produced matched adversaries for 15 tasks. Historical credit yielded mean inflation of .252, conditional attack success of 10/15, end-to-end yield of 10/18, and detected none of 14 strict rollbacks. Semantic LLM judges were more resistant but remained vulnerable, especially under evaluator-targeted attacks, while PB-CSTE current-state controls, defined as exact functions of the certified components, yielded zero inflation by construction. PartHackBench thus provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.