arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BenchShield:面向LLM智能体评估基础设施中奖励完整性的形式化模型支撑检测层

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He, Xiao Cheng, Jicheng Wang, Wenbo Chen, Alex Yates, Yinzhe Zhao, Bingran You, Yuan Gao, Ayush Munot, Shubham Gaur, Zhe Ye, Hao Wang, Xiangyi Li, Dawn Song, Christophe Hauser

arXiv 2609.11028首次发表:更新:

发表机构

Dartmouth College; Ohio State University; RLWRLD; University of New South Wales; University of California, Davis; Macquarie University; Amazon; BenchFlow; University of Washington; UC Berkeley(达特茅斯学院; 俄亥俄州立大学; RLWRLD; 新南威尔士大学; 加州大学戴维斯分校; 麦考瑞大学; 亚马逊; BenchFlow; 华盛顿大学; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体评估中的奖励黑客问题,BenchShield基于生命周期模型提供静态与运行时双重分析,显著提升检测召回率并降低成本。

AI 中文摘要

LM智能体基准测试日益作为交互式评估基础设施发挥作用。智能体观察状态、调用工具、修改工作空间、提交工件,并从结果流程中接收奖励。这种交互性使评估容易受到奖励黑客攻击:智能体通过利用与奖励相关的轨迹而非解决预期任务来提高其测量得分。现有防御措施主要依赖特定任务的补丁、提示指令或事后检测器,它们无法提供可复用的证据来证明具体运行保持在预期评估边界内。本文提出BenchShield,一个面向LLM智能体评估中奖励完整性的模型支撑检测层。BenchShield将检测基于评估的奖励相关事件的有限生命周期模型。在基准基础设施内,两种互补的分析在此模型上运行。静态的、阶段感知的污点分析在运行前暴露奖励黑客路径;其运行时对应部分利用基础设施侧证据归因具体智能体使用并发出基于证据的声明。我们构建了BenchShield Trajectories,一个包含来自三个基准的超过31,000次公共智能体运行中456条经裁决轨迹的人工标注语料库。与相同任务和模型上的智能体可黑客性扫描器基线相比,BenchShield将全链召回率从23-94%提升至77-100%,同向量覆盖率从16-56%提升至43-78%,并将每任务成本降低高达65%。其运行时分析在从基础设施侧证据检测奖励黑客方面达到96%的准确率。

英文摘要

LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not provide reusable evidence that a concrete run remained within its intended evaluation boundary. This paper presents BenchShield, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation. BenchShield grounds detection in a finite lifecycle model of an evaluation's reward-relevant events. Within the benchmark infrastructure, two complementary analyses operate over this model. A static, phase-aware taint analysis exposes reward-hacking paths before a run. Its runtime counterpart uses infrastructure-side evidence to attribute concrete agent use and emit evidence-backed claims. We construct BenchShield Trajectories, a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks. Compared with an agentic hackability scanner baseline on the same tasks and model, BenchShield improves full-chain recall from 23-94% to 77-100%, same-vector coverage from 16-56% to 43-78%, and reduces per-task cost by up to 65%. Its runtime analysis achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑