arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HANDBOOK.md:长上下文代理指令跟随基准测试

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen

arXiv 2607.25398首次发表:更新:

发表机构

Surge AI(Surge AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出HANDBOOK.md基准测试,模拟企业员工遵循公司手册,含65个代理任务跨越多领域多公司。任务修改基础手册防记忆,评分具确定性。测试发现模型配置通过率低,代理常违背策略,还发布了所有任务、环境及评估工具。

AI 中文摘要

语言模型代理越来越多地在理解指令的情况下部署:系统提示、策略文件或技能文档被置于上下文中,代理被信任遵循其后续的每一个行动。现有基准测试很少直接测试这种部署模式;它们衡量代理是否能完成任务,而非长且有约束力的策略文档是否能在扩展的工具使用范围内实际约束其行为。我们提出了这个http URL,一个基于企业员工遵循公司手册方式构建的包含65个代理任务的基准测试。每个任务将代理置于一个独立的公司环境中,该环境有文件工作区以及通过模型上下文协议暴露的模拟邮件、聊天、日历、问题跟踪和商业服务,并指示它按照一份20至124页的专家编写的标准操作程序进行日常专业工作。任务跨越五个领域(金融、医疗计费、保险、物流和人力资源)以及十个虚构公司。为防止记忆,每个任务修改十个基础手册之一,改变评分所依据的具体规则和阈值,所以没有两个任务共享相同策略。评分完全是确定性的:每个任务都有一套编程标准(共824条)的评分细则,检查所需行动是否发生以及禁止行动是否未发生。在严格评分下,即只有每个标准都满足试验才通过,三十个评估模型配置中最好的通过了36.2%的试验,大多数前沿配置仍低于25%。失败呈现出一致模式:代理让环境中看似合理的请求凌驾于既定策略之上,执行所需检查后却违背结果行动,在长时间跨度中丢失规则细节,并报告未实现的合规情况。我们发布了所有任务、环境和评估工具。

英文摘要

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constrains its behavior over an extended tool-use horizon. We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks. Each task places an agent in a self-contained company environment (a file workspace with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol) and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20-124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and 10 fictional companies. To resist memorization, every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the strongest evaluated model passes 36.2% of trials, and most frontier models remain below 25%. Failures follow consistent patterns: agents let a plausible but unauthorized in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release the tasks, environments, and evaluation harness.

Comments16 pages, 3 figures, 5 tables. Accepted to the Workshop on Agent Behavior (WAB) at COLM 2026. Benchmark, environments, and evaluation harness: https://github.com/surge-ai/handbook

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑