arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12654cs.AIcs.CLcs.LG

SteerBench-Work:动作边界处智能体决策的基准测试

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

Oguz Serdar, Cuneyt Mertayak

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出职场智能体动作边界决策基准SteerBench-Work,发现模型在该任务中存在过度拒绝的显著偏差,且通用能力与引导校准并不等同。

中文摘要 AI 辅助

长期运行的大语言模型(LLM)智能体通过工具执行操作,单一步骤可完成发送邮件、合并拉取请求或转账等任务。决策引导是在该边界处的提交前选择:继续执行,还是暂停以等待人工或政策审查。我们推出SteerBench-Work,这是一个以事件为锚点的双向基准测试,用于评估开发者运营、客户服务、金融、法律、医疗、人力资源及安全领域职场智能体的此类决策。v2026-05版本包含106个锚定公开事件的场景、配对的证据反转镜像场景,以及校准控制项,标签在“继续”与“暂停”间近乎均匀划分,使两类错误方向获得几乎相同的测试机会。模型会接收拟执行的操作及可用证据,返回门控决策,并根据其是否正确跨越或守住边界进行评分。在30种模型条件下,错误几乎全部集中在一个方向:模型在28.1%的机会中错误地暂停已获授权、证据已清除的工作,在1.0%的机会中错误地允许不安全工作。最困难的案例是风险已解决的提交,其中已签名或结构化证据已清除真实风险触发因素,模型在著名事件的证据反转镜像场景上的得分(63.8%)明显低于在事件本身的得分(98.5%)。通用能力与引导校准并非同一概念:更高能力的模型常在提交边界过度拒绝,更多推理能力可修复薄弱的门控决策,同时保持校准良好的门控决策稳定。公开排行榜位于this http URL。

英文摘要

Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an incident-anchored, bidirectional benchmark for that decision in workplace agents across developer operations, customer service, finance, legal, medical, HR, and security. Release v2026-05 contains 106 scenarios anchored in public incidents, paired evidence-reversed mirrors, and calibration controls, with labels split nearly evenly between proceed and hold so the two error directions get near-identical numbers of chances. A model sees the proposed action and the available evidence, returns a gate decision, and is scored on whether it crosses or holds the boundary correctly. Across 30 model conditions the failures run almost entirely in one direction: models wrongly hold authorized, evidence-cleared work on 28.1% of opportunities and wrongly allow unsafe work on 1.0%. The hardest cases are risk-resolved commits, where signed or structured evidence has already cleared a real risk trigger, and models score markedly worse on evidence-reversed mirrors of famous incidents (63.8%) than on the incidents themselves (98.5%). General capability is not the same as steering calibration: higher-capability models often over-refuse at the commit boundary, and more reasoning can repair a weak gate while leaving a calibrated one flat. The public leaderboard is at steerbench.com.

发表机构

  • AgentDock

机构由 AI 辅助整理,请以论文原文为准。

↑