发表机构
College of Computing, Georgia Institute of Technology; Department of Computer Science, University of Colorado Boulder(佐治亚理工学院计算学院; 科罗拉多大学博尔德分校计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在GPT-5.6上开展预先指定的等效性研究,发现调整推理工作量未导致未授权工具使用,且规则探测率上升的模式与针对性搜索假设不符。
AI 中文摘要
通过工具调用执行多步骤工作流的语言模型智能体在访问控制策略下运行,该策略限制每个角色可执行的操作。为了控制成本和延迟,运营者会调整为这些智能体提供服务的API所暴露的推理工作量参数。该参数是否也会改变未授权工具使用的发生率,尚未在单一模型内通过直接操作进行测试。我们在GPT-5.6内部针对TRIO-20的14个验证场景改变了推理工作量(低、最大),TRIO-20是由20个匹配的工作场所三元组组成的套件,其中禁止的工具调用在环境中有效且其对目标指标的影响已说明、有效但仅可通过规则检查发现或无效。这三种条件源自同一代码库,仅在两个配置字段上存在差异,提示和工具集完全相同。所有分析均在验证数据收集前的固定计划中预先指定。在840条轨迹和两个模型层级中,未发生任何未授权工具调用。精确单侧95%置信限将每个组的违规率限制在3.50%以下(Terra,n=84)和5.21%以下(Sol,n=56)。交互作用估计量在Terra上的同时精确95%区间为±4.34个百分点,处于±7.01个百分点的等效性边际内。提高推理工作量确实改变了行为,但仅在规则检查方面:所有条件下的规则探测率均上升,尤其是在探测无工具性回报的情况下,这一模式与针对性搜索的假设不一致(-14.3个百分点,95% CI -27.4至+1.2)。原始轨迹已发布在此httpsURL。
英文摘要
Purpose: Developers choose how much a language-model agent reasons before it acts. Some pick a high level for fear that a low one is not enough, and pay for it in tokens, time and overthinking. We tested whether a low level is enough for routine office work. Methods: Two tiers of GPT-5.6 did 14 routine office tasks at the low and the max reasoning level, 840 runs in all. To make the tasks harder, each one sets a target that the rules make impossible to reach and offers a forbidden tool that would reach it. Some versions also tell the agent that the forbidden tool counts. A program checked every run from its final state. Results: At both levels the agent always followed the rules and never used the forbidden tool, including 30 runs in which it was told that the tool counts while it could still have used it. The low level used about 43% fewer output tokens and 20% less time. Conclusion: For routine office work, a low reasoning level is enough.