发表机构
University of the Chinese Academy of Sciences; University of Birmingham(中国科学院大学; 伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SAKIKO框架通过方向性错误发现和目的地解析验证,审计LLM内部干预,证明行为改变不等于修复,需结果解析裁决。
AI 中文摘要
在调用外部工具之前,智能体LLM必须在K路动作空间中进行选择:执行调用、寻求澄清、直接回答或拒绝。虽然内部激活引导可以改变这些执行前的决策,但传统的聚合指标掩盖了改变后的状态落在何处以及它们造成的附带损害。我们提出了SAKIKO,一个审计框架,通过方向性错误发现、路由器条件干预、目的地解析验证和前瞻性冻结统计许可来形式化表示修复。在When2Call和MetaTool上的七个LLM中,通道键控干预在五个模型中诱导了方向特定的净增益;在三个密封评估中,59个预算匹配的随机方向中没有一个匹配校准的目标增益。关键的是,目的地审计表明行为移动不等于修复:一个实现+55净增益的干预破坏了它触及的超过一半的基线正确决策,而Qwen3-4B和Gemma-2-9B上的有希望的估计由于有限样本不确定性而被正式拒绝。SAKIKO确立了在声称内部修复之前进行结果解析裁决的必要性。代码:此HTTPS URL。
英文摘要
Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: https://github.com/ruizheliUOA/mechanistic-tool-use-llm.
CommentsPreprint