AI 中文总结
研究大语言模型工具调用中参数的不同风险,提出角色分层共形风险控制方法,为语义参数角色设阈值和预算,实验表明该方法能有效实现特定角色预算合规,证明应在语义角色层面认证结构化工具调用。
AI 中文摘要
语言模型代理通过结构化工具调用发挥作用,其参数具有不同风险。现有统计方法通常控制整个行动的风险,使罕见高风险领域的失败被良性参数掩盖。我们引入角色分层的逐字段共形风险控制,为语义参数角色设置单独阈值和风险预算。通过实验表明该方法能在多种情况下实现最一致的特定角色预算合规,结果表明结构化工具调用应在语义角色层面而非整个行动层面进行认证。
英文摘要
Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shape an email body but should never set a recipient, account, command, or credential. Existing conformal risk control methods certify a tool call as a whole, so a failure in one rare high-risk field can be averaged away by the many benign arguments around it, leaving the argument that causes harm uncertified. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and assigns a separate threshold and risk budget to each semantic argument role. We show that aggregate certification pays a price of coarseness, tightening a rare role's effective budget in proportion to how often that role appears, whereas role-stratified calibration certifies each sufficiently sampled role directly with a finite-sample guarantee and pools the rarest roles. Across AgentDojo and InjecAgent with six language models, our method achieves the most consistent role-specific budget compliance among the methods we evaluate under model and attack transfer, detector noise, gradual drift, unseen tool suites, and adaptive attacks, providing formal per-role guarantees under exchangeability or after recalibration. These results suggest that structured tool calls should be certified at the semantic-role level, not the whole action.