SINGED:正确输出不能证明LLM智能体的安全执行
SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents
浏览论文内容
中文总结 AI 辅助
SINGED基准揭示LLM智能体可能选择功能伪造的第三方工件,正确输出不保证安全执行,评估需关联输出与执行路径。
中文摘要 AI 辅助
使用工具的语言模型智能体会选择并执行第三方工件。不同的实现可以在返回所请求输出的同时产生隐藏的执行效果,而基于任务、攻击或选择的评估可能会遗漏这些效果。我们研究了功能性伪造:这些实现与良性替代品在请求的输出上匹配,但添加了任务契约所禁止的效果。我们引入了SINGED(LLM智能体执行决策中的源完整性与不可识别性差距),一个受控基准,涵盖五个主要任务族和两个保留任务族。它变化了显示排名、证据深度、决策策略、模型发布和智能体配置,同时任务和过程预言机验证工件和执行路径。在7,549次审计试验中,随机排名研究发现排名第一的试验中有45%(27/60)存在伪造执行,而在后续排名中则没有。跨候选比较消除了浅层失败,并将分层失败从15.7%降至4.2%,但留下了依赖失败;其益处对于未见效果和公共包结构不确定。此外,当良性替代品可用时,七个没有伪造执行的版本在移除替代品后,在55/175个单源单元中执行了伪造。因此,SINGED揭示了排名、证据和选择敏感的结果到执行的差距:评估必须将正确输出与执行路径联系起来。
英文摘要
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, while task and process oracles verify the artifact and execution path. Across 7,549 audited trials, the randomized-rank study finds counterfeit execution in 45% (27/60) of rank-one trials and none at later ranks. Cross-candidate comparison eliminates shallow failures and reduces layered failures from 15.7% to 4.2%, but leaves dependency failures; its benefit is uncertain on unseen effects and public-package structures. Moreover, seven releases with no counterfeit executions when benign alternatives are available execute the counterfeit in 55/175 single-source cells after alternatives are removed. SINGED thus exposes a rank-, evidence-, and choice-sensitive outcome-to-execution gap: evaluation must connect correct outputs to execution paths.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Hangzhou Dianzi University(杭州电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。