AI 中文总结
研究评估大语言模型智能体遵循逻辑关系情况,引入SkillLogic框架,构建SLBench基准,扫描超500个公共技能,评估发现不安全率高,人工审核分析失败原因,还表明SLGuard可减少违规,确立遵循逻辑关系是技能引导智能体的可靠性挑战。
AI 中文摘要
智能体技能通过可重复使用的程序、工具和特定领域工作流程扩展了大语言模型智能体,但其安全性取决于解决交互指令之间的依赖关系。我们引入了SkillLogic,这是一个用于分析技能文件中的逻辑关系并从中构建可执行测试的框架。我们的分类法涵盖八种关系类型。使用SkillLogic,我们扫描了5000多个公共技能,发现70%至少包含一种逻辑关系。然后我们构建了SLBench。评估结果显示不安全率高达70%。人工审核将失败归因于智能体能力差距和技能文本显著性低。我们还表明,SLGuard可将目标案例中的违规行为减少63%。我们的结果将遵循逻辑关系确立为技能引导智能体一项独特的可靠性挑战。
英文摘要
Agent skills extend LLM agents with reusable procedures, tools, and domain-specific workflows, but their safety depends on resolving dependencies among interacting instructions. We introduce SkillLogic, a framework for analyzing logical relations in skill files and constructing executable tests from them. Our taxonomy covers eight relation types, including preconditions that gate valid actions, constraints that limit how allowed actions may be performed, and fallbacks that specify recovery behavior after failure. Using SkillLogic, we scan over 5000 public skills and find that 70% contain at least one logical relation. We then construct SLBench, an 86-case executable benchmark from high-confidence, high-impact, and locally testable relations. Evaluating Codex and Claude Code across six LLM backbones shows unsafe rates up to 70%, with violations leading to privacy leaks, unsafe configuration changes, and incomplete cleanup. The human audit attributes failures to both agent capability gaps and low-salience skill text. We further show that SLGuard, a lightweight inference-time scaffold, reduces violations by 63% on targeted cases. Our results establish logical-relation following as a distinct reliability challenge for skill-guided agents.