发表机构
Huazhong University of Science and Technology; Nanyang Technological University(华中科技大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TRUSS是一种证据引导框架,可生成安全可靠的智能体技能,在多基准测试中实现了漏洞检测100%的精度召回率,同时显著提升任务有效性与安全率。
AI 中文摘要
智能体技能将可复用的自然语言过程与可执行资源封装,使软件智能体无需进行模型适配即可获取特定任务能力。自动生成此类技能可提升任务性能,但仅从其制品或最终任务结果评估候选技能,无法明确配备该技能的智能体将执行哪些动作,以及这些动作会产生哪些副作用。本文提出TRUSS,一种用于生成功能有效且安全可靠的智能体技能的证据引导框架。TRUSS首先对照源证据和领域证据检查功能声明,同时在九项预定义安全属性下评估完整制品。通过该静态关卡的候选技能会被加载到可控执行环境内的影子智能体中,其中中介工具将请求的动作暴露给策略执行模块,并将其结果记录为保留来源信息的执行轨迹。功能故障和属性违规会被关联到对应的技能内容,用于指导迭代优化。我们在168个SkillInject制品、155个SkillSafetyBench案例以及SkillGenBench的全部187个任务上对TRUSS进行评估。TRUSS在漏洞检测中达到100.00%的精度和召回率;修复环节中,使用GPT 5.5时攻击成功率从38.71%降至19.35%,使用GPT 5.4时从46.45%降至29.68%,且无攻击退化现象。在技能生成方面,TRUSS将任务有效性从无技能时的17.11%提升至52.94%,同时将基准安全率从50.80%提升至100.00%。这些结果表明,执行证据可揭示制品检查遗漏的行为故障,并能引导技能生成朝着经联合验证的功能与安全结果方向发展。
英文摘要
Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate solely from its artifact or final task outcome leaves unresolved which actions the equipped agent will perform and which side effects those actions will produce. We present TRUSS, an evidence guided framework for generating functionally effective and safety reliable Agent Skills. TRUSS first inspects functional claims against source and domain evidence while evaluating the complete artifact under nine predefined safety properties. Candidates admitted by this static gate are loaded by a shadow agent inside a Controllable Execution Environment, where brokered tools expose requested actions to policy enforcement and record their results as provenance preserving execution traces. Functional failures and property violations are linked back to the responsible Skill content and used to guide iterative refinement. We evaluate TRUSS on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 tasks in SkillGenBench. TRUSS achieves 100.00\% precision and recall in vulnerability detection. Repair reduces attack success from 38.71\% to 19.35\% with GPT 5.5 and from 46.45\% to 29.68\% with GPT 5.4, with zero attack regression. For Skill generation, TRUSS raises task effectiveness from 17.11\% without Skills to 52.94\%, while increasing the benchmark Security rate from 50.80\% to 100.00\%. These results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.