arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SKILLLITE:基于证据引导的恶意技能审计与紧凑型大语言模型

SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs

Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam

arXiv 2609.36879首次发表:更新:

发表机构

Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对紧凑型LLM难以审计复杂技能包中隐式恶意行为的问题,提出证据引导的智能体框架SKILLLITE,通过提取安全行为并推断功能上下文实现高效检测,性能优于现有基线且延迟低。

AI 中文摘要

随着基于大语言模型(LLM)的智能体执行日益复杂的任务,Agent Skills(智能体技能)已成为扩展其能力的一种灵活机制。Agent Skill(智能体技能)将任务特定指令与可执行组件和辅助资源打包,以提供专门功能。然而,第三方技能的日益普及引入了新的供应链攻击面。恶意技能可能嵌入有害行为,滥用智能体权限,并危及智能体执行环境或可访问资源。尽管近期基于LLM的恶意技能审计方法取得了有前景的性能,但它们通常依赖能力强大的商业LLM。如何在安全敏感和资源受限的环境中,利用紧凑型、可本地部署的LLM实现有效审计,在很大程度上仍未得到探索。我们的调查揭示,紧凑型LLM难以识别隐藏在复杂技能包中的恶意行为。这一困难源于此类行为的隐式性质以及紧凑型LLM有限的推理能力。为应对这些挑战,我们提出了SKILLLITE,一个用于恶意技能检测的基于证据引导的智能体框架。SKILLLITE有效提取安全相关行为,并从复杂技能包中推断预期功能。然后,它利用紧凑型LLM,基于观察到的行为及其功能上下文评估技能的恶意性。实验表明,SKILLLITE在不同紧凑型LLM骨干网络上提升了恶意技能检测性能,并超越了现有代表性审计基线。其有效性可泛化到行为上已确认的现实世界恶意技能。同时,SKILLLITE保持低推理延迟,支持其实用部署。

英文摘要

As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑