arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05204cs.AIcs.LG

SkillTrace:面向LLM智能体技能复用的多溯源审计

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SKILLTRACE框架,通过提取三种溯源审计LLM智能体技能复用,在SKILLTRACE-BENCH上取得高指标,野外审计显示其能生成更具操作性的复用审查队列。

中文摘要 AI 辅助

LLM智能体生态正围绕可复用技能快速发展,这些技能是包含元数据、自然语言指令、代码、工具、参考资料及操作工作流的混合模态包。当技能成为市场制品后,对其复用的审计不再等同于普通代码克隆检测问题。现有检测器针对单模态源代码或整包相似性,但技能复用证据分布在创作文本、实现片段及操作结构中,因此可能遗漏仅保留技能某一部分的复用。本文提出SKILLTRACE,一种面向LLM智能体技能复用的多溯源审计框架。SKILLTRACE提取三种溯源:表达式、实现及操作溯源,将操作溯源表示为技能操作图(SOG),捕捉激活、过程及资源流结构。仅在录入时由LLM辅助一次操作溯源提取;审计时SKILLTRACE确定性比较缓存的溯源,针对相同功能的严格负样本校准各溯源,并报告哪一溯源支持复用决策。在SKILLTRACE-BENCH(含100个市场基准的820个转换复用正例及751个负对照)上,SKILLTRACE达到AUROC 0.938、F1 0.898;对36446个技能的野外审计进一步显示,溯源归因证据能生成超出仓库级基线的可操作复用审查队列。

英文摘要

LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. Existing detectors target single-modality source code or whole-package similarity, yet skill reuse evidence is distributed across authored text, implementation fragments, and operational structure. As a result, they can miss reuse that preserves only one part of a skill. We present SKILLTRACE, a multi-trace provenance auditing framework for LLM-agent skill reuse. SKILLTRACE extracts three provenance traces: Expression, Implementation, and Operational. It represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure. An LLM assists only the Operational-trace extraction, once at ingestion; at audit time SKILLTRACE compares cached traces deterministically, calibrates each trace against same-function strict negatives, and reports which trace supports a reuse decision. On SKILLTRACE-BENCH, with 820 transformed reuse positives over 100 marketplace anchors and 751 negative controls, SKILLTRACE achieves AUROC 0.938 and F1 0.898. A 36,446-skill wild audit further shows that trace-attributed evidence surfaces actionable reuse review queues beyond repository-level baselines.

↑