arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10534cs.AIcs.CRcs.LG

智能体技能中的跨层错位检测:一种渐进式加载感知对比学习方法

Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach

  • Indiana University Bloomington(印第安纳大学布卢明顿分校)

机构由 AI 辅助整理,请以论文原文为准。

Chengjun Zhang, Yang Gao, Jianna Hur, Jingjing Zhang, Sagar Samtani

AI总结:

研究智能体技能跨层错位问题,提出渐进式加载感知分层对比学习框架PL-HCL,通过建模技能分层结构和学习跨层一致性来检测错位,使用相关语料库与挑战集,提升宏F1指标,为用户和运营商提供筛选工具与设计原则。

AI中文摘要:

大语言模型(LLM)智能体越来越多地通过智能体技能进行扩展,智能体技能是用于运行时使用的可重复使用工件,它打包自然语言元数据、程序指令和执行时资源。随着开源技能市场的扩大,用户和智能体越来越依赖简短元数据来选择第三方技能,这使得难以检测技能描述与其真实行为之间的不一致,即跨层错位问题。为解决此问题,我们提出渐进式加载感知分层对比学习(PL-HCL),这是一个基于LLM的框架,通过对智能体技能的分层结构进行建模并学习跨层一致性来检测错位。使用超过264,000个开源技能的归一化语料库和人工验证的挑战集,PL-HCL将未调整基线的宏F1从约0.45提高到评估的LLM主干上的0.87-0.89。该方法为用户和运营商提供了有效的筛选工具,以及检测分层数字工件中不一致性的设计原则。

英文摘要:

Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand, users and agents increasingly rely on brief metadata to select third-party skills, making it difficult to detect inconsistencies between a skill's description and its true behavior, a problem we call cross-layer misalignment. To address this issue, we propose Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL), an LLM-based framework that detects misalignment by modeling the layered structure of Agent Skills and learning cross-layer consistency. Using a normalized corpus of over 264,000 open-source skills and a human-verified challenge set, PL-HCL improves Macro-F1 from approximately 0.45 for unadapted baselines to 0.87-0.89 across evaluated LLM backbones. This approach offers an effective screening tool for users and operators, as well as design principles for detecting inconsistencies in layered digital artifacts.

补充信息

↑