发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长SKILL合规性检测的成本与精度矛盾,提出SkillCDG框架,结合两级检索与依赖闭包实现检测,在多数据集上优于基线,还发现尺度趋势以优化小模型性能。
AI 中文摘要
企业业务场景日益复杂,推动了智能体系统中长SKILL文档的广泛应用,给合规性检测带来新挑战:大模型会产生大量推理成本,而小模型可能无法维持检测精度。为解决这一差距,我们提出SkillCDG,一种基于图的长SKILL合规性检测框架。SkillCDG将复杂业务策略表示为两层约束依赖图:上层索引SKILL描述以实现场景路由,下层捕获每个SKILL内原子约束间的依赖关系。推理过程中,两级检索结合依赖闭包支持合规性判断与来源可追溯性。我们在三个企业数据集和两个受控公共基准变体上对该框架进行全面评估。实验结果表明,SkillCDG在检测F1分数上比基线方法高出最多12.8个百分点,同时将token消耗最多降低64.3%。此外,我们进一步研究了策略图复杂度、模型规模与检测性能之间的内在关系。在单一模型家族的四个检查点上开展的对比实验验证了一条简洁有效的尺度趋势:端到端检测正确性呈现出复杂度差异化的尺度模式,且由约束依赖图推导的复杂度指标可有效量化实例难度与模型的性能提升潜力。利用这一有洞见的尺度趋势,我们开展自适应训练样本选择并采用在线策略蒸馏,以高效增强小模型的合规性检测能力。
英文摘要
The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new challenges for compliance detection: large models incur substantial inference costs, while small models may fail to maintain detection accuracy. To address this gap, we propose SkillCDG, a graph-based framework for long SKILL compliance detection. SkillCDG represents complex business policies as a two-layer constraint dependency graph, where the upper layer indexes SKILL descriptions for scenario routing and the lower layer captures dependencies among atomic constraints within each SKILL. During inference, two-level retrieval followed by dependency closure supports compliance judgment and source traceability. We comprehensively evaluate the framework on three enterprise datasets and two controlled public benchmark variants. Experimental results demonstrate that SkillCDG outperforms baseline methods by up to 12.8 percentage points in detection F1 score, while reducing token consumption by a maximum 64.3\%. Moreover, we further investigate the inherent relationships among policy-graph complexity, model scale, and detection performance. Comparative experiments conducted on four checkpoints from a single model family validate a concise and effective scaling trend: end-to-end detection correctness exhibits a complexity-differentiated scaling pattern, and the complexity metric derived from the constraint dependency graph can effectively quantify instance difficulty and the performance improvement potential of models. Leveraging this insightful scaling trend, we conduct adaptive training sample selection and adopt on-policy distillation to efficiently enhance the compliance detection capability of small-scale models.