AI 中文总结
本文提出四维MCRI框架及基于大模型的MCRI-Eval评估方法,利用信息增益和行为约束,在三个基准上显著提升技能选择的下游性能排名,为技能评估提供预执行信号。
AI 中文摘要
随着智能体从单一工具系统演变为模块化、复合架构,技能正成为能力开发和分发的重要机制。然而,学术界缺乏一个结构化的框架来系统地分析和评估技能。借鉴信息增益和行为约束,我们提出了四维MCRI框架,并将其实现为MCRI-Eval,一种基于大型语言模型的评估方法。我们使用来自OpenClaw技能中心的63,812个公共技能对MCRI-Eval进行了评估,在BigCodeBench、BFCL-Fundamental和Mind2Web上进行了58,275次技能条件下的模型执行。MCRI-Eval得分与社区流行度信号正相关,并在所评估的方法中实现了最高的下游排名一致性。MCRI-Eval还提高了所有三个基准上的前1名技能选择:与每个基准上最强的基线相比,MCRI-Eval选择的技能在BigCodeBench、BFCL-Fundamental和Mind2Web上的下游性能排名分别提升了17.7、22.8和19.6个百分点。这些结果表明,MCRI-Eval为在昂贵的基于执行的评估之前优先选择有前景的技能提供了一个有用的预执行信号。
英文摘要
As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
Comments24PAGES