发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出M-SQE框架,通过理论视图和行动视图统一评分,评估检索后技能质量,提升低资源语言任务成功率,促进智能体技能使用的语言平等。
AI 中文摘要
智能体技能是可复用的程序化文档,将LLM智能体扩展到其参数记忆之外,已成为在现实任务中部署智能体的重要接口。围绕该接口构建的社区维护技能库正在快速增长。然而,这一生态系统仍深度以英语为中心:我们的审计发现,斯瓦希里语和印地语等低资源语言没有语言内技能内容,因此检索通常会返回与查询语言不同的技能,从而降低准确率和召回率。一种实用的解决方案是为检索合成语言内技能,但其质量可能不可靠,因此仅凭该场景下的相关性往往会浮现出相关但不可用的候选。为解决此问题,我们提出M-SQE,一种检索后的多语言技能质量估计框架,通过理论视图评估内在质量,通过行动视图评估基于任务的实用性,并将两者统一为领域条件化的最终分数。我们在三个技能使用领域评估M-SQE:通用、工具使用和文化任务。实证中,我们构建了反映当前生态系统的三层候选技能池,其中M-SQE的任务成功率在三种不同检索器上平均至少超过现有基线+3.5个百分点。特别地,M-SQE对最低资源语言的提升最大(印地语+12.9个百分点,斯瓦希里语+5.6个百分点),并在所有六个文化区域取得强劲性能,从而推动智能体技能使用走向语言和文化平等。
英文摘要
Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However, this ecosystem remains deeply English-centric: our audit finds that low-resource languages such as Swahili and Hindi have no in-language skill content, so retrieval often returns a skill written in a different language than the query, degrading accuracy and recall. A practical solution is to synthesize in-language skills for retrieval but the quality can be unreliable, so relevance in this setting alone often surfaces a related but unusable candidate. To address this, we propose M-SQE, a post-retrieval Multilingual Skill Quality Estimation framework that scores candidates via a Theory view for intrinsic quality and an Action view for task-grounded utility, unified into a domain-conditioned final score. We evaluate M-SQE across three skill-use domains: general, tool-use, and cultural tasks. Empirically, we build three-layer candidate skill pools mirroring today's ecosystem, where M-SQE's task success exceeds existing baseline's average by at least +3.5 points across three different retrievers. Particularly, M-SQE lifts the lowest-resource languages most (+12.9pp on Hindi and +5.6pp on Swahili) and achieves strong performance across all six culture regions, thereby moving agentic skill use toward linguistic and cultural equality.
Comments17 pages, 6 figures