arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

M-SQE:多语言技能质量估计以增强智能体技能使用中的语言平等性

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

Yilun Liu, Shimin Tao, Minggui He, Chenxin Liu, Li Zhang, Chen Liu, Miao Zhang, Jiaxin Guo, Min Zhang, Liqun Deng, Xiaojun Meng, Daimeng Wei

arXiv 2609.18445首次发表:更新:

发表机构

Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出M-SQE框架,通过理论视图和行动视图统一评分,评估检索后技能质量,提升低资源语言任务成功率,促进智能体技能使用的语言平等。

AI 中文摘要

智能体技能是可复用的程序化文档,将LLM智能体扩展到其参数记忆之外,已成为在现实任务中部署智能体的重要接口。围绕该接口构建的社区维护技能库正在快速增长。然而,这一生态系统仍深度以英语为中心:我们的审计发现,斯瓦希里语和印地语等低资源语言没有语言内技能内容,因此检索通常会返回与查询语言不同的技能,从而降低准确率和召回率。一种实用的解决方案是为检索合成语言内技能,但其质量可能不可靠,因此仅凭该场景下的相关性往往会浮现出相关但不可用的候选。为解决此问题,我们提出M-SQE,一种检索后的多语言技能质量估计框架,通过理论视图评估内在质量,通过行动视图评估基于任务的实用性,并将两者统一为领域条件化的最终分数。我们在三个技能使用领域评估M-SQE:通用、工具使用和文化任务。实证中,我们构建了反映当前生态系统的三层候选技能池,其中M-SQE的任务成功率在三种不同检索器上平均至少超过现有基线+3.5个百分点。特别地,M-SQE对最低资源语言的提升最大(印地语+12.9个百分点,斯瓦希里语+5.6个百分点),并在所有六个文化区域取得强劲性能,从而推动智能体技能使用走向语言和文化平等。

英文摘要

Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interface are growing rapidly. However, this ecosystem remains deeply English-centric: our audit finds that low-resource languages such as Swahili and Hindi have no in-language skill content, so retrieval often returns a skill written in a different language than the query, degrading accuracy and recall. A practical solution is to synthesize in-language skills for retrieval but the quality can be unreliable, so relevance in this setting alone often surfaces a related but unusable candidate. To address this, we propose M-SQE, a post-retrieval Multilingual Skill Quality Estimation framework that scores candidates via a Theory view for intrinsic quality and an Action view for task-grounded utility, unified into a domain-conditioned final score. We evaluate M-SQE across three skill-use domains: general, tool-use, and cultural tasks. Empirically, we build three-layer candidate skill pools mirroring today's ecosystem, where M-SQE's task success exceeds existing baseline's average by at least +3.5 points across three different retrievers. Particularly, M-SQE lifts the lowest-resource languages most (+12.9pp on Hindi and +5.6pp on Swahili) and achieves strong performance across all six culture regions, thereby moving agentic skill use toward linguistic and cultural equality.

Comments17 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑