多语言多智能体规划失败的可操作诊断
An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
浏览论文内容
中文总结 AI 辅助
本研究针对多语言多智能体系统在非英语场景下的规划失败问题,提出规划-接地失败分类法,开发TART方法,在多语言任务中显著提升了现有系统的性能。
中文摘要 AI 辅助
多语言多智能体系统在非英语场景下性能大幅下降,而现有研究很少关注用户请求转换为可执行计划时任务关键信息的丢失问题。本研究将多智能体系统中的规划器视为请求到动作的接口,从真实世界任务执行失败案例中推导得出规划-接地失败的可操作分类法。基于大语言模型(LLM)的分析显示,随着语言资源可用性降低,这类失败在不成功执行中所占比例不断上升,在低资源语言中影响最为显著。为验证该分类法是否可用于缓解问题,本文提出TART(Taxonomy-Guided Actionable Representation,即分类法引导的可操作表示),该方法将分类法的关键方面明确呈现给规划器和下游子智能体。在多种语言、三种LLM主干、两个数据集及两种智能体配置下,TART均能稳定提升性能;在多语言GAIA数据集上,它将某一当前最优系统的准确率在涵盖低资源到高资源设置的11种语言中平均提高了5.6个百分点。
英文摘要
Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are converted into executable plans. We study the planner in a multi-agent system as the request-to-action interface and derive an actionable taxonomy of planning-grounding failures from failed real-world task executions. LLM-based analysis shows that these failures constitute an increasing share of unsuccessful executions as language-resource availability declines, with the strongest effects in low-resource languages. To test whether the taxonomy supports mitigation, we introduce TART, Taxonomy-Guided Actionable Representation, that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents. Across multiple languages, three LLM backbones, two datasets, and two agentic configurations, TART consistently improves performance. On multilingual GAIA, it raises a state-of-the-art system's accuracy by 5.6 percentage points averaged across eleven languages spanning low- to high-resource settings.
发表机构
- Fujitsu Research of Europe(富士通欧洲研究院)
- Cohere(科here公司)
机构由 AI 辅助整理,请以论文原文为准。