arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16986cs.MA

ToMAS:基于多智能体LLM失败事件的试点失败驱动心智理论基准

ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

Muhammad Ashar Ishfaq, Glaucia Melo

首次发表
浏览论文内容

中文总结 AI 辅助

ToMAS提出将多智能体LLM失败案例转化为伙伴状态推理条目的基准,通过转换规则生成训练数据,并揭示GRPO实验因LoRA更新过小而无效,为后续研究指明方向。

中文摘要 AI 辅助

基于LLM的多智能体系统即使在通信成功的情况下也可能失败,因为智能体未能正确追踪同伴的角色、知识或意图。我们研究了此类智能体间错位案例(在MAST-Data中标记为FC2)是否可转化为功能性伙伴状态推理条目。ToMAS对诊断出的执行轨迹应用了四条明确的转换标准。对242条符合条件的非AG2训练轨迹进行完整转换后,产生了39个CLEAN条目。在18条轨迹的可靠性试点中,两名标注者达到了94.4%的原始一致率和Cohen's kappa = 0.92。随后,我们将转换后的条目作为二元奖励,在Qwen2.5-1.5B上进行了一项小规模GRPO可行性实验。在28项留出Magentic GAIA诊断集上,每个评估条件在相同的2/28条目上超过了ROUGE-L阈值。事后适配器检查揭示了原因:在所使用的学习率下,LoRA更新在数值上可忽略不计(最大绝对Delta W约为7e-6),因此所有条件的解码结果与未训练检查点完全相同。因此,该实验未显示出训练效果,也无法确立训练效果;它报告了一个可执行的流程,并指出了任何结论性研究必须解决的两个局限性:训练与评估条目之间的来源差距,以及词汇重叠评分。ToMAS为将诊断出的协调失败转化为可训练的伙伴状态推理条目提供了初步的规则和流程,并确定了进行结论性匹配域评估的要求。

英文摘要

LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers' roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies four explicit convertibility criteria to diagnosed execution traces. A full conversion pass over 242 eligible non-AG2 training traces produced 39 CLEAN items. In an 18-trace reliability pilot, two annotators achieved 94.4% raw agreement and Cohen's kappa = 0.92. We then used the converted items as binary rewards in a small-scale GRPO feasibility experiment with Qwen2.5-1.5B. On a 28-item held-out Magentic GAIA diagnostic, every evaluated condition exceeded the ROUGE-L threshold on the same 2 of 28 items. Post-hoc adapter checks show why: under the learning rate used, the LoRA update remained numerically negligible (max abs Delta W about 7e-6), so all conditions decode identically to the untrained checkpoint. The experiment therefore does not show a training effect and cannot establish one; it reports an executable pipeline together with two limitations that any conclusive study must address: a provenance gap between the training and evaluation items, and lexical-overlap scoring. ToMAS provides a preliminary rubric and pipeline for converting diagnosed coordination failures into trainable partner-state reasoning items and identifies the requirements for a conclusive matched-domain evaluation.

发表机构

  • The Islamia University of Bahawalpur(巴哈瓦尔布尔伊斯兰大学)
  • Toronto Metropolitan University(多伦多都会大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑