arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DREAM:多模态多智能体辩论中的动态分辨率分配

DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

Khanh-Binh Nguyen, Van Dai Do, Tien Anh Nguyen, Svetha Venkatesh, Hung Le

arXiv 2610.05615首次发表:更新:

发表机构

Deakin University; Applied Artificial Intelligence Initiative(迪肯大学; 应用人工智能计划)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DREAM通过动态分辨率分配和不确定性引导回滚聚合,解决多模态多智能体辩论中的视觉尺度差异与群体思维问题,在六个数据集上提升准确率1.5-3.2%。

AI 中文摘要

多智能体辩论(MAD)已成为提升大型语言模型(LLMs)推理能力的有效范式,并日益扩展到多模态场景。然而,现有的多模态MAD框架通常让智能体暴露于相同的固定视觉输入,忽略了不同样本和智能体之间所需视觉尺度的显著差异。此外,这些框架经常遭受群体思维(groupthink)现象的影响,即智能体过早放弃正确推理,转而迎合自信但产生幻觉的同伴回应。为解决这些瓶颈,我们提出了DREAM(多模态多智能体辩论中的动态分辨率分配),其通过两个核心组件运作:(1)动态分辨率分配,一个零样本探测轮,智能体测试多种分辨率,使用平均归一化对数似然(ANLL)量化不确定性,并利用自适应阈值为每个智能体分配其经验最优分辨率;(2)不确定性引导的回滚聚合,通过跟踪每个智能体在轮次间的不确定性,并恢复被群体压力覆盖的早期低不确定性答案,来对抗群体思维。在六个多模态数据集上,DREAM在无需数据集特定调优的情况下,相较于多智能体辩论基线,在准确率-令牌权衡上提升了1.5-3.2%的准确率。

英文摘要

Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑