发表机构
FPT Software; University of New Brunswick; University of Science and Technology of Hanoi(FPT软件公司; 新不伦瑞克大学; 河内科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出自我进化多智能体框架SEMV,通过争议记忆巩固实现多媒体验证,在COSMOS上达91.88%准确率,负迁移降至0.2%,并支持可追溯、可修订决策。
AI 中文摘要
多媒体验证不仅需要准确的决策,还需要可追溯的证据、可靠的人工修正以及先前经验的安全重用。现有系统往往缺乏明确的机制来修正中间推理或防止有害知识迁移。我们提出了SEMV(自我进化多媒体验证),一个自我进化的多智能体框架,将带有出处论证作为证据、推理、人工争议和记忆之间的接口。SEMV结合了基于竞技场的定量双极论证(A-QBAF)、因果和范围限定修订,以及带有显式冲突保留的验证门控记忆巩固。在COSMOS基准上,SEMV达到91.88%的准确率,而最强可比基线为89.10%。验证记忆将负迁移从5.7%降至0.2%。在CTR基准上,该基准由审稿人争议构建,范围限定因果修订纠正了96.7%的初始错误,同时节省了52.8%的计算量。MV2026大挑战数据集进一步支持基于证据、时间一致的报告。这些结果表明,SEMV可以通过验证经验进化,同时保持累积知识和后续决策的可追溯性、可修订性和可争议性。
英文摘要
Multimedia verification requires not only accurate decisions but also traceable evidence, reliable human correction, and safe reuse of prior experience. Existing systems often lack explicit mechanisms for revising intermediate reasoning or preventing harmful knowledge transfer. We present SEMV (Self-Evolving Multimedia Verification), a self-evolving multi-agent framework that treats provenance-bearing arguments as the interface between evidence, reasoning, human contestation, and memory. SEMV combines arena-based quantitative bipolar argumentation (A-QBAF), causal and scoped revision, and verification-gated memory consolidation with explicit conflict retention. On COSMOS benchmark, SEMV achieves 91.88% accuracy versus 89.10% for the strongest comparable baseline. Verified memory reduces negative transfer from 5.7% to 0.2%. On CTR benchmark, constructed from reviewer contestations, scoped causal revision corrects 96.7% of initial errors while saving 52.8% compute. MV2026 Grand Challenge dataset further supports evidence-grounded, temporally consistent reporting. These results show that SEMV can evolve through verified experience while keeping accumulated knowledge and subsequent decisions traceable, revisable, and contestable.