发表机构
Communication University of China; Australian Institute for Machine Learning, Adelaide University; University of New South Wales; University of Auckland; Huazhong University of Science and Technology(中国传媒大学; 阿德莱德大学澳大利亚机器学习研究所; 新南威尔士大学; 奥克兰大学; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对漫画视觉问答的分镜叙事结构问题,提出无监督框架ManGo,通过主动叙事草图结合双奖励的组相对策略训练,在标准基准上取得最优性能。
AI 中文摘要
漫画视觉问答要求模型回答基于分镜的视觉叙事问题,相关证据分布在有序分镜、嵌入文本、重复出现的角色以及隐含的事件转换中。这种结构使得被动的页面编码不足,因为模型必须识别要检查的分镜、保留哪些线索,以及何时积累的证据足以回答问题。我们提出ManGo(Manga Active Narrative Grounding Optimization),这是一个用于主动漫画视觉问答的无监督框架。ManGo引入主动叙事草图(ANS),它迭代选择分镜、提取简洁的定位线索,并决定何时停止,在生成答案前形成一个紧凑的、面向问题的证据草图。为了在没有人工标注答案或理由路径的情况下优化这种行为,ManGo采样多个ANS展开,并应用具有两个奖励的组相对训练:来自列表式自排序的答案偏好,以及来自稳定有序分镜轨迹的路径一致性。组合奖励通过组相对策略训练进行优化,鼓励模型同时改进最终答案和支持答案的分镜级证据路径。在标准漫画理解基准上的实验表明,ManGo在不同设置下实现了最先进的性能。
英文摘要
Manga visual question answering requires models to answer questions over panel-based visual narratives, where relevant evidence is distributed across ordered panels, embedded text, recurring characters, and implicit event transitions. This structure makes passive page encoding insufficient, as the model must identify which panels to inspect, what clues to retain, and when the accumulated evidence is sufficient for answering. We propose ManGo (Manga Active Narrative Grounding Optimization), an unsupervised framework for active manga visual question answering. ManGo introduces Active Narrative Sketching (ANS), which iteratively selects panels, extracts concise grounded clues, and decides when to stop, forming a compact question-directed evidence sketch before answer generation. To optimize this behavior without human-annotated answers or rationale paths, ManGo samples multiple ANS rollouts and applies group-relative training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. The combined reward is optimized with group-relative policy training, encouraging the model to improve both final answers and the panel-level evidence paths that support them. Experiments on standard manga understanding benchmarks show that ManGo achieves state-of-the-art performance across different settings.
Comments16 pages, 9 figures, 7 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026