arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18958cs.MMcs.AI

转录,然后推理:多模态评论的两遍分解

Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review

Bojie Li, Noah Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态模型单遍评论长内容时丢失信息的问题,提出先转录后评论的两遍分解方法,显著提升忠实度与覆盖率。

中文摘要 AI 辅助

使用多模态模型评论长录音或文档的自然方式是,将原始资料交给模型,并在一次调用中请求评论。我们表明这种方式会悄然失败:模型会满足于次优结果,丢弃大约三分之一的内容,并修饰其余部分。这种失败并非感知问题——当同一个模型被简单要求转录资料时,几乎所有被丢弃的内容都会重新出现。瓶颈在于负载下的生成:单遍处理无法同时感知、推理并撰写一篇长而忠实的评论,因为同时执行这三项任务会竞争同一个输出。我们排除了明显的替代方案。这不是模态问题:模型阅读文本和同一文本的图像表现同样出色。这也不仅仅是更深入思考的问题:给单遍处理一个更大的推理预算并不能恢复丢失的内容,因为模型会将预算用于规划评论,而不是写下资料。有效的方法是将工作分配给两个相同权重的遍次——先转录,然后评论转录内容——这样每一步都有自己完整的输出预算。这种先转录后评论的分解方法在21个来源的测试套件中提高了忠实度和覆盖率。这种收益并非均匀:我们观察到,在一遍基线最弱的地方,它帮助最大,而在基线已经很强的地方,帮助很小,这种模式也随来源的长度和模态而变化。分解带来了两种失败模式——评论遍次在非常长的来源上空间不足,以及一旦基础来源被移除,从记忆中虚构内容。

英文摘要

The natural way to review a long recording or document with a multimodal model is to hand it the raw source and ask for a review in one call. We show that this quietly fails: the model satisfices, dropping roughly a third of the content and embellishing the rest. The failure is not perception--almost all of the dropped content reappears when the same model is simply asked to transcribe the source. The bottleneck is generation under load: a single pass cannot perceive, reason over, and write a long faithful review at the same time, because doing all three competes for one output. We rule out the obvious alternatives. It is not the modality: models read text and an image of the same text equally well. And it is not merely a matter of thinking harder: giving the single pass a far larger reasoning budget does not recover the lost content, because the model spends that budget planning a review rather than writing the source down. What works is to split the labor across two same-weights passes--first transcribe, then review the transcript--so each step gets a full output budget of its own. This transcribe-then-review decomposition improves both faithfulness and coverage across a 21-source suite. The benefit is not uniform: we observe that it helps most where the one-pass baseline is weakest and little where that baseline is already strong, a pattern that also tracks the source's length and modality. Decomposition comes with two failure modes--the review pass running out of room on very long sources, and confabulating from memory once the grounding source is removed.

发表机构

  • Pine AI
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑