arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProCap:用于忠实和全面视频字幕的突出引导对象校正

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Debjyoti Das Adhikary, Aritra Hazra, Partha Pratim Chakrabarti

arXiv 2607.21022首次发表:更新:

发表机构

Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur(印度理工学院卡拉格布尔分校计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视频字幕质量提升问题,提出突出感知的迭代事后校正框架。通过轻量级评分机制排序对象,经提示驱动细化循环多轮注入相关对象。在多数据集验证,相比基线提升完整性、减少幻觉,提供了轻量级且模型无关的视频字幕优化途径。

AI 中文摘要

提高视频字幕质量通常需要重新训练大型视觉语言模型,这既昂贵又不切实际。现有的免训练替代方法通过检测到的对象来抑制幻觉,但只进行一次固定的校正,没有区分对象的重要性,导致语义重要内容被遗漏。我们提出了一个突出感知的迭代事后校正框架,在不修改字幕模型参数的情况下克服了这两个限制:一个轻量级评分机制按空间显著性、时间持久性和关系动态对检测到的对象进行排序,一个迭代的、提示驱动的细化循环使用此排序在多轮中逐步将缺失但上下文相关的对象注入字幕。我们在MSVD和MSR-VTT上使用基于对象的自动指标、110名参与者的人类研究以及与ChatGPT和Gemini的定性比较来验证该框架;在人类评估中,相对于强大的预训练字幕基线,该框架将感知完整性提高了48%,将幻觉减少了45%,且无需重新训练或参考字幕。这些结果表明突出引导的迭代校正为更完整和可信的视频字幕提供了一条轻量级、可扩展且与模型无关的途径,与可访问性、检索和其他多媒体理解应用直接相关。

英文摘要

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted. We propose a prominence-aware, iterative post-hoc rectification framework that overcomes both limitations without modifying the underlying captioning model's parameters: a lightweight scoring mechanism ranks detected objects by spatial saliency, temporal persistence, and relational dynamics, and an iterative, prompt-driven refinement loop uses this ranking to progressively inject missing yet contextually relevant objects into the caption over multiple rounds. We validate the framework on MSVD and MSR-VTT using object-grounded automatic metrics, a 110-participant human study, and qualitative comparison against ChatGPT and Gemini; in human evaluation, the framework raises perceived completeness by up to 48% and reduces hallucination by up to 45% relative to a strong pretrained captioning baseline, all without retraining or reference captions. These results position prominence-guided iterative rectification as a lightweight, scalable, and model-agnostic route to more complete and trustworthy video captioning, with direct relevance to accessibility, retrieval, and other multimedia understanding applications.

Comments10 pages, 7 figures, 5 tables. Submitted to IEEE Transactions on Multimedia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑