PercepCap:具有结构化时空感知的视频字幕生成器
PercepCap: Video Captioner with Structured Spatio-Temporal Perception
浏览论文内容
中文总结 AI 辅助
研究视频字幕生成问题,提出PercepCap框架,遵循感知-描述生成链,设计两阶段训练策略及相关数据构建方法,在评估中优于基线,提升视频字幕生成质量。
中文摘要 AI 辅助
视频字幕需要对视频进行细粒度的时空理解,包括物体位置的空间感知和事件发生时间的时间感知。现有多模态语言模型通常直接从视频输入生成字幕,而不展示描述背后的感知证据。因此,时空感知错误只能在最终字幕中观察到,难以直接识别潜在的感知错误。为了解决这些问题,我们提出了PercepCap,这是一个感知感知视频字幕框架,在生成最终字幕之前使感知证据明确。具体来说,PercepCap遵循感知-描述生成链,模型首先生成包含物体轨迹和时间事件的时空感知轨迹,然后根据感知到的证据生成最终字幕。为了支持这一点,我们设计了一个两阶段训练策略。感知-然后-描述监督微调使模型从仅字幕生成适应到提出的感知-描述链,而感知基础强化学习通过对感知链和最终字幕的联合奖励来优化感知轨迹和字幕质量。为了支持我们的两阶段训练,我们引入了字幕锚定感知数据构建。该管道通过首先生成仅字幕描述,提取其中提到的物体和事件,并将它们用框和时间戳重新定位到视频中,来构建SFT和RL训练数据。这产生了字幕对齐的感知数据,提供了可靠的训练地面真值,确保明确的感知轨迹和最终字幕指的是相同的物体和事件。在直接字幕和字幕到问答评估中,PercepCap始终优于Qwen3-VL基线,并展示了领先的字幕质量。
英文摘要
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without exposing the perceptual evidence behind descriptions. As a result, mistakes in spatiotemporal perception are only observed in the final caption, making it difficult to identify the underlying perceptual errors directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatiotemporal perception trace comprising object trajectories and temporal events, and then generates the final caption conditioned on the perceived evidence. To support this, we design a two-stage training strategy. Perceive-then-Describe Supervised Fine-tuning adapts the model from caption-only generation to the proposed perceive-describe chain, while Perception-Grounded Reinforcement Learning optimizes perception trace and caption quality with joint rewards over perception chain and the final caption. To support our two-stage training, we introduce Caption-Anchored Perception Data Construction. This pipeline builds the SFT and RL training data by first generating a caption-only description, extracting the objects and events it mentions, and grounding them back in the video with boxes and timestamps. This yields caption-aligned perception data that provides solid training ground truth, ensuring that the explicit perception trace and final caption refer to the same objects and events. Across direct caption and caption-to-QA evaluation, PercepCap consistently improves upon the Qwen3-VL baseline and demonstrates leading caption quality.
发表机构
- Nanjing Univerisity(南京大学)
- Kuaishou Technology(快手科技)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。