AI 中文总结
研究视频-音频一致性问题,核心方法是将其作为有限轨迹模态监测,通过公式库规定多种一致性,证书记录关键信息,工件进行视频解码等操作,贡献是提供可重现见证及相关处理流程。
AI 中文摘要
多模态视频系统通常会暴露片段级分数,却隐藏导致视频不一致的局部时间故障。我们将视频-音频一致性表述为对同步的视觉、音频和字幕/OCR原子的有限轨迹模态监测。公式库规定了语音-说话者一致性、视听事件一致性、字幕-视频一致性、场景连续性以及编辑引起的时间冲击。证书记录轨迹哈希、公式标识符、判定结果、违规索引、缺陷分数、首个反例窗口、执行引擎和证书哈希。该证书并非探测器正确性的证明,而是有限轨迹检查器在原子轨迹固定后能独立重建的可重现见证。工件解码真实MP4视频,提取CLIP视觉原子、AST音频原子,进行密集反事实扰动扫描,发出CSV轨迹和JSON证书,并从这些文件渲染手稿图。
英文摘要
Multimodal video systems often expose clip-level scores while hiding the local temporal failure that makes a video inconsistent. We formulate video-audio consistency as finite-trace modal monitoring over synchronized visual, audio, and subtitle/OCR atoms. A formula library specifies speech-speaker agreement, audio-visual event agreement, subtitle-video agreement, scene continuity, and edit-induced temporal shock. A certificate records the trace hash, formula identifier, verdict, violating indices, defect score, first counterexample window, execution engine, and certificate hash. The certificate is not a proof that a detector is correct; it is a reproducible witness that the finite-trace checker can independently reconstruct once the atom trace is fixed. The artifact decodes real MP4 video, extracts CLIP visual atoms, extracts AST audio atoms, runs dense counterfactual perturbation sweeps, emits CSV traces and JSON certificates, and renders manuscript figures from those files.
Comments15 pages, 5 figures, 4 tables; ancillary files https://huggingface.co/datasets/Lightcap/pcmt-artifact