arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28564cs.CRcs.LG

不要读日志:执行痕迹污染视频生成智能体中的验证器

Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents

Jian Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频生成智能体中执行痕迹对多模态验证器的污染,发现痕迹文本可大幅扭曲视觉判定,且无法通过指令消除,形成修复循环的通过率上限。

中文摘要 AI 辅助

智能体视频生成系统在生成器与验证器之间形成了一个闭环:一个大型语言模型规划镜头,调用文生视频模型,并由一个多模态评判器决定结果是否满足请求。为了诊断长工作流中的失败位置,最近的框架刻意向评判器展示比视频更多的信息——智能体的执行痕迹、其计划、其合成的旁白。我们探究这些辅助文本是否会在保持画面不变的情况下,移动评判器对纯视觉需求的判定。在一个包含109个带人工标注的两事件生成片段基准上,其中请求的事件要么明显完成要么明显缺失,报告成功工具调用的痕迹使得三个开放权重的Qwen-VL评判器(7B、8B、32B)对失败片段的接受率从无文本时的7%–19%上升至78%–90%,而相互矛盾的痕迹则使它们对正确片段的拒绝率高达100%;指示“仅使用画面”并不能消除该效应。前沿的封闭评判器在相同片段上基本不受影响,表明该漏洞是评判器对工具日志的学习信任的属性,而非任务本身的属性。源自计划的文本不携带任何片段特定信息,因此只能移动评判器的工作点,而在修复循环中,这种移动成为真实通过率的上限,任何修复策略都无法超越;该上限与模拟结果精确到小数点后两位。在循环中,污染无需任何对抗性智能体即可被利用:一个总是重新生成的诚实LLM规划器最终获得评判器通过率1.00和人工标注通过率0.28,而一个廉价检查器将其判定写入痕迹的流水线,将该检查器的错误洗白为更强的最终评判器的接受(0.69次错误接受)。

英文摘要

Agentic video-generation systems close a loop between a generator and a verifier: an LLM plans shots, calls a text-to-video model, and a multimodal judge decides whether the result satisfies the request. To diagnose where a long workflow fails, recent harnesses deliberately show the judge more than the video-the agent's execution trace, its plan, the narration it synthesized. We ask whether this auxiliary text moves the judge's verdict on purely \emph{visual} requirements, holding the frames fixed. On a benchmark of 109 generated two-event clips with manual labels, in which the requested event is either visibly completed or visibly missing, a trace that reports a successful tool call makes three open-weight Qwen-VL judges (7B, 8B, 32B) accept $78$--$90\%$ of the failures, up from $7$--$19\%$ without text, and a contradicting trace makes them reject up to $100\%$ of correct clips; an instruction to ``use only the frames'' does not remove the effect. Frontier closed judges are essentially unmoved on the same clips, showing that the vulnerability is a property of the judge's learned trust in tool logs rather than of the task. Plan-derived text carries no clip-specific information, so it can only shift a judge's operating point, and in a repair loop that shift becomes a cap on the true pass rate that no repair policy can exceed; the cap matches simulation to two decimals. In the loop, contamination is exploited without any adversarial agent: an honest LLM planner that always regenerates ends with a judge pass rate of $1.00$ and a human-labelled pass rate of $0.28$, and a pipeline in which a cheap checker writes its verdict into the trace launders that checker's errors into a stronger final judge ($0.69$ false accepts).

发表机构

  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑