发表机构
National University of Singapore; University of New South Wales; Fuzhou University(新加坡国立大学; 新南威尔士大学; 福州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示文本到视频安全过滤器因忽略帧顺序而存在时间调节缺口,并证明顺序感知检测器能有效识别此类有害片段。
AI 中文摘要
文本到视频(T2V)服务沿用了图像生成的安全机制,将关键词提示过滤器与逐帧检查器配对,一旦某个采样帧看起来不安全,该检查器就会阻止整个片段。这种机制存在视频特有的盲区。我们证明,任何忽略帧顺序的审核器,只要接受了该片段的无害洗牌版本,就必然会接受有害片段,因此仅由顺序承载的危害得以逃脱。实验上,未修改的基准提示已在四个保留种子上,对32.7%的连续动作目标使片段落入此调节缺口,而改写、场景分割和反馈驱动的提示搜索均未显示显著改进(配对McNemar检验p≥0.12),因此无需提示工程即可暴露此漏洞。对全部97个渲染帧进行密集评分显示,约三分之一的已交付片段仅是隐藏了不安全帧,而其余片段即使每帧都通过,作为有序视频仍然有害,未修改提示在四分之一的连续动作目标上达到了这种无视顺序的残余。我们还记录了一个测量陷阱,因为对搜索到的提示在其自身渲染种子上评分,会将7.5%的每代生成率虚增为46.7%。一项用户研究证实,人们将这些片段解读为有害,而将其洗牌版本解读为安全。解决方案是读取帧顺序,一个顺序感知检测器以AUC 0.74将这些片段与其自身洗牌版本区分开来,而逐帧检查则处于随机水平,这正是已部署审核所丢弃的信号。
英文摘要
Text-to-video (T2V) services inherit their safety stack from image generation, pairing a keyword prompt filter with a per-frame checker that blocks a clip whenever one sampled frame looks unsafe. This stack has a blind spot unique to video. We prove that any moderator ignoring frame order accepts a harmful clip whenever it accepts that clip's benign shuffle, so harm carried by the ordering alone escapes. Empirically, the unmodified benchmark prompt already lands a clip in this moderation gap on 32.7% of Sequential-Action targets over four held-out seeds, and paraphrasing, scene splitting, and a feedback-driven prompt search show no significant improvement (paired McNemar $p\ge0.12$), so prompt engineering is not needed to expose the vulnerability. Dense-scoring all 97 rendered frames shows that about a third of the delivered clips merely hide an unsafe frame, while the rest stay harmful as ordered videos even though every frame passes, an order-blind residual the unmodified prompt reaches on a quarter of Sequential-Action targets. We also document a measurement pitfall, since scoring a searched prompt on its own render seed inflates a 7.5% per-generation rate into an apparent 46.7%. A user study confirms that people read these clips as harmful and their shuffles as safe. The fix is to read frame order, and an order-aware detector separates these clips from their own shuffles at AUC 0.74 where per-frame checking sits at chance, which is the signal deployed moderation throws away.
Comments12 pages, 2 figures