发表机构
Oregon State University(俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次系统性探究文本到视频(T2V)扩散模型对硬件故障的鲁棒性,发现单个故障可降低性能、改变语义,内存故障危害更大,揭示了已部署T2V系统的可靠性风险。
AI 中文摘要
本文首次系统性研究了文本到视频(T2V)扩散模型在随机硬件级故障下的鲁棒性。尽管T2V模型因能生成高质量、时间连贯且逼真的视频而被广泛用于自动化视频生成,但其迭代去噪过程与时空依赖关系会引入独特的故障模式。我们开展了涵盖计算故障与内存故障的大规模故障注入研究,涉及3种T2V模型及1个代表性基准。结果显示:(1)单个故障可使整体性能下降达3.7%,语义正确性比感知质量受影响更严重;(2)内存故障比计算故障破坏性更强,高阶指数位尤其易受攻击,且广泛使用的bfloat16格式比替代格式更易受影响;(3)7%至28%的故障会引发可见伪影,包括添加物体等语义变化,表明单个故障足以改变输出语义。我们的发现揭示了已部署T2V系统的可靠性风险,并为提升故障鲁棒性的后续研究提供了动力。代码:[链接]。
英文摘要
We present the first systematic study of the resilience of text-to-video (T2V) diffusion models under random hardware-level faults. While T2V models are widely used for automated video generation due to their ability to produce high-quality, temporally coherent, and realistic videos, their iterative denoising process and spatiotemporal dependencies introduce unique failure modes. We perform an extensive fault-injection study covering both computational and memory faults across three T2V models and a representative benchmark. Our results show that (1) a single fault can degrade overall performance by up to 3.7\%, with semantic correctness more affected than perceptual quality; (2) memory faults are more damaging than computational faults, high-order exponent bits are particularly vulnerable, and the widely-used bfloat16 is more susceptible than alternative formats; and (3) 7-28\% of faults cause visible artifacts, including semantic changes such as added objects, suggesting that single faults are sufficient to alter output semantics. Our findings reveal reliability risks in deployed T2V systems and motivate further research on improving fault resilience. Code: \href{https://github.com/ztcoalson/T2V-Resilience}{https://github.com/ztcoalson/T2V-Resilience}.
CommentsAccepted to ICML 2026 Workshop on From Frames to Stories (F2S)