发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到音视频生成的安全评估空白,提出首个专门基准AV-SafetyBench,含四轴13类分类和5200个提示词,通过三视角评估揭示音视频联合风险,实验表明仅视频评估会遗漏大量不安全输出。
AI 中文摘要
最近的文本到音视频(T2AV)模型能够从单个文本提示词联合生成视频、语音、音效和环境音。这一能力给安全评估带来了新的挑战,因为不安全内容可能通过音轨传播,或仅在视觉和音轨联合解读时出现。现有的安全基准大多只单独关注生成的视频或生成的音频,因此无法捕捉这些风险。为弥补这一空白,我们引入了AV-SafetyBench,这是首个专门为T2AV生成开发的安全基准。AV-SafetyBench包含一个四轴、13类别的分类体系,以及5,200个经过人工审核的提示词,这些提示词指定了视觉场景、语音和非语音音频。我们的评估协议在三种视角下评估每个输出:完整音视频(Full-AV)、仅视频(Video-Only)和仅音频(Audio-Only)。然后,它使用仅视频和仅音频的判断将Full-AV不安全输出归因于四种风险来源之一:仅视频(Video-Only)、仅音频(Audio-Only)、音视频共有(AV-Both)或音视频联合(AV-Joint)。我们评估了五个开源T2AV模型,并针对人工标注验证了自动化的Full-AV判断。在这五个模型中,Full-AV不安全率在25.1%至49.4%之间。除了这些总体比率外,风险来源分析显示,对于五个模型中的四个,仅音频和音视频联合案例(即仅视频评估遗漏的不安全输出)占可归因风险来源的Full-AV不安全输出的41.6%至48.3%。在跨模态危害涌现类别中,音视频联合占具有已归因风险来源的不安全输出的87.5%。这些发现共同证明了AV-SafetyBench在评估T2AV安全(涵盖视觉和音频模态及其交互)方面的价值。
英文摘要
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a single text prompt. This capability poses new challenges for safety evaluation, as unsafe content may be conveyed through the audio track or arise only when the visual and audio tracks are interpreted jointly. Existing safety benchmarks largely focus on either generated video or generated audio in isolation and are therefore not designed to capture these risks. To close this gap, we introduce AV-SafetyBench, the first safety benchmark developed specifically for T2AV generation. AV-SafetyBench comprises a four-axis, 13-category taxonomy and 5,200 manually reviewed prompts that specify visual scenes, speech, and non-speech audio. Our evaluation protocol assesses each output under three views: Full-AV, Video-Only, and Audio-Only. It then uses the Video-Only and Audio-Only judgments to assign Full-AV unsafe outputs to one of four risk sources: Video-Only, Audio-Only, AV-Both, or AV-Joint. We evaluate five open-source T2AV models and validate the automated Full-AV judgments against human annotations. Across the five models, Full-AV Unsafe Rates range from 25.1% to 49.4%. Beyond these aggregate rates, risk-source analysis reveals that, for four of the five models, Audio-Only and AV-Joint cases-unsafe outputs missed by video-only evaluation-account for 41.6-48.3% of Full-AV unsafe outputs for which a risk source could be assigned. In the Cross-Modal Harm Emergence category, AV-Joint accounts for 87.5% of unsafe outputs withan assigned risk source. Together, these findings demonstrate the value of AV-SafetyBench for evaluating T2AV safety across the visual and audio modalities and their interaction.
Comments34 pages, 19 figures, 17 tables