arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于推理的防护栏何时效率不高?ResponseGuard:用于实时审核的快速视觉语言防护工具

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

Dongbin Na

arXiv 2607.21401首次发表:更新:

发表机构

POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言防护栏筛查回复是否需推理,提出无链式推理的ResponseGuard,通过对请求、回复和图像的单一合并表示一次前向传递读取有害裁决,在回复有害性检测上优于基于推理的防护栏,且时间成本低,还发现了差距原因及推理防护栏注意力问题。

AI 中文摘要

视觉语言人工智能助手以生成的令牌流形式返回答案,因此监控答案的安全防护工具必须跟上该流,并在用户读取之前阻止有害回复。近期的视觉语言防护栏在做出裁决前会生成一系列推理步骤,认为逐步推理能产生更安全的防护,但这种设计使防护栏笨重且缓慢。本文提出视觉语言防护栏筛查回复是否真的需要推理的问题,并给出了一种无链式推理的防护工具ResponseGuard。它通过对请求、回复和图像的单一合并表示进行一次前向传递来读取有害裁决。在标准多模态防护栏基准测试中,2B的ResponseGuard在回复有害性检测方面优于近期基于3B推理的视觉语言防护栏,无需任何推理且时间成本降低约150倍。在请求有害性方面,推理防护栏仍保持总体领先,差距主要在仅图像的单元格上。研究发现差距可能源于两种设计都使用的冻结视觉编码器,而非缺少链式推理。此外,推理防护栏几乎没有将裁决注意力指向图像。基于单次检测,ResponseGuard可以逐句筛查答案流并在有害答案完成前阻止。对于视觉语言模型的回复防护,校准后的单次标签可能提供足够的安全信号。研究还完全发布了所有源代码、训练模型和数据集。

英文摘要

A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up with the stream and stop a harmful reply before a user reads it. Recent vision-language guardrails instead generate a chain of thought before they issue a verdict. They believe that step-by-step reasoning yields a safer guard. This design makes the guard heavy and slow, since the model must decode many tokens for harmfulness detection. We pose the question of whether a vision-language guard really needs to reason in order to screen a response. We answer with a guard that has no chain. ResponseGuard reads a harmful verdict from a single pooled representation of the request, the response, and the image in one forward pass. Across a standard multimodal guardrail benchmark, our 2B ResponseGuard outperforms a recent 3B reasoning-based vision-language guard on response harmfulness detection, without any reasoning and at about 150 times lower time cost. On request harmfulness the reasoning guard retains an overall lead, and the remaining gap on both tracks sits on the image-only cells. We observe that the gap may stem from the frozen vision encoders that both designs use rather than from the missing chain. We have also found the reasoning guard directs almost none of its verdict attention to the image. Based on a single-pass detection, ResponseGuard can screen an answer sentence by sentence as it streams and stop a harmful answer before it finishes. For guarding the response of a vision-language model, a calibrated single-pass label may provide a sufficient safety signal. We fully release all source code, trained models, and datasets at https://github.com/ndb796/ResponseGuard.

Comments8 pages, 6 figures, 3 tables. Project page: https://ndb796.github.io/ResponseGuard ; Code: https://github.com/ndb796/ResponseGuard

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑