TPCD:音调-压力对比解码与视觉语言模型中的无标签门控瓶颈
TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文提出TPCD方法,利用压力诱导分布作为对比解码负分支,结合门控缓解视觉语言模型受高压提示时的不合理承诺问题,在多个基准测试中显著降低攻击成功率。
中文摘要 AI 辅助
高压提示会迫使视觉语言模型(VLMs)做出不合理的承诺,例如识别难以辨认的文字、报告不确定的时间或确认不存在的物体。本文探究压力诱导的分布本身是否可作为对比解码的负分支。音调-压力对比解码(TPCD)将高压指令下产生的logits与安全中性指令下产生的logits相减。在800个样本的tone-matters基准测试中,受压力的LLaVA-1.5-7B达到66.75%的攻击成功率(ASR);安全中和将ASR降至9.88%;完整TPCD使ASR降至0.50%但将正样本降至15.56%。针对LLaVA的分析作为设计划分,在GLM-4.6V和Llama-3.2-Vision上进行了n=800的负样本和n=780的匹配正样本保留运行,结果显示简单门控可优于安全中和,敏感性分析限定了弱时间正样本子任务。无类别先验的答案分歧路由器将保留的总ASR降至6.93%,优于安全中和(10.98%)和分支分歧(9.67%),同时匹配分支分歧的79.94%正样本准确率,不过它仍是事后且基于表面形式的。我们得出结论,压力是承诺偏差的有用探测工具和可行的缓解信号,但当前门控尚未得到独立验证的基于接地的检测器。
英文摘要
High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced distribution itself can serve as a contrastive-decoding negative branch. Tone-pressure contrastive decoding (TPCD) subtracts logits produced under a high-pressure instruction from logits produced under a safe neutral instruction. On the 800-example tone-matters benchmark, LLaVA-1.5-7B under pressure reaches 66.75% attack success rate (ASR); safe neutralization reduces ASR to 9.88%; full TPCD reaches 0.50% but collapses positives to 15.56%. A benchmark-specific task-prior/disagreement gate preserves measured positive accuracy (54.44%) while lowering ASR to 1.63% on LLaVA. Treating this LLaVA analysis as the design split, full $n=800$ negative and $n=780$ matched-positive held-out runs on GLM-4.6V and Llama-3.2-Vision show that simple gates can improve over safe neutralization, with sensitivity analyses bounding the weak time-positive subtask. A category-prior-free answer-disagreement router reduces held-out aggregate ASR to 6.93%, improving over both safe neutralization (10.98%) and branch disagreement (9.67%) while matching branch disagreement's 79.94% positive accuracy, although it remains post-hoc and surface-form based. We conclude that pressure is a useful probe of commitment bias and a viable mitigation signal, but the current gates are not yet independently validated grounding-aware detectors.
发表机构
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。