发表机构
University of Surrey(萨里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过稀疏自编码器和因果干预,揭示视觉语言模型在有害模因检测中的失败主要源于证据路由缺口而非表征缺失,并证明校准路由可恢复大部分性能差距。
AI 中文摘要
当大型视觉语言模型错误分类有害模因时,该失败可能反映内部证据缺失,或无法将已表征的证据路由到输出端。我们使用稀疏自编码器、角色条件探针、因果干预和恢复实验,在Gemma-3和Qwen3.5上区分这些情况,涵盖六个有害内容基准,并附加西班牙语和印地语-英语代码混合评估。稀疏读出在所有六个主要二分类任务上优于原生预测:Qwen平均宏F1从原生0.432提升至0.740,而残差重建达到0.486;Gemma从0.532提升至0.714。这些差异反映监督可访问性,而非预先存在的原生决策规则,且最具影响力的token角色因任务而异。在评估的评分尺度下,Qwen静默特征消融的探针敏感性高出24-63倍,而字面是/否任务上的路由特征修补的输出敏感性高出16-140倍。仅校准路由恢复平均差距的93.3%,探针蒸馏的LoRA改善原生预测,尽管共享多任务适应导致负迁移。对Gemma-3-12B在Facebook仇恨模因上的案例研究发现分布式秩32图像-提示交互,达到0.756对比原生0.685的宏F1。鲁棒性控制表明该信号超越英语,不仅由伴随的OCR解释,且依赖于配对的视觉证据。因此,路由而非仅表征,是有害模因分类中反复出现的瓶颈。
英文摘要
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ for native macro-F1, residual reconstruction reaches $0.486$, and Gemma improves from $0.532$ to $0.714$. These gains measure how accessible the label is to a supervised readout; they do not show that the model's native generation already applies such a decision rule. Under the evaluated scales, Qwen silent-feature ablation is $24-63$ times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is $16-140$ times more output-sensitive. Native-only threshold calibration explains much, but not all of the gap: on five tasks with matched probe scores, it recovers $69.8$\% of the raw native-to-probe difference, while direct routing adds $0.094$ mean macro-F1 beyond calibrated native scoring. Joint gold-label, probe-KL, and pairwise LoRA supervision improves dedicated FHM prediction, but a gold-only adapter performs better on the shared seven-task mean. A case study of Gemma-3-12B on the Facebook Hateful Memes dataset finds a distributed rank-32 image-prompt interaction, reaching $0.756$ versus $0.685$ native macro-F1. Robustness controls show that the signal is not explained solely by accompanying OCR and depends on paired visual evidence, and that it extends beyond English. In many of the errors we study, the evidence is represented but does not reach the answer; therefore, routing is a common bottleneck in harmful meme classification.
Comments42 pages, 9 figures