发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对VLM灾害评估召回率低的缺陷,通过机械可解释性定位情感回路,注入恐惧方向引导准则转变,在Qwen2.5-VL-7B-Instruct的SeisMLLM流水线中使红色召回率提升至75.7%,实现推理时无需重训的决策调整。
AI 中文摘要
视觉语言模型(VLM)在灾后灾害评估中展现出巨大潜力,但存在一个反复出现的缺陷:它们不愿判定为灾害;即即便整体准确率看似足够,召回率仍较低。本研究采用信号检测理论(SDT)将决策行为分解为感知能力和决策准则位置,以探究该缺陷。受“恐惧使人类风险规避”这一发现启发,我们提出一种纠正过度保守决策策略的新方法,并探究与恐惧相关的情感表征是否可被因果操纵,从而类似地改变VLM的决策倾向。我们利用机械可解释性在模型中定位到一个因果相关的情感回路,并通过激活引导操纵该回路,同时观察其对下游预测的影响。该方法在基于Qwen2.5-VL-7B-Instruct构建的两阶段SeisMLLM流水线中进行测试,该流水线在SeisMLLM-1K测试分割中仅标记了27.0%的真正不安全建筑,且从未发出错误的“红色”(Red)判定,其SDT准则为c=+1.354,尽管证据质量充足(d'=1.521)。我们在富含情感的自然场景上定位到一个情感方向,通过稀疏神经元敲除和对保留情感数据的分布式引导对其进行因果验证,随后将其注入建筑任务。恐惧方向的注入使“红色”召回率提升至75.7%(p<0.001),而减去同一方向则完全抑制了“红色”预测,相比之下,与范数匹配的随机方向和匹配的快乐方向均无显著影响。该机制是准则的转变(c=-1.515),而辨别力未得到提升(d'=-0.493)。这些结果表明,VLM的决策可在推理时无需重新训练即可调整,并证明机械可解释性可用于诊断和控制工程应用中VLM的决策行为。
英文摘要
Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate. This study examines that deficiency by using signal detection theory to decompose the decision behavior into perceptual capability and decision-criterion placement. We then propose a novel method for correcting the over-conservative decision policy, inspired by the finding that fear makes humans risk-averse, and ask whether an affective representation associated with fear can be causally manipulated to similarly alter a VLM's decision tendency. Using mechanistic interpretability, we localize a causally implicated affective circuit in the model and use activation steering to manipulate it while observing the effect on downstream prediction. The method is tested on a two-stage SeisMLLM pipeline built on Qwen2.5-VL-7B-Instruct, which flags only 27.0% of genuinely unsafe buildings on the SeisMLLM-1K test split and never issues a false Red, an SDT criterion of c = +1.354, despite adequate evidence quality (d' = 1.521). An affective direction is localized on emotion-rich natural scenes, causally validated by sparse-neuron knockout and distributed steering on held-out emotion data, and then injected into the building task. Fear-direction injection raises Red recall to 75.7% (p<0.001), and subtracting the same direction suppresses Red predictions entirely, whereas norm-matched random and matched happiness directions show no significant effect. The mechanism is a shift in criterion (c=-1.515) while discrimination is not improved (d'=-0.493). These results show that VLM decisions can be adjusted at inference time without retraining and demonstrate how mechanistic interpretability can be used to diagnose and control VLM decision behaviors in engineering applications.