从“假设”到“事实”:面向视觉脑解码的反事实思维启发语义对齐
From "What-If" to "What-Is": Counterfactual Thinking-Inspired Semantic Alignment for Visual Brain Decoding
AI总结:
提出ConceptAlign框架,通过反事实语义对齐解决视觉脑解码的语义错误问题,在自然场景数据集上相比MindEye2提升了重构、语义区分等指标。
AI中文摘要:
视觉脑解码通过fMRI等神经测量手段重构人所感知的视觉内容,为研究大脑如何表征视觉信息提供了计算方法。近期多模态表征与扩散先验提升了重构的真实性,但视觉上合理的重构可能包含错误的物体、属性或关系,因为强生成先验会补全解码表征未充分指定的内容。传统重构指标主要评估最终图像,可能掩盖此类语义错误。我们提出ConceptAlign,一种面向视觉脑解码的反事实语义对齐框架。ConceptAlign汇聚解码的视觉token并将其投影到冻结的文本嵌入空间,使表征与真实描述对齐,同时将其与保留场景的近错候选区分开。这些由大语言模型(LLM)离线生成的候选会修改一个关键物体、属性或关系,同时保留场景。基于间隔的目标函数学习观测刺激与合理但错误解释之间的细粒度语义边界,推理时无需调用LLM。我们引入系统的三级语义评估框架,涵盖基础可辨别性、反事实描述区分和表征几何。在自然场景数据集(Natural Scenes Dataset)上的实验表明,ConceptAlign相比MindEye2主干网络,提升了重构指标、反事实语义区分度和表征对齐度。匹配的负源消融实验、独立的LLM生成及人工编写的候选,以及人工评估均验证了该监督的有效性与鲁棒性,在细粒度冲突、少数据解码和跨被试结构方面呈现良好表现。
英文摘要:
Visual brain decoding reconstructs visual content perceived by a person from neural measurements such as fMRI, providing a computational approach to studying how visual information is represented in the brain. Recent multimodal representations and diffusion priors have improved reconstruction realism. However, visually plausible reconstructions may contain incorrect objects, attributes, or relations because a strong generative prior can complete content not sufficiently specified by the decoded representation. Conventional reconstruction metrics mainly assess the final image and may therefore obscure such semantic errors. We propose ConceptAlign, a counterfactual semantic alignment framework for visual brain decoding. ConceptAlign pools decoded visual tokens and projects them into a frozen text-embedding space, aligning the representation with the ground-truth caption while separating it from scene-preserving near-miss alternatives. Generated offline by an LLM, these alternatives modify one critical object, attribute, or relation while retaining the scene. A margin-based objective learns fine-grained semantic boundaries between the observed stimulus and plausible but incorrect interpretations without requiring LLM calls during inference. We introduce a systematic three-level semantic evaluation framework covering foundational discriminability, counterfactual description discrimination, and representational geometry. Experiments on the Natural Scenes Dataset show that ConceptAlign improves reconstruction measures, counterfactual semantic discrimination, and representational alignment over the MindEye2 backbone. Matched negative-source ablations, independent LLM and human-written alternatives, and human evaluation support the effectiveness and robustness of the supervision, with favorable patterns in fine-grained conflicts, limited-data decoding, and cross-subject structure.