arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迷失在视觉翻译中:用于脑电到图像重建的基于视觉语言模型的感知语义连贯框架

Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction

Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik

arXiv 2607.12364首次发表:更新:

发表机构

Mahindra University; MU-VT Interdisciplinary Advanced Research Centre for Transformative Technologies, Mahindra University(马欣德拉大学; 马欣德拉大学MU-VT变革性技术跨学科高级研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究脑电到图像重建中视觉保真度与语义可恢复性评估问题,引入基于四个视觉语言模型的框架,通过结构化问题评估图像对,产生宽容感知和语义对齐分数,提炼出脑机接口连贯分数,有效评估感知语义可恢复性。

AI 中文摘要

脑电到图像的评估应区分视觉保真度和可恢复的意义。然而,从脑电得出的重建图像模糊、扭曲且细节少,导致结构相似性(SSIM)、学习感知图像补丁相似度(LPIPS)和对比语言图像预训练(CLIP)对语义上可恢复的输出进行惩罚或奖励看似合理但错误的输出。我们使用语义探针、字幕粗糙度和盲点率以及受控退化分析了来自ATM、ENIGMA、BrainVis和DreamDiffusion的6855个真实/重建图像对。像素指标与语义一致性的相关性接近零,而表示指标则混淆了感知和语义错误。因此,我们引入了一个脑机接口感知框架,其中四个视觉语言模型通过结构化问题评估图像对,产生宽容感知对齐分数(T-PAS)和宽容语义对齐分数(T-SAS)。它们的共识被提炼为脑机接口连贯分数(BCS),这是一个紧凑的评估器,在我们的数据上实现了T-PAS平均绝对误差(MAE)为0.079(r = 0.700)和T-SAS MAE为0.082(r = 0.850)。人类验证显示出高度可靠的联合连贯判断,科恩kappa系数为0.882±0.174,克里彭多夫阿尔法系数为0.882,支持感知语义可恢复性优于一般视觉相似性。代码和资源可在该https网址获取。

英文摘要

EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted, and low-detail, causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible but incorrect ones. We analyze 6,855 ground-truth/reconstruction pairs from ATM, ENIGMA, BrainVis, and DreamDiffusion using semantic probes, caption harshness and blind-spot rates, and controlled degradations. Pixel metrics show near-zero correlation with semantic consistency, while representation metrics conflate perceptual and semantic errors. We therefore introduce a BCI-aware framework in which four VLMs assess image pairs through structured questions, producing Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS). Their consensus is distilled into the BCI-Coherence Score (BCS), a compact evaluator achieving a T-PAS MAE of 0.079 (r = 0.700) and a T-SAS MAE of 0.082 (r = 0.850) on our data. Human validation shows highly reliable joint coherence judgments, with Cohen's kappa = 0.882 +/- 0.174 and Krippendorff's alpha = 0.882, supporting perceptual-semantic recoverability over generic visual similarity. Code and resources are available at https://sukt03.github.io/BCS/.

Comments27 pages, 3 figures, 13 tables. Accepted at the 5th International Workshop on Human Brain and Artificial Intelligence (HBAI 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑