arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16975cs.CLcs.LG

用于脑-语言对应关系的边际正则化结构化语义对齐

Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

  • School of Computer Science and Technology, Northwestern Polytechnical University(西北工业大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Jiaqi Wang, Huawen Hu, Shu Zhang

AI总结:

本文针对脑-语言解码的模糊性问题,提出MD-SigLIP框架,通过边际正则化结构化语义对齐实现基于检索的解码,在全词汇与子集评估下均取得最优检索性能。

AI中文摘要:

随着大型语言模型的快速发展,脑-语言解码已取得显著进展,但目前仍不清楚解码内容是真正反映神经表征,还是在很大程度上由语言模型自身重构,这种模糊性限制了解释性并阻碍了对内在脑-语言对应关系的研究。为应对这一挑战,本文提出MD-SigLIP,这一边际正则化结构化语义对齐框架在共享语义空间中直接对齐脑嵌入与文本嵌入,支持基于检索的解码,该公式可明确建模神经表征与语言语义之间的对应关系。基于感知重复的sigmoid对比学习,本文引入列表式边际正则化项,以在正语义簇与负样本之间施加结构化排序约束。通过同时建模多正语义结构与基于边际的排序,该方法可捕捉神经信号所反映的语言嵌入的流形结构。实验结果表明,在全词汇与子集评估设置下,该方法均实现了最优的检索性能。

英文摘要:

With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspondence. To address this challenge, we propose MD-SigLIP. This margin-regularized structured semantic alignment framework directly aligns brain embeddings with text embeddings in a shared semantic space, enabling retrieval-based decoding. This formulation enables explicit modeling of the correspondence between neural representations and language semantics. Building upon duplicate-aware sigmoid contrastive learning, we introduce a listwise margin-regularized term that enforces structured ranking constraints between positive semantic clusters and negative samples. By modeling multi-positive semantic structure and margin-based ordering simultaneously, the method captures the manifold organization of language embeddings reflected in neural signals. Experiments demonstrate state-of-the-art retrieval performance under both full-vocabulary and subset evaluation settings.

↑