发表机构
Fuzhou University; Chinese Academy of Sciences; Peking University; Alibaba Group; University of Science and Technology of China; Fullive Innovation (Beijing) AI Technology Co., Ltd.; Baidu; Wuhan University(福州大学; 中国科学院; 北京大学; 阿里巴巴集团; 中国科学技术大学; 福莱创新(北京)人工智能科技有限公司; 百度; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DiffImaginE将多模态命名实体识别的类型验证建模为条件潜在扩散推理,在Twitter-2015和Twitter-2017数据集上,相比确定性对照模型取得了一致的性能提升。
AI 中文摘要
多模态命名实体识别(MNER)用于判断每个候选文本片段和实体类型假设是否得到文本与视觉证据的共同支持。现有的“想象-对比”验证器将每个(文本片段,类型)对映射为一个预测的视觉特征,将多样的视觉实现压缩为单个原型,提供兼容性评分但缺乏显式概率语义。本文提出DiffImaginE,将MNER类型验证建模为条件潜在扩散推理:给定文本片段定位的视觉证据,类型条件去噪器预测注入其标准化潜在空间的噪声,得到的去噪误差为类型条件负对数似然提供了ELBO一致的替代指标,可通过竞争类型假设对观测的解释能力进行排序。DiffImaginE保留标准多模态编码器栈,用无分类器引导扩散评分器替代确定性验证器,采用Min-SNR权重训练;直接将每个类型的扩散评分作为分类逻辑进行监督,学习跨噪声水平的聚合,并使用对偶采样降低蒙特卡洛对比方差。分析表明,无分类器引导可锐化诱导的类型后验,并确定对偶配对在同等去噪器成本下何时能降低方差。在Twitter-2015和Twitter-2017上的实验显示,在相同编码器、辅助目标和评估协议下,与匹配的确定性ImaginE对照模型相比,DiffImaginE取得了一致的性能提升,且得到了 ablation 实验和配对显著性检验的支持。
英文摘要
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.