arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-22 至 2025-12-22 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 5 篇

2406.14862 2025-12-22 cs.LG cs.CL cs.CV 81%

LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models

LatentExplainer: 通过多模态大语言模型解释深度生成模型中的潜在表示

Mengdan Zhu, Raasikh Kanjiani, Jiahui Lu, Andrew Choi, Qirui Ye, Liang Zhao

机构 * Emory University(埃默里大学) University College London(伦敦大学学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.CV、cs.LG

AI总结 LatentExplainer通过多模态大语言模型为深度生成模型的潜在变量生成语义解释,提升模型可解释性。

Comments Accepted to CIKM 2025 Full Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17902 2025-12-22 cs.CV cs.AI cs.CR 73%

Adversarial Robustness of Vision in Open Foundation Models

开放基础模型中视觉的对抗鲁棒性

Jonathon Fox, William J Buchanan, Pavlos Papadopoulos

机构 * Blockpass ID Lab(Blockpass ID 实验室) Edinburgh Napier University(爱丁堡纳皮尔大学)

专题命中 幻觉与鲁棒性 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了LLaVA-1.5-13B和Llama 3.2 Vision-8B-2在视觉输入下的对抗鲁棒性,发现Llama 3.2 Vision在高扰动水平下性能下降更小,表明视觉模态是降级开放权重VLMs的可行攻击向量。

Journal ref IEEE Access, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15528 2025-12-22 cs.CV 70%

EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration

EmoCaliber: 通过置信度 verbalization 和校准推进可靠的视觉情绪理解

Daiqing Wu, Dongbao Yang, Can Ma, Yu Zhou

机构 * IIE, Chinese Academy of Sciences(中国科学院信息研究所) Nankai University(南开大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 EmoCaliber通过置信度 verbalization 和校准提升视觉情绪理解的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15068 2025-12-22 cs.LG cs.AI cs.CL 62%

The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems

语义幻觉:基于嵌入的幻觉检测在RAG系统中的认证极限

Debu Sinha

机构 * Independent Researcher(独立研究者)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI、cs.LG

AI总结 研究揭示了基于嵌入的幻觉检测在RAG系统中存在语义幻觉问题,通过符合预测方法发现真实幻觉检测的挑战,证明需通过推理而非表面语义解决。

Comments 12 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09814 2025-12-22 cs.CV 57%

On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness

CLIP纹理-形状偏倚的动态演变及其与人类对齐和模型鲁棒性的关系

Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Alexandra Gómez-Villa, Jorge Vila-Tomás, Valero Laparra, Jesus Malo

机构 * Image Processing Lab, Universitat de Valencia, Spain(瓦伦西亚大学图像处理实验室) Centro de Biomateriales e Ingenieria Tisular, Universitat Politecnica de Valencia, Spain(瓦伦西亚理工大学生物材料与组织工程中心) Computer Vision Center, Spain(西班牙计算机视觉中心) Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 本文研究CLIP模型在训练过程中纹理-形状偏倚的演变,揭示其与人类感知对齐及模型鲁棒性的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏