arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-26 至 2025-11-26 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 7 篇

2511.19834 2025-11-26 cs.CV 83%

Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation

基于多模态检索增强生成的大型语言模型辅助Birt-Hogg-Dube综合征诊断

Haoqing Li, Jun Shi, Xianmeng Chen, Qiwei Jia, Rui Wang, Wei Wei, Hong An, Xiaowen Hu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Department of Pulmonary and Critical Care Medicine(呼吸与危重症医学科) Center for Diagnosis and Management of Rare Diseases(罕见病诊断与管理中心) the First Affiliated Hospital of USTC(USTC第一附属医院) Division of Life Sciences and Medicine(生命科学与医学系) USTC WanNan Medical College(皖南医学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本研究提出BHD-RAG框架,通过整合多模态检索增强生成与领域专业知识,提升Birt-Hogg-Dube综合征的CT影像诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19509 2025-11-26 cs.LG 82%

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

TouchFormer: 一种基于变换器的鲁棒多模态材料感知框架

Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong, Xuexin Liu, Zhuojun Zou, Haoyue Yang, Lin Shu, Jie Hao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 TouchFormer通过模态自适应门控和注意力机制提升多模态材料感知的鲁棒性,并在细粒度分类任务中实现性能提升。

Comments 9 pages, 7 figures, Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20030 2025-11-26 cs.LG 78%

Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph Filtering

多模态属性图的跨对比聚类:带有双图过滤的多视图图聚类

Haoran Zheng, Renchi Yang, Hongtao Wang, Jianliang Xu

机构 * Hong Kong Baptist University(香港 Baptist 大学)

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文提出双图过滤方案,通过特征去噪和三交叉对比学习提升多模态属性图的聚类性能。

Comments Accepted by SIGKDD 2026. The code is available at https://github.com/HaoranZ99/DGF

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02025 2025-11-26 cs.SE 78%

SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash Data

SAFE: 借助LLM从多模态碰撞数据中驱动场景的ADS测试

Siwei Luo, Yang Zhang, Yao Deng, Linfeng Liang, Xi Zheng

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 SAFE通过多模态提取和LLM技术,提升从碰撞数据中重建真实场景和检测安全违规的能力,实现高准确率和高效生成。

Comments The paper has been accepted for publication in the proceedings of the IEEE/ACM 48th International Conference on Software Engineering to be held 12-18 April 2026 (ICSE2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19200 2025-11-26 cs.CV 57%

Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?

现代视觉模型能否理解物体与相似物之间的差异?

Itay Cohen, Ethan Fetaya, Amir Rosenfeld

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文研究了现代视觉模型能否区分真实物体与相似物,通过构建RoLA数据集并改进CLIP模型的嵌入空间方向,提升跨模态检索和描述生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11916 2025-11-26 cs.CV 57%

NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition

NeuroGaze-Distill:基于脑科学的蒸馏与抑郁启发的几何先验用于鲁棒面部情绪识别

Zilin Li, Weiwei Xu, Xuanqi Zhao, Yiran Zhu

机构 * School of Information and Intelligent Science(信息与智能科学学院) Department of Computer(计算机系) North China Electric Power University (BaoDing)(华北电力大学(保定))

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 NeuroGaze-Distill通过脑科学先验和抑郁启发的几何先验提升面部情绪识别的鲁棒性,采用跨模态蒸馏框架实现无需生物信号的部署。

Comments Preprint. Vision-only deployment; EEG used to form static prototypes. Includes appendix, 7 figures and 3 tables. Considering submission to ICLR 2026. Revision note: This version corrects inaccuracies in the authors' institutional affiliations. No technical content has been modified

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20227 2025-11-26 cs.IR 50%

HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents

HKRAG:面向视觉丰富文档的综合知识检索增强生成

Anyang Tong, Xiang Niu, ZhiPing Liu, Chang Tian, Yanyan Wei, Zenglin Shi, Meng Wang

专题命中 跨模态检索 :multimodal(abstract)

AI总结 HKRAG通过综合检索和生成机制,提升视觉丰富文档中显著与细节知识的检索与生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏