arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-24 至 2025-10-24 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2510.20291 2025-10-24 cs.CV cs.AI 82%

A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization

LinFeng Li, Jian Zhao, Zepeng Yang, Yuhang Song, Bojun Lin, Tianle Zhang, Yuchen Yuan, Chi Zhang, Xuelong Li

机构 * The Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) East China Normal University(东华师范大学) National Tsing Hua University(国立清华大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Journal ref IROS 2025 Robosense Cross-Modal Drone Navigation Challenge first place

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15235 2025-10-24 cs.CV cs.CL 62%

ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding

Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, Xinghao Chen

机构 * Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22651 2025-10-24 cs.CV cs.CL cs.LG 62%

Sherlock: Self-Correcting Reasoning in Vision-Language Models

Yi Ding, Ruqi Zhang

机构 * Department of Computer Science, Purdue University, USA(计算机科学系,普渡大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Published at NeurIPS 2025, 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 62%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20696 2025-10-24 cs.CV 57%

Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward

Jing Bi, Guangyu Sun, Ali Vosoughi, Chen Chen, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21401 2025-10-24 cs.CV 57%

JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation

Md Jueal Mia, M. Hadi Amini

机构 * Knight Foundation School of Computing and Information Sciences (KFSCIS)(骑士基金会计算与信息科学学院) Florida International University(佛罗里达国际大学) Sustainability, Optimization, and Learning for InterDependent networks laboratory (solid lab)(可持续性、优化与互依赖网络学习实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20278 2025-10-24 cs.LG 50%

KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models

Guangyu Dai, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏