arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-10 至 2025-12-10 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 5 篇

2510.14885 2025-12-10 cs.CV cs.CL 84%

You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction

你可能可以自由发言:通过答案提取提升多模态大语言模型的细粒度视觉识别能力

Logan Lawrence, Oindrila Saha, Megan Wei, Chen Sun, Subhransu Maji, Grant Van Horn

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校) Brown University(布朗大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 本文提出nlg2choice方法,通过两阶段策略提升多模态大语言模型在细粒度视觉识别任务中的表现,通过开放性问题和受限解码提高检索效率。

Comments Accepted to WACV26. 12 pages, 8 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22930 2025-12-10 cs.CV 83%

Towards Explainable Bilingual Multimodal Misinformation Detection and Localization

迈向可解释的双语多模态虚假信息检测与定位

Yiwei He, Zhenglin Huang, Haiquan Wen, Tianxiao Li, Yi Dong, Hao Fei, Baoyuan Wu, Guangliang Cheng

机构 * University of Liverpool, United Kingdom(利物浦大学) National University of Singapore, Singapore(新加坡国立大学) The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 BiMi提出双语多模态框架,通过联合定位、一致性检测和自然语言解释,提升多语言虚假信息检测的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17537 2025-12-10 cs.AI cs.CL cs.CV 67%

CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale

CLIBD:融合视觉与基因组学用于大规模生物多样性监测

ZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum, Scott C. Lowe, Graham W. Taylor, Angel X. Chang

机构 * Simon Fraser University(西蒙弗雷泽大学) Aalborg University(奥胡斯大学) Vector Institute(向量研究所) University of Guelph(圭尔夫大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 CLIBD通过融合视觉与基因组学数据,利用对比学习实现对昆虫物种的高效分类,提升生物多样性监测的准确性。

Comments Add Variations of DNA encoding

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08673 2025-12-10 cs.CV 57%

Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds

双分支中心-周围对比:重新思考3D点云中的对比学习

Shaofeng Zhang, Xuanqi Chen, Xiangdong Zhang, Sitong Wu, Junchi Yan

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Chinese University of Hong Kong(香港中文大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出双分支中心-周围对比框架,通过双分支输入和补丁级对比损失,提升3D点云对比学习性能,达到SOTA效果。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08330 2025-12-10 cs.CV 57%

PointDico: Contrastive 3D Representation Learning Guided by Diffusion Models

PointDico: 通过扩散模型引导的对比3D表示学习

Pengbo Li, Yiding Sun, Haozhe Cheng

机构 * International School Beijing University of Posts(国际学校 北京邮电大学) School of Software Engineering Xi'an Jiaotong University Xi'an, China(软件工程学院 西安交通大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 PointDico通过融合扩散模型和对比学习的方法,实现了3D表示学习的突破,达到了ScanObjectNN和ShapeNetPart上的新高精度。

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏