arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-29 至 2025-10-29 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2510.24331 2025-10-29 cs.LG cs.CV 79%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24446 2025-10-29 cs.CL cs.CV 62%

SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space

Viktoriia Zinkovich, Anton Antonov, Andrei Spiridonov, Denis Shepelev, Andrey Moskalenko, Daria Pugacheva, Elena Tutubalina, Andrey Kuznetsov, Vlad Shakhuro

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24321 2025-10-29 cs.CV cs.AI 62%

Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning

Ivica Dimitrovski, Vlatko Spasev, Ivan Kitanovski

机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) University Ss Cyril and Methodius(西里尔与美多西乌斯大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05229 2025-10-29 cs.CV cs.MM 62%

Does CLIP perceive art the same way we do?

Andrea Asperti, Leonardo Dessì, Maria Chiara Tonetti, Nico Wu

机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24650 2025-10-29 cs.AI 57%

Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning

Nitin Rai, Daeun, Choi, Nathan S. Boyd, Arnold W. Schumann

机构 * Department of Horticultural Sciences(园艺科学系) Gulf Coast Research and Education Center(墨西哥湾沿岸研究与教育中心) University of Florida(佛罗里达大学) Department of Agricultural and Biological Engineering(农业与生物工程系) Department of Soil, Water, and Ecosystem Sciences(土壤、水与生态系统科学系) Citrus Research and Education Center(柑橘研究与教育中心)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 26 pages, 8 figures, and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07046 2025-10-29 cs.CV 57%

RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning

Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang

机构 * University of Chinese Academy of Science, Beijing,100190, China(中国科学院大学) Macquarie University(麦考瑞大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Published in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23894 2025-10-29 cs.CV 57%

Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation

Jinxin Zhou, Jiachen Jiang, Zhihui Zhu

机构 * Department of Computer Science(计算机科学系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 23 pages, 10 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15256 2025-10-29 cs.CV 57%

Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology Images

Jinsol Song, Jiamu Wang, Anh Tien Nguyen, Keunho Byeon, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * Korea University(韩国大学) The Catholic University of Korea(韩国天主大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa

详情

展开后加载摘要…

URL PDF HTML 收藏