arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2508.11452 2025-09-03 cs.AI cs.CL cs.HC 62%

Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Kangyu Wang, Hongliang He, Lin Liu, Ruiqi Liang, Zhenzhong Lan, Jianguo Li

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Westlake University(西湖大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Our platform is publicly accessible at https://www.tbox.cn/about/model-ranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09040 2025-08-26 cs.RO cs.AI cs.CV cs.LG 62%

RT-Cache: Training-Free Retrieval for Real-Time Manipulation

Owen Kwon, Abraham George, Alison Bartsch, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures. 2025 IEEE-RAS 24th International Conference on Humanoid Robots

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09545 2025-08-26 cs.CL cs.AI cs.CY 62%

Does GPT-4 surpass human performance in linguistic pragmatics?

Ljubisa Bojic, Predrag Kovacevic, Milan Cabarkapa

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 62%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00589 2025-08-13 cs.CV cs.CL cs.IR cs.RO 62%

Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving

Stefan Englmeier, Max A. Büttner, Katharina Winter, Fabian B. Flohr

机构 * Munich University of Applied Sciences(慕尼黑应用科学大学) Intelligent Vehicles Lab (IVL)(智能车辆实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://iv.ee.hm.edu/contextmotionclip/; This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03644 2025-08-06 cs.CL cs.CV cs.IR 62%

Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?

Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen, Junjie Yang, Yao Wan, Weiwei Lin

机构 * South China University of Technology(华南理工大学) Huazhong University of Science and Technology(华中科技大学) University of Maryland(马里兰大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments In submission. Project website: https://double-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03262 2025-08-06 cs.CL cs.AI 62%

Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?

Junhyuk Choi, Hyeonchu Park, Haemin Lee, Hyebeen Shin, Hyun Joung Jin, Bugeun Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20110 2025-07-29 cs.CV cs.AI cs.LG 62%

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Shiyu Liu, Lianlei Shan

机构 * School of Electrical and Electronic Engineering(电气与电子工程学院) Nanyang Technological University(南洋理工大学) School of Computer Science and Technology(计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments **14 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15639 2025-07-22 cs.CL cs.AI cs.LG 62%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 跨模态检索 :MLLM(abstract);分类 cs.CL、cs.AI

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12425 2025-07-17 cs.CL cs.AI cs.CE cs.IR 62%

Advancing Retrieval-Augmented Generation for Structured Enterprise and Internal Data

Chandana Cheerla

机构 * IIT Roorkee(罗尔基大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09482 2025-07-15 cs.CL cs.AI cs.HC 62%

ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning

Changli Wang, Rui Wu, Fang Yin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07748 2025-07-11 cs.CL cs.AI 62%

When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

Peizhang Shao, Linrui Xu, Jinxi Wang, Wei Zhou, Xingyu Wu

机构 * School of Law, China University of Political Science and Law(中国政法大学法学院) Zhejiang University of Finance and Economics Dongfang College(浙江财经大学东洋学院) School of Information Management for Law, China University of Political Science and Law(中国政法大学信息管理法学院) Department of Artificial Intelligence, Chung-Ang University(成均馆大学人工智能系) Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University(香港理工大学数据科学与人工智能系)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15379 2025-07-11 cs.IR cs.AI cs.CV 62%

Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval

Zijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard Mccreadie, Paul Henderson

机构 * Hunan University(湖南大学) University of Glasgow(格拉斯哥大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05513 2025-07-09 cs.CV cs.AI 62%

Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model

Mengyao Xu, Gabriel Moreira, Ronay Ak, Radek Osmulski, Yauhen Babakhin, Zhiding Yu, Benedikt Schifferer, Even Oldridge

机构 * Proceedings of Conference XXX(会议论文集)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20964 2025-07-09 cs.CV cs.AI 62%

Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering

Zhengxuan Zhang, Yin Wu, Yuyu Luo, Nan Tang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19834 2025-06-18 cs.LG cs.CV cs.MM 62%

Knowledge Bridger: Towards Training-free Missing Modality Completion

Guanzhou Ke, Shengfeng He, Xiao Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Yi Xie, HeXing Su

机构 * Beijing Jiaotong University(北京交通大学) Singapore Management University(新加坡国立大学) Southeast University(东南大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Harbin Institute of Technology(哈尔滨工业大学) Nanjing University of Science and Technology(南京理工大学) South China University of Technology(华南理工大学) Xiamen Institute of Technology(厦门理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07785 2025-06-10 cs.CV cs.AI cs.LG 62%

Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger

Qi Yang, Chenghao Zhang, Lubin Fan, Kun Ding, Jieping Ye, Shiming Xiang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所MAIS部) Alibaba Cloud Computing, China(阿里巴巴云计算)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICML 2025 Spotlight. 22 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16809 2025-06-10 cs.CV cs.MM 62%

Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities

Junze Wang, Lei Fan, Weipeng Jing, Donglin Di, Yang Song, Sidong Liu, Cong Cong

机构 * College of Computer and Control Engineering, Northeast Forestry University(东北林业大学计算机与控制工程学院) The Centre for Healthy Brain Ageing (CHeBA), UNSW(健康脑年龄中心(CHeBA),UNSW) School of Computer Science and Engineering, UNSW(计算机科学与工程学院,UNSW) School of Software, Tsinghua University(清华大学软件学院) Centre for Health Informatics, Macquarie University(健康信息学中心,麦考瑞大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments MICCAI 2025 Early Accept. The code is available at https://github.com/reeive/ReHyDIL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06208 2025-06-09 cs.CL cs.AI 62%

Building Models of Neurological Language

Henry Watkins

机构 * UCL Queen Square Institute of Neurology(伦敦大学学院女王广场神经病学研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23809 2025-06-04 cs.CL cs.AI cs.IR 62%

LLM-Driven E-Commerce Marketing Content Optimization: Balancing Creativity and Conversion

Haowei Yang, Haotian Lyu, Tianle Zhang, Dingzhou Wang, Yushang Zhao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01388 2025-06-03 cs.CV cs.AI 62%

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

Yihao Ding, Soyeon Caren Han, Yan Li, Josiah Poon

机构 * The University of Melbourne(墨尔本大学) The University of Sydney(悉尼大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at IJCAI 2025 Demonstrations Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23045 2025-05-30 cs.CV cs.AI 62%

Multi-Sourced Compositional Generalization in Visual Question Answering

Chuanhao Li, Wenbo Ye, Zhen Li, Yuwei Wu, Yunde Jia

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07688 2025-05-29 cs.CV cs.AI 62%

ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG

Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Zian Guan, Bin Chen, Yuhao Wang, Xu Jia, Yuxiang Cai, Yongheng Shang, Jianwei Yin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究院) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) Polytechnic Institute of Zhejiang University(浙江大学 polytechnic 院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15876 2025-05-29 cs.CV cs.AI cs.LG 62%

End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding

Kwanyoung Kim, Yujin Oh, Sangjoon Park, Hwa Kyung Byun, Joongyo Lee, Jin Sung Kim, Yong Bae Kim, Jong Chul Ye

机构 * Samsung Research(三星研究院) Center for Advanced Medical Computing and Analysis (CAMCA)(先进医学计算与分析中心) Massachusetts General Hospital (MGH) and Harvard Medical School(麻省总医院和哈佛医学院) Yonsei University College of Medicine(延世大学医学院) Yonsei University(延世大学) Institute for Innovation in Digital Healthcare(数字医疗创新研究所) Yongin Severance Hospital(Yongin Severance医院) Gachon University Gil Hospital(高仁大学Gil医院) Oncosoft Inc(Oncosoft公司) Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)(人工智能研究生院,韩国科学技术院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for Medical Image Analysis 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17796 2025-05-26 cs.CV cs.AI cs.IR 62%

DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval

Yuxin Yang, Yinan Zhou, Yuxin Chen, Ziqi Zhang, Zongyang Ma, Chunfeng Yuan, Bing Li, Lin Song, Jun Gao, Peng Li, Weiming Hu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17085 2025-05-26 cs.CR cs.AI cs.CL 62%

GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis

Kaibo Huang, Zipei Zhang, Yukun Wei, TianXin Zhang, Zhongliang Yang, Linna Zhou

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing IntokenTech Co., Ltd.(北京IntokenTech公司)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17058 2025-05-26 cs.CL cs.AI 62%

DO-RAG: A Domain-Specific QA Framework Using Knowledge Graph-Enhanced Retrieval-Augmented Generation

David Osei Opoku, Ming Sheng, Yong Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Beijing National Research Center for Information Science and Technology - Tsinghua University(北京信息科学与技术国家研究中心-清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 5 figures;

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16193 2025-05-23 cs.CL cs.CV 62%

An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability

Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01222 2025-05-23 cs.CV cs.CL 62%

Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

Wenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang, Li Shen, Yong Luo, Bo Du, Dacheng Tao

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09243 2025-05-19 cs.RO cs.AI cs.CV 62%

GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation

Ruihai Wu, Ziyu Zhu, Yuran Wang, Yue Chen, Jiarui Wang, Hao Dong

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏