arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2510.10194 2025-12-02 cs.CV 57%

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07062 2025-12-02 cs.AI 57%

Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision

通过噪声长描述监督提升城市影像区域表示学习

Yimei Zhang, Guojiang Shen, Kaili Ning, Tongwei Ren, Xuebo Qiu, Mengmeng Wang, Xiangjie Kong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出UrbanLN框架,通过长文本意识和噪声抑制提升城市影像区域表示学习,有效解决细粒度特征对齐与噪声干扰问题。

Comments Accepted as a full paper by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00293 2025-12-02 cs.LG cs.AI 57%

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

FiCoTS: 细到粗的LLM增强层次跨模态交互用于时间序列预测

Yafei Lyu, Hao Zhou, Lu Zhang, Xu Yang, Zhiyong Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所MAIS) Great Bay University(大亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 FiCoTS通过细到粗的LLM增强层次跨模态交互框架,提升多模态时间序列预测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22172 2025-12-01 cs.CV 57%

Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning

引导内在之眼:一种用于层次化和灵活的视觉基础推理的框架

Zhaoyang Wei, Wenchao Ding, Yanchao Hao, Xi Chen

机构 * Basic Algorithm Center, PCG, Tencent(腾讯基本算法中心、PCG、腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 GRiP通过认知增强的强化学习框架,提升视觉基础推理的鲁棒性和灵活性,实现复杂场景下的高性能表现。

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22107 2025-12-01 cs.CV 57%

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

HyperST:面向空间转录组预测的分层双曲学习

Chen Zhang, Yilu An, Ying Chen, Hao Li, Xitong Ling, Lihao Liu, Junjun He, Yuxiang Lin, Zihui Wang, Rongshan Yu

机构 * Xiamen University(厦门大学) Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 HyperST通过在双曲空间中建模数据的层次结构,实现空间转录组预测的多级图像-基因表示学习,提升跨模态预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21741 2025-12-01 cs.CL cs.LG stat.ML 57%

A Multiscale Geometric Method for Capturing Relational Topic Alignment

一种多尺度几何方法用于捕捉关系主题对齐

Conrad D. Hougen, Karl T. Pazdernik, Alfred O. Hero

机构 * University of Michigan(密歇根大学) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出一种多尺度几何方法,结合多模态文本和合著者网络数据,以捕捉关系主题对齐并识别罕见话题结构。

Comments 5 pages, 3 figures, 2025 IEEE International Workshop on Computational Advances in Multi-Sensor Adaptive Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21389 2025-11-27 cs.IR cs.AI 57%

FITRep: Attention-Guided Item Representation via MLLMs

FITRep: 通过大语言模型实现的注意力引导的项目表示

Guoxiao Zhang, Ao Li, Tan Qu, Qianlong Xie, Xingxing Wang

机构 * Meituan(美团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 FITRep通过引入注意力引导的白盒表示框架,利用多模态大语言模型实现细粒度项目去重,提升了广告点击率和每千次展示成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14264 2025-11-27 cs.CV 57%

Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models

Directed-Tokens: 一种稳健的多模态对齐方法用于大语言-视觉模型

Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj, Khoa Luu

机构 * CVIU Lab, University of Arkansas, USA(美国亚拉巴马大学CVIU实验室) Vietnam National University, Ho Chi Minh City University of Science, Vietnam(越南国家大学胡志明市科技大学) Carnegie Mellon University, USA(美国卡内基梅隆大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Directed-Tokens方法,通过引入图像和文本顺序重建任务及新的损失函数,提升大语言-视觉模型的鲁棒性和跨模态对齐能力,在多个基准上取得最佳性能。

Comments Accepted to NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18865 2025-11-25 cs.CV 57%

DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection

DualGazeNet: 一种生物启发的双目注视查询网络用于显著物体检测

Yu Zhang, Haoan Ping, Yuchen Li, Zhenshan Bing, Fuchun Sun, Alois Knoll

机构 * School of Computation, Information and Technology, Technical University of Munich(计算信息与技术学院,慕尼黑技术大学) Department of Informatics, Technical University of Munich(信息学院,慕尼黑技术大学) University Research and Innovation Center, Obuda University(奥布达大学研究与创新中心) State Key Laboratory for Novel Software Technology, Nanjing University Suzhou Campus(新型软件技术国家重点实验室,南京大学苏州校区) School of Science and Technology, Nanjing University Suzhou Campus(科学与技术学院,南京大学苏州校区) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 DualGazeNet通过生物启发的纯Transformer框架,实现了显著物体检测的高效准确性和计算效率,超越了多种现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18028 2025-11-25 cs.CV 57%

MambaX: Image Super-Resolution with State Predictive Control

MambaX:基于状态预测控制的图像超分辨率

Chenyu Li, Danfeng Hong, Bing Zhang, Zhaojie Pan, Naoto Yokoya, Jocelyn Chanussot

机构 * Southeast University(东南大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) College of Resources and Environment, University of Chinese Academy of Sciences(中国科学院大学资源与环境学院) School of Mathematics, Southeast University(东南大学数学学院) Department of Complexity Science and Engineering, Graduate School of Frontier Sciences, the University of Tokyo(东京大学前沿科学研究院复杂科学与工程部门) Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、Inria、CNRS、Grenoble INP、LJK)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 MambaX通过动态状态预测控制和多模态融合方法提升图像超分辨率性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17967 2025-11-25 cs.CV 57%

CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking

CADTrack: 基于可变形对齐的上下文聚合学习用于鲁棒的RGBT跟踪

Hao Li, Yuhao Wang, Xiantao Hu, Wenning Hao, Pingping Zhang, Dong Wang, Huchuan Lu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CADTrack通过上下文聚合与可变形对齐框架,提升RGBT跟踪的鲁棒性和准确性。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17668 2025-11-25 cs.CV 57%

MedPEFT-CL: Dual-Phase Parameter-Efficient Continual Learning with Medical Semantic Adapter and Bidirectional Memory Consolidation

MedPEFT-CL:基于医学语义适配器和双向记忆巩固的双阶段参数高效持续学习

Ziyuan Gao

机构 * University College London(伦敦大学学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 MedPEFT-CL通过语义驱动适配器和双向记忆协调,实现医学视觉-语言任务中的参数高效持续学习,有效缓解灾难性遗忘。

Comments Accepted by WACV 2026 (round 2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17358 2025-11-24 cs.CL 57%

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

不要学习,而是依托:自然语言推理与视觉依托的案例

Daniil Ignatev, Ayman Santeer, Albert Gatt, Denis Paperno

机构 * Utrecht University(乌特勒支大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15967 2025-11-21 cs.CV 57%

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

InfoCLIP: 通过信息论对齐转移连接视觉语言预训练与开放词汇语义分割

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang, Lan Ma, Yuan Gao, Jiangyong Ying, Yudeng Xin

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

AI总结 InfoCLIP通过信息论对齐转移提升开放词汇语义分割的性能,有效解决预训练CLIP在微调过程中的过拟合问题。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15669 2025-11-21 cs.CV 57%

UINO-FSS: Unifying Representation Learning and Few-shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelation

UINO-FSS: 通过层次化蒸馏和Mamba-超相关性统一表示学习与少样本分割

Wei Zhuo, Zhiyue Tang, Wufeng Xue, Hao Ding, Junkai Ji, Linlin Shen

机构 * School of Artificial Intelligence and the National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(人工智能学院和大数据系统计算技术国家工程实验室,深圳大学) Guangdong Provincial Key Laboratory of Intelligent Information Processing, China(广东省智能信息处理重点实验室,中国) School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University(生物医学工程学院,深圳大学医学院,深圳大学) Department of Computer Science, University of Nottingham Ningbo China(计算机科学系,宁波大学中国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 UINO-FSS通过层次化蒸馏和Mamba-超相关性整合不同基础模型知识,实现少样本分割的统一学习框架,取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07033 2025-11-20 cs.CV 57%

One Latent Space to Rule All Degradations: Unifying Restoration Knowledge for Image Fusion

Haolong Ma, Hui Li, Chunyang Cheng, Zeyang Zhang, Xiaoqing Luo, Xiaoning Song, Xiao-Jun Wu

机构 * Jiangnan University(江南大学) Suzhou University of Science and Technology(苏州科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15016 2025-11-20 cs.CV 57%

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

Zhenyu Cui, Jiahuan Zhou, Yuxin Peng

机构 * Zhenyu Cui, Jiahuan Zhou, Yuxin Peng(作者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05923 2025-11-20 cs.CV 57%

Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation

Qiming Li, Zekai Ye, Xiaocheng Feng, Weihong Zhong, Weitao Ma, Xiachong Feng

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 57%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18757 2025-11-20 cs.CV 57%

ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * CCDS, NTU, Singapore(南洋理工大学新加坡分校) CCST, ZJUT, China(浙江工业大学计算机科学与技术学院) Terminus AI Lab, UCAS, China(中国科学院大学人工智能实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14698 2025-11-19 cs.CV cs.LG eess.SP 57%

HyMAD: A Hybrid Multi-Activity Detection Approach for Border Surveillance and Monitoring

Sriram Srinivasan, Srinivasan Aruchamy, Siva Ram Krisha Vadali

机构 * Sriram Srinivasan(独立研究者) Srinivasan Aruchamy(独立研究者) Siva Ram Krishna Vadali(独立研究者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Multi-label seismic signal classification using novel attention-based feature fusion. Submitting to cs.CV due to relevance to general pattern recognition and time-frequency (spectrogram) analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14064 2025-11-19 cs.LG cs.AI stat.ME 57%

CafeMed: Causal Attention Fusion Enhanced Medication Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

机构 * Department of Computer Science and Engineering, Hanyang University(韩阳大学计算机科学与工程系) Department of Applied Artificial Intelligence, Hanyang University(韩阳大学应用人工智能系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments Accepted by BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13150 2025-11-18 cs.CV 57%

Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification

Rifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13047 2025-11-18 cs.CV cs.RO 57%

DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation

Yan Gong, Jianli Lu, Yongsheng Gao, Jie Zhao, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学) Institute for Infocomm Research, A*STAR(信息通信研究机构,A*STAR) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12525 2025-11-18 cs.CV 57%

MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics

Jing Li, Yifan Wang, Jiafeng Yan, Renlong Zhang, Bin Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 7 figures. Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12368 2025-11-18 cs.CV 57%

Fast Reasoning Segmentation for Images and Videos

Yiqing Shen, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰·霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12365 2025-11-18 cs.CV 57%

Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning

Yiqing Shen, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12220 2025-11-18 cs.CV cs.LG 57%

Suppressing VLM Hallucinations with Spectral Representation Filtering

Ameen Ali, Tamim Zoabi, Lior Wolf

机构 * Tel Aviv University(特拉维夫大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06298 2025-11-18 cs.CV 57%

SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection

Xin Zuo, Chenyu Qu, Haibo Zhan, Jifeng Shen, Wankou Yang

机构 * School of Computer, Jiangsu University of Science and Technology(江苏科技大学计算机学院) School of Electrical and Information Engineering, Jiangsu University(江苏大学电气与信息工程学院) School of Automation, Southeast University(东南大学自动化学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 11 pages,8 figures, accepted by IEEE TGRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11750 2025-11-18 cs.LG cs.AI 57%

IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation

Hanting Yan, Pan Mu, Shiqi Zhang, Yuchao Zhu, Jinglin Zhang, Cong Bai

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏