arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-24 至 2026-02-24 共收录 107 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 22 篇

2511.06450 2026-02-24 cs.CV cs.LG 79%

Countering Multi-modal Representation Collapse through Rank-targeted Fusion

通过秩目标融合对抗多模态表示崩溃

Seulgi Kim, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出Rank-enhancing Token Fuser框架,通过提升有效秩对抗多模态表示崩溃,验证了深度与RGB融合的平衡性,并在动作预测任务中取得显著性能提升。

Comments Accepted in 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10652 2026-02-24 cs.CL 79%

ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images

ViTextVQA: 一个大规模视觉问答数据集和一种新的多模态特征融合方法用于图像中的越南语文本理解

Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

机构 * Faculty of Information Science and Engineering, University of Information Technology(信息科学与工程学院,信息科技大学) Vietnam National University(越南国家大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 ViTextVQA是一个大规模视觉问答数据集,提出了一种新的多模态特征融合方法,用于提升图像中越南语文本的理解能力。

Comments International Journal of Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18752 2026-02-24 cs.CV cs.GR 79%

Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement

多模态大模型中ID一致性优化:通过对齐、纠缠与解纠缠进行面部修复

Yuran Dong, Hang Dai, Mang Ye

机构 * National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 EditedID通过引入对齐、解纠缠和纠缠机制,提升多模态大模型中面部身份一致性的修复能力。

Comments ICLR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18863 2026-02-24 eess.IV cs.CV cs.LG cs.MM 73%

TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking

TIACam: 基于自增强的文本锚定不变特征学习用于抗相机零水印

Abdullah All Tanvir, Agnibh Dasgupta, Xin Zhong

机构 * Department of Computer Science University of Nebraska Omaha(计算机科学系 内布拉斯加大学奥马哈分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 TIACam通过文本锚定不变特征学习和自增强技术,实现抗相机重拍的零水印系统,提升特征稳定性和水印提取准确性。

Comments This paper is accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19367 2026-02-24 cs.AI cs.CV 62%

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

时间序列、视觉与语言:探讨对比表示空间中对齐的极限

Pratham Yashwante, Rose Yu

机构 * University of California San Diego, USA(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究探讨了时间序列、视觉和语言在对比表示空间中的对齐问题,发现模型规模越大对齐性越强,但对齐是不对称的,图像可作为中介,文本和视觉信息的密度影响对齐效果。

Comments 24 Figures, 12 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05992 2026-02-24 cs.CV cs.AI 62%

Exploring Partial Multi-Label Learning via Integrating Semantic Co-occurrence Knowledge

探索通过整合语义共现知识的半多标签学习

Xin Wu, Fei Teng, Yue Feng, Kaibo Shi, Zhuosheng Lin, Ji Zhang, James Wang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学) Engineering Research Center of Sustainable Urban Intelligent Transportation, Ministry of Education(可持续城市智能交通工程研究中心,教育部) School of Engineering, Swinburne University of Technology(工程学院,斯winburne大学) School of Electronic and Information Engineering, Wuyi University(电子与信息工程学院,五邑大学) College of Electrical Engineering, Sichuan University(电气工程学院,四川大学) School of Computer Science, Chengdu University(计算机科学学院,成都大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SCINet框架,通过整合语义共现知识,提升半多标签学习的准确性与效果。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19870 2026-02-24 cs.CV 57%

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

ApET:基于近似误差的令牌压缩用于高效的视觉语言模型

Qiankun Ma, Ziyao Zhang, Haofei Wang, Jie Chen, Zhen Song, Hairong Zheng

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Peng Cheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Harbin Institute of Technology(哈尔滨工业大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 ApET通过近似误差指导的令牌压缩,在不使用注意力机制的情况下高效压缩视觉语言模型的令牌预算,提升推理效率。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19615 2026-02-24 cs.CV 57%

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

清晰可见,自信推理:用于视觉语言模型盲点的即插即用修复方法

Xin Hu, Haomiao Ni, Yunbei Zhang, Jihun Hamm, Zechen Li, Zhengming Ding

机构 * Department of Computer Science, Tulane University(路易斯安那大学计算机科学系) Department of Computer Science, University of Memphis(密苏里大学计算机科学系)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种无需微调的即插即用模块,通过细化视觉标记和丰富文本提示,提升VLM对罕见物体的推理能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12047 2026-02-24 cs.CV 57%

PSGait: Gait Recognition using Parsing Skeleton

PSGait: 基于解析骨架的步态识别

Hangrui Xu, Zhengxian Wu, Chuanrui Zhang, Zhuohong Chen, Zhifang Liu, Peng Jiao, Haoqian Wang

机构 * The Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院) School of Computer Science and Information Engineering, Hefei University of Technology, China(合肥工业大学计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 PSGait通过解析骨架与轮廓融合的方法提升步态识别的准确性和泛化能力,实现15.7%的精度提升。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17195 2026-02-24 cs.RO 50%

Depth-PC: A Visual Servo Framework Integrated with Cross-Modality Fusion for Sim2Real Transfer

Depth-PC: 一种集成跨模态融合的视觉伺服框架用于仿真到现实迁移

Haoyu Zhang, Yang Liu, Yimu Jiang, Weiyang Lin, Chao Ye

机构 * Research Institute of Intelligent Control and Systems, Harbin Institute of Technology(智能控制与系统研究所,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 Depth-PC通过跨模态融合和图神经网络实现零样本仿真到现实迁移的视觉伺服框架

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 7 篇

2602.19188 2026-02-24 cs.CV 83%

PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration

PositionOCR: 通过混合专家整合增强多模态模型的位置意识

Chen Duan, Zhentao Guo, Pei Fu, Zining Wang, Kai Zhou, Pengfei Yan

专题命中 其他多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 PositionOCR通过整合文本定位专家与LLM的上下文能力,提升多模态模型在文本定位和定位任务中的位置准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18869 2026-02-24 cs.CV 70%

Enhancing 3D LiDAR Segmentation by Shaping Dense and Accurate 2D Semantic Predictions

通过塑造密集且准确的2D语义预测来增强3D激光雷达分割

Xiaoyu Dong, Tiankui Xian, Wanshui Gan, Naoto Yokoya

机构 * The University of Tokyo(东京大学)

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MM2D3D模型,通过多模态引导滤波和动态跨伪监督提升2D预测质量,从而增强3D激光雷达分割的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19319 2026-02-24 cs.MM cs.AI cs.CR cs.DB cs.DC 62%

Health+: Empowering Individuals via Unifying Health Data

Health+: 通过统一健康数据赋能个体

Sujaya Maiyya, Shantanu Sharma, Avinash Kumar

机构 * University of Waterloo(滑铁卢大学) New Jersey Institute of Technology(新泽西理工学院) Independent Researcher(独立研究者)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 Health+ 是一个以用户为中心的多模态健康数据管理系统,通过统一数据和智能推荐,提升个体对健康信息的控制与管理能力。

Comments This paper has been accepted in ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06292 2026-02-24 eess.IV cs.CV 57%

Zero-shot Multi-Contrast Brain MRI Registration by Intensity Randomizing T1-weighted MRI (LUMIR25)

无监督多对比脑MRI配准:通过强度随机化T1加权MRI(LUMIR25)

Hengjie Liu, Yimeng Dou, Di Xu, Xinyi Fu, Dan Ruan, Ke Sheng

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于强度随机化的多对比度脑MRI配准方法,通过多模态损失、强度随机化和轻量级优化策略,在无需显式图像合成的情况下实现了跨对比度的稳健泛化。

Comments Submitted to and reviewed by Learn2Reg MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20132 2026-02-24 cs.LG 50%

LAD: Learning Advantage Distribution for Reasoning

LAD: 为推理学习优势分布

Wendi Li, Sharon Li

专题命中 其他多模态 :multimodal(abstract)

AI总结 LAD通过学习优势分布提升大型模型推理的多样性和准确性,无需额外训练成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19667 2026-02-24 cond-mat.mtrl-sci cs.SE physics.data-an 50%

Dara: Automated multiple-hypothesis phase identification and refinement from powder X-ray diffraction

Dara:从粉末X射线衍射自动识别和优化多假设相

Yuxing Fei, Matthew J. McDermott, Christopher L. Rom, Shilong Wang, Gerbrand Ceder

专题命中 其他多模态 :multimodal(abstract)

AI总结 Dara通过自动化多假设相识别和优化,提升粉末X射线衍射分析的可靠性和准确性,推动全自动材料发现。

Journal ref Chem. Mater. 2026, 38, 3, 1364-1376

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10758 2026-02-24 stat.ML cs.LG stat.CO 50%

Stochastic Localization via Iterative Posterior Sampling

通过迭代后验采样实现随机定位

Louis Grenioux, Maxence Noble, Marylou Gabrié, Alain Oliviero Durmus

机构 * CMAP, CNRS, École polytechnique, Institut Polytechnique de Paris(CMAP、法国国家科学研究中心、巴黎高等学院、巴黎理工学院)

专题命中 其他多模态 :multi-modal(abstract)

AI总结 本文提出SLIPS方法,通过迭代后验采样实现随机定位,用于从无规范目标密度中采样,适用于多模分布的基准测试。

Comments Accepted at ICML 2024, improved assumption A0 (and consequences), fixed corollary 11

详情

展开后加载摘要…

URL PDF HTML 收藏