arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-12 至 2026-02-12 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 13 篇

2509.14671 2026-02-12 cs.CL cs.AI cs.LG 86%

TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding

TableDART: 表格理解的动态自适应多模态路由

Xiaobo Xing, Wei Yuan, Tong Chen, Quoc Viet Hung Nguyen, Xiangliang Zhang, Hongzhi Yin

机构 * The University of Queensland, Australia(昆士兰大学) Griffith University, Australia(格里菲斯大学) University of Notre Dame, USA(诺丁汉大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);cross-modal(abstract)

AI总结 TableDART通过动态选择文本、图像或融合视角,提升表格理解的准确性和效率,避免昂贵的多模态模型微调。

Comments Accepted to ICLR 2026. 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10138 2026-02-12 cs.CV cs.AI cs.CL cs.LG 85%

Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement

多模态信息融合用于图表理解:MLLMs的综述——演变、局限与认知增强

Zhihang Yi, Jian Zhao, Jiancheng Lv, Tao Wang

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Engineering Research Center of Machine Learning(机器学习工程研究中心) China Telecom Institute of AI(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了多模态大语言模型在图表理解中的应用,分析了其演变、局限及未来发展方向,旨在推动更稳健的系统发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22463 2026-02-12 cs.MM 83%

Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation

正交解缠与投影特征对齐用于对话中多模态情绪识别

Xinyi Che, Wenbo Wang, Jian Guan, Qijun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出OD-PFA框架,通过正交解缠与投影特征对齐技术提升对话中多模态情绪识别性能。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02172 2026-02-12 cs.CV cs.AI cs.MM 82%

GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting

GaussianCross: 通过高斯点撒技术实现跨模态自监督3D表示学习

Lei Yao, Yi Wang, Yi Zhang, Moyun Liu, Lap-Pui Chau

机构 * Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 GaussianCross通过高斯点撒技术实现跨模态自监督3D表示学习,提升3D点云的表示质量和泛化能力。

Comments 14 pages, 8 figures, accepted by MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10161 2026-02-12 cs.CR cs.AI cs.CL 79%

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐

Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出OmniSteer方法,通过提取黄金拒绝向量和轻量级适配器提升多模态模型的安全性与通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12768 2026-02-12 cs.CV cs.AI cs.LG 73%

CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence

CoRe3D:协作推理作为3D智能的基础

Tianjiao Yu, Xinzhuo Li, Yifan Shen, Yuanzhe Liu, Ismini Lourentzou

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 CoRe3D通过协作推理框架实现3D内容生成,结合语义和空间推理提升3D输出的一致性与描述一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 70%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09651 2026-02-12 cs.CV cs.AI cs.LG 62%

Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era

地理空间表示学习:从深度学习到大语言模型时代的一次综述

Xixuan Hao, Yutian Jiang, Xingchen Zou, Jiabo Liu, Yifang Yin, Song Gao, Flora Salim, Tianrui Li, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The University of New South Wales(新南威尔士大学) Southwest Jiaotong University(西南交通大学) Institute for Infocomm Research (I$^2$R), A*STAR(信息通信研究院(I$^2$R),A*STAR)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了从深度学习到大语言模型时代的地理空间表示学习,探讨了其方法、应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10845 2026-02-12 cs.AI cs.LG 57%

SynergyKGC: Reconciling Topological Heterogeneity in Knowledge Graph Completion via Topology-Aware Synergy

SynergyKGC: 通过拓扑感知协同解决知识图谱补全中的拓扑异质性

Xuecheng Zou, Yu Tang, Bingbing Wang

机构 * School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) School of Mathematical Sciences, Soochow University(数学科学学院,苏州大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 SynergyKGC通过拓扑感知协同解决知识图谱补全中的拓扑异质性问题,提升KGC命中率。

Comments 10 pages, 5 tables, 7 figures. This work introduces the Active Synergy mechanism and Identity Anchoring for Knowledge Graph Completion. Code: https://github.com/XuechengZou-2001/SynergyKGC-main

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10740 2026-02-12 cs.CL 57%

Reinforced Curriculum Pre-Alignment for Domain-Adaptive VLMs

强化课程预对齐用于领域自适应视觉-语言模型

Yuming Yan, Shuo Yang, Kai Tang, Sihong Chen, Yang Zhang, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu, Edith C. H. Ngai

机构 * Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出RCPA方法,通过课程意识的渐进调节机制,在领域自适应中平衡领域知识获取与通用能力保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10494 2026-02-12 cs.CL 57%

Canvas-of-Thought: Grounding Reasoning via Mutable Structured States

Canvas-of-Thought:通过可变的结构化状态进行推理

Lingzhuang Sun, Yuxia Zhu, Ruitong Liu, Hao Liang, Zheng Sun, Caijun Jia, Honghao He, Yuchen Wu, Siyuan Li, Jingxuan Wei, Xiangxiang Zhang, Bihui Yu, Wentao Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) New York University(纽约大学) Westlake University(西湖大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 Canvas-CoT通过引入HTML Canvas实现可变结构化状态,提升多模态推理效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10053 2026-02-12 cs.CV 57%

DiCo: Disentangled Concept Representation for Text-to-image Person Re-identification

DiCo: 用于文本到图像人物重识别的解耦概念表示

Giyeol Kim, Chanho Eom

机构 * organization= Department of Imaging Science, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea organization= Department of Metaverse Convergence, Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University , addressline= , city= Seoul , postcode= 06974 , state= , country= South Korea

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 DiCo通过解耦概念表示方法,提升文本到图像人物重识别的跨模态对齐和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15147 2026-02-12 cs.CV 57%

From Pixels to Images: A Structural Survey of Deep Learning Paradigms in Remote Sensing Image Semantic Segmentation

从像素到图像:深度学习在遥感图像语义分割中的结构调查

Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang

机构 * College of Science and Engineering and Centre for AI and Data Science Innovation, James Cook University(科学与工程学院和人工智能与数据科学创新中心,詹姆斯库克大学) Department of Forest and Wildlife Ecology, University of Wisconsin-Madison(森林与野生动物生态学系,威斯康星大学麦迪逊分校) School of Computing, Engineering and Mathematical Sciences, La Trobe University(计算、工程与数学科学学院,拉特罗布大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文系统回顾了深度学习在遥感图像语义分割中的结构演变,从像素到图像的层次化方法,涵盖多种技术并提供可复现的代码库。

Comments 34 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏