arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-27 至 2026-01-27 共收录 116 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 24 篇

2510.05840 2026-01-27 cs.LG 78%

Multimodal Trajectory Representation Learning for Travel Time Estimation

多模态轨迹表示学习用于旅行时间估计

Zhi Liu, Xuyuan Hu, Xiao Han, Zhehao Dai, Zhaolin Deng, Guojiang Shen, Xiangjie Kong

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University of Technology College of Computer Science(浙江工业大学计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出多模态动态轨迹整合框架,通过整合GPS、网格轨迹和道路网络约束,提升旅行时间估计的性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12869 2026-01-27 cs.LG cs.AI cs.DC cs.IT cs.MA math.IT 70%

On the Fundamental Limits of LLMs at Scale

在大规模下的大语言模型根本限制

Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal, Zeeshan Memon, Muhammad Ibtsaam Qadir, Sagnik Bhattacharya, Hassan Rizwan, Abhiram R. Gorle, Maahe Zehra Kazmi, Nukhba Amir, Ali Subhan, Muhammad Usman Rafique, Zihao He, Pulkit Mehta, Muhammad Ali Jamshed, John M. Cioffi

机构 * Stanford University(斯坦福大学) The University of Oklahoma(俄克拉荷马大学) Emory University(埃默里大学) Purdue University(普渡大学) UC Riverside(加州大学河滨分校) UC Berkeley(加州大学伯克利分校) Khyber Medical University(克希伯医学大学) Universtat Pompeu Fabra(庞培法华大学) Zoox(Zoox公司) Meta Google DeepMind(谷歌DeepMind) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Glasgow(格拉斯哥大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文探讨了大规模大语言模型的根本限制,提出统一框架分析计算、信息和学习的基础限制,并提供缓解方法。

Comments Submitted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09828 2026-01-27 cs.CV cs.LG cs.RO 70%

DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception

DGFusion:基于深度的传感器融合用于鲁棒的语义感知

Tim Broedermannn, Christos Sakaridis, Luigi Piccinelli, Wim Abbeloos, Luc Van Gool

机构 * Computer Vision Laboratory, ETH Zurich(计算机视觉实验室,苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 DGFusion通过整合深度信息提升多模态传感器融合,实现自动驾驶中的鲁棒语义感知。

Comments Code and models are available at https://github.com/timbroed/DGFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23118 2026-01-27 cs.CV 70%

Quantizing Space and Time: Fusing Time Series and Images for Earth Observation

量化空间与时间:融合时间序列和图像用于地球观测

Gianfranco Basile, Johannes Jakubik, Benedikt Blumenstiel, Thomas Brunschwiler, Juan Bernabe Moreno

机构 * IBM Research Europe(IBM欧洲研究院) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种任务无关的多模态融合框架,通过时间序列和图像的统一表示空间提升地球观测任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17073 2026-01-27 cs.LG cs.CV stat.ML 70%

Attention-Based Variational Framework for Joint and Individual Components Learning with Applications in Brain Network Analysis

基于注意力的变分框架用于联合和个体成分学习及其在脑网络分析中的应用

Yifei Zhang, Meimei Liu, Zhengwu Zhang

机构 * Department of Biostatistics, Yale School of Public Health(生物统计学系,耶鲁公共卫生学院) Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工大学) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(统计学与运筹学系,北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CM-JIVNet,一种基于注意力的变分框架,用于联合和个体成分学习,以提升脑网络分析中的跨模态重建和行为预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18625 2026-01-27 cs.CV 57%

CONQUER: Context-Aware Representation with Query Enhancement for Text-Based Person Search

基于查询增强的上下文感知表示方法用于基于文本的人检索

Zequn Xie

机构 * Zhejiang University, Hangzhou, China(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CONQUER通过增强跨模态对齐和自适应查询细化,提升基于文本的人检索的准确性和鲁棒性。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18065 2026-01-27 cs.CL 57%

Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models

grounded concreteness: 人类-like 的 concreteness 敏感性在 vision-language 模型中

Aryan Roy, Zekun Wang, Christopher J. MacLellan

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文研究了视觉语言模型在纯文本提示下对concreteness的敏感性,并发现其在更具体的输入上表现更优,具有更清晰的表示和更符合人类规范的判断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11609 2026-01-27 stat.ML cs.AI cs.LG stat.ME 57%

Towards Interpretable Deep Generative Models via Causal Representation Learning

通过因果表示学习实现可解释的深度生成模型

Gemma E. Moran, Bryon Aragam

机构 * Department of Statistics, Rutgers University(统计学系,罗格斯大学) Booth School of Business, University of Chicago(商学院,芝加哥大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出通过因果表示学习实现可解释的深度生成模型,结合潜变量模型、因果图模型和非参数统计,旨在提升生成模型的可解释性和迁移能力。

Comments Accepted in Journal of the American Statistical Association: Special Issue on AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17342 2026-01-27 cs.CV 57%

STARS: Shared-specific Translation and Alignment for missing-modality Remote Sensing Semantic Segmentation

STARS: 共享-特定翻译与对齐用于缺失模态遥感语义分割

Tong Wang, Xiaodong Zhang, Guanzhou Chen, Jiaqi Wang, Chenxi Liu, Xiaoliang Tan, Wenchao Guo, Xuyang Li, Xuanrui Wang, Zifan Wang

机构 * State Key Laboratory of information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学) Electronic Information School, Wuhan University(电子信息学院,武汉大学) Hubei FreerTech Co. Ltd(湖北弗瑞特科技有限公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 STARS通过共享-特定翻译与对齐机制,解决多模态遥感中缺失模态带来的语义分割问题,提升鲁棒性和类别识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13390 2026-01-27 cs.CV 57%

Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment

通过大模型感知的语义蒸馏与对齐实现WiFi手势识别的泛化

Feng-Qi Cui, Yu-Tong Guo, Tianyue Zheng, Jinyang Huang

机构 * Institute of Advanced Technology, University of Science and Technology of China(科学技术大学先进技术研究院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 GLSDA通过大模型感知的语义蒸馏与对齐提升WiFi手势识别的泛化能力,实现领域内和跨领域任务的高性能识别。

Comments Accepted by IEEE ICPADS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16414 2026-01-27 q-bio.NC cs.CV eess.IV 57%

NeuroKoop: Neural Koopman Fusion of Structural-Functional Connectomes for Identifying Prenatal Drug Exposure in Adolescents

NeuroKoop:神经Koopman融合结构-功能连接组用于识别青少年孕期药物暴露

Badhan Mazumder, Aline Kotoski, Vince D. Calhoun, Dong Hye Ye

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) Neuroscience Institute, Georgia State University(神经科学研究所,佐治亚州立大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS)(跨机构神经影像与数据科学转化研究中心(TReNDS)) Georgia State University, Georgia Institute of Technology, and Emory University(佐治亚州立大学、佐治亚理工学院和埃默里大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 NeuroKoop通过神经Koopman算子融合结构-功能连接组,提升青少年孕期药物暴露识别的准确性和鲁棒性。

Comments Published in the Proceedings of the 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI). IEEE Xplore. DOI: 10.1109/BHI67747.2025.11269557

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17089 2026-01-27 cs.CV 57%

GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing

GRASP: 基于区域感知的稀疏提示引导方法用于适应遥感图像的多模态大语言模型

Qigan Sun, Chaoning Zhang, Jianwei Zhang, Xudong Wang, Jiehui Xie, Pengcheng Zheng, Haoyu Wang, Sungyoung Lee, Chi-lok Andy Tai, Yang Yang, Heng Tao Shen

机构 * School of Computing, Kyung Hee University(京畿大学计算机学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Computer Science and Information Engineering, Harbin Normal University(哈尔滨师范大学计算机科学与信息工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及继续教育学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 GRASP通过引导区域感知的稀疏提示方法,提升多模态大语言模型在遥感图像任务中的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18326 2026-01-27 cs.LG 50%

Cognitive Fusion of ZC Sequences and Time-Frequency Images for Out-of-Distribution Detection of Drone Signals

基于ZC序列和时频图像的认知融合用于无人机信号的分布外检测

Jie Li, Jing Li, Lu Lv, Zhanyu Ju, Fengkui Gong

机构 * State Key Laboratory of Integrated Services Network(集成服务网络国家重点实验室) Xidian University(西安电子科技大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出基于ZC序列和时频图像的认知融合算法,用于提升无人机信号分布外检测的性能,实验表明其在识别和检测指标上均有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16516 2026-01-27 cs.LG 50%

Rethinking Large Language Models For Irregular Time Series Classification In Critical Care

重新思考用于危重监护中不规则时间序列分类的大型语言模型

Feixiang Zheng, Yu Wu, Cecilia Mascolo, Ting Dang

机构 * The University of Melbourne, Australia(墨尔本大学) University of Cambridge, UK(剑桥大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文研究了LLMs在不规则ICU时间序列分类中的有效性,发现编码器设计比对齐策略更重要,但LLMs在训练时间和少样本学习中表现欠佳。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19034 2026-01-27 eess.SP 50%

Fast Vortex Beam Alignment for OAM Mode Multiplexing in LOS MIMO Networks

快速涡旋束对齐用于 LOS MIMO 网络中的 OAM 模式复用

Poorya Mollahosseini, Yasaman Ghasempour

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 OrthoVortex 提出了一种快速且精确的 OAM 模式对齐方法,通过跨模式相位识别实现快速对齐,提升 LOS MIMO 网络的容量和信干比。

Comments 13 pages, 12 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 11 篇

2601.18240 2026-01-27 cs.CV 70%

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

V-Loop:用于医学视觉问答中幻觉检测的视觉逻辑循环验证

Mengyuan Jin, Zehui Liao, Yong Xia

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 V-Loop通过双向推理和视觉逻辑循环验证,提升医学视觉问答中幻觉检测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18222 2026-01-27 cs.CV 57%

HomoFM: Deep Homography Estimation with Flow Matching

HomoFM:基于流匹配的深度视角估计

Mengfan He, Liangzheng Sun, Chunyu Li, Ziyang Meng

机构 * Department of Precision Instrument, Tsinghua University(清华大学精密仪器系) School of Instrument Science and Opto-Electronics Engineering, Beijing Information Science and Technology University(北京信息科技大学仪器科学与光电工程学院) School of Aerospace Engineering, Beijing Institute of Technology(北京理工大学航空航天工程学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 HomoFM通过引入流匹配技术,首次将生成建模中的速度场学习应用于视角估计,提升估计精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18619 2026-01-27 cs.AI 57%

Visual Attention Reasoning via Hierarchical Search and Self-Verification

通过分层搜索与自验证的视觉注意力推理

Wei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang, Ming Zhu, Haichuan Tang, Xuelong Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出通过分层搜索与自验证的视觉注意力推理框架,有效提升多模态大语言模型的视觉定位和推理能力,显著降低幻觉发生率。

Comments The paper is withdrawn by the authors after discovering a flaw in the theoretical derivation presented in the Method section. This incorrect step leads to conclusions that are not supported by the corrected derivation. The authors plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07137 2026-01-27 cs.LG cs.AI 57%

A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications

混合专家的全面综述:算法、理论与应用

Siyuan Mu, Sen Lin

机构 * University of Houston(德克萨斯大学休斯敦分校) Department of Computer Science, University of Houston(德克萨斯大学休斯敦分校计算机科学系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了混合专家模型在算法、理论和应用方面的最新进展,探讨了其在持续学习、元学习等机器学习范式中的设计,并总结了其在计算机视觉和自然语言处理中的应用。

Comments 29 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18193 2026-01-27 cs.HC 50%

InkIdeator: Supporting Chinese-Style Visual Design Ideation via AI-Infused Exploration of Chinese Paintings

InkIdeator: 通过融合AI的中国绘画探索支持中文风格视觉设计构思

Shiwei Wu, Ziyao Gao, Zhendong He, Zongtan He, Zhupeng Huang, Xia Chen, Wei Zeng, Xiaojuan Ma, Zhenhui Peng

专题命中 其他多模态 :multi-modal(abstract)

AI总结 InkIdeator通过融合AI技术,利用多模态大模型分析中国绘画中的文化维度,为中文风格视觉设计提供构思支持,帮助用户高效探索文化符号和可视化想法。

Comments 21 pages, 15 figures, CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17890 2026-01-27 cs.HC 50%

Physiological and Behavioral Modeling of Stress and Cognitive Load in Web-Based Question Answering

基于Web问答的应激与认知负荷的生理与行为建模

Ailin Liu, Francesco Chiossi, Felix Henninger, Lisa Bondo Andersen, Tobias Wistuba, Sonja Greven, Frauke Kreuter, Fiona Draxler

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文通过多模态数据建模,探索基于网络问答的应激与认知负荷的快速检测方法,为更适应的调查界面提供可行性证明。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17676 2026-01-27 cs.HC 50%

GazeSummary: Exploring Gaze as an Implicit Prompt for Personalization in Text-based LLM Tasks

GazeSummary: 探索目光作为隐式提示在基于文本的LLM任务中的个性化应用

Jiexin Ding, Yizhuo Zhang, Xinyun Liu, Ke chen, Yuntao Wang, Shwetak Patel, Akshay Gadre

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文研究了利用用户目光作为隐式提示,通过LLM生成个性化文本摘要的方法,并验证了其在真实阅读任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17440 2026-01-27 cs.RO 50%

PILOT: A Perceptive Integrated Low-level Controller for Loco-manipulation over Unstructured Scenes

PILOT:一种用于非结构化场景中人形机器人类人协作的感知式低层控制器

Xinru Cui, Linxi Feng, Yixuan Zhou, Haoqi Han, Zhe Liu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing(自动化与智能感知学院) Shanghai Jiao Tong University(上海交通大学) Global College(全球学院) school of Computer Science(计算机科学学院) Shanghai Key Laboratory of Navigation and Location Based Services(导航与位置基于服务上海市重点实验室)

专题命中 其他多模态 :cross-modal(abstract)

AI总结 PILOT通过统一的强化学习框架,实现了感知式人形机器人移动与操作的稳定控制,提升了在复杂非结构化环境中的任务执行能力。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17296 2026-01-27 econ.EM 50%

Recovering Counterfactual Distributions via Wasserstein GANs

通过Wasserstein GANs恢复反事实分布

Xinran Liu

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出基于Wasserstein GANs的稳健估计方法,用于恢复反事实分布,解决了传统DSC在支持不匹配和多模态结果下的稳定性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20280 2026-01-27 math.DS 50%

Rigidité, expansion et entropie en dynamique non-archimédienne (Rigidity, expansion and entropy in non-Archimedean dynamics)

刚性、扩张与熵在非阿基米德动力学中

Charles Favre, Juan Rivera-Letelier

专题命中 其他多模态 :multimodal(abstract)

AI总结 该研究在非阿基米德动力学中证明了有理映射的刚性性质,表明其拓扑熵为整数对数,并建立了李亚普诺夫指数与野分支位置的关系。

Comments 60 pages, in french. Comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17012 2026-01-27 cs.CY cs.HC 50%

The Digital Divide in Geriatric Care: Why Usability, Not Access, is the Real Problem

老龄化护理中的数字鸿沟:为什么是可用性,而不是访问性才是真正的难题

Christine Ine

专题命中 其他多模态 :multimodal(abstract)

AI总结 本研究指出,老年人数字鸿沟的核心问题在于可用性设计,而非访问性,强调以用户为中心的设计改进对提升老年人数字健康采用至关重要。

Comments Page Count - 15 Word Count - 3671

详情

展开后加载摘要…

URL PDF HTML 收藏