arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-19 至 2026-02-19 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 7 篇

2602.16245 2026-02-19 cs.CV 79%

HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis

HyPCA-Net:在医学图像分析中推进多模态融合

J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat, A. Alavi, S. Mistry, N. Zaidi

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) RoentGen Health(RoentGen健康公司) Lulea University of Technology(卢勒奥大学) QUT(昆士兰科技大学) RMIT University(皇家墨尔本理工大学) Curtin University(Curtin大学) Deakin University(德肯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HyPCA-Net通过高效残差注意力模块和双视角级联注意力模块,提升多模态医学图像分析的性能与效率。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15896 2026-02-19 cs.CL 79%

Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens

每一点帮助都有价值:通过细粒度可迁移多模态标记构建知识图谱基础模型

Yichi Zhang, Zhuo Chen, Lingbing Guo, Wen Zhang, Huajun Chen

机构 * Zhejiang University, Zhejiang, China(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文提出TOFU模型,通过细粒度可迁移多模态标记提升多模态知识图谱推理的跨图谱迁移能力。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08755 2026-02-19 cs.LG 78%

Align and Adapt: Multimodal Multiview Human Activity Recognition under Arbitrary View Combinations

对齐与适应:在任意视图组合下的人类活动识别多模态多视角学习

Duc-Anh Nguyen, Nhien-An Le-Khac

机构 * University College Dublin(都柏林大学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 AliAd通过结合多视角对比学习和专家混合模块,实现任意视图组合下的灵活多模态人类活动识别,提升模型适应性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08177 2026-02-19 cs.CV cs.AI 73%

MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision

MedReasoner: 通过强化学习实现从临床思维到像素级精度的推理 grounding

Zhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu, Kongming Liang, Zhanyu Ma

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 MedReasoner 通过强化学习实现从临床推理到像素级精度的医学影像 grounding,提出统一医学推理 grounding 任务和 U-MRG-14K 数据集,展示其在医学影像分析中的优越性能。

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16681 2026-02-19 cs.CV 57%

VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detection

VETime: 基于视觉增强的零样本时间序列异常检测

Yingyuan Yang, Tian Lan, Yifei Gao, Yimeng Lu, Wenjun He, Meng Wang, Chenghao Liu, Chen Zhang

机构 * Department of Industrial Engineering, Tsinghua University, Beijing, China(清华大学工业工程系) Lab, Huawei Technologies Ltd, Beijing, China(华为技术有限公司2012实验室) Datadog AI Research(Datadog人工智能研究)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 VETime通过融合视觉与时间模态,提出零样本时间序列异常检测框架,实现高精度定位与低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15857 2026-02-19 cs.CL 57%

Multi-source Heterogeneous Public Opinion Analysis via Collaborative Reasoning and Adaptive Fusion: A Systematically Integrated Approach

通过协同推理与自适应融合进行多源异构公共意见分析:一种系统整合方法

Yi Liu

机构 * Yi Liu School of Software Xi’an Jiaotong University(刘毅 软件学院 西安交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出CRAF框架,通过协同推理与自适应融合整合传统方法与LLMs,提升多源异构公共意见分析的跨平台适应性与性能

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16584 2026-02-19 q-bio.NC 50%

The Representational Alignment Hypothesis: Evidence for and Consequences of Invariant Semantic Structure Across Embedding Modalities

表征对齐假说:跨嵌入模态的不变语义结构的证据及其后果

Akhil Ramidi, Kevin Scharp

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 该研究提出表征对齐假说,探讨跨模态嵌入的不变语义结构,质疑其根本性,并提出新的哲学视角和区分方法。

Comments 23 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏