arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-29 至 2025-12-29 共收录 53 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 10 篇

2512.19415 2025-12-29 cs.CV 57%

Non-Contrast CT Esophageal Varices Grading through Clinical Prior-Enhanced Multi-Organ Analysis

通过临床先验增强的多器官分析进行非对比CT食管静脉曲张分级

Xiaoming Zhang, Chunli Li, Jiacheng Hao, Yuan Gao, Danyang Tu, Jianyi Qiao, Xiaoli Yin, Le Lu, Ling Zhang, Ke Yan, Yang Hou, Yu Shi

机构 * Department of Radiology, Shengjing Hospital of China Medical University, 110004, Shenyang, China(中国医科大学附属盛京医院放射科) DAMO Academy, Alibaba Group(阿里巴巴集团 DAMO 院) School of Biomedical Engineering, Tsinghua University, 100084, Beijing, China(清华大学生物医学工程学院) Faculty of Science and Engineering, Sorbonne University, 75005, Paris, France(索邦大学科学与工程学院) Hupan Lab, 310023, Hangzhou, China(华平实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 MOON++通过结合临床先验知识的多器官分析,实现了非对比CT食管静脉曲张的高效分级。

Comments Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21482 2025-12-29 cs.AI 57%

LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis

LogicLens:面向文本导向伪造分析的视觉-逻辑协同推理

Fanwei Zeng, Changtao Miao, Jing Huang, Zhiya Tan, Shutao Gong, Xiaoming Yu, Yang Wang, Huazhe Tan, Weibin Yao, Jianshu Li

机构 * Ant Group(蚂蚁集团) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :MLLM(abstract);分类 cs.AI

AI总结 LogicLens通过视觉-逻辑协同推理框架,提升文本导向伪造分析的准确性和鲁棒性,其在多个基准测试中表现优异。

Comments 11 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21402 2025-12-29 cs.CV 57%

Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation

理解病毒性:一种基于Rubric的视觉-语言模型框架用于短形式教育娱乐内容评估

Arnav Gupta, Gurekas Singh Sahney, Hardik Rathi, Abhishek Chandwani, Ishaan Gupta, Pratik Narang, Dhruv Kumar

机构 * Birla Institute of Technology and Science, Pilani(比拉理工学院和科学学院,比里尼)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于Rubric的视觉-语言模型框架,用于评估短形式教育娱乐视频的病毒性,通过提取音频视觉特征并训练回归评估器,实现可解释且可扩展的参与度预测。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19855 2025-12-29 cs.LG cs.HC 50%

Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning

在LLM中诱导因果世界模型以实现零样本物理推理

Aditya Sharma, Ananya Gupta, Chengyu Wang, Chiamaka Adebayo, Jakub Kowalski

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Bombay(印度理工学院孟买分校) Department of Computer Science and Automation(计算机科学与自动化系) Indian Institute of Science(印度科学研究所) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学) University of Lagos(拉各斯大学) Faculty of Mathematics, Informatics, and Mechanics(数学、信息学与力学学院) University of Warsaw(华沙大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出CWMI框架,通过引入因果物理模块和因果干预损失,使LLM具备因果推理能力,在零样本物理推理任务中取得显著成效。

Comments 12 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2512.22009 2025-12-29 cs.CV 70%

iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception

iSHIFT: 轻量级慢-快 GUI 代理与自适应感知

Sarthak Mehrotra, Sairam V C Rebbapragada, Mani Hemanth Reddy Bonthu, Vineeth N Balasubramanian

机构 * Indian Institute of Technology, Bombay(印度理工学院,孟买) Indian Institute of Technology, Hyderabad(印度理工学院,海得拉巴)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 iSHIFT是一种轻量级GUI代理,通过慢-快混合推理和灵活标记实现高效与精确的视觉交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21220 2025-12-29 cs.AI cs.CV cs.RO 62%

RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic

RoboSafe: 通过可执行的安全逻辑保障具身智能体

Le Wang, Zonghao Ying, Xiao Yang, Quanchen Zou, Zhenfei Yin, Tianlin Li, Jian Yang, Yaodong Yang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航大学) AI Security Lab(360AI安全实验室) The University of Sydney(悉尼大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 RoboSafe通过可执行的安全逻辑保障具身智能体,减少危险行为并保持任务性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21598 2025-12-29 cs.CV 57%

From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement

从浅层幽默到隐喻:通过LMM代理自我改进实现无标签有害迷因检测

Jian Lang, Rongpei Hong, Ting Zhong, Leiting Chen, Qiang Gao, Fan Zhou

机构 * University of Electronic Science and Technology of China(电子科技大学) Southwestern University of Finance and Economics(西南财经大学) Intelligent Digital Media Technology Key Laboratory of Sichuan Province(四川省智能数字媒体技术重点实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 ALARM通过LMM代理自我改进实现无标签有害迷因检测,利用浅层迷因信息提升对复杂迷因的识别能力。

Comments 12 pages. Accepted by KDD 2026 research track. Codes are released at https://github.com/Jian-Lang/ALARM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13030 2025-12-29 cs.CV cs.LG cs.RO 57%

Motus: A Unified Latent Action World Model

Motus:一个统一的潜在动作世界模型

Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, Jun Zhu

机构 * Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,THBI实验室,清华-博世联合机器学习中心,清华大学) Peking University(北京大学) Horizon Robotics

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 Motus通过统一的潜在动作世界模型整合理解、视频生成和动作模块,利用光流和三阶段训练流程提升大规模动作预训练效果,实现模拟与现实场景中的性能突破。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 14 篇

2512.21508 2025-12-29 cs.CV 83%

Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification

固定预算参数高效训练结合冻结编码器提升多模态胸片分类

Md Ashik Khan, Md Nahid Siddique

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India(计算机科学与工程系,印度理工学院Kharagpur分校) Knight Foundation School of Computing and Information Sciences, Florida International University, Florida, USA(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究通过冻结编码器的参数高效训练策略,在降低计算成本的同时提升了多模态胸片分类的性能。

Comments Accepted at the 2025 28th International Conference on Computer and Information Technology (ICCIT). 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21760 2025-12-29 cs.CV cs.AI 81%

A-QCF-Net: An Adaptive Quaternion Cross-Fusion Network for Multimodal Liver Tumor Segmentation from Unpaired Datasets

A-QCF-Net:一种自适应四元数交叉融合网络用于从无配对数据集进行多模态肝肿瘤分割

Arunkumar V, Firos V M, Senthilkumar S, Gangadharan G R

机构 * University College of Engineering, Bharathidasan Institute of Technology Campus, Anna University(安娜大学工程学院,巴拉特拉桑理工学院校区,安娜大学) National Institute of Technology(国家理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 A-QCF-Net通过自适应四元数交叉融合网络,实现从无配对CT和MRI数据集中的肝肿瘤分割,提升分割精度并验证模型的临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21916 2025-12-29 cs.CV 79%

Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition

补丁作为节点:面向多模态动作识别的人本图表示学习

Zeyu Liang, Hailun Xia, Naichuan Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PAN框架,通过人本图表示学习提升多模态动作识别性能,结合双路径和统一网络结构实现高效融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21897 2025-12-29 cs.LG cs.AI 79%

MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction

MMCTOP: 一种用于临床试验结果预测的多模态文本化与专家混合框架

Carolina Aparício, Qi Shi, Bo Wen, Tesfaye Yadete, Qiwei Han

机构 * Nova School of Business and Economics(诺瓦商学院) Hogarthian Technologies(霍加斯技术) School of Medicine(医学院) Oregon Health & Science University(俄勒冈健康与科学大学) Cleveland Clinic(克利夫兰诊所) IBM Research(IBM研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MMCTOP通过多模态文本化与专家混合框架,提升临床试验结果预测的精度与稳定性。

Comments 15 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19443 2025-12-29 cs.CV 79%

D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning

D2Pruner: 用于MLLM标记剪枝的去偏重要与结构多样性

Evelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen, Ke Yan, Shouhong Ding, Biqing Qi, Linfeng Zhang

机构 * Tencent YouTu Lab(腾讯YouTu实验室)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 D2Pruner通过结合去偏重要与结构剪枝机制,有效提升MLLM标记剪枝的效率和保真度,尤其在细粒度定位任务中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09626 2025-12-29 cs.SI cs.AI cs.LG 79%

Certainly Bot Or Not? Trustworthy Social Bot Detection via Robust Multi-Modal Neural Processes

确定是机器人还是不是?通过鲁棒多模态神经过程进行可信的社交机器人检测

Qi Wu, Yingguang Yang, hao liu, Hao Peng, Buyun He, Yutong Xia, Yong Liao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本研究提出鲁棒多模态神经过程框架,通过增强多模态神经过程的鲁棒性来检测社交机器人,同时提升不确定性估计能力。

Comments We withdraw this paper due to an error identified in the experimental setup. Specifically, the evaluation protocol described in Section 4 does not correctly reflect the intended experimental design, which may affect the validity of the reported results. To avoid potential misunderstanding by readers, we choose to withdraw this version and revise the work before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20146 2025-12-29 cs.CV 79%

AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection

AlignFreeNet: 跨模态预对齐是否必要?一种端到端无对齐的轻量级网络用于可见-红外目标检测

Dingkun Zhu, Haote Zhang, Lipeng Gu, Wuzhou Quan, Fu Lee Wang, Honghui Fan, Jiali Tang, Haoran Xie, Xiaoping Zhang, Mingqiang Wei

机构 * School of Computer Science, Jiangsu University of Technology(江苏科技大学计算机科学学院) School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院) School of Data Science, Lingnan University(岭南大学数据科学学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 AlignFreeNet通过无对齐融合范式,提出VCC和FCF模块,有效缓解可见-红外目标检测中的跨模态错位问题,实现端到端轻量级网络的高鲁棒性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22102 2025-12-29 cs.LG 78%

Explainable Multimodal Regression via Information Decomposition

通过信息分解实现可解释的多模态回归

Zhaozhao Ma, Shujian Yu

机构 * Zhejiang University(浙江大学) Georgia Institute of Technology(佐治亚理工学院) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) UiT - The Arctic University of Norway(挪威北极大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出基于部分信息分解的多模态回归框架,通过分解模态特定表示为唯一、冗余和协同组件,提升预测准确性和可解释性,并在多个数据集上验证其有效性。

Comments Project Page: https://github.com/zhaozhaoma/PIDReg

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21543 2025-12-29 cs.IR 78%

CEMG: Collaborative-Enhanced Multimodal Generative Recommendation

CEMG: 基于协作增强的多模态生成推荐

Yuzhen Lin, Hongyi Chen, Xuanjing Chen, Shaowen Wang, Ivonne Xu, Dongming Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 CEMG通过多模态融合层和残差量化变分自编码器,提升多模态生成推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21506 2025-12-29 cs.LG cs.AI cs.CL cs.HC 76%

MotionTeller: Multi-modal Integration of Wearable Time-Series with LLMs for Health and Behavioral Understanding

MotionTeller: 多模态整合可穿戴时间序列与大语言模型用于健康和行为理解

Aiwei Zhang, Arvind Pillai, Andrew Campbell, Nicholas C. Jacobson

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CL、cs.AI

AI总结 MotionTeller通过整合可穿戴时间序列与大语言模型,实现高精度的自然语言行为摘要生成,提升健康和行为理解的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21476 2025-12-29 cs.CV cs.AI 62%

GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification

GPF-Net:门控渐进融合学习用于息肉重识别

Suncheng Xiang, Xiaoyang Wang, Junjie Jiang, Hejia Wang, Dahong Qian

机构 * Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Shanghai Fifth People's Hospital(上海第五人民医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 GPF-Net通过门控渐进融合学习提升息肉重识别性能,结合多模态融合策略优于现有单模态模型。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21452 2025-12-29 cs.CV cs.AI 62%

Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism

基于特征融合与注意力机制的智能GPR道路隐缺陷图像识别

Haotian Lv, Yuhui Zhang, Jiangbo Dai, Hanli Wu, Jiaji Wang, Dawei Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出基于特征融合与注意力机制的GPR道路隐缺陷图像识别框架,通过数据增强、多模态特征融合和迁移学习提升检测精度与鲁棒性。

Comments Accepted for publication in *IEEE Transactions on Geoscience and Remote Sensing*

Journal ref IEEE Transactions on Geoscience and Remote Sensing, 2025, 63, 5213217

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21863 2025-12-29 cs.IR cs.MM 57%

Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion

冻结的大型视频语言模型用于微视频推荐:特征提取与融合的系统研究

Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.MM

AI总结 本文通过系统研究冻结LVLMs的特征提取与融合策略,提出DFF框架,证明中间隐藏状态优于标题表示,ID嵌入融合优于替换,并在微视频推荐中取得最佳性能。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21856 2025-12-29 cs.CV 57%

Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection

突破对齐障碍:TPS驱动的语义相关性学习用于无对齐RGB-T显著目标检测

Lupiao Hu, Fasheng Wang, Fangmei Chen, Fuming Sun, Haojie Li

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出TPS-SCL网络,通过双流MobileViT和高效Mamba机制,结合语义相关性约束和薄板样条对齐模块,有效解决未对齐RGB-T图像对中的显著目标检测问题。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 1 篇

2512.21380 2025-12-29 cs.SI cs.CY 71%

SENTINEL: A Multi-Modal Early Detection Framework for Emerging Cyber Threats using Telegram

SENTINEL: 一种利用Telegram的多模态早期检测框架用于新兴网络威胁

Mohammad Hammas Saeed, Howie Huang

专题命中 其他多模态 :multi-modal(title)

AI总结 SENTINEL通过多模态信号分析Telegram中的网络安全讨论,实现对新兴网络威胁的早期检测,F1值达0.89。

详情

展开后加载摘要…

URL PDF HTML 收藏