arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-30 至 2025-12-30 共收录 79 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 10 篇

2512.23158 2025-12-30 eess.SY cs.RO cs.SY 50%

Breaking Symmetry-Induced Degeneracy in Multi-Agent Ergodic Coverage via Stochastic Spectral Control

通过随机频谱控制打破多智能体等耗覆盖中的对称诱导退化

Kooktae Lee, Julian Martinez

机构 * Department of Mechanical Engineering, New Mexico Institute of Mining and Technology(机械工程系,新墨西哥矿业与技术研究所)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出随机频谱控制方法,通过引入随机扰动和收缩项,解决多智能体等耗覆盖中因对称性导致的梯度抵消问题,确保轨迹有界并避免停滞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22968 2025-12-30 eess.SY cs.SY 50%

A Bezier Curve Based Approach to the Convexification of the AC Optimal Power Flow Problem

基于贝塞尔曲线的AC最优潮流问题凸化方法

Carlos Arturo Saldarriaga-Cortes, Carlos Adrian Correa-Florez, Maximiliano Bueno-Lopez, Maria Victoria Gasca-Segura

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出基于贝塞尔曲线的AC最优潮流问题凸化方法,通过引入辅助变量和对数变换,实现高效且准确的电力系统优化。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19854 2025-12-30 cs.RO cs.HC 50%

Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models

思考、行动、学习:一种利用闭环大语言模型的自主机器人代理框架

Anjali R. Menon, Rohit K. Sharma, Priya Singh, Chengyu Wang, Aurora M. Ferreira, Mateja Novak

机构 * Dept. of Electronics & Comm. Eng.(电子与通信工程系) Government Engineering College(政府工程学院) Dept. of Electrical & Electronics Eng.(电气与电子工程系) Poornima College of Engineering(波奥尼玛工程学院) Dept. of Electronics & Telecom. Eng.(电子与电信工程系) Shivaji University College of Eng.(希瓦吉大学工程学院) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学) Dept. of Electrical Eng.(电气工程系) Instituto Federal do Maranhão(马里兰联邦学院) Dept. of Electrical & Computer Eng.(电气与计算机工程系) Technical University of Košice(科希丘夫技术大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出T-A-L框架,通过闭环大语言模型实现机器人自主学习与策略优化,显著提升复杂任务的成功率和泛化能力。

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 14 篇

2512.23056 2025-12-30 cs.LG physics.comp-ph 88%

PI-MFM: Physics-informed multimodal foundation model for solving partial differential equations

PI-MFM:基于物理的多模态基础模型用于求解偏微分方程

Min Zhu, Jingmin Sun, Zecheng Zhang, Hayden Schaeffer, Lu Lu

机构 * Department of Statistics and Data Science, Yale University(统计与数据科学系,耶鲁大学) Department of Applied Mathematics and Statistics, Johns Hopkins University(应用数学与统计学系,约翰霍普金斯大学) Department of Applied Computational Mathematics and Statistics, University of Notre Dame(应用计算数学与统计学系,圣母大学) Department of Mathematics, University of California Los Angeles(数学系,加州大学洛杉矶分校) Department of Chemical and Environmental Engineering, Yale University(化学与环境工程系,耶鲁大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract)

AI总结 PI-MFM是一种基于物理的多模态基础模型,通过强制执行偏微分方程在预训练和适应过程中,提高求解PDE的效率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22605 2025-12-30 cs.AI cs.CV 84%

Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation

学习多模态移动动态以实现通用的下一站推荐

Junshu Dai, Yu Wang, Tongya Zheng, Wei Ji, Qinghong Guo, Ji Cao, Jie Song, Canghong Jin, Mingli Song

机构 * Zhejiang University(浙江大学) Hangzhou City University(杭州市大学) Nanjing University(南京大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态移动(M^3ob)方法,通过构建统一时空关系图和门控机制,提升位置推荐任务的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 82%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22878 2025-12-30 cs.CV cs.AI 81%

SwinTF3D: A Lightweight Multimodal Fusion Approach for Text-Guided 3D Medical Image Segmentation

SwinTF3D: 一种轻量级多模态融合方法用于文本引导的3D医学图像分割

Hasan Faraz Khan, Noor Fatima, Muzammil Behzad

机构 * King Fahd University of Petroleum(国王法赫德石油大学) SDAIA-KFUPM Joint Research Center for Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SwinTF3D通过统一视觉和语言表示,实现文本引导的3D医学图像分割,具备高效且适应性强的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22545 2025-12-30 cs.CV cs.AI 81%

Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains

跨多模态领域自奖励的推理框架

Jesen Zhang, Ningyuan Liu, Kaitong Cai, Sidi Liu, Jing Yang, Ziliang Chen, Xiaofei Sun, Keze Wang

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SR-MCR通过自奖励机制提升多模态推理的连贯性和准确性,在视觉基准测试中取得领先性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22503 2025-12-30 cs.CV 79%

SCAFusion: A Multimodal 3D Detection Framework for Small Object Detection in Lunar Surface Exploration

SCAFusion: 一种针对月球表面探索的小目标多模态3D检测框架

Xin Chen, Kang Luo, Yangyi Xiao, Hesheng Wang

机构 * Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, Key Laboratory of Marine Intelligent Equipment and System of Ministry of Education, Shanghai Engineering Research Center of Intelligent Control and Management, Shanghai Jiao Tong University(自动化系、教育部系统控制与信息处理重点实验室、教育部海洋智能装备与系统重点实验室、上海智能控制与管理工程研究中心、上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SCAFusion提出了一种针对月球表面探索的小目标多模态3D检测框架,通过改进的特征对齐和坐标注意机制提升小目标检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10731 2025-12-30 cs.CV 70%

Removal then Selection: A Coarse-to-Fine Fusion Perspective for RGB-Infrared Object Detection

去除后再选择:一种从粗到细融合视角的RGB红外目标检测

Tianyi Zhao, Maoxun Yuan, Feng Jiang, Nan Wang, Xingxing Wei

机构 * Institute of Artificial Intelligence, Hangzhou Innovation Institute, Beihang University, Beijing(北京航空航天大学人工智能研究院,杭州创新院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Beijing Institute of Control and Electronic Technology(北京控制与电子技术研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种从粗到细的融合视角,通过去除冗余信息和动态选择特征来提升RGB红外目标检测的性能。

Comments 11pages, 10figures

Journal ref IEEE Transactions on Intelligent Transportation Systems, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16124 2025-12-30 cs.HC cs.LG 67%

ZIA: A Theoretical Framework for Zero-Input AI

ZIA:零输入AI的理论框架

Aditi De

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥基分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract)

AI总结 ZIA提出了一种基于多模态融合的零输入AI框架,通过整合生物信号和上下文数据实现前瞻性意图预测,提升实时推理效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18262 2025-12-30 cs.RO cs.AI cs.CV cs.HC cs.LG 62%

ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement

ReSemAct:通过语义结构化和效用细化推进细粒度机器人操作

Chenyu Su, Weiwei Shang, Chen Qian, Fei Zhang, Shuang Cong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ReSemAct 通过语义结构化和效用细化方法,在细粒度机器人操作中实现更精确的效用目标生成与动态环境适应。

Comments Code and videos: https://github.com/scy-v/ReSemAct and https://resemact.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22217 2025-12-30 cs.CV cs.AI 62%

VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition

VLM-PAR:一种用于行人属性识别的视觉语言模型

Abdellah Zakaria Sellam, Salah Eddine Bekhouche, Fadi Dornaika, Cosimo Distante, Abdenour Hadid

机构 * Department of Innovation Engineering(创新工程系) University of Salento, Italy(意大利萨伦托大学) Institute of Applied Sciences and Intelligent Systems - CNR(应用科学与智能系统研究所 - CNR) University of the Basque Country UPV/EHU(巴斯克国家大学UPV/EHU) IKERBASQUE, Basque Foundation for Science(伊基塔斯克巴塞克基金会) Sorbonne University Abu Dhabi(索邦大学阿布扎比分校)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 VLM-PAR通过整合大规模视觉语言预训练与跨模态细化,提升行人属性识别在类别不平衡和泛化挑战中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22188 2025-12-30 cs.CV cs.AI 62%

HookMIL: Revisiting Context Modeling in Multiple Instance Learning for Computational Pathology

HookMIL: 重新审视多实例学习中的上下文建模以用于计算病理学

Xitong Ling, Minxi Ouyang, Xiaoxiao Li, Jiawen Li, Ying Chen, Yuxuan Sun, Xinrui Chen, Tian Guan, Xiaoping Liu, Yonghong He

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Informatics, Xiamen University(厦门大学信息学院) School of Engineering, Westlake University(西湖大学工程学院) Zhongnan Hospital, Wuhan University(武汉大学中南医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 HookMIL通过引入可学习的钩子标记和多样性损失,提升多实例学习在计算病理学中的效率与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22981 2025-12-30 cs.CV 57%

Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation

具有空间意识的对称对齐用于文本引导的医学图像分割

Linglin Liao, Qichuan Geng, Yu Liu

机构 * School of Information Engineering, Capital Normal University Beijing, China(信息工程学院,首都师范大学北京,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出SSA框架,通过空间感知对称对齐机制和复合方向引导策略,提升文本引导的医学图像分割性能,尤其在准确分割具有空间约束的病变方面表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22800 2025-12-30 cs.CV 57%

Medical Scene Reconstruction and Segmentation based on 3D Gaussian Representation

基于3D高斯表示的医学场景重建与分割

Bin Liu, Wenyan Tian, Huangxin Fu, Zizheng Li, Zhifen He, Bo Li

机构 * Nanchang Hangkong University(南昌航空大学) Guilin University Of Electronic Technology(桂林电子科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于3D高斯和三平面表示的医学图像重建方法,解决传统方法在稀疏切片中的结构不连续和细节丢失问题,提升重建效率和图像质量。

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22135 2025-12-30 cs.DC cs.AI cs.HC cs.MA 57%

SoDA: An Efficient Interaction Paradigm for the Agentic Web

SoDA:面向代理网络的高效交互范式

Zicai Cui, Zhouyuan Jian, Weiwen Liu, Weinan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 SoDA提出了一种面向代理网络的高效交互范式,通过解耦记忆与应用逻辑,降低数据锁定和认知过载,提升信息处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 2 篇

2512.22603 2025-12-30 cs.CL 79%

Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis

结构化提示与大语言模型集成用于多模态对话基于方面的情感分析

Zhiqiang Gao, Shihao Gao, Zixing Zhang, Yihao Guo, Hongyu Chen, Jing Han

机构 * Hunan University(湖南大学) University of Cambridge(剑桥大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出结构化提示与大语言模型集成方法,用于多模态对话基于方面的情感分析,有效提升情感识别与翻转检测的准确性。

Journal ref ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23454 2025-12-30 cs.CV 57%

Automated river gauge plate reading using a hybrid object detection and generative AI framework in the Limpopo River Basin

利用混合目标检测和生成式AI框架实现利比里河盆地的自动水位计读数

Kayathri Vigneswaran, Hugo Retief, Jai Clifford Holmes, Mariangel Garcia Andarcia, Hansaka Tennakoon

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本研究提出结合目标检测和生成式AI的混合框架,用于自动读取河流水位计,提升水文监测的精度和效率。

Comments 11 pages, 14 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏