arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-16 至 2025-12-16 共收录 87 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2512.12921 2025-12-16 cs.CR cs.AI 57%

Cisco Integrated AI Security and Safety Framework Report

思科集成AI安全与安全框架报告

Amy Chang, Tiffany Saade, Sanket Mendapara, Adam Swanda, Ankit Garg

机构 * Cisco AI Threat and Security Research(思科人工智能威胁与安全研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出思科集成AI安全与安全框架,旨在统一分类和操作化AI风险,涵盖安全与安全,适用于威胁识别、风险优先级排序等,具有全面性和扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12630 2025-12-16 cs.HC cs.AI 57%

ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists

ORIBA:探索基于大语言模型的对话机器人作为原创角色艺术家创造力支持工具

Yuqian Sun, Xingyu Li, Shunyu Yao, Noura Howell, Tristan Braud, Chang Hee Lee, Ali Asadipour

机构 * Computer Science Research Centre, Royal College of Art(皇家艺术学院计算机科学研究中心) Digital Media, School of Literature, Media, and Communication, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Princeton University(普林斯顿大学) Digital Media, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Division of Integrative Systems and Design, The Hong Kong University of Science and Technology(香港科学大学整合系统与设计 division) Industrial Design Department, College of Engineering, KAIST(韩国科学技术院工程学院工业设计系)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

AI总结 ORIBA通过大语言模型支持原创角色艺术家的创意过程,平衡AI辅助与创意自主权。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02580 2025-12-16 cs.CL 57%

From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks

从模仿到辨别:一种通用的课程优势机制,增强跨领域推理任务

Changpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang, Yang Li, Qiliang Liang, Hongzhen Wang, Shuai Nie, Jiaming Xu, Runyu Shi, Ying Huang, Guoquan Zhang

机构 * \equalcontrib(1号机构)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 CAPO是一种基于优势信号的自适应课程机制,通过引导模仿学习和引入负信号提升跨领域推理任务的泛化能力。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 15 篇

2512.12822 2025-12-16 cs.CV cs.AI 84%

Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding

Lemon:一种统一且可扩展的3D多模态模型用于通用空间理解

Yongyuan Liang, Xiyao Wang, Yuanchen Ju, Jianwei Yang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 Lemon提出一种统一的Transformer架构,通过联合处理3D点云块和语言标记,实现空间-语言融合,提升3D多模态模型的可扩展性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04870 2025-12-16 eess.IV cs.CV 83%

Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images

多模态不确定性鲁棒树冠覆盖分割用于高分辨率遥感图像

Yuanyuan Gui, Wei Li, Yinjian Wang, Xiang-Gen Xia, Mauro Marty, Christian Ginzler, Zuyuan Wang

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(空间智能信息处理国家重点实验室) Department of Electrical and Computer Engineering, University of Delaware(电气与计算机工程系,德雷塞尔大学) Swiss Federal Institute for Forest, Snow, and Landscape Research WSL(瑞士森林、雪和景观研究联邦 institute WSL)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MURTreeFormer通过多模态分割框架降低不确定性,提升高分辨率遥感图像中树冠分割的鲁棒性与准确性。

Journal ref IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13261 2025-12-16 cs.CV 83%

Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

解锁基于强化学习的多模态推理中难度先验的潜力

Mingrui Chen, Haogeng Liu, Hao Liang, Huaibo Huang, Wentao Zhang, Ran He

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) NLPR&MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文通过建模问题难度先验信息,改进基于强化学习的多模态推理性能,通过数据筛选、优势分化和难度提示提升模型推理深度和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11901 2025-12-16 cs.CV cs.LG 83%

CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities

CLARGA:任意模态集合上的多模态图表示学习

Santosh Patapati

机构 * Santosh Patapati(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 CLARGA是一种通用的多模态融合架构,通过构建注意力加权图实现多模态表示学习,适用于任意模态集合,具有高效的融合能力和良好的鲁棒性。

Comments WACV; Supplementary material is available on CVF proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12657 2025-12-16 cs.CV 79%

Cross-modal Fundus Image Registration under Large FoV Disparity

跨模态视网膜图像在大视野差异下的配准

Hongyang Li, Junyi Tao, Qijie Wei, Ningzhi Yang, Meng Wang, Weihong Yu, Xirong Li

机构 * Renmin University of China, Beijing, China(中国人民大学) Peking Union Medical College Hospital, Beijing, China(北京友谊医院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出CARe方法,用于解决大视野差异下的跨模态视网膜图像配准问题,通过裁剪和双拟合对齐改进配准效果。

Comments Accepted as a regular paper at MMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20991 2025-12-16 cs.CV 79%

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

MR-COSMO:一种用于查询驱动3D分割的视觉-文本记忆召回与直接跨模态对齐方法

Chade Li, Pengju Zhang, Yihong Wu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MR-COSMO通过视觉-文本记忆召回与直接跨模态对齐方法,在查询驱动的3D分割中实现几何与语义特征的精确融合,提升点云分割性能。

Comments Accepted by AAAI 2026. Copyright (c) 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12881 2025-12-16 cs.LG q-bio.NC stat.ML 78%

Unsupervised learning of multiscale switching dynamical system models from multimodal neural data

从多模态神经数据中无监督学习多尺度切换动力学系统模型

DongKyu Kim, Han-Lin Hsieh, Maryam M. Shanechi

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种无监督学习方法,通过多尺度神经观测学习切换多尺度动力学系统模型,以更准确解码行为并提升脑机接口性能。

Comments 30 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12842 2025-12-16 cs.RO cs.AI cs.LG 70%

SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding

SAGA:通过结构化可及性 grounding 实现开放世界移动操作

Kuan Fang, Yuxin Chen, Xinghao Zhu, Farzad Niroui, Lingfeng Sun, Jiuguang Wang

机构 * RAI Institute(RAI研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 SAGA通过结构化可及性接地实现开放世界移动操作,能有效处理多种任务形式并实现零样本执行。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01042 2025-12-16 cs.CV 70%

WCCNet: Wavelet-context Cooperative Network for Efficient Multispectral Pedestrian Detection

WCCNet:小波-上下文协作网络用于高效多光谱行人检测

Xingjian Wang, Li Chai, Jiming Chen, Zhiguo Shi

机构 * the College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院) the College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 WCCNet通过低计算成本的多光谱特征提取与跨模态融合,提升自动驾驶中的行人检测效率与精度。

Comments 35 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12690 2025-12-16 cs.LG cs.CL cs.CV 62%

Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning

重新评估监督微调的作用:VLM推理中的实证研究

Yongcan Yu, Lingxiao He, Shuo Lu, Lijun Sheng, Yinuo Xu, Yanbo Wang, Kuangpu Guo, Jianjie Cheng, Meng Wang, Qianlong Xie, Xingxing Wang, Dapeng Hu, Jian Liang

机构 * NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(人工智能研究院 & 模式识别与人工智能研究所,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Meituan(美团)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本研究通过实证分析发现,监督微调在VLM推理中具有重要作用,挑战了强化学习优于监督微调的主流观点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11849 2025-12-16 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 62%

LocoMamba: Vision-Driven Locomotion via End-to-End Deep Reinforcement Learning with Mamba

LocoMamba:基于端到端深度强化学习的视觉驱动运动控制框架

Yinuo Wang, Gavin Tao

机构 * School of Computer Science and Statistics, Trinity College Dublin(计算机科学与统计学系,三一学院都柏林)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 LocoMamba通过端到端深度强化学习实现视觉驱动的运动控制,利用Mamba模型提升序列建模效率,有效捕捉长程依赖并增强训练效率。

Comments 14 pages. This paper has been published in Advanced Engineering Informatics. Please cite the journal version: DOI: 10.1016/j.aei.2025.104230

Journal ref Advanced Engineering Informatics, Vol. 70, Art. no. 104230 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13240 2025-12-16 cs.AI cs.LG 57%

Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection

反射偏好优化(RPO):通过提示引导的反思增强策略对齐

Zihui Zhao, Zechang Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 RPO通过提示引导的反思增强策略对齐,减少幻觉并提升多模态基准性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12658 2025-12-16 cs.CV 57%

CogDoc: Towards Unified thinking in Documents

CogDoc: 向文档中的统一思维迈进

Qixin Xu, Haozhe Wang, Che Liu, Fangzhen Lin, Wenhu Chen

机构 * Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学) University of Waterloo(滑铁卢大学) Imperial College London(伦敦帝国学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 CogDoc提出一种统一的粗到细思维框架,通过直接强化学习提升文档推理性能,优于现有方法并在视觉丰富任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09350 2025-12-16 cs.CV 57%

TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment

TextGuider: 通过注意力对齐实现文本渲染的无训练指导

Kanghyun Baek, Sangyub Lee, Jin Young Choi, Jaewoo Song, Daemin Park, Jooyoung Choi, Chaehun Shin, Bohyung Han, Sungroh Yoon

机构 * Interdisciplinary Program in Artificial Intelligence, Seoul National University(人工智能交叉学科项目,首尔国立大学) Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Global Technology Research, Samsung Electronics(三星电子全球技术研究) AIIS, ASRI, INMC, ISRC, Seoul National University(AIIS、ASRI、INMC、ISRC、首尔国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 TextGuider通过注意力对齐实现文本渲染的无训练指导,提升文本生成的准确性和召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12134 2025-12-16 q-bio.BM cs.LG 50%

Modeling Dabrafenib Response Using Multi-Omics Modality Fusion and Protein Network Embeddings Based on Graph Convolutional Networks

基于图卷积网络的多组学模态融合与蛋白质网络嵌入建模达拉菲尼反应

La Ode Aman, A Mu'thi Andy Suryadi, Dizky Ramadani Putri Papeo, Hamsidar Hasan, Ariani H Hutuba, Netty Ino Ischak, Yuszda K. Salimi

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本研究通过多组学融合与图卷积网络提升达拉菲尼反应预测,揭示互补分子决定因素。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 9 篇

2512.12461 2025-12-16 cs.LG cs.AI q-bio.NC 79%

Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling

跨模态表征知识蒸馏用于增强基于尖峰的LFP建模

Eray Erturk, Saba Hashemi, Maryam M. Shanechi

机构 * Ming Hsieh Department of Electrical and Computer Engineering(明希斯电气与计算机工程系) Thomas Lord Department of Computer Science(托马斯·劳德计算机科学系) Alfred E. Mann Department of Biomedical Engineering(阿尔弗雷德·E·曼生物医学工程系) Neuroscience Graduate Program University of Southern California(神经科学研究生项目美国南加州大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出跨模态知识蒸馏框架,通过将预训练的尖峰模型知识转移至LFP模型,提升LFP建模的准确性和泛化能力。

Comments Published at the 39th Annual Conference on Neural Information Processing Systems 2025. Code is available at https://github.com/ShanechiLab/CrossModalDistillation

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13728 2025-12-16 eess.SY cs.SY eess.SP 78%

BioGAP-Ultra: A Modular Edge-AI Platform for Wearable Multimodal Biosignal Acquisition and Processing

BioGAP-Ultra:一种模块化边缘AI平台,用于可穿戴多模态生物信号采集与处理

Sebastian Frey, Giusy Spacone, Andrea Cossettini, Marco Guermandi, Philipp Schilk, Luca Benini, Victor Kartsch

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 BioGAP-Ultra是一款模块化边缘AI平台,支持多模态生物信号采集与处理,通过提升存储、连接和信号模式实现高效低功耗的可穿戴设备应用。

Comments 17 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09359 2025-12-16 cs.HC 71%

Smart Device Development for Gait Monitoring: Multimodal Feedback in an Interactive Foot Orthosis, Walking Aid, and Mobile Application

智能设备开发用于步态监测:交互式足垫、行走辅助器和移动应用中的多模式反馈

Stefan Resch, André Kousha, Anna Carroll, Noah Severinghaus, Felix Rehberg, Marco Zatschker, Yunus Söyleyici, Daniel Sanchez-Morillo

专题命中 其他多模态 :multimodal(title)

AI总结 本文提出了一种结合智能足垫和前臂拐杖的模块化传感器系统,通过多模式触觉反馈提升步态监测的实用性与用户适应性。

Comments This work has been published in Technologies (MDPI)

Journal ref Technologies. 2025; 13(12):588

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18577 2025-12-16 q-fin.CP cs.AI cs.LG 57%

Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges

用基础模型推进金融工程:进展、应用与挑战

Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Xiaoyu Wang, Henglin Liu, Chuang Li, Kecheng Jiao, Jixuan Ying, Yang Veronica Liu, Qiang Yang, Xiu Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了金融基础模型(FFMs)的进展、应用与挑战,涵盖三种关键模态,并探讨了数据可用性、算法可扩展性和基础设施限制等关键问题。

Comments Accepted by [J]. Engineering, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12821 2025-12-16 cs.LG cs.AI physics.data-an 57%

On the continuity of flows

关于流的连续性

Congzhou M Sha

机构 * Penn State College of Medicine(宾夕法尼亚州立大学医学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了流匹配中因分布拓扑不匹配导致的连续性问题,揭示了最优速度场的跳跃不连续性现象,并探讨了其对流匹配方法的影响。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12773 2025-12-16 cs.HC cs.AI 57%

Designing The Drive: Enhancing User Experience through Adaptive Interfaces in Autonomous Vehicles

设计驱动:通过自适应界面提升自动驾驶车辆中的用户体验

Reeteesha Roy

机构 * VIT Bhopal University(维特理工学院博帕尔分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文探讨了通过自适应界面提升自动驾驶车辆用户体验的方法,强调透明性和用户控制在增强信任和满意度中的作用。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12793 2025-12-16 cs.CV 57%

ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution

ViCO: 一种面向语义感知动态高分辨率的训练策略

Long Cui, Weiyun Wang, Jie Shao, Zichen Wen, Gen Luo, Linfeng Zhang, Yanting Zhang, Yu Qiao, Wenhai Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Nanjing University(南京大学) Donghua University(东华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 ViCO通过动态调整视觉标记数量以适应图像语义复杂度,有效降低推理成本的同时保持模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18319 2025-12-16 stat.AP 50%

Sequential Design for the Efficient Estimation of Offshore Structure Failure Probability

用于高效估计海上结构失效概率的序列设计

Matthew Speers, Jonathan Angus Tawn, Philip Jonathan

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出IS-PT和AGE两种方法,用于高效估计海上结构在极端环境下的失效概率,其中IS-PT在计算成本较低时表现更可靠,而AGE在已知权重参数时可进一步降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04140 2025-12-16 cs.LG 50%

TreeFlow: Going beyond Tree-based Gaussian Probabilistic Regression

TreeFlow: 超越树形的高斯概率回归

Patryk Wielopolski, Maciej Zięba

机构 * Wrocław University of Science and Technology(沃拉夫大学科学与技术学院)

专题命中 其他多模态 :multi-modal(abstract)

AI总结 TreeFlow结合树集成与归一化流,实现对回归输出复杂分布的建模,取得多模态数据集上的SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏