arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

1604.01683 2016-04-07 cs.CV 57%

Fusing Face and Periocular biometrics using Canonical correlation analysis

N. S. Lakshmiprabha

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.01359 2016-03-07 stat.ML cs.CV cs.LG 57%

Learning deep representation of multityped objects and tasks

Truyen Tran, Dinh Phung, Svetha Venkatesh

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.0226 2015-10-28 cs.CV 57%

Non-parametric Image Registration of Airborne LiDAR, Hyperspectral and Photographic Imagery of Forests

Juheon Lee, Xiaohao Cai, Carola-Bibiane Schonlieb, David Coomes

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.06073 2015-03-02 cs.CV 57%

Study on Sparse Representation based Classification for Biometric Verification

Zengxi Huang, Yiguang Liu, Xiaoming Wang, Jinrong Hu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1407.6748 2014-07-28 cs.CV 57%

Enhancing the Accuracy of Biometric Feature Extraction Fusion Using Gabor Filter and Mahalanobis Distance Algorithm

Ayodeji S. Makinde, Yaw Nkansah-Gyekye, Loserian S. Laizer

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Focused on extraction of feature from two different modalities (face and fingerprint) using Gabor filter

详情

展开后加载摘要…

URL PDF HTML 收藏
1404.7796 2014-06-20 stat.ML cs.LG cs.MM 57%

Majority Vote of Diverse Classifiers for Late Fusion

Emilie Morvant, Amaury Habrard, Stéphane Ayache

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.MM

Comments IAPR Joint International Workshops on Statistical Techniques in Pattern Recognition and Structural and Syntactic Pattern Recignition, Joensuu : Finland (2014)

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.5732 2014-05-27 cs.CV 57%

Self-tuned Visual Subclass Learning with Shared Samples An Incremental Approach

Hossein Azizpour, Stefan Carlsson

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Updated ICCV 2013 submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1312.6171 2014-01-14 cs.NE cs.CV cs.LG 57%

Learning Paired-associate Images with An Unsupervised Deep Learning Architecture

Ti Wang, Daniel L. Silver

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 9 pages, for ICLR-2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1307.7351 2013-08-02 cs.AI cs.RO 57%

Knowledge Representation for Robots through Human-Robot Interaction

Emanuele Bastianelli, Domenico Bloisi, Roberto Capobianco, Guglielmo Gemignani, Luca Iocchi, Daniele Nardi

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments Knowledge Representation and Reasoning in Robotics Workshop at ICLP 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1007.0412 2010-07-05 cs.AI 57%

Improving Iris Recognition Accuracy By Score Based Fusion Method

Ujwalla Gawande, Mukesh Zaveri, Avichal Kapur

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments http://ijict.org/index.php/ijoat/article/view/improving-iris-recognition

Journal ref International Journal of Advancements in Technology, Vol 1, No 1 (2010)

详情

展开后加载摘要…

URL PDF HTML 收藏
0709.1099 2009-12-01 cs.AI cs.RO 57%

Multi-Sensor Fusion Method using Dynamic Bayesian Network for Precise Vehicle Localization and Road Matching

Cherif Smaili, Maan El Badaoui El Najjar, François Charpillet

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22271 2026-07-27 cond-mat.mtrl-sci 新提交 56%

SAGE-Net: Semantics-Augmented Geometric Encoder for Material Property Prediction

SAGE-Net:用于材料属性预测的语义增强几何编码器

Guanghui Zhang, Yuxuan Yao, Kieran B. Spooner, Jun Yin, Dan Han, David O. Scanlon, Lijun Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(comments)

AI总结 研究针对材料属性预测,提出语义增强几何编码器网络SAGE-Net,通过语义引导消息传递将晶体学语义注入几何消息传递,在多基准测试中表现出色,能有效捕捉晶体学特征,是深度集成多模态材料学习的通用可转移框架。

Comments 29 pages, 5 figures, multi-modal network

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19840 2025-07-29 cs.CV cs.AI cs.CL 56%

AutoSign: Direct Pose-to-Text Translation for Continuous Sign Language Recognition

Samuel Ebimobowei Johnny, Blessed Guda, Andrew Blayama Stephen, Assane Gueye

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态训练与对齐 :分类 cs.CV、cs.CL、cs.AI;multimodal(comments)

Comments Paper to appear at the 1st Workshop in Multimodal Sign Language Recognition at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21129 2026-08-25 astro-ph.EP astro-ph.IM 版本更新 50%

Transit Searches for Habitable-Zone Exoplanets with Artificial Intelligence

利用人工智能搜索宜居带系外行星的凌星观测

Qingtian Liu

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本综述探讨人工智能在宜居带系外行星凌星信号搜索各环节的进展,指出其可提升分析效率、恢复弱信号,结合物理建模等能提高发现效率与可靠性。

Comments This paper has been withdrawn by the authors for administrative reasons related to our institutional research policy

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25194 2026-08-25 cs.LG 版本更新 50%

Localize and Neutralize: Gradient-guided Token Suppression against Visual Prompt Injection Attack

先定位再中和:梯度引导的令牌抑制对抗视觉提示注入攻击

Dongpeng Zhang, Ke Ma, Yangbangyan Jiang, Gaozheng Pei, Longtao Huang, Qianqian Xu, Qingming Huang

机构 * School of Advanced Interdisciplinary Sciences, UCAS(UCAS交叉学科研究院) School of Electronic, Electrical and Communication Engineering, UCAS(UCAS电子电气与通信工程学院) State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室) Alibaba Group(阿里巴巴集团) School of Computer Science and Technology, UCAS(UCAS计算机科学与技术学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Key Laboratory of Big Data Mining and Knowledge Management, UCAS(UCAS大数据挖掘与知识管理重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 针对多模态大语言模型的视觉提示注入攻击,提出梯度令牌掩码(GTM)方法,通过梯度分析定位关键图像令牌并掩码中和,将攻击成功率降至接近零且计算开销极小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19281 2026-08-21 cs.RO cs.HC 新提交 50%

APPROVE: Visual End-User-in-the-Loop Robot Programming with LLMs

APPROVE:结合大型语言模型的面向视觉端用户的机器人编程

Bijan Kavousian, Miray Özakkas, Josefine Monnet, Oliver Petrovic, Christian Brecher

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 APPROVE是一种结合LLMs的多模态端用户机器人编程框架,通过Blockly可视化程序并支持用户确认、修改或拒绝,还可存储复用已确认功能,解决了现有系统透明度不足、意图对齐差、复用有限的问题,提升了用户信任与编程灵活性。

Comments Accepted for publication in Procedia CIRP, Proceedings of the 20th CIRP Conference on Intelligent Computation in Manufacturing Engineering (ICME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18410 2026-08-20 cs.LG 新提交 50%

Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

用于高效视觉-语言-动作策略的角色条件子令牌路由

Wei Jiang, Wei Wang

机构 * Futurewei Technologies(未来智能技术公司)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 该研究针对VLA模型推理成本高的问题,提出RoleSub方法,通过角色条件子令牌路由压缩视觉和语言表示,在匹配视觉-KV预算下多数场景优于仅令牌控制,结合压缩后总KV仅为原始的9.2-11.3%且控制性能强。

Comments 12 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17667 2026-08-19 math.ST stat.TH 版本更新 50%

Rapid Bayesian Computation and Estimation for Neural Networks via Log-Concave Coupling

Curtis McDonald, Andrew R. Barron

专题命中 多模态训练与对齐 :multi-modal(abstract)

Journal ref Math. Stat. Learn. (2026), published online first

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16005 2026-08-18 cs.LG 新提交 50%

Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment

基于检索的双融合与相似度感知对比的分子-文本对齐方法

Shunshun Gu, Shengqi Qiu, Hang Zhou, Xiao Luo

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 针对现有分子-文本对齐方法忽略子结构与文本细粒度语义关系的问题,提出RISEN方法,通过检索构建孪生分子并结合相似度感知对比学习,在基准数据集上取得更优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15748 2026-08-18 cs.RO 新提交 50%

Making two action heads agree: coordination mechanisms and a runtime collapse certificate for flow-matching policies

让两个动作头达成一致:流匹配策略的协调机制与运行时崩溃证书

Jinhui Sun, Wei Zhou, Bowen Yang, Xinliang Xiao, Li Yang

机构 * School of Automation, Nanjing University of Science and Technology(南京理工大学自动化学院)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 该研究针对流匹配策略的双动作头多模态误报问题,提出四类协调机制并推导机会校正协调边界,在LIBERO-Plus等测试中验证了残差为最强故障信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15516 2026-08-18 cs.LG 新提交 50%

UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity

UniFed-VLM:面向多异质性视觉语言模型的联邦指令调优

Pengyu Wang, Baochen Xiong, Xiaoshan Yang, Yifan Xu, Zhang Qimeng, Haifeng Chen, Changsheng Xu

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 UniFed-VLM是解决任务、模态、模型架构联合异质性的VLM联邦指令调优框架,含FedCSA与TCoD组件,在多基准数据集上优于现有FL方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15310 2026-08-18 cs.CR cs.LG 新提交 50%

FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting

FedADB:用于缓解遗忘的类别锚驱动双分支联邦学习

Zhenyan Liu, Hua Zhang, Haoran Gao, Qi Li, Hongliang Zhu, Huiyu Zhou, Zongliang Shen, Yanxin Xu, Jiahui Wang

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 针对联邦学习中跨客户端数据异质性导致的全局知识遗忘问题,提出FedADB框架,通过类别锚与双分支机制平衡全局一致性与本地优化,在多数据集上提升了准确率与收敛速度。

Comments Accepted to appear in Proceedings of the 34th ACM International Conference on Multimedia (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13911 2026-08-17 cs.LG 新提交 50%

MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity

MedMix:模态异质性下的专业化一致联邦稀疏混合专家模型

Adiba Orzikulova, Dong Min Kim, Jaehong Yoon, Sung-Ju Lee

机构 * KAIST(韩国科学技术院) NTU Singapore(新加坡南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 针对联邦多模态医疗AI的客户端与样本级模态异质性问题,提出MedMix框架,通过模态上下文感知路由、共识引导路由对齐与客户端自适应专家聚合,在多模态医疗数据集上取得最优平均F1值,严重异质性下提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11364 2026-08-13 cs.HC 新提交 50%

QUARTZ: Qualitative Understanding via Accessible Representation and Visualization

QUARTZ:通过可访问表示与可视化实现定性理解

Omar Khan, JooYoung Seo

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 该研究针对盲人和低视力研究者无法访问定性数据可视化的问题,提出基于网络的QUARTZ系统,通过用户研究揭示相关障碍并提出设计准则,以提升该群体参与定性研究的可访问性。

Comments 24 pages, accepted at ASSETS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10350 2026-08-12 cond-mat.str-el cond-mat.mtrl-sci 新提交 50%

Observation geometry for uncertainty-aware Hamiltonian inference and experimental design in quantum magnets

量子磁体中不确定性感知哈密顿量推断与实验设计的观测几何

Roy Liu, Venugopal Ranganathan, David Dahlbom, Shizhou Xu, Tianyu Zhang, Yuan Ni, Daniel M. Pajerowski, Garrett Granroth, Thomas Strohmer, Matthew B. Stone, Andrew F. May, Mark D. Lumsden, Joshua J. Turner, Yongqiang Cheng, Zhantao Chen

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 该研究提出AI赋能的不确定性感知哈密顿量推断与自适应实验设计框架,结合哈密顿量条件神经代理等方法,通过NiPS₃的中子散射测量验证,为量子材料表征与实验设计提供通用策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10959 2026-08-12 cs.LG 版本更新 50%

Population-Aware Physics-Informed Neural Particle Flow for Robust Spacecraft Bayesian Navigation

群体感知的物理信息神经粒子流用于贝叶斯更新

Batu Candan, Simone Servadio

机构 * Iowa State University(爱荷华州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 提出群体感知的物理信息神经粒子流(PA-PINPF),通过群体编码器增强粒子更新,在贝叶斯后验传输中优于标准方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23035 2026-08-12 cs.LG 版本更新 50%

Learning Disease-Sensitive Latent Interaction Graphs From Noisy Cardiac Flow Measurements

从噪声心脏血流测量中学习疾病敏感的潜在交互图

Viraj Patel, Marko Grujic, Philipp Aigner, Theodor Abart, Marcus Granegger, Deblina Bhattacharjee, Katharine Fraser

机构 * Department of Computer Science, University of Bath(巴斯大学计算机科学系) Department of Mechanical Engineering, University of Bath(巴斯大学机械工程系) Christian Doppler Laboratory for Mechanical Circulatory Support, Department of Cardiac and Thoracic Aortic Surgery, Medical University of Vienna(维也纳医学大学心脏和胸主动脉外科系基督教多普勒实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 本文提出了一种物理指导的潜在关系框架,用于从噪声心脏血流测量中学习疾病敏感的潜在交互图,以捕捉心脏疾病和干预的稳健标志。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09166 2026-08-11 cs.RO cs.LG stat.ME 新提交 50%

Particle-Based Conformal Prediction for Contact-Aware Uncertainty Calibration in Stratified Configuration Spaces

分层构型空间中接触感知不确定性校准的基于粒子的保形预测

Luís Marques, Kristian Popov, Dmitry Berenson

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 针对机器人接触障碍物时运动模型不准确、未来构型分布异常的问题,提出CaPTURe算法,实现了接触与非接触场景下的指定覆盖率要求,任务成功率较最佳基线提升30%。

Comments 31 pages, 10 figures, 5 tables. Accepted at COPA 2026 (Conformal and Probabilistic Prediction with Applications). Project page: https://um-arm-lab.github.io/capture/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01414 2026-08-11 cs.RO 版本更新 50%

Learning When to See and When to Feel: Adaptive Vision-Torque Fusion for Contact-Aware Manipulation

学习何时视觉何时触觉:适应性视觉-扭矩融合用于接触感知操控

Jiuzhou Lei, Chang Liu, Yu She, Xiao Liang, Minghui Zheng

机构 * Edwardson School of Industrial Engineering, Purdue University(普渡大学爱德华森工业工程学院) Zachry Department of Civil and Environmental Engineering, Texas A&M University(德州农工大学扎克里土木与环境工程系)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出适应性视觉-扭矩融合策略,通过比较不同融合方法提升机器人操控性能,实验显示方法在成功率上优于基线14%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07183 2026-08-10 cs.LG 新提交 50%

Conformal Fusion Under Missing Modalities

缺失模态下的共形融合

Alireza Moayedikia

机构 * Swinburne University of Technology(斯威本科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出MCCF架构,解决缺失模态下的多模态融合问题,该架构兼具模态缺失鲁棒性与校准不确定性,在多基准测试中表现出良好的覆盖率与精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏