arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4892 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4892 篇

1602.02343 2016-02-23 cs.CV 79%

Eye-CU: Sleep Pose Classification for Healthcare using Multimodal Multiview Data

Carlos Torres, Victor Fragoso, Scott D. Hammond, Jeffrey C. Fried, B. S. Manjunath

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Ten-page manuscript including references and ten figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.00457 2015-08-04 cs.NE cs.AI q-bio.QM 79%

Evolutionary Multimodal Optimization: A Short Survey

Ka-Chun Wong

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1111.1461 2011-11-08 cs.CV 79%

Multimodal diff-hash

Michael M. Bronstein

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.6401 2011-10-03 cs.LO cs.AI math.LO 79%

An Interpretation of Belief Functions by means of a Probabilistic Multi-modal Logic

Frederic Dambreville

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
0905.4369 2009-12-01 cs.AI cs.LO 79%

Automating Quantified Multimodal Logics in Simple Type Theory -- A Case Study

Christoph Benzmueller

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments ii + 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01539 2026-07-07 hep-ph hep-ex hep-th nucl-ex nucl-th 版本更新 79%

Multimodal Fragmentation of All-Heavy Pentaquarks: Uncertainty-Aware Predictions for Hadron Colliders

所有重子五夸克的多模碎片化:用于强子对撞机的不确定性感知预测

Francesco Giovanni Celiberto

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出考虑不确定性的主导功率碎片化描述,用于研究强子对撞机中所有charm五夸克态的碎片化,构建多模碎片函数集PQ5Q1.1,结合微扰与非微扰不确定性,优化初始尺度输入以描述多夸克和双夸克产生机制,并应用于HL-LHC及未来FCC。

Comments 51 pages, 9 figures, 1 table, 264 references, version published in Phys. Rev. D. One novel set of multimodal (direct multicharm and diquark-like initial-scale inputs) PQ5Q1.1 NLO collinear fragmentation functions. Includes F-MHOU and F-NPWF uncertainty replicas, threshold-aware HF-NRevo DGLAP evolution, and LHAPDF release at https://github.com/FGCeliberto/Collinear_FFs

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16274 2025-05-23 cs.CY 79%

Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse

Wei Meng

专题命中 其他多模态 :multimodal(title,abstract)

Comments It integrates artificial intelligence, multimodal cognitive modeling, political behavior simulation, and data visualization, making it an interdisciplinary AI-based modeling study

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00939 2024-07-02 cs.NE math.OC 79%

Modified CMA-ES Algorithm for Multi-Modal Optimization: Incorporating Niching Strategies and Dynamic Adaptation Mechanism

Wathsala Karunarathne, Indu Bala, Dikshit Chauhan, Matthew Roughan, Lewis Mitchell

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(comments)

Comments 15 pages, 1 figure, 16 tables. Submitted for GECCO 2024 competition on Benchmarking Niching Methods for Multimodal Optimization

详情

展开后加载摘要…

URL PDF HTML 收藏
1406.2539 2014-06-11 cs.NE 79%

Maximizing Diversity for Multimodal Optimization

Fabricio Olivetti de Franca

专题命中 其他多模态 :multimodal(title,abstract)

Comments submitted to PPSN'14 Workshop Advances in Multimodal Optimization

详情

展开后加载摘要…

URL PDF HTML 收藏
1003.1466 2010-03-09 math.OC 79%

Firefly Algorithms for Multimodal Optimization

Xin-She Yang

专题命中 其他多模态 :multimodal(title,abstract)

Journal ref X.-S. Yang, Firefly algorithms for multimodal optimization, in: Stochastic Algorithms: Foundations and Applications, SAGA 2009, Lecture Notes in Computer Sciences, Vol. 5792, pp. 169-178 (2009).

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0311029 2009-12-01 cs.IR cs.PL 79%

Staging Transformations for Multimodal Web Interaction Management

Michael Narayan, Chris Williams, Saverio Perugini, Naren Ramakrishnan

专题命中 其他多模态 :multimodal(title,abstract)

Comments Describes framework and software architecture for multimodal web interaction management

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07104 2026-08-21 cs.CV cs.AI 版本更新 79%

Extended to Reality: Prompt Injection in 3D Environments

扩展到现实:3D环境中的提示注入

Zhuoheng Li, Ying Chen

专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究提出PI3D攻击,通过在3D环境中放置带文本的物理物体来注入提示,挑战多模态大语言模型的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18320 2026-04-21 cs.CV cs.AI 79%

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations

EVE: 通过可执行视觉变换实现多模态大语言模型的可验证自进化

Yongrui Heng, Chaoya Jiang, Han Yang, Shikun Zhang, Wei Ye

机构 * National Engineering Research Center for Software Engineering, Peking University(软件工程国家级工程研究中心,北京大学) School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出EVE框架,通过可执行视觉变换实现多模态大语言模型的可验证自进化,解决伪标签方法和模板方法的不足,通过双策略架构和多维奖励系统提升模型自进化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03307 2026-04-17 cs.CV cs.AI 79%

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

V-Reflection:将多模态大语言模型从被动观察者转变为主动提问者

Jiazhou Zhou, Yucheng Chen, Hongyang Li, Qing Jiang, Hu Zhou, Ying-Cong Chen, Lei Zhang

机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)人工智能方向) International Digital Economy Academy(国际数字经济学院) MedVisAI Lab, Lee Kong Chian School of Medicine, Nanyang Technological University, and Centre of AI in Medicine(医学视觉人工智能实验室,南洋理工大学Lee Kong Chian医学院,以及人工智能在医学中的中心) South China University of Technology(华南理工大学) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系)

专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出V-Reflection框架,通过'思考后再观察'的视觉反思机制,使多模态大语言模型能主动提问视觉特征空间,提升细粒度任务的感知能力。

Comments Main paper 14 pages with supplementary 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08751 2025-05-14 cs.CL cs.CV cs.LG 79%

Aya Vision: Advancing the Frontier of Multilingual Multimodality

Saurabh Dash, Yiyang Nan, John Dang, Arash Ahmadian, Shivalika Singh, Madeline Smith, Bharat Venkitesh, Vlad Shmyhlo, Viraat Aryabumi, Walter Beller-Morales, Jeremy Pekmez, Jason Ozuzu, Pierre Richemond, Acyr Locatelli, Nick Frosst, Phil Blunsom, Aidan Gomez, Ivan Zhang, Marzieh Fadaee, Manoj Govindassamy, Sudip Roy, Matthias Gallé, Beyza Ermis, Ahmet Üstün, Sara Hooker

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17971 2026-08-25 cs.HC cs.ET 78%

Generating Multimodal Textures with a Soft Hydro-Pneumatic Haptic Ring

使用软水气动触觉环生成多模态纹理

Ana Sanz Cozcolluela, Koen Wosten, Yasemin Vardar

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 提出一种软触觉环,结合气动和液压驱动,通过数据驱动渲染方法在近端指骨上生成粗糙度、热感和软度等多模态纹理,用户研究显示平均纹理匹配精度为68%。

Comments 29 pages, 19 figures, journal

Journal ref International Journal of Human-Computer Studies. volume 212, pages 103814, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05708 2026-08-25 q-bio.QM 版本更新 78%

HuBMAP Data Portal: a resource for multimodal spatial and single-cell data of healthy human tissues

HuBMAP 数据门户:健康人体组织多模态空间和单细胞数据的资源

Morgan L. Turner, Thomas C. Smits, Tiffany S. Liaw, Brendan Honick, Bill Shirey, Lisa Choy, Nikolay Akhmetov, Shaokun An, David Betancur, Dominic Bordelon, Karl Burke, Ivan Cao-Berg, John Conroy, Chris Csonka, Penny Cuda, Sean Donahue, Stephen Fisher, Derek Furst, Ed Hanna, Josef Hardi, Tabassum Kakar, Mark S. Keller, Devin Lange, Xiang Li, Yan Ma, Alison McWilliams, Austen Money, Richard Morgan, Eric Moerth, Juan Muerto, Mark A. Musen, Emily Nic, Martin J. O'Connor, Gesina Phillips, Alexander J. Ropelewski, Ryan Sablosky, Sravani Saripalli, Max Sibilla, Derek Simmel, Alan Simmons, Xu Tang, Joel Welling, Zhou Yuan, Martin Hemberg, HuBMAP Consortium, Matthew Ruffalo, Jonathan Silverstein, Philip Blood, Nils Gehlenborg

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 介绍 HuBMAP 数据门户,一个整合健康人体组织多模态空间和单细胞数据的综合平台,支持搜索、可视化和分析,并通过统一处理流程确保数据可比性。

Comments 41 pages; 8 figures; 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18953 2026-08-20 stat.ME 新提交 78%

Mixed Membership Model of Low-rank Matrices with Multimodal Extension

具有多模态扩展的低秩矩阵混合隶属度模型

David Snider, Zhongyuan Lyu, Jian Kang, Yuqi Gu

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 针对矩阵值数据提出带多模态扩展的低秩混合隶属度模型,兼具可解释群体原型与连续受试者隶属度,在模拟及人类连接组计划数据中表现优异,达极小极大最优重构率。

Comments 75 pages, 13 figures, 3 tables, including Supplementary Material

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05163 2026-08-19 physics.optics cond-mat.mes-hall quant-ph 版本更新 78%

Topology and criticality in non-Hermitian multimodal optical resonators through engineered losses

通过可控损耗实现非厄米多模光学谐振器中的拓扑与临界性

Elizabeth Louis Pereira, Hongwei Li, Andrea Blanco-Redondo, Jose L. Lado

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究提出多模非厄米晶格模型,证实其可呈现拓扑模式与临界性,分析其对损耗波动等的鲁棒性,发现可通过外部规范场调控局域化特性,为可控非厄米拓扑设计提供新策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14951 2026-08-18 cs.LG eess.IV q-bio.QM stat.ML 新提交 78%

PathFinder: Joint Decompositions of Linked Multimodal Datasets

PathFinder:链接多模态数据集的联合分解方法

Ying-Qiu Zheng, Alex Fung, Stephen M Smith, Rogier B Mars, Saad Jbabdi

机构 * Oxford Centre for Integrative Neuroimaging(牛津整合神经成像中心) University of Oxford(牛津大学)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 PathFinder是一种通用的联合矩阵分解框架,可分析未必共享同一维度的多模态数据集,能发现跨模态等的共同模式并预测缺失数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14297 2026-08-17 math.LO cs.LO 新提交 78%

Multimodal Logic Programming with Full Formulas

带完整公式的多模态逻辑程序设计

Kenji Tokuo

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出一阶多模态逻辑程序设计系统MMLP,其支持任意公式作为程序与查询,基于带模态原理的希尔伯特系统给出声明语义,采用嵌套证明演算,合一算法控制特征参数作用域,性能优于相关系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18990 2026-08-17 physics.med-ph cs.NA math.NA 版本更新 78%

Quantitative Multi-Modal Optical Coherence Photoacoustic Elastography

定量多模态光学相干光声弹性成像

Ekaterina Sherina, Lisa Krainz, Wolfgang Drexler, Otmar Scherzer

专题命中 其他多模态 :multi-modal(title,abstract)

AI总结 提出结合光学相干断层扫描(OCT)和光声断层扫描(PAT)的多模态弹性成像框架,通过混合反演算法融合互补信息,在硅胶体模上实现更高信噪比和更准确的刚度估计。

Comments 10 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02998 2026-08-12 cs.CY cs.CR 78%

Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction

Karan Patel, Yu-Zheng Lin, Gaurangi Raul, Bono Po-Jen Shih, Matthew W. Redondo, Banafsheh Saber Latibari, Jesus Pacheco, Soheil Salehi, Pratik Satam

专题命中 其他多模态 :multimodal(title,abstract)

Comments 9 pages, 3 figures, accepted by IEEE FIE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07795 2026-08-11 stat.ML cs.LG stat.ME 新提交 78%

Conformal Calibration for Multi-Modal Regression with Missing Modalities

针对模态缺失的多模态回归的共形校准

Ilia Azizi

专题命中 其他多模态 :multi-modal(title,abstract)

AI总结 针对多模态回归中模态缺失或分歧导致预测区间难校准的问题,提出感知模态的共形校准层,在多组实验中提升了区间性能,尤其在模态缺失时恢复了覆盖率。

Comments Published at the Symposium on Conformal and Probabilistic Prediction with Applications (COPA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08071 2026-08-11 cond-mat.mtrl-sci 新提交 78%

Multimodal deep learning framework to predict strain localization of Mg/LPSO two-phase alloys

预测Mg/LPSO两相合金应变局域化的多模态深度学习框架

Daiki Kuriki, Fabien Briffod, Takayuki Shiraiwa, Manabu Enoki

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究采用多模态深度学习框架,结合三种微观结构描述符,从三维图像预测Mg/LPSO两相合金的压缩局域应变分布,验证了方法的有效性。

Comments Published in Acta Materialia, Volume 281, 120398 (2024)

Journal ref Acta Materialia 281 (2024) 120398

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04430 2026-08-11 math.NA cs.NA stat.ML 版本更新 78%

An adaptive split-combine Gaussian mixture filter for nonlinear and multimodal state estimation

面向非线性与多模态状态估计的自适应拆分-合并高斯混合滤波器

San Kim, Won Chang, Daniel B. Forger, Dae Wook Kim

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文针对非线性多模态状态估计问题,提出自适应拆分-合并高斯混合滤波器(AMF),通过拆分合并高斯粒子实现高效准确的PDF估计,在多基准测试中性能优于基线滤波器,还提出了其并行实现方案。

Comments 60 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06707 2026-08-10 cs.RO 新提交 78%

Hoverflie: An empirical investigation of rotor shrouds to transform micro air vehicles into multi-modal hovercraft

Hoverflie:对转子导流罩将微型飞行器转变为多模式气垫船的实验研究

Mrinmoy Modak, Daniel S. Drew

机构 * University of Hawaii at Manoa(夏威夷大学马诺阿分校)

专题命中 其他多模态 :multi-modal(title,abstract)

AI总结 本研究通过设计定制转子导流罩系统,将Crazyflie 2.1微型飞行器转变为多模式机器人,开发经验模型优化导流罩配置,提升了地面效应性能并延长了续航时间,为相关后续研究提供了基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06494 2026-08-10 astro-ph.EP astro-ph.IM physics.app-ph 新提交 78%

Asteroid Disruption and Deflection Simulations for Multi-Modal Planetary Defense

用于多模态行星防御的小行星破坏与偏转模拟

Alexander N. Cohen, Philip Lubin, Darrel Robertson, Mark Boslough, Sasha Egan, Brin Bailey, Angela M. Stickle, Elizabeth A. Silber, Peter Meinhold, Dharv Patel

专题命中 其他多模态 :multi-modal(title,abstract)

AI总结 该研究针对小行星行星防御的终端场景问题,采用PI方法结合ALE3D模拟,证实20-100米级碎石堆小行星可通过特定参数的超高速钨侵彻体撞击有效减缓,为多模态行星防御提供可行方案。

Comments 9 pages, 8 figures

Journal ref Acta Astronautica, Volume 225, December 2024, Pages 960-967

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04910 2026-08-06 cs.HC 新提交 78%

AutoCue: Multimodal LLM-Assisted Externalization of Implicit Inputs as Instructional Visual Cues in Screencast Tutorials

AutoCue:多模态大语言模型辅助将隐式输入外部化为录屏教程中的教学视觉提示

Shengyang Luo, Shengyao Luo, Xiaolei Guo, Fengze Zhang, James Liang, Yingjie Victor Chen

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 AutoCue是多模态大语言模型辅助的教程增强流水线,可将录屏教程中隐式输入外部化为视觉提示,在Autodesk Maya上的评估显示其能缩短任务时间、减少交互中断并提升学习者体验

Comments Accepted to Graphics Interface 2026 (GI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03132 2026-08-05 cs.HC 新提交 78%

Understanding Organizational Strategies Across Multimodal Artifacts in Immersive Computational Notebooks

沉浸式计算笔记本中跨多模态制品的组织策略理解

Sungwon In, Minju Baeck, Yalong Yang, Sang Ho Yoon, Woontack Woo, Mallesham Dasari

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究针对沉浸式计算笔记本(ICoN)中多模态制品的组织策略开展用户研究,发现参与者主要采用基于深度的布局,空间组织围绕基于单元的制品构建,弥补了该领域多模态组织策略研究的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏