arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2507.20993 2026-07-29 cs.LG cs.AI stat.ML 版本更新 79%

Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records

借助注释的学习:从多模态电子健康记录中学习治疗策略

Henri Arno, Thomas Demeester

机构 * Department of Information Technology Ghent University - imec(信息科技系根特大学 - imec)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出AACE方法,通过专家注释辅助学习多模态EHR中的因果策略,提升治疗决策效果。

Comments Accepted at Machine Learning for Healthcare (MLHC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24447 2026-07-28 cs.CV 新提交 79%

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

RP - OPSD:用于多模态大语言模型的分辨率特权在线自蒸馏

Qihui Zhu, Yuchen Wang, Zijian Wen, Tao Zhang, Mengjie Zhang, Yang Liu, Shuangwu Chen, Siying Wu, Jian Yang, Xiaofeng Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对多模态大语言模型在线自蒸馏问题,利用同一图像高低分辨率视图信息差提出RP - OPSD,学生策略用低分辨率图像生成轨迹,教师策略用原始分辨率图像监督,实验表明该方法能提升性能并加快训练速度,为在线自蒸馏提供有效途径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23658 2026-07-28 cs.CV 新提交 79%

XMatchAD: A Cross-Modal Matching Perspective on Reconstruction-based Anomaly Detection

XMatchAD:基于重建的异常检测的跨模态匹配视角

Mingxiu Cai, Zhe Zhang, Gaochang Wu, Tianyou Chai

机构 * Northeastern University(东北大学) State Key Laboratory of Synthetical Automation for Process Industries(流程工业综合自动化国家重点实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对基于重建的无监督异常检测方法难以捕捉细微异常、异常边界模糊的问题,提出XMatchAD框架,从伪跨模态匹配视角利用输入与重建图像的匹配关系,通过多步骤提升异常检测和定位精度,性能优于现有方法。

Comments accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23245 2026-07-28 cs.MM cs.LG 新提交 79%

FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities

FedTaste:用于具有缺失模态的多模态联邦学习的拓扑感知结构转移

Haochen Liang, Jie Zhang, Hideya Ochiai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 针对多模态联邦学习中模态缺失和数据分布问题,提出FedTaste框架,利用冻结基础模型提取联合多模态拓扑,结合模态自适应结构提示等方法,避免显式模态插补,在多数据集和非IID设置下性能优越且通信开销低。

Comments Accepted to ACM Multimedia (ACM MM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22779 2026-07-28 cs.LG cs.AI 新提交 79%

Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control

使用基于查询的变压器进行多模态表面肌电手势识别以用于假肢控制

Federico Del Pup, Elisa Tentori, Manfredo Atzori

机构 * Department of Information Engineering, University of Padua(帕多瓦大学信息工程系) Padova Neuroscience Center, University of Padua(帕多瓦大学帕多瓦神经科学中心) Department of Biomedical Sciences, University of Padua(帕多瓦大学生物医学科学系) Department of Neuroscience, University of Padua(帕多瓦大学神经科学系) Information Systems Institute, University of Applied Sciences Western Switzerland (HES-SO Valais)(瑞士西部应用科学大学(HES-SO瓦莱州)信息系统研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对表面肌电手势识别用于假肢控制时当前架构的局限,提出EMG-CrossFormer,通过级联交叉注意力融合层组合单模态编码器表示,经实验验证该方法能提升仅sEMG解码及多模态融合下的手势识别性能。

Comments GitHub repository: see https://github.com/deepPNClab/emg-crossformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22702 2026-07-28 cs.CV cs.LG 新提交 79%

MIME: Multimodal Interactive Motion Encoder

MIME:多模态交互式运动编码器

Addison Zucek, Prerit Gupta, Kamila Kuatova, Aniket Bera

机构 * Purdue University(普渡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对动画、AR/VR等场景中多人交互的文本-运动表示学习问题,提出多模态交互式运动编码器MIME,通过特定方法捕捉结构并训练,在文本-运动检索任务中表现出色,还能跨数据集支持下游运动生成。

Comments Under review at WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16170 2026-07-28 cs.IR cs.AI 版本更新 79%

EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation

EGRA:迈向用于多模态推荐的增强行为图和表示对齐

Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Yongjie Wang, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 针对多模态推荐现有方法不足,EGRA通过纳入预训练模型生成表示构建的项-项图减轻稀疏性,引入双层动态对齐加权机制提升模态-行为表示对齐,实验证明其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22422 2026-07-27 physics.flu-dyn cs.AI 新提交 79%

PRIMS: Physics-guided Representation for Fluid Identification in Multimodal Sensing

PRIMS:多模态传感中用于流体识别的物理引导表示

Hai-Long Nguyen, Trung Thanh Nguyen, Lars Holm, Dennis Alveringh, Duc Viet Le

机构 * University of Twente(特文特大学) Nagoya University(名古屋大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对微流体应用中流体识别挑战,提出PRIMS物理感知多模态Transformer,通过三个模块集成物理知识于表示学习和注意力机制,实现可解释、数据高效的流体分类,实验显示其性能优越且鲁棒性强。

Comments 2026 European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21660 2026-07-27 cond-mat.mtrl-sci cs.AI cs.LG 新提交 79%

Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives

用于材料预测和设计的生成式和多模态人工智能:进展、挑战与展望

Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, Nicola A. Morley, Robert D. J. Oliver, Alexandra J. Ramadan, Delvin Ce Zhang, Katerina A. Christofidou, Haiping Lu

机构 * School of Computer Science, University of Sheffield, Sheffield, UK(计算机科学学院,谢菲尔德大学) Centre for Machine Intelligence, University of Sheffield, Sheffield, UK(智能机器研究中心,谢菲尔德大学) School of Chemical, Materials and Biological Engineering, University of Sheffield, Sheffield, UK(化学、材料与生物工程学院,谢菲尔德大学) Henry Royce Institute, Royce Discovery Centre, University of Sheffield, Sheffield, UK(亨利·罗伊奇研究所,谢菲尔德大学) Materials Nexus Ltd., Salisbury House, Cambridge, UK(材料 nexus 有限公司,剑桥,英国) School of Mathematical and Physical Sciences, University of Sheffield, Sheffield, UK(数学与物理科学学院,谢菲尔德大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究探讨生成式和多模态人工智能在材料预测与设计中的应用,引入材料属性层次结构框架,指出当前证据局限,强调需全社区标准支持多模态相关建模与基准测试,以设计具科学和实际新颖性、可实验实现的材料。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21347 2026-07-24 cs.CV 新提交 79%

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

质量感知多模态融合揭示效价-唤醒特征中的隐含身份

Jisu Kim, Benjamin S. Riggan

机构 * University of Nebraska-Lincoln(内布拉斯加大学林肯分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对传统人脸识别在无约束环境中表现不佳的问题,提出质量感知自适应融合方法 QAAF 用于多模态效价-唤醒估计,在 VA 估计任务中取得良好效果,且经 VA 训练的特征在编码身份上表现出色,确立了多模态 VA 估计为传统人脸识别的互补软生物特征模态。

Comments 10 pages, 3 figures, 6 tables. Accepted for publication at IEEE International Joint Conference on Biometrics (IJCB), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20511 2026-07-24 cs.AI 新提交 79%

SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning

SiGMA:用于多模态持续指令微调的符号引导合并与适配

Keonhee Park, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究多模态持续指令微调中存在的问题,提出SiGMA框架,通过训练时符号引导自适应调优、推理时符号引导合并减轻负面干扰,在相关基准实验中显著减少干扰且性能优于现有方法。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19600 2026-07-23 stat.ME cs.CV cs.LG q-bio.QM stat.ML 新提交 79%

Deep Shape Regression for Planar Curves with Multimodal Covariates

具有多模态协变量的平面曲线深度形状回归

Manuel Pfeuffer, Roshan Prakash Rane, Hadya Yassin, Kerstin Ritter, Sonja Greven

机构 * Humboldt-Universität zu Berlin, Berlin, Germany(柏林洪堡大学) Hertie Institute for AI in Brain Health, Universität Tübingen, Tübingen, Germany(人工智能与脑健康赫特研究所) Universität Tübingen, Tübingen, Germany(图宾根大学) Hasso Plattner Institut, Universität Potsdam, Potsdam, Germany(哈索·普拉特纳研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对平面曲线提出深度形状回归模型,允许多模态高维协变量。用复值函数表示曲线,提出含模态特定编码器的协方差平滑器,模型具多种不变性,还提供弹性均值估计算法,经模拟和实际应用验证了方法有效性。

Comments 17 pages, 4 figures, 1 algorithm. Submitted to the ShapeMI Workshop, MICCAI 2026. Code is available at https://github.com/mpff/dnn-shapes

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04733 2026-07-22 cs.CL cs.LG 版本更新 79%

LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure

LP-SFT:通过多模态熵结构进行局部保持监督微调

Yueyang Wang, Baolong Bi, Shuo Lu, Jingyuan Zhang, Jiajun Shi

机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 研究监督微调中存在的问题,利用香农和雷尼熵分析揭示预训练模型的多模态熵结构,提出LP-SFT目标,在多领域实验中提升性能,平衡准确率和k准确率。

Comments 20 pages, 3 figures. Code is available at https://github.com/Wakaka161/LP-SFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17351 2026-07-21 cs.AI cs.RO 新提交 79%

DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception

深度雷达:用于自动驾驶车辆感知的端到端MIMO雷达设计与多模态融合

Eli Goldenshluger, Barak Pinkovich, Chaim Baskin

机构 * Technion–Israel Institute of Technology(以色列理工学院) Ben-Gurion University of the Negev(内盖夫本-古里安大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 深度雷达以雷达为中心,通过端到端学习稀疏采集模式,共同设计雷达传感与多模态3D检测。其可学习的MIMO设计模块在融合网络中训练,由其他传感器监督。在RADIal数据集上评估,能发现稀疏配置,降低成本与复杂性,表明最优设计取决于融合堆栈和感知任务。

Comments Accepted for publication at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17257 2026-07-21 cs.RO cs.AI 新提交 79%

Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion

通过延迟感知引导融合实现异步多模态扩散策略组合

Zihao He, Hongjie Fang, Shirun Tang, Cewu Lu, Haoshu Fang

机构 * Shanghai Jiao Tong University(上海交通大学) Noematrix(无矩阵) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对多模态扩散策略在模态差异下的融合问题,提出LAG - Fusion框架,通过延迟感知引导融合实现异步策略组合,推导参考帧重定位规则,在接触丰富操纵实验中提升了策略响应性和任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16366 2026-07-21 cs.RO cs.AI 新提交 79%

PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments

PRISM:非结构化环境中用于漫游车导航的多模态地形映射

Raul Castilla-Arquillo, Carlos Perez-del-Pulgar, Levin Gerdes, Alfonso Garcia-Cerezo, Miguel A. Olivares-Mendez

机构 * SnT, University of Luxembourg(卢森堡大学安全、可靠性和信任跨学科中心) Department of Automation and Systems Engineering, Universidad de Málaga, Andalucía Tech(马拉加大学自动化与系统工程系,安达卢西亚技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对非结构化环境中漫游车导航问题,提出PRISM多模态感知系统,利用定制传感器套件和OmniUnet网络,经新数据集验证及现场实验,可在资源受限设备上高效生成可通行性地图,实现漫游车自主导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16233 2026-07-21 cs.LG cs.AI 新提交 79%

Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction

用于乳腺癌亚型分类和生存预测的具有对比多任务学习的令牌级跨模态变换器

Suxing Liu Byungwon Min

机构 * Jiangxi Arts & Ceramics Technology Institute(江西艺术陶瓷科技职业学院) Mokwon University(木浦大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 针对整合异质模态进行癌症亚型分类和生存预测的挑战,提出令牌级跨模态变换器及对比多任务学习方法,克服现有方法在模态交互、融合方式及目标优化上的局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30994 2026-07-21 cs.MM 版本更新 79%

Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis

动态交互感知与因果解耦的多模态情感分析框架

Guangyuan Dong, Ziwei Hong, Shenghao Liu, Chenyu Wu, Yuanyuan Fang, Zihao Li, Xudong Zhang, Bingchen Liu, Yuchen Zhang, Haitao Ding, Zhenzhou Zhou, Ziyu Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 提出动态多模态因果解耦与自适应融合框架(MCAF),通过因果干预模块解耦语言模态的偏见,动态路由器实时评估交互状态并自适应融合,以及条件扩散去噪模块过滤噪声,在CMU-MOSI和CMU-MOSEI上取得最优分类性能。

Comments This preprint is withdrawn for unauthorized posting and incorrect author metadata.It was uploaded without full consent of all co-authors, with wrong name and affiliation information. We withdraw it to avoid copyright disputes. This corrects submission irregularities only, not academic content or conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24679 2026-07-21 cs.AI eess.SP 版本更新 79%

Multi-modal cross-domain mixed fusion model with dual disentanglement for fault diagnosis under unseen working conditions

用于未知工况下故障诊断的具有双重解缠的多模态跨域混合融合模型

Pengcheng Xia, Yixiang Huang, Chengjin Qin, Chengliang Liu

机构 * State Key Laboratory of Mechanical System and Vibration(机械系统与振动国家重点实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 针对未知工况下故障诊断问题,提出具有双重解缠的多模态跨域混合融合模型,通过双重解缠框架、跨域混合融合策略和三模态融合机制,实现多模态表示学习与域泛化,实验验证该方法优于先进方法及各组件有效性。

Comments Accepted for publication in Mechanical Systems and Signal Processing

Journal ref Mechanical Systems and Signal Processing, Volume 258, 2026, 114693

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18931 2026-07-21 cs.CV 版本更新 79%

Enhancing Vision Foundation Models via Multimodal Continual Pre-Training

通过多模态持续预训练增强视觉基础模型

Yitong Chen, Lingchen Meng, Wujian Peng, Jun Tao, Chenjie Xu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Baosight Software Co., Ltd.(上海博视软件有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究通过多模态持续预训练增强视觉基础模型,提出M-CPT框架,用持续位置嵌入处理不同分辨率视觉输入,通过特征对齐目标提升视觉与文本表示一致性,实验证明该框架能提升多模态理解性能并保持标准视觉任务表现。

Comments Code is available in https://github.com/ShareLab-SII/CoMP-MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18824 2026-07-20 cs.CV cs.LG 版本更新 79%

Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos

他们将去哪里?从自我中心视频建模多模态行人机动

Yuxuan Xie, Nicolas Pugeault, Chongfeng Wei, Hubert P. H. Shum, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算机科学学院) James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) Department of Computer Science, Durham University(杜伦大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出MMPM框架,通过行为感知交互模块和基于CVAE的模态感知轨迹预测器,分别建模行人过马路和不过马路两种模式,提升自我中心视角下多模态轨迹预测准确性。

Comments Accepted at The IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27556 2026-07-20 cs.CV 版本更新 79%

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

迈向领域通用的开放词汇物体检测:一种渐进领域不变跨模态对齐方法

Xiaoran Xu, Xiaoshan Yang, Jiangang Yang, Yifan Xu, Jian Liu, Changsheng Xu

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室) Institute of Microelectronics, University of Chinese Academy of Sciences(中国科学院大学微电子学院) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学前沿交叉科学学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PICA方法,通过多级模糊性和信号强度课程,解决开放词汇检测中领域变化导致的跨模态空间崩溃问题,提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04043 2026-07-17 cs.CV 版本更新 79%

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting

AnyStyle: 单次传递多模态风格化用于3D高斯点云

Joanna Kaleta, Bartosz Świrta, Kacper Kania, Tomasz Trzciński, Przemysław Spurek, Marek Kowalski

机构 * Warsaw University of Technology(华沙技术大学) Sano Centre for Computational Medicine(Sano计算医学中心) Jagiellonian University(雅盖隆大学) Microsoft(微软公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 AnyStyle通过多模态条件化实现姿态无关的零样本风格化,支持文本和视觉输入,提升3D重建的风格可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13931 2026-07-16 cs.CV 新提交 79%

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

SIVA-RL:用于多模态强化学习的灵敏度不变视觉对齐

Cheng Tang, Junzhi Ning, Min Cen, Wei Li, Xinyi Zeng, Pinxian Zeng, Rongbin Li, Qiming Zhu, Yuqiang Li, Junjun He, Yirong Chen, Ming Hu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Sichuan University(四川大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Macau(澳门大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究多模态强化学习中视觉语言模型预测与视觉证据结合问题,提出SIVA-RL框架,通过特定方法构建局部干预并以奖励下降为权重驱动对齐,在多基准测试中相比基线改进了模型。

Comments 27 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29909 2026-07-16 cs.CV 版本更新 79%

Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification

Traffic-CBM:一种结构可解释的多模态加密流量分类框架

Honglei Jin, Wenshuo Chen, Shaofeng Liang, Haozhe Jia, Runwei Guan, Menshuo Zhao, Shuxu Jin, Songning Lai, Yutao Yue

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出Traffic-CBM框架,通过统一层次化概念空间组织多模态流量证据,实现结构可解释的加密流量分类,性能与端到端模型相当且解释性更强。

Comments 14 pages, figures and tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12372 2026-07-15 cs.CV 新提交 79%

UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

UMSS:迈向无监督多模态语义分割

Haitian Zhang, Thai Duy Nguyen, Xiangyuan Wang, Mohan Liu, Lin Wang

机构 * EmPACT Lab, School of EEE, Nanyang Technological University(电气与电子工程学院电磁脉冲与天线研究室,南洋理工大学) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文针对无监督多模态语义分割问题,提出基于DINOv3的UniM2框架,通过跨模态对应协同学习统一潜在空间提取语义线索,并引入跨模态协调器缓解冲突,实验证明该框架相比现有框架有明显优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11131 2026-07-14 cs.CL 新提交 79%

TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding

TIGER:用于多模态推测解码的文本条件视觉门控路由与接受对齐

Quynh Vo, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan, Thong Nguyen

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Center of AI Research, VinUniversity(Vin大学人工智能研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 针对多模态推测解码中草稿模型视觉关键内容易偏差及现有方法不足的问题,提出TIGER框架,基于文本状态动态选视觉令牌,用接受对齐分组策略训练优化草稿模型,实验证明其在多方面取得良好效果。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11106 2026-07-14 cs.CV 新提交 79%

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

超越视觉:通过自我调节的隐式视觉工具实现高效多模态推理

Xiuwei Chen, Quanlin Chen, Wentao Hu, Zisheng Chen, Kun Xiang, Zehua Ma, Mingyang Zhang, Jianhua Han, Hanhui Li, Hang Xu, Xiaodan Liang

机构 * Sun Yat-sen University(中山大学) The Hong Kong Polytechnic University(香港理工大学) Yinwang Intelligent Technology Co., Ltd.(银望智能科技有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对多模态大语言模型推理效率问题,提出超越视觉(BEE)的隐式视觉工具范式,通过将视觉工具调用行为纳入训练目标,经两阶段训练,包括形式化思维链监督微调与自我调节奖励驱动对齐,提升了模型在细粒度视觉感知任务中的性能和推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10798 2026-07-14 cs.CL 新提交 79%

Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG

融合前的信任:用于受污染多模态RAG的QIMG-7和源感知分辨率

Saadeldine Eletter, Owais Aijaz, Preslav Nakov

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 研究多模态检索增强生成中受污染内容问题,提出QIMG-7基准。针对朴素多模态融合脆弱问题,提出源感知信任分辨率(SATR),Field-Selector变体效果最佳,结果支持选择性信任,显式文本可靠性建模是收益主要驱动因素。

Comments 23 pages, 6 figures, 23 tables. Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10599 2026-07-14 cs.AI eess.SP 新提交 79%

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

MRUF:用于鲁棒多模态情感分析的具有不确定性感知融合的多粒度路由

Haoran Ma, Yinfeng Yu, Liejun Wang

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语言认知计算国际联合研究实验室) Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center(新疆多模态智能处理与信息安全工程技术研究中心) Pengcheng Laboratory Xinjiang Network Node(鹏城实验室新疆网络节点) Embodied Intelligence Joint Laboratory(具身智能联合实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 针对多模态情感分析中模态质量不一问题,提出MRUF方法,结合多粒度路由与不确定性感知校准,总结情感表示、进行路由并监督模态重要性估计,预测模态不确定性并细化模态门,实验显示相比基线有改进,验证了高不确定性模态获低融合权重。

Comments Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems and Man and and Cybernetics 2026 (IEEE SMC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏