arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-02 至 2025-12-02 共收录 112 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2512.01949 2025-12-02 cs.CV 79%

Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models

脚本:图结构和查询条件的语义令牌修剪用于多模态大语言模型

Zhongyu Yang, Dannong Xu, Wei Pang, Yingfang Yuan

机构 * BCML, Heriot-Watt University(赫瑞瓦德大学BCML中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 Script通过图结构和查询条件的语义令牌修剪,提升多模态大语言模型的效率和准确性,实现显著的性能提升。

Comments Published in Transactions on Machine Learning Research, Project in https://01yzzyu.github.io/script.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20469 2025-12-02 q-bio.QM cs.CV 79%

Prediction of Distant Metastasis in Head and Neck Cancer Patients Using Tumor and Peritumoral Multi-Modal Deep Learning

利用肿瘤及周围多模态深度学习预测头颈癌患者远端转移

Nuo Tong, Changhao Liu, Zizhao Tang, Feifan Sun, Yingping Li, Shuiping Gou, Mei Shi

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本研究提出多模态深度学习模型,结合CT影像、放射组学和临床数据,用于预测头颈癌患者远端转移风险,通过多模态融合显著提高预测性能。

Comments 23 pages, 6 figures, 7 tables. Nuo Tong and Changhao Liu contributed equally. Corresponding Authors: Shuiping Gou and Mei Shi

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01410 2025-12-02 cs.CL 79%

DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis

DyFuLM:一种用于情感分析的先进多模态框架

Ruohan Zhou, Jiachen Yuan, Churui Yang, Wenzheng Huang, Guoyan Zhang, Shiyao Wei, Jiazhen Hu, Ning Xin, Md Maruf Hasan

机构 * Department of Applied Mathematics, Xi'an Jiaotong-Liverpool University(应用数学系,西安交通大学-利物浦大学) School of AI and Advanced Computing, XJTLU Entrepreneur College (Taicang)(人工智能与先进计算学院,XJTLU创业学院(太仓)) Department of Intelligent Science, Xi'an Jiaotong-Liverpool University(智能科学系,西安交通大学-利物浦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 DyFuLM通过动态融合和门控聚合模块提升多模态情感分析的准确率与稳定性

Comments 8 pages, 6 figures, preprint. Under review for a suitable AI conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00596 2025-12-02 cs.IR cs.AI 79%

DLRREC: Denoising Latent Representations via Multi-Modal Knowledge Fusion in Deep Recommender Systems

DLRREC: 通过深度融合多模态知识在深度推荐系统中进行潜在表示去噪

Jiahao Tian, Zhenkai Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 DLRREC通过深度融合多模态和协同知识,提升深度推荐系统中潜在表示的去噪能力,从而实现更精确的推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01750 2025-12-02 eess.SP cs.LG 78%

Multimodal Mixture-of-Experts for ISAC in Low-Altitude Wireless Networks

多模态专家混合模型用于低空无线网络中的ISAC

Kai Zhang, Wentao Yu, Hengtao He, Shenghui Song, Jun Zhang, Khaled B. Letaief

机构 * IEEE

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种多模态专家混合模型,用于提升低空无线网络中ISAC的性能,通过自适应融合策略提高环境感知和通信效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00042 2025-12-02 cs.CV cs.AI cs.CL cs.CY 67%

Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions

弥合差距:面向标准化考试题目的视觉语言模型数据驱动微调

Egemen Sert, Şeyda Ertekin

机构 * organization= Department of Computer Engineering, Middle East Technical University (METU) , city= Ankara , country= Türkiye organization= METU-DTX Digital Transformation \& Innovation Centre, METU , city= Ankara , country= Türkiye

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究通过高质量数据和优化语法提升视觉语言模型在标准化考试题目的多模态推理性能,达到接近SOTA的水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10194 2025-12-02 cs.CV 57%

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07062 2025-12-02 cs.AI 57%

Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision

通过噪声长描述监督提升城市影像区域表示学习

Yimei Zhang, Guojiang Shen, Kaili Ning, Tongwei Ren, Xuebo Qiu, Mengmeng Wang, Xiangjie Kong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出UrbanLN框架,通过长文本意识和噪声抑制提升城市影像区域表示学习,有效解决细粒度特征对齐与噪声干扰问题。

Comments Accepted as a full paper by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00293 2025-12-02 cs.LG cs.AI 57%

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

FiCoTS: 细到粗的LLM增强层次跨模态交互用于时间序列预测

Yafei Lyu, Hao Zhou, Lu Zhang, Xu Yang, Zhiyong Liu

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所MAIS) Great Bay University(大亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 FiCoTS通过细到粗的LLM增强层次跨模态交互框架,提升多模态时间序列预测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01358 2025-12-02 cs.RO cs.LG 50%

Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1

模态增强的基座机器人策略微调用于GR1和G1跨躯体操作

Junsung Park, Hogun Kee, Songhwai Oh

机构 * Department of Electrical and Computer Engineering, Seoul National University(电气与计算机工程系,首尔国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出了一种模态增强的微调方法,通过多模态数据提升机器人策略在不同躯体上的性能,显著提高了任务成功率。

Comments 8 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00063 2025-12-02 q-bio.NC 50%

Exploring the changes in brain network SC-FC coupling patterns of partial sleep deprivation based on DTI-fMRI fusion analysis

基于DTI-fMRI融合分析探讨部分睡眠剥夺对脑网络SC-FC耦合模式变化的影响

Mengyuan Liu, Jing Hu, Zhenzhen Ru, Ruomeng Quan, Xu Zhang, Ning Qiang, Jin Li

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本研究通过DTI-fMRI融合分析,探讨部分睡眠剥夺对脑网络SC-FC耦合模式的影响,发现睡眠剥夺导致神经网络结构和功能连接的异常,提出SC-FC耦合方法作为睡眠相关情绪失调的新型生物标记物。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 11 篇

2504.05187 2025-12-02 cs.NI cs.AI cs.LG 83%

Resource-Efficient Beam Prediction in mmWave Communications with Multimodal Realistic Simulation Framework

毫米波通信中基于多模态真实模拟框架的高效波束预测

Yu Min Park, Yan Kyaw Tun, Eui-Nam Huh, Walid Saad, Choong Seon Hong

机构 * Department of Computer Science and Engineering, Kyung Hee University(计算机科学与工程系,庆尚大学) Bradley Department of Electrical and Computer Engineering, Virginia Tech(电气与计算机工程系,弗吉尼亚理工大学) Department of Electronic Systems, Aalborg University(电子系统系,奥尔堡大学)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种基于多模态真实模拟框架的资源高效学习框架,通过跨模态关系知识蒸馏算法,利用雷达数据实现高精度波束预测,同时降低计算复杂性和对多模态传感器数据的依赖。

Comments 13 pages, 9 figures, Submitted to IEEE Transactions on Mobile Computing on Dec. 01, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10044 2025-12-02 cs.CR cs.AI 79%

Large Language Models for Power System Security: A Novel Multi-Modal Approach for Anomaly Detection in Energy Management Systems

用于电力系统安全的大型语言模型:一种用于能源管理系统异常检测的新型多模态方法

Aydin Zaboli, Junho Hong, Alexandru Stefanov, Chen-Ching Liu, Chul-Sang Hwang

机构 * Department of Electrical and Computer Engineering, University of Michigan -- Dearborn, MI, 48128 USA.(电气与计算机工程系,密歇根大学迪尔伯恩分校) Department of Electrical Sustainable Energy, Technische Universiteit Delft, 2628 CD Delft, Netherlands.(可持续能源系,代尔夫特理工大学) Bradley Department of Electrical and Computer Engineering, Virginia Polytechnic Institute and State University, Blacksburg, VA 24061, USA.(布雷德利电气与计算机工程系,弗吉尼亚理工学院和州立大学) Smart Grid Research Division System Reliability Research Team, Korea Electrotechnology Research Institute (KERI), Gwangju-si, 61751, South Korea.(智能电网研究分会系统可靠性研究团队,韩国电力技术研究所(KERI))

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于大型语言模型的多模态方法,用于电力系统中能源管理系统的安全防护和异常检测。

Comments 10 Figures; 6 Tables; Accepted, IEEE ACCESS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15858 2025-12-02 cs.IR cs.LG 78%

Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings

利用多模态嵌入优化电子商务中的产品去重

Aysenur Kulunk, Berk Taskin, M. Furkan Eseoglu, H. Bahadir Sahin

机构 * Data Analytics Department Hepsiburada Istanbul, Turkey(数据分析师部门 赫斯伯达 特拉布宗 土耳其)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究提出了一种基于多模态嵌入的电子商务产品去重方法,结合BERT和MaskedAutoEncoders,利用Milvus数据库实现高效高精度的相似性搜索,F1分数达0.90,优于现有解决方案。

Comments 8 pages, accepted to 2025 IEEE International Conference on Big Data, Industrial and Goverment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19476 2025-12-02 cs.RO 71%

Gentle Object Retraction in Dense Clutter Using Multimodal Force Sensing and Imitation Learning

在密集障碍物中使用多模态力感知和模仿学习实现温和的对象回退

Dane Brouwer, Joshua Citron, Heather Nolte, Jeannette Bohg, Mark Cutkosky

机构 * Department of Mechanical Engineering, Stanford University, USA(机械工程系,斯坦福大学) Department of Computer Science, Stanford University, USA(计算机科学系,斯坦福大学)

专题命中 其他多模态 :multimodal(title)

AI总结 本研究通过多模态力感知和模仿学习,实现机器人在密集障碍物中温和地提取物体,显著提升成功率和效率。

Comments Accepted in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00019 2025-12-02 cs.RO cs.AI cs.CV 62%

A Comprehensive Survey on Surgical Digital Twin

外科数字孪生的全面综述

Afsah Sharaf Khan, Falong Fan, Doohwan DH Kim, Abdurrahman Alshareef, Dong Chen, Justin Kim, Ernest Carter, Bo Liu, Jerzy W. Rozenblit, Bernard Zeigler

机构 * IEEE

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了外科数字孪生的技术现状与挑战,提出分类方法并识别了验证、安全性和数据治理等开放问题,旨在推动其在临床中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01262 2025-12-02 cs.SI cs.AI cs.ET cs.LG 57%

Social Media Data Mining of Human Behaviour during Bushfire Evacuation

社交媒体中人类在森林火灾疏散中的行为挖掘

Junfeng Wu, Xiangmin Zhou, Erica Kuligowski, Dhirendra Singh, Enrico Ronchi, Max Kinateder

机构 * RMIT University(皇家墨尔本理工大学) CSIRO(澳大利亚国家科学研究院) Lund University(吕勒奥大学) National Research Council Canada(加拿大国家研究理事会)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了利用社交媒体数据挖掘森林火灾疏散行为的挑战与方法,提出未来应用及开放问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07820 2025-12-02 cs.AI cs.LG 57%

AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift

AI应更善于感知,而非仅仅更大:适应性感知作为范式转变

Eunsu Baek, Keondo Park, Jeonggil Ko, Min-hwan Oh, Taesik Gong, Hyung-Sin Kim

机构 * SNU(首尔国立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出适应性感知作为AI范式转变,通过动态调节传感器参数以提升效率和公平性,推动可持续且稳健的人工智能发展。

Comments Published in NeurIPS 2025 (Position Paper Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00366 2025-12-02 cs.LG cs.AI 57%

S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting

S^2-KD:语义-频谱知识蒸馏时空预测

Wenshuo Wang, Yaomin Shen, Yingjie Tan, Yihao Chen

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 S^2-KD通过结合语义与频谱知识蒸馏,提升时空预测模型的性能,使其在复杂场景中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22960 2025-12-02 cs.NE cond-mat.other 50%

Hybrid Particle Swarm Optimization for Fast and Reliable Parameter Extraction in Thermoreflectance

混合粒子群优化用于热反射法中快速可靠的参数提取

Bingjia Xiao, Tao Chen, Wenbin Zhang, Xin Qian, Puqing Jiang

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出混合粒子群优化算法,用于提高热反射法中参数提取的收敛速度和准确性,HPSO在60秒内达到目标精度的试验比例达80%,表现优于其他方法。

Comments 28 pages, 8 figures

Journal ref International Journal of Heat and Mass Transfer 256, 128109 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00835 2025-12-02 cs.LG 50%

Uncertainty Quantification for Deep Regression using Contextualised Normalizing Flows

基于上下文化规范化流的深度回归不确定性量化

Adriel Sosa Marco, John Daniel Kirwan, Alexia Toumpa, Simos Gerasimou

机构 * Arquimea Research Center(阿奎米亚研究中心) Department of Computer Science, University of York(约克大学计算机科学系) Department of Elect. Eng., and Computer Science and Eng., Cyprus University of Technology(塞浦路斯技术大学电子工程与计算机科学系)

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出MCNF方法,通过上下文化规范化流实现深度回归模型的不确定性量化,生成预测区间和完整条件分布,无需重新训练模型,提升高风险领域的决策可靠性。

Journal ref Neural Information Processing Systems (NeurIPS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00090 2025-12-02 q-bio.NC 50%

Biomimetic Metamaterial-based Interface for Decoding Heterogeneous Mechanodermal Activity

仿生元材料基于接口用于解码异质机械皮肤活动

Muzi Xu, Jiaqi Zhang, Chaoqun Dong, Zibo Zhang, Duanyang Li, Wentian Yi, Miaomiao Zou, Chenyu Tang, George G. Malliaras, Luigi G. Occhipinti

专题命中 其他多模态 :multimodal(abstract)

AI总结 本研究提出一种仿生元材料接口,用于解码异质机械皮肤活动,通过捕捉和解码多种生理状态,应用于医疗监测和人机交互。

Comments 26 pages, 5 figures, 55 references

详情

展开后加载摘要…

URL PDF HTML 收藏