arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-03 至 2026-02-03 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 9 篇

2506.09114 2026-02-03 cs.LG 82%

TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval

TRACE: 在上下文中对时间序列进行建模以实现多模态嵌入与检索

Jialin Chen, Ziyu Zhao, Gaukhar Nurbek, Aosong Feng, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, Rex Ying

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 TRACE通过在上下文中对时间序列进行建模,实现多模态嵌入与检索,提升下游任务的预测精度和可解释性,同时作为强大的独立编码器优化上下文感知表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02004 2026-02-03 cs.CV cs.AI 81%

ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理

Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) West China Biomedical Big Data Center(西京生物大数据中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ClueTracer通过问题到视觉线索追踪,无训练抑制多模态推理中的幻觉,提升推理和非推理任务性能。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01561 2026-02-03 cs.CV cs.AI 81%

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

多模态UNcommonsense:从奇特到普通和从普通到奇特

Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu

机构 * Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 多模态UNcommonsense通过R-ICL框架提升模型在非典型场景下的推理能力,实现从奇特到普通和从普通到奇特的转换。

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01594 2026-02-03 cs.CV 79%

UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception

UV-M3TL: 一种统一且多功能的多模态多任务学习框架用于辅助驾驶感知

Wenzhuo Liu, Qiannan Guo, Zhen Wang, Wenshuo Wang, Lei Yang, Yicheng Qiao, Lening Wang, Zhiwei Li, Chen Lv, Shanghang Zhang, Junqiang Xi, Huaping Liu

机构 * Energy and Transportation Domain, Beijing Institute of Technology(能源与交通领域,北京理工大学) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室和清华大学计算机科学与技术系) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学) School of Transportation Science and Engineering and the State Key Lab of Intelligent Transportation System, Beihang University(交通运输科学与工程学院和智能交通系统国家重点实验室,北京航空航天大学) Beijing University of Chemical Technology(北京化工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UV-M3TL通过双分支结构和自适应损失机制,实现多模态多任务学习,提升辅助驾驶感知的性能与多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21033 2026-02-03 cs.SD cs.AI 70%

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

SupCLAP:通过支持向量正则化控制音频-文本对比学习中的优化轨迹漂移

Jiehui Luo, Yuguo Yin, Yuxin Xie, Jinghan Ru, Xianwei Zhuang, Minghua He, Aofan Liu, Zihan Xiong, Dongchao Yang

机构 * Peking University(北京大学) Central Conservatory of Music(中央音乐学院) The Chinese University of Hong Kong(香港中文大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 SupCLAP通过支持向量正则化有效控制音频-文本对比学习中的优化轨迹漂移,提升多模态学习的稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00131 2026-02-03 cs.CV cs.RO 70%

PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living

PovNet+: 一种深度学习架构用于社交辅助机器人学习和协助多种日常活动

Fraser Robinson, Souren Pashangpour, Matthew Lisondra, Goldie Nejat

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 PovNet+是一种多模态深度学习架构,用于社交辅助机器人识别多种日常活动并主动发起辅助行为,提升了ADL分类准确率和人机交互能力。

Comments Submitted to Advanced Robotics (Taylor & Francis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00621 2026-02-03 cs.CV 57%

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

通过对比神经元引导实现大型视觉语言模型中幻觉分析与缓解

Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University, Xi'an, China(西安电子科技大学电子工程学院) School of Computer Science and Technology, Xidian University, Xi'an, China(西安电子科技大学计算机科学与技术学院) College of Computer and Information, Hohai University, Nanjing, China(河海大学计算机与信息学院) Institute for Infocomm Research, A*STAR, Singapore(新加坡资讯研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 通过对比神经元引导方法,分析并缓解大型视觉语言模型中的幻觉问题,提升视觉表示的稳健性和语义基础性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00547 2026-02-03 cs.LG cs.AI 57%

Contrastive Domain Generalization for Cross-Instrument Molecular Identification in Mass Spectrometry

对比域泛化用于质谱中跨仪器分子识别

Seunghyun Yoo, Sanghong Kim, Namkyung Yoon, Hwangnam Kim

机构 * School of Electrical Engineering, Korea University, Seoul 02841, Korea(韩国大学电气工程学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种跨模态对齐框架,通过将质谱直接映射到预训练化学语言模型的分子结构嵌入空间,提升跨仪器分子识别的泛化能力。

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05266 2026-02-03 cs.AR cs.CL cs.LG 57%

Understanding and Mitigating Errors of LLM-Generated RTL Code

理解并缓解LLM生成的RTL代码错误

Jiazheng Zhang, Cheng Liu, Long Cheng, Xiaowei Li, Huawei Li

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出基于LLM的框架,通过检索增强生成、规则检查、多模态转换和迭代仿真调试,显著提升了RTL代码生成的准确性。

Comments Accepted by IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏