arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-19 至 2026-01-19 共收录 35 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2407.20114 2026-01-19 cs.IR cs.AI cs.CV 84%

FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis

FiCo-ITR:弥合细粒度与粗粒度图像-文本检索以进行性能比较分析

Mikel Williams-Lekuona, Georgina Cosma

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 FiCo-ITR通过标准化评估方法,对比细粒度与粗粒度图像-文本检索模型的性能与效率权衡,为模型选择和混合系统研究提供依据。

Comments Published at the International Journal of Multimedia Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11243 2026-01-19 cs.CV 79%

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

图像-文本知识建模用于无监督多场景人物重识别

Zhiqi Pang, Lingling Zhao, Yang Liu, Chunyu Wang, Gaurav Sharma

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

AI总结 本文提出图像-文本知识建模用于无监督多场景人物重识别,通过三阶段框架提升跨场景识别性能。

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11464 2026-01-19 cs.CV cs.AI cs.CL cs.LG 67%

MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models

MHA2MLA-VLM: 使DeepSeek的经济型多头潜在注意力在视觉-语言模型中生效

Xiaoran Fan, Zhichao Sun, Tao Ji, Lixing Shen, Tao Gui

机构 * DeepSeek

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MHA2MLA-VLM通过参数高效的方法将现有视觉-语言模型转换为多头潜在注意力架构,减少缓存足迹并提升推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10835 2026-01-19 cs.CV cs.AI 62%

Can Vision-Language Models Understand Construction Workers? An Exploratory Study

视觉-语言模型能理解建筑工人吗?一项探索性研究

Hieu Bui, Nathaniel E. Chodosh, Arash Tavakoli

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Villanova University(维拉诺瓦大学) Department of Computing Sciences(计算科学系) Department of Civil and Environmental Engineering(土木与环境工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究评估了三种视觉-语言模型在识别建筑工人行为和情绪方面的性能,发现GPT-4o表现最佳,但需进一步改进以提高实际应用可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2506.05879 2026-01-19 cs.HC 82%

Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions

与语言病理学家在亲子互动中的人机对齐多模态大语言模型

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 本研究通过多模态大语言模型与语言病理学家的对齐,开发了支持亲子互动分析的系统,实现了85%的感知线索提取准确率和75%的判断精确率,并提出了行为观察-判断系统的构建指南。

Comments This is an earlier version of the work released in May 2025. The version accepted at CHI 2026 is available as a separate preprint at arXiv:2511.04366

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01166 2026-01-19 cs.SD 78%

Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR

听更清晰,用更少:多模态检索与选择增强的对话式LLM基于ASR

Bingshen Mu, Hexin Liu, Hongfei Xue, Kun Wei, Lei Xie

专题命中 音频语音多模态 :multi-modal(title,abstract)

AI总结 本文提出MARS方法,通过多模态检索与选择提升对话式LLM-ASR的准确性,仅用1.5K小时数据即可超越训练于179K小时数据的顶级系统。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12537 2026-01-19 cs.CL cs.AI eess.AS 67%

What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study

什么使LLM为中心的语音生成中的良好语音分词器?系统研究

Xiaoran Fan, Zhichao Sun, Yangfan Gao, Jingfei Xiong, Hang Yan, Yifei Cao, Jiajun Sun, Shuo Li, Zhihao Zhang, Zhiheng Xi, Yuhao Zhou, Senjie Jin, Changhao Jiang, Junjie Ye, Ming Zhang, Rui Zheng, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Ji, Tao Gui

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文研究了LLM为中心的语音生成中语音分词器设计的影响,通过引入多令牌预测和说话人感知生成,提升了语音生成的质量和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11262 2026-01-19 cs.SD cs.IR cs.LG 50%

Scalable Music Cover Retrieval Using Lyrics-Aligned Audio Embeddings

基于歌词对齐音频嵌入的可扩展音乐封面检索

Joanne Affolter, Benjamin Martin, Elena V. Epure, Gabriel Meseguer-Brocal, Frédéric Kaplan

机构 * Deezer Research, Paris, France(DeepZoom研究机构,法国巴黎) EPFL, Lausanne, Switzerland(瑞士洛桑联邦理工学院)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 LIVI通过利用歌词信息提升音乐封面检索的准确性和效率,减少对复杂模型的依赖。

Comments Published at ECIR 2026 (European Conference of Information Retrieval)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2601.11359 2026-01-19 cs.CV cs.AI 62%

Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding

Think-Clip-Sample: 慢-快帧选择用于视频理解

Wenhui Tan, Ruihua Song, Jiaze Li, Jianzhong Ju, Zhenbo Luo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 Think-Clip-Sample通过多查询推理和慢-快采样提升长视频理解的效率和效果

Comments Accepted by ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11220 2026-01-19 cs.CL 57%

MultiCaption: Detecting disinformation using multilingual visual claims

多 caption:利用多语言视觉主张检测虚假信息

Rafael Martins Frade, Rrubaa Panchendrarajan, Arkaitz Zubiaga

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 MultiCaption通过多语言视觉主张数据集提升多模态虚假信息检测性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02937 2026-01-19 cs.LG cs.AI 57%

Towards Explainable Traffic Flow Prediction with Large Language Models

面向大语言模型的可解释交通流预测

Xusen Guo, Qiming Zhang, Junyue Jiang, Mingxing Peng, Meixin Zhu, Hao, Yang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) Johns Hopkins University(约翰霍普金斯大学) Department of Civil and System Engineering(土木与系统工程系)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出xTP-LLM模型,利用大语言模型生成可解释的交通流预测,首次将LLM应用于交通预测的可解释性研究。

Comments 31pages, 16 figures

Journal ref Communications in Transportation Research, vol. 4, 100150, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10986 2026-01-19 cs.LG 50%

StellarF: A Physics-Informed LoRA Framework for Stellar Flare Forecasting with Historical & Statistical Data

StellarF:一种结合物理知识的LoRA框架,用于利用历史和统计数据进行恒星耀斑预测

Tianyu Su, Zhiqiang Zou, Qingyu Lu, Feng Zhang, Ali Luo, Xiao Kong, Min Li

机构 * School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机科学学院) Jiangsu Key Laboratory of Big Data Security and Intelligent Processing(江苏大数据安全与智能处理重点实验室) University of Chinese Academy of Sciences(中国科学院大学) CAS Key Laboratory of Optical Astronomy, National Astronomical Observatories(中国科学院国家天文台光学天文重点实验室) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学学院)

专题命中 视频多模态 :multimodal(abstract)

AI总结 StellarF通过结合物理知识和LoRA框架,利用历史和统计数据提升恒星耀斑预测的准确性与物理可解释性。

Comments 12 pages, 8 figures (5 main, 3 appendix), 7 tables (2 main, 5 appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2601.11151 2026-01-19 cs.IR cs.AI 88%

Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation

跨模态注意力网络与双图学习在多模态推荐中的应用

Ji Dai, Quan Fang, Jun Hu, Desheng Cai, Yang Yang, Can Zhao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学) Tianjin University of Technology(天津工业大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管重大科技专项实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 CRANE通过双图学习和递归注意力机制,解决多模态推荐中的浅层融合和不对称特征处理问题,提升推荐性能。

Comments Accepted to ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11006 2026-01-19 cs.LG 78%

Backdoor Attacks on Multi-modal Contrastive Learning

多模态对比学习中的后门攻击

Simi D Kuniyilh, Rita Machacy

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract)

AI总结 本文研究了多模态对比学习中的后门攻击问题,分析了攻击方法、威胁模型及防御措施,并探讨了该领域的发展挑战与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11028 2026-01-19 cs.LG 71%

AVP-Pro: An Adaptive Multi-Modal Fusion and Contrastive Learning Approach for Comprehensive Two-Stage Antiviral Peptide Identification

AVP-Pro: 一种自适应多模态融合与对比学习方法用于全面两阶段抗病毒肽鉴定

Xinru Wen, Weizhong Lin, zi liu, Xuan Xiao

专题命中 跨模态检索 :multi-modal(title)

AI总结 AVP-Pro通过自适应多模态融合和对比学习方法,实现了抗病毒肽的高效两阶段鉴定,提升了模型判别能力和小样本下的分类精度。

Comments arXiv admin note: substantial text overlap with arXiv:2512.21544

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2506.12198 2026-01-19 cs.CV 83%

ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

ViSTA: 基于多模态适配器的文本到图像扩散模型用于视觉叙事

Sibo Dong, Ismail Shaheen, Maggie Shen, Rupayan Mallick, Sarah Adel Bargal

机构 * Department of Computer Science, Georgetown University(计算机科学系,乔治城大学)

专题命中 多模态生成 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 ViSTA通过多模态历史适配器提升文本到图像扩散模型在视觉叙事中的生成一致性与文本对齐能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09145 2026-01-19 cs.LG cs.CL cs.CV 81%

MoLAN: A Unified Modality-Aware Noise Dynamic Editing Framework for Multimodal Sentiment Analysis

MoLAN: 一种统一的多模态感知噪声动态编辑框架用于多模态情感分析

Xingle Xu, Yongkang Liu, Dexian Cai, Shi Feng, Xiaocui Yang, Daling Wang, Yifei Zhang

机构 * Northeastern University, China(东北大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 MoLAN提出了一种统一的多模态感知噪声动态编辑框架,通过动态分配去噪强度以提升多模态情感分析的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10527 2026-01-19 cs.AI cs.CL cs.CV cs.LG 67%

A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

GPT-5.2、Gemini 3 Pro、Qwen3-VL、Grok 4.1 Fast、Nano Banana Pro 和 Seedream 4.5 的安全报告

Xingjun Ma, Yixu Wang, Hengyuan Xu, Yutao Wu, Yifan Ding, Yunhan Zhao, Zilong Wang, Jiabin Hua, Ming Wen, Jianan Liu, Ranjie Duan, Yifeng Gao, Yingshui Tan, Yunhao Chen, Hui Xue, Xin Wang, Wei Cheng, Jingjing Chen, Zuxuan Wu, Bo Li, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Innovation institute(上海创新研究院) Deakin University(德克萨斯大学) UIUC(伊利诺伊大学香槟分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本报告对六个前沿模型进行综合安全性评估,揭示其在安全性和鲁棒性方面的差异,强调需要标准化评估以指导负责任的部署。

Comments 41 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11522 2026-01-19 cs.CV 57%

UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation

UniX: 统一自回归与扩散以实现胸部X光理解与生成

Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao, Hao Yan, Xiao He, Lei Chen, Zhou Wei, Yong Luo, Zengmao Wang, Lefei Zhang, Dacheng Tao, Bo Du

机构 * Wuhan University(武汉大学) Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 UniX通过统一自回归与扩散模型,实现了胸部X光的高效理解和生成,显著提升了性能并减少了参数使用。

Comments Codes and models are available at https://github.com/ZrH42/UniX

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11194 2026-01-19 cs.CV 57%

ATATA: One Algorithm to Align Them All

ATATA:一切对齐的统一算法

Boyi Pang, Savva Ignatyev, Vladimir Ippolitov, Ramil Khafizov, Yurii Melnik, Oleg Voynov, Maksim Nakhodnov, Aibek Alanov, Xiaopeng Fan, Peter Wonka, Evgeny Burnaev

机构 * Harbin Institute of Technology(哈尔滨工业大学) Applied AI Institute(应用人工智能研究所) AXXX FusionBrain Lab(融合大脑实验室) MSU(莫斯科州立大学) Constructor University(建设大学) HSE University(高等经济大学) Peng Cheng Laboratory(鹏城实验室) HIT Suzhou Research Institute(哈尔滨工业大学苏州研究 institute) KAUST(科威特大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 ATATA提出一种基于Rectified Flow模型的多模态算法,实现高效结构对齐的样本生成,提升图像、视频和3D生成的质量与速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10781 2026-01-19 cs.CV 57%

Future Optical Flow Prediction Improves Robot Control & Video Generation

未来光流预测改进机器人控制与视频生成

Kanchana Ranasinghe, Honglu Zhou, Yu Fang, Luyu Yang, Le Xue, Ran Xu, Caiming Xiong, Silvio Savarese, Michael S Ryoo, Juan Carlos Niebles

机构 * Salesforce AI Research(Salesforce AI研究院) Stony Brook University(石溪大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 FOFPred通过统一的视觉-语言模型和扩散架构,实现高效的未来光流预测,提升机器人控制与视频生成的性能。

Comments Project Site (Code, Models, Demo): https://fofpred.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02358 2026-01-19 cs.CV 57%

VINO: A Unified Visual Generator with Interleaved OmniModal Context

VINO:一个具有交错多模态上下文的统一视觉生成器

Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 VINO通过统一的视觉生成框架实现图像和视频的生成与编辑,利用交错多模态上下文条件处理,提升多任务视觉创作能力。

Comments Project page: https://sotamak1r.github.io/VINO-web/

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 4 篇

2312.03543 2026-01-19 cs.CV cs.AI 88%

GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models

GPT-4增强的多模态接地用于自动驾驶:利用跨模态注意力与大语言模型

Haicheng Liao, Huanming Shen, Zhenning Li, Chengyue Wang, Guofa Li, Yiming Bie, Chengzhong Xu

机构 * Univ City(大学城市)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种结合GPT-4的大语言模型的多模态接地框架,用于提升自动驾驶中的视觉理解和指令执行能力。

Journal ref Communications in Transportation Research 4 (2024) 100116

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08779 2026-01-19 cs.CV 79%

BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models

BBQ-V:评估大型多模态模型中视觉刻板印象偏见的基准

Vishal Narnaware, Ashmal Vayani, Rohit Gupta, Sirnam Swetha, Mubarak Shah

机构 * Institute of Artificial Intelligence, University of Central Florida(人工智能研究所,中央佛罗里达大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 BBQ-V通过真实多演员图像和多样类别评估LMMs的刻板印象偏见,揭示了主流模型在推理链中的偏见问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10962 2026-01-19 cs.CV 79%

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

V2X-Radar: 一种包含4D雷达的多模态数据集用于协作感知

Lei Yang, Xinyu Zhang, Jun Li, Chen Wang, Jiaqi Ma, Zhiying Song, Tong Zhao, Ziying Song, Li Wang, Mo Zhou, Yang Shen, Kai Wu, Chen Lv

机构 * School of Vehicle and Mobility, Tsinghua University(车辆与移动性学院,清华大学) Nanyang Technological University(南洋理工大学) CUMTB(中国交通车辆技术研究所) University of California, Los Angeles(加州大学洛杉矶分校) Beijing Jiaotong University(北京交通大学) ByteDance(字节跳动)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 V2X-Radar是首个包含4D雷达的多模态数据集,用于提升自动驾驶中的协作感知能力。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06404 2026-01-19 cs.CV 57%

InfoAffect: Affective Annotations of Infographics in Information Spread

InfoAffect: 信息传播中信息图的情感标注

Zihang Fu, Yunchao Wang, Chenyu Huang, Guodao Sun, Ronghua Liang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 InfoAffect数据集通过多模态大语言模型和递归排名融合算法,实现了信息图情感标注的高准确度研究。

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2510.05611 2026-01-19 cs.CL cs.AI 73%

MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction

MADIAVE:多智能体辩论用于隐式属性值提取

Wei-Chieh Huang, Cornelia Caragea

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

AI总结 MADIAVE通过多智能体辩论框架提升隐式属性值提取的准确性和鲁棒性,适用于多模态电子商务场景。

Comments Accepted by EACL 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01439 2026-01-19 cs.AI cs.RO 57%

Probabilistic Mission Design for Neuro-Symbolic Unmanned Aircraft Systems

基于概率的方法用于神经符号无人机系统任务设计

Simon Kohaut, Benedict Flade, Daniel Ochs, Devendra Singh Dhami, Julian Eggert, Kristian Kersting

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出概率任务设计(ProMis)方法,结合神经符号技术,用于在法律框架内导航无人驾驶航空器,通过混合概率逻辑程序和机器学习模型提升任务规划的可靠性和适应性。

Comments arXiv admin note: text overlap with arXiv:2406.03454

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 4 篇

2601.09105 2026-01-19 cs.AI cs.CL cs.CV 90%

AviationLMM: A Large Multimodal Foundation Model for Civil Aviation

AviationLMM:民用航空的大规模多模态基础模型

Wenbin Li, Jingling Wu, Xiaoyong Lin. Jing Chen, Cong Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 AviationLMM旨在通过统一民用航空的异构数据流,提升航空安全与效率,推动可信且隐私保护的航空人工智能生态系统发展。

Comments Accepted by 2025 7th International Conference on Interdisciplinary Computer Science and Engineering (ICICSE 2025), Chongqing, China; 9 pages,1 figure,5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07399 2026-01-19 cs.IR 78%

Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities

多模态嵌入真的能提升推荐效果吗?对整体与个体模态的深入探究

Yu Ye, Junchen Fu, Yu Song, Kaiwen Zheng, Joemon M. Jose

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 研究探讨多模态嵌入对推荐效果的实际影响,发现整体嵌入提升效果有限,但文本模态单独使用效果显著,图像模态则不具优势。

Comments Accepted by ECIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏