arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2604.12081 2026-04-15 cs.AI 79%

Human-Inspired Context-Selective Multimodal Memory for Social Robots

以人为本的上下文选择性多模态记忆用于社交机器人

Hangyeol Kang, Slava Voloshynovskiy, Nadia Magnenat Thalmann

机构 * Department of Computer Science, University of Geneva(日内瓦大学计算机科学系) MIRALab, University of Geneva(日内瓦大学MIRALab)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种以人为本的多模态记忆架构,用于社交机器人,通过捕捉和检索文本和视觉事件痕迹,提升个性化和上下文感知的交互能力。

Comments Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11095 2026-04-14 cs.LG cs.AI 79%

Bottleneck Tokens for Unified Multimodal Retrieval

瓶颈令牌用于统一多模态检索

Siyu Sun, Jing Ren, Zhaohe Liao, Dongxiao Mao, Xiangyuan Ren, Yiyi Zhang, Haohua Zhao, Weixiong Lin, Jiang Shaohua, Liqing Zhang, Yuchao Zheng

机构 * Shanghai Jiao Tong University(上海交通大学) ByteDance(字节跳动)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出瓶颈令牌和生成信息压缩方法,解决多模态大语言模型在统一检索中的隐式池化和对比微调问题,提升语义压缩效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10071 2026-04-14 cs.CV 79%

Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation

聚光与阴影:基于注意力的双锚点反思解码用于MLLM幻觉缓解

Yebo Wu, Han Jin, Zhijiang Guo, Li Li

机构 * State Key Laboratory of IOTSC, University of Macau(澳门大学物联网国家重点实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 跨模态检索 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出双锚点反思解码框架DaID,通过动态校准每个token生成来缓解MLLM幻觉,利用视觉注意力分布指导双锚点选择,提升推理能力。

Comments Accepted for Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07863 2026-04-10 cs.IR cs.AI 79%

Task-Adaptive Retrieval over Agentic Multi-Modal Web Histories via Learned Graph Memory

基于学习图记忆的代理多模态网络历史任务自适应检索

Saman Forouzandeh, Kamal Berahmand, Mahdi Jalili

机构 * School of Engineering, Royal Melbourne Institute of Technology University(皇家墨尔本理工大学工程学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出ACGM方法,通过任务自适应图记忆检索代理历史,利用策略梯度优化提升检索质量,在多个数据集上取得显著成果。

Comments The 49th International ACM SIGIR Conference on Research and Development in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06179 2026-04-09 cs.IR cs.CL 79%

ARIA: Adaptive Retrieval Intelligence Assistant -- A Multimodal RAG Framework for Domain-Specific Engineering Education

ARIA:自适应检索智能助手——面向领域特定工程教育的多模态RAG框架

Yue Luo, Dibakar Roy Sarkar, Rachel Herring Sangree, Somdatta Goswami

机构 * Dalian University of Technology(大连理工大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 ARIA通过多模态内容提取管道和e5-large-v2模型,实现领域特定工程教育的智能教学助手,展现高精度和教学一致性,验证了其在课程相关问题上的高准确率和响应质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29441 2026-04-09 cs.CV 79%

EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images

地球嵌入探索器:一种用于全球卫星图像跨模态检索的Web应用

Yijie Zheng, Weijie Wu, Bingyue Wu, Long Zhao, Guoqing Li, Mikolaj Czerkawski, Konstantin Klemmer

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) University of Chinese Academy of Sciences(中国科学院大学) Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences(中国科学院地理科学与资源研究所) Asterisk Labs(Asterisk实验室) LGND AI, Inc.(LGND AI公司) University College London(伦敦大学学院)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文介绍EarthEmbeddingExplorer,一种Web应用,通过跨模态查询实现全球卫星图像的动态检索,帮助研究人员将研究成果转化为实际应用。

Comments ICLR 2026 Workshop ML4RS Tutorial Track (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05773 2026-04-08 cs.CV 79%

PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization

PDMP:通过性能主导模态优先级重思平衡多模态学习

Shicai Wei, Chunbo Luo, Qiang Zhu, Yang Luo

机构 * Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China(电子科技大学智能协同计算实验室) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) Peng Cheng Laboratory(鹏城实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PDMP策略,通过性能主导模态优先级提升多模态学习效果,解决传统方法中因模态不平衡导致的优化不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05038 2026-04-08 cs.CL 79%

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization

引导查询细化:多模态混合检索与测试时优化

Omri Uzan, Asaf Yehudai, Roi pony, Eyal Shnarch, Ariel Gera

机构 * Stanford University(斯坦福大学) IBM Research(IBM研究院) The Hebrew University of Jerusalem(耶路撒冷希伯来大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出引导查询细化方法,通过测试时优化提升多模态检索性能,使视觉中心模型在效率和性能上达到更优平衡。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01247 2026-04-03 cs.SD eess.AS 79%

Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS

结合掩码语言建模与跨模态对比学习的语调感知语音合成

Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov, Nikita Vasiliev, Mikhail Gorodnichev, Grach Mkrtchian

机构 * MTUCI(莫斯科电信与信息技术大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 eess.AS

AI总结 本文研究了基于扩散模型的语音合成中多阶段预训练对语调建模的影响,通过掩码语言建模与混合音素对比学习相结合,提升了生成质量与感知指标。

Comments This paper has been submitted to Interspeech 2026 for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29376 2026-04-03 cs.CV 79%

Assessing Multimodal Chronic Wound Embeddings with Expert Triplet Agreement

基于专家三元组一致性的多模态慢性伤口嵌入评估

Fabian Kabus, Julia Hindel, Jelena Bratulić, Meropi Karakioulaki, Ayush Gupta, Cristina Has, Thomas Brox, Abhinav Valada, Harald Binder

机构 * Institute of Medical Biometry and Statistics (IMBI), Medical Faculty and Medical Center, University of Freiburg(弗莱堡大学医学院与医学中心医学统计与生物统计研究所) Department of Computer Science, Faculty of Engineering, University of Freiburg(弗莱堡大学工程学院计算机科学系) Department of Dermatology, Medical Faculty and Medical Center, University of Freiburg(弗莱堡大学医学院与医学中心皮肤科)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过专家三元组比较评估嵌入空间,引入TriDerm框架整合图像、边界掩码和专家报告,融合视觉与文本模态提升专家一致性至73.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03380 2026-04-03 cs.CV 79%

Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization

通过链看穿:通过CoT压缩和对比偏好优化缓解多模态推理模型的幻觉

Hao Fang, Jinyu Li, Jiawei Kong, Tianqu Zhuang, Kuofeng Gao, Bin Chen, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出C3PO框架,通过CoT压缩和对比偏好优化缓解多模态推理模型的幻觉问题,通过过滤冗余思考 token 提升信号效率,并利用高质量反馈和定制诱导器增强对比学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29631 2026-04-01 cs.CV cs.DC cs.IR 79%

Storing Less, Finding More: How Novelty Filtering Improves Cross-Modal Retrieval on Edge Cameras

存储更少,查找更多:新颖性过滤如何改进边缘摄像头上的跨模态检索

Sherif Abdelwahab

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种流式检索架构,通过边缘设备的epsilon-net过滤器保留语义新颖帧,构建去噪嵌入索引,并结合跨模态适配器和云重排序器,提升边缘摄像头跨模态检索性能。

Comments 6 pages, 3 figures, 5 tables; supplementary video included as ancillary file

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29259 2026-04-01 cs.IR cs.CL 79%

Aligning Multimodal Sequential Recommendations via Robust Direct Preference Optimization with Sparse MoE

通过鲁棒直接偏好优化与稀疏MoE对齐多模态序列推荐

Hejin Huang, Jusheng Zhang, Kaitong Cai, Jian Wang, Rong Pan

机构 * Sun Yat-sen University(中山大学) Snap Inc(Snap公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文研究了在隐式反馈下直接偏好优化的行为,通过系统实验比较了常见负样本选择策略及其与DPO训练的交互,发现用动态top-K候选池进行随机采样可提升排名性能,原因在于减少错误抑制梯度和保留信息硬信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26695 2026-03-31 eess.SP cs.AI cs.LG eess.IV quant-ph 79%

Complementarity-Preserving Generative Theory for Multimodal ECG Synthesis: A Quantum-Inspired Approach

保持互补性的多模态ECG生成理论:一种受量子启发的方法

Timothy Oladunni, Farouk Ganiyu-Adewumi, Clyde Baidoo, Kyndal Maclin

机构 * Morgan State University(摩根州立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出保持互补性的生成理论,通过量子启发框架Q-CFD-GAN生成多模态ECG数据,减少潜在嵌入方差并恢复三领域互补性,提升临床应用的生理意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26669 2026-03-31 cs.IR cs.AI cs.LG 79%

ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval

ReCQR:将对话查询重写纳入多模态图像检索以提高性能

Yuan Hu, ZhiYu Cao, PeiFeng Li, QiaoMing Zhu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出ReCQR任务,通过构建多轮对话查询重写数据集,提升多模态图像检索的准确性和用户查询建模能力。

Comments 4 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24736 2026-03-27 cs.AI cs.LG 79%

AutoSAM: an Agentic Framework for Automating Input File Generation for the SAM Code with Multi-Modal Retrieval-Augmented Generation

AutoSAM:一种用于自动化生成SAM代码输入文件的代理框架,结合多模态检索增强生成

Zaid Abulawi, Zavier Ndum Ndum, Eric Cervi, Rui Hu, Yang Liu

机构 * Department of Nuclear Engineering, Texas A\&M University. Engineering Division, Argonne National Laboratory

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 AutoSAM通过多模态检索增强生成技术,自动化生成SAM代码输入文件,解决异构工程文档中提取设计数据并转换为求解器语法的难题,实现100%结构化输入利用和88%PDF文本提取。

Comments 34 Pages, 14 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22821 2026-03-25 cs.CV 79%

Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference

跨切片知识迁移 via 遮蔽多模态异质图对比学习用于空间基因表达推断

Zhiceng Shi, Changmiao Wang, Jun Wan, Wenwen Min

机构 * Yunnan University(云南大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Zhongnan University of Economics and Law(中南财经政法大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出SpaHGC模型,通过多模态异质图对比学习,结合局部空间上下文和跨切片相似性,提升空间基因表达预测精度,优于现有九种方法。

Comments Accepted by CVPR-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04846 2026-03-24 cs.CV 79%

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

多范式协同对抗攻击多模态大语言模型

Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对多模态大语言模型的多范式协同对抗攻击方法,通过聚合视觉和语言特征进行联合优化,提升对抗示例的可转移性,实验表明优于现有方法。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 79%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20970 2026-03-24 cs.CV 79%

GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies

GraPHFormer:一种多模态图持久同调变换器,用于神经科学形态学分析

Uzair Shah, Marco Agus, Mahmoud Gamal, Mahmood Alzubaidi, Corrado Cali, Pierre J. Magistretti, Abdesselam Bouzerdoum, Mowafa Househ

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学) University of Turin(都灵大学) BESE, King Abdullah University of Science and Technology(贝赛,国王阿卜杜勒阿齐兹大学科学与技术学院) University of Wollongong(沃林根大学) Neuroscience Institute Cavalieri Ottolenghi(卡瓦利埃-奥托伦奇神经科学研究所) Université Grenoble-Alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 GraPHFormer通过CLIP式对比学习统一拓扑和图结构分析,利用持久图像编码和树状LSTM编码器,实现对神经形态的高精度识别与分类,优于传统方法。

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20181 2026-03-23 cs.CR cs.AI 79%

Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning

通过多模态对比学习提升网络安全任务的泛化能力

Jianan Huang, Rodolfo V. Valentim, Luca Vassio, Matteo Boffa, Marco Mellia, Idilio Drago, Dario Rossi

机构 * Huawei Paris Research Center, France(华为巴黎研究中心,法国)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出多模态对比学习框架,通过文本指导 payloads 分类,提升网络安全任务的泛化能力,并在合成基准和真实数据集上验证效果。

Comments Submitted to Euro S&P - 5th International Workshop on Designing and Measuring Security in Systems with AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20016 2026-03-23 cs.CV 79%

CFCML: A Coarse-to-Fine Crossmodal Learning Framework For Disease Diagnosis Using Multimodal Images and Tabular Data

CFCML:一种用于多模态图像和表格数据疾病诊断的粗到细跨模态学习框架

Tianling Liu, Hongying Liu, Fanhua Shang, Lequan Yu, Tong Han, Liang Wan

机构 * College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) Medical School of Tianjin University(天津大学医学院) Peng Cheng Lab(鹏城实验室) Department of Statistics and Actuarial Science, School of Computing and Data Science, The University of Hong Kong(统计与精算系,计算与数据科学学院,香港大学) Department of Radiology, Tianjin Huanhu Hospital(天津华医院放射科) Tianjin Key Laboratory of Cerebral Vascular and Neurodegenerative Diseases(天津脑血管与神经退行性疾病重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CFCML框架,通过粗到细的跨模态学习逐步缩小多模态数据间的模态差距,提升诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15623 2026-03-18 cs.IR cs.AI 79%

Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval

Finder:一种多模态AI驱动的制药数据检索框架

Suyash Mishra, Srikanth Patil, Satyanarayan Pati, Sagar Sahu, Baddu Narendra

机构 * Researcher, Global Product Strategy F. Hoffmann-La Roche Ltd. Basel, Switzerland(全球产品战略研究员 罗氏有限公司 巴塞尔,瑞士) Associate Vice President (Gen AI) Involead Services Pvt Ltd. Pune, India(高级副总裁(生成式人工智能) Involead服务私人有限公司 普纳,印度) Lead Data Scientist Involead Services Pvt Ltd. Delhi, India(首席数据科学家 Involead服务私人有限公司 德里,印度) Data Scientist Involead Services Pvt Ltd. Bhubaneswar, India(数据科学家 Involead服务私人有限公司 奇塔拉,印度)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 Finder利用混合向量搜索统一文本、图像、音频和视频的检索,通过稀疏词汇和密集语义模型提升多模态内容处理能力,支持自然语言推理搜索,已处理超过29万文档、3.1万视频和1192音频文件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15341 2026-03-17 cs.AI cs.HC cs.MA 79%

Intelligent Co-Design: An Interactive LLM Framework for Interior Spatial Design via Multi-Modal Agents

智能协同设计:一种基于多模态代理的交互式LLM框架用于室内空间设计

Ren Jian Lim, Rushi Dai

机构 * Hong Kong Center for Construction Robotics(香港建设机器人中心)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于LLM的多模态多代理框架,通过自然语言描述和图像动态生成3D设计,提升用户参与度和设计效率。

Comments 25 pages, 20 figures; accepted for publication in the Proceedings of ACADIA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19554 2026-03-17 cs.LG cs.AI 79%

CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning

CARE:对比锚定反射用于可验证多模态推理

Yongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han, Haokun Lin, Yingying Zhu, Xiaojun Chang, Xiaodan Liang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CARE通过对比锚定反射框架,将错误转化为监督信号,提升多模态推理的准确性和训练平滑度,在六个可验证视觉推理基准上提升4.6个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04268 2026-03-12 cs.CV 79%

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

KVSmooth: 通过键值平滑缓解多模态大语言模型中的幻觉

Siyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun He

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 KVSmooth通过键值平滑技术有效缓解多模态大语言模型中的幻觉问题,提升生成精度和召回率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07997 2026-03-10 cs.AI 79%

CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval

CMMR-VLN:通过持续多模态记忆检索实现视觉与语言导航

Haozhou Li, Xiangyu Dong, Huiyan Jiang, Yaoming Zhou, Xiaoguang Ma

机构 * Foshan Graduate School of Innovation at Northeastern University(东北大学创新研究生院) Faculty of Robot Science and Engineering at Northeastern University(东北大学机器人科学与工程学院) College of Software at Northeastern University(东北大学软件学院) School of Aeronautic Science and Engineering at Beihang University(北航航空科学与工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CMMR-VLN通过引入持续多模态记忆检索机制,提升视觉与语言导航任务中对先前经验的选择性利用能力,显著提高导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 79%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06982 2026-03-10 cs.CV cs.IR 79%

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

优化多模态模型用于基于图像的形状检索:预对齐和硬对比学习的作用

Paul Julius Kühn, Cedric Spengler, Michael Weinmann, Arjan Kuijper, Saptarshi Neil Sinha

机构 * Fraunhofer IGD(弗劳恩霍夫图像研究中心) Delft University of Technology(代尔夫特理工大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出通过预对齐和硬对比学习优化多模态模型,提升基于图像的形状检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06503 2026-03-09 cs.CL 79%

Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing

超越行到推理:面向多模态电子表格理解与编辑的智能检索

Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 BRTR通过迭代工具调用框架实现多模态电子表格的端到端理解与编辑,取得领先性能。

详情

展开后加载摘要…

URL PDF HTML 收藏