arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2512.19972 2025-12-24 cs.DC 50%

Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions

重新思考协同机器学习中的知识蒸馏:记忆、知识及其交互

Pengchao Han, Xi Huang, Yi Fang, Guojun Han

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 本文重新审视协同机器学习中的知识蒸馏,探讨记忆与知识的交互作用,分析不同学习模式下的挑战与未来方向。

Comments Published in IEEE TNSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01146 2025-12-19 q-bio.OT 50%

Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications

生物医学中的检索增强生成:技术、数据集和临床应用的综述

Jiawei He, Boya Zhang, Hossein Rouhizadeh, Yingjian Chen, Rui Yang, Jin Lu, Xudong Chen, Nan Liu, Douglas Teodoro

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文综述了生物医学中检索增强生成技术的发展,探讨了其在临床应用中的挑战与未来发展方向。

Comments 49 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09548 2025-12-11 cs.MA 50%

Supporting Dynamic Agentic Workloads: How Data and Agents Interact

支持动态代理工作负载:数据与代理如何相互作用

Ioana Giurgiu, Michael E. Nidd

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 本文提出了一种以代理为中心的数据织物,旨在高效支持动态、多模态的代理工作负载,通过注意力引导的数据检索、语义微缓存等机制提升数据处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01653 2025-12-02 eess.SP cs.LG 50%

Cuffless Blood Pressure Estimation from Six Wearable Sensor Modalities in Multi-Motion-State Scenarios

无袖带血压估计:基于六种可穿戴传感器模态在多运动状态场景中的应用

Yiqiao Chen, Fazheng Xu, Zijian Huang, Juchi He, Zhenghui Feng

机构 * Faculty of Frontier Sciences, Harbin Institute of Technology, Shenzhen(前沿科学学院,哈尔滨工业大学深圳校区) University of Melbourne(墨尔本大学) University of New South Wales(新南威尔士大学)

专题命中 跨模态检索 :cross-modal(abstract)

AI总结 本文提出了一种基于六种可穿戴传感器模态的无袖带血压估计框架,通过联合利用ECG、PPG、压力、温度和运动传感器数据,实现多运动状态下的血压准确估计。

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06478 2025-12-02 eess.SP 50%

Retrieving Filter Spectra in CNN for Explainable Sleep Stage Classification

在CNN中检索滤波器光谱以实现可解释的睡眠阶段分类

Stephan Goerttler, Yucheng Wang, Fei He, Min Wu

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本研究提出了一种在CNN中检索滤波器光谱的工具,用于提高睡眠阶段分类的可解释性,通过分析EEG通道的光谱信息来增强模型性能。

Comments 6 pages, 3 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00379 2025-12-02 q-bio.BM cs.LG 50%

EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants

EnzyCLIP:一种基于对比学习的跨注意力双编码框架,用于预测酶动力学常数

Anas Aziz Khan, Md Shah Fahad, Priyanka, Ramesh Chandra, Guransh Singh

机构 * SCOPE Vellore Institute of Technology(维洛雷理工学院) BIT Department of Computer Science(计算机科学系) Department of Bioengineering and Biotechnology(生物工程与生物技术系)

专题命中 跨模态检索 :multimodal(abstract)

AI总结 EnzyCLIP通过对比学习和跨注意力机制,结合蛋白质序列和底物分子结构预测酶动力学参数,提升Kcat和Km预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20227 2025-11-26 cs.IR 50%

HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents

HKRAG:面向视觉丰富文档的综合知识检索增强生成

Anyang Tong, Xiang Niu, ZhiPing Liu, Chang Tian, Yanyan Wei, Zenglin Shi, Meng Wang

专题命中 跨模态检索 :multimodal(abstract)

AI总结 HKRAG通过综合检索和生成机制,提升视觉丰富文档中显著与细节知识的检索与生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18598 2025-11-25 q-bio.OT 50%

Assessing Gaze and Pointing: Human Cue Interpretation by Indian Free-Ranging Dogs in a Food Retrieval Task

评估目光与指认:印度自由放养狗在食物获取任务中的人类提示解读

Srijaya Nandi, Dipanjan Roy, Aesha Lahiri, Anamitra Roy, Anindita Bhadra

专题命中 跨模态检索 :multimodal(abstract)

AI总结 研究发现印度自由放养狗在结合指认和目光提示时能准确找到隐藏食物,但单一或冲突提示下表现无显著差异,且狗的气质影响其参与意愿和接近延迟,但不影响选择准确性。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10584 2025-11-25 cs.IR 50%

DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System

DAS: 基于双对齐语义ID的工业推荐系统

Wencai Ye, Mingjie Sun, Shaoyun Shi, Peng Wang, Wenjin Wu, Peng Jiang

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 DAS通过双对齐语义ID方法,提升推荐系统中多模态内容整合与协同信号对齐效率,有效解决信息损失与灵活性问题。

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16950 2025-11-24 physics.optics 50%

Single-Axis Ptychographic Coherent Diffractive Imaging for Spectroscopic and Wavefront Retrieval

单轴投影衍射成像用于光谱和波前恢复

Qijun You, Lingshuo Meng, Fangrui Quan, Wei Cao

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 单轴投影衍射成像技术通过单轴扫描提高通量,实现光谱和波前的同步成像,用于高效多模式实时成像。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11912 2025-11-18 cs.LG cs.CR 50%

A Systematic Study of Model Extraction Attacks on Graph Foundation Models

Haoyan Xu, Ruizhi Qian, Jiate Li, Yushun Dong, Minghao Lin, Hanson Yan, Zhengtao Yao, Qinghua Liu, Junhao Dong, Ruopeng Huang, Yue Zhao, Mengyuan Li

机构 * University of Southern California(南加州大学) Florida State University(佛罗里达州立大学) The Ohio State University(俄亥俄州立大学) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14772 2025-11-12 cs.HC 50%

UMind: A Unified Multitask Network for Zero-Shot M/EEG Visual Decoding

Chengjian Xu, Yonghao Song, Zelin Liao, Haochuan Zhang, Qiong Wang, Qingqing Zheng

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25718 2025-10-30 cs.IR cs.DL 50%

Retrieval-Augmented Search for Large-Scale Map Collections with ColPali

Jamie Mahowald, Benjamin Charles Germain Lee

专题命中 跨模态检索 :multimodal(abstract)

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22613 2025-10-28 cs.SE 50%

DynaCausal: Dynamic Causality-Aware Root Cause Analysis for Distributed Microservices

Songhan Zhang, Aoyang Fang, Yifan Yang, Ruiyi Cheng, Xiaoying Tang, Pinjia He

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07759 2025-10-28 cs.IR 50%

A Survey of Long-Document Retrieval in the PLM and LLM Era

Minghan Li, Miyang Luo, Tianrui Lv, Yishuai Zhang, Siqi Zhao, Ercong Nie, Guodong Zhou

专题命中 跨模态检索 :multimodal(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16803 2025-10-21 cs.IR 50%

An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision

Zishuai Zhang, Sihao Yu, Wenyi Xie, Ying Nie, Junfeng Wang, Zhiming Zheng, Dawei Yin, Hainan Zhang

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13406 2025-10-16 cs.LG 50%

When Embedding Models Meet: Procrustes Bounds and Applications

Lucas Maystre, Alvaro Ortega Gonzalez, Charles Park, Rares Dolga, Tudor Berariu, Yu Zhao, Kamil Ciosek

机构 * UiPath Spotify

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12213 2025-10-15 astro-ph.IM 50%

Encapsulating Textual Contents into a MOC data Structure for Advanced Applications

Giuseppe Greco, Thomas Boch, Pierre Fernique, Manon Marchand, Mark Allen, Francois Xavier Pineau, Matthieu Baumann, Marco Molinaro, Roberto De Pietri, Marica Branchesi, Steven Schramm, Gergely Dalya, Elahe Khalouei, Barbara Patricelli, Giulia Stratta

专题命中 跨模态检索 :multimodal(abstract)

Comments Published in Astronomy and Computing; 11 pages, 4 figures

Journal ref Astron. Comput. 54 (2026) 101014

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01853 2025-10-07 cs.LG cs.LO 50%

Learning Representations Through Contrastive Neural Model Checking

Vladimir Krsmanovic, Matthias Cosler, Mohamed Ghanem, Bernd Finkbeiner

机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全研究中心)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23543 2025-09-30 q-bio.GN cs.NE q-bio.MN 50%

Contrastive Learning Enhances Language Model Based Cell Embeddings for Low-Sample Single Cell Transcriptomics

Luxuan Zhang, Douglas Jiang, Qinglong Wang, Haoqi Sun, Feng Tian

专题命中 跨模态检索 :multimodal(abstract)

Comments 14 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21507 2025-09-29 cs.CE 50%

QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance

Haoxue Wang, Keli Wen, Yuante Li, Qiancheng Qu, Xiangxu Mu, Xinjie Shen, Jiaqi Gao, Chenyang Chang, Chuhan Xie, San Yu Cheung, Zhuoyuan Hu, Xinyu Wang, Sirui Bi, Bi'an Du

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14031 2025-09-22 cs.NE cs.LG 50%

Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms

Shreya Saha, Ishaan Chadha, Meenakshi Khosla

机构 * Electrical and Computer Engineering University of California, San Diego(电气与计算机工程大学加州大学圣地亚哥分校) Halıcıoğlu Data Science Institute University of California, San Diego(Halıcıoğlu数据科学研究所大学加州大学圣地亚哥分校) Department of Cognitive Science, Department of Computer Science and Engineering University of California, San Diego(认知科学系计算机科学与工程系大学加州大学圣地亚哥分校)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13326 2025-09-18 cs.HC cs.LG 50%

LLM Chatbot-Creation Approaches

Hemil Mehta, Tanvi Raut, Kohav Yadav, Edward F. Gehringer

专题命中 跨模态检索 :multimodal(abstract)

Comments Forthcoming in Frontiers in Education (FIE 2025), Nashville, Tennessee, USA, Nov 2-5, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12824 2025-09-18 cs.IR 50%

DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval

Zechao Liu, Zheng Zhou, Xiangkun Chen, Tao Liang, Dapeng Lang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04836 2025-09-08 cs.RO 50%

COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks

Dongping Li, Shaoting Peng, John Pohovey, Katherine Rose Driggs-Campbell

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Department of Electrical and Computer Engineering(电气与计算机工程系) ZJU-UIUC Institute(浙大-伊利诺伊大学联合学院)

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09647 2025-09-08 cs.RO 50%

Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic Map

Taichi Sakaguchi, Akira Taniguchi, Yoshinobu Hagiwara, Lotfi El Hafi, Shoichi Hasegawa, Tadahiro Taniguchi

机构 * Ritsumeikan University(立命馆大学) Soka University(早稻田大学) Kyoto University(京都大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments See website at https://emergentsystemlabstudent.github.io/MultiViewRetrieve/. Accepted to IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01184 2025-09-03 cs.IR 50%

MARS: Modality-Aligned Retrieval for Sequence Augmented CTR Prediction

Yutian Xiao, Shukuan Wang, Binhao Wang, Zhao Zhang, Yanze Zhang, Shanqi Liu, Chao Feng, Xiang Li, Fuzhen Zhuang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09128 2025-09-03 stat.ML cs.LG 50%

A Generalization Theory for Zero-Shot Prediction

Ronak Mehta, Zaid Harchaoui

机构 * University of Washington(华盛顿大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments Published at ICML '25 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11535 2025-08-29 stat.ML cs.LG cs.SY eess.SY stat.CO 50%

Canonical Bayesian Linear System Identification

Andrey Bryutkin, Matthew E. Levine, Iñigo Urteaga, Youssef Marzouk

机构 * Massachusetts Institute of Technology(麻省理工学院) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Basis Research Institute(Basis研究机构) Eric and Wendy Schmidt Center(埃里克和温迪·施密特中心) BCAM (Basque Center for Applied Mathematics)(BCAM(巴斯克应用数学中心)) Ikerbasque (Basque Foundation for Science)(Ikerbasque(巴斯克科学基金会))

专题命中 跨模态检索 :multi-modal(abstract)

Comments 46 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19942 2025-08-28 cs.HC 50%

Socially Interactive Agents for Preserving and Transferring Tacit Knowledge in Organizations

Martin Benderoth, Patrick Gebhard, Christian Keller, C. Benjamin Nakhosteen, Stefan Schaffer, Tanja Schneeberger

专题命中 跨模态检索 :multimodal(abstract)

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏