发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学时间序列仅作数值表示而忽视形态信息的问题,提出ViRe方法,利用冻结VLM生成视觉查询引导跨模态检索,在六个基准上相对提升6.42%。
AI 中文摘要
医学时间序列(MedTS)支撑着许多临床分类任务,然而现有方法通常仅将其表示为数值序列,未能充分利用在波形检查中显而易见的形态学信息。为弥合这一差距,我们引入了视觉信息检索(ViRe),该方法利用冻结的VLM派生的波形表示作为形态学感知的查询(Query),以引导从原始数值MedTS特征中进行检索。具体而言,使用预训练的视觉-语言模型(VLMs)提取视觉查询(Vision Query),以从波形图中获得形态学感知的先验信息。随后,一种定制的基于注意力的跨模态检索机制利用视觉查询从数值表示中选择与形态学相关的时间与通道证据。ViRe在十个既有基线上展现出强大的有效性,在六个公开基准上相较于先前最先进方法总体相对提升了6.42%。代码、训练脚本和可复现材料已在GitHub仓库中公开提供:this https URL。
英文摘要
Medical time series (MedTS) underpin many clinical classification tasks, yet existing methods usually represent them only as numerical sequences and underuse the morphology that is explicit in waveform inspection. To bridge this gap, we introduce Vision-Informed Retrieval (ViRe), which uses a frozen VLM-derived waveform representation as a morphology-aware Query to guide retrieval from raw numerical MedTS features. Specifically, a Vision Query is extracted using pre-trained vision-language models (VLMs) to obtain morphology-aware priors from waveform plots. A tailored attention-based cross-modal retrieval mechanism then uses the Vision Query to select morphology-relevant temporal and channel evidence from the numerical representation. ViRe demonstrates strong effectiveness against ten established baselines, yielding an overall 6.42% relative improvement over the previous state of the art across six public benchmarks. Code, training scripts, and reproducibility materials are publicly available in the GitHub Repository: https://github.com/Levi-Ackman/ViRe.
CommentsAccepted by NeurIPS 2026