arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检索增强型视觉提示:引导基础模型用于双光子成像

Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto, Daniela Giordano, Concetto Spampinato, Federica Proietto Salanitri

arXiv 2608.21970首次发表:更新:

发表机构

University of Catania(卡塔尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出RAVP框架,通过在推理阶段注入外部视觉记忆引导基础模型,经Allen Brain Observatory实验验证,可提升零样本神经元检测与实例分割性能,单示例提示效果优于多示例。

AI 中文摘要

双光子钙成像对基础模型而言是极具挑战性的场景:不同记录和实验条件下的图像外观差异显著,标注稀缺,且常需快速适配。我们未通过微调调整模型权重,而是探究能否在推理阶段通过将外部视觉记忆直接注入模型输入来引导基础模型。我们以SAM 3实现该思路,并提出检索增强型视觉提示(Retrieval-Augmented Visual Prompting, RAVP)框架,该框架中每个目标图像块会被补充检索到的带标注示例,其边界框被用作概念提示。RAVP将检索转化为视觉提示的一种形式,仅通过输入设计即可实现适配。我们研究了多种示例选择策略,包括荧光引导启发式方法,以及用于估计哪个示例对目标图像块最具信息性的轻量召回预测器。在Allen Brain Observatory上的实验表明,带示例增强的推理始终能提升零样本神经元检测和实例分割性能。消融研究进一步显示,精心挑选的单个示例比用多个检索到的示例进行提示更有效。这些结果表明,推理阶段的视觉记忆注入是针对生物医学成像领域基础模型的一种简单且有效的参数替代方案。

英文摘要

Two-photon calcium imaging presents a challenging setting for foundation models: image appearance varies substantially across recordings and experimental conditions, annotations are scarce, and rapid adaptation is often needed. Rather than adapting model weights through fine-tuning, we ask whether a foundation model can be guided at inference time by injecting external visual memory directly into its input. We implement this idea with SAM 3 and introduce Retrieval-Augmented Visual Prompting (RAVP), a framework in which each target tile is augmented with a retrieved annotated exemplar whose bounding box is used as a concept prompt. RAVP turns retrieval into a form of visual prompting and enables adaptation through input design alone. We study multiple exemplar selection strategies, including fluorescence-guided heuristics and a lightweight recall predictor trained to estimate which exemplar is most informative for a target tile. Experiments on the Allen Brain Observatory show that exemplar-augmented inference consistently strengthens zero-shot neuron detection and instance segmentation. Ablation studies further show that a single carefully selected exemplar is more effective than prompting with multiple retrieved examples. These results position inference-time visual memory injection as a simple and effective alternative to parameter adaptation for foundation models in specialized biomedical imaging.

Comments14 pages, 4 figures, 6 tables. Supplementary material included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑