arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RefineRank:用于外科时空定位的联合框优化与排序

RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding

Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque

arXiv 2608.23928首次发表:更新:

发表机构

UCL Hawkes Institute, University College London; King’s College London; School of Medicine, Nankai University; University of Manchester(伦敦大学学院霍克斯研究所; 伦敦国王学院; 南开大学医学院; 曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RefineRank通过RefineNet模块结合冻结医学视觉语言模型与开放集检测器,实现外科时空定位的联合框优化与排序,在MedVidBench上取得较高STG mIoU,提升了定位性能且无需重训主干。

AI 中文摘要

外科时空定位(STG)要求在手术视频的每个时间点,根据程序性问题定位问题所指向的物体。现有方法存在权衡:视觉语言模型能理解问题上下文,但生成的坐标不够精确;而开放集检测器能提供定位的候选框,但其置信度无法反映哪个框能回答问题。我们提出RefineRank,在候选框层面缩小这一差距。一个紧凑的可训练模块RefineNet,将冻结的医学视觉语言模型的语言和区域特征,与冻结的开放集检测器的候选框结合:它为每个候选框预测有界坐标修正和质量分数,固定解码规则返回分数最高的原始框或优化后框。在MedVidBench官方排名(验证集)中,RefineRank的STG mIoU达0.421,是显示的最高STG分数,其全局多指标排名为11。在独立训练和评估视频的受控评估中,坐标修正将候选框的上限从0.6772提升至0.7302,而通过RefineNet分数对原始框和优化后框的联合池进行排序,使STG mIoU从0.2719提升至0.4534,而在同一池上单独训练的选择器最多达到0.4186。这些结果表明,小型框级模块可协调问题理解与精确定位,无需重新训练任一主干网络。代码可在[this https URL]获取。

英文摘要

Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the question asks about. Existing approaches face a trade-off: vision language models understand the question context but produce imprecise coordinates, whereas open-set detectors provide localized candidate boxes whose confidence does not reflect which box answers the question. We introduce RefineRank, which closes this gap at the candidate-box level. A compact trainable module, RefineNet, combines the language and regional features of a frozen medical vision language model with the proposals of a frozen open-set detector: it predicts a bounded coordinate correction and a quality score for every candidate box, and a fixed decoding rule returns the original or refined box with the highest score. On the MedVidBench Official Rankings (Verified), RefineRank records 0.421 STG mIoU, the highest displayed STG score, while its global multi-metric rank is 11. In a controlled evaluation on separate training and evaluation videos, coordinate correction raises the candidate oracle upper bound from 0.6772 to 0.7302, and ranking the joint pool of original and refined candidates by their RefineNet scores improves STG mIoU from 0.2719 to 0.4534, whereas separately trained selectors over the same pool reach at most 0.4186. These results show that a small box-level module can reconcile question understanding with precise localization without retraining either backbone. Code is available at [https://github.com/linzhe001/RefineRank](https://github.com/linzhe001/RefineRank).

Comments17 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑