STSG-VQA:基于手术时空场景图的可证据化时间问答
STSG-VQA: Evidence-Grounded Temporal Question Answering from Surgical Spatio-Temporal Scene Graphs
浏览论文内容
中文总结 AI 辅助
提出多层级结构化时间监督方法构建手术时空场景图,生成带证据的STSG-VQA基准(18,458个问答对),微调手术VLM在时间推理上较基线显著提升。
中文摘要 AI 辅助
尽管手术视觉语言模型(VLMs)近期取得了进展,但由于现有监督主要基于帧,时间推理仍然受限。帧级场景图(SGs)已被证明能有效提供手术环境的结构化表示,但并未显式建模手术工作流的动态变化。为显式建模手术状态随时间的演变,我们引入了一种多层级结构化时间监督方法,该方法通过对象级连续性、事件级交互连续性和流程级连通性来增强帧级手术场景图。随后,我们在生成的时空场景图(STSGs)上执行时间查询,以生成基于证据的问答对,这些问答对共同构成STSG-VQA基准。每个问题都关联到用于推导其参考答案的时间区间和STSG证据,从而实现可追溯的验证。该基准包含18,458个问答对,涵盖七种时间类别。使用STSG派生的监督对Qwen3-VL-4B和Hulu-Med-4B进行微调,在问题级微观准确率上分别比其零样本基线提高了24.39和19.56个百分点,比静态场景图监督分别提高了16.50和14.25个百分点。这些提升跨越所有时间类别,表明STSG派生的监督有助于手术VLM对时间上有依据的交互进行推理,而非孤立帧。代码和数据集将在录用后公开。
英文摘要
Despite recent advances in surgical vision-language models (VLMs), temporal reasoning remains limited because existing supervision is largely frame-centric. Frame-level scene graphs (SGs) have proven effective in providing structured representations of surgical environments but do not explicitly model the dynamics of surgical workflows. To explicitly model how surgical states evolve across time, we introduce a multi-level structured temporal supervision methodology that augments frame-level surgical SGs with object-level continuity, event-level interaction continuity, and procedure-level connectivity. We then execute temporal queries over the resulting spatio-temporal scene graphs (STSGs) to generate evidence-grounded question-answer pairs, which together form the STSG-VQA benchmark. Each question is linked to the temporal interval and STSG evidence used to derive its reference answer, enabling traceable verification. The benchmark contains 18,458 question-answer pairs across seven temporal categories. Fine-tuning Qwen3-VL-4B and Hulu-Med-4B with STSG-derived supervision improves question-level micro accuracy by 24.39 and 19.56 percentage points over their zero-shot baselines and by 16.50 and 14.25 points over static scene-graph supervision, respectively. These gains span all temporal categories, indicating that STSG-derived supervision helps surgical VLMs reason over temporally grounded interactions rather than isolated frames. The code and dataset will be made publicly available upon acceptance.
发表机构
- School of Computer Science, University of Leeds(利兹大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。