arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31524cs.CV

结构化推理智能体框架用于可解释的安全关键视图评估

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

  • University of Nottingham(诺丁汉大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Qing Xu, Yuxiang Luo, Zhen Chen

AI总结:

提出ReasonCVS,一种基于视觉语言模型的结构化推理智能体框架,通过解剖场景图抽象和理性感知推理分解CVS评估,在Endoscapes-CVS201上以68.1% mAP实现可解释的细粒度手术安全评估。

AI中文摘要:

手术场景理解对于计算机辅助干预至关重要,然而腹腔镜胆囊切除术仍面临肝胆囊三角区复杂解剖结构和胆管损伤风险的挑战。现有的安全关键视图(CVS)评估方法通常将其视为整体预测任务,直接将视觉特征映射到标准级标签。这种黑箱范式缺乏对解剖关系的显式推理,限制了可解释性和组合泛化能力。为解决此问题,我们提出ReasonCVS,一种由视觉语言模型(VLMs)驱动的结构化推理智能体框架,将CVS评估分解为显式、细粒度的解剖验证。具体而言,我们设计了解剖场景图抽象(ASGA),将解剖实体及其空间关系组织成结构化表示。为实现此操作,我们引入了基于理性感知的推理智能体,由通过理性蒸馏微调的大语言模型(LLM)驱动。作为严格的核心决策者,它调用VLM驱动的子标准验证器作为专门感知工具来解析图并独立评估各个子标准。通过校准的软推理,该智能体综合工具收集的分布式观察,产生最终判定及可追溯的临床理性依据。在Endoscapes-CVS201基准上的大量实验表明,ReasonCVS实现了优越性能(68.1% mAP),优于现有最先进方法,同时为可靠的手术评估提供可解释的、标准级解释。

英文摘要:

Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety (CVS) assessment typically treat it as a holistic prediction task, mapping visual features directly to criterion-level labels. This black-box paradigm lacks explicit reasoning about anatomical relationships, limiting both interpretability and compositional generalization. To address this, we propose ReasonCVS, a structured reasoning agentic framework empowered by Vision-Language Models (VLMs) that decomposes CVS assessment into explicit, fine-grained anatomical verification. Specifically, we devise an Anatomical Scene Graph Abstraction (ASGA) that organizes anatomical entities and their spatial relationships into a structured representation. To operationalize this, we introduce a Rationale-Aware Reasoning Agent, powered by a Large Language Model (LLM) fine-tuned via rationale distillation. Functioning as a strict central decision-maker, it invokes VLM-driven Sub-criterion Verifier as a specialized perceptual tool to parse the graph and independently evaluate individual sub-criteria. Through calibrated soft reasoning, this agent synthesizes the tool-gathered distributed observations, yielding a final verdict alongside a traceable clinical rationale. Extensive experiments on the Endoscapes-CVS201 benchmark demonstrate that ReasonCVS achieves superior performance (68.1\% mAP) over state-of-the-art while providing interpretable, criterion-level explanations for reliable surgical assessment.

↑