arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14015cs.CVcs.AI

MedClaw:用于长程手术视频推理的启发式智能体框架

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

  • School of Automation and Intelligence, Beijing Jiaotong University(北京交通大学自动化与智能学院)
  • University of Maryland(马里兰大学)
  • Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州高等研究院)
  • University of British Columbia(不列颠哥伦比亚大学)
  • Beijing Tiantan Hospital, Capital Medical University(首都医科大学附属北京天坛医院)

机构由 AI 辅助整理,请以论文原文为准。

Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang

AI总结:

本研究提出MedClaw智能体框架,通过分离推理与感知、启发式技能蒸馏循环,在仅需约100个标注示例的情况下,于自建及公开手术视频基准MedClawBench上,显著优于现有一次性VLM和通用视频智能体。

AI中文摘要:

理解时长数十分钟的手术视频需要长程时间推理,即通过将问题与分布在时间维度上的视觉证据关联,来回答手术过程中“之前发生了什么”“之后发生了什么”或“不同阶段之间发生了什么”这类问题。现有方法在这方面表现不佳:一次性视觉语言模型(VLM)会压缩整个手术过程以适配其上下文窗口,丢失“之前”或“之后”问题依赖的细节;而专注于训练模型关注位置的视频智能体则数据需求量大,且难以迁移到域外手术场景。我们构建了一个将推理与感知分离、通过演化上下文而非优化权重来提升性能的智能体框架。纯文本的协调器规划需收集的证据,并生成可审计的工具调用序列,而冻结的视觉语言子智能体则针对像素执行每次调用,包括查看、裁剪、检查帧以及检索外部知识。我们进一步提出了无梯度、奖励门控的启发式技能蒸馏循环,该循环挖掘智能体自身低分轨迹,仅当候选技能提升验证奖励时才保留,从而生成可复用的检索技能,尤其是定向重看技能。该循环通过扩展外部技能库而非调整权重,仅需约100个标注示例即可完成适配,远少于监督微调或强化微调所需的样本量。为评估该智能体,我们构建了MedClawBench,这是一个去泄露、基于医生意见的基准,包含针对自建长时神经外科手术录音和保留的公开讲座视频测试集的1123个问题。在两个数据集及全部四个评估维度上,我们的智能体均持续优于一次性VLM和通用视频智能体框架,在长时、域外神经外科手术视频上的提升最为显著。项目页面:this https URL

英文摘要:

Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time. Existing approaches handle this poorly: a one-shot vision-language model (VLM) compresses the whole procedure to fit its context window and loses the detail a "before" or "after" question depends on, while video agents that train the model where to look are data-hungry and transfer poorly to out-of-domain surgery. We build an agent harness that separates reasoning from perception and improves by evolving context rather than optimizing weights. A text-only orchestrator plans which evidence to gather and issues an auditable sequence of tool calls, while frozen vision-language sub-agents execute each call over the pixels, viewing, cropping, inspecting frames, and retrieving external knowledge. We further propose a gradient-free, reward-gated Heuristic Skill Distillation loop that mines the agent's own low-scoring traces and keeps a candidate skill only when it raises a validation reward, yielding reusable retrieval skills, notably directed re-look. Growing an external skill library rather than tuning weights, the loop adapts from only about 100 labeled examples, far fewer than supervised or reinforcement fine-tuning requires. To evaluate this agent, we introduce MedClawBench, a de-leaked, doctor-grounded benchmark of 1,123 questions over self-built long neurosurgery recordings and a held-out public lecture-video test split. Across both datasets and all four evaluation dimensions, our agent consistently outperforms one-shot VLMs and general video-agent frameworks, with the largest gains on the long, out-of-domain neurosurgery videos. Project page: https://fyycs.github.io/medclaw/.

↑