MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
MedClaw:用于长程手术视频推理的启发式智能体框架
Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
机构
*
School of Automation and Intelligence, Beijing Jiaotong University(北京交通大学自动化与智能学院)
;
University of Maryland(马里兰大学)
;
Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州高等研究院)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Beijing Tiantan Hospital, Capital Medical University(首都医科大学附属北京天坛医院)
机构
*
Centennial High School, Frisco, Texas, USA(Centennial High School, Texas, USA)
;
Lebanon Trail High School, Frisco, Texas, USA(Lebanon Trail High School, Texas, USA)
;
West Windsor-Plainsboro High School, Princeton Junction, New Jersey, USA(West Windsor-Plainsboro High School, New Jersey, USA)
;
Algoverse AI Research, Palo Alto, California, USA(Algoververse AI Research, California, USA)
专题命中
视觉推理
:VLM(title_cn);vision-language model(abstract);vision language model(abstract);visual reasoning(abstract)
Comments14 pages, 4 figures, 8 tables. Presented at the 39th Conference on Neural Information Processing Systems Workshop: VLM4RWD. Presented at the 43th International Conference on Machine Learning Workshops: ICML 2026 CTB, ICML 2026 FAGEN, ICML 2026 EMM-QA. Authors Aahana Basappa and Pranay Goel contributed equally. Code: https://github.com/AahanaB24/AMVICC, Data: https://doi.org/10.5281/zenodo.17646068
Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation
先看后思:解耦感知与推理以实现抗捷径的多模态在策略自蒸馏
Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han
机构
*
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉推理
:MLLM(summary_cn,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述
Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv
机构
*
School of Computer Science, Sichuan University(四川大学计算机学院)
;
School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院)
;
Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI、cs.LG
机构
*
Fudan University(复旦大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信)
;
Tianjin University(天津大学)
;
Northwestern Polytechnical University(西北工业大学)
;
Tsinghua University(清华大学)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);vision-language model(abstract);VLM(abstract_cn);grounding(abstract)
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
检索、整合与综合:空间-语义 grounded 的潜在视觉推理
Jin Cui, Xinyue Long, Xunyong Zhang, Yadong Zhang, Chuanchang Su, Jingye Gan, Boran Zhao, Pengju Ren
机构
*
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)
专题命中
视觉推理
:visual reasoning(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract)
VisPlay: Self-Evolving Vision-Language Models from Images
VisPlay: 从图像中自我进化视觉-语言模型
Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, Yonghui Yang
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of Maryland(马里兰大学)
;
National University of Singapore(新加坡国立大学)
AVIS: Adaptive Test-Time Scaling for Vision-Language Models
AVIS: 视觉语言模型的自适应测试时缩放
Ahmadreza Jeddi, Minh Ngoc Le, Amirhossein Kazerouni, Hakki Can Karaimer, Hue Nguyen, Iqbal Mohomed, Michael Brudno, Alex Levinshtein, Konstantinos G. Derpanis, Babak Taati, Radek Grzeszczuk
机构
*
AI Center-Toronto, Samsung Electronics(三星电子多伦多AI中心)
;
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
York University(约克大学)
机构
*
BARCELONA Supercomputing Center(巴塞罗那超级计算中心)
;
Universidad Pontificia Comillas(庞培法布拉大学)
;
Chung-Ang University(中央大学)
;
LUXEMBOURG Institute of Science and Technology(卢森堡科学技术研究所)
CommentsAccepted as a book chapter in "Advances in Global Applied Artificial Intelligence" (G. A. Tsihrintzis, M. Virvou, N. G. Bourbakis, L. C. Jain, Eds.), authenticated version will be published in Springer series: Learning and Analytics in Intelligent Systems
机构
*
Qwen Large Model Application Team, Alibaba(阿里云大模型应用团队)
;
Alibaba Zhejiang University(阿里巴巴浙江大学)
;
University of Waterloo(多伦多大学)
;
Vector Institute(向量研究所)
;
Nanjing University(南京大学)
;
Binjiang Institute of Zhejiang University(浙江大学滨江学院)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
EndoCoT:扩散模型中的内生思维链推理扩展
Xuanlang Dai, Yujie Zhou, Long Xing, Jiazi Bu, Xilin Wei, Yuhong Liu, Beichen Zhang, Kai Chen, Yuhang Zang
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Xi’an Jiaotong University(西安交通大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Fudan University(复旦大学)
;
The Chinese University of Hong Kong(香港中文大学)
专题命中
视觉推理
:MLLM(summary_cn,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV