Comments5 pages, 4 figures, 3 tables, presented at 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016), New York, NY. arXiv admin note: substantial text overlap with arXiv:1606.03556
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
用于CT扫描中可靠且可审计的空间关系验证的模块化智能体
Simon Vincent Abel, Heiko Hillenhagen, Michael Götz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
机构
*
Ulm University(乌尔姆大学)
;
Ulm University Hospital(乌尔姆大学医院)
;
Technical University of Munich (TUM)(慕尼黑工业大学(TUM))
;
TUM University Hospital(慕尼黑工业大学医院)
;
Department of Radiation Oncology, TUM University Hospital(慕尼黑工业大学医院放射肿瘤科)
Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
拼接与讲述:一种结构化多模态数据增强方法用于空间理解
Hang Yin, Xiaomin He, PeiWen Yuan, Yiwei Li, Jiayi Shi, Wenxiao Fan, Shaoxiong Feng, Kan Li
机构
*
School of Computer Science, Beijing Institute of Technology(北京理工大学计算机科学学院)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Xiaohongshu Inc(小红书公司)
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion:用于整体医学理解的以视觉为中心的多模态大语言模型系统
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
机构
*
DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
;
Hupan Laboratory(湖畔实验室)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University(宁波大学附属阳明医院放射科)
;
Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学)
;
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学)
;
School of Software, Tsinghua University(清华大学软件学院)
;
Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
PathScale-R1:用于病理图像分析的跨尺度推理
Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu
机构
*
National University of Singapore(新加坡国立大学)
;
PuzzleLogic Pte Ltd(拼图逻辑私人有限公司)
;
Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院)
Enhancing Pathological VLMs with Cross-scale Reasoning
增强病理视觉语言模型的跨尺度推理能力
Chi Phan, Tianyi Zhang, Qiaochu Xue, Yufeng Wu, Dan Hu, Zeyu Liu, Sudong Wang, Yueming Jin
机构
*
Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系)
;
PuzzleLogic Pte Ltd(PuzzleLogic私人有限公司)
;
Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院病理科暨福建省肿瘤医院)
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
机构
*
Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系)
;
Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所)
;
Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
College of Software, Nankai University, China(南开大学软件学院)
;
The University of Hong Kong, China(香港大学)
;
NKIARI, Shenzhen Futian, China(NKIARI,深圳福田,中国)
;
AAIS & VCIP, Nankai University, China(AAIS与VCIP,南开大学,中国)
;
A*STAR Institute of Advanced Intelligence and Computing, Singapore(新加坡A*STAR先进智能与计算研究所)
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
关注、变换或静默:面向高效多模态大语言模型推理的算子级视觉跳跃
Zhaoyang Luo, Runmin Dong, Miao Yang, Fan Wei, Yushan Lai, Bin Luo, Haohuan Fu
机构
*
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
Sun Yat-sen University(中山大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
;
Tsinghua University(清华大学)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Korea University(高丽大学)
;
Upstage AI
;
Kyung Hee University(庆熙大学)
;
KAIST(韩国科学技术院)
;
Hanyang University College of Medicine(汉阳大学医学院)
;
AIGEN Sciences
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
基于LLM增强优化的无人机低空经济网络高效机载视觉-语言推理
Yang Li, Ruichen Zhang, Yinqiu Liu, Guangyuan Liu, Abbas Jamalipour, Xianbin Wang, Dong In Kim
机构
*
College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院、新加坡国立科技大学)
;
The University of Sydney, Sydney, Australia(悉尼大学、澳大利亚悉尼)
;
Department of Electrical and Computer Engineering, Western University, London, Canada(电气与计算机工程系、西方大学、加拿大伦敦)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University, South Korea(电气与计算机工程系、全州大学、韩国)
机构
*
Computer Aided Medical Procedures (CAMP)(计算机辅助医疗程序)
;
TU Munich, Germany(慕尼黑工业大学,德国)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Munich, Germany(慕尼黑,德国)
;
Zhongshan Hospital, Fudan University, China(复旦大学中山医院)
;
The University of Hong Kong, Hongkong, China(香港大学,香港,中国)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
迈向具有空间定位病变证据的临床可解释性眼科VQA
Xingyue Wang, Bo Liu, Meng Wang, Zhixuan Zhang, Chengcheng Zhu, Huazhu Fu, Jiang Liu
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系)
;
The Hong Kong Polytechnic University(香港理工大学)
;
National University of Singapore(新加坡国立大学)
;
University of Washington(华盛顿大学)
;
Institute of High Performance Computing, Agency for Science, Technology and Research(科技研究局高性能计算研究所)