Unbiased Visual Reasoning with Controlled Visual Inputs
通过受控视觉输入实现无偏视觉推理
Zhaonan Li, Shijie Lu, Fei Wang, Jacob Dineen, Xiao Ye, Zhikun Xu, Siyi Liu, Young Min Cho, Bangzheng Li, Daniel Chang, Kenny Nguyen, Qizheng Yang, Muhao Chen, Ben Zhou
机构
*
Arizona State University(亚利桑那州立大学)
;
University of Southern California(南加州大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of California, Davis(加州大学戴维斯分校)
机构
*
Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学庆元研究院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Department of Pathology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生命科学与医学学院病理科)
;
Intelligent Pathology Institute, Division of Life Sciences and Medicine(生命科学与医学学院智能病理研究所)
;
Department of Pathology, Fudan University Shanghai Cancer Center(复旦大学上海癌症中心病理科)
;
Department of Oncology, Shanghai Medical College, Fudan University(复旦大学上海医学院肿瘤科)
;
Institute of Pathology, Fudan University(复旦大学病理研究所)
;
Department of Pathology, The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院病理科)
CommentsThis paper is a development of the visual riddle game with Human-AI interaction, entitled "GuessWhat - Riddle Eye with AI", developed by Ciprian Constantinescu (POLItEHNICA Bucharest), Serena Stan (Instituto Cervantes Bucarest) and Marius Leordeanu (POLITEHNICA Bucharest), which was the winner (1st place) of the NeoArt Connect NAC 2025 Scholarship Program
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
MM-UAVBench: 多模态大语言模型在低空无人机场景中的感知、推理与规划能力如何?
Shiqi Dai, Zizhi Ma, Zhicong Luo, Xuesong Yang, Yibin Huang, Wanyue Zhang, Chi Chen, Zonghao Guo, Wang Xu, Yufei Sun, Maosong Sun
机构
*
Tsinghua University(清华大学)
;
Nankai University(南开大学)
;
Northwest Polytechnical University(西北工业大学)
;
Chinese Academy of Sciences(中国科学院)
;
Harbin Institute of Technology(哈尔滨工业大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
机构
*
Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);分类 cs.CV、cs.LG