CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
CoTZero:通过分层合成CoT实现无标注的人类级视觉推理
Chengyi Du, Yazhe Niu, Dazhong Shen, Luxin Xu
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong MMLab(香港中文大学 MMLab)
;
The College of Computer Science and Technology(计算机科学与技术学院)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
机构
*
Peking University(北京大学)
;
Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院)
;
National University of Singapore(新加坡国立大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Science and Technology of China(中国科学技术大学)
;
Cornell University(康奈尔大学)
;
Hong Kong Polytechnic University(香港理工大学)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
CommentsSubmitted to CVPR 2026. Introduces the QVLM architecture and the SQuID dataset for quantitative geospatial reasoning. Dataset DOI: 10.57967/hf/7565
机构
*
University of Manchester(曼彻斯特大学)
;
Queen Mary University of London(伦敦大学玛丽女王学院)
;
Hongkong University of Science and Technology(香港科学与技术大学)
;
Nanjing University(南京大学)
;
Dartmouth College(达特茅斯学院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
Shenzhen Stomatology Hospital (Pingshan) of Southern Medical University(南方医科大学深圳口腔医院(平山))
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
State Key Laboratory of Membrane Biology, Beijing Key Laboratory of Cardiometabolic Molecular Medicine, Institute of Molecular Medicine, National Biomedical Imaging Center, School of Future Technology, Peking University(北京大学膜生物学国家重点实验室、北京心代谢分子医学重点实验室、分子医学研究院、国家生物医学成像中心、未来技术学院)
;
Freedom AI
;
Division of Applied Oral Sciences & Community Dental Care Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院应用口腔科学与社区牙科护理系)
;
Beijing Institute of Collaborative Innovation(北京协同创新研究院)
;
National Health Data Institute, Shenzhen(深圳国家健康数据研究院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Shenzhen Institute of Big Data(深圳大数据研究院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro
机构
*
Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))
;
Visual Intelligence Research Section, Superintelligence Creative Research Laboratory, Electronics and Telecommunications Research Institute (ETRI)(视觉智能研究部,超智能创意研究实验室,电子电信研究院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
National University of Singapore(新加坡国立大学)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
Zhejiang University(浙江大学)
;
DeepWisdom(深智科技)