From Perception to Action: An Interactive Benchmark for Vision Reasoning
从感知到行动:一个交互式基准测试用于视觉推理
Yuhao Wu, Maojia Song, Yihuai Lan, Lei Wang, Zhiqiang Hu, Yao Xiao, Heng Zhou, Weihua Zheng, Dylan Raharja, Soujanya Poria, Roy Ka-Wei Lee
机构
*
Singapore University of Technology(新加坡科技设计大学)
;
Singapore Management University (SMU), Singapore(新加坡管理学院)
;
Nanyang Technological University (NTU), Singapore(南洋理工大学)
;
University of Science(科学大学)
机构
*
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
LimX Dynamic(LimX动态)
;
Nanjing University(南京大学)
;
Nanyang Technological University(南洋理工大学)
;
Zhejiang University(浙江大学)
专题命中
GUI与屏幕智能体
:vision language model(abstract);分类 cs.CV
Physics-based phenomenological characterization of cross-modal bias in multimodal models
基于物理现象的多模态模型跨模态偏差表征
Hyeongmo Kim, Sohyun Kang, Yerin Choi, Seungyeon Ji, Junhyuk Woo, Hyunsuk Chung, Soyeon Caren Han, Kyungreem Han
机构
*
B rain Science Institute(脑科学研究院)
;
Korea Institute of Science and Technology(韩国科学技术院)
;
Department of Physics and Astronomy(物理与天文学系)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
University of Science and Technology KIST School(科学技术KIST学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
AI总结
本文提出基于物理现象的多模态模型跨模态偏差表征方法,揭示多模态输入可能强化模态主导性。
CommentsBest Paper Award at BiasinAI track in AAAI2026
机构
*
School of Big Data & Software Engineering, Chongqing University(大数据与软件工程学院,重庆大学)
;
School of Computer Science & Technology, Chongqing University(计算机科学与技术学院,重庆大学)
机构
*
Software School, Shandong University(山东大学软件学院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)
机构
*
RIKEN Center for Advanced Intelligence Project, Japan(日本先进人工智能研究中心)
;
The University of Tokyo, Japan(东京大学)
;
Southeast University, China(东南大学)
;
China University of Mining(中国矿业大学)
专题命中
其他VLM
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV