EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
EgoDyn-Bench:评估面向自动驾驶的视觉中心基础模型中的自我运动理解
Finn Rasmus Schäfer, Yuan Gao, Dingrui Wang, Thomas Stauner, Stephan Günnemann, Mattia Piccinini, Sebastian Schmidt, Johannes Betz
机构
*
Professorship of Autonomous Vehicle Systems, Technical University of Munich, Munich, Germany(自动驾驶车辆系统教授职位,慕尼黑技术大学,德国慕尼黑)
;
Bayerische Motoren Werke AG, Munich, Germany(巴伐利亚发动机有限公司,德国慕尼黑)
;
Data Analytics and Machine Learning Group, Technical University of Munich, Munich, Germany(数据分析与机器学习小组,慕尼黑技术大学,德国慕尼黑)
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench:当代大型多模态模型的一个不可能的视觉基准测试
Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Yan, Wenye Lin, Gyungin Shin, Qiaochu Yang, Anh Totti Nguyen, David I. Atkinson, Aaditya Baranwal, Alexandru Coca, Mikah Dang, Sebastian Dziadzio, Jakob D. Kunz, Kaiqu Liang, Alexander Lo, Brian Pulfer, Steven Walton, Charig Yang, Kai Han, Samuel Albanie
机构
*
University of Cambridge(剑桥大学)
;
University of Alberta(阿尔伯塔大学)
;
The University of Hong Kong(香港大学)
;
University of Oxford(牛津大学)
;
Northeastern University(东北大学)
;
Astadeus
;
College of Southern Maryland(马里兰州南部学院)
;
University of Geneva(日内瓦大学)
;
University of Oregon(俄勒冈大学)
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
OPPO AI Center, OPPO Inc.(OPPO公司OPPO人工智能中心)
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
视觉语言动作模型所言即所指?论忠实性在具身推理中的作用
Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone
机构
*
Stanford University(斯坦福大学)
;
Politecnico di Milano(米兰理工大学)
;
KTH Royal Institute of Technology(皇家理工学院)
;
Italian Institute of Artificial Intelligence (AI4I)(意大利人工智能研究所)
;
NVIDIA Research(英伟达研究院)
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
ACE:通过零样本工作流推理实现具身操纵的智能体控制
Iok Tong Lei, QianZhi Li, Ying Jie Yap, Yujie Zhang, Rui Zhong, Haichao Gui, Xiaolong Liu, Zhidong Deng
机构
*
Department of Computer Science, Tsinghua University(清华大学计算机科学系)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
;
Vast Intelligence Lab(旷视智能实验室)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU:评估多模态语言模型中的视觉知识理解
Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
EduArt:评估大型语言模型艺术史知识的教育级基准
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
University of Bologna(博洛尼亚大学)
;
Villa i Tatti – The Harvard University Center for Italian Renaissance Studies(哈佛大学意大利文艺复兴研究中心(I Tatti))
;
University of Copenhagen(哥本哈根大学)
OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping
OmniView-Space: 通过多视角空间映射增强空间推理
Xudong Li, Mengdan Zhang, Peixian Chen, Jiaxi Tan, Zihao Huang, Jingyuan Zheng, Yan Zhang, Xiawu Zheng, Xing Sun, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Beijing Institute of Technology(北京理工大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning
RoboSpatial Challenge at CVPR 2026技术报告:选择性推理激活与参考框架消歧用于具身空间推理
RASR: Retrieval-Augmented Semantic Reasoning for Fake News Video Detection
RASR:基于检索的语义推理用于虚假新闻视频检测
Hui Li, Peien Ding, Jun Li, Guoqi Ma, Zhanyu Liu, Ge Xu, Junfeng Yao, Jinsong Su
机构
*
School of Informatics, Xiamen University(厦门大学信息学院)
;
School of Computer Science and Information Security, Guilin University of Electronic Technology(桂林电子科技大学计算机科学与信息安全学院)
;
School of Computer and Big Data, Minjiang University(闽江学院计算机与大数据学院)
;
Xiamen University(厦门大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
机构
*
Zhejiang University(浙江大学)
;
University of Science and Technology of China(中国科学技术大学)
;
East China Normal University(华东师范大学)
;
Zhejiang Provincial People’s Hospital(浙江省人民医院)
;
National University of Singapore(新加坡国立大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV