BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding
BARISTA:一种多任务第一人称视角基准,用于组合视觉理解
Patrick Knab, Orgest Xhelili, Inis Buzi, Drago Andres Guggiana Nilo, Mohd Saquib Khan, Lorenz Kolb, Manuel Scherzer, Kerem Yildirir, Christian Bartelt, Philipp Johannes Schubert
机构
*
Ramblr.ai Research(Ramblr.ai 研究院)
;
Technical University of Clausthal(Clausthal 技术大学)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Kling Team, Kuaishou Technology(快手科技 Kling 团队)
;
Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)
专题命中
视觉推理
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
IGV-RRT: Prior-Real-Time Observation Fusion for Active Object Search in Changing Environments
IGV-RRT:面向动态环境主动目标搜索的先验-实时观测融合
Wei Zhang, Ping Gong, Yujie Wang, Leilei Yao, Minghui Bai, Rongfeng Ye, Yinchuan Wang, Yachao Wang, Chen Sun, Chaoqun Wang
机构
*
The School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)
;
Department of Data and Systems Engineering, HKU(数据与系统工程系,香港大学)
专题命中
视觉推理
:VLM(abstract,abstract_cn);vision language model(abstract)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
;
Zhongguancun Academy(中关村学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
CheXTemporal:用于胸部X光影像中时间感知推理的数据集
Eva Prakash, Yunhe Gao, Chong Wang, Justin Xu, Neal Prakash, Arne Michalson, Seena Dehkharghani, Eun Kyoung Hong, Julie Bauml, Roger Boodoo, Jean-Benoit Delbrouck, Sophie Ostmeier, Curtis Langlotz
机构
*
Stanford University(斯坦福大学)
;
University of Oxford(牛津大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
HOPPR
;
University Hospital Zurich(苏黎世大学医院)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
学会思考:通过视觉感知的自我改进训练提升多模态推理
Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao
机构
*
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence(计算机学院、多媒体软件国家工程研究中心、人工智能研究院)
;
Hubei Key Laboratory of Multimedia(湖北多媒体重点实验室)
;
Network Communication Engineering, Wuhan University, China(网络通信工程、武汉大学,中国)
;
The University of Sydney, Australia(悉尼大学,澳大利亚)
;
Nanyang Technological University, Singapore(南洋理工大学,新加坡)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
机构
*
Beijing Institute of Technology(北京理工大学)
;
AMAP, Alibaba Group(阿里集团AMAP)
;
City University of Hong Kong(香港城市大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Yangtze Delta Region Academy of Beijing Institude of Technology, Jiaxing, China(北京理工大学扬子江地区学院,嘉兴,中国)
CommentsDORA stress-tests LLM agents on real-world disaster operations that demand comprehensive orchestration of 108 specialized tools over heterogeneous geospatial data