ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
ViSTR-Bench:多模态大语言模型能否从动态场景中的连续视觉线索进行推理?
Han Li, Si Liu, Zehao Huang, Dongxin Lyu, Longfei Xu, Jiahui Fu, Daxin Tian, Yuliang Xiu, Naiyan Wang
机构
*
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
Zhongguancun Academy(中关村科学城)
;
School of Engineering, Westlake University(西湖大学工学院)
;
School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Vision-Language-Policy Model for Dynamic Robot Task Planning
用于动态机器人任务规划的视觉-语言-策略模型
Jin Wang, Kim Tien Ly, Jacques Cloete, Jin Jin, Nikos Tsagarakis, Ioannis Havoutis
机构
*
Dynamic Robot Systems Group, Oxford Robotics Institute, University of Oxford(牛津大学机器人研究所动态机器人系统组)
;
Humanoids and Human-Centered Mechatronics (HHCM), Istituto Italiano di Tecnologia(意大利技术研究所人形与以人为中心的机电系统)
机构
*
School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
超越语义的艺术:用于多关系表示的层状信息对比学习
Ludovica Schaerf, Antonio Purificato, Piera Riccio, Fabrizio Silvestri, Noa Garcia
机构
*
University of Zurich(苏黎世大学)
;
Max Planck Institute Bibliotheca Hertziana(马克斯·普朗克赫兹iana图书馆研究所)
;
Sapienza University of Rome(罗马第一大学)
;
Amazon(亚马逊)
;
University of Amsterdam(阿姆斯特丹大学)
;
The University of Osaka(大阪大学)
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
TSRouter:用于时间序列推理的动态模态-模型选择
Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou
机构
*
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
MBZUAI(Mohamed Bin Zayed University of Artificial Intelligence)
DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
MAGIS:基于证据的多智能体推理用于可解释的斜视临床决策
Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan
机构
*
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院)
;
Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong(汕头大学·香港中文大学联合汕头国际眼科中心)
;
School of Artificial Intelligence, Guangzhou City Polytechnic(广州城市职业学院人工智能学院)
;
Medical College, Shantou University(汕头大学医学院)
;
College of Engineering, Shantou University(汕头大学工学院)
;
Department of Ophthalmology, Xinhua Hospital Affiliated to Shanghai Jiaotong University School of Medicine(上海交通大学医学院附属新华医院眼科)
;
Shenzhen Loop Area Institute(深圳河套学院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots
AEGIS:用于开源液体处理机器人的检测感知协议验证和运行时监测
Priyanka V. Setty, Arvind Ramanathan, Ian Foster, Rick Stevens
机构
*
Data Science and Learning Division, Argonne National Laboratory(阿贡国家实验室数据科学与学习部)
;
Department of Computer Science, University of Chicago(芝加哥大学计算机科学系)
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
VCG-Bench:迈向统一的视觉导向基准,用于结构化生成与编辑
Xiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang, Yuyao Zhai, Kaitao Lin, Liang Chen, Gai Yuhang, Yuyu Luo, Qiang Wang, Xiaowen Chu
机构
*
The Hong Kong University of Science and Technology (GuangZhou)(香港科学与技术大学(广州))
;
Huawei Technologies Co., Ltd(华为技术有限公司)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
South China University of Technology(华南理工大学)