机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Washington University in St. Louis(华盛顿大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Department of Surgical Oncology and General Surgery, Key Laboratory of Precision Diagnosis and Treatment of Gastrointestinal Tumours, Ministry of Education, The First Hospital of China Medical University(外科肿瘤科和普通外科,国家教育委员会胃肠道肿瘤精准诊断与治疗重点实验室,中国医科大学第一医院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Sensetime Research(商汤科技研究院)
机构
*
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院)
;
Department of Automation, Tsinghua University, China(清华大学自动化系)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
面向细粒度图像编辑评估的人类对齐MLLM评判:一个基准、框架和分析
Runzhou Liu, Hailey Weingord, Sejal Mittal, Prakhar Dungarwal, Anusha Nandula, Bo Ni, Samyadeep Basu, Hongjie Chen, Nesreen K. Ahmed, Li Li, Jiayi Zhang, Koustava Goswami, Subhojyoti Mukherjee, Branislav Kveton, Puneet Mathur, Franck Dernoncourt, Yue Zhao, Yu Wang, Ryan A. Rossi, Zhengzhong Tu, Hongru Du
机构
*
University of Virginia(弗吉尼亚大学)
;
Columbia University(哥伦比亚大学)
;
Vanderbilt University(范德比大学)
;
Adobe Research(Adobe研究)
;
Dolby Laboratories(杜比实验室)
;
Cisco Research(思科研究)
;
University of Southern California(南加州大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
University of Oregon(俄勒冈大学)
;
Texas A&M University(德克萨斯大学)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Invert4TVG: 一种具有逆向任务的时序视频定位框架,以保持动作理解能力
Zhaoyu Chen, Hongnan Lin, Yongwei Nie, Fei Ma, Xuemiao Xu, Fei Yu, Chengjiang Long
机构
*
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
;
ByteDance Inc.(字节跳动公司)
What Matters in Building Vision-Language-Action Models for Generalist Robots
在通用机器人中构建视觉-语言-动作模型所关注的关键因素
Xinghang Li, Peiyan Li, Long Qian, Minghuan Liu, Dong Wang, Jirong Liu, Bingyi Kang, Xiao Ma, Xinlong Wang, Di Guo, Tao Kong, Hanbo Zhang, Huaping Liu
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
ByteDance Research(字节跳动研究院)
;
CASIA MAIS-NLPR
;
Shanghai Jiao Tong University(上海交通大学)
;
National University of Singapore(新加坡国立大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中
GUI与屏幕智能体
:vision language model(abstract);VLM(abstract);分类 cs.CV
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Southeast University(东南大学)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
专题命中
VLM训练与架构
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
机构
*
College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学)
;
College of Computer Science, Wuhan University(计算机科学学院,武汉大学)
;
College of Computer Science, Tongji University(计算机科学学院,同济大学)
;
State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)