机构
*
Bytedance(字节跳动)
;
Peking University(北京大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The University of Hong Kong(香港大学)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学)
;
Stanford University(斯坦福大学)
;
University of Michigan(密歇根大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song, Tianxiang Hu, Chenyi Zhou, Bin Pu, Yan Zhang, Zhibo Yang, Yang Feng, Joey Tianyi Zhou, Jin Hao, Zijian Chen, Ruijia Wu, Tao Tang, Junhui Lv, Hongxia Xu, Hongwei Wang, Jun Xiao, Bin Feng, Fudong Zhu, Kenli Li, Weidi Xie, Jimeng Sun, Jian Wu, Zuozhu Liu
机构
*
College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute(浙江大学计算机科学与技术学院)
;
Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine(浙江大学口腔医院)
;
Alibaba Inc(阿里巴巴集团)
;
College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)
;
Angelalign Technology Inc.(Angelalign技术有限公司)
;
CFAR & IHPC, Agency for Science, Technology and Research(CFAR与IHPC,新加坡科技研究局)
;
Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University(上海第九人民医院正畸科,上海交通大学口腔医学院)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
Shreyas C. Dhake, Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque
机构
*
UCL Hawkes Institute(UCL哈维斯研究所)
;
University College London(伦敦大学学院)
;
Dept of Medical Physics & Biomedical Engineering(医学物理与生物医学工程系)
;
UCL(伦敦大学学院)
;
Dept of Computer Science(计算机科学系)
;
National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院)
;
Division of Informatics, Imaging and Data Science(信息学、成像与数据科学 division)
机构
*
Northeastern University(东北大学)
;
Microsoft Research(微软研究院)
;
University of Southern California(南加州大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
University of Oxford(牛津大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Beijing University of Technology(北京工业大学)
;
Tsinghua University(清华大学)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉定位与Grounding
:vision language model(abstract);分类 cs.CV
CommentsThis paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper
Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control
Rewida Ali, Cristian C. Beltran-Hernandez, Weiwei Wan, Kensuke Harada
机构
*
Department of Systems Innovation, Graduate School of Engineering Science, Osaka University(大阪大学系统创新部门,工学研究科)
;
OMRON SINIC X Corporation(OMRON SINIC X公司)
;
The National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)
ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Anurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong, Juan Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan
SENT Map -- Semantically Enhanced Topological Maps with Foundation Models
Raj Surya Rajendran Kathirvel, Zach A Chavis, Stephen J. Guy, Karthik Desingh
机构
*
Minnesota Robotics Institute (MnRI)(明尼苏达州机器人研究所)
;
Department of Computer Science and Engineering (CS&E)(计算机科学与工程系)
;
University of Minnesota(明尼苏达大学)
专题命中
视觉定位与Grounding
:grounding(abstract)
CommentsAccepted at ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
University of Oxford(牛津大学)
;
Xi’an Jiaotong University(西安交通大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
City University of Hong Kong(香港城市大学)
;
Beijing University of Technology(北京理工大学)
;
Duke University(杜克大学)
;
X-Humanoid Project(X-Humanoid 项目)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.CV
Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks
Shipeng Cen, Ying Tan
机构
*
School of Intelligence Science
;
Technology, Institute for Artificial Intellignce, Peking University, Beijing, China
;
Technology, Institute for Artificial Intellignce, National Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China