机构
*
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室)
;
Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)
;
Xi’an Jiaotong University(西安交通大学)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
为什么强化学习比监督微调在泛化能力上更优?一种以数据为中心的视觉语言模型后训练视角
Aojun Lu, Tao Feng, Hangjie Yuan, Wei Li, Yanan Sun
机构
*
College of Computer Science, Sichuan University(四川大学计算机科学学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
Innovator-VL:一种用于科学发现的多模态大语言模型
Zichen Wen, Boxue Yang, Shuang Chen, Yaojie Zhang, Yuhang Han, Junlong Ke, Cong Wang, Yicheng Fu, Jiawang Zhao, Jiangchao Yao, Xi Fang, Zhen Wang, Henxing Cai, Lin Yao, Zhifeng Gao, Yanhui Hong, Nang Yuan, Yixuan Li, Guojiang Zhao, Haoyi Tao, Nan Wang, Han Lyu, Guolin Ke, Ning Liao, Xiaoxing Wang, Kai Chen, Zhiyu Li, Feiyu Xiong, Sihan Hu, Kun Chen, Yanfeng Wang, Weinan E, Linfeng Zhang, Linfeng Zhang
机构
*
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou, Zihan You, Hangjun Ye, Jie Ma, Long Chen, Yan Wang
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Xiaomi EV(小米电动车)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
专题命中
VLM训练与架构
:vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI
机构
*
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))
;
Hong Kong Baptist University, China(香港 Baptist大学)
;
City University of Hong Kong, China(香港城市大学)
机构
*
Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学)
;
Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学)
;
Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科)
;
Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院)
;
School of Medicine Shaoxing University(绍兴大学医学院)
专题命中
VLM训练与架构
:multimodal large language model(title);LLaVA(abstract);分类 cs.AI、cs.LG
AI总结
Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
令牌空间中的扩展能力:对大视觉语言模型的分析
Tenghui Li, Guoxu Zhou, Xuyang Zhao, Qibin Zhao
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
Key Laboratory of Intelligent Detection and the Internet of Things in Manufacturing, Ministry of Education(教育部智能制造智能检测与物联网重点实验室)
;
Guangdong Provincial Key Laboratory of Intelligent Systems and Optimization Integration(广东省智能系统与优化集成重点实验室)
;
Medical Science Data-driven Mathematics Team, RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences(RIKEN跨学科理论与数学科学中心医学科学数据驱动数学团队)
;
Medical Data Mathematical Reasoning Special Team, RIKEN Center for Integrative Medical Sciences(RIKEN整合医学科学中心医学数据数学推理特别团队)
;
Department of Artificial Intelligence Medicine, Chiba University(千叶大学人工智能医学系)
;
Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(RIKEN高级人工智能项目中心张量学习团队)
专题命中
VLM训练与架构
:vision language model(title);vision-language model(abstract);分类 cs.AI、cs.LG
Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions
弥合差距:面向标准化考试题目的视觉语言模型数据驱动微调
Egemen Sert, Şeyda Ertekin
机构
*
organization= Department of Computer Engineering, Middle East Technical University (METU) , city= Ankara , country= Türkiye
;
organization= METU-DTX Digital Transformation \& Innovation Centre, METU , city= Ankara , country= Türkiye
专题命中
VLM训练与架构
:vision language model(title,abstract);分类 cs.CV、cs.AI
机构
*
Morgan Stanley(摩根士丹利)
;
Clemson University(克莱姆森大学)
;
Arizona State University(亚利桑那州立大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
University of Arizona(亚利桑那大学)
;
University of Notre Dame(圣母大学)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
释放视觉-语言模型在长尾多标签视觉识别中的潜力
Wei Tang, Zuo-Zheng Wang, Kun Zhang, Tong Wei, Min-Ling Zhang
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of Computer Network and Information Integration (Southeast University)(计算机网络与信息集成重点实验室)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(Mohamed bin Zayed人工智能大学)
;
Carnegie Mellon University(卡内基梅隆大学)