Scaling Capability in Token Space: An Analysis of Large Vision Language Model
令牌空间中的扩展能力:对大视觉语言模型的分析
机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) ; Key Laboratory of Intelligent Detection and the Internet of Things in Manufacturing, Ministry of Education(教育部智能制造智能检测与物联网重点实验室) ; Guangdong Provincial Key Laboratory of Intelligent Systems and Optimization Integration(广东省智能系统与优化集成重点实验室) ; Medical Science Data-driven Mathematics Team, RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences(RIKEN跨学科理论与数学科学中心医学科学数据驱动数学团队) ; Medical Data Mathematical Reasoning Special Team, RIKEN Center for Integrative Medical Sciences(RIKEN整合医学科学中心医学数据数学推理特别团队) ; Department of Artificial Intelligence Medicine, Chiba University(千叶大学人工智能医学系) ; Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(RIKEN高级人工智能项目中心张量学习团队)
AI总结 本研究通过理论分析和实证验证,揭示了视觉语言模型在视觉令牌数量上的扩展规律,发现不同数量的视觉令牌对应不同的扩展模式,并提出了扩展指数与视觉令牌表示相关结构的关系。
Journal ref Journal of Machine Learning Research, volume 26, number 253, page 1--61, 2025