SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
SoPE: 基于球坐标的位置嵌入:增强3D大视觉-语言模型的空间感知
Guanting Ye, Qiyan Zhao, Wenhao Yu, Liangyu Yuan, Mingkai Li, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Qing Jiang, Ka-Veng Yuen
机构
*
University of Macau(澳门大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Hefei University of Technology(合肥工业大学)
;
National University of Singapore(新加坡国立大学)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
学习关键要素:通过相对误差驱动的样本选择进行优先概念学习
Shivam Chandhok, Qian Yang, Oscar Manas, Kanishk Jain, Leonid Sigal, Aishwarya Agrawal
机构
*
Mila - Québec AI Institute(魁北克人工智能研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Université de Montréal(蒙特利尔大学)
;
Vector Institute for AI(人工智能矢量研究所)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
机构
*
Shenzhen Technology University(深圳技术大学)
;
Shenzhen University of Information Technology(深圳信息科技学院)
;
University of Washington(华盛顿大学)
;
Foundation Model Team, Meituan(美团基础模型团队)
;
National University of Singapore(新加坡国立大学)
;
Tongji University(同济大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Technical University of Darmstadt(德累斯顿技术大学)
;
Max-Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
University of Tuebingen(图宾根大学)
;
Tubingen AI Center(图宾根人工智能中心)
;
Max Planck Institute for Informatics(马克斯·普朗克信息研究所)
Sparrow: Text-Anchored Window Attention with Visual-Semantic Glimpsing for Speculative Decoding in Video LLMs
Sparrow: 一种基于文本锚定窗口注意力与视觉语义窥视的推测解码方法用于视频大语言模型
Libo Zhang, Zhaoning Zhang, Wangyang Hong, Peng Qiao, Dongsheng Li
机构
*
National Key Laboratory of Parallel and Distributed Computing(平行与分布式计算国家重点实验室)
;
College of Computer Science and Technology(计算机科学与技术学院)
;
National University of Defense Technology(国防科技大学)
EIDOS: Latent-Space Predictive Learning for Time Series Foundation Models
EIDOS:用于时间序列基础模型的潜在空间预测学习
Xinxing Zhou, Qingren Yao, Yiji Zhao, Chenghao Liu, Flora Salim, Xiaojie Yuan, Yanlong Wen, Ming Jin
机构
*
Nankai University(南开大学)
;
Eindhoven University of Technology(埃因霍温理工大学)
;
Yunnan University(云南大学)
;
University of New South Wales(新南威尔士大学)
;
Griffith University(格里菲斯大学)
机构
*
School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)
;
School of Computing, National University of Singapore(计算学院,新加坡国立大学)
LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing
基于LLM的诊断证据对齐用于有限配对情况下的医学视觉-语言预训练
Huimin Yan, Liang Bai, Xian Yang, Long Chen
机构
*
Institute of Intelligent Information Processing, Shanxi University(山西大学智能信息处理研究所)
;
Alliance Manchester Business School, The University of Manchester(曼彻斯特大学阿利安斯商学院)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
对高歧义时空视频 footage 的混合标注流程的评估
Juan Gutiérrez, Victor Gutiérrez, Ángel Mora, Silvia Rodriguez, José Luis Blanco
机构
*
IPTC(信息处理与电信中心)
;
Telecommunications Center, Escuela Técnica Superior de Ingenieros de Telecomunicación, Av. Complutense 30, 28040 Madrid, Spain(电信中心,电信工程师高级学院,Complutense大道30号,28040马德里,西班牙)
How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
视觉标记能承载多少信息?VLMs中识别限制的缩放定律
Shuxin Zhuang, Zi Liang, Runsheng Yu, Hongzong Li, Rong Feng, Shiqin Tang, Youzhi Zhang
机构
*
City University of Hong Kong(香港城市大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
Centre for Artificial Intelligence and Robotics, Chinese Academy of Sciences(中国科学院人工智能与机器人研究院)