Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
基于检测器的视频大语言模型用于高效的时空定位
Shida Gao, Feng Xue, Xiangfeng Wang, Anlong Ming, Zhaowen Lin, Haiyang Zhang, Teng Long, Nicu Sebe, Yihua Shao, Haozhe Wang, Wei Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Trento(特伦特大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hong Kong University of Science and Technology(香港科技大学)
;
ZTE Corporation(中兴通讯)
专题命中
视觉定位与Grounding
:grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
机构
*
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
Zhengzhou Advanced Research Institute of Harbin Institute of Technology(哈尔滨工业大学郑州先进研究院)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(深圳先进技术大学人工智能研究院)
Journal refProceedings of the 9th International Conference on Medical Imaging with Deep Learning, Proceedings of Machine Learning Research 315 (2026) 2958-2986
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
视觉语言模型能预测未来状态吗?从逆动力学引导世界模型
Yifu Qiu, Yftah Ziser, Anna Korhonen, Shay B. Cohen, Edoardo M. Ponti
机构
*
Institute for Language, Cognition and Computation, University of Edinburgh(语言、认知与计算研究所,爱丁堡大学)
;
Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)
;
NVIDIA(NVIDIA公司)
;
University of Groningen(格罗宁根大学)
Comments24 pages, 15 figures, 4 tables. Model weights at https://huggingface.co/NCSOFT/VARCO-VISION-14B. Benchmarks released at NCSOFT's HuggingFace repositories (K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W). VARCO-VISION is an open-source Korean-English VLM with OCR, grounding, and referring capabilities
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
LocateAnything: 基于并行框解码的快速高质量视觉定位
Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei, Yangzhou Liu, Zhiqi Li, Yunze Man, Guo Chen, Andrew Tao, Guilin Liu, Jan Kautz, Lei Zhang, Zhiding Yu
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Princeton University(普林斯顿大学)
;
Nanjing University(南京大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
National Institute of Health Data Science, Peking University(北京大学健康数据科学国家研究院)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Tianjin Institute of Cardiology, the Second Hospital of Tianjin Medical University(天津医科大学第二医院心内科)
;
National University of Singapore(新加坡国立大学)
;
Jarvis Lab, Tencent(腾讯 Jarvis实验室)
;
HeartVoice Medical Technology(HeartVoice医疗科技)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract)
Language Movement Primitives: Grounding Language Models in Robot Motion
语言运动基元:将语言模型锚定在机器人运动中
Yinlong Dai, Benjamin A. Christie, Daniel J. Evans, Dylan P. Losey, Simon Stepputtis
机构
*
Collab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(合作组,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
;
TEA Lab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(TEA实验室,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
通过对象中心和几何约束实现抗杂波的视觉-语言-动作模型
Khoa Vo, Taisei Hanyu, Yuki Ikebe, Trong Thang Pham, Nhat Chung, Minh Nhat Vu, Duy Nguyen Ho Minh, Anh Nguyen, Anthony Gunderman, Chase Rainwater, Ngan Le
机构
*
University of Arkansas(阿拉巴马大学)
;
National University of Singapore(新加坡国立大学)
;
TU Wien(维也纳技术大学)
;
Max Planck Research School for Intelligent Systems(智能系统马克斯·普朗克研究学校)
;
University of Stuttgart(斯图加特大学)
;
University of Liverpool(利物浦大学)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
Soo Yong Kim, Suin Cho, Vincent-Daniel Yun, Gyeongyeon Hwang
机构
*
A.I.MATICS Inc(A.I.MATICS公司)
;
Boston University(波士顿大学)
;
University of Southern California(南加州大学)
;
Heuron(Heuron公司)
;
MODULABS, Open Neural Networks Research Lab(MODULABS,开放神经网络研究实验室)
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
MV-GEL:语言驱动的多视图网格几何实体定位
Kartik Bali, Roland Aydin
机构
*
Helmholtz Zentrum Hereon(亥姆霍兹赫里恩研究中心)
;
Institute for Continuum and Material Mechanics, Hamburg University of Technology(汉堡工业大学连续介质与材料力学研究所)
;
German Research Center for Artificial Intelligence(德国人工智能研究中心)
专题命中
视觉定位与Grounding
:VLM(summary_cn,abstract);vision language model(abstract);grounding(abstract);分类 cs.CV