Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding
金点狙击手:VLM中的自引导视觉推理用于细粒度动作理解
Haodi Liu, Xinhang Yang, Kunda Yan, Sen Cui, Zeyu Zhang, Changshui Zhang
机构
*
Beijing National Research Center for Information Science and Technology (BNRist), Department of Automation, Tsinghua University(清华大学自动化系北京信息科学与技术国家研究中心)
;
State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院通用人工智能国家重点实验室)
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
FOCA: 面向未来的条件化用于数据高效的视觉-语言-动作适应
Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen, Trong-Bao Ho, Doanh Le, Tan Q. Nguyen, Thien-Loc Ha, Nhiem Tran, Bao Thach, Nhat X. Tran, Tuan A. Tran, Artur Habuda, Philip Lund Møller, Tran Nguyen Le, Daniel Sonntag, Matthias Niepert, Khoa D. Doan, Vu Duong, Hung Ngo, Minh N. Vu, Duy M. H. Nguyen, An Thai Le, Ngo Anh Vien
机构
*
Center for AI Research, VinUniversity, Vietnam
;
University of Utah, USA
;
German Research Center for Artificial Intelligence (DFKI)
;
Technical University of Denmark, Denmark
;
University of Oldenburg, Germany
;
University of Stuttgart, Germany
;
Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany
NEST: Narrative Event Structures in Time for Long Video Understanding
NEST:面向长视频理解的时间叙事事件结构
Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker, Hani Alomari, Chia-Wei Tang, Anushka Sivakumar, Zaber Ibn Abdul Hakim, Shaurya Mallampati, Chris Thomas
机构
*
Department of Computer Science, Virginia Tech(弗吉尼亚理工大学计算机科学系)
FEMOT: Multi-Object Tracking using Frame and Event Cameras
FEMOT: 使用帧和事件摄像机的多目标跟踪
Shiao Wang, Xiao Wang, Chao Wang, Yitao Li, Menghao Liu, Bo Jiang, Yaowei Wang, Yonghong Tian, Jin Tang
机构
*
School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)
;
Peng Cheng Laboratory(鹏城实验室)
;
National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理全国重点实验室)
;
School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University(北京大学深圳研究生院电子与计算机工程学院)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
CommentsAccepted for publication at the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026). 6 pages, 3 figures, 1 table
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
GIRL-DETR: 梯度隔离强化学习用于视频时刻检索
Shihang Zhang, Mingjin Kuai, Ye Wei, Zhen Zhang, Wei Ji
机构
*
College of Electronics and Information Engineering, Sichuan University(四川大学电子信息工程学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
从博弈视角重新思考弱监督视频时间定位
Xiang Fang, Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen, Jianfeng Dong, Keke Tang, Pan Zhou, Yu Cheng, Daizong Liu
机构
*
Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室)
;
Hubei Engineering Research Center on Big Data Security(大数据安全工程研究中心)
;
School of Cyber Science and Engineering(网络安全科学与工程学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
University of Central Florida(佛罗里达中央大学)
;
Zhejiang Gongshang University(浙江工商大学)
;
Guangzhou University(广州大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
看见 vs. 相信:评估开源多模态大模型在反直觉场景中的语言偏见
Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
机构
*
Zhejiang University(浙江大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))