TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
TAR:基于时间锚的视频时间定位推理
Chaohong Guo, Xun Mo, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long
机构
*
South China University of Technology(华南理工大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))
;
Bytedance Inc.(字节跳动有限公司)
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架
Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
机构
*
organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom
;
organization= SpaceTimeLab, Department of Civil, Environmental
;
Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China
;
organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom
;
organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
China University of Petroleum (Beijing)(中国石油大学(北京))
;
Hainan Institute of China University of Petroleum (Beijing)(中国石油大学(北京)海南学院)
;
South China Normal University(华南师范大学)
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
学习聚焦与精确裁剪:一种带有信息缺口和接地损失的强化学习框架用于多模态大语言模型
Xuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen, Yao Liu, Yue Wu, Tao Gong, Qi Chu, Nenghai Yu
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络空间安全学院)
;
Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院)
;
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
HM-Bench:多模态大语言模型在高光谱遥感中的综合基准
Xinyu Zhang, Zurong Mai, Qingmei Li, Zjin Liao, Yibin Wen, Yuhang Chen, Xiaoya Fan, Chan Tsz Ho, Bi Tianyuan, Haoyuan Liang, Ruifeng Su, Zihao Qian, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
机构
*
Sun Yat-sen University(中山大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
China Agricultural University(中国农业大学)
;
Southwest Jiaotong University(西南交通大学)
;
Southwest University(西南大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
机构
*
Tsinghua University(清华大学)
;
Sun Yat-sen University(中山大学)
;
Independent Researcher(独立研究员)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Guangdong Laboratory of AI and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
机构
*
School of Mathematics, Shandong University(山东大学数学学院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Yeshiva University(叶史瓦大学)
;
Qilu University of Technology(齐鲁工业大学)
;
Shandong Normal University(山东师范大学)
;
Chuzhou University(滁州学院)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
基于可定位性的自适应推理图像地理定位方法
Bo Yu, Fengze Yang, Yiming Liu, Chao Wang, Xuewen Luo, Taozhe Li, Ruimin Ke, Xiaofan Zhou, Chenxi Liu
机构
*
The University of Utah(犹他大学)
;
The University of Oklahoma(俄克拉荷马大学)
;
Worcester Polytechnic Institute(沃斯特理工学院)
;
Rensselaer Polytechnic Institute(雷士利理工学院)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)