OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding
OptiSAR-Net++:跨领域遥感视觉定位的大型基准和无Transformer框架
Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
机构
*
South China Normal University(华南师范大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Tongji University(同济大学)
;
Shanghai Research Institute for Intelligent Autonomous Systems(上海自主智能无人系统科学中心)
;
State Key Laboratory of Intelligent Autonomous Systems(智能自主系统国家重点实验室)
;
Frontiers Science Center for Intelligent Autonomous Systems(智能自主系统前沿科学中心)
机构
*
School of Computer, National University of Defense Technology, Changsha, China(国防科技大学计算机学院,中国长沙)
;
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China(中山大学深圳校区计算机科学与技术学院,中国深圳)
;
Hong Kong Polytechnic University, Department of Computing, Hong Kong(香港理工大学计算学院,香港)
;
College of Computer Science and Electronic Engineering, Hunan University, Changsha, China(湖南大学计算机科学与电子工程学院,中国长沙)
A Large-Scale Remote Sensing Dataset and VLM-based Algorithm for Fine-Grained Road Hierarchy Classification
一个大规模遥感数据集和基于视觉-语言的算法用于细粒度道路层级分类
Ting Han, Xiangyi Xie, Yiping Chen, Yumeng Du, Jin Ma, Aiguang Li, Jiaan Liu, Yin Gao
机构
*
School of Geospatial Engineering and Science(地理工程与科学学院)
;
Laboratory of Intelligent Collaborative Computing(智能协同计算实验室)
;
United Nations University Institute in Macau(联合国澳门研究所)
;
Moganshan Geospatial Information Laboratory(莫干山地理信息实验室)
Think and Answer ME: Benchmarking and Exploring Multi-Entity Reasoning Grounding in Remote Sensing
思考并回答ME:基准测试和探索遥感中的多实体推理 grounding
Shuchang Lyu, Haiquan Wen, Guangliang Cheng, Meng Li, Zheng Zhou, You Zhou, Dingding Yao, Zhenwei Shi
机构
*
Beihang University, Beijing, China(北京航空航天大学)
;
University of Liverpool, Liverpool, UK(利物浦大学)
;
Institute of Acoustics, Chinese Academy of Sciences, Beijing, China(中国科学院声学研究所)
Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
重新思考MLLM本身作为分割器:仅用一个分割标记
Anqi Zhang, Xiaokang Ji, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
机构
*
Beijing Institute of Technology(北京理工大学)
;
University of Birmingham(伯明翰大学)
;
Beijing Jiaotong University(北京交通大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
ForensicZip: 更多令牌更好但并非必要在取证视觉-语言模型中
Yingxin Lai, Zitong Yu, Jun Wang, Linlin Shen, Yong Xu, Xiaochun Cao
机构
*
Great Bay University(大湾大学)
;
Shenzhen University(深圳大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络安全科学与技术学院)
专题命中
视觉定位与Grounding
:vision-language model(title);multimodal large language model(abstract);分类 cs.CV
Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time
用眼睛倾听:跨时空的自体视觉共指基准测试
Weijie Zhou, Xuantang Xiong, Zhenlin Hu, Xiaomeng Zhu, Chaoyang Zhao, Honghui Dong, Zhengyou Zhang, Ming Tang, Jinqiao Wang
机构
*
Beijing Jiaotong University(北京交通大学)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences (CASIA)(基础模型研究中心、自动化研究所、中国科学院(CASIA))
;
Tencent Robotics X(腾讯机器人X)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST)(计算机科学与工程系、香港科学与技术大学(HKUST))
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
T2SGrid: 视频时间定位的时空网格化
Chaohong Guo, Yihan He, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long
机构
*
South China University of Technology(南方科技大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室)
;
Bytedance Inc(字节跳动公司)
机构
*
Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(电子科技大学复杂系统建模与仿真实验室,计算机科学与技术学院)
;
Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江省空间信息感知与传输重点实验室,电子科技大学)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系)