ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)
An empirical study of the effect of video encoders on Temporal Video Grounding
Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Felipe Bravo-Marquez
机构
*
Department of Computer Science, University of Chile(计算机科学系,智利大学)
;
CENIA and IMFD(CENIA和IMFD)
;
Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)
;
National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究所)
机构
*
aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts
;
Telecommunications, China [1ex] dKey Laboratory of Computing Power Network
;
Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet
;
Service Computing, Shandong Fundamental Research Center for Computer Science, China
Jingchao Wang, Wenlong Zhang, Dingjiang Huang, Hong Wang, Yefeng Zheng
机构
*
School of Data Science(数据科学学院)
;
Engineering East China Normal University Shanghai, China(工程 东华师范大学 上海中国)
;
OpenScience Lab Shanghai AI Laboratory Shanghai, China(OpenScience Lab 上海AI实验室 上海中国)
;
School of Life Science(生命科学学院)
;
Technology Xi'an Jiaotong University Xi'an, China(技术 西安交通大学 西安中国)
;
Medical Artificial Intelligence Laboratory Westlake University Hangzhou, China(医学人工智能实验室 西湖大学 杭州中国)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
Seunghun Lee, Jiwan Seo, Jeonghoon Kim, Sungho Moon, Siwon Kim, Haeun Yun, Hyogyeong Jeon, Wonhyeok Choi, Jaehoon Jeong, Zane Durante, Sang Hyun Park, Sunghoon Im
机构
*
Sun Yat-sen University(中山大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Key Laboratory of Machine Intelligence and Advanced Computing(人工智能与先进计算重点实验室)
机构
*
University of California, San Diego(加州大学圣地亚哥分校)
;
ByteDance(字节跳动)
;
The University of Queensland(昆士兰大学)
;
University of Southern California(南加州大学)
;
University at Buffalo(布法罗大学)
;
University of California, Merced(加州大学默塞德分校)
Investigating Traffic Accident Detection Using Multimodal Large Language Models
Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig
机构
*
Embedded Systems Group (Dept.-E)(嵌入式系统组)
;
Virtual Vehicle Research GmbH(虚拟车辆研究公司)
;
Control Systems Group (Dept.-E)(控制系统组)
;
Institute of Visual Computing(视觉计算研究所)
;
Graz University of Technology(格拉茨技术大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);分类 cs.CV
CommentsAccepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
Zijun Lin, Shuting He, Cheston Tan, Bihan Wen
机构
*
Nanyang Technological University(南洋理工大学)
;
Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR)
;
MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(计算与经济学交叉研究关键实验室,上海财经大学)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
Fan Yang, Yousong Zhu, Xin Li, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang, Ming Tang, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Science(中国科学院大学人工智能学院)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
;
Wuhan AI Research, Wuhan, China(武汉人工智能研究所)
专题命中
视觉定位与Grounding
:vision-language model(title);vision language model(abstract);分类 cs.CV
Journal refNon-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada