机构
*
School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
;
Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心)
;
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院计算与智能研究所)
;
Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University(贵州大学文字计算与认知智能教育部工程研究中心)
;
Chinese University of Hong Kong(香港中文大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
;
ByteDance Seed (China)(字节跳动种子(中国))
T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting
Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound
机构
*
University of Nottingham(诺丁汉大学)
;
University of St Andrews(圣安德鲁大学)
;
City University of Hong Kong(香港城市大学)
;
Nanyang Technology University(南洋理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models
Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu
机构
*
Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院)
;
Agency for Science, Technology and Research(科技研究局)
;
School of Information Technology(信息技术学院)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
Hyeonsu Kang, Emily Bao, Anjan Goswami
专题命中
图文多模态
:multimodal(abstract);分类 cs.CV、cs.AI
Comments39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang, Man Lan
机构
*
School of Computer Science and Technology, East China Normal University, China(上海师范大学计算机科学与技术学院)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)
机构
*
School of Computer Science, University of South China(南方大学计算机科学学院)
;
New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA)
;
School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学)
;
Department of Computer Science and Engineering, University of California, Merced(加州大学默塞德分校计算机科学与工程系)
;
Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)
机构
*
School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学)
;
School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)
;
National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学)
;
School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学)
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
Brandon James Carone, Iran R. Roman, Pablo Ripollés
机构
*
Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室)
;
Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
Yaru Chen, Faegheh Sardari, Peiliang Zhang, Ruohao Guo, Yang Xiang, Zhenbo Li, Wenwu Wang
机构
*
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(视觉、语音和信号处理中心(CVSSP),萨里大学)
;
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学)
;
National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,北京大学智能科学与技术学院)
;
College of Information and Electrical Engineering, China Agricultural University(信息与电子工程学院,中国农业大学)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong
机构
*
Beihang University(北京航空航天大学)
;
University of Science and Technology Beijing(北京科技大学)
;
School of Computer Science and Engineering(计算机科学与工程学院)
;
State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)