Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language
更少步骤,更优性能:基于语言的高效跨模态视频片段修剪用于视频时刻检索
Xiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou, Zichuan Xu, Wenzheng Xu, Junyang Chen, Renfu Li
机构
*
Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science of Technology(湖北大数据安全工程研究中心,网络安全学院,华中科技大学)
;
Peking University(北京大学)
;
Henan University(河南大学)
;
Dalian University of Technology(大连理工大学)
;
Sichuan University(四川大学)
;
Shenzhen University(深圳大学)
;
Huazhong University of Science and Technology(华中科技大学)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
如何以及想象什么?统一多模态模型中的视觉思维用于跨视角空间推理
Qian Yang, Ankur Sikarwar, Huy Le, Le Zhang, Zhuan Shi, Perouz Taslakian, Aishwarya Agrawal
机构
*
Mila - Québec AI Institute(蒙特利尔AI研究所)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
ServiceNow AI Research(ServiceNow人工智能研究)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models
MAIL++: 视觉语言模型的多模态双向智能体层
Kaixiang Chen, Pengfei Fang, Hui Xue
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用国家重点实验室(东南大学),中华人民共和国教育部,中国)
机构
*
Sun Yat-sen University(中山大学)
;
Shandong Normal University(山东师范大学)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
Southeast University(东南大学)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
Cattle-CLIP:一种用于从视频中识别牛行为的多模态框架
Huimin Liu, Jing Gao, Daria Baran, AxelX Montout, Neill W Campbell, Andrew W Dowsey
机构
*
Bristol Veterinary School, University of Bristol(布里斯托尔大学布里斯托尔兽医学院)
;
School of Computer Science, University of Bristol(布里斯托尔大学计算机科学学院)
;
Joint Institute of Qingdao Huanghai University and WeITec, Qingdao Huanghai University(青岛黄海学院与WeITec联合学院,青岛黄海学院)
Camouflage-aware Image-Text Retrieval via Expert Collaboration
通过专家协作实现伪装意识的图像-文本检索
Yao Jiang, Zhongkuan Mao, Xuan Wu, Keren Fu, Qijun Zhao
机构
*
College of Computer Science, Sichuan University(四川大学计算机科学学院)
;
National Key Lab of Fundamental Science on Synthetic Vision, Sichuan University(四川大学视觉合成图形图像技术国家级重点实验室)
Huy Hoang Nguyen, Cédric Jung, Shirin Salehi, Tobias Glück, Anke Schmeink, Andreas Kugi
机构
*
AIT Austrian Institute of Technology(奥地利理工学院)
;
Automation & Control Institute, Technical University of Vienna(维也纳技术大学自动化与控制研究所)
;
Chair of Information Theory and Data Analytics (INDA), RWTH Aachen University(亚琛工业大学信息理论与数据分析教席)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
面向可扩展多模态推理的推理对齐感知解耦
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Xin Jin, Zhenguo Li, James T. Kwok, Yu Zhang
机构
*
Southern University of Science and Technology(南方科技大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Huawei Cloud Project(华为云项目)
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
LLaVAShield: 保障视觉语言模型中的多模态多轮对话安全
Guolei Huang, Qinzhi Peng, Gan Xu, Yao Huang, Yuxuan Lu, Yongjun Shen
机构
*
Southeast University(东南大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Zhejiang University of Technology(浙江工业大学)
;
Tsinghua University(清华大学)
;
RealAI
MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection
MGCR-Net:多模态图条件视觉-语言重建网络用于遥感变化检测
Chengming Wang, Guodong Fan, Jinjiang Li, Min Gan, C. L. Philip Chen
机构
*
School of Computer Science and Technology, Shandong Technology and Business University(山东科技职业大学计算机科学与技术学院)
;
School of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)