Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Invert4TVG: 一种具有逆向任务的时序视频定位框架,以保持动作理解能力
Zhaoyu Chen, Hongnan Lin, Yongwei Nie, Fei Ma, Xuemiao Xu, Fei Yu, Chengjiang Long
机构
*
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
;
ByteDance Inc.(字节跳动公司)
机构
*
Department of Automation, Tsinghua University, China(清华大学自动化系)
;
University of Electronic Science and Technology of China, China(电子科技大学)
;
University of Copenhagen, Denmark(哥本哈根大学)
;
Lingsu Lab, China(灵素实验室)
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
MARE: 多模态对齐与强化学习用于通过视觉-语言模型的可解释深度伪造检测
Wenbo Xu, Wei Lu, Xiangyang Luo, Jiantao Zhou
机构
*
School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University, Guangzhou 510006, China(计算机科学与工程学院,信息技术MOE实验室,广东省信息安全技术重点实验室,中山大学,广州510006,中国)
;
State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室)
;
Department of Computer and Information Science, University of Macau.(计算机与信息科学系,澳门大学)
Minh Duc Chu, Kshitij Pawar, Zihao He, Roxanna Sharifi, Ross Sonnenblick, Magdalayna Curry, Laura D'Adamo, Lindsay Young, Stuart B Murray, Kristina Lerman
机构
*
USC Information Sciences Institute(USC信息科学研究所)
;
Keck School of Medicine, USC(USC凯克医学院)
;
Department of Clinical Psychology, Drexel University(德雷塞尔大学临床心理学系)
;
Department of Psychiatry and Biobehavioral Sciences, UCLA(UCLA精神病学与生物行为科学系)
Uni-FinLLM: A Unified Multimodal Large Language Model with Modular Task Heads for Micro-Level Stock Prediction and Macro-Level Systemic Risk Assessment