机构
*
LMU Munich(慕尼黑大学)
;
Harvard University(哈佛大学)
;
University of Cambridge(剑桥大学)
;
Mina AI
;
Konrad Zuse School of Excellence in Reliable AI (relAI)(康拉德·楚泽可靠人工智能卓越学校(relAI))
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
通过定制代理推理增强VLMs在少样本多模态时间序列分类中的能力
Lin Li, Jiawei Huang, Qihao Quan, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Wenjie Feng, Jian Lou, See-Kiong Ng
机构
*
Sun Yat-sen University(中山大学)
;
Xiaomi Corporation(小米公司)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
CMTA:利用跨模态时间特征进行通用的AI生成视频检测
Hang Wang, Chao Shen, Chenhao Lin, Minghui Yang, Lei Zhang, Cong Wang
机构
*
Xi’an Jiaotong University(西安交通大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Guangdong OPPO Mobile Communications Co., Ltd.(广东OPPO移动通信有限公司)
;
City University of Hong Kong(香港城市大学)
CommentsThis paper has been withdrawn by the author. After further review, the author believes that the current version does not meet the desired standards and plans to revise the work before any potential resubmission
A multimodal and temporal foundation model for virtual patient representations at healthcare system scale
面向医疗系统规模的多模态与时间基础模型:虚拟患者表示
Andrew Zhang, Tong Ding, Sophia J. Wagner, Caiwei Tian, Ming Y. Lu, Rowland Pettit, Joshua E. Lewis, Alexandre Misrahi, Dandan Mo, Long Phi Le, Faisal Mahmood
机构
*
Department of Pathology, Mass General Brigham, Harvard Medical School(病理学系,马萨诸塞州总医院与哈佛医学院)
;
Cancer Program, Broad Institute of Harvard and MIT(癌症计划,哈佛-麻省理工Broad研究所)
;
Data Science Program, Dana-Farber Cancer Institute(数据科学计划,达纳-法伯癌症研究所)
;
Health Sciences and Technology, Harvard-MIT(健康科学与技术,哈佛-麻省理工)
;
Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛约翰·A·保罗森工程与应用科学学院,哈佛大学)
;
Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院)
;
Electrical Engineering and Computer Science, Massachusetts Institute of Technology (MIT)(电气工程与计算机科学,麻省理工学院(MIT))
;
School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland(计算机与通信科学学院,EPFL,瑞士洛桑)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
通过视觉记忆机制扩展多模态大语言模型的长视频理解
Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen, Chao Chang, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
跨模态可迁移的图像到视频攻击视频质量度量
Georgii Gotin, Ekaterina Shumitskaya, Anastasia Antsiferova, Dmitriy Vatolin
机构
*
Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学)
;
ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统研究所在可信人工智能研究中心)
;
MSU Institute for Artificial Intelligence(莫斯科大学人工智能研究所)
;
Laboratory of Innovative Technologies for Processing Video Content(视频内容处理创新技术实验室)