Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences
顺序很重要:LVLMs作为图像序列时间推理的评判者
Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes
机构
*
University of Bologna(博洛尼亚大学)
;
NOVA School of Science and Technology(NOVA科技学院)
;
NOVA Laboratory for Computer Science and Informatics(NOVA计算机科学与信息实验室)
机构
*
Shandong University(山东大学)
;
City University of Hong Kong(香港城市大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Shandong Jianzhu University(山东建筑大学)
机构
*
South China University of Technology(华南理工大学)
;
Westlake University(西湖大学)
;
Johns Hopkins University(约翰·霍普金斯大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
机构
*
Dalian University of Technology(大连理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Zhejiang University(浙江大学)
;
WeChat, Tencent Inc.(腾讯微信)
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
推进面向艺术字的场景文本识别:数据集与方法
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Haojie Zhang, Chong Sun, Chen Li, Jing Lyu, Zhineng Chen
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所)
;
Shanghai Key Laboratory of Multimodal Embodied AI, Fudan University(复旦大学上海市多模态具身人工智能重点实验室)
;
WeChat Vision, Tencent Inc.(腾讯微信视觉团队)
;
South China University of Technology(华南理工大学)
机构
*
School of Biomedical Engineering \& State Key Laboratory of Advanced Medical Materials
;
Devices, ShanghaiTech University, Shanghai, China
;
National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, Wuhan, China
;
School of Electronic Information
;
Communications, Huazhong University of Science
;
Department of Computer Science \& Engineering, Texas A\&M University, USA
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
MirrorCheck: 视觉-语言模型的高效对抗防御
Samar Fares, Klea Ziu, Toluwani Aremu, Nikita Durasov, Martin Takáč, Pascal Fua, Ivan Laptev, Karthik Nandakumar
机构
*
Mohamed Bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能大学)
;
NVIDIA
;
École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
;
Michigan State University(密歇根州立大学)
MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
MCR-VQGAN:一种用于阿尔茨海默病成像的可扩展且经济高效的Tau PET合成方法
Jin Young Kim, Jeremy Hudson, Jeongchul Kim, Qing Lyu, Christopher T. Whitlow
机构
*
Department of Biomedical Engineering, Wake Forest University School of Medicine(生物医学工程系,威克森林大学医学院)
;
Department of Radiology, Wake Forest School of Medicine(放射学系,威克森林医学院)
;
Department of Radiology and Biomedical Imaging, Yale School of Medicine(放射学与生物医学成像系,耶鲁医学院)