Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs
图像胜过言语吗?探讨文本误导在VLMs中的影响
Chi Zhang, Wenxuan Ding, Jiale Liu, Mingrui Wu, Qingyun Wu, Ray Mooney
机构
*
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
New York University(纽约大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
AG2ai, Inc.(AG2ai公司)
MV-S2V: Multi-View Subject-Consistent Video Generation
MV-S2V:多视图主体一致视频生成
Ziyang Song, Xinyu Gong, Bangya Liu, Zelin Zhao
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Texas at Austin(德克萨斯大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Georgia Institute of Technology(佐治亚理工学院)
机构
*
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai 200240, China(上海交通大学自动化与智能感知学院)
;
Department of Civil, Architectural, and Environmental Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校土木、建筑与环境工程系)
;
College of Electrical Engineering, Zhejiang University(浙江大学电气工程学院)
;
School of Automotive Studies, Tongji University(同济大学汽车学院)
;
Key Laboratory of Road and Traffic Engineering of Ministry of Education, Tongji University(同济大学交通工程教育部长三角交通工程重点实验室)
;
School of Architecture, The University of Texas at Austin(德克萨斯大学奥斯汀分校建筑系)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
ReVision:一个用于隐私保护任务导向视觉指令重写的数据集和基线VLM
Abhijit Mishra, Mingda Li, Hsiang Fu, Richard Noh, Minji Kim
机构
*
School of Information(信息学院)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Department of Statistics and Data Science(统计与数据科学系)
;
Yale University(耶鲁大学)
;
School of Computing and Augmented Intelligence(计算与增强智能学院)