Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models
Timage: 一种用于微调视觉语言模型的文本嵌入图像生成范式
Yifeng Wu, Huimin Huang, Ruiluo Wu, Chunyi Lin, Guanhua Chen, Xian Wu, Wang Song, Ruize Han
机构
*
Fudan University(复旦大学)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
;
Tencent Jarvis Lab(腾讯贾维斯实验室)
;
Southern University of Science and Technology(南方科技大学)
专题命中
指令微调
:language model(title,abstract);large language model(abstract)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
Northeastern University(东北大学)
;
Tsinghua University(清华大学)
机构
*
Indian Institute of Technology Delhi, India(印度理工学院德里分校)
;
NVIDIA AI Technology Center, India(NVIDIA AI技术中心)
;
Jawaharlal Nehru University, India(贾瓦哈拉尔·尼赫鲁大学)
HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection
HydraPrompt: 面向合成图像检测的视觉语言模型自适应非对称框架
Senyuan Shi, Hao Tan, Zichang Tan, Shuhan Feng, Ajian Liu, Sergio Escalera, Jun Wan
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
School of Advanced Interdisciplinary Sciences (SAIS), University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
;
Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
;
University of Barcelona(巴塞罗那大学)
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model
看清并非学习:保护多模态数据免受大视觉语言模型的未经授权微调
Chengshuai Zhao, Zhen Tan, Dawei Li, Zhiyuan Yu, Huan Liu
机构
*
School of Computing
;
Augmented Intelligence, Arizona State University, Tempe, AZ, USA
;
Department of Computer Science
;
Engineering, Texas A\&M University, College Station, TX, USA
机构
*
Shandong Provincial Key Laboratory of New Power Distribution & Utilization Technology and Equipment, School of Electrical and Electronic Engineering, Shandong University of Technology(山东省新型电力分布与利用技术及设备重点实验室,山东理工大学电气与电子工程学院)
;
School of Computer Science and Technology, Shandong University of Technology(山东理工大学计算机科学与技术学院)
;
School of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)
;
Pillar of Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计 pillar)
;
School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院)
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following
无需指令的大型视觉语言模型在医疗指令遵循中的调优
Myeongkyun Kang, Soopil Kim, Xiaoxiao Li, Sang Hyun Park
机构
*
Department of Electrical and Computer Engineering, The University of British Columbia(英属哥伦比亚大学电气与计算机工程系)
;
Department of Robotics and Mechatronics Engineering, Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院机器人与机电工程系)
;
Division of Intelligent Robot, Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院智能机器人系)
;
Vector Institute(向量研究所)
;
Department of Computer Science and Engineering, Pohang University of Science and Technology (POSTECH)(釜山科学技术大学计算机科学与工程系)