ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
ForensicZip: 更多令牌更好但并非必要在取证视觉-语言模型中
Yingxin Lai, Zitong Yu, Jun Wang, Linlin Shen, Yong Xu, Xiaochun Cao
机构
*
Great Bay University(大湾大学)
;
Shenzhen University(深圳大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络安全科学与技术学院)
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
GeoAlignCLIP: 通过多粒度一致性学习增强遥感中的细粒度视觉-语言对齐
Xiao Yang, Ronghao Fu, Zhuoran Duan, Zhiwen Lin, Xueyan Liu, Bo Yang
机构
*
College of Computer Science and Technology, Jilin University, Changchun 130012, China(吉林大学计算机科学与技术学院,长春130012,中国)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University(教育部符号计算与知识工程重点实验室,吉林大学)
机构
*
Institute of Trustworthy Embodied AI, Fudan University(可信具身AI研究院,复旦大学)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身AI重点实验室)
;
City University of Hong Kong(香港城市大学)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
SPEX:一种用于光谱遥感图像土地覆盖提取的视觉-语言模型
Dongchen Si, Di Wang, Erzhong Gao, Xiaolei Qin, Liu Zhao, Jing Zhang, Minqiang Xu, Jianbo Zhan, Jianshe Wang, Lin Liu, Bo Du, Liangpei Zhang
机构
*
College of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
iFlytek Co., Ltd.(iFlytek公司)
;
National Engineering Research Center of Speech and Language Information Processing(语音与语言信息处理国家工程研究中心)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Zhongguancun Academy(中关村学院)
;
National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心)
;
Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(湖北省多媒体与网络通信工程重点实验室,武汉大学)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学)
Vision-Language Feature Alignment for Road Anomaly Segmentation
视觉-语言特征对齐用于道路异常分割
Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue
机构
*
School of Computer Science, Fudan University(复旦大学计算机科学学院)
;
Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)
MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
MedGPT-oss: 为生物医学训练一个通用的视觉-语言模型
Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
机构
*
Department of Computer Science and Engineering, Lehigh University(莱斯大学计算机科学与工程系)
;
Department of Computer Science and Engineering, University of Notre Dame(圣母大学计算机科学与工程系)
;
Department of Health Outcomes & Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系)
;
Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
AI Technology Center, NVIDIA(NVIDIA人工智能技术中心)
CommentsThis is the extended version of the paper accepted in ICASSP'26, which will be publicly available in May. Authors' contributions may vary among the versions