Large Vision-Language Models (LVLMs) have achieved remarkable success across diverse multimodal tasks, yet their practical deployment remains constrained by the computational burden arising from lengthy visual tokens. While visual token pruning has emerged as a promising solution, existing methods suffer from a fundamental limitation: once tokens are pruned at a specific layer, they become inaccessible to all subsequent layers, leading to premature information loss that can compromise model performance. Through empirical studies, we observe that different layers exhibit distinct visual region focus, indicating a varying optimal token subset across layers. Motivated by this insight, we propose Adaptive Layer-wise Visual Token Selection (ALVTS), a novel framework that breaks away from the conventional static token pruning paradigm. ALVTS incorporates a lightweight token selector to identify and route important tokens for further processing, while allowing less important tokens to skip the layer, thus minimizing computational redundancy. These two streams of tokens are seamlessly reintegrated before being fed into subsequent layers, facilitating adaptive compression across the entire model. Grounded in our importance consistency constrained low-rank approximation, the proposed token selection module closely emulates the full attention mechanism, effectively capturing its essential patterns without requiring model retraining. Extensive experiments on LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL validate the effectiveness of our method. With an 89% token compression ratio, ALVTS retains 96.7% of the original model's accuracy, achieving a superior efficiency-accuracy trade-off for LVLM inference.
JDCNet: Confidence-Gated Privileged-Modality Distillation for Cost-Preserving X-ray Inference
JDCNet:基于置信度门控的特权模态蒸馏用于成本保持的X射线推断
Bo Ma, Jinsong Wu, Weiqi Yan, Hongjiang Wei, Kun Liu
机构
*
Auckland University of Technology(奥克兰技术大学)
;
Guilin University of Electronic Technology(桂林电子科技大学)
;
Hikvision Technology Co., Ltd(海康威视技术有限公司)
;
Hebei University of Technology(河北工业大学)
我们研究了一个系统层面的视觉推断问题:在训练时使用昂贵的特权模态,同时保持固定成本的单模态部署路径。我们提出了JDCNet,一种置信度门控的CT到X射线蒸馏框架,在训练时,CT教师仅在教师置信度超过阈值的训练样本上提供辅助的硬目标或温度缩放目标;在部署时,学生仅使用X射线输入,并匹配监督X射线基线的参数、MAC和延迟配置。在510名患者配对的BIMCV队列中,经过患者层面的5折交叉验证,两种JDCNet配置在固定转移门下优于监督的ResNet-18基线:3切片软KL监督产生ΔBA=+0.035(95% CI [+0.011, +0.057]),中切片硬监督产生+0.033([+0.007, +0.058])。在相同的划分和门下,logit蒸馏、门控logit蒸馏、对比对齐、注意力转移、特征提示、BiomedCLIP微调以及模块增强变体均未通过。置信度门控的辅助目标因此比均匀软化的CT日志更可转移;证据仅限于一个配对队列,因此在任何部署之前都需要外部配对队列的重复验证。
英文摘要
We study a systems-level visual inference problem: using an expensive privileged modality during training while preserving a fixed-cost, single-modality deployment path. We present JDCNet, a confidence-gated CT-to-X-ray distillation framework in which the CT teacher supplies an auxiliary hard or temperature-scaled target only on training samples whose teacher confidence exceeds a threshold; at deployment the student takes X-ray input alone and matches the parameter, MAC, and latency profile of the supervised X-ray baseline. On a 510-patient same-patient paired BIMCV cohort with patient-level 5-fold cross-validation, two JDCNet configurations clear a fixed transfer gate against the supervised ResNet-18 baseline: 3-slice soft-KL supervision yields $Δ\mathrm{BA}{=}{+}0.035$ ($95\%$ CI $[{+}0.011,{+}0.057]$) and mid-slice hard supervision yields $+0.033$ ($[{+}0.007,{+}0.058]$). Under the same splits and gate, logit distillation, gated logit distillation, contrastive alignment, attention transfer, feature hints, BiomedCLIP fine-tuning, and a module-augmented variant do not pass. Confidence-gated auxiliary targets are therefore a more transferable channel than uniformly softened CT logits; the evidence is bounded to one paired cohort, so external paired-cohort replication is required before any deployment claim.
A multi-center analysis of deep learning methods for video polyp detection and segmentation
多中心分析深度学习方法在视频息肉检测和分割中的应用
Noha Ghatwary, Pedro Chavarias Solano, Mohamed Ramzy Ibrahim, Adrian Krenzer, Frank Puppe, Stefano Realdon, Renato Cannizzaro, Jiacheng Wang, Liansheng Wang, Thuy Nuong Tran, Lena Maier-Hein, Amine Yamlahi, Patrick Godau, Quan He, Qiming Wan, Mariia Kokshaikyna, Mariia Dobko, Haili Ye, Heng Li, Ragu B, Antony Raj, Hanaa Nagdy, Osama E Salem, James E. East, Dominique Lamarque, Thomas de Lange, Sharib Ali
机构
*
Computer Engineering Department, Arab Academy for Science and Technology(阿拉伯科学与技术学院计算机工程系)
;
AI in Medicine and Surgery Group, School of Computer Science, University of Leeds(莱斯特大学计算机科学学院医学与外科AI小组)
;
Department of Artificial Intelligence and Knowledge Systems, University of Würzburg(维尔茨堡大学人工智能与知识系统系)
;
CRO Centro Riferimento Oncologico IRCCS(肿瘤研究中心IRCCS)
;
Department of Computer Science at School of Informatics, Xiamen University(厦门大学信息学院计算机科学系)
;
Div. Intelligent Medical Systems, German Cancer Research Center (DKFZ)(德国癌症研究中心(DKFZ)智能医疗系统部)
;
Hangzhou Hikvision Digital Technology Co.,ltd(杭州海康威视数字技术有限公司)
;
The Machine Learning Lab, Ukrainian Catholic University(乌克兰天主教大学机器学习实验室)
Colonic polyps are well-recognized precursors to colorectal cancer (CRC), typically detected during colonoscopy. However, the variability in appearance, location, and size of these polyps complicates their detection and removal, leading to challenges in effective surveillance, intervention, and subsequently CRC prevention. The processes of colonoscopy surveillance and polyp removal are highly reliant on the expertise of gastroenterologists and occur within the complexities of the colonic structure. As a result, there is a high rate of missed detections and incomplete removal of colonic polyps, which can adversely impact patient outcomes. Recently, automated methods that use machine learning have been developed to enhance polyps detection and segmentation, thus helping clinical processes and reducing missed rates. These advancements highlight the potential for improving diagnostic accuracy in real-time applications, which ultimately facilitates more effective patient management. Furthermore, integrating sequence data and temporal information could significantly enhance the precision of these methods by capturing the dynamic nature of polyp growth and the changes that occur over time. To rigorously investigate these challenges, data scientists and experts gastroenterologists collaborated to compile a comprehensive dataset that spans multiple centers and diverse populations. This initiative aims to underscore the critical importance of incorporating sequence data and temporal information in the development of robust automated detection and segmentation methods. This study evaluates the applicability of deep learning techniques developed in real-time clinical colonoscopy tasks using sequence data, highlighting the critical role of temporal relationships between frames in improving diagnostic precision.
The rapid progress of diffusion models highlights the growing need for detecting generated images. Previous research demonstrates that incorporating diffusion-based measurements, such as reconstruction error, can enhance the generalizability of detectors. However, ignoring the differing impacts of aleatoric and epistemic uncertainty on reconstruction error can undermine detection performance. Aleatoric uncertainty, arising from inherent data noise, creates ambiguity that impedes accurate detection of generated images. As it reflects random variations within the data (e.g., noise in natural textures), it does not help distinguish generated images. In contrast, epistemic uncertainty, which represents the model's lack of knowledge about unfamiliar patterns, supports detection. In this paper, we propose a novel framework, Diffusion Epistemic Uncertainty with Asymmetric Learning~(DEUA), for detecting diffusion-generated images. We introduce Diffusion Epistemic Uncertainty~(DEU) estimation via the Laplace approximation to assess the proximity of data to the manifold of diffusion-generated samples. Additionally, an asymmetric loss function is introduced to train a balanced classifier with larger margins, further enhancing generalizability. Extensive experiments on large-scale benchmarks validate the state-of-the-art performance of our method.
The Role of Entropy in Visual Grounding: Analysis and Optimization
Shuo Li, Jiajun Sun, Zhihao Zhang, Xiaoran Fan, Senjie Jin, Hui Li, Yuming Yang, Junjie Ye, Lixing Shen, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang
机构
*
Fudan University(复旦大学)
;
Hikvision Research Institute(海康威视研究院)
详情
英文摘要
Recent advances in fine-tuning multimodal large language models (MLLMs) using reinforcement learning have achieved remarkable progress, particularly with the introduction of various entropy control techniques. However, the role and characteristics of entropy in perception-oriented tasks like visual grounding, as well as effective strategies for controlling it, remain largely unexplored. To address this issue, we focus on the visual grounding task and analyze the role and characteristics of entropy in comparison to reasoning tasks. Building on these findings, we introduce ECVGPO (Entropy Control Visual Grounding Policy Optimization), an interpretable algorithm designed for effective entropy regulation. Through entropy control, the trade-off between exploration and exploitation is better balanced. Experiments show that ECVGPO achieves broad improvements across various benchmarks and models.
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
Jiayi Zhu, Yihao Huang, Yue Cao, Xiaojun Jia, Qing Guo, Felix Juefei-Xu, Geguang Pu, Bin Wang
机构
*
Hangzhou Institute of Technology, Xidian University, China(杭州科技学院,西安电子科技大学,中国)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
;
Nankai University(南开大学)
;
New York University(纽约大学)
;
East China Normal University(华东师范大学)
;
Hikvision Digital Technology Co., Ltd(海康威视数字技术有限公司)
;
Shanghai Industrial Control Safety Innovation Tech. Co., Ltd(上海工业控制安全创新技术有限公司)
详情
英文摘要
Large Visual Language Models (LVLMs) now pose a serious yet overlooked privacy threat, as they can infer a social media user's geolocation directly from shared images, leading to unintended privacy leakage. While adversarial image perturbations provide a potential direction for geo-privacy protection, they require relatively strong distortions to be effective against LVLMs, which noticeably degrade visual quality and diminish an image's value for sharing. To overcome this limitation, we identify typographical attacks as a promising direction for protecting geo-privacy by adding text extension outside the visual content. We further investigate which textual semantics are effective in disrupting geolocation inference and design a two-stage, semantics-aware typographical attack that generates deceptive text to protect user privacy. Extensive experiments across three datasets demonstrate that our approach significantly reduces geolocation prediction accuracy of five state-of-the-art commercial LVLMs, establishing a practical and visually-preserving protection strategy against emerging geo-privacy threats.
1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
Zeliang Zong, Kai Zhang, Zheyang Li, Wenming Tan, Ye Ren, Yiyan Zhai, Jilin Hu
机构
*
Hikvision Research Institute(海康威视研究院)
;
School of Data Science and Engineering, East China Normal University(华东师范大学数据科学与工程学院)
Comments15 pages, 6 figures, EMNLP 2025 findings
详情
英文摘要
Large Language Models (LLMs) have demonstrated remarkable proficiency in language comprehension and generation; however, their widespread adoption is constrained by substantial bandwidth and computational demands. While pruning and low-rank approximation have each demonstrated promising performance individually, their synergy for LLMs remains underexplored. We introduce \underline{S}ynergistic \underline{S}parse and \underline{L}ow-Rank \underline{C}ompression (SSLC) methods for LLMs, which leverages the strengths of both techniques: low-rank approximation compresses the model by retaining its essential structure with minimal information loss, whereas sparse optimization eliminates non-essential weights, preserving those crucial for generalization. Based on theoretical analysis, we first formulate the low-rank approximation and sparse optimization as a unified problem and solve it by iterative optimization algorithm. Experiments on LLaMA and Qwen2.5 models (7B-70B) show that SSLC, without any additional training steps, consistently surpasses standalone methods, achieving state-of-the-arts results. Notably, SSLC compresses Qwen2.5 by 50\% with no performance drop and achieves at least 1.63$\times$ speedup, offering a practical solution for efficient LLM deployment.
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Tao Ji, Bin Guo, Yuanbin Wu, Qipeng Guo, Lixing Shen, Zhan Chen, Xipeng Qiu, Qi Zhang, Tao Gui
机构
*
School of Computer Science, Fudan University(复旦大学计算机科学学院)
;
School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
;
Institute of Modern Languages and Linguistics, Fudan University(复旦大学现代语言与语言学研究所)
;
Institute of Trustworthy Embodied Artificial Intelligence, Fudan University(复旦大学可信具身人工智能研究所)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Pengcheng Laboratory(鹏城实验室)
;
Hikvision Inc(海康威视公司)
Comments16 pages, 8 figures; Accepted to ACL 2025
详情
英文摘要
Multi-head Latent Attention (MLA) is an innovative architecture proposed by DeepSeek, designed to ensure efficient and economical inference by significantly compressing the Key-Value (KV) cache into a latent vector. Compared to MLA, standard LLMs employing Multi-Head Attention (MHA) and its variants such as Grouped-Query Attention (GQA) exhibit significant cost disadvantages. Enabling well-trained LLMs (e.g., Llama) to rapidly adapt to MLA without pre-training from scratch is both meaningful and challenging. This paper proposes the first data-efficient fine-tuning method for transitioning from MHA to MLA (MHA2MLA), which includes two key components: for partial-RoPE, we remove RoPE from dimensions of queries and keys that contribute less to the attention scores, for low-rank approximation, we introduce joint SVD approximations based on the pre-trained parameters of keys and values. These carefully designed strategies enable MHA2MLA to recover performance using only a small fraction (0.3% to 0.6%) of the data, significantly reducing inference costs while seamlessly integrating with compression techniques such as KV cache quantization. For example, the KV cache size of Llama2-7B is reduced by 92.19%, with only a 0.5% drop in LongBench performance.