arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARVE:用于高效3D医学体积理解的视觉证据跨切片各向异性重分配

CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding

Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang

arXiv 2608.04515首次发表:更新:

AI 中文总结

针对3D医学体积理解中切片式MLLM的视觉令牌冗余问题,提出无需训练的CARVE框架,通过跨切片各向异性重分配压缩80%令牌,在AMOS-MM等基准上性能优于现有方法。

AI 中文摘要

基于切片的多模态大型语言模型(MLLM)通过将3D体积表示为2D切片序列,利用成熟的2D编码器。然而,这种按切片划分的形式会产生数千个视觉令牌,给大语言模型(LLM)主干带来负担,其中许多令牌捕获了相邻切片间重叠的视觉证据。为探究不断增加的视觉令牌预算对性能的提升效果,我们在两个3D医学视觉问答(VQA)基准上进行了缩放分析,发现收益递减:成本持续上升而准确率趋于饱和,且在可比预算下,提高平面内分辨率比增加切片更有效。因此,预算应更有选择性地分配,而非单纯扩大;但大多数令牌压缩方法是为2D图像或视频设计的,其冗余源于空间布局或时间运动,而非深度轴上的近重复内容。我们提出CARVE,这是一种无需训练的框架,在LLM推理前压缩视觉令牌,并将令牌减少问题转化为受预算约束的2.5D分配。CARVE将深度轴划分为连贯窗口,根据归一化跨切片证据非均匀分配令牌。在共享预算下,CARVE在代表性切片上构建空间锚点,从完整体积中检索局部变化的证据,再将每个窗口内剩余符合条件的令牌合并到附近锚点。在Hulu-Med-7B上移除约80%的视觉令牌后,CARVE在所有AMOS-MM报告生成指标上优于所有压缩基线,比最强基线的全令牌质量保留率高6.2个百分点,并在三个VQA基准上保留了全令牌性能的98.1%。

英文摘要

Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual evidence across adjacent slices. To understand how effectively a growing visual token budget improves performance, we perform scaling analyses on two 3D medical VQA benchmarks and find diminishing returns: cost keeps rising while accuracy saturates, and improving in-plane resolution is more effective than adding slices at comparable budgets. The budget should therefore be allocated more selectively rather than simply enlarged, yet most token compression methods are designed for 2D images or videos, where redundancy arises from spatial layout or temporal motion rather than from near-duplicate content along the depth axis. We present CARVE, a training-free framework that compresses visual tokens prior to LLM inference and casts token reduction as budget-constrained 2.5D allocation. CARVE partitions the depth axis into coherent windows and allocates tokens non-uniformly according to normalized cross-slice evidence. Under a shared budget, CARVE builds spatial anchors on representative slices and retrieves locally varying evidence from the full volume, then merges remaining eligible tokens into nearby anchors within each window. Removing roughly 80% of the visual tokens on Hulu-Med-7B, CARVE leads all compression baselines on every AMOS-MM report-generation metric, with 6.2 points higher retention of full-token quality than the strongest baseline, and preserves 98.1% of full-token performance across three VQA benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑