发表机构
Xiamen University(厦门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FastJEV针对JEV推理中的冗余问题,通过共享上下文锚定、候选前缀共享和决策引导剪枝,在不额外训练的情况下降低候选深度,同时保留高任务分数,提升JEV推理效率。
AI 中文摘要
JEV模型通过直接对候选进行评分来做出多模态决策。尽管公共上下文仅编码一次,但候选评估仍可能重复匹配标记历史、复制推理状态并执行完整的主干网络。在本文中,我们研究这些冗余来源并提出用于紧凑候选评估的FastJEV。我们联合组织历史复用和状态存储,因为共享计算需要为后续分支保留状态。我们首先引入共享上下文锚定以复用循环初始状态并省略未使用的最终循环缓存。我们通过候选前缀共享扩展这种复用,保留后续分支所需的中间状态。为进一步减少这些路径的深度,我们应用基于在小型未标记集上测量的相对分数变化的决策引导剪枝。我们的方法保留完整的上下文编码和所有候选,无需额外训练。我们在三种OmniJev模型规模上评估FastJEV,涉及五个公共基准和重构的LIBERO-10离线问题。在选定的剪枝预算下,完整方法将候选深度降低43.75%至45.83%,同时在六个评估集上平均保留原始任务分数的93.66%至97.52%。通过对照实验,我们展示了候选重叠和分支结构如何影响历史复用的执行成本。在我们的实现中,候选前缀共享可减少重复计算,同时增加延迟。这些发现促使人们共同设计共享粒度和执行调度以实现高效的JEV推理。
英文摘要
JEV models make multimodal decisions by directly scoring candidates. Although the common context is encoded once, candidate evaluation can still repeat matching token histories, duplicate inference states, and execute the full backbone. In this paper, we study these sources of redundancy and present FastJEV for compact candidate evaluation. We jointly organize history reuse and state storage, since sharing computation requires preserving states for later branches. We first introduce shared context anchoring to reuse recurrent initial states and omit unused final recurrent caches. We extend this reuse through candidate prefix sharing, retaining the intermediate states needed by subsequent branches. To further reduce the depth of these paths, we apply decision guided pruning based on relative score changes measured on a small unlabeled set. Our method retains full context encoding and all candidates without additional training. We evaluate FastJEV across three OmniJev model sizes on five public benchmarks and reconstructed LIBERO-10 offline questions. At the selected pruning budgets, the complete method reduces candidate depth by 43.75% to 45.83%, while retaining 93.66% to 97.52% of the original task scores on average across the six evaluation sets. Through controlled experiments, we show how candidate overlap and branching structure affect the execution cost of history reuse. In our implementation, candidate prefix sharing can reduce repeated computation while increasing latency. These findings motivate designing sharing granularity and execution schedules together for efficient JEV inference.