发表机构
Nanyang Normal University; Tianjin Normal University; Linyi University(南阳师范学院; 天津师范大学; 临沂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ACV-Gate框架,通过近似全预算反事实评估并选择性细化候选,在生成式图像通信中显著降低编码端计算量,同时提升重建质量。
AI 中文摘要
生成式图像通信在有限的数据包预算下传输紧凑的语义令牌,其中令牌选择直接影响完整数据包解码后的最终重建质量。然而,准确估计每个候选令牌的终端价值需要重复的接收端重建,导致编码端计算量巨大。为解决此问题,我们提出了ACV-Gate,一种自适应候选评估框架,学习近似全预算反事实评估,并选择性地将精确评估分配给最具信息量的候选。具体而言,使用终端优势和遗憾训练一个集合感知学生模型,直接预测候选排名,而选择性细化机制仅评估包含Local-MDL和直接动作的有界候选集;基于成本的阈值进一步实现对平均评估工作量的显式控制。在CIFAR-10上的实验表明,ACV-Gate在显著减少候选评估的同时持续提高重建质量;在0.20 bpp下,主要自适应配置相比LocalMDL将PSNR提高了0.636 dB,每张图像仅需2.13次候选评估,相当于Exact-Full专家所需调用次数的27.60%。匹配候选比较、同步GPU测量以及在STL-10和384*384尺度迁移上的评估进一步证明了一致的质量计算权衡,尤其在低比特率下收益显著。这些结果表明,将终端价值学习与选择性候选评估相结合,为在数据包受限的生成式图像通信中分配编码器计算提供了一种有效且可控的机制。
英文摘要
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates. Specifically, a set-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local-MDL and direct actions; cost-based thresholds further enable explicit control of the average evaluation workload. Experiments on CIFAR-10 show that ACV-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0.20 bpp, the primary adaptive configuration improves PSNR over LocalMDL by 0.636 dB with only 2.13 candidate evaluations per image, corresponding to 27.60% of the calls required by the Exact-Full expert. Matched-candidate comparisons, synchronized GPU measurements, and evaluations on STL-10 and 384 *384 scale transfer further demonstrate consistent quality computation trade-offs, with particularly pronounced gains at low bit rates. These results show that combining terminal-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet-constrained generative image communication.
CommentsVisual token communication, counterfactual evaluation, selective computation, knowledge distillation, resource allocation