通过生成式AI层对端到端视频流管道的感知细化
Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers
浏览论文内容
中文总结 AI 辅助
PRESLEY通过生成式AI层在视频流中自适应退化并重建不关注区域,实现显著比特率节省(最高-29.4% BD-rate),同时保持前景保真度,并证明恢复后损伤可预测,为联合优化提供路线图。
中文摘要 AI 辅助
传统编解码器对帧的每个区域一视同仁;而生成式层可以反而降低观看者最不关注的区域的质量,并在客户端对其进行重建。我们提出了PRESLEY,它扩展了之前的会议工作ELVIS,通过用可移除掩码下的自适应原位退化取代破坏性块移除,在比特打包的边信道中传递每块强度信号,并基于传输的视觉先验条件而非无条件修复的生成骨干网络进行恢复。我们将问题分解为三个目标:选择要退化的块、退化它们以使编码器花费更少的比特,以及恢复它们。在与前身相同码率下,PRESLEY在跨越多个编解码器和数据集家族的13个码率阶梯上,实现了交付背景质量的平均BD-rate降低-56.4%的决定性改进。与原始基线相比,PRESLEY定义了生成式传输的工作机制:在目标比特匮乏机制中,实现显著的比特率节省(高达-29.4% BD-rate)和优越的背景质量(17/23个序列),同时保持前景保真度比特精确。我们进一步绘制了此类架构中理论余量的位置。使用精确的留一超级块排除组合预言机作为加性经验界限,我们表明现有的复杂度启发式方法已经捕获了83.3%的比特成本节省,将剩余的成本轴余量限制在总比特率的约5%。然后,我们识别并建模了主要未解决的轴——恢复后损伤——其分散范围广泛(4.9-8.4 dB)。我们证明这种损伤在传输前是可预测的(留出集rho = +0.400),确立了传输时可恢复性建模的可行性,并为联合率-失真-恢复选择规则定义了路线图。
英文摘要
Traditional codecs treat every region of a frame alike; a generative layer can instead degrade the regions a viewer attends to least and reconstruct them at the client. We present PRESLEY, which extends the prior conference work ELVIS by replacing destructive block removal with adaptive in-place degradation under a removability mask, signaling per-block strength in a bit-packed side channel, and restoring via generative backbones conditioned on transmitted visual priors rather than unconditioned in-painting. We separate the problem into three goals: choosing which blocks to degrade, degrading them so the encoder spends fewer bits, and restoring them. Against its predecessor at matched rate, PRESLEY achieves a decisive mean -56.4% BD-rate reduction on delivered background quality across 13 rate ladders spanning multiple codecs and dataset families. Against pristine baselines, PRESLEY defines the operating regime of generative transport: delivering substantial bitrate savings (up to -29.4% BD-rate) and superior background quality (17/23 sequences) in the target bit-starved regime, while maintaining foreground fidelity bit-exact. We further map where the theoretical headroom in this class of architecture lies. Using an exact leave-one-superblock-out combinatorial oracle as an additive empirical bound, we show that existing complexity heuristics already capture 83.3% of bit-cost savings, bounding remaining cost-axis headroom at about 5% of total bitrate. We then identify and model the primary unaddressed axis -- post-restoration damage -- which disperses widely (4.9-8.4 dB). We prove that this damage is predictable before transmission (held-out rho = +0.400), establishing the feasibility of transmit-time restorability modeling and defining the roadmap for joint rate-distortion-restoration selection rules.