arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniTAC:基于加权失真度量的通用任务感知压缩

UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures

Homa Esfahanizadeh, Matin Mortaheb, Adeel Mahmood, Jinfeng Du, Harish Viswanathan

arXiv 2608.16696首次发表:更新:

发表机构

Nokia Bell Labs(诺基亚贝尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UniTAC是一种无需重新训练的通用任务感知图像编解码器,通过任务重要性向量调节,在0.034 bpp时准确率达91.4%,优于通用编解码器,接近任务专用编解码器。

AI 中文摘要

自动驾驶汽车、机器人等物理AI系统需在带宽、延迟和能量预算紧张的情况下及时交换高维传感信号。由于驱动下游决策的任务会随时间变化,特定任务的编解码器脆弱,且在实际场景中为每个任务重新训练不可行。我们提出UniTAC,一种单一的学习型图像编解码器,可在通用(任务无关)到任务专用的操作范围内运行,无需重新训练即可在运行时重新适配。该任务被抽象为每个组件的重要性向量,例如从任何下游模型的梯度归因中推导,并作为低开销辅助信息传输,以调节编码器和解码器。UniTAC在广泛的此类随机向量族上针对加权重构失真进行一次训练,保持固定的骨干网络和单一的人类可见重构,通过交换注入的向量来控制其对活动任务的保真度。我们分析了潜在的加权率失真问题,表征了对角加权失真何时与任务一致,以及权重与任务敏感性的关系。在此指导下,我们设计了一种Vision Transformer(ViT)编解码器,其令牌级调节天然实现了这种权重驱动的代码。在0.034 bpp的本地化任务中,单个UniTAC模型达到91.4%的准确率,仅比基于任务的编解码器(93.3%)低1.9%,高于通用编解码器(76.9%)。

英文摘要

Lossy compression is conventionally driven by a task-agnostic distortion (e.g., MSE or MS-SSIM), yet in many emerging applications the receiver cares not about uniform fidelity but about a downstream task whose relevant content varies across the signal and evolves over time. We formulate task-aware compression as a weighted rate-distortion problem, in which a single codec is driven by a separable, per-component weighted distortion whose weights encode task importance and may depend on the source. We introduce task consistency, i.e., that minimizing the weighted distortion also minimizes the true task loss, and characterize when it holds: for linear tasks, the task loss admits a weighted-MSE form with signal-independent weights under suitable cross-term conditions, while for nonlinear tasks, an integrated-gradients analysis motivates separable task-aware weights. We show how task symmetry and irrelevance further constrain the admissible weights. Guided by this theory, we realize the weight-conditioned code in a single learned Vision Transformer (ViT) codec whose token-level attention natively consumes a per-component importance vector, so one fixed backbone is re-targeted at runtime, from universal (task-agnostic) to task-specialized operation, purely by swapping the injected weights, without retraining, while producing a single human-viewable reconstruction steered to the active task. On downstream face-analysis tasks, a single model reaches 91.4% accuracy at 0.034 bpp on a localized task, within 1.9% of a task-specific codec (93.3%) and well above a universal codec (76.9%). Such task-adaptive compression suits bandwidth-constrained perception systems, e.g., in Physical AI, where the active task drifts and per-task retraining is infeasible.

Comments13 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑