发表机构
Kabale University(卡巴莱大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究下游任务未知时系统压缩所需保留的状态量,提出精确前沿公式,证明最优建议分区为强NP困难,并通过三个实例展示任务信息可大幅减少所需状态。
AI 中文摘要
一个系统可能在其下游任务完全明确之前就被压缩。我们询问,在有限的事前任务信息下,需要保留多少状态,又能节省多少状态。对于有限族线性任务,在状态形成之前揭示一个任务消息,而确切任务仅在之后揭示。对于大小为 $K$ 的建议字母表,精确前沿为 \\[ p^*(K)= \min_{\substack{\Pcal\text{ partition of }\U\\\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C), \\] 通过设置 $K=\min(2^b,|\U|)$ 得到 $b$ 位前沿。因此,事前任务信息通过其联合任务算子具有低秩的分区来减少状态。我们还给出了一个近似的奇异值前沿、一个公共核心下界和精确直和律,以及寻找最优建议分区的强 NP 困难性。该困难性在每个固定的正近似容差下持续存在。三个例子说明了该结果。一个良态的 softmax 注意力构造在九位解析 $512$ 个延续之一时,给出了精确的 $524{,}288\to1{,}024$ 坐标前沿。一个域分解的数字孪生产生了一个接口加局部状态律和一个针对异构区域的加权分区问题。一个层次化多任务模型给出了一个两阶段前沿,其中三位将所需状态从 $3136$ 个坐标减少到 $448$ 个坐标,进一步的任务信息接近不可约的 $328$ 坐标单任务下限。
英文摘要
A state may need to be compressed before its exact downstream task is known. For a finite family of linear task operators $\{T_u\}_{u\in\U}$, we characterize exactly how much state must survive when only one of $K$ coarse task messages is available at compression time: $$p^*(K)= \min_{\substack{\Pcal\text{ partition of }\U\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C),$$ where $T_C$ stacks the tasks that remain unresolved within cell $C$. Thus task information reduces retained state precisely by separating tasks whose joint task-visible subspace is expensive. We derive the dual bit frontier, an irreducible common-core floor, and a sharp singular-value characterization for nonzero tolerance. For overlapping task families, a cumulative overlap deficit places the exact frontier within an additive $Δ$ of weighted load balancing. Optimal partitioning is strongly NP-hard, yet direct-sum families admit classical approximation guarantees and a PTAS. We stress-test the theory on learned task operators from CIFAR-10, Fashion-MNIST, SVHN, and Burgers observables. For Burgers with $K=2$ and $\eps=0.35$, a constructive partition attains $p=4$, while balanced random partitions require median $p=6$, with held-out worst standardized RMSE $0.0368$. The results distinguish uniform task sufficiency from distributional and decision-level performance.
Comments12 pages, 2 figures, 1 table. Substantially expanded version: added sharp approximate bounds, overlap-deficit theory, approximation guarantees, and empirical stress tests on learned representations and Burgers observables