arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

二值量化神经网络训练以输入和输出维度为参数是 W[1]-困难的

Binary Quantized Neural Network Training Is W[1]-Hard Parameterized by Input and Output Dimensions

Tao Jiang, Minbo Gao, Shaowei Cai

arXiv 2609.27932首次发表:更新:

发表机构

Key Laboratory of System Software (Chinese Academy of Sciences); State Key Laboratory of Computer Science; Institute of Software, Chinese Academy of Sciences; School of Computer Science and Technology, University of Chinese Academy of Sciences(系统软件重点实验室(中国科学院); 计算机科学国家重点实验室; 中国科学院软件研究所; 中国科学院大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明二值量化神经网络训练(2-QNNT)以输入和输出维度之和为参数时是 W[1]-困难的,即使在零误差和特定前缀链数据上也成立,并基于指数时间假说排除了 FPT 算法,核心归约利用单翻转路由等价性从边不相交路径问题构造。\n

AI 中文摘要

Ganian 等人(ICLR 2026)证明了当以架构树宽、输入维度 $\alpha$ 和输出维度 $\omega$ 为联合参数时,量化神经网络训练是固定参数易处理的,并留下了一个开放问题:仅以 $\alpha+\omega$ 为参数是否也能得到固定参数易处理性。我们证明了 2-QNNT 以 $\alpha+\omega$ 为参数是 W[1]-困难的。该困难性在 $D_k=\{(\xi^{(r)},\xi^{(r)}):0\le r\le k\}$ 上以零误差成立,其中每个输入等于其目标,$|D_k|=\alpha=\omega=k+1$,且这些样本构成一个坐标方向的前缀链。该困难性在将所有非源偏置固定为零时仍然成立。在指数时间假说下,不存在任何可计算函数 $f$ 使得算法在 $f(\alpha+\omega)|I|^{o(\alpha+\omega)}$ 时间内运行。该归约从有向无环图中的边不相交路径问题出发,通过有向线图将边容量转换为顶点容量,并将结果规范化为有效的分层架构。关键的结构步骤是一个单翻转路由等价性:在前缀链输入上,非负二值权重使每个激活函数单调,且每个所需的输出转换都有一个权重为 1 的前驱节点做出相同的转换。反向迭代这一关系可从唯一变化的输入中提取出一条路径,而不同的转换产生顶点不相交的路径。特别地,这些输入上的每个神经元只有 $k+1$ 种可能的激活轮廓。

英文摘要

Ganian et al. (ICLR 2026) proved that quantized neural network training is fixed-parameter tractable when parameterized jointly by architecture treewidth, input dimension $α$, and output dimension $ω$, and left open whether $α+ω$ alone yields fixed-parameter tractability. We prove that 2-QNNT is W[1]-hard parameterized by $α+ω$. The hardness already holds with zero error on $D_k=\{(ξ^{(r)},ξ^{(r)}):0\le r\le k\}$, where every input equals its target, $|D_k|=α=ω=k+1$, and the examples form a coordinatewise prefix chain. It also holds when every non-source bias is fixed to zero. Under the Exponential Time Hypothesis, no algorithm runs in $f(α+ω)|I|^{o(α+ω)}$ for any computable $f$. The reduction starts from DAG edge-disjoint paths, converts edge capacity to vertex capacity with a directed line graph, and normalizes the result into a valid layered architecture. The key structural step is a one-flip routing equivalence: on the prefix-chain inputs, nonnegative binary weights make every activation monotone, and each required output transition has a weight-one predecessor making the same transition. Iterating this relation backward extracts a path from the unique changing input, while different transitions yield vertex-disjoint paths. In particular, every neuron on these inputs has only $k+1$ possible activation profiles.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑