AI 中文总结
针对可实现SVM,通过确定性删除问题的熵归纳与KKT恒等式,证明了基于锐利间隔的高概率泛化界,其复杂度由经验间隔和训练半径决定。
AI 中文摘要
设精确的齐次硬间隔支持向量机在实希尔伯特空间上的一个Borel概率律的\\(m\\)个独立观测上训练。我们证明,将得分为零计为错误时,存在一个通用数值常数\\(C\\),使得\\[ \Pp\left( \gamma_m>0,\quad \Risk(u_m)> \frac{C}{m} \left( K_m+\log\frac1\delta \right) \right) \le \delta. \\] 这里\\(\gamma_m\\)是经验齐次间隔,\\(u_m\\)是精确的最小范数单位间隔分离器,\\(r_m\\)是最大的训练半径,且在\\(\{\gamma_m>0\}\\)上\\(K_m:=r_m^2\norm{u_m}^2=r_m^2/\gamma_m^2\\)。证明由一个确定性删除问题驱动。给定单位球中的向量\\(x_1,\ldots,x_n\\),删除一组约束\\(B\\),并令\\(u_B\\)为满足所有保留的单位间隔约束的离原点最近的点。假设\\(\norm{u_B}^2\le k\\)且每个被删除的向量在\\(u_B\\)下的得分非正。我们证明,这样的基数\\(q\\)的删除集族的大小至多为\\(\exp(8k+2q)\\)。概念性步骤是从\\(u_B\\)的KKT表示获得的精确恒等式。对于随机删除集,该恒等式将分离器的均方散布转化为得分缺口的加权和。因此,它强制一个坐标,其删除状态将两个条件均值分开一个数量级大的量。揭示该坐标会减少条件分离器方差,足以控制分裂的二元熵。熵归纳给出删除计数,而精确的阶乘幽灵样本恒等式将该计数转化为所述的高概率SVM界。
英文摘要
Let the exact homogeneous hard-margin support vector machine be trained on \(m\) independent observations from a Borel probability law on a real Hilbert space. We prove that, with score zero counted as an error, there is a universal numerical constant \(C\) such that \[ \Pp\left( γ_m>0,\quad \Risk(u_m)> \frac{C}{m} \left( K_m+\log\frac1δ \right) \right) \le δ. \] Here \(γ_m\) is the empirical homogeneous margin, \(u_m\) is the exact minimum-norm unit-margin separator, \(r_m\) is the largest training radius, and \(K_m:=r_m^2\norm{u_m}^2=r_m^2/γ_m^2\) on \(\{γ_m>0\}\). The proof is driven by a deterministic deletion problem. Given vectors \(x_1,\ldots,x_n\) in the unit ball, delete a set \(B\) of constraints and let \(u_B\) be the closest point to the origin that satisfies every retained unit-margin constraint. Suppose that \(\norm{u_B}^2\le k\) and that every deleted vector has nonpositive score under \(u_B\). We prove that a family of such deletion sets of cardinality \(q\) has size at most \(\exp(8k+2q)\). The conceptual step is an exact identity obtained from the KKT representation of \(u_B\). For a random deletion set, the identity converts the mean squared spread of the separators into a weighted sum of score deficits. It therefore forces a coordinate whose deletion status separates the two conditional means by a quantitatively large amount. Revealing that coordinate decreases the conditional separator variance enough to control the binary entropy of the split. An entropy induction gives the deletion count, and an exact factorial ghost-sample identity converts that count into the stated high-probability SVM bound.