arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

计算阈值下的自适应与异方差线性回归算法

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold

Spencer Compton, Tselil Schramm

arXiv 2608.18402首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对异方差、自适应线性回归模型,设计多项式时间估计器并给出计算下界,还提出植入式线性回归问题以探究信息-计算间隙。

AI 中文摘要

我们研究存在多样且未知标签噪声时的有限样本线性回归,聚焦于异方差与自适应线性回归模型。异方差线性回归对应标签质量各异的场景:我们接收n个样本对$(X_i,Y_i)$,其中标签满足$Y_i=X_i^\top\beta+\boldsymbol{\beta}$,其中$\boldsymbol{\beta}\tilde{N}(0,\boldsymbol{\beta}^2)$,且方差对估计器未知。该问题难度的一个自然衡量是满足$\boldsymbol{\beta}^2\tilde{1}$的样本数m(m越大问题越易)。当$m\tilde{d}^{3/4}n^{1/4}$时,我们得到多项式时间估计器,其误差率为$\tilde{O}((nd^3/m^4)^{1/6})$,同时给出几乎匹配的下界。对于$d=O(1)$,我们的估计器在$m\tilde{n}^{1/4}$时误差达到$o(1)$,而$L_1$回归及其他传统方法需要$m\tilde{n}^{1/2}$。在自适应线性回归中,误差从未知分布p中独立同分布抽取,目标是设计通用估计器,其性能接近知晓p的最优定制估计器。我们引入了一个(计算低效的)自适应估计器:只要p是k个对称对数凹密度的混合,该估计器的误差就与知晓p的最优估计器相当,且需要$\tilde{\beta}(n/k)$个样本。当k=1时,我们证明$L_q$回归(q依赖于数据)可给出多项式时间估计器。最后,为研究两个问题的计算极限,我们引入植入式线性回归问题:其中$X_i\tilde{N}(0,I_d)$,m个未知样本无噪声,其余样本的误差$\boldsymbol{\beta}\tilde{N}(0,1)$。我们推测,当m在$m=d+1$与$m\tilde{d}^{3/4}n^{1/4}$之间时,恢复$\beta$至误差$\tilde{\beta}\tilde{d/n}$(或完全恢复)可能存在信息-计算间隙,这由我们的几乎匹配的多项式时间估计器与统计查询(SQ)下界所暗示。

英文摘要

We study finite-sample linear regression in the presence of varied and unknown label noise, focusing on the heteroskedastic and adaptive linear regression models. Heteroskedastic linear regression models settings where the labels are of varying quality. We receive $n$ pairs $(X_i,Y_i)$ with labels $Y_i=X_i^\topβ+\varepsilon_i$, where $\varepsilon_i\sim N(0,σ_i^2)$ and the variances are unknown to the estimator. One natural measurement of the difficulty of this problem is the number of samples $m$ for which $σ_i^2\le1$ (larger $m$ is easier). We obtain a polynomial-time estimator with rate $\tilde{O}((nd^3/m^4)^{1/6})$ when $m\gg d^{3/4}n^{1/4}$, as well as nearly-matching lower bounds. For $d=O(1)$, our estimator achieves error $o(1)$ when $m\gg n^{1/4}$, whereas $L_1$ regression and other traditional approaches require $m\gg n^{1/2}$. In adaptive linear regression, the errors are drawn i.i.d. from an unknown distribution $p$, and our goal is to design a generic estimator that performs nearly as well as the best custom estimator that knows $p$. We introduce a (computationally inefficient) adaptive estimator that, so long as $p$ is a mixture of $k$ symmetric log-concave densities, achieves error comparable with the optimal estimator that knows $p$ and has $\tildeΘ(n/k)$ samples. For $k=1$, we show that $L_q$ regression (with data-dependent $q$) gives a polynomial-time estimator. Finally, to study the computational limits of both problems, we introduce the planted linear regression problem, where $X_i\sim N(0,I_d)$, $m$ unknown samples are noiseless, and the rest have error $\varepsilon_i\sim N(0,1)$. We conjecture that recovering $β$ up to error $\ll\sqrt{d/n}$ (or exactly) may have an information-computation gap between $m=d+1$ and $m\sim d^{3/4}n^{1/4}$, as is suggested by our near-matching polynomial-time estimator and statistical query (SQ) lower bound.

Commentsshortened arxiv abstract

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑