发表机构
University of North Carolina at Charlotte; Carnegie Mellon University(北卡罗来纳大学夏洛特分校; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对未知聚合下两分辨率监督学习,提出估计与跟踪策略,推导盈亏平衡条件并证明最优性,实现成本与信息的有效平衡。
AI 中文摘要
现代学习系统通常以多种分辨率获取监督信息,在标注成本与信息含量之间进行权衡。我们研究成本感知的两分辨率学习问题,其中昂贵的细粒度标签揭示向量响应,而较便宜的粗粒度标签揭示由未知权重形成的标量聚合,目标仍然是完整的响应。挑战在于,未知聚合会改变粗粒度数据能够识别的方向,因此粗粒度监督的价值取决于成本、噪声和可识别性三者的共同作用。我们刻画了这一信息几何结构,并开发了一种估计与跟踪策略,该策略学习聚合规则并跟踪最优分辨率组合。我们推导出粗粒度监督的闭式盈亏平衡条件,并证明在线策略能达到最优的前导累积风险系数,同时匹配相应的局部渐近极小极大下界。合成实验支持了所预测的全细粒度/混合转换,显示在线学习器接近预言份额基准,并证明当粗粒度监督足够有利时,在有限预算下优于全细粒度采集。我们的结果为在监督分辨率之间平衡信息与标注成本提供了一种原则性方法。
英文摘要
Modern learning systems often acquire supervision at multiple resolutions, trading annotation cost against information content. We study cost-aware two-resolution learning, where expensive fine labels reveal a vector response and cheaper coarse labels reveal a scalar aggregate formed with unknown weights, while the target remains the full response. The challenge is that unknown aggregation changes which directions coarse data can identify, so the value of coarse supervision depends jointly on cost, noise, and identification. We characterize this information geometry and develop an estimate-and-track policy that learns the aggregation rule and tracks the optimal resolution mix. We derive a closed-form break-even condition for coarse supervision and prove that the online policy attains the optimal leading cumulative-risk coefficient, with a matching local asymptotic minimax lower bound. Synthetic experiments support the predicted all-fine/mixed transition, show the online learner approaching the oracle-share benchmark, and demonstrate a finite-budget gain over all-fine acquisition when coarse supervision is sufficiently favorable. Our results provide a principled way to balance information and annotation cost across supervision resolutions.
Comments34 pages, 6 Figures