arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于查询-值模型的多实例学习的基于期望最大化的迭代

EM-based iterations for multiple instance learning on a query-value model

Ethan Levien

arXiv 2607.17404首次发表:更新:

AI 中文总结

研究多实例回归中基于查询-值模型的学习问题,推导无噪声极限下的参数化迭代族,推广EM-DD算法,得出查询和值向量最大似然估计器的集中结果,表明值向量随机初始化平均指向正确方向,多项式数量包可使EM算法高效收敛。

AI 中文摘要

在多实例回归中,数据被组织成包(特征空间中的实例集合),目标是学习一个为包分配标签的映射。典型假设是特征空间中存在一个所谓的概念点,与它的接近程度决定包的标签。受基于注意力的现代多实例回归架构的启发,我们研究了一个将概念点和标签方案解耦的softmax模型。两者分别由特征空间中的查询方向和值方向确定。这个问题分离了从包级监督中学习查询向量和值向量的基本挑战。我们从这个模型在无噪声极限下推导出一个参数化的迭代族,它推广了一种称为EM-DD算法的方法。然后,我们得出从随机选择的实例中获得的查询向量和值向量的最大似然估计器的集中结果。我们对值向量的结果表明,值向量的单个随机初始化平均已经指向正确方向,因此多项式数量的包(与每个包中的实例数量和特征维度有关)足以使EM算法以高概率在$O(1)$步内收敛。该分析的一个关键方面是经验协方差矩阵的集中与选择规则产生的极值统计之间的相互作用。

英文摘要

In multiple instance regression (MIR) data are organized into bags (collections of instances in feature space) and the goal is to learn a mapping that assigns labels to bags. A typical assumption is that there is a so-called concept point in feature space, the proximity to which dictates the bag label. Motivated by modern MIR architectures which are based on attention, we study a softmax model that decouples the concept point and the labeling scheme. The two are respectively determined by a query direction and a value direction value in feature space. This problem isolates a basic challenge of learning both the query and value vectors from bag-level supervision. From this model we derive a parametric family of iterations in the noiseless limit, which generalizes a method known as the EM-DD algorithm. We then derive concentration results for the MLE estimators of the query and value vectors obtained from a random selection of instances. Our result for the value vector shows that a single random initialization of the value vector already points in the correct direction on average, so that a polynomial (in the number of instances per bag and the feature dimension) number of bags is enough for the EM algorithm to converge in $O(1)$ steps with high probability. A key aspect of this analysis is the interplay between concentration of empirical covariance matrices and extremal statistics arising from the selection rule.

Comments20 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑