从内部表征到通过预测误差改进模型
From internal representations to model improvement through prediction errors
浏览论文内容
中文总结 AI 辅助
本研究提出一种利用模型内部特征关联预测误差及其性能影响的数据选择方法,无需候选图像标签,在目标检测任务上显著优于基于特征稀有性的选择,并排名前列。
中文摘要 AI 辅助
在标注预算有限的情况下,选择哪些图像进行标注决定了模型能改进多少。使用来自单独训练模型的特征或视觉语言模型编写的场景描述的数据选择方法已取得成功,但这些信号并未直接捕捉到被改进模型的变化。目标模型自身的内部特征反映了其迄今所学到的内容,并随重新训练而变化,使其成为选择下一批训练数据的自然线索。然而,仅凭特征稀有性并不能揭示对性能至关重要的误差。在此,我们将内部特征与预测误差及其对性能的预期影响联系起来,并在不使用候选图像标签的情况下选择图像进行标注和重新训练。我们使用一个目标检测器在两个数据集和两对随机种子上评估了该方法。加入内部特征在16个条件中的15个条件下改善了对预测误差的识别。当性能在连续标注轮次上取平均时,该方法在所有四个评估设置中均优于仅基于特征稀有性的选择,并在六种方法中排名前二。在其他条件保持不变的情况下,重新训练后的性能再次高于基于稀有性的选择,尽管后者收集了更多误差。随着重新训练的延长,所提方法在六种方法中排名第一。这些结果表明,将模型的内部特征与其误差及误差对性能的影响联系起来,可能有助于选择能提升性能的训练图像,从而使模型的当前状态指导接下来标注哪些图像。
英文摘要
With limited annotation budgets, choosing which images to label determines how much a model improves. Data-selection methods that use features from a separately trained model, or scene descriptions written by vision-language models, have been successful, but those signals do not directly capture changes in the model being improved. The target model's own internal features reflect what it has learned so far and change with retraining, making them a natural cue for choosing the next training data. However, feature rarity alone does not reveal the errors that matter for performance. Here we link internal features to prediction errors and their expected impact on performance and select images for labeling and retraining without using labels for candidate images. We evaluated the method with an object detector on two datasets and two pairs of random seeds. Adding internal features improved the identification of prediction errors in 15 of 16 conditions. When performance was averaged over successive labeling rounds, the method outperformed selection based only on feature rarity in all four evaluation settings and ranked among the top two of six methods. With other conditions held fixed, performance after retraining was again higher than with rarity-based selection, even though the latter collected more errors. With longer retraining, the proposed method ranked first among six methods. These results suggest that linking a model's internal features to its errors and their effects on performance may help select training images that improve performance, thereby allowing the model's current state to guide which images are labeled next.
发表机构
- Adansons Corp.(Adansons公司)
- Tohoku University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。