arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25487cs.CR

识别Google Play应用移除预测中的疑似错误标注应用:标签噪声检测方法的实证比较

Identifying Suspected Mislabeled Apps in Google Play Application Removal Prediction: An Empirical Comparison of Label Noise Detection Methods

Deborah Dobles Montalvan, F. Mohsen, H. de Weerd

首次发表
浏览论文内容

中文总结 AI 辅助

本研究比较三种标签噪声检测方法在Google Play应用移除预测中的应用,发现被标记的疑似错误标注应用移除后无益,但有助于刻画标签噪声特征。

中文摘要 AI 辅助

预测哪些Google Play应用将被移除的模型,其训练标签仅记录应用在后续观察时是否仍在商店中。消失的应用被标记为已移除,存在的应用被标记为稳定,但两者均未记录原因。自愿下架和政策性移除都会产生“已移除”标签,而未被捕获的垃圾应用则保持“稳定”标签。本研究将这种不匹配称为标签噪声。来自不同方法论家族的三种检测器被应用于Mohsen、Karastoyanova和Azzopardi(2022)的870,514个应用:隔离森林(Isolation Forest),标记在特征空间中异常的应用;邻域不一致(Neighborhood Disagreement),标记其最近邻携带相反标签的应用;以及预测不一致(Prediction Inconsistency),标记分类器与数据标签不同的应用。被所有三种检测器标记的应用(即重叠部分)在默认设置下数量为7,598个,是最强的错误标注候选。随之产生两个问题。第一,移除被标记的应用是否能改进模型?不能。没有任何检测器、重叠集或并集能超越基线,且性能损失随移除数量增加而增大。第二,被标记的应用在VirusTotal和Quark Engine确认标签的应用中是否出现频率低于预期?在已确认移除的应用中确实如此,随着阈值收紧,其出现频率降至预期率的0.43倍,而在稳定侧的多余现象在考虑被扫描应用的年龄后消失。仅使用3,021个可训练的重叠应用训练的模型达到的测试AUC为0.2518,远低于随机水平,因此这些应用中特征与标签的关系与其余数据相反。被标记的应用在两个方向上都出错:类似垃圾邮件的被弃应用携带“稳定”标签,而看起来健康的应用携带“已移除”标签。这些检测器的价值在于刻画这种标签噪声的特征。它们定位了一小部分候选,但无法从中获益地移除它们。

英文摘要

Models that predict which Google Play apps will be removed are trained on labels that record only whether an app was still in the store at a later observation. A disappeared app is labeled removed and a present one stable, but neither records why. A voluntary withdrawal and a policy takedown both produce removed, and an uncaught spam app keeps stable. This work calls that mismatch label noise. Three detectors from different methodological families are applied to the 870,514 apps of Mohsen, Karastoyanova, and Azzopardi (2022): Isolation Forest, flagging apps unusual in the feature space, Neighborhood Disagreement, flagging apps whose nearest neighbors carry the opposite label, and Prediction Inconsistency, flagging apps a classifier labels differently from the data. The apps flagged by all three, the overlap, number 7,598 at default settings and are the strongest mislabeling candidates. Two questions follow. First, does removing flagged apps improve the model? It does not. No detector, overlap, or union beats the baseline, and the loss grows with the number removed. Second, do flagged apps appear less often than expected among apps whose label VirusTotal and Quark Engine confirm? Among confirmed removals they do, falling to 0.43 times the expected rate as the threshold tightens, while an excess on the stable side disappears once the age of the scanned apps is accounted for. A model trained on only the 3,021 trainable overlap apps reaches a test AUC of 0.2518, far below chance, so the relationship between features and labels there runs opposite to the rest of the data. The flagged apps run wrong in both directions: abandoned apps that resemble spam carry stable, while apps that look healthy carry removed. The value of the detectors lies in characterizing this label noise. They locate a small set of candidates they cannot profitably remove.

发表机构

  • University of Groningen(格罗宁根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑