arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不完美数据下的机器学习:挑战与方法

Machine Learning under Imperfect Data: Challenges and Methods

Masoumeh Zareapoor

arXiv 2609.13914首次发表:更新:

发表机构

Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文综述了不完美数据(缺失、不均衡、弱监督、分布偏移)下机器学习的挑战,归纳为信息丢失、经验风险偏差、监督模糊和表示不稳定四种机制,并探讨了相应方法及未来方向。

AI 中文摘要

机器学习模型通常在假设训练和测试数据足够完整、平衡、有标注且来自兼容分布的条件下开发。在实践中,这些条件中的一个或多个经常被违反。测量值可能缺失或损坏,稀有类别可能代表性不足,监督可能较弱,部署环境可能不同于训练环境。这些不完美通常被视为独立的技术问题,尽管它们通过少量共享机制改变学习:信息丢失、经验风险偏差、监督模糊和表示不稳定。这篇短综述围绕这些机制组织代表性方法。它回顾了重建与生成、再平衡与表示校准、有限监督下的学习、跨领域和跨模态的适应,以及分布变化下的可靠性。讨论强调了合理重建、基准特定修正和缺乏可信反馈的适应的局限性。它最后提出了证据感知学习、保留不确定性的预测以及将视觉合理性与决策效用分开的评估方向。

英文摘要

Machine-learning models are commonly developed under an assumption that training and test data are sufficiently complete, balanced, labelled, and drawn from compatible distributions. In practice, one or more of these conditions is often violated. Measurements may be missing or corrupted, rare classes may be poorly represented, supervision may be weak, and the deployment environment may differ from the training environment. These imperfections are usually treated as separate technical problems, although they alter learning through a small number of shared mechanisms: loss of information, biased empirical risk, ambiguous supervision, and unstable representations. This short survey organises representative methods around these mechanisms. It reviews reconstruction and generation, rebalancing and representation calibration, learning with limited supervision, adaptation across domains and modalities, and reliability under distribution change. The discussion highlights the limits of plausible reconstruction, benchmark-specific correction, and adaptation without trustworthy feedback. It concludes with directions for evidence-aware learning, uncertainty-preserving prediction, and evaluation that separates visual plausibility from decision utility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑