基于压缩的机器学习导论
An Introduction to Compression-Based Machine Learning
浏览论文内容
中文总结 AI 辅助
本文系统梳理压缩与机器学习的双向转换关系,提出并验证基于压缩的机器学习设计框架,在恶意软件任务上表现优于传统基线,准确率提升可达0.62。
中文摘要 AI 辅助
任何无损压缩算法(如gzip)都可以通过归一化压缩距离或最小描述长度原理转换为机器学习方法。任何自回归模型都可以通过熵编码转换为无损压缩方法。这种看似循环的依赖关系在现代人工智能和机器学习中具有未实现的潜力,我们调查并形式化了用于利用压缩进行机器学习的各种策略。我们引入并实证验证了一个基于压缩的机器学习设计框架,发现基于压缩的方法与传统基线相比具有竞争力,在恶意软件检测上明显更强。我们发现,改变这些设计选择可带来高达0.62的准确率提升。
英文摘要
Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly circular dependence has unrealized potential in modern artificial intelligence and machine learning, and we survey and formalize the various strategies that have been used to leverage compression for machine learning. We introduce and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware. We find that varying these design choices yields accuracy gains of up to 0.62.
发表机构
- University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
- CrowdStrike(CrowdStrike公司)
- Syracuse University(雪城大学)
机构由 AI 辅助整理,请以论文原文为准。