arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09913cs.DB

合成数据压缩的案例

The Case for Synthesized Data Compression

Ivan Iachnyk, Mihail Stoian, Luisa Neubauer, Ryan Marcus, Andreas Kipf

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于智能体的数据压缩算法合成方法,针对特定数据集定制压缩方案,在构建时完全验证,实验显示压缩因子比通用方法提升近2倍。

中文摘要 AI 辅助

压缩是数据管理中的基本工具。如今的压缩方案通常由人工设计,适用于多种不同类型的数据(例如游程编码或算术编码),导致出现了大量通用算法,这些算法在许多数据集上表现良好,但可能对任何特定数据集都不是最优的。受近期智能体编码进展的启发,我们提出在逐数据集的基础上合成定制化的数据压缩算法。与先前使用智能体合成数据库组件的工作不同,数据特定的压缩算法可以在构建时完全验证,从而消除了大多数正确性问题。现代文件格式,如AnyBlox和F3,允许将合成代码与数据一起包含。实验上,我们表明我们的合成智能体可以扩展压缩因子/解压速度的帕累托前沿,在压缩因子方面比通用方法高出近2倍。

英文摘要

Compression is a fundamental tool in data management. Today's compression schemes are generally designed by hand to apply to many different types of data (e.g., run length encoding or arithmetic encoding), leading to a zoo of general-purpose algorithms which work well for many datasets, but may not be optimal for any specific dataset. Inspired by recent advancements in agentic coding, we propose synthesizing custom-tailored data compression algorithms on a per-dataset basis. Unlike prior work using agents to synthesize database components, a data-specific compression algorithm can be fully verified at construction time, alleviating most correctness concerns. Modern file formats, like AnyBlox and F3, allow including the synthesized code alongside the data itself. Experimentally, we show that our synthesis agent can expand the compression-factor/decompression-speed Pareto front, outperforming general-purpose methods by nearly $2\times$ in compression factor.

发表机构

  • University of Technology Nuremberg(纽伦堡工业大学)
  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑