发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SAIVE通过直方图的直方图过滤准则,基于Zipf-Mandelbrot幂律分布,低成本筛选数据湖中适合AI分析的高价值字段。
AI 中文摘要
数据湖存储大量遥测数据,来自网络传感器、主机和应用程序的日志可能包含每个事件的数百个字段。大型企业因此面临数据湖无法被AI高效分析的问题。聚合分析关注行为随时间发生的持续性变化。数据湖中的许多字段和列并不有用,因为它们包含的信息不够多样或集中,无法支持AI分析。SAIVE是一种简单的方法,用于检查大型表中的少量行,并应用直方图的直方图过滤标准来选择那些更可能为AI分析产生有用结果的字段。本文通过假设底层数据服从Zipf-Mandelbrot幂律分布,为SAIVE启发式方法提供了原理性基础。将Zipf-Mandelbrot指数alpha限制在合理范围内,提供了一种实用、廉价、无需专家的过滤器,用于在大型数据集中选择AI有价值实体。
英文摘要
Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.
CommentsTo be presented at IEEE MIT URTC 2026