AI 中文总结
研究针对IMDb真实世界数据集无比例因子问题,提出结合大语言模型语义字典和确定性时间图关系数据生成的好莱坞数据集生成方法,实验表明其引发的基数估计误差与原数据集相当或更大,还能测试估计器泛化能力。
AI 中文摘要
在过去十年中,JOB基准测试的IMDb真实世界数据集因其能够对传统估计器和学习估计器进行压力测试,而被广泛用于基数估计研究。然而,与合成的TPC系列不同,它没有比例因子,只是一个简单的转储。我们引入了好莱坞,这是一个与IMDb兼容的合成基准生成器,它将大语言模型生成的语义字典与基于确定性时间图的关系数据生成相结合。我们分析了一个初步的好莱坞200K数据集,它包含200,000部主要电影、生成的系列和剧集标题行、1970万条IMDb风格的行,以及213个非零的JOB-Light、JOB和JOB-Complex查询。对两个开放系统的实验表明,好莱坞引发的基数估计误差与在原始IMDb数据集上观察到的误差相当或更大。此次发布包括生成设置、提示/大语言模型输出出处以及适配的SQL和标签,能够测试基数估计器是否能在固定电影快照和分布之外进行泛化。
英文摘要
The IMDb real-world dataset of the JOB benchmark has been extensively used in the last decade as part of the research line on cardinality estimation, given its ability to stress test both traditional and learned estimators. However, unlike the synthetic TPC family, it does not come with a scale factor, being a simple dump. We introduce Hollywood, a synthetic IMDb-compatible benchmark generator that combines LLM-generated semantic dictionaries with deterministic temporal-graph-based relational data generation. We analyze a preliminary Hollywood-200K, which contains 200,000 primary movies, generated series and episode title rows, 19.7M IMDb-style rows, and 213 nonzero JOB-Light, JOB, and JOB-Complex queries. Experiments with two open systems demonstrate that Hollywood induces cardinality estimation errors comparable to or exceeding those observed on the original IMDb dataset. The release includes generation settings and prompt/LLM-output provenance together with adapted SQL and labels, enabling tests of whether cardinality estimators generalize beyond a fixed movie snapshot and distribution.
CommentsAccepted at the Seventh International Workshop on Applied AI for Database Systems and Applications (AIDB 2026)