AI 中文总结
针对现有数据单元测试框架忽略下游任务代码语义的问题,提出复合AI系统PrismaDV,联合分析数据与任务代码生成任务感知数据单元测试,通过交互式界面完成验证与优化。
AI 中文摘要
数据是现代企业和机构的核心资源,数据管道中传播的数据错误会对生产造成严重影响,因此数据验证对于确保下游应用的可靠性至关重要,这催生了数据单元测试的开发,这类可执行程序会在数据通过大型数据管道移动前对其进行测试。然而,现有框架仅从观测数据中推导数据单元测试,忽略了下游使用数据的代码的语义。为此,我们提出PrismaDV,一种复合AI系统,它通过联合分析数据和下游任务代码,为表格数据综合任务感知的数据单元测试。PrismaDV将测试生成分解为多个由大语言模型(LLM)驱动的步骤:数据剖析、列访问检测、任务代码中的数据流分析以及隐式数据假设的推断。随后,它合成数据单元测试的代码,并维护一个内部“数据-代码假设图”,将生成的数据约束与任务的源代码关联起来。我们通过一个基于网络的交互式界面展示PrismaDV,参与者在该界面上用5个包含60个下游任务的真实数据集运行该系统,综合、检查和完善关于数据的自然语言假设以及可执行数据约束。该界面允许参与者导航数据-代码假设图、在错误数据批次上将任务感知的数据单元测试与任务不可知的基线进行比较,并交互式编辑假设和数据约束。此外,参与者还可以观察自定义提示优化器如何随时间使系统适应特定数据集。
英文摘要
Data is a central resource for modern enterprises and institutions, and data errors propagating through data pipelines lead to serious impact in production. Therefore, data validation is essential for ensuring the reliability of downstream applications. This led to the development of data unit tests, executable programs that test data before moving it around through large data pipelines. However, existing frameworks derive data unit tests from observed data alone, ignoring the semantics of the code that consumes the data downstream. To this end, we present PrismaDV, a compound AI system that synthesizes task-aware data unit tests for tabular data by jointly analyzing data and downstream task code. PrismaDV decomposes the test generation into multiple LLM-powered steps: data profiling, detection of column accesses, data flow analysis in the task code, and the inference of implicit data assumptions. It subsequently synthesizes code for the data unit test, and maintains an internal ``data-code assumption graph'' that links generated data constraints back to the task's source code. We demonstrate PrismaDV through an interactive web-based interface where attendees run the system on five real-world datasets with 60 downstream tasks, synthesize, inspect and refine both natural language assumptions about the data and executable data constraints. The interface allows attendees to navigate the data-code assumption graph, compare task-aware data unit tests against task-agnostic baselines on erroneous data batches, and interactively edit assumptions and data constraints. Furthermore, attendees can observe how a custom prompt optimizer adapts the system to specific datasets over time.
CommentsAccepted as a demo paper at CIKM 2026