arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15109cs.AI

基于大语言模型智能体的列间约束发现的约束感知合成表格数据生成

Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents

Jianxing Zhao, Mao Guan, Dongyu Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对合成表格数据易违反领域约束的问题,提出基于大语言模型智能体的列间约束发现的统一工作流,结合后处理器修复输出,提升了约束合规性与下游效用。

中文摘要 AI 辅助

生成结构有效的合成表格数据仍然存在困难:具有高统计保真度和下游效用的输出仍可能违反具有语义意义的领域约束。我们研究三类互补的列间约束族——方程、线性不等式和逻辑依赖关系的发现与执行。我们的统一工具化工作流将这三类约束都表示为机器可执行的假设,并应用通用接口实现全表验证、确定性诊断和反例引导的修正。与生成器无关的后处理器协调对来自未改变的表格生成器的输出进行特定于族的修复。在精心策划的行为审计和端到端评估中,完整工作流相对于一次性直接提示提高了保留的约束的检测率,而后处理对于每个保留的适用约束都产生了零测量违规,在大多数数据集上提高了下游效用,并在很大程度上保留了单变量边际分布。

英文摘要

Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.

发表机构

  • University of California, Davis(加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

↑