arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LAB-Tab:面向小样本表格生成的大语言模型增强型贝叶斯网络适配方法

LAB-Tab: LLM-Augmented Bayesian Network Adaptation for Few-Shot Tabular Generation

Zijian Shen, Taijie Chen, Bin Zhou, Ziyang Jiang, Jintao Ke

arXiv 2608.01879首次发表:更新:

AI 中文总结

针对小样本表格生成的分布偏移问题,提出LLM增强的贝叶斯网络适配框架LAB-Tab,在六个ACS分布偏移场景中表现优于基准方法,性能提升显著。

AI 中文摘要

当目标域数据稀缺时,表格数据生成可支撑分析与决策,但收集完整的目标样本往往成本高昂。一种实用却未被充分探索的场景是仅提供少量目标记录,同时提供来自相关域的更丰富源数据。现有的小样本表格生成器要么直接拟合稀疏的目标统计量,可能对偶然模式过拟合;要么复用源域生成器,可能保留在目标域不再成立的依赖关系。为解决该问题,我们提出LAB-Tab,一种面向源感知小样本表格生成的大语言模型(LLM)增强型贝叶斯网络(BN)适配框架。LAB-Tab首先从源数据拟合BN,随后利用LLM提出源BN图中不存在的合理目标域BN边,该步骤将语义与弱统计证据转化为显式结构假设,从而在源拟合图之外扩展了可编辑边空间。由于提出的边可能存在噪声并与现有依赖交互,近端策略优化(PPO)策略通过边级操作校准增强型BN的边,操作包括保留、弱化、强化、翻转和弃权(不执行)。PPO策略使用结合分布对齐、下游效用及目标相关依赖保留的奖励进行训练,随后对适配后的BN采样以合成目标域表格。在由三项美国人口普查(ACS)预测任务构建的六个源-目标分布偏移场景中,当目标数据预算为10%时,LAB-Tab取得最佳性能,在六个单独场景中领先四个,相比最强基准将宏观综合得分降低33.8%,同时在保持竞争力的特征-标签保留的情况下,获得最佳宏观JSD、WAPE和UtilityGap。

英文摘要

Tabular data generation supports analysis and decision-making when target-domain data are scarce, yet collecting complete target samples is often costly. A practical but underexplored setting provides only a few target records together with richer source data from a related domain. Existing few-shot tabular generators often either fit sparse target statistics directly, which can overfit incidental patterns, or reuse source-domain generators, which may preserve dependencies that no longer hold in the target domain. To address this problem, we propose LAB-Tab, an LLM-augmented Bayesian network (BN) adaptation framework for source-aware few-shot tabular generation. LAB-Tab first fits a BN from source data and then uses an LLM to propose plausible target-domain BN edges that are absent from the source BN graph. This step converts semantic and weak statistical evidence into explicit structural hypotheses, thereby expanding the editable edge space beyond the source-fitted graph. Because the proposed edges may be noisy and interact with existing dependencies, a PPO policy calibrates edges in the augmented BN through edge-level actions, including keep, weaken, strengthen, flip, and deactivate. The PPO policy is trained with a reward that combines distributional alignment, downstream utility, and preservation of target-relevant dependencies. The adapted BN is then sampled to synthesize target-domain tables. Across six source--target distribution-shift scenarios built from three US Census (ACS) prediction tasks, LAB-Tab achieves the best performance at the 10% target-data budget, leads four of the six individual scenarios, and reduces the macro Overall score by 33.8% relative to the strongest baseline. It also obtains the best macro JSD, WAPE, and UtilityGap while maintaining competitive feature--label preservation.

Comments21 pages, 6 figures. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑