arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEBGen:一种用于少样本出行调查数据生成的LLM增强贝叶斯网络框架

LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation

Zijian Shen, Bin Zhou, Jiguang Wang, Ya Zhao, Jintao Ke

arXiv 2609.08288首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LEBGen提出LLM增强贝叶斯网络框架,利用大语言模型知识细化结构,从少样本生成出行调查数据,显著提升分布与依赖保真度。

AI 中文摘要

出行调查数据对于交通规划和出行行为分析至关重要,然而收集大规模具有代表性的样本既昂贵又耗时。一种实用的替代方案是从少样本中生成合成调查记录。然而,此类样本对异质性出行者群体的覆盖不完整,且证据不足以恢复人口统计特征与出行行为之间的复杂依赖关系。现有方法存在互补的局限性。概率生成模型(如贝叶斯网络(BN))提供显式的分布控制,但从少样本中学习的结构可能遗漏有意义的依赖关系或保留虚假的依赖关系。大语言模型(LLM)可以通过提供补充有限统计证据的行为知识,帮助解决BN结构学习中的这些困难。因此,我们提出了LEBGen,一种LLM增强的BN框架,利用这些知识来细化网络结构,以用于少样本出行调查数据生成。具体而言,LLM首先从人口统计属性和出行行为统计中识别出行者画像,然后恢复被画像增强的BN结构遗漏的依赖关系,并剪除虚假的依赖关系。细化后的BN仅从观测数据中参数化以生成合成记录。在2022年香港出行特征调查的2%少样本设置下,LEBGen将平均边际Jensen-Shannon散度从0.0671降至0.0091,并将平均绝对Cramer's V误差相对于最佳基线降低了14.3%,显著提高了分布保真度和依赖保真度。

英文摘要

Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑