arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32546cs.LGcs.CL

共享自回归上下文会扭曲合成数据中的变量关系

Shared Autoregressive Context Can Distort Relationships in Synthetic Data

Thomas S. Robinson

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现大型语言模型在一次自回归生成中共享上下文会扭曲合成数据中的变量关系,通过受控实验证实该效应源于答案历史,并强调请求构建影响数据有效性。

中文摘要 AI 辅助

大型语言模型可以在一次自回归生成中产生多条记录,使较早的答案可作为后续记录的上下文。本文通过在合成调查受访者上的受控测试表明,这种共享完成批处理会扭曲所得合成数据中变量间的关系。在一项针对2,000个欧洲社会调查档案的匹配实验中,在保持档案、示例、问题和解码参数固定的情况下,每次请求生成10个而非1个受访者,会使国内相关性中的平均绝对误差在三个随机种子下增加48-58%(Qwen3.8-27B)和114-127%(Llama-3.3-70B-Instruct)。这种扭曲主要反映了关系强度的夸大,同时与人类对相关性的排序保持相当一致。受控干预确立了答案历史作为因果渠道:在档案和边际分布固定的情况下,重新配对相同的前置值会改变后续生成响应中的相关性。隐藏前置答案在测试设置中降低了相关性误差,但恶化了边际准确性。跨社会态度、健康和经济数据的探索性修正同样表明,较低的相关性误差可以与更差的边际分布和回归估计共存。因此,请求构建是数据生成过程的一部分,合成数据的有效性必须根据生成数据所支持的分析进行评估。

英文摘要

Large language models can generate several records within one autoregressive completion, making earlier answers available as context for later records. This paper shows that such shared-completion batching can distort relationships among variables in the resulting synthetic data, using controlled tests on synthetic survey respondents. In a matched experiment on 2,000 European Social Survey profiles, generating ten rather than one respondent per request increases mean absolute error in within-country correlations by 48-58% for Qwen3.8-27B and 114-127% for Llama-3.3-70B-Instruct across three seeds, holding profiles, examples, questions and decoding parameters fixed. The distortion primarily reflects exaggerated relationship strength, while retaining substantial agreement with the human ordering of correlations. Controlled interventions establish answer history as a causal channel: re-pairing the same preceding values, with profiles and marginal distributions fixed, changes correlations among subsequently generated responses. Hiding preceding answers reduces correlation error in the tested settings but worsens marginal accuracy. Exploratory corrections across social-attitude, health and economic data likewise show that lower correlation error can coexist with worse marginal distributions and regression estimates. Request construction is therefore part of the data-generating process, and synthetic-data validity must be evaluated against the analyses the generated data are intended to support.

补充信息

↑