arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越原始文本记录:面向基于大语言模型的数字孪生的结构化角色提取

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

Iris Ye, Tianze Deng, Ozan Candogan

arXiv 2608.20344首次发表:更新:

发表机构

University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对基于LLM的数字孪生,提出结构化角色提取方法,发现信息结构是性能关键,固定结构难通用,自动结构发现流水线可恢复性能并提升准确率。

AI 中文摘要

基于大语言模型(LLM)的“数字孪生”旨在,给定个体过往行为的某种表征,模拟该个体在新环境中的行为表现或对新问题的回应。常见方法是从调查文本记录或摘要回应中构建该表征。已有研究表明,将长文本记录压缩为LLM生成的较短摘要不会显著降低预测准确率,这说明信息量并非主要瓶颈。在本研究中,我们认为关键限制在于结构:即角色信息在提供给模拟器模型之前的组织方式。我们通过对比非结构化摘要与结构化角色表征来研究这一问题。首先,我们引入了一个基于消费者行为理论的手工构建的模式(BDE:背景、决策程序、评估),并在同构基准(Twin-2K-500)上证明其相比原始文本记录可提升1.91个百分点的预测准确率,在gpt-5.4-mini和Qwen3-8B上进行鲁棒性检查时也取得了类似提升。然而,这种固定结构无法在更多异构任务上通用,在这些任务中其性能与原始文本记录基线无统计显著差异。为解决这一局限,我们提出了一种自动结构发现流水线,其中LLM会迭代提出并优化任务特定的角色结构与提取提示。在包含13项不同子研究的基准上,该方法恢复了性能,相比原始文本记录基线平均准确率提升1.91个百分点,并消除了固定模式下观察到的显著损失。总体而言,我们的结果表明,基于LLM的数字孪生的主要约束并非提供的信息量,而是信息的组织方式——且最优结构取决于具体任务。

英文摘要

LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.

CommentsPreprint. Submitted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑