个体文本语料库预测用户特定知识:个体化知识模拟的基准
Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation
浏览论文内容
中文总结 AI 辅助
本研究利用搜索历史构建个体文本语料库,通过微调Qwen3-1.7B和检索增强生成模拟个体知识,在公开题目上超越常模,但非公开题目表现不佳,并提出基于熵的评估基准。
中文摘要 AI 辅助
本研究探讨是否可以利用来自搜索历史的个体文本语料库(ICs)来模拟个体知识。我们收集了316名成年人的ICs,这些成年人回答了36道多项选择知识题,并在此任务上比较了几种大型语言模型(LLMs),其中只有Qwen3-1.7B被证明可行。通过低秩适配(LoRA)进行任务特定微调后,Qwen3-1.7B在公开可用的题目上表现优于参与者和一个具有代表性的德国常模样本。然而,在非公开问题上,该LLM的表现不如我们的参与者,这暗示公开问题可能存在训练数据污染。当将ICs整合到检索增强生成中以预测个体反应时,LLM与参与者的匹配准确率显著高于随机水平,这表明存在可检测的个体知识信号。然而,分配给参与者答案的概率较低,远低于正确答案的概率,表明对个体反应模式的校准不佳。知识差距预测并非最优,尽管对于超过五百万词元的语料库有所改善。我们讨论了基于熵的评估基准作为个体化知识模拟的校准指标。
英文摘要
This study examines whether individual text corpora (ICs) from search histories can be used to simulate individual knowledge. We collected ICs from 316 adults, who answered 36 multiple-choice knowledge items, and compared several large language models (LLMs) on this task, of which only Qwen3-1.7B proved viable. After task-specific fine-tuning via Low-Rank Adaptation (LoRA), Qwen3-1.7B outperformed both participants and a representative German norm sample on publicly available items. On non-public questions, however, the LLM performed worse than our participants, suggesting possible training data contamination for the public questions. When integrating ICs into retrieval-augmented generation to predict individual responses, LLM-participant Match accuracies significantly exceeded chance, which demonstrates a detectable individual knowledge signal. The probabilities assigned to the participants' answers were, however, low and far below the probability of correct answers, indicating poor calibration toward individual response patterns. Knowledge-gap prediction was sub-optimal, though it improved for corpora exceeding five million tokens. We discuss our entropy based evaluation benchmarks as calibration indices for individualized knowledge simulation.
发表机构
- Bergische Universität Wuppertal(伍珀塔尔大学)
- Goethe-Universität Frankfurt(法兰克福歌德大学)
机构由 AI 辅助整理,请以论文原文为准。