arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10155cs.CLcs.IR

从检索到权重:基于个体文本语料库的小语言模型参数化个性化

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

  • Bergische Universität Wuppertal(伍珀塔尔大学)
  • Goethe-Universität Frankfurt(法兰克福大学)

机构由 AI 辅助整理,请以论文原文为准。

Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Markus J. Hofmann

中文总结 AI 辅助

本研究通过DoRA微调将个体文本语料库整合进小语言模型权重,实现参数化个性化,实验表明适配器能有效拟合个体文本,为个性化辅导代理奠定基础。

中文摘要 AI 辅助

我们从认知模拟的视角出发,在多项选择问答中处理情景记忆和语义记忆,通过将个体文本语料库(ITC)中的文本整合到检索增强生成和DoRA微调中。我们通过网络爬取515名参与者的搜索历史,这些参与者回答了36个多项选择知识题,并分析了一个包含150名参与者的分层子样本。对于每位参与者,一个DoRA适配器将其ITC整合到一个小语言模型(SLM)中,该模型的基线正确率低于参与者最低四分位数。该适配器可测量地将ITC写入权重:它对自己参与者的留出文本的拟合程度优于其他参与者的文本(dz=1.27),这种个体性效应随ITC大小的排序而增强。然而,在泛化知识测试中,适配器增加的是知识而非与个体的对齐:对数损失匹配有所改善,而在偏差校正的PMI读出下的匹配准确率并未提高,且检索并未带来额外提升。我们的结果表明,ITC可以被整合到SLM的权重中,这为个性化辅导代理提供了令人鼓舞的基础,我们讨论了如何从这一基础出发,在个体层面实现对情景记忆和语义记忆的现实模拟。

英文摘要

We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice knowledge items and analyze a stratified subsample of 150 participants. For each participant, one DoRA adapter consolidates their ITC into a small language model (SLM) whose baseline correctness falls below the participants' lowest quartile. The adapter measurably writes the ITC into the weights: it fits its own participant's held-out text better than other participants' texts (dz =1.27), an individuality effect that increases with ITC size in rank order. On the generalized knowledge test, however, the adapter adds knowledge rather than alignment with the individual: log-loss match improves, whereas match accuracy under a bias-corrected PMI readout does not, and retrieval adds nothing on top. Our results demonstrate that ITCs can be consolidated into the weights of SLMs, an encouraging basis for individualized tutoring agents, and we discuss how to move from there toward a realistic simulation of episodic and semantic memory at the individual level.

↑