arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21861cs.CLcs.AI

容量之上的数据质量:将文档内化到LoRA适配器中用于闭卷问答

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA

Joan Figuerola Hurtado

首次发表
浏览论文内容

中文总结 AI 辅助

研究将文档内化到LoRA适配器用于闭卷问答,发现适配器容量足够时,训练数据质量主导闭卷准确率,一次整理大幅提升准确率,确认容量趋势及秩与学习率耦合,内化适配器在低延迟下表现优于检索基线,还报告了调试LLM训练的过程。

中文摘要 AI 辅助

我们研究通过LoRA将文档直接烘焙到4位Gemma - 4 - e4b模型的权重中,使系统能够闭卷回答关于语料库的问题,即无需检索且无上下文窗口预算。在从单文档到99文档语料库的约100次训练运行中,我们发现一旦适配器容量足够,训练数据质量是闭卷准确率的主导因素,超过LoRA秩、学习率和两种替代架构的组合;容量本身是一个硬门槛,低于此门槛数据干预无效。一次整理(将标准答案缩短为规范的1 - 6词跨度并删除琐事)使15文档语料库的闭卷准确率从57.7%提高到85.7%。我们确认了容量趋势(秩必须随语料库大小增长)以及秩与学习率之间的耦合,我们最初误诊了这种耦合。在15文档切片上,我们添加了实际检索基线:内化适配器(召回率84.2%)在更低延迟下击败了带有基础阅读器的BM25 - RAG管道(58.9%),甚至击败了实际的黄金块预言机(65.6%)。我们报告了整个过程,包括三次误诊,作为实证调试LLM训练的案例研究。

英文摘要

We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book: no retrieval and no context-window budget. Across roughly 100 training runs from single documents to a 99-document corpus, we find that once adapter capacity is adequate, training-data quality is the dominant lever on closed-book accuracy, outweighing LoRA rank, learning rate, and two alternative architectures combined; capacity itself is a hard gate below which no data intervention helps. A single curation pass (shortening gold answers to canonical 1-6 word spans and dropping trivia) moved closed-book accuracy from 57.7% to 85.7% on a 15-document corpus, a larger jump than any architectural change. We confirm a capacity trend (rank must grow with corpus size) entangled with a coupling between rank and learning rate that we initially misdiagnosed. On a 15-document slice we add a real retrieval baseline: the internalized adapter (84.2% recall) beats a BM25-RAG pipeline with a base reader (58.9%) and even a realistic gold-chunk oracle (65.6%) at lower latency. We report the full arc, including three misdiagnoses, as a case study in debugging LLM training empirically.

↑