AI 中文总结
本研究针对一名用户的LINE对话历史开展检索增强生成的初始案例,构建三种搜索表示并对比多种检索方法,发现特定混合检索配置在100个评估问题上的检索性能最优,但研究存在单用户单标注等局限性。
AI 中文摘要
作为大型语言模型(LLM)个人记忆检索增强生成(RAG)的初始步骤,本研究针对一名用户的LINE对话历史开展了仅检索的案例研究。我们将358896条消息分割为22329个时间连贯的块,并构建了三种搜索表示:raw_text(原始文本)、生成的摘要,以及embedding_text(嵌入文本,结合摘要、原始文本摘录及其他固定文本)。我们在由一名标注者验证的100个评估问题上,比较了BM25、稠密向量检索和线性混合检索。在单一检索器中,embedding_text_bm25(结合embedding_text与BM25的检索器)达到最高点估计值,Recall@5为0.584。随后,我们在同一评估集上探索了6种检索器配对和21种权重,共126种配置。所选的embedding_text_bm25与embedding_text_vector(结合embedding_text与稠密向量的检索器)在beta=0.45时的组合,实现了Recall@5=0.697、MRR@5=0.595、nDCG@5=0.575。其Recall@5较embedding_text_bm25高出0.113,问题层面的配对百分位数自助法95%置信区间为[0.048, 0.184]。该区间以固定100个问题上所选配置为条件,未考虑配置选择或权重搜索带来的不确定性。其与beta=0.50时基于摘要的混合检索的差值为0.050,95%置信区间为[-0.013, 0.115],故无法确定明确差异。17个聚合问题的点估计值也低于其他问题类型,表明当证据分布在多个时间点和对话中时,平滑块级检索效果不佳。本评估是一项探索性的单用户、单标注者研究,采用了用于配置搜索的同一问题集,未评估最终答案生成或对未见问题的泛化能力。
英文摘要
As an initial step toward personal memory retrieval-augmented generation (RAG) for large language models (LLMs), this study presents a retrieval-only case study over one user's LINE conversation history. We segmented 358,896 messages into 22,329 temporally coherent chunks and constructed three search representations: raw_text, a generated summary, and embedding_text, which combines a summary with a raw-text excerpt and other fixed text. We compared BM25, dense vector retrieval, and linear hybrid retrieval on 100 evaluation questions verified by a single annotator. Among individual retrievers, embedding_text_bm25 achieved the highest point estimate, with Recall@5 of 0.584. We then explored six retriever pairings and 21 weights, for 126 configurations on the same evaluation set. The selected combination of embedding_text_bm25 and embedding_text_vector at beta = 0.45 achieved Recall@5 = 0.697, MRR@5 = 0.595, and nDCG@5 = 0.575. Its Recall@5 exceeded that of embedding_text_bm25 by 0.113, with a question-level paired percentile-bootstrap 95% confidence interval of [0.048, 0.184]. This interval is conditional on fixing the configuration selected on the same 100 questions and does not account for uncertainty from configuration selection or weight search. The difference from a summary-based hybrid at beta = 0.50 was 0.050, with a 95% confidence interval of [-0.013, 0.115], so no clear difference could be established. The 17 aggregate questions also yielded lower point estimates than the other question types, suggesting that flat chunk-level retrieval struggles when evidence is distributed across multiple times and conversations. This evaluation is an exploratory single-user, single-annotator study conducted on the same question set used for configuration search; it does not evaluate final answer generation or generalization to unseen questions.
Comments16 pages, 6 figures, 10 tables. Exploratory single-user case study