arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从知识获取到来源学习:发展来源特定能力

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

arXiv 2610.02150首次发表:更新:

发表机构

Georgia Institute of Technology; University of Illinois at Urbana-Champaign; University of California, Los Angeles(佐治亚理工学院; 伊利诺伊大学厄巴纳-香槟分校; 加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体重复使用外部来源时缺乏渐进理解的问题,提出SourceLearn,结合自我导向与任务引导两种学习机制构建持久来源模型,在五个基准上取得最优性能。

AI 中文摘要

大型语言模型(LLM)智能体日益依赖持久的外部来源来解决一系列知识密集型任务。现有方法改善了来源内容的访问和组织方式,而智能体记忆系统则保留先前交互中的可复用知识,但对同一来源的重复使用在很大程度上仍被视为重复访问,而非逐步改进对该来源理解的机会。我们研究来源学习:在持久权威来源上发展可复用的来源特定能力。我们用持久来源模型表示这种能力,该模型捕获对来源的可复用理解,包括其知识如何被结构化、解释和应用。为构建并逐步优化此类模型,我们提出SourceLearn,它结合了两种互补的学习机制。自我导向来源学习识别哪些内容尚未被完全理解,并自适应地重新访问来源,而任务引导来源学习利用下游经验揭示来源知识在组织方式上的局部表征缺口和反复出现的需求。在两种情况下,学习信号决定应重新考虑的内容,而持久更新则从权威来源中重建。在五个基准和三个LLM后端上,SourceLearn在15个设置中的13个中取得了最佳性能,相较于Hybrid RAG提升高达22.6个百分点,并在静态来源表示和基于经验的记忆基线上实现了显著的总体改进。

英文摘要

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.

CommentsWebsite: https://sourcelearn.github.io/ Code: https://github.com/luchengfu6/SourceLearn

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑