arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识即技能:面向LLM智能体的自主知识库使用的结构化设计

Knowledge-as-Skill: A Structural Design for Autonomous Knowledge-Base Use by LLM Agents

Jiangxu Wu

arXiv 2609.25991首次发表:更新:

AI 中文总结

提出知识即技能方案,通过三层结构使知识库可发现可导航,在WixQA基准上提升事实性与上下文召回率,实现LLM智能体自主知识库使用。

AI 中文摘要

检索增强生成(RAG)使大型语言模型(LLM)能够访问外部知识,但其传统的检索-拼接-生成流程替模型做出了检索决策。随着工具使用和智能体循环变得更加可靠,智能体可以自行决定是否检索、检查什么以及何时停止。这种转变暴露了一个新的瓶颈:智能体可能不知道知识库中包含什么。传统知识库将文档呈现为匿名的文本块,关于范围、用途、来源或关系的信息有限。我们提出了知识即技能(Knowledge-as-Skill),一种使知识库可发现、可导航且自描述的组织方案。它包含三层:以this http URL为中心发现层;每个目录对应一个this http URL的导航层;以及包含带有YAML前置元数据(主题、类型、来源和生命周期)文档的知识层。该设计遵循开放知识格式(OKF)和技能协议,无需修改智能体框架。我们还提供了知识即技能管道,用于将异构的PDF、Word文件、网页导出和笔记集合转换为这种结构。在WixQA企业客户支持基准的初步评估中,我们的设置获得了0.889的事实性和0.816的上下文召回率,而报告的Corpus2Skill值分别为0.767和0.708。它在忠实度上略低,上下文精确度较低,交互轮次更多。由于模型、提示和知识包构建不同,这些结果是跨工作的方向性证据,而非受控比较。

英文摘要

Retrieval-augmented generation (RAG) gives large language models (LLMs) access to external knowledge, but its conventional retrieve-concatenate-generate pipeline makes retrieval decisions on behalf of the model. As tool use and agent loops become more reliable, an agent can decide whether to retrieve, what to inspect, and when to stop. This shift exposes a new bottleneck: the agent may not know what a knowledge base contains. Traditional knowledge bases expose documents as anonymous text chunks with limited information about scope, purpose, provenance, or relations. We propose Knowledge-as-Skill, an organization scheme that makes a knowledge base discoverable, navigable, and self-descriptive. It has three layers: a discovery layer centered on SKILL.md; a navigation layer with one index.md per directory; and a knowledge layer containing documents with YAML frontmatter for topic, type, provenance, and lifecycle. The design follows the Open Knowledge Format (OKF) and the Skill protocol without modifying the agent framework. We also provide knowledge-as-skill, a pipeline for converting heterogeneous collections of PDFs, Word files, web exports, and notes into this structure. In a preliminary evaluation on the WixQA enterprise customer-support benchmark, our setup obtains 0.889 Factuality and 0.816 Context Recall, compared with reported Corpus2Skill values of 0.767 and 0.708. It obtains slightly lower Faithfulness, lower Context Precision, and more interaction turns. Because the models, prompts, and knowledge-package construction differ, these results are directional cross-work evidence rather than a controlled comparison.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑