发表机构
PleIAs; Lattice, ENS-PSL; Sorbonne Center for Artificial Intelligence; Sciences Po Médialab; Paris Dauphine-PSL(PleIAs; Lattice,巴黎高等师范学院-PSL; 索邦人工智能中心; 巴黎政治学院媒介实验室; 巴黎多菲纳-PSL大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Wikidata搜索轨迹数据集及递归语言模型框架,通过结构可控的多跳问题训练图谱搜索智能体,实验表明该框架显著提升开放与封闭模型的搜索准确率。
AI 中文摘要
Wikidata是最大的开放知识库之一,然而回答其上的复杂问题仍需要编写SPARQL查询,该查询需命名正确的实体和属性并链接它们的关系。语言模型提供了一种自然语言替代方案,但主要依赖记忆作答,这对于不太知名的实体而言可靠性最低。我们研究通过探索图谱来作答的智能体,并认为两个障碍限制了它们:缺乏记录求解器如何探索的训练数据,以及将大型图谱结果直接添加到模型上下文中的接口。我们测试了三个假设:图谱搜索的难度可以通过问题的结构而非仅通过生僻实体或措辞来控制;长程搜索中的许多失败源于检索证据的管理方式而非模型本身;在合适的环境中,开放权重模型可以媲美商业封闭模型。我们在冻结的Wikidata快照上构建多跳问题,方法是用嵌套条件替换命名实体,并在每次扩展后检查目标保持唯一且每个新条件都是必要的。我们发布了10,235条覆盖单实体和多跳问题的求解轨迹,以及生成这些轨迹的递归语言模型(RLM)框架,其中模型批量调用图谱、将结果保存在持久化Python状态中,并通过子调用解释所选证据。在100个问题上,该框架在两种接口下均提升了我们运行的两个模型的性能,相比直接调用相同函数的工具调用:gpt-6-luna从49个正确答案提升至61个,多跳准确率翻倍;Qwen3.8-27B(一个在单GPU上服务的开放权重模型)从60提升至74。
英文摘要
Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less prominent entities. We study agents that instead answer by exploring the graph, and argue that two obstacles limit them: the lack of training data recording how a solver explores, and interfaces that add large graph results directly to the model's context. We test three hypotheses: that the difficulty of graph search can be controlled through the structure of a question rather than only through obscure entities or wording; that much of the failure on long-horizon search comes from how retrieved evidence is managed rather than from the model itself; and that, in a suitable environment, open-weight models can match commercial closed ones. We construct multi-hop questions on a frozen Wikidata snapshot by replacing named entities with nested conditions, checking after each expansion that the target remains unique and that every new condition is necessary. We release 10,235 solving traces over single-entity and multi-hop questions, together with the recursive language model (RLM) harness that produced them, in which models batch graph calls, keep results in persistent Python state and interpret selected evidence through sub-calls. On 100 questions, the harness improves both models we ran under both interfaces compared with direct tool calling over the same functions: gpt-6-luna rises from 49 to 61 correct answers, doubling its multi-hop accuracy, and Qwen3.8-27B, an open-weight model served on a single GPU, from 60 to 74.
CommentsTechnical Report