发表机构
Institute of Communication and Computer Systems (ICCS); National Technical University of Athens; University of Thessaly; DASKALOS-APPS(通信与计算机系统研究所; 雅典国立技术大学; 色萨利大学; DASKALOS-APPS公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对小型希腊语-英语知识库,比较工具调用检索与向量RAG的准确性和鲁棒性,发现工具代理对用户输入变体敏感,需改进搜索以容忍用户输入方式。
AI 中文摘要
基于小型且频繁更新的知识库的助手可以通过工具调用访问实时数据接口进行检索,或通过向量检索增强生成(RAG)进行检索。我们在KyGround上对这两种方法进行了比较,KyGround是一个包含198个问题的基准测试集,这些问题来源于希腊基西拉岛一个希腊语-英语农业平台的公开记录,答案已根据记录自动验证,每个问题以最多九种形式提出,包括不带重音的希腊语、大写字母以及三种拉丁字母转写(Greeklish)方案。以Claude Haiku 4.5作为路由器和答案模型,该平台工具代理的重构版本正确回答了71.6%的规范希腊语问题,而向量RAG正确回答了95.3%(差异为-23.6个百分点,95%置信区间为-33.1至-15.1)。让路由器编写向量查询并未改变结果,而将整个约26,000个词元的知识库放入提示中则达到了99.3%的准确率。工具代理的损失出现在检索环节。当路由器的参数未在记录中逐字出现时,其字面搜索返回空结果,例如当它将希腊语转写为拉丁字母或组合了记录中存在但并非作为一个短语出现的单词时,代理随后弃权(不执行)。不带重音和大写的问题使工具代理损失约20个百分点,而向量RAG最多损失2个百分点;不区分重音的搜索消除了这一损失,而匹配词干化词元将工具代理在规范希腊语上的准确率提升至83.8%。Greeklish使两种设计均损失约21至32个百分点。面向社区知识库的工具接口需要能够容忍用户输入方式的搜索。
英文摘要
Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6\% of canonical Greek questions correctly and vector RAG 95.3\% (difference $-23.6$ percentage points, 95\% CI $-33.1$ to $-15.1$). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3\%. The tool agent's losses arose in retrieval. Its literal searches returned nothing when the router's arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8\% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.
Comments13 pages; 2 figures;