arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Wontopos Tablet 2:无需词汇匹配的多语言多模态记忆检索测量

Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching

Sunwoo Kim

arXiv 2608.23920首次发表:更新:

发表机构

Wontopos L.L.C.(万托波斯有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文测量Wontopos Tablet 2的多语言多模态记忆检索性能,其无需词汇匹配,在文本、跨语言及多模态基准上表现优于BM25等方法,但存在低资源语言性能下降等问题。

AI 中文摘要

我们对Wontopos Tablet 2(一种用于语言模型的生产级长期记忆引擎)进行了测量,测量涵盖该领域已使用的文本基准,以及完全无文本存储照片的跨语言检索。其检索路径不包含词汇匹配、关键词评分,也不包含自身的语言模型。在LongMemEval-S(500个问题)上,它的得分为95.7%[93.4, 97.1];在BEAM-1M(700个问题,221万条存储记忆)上,得分为67.5%[64.8, 70.2]。这些是问题采样区间,而非运行间波动,运行间波动的幅度比该区间小一个数量级。本文大部分内容围绕这些指标单独使用时的局限性展开:在保持引擎、语料库、设置和评估标准固定的情况下,仅改变读取器,LongMemEval-S的得分会变动2.0个百分点;仅改变重问预算,BEAM-1M的得分会变动8.9个百分点。这些变动均未在我们对比的相关报告中提及,且后者的变动幅度超过了多数报告中的得分差距,因此我们仅提供该得分表作为定位参考,而非排名。针对多模态维度,我们设置了两项对照实验:与尽可能优化配置的BM25相比,在70个存储-查询语言单元中,我们的方法达到了95.2%的平均召回率@5,而BM25仅达到19.0%,且在54个单元中得分为0;对于无标题照片,词汇方法根本没有可评分的文档。在300张来自Crossmodal-3600的无标题照片(覆盖14种语言)上,开放密集基线实验显示,密集性并未带来语言独立性:某基线使用相同图像向量时,英语得分91.0%,俄语仅得4.7%;某多语言变体在泰卢固语和斯瓦希语上表现崩溃。我们的跨语言得分波动为14.0,而对应基线的波动为27.5和27.7。针对我们的方法,有三项需注意的结果(按同等权重报告):低资源语言的性能大幅下降(斯瓦希语53.0%,泰卢固语64.0%);附加标题会使跨语言检索得分降低11.4个百分点;我们自身检索流程某一阶段的设置缺失,导致韩语的top-1准确率损失37个百分点,却未影响其他9种语言。

英文摘要

We measure tablet-2, a production long-term memory engine for language models, on the text benchmarks the field already uses and on cross-lingual retrieval of photographs stored with no text at all. Its retrieval path contains no lexical matching, no keyword scoring, and no language model of its own. On LongMemEval-S (500 questions) it scores 95.7% [93.4, 97.1]; on BEAM-1M (700 questions, 2.21M stored memories) 67.5% [64.8, 70.2]. Those are question-sampling intervals, not the run-to-run spread, which is an order of magnitude narrower. Most of the paper is about how little they mean alone. Holding engine, corpus, settings and judge fixed, changing only the reader moves LongMemEval-S by 2.0 points; changing only the re-ask budget moves BEAM-1M by 8.9. Neither is stated in the reports we compare against, and the second exceeds most gaps there, so we give that table as a placement and not a ranking. For the multimodal axis we run two controls. Against BM25, configured as strongly as we could, we reach 95.2% mean recall@5 over 70 store-and-query language cells where BM25 reaches 19.0% and is exactly zero in 54. On captionless photographs a lexical method has no document to score at all. Open dense baselines on 300 Crossmodal-3600 photographs in 14 languages show that density confers no language independence: one scores 91.0% on English and 4.7% on Russian from identical image vectors, and a multilingual variant collapses on Telugu and Swahili. Our spread across languages is 14.0 against their 27.5 and 27.7. Three results run against us and are reported at equal weight: low-resource languages degrade sharply (Swahili 53.0%, Telugu 64.0%), attaching captions lowers cross-lingual retrieval by 11.4 points, and one setting omitted into one stage of our own retrieval cost 37 points of Korean top-1 accuracy while leaving nine languages untouched.

Comments42 pages, 8 figures. Harness and per-question records released

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑