arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26295cs.CLcs.AIcs.MAcs.SE

MemToC:大型语言模型中记忆-工具冲突解决的基准测试

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov, Alexander Panchenko, Ilseyar Alimova

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出MemToC基准,测试工具增强型LLM在记忆与工具返回冲突时的仲裁能力,发现指令调优模型在多数情况遵循工具,微调可改善仲裁但需兼顾工具使用与弃权表现。

中文摘要 AI 辅助

工具增强型大型语言模型(LLM)在工具返回结果与其参数化记忆存在冲突时,必须在两个易出错的来源间进行仲裁,但现有评估仅测量来源偏好,未确定来源的正确性。本文引入MemToC,这是一个用于工具返回后仲裁的受控基准,配备可执行工具。MemToC由6504个评估回合构成,这些回合基于542个经过质量控制的事实问题、独立引出的模型特定闭卷答案,以及已知正确性的受控工具返回结果构建。这些组件实例化了四种来源正确性情况;工具错误和无工具条件是单独的对照项。在五个开放权重的7B至9B模型中,工具返回结果在引出的闭卷答案中占据主导地位。四个指令调优模型在符合条件的案例中,仅在6.5%-17.1%的案例中保留了针对错误工具的已验证正确答案,在86.0%-93.1%的案例中遵循了正确工具,在两个来源均错误的案例中,有78.4%-86.0%的案例重复了工具返回结果。在问题和回合内容保持固定的情况下,三种指令措辞变体之间没有稳定的跨模型排序。我们使用ToolHop上的链级交叉拟合,将提示学习与SFT(监督微调)和DPO(直接偏好优化)进行比较,因此共享底层事实的问题永远不会跨越训练和评估。我们应用不对称成功标准:正确答案保留必须在不降低正确工具遵循的情况下得到改善。SFT和DPO在四个指令调优骨干中的两个上满足该标准。改进很少是纯粹的:20个测试的方法-模型组合中有19个在工具错误或无法回答的输入后减少了弃权(不执行)。MemToC之外的迁移是积极的但部分的,取决于模型和呈现框架。基于正确性的仲裁可以通过微调得到改善,但收益必须与正确工具使用、弃权(不执行)以及对表述的鲁棒性共同评估。

英文摘要

Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce MemToC, a controlled benchmark for post-tool-return arbitration with executable tools. MemToC comprises 6,504 evaluation episodes constructed from 542 quality-controlled factual questions, independently elicited model-specific closed-book answers, and controlled tool returns of known correctness. These components instantiate four source-correctness cases; tool-error and no-tool conditions are separate controls. Across five open-weight 7-9B models, tool returns strongly dominate elicited closed-book answers. The four instruction-tuned models retain a verified-correct answer against an incorrect tool in only 6.5-17.1% of eligible cases, follow a correct tool in 86.0-93.1%, and repeat the tool return in 78.4-86.0% of cases where both sources are wrong. No cross-model ordering remains stable across three instruction-wording variants with the question and episode content held fixed. We compare prompting with SFT and DPO using chain-level cross-fitting over ToolHop, so questions sharing an underlying fact never straddle training and evaluation. We apply an asymmetric success criterion: correct-answer retention must improve without a detected reduction in correct-tool following. SFT and DPO meet this criterion on the same two of four instruction-tuned backbones. Improvements rarely come cleanly: 19 of 20 tested method-model combinations reduce abstention after tool errors or on unanswerable inputs. Transfer beyond MemToC is positive but partial and depends on the model and presentation frame. Correctness-conditioned arbitration can be improved through fine-tuning, but gains must be evaluated jointly with correct tool use, abstention, and robustness to formulation.

发表机构

  • Central University(中央大学)
  • Ural Federal University(乌拉尔联邦大学)
  • Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究院)
  • AIRI(人工智能研究院(AIRI))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑