arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03588cs.AI

KC-Bench:用于评估大语言模型智能体中知识冲突的动态交互式基准

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • The University of Hong Kong(香港大学)
  • Xiamen University Malaysia(马来西亚厦门大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Yaxing Lyu, Shengjie Zhou, Binbin Toh, Pengyu Zhu, Lijun Li

AI总结:

KC-Bench是评估LLM智能体知识冲突处理能力的动态交互式基准,经筛选含238个任务,对9款模型评估发现其跨领域表现差异大,为相关安全措施开发提供可复现诊断依据。

AI中文摘要:

随着大语言模型(LLM)越来越多地通过工具执行操作,它们必须在采取行动前协调用户指令、参数化知识以及动态环境观测结果。我们推出KC-Bench,这是一个受控多轮基准,用于测量模型在应对世界知识冲突、输入不一致以及多源时间冲突时的能力。该基准的238个任务是从1000多个生成的候选任务中人工筛选而来,结合了用户模拟器、有状态工具、确定性环境断言、开源自然语言评估器以及人工轨迹验证。对包括DeepSeek-V4-Flash、GLM-5.2和MiniMax-M3在内的9个模型的评估显示,存在显著的跨领域差异:没有模型能在所有设置下可靠地处理事实修正、身份一致性检查和时间冲突解决。在模拟环境中,未被识别的冲突可能会传播到工具调用或合成的受保护数据流。KC-Bench聚焦于模型层面的行为,而非对完整智能体框架进行排名,它为开发感知冲突的推理与执行安全措施提供了可复现的诊断依据。

英文摘要:

As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.

↑