arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30322cs.AIcs.CL

无知还是无能?构建面向大语言模型智能体的知识门控可验证任务

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang, Hongquan Zhu, Qiufei Hu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型智能体任务缺乏对约定访问控制的问题,提出知识门控任务构建协议,经实验验证其有效性并发布部分任务套件与工具。

中文摘要 AI 辅助

专业智能体任务往往依赖于公开语料库中未包含的约定,但基准测试很少控制智能体是否能访问这些约定。我们提出一种知识门控任务构建协议,该协议将任务指令与包含私有约定、参考表和实用算子的紧凑人工制品分离。构建时的来源、提供人工制品与 withheld-artefact 条件下字节完全相同的任务指令、泄露审计及可执行见证,使对人工制品的依赖明确且可测试。在15个校准任务中,某前沿智能体配置在有人工制品时通过率达68.0%,无人工制品时为0%;在其中1个任务中,看似合理但不正确的人工制品在5次试验中也产生0%的结果。确定性求解器和规则语料库为结构化任务提供精确真值,而命名的标准级 rubric 支持无法由单个可执行预言机检查的输出。配置相对校准筛选保留了7个满足我们5次试验经验知识门控筛选的任务。这些实验验证了构建协议的行为,但未证实保留的任务在训练后会得到改进。我们在该 https URL 公开发布部分任务套件及支持工具。

英文摘要

Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators. Construction-time provenance, byte-identical task instructions across the provided- and withheld-artefact conditions, leak audits, and executable witnesses make dependence on the artefact explicit and testable. Across fifteen calibration tasks, one frontier agent configuration achieves a 68.0% pass rate with the artefact and 0% without it; on one task, a plausible but incorrect artefact also yields 0% across five trials. Deterministic solvers and rule corpora provide exact ground truth for structured tasks, while named criterion-level rubrics support outputs that cannot be checked by a single executable oracle. A configuration-relative calibration screen retains seven tasks satisfying our five-trial empirical knowledge-gating screen. These experiments validate the behavior of the construction protocol; they do not establish that the retained tasks improve post-training. We publicly release part of the task suite and supporting tooling at https://github.com/DatagridsAI/Knowledge-Gated-Task-Construction.

发表机构

  • DataGrids
  • Shanghai University(上海大学)
  • Peking University(北京大学)
  • Northwestern Polytechnical University(西北工业大学)
  • Nanyang Technological University(南洋理工大学)
  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑