你懂我意思:智能体对话指代消歧基准
You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding
浏览论文内容
中文总结 AI 辅助
本研究构建了基于GitHub代码仓库的RepoRef基准,针对对话指代消歧任务,发现现有最优智能体仅达67%成功率,为多工具环境下智能体信息处理研究提供了基准。
中文摘要 AI 辅助
协作对话中常包含指代对象为间接提及而非明确命名的表述,如解析“这看起来是昨天讨论的修复方案”需要结合对话上下文与可通过API或用户界面访问的周边工作区证据。我们将该问题形式化为对话指代消歧(Conversational Reference Grounding, CoRG),即利用给定工具集将对话中的指代解析为说话者意图的唯一外部实体。CoRG具有挑战性,因为它结合了分布在对话和外部工作区中的词汇、语义和时间线索,智能体需将这些异构信号转化为有效的工具使用:制定策略、发现合理候选、检查其元数据和内容、排除相近替代项。我们通过RepoRef研究CoRG,这是一个包含400个开发者对话片段的基准,基于92个代码仓库的GitHub议题、拉取请求和提交构建。与单步检索任务不同,RepoRef常需要多步工具使用。我们的结果显示,当前智能体在CoRG任务上仍具挑战性,即使是最优智能体也仅达到67.0%的成功率,仍有三分之一的指代无法解析。这些发现表明,CoRG是研究智能体在现实多工具环境中如何搜索、检查和验证信息的具体基准。
英文摘要
Collaborative conversations frequently contain references whose targets are indirect rather than named: resolving "this looks like the fix discussed yesterday" requires combining conversational context with evidence from the surrounding workspace which is accessible through APIs or user interfaces. We formalize this problem as Conversational Reference Grounding (CoRG): using a given set of tools to resolve a reference in conversation to the unique external item intended by the speaker. CoRG is challenging because it combines lexical, semantic, and temporal cues distributed across the conversation and the external workspace. Agents must translate these heterogeneous signals into effective tool use: formulating strategies, discovering plausible candidates, inspecting their metadata and content, and ruling out close alternatives. We study CoRG through RepoRef, a benchmark of 400 developer-chat segments grounded in GitHub issues, pull requests, and commits across 92 repositories. Unlike single-shot retrieval tasks, RepoRef often requires multi-step tool use. Our results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved. These findings position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.
发表机构
- Bar-Ilan University(巴伊兰大学)
- Allen Institute for AI(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。