arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于表格数据的对话式任务歧义消除:感知泄漏的公式化、基准套件与训练

Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training

Nafiseh Ghoroghchian, Luis Scoccola, Tina Sedaghat, Omid Vaheb, Hannah Chen, Dino D'Agostino, Keyvan Golestan

arXiv 2610.10740首次发表:更新:

发表机构

Layer 6 AI(Layer 6 AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对表格数据的对话式任务歧义消除,提出感知泄漏的框架,构建AmbiTab基准套件,用强化学习训练提问策略,提升了歧义消除指标与任务成功度,还可诊断神谕泄漏。

AI 中文摘要

基于表格数据的对话式任务歧义消除利用对话来解决用户预期任务中缺失的信息,之后再针对表格或数据库生成解决方案。现有评估与训练缺乏感知泄漏的基础,任务成功度混合了智能体的歧义消除与解决方案生成能力,还可能反映出神谕泄漏,即用户模拟器披露了真实用户不会提供的额外信息。现有数据集也缺乏歧义与访问边界的共享表示。我们引入了歧义可验证任务的概念,该概念将歧义与解决方案形式化,将智能体分解为提问策略与解决方案策略,将环境分解为神谕与验证器。此框架提供了将任务歧义消除与解决方案生成分开评估的基线和指标、神谕泄漏的形式化定义、无评判者的泄漏诊断方法,以及针对提问策略的训练目标。我们在文本转SQL任务中实例化该框架,构建了AmbiTab基准套件,该套件在指定智能体、神谕和验证器可访问内容的统一表示下整合了六个歧义数据集。我们评估了澄清策略与神谕泄漏,并使用强化学习训练提问策略。训练后的提问者在所有六个数据集上提升了我们的歧义消除指标,在五个数据集上提升了任务成功度,且我们的泄漏诊断方法可衡量训练对神谕泄漏的影响。

英文摘要

Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.

Comments39 pages (9 main, 30 appendix), 12 figures (5 main, 7 appendix)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑