发表机构
ITMO University(ITMO大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出多轮不可回答性基准与评估框架,验证潜在欠指定信号的可迁移性,并表明基于校准探测器的门控能精准触发,但端任务提升有限,差距主要在澄清信息的使用。
AI 中文摘要
大型语言模型通常会回答那些无法从给定信息中回答的问题,并且在对话中,它们在信息尚未充分时便过早作答。不可回答性可从隐藏状态中线性解码,但尚不清楚其各种形式是否共享同一表征,以及该信号在对话中是否有用。我们构建了一个带轮次标注的多轮基准(423段对话,1,661个标注的轮次状态)和一个带有模拟用户的评估框架,该模拟用户会回答澄清性问题,并使用六个数据集和六个开放权重的大型语言模型来测试不可回答性探测器的迁移能力。探测器在共享不可回答性基础的数据库之间稳健迁移:数学中缺失信息(AUROC 0.77-0.97)和段落中缺失信息(SQuAD 2.0<->MuSiQue,0.77-0.90)。针对认知性“已知未知”的探测器迁移到数学领域时表现不佳,但这种分离在词汇控制下会减弱,并随层和坐标系变化,因此该问题仍未解决。单轮探测器在零样本情况下无法检测对话何时变得可回答;结构内探测器能够恢复该能力,但并未优于词袋分类器。一个基于校准探测器的门控,无需模型微调,在欠指定轮次上的触发精确度远高于随机水平,其最终任务成功率与使用真实标签的门控相差在0.08以内。然而,在四个模型中,它并未可靠地优于普通生成或提示式整合。剩余的差距主要在于模型如何使用澄清信息,而非检测本身。
英文摘要
Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
Comments15 pages, 3 figures, 10 tables. Under review