TPBench:用于对话压缩的转折点基准
TPBench: A Turning-Point Benchmark for Dialogue Compression
浏览论文内容
中文总结 AI 辅助
TPBench是一个针对对话压缩的转折点基准,通过三个探针评估压缩方法在保留初始目标和当前槽位值上的表现,发现现有压缩方法在联合探针上低于完整上下文。
中文摘要 AI 辅助
压缩器可以保留对话中的事实,但仍可能丢弃改变这些事实的对话轮次。用户纠正价格、逆转选择或添加约束。我们将这种失败称为转折点驱逐。单一的总体保留分数会掩盖这一问题,因为该分数混合了用户最初的需求和用户当前的需求。我们引入了TPBench,它在共享的名义保留预算下评估三个互补的信息目标。P1询问用户的初始目标。P2询问用户修订过的槽位的当前值。P3在带有后期标注槽位更新的对话中同时询问两者。当前值的答案来自MultiWOZ和SGD的人工对话状态标注。初始目标的答案是第一个用户轮次的第一句话。两者都不需要新的众包。特定探针的评估对压缩方法的排名不同。在保留比例为0.30的联合探针上,每个测试的压缩方法在使用主要Llama阅读器时均低于完整上下文。删除携带更新的对话轮次会显著降低当前值准确率,而删除一个匹配的不相关轮次则使其保持不变。Mistral阅读器重复了P2/P3排名和联合探针差距。当前值恢复在额外的语料库LongMemEval-KU和中文RiSAWOZ上进行了测试:完整上下文具有最高的准确率,而近期性在两个评估中具有最高的压缩方法平均值。
英文摘要
A compressor can keep the facts of a dialogue and still drop the turn that changed them. A user corrects a price, reverses a choice, or adds a constraint. We call this failure turning-point eviction. One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now. We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets. P1 asks for the user's initial goal. P2 asks for the current value of a slot the user revised. P3 asks for both, in dialogues with a late annotated slot update. The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD. The initial-goal answer is the first sentence of the first user turn. Neither requires new crowdsourcing. The probe-specific evaluations rank compression methods differently. On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader. Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged. A Mistral reader repeats the P2/P3 rankings and the joint-probe gap. Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ: full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations.
发表机构
- Korea Institute of Energy Technology (KENTECH)(韩国能源技术研究院)
机构由 AI 辅助整理,请以论文原文为准。