TWIST:一个针对对话记忆中干预质量提出的基准,并带有经人工验证的草稿对齐
TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
浏览论文内容
中文总结 AI 辅助
TWIST基准评估对话记忆系统在信念变化点的干预质量,通过四轨道和硬负样本控制,经人工验证,揭示系统在矛盾检测与避免过度标记间的权衡。
中文摘要 AI 辅助
长对话记忆基准日益测试回忆和提示性知识更新,近期工作则研究不断演变的用户信念和记忆状态。TWIST是一个针对互补且未测量属性的基准套件:干预质量——即部署的记忆系统,通过其自身的摄取/回忆/审查界面运作,在信念变化点上是否行为正确。四个轨道涵盖无提示的紧张检测、根据记录审查出站草稿、在保留取代历史的同时用当前信念回答,以及管理敏感回忆。该套件扩展了LoCoMo的语料库和框架,将每个检测/阻止指标与匹配的不过度检测控制配对:表面匹配的硬负样本为虚假干预定价,因此任何轨道都不能通过标记所有内容来作弊。基准本身首先得到验证:独立、金标准盲双重标注并附有裁决、法官诱饵校准和可分离性审计。在经人工验证的轨道B v1.0关键集(161项,裁决后kappa=0.85)上,没有测试配置同时实现高矛盾回忆、高硬负样本特异性和高归因:扁平RAG基线检测到0.76-0.97的真实矛盾,但根据后端不同,错误标记了16-43%的表面匹配安全草稿,而部署的面向连贯性系统几乎从不过度标记(0.98-1.00特异性),却捕获了42%的真实矛盾——这是仅回忆分数无法看到的权衡。一个13配置基线阶梯定位了原因:每个金标准矛盾仅凭其证据即可检测(回忆1.000),校准模型在给定完整转录时几乎解决该轨道——与大量检索覆盖缺口一致——而仅草稿的下限揭示了模型相关的风格先验。系统的TWIST概况,连同其回忆分数,衡量记忆是否知道何时干预以及何时不干预。
英文摘要
Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo's corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything. The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions -- a trade-off no recall-only score can see. A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript -- consistent with substantial retrieval-coverage gaps -- and draft-only floors reveal model-dependent style priors. A system's TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.
发表机构
- MindTwin
机构由 AI 辅助整理,请以论文原文为准。