通过未来反馈预测实现开放式对话技能的可验证自我进化
Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
浏览论文内容
中文总结 AI 辅助
研究开放式对话技能自我进化难题,提出未来反馈技能进化方法,通过预测答案对用户后续信号的影响进行验证,转化为固定离线学习目标,在销售助手数据集上获超75%准确率,实现可重复技能进化。
中文摘要 AI 辅助
文本技能是提升固定语言模型智能体的轻量级方法,但其自我进化通常需要稳定的验证信号。在数学或代码领域,答案改变后可检查,而在开放式对话中,改变助手回复会改变用户的后续反应,导致记录的反应无法直接评估反事实回复。我们提出未来反馈技能进化方法,将自我进化从规定当前答案转向预测观察到的答案是否会导致积极或消极的后续用户信号。此预测任务可在固定记录元组上验证,支持验证门控文本优化。进化后的反馈技能捕捉可解释的响应质量标准,可作为答案技能的诊断和优化目标。在专有隐私保护销售助手数据集上,经仔细质量过滤和平衡的已解决/未解决分割,预测准确率超过75%。核心贡献是将动态对话反馈转化为固定离线学习目标,实现可重复技能进化,无需将每个候选技能投入实时交互。我们讨论了观察验证与反事实有效性的界限,并将该方法定位为离线优化阶段,而非最终人工或在线评估的替代方法。
英文摘要
Textual skills provide a lightweight way to improve frozen language-model agents, but their self-evolution normally requires a stable validation signal. Such signals are natural in mathematics or code, where an answer can be checked after it changes, yet are problematic in open-ended dialogue: changing the assistant response also changes the user's next reaction, so a logged reaction cannot directly evaluate a counterfactual response. We propose future-feedback skill evolution, which first redirects self-evolution from prescribing the current answer to predicting whether the observed answer will lead to a positive or negative subsequent user signal. This prediction task is verifiable on fixed logged tuples and therefore supports validation-gated textual optimization. The evolved feedback skill captures interpretable criteria for response quality and can subsequently serve as a diagnostic and optimization target for answer skills. On a proprietary, privacy-preserving sales-assistant dataset, careful quality filtering and a balanced resolved/unresolved split yield more than 75% prediction accuracy. Beyond this result, the central contribution is a formulation that converts otherwise moving conversational feedback into a fixed offline learning target, enabling reproducible skill evolution without placing every candidate skill in live traffic. We discuss the boundary between observational verification and counterfactual validity, and position the method as an offline optimization stage rather than a replacement for final human or online evaluation.
发表机构
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。