发表机构
University of Minnesota, Twin Cities(明尼苏达大学双城分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PragAlign框架,采用生成-评估-修正循环实现受控合成对话生成,在800组对话规格上获99.50%评估接受率,可有效提升多交际约束满足度,但情感实现等仍为未解决挑战。
AI 中文摘要
合成对话生成可支持隐私受限服务场景的研究,但生成对话必须保留交际意图、情感意义与自然对话流。本文提出PragAlign,一种面向受控合成对话生成的反馈引导框架,其以服务上下文、目标意图与目标情感为条件,辅以特质风格控制。PragAlign采用“生成-评估-修正”循环:基于大语言模型(LLM)的评估器对意图对齐度、情感对齐度、连贯性、流畅性及综合质量打分,随后提供针对各准则的反馈,最多支持三轮优化。在800组匹配的对话规格上,PragAlign获得评估器定义的99.50%接受率,而一次性生成的接受率为72.25%,无结构化反馈的重复生成接受率为95.88%。这表明重复尝试是一次性生成性能提升的主要原因,而结构化反馈主要改善最后一英里的多约束满足度,而非整体平均质量。优化增益集中于情感对齐,这也是消融研究中的主要失败模式。对1200个生成对话的独立人工评估显示,标注者可高度识别意图表达与对话流,而情感恰当性则较不稳定且更具主观性。这些结果表明,PragAlign是提升评估器定义的交际约束满足度的质量控制框架,同时也说明情感实现与独立的人工感知质量仍是未解决的挑战。
英文摘要
Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a feedback-guided framework for controlled synthetic dialogue generation conditioned on service context, target intent, and target emotion, with auxiliary trait-style controls. PragAlign uses a generate--evaluate--revise loop in which an LLM-based evaluator scores intent alignment, emotion alignment, coherence, fluency, and aggregate quality, then provides criterion-specific feedback for up to three refinement rounds. On 800 matched dialogue specifications, PragAlign achieves 99.50\% evaluator-defined acceptance, compared with 72.25\% for one-shot generation and 95.88\% for repeated generation without structured feedback. This indicates that repeated attempts account for much of the gain over one-shot generation, while structured feedback primarily improves last-mile multi-constraint satisfaction rather than broad average quality. Refinement gains are concentrated in emotion alignment, which is also the dominant failure mode in ablations. A separate human evaluation of 1,200 generated dialogues shows that intent expression and dialogue flow are highly recognizable to annotators, while emotion appropriateness is less stable and more subjective. These results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.