arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19022cs.CLcs.AI

TalkMatrix:生成既一致又多样化的角色对话

TalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse

Ayuto Tsutsumi, Yuu Jinnai

首次发表
浏览论文内容

中文总结 AI 辅助

TalkMatrix通过结构化多提示多补全选择,利用四个嵌入目标及两级极小极大优化,为角色对话生成既一致又多样化的完整矩阵,实验验证其优于独立选择基线。

中文摘要 AI 辅助

基于候选的解码通常为每个提示独立选择一个补全,但许多应用需要一组满足全局、不可分解要求的输出。我们将此设置形式化为结构化多提示、多补全选择问题:给定每个提示的候选池,为每个提示选择一个补全以优化集合级目标。我们在角色对话中实例化该问题,其中每个角色应在不同情境下保持一致,每句台词应适合其情境,并且角色与情境应保持可区分。我们的方法TalkMatrix为每个角色-情境对生成多个候选,并使用四个基于嵌入的一致性和多样性目标联合选择一个完整矩阵。由于加权和可能通过牺牲某一维度来改进其他维度,TalkMatrix通过两级极小极大公式最大化表现最差的目标。我们使用多起点坐标上升法近似优化所得离散目标,并将其与局部、部分矩阵和通用组合搜索基线进行比较。我们在50个合成角色扮演场景和25个精心策划的棋盘游戏场景(其中多个角色在预定义情境中互动)上进行实验。一个LLM作为评判者,将矩阵级选择评为高于随机和独立单元格级选择基线。这些结果表明结构化选择对于全局控制的对话生成的价值,而我们的实证验证仍特定于角色扮演场景。

英文摘要

Candidate-based decoding typically selects a completion for each prompt independently, but many applications require a collection of outputs that satisfies global, non-decomposable requirements. We formulate this setting as structured multi-prompt, multi-completion selection: given a candidate pool for every prompt, select one completion per prompt to optimize a collection-level objective. We instantiate the problem in character dialogue, where each character should remain consistent across situations, each line should fit its situation, and characters and situations should remain distinguishable. Our method, TalkMatrix, generates multiple candidates for every character--situation pair and jointly selects a complete matrix using four embedding-based consistency and diversity objectives. Because a weighted sum can improve some dimensions by sacrificing another, TalkMatrix maximizes the worst-performing objective through a two-level minimax formulation. We approximately optimize the resulting discrete objective with multi-start coordinate ascent, and compare it with local, partial-matrix, and generic combinatorial search baselines. We run experiments on $50$ synthetic role-playing scenarios and $25$ curated board game scenarios where multiple characters interact in predefined situations. An LLM-as-a-judge rates matrix-level selection higher than random and independent cell-level selection baselines. These results show the value of structured selection for globally controlled dialogue generation, while our empirical validation remains specific to role-playing scenarios.

↑