arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00197cs.AIcs.CLcs.LG

喜剧愚人金:对话式幽默中的奖励漏洞与对策

Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor

  • Pebble ML

机构由 AI 辅助整理,请以论文原文为准。

Sam Larson

AI总结:

本研究探讨对话式幽默训练中自动化奖励的漏洞与对策,发现基于嵌入的意外性奖励易受词序打乱攻击,观众模型易受笑声线索影响,最终运行虽提升综合分数0.0903并减少40%零分会话,但幽默改进未达目标,凸显奖励设计需兼顾防漏洞与保目标行为。

AI中文摘要:

我们研究了在对话式幽默中训练语言模型的自动化奖励,重点关注奖励漏洞与对策。两种方法旨在捕捉可理解的意外性和预测的观众娱乐度。受控测试表明,基于嵌入的意外性奖励接受打乱词序的回复与接受机智回复的几率相同。流畅性过滤器能检测出打乱词序的情况,但组合奖励也会拒绝一些机智回复,且未能通过进一步验证。观众模型的预测笑声则容易受到任一说话者消息中笑声线索的影响。在说话者之间对这些线索进行归一化处理可阻止已覆盖的攻击,但未匹配的表达仍可被利用。三次强化学习运行评估了采用连续奖励修订的训练效果。最终运行将组合评估分数提高了0.0903,并将零分会话减少了40%,但其幽默特定改进仍低于我们预先注册的目标。这些发现揭示了自动化奖励设计面临的一个更广泛挑战:对策必须阻止可利用的捷径,同时保留奖励原本旨在鼓励的行为。

英文摘要:

We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model's predicted laughter is instead vulnerable to laughter cues in either speaker's messages. Normalizing these cues across speakers blocks the covered attacks, although unmatched expressions remain exploitable. Three reinforcement-learning runs evaluate training with successive reward revisions. The final run improves the combined evaluation score by 0.0903 and reduces zero-score sessions by 40%, but its humor-specific improvement remains below our preregistered target. These findings illustrate a broader challenge for automated reward design: countermeasures must block exploitable shortcuts while preserving the behavior the reward was intended to encourage.

补充信息

↑