玩笑之外:测量双重含义的语义距离
Jokes Aside: Measuring the Semantic Distance of Double Meanings
浏览论文内容
中文总结 AI 辅助
本研究基于词嵌入重新审视现有幽默指标并引入对称性指标,在三个数据集上验证后发现,基于这些指标训练的模型预测幽默评分表现不佳,但对称性指标与高分笑话相关。
中文摘要 AI 辅助
大型语言模型极大丰富了计算幽默研究的工具包,尤其是在笑话与双关语的自动生成方面。一项关键创新——上下文嵌入向量,为重新审视和完善早期假设提供了新机遇。值得注意的是,Petrovic和Matthews(2013)提出了基于“我喜欢我的X,就像我喜欢我的Y,Z”这一模式的笑话生成模型(例如“我喜欢我的冰,就像我喜欢我的梦,被碾碎的”),他们提出笑话的趣味性随以下因素增加:a)Z与X、Y的频繁关联,b)Z的稀有性,c)Z的歧义性,d)X与Y之间的含义距离。在此基础上,Winters等人(2019)提出了一组基于Google Ngrams和Word2Vector的指标。本研究使用词嵌入重新审视了其中5个指标中的3个:明显度、兼容性和比较度;同时首次引入了对称性这一指标,定义为Z与X、Y的接近程度。研究使用两个模型收集嵌入向量:OpenAI text-embedding-3-small和MiniLM all-MiniLM-L6-v2,涉及三个数据集:JokeJudger、Expunations和rJokes。后两个数据集(Expunations和rJokes)通过添加成对句子进行了扩展,这些成对句子捕捉了每个笑话核心歧义表达的两种不同含义。结果显示,基于所提出指标训练的模型在预测幽默评分时表现不佳:在JokeJudger上,最佳模型达到57.1%的准确率,低于61.5%的基线;在Expunations和rJokes上的表现甚至更低。不过,对称性指标似乎始终与评分更高的笑话相关联,表明它可能捕捉到幽默的一个必要(虽非充分)属性。
英文摘要
Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to revisit and refine earlier hypotheses. Notably, Petrovic and Matthews (2013) proposed a joke generation model based on the scheme "I like my X like I like my Y, Z" (e.g. "I like my ice like I like my dreams, crushed"). They suggested that joke hilarity increases with: a) frequent association of Z with X and Y, b) rarity of Z, c) ambiguity of Z, and d) meaning distance between X and Y. Building on this, Winters et al. (2019) proposed a set of metrics, based on Google Ngrams and Word2Vector. In this work, three out of their five metrics are revisited with word embeddings: obviousness, compatibility, and comparison. Another measure, symmetry, defined as closeness of Z to both X and Y, is introduced here for the first time. Two models were used to collect the embedding vectors (OpenAI text-embedding-3-small and MiniLM all-MiniLM-L6-v2) on three datasets: JokeJudger, Expunations, and rJokes. The last two datasets, Expunations, and rJokes, were expanded by adding paired sentences that captured the ambiguous expression at the core of each joke in its two different meanings. Results revealed that models trained on the proposed metrics performed poorly in predicting humor ratings: on JokeJudger, the best model achieved 57.1% accuracy, below the 61.5% baseline, while performance on Expunations and rJokes was even lower. Nevertheless, the symmetry metric seems consistently associated with higher-rated jokes, suggesting it may capture a necessary -though not sufficient- property of humor.