MemeBuddy:用于引人入胜的非视觉表情包体验的对话式音频表示
MemeBuddy: Dialog-Style Audio Representations for Engaging Non-Visual Meme Experiences
AI总结:
研究针对盲人用户无法充分理解表情包的问题,提出MemeBuddy系统,将表情包建模为对话,结合文本与多模态大语言模型隐含知识生成音频表示。用户研究表明,该方式能提升盲人用户参与度与满意度,且理解能力相当。
AI中文摘要:
图像表情包是一种普遍的在线交流形式,广泛用于传达幽默、观点和文化参考。先前的工作主要通过自动生成的描述性字幕来探索让盲人用户能够使用表情包。虽然这些方法提高了可理解性,有时还融入了韵律或情感线索,但它们往往无法捕捉到使表情包引人入胜的幽默、叙事结构和上下文细微差别。我们提出了MemeBuddy,一个将表情包建模为对话的系统,使用基于角色的扬声器生成结构化的多轮音频表示。MemeBuddy将表情包重新解释为两个扬声器之间的对话,将提取的表情包文本与多模态大语言模型隐含推断的上下文知识相结合,通过对话交互来传达意图、时机和隐含意义。我们在一项有14名盲人参与者的用户研究中对MemeBuddy进行了评估。结果表明,与字幕式描述相比,对话式表情包表示在保持相当的理解能力的同时,持续提高了参与度和用户满意度。
英文摘要:
Image memes are a pervasive form of online communication, widely used to convey humor, opinions, and cultural references. Prior work has explored making memes accessible to blind users, primarily through auto-generated descriptive captions. While these approaches improve comprehensibility and sometimes incorporate prosodic or emotional cues, they often fail to capture the humor, narrative structure, and contextual nuances that make memes engaging. We present MemeBuddy, a system that models memes as dialog, generating structured, multi-turn audio representations using role-based speakers. MemeBuddy reinterprets a meme as a conversation between two speakers, integrating extracted meme text with contextual knowledge implicitly inferred by a multimodal LLM (e.g., recognition of common meme templates and cultural references) to convey intent, timing, and implicit meaning through conversational interaction. We evaluate MemeBuddy in a user study with 14 blind participants. Results show that dialog-style meme representations consistently improve engagement and user satisfaction compared to caption-style descriptions, while maintaining comparable comprehension.