arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32983cs.HC

VocalEyes:通过对话内注册实现说话人感知的增强现实字幕

VocalEyes: Speaker-Aware Augmented Reality Captioning through In-Conversation Registration

Yuxiao Wang, Xulong Tang, Chen Chen, Rawan alghofaili

首次发表
浏览论文内容

中文总结 AI 辅助

VocalEyes通过对话内注册创建说话人档案,在AR字幕中保留说话人归属,将跟踪准确率从47.2%提升至87.3%,并降低用户工作量。

中文摘要 AI 辅助

共置的增强现实(AR)字幕使语音变得可读,但它们可能将话语与其产生者分离。在不熟悉的群体中,失去该来源会使即时回应和后续回顾变得复杂:用户不仅需要恢复说了什么,还需要知道是谁说的。传统的说话人日志返回匿名聚类,而说话人识别通常假设会议前注册。我们构建了VocalEyes,一个说话人感知的AR字幕系统,该系统从自然的自我介绍中创建具名的语音档案。该界面协调了说话人归属的字幕、固定的档案卡片以及标记说话人脸部的视觉提示。在一项有20名报告正常听力的参与者参与的受控受试者内研究中,VocalEyes以88.0%的准确率识别说话人,并将参与者的说话人跟踪准确率从仅字幕AR的47.2%提升到完整界面的87.3%。参与者还报告使用完整界面时工作量更低。这些发现表明,在脚本化的小组会议中,对话内注册如何在实时字幕和会议记录中保持说话人归属。

英文摘要

Co-located augmented reality (AR) captions make speech readable, but they can separate an utterance from the person who produced it. In unfamiliar groups, losing that source complicates immediate responses and later review: users must recover not only what was said, but also who said it. Conventional diarization returns anonymous clusters, while speaker recognition typically assumes pre-meeting enrollment. We built VocalEyes, a speaker-aware AR captioning system that creates named voice profiles from natural self-introductions. The interface coordinates speaker-attributed captions, a fixed profile card, and a visual cue that marks the articulating face. In a controlled within-subjects study with 20 participants who reported typical hearing, VocalEyes identified speakers with 88.0% accuracy and increased participant speaker-tracking accuracy from 47.2% with caption-only AR to 87.3% with the complete interface. Participants also reported lower workload with the complete interface. These findings show how in-conversation registration can preserve speaker attribution across live captions and meeting records in scripted small-group meetings.

发表机构

  • University of Texas at Dallas(达拉斯大学)
  • Florida International University(佛罗里达国际大学)

机构由 AI 辅助整理,请以论文原文为准。

↑