迈向参与式语音数据集策展:一项酷儿案例研究与概念框架
Towards participatory speech dataset curation: A queer case study and conceptual framework
- University of Calgary(卡尔加里大学)
- Meta FAIR
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过酷儿社区案例,提出一个由边缘化社区主导的参与式语音数据策展概念框架,包含社区定义、项目制定、参与模式和个人自主性四个重叠双向过程。
AI中文摘要:
在本文中,我们通过LGBTQIA+(即酷儿)社区的案例研究,论证了建立参与式语音数据集创建框架的必要性——该社区对人工智能有明确的担忧,并遭受过相关伤害,包括有人试图开发所谓的“同性恋雷达”技术,据称能识别个体的酷儿身份。我们回顾了常见的语音数据收集实践,探讨了这些方法为何可能不适合与酷儿说话者互动,并讨论了先前在酷儿社区参与下进行的参与式人工智能努力,以及针对其他边缘化社区进行语音数据收集的参与式尝试。基于这一回顾,我们借鉴协同设计与知识共享的见解,为边缘化社区构建了一个由他们主导、为他们服务、与他们共同进行的参与式语音数据策展概念框架。我们提出了一个由社区定义、项目制定、参与模式和个人自主性等重叠且双向的过程组成的框架。
英文摘要:
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.