AI 中文总结
研究社交媒体内容审核模式变化,探讨“策略即提示”方法在内容审核中应用,阐述其技术与治理特性、局限性及风险,提出有效提示治理考量,指出仅编写提示不适用于确保社区治理。
AI 中文摘要
内容审核实践和治理模式正在迅速变化,社交媒体公司减少了集中聘请人类审核员作为“专家”的做法,转而更关注社区方法,依靠志愿者提供准确信息并做出正确决策。在分散式审核中,社区一直依赖志愿者、更新社区指南及内部讨论。对于这两种内容审核模式,人工智能似乎有助于减轻时间、心理健康和准确性方面的审核负担。一种在内容审核中应用人工智能的可能方式是“策略即提示”方法,即将策略制定为自然语言提示并传递给大语言模型,该模型随后辅助审核任务。本文简要阐述了此方法的技术和治理特性,并认为其局限性会导致特定风险和危害,必须加以解决。为减轻这些问题,我们提出了关于更有效提示治理的多项考量,但最终发现仅编写提示并不适合确保有意义的社区治理。
英文摘要
Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.
CommentsAccepted at the Mensch und Computer (MuC) 2026 Workshop "Human-Centered Content Moderation: Expertise, Context & Evaluation"