过滤攻击性内容会改变其可见性,但不会改变用户行为:在Nextdoor上针对20万用户进行的两项随机对照试验
Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor
浏览论文内容
中文总结 AI 辅助
研究在Nextdoor平台针对20万用户进行两项随机对照试验,探究减少攻击性内容可见性的干预措施效果。通过不同过滤机制,虽大幅降低了攻击性内容可见性,但未改变用户行为,凸显了在线平台管理的复杂性。
中文摘要 AI 辅助
我们研究了在本地社交平台Nextdoor上减少攻击性内容可见性的干预措施的有效性。内容过滤(隐藏或降低违反平台规则但未明确违规的攻击性内容的排名)几乎在每个主要平台都有应用,但几乎没有关于其是否改变用户行为的实地证据。我们报告了两项大规模随机对照试验,每项试验涉及10万用户。研究1(2022年)测试了应用于帖子评论的举报触发过滤器,攻击性评论的浏览量适度降低了12%;在另外十一项平台行为指标上未发现显著影响。研究2(2023 - 2024年)通过在创建时用谷歌拼图的Perspective API对帖子和评论进行主动评分并从信息流中过滤,弥补了研究1的核心局限。这使得攻击性帖子的浏览量几乎完全(95%)降低,但在另外十三项指标上仍未发现显著影响。在跨越不同内容类型、过滤机制、分类器和国家的两项独立试验中,尽管过滤强度从12%提高到95%,但过滤可靠地降低了攻击性内容的可见性,而未改变平台访问量、内容消费或内容生产。这些一致的零结果为一种普遍存在的干预措施提供了罕见的实地证据,并凸显了有效管理在线平台的复杂性。
英文摘要
We investigate the effectiveness of interventions that reduce the visibility of offensive content on the local social platform Nextdoor. Content filtering -- hiding or downranking offensive content that brushes against a platform's rules without clearly breaking them -- is deployed across virtually every major platform, yet almost no field evidence exists on whether it changes user behavior. We report two large-scale randomized controlled trials, each involving 100,000 users. Study 1 (2022) tested a report-triggered filter applied to comments in post threads and produced a modest 12% reduction in views of offensive comments; across eleven further measures of platform behavior we found no significant effects. Study 2 (2023-2024) remedied Study 1's central limitation -- a weak manipulation driven by slow, report-based eligibility -- by proactively scoring posts and comments at creation with Google Jigsaw's Perspective API and filtering them from the newsfeed. This produced a near-complete (95%) reduction in views of offensive posts, yet across thirteen further measures we again found no significant effects. Across two independent trials spanning different content types, filtering mechanisms, classifiers, and countries -- and despite manipulation strength rising from 12% to 95% -- filtering reliably reduced the visibility of offensive content without altering platform visitation, content consumption, or content production. These convergent null results provide rare field evidence on a ubiquitous intervention and underscore the complexity of effectively moderating online platforms.