arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18822cs.CY

控制理论视角下的内容审核

Control-Theoretic Content Moderation

Benedetta Tessa, Serena Tardelli, Marco Avvenuti, Anna Monreale, Stefano Cresci

AI总结:

本文提出将内容审核建模为全局自适应序贯决策问题,引入控制理论框架组合干预措施,模拟实验表明其优于基线策略,能更好平衡目标并有效应对有害性激增。

AI中文摘要:

大量文献在局部层面研究内容审核,即在单个审核决策的层面上,例如通过测量或预测特定干预措施的效果。然而,关于如何将这些决策组合成有效的平台级审核策略的问题,相对而言尚未得到充分探索。我们通过将内容审核表述为一个全局、自适应且序贯的决策过程来解决后一个问题,在该过程中,异构干预措施必须共同平衡多个相互竞争的目标。借鉴反馈控制理论,我们引入了一个通用的控制理论框架,用于根据审核行动对不断演化的平台的预期效果来组合这些行动。我们在大规模、基于经验模拟的环境中实例化了该框架,并将两种控制理论审核器与若干基线策略和局部策略进行了比较。当审核旨在将相互竞争的平台级属性维持在期望状态附近时,控制理论方法取得了最佳的整体性能。它们还更有选择性地使用严厉干预措施,并在外部有害性激增后能更有效地恢复。这些结果证明了将内容审核视为一个全局、自适应且序贯的决策问题的优势。

英文摘要:

A sizable literature studies content moderation locally, at the level of individual moderation decisions, for example by measuring or predicting the effects of specific interventions. However, the problem of how such decisions should be combined into effective platform-level moderation strategies is comparatively unexplored. We address this latter problem by formulating content moderation as a global, adaptive, and sequential decision process in which heterogeneous interventions must jointly balance multiple competing objectives. Drawing on feedback control, we introduce a general control-theoretic framework for composing moderation actions according to their expected effects on an evolving platform. We instantiate the framework in large-scale, empirically grounded simulations and compare two control-theoretic moderators against several baselines and local strategies. When moderation aims to maintain competing platform-level properties around desired conditions, the control-theoretic approaches achieve the best overall performance. They also use severe interventions more selectively and recover more effectively after external surges of harmfulness. These results demonstrate the advantages of studying content moderation as a global, adaptive, and sequential decision problem.

↑