发表机构
Microsoft Responsible AI; Goodfire(微软负责任人工智能部门; 古德法尔公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对内容审核部署中的过滤点与响应方式,对比不同方案的有用性、有害暴露等指标,发现响应重写可平衡流量恢复与安全,支持按部署约束选择审核配置。
AI 中文摘要
内容审核分类器通常是单独评估的,但部署时需要选择干预的位置以及标记后采取的措施。我们使用两个端到端的客户结果指标而非组件准确率来评估这些选择:有用性(即显示的非有害、相关响应的对话轮次占比)和有害暴露(即显示的有害响应的对话轮次占比),延迟和错误率作为辅助诊断指标。我们在人工标注的产品基准和公开的ToxicChat评估中,比较仅输入过滤、仅响应过滤、输入+响应硬阻断三种方案。在评估的操作点下,仅响应过滤在两种设置中均实现了仅过滤方案的最高有用性,而输入+响应方案的有害暴露更低。将仅响应阻断替换为响应+重写,可恢复大部分被阻断的流量,且所选配置的观察到的有害暴露数量与仅响应阻断相同;但这种相等性并非等价结果。在可比测量结果下,探针路由相比LLM路由大幅降低了条件路由与生成时间。针对性的输出审查显示,重写通过泛化触发语言同时保留良性意图和安全重定向,平衡了过滤通过性与有用性;不过部分敏感领域的输出仍会遗漏潜在的安全相关支持信息。这些结果支持根据部署特定的安全与延迟约束来比较审核配置,而非采用通用的过滤点规则。代码和公开成果可在该https URL获取。
英文摘要
Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate these choices using two end-to-end customer-outcome metrics rather than component accuracy: Usefulness, the fraction of turns with a shown, non-harmful, relevant response, and Harmful Exposure, the fraction with a shown harmful response. Latency and error rates are diagnostics. We compare Input only, Response only, and Input + response hard blocking on a human-labelled product benchmark and public ToxicChat evaluation. At the evaluated operating points, Response only achieves the highest filter-only Usefulness in both settings, while Input + response achieves lower Harmful Exposure. Replacing Response only blocking with Response + rewrite recovers most blocked traffic and yields the same observed Harmful Exposure count as Response only blocking for the selected configuration; this equality is not an equivalence result. Probe routing substantially reduces conditional route-and-generation time relative to LLM routing at comparable measured outcomes. A focused output review shows how rewrites balance filter passage with usefulness by generalizing triggering language while retaining benign intent and safe redirection; some sensitive-domain outputs nevertheless omit potentially safety-relevant support information. These results support comparing moderation configurations under deployment-specific safety and latency constraints rather than applying a universal placement rule. Code and public artifacts are available at https://github.com/microsoft/mod-frontier