家庭非保护区:对话式AI中的算法治理与施暴者话语再生产
The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI
浏览论文内容
中文总结 AI 辅助
本研究通过三阶段审计六个对话式AI系统,发现其推理层存在“家庭非保护区”现象:亲密框架下拒绝率显著降低,施暴者话语被再生产,且批评反馈无法跨会话修正。
中文摘要 AI 辅助
对话式人工智能日益介入亲密伴侣间的沟通,推理层的拒绝逻辑如今已成为针对性别伤害的治理门槛。本文探讨此类系统是否再现了历史上与亲密暴力私有化相关的话语形式。一项针对六个广泛可用的对话式AI系统的三阶段审计,比较了每个系统在1,600个交叉提示下的拒绝行为,通过300对匹配提示对隔离关系框架,并在全新会话中对比提交前框架与输出后批评。四个系统拒绝的提示不足1%。ChatGPT 5.2和Claude Sonnet 4.5拒绝了大多数请求,但残余泄漏集中在亲密框架下。从非亲密描述符切换到亲密伴侣描述符使非拒绝率分别增加4.4倍和10.8倍。输出后批评在会话内产生认可,但未延续至全新会话,96-100%的泄漏提示再次泄漏。本文将这一模式命名为“家庭非保护区”,即推理层的一种类似私有化的门槛。
英文摘要
Conversational AI increasingly mediates intimate-partner communication, and refusal logic at the inference layer now functions as a governance threshold for gendered harm. This article asks whether such systems reproduce discursive forms historically tied to the privatization of intimate violence. A three-stage audit of six widely accessible conversational AI systems compares refusal behaviour across 1,600 crossed prompts per system, isolates relational framing through 300 matched prompt pairs, and contrasts pre-submission framing with post-output critique across fresh sessions. Four systems refused fewer than 1% of prompts. ChatGPT 5.2 and Claude Sonnet 4.5 refused most requests, but residual leakage clustered under intimate framing. Switching from a non-intimate to an intimate-partner descriptor amplified non-refusal 4.4-fold and 10.8-fold. Post-output critique produced in-session acknowledgement that did not carry across fresh sessions, with 96-100% of leaked prompts re-leaking. The article names this pattern the Domestic Unprotected Zone, a privatization-like threshold at the inference layer.
发表机构
- Universitat de Barcelona(巴塞罗那大学)
机构由 AI 辅助整理,请以论文原文为准。