发表机构
University of Bucharest(布加勒斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种可工程化的人工偏袒机制,通过关怀与特殊性两个组件在对话智能体中实现持久关系,并验证其行为特征,为未来感知研究提供候选方案。
AI 中文摘要
我们研究了一种在对话智能体中实现人工偏袒的行为机制。本文不报告人类受试者数据,也不对依恋、信任、感知心智或机器内在性作出任何声明。偏袒包含两个组成部分:关怀,定义为在保证的每用户服务底线之上分配有限的交互盈余;以及特殊性,定义为在固定推理机制上根据关系历史累积的每用户状态估计。我们将两者实现为基于一个开放权重基础模型的持久化层,并在具有已知隐藏状态调度的构造多会话关系上进行评估。四通道差异比较在相同探针上使用配对生成种子对比已知用户和陌生用户条件;陌生用户提示未进行长度匹配,且一个通道(主动性)是分配器输出而非生成行为。完整机制再现了解析指定曲线的序数形状,但未再现其幅度,并且是唯一在校准电池上满足协议定义的联合特征的方案。初始机制因关系内容降低主题相关性而未通过探针质量等价性检验。经过防护的修订版本在校准电池上满足等价性,但在未用于校准的第二个电池上未满足,其得分高于未修改基线。在该第二个电池上,特异性边际为0.0198,低于0.02的协议阈值。配对平均右用户-错误用户估计准确度对比为0.375,但并非每次运行均为正。因此,该机制是后续感知研究的候选方案,而非用户感知到爱的证据。版本化协议记录未经独立时间戳标记,且未被描述为预注册。伦理约束包括陌生用户待遇底线、有限盈余、披露和去强化化。
英文摘要
We study a behavioral mechanism for artificial partiality in conversational agents. The paper reports no human-subjects data and makes no claim about attachment, trust, perceived mind, or machine interiority. Partiality has two components: caring, defined as allocation of a finite interaction surplus above a guaranteed per-user service floor, and particularity, defined as a per-user state estimate accumulated from relationship history on fixed inference machinery. We implement both as a persistence layer over one open-weight base model and evaluate them on constructed multi-session relationships with known hidden-state schedules. A four-channel divergence compares known-user and stranger conditions on identical probes with paired generation seeds; the stranger prompt is not length-matched, and one channel (initiative) is an allocator output rather than generated behavior. The full mechanism reproduces the ordinal shape, but not the magnitude, of an analytically specified curve and is the only arm satisfying the protocol-defined joint signature on the calibrated battery. The initial mechanism fails probe-quality equivalence because relationship content reduces topical relevance. A guarded revision satisfies equivalence on the calibrated battery but not on a second battery not used in calibration, where it scores higher than the unmodified baseline. On that second battery, the specificity margin is 0.0198, below the 0.02 protocol threshold. The paired mean right-user--wrong-user estimation-accuracy contrast is 0.375 but is not positive in every run. The mechanism is therefore a candidate for a later perception study, not evidence that users perceive love. The versioned protocol record is not independently time-stamped and is not described as a preregistration. Ethical constraints include a stranger-treatment floor, a bounded surplus, disclosure, and de-intensification.
Comments16 pages, 1 figure, 4 tables. Companion framework paper: arXiv:2607.15883 Code and data archived at doi:10.5281/zenodo.21462976