arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TD-DPO:用于减轻临床自闭症干预对话中谄媚行为的差异感知偏好优化

TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

Shuzhong Lai, Junhong Lai, Chenxi Li, Qing Zhou, Haifeng Li, Gang Pan, Lin Yao, Yueming Wang

arXiv 2607.18304首次发表:更新:

发表机构

Nanhu Brain-Computer Interface Institute; MOE Frontiers Science Center for Brain and Brain-Machine Integration, Zhejiang University; College of Computer Science and Technology, Zhejiang University; Children’s Hospital Zhejiang University School of Medicine; State Key Laboratory of Brain-Machine Intelligence; Department of Neurobiology, Affiliated Mental Health Center and Hangzhou Seventh People’s Hospital, Zhejiang University School of Medicine(南湖脑机接口研究所; 浙江大学脑与脑机融合教育部前沿科学中心; 浙江大学计算机科学与技术学院; 浙江大学医学院附属儿童医院; 脑机智能技术国家重点实验室; 浙江大学医学院附属精神卫生中心和杭州市第七人民医院神经生物学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型在自闭症干预对话中的谄媚问题,提出最小编辑数据增强策略及令牌级差异直接偏好优化方法,通过加权差异令牌、降权共享令牌抑制背景漂移,经实验验证该方法能在减轻谄媚与保留干预能力间达较好平衡。

AI 中文摘要

大语言模型的谄媚行为会增加自闭症儿童干预对话中的安全风险。监督微调虽能有所减少,但仅依赖正例往往不足以识别和纠正失败模式。研究发现谄媚行为常局限于模型响应的有限跨度内,序列级偏好优化会过度更新无关令牌并降低干预能力。为此提出最小编辑数据增强(MEDA)策略构建偏好对,以及令牌级差异直接偏好优化(TD-DPO),对差异令牌加权、共享令牌降权以抑制背景漂移。多骨干和评估器的广泛实验表明,TD-DPO在减轻谄媚和保留干预能力间实现更好权衡,凸显其作为自闭症干预实用对齐方法的潜力。

英文摘要

The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often insufficient to identify and correct failure patterns. We observe that sycophancy behaviors can often be localized to a limited span within the model response. In this regime, sequence-level preference optimization can over-update preference-irrelevant tokens and degrade intervention ability. To address this, we propose the \textbf{M}inimal \textbf{E}dit \textbf{D}ata \textbf{A}ugmentation (MEDA) strategy to construct controlled, stable, minimal edit preference pairs and \textbf{T}oken-level \textbf{D}ifference \textbf{D}irect \textbf{P}reference \textbf{O}ptimization (TD-DPO), which upweights difference tokens between chosen and rejected responses while downweighting shared tokens to suppress background drift. Extensive experiments across multiple backbones and evaluators show that TD-DPO achieves a better trade-off between sycophancy mitigation and intervention ability retention in our offline settings, highlighting its potential as a practical alignment approach for autism intervention.

Commentscommited to EMNLP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑