arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24532cs.HC

对抗人格漂移的提示:比较LLM模拟对话中的干预时机与内容

Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations

发表机构柏林祖斯研究所 · 柏林工业大学 · 韦曾鲍姆研究所
另 1 家 · 查看机构详情
  • Zuse Institute Berlin(柏林祖斯研究所)
  • TU Berlin(柏林工业大学)
  • Weizenbaum Institute(韦曾鲍姆研究所)
  • HU Berlin(柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Nicolas Leins, Jennifer Haase, Varvara Geronimus, Jana Gonnermann-Müller, Sebastian Pokutta

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估五种提示级干预机制对抗LLM模拟学生人格漂移的效果,发现行为特定指令降漂移率87%,但自适应时机未优于静态调度。

中文摘要 AI 辅助

使用大语言模型(LLM)模拟学生人格能够实现对教育系统的可扩展评估。然而,行为漂移,即人格一致性的逐渐下降,可能在长时间对话中出现,从而限制此类模拟的有效性。我们评估了五种提示级机制,采用独立的监控和干预流程。在跨越四个LLM和两种ADHD人格强度的1200次28轮对话中,我们变化了干预时机(静态与自适应)和注入内容(重新注入与反思性提醒),并新增了一种自适应条件,其中监控器生成行为特定指令。相对于无干预,重新注入将LLM评定的漂移率降低了35%至38%,反思性提醒降低了22%至27%,而行为特定指令降低了87%。但没有任何方法完全消除漂移。我们没有发现自适应时机优于静态调度的证据。因此,监控似乎更有用于决定纠正“什么”而非“何时”干预,尽管行为特定指令需要组件级测试。

英文摘要

Simulating student personas with large language models (LLMs) enables scalable evaluation of educational systems. However, behavioral drift, a progressive decline in persona consistency, can emerge over extended conversations, limiting the validity of such simulations. We evaluate five prompt-level mechanisms using separate monitoring and intervention pipelines. Across 1,200 28-turn conversations spanning four LLMs and two ADHD persona intensities, we varied when to intervene (static vs. adaptive) and what to inject (reinjection vs. reflective reminder), plus a novel adaptive condition in which a monitor generates behavior-specific instructions. Relative to no intervention, reinjection reduced the modeled rate of LLM-rated drift by 35--38\%, reflective reminders by 22--27\%, and behavior-specific instruction by 87\%. None eliminated drift. We found no evidence that adaptive timing outperformed static scheduling. Monitoring therefore appears more useful for deciding \textit{what} to correct than \textit{when} to intervene, although behavior-specific instruction requires component-level testing.

↑