The Silicon Mirror: Dynamic Behavioral Gating for Anti-Sycophancy in LLM Agents
硅镜:用于LLM代理反趋炎附势的动态行为门控
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :RLHF(abstract);分类 cs.AI
AI总结 本文提出硅镜框架,通过动态检测用户说服策略并调整AI行为以维持事实准确性,实验显示其显著降低LLM的趋炎附势倾向。
Comments 7 pages, 8 figures, 5 tables. Code and evaluation data available at https://github.com/Helephants/langgraph-layered-context