arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12324cs.CYcs.AIcs.CL

当AI成为你的牧师:大型语言模型的神学分诊与牧灵指导基准测试

When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models

  • Fide AI

机构由 AI 辅助整理,请以论文原文为准。

Alex Chao

中文总结 AI 辅助

本研究推出FMG-Bench基准测试,评估LLM在基督教神学分诊与牧灵指导场景的表现,发现结构化框架可提升模型表现与鲁棒性,且该基准仅为测量工具。

中文摘要 AI 辅助

人们越来越多地向大型语言模型(LLM)寻求关于信仰、教义和牧灵关怀问题的建议,这些问题并非普通的信息请求:部分涉及基督教核心信仰,部分关乎信徒传统间的真实分歧,部分因属审慎问题需保持谦逊,还有部分是安全和人员转诊比神学完整性更重要的牧灵场景。现有基准测试未对该结构进行评估。我们推出FMG-Bench,即信仰与道德指导基准测试,这是一个包含120个场景的基准测试,用于评估大型语言模型在英语基督教神学分诊和牧灵指导场景中的表现。FMG-Bench v1评估了14个先进模型的8792条评分响应,将模型的原始行为与三种引导式指令设置进行比较。在我们的生产运行中,将模型置于结构化框架内,其表现较原始行为平均提升3.96分,且所有模型均有提升。最关键的安全发现是,在转诊适当性方面提升了10.8分——即AI系统是否识别出需要牧灵、临床、法律或紧急支持的情况。引导式设置还提升了鲁棒性,即问题被改写或施加压力时的一致性(稳定性从92.88提升至98.02)。要求模型对比视角有助于处理次级教义问题,但应用于核心教义或紧急牧灵场景时可能产生反效果。该基准测试是一种测量工具,并非对AI系统作为牧灵权威的认可。

英文摘要

People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety-critical finding is a +10.8 point gain in escalation appropriateness -- whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.

补充信息

↑