arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向沉思型大语言模型:心理健康领域中评估与增强大语言模型对齐性的模块化框架

Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

Asher Sprigler, Yang-Yang Feng, Iftach Amir, Jonathan E. Bogard, Todd S Braver, Yi Ding, David Kinney, Yixue Zhao

arXiv 2607.10871首次发表:更新:

发表机构

Purdue University; Washington University in St. Louis; Yixue Research Institute(普渡大学; 圣路易斯华盛顿大学; 易学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对新模型等涌现使评估沉思原则增强大语言模型对齐性具挑战的问题,提出模块化可扩展评估框架,能集成新元素并重现先进结果,支持交叉评估,其提示模块便于纳入伦理视角,为跨学科研究及有益人机生态系统奠定基础。

AI 中文摘要

沉思传统长期以来指导着道德行为和亲社会互动,近期研究表明沉思原则(如正念、慈悲、非二元推理)可能为大语言模型(LLMs)对齐提供有前景的范式,以改善合作并减少伦理违规。但随着新模型、评估指标和基准迅速涌现,系统评估沉思原则在不同场景下是否以及如何增强LLMs对齐性仍具挑战,现有方法往往临时且缺乏通用性。我们提出一个模块化、可扩展的评估框架,最初针对心理健康领域,通过可重复使用的管道实现新模型、指标和基准的无缝集成。该框架目前能重现现有最先进结果,并支持通过灵活组合模型、指标和基准进行系统交叉评估,其即插即用的提示模块为纳入沉思原则等伦理视角提供了原则性途径。尽管最初聚焦心理健康,但该框架与领域无关,自然可扩展到决策、道德推理和人机协作等领域;通过将计算评估与以人为本的伦理推理相联系,为跨学科研究奠定基础,朝着强大、可信且有益社会的人机生态系统发展。

英文摘要

Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language models (LLMs), improving cooperation and reducing ethical violations in LLM outputs. However, as new models, evaluation metrics, and benchmarks emerge rapidly, it remains challenging to systematically assess whether and how contemplative principles enhance LLM alignment across diverse and evolving scenarios, and existing approaches are often ad hoc and fail to generalize. We present a modular, extensible evaluation framework, initially targeted at the mental health domain, that enables seamless integration of new models, metrics, and benchmarks through a reusable pipeline. The framework currently reproduces existing state-of-the-art results and supports systematic cross-evaluation by flexibly mixing and matching models, metrics, and benchmarks, enabling fair comparison and deeper insight. Its plug-and-play prompting module offers a principled pathway for incorporating ethical perspectives such as contemplative principles, allowing domain experts to define alignment criteria without requiring technical expertise. Although initially focused on mental health, the framework is domain-agnostic and extends naturally to areas such as decision-making, moral reasoning, and human-AI collaboration. By bridging computational evaluation with human-centered ethical reasoning, this work lays the groundwork for interdisciplinary research spanning cognitive science, behavioral economics, philosophy, and system design, toward robust, trustworthy, and socially beneficial human-AI ecosystems.

CommentsAccepted as an oral presentation at HARMONY 2026 (Human-centered AI Research for Mental Health, an Open Networking Symposium), co-located with IEEE/ACM Conference on Connected Health: Applications, Systems, and Engineering Technologies (CHASE 2026) held in Pittsburgh, August 6, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑