基于LLM的架构感知分裂学习用于跨异构调查的隐私保护心理困扰预测
LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys
浏览论文内容
中文总结 AI 辅助
提出一种基于LLM的架构感知分裂学习框架,通过语义编码统一异构调查数据,在保护隐私的同时实现高效协作学习,显著提升准确率并降低客户端计算开销。
中文摘要 AI 辅助
社会和生活复杂性的增加与全球心理困扰患病率的上升有关。教育机构、工作场所、诊所等收集大量心理健康调查数据,以了解和减轻这一负担。对此类数据的协作分析可以产生有效的、可泛化的预测模型。隐私约束和不同的调查设计(即不同的问题、量表、格式)阻碍了直接整合。我们提出了一种保留隐私的架构感知分裂学习(SL)框架,使用大型语言模型(LLM)作为共享语义编码器,以协调跨机构的异构调查架构。我们将每条调查记录序列化为自然语言描述,将不同的调查架构统一为通用格式。LLM通过低秩适配(LoRA)针对心理困扰评估进行微调,并分区部署在客户端和服务器上。客户端在本地保留原始调查响应,仅运行轻量级前端,因此原始记录永远不会离开收集它们的机构。资源密集型的主干在服务器上运行,最小化客户端计算。使用LLaMA-3.2-3B-Instruct,该框架仅用2,000个训练样本就达到了0.708的平均ANLS,在九种设置中的八种中超过了联邦学习(FL),并将每个客户端的计算量减少了三个数量级,同时泛化到未见过的数据集。总体而言,它实现了从异构心理健康调查数据中进行准确、隐私保护和资源高效的协作学习。
英文摘要
Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generalizable predictive models. Privacy constraints and varied survey designs (i.e., different questions, scales, and formats) hinder direct integration. We propose a schema-aware split learning (SL) framework that preserves privacy, using a large language model (LLM) as a shared semantic encoder to harmonize heterogeneous survey schemas across institutions. We serialize each survey record into a natural-language description, unifying disparate survey schemas into a common format. The LLM is fine-tuned for mental distress assessment via Low-Rank Adaptation (LoRA) and partitioned across client and server. Clients retain the raw survey responses locally and run only a lightweight front-end, so original records never leave the institution that collected them. The resource-intensive backbone runs on the server, minimizing client-side computation. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with only 2,000 training samples, surpasses federated learning (FL) in eight of nine settings, and cuts per-client computation by three orders of magnitude, while generalizing to unseen datasets. Overall, it enables accurate, privacy-preserving, and resource-efficient collaborative learning from heterogeneous mental health survey data.
发表机构
- Southern Illinois University Carbondale(南伊利诺伊大学卡本代尔分校)
机构由 AI 辅助整理,请以论文原文为准。