自适应生态瞬时评估与混合语言模型:形成性专家评审与回顾性评估
Adaptive Ecological Momentary Assessment with a Hybrid Language Model: Formative Expert Review and Retrospective Evaluation
浏览论文内容
中文总结 AI 辅助
提出EMA-E4B混合框架,结合岭回归与Gemma语言模型实现自适应生态瞬时评估,专家评审偏好完整响应,行动得分与参考相当。
中文摘要 AI 辅助
生态瞬时评估(EMA)用于测量日常生活中的体验,但固定的问卷和时间表收集的信息价值参差不齐,并可能打扰参与者。我们提出并回顾性评估了EMA-E4B,一个用于问题选择和提示时机的混合框架。独立的岭回归模型提出项目集和延迟;一个有监督的Gemma 4 E4B语言层生成最终的结构化响应和解释。评估区分了代理行动性能、输出符合性和响应质量的形成性判断。数据包含来自79名参与者的4,372条记录,并在参与者分离划分下产生3,516个序列案例。一位相关领域专家在20次决定性比较中的15次中偏好完整的混合响应,在22次评审中有两次平局。在75个重复使用的开发案例中,混合问题效用和时机相似度分别为0.832和0.818;仅头部模型分别达到0.852和0.818。一个单独的60案例比较,在相同头部下使用未触及的E4B,行动差异为-0.0031和-0.0105。因此,语言层产生的结构化响应的行动得分与参考配置相当或略低,而专家反馈偏好完整的混合响应。这些观察建立了一个具体、可检查的框架,并阐明了行动评分和响应评审的不同角色。重复的自适应管理和对测量及参与者负担的实际影响仍是未来研究。
英文摘要
Ecological momentary assessment (EMA) measures experience in daily life, but fixed questionnaires and schedules collect information of uneven value and can interrupt participants. We present and retrospectively evaluate EMA-E4B, a hybrid framework for question selection and prompt timing. Separate ridge models propose an item set and delay; a supervised Gemma 4 E4B language layer produces the final structured response and explanation. Evaluation distinguishes proxy action performance, output conformity, and formative judgments of response quality. The data contain 4,372 records from 79 participants and yield 3,516 sequential cases under a participant separated split. One involved domain expert preferred the complete hybrid response in 15 of 20 decisive comparisons, with two ties among 22 reviews. On 75 reused development cases, hybrid question utility and timing similarity were 0.832 and 0.818; the head alone reached 0.852 and 0.818. A separate 60 case comparison with untouched E4B under the same head gave action differences of -0.0031 and -0.0105. Thus, the language layer produced structured responses with action scores comparable to or slightly below the reference configurations, while the expert feedback favored the complete hybrid response. These observations establish a concrete, inspectable framework and clarify the distinct roles of action scoring and response review. Repeated adaptive administration and practical effects on measurement and participant burden remain future research.
发表机构
- School of Electrical and Computer Engineering, University of Oklahoma(俄克拉荷马大学电气与计算机工程学院)
- School of Psychology, Georgia Institute of Technology(佐治亚理工学院心理学院)
机构由 AI 辅助整理,请以论文原文为准。