arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00073cs.HC

家庭能源管理会话系统综合评价框架

A Comprehensive Evaluation Framework for Conversational Home Energy Management Systems

Wooyoung Jung

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM集成家庭能源管理系统的多轮对话评估缺失问题,提出基于目标-问题-度量的23项指标框架,经970段对话验证,可区分不同系统配置的优劣。

中文摘要 AI 辅助

家庭能源管理(HEM)日益增长的复杂性要求先进的系统能够引导居住者做出反映其背景、偏好和情境的明智能源决策。集成大语言模型(LLM)的HEM系统(HEMS)已展现出潜力,但以往研究依赖于以响应准确性为主要指标的单轮或单任务评估。这些系统能否在现实使用中典型的扩展多轮对话中提供有效交互,仍是一个悬而未决的问题。本研究引入了一个基于目标-问题-度量方法论的LLM集成HEMS综合评价框架,该框架分为五个类别:任务性能、事实准确性、交互质量、控制能力和系统效率。研究提出了跨多轮对话的共23项指标,并采用LLM作为评审的流水线以实现可扩展的自动评分。其可靠性通过与三位训练有素的人工编码员进行验证:经过迭代评分标准校准后,十五项LLM评分指标中有十二项达到强一致性(ICC >= 0.73),其中三项达到完全一致,其余三项在人工评分中表现出接近零的方差,因此通过平均绝对误差(0.04-0.28)进行报告。为展示该框架的有效性,生成了970段对话——涵盖16个场景和五种人物角色——并跨四种会话式HEMS配置进行了评估,这些配置跨越从使用原始能源数据的普通LLM到多智能体HEMS的复杂度梯度。该框架在多个评估维度上区分了四种配置,揭示了它们各自的优缺点。本研究通过提供一种可复现、多维度、全面评估持续且情境感知系统性能的评估方法,为会话式HEMS领域做出了贡献。

英文摘要

The growing complexity in home energy management (HEM) demands advanced systems that guide occupants toward informed energy decisions reflecting their background, preferences, and context. Large language model (LLM)-integrated HEM systems (HEMS) have demonstrated promise, but previous studies relied on single-turn or single-task evaluations with response accuracy as the primary metric. Whether such systems deliver effective interactions across the extended multi-turn dialogues typical of real-world use remains an open question. This study introduces a comprehensive evaluation framework of LLM-integrated HEMS derived from the Goal-Question-Metric methodology, organized across five categories: task performance, factual accuracy, interaction quality, control capability, and system efficiency. A total of 23 metrics across multi-turn conversations are proposed and an LLM-as-judge pipeline is employed to enable scalable automated scoring. Its reliability is validated against three trained human coders: after iterative rubric calibration, twelve of the fifteen LLM-scored metrics reached strong agreement (ICC >= 0.73), three of them perfect, while the remaining three exhibited near-zero variance in human scores and are instead reported via mean absolute error (0.04-0.28). To demonstrate the framework's effectiveness, 970 dialogues -- 16 scenarios and five personas -- were generated and evaluated across four conversational HEMS configurations spanning a sophistication gradient, from a vanilla LLM with raw energy data to a multi-agent HEMS. The framework distinguished the four configurations across multiple evaluation dimensions, revealing their respective strengths and weaknesses. This study contributes to conversational HEMS by providing a reproducible, multi-dimensional evaluation methodology that comprehensively assesses sustained, context-aware system performance.

发表机构

  • University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑