MonitrLLM:面向大型语言模型的社区中心评估基础设施
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出开源LLM评估基础设施MonitrLLM,将对话记录、用户任务意图与结果关联,试点显示多轮对话失败率为单轮2.5倍,为LLM评估提供新方法。
中文摘要 AI 辅助
基准套件评估模型在受控任务上的能力;大规模对话语料库捕获无用户反馈的自然使用场景;界面内反馈机制记录无任务目的的满意度。这些共同导致了大型语言模型(LLM)评估中存在关键缺口:现有基础设施未将交互轨迹与用户定义的结果常规关联。我们提出MonitrLLM,这是面向社区中心LLM评估的开源基础设施,它将完整对话记录与用户报告的任务意图及结果评估关联,将三者均视为主要评估信号而非可选元数据。为验证该方法的价值,我们开展了为期两周的可行性试点,26名大学生使用ChatGPT,收集到206份含完整对话记录的评估报告。试点结果表明关联对话轨迹与用户报告结果的价值:例如,尽管参与者对LLM交互的平均满意度达4.19/5,但在目标任务上仍存在23.1%的较高失败率;此外,多轮对话的失败率是单轮对话的2.5倍,该模式将多轮交互重新定义为困难的信号而非参与度的体现。最后,我们讨论将直接用户反馈与观测数据结合以实现稳健LLM评估的价值,以及达成该目标的基础设施可能性。
英文摘要
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.