arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22459cs.HC

“我想要被推动,我想要成长”:让社会工作者能够设计工作中大型语言模型(LLM)增强的评估方案

"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work

Anna Kawakami, Chloe Qianhui Zhao, Renee Shelby, Fernando Diaz, Haiyi Zhu, Kenneth Holstein

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出工作者驱动的AI测量方法,通过与19名校社工作者开展8场研讨会,设计LLM评估基准,验证其可区分6种先进LLM性能,该方法可作为自上而下AI评估的补充。

中文摘要 AI 辅助

越来越多的工作者被要求采用人工智能(AI)系统辅助工作,但在定义什么是有意义的AI增强应呈现的形态、以及如何对其进行评估方面,他们几乎没有话语权。本文中,我们提出了工作者驱动的AI测量——一种自下而上的AI评估方法,工作者可通过协作来决定AI应增强哪些任务、“成功的”增强是什么样的、以及应如何对其进行测量。我们通过与来自当地学校社会工作组织的19名工作者开展的案例研究,探索如何为这一过程提供支持。通过一系列共8场研讨会,工作者迭代制定了用于AI评估的自身测量目标,将这些目标系统化,随后设计了一个基准,以捕捉大型语言模型(LLM)在日常工作场景中,能在多大程度上“推动”工作者反思自身的假设与偏见。工作者基于自身的专业知识与生活经验,协作设计并优化了一个“LLM作为评判者”的 rubric( rubric 保留英文原名,指评分标准细则)。在对工作者创建的基准进行验证时,我们发现工作者与LLM评判者的评分之间存在高度一致性,且该基准能够区分6种最先进的LLM的性能表现。基于我们的案例研究,我们探讨了未来工作的机遇,以支持工作者驱动的AI测量,将其作为现有自上而下AI评估方法的补充途径。

英文摘要

Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what "successful" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can "challenge" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their professional and lived expertise. In validations of the worker-created benchmark, we find that there is strong agreement between worker and LLM judge ratings and that the resulting benchmark can differentiate performance across six state-of-the-art LLMs. Based on our case study, we discuss opportunities for future work to support worker-driven AI measurement as a complementary approach to existing top-down AI evaluation approaches.

补充信息

↑