arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15864cs.HC

迈向可扩展的持久技能测量

Towards Scalable Measurement of Durable Skills

Amir Globerson, Amy Keeling, Anisha Choudhury, Anna Iurchenko, Aviad Segal, Avinatan Hassidim, Ayça Çakmakli, Ben Gomes, Benn Witt, Cathy Cheunga, Cristine Lega… 展开作者

Amir Globerson, Amy Keeling, Anisha Choudhury, Anna Iurchenko, Aviad Segal, Avinatan Hassidim, Ayça Çakmakli, Ben Gomes, Benn Witt, Cathy Cheunga, Cristine Legare, Diana Akrong, Eliad Carmi, Elisabeth Bauer, Gal Elidan, Hadas Gelbart, Hairong Mu, Katherine Chou, Lev Borovoi, Nir Kerem, Niv Efron, Noa Kerrem Gilo, Preeti Singh, Rajvi Kapadia, Rena Levitt, Roni Rabin, Ronit Levavi Morad, Rotem Yulzary, Shashank Agarwal, Sophie Allweis, Tracey Lee-Joe, Tzvika Stein, Yael Bar Moshe, Yael Haramaty, Yaniv Carmel, Yishay Mor, Yoav Bar Sinai, Yoav Bergner, Yossi Matias, Yuri Lev

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出利用LLM框架,通过AI队友对话模拟人际互动,结合执行LLM引导证据和AI评估器,实现持久技能的可扩展、可控测量,并验证其有效性与专家评分一致性。

中文摘要 AI 辅助

持久技能,如协作、创造力和批判性思维,对现代职场的成功至关重要。然而,测量这些技能仍然是一个持续的挑战。此外,由于未被测量的内容往往不被教授,这些技能在主流教育课程中常常被忽视。设计有效的评估方法需要平衡两个经常冲突的要求:生态效度和心理测量严谨性。一方面,评估环境应模拟真实世界中人与人之间的自然互动。另一方面,它应具有可扩展性、可控性和可重复性。在此,我们认为LLM可用于更好地实现这两个目标。具体而言,我们开发了一个框架,其中受试者与AI队友进行对话,这种方式类似于人际互动以保持真实性,同时提供信息丰富且稳健评估所需的心理测量控制。重要的是,AI参与者不仅作为队友,而且在“执行LLM”设置中,引导对话以引出高密度的可观察技能熟练度证据。我们辅以一个AI评估器,用于在此类互动中测量技能熟练度。我们基于人类参与者与我们的AI框架互动的转录文本,针对多种持久技能评估了我们的评估协议。对于创造力技能,我们进一步展示了自动评分器在评估真实学生执行的复杂任务中的有效性。我们的分析表明,使用执行LLM显著增加了引出的证据,并且LLM自动评分对话与专家标注者的评分基本一致。这项研究展示了编排式LLM方法在以可扩展和可控方式测量复杂社会与认知构念方面的实用性。

英文摘要

Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.

发表机构

  • Google Research
  • OpenMic
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

↑