arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁在打分?用Attune实现对大语言模型驱动打分的交互式调控

Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune

Bhavya Chopra, Meng Chen, Rebecca Dang, Chanbin Park, Shreya Shankar, Sepanta Zeighami, Bjoern Hartmann, Aditya Parameswaran

arXiv 2608.14948首次发表:更新:

AI 中文总结

研究针对LLM打分存在的全局理解与局部一致性不足问题,提出Attune混合启动系统,通过成对比较推导打分规则并支持用户调控,经技术评估与用户研究验证了其有效性。

AI 中文摘要

大语言模型(LLM)正越来越多地被用于大规模对文本记录打分(例如将求职者简历按1-5分制评级)。然而,现有的LLM驱动打分方法未考虑到有效打分既需要对记录的整体理解,也需要对相似记录做出局部一致的判断。我们提出Attune,这是一个可调控的LLM驱动打分混合启动系统。给定任务描述和打分范围,Attune首先在记录间进行成对比较以形成全局理解,随后将这些比较转化为一致的打分分配——在此过程中自下而上地推导打分标准和规则。这些内容作为打分逻辑的共享表示,用户可对其进行检查和编辑。基于一项针对12名受试者的形成性研究的见解,Attune的界面引入了新颖的调控交互方式,允许用户确定性地优化打分逻辑。用户可提供示例、直接编辑标准、规则或目标分布,以及给出自然语言反馈,所有优化都会编译为指导重新打分的约束条件。我们通过三项工作负载的技术评估,以及针对医疗、法律、教育和AI评估领域专家(共8名受试者)的用户研究,对我们的方法进行了验证。

英文摘要

Large language models (LLMs) are increasingly used to score text records at scale (e.g., rating candidate resumes on a 1-5 scale). However, existing LLM-powered approaches do not account for the fact that effective scoring requires both holistic understanding of records and locally consistent judgments across similar ones. We present Attune, a mixed-initiative system for steerable LLM-powered scoring. Given a task description and scoring range, Attune performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process. These serve as shared representations of scoring logic that users can inspect and edit. Based on insights from a formative study (n = 12), Attune's interface introduces novel steering interactions that allow users to deterministically refine scoring logic. Users can provide examples, directly edit criteria, rules, or target distributions, and give natural language feedback-with all refinements compiling into constraints that guide re-scoring. We validate our approach through a technical evaluation across three workloads and a user study with domain experts (n = 8) in healthcare, law, education, and AI evaluation.

Comments18 pages, To appear at ACM UIST 2026

DOI:10.1145/3830398.3830567

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑