AI 中文总结
研究针对LLM打分存在的全局理解与局部一致性不足问题,提出Attune混合启动系统,通过成对比较推导打分规则并支持用户调控,经技术评估与用户研究验证了其有效性。
AI 中文摘要
大语言模型(LLM)正越来越多地被用于大规模对文本记录打分(例如将求职者简历按1-5分制评级)。然而,现有的LLM驱动打分方法未考虑到有效打分既需要对记录的整体理解,也需要对相似记录做出局部一致的判断。我们提出Attune,这是一个可调控的LLM驱动打分混合启动系统。给定任务描述和打分范围,Attune首先在记录间进行成对比较以形成全局理解,随后将这些比较转化为一致的打分分配——在此过程中自下而上地推导打分标准和规则。这些内容作为打分逻辑的共享表示,用户可对其进行检查和编辑。基于一项针对12名受试者的形成性研究的见解,Attune的界面引入了新颖的调控交互方式,允许用户确定性地优化打分逻辑。用户可提供示例、直接编辑标准、规则或目标分布,以及给出自然语言反馈,所有优化都会编译为指导重新打分的约束条件。我们通过三项工作负载的技术评估,以及针对医疗、法律、教育和AI评估领域专家(共8名受试者)的用户研究,对我们的方法进行了验证。
英文摘要
Large language models (LLMs) are increasingly used to score text records at scale (e.g., rating candidate resumes on a 1-5 scale). However, existing LLM-powered approaches do not account for the fact that effective scoring requires both holistic understanding of records and locally consistent judgments across similar ones. We present Attune, a mixed-initiative system for steerable LLM-powered scoring. Given a task description and scoring range, Attune performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process. These serve as shared representations of scoring logic that users can inspect and edit. Based on insights from a formative study (n = 12), Attune's interface introduces novel steering interactions that allow users to deterministically refine scoring logic. Users can provide examples, directly edit criteria, rules, or target distributions, and give natural language feedback-with all refinements compiling into constraints that guide re-scoring. We validate our approach through a technical evaluation across three workloads and a user study with domain experts (n = 8) in healthcare, law, education, and AI evaluation.
Comments18 pages, To appear at ACM UIST 2026