arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

万花筒计划:面向现实世界人工智能应用的情境化、符合人类需求的评估

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee

arXiv 2607.14673首次发表:更新:

发表机构

GovTech; National University of Singapore; University of British Columbia(新加坡政府科技局; 新加坡国立大学; 英属哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该项目旨在解决现实世界AI应用评估瓶颈,提出万花筒计划,通过集成基于角色的测试生成、情境化评分标准和人工审查,实现可靠性门控自动评分,经试点和实验验证了其在端到端可靠自动评分方面的有用特征。

AI 中文摘要

评估是现实世界人工智能应用的部署瓶颈:公共基准很少能匹配团队的用户、情境或政策,人工审查往往难以扩展。受我们在公共部门人工智能应用工作的启发,该项目解决了应用必须满足当地政策和治理要求时反复出现的评估挑战。我们提出了万花筒计划,这是一种用于情境功能评估的集成工作流程,它将基于角色的测试生成、情境化评分标准和人工审查联系起来,以实现可靠性门控自动评分。生成的测试用例根据特定于应用的评分标准进行评分;人工注释提供可审查的标签;只有当大语言模型判断与这些标签的一致性达到配置的阈值时,才会自动进行评分。因此,万花筒计划是产品团队实用、可检查、迭代的工作流程。我们报告了在四个组织用例上进行的为期三周的试点以及针对跨越四个领域和14个评估维度的108个带注释问答对进行的自定义评分标准判断实验的早期证据。结果突出了端到端可靠自动评分的有用特征。

英文摘要

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑