arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23766cs.CLcs.AI

哪些内容会进入专家评审?AI辅助项目开发中的表征、结构筛选与候选形式依赖性

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

Christopher Brooks

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过两项针对32000个大五人格项目的计算机模拟研究,揭示AI辅助项目开发中计算评估器的表征、筛选策略会影响进入专家评审的内容,其并非中立基础设施,而是测量设计的可修改部分。

中文摘要 AI 辅助

在AI辅助项目生成与专家评审之间存在一个计算评估器,其决策通常被视为技术层面的初步环节。然而,表征、结构简化与选择策略决定了心理测量学家最终能接收到哪些项目和证据。我们针对32000个筛选出的大五人格(Big Five)项目开展了两项关联的计算机模拟研究,追踪了从语义表征到结构评估再到候选形式构建的固定源群体。语义几何层面的广泛一致性掩盖了关键的局部差异:相同措辞会产生不同的构念证据,不同的项目会留存下来,即便社区对应关系有所改善,预期属性也可能消失。这些敏感性在不同生成源群体间也存在差异。在最终评审边界处,两种资格政策均覆盖了所有可评估形式的每个内容单元,但呈现的措辞不同。在各类嵌入配置中,包容性主要形式的40个项目中仅平均共享6个,反映了在结构证据和排序中改变表征所产生的全部下游影响。因此,全局摘要和完整形式的表面稳定性掩盖了心理测量学家所接触内容的不稳定性。计算评估器并非生成与专业知识之间的中立基础设施,而是测量设计中可检查和可修改的组成部分。

英文摘要

Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.

发表机构

  • School of Information, University of Michigan(密歇根大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑