arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为教学提示的数据注释:从主观标签到批判性思维

Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking

Ralf Raumanns, Theresa Elstner, Louis Ferger-Andrews, Louise M. Carlsen, Martin Potthast, Gerard Schouten, Josien P. W. Pluim, Veronika Cheplygina

arXiv 2607.20149首次发表:更新:

发表机构

Fontys University of Applied Sciences; Eindhoven University of Technology; Kassel University; IT University of Copenhagen(方蒂斯应用科学大学; 埃因霍温理工大学; 卡塞尔大学; 哥本哈根信息技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨人工数据标注任务对学生理解标注主观性的作用,通过在两所大学开展标注活动及调查,发现能提升学生对课程内容熟悉度等,但存在情绪不安等问题,建议改进材料模糊性等,强调人工标注可让学生理解人类判断对模型的影响及分歧的意义。

AI 中文摘要

机器学习课程常使用预标记数据集,掩盖了人工标注的主观性,使学生过度信任人工智能数据和模型,低估解释多样性。我们调查人工数据标注任务能否让学生了解主观标注。在荷兰的方泰斯大学和丹麦的哥本哈根信息技术大学开展标注活动,让学生对皮肤病变图像的毛发覆盖情况进行三点量表标注。收集43名参与者的调查问卷,了解他们对标注模糊性、数据质量、偏差、公平性、实施障碍和教学效果的理解。主要发现:所有概念上自我报告的课程内容熟悉度大幅提高。大多学生认识到个人解读影响标注。学生认为该活动在理解偏差方面比传统讲座更有效,参与者有更强学习动力。主要缺点:查看医学图像带来的情绪不安是主要问题。许多学生仍要求更清晰指导方针以减少分歧,表明他们未将不同观点的分歧视为学习特点而非问题。对未来迭代的建议:确保材料中有足够的解释模糊性。减少重复标注工作量。减轻敏感内容带来的情绪不安。明确将分歧视为学习机会而非要解决的问题。人工数据标注能有效让学生明白人类判断塑造模型行为,分歧反映领域复杂性而非仅仅是噪声。

英文摘要

Machine learning courses often use pre-labelled datasets, hiding the subjectivity of human annotation. This produces an overly trusting view of data and AI models in students, at the expense of interpretive diversity and contestability of algorithmic outputs. We investigated whether manual data annotation tasks teach students about subjective labelling. Study Design: An annotation activity was implemented at two universities: Fontys (Netherlands) and IT University Copenhagen (Denmark). Students annotated skin lesion images for hair coverage on a 3-point scale. Surveys were collected from 43 participants, measuring their understanding of annotation ambiguity, data quality, bias, fairness, implementation barriers, and pedagogical effectiveness. Key Findings: Self-reported familiarity with the course content increased substantially across all concepts. Most students recognised that personal interpretation affects annotations. Students rated the activity as more effective than traditional lectures in understanding bias. Participants were motivated to learn more. Main Drawbacks: Emotional discomfort from viewing medical images was the primary issue. Many students still requested clearer guidelines to reduce disagreement, suggesting they had not yet internalised that disagreement arising from different perspectives is a feature, not a bug. Recommendations for Future Iterations: Ensure sufficient interpretive ambiguity in materials. Reduce repetitive annotation workload. Mitigate emotional discomfort from sensitive content. Explicitly frame disagreement as a learning opportunity rather than a problem to solve. Manual data annotations effectively teach students that human judgement shapes model behaviour and that disagreement reflects domain complexity, not just noise.

Comments24 pages, 6 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑