ACL 负责任 NLP 清单是一种应付差事的形式主义吗?EMNLP 2025 的大规模分析
Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025
- University of Aberdeen(阿伯丁大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过分析 EMNLP 2025 73922 份清单回答及理由,发现清单存在逻辑矛盾、理由质量差等问题,提出强制最低字数等优化建议。
AI中文摘要:
负责任 NLP 实践包含 a)透明度、b)伦理、c)社会影响三个方面,负责任 NLP 清单旨在推动这些目标并促进负责任的实践。近期 ACL 发布了 EMNLP 2025 清单,以辅助当前研究实践的透明度,本研究以此为核心展开。我们整理并发布了首批两个数据集:a)EMNLP 2025 主会论文集(Main)和发现论文集(Finding)所有清单的回答及理由;b)关联论文章节的清单参考。我们还通过检查 73922 个清单回答及理由,完成了对近期 EMNLP 清单的首次分析。针对主会论文集,我们发现作者将清单中的伦理问题与论文主体内容割裂开来,模仿了伦理问题是事后补充的趋势。随后我们检查了“NO”回答,发现 44.9% 的理由质量较差或缺乏诚意,表现为简短或空泛。我们还发现清单设计和作者付出存在重大问题:6% 的清单包含父项与子项回答之间的逻辑矛盾。此外,我们发现存在表面遵守负责任伦理的证据,53% 的作者对其工作的潜在风险或社会影响不予认可,而这类内容本不应缺失。我们将此与发现论文集对比,发现两个论文集存在相似趋势。最后,我们讨论了清单设计的影响并为未来清单迭代提供建议,包括:a)强制最低字数要求;b)加强对应用风险的审查。
英文摘要:
Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications from the EMNLP 2025 Main and Finding tracks; b) checklist reference linking to paper sections. We also provide the first analysis of recent EMNLP Checklists, by examining $73,922$ responses and justifications to them. For the Main track, we find that authors isolate ethics questions of the Checklist from the paper's bulk, mimicking the trend of ethics being an afterthought. We then examine \texttt{NO} responses. We find $44.9\%$ of justifications are poor or bad-faith, being brief or empty. Then, we find significant issues with the checklist design and effort of authors, namely that $6\%$ of all checklists contained logical contradictions between parent and child responses. We also find evidence of surface compliance for responsible ethics, with $53\%$ authors dismissing potential risks or social impacts of their work, for which there should be none. We compare this to the Findings track, noticing a similar trend in both tracks. Lastly, we discuss the implications of the checklist design and provide recommendations for future checklist iterations. Including: a) enforcing a minimum word count, b) enforcing more scrutiny on the risks of appliances.