发表机构
Chulalongkorn University; Google Research(朱拉隆功大学; 谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过人机协作的迭代混合定性编码框架,大规模分析2020至2025年ACL和EMNLP论文中的局限性部分,揭示自我报告趋势、相关属性及写作模式,为NLP社区提供批判性反思。
AI 中文摘要
自2022年底以来,局限性部分已成为许多顶级NLP会议论文的强制要求。这些会议接收论文数量的不断增加,产生了大量自我报告的局限性文本,这些文本无法全部人工审阅,且至今仍未被系统性地分析。因此,在本文中,我们对2020年至2025年间发表于ACL和EMNLP会议的论文中的局限性部分进行了大规模分析,以了解研究者如何披露自身工作的不足。为此,我们实现了一个新颖的人机协作框架,用于迭代式混合定性编码。该框架使我们能够调查自我报告局限性随时间变化的趋势、其与特定论文属性的相关性,以及围绕这些披露反复出现的写作模式。我们的研究结果对NLP社区研究者所报告的多样化挑战以及其自我报告实践提供了批判性反思。
英文摘要
Since late 2022, a Limitations section has become mandatory at many top-tier NLP conferences. The growing number of accepted papers at these venues has resulted in a vast corpus of self-reported limitations that cannot all be manually reviewed, yet remains systematically unanalyzed. Therefore, in this paper, we conduct a large-scale analysis of the Limitations sections from ACL and EMNLP papers published between 2020 and 2025 to understand what researchers disclose about their own work. To do so, we implement a novel human-AI framework for iterative hybrid qualitative coding. This framework enables us to investigate trends in self-reported limitations over time, their correlations with specific paper attributes, and the writing patterns that recur around these disclosures. Our findings offer a critical reflection on the diverse reported challenges as well as the self-reporting practices of researchers in the NLP community.
CommentsEMNLP 2026 Findings