Alteron:用于NLP分类器版本间行为回归测试的工具
Alteron: A Tool for Behavioral Regression Testing Across NLP Classifier Versions
浏览论文内容
中文总结 AI 辅助
该研究提出工具Alteron,通过变形测试检测NLP分类器版本间的行为回归,在评估中识别出16个行为回归,其中11个为发布阻塞性问题,可揭示基准指标遗漏的故障。
中文摘要 AI 辅助
评估不断演进的自然语言处理(NLP)模型对于确保其在更新后行为可靠至关重要,但标准基准指标无法完全捕捉模型版本间的行为变化。现有工作主要关注孤立测试模型,而非在持续集成工作流中比较连续版本。我们提出Alteron,一种利用变形测试检测NLP模型版本间行为回归的工具。Alteron从带标签的源示例构建测试语料库,并通过变形转换后的输入比较模型版本。在涵盖10个变形关系(MRs)、4个模型版本和3次模型更新转换的评估中,Alteron识别出16个行为回归,其中11个为发布阻塞性问题。结果表明,常见模型更新可在保持整体任务性能的同时引入不良行为变化,且跨模型版本的行为检查能揭示仅聚合基准指标无法捕捉的故障。该工具为开源,可在指定网址获取,另有屏幕录制演示可在指定网址获取。
英文摘要
Evaluating evolving Natural Language Processing (NLP) models is important for ensuring reliable behavior across updates, but standard benchmark metrics do not fully capture how model behavior changes across versions. Existing work has focused mainly on testing models in isolation rather than comparing successive versions in continuous integration workflows. We present Alteron, a tool for detecting behavioral regressions across NLP model versions with metamorphic testing. Alteron constructs a test corpus from labeled source examples and compares model versions on metamorphically transformed inputs. In an evaluation spanning 10 metamorphic relations (MRs), 4 model versions, and 3 model-update transitions, Alteron identified 16 behavioral regressions, 11 of which were release-blocking. The results show that common model updates can preserve overall task performance while still introducing undesirable behavior changes, and that behavioral checks across model versions can reveal failures that aggregate benchmark metrics alone do not capture. The tool is open-source and available at https://github.com/shazzad5709/alteron. A screencast demonstration is available at https://youtu.be/szwiWW5O4do.