使用差异感知特征进行部署风险评估:以Prime Video为例
Deployment Risk Assessment Using Diff-Aware Features: A Case Study at Prime Video
浏览论文内容
中文总结 AI 辅助
研究针对Prime Video代码部署风险评估难题,引入以差异感知特征为中心的框架,系统识别风险预测指标,用大语言模型提取特征,在两数据集评估,最佳模型有较好检测结果,证明该框架能有效评估风险并避免隐私问题。
中文摘要 AI 辅助
在亚马逊Prime Video,我们面临着在直播活动和快速功能发布期间管理代码部署而不导致服务中断的关键运营挑战。当前的变更控制方法使用全面的部署冻结,不顾风险阻止所有变更,给开发者带来大量工作。虽然先前研究探索了风险变更预测器,但依赖开发者特定元数据或大量历史数据,引发隐私问题并限制新项目适用性。我们引入了一个以差异感知特征为中心的框架,这些特征直接从代码修改中得出。我们的关键贡献是系统识别风险预测所需的定量指标(代码级和变更级指标)和定性指标(编码风格违规、变更类型分类)。我们使用大语言模型作为多语言特征提取器,证明其在代码分析中的有效性。我们在两个数据集上评估了框架,最佳模型在检测风险代码变更时平均召回率为0.83,F1分数为0.81。消融分析表明变更级数量指标是有噪声的预测器,而结构代码复杂度提供更强风险信号。这些结果表明精心的特征选择能在不同编程语言和组织环境中有效进行变更风险评估,同时避免隐私问题。
英文摘要
At Amazon Prime Video, we face the critical operational challenge of managing code deployments during live events and rapid feature releases without causing service outages. Current change control approaches use blanket deployment freezes that block all changes regardless of risk, creating significant developer toil. While prior research has explored risky change predictors, these rely on developer-specific metadata or extensive historical data, raising privacy concerns and limiting applicability to new projects. We introduce a framework centered on diff-aware features, characteristics derived directly from code modifications. Our key contribution is the systematic identification of which quantitative metrics (code-level and change-level metrics) and qualitative indicators (coding style violations, change type classification) are necessary for risk prediction. We employ LLMs as multi-language feature extractors, demonstrating their effectiveness for code analysis beyond generation tasks and eliminating the need for language-specific tooling. We evaluated our framework on two datasets: Prime Video's production environment and the public ApacheJIT dataset. Our best-performing model achieves an average recall of 0.83 and F1 score of 0.81 across both datasets for detecting risky code changes. Notably, ablation analysis reveals that change-level volume metrics (e.g., lines added/deleted) are noisy predictors, while structural code complexity provides a substantially stronger risk signal. These results demonstrate that thoughtful feature curation enables effective change risk assessment across different programming languages and organizational contexts while avoiding privacy concerns.
发表机构
- University of California(加州大学)
机构由 AI 辅助整理,请以论文原文为准。