HARGO:面向HPC任务的LLM RL后训练的异质性感知奖励引导优化
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
浏览论文内容
中文总结 AI 辅助
针对HPC任务的LLM RL后训练的异质性问题,提出HARGO方法,通过置信度调制优势实现逐响应重要性加权,在四项HPC任务的三项核心指标上均取得最优性能。
中文摘要 AI 辅助
监督微调(SFT)可让大语言模型(LLM)掌握高性能计算(HPC)任务的领域知识,比如数据竞争检测、基准问答等。但仅有知识无法保证任务适配行为:一款对C/C++数据竞争样本分类准确率达88.65%的SFT模型,回答事实查询时表述冗长不精准,65.9%的MLPerf响应字符数超40。强化学习(RL)后训练通过优化任务特定奖励而非 token 级模仿来弥补这一差距。然而HPC任务存在极强异质性:二分类、事实问答、语义生成的答案长度差异达58倍,涵盖三种不同奖励分布,且SFT准确率差异显著,这使得GRPO等均匀权重RL方法表现欠佳。我们提出HARGO(Heterogeneity-Aware Reward-Guided Optimization,异质性感知奖励引导优化),它引入基于置信度调制优势的逐响应重要性加权:通过组级奖励对比计算判别信号,通过参考模型对数概率计算置信度信号,随后在计算逐响应权重前调制优势,无需任务类型标签。在四项HPC任务及九种方法的对比中,HARGO在三项核心指标上均取得最优性能:WinRate为54.62%,数据竞争F1为91.30%,PLP相似度为0.8558。 ablation 实验证实两种信号均有互补贡献,HARGO在异构HPC任务的对比方法中实现了最佳整体对齐质量。
英文摘要
Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering. However, knowledge alone does not guarantee task-appropriate behavior: the same SFT model that correctly classifies 88.65\% of C/C++ data race samples produces verbose, imprecise answers to factual queries, with 65.9\% of MLPerf responses exceeding 40 characters. Reinforcement learning (RL) post-training addresses this gap by optimizing for task-specific rewards rather than token-level imitation. Yet HPC tasks exhibit extreme heterogeneity, with binary classification, factual QA, and semantic generation differing by 58x in answer length, spanning three distinct reward distributions, and showing widely varying SFT accuracy. This makes uniform-weight RL methods such as GRPO suboptimal. We propose HARGO, Heterogeneity-Aware Reward-Guided Optimization, which introduces per-response importance weighting via confidence-modulated advantage: computing a discrimination signal from group-level reward contrast and a confidence signal from reference model log-probabilities, then modulating the advantage before computing per-response weights, without requiring task-type labels. Across four HPC tasks and nine methods, HARGO achieves the best performance on all three primary metrics: WinRate 54.62\%, Data Race F1 91.30\%, and PLP Similarity 0.8558. Ablation confirms complementary contributions from both signals. HARGO establishes the best overall alignment quality among compared methods for heterogeneous HPC tasks.