发表机构
University of Science and Technology of China; Singapore Management University(中国科学技术大学; 新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对HOBRE的三项任务,首次系统研究LLM-as-a-Judge评估范式,推出无参考基准BinJudgeBench,提出自适应选择最优配置的BinJudge,提升评估相关性并降低API成本。
AI 中文摘要
面向人类的二进制逆向工程(HOBRE)旨在将反编译伪代码转换为更便于人类理解的表示形式,从而减轻逆向分析的认知负担并提升效率。然而,可靠评估HOBRE输出仍是一项基础挑战:人工评估成本高、耗时且难以规模化,而现有自动指标要么需要可执行测试用例和运行时环境(对于真实世界二进制文件通常不可用),要么依赖高质量源代码参考(通常无法获取,且无法捕捉语义等价但词汇多样的输出)。尽管“大语言模型作为评判者(LLM-as-a-Judge)”范式天然适用于HOBRE评估,但其有效性仍未得到充分探索。本文针对HOBRE的三项代表性任务(函数名恢复、二进制代码摘要、反编译优化),首次系统研究了LLM-as-a-Judge范式。我们推出BinJudgeBench,这是首个基于多维度人工评判的专家标注无参考评估基准,其中LLM-as-a-Judge与人工评判的平均相关性达63.20%,优于传统自动指标的35.04%。通过分析主干大语言模型、提示策略和解码温度的评判配置,我们发现不存在“一刀切”的配置,最优设置因任务和单个样本而异。为解决这一问题,我们提出BinJudge,其采用轻量路由机制为每个任务和样本自适应选择最优评判配置。BinJudge将与人工专家的相关性提升4.5%至24.7%,并将API成本降至静态最优配置的0.06倍至0.84倍,为HOBRE提供了可规模化、高性价比且高保真的自动评估方案。
英文摘要
Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency. However, reliably evaluating HOBRE outputs remains a fundamental challenge: human evaluation is costly, time-consuming, and difficult to scale, while existing automated metrics either require executable test cases and runtime environments that are often unavailable for real-world binaries, or rely on high-quality source code references that are typically inaccessible and fail to capture semantically equivalent but lexically diverse outputs. Although LLM-as-a-Judge paradigm is naturally well-suited to HOBRE evaluation, its effectiveness remains underexplored. This paper presents the first systematic investigation of the LLM-as-a-Judge paradigm for HOBRE across three representative tasks: function name recovery, binary code summarization, and decompilation optimization. We introduce BinJudgeBench, the first expert-annotated, reference-free evaluation benchmark based on multi-dimensional human judgment, where LLM-as-a-Judge achieves an average correlation of 63.20\% with human judgment, outperforming traditional automated metrics at 35.04\%. By analyzing judge configurations across backbone LLMs, prompting strategies, and decoding temperatures, we find that no ``one-size-fits-all'' configuration exists, as the optimal setup varies across tasks and individual samples. To address this, we propose BinJudge, which employs a lightweight routing mechanism to adaptively select the optimal judge configuration for each task and sample. BinJudge improves correlation with human experts by 4.5\%-24.7\% and reduces API cost to 0.06$\times$-0.84$\times$ of that of static best configurations, providing a scalable, cost-effective, and high-fidelity automated evaluation scheme for HOBRE.
CommentsAccepted by the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)