腾讯工作伙伴基准测试:一个具有抗污染任务构建的多领域编码智能体基准测试
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
- Tencent(腾讯)
- Youtu Lab(腾讯优图实验室)
- Keen Security Lab(腾讯安全科恩实验室)
- Workbuddy(腾讯工作伙伴)
- Yunding Security Lab(腾讯云鼎实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
介绍腾讯工作伙伴基准测试这一多领域编码智能体评估套件,核心是统一评估框架,任务经反向设计防污染,数据集公开,四个子集探测工作不同方面,以统一格式和协议运行,还给出跨模型排行榜。
AI中文摘要:
我们介绍了腾讯工作伙伴基准测试,这是一个用于编码智能体的多领域评估套件。本报告记录了其构建方法、评分协议和跨模型排行榜。其核心是一个统一的评估框架,用于在代码、网络、办公和安全四个工作领域构建和运行基于分布的编码智能体任务。每个任务都从真实提交、拉取请求或业务场景反向设计而来,并改写为简短、口语化的角色扮演请求,以防止通过网络搜索底层问题来恢复任务提示。由于数据集是公开发布的,抗污染依赖于这种构建方式和数据集版本控制而非保密性。四个子集探测实际工作的互补方面,都采用统一格式并在统一可重复协议下运行。报告了跨多个模型家族的跨模型排行榜。
英文摘要:
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.