训练了但没学会:面向作为前部署工程师的LLM智能体的后训练交付基准
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
- University of California, Berkeley(加州大学伯克利分校)
- Imperial College London(伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM智能体作为前部署工程师的后训练交付场景,提出一个十阶段受治理交付基准,揭示“训练了但没学会”的隐性失败,并用验收门控和检测器确保交付可信,实验覆盖四个前沿智能体与人类对照。
AI中文摘要:
后训练正在成为一种服务(PTaaS):客户向操作员提供数据和目标,一名前部署工程师(FDE)在预算、人工审批门控和可复现性要求下,返回一个微调、评估并部署好的模型。让LLM智能体担任FDE角色引发了一个现有基准无法回答的问题:不是智能体能否提升指标,而是它能否被信任去完成交付。我们在一个受治理的交付平面上回答这个问题,智能体驱动十个阶段,一个预言机根据平台记录的事实对每个阶段进行评分。核心的隐性失败是“训练了但没学会”(TBDL)的运行:损失下降,所有信号保持绿色,但交付的模型并不比基础模型更好。操作员运行的验收门控在付款前捕获每一次此类运行,一个在已知损坏运行上校准的检测器在运行中途标记严重损坏。我们在计量的L40S、A100和H200 GPU上,对8B到70B的开源基础模型,端到端运行了四个前沿智能体(Claude Opus 5、GPT-5.6-luna、Gemini 3.7 Flash、DeepSeek V4-Pro),在评分前认证每个场景。我们还在同一预言机下运行了一个人类FDE对照组,并将每个智能体与其进行比较。
英文摘要:
Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a question existing benchmarks cannot answer: not whether an agent can raise a metric, but whether it can be trusted to deliver. We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that trains but does not learn (TBDL): loss falls, every signal stays green, and the delivered model is no better than the base. An operator-run acceptance gate catches every such run before payment, and a detector calibrated on known-corrupted runs flags severe corruption mid-run. We ran four frontier agents (Claude Opus 5, GPT-5.6-luna, Gemini 3.7 Flash, DeepSeek V4-Pro) end to end on metered L40S, A100, and H200 GPUs across 8B to 70B open bases, certifying every scenario before scoring. We also ran a human FDE arm under the same oracle and compare every agent against it.