基于缩放假设的无标签智能体学习控制器能力评估
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
浏览论文内容
中文总结 AI 辅助
该研究提出基于缩放假设的无标签智能体学习控制器评估框架,以教师模型与学生模型的收敛度为评分指标,验证其可作为标签缺失时控制器真实增益的有效代理。
中文摘要 AI 辅助
智能体“持续学习控制器”是将大语言模型(LLM)与检索或记忆结合、无需重新训练即可从反馈中改进的系统,在网络安全领域的价值日益凸显。但传统上其价值通过与带标签基准的增益来衡量,这种方法在实际安全场景中往往失效,因为基准标签稀缺、过时且不具代表性,从业者无法判断某个控制器是否有帮助,或两个控制器中哪个更适合任务。传统的LLM作为评判者无法提供有效信号,因为它并不比被评估的智能体更强,且在稀缺、零散、有偏差的标签上进行知识蒸馏不可靠。我们提出一种基于缩放假设、无需带标签基准即可端到端评估学习控制器的框架:更强的教师模型向配备持续学习控制器的较小学生模型提供稀疏采样的修正,我们通过学生模型随时间向教师模型的收敛程度来对控制器评分。在安全任务、模型家族和控制器设计上,我们发现相对于教师模型的改进与相对于保留金标准的改进相关,验证了标签缺失时,教师相对提升可作为控制器真实增益的代理。我们进一步发现,同等规模模型间的LLM作为评判者无法产生可用信号。这些结果表明,当人类提供相同类型的稀疏、高精度修正时,教师规模的模型可通过同一控制器得到改进。
英文摘要
Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. But their value is conventionally measured by gains against labeled benchmarks, an approach that often fails in operational security settings. Benchmark labels are scarce, stale, and unrepresentative, so a practitioner often cannot tell whether a given harness helps at all or which of two is better for their task. Traditional LLM-as-a-judge offers little signal because it is no stronger than the agent it evaluates, and distillation is unreliable on scarce, sporadic, and biased labels. We propose a framework for evaluating learning harnesses end-to-end without a labeled benchmark, grounded in the scaling hypothesis. A stronger teacher model provides sparsely sampled corrections to a smaller student with a continual learning harness. We score a harness by how much its student converges toward the teacher over time. Across security tasks, model families, and harness designs, we show that improvement relative to the teacher correlates with improvement relative to a held-out gold standard, validating teacher-relative lift as a proxy for true harness uplift when labels are absent. We further show that LLM-as-a-judge between similarly powered models yields no usable signal. These results suggest that a teacher-sized model can be improved through the same harness when humans provide the same kind of sparse, high-precision corrections.
发表机构
- Georgian
机构由 AI 辅助整理,请以论文原文为准。