从轨迹到证据:面向工业研究智能体的可审计实验记录
From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
- Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对工业研究智能体的实验轨迹无法自动成为有效证据的问题,提出基于证据的框架,经实验验证该框架可生成可审计记录,且生成的候选方案相对基准实现了积极在线提升。
AI中文摘要:
研究智能体在工业推荐场景中日益开展多轮机器学习实验,并保留所产生的轨迹以指导后续决策。然而,完整的轨迹并非自动成为证据:生成的人工制品可能缺乏支撑或不完整,执行的轮次可能无效或存在混杂,后续修改可能掩盖早期发现。我们研究“轨迹到证据的转换”,探究完整的研究过程实际确立了什么。我们引入了一个基于证据的框架,该框架将对关键人工制品的有限验证与执行后的声明限定相结合。一个上下文隔离的生成-验证-修复流程会在发布前检查人工制品是否存在证据违规和缺失的下游要求。执行后,有效性和归因检查会整合多轮证据,将干预级声明限定为可操作的修复、诊断性保护或被保留的发现,并将被认可的声明保存为具有明确来源和适用边界的可审计记录。随后,混合大语言模型(LLM)辅助控制器会基于可用的目标证据应用、推迟或拒绝记录。记录审计可确定哪些声明通过了限定,而下游诊断则将肯定性适用性判断识别为被测控制器的瓶颈。在从论文到目标的适配中,后续轮次通常优于第一轮,而最终轮次的表现往往不如早期的最佳轮次,这暴露出轨迹演化的非单调性。通过完整工作流程生成的候选方案相对于部署的基准方案产生了积极的在线提升效果。
英文摘要:
Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study \textbf{trajectory-to-evidence conversion}, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate--verify--repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.