辩论到技能:面向工业查询到智能体标注的能力边界过程监督
Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation
- Baidu, Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对工业查询到智能体匹配中主题相关性与可执行能力混淆的问题,提出能力边界过程监督方法辩论到技能,通过决策原则与结构化审议提升灰色地带案例的匹配效果。
AI中文摘要:
工业查询到智能体的匹配在主题相关性被误认为可执行能力时会失败,尤其是在长尾和边界敏感请求上。我们将标注形式化为能力边界过程监督,并通过辩论到技能(Debate-to-Skill)方法实例化,该方法使用可复用的决策原则、结构化审议、基于验证器的结论提取以及分歧驱动的改进。在工业查询到智能体基准(Query2Agent)上,我们将辩论到技能与直接标签监督、推理监督微调(reasoning-SFT)及结构消融进行比较。结果检验了收益是否来自对能力关键决策过程本身的监督,尤其是在语义相关性与可执行能力相背离的灰色地带案例中。
英文摘要:
Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as \emph{capability-bound process supervision} and instantiate it with Debate-to-Skill, which uses reusable decision principles, structured deliberation, verifier-based verdict extraction, and disagreement-driven refinement. On an industrial Query2Agent benchmark, we compare Debate-to-Skill with direct-label supervision, reasoning-SFT, and structural ablations. The results test whether gains come from supervising the capability-critical decision process itself, especially on grey-zone cases where semantic relatedness and executable capability diverge.