发表机构
INQUIRE Lab, School of Electrical and Computer Engineering; University of Oklahoma(INQUIRE实验室,电气与计算机工程学院; 俄克拉荷马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过快棋和段落重排序任务,验证了将概率模型 Jev 的判断嵌入搜索并蒸馏为轻量评估器,同时结合大语言模型教师进行多教师蒸馏,可显著提升决策性能。
AI 中文摘要
仅输出概率的模型(TypeSafe 称之为系统一模型)能在毫秒内为固定选项返回校准概率,且不生成任何文本。我们通过两个需要在严格约束下做出决策的任务来研究这样一个模型——Jev。在快棋(子弹棋)中,一个将 Jev 的判断嵌入 Stockfish 搜索,并结合开局库和残局库的机器人,其 Lichess 快棋等级分在与其他机器人的对战中攀升至 2200 以上。由于实时模型调用对搜索而言太慢,我们将成对判断蒸馏成一个紧凑的评估器,使其能在每个局面运行。随后,我们探讨在有一个大语言模型 Qwen3-32B 作为第二教师时,如何最优地分配固定的标注预算。在国际象棋中,平均两位教师的标签比将全部预算单独用于 Qwen 高出 9.6 Elo(95% 置信区间为 4.3 至 14.9),且该增益在新开局中可复现;从同一教师获取第二个答案无法替代,而在所测试的模型中,Jev 是 Qwen 的最强搭档。在段落重排序中,仅使用 Jev 的标签训练的重排序器,其得分与使用 Qwen 训练的不相上下,而仅需 21 分钟的 API 调用,而非 5.1 GPU 小时;添加 Qwen 最多只能在排序质量上提升千分之几。搜索提供了前瞻能力,蒸馏使判断成本足够低以用于每个局面,而大语言模型搭档在国际象棋中带来了收益。
英文摘要
Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges' labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev's labels alone train a reranker that scores as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.