arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00385cs.LGcs.AIcs.CLstat.ML

FAER:面向语言模型后训练的可审计、效用对齐的轨迹重放

FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training

Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, Jun Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

针对重放选择与学习效用脱节的问题,提出可审计的FAER框架,通过学习器感知选择器与审计契约提升后训练质量,在GSM8K上达到0.6624。

中文摘要 AI 辅助

重放选择器通常根据格式反馈、置信度、新鲜度或响应长度对缓存轨迹进行排序,尽管缓存级别的正确性与下游学习者的效用是不同的目标。我们形式化了这种选择与学习之间的差距,并引入FAER作为一个可审计的全轨迹重放框架。其免训练的固定选择器是一个协议基线;FAER-UTILITY是学习器感知的选择器,在不相交的校准块上拟合。归一化梯度对齐作为基线被报告,而一次性的优化器感知虚拟更新提供了一个幅度感知的效用面。审计契约在评估标签加入之前冻结观察到的字段和重放轨迹。在GSM8K上使用Qwen2.5-1.5B-Instruct,匹配的学习者研究报告固定选择器的质量为0.6329,而均匀选择为0.5482,格式反馈在128次更新下为0.6037。仅元数据的交叉拟合校准在八个种子上达到0.6476±0.0139(中位数0.6481;配对95%区间[+0.079,+0.122]),目标运行令牌为63,276;其记录的总成本为189,642个令牌和3.48个GPU小时,包括校准。完成的FAER-UTILITY行在62,844个目标运行令牌和4.26个GPU小时下达到0.6624。格式反馈选择的记录正确率为0.6953,而固定选择器为0.3594,尽管下游排名不同。完成的比较面报告了学习器感知的消融、同种子差距、策略优化行和严格的零样本迁移。

英文摘要

Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑