arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OTROPE:基于最优传输的大语言模型鲁棒离线评估

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

Liner Xiang, Wenbo Zhang, Hengrui Cai

arXiv 2609.36264首次发表:更新:

发表机构

University of California, Irvine(加州大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM离线评估中标注稀缺、分布偏移和黑盒不可似然的问题,提出基于最优传输的OTROPE方法,通过语义空间分布校正结合双重稳健估计,实现无需行为模型或密度比的鲁棒评估,实验证明其优于基线且可提升弱评估器集成性能。

AI 中文摘要

大语言模型(LLM)的可靠评估对其开发与部署至关重要,然而在线评估往往成本高昂、风险较大且难以安全执行。我们研究LLM的离线策略评估(off-policy evaluation),其中使用来自行为模型的有限人工标注数据来评估更新的目标LLM。该设置具有挑战性,因为标注稀缺、行为与目标分布偏移常见,且黑盒LLM的响应似然通常不可用。我们提出基于最优传输的鲁棒离线评估(OTROPE),这是一种无似然评估方法,通过最优传输在语义空间中进行分布校正,将带标注的行为策略样本与未标注的目标策略样本对齐。OTROPE将校正后的人工标注残差与代理预测器相结合,产生一种无需行为策略建模或密度比估计的双重稳健风格评估。我们从理论上刻画了基线评估器在LLM分布偏移下失败的原因,并建立了当重加权行为分布或代理预测器收敛时OTROPE的一致性和收敛速率。在合成和真实LLM评估任务上的实验表明,OTROPE始终优于基线,同时使较弱的LLM评估器集成能够接近甚至超越较强的评估器。代码可在该https URL获取。

英文摘要

Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑