arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向高流量应用中高效个性化主观判断的多维比较量表构建

Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

Xianglong Shi, Shifeng Liu, Sirui Zhao, Shengming Yuan, Enhong Chen

arXiv 2609.33282首次发表:更新:

AI 中文总结

本文提出多维比较量表构建框架,通过成对比较与稀疏Elo优化降低成本,并推出SubJudge模型实现个性化评分的O(1)推理,在多项指标上媲美前沿大模型。

AI 中文摘要

主观判断在许多高流量应用中处于核心地位,但主观强度难以量化,且不同个体之间的感知差异显著。为应对这些挑战,我们提出了一种用于多维量表构建的成对比较框架。通过沿案例维度和画像维度比较案例-个体对,该框架构建出能够同时捕捉细粒度强度和个体差异的相对量表。为支持实际的高流量部署,我们优化了离线量表构建和在线推理两个环节。在量表构建方面,我们将稀疏Elo比较与多裁判投票相结合,对于N个对象和每个对象K个对手的预算,将比较成本从O(N^2)降至O(NK),同时限制对任何单一裁判的依赖。在推理方面,我们提出了SubJudge,一个用于个性化评分的System One模型,采用批量偏好优化(BPO)。利用Bradley-Terry比较,BPO训练模型学习相对排序,SubJudge从首个响应位置的数字令牌概率中读取连续分数,每个标准仅需一次前向传播,将推理复杂度降至O(1)。在PluriHarms和iNews上的实验表明,我们的9B模型在多项指标上达到或超越了所评估的前沿LLM。在H100 GPU上,与不同思考预算下的Qwen3.5-9B相比,SubJudge的平均推理延迟实现了约1.29倍至261倍的加速。代码可在该https URL获取。

英文摘要

Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from $O(N^2)$ to $O(NK)$ for $N$ objects and a budget of $K$ opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to $O(1)$. Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately $1.29\times$ to $261\times$ speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at https://github.com/Longchentong/SubJudge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑