AI 中文总结
本研究将标题选择建模为赢家通吃分类任务,对比发现用于标题点击率预测的密集嵌入回归模型,性能优于经LoRA微调的0.6B规模生成式语言模型,凸显了嵌入回归模型在高吞吐量内容排序中的应用潜力。
AI 中文摘要
优化数字内容标题的点击率(CTR)是在线媒体和推荐系统中的重要问题。尽管大语言模型(LLM)展现出强大的生成能力,但它们在判别式排序任务(如从候选集中选择表现最佳的标题)中的有效性仍有待深入研究。本研究将标题选择建模为赢家通吃分类问题,对比了经低秩适配(LoRA)微调的因果语言模型LOLAQwen(0.6B)与密集嵌入回归模型在标题选择任务中的表现,采用包含3263个A/B测试标题组的数据集进行评估。性能以Top-1准确率衡量,即模型正确识别每组最高表现标题的比例。嵌入回归模型的Top-1准确率达42.79%,而经LoRA微调的语言模型仅为35.70%。结果表明,在该标题选择任务中,轻量级判别式方法的表现优于采用参数高效适配微调的小型生成式语言模型;研究结果凸显了基于嵌入的回归模型作为高效替代方案,在高吞吐量内容排序应用中替代生成式模型的潜力。
英文摘要
Optimizing digital content headlines for click-through rate (CTR) is an important problem in online media and recommendation systems. While large language models (LLMs) have demonstrated strong generative capabilities, their effectiveness for discriminative ranking tasks, such as selecting the highest-performing headline from a set of candidates, remains less well understood. In this work, we compare a LoRA-fine-tuned causal language model, LOLAQwen (0.6B), with a dense embedding regression model for headline selection. We formulate headline selection as a winner-take-all classification problem and evaluate both approaches using a dataset of 3,263 A/B-tested headline groups. Performance is measured using Top-1 accuracy, defined as the proportion of groups for which the model correctly identifies the highest-performing headline. The embedding regression model achieves a Top-1 accuracy of 42.79%, compared with 35.70% for the LoRA-fine-tuned language model. These results indicate that, for this headline selection task, a lightweight discriminative approach can outperform a small generative language model fine-tuned using parameter-efficient adaptation. The findings highlight the potential of embedding-based regression models as efficient alternatives to generative models for high-throughput content ranking applications.