预测Bluesky新帖子的自定义信息流返回:一项前瞻性研究
Predicting Custom-Feed Returns for New Bluesky Posts: A Prospective Study
浏览论文内容
中文总结 AI 辅助
该研究针对Bluesky自定义信息流的新帖子冷启动路由任务,构建了先收集后标注的基准数据集,实验显示LambdaRank在相关指标上表现最优。
中文摘要 AI 辅助
传统冷启动推荐方法针对新用户或新引入的项目。Bluesky自定义信息流营造了不同的场景:独立运营的信息流从共享公共流中筛选内容。在此场景中,新发布的帖子是冷启动对象,而信息流则作为候选。我们提出了一种冷启动路由任务,其中新接收的公共帖子作为查询,对监控面板中所有可排名的信息流进行排名,以确定每个信息流是否会后续返回该帖子。我们构建了一个仍在完善的先收集后标注基准数据集。该收集的数据集涵盖了固定的5000个监控信息流面板,包含1780.4万条公共帖子、186.5万条可观测的帖子-信息流返回记录以及625083次有效的信息流轮询。标签记录了帖子在发布后24小时内的至少一次轮询中,是否在信息流的AppView前50个结果中被观测到。当前实验使用两个不重叠的24小时测试折,每个测试折搭配一个24小时训练窗口,且间隔24小时的结果可用期。评估以满足指标资格标准的602186条测试帖子为条件,这些帖子占全部6661658条测试帖子的9.04%。在两个折中,LambdaRank在评估模型中取得了最佳的同折均值:截断Recall@10为0.7361,NDCG@10为0.6127,Hit@10为0.7749。
英文摘要
The conventional approach to cold-start recommendation addresses new users or newly introduced items. Bluesky custom feeds create a different setting: independently operated feeds filter content from a shared public stream. In this setting, newly published posts are the cold-start objects, while the feeds serve as candidates. We propose a cold-start routing task in which a newly ingested public post is the query and all rankable feeds in the monitored panel are ranked according to whether each will subsequently return it. We build a still-evolving collect-first, label-later benchmark dataset. The collected dataset covers a fixed panel of 5,000 monitored feeds and contains 17.804 million public posts, 1.865 million observable post--feed return records, and 625,083 valid feed polls. The labels record whether a post is observed among a feed's AppView Top-50 results in at least one poll during the 24 hours after publication. The current experiments use two disjoint 24-hour test folds, each paired with a 24-hour training window and separated by a 24-hour outcome-availability gap. Evaluation is conditional on the 602,186 test posts that have at least one positive observed label and satisfy the metric eligibility criteria; these posts account for 9.04% of all 6,661,658 test posts. Across the two folds, LambdaRank achieves the best equal-fold mean values among the evaluated models: 0.7361 for capped Recall@10, 0.6127 for NDCG@10, and 0.7749 for Hit@10.