AI 中文总结
本研究通过行为测试评估证据阅读诊断特征是否有助于选择小型LLM推荐接口,实验显示这些特征未显著提升NDCG@5,但相似用户证据可改善提示效果。
AI 中文摘要
行为测试衡量语言模型如何阅读证据。我们探究这些测量是否有助于选择推荐接口。我们在四个推荐领域评估了六个经过指令微调的小型检查点,采用时间顺序评估,涉及3,426名评估用户。每个请求对八个候选项目进行排名。基线选择器在仅历史提示、带协作证据的提示和分数融合之间进行选择。它使用可观察特征和六个稳定性提示,这些提示在措辞和候选顺序上有所变化。增强选择器添加了来自六个证据阅读提示的特征,这些提示要求模型比较支持计数。在开发(验证)数据上为每个领域和检查点选择一次接口,其NDCG@5得分为0.5524,而基线选择器为0.5447,增强选择器为0.5428。添加诊断特征使NDCG@5变化了-0.0019(95%置信区间[-0.0046, 0.0004])。该区间包含零,且其上界低于分析计划中0.005的改进目标。匹配选择器的超参数后,区间上界仍低于该目标。来自检索到的相似用户的证据比使用随机选择的、按活动匹配的用户的对照在提示上提高了0.0999 NDCG@5。证据阅读测试还揭示了答案位置和并列响应偏差。这些结果涉及所测试的选择器和候选集。它们说明了为什么诊断测量应通过它们是否在现有特征和固定接口之外改善推荐选择来评估。
英文摘要
Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate six small instruction-tuned checkpoints across four recommendation domains with chronological evaluation and 3,426 evaluation users. Each request ranks eight candidates. A baseline selector chooses among history-only prompting, prompting with collaborative evidence, and score fusion. It uses observable features and six stability prompts that vary wording and candidate order. An augmented selector adds features from six evidence-reading prompts that ask the model to compare support counts. An interface chosen once on development (validation) data for each domain and checkpoint scores 0.5524 NDCG@5, compared with 0.5447 for the baseline selector and 0.5428 for the augmented selector. Adding the diagnostic features changes NDCG@5 by -0.0019 (95% interval [-0.0046, 0.0004]). The interval includes zero, and its upper bound is below the analysis plan's 0.005 improvement target. Matching the selectors' hyperparameters also leaves the interval upper bound below that target. Evidence from retrieved similar users improves prompting by 0.0999 NDCG@5 over a control using randomly selected users matched for activity. The evidence-reading tests also reveal answer-position and tie-response biases. These results concern the tested selectors and candidate sets. They illustrate why diagnostic measurements should be evaluated by whether they improve recommendation choices beyond existing features and a fixed interface.