AI 中文总结
研究大语言模型个性化中监督微调与上下文学习的选择问题,通过构建框架分析用户面临的统计 - 经济权衡,得出不同方法占优情况、资源消耗特性及平台策略等结论,实验与平台调研验证了相关理论。
AI 中文摘要
大语言模型(LLMs)变革了人工智能服务,但出现了关键矛盾:个性化虽提升模型性能,却消耗用户须共享的稀缺计算资源。用户何时应投资昂贵的监督微调(SFT)而非轻量级的上下文学习(ICL)?其他用户的个性化选择造成的拥塞如何重塑这些激励?平台提供多种个性化算法时应采取什么策略?我们为大语言模型服务开发了一个易于处理的框架,捕捉用户面临的统计 - 经济权衡。分析得出几个惊人见解。首先,ICL和SFT在不同情况下占优,由预训练覆盖范围和数据信噪比的相互作用决定,但拥塞会改变这些排名。其次,均衡资源消耗呈现明显的非单调性:提高预训练精度可减少拥塞,而更广泛的预训练覆盖范围和更难的任务有时会增加拥塞。第三,我们证明提供两种个性化方法绝不会损害平台的最大利润,尽管可能增加计算负载。对GPT - 2进行的线性回归任务实验验证了我们关于算法性能的理论预测。对21个主要人工智能平台文档的审查表明,提供SFT和ICL的平台份额从2021年的9.5%增至2025年的71.4%,与我们的平台设计启示一致。
英文摘要
Large Language Models (LLMs) have revolutionized AI services, but a critical tension emerges: while personalization improves model performance, it consumes scarce computational resources that users must share. When should a user invest in expensive Supervised Fine-Tuning (SFT) versus lightweight In-Context Learning (ICL)? How does congestion from other users' personalization choices reshape these incentives? And what strategies should platforms adopt when offering multiple personalization algorithms? We develop a tractable framework for LLM serving that captures the statistical-economic trade-offs users face. Our analysis yields several surprising insights. First, we show that ICL and SFT dominate in different regimes, determined by an interplay between pretraining coverage and data signal-to-noise ratios, but congestion can flip these rankings. Second, equilibrium resource consumption exhibits pronounced non-monotonicity: improving pretraining precision reduces the congestion, while broader pretraining coverage and harder tasks sometimes increase it. Third, we prove that offering both personalization methods never hurts the platform's maximal profits, despite potentially increasing computational load. Experiments with GPT-2 on linear regression tasks validate our theoretical predictions about algorithm performance. Complementing these results, our review of documentation from 21 major AI platforms shows that the share offering both SFT and ICL increased from 9.5% in 2021 to 71.4% in 2025, consistent with our platform-design implications.