AI 中文总结
针对LLM不同表述问题的答案质量不均问题,提出约束混合策略GroupDRO框架选择系统提示,在双语医疗和消费金融基准上使LLM的最差损失平均降低13.7%。
AI 中文摘要
大型语言模型(LLM)越来越多地用于信息检索,但语义等效的不同表述问题可能会收到质量差异显著的答案。系统提示被广泛用于引导响应行为,但通常针对平均情况质量优化,因此部分问题表述仍可能收到不完整或低质量的答案。为解决该问题,我们提出一种用于系统提示选择的约束混合策略GroupDRO框架。该框架不优化系统提示文本,而是为现有系统提示池中各系统提示分配权重,以最小化评估指标和各组间的最坏情况信息质量损失,同时约束平均损失保持接近基于平均选择的损失。由于池生成与选择解耦,该方法适用于任何系统提示池,可利用互补系统提示的集成而非单个提示。在两个双语医疗和消费金融基准上对五个LLM开展实验,与无缓解措施相比,该约束方法平均降低总体均值、最差25%均值和最差损失分别达13.1%、13.2%和13.7%,同时保持总体质量接近平均选择水平。其多提示权重揭示了指标-组对间的互补性,代码和数据可在该https链接获取。
英文摘要
Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior, but they are typically optimized for average-case quality, so some question phrasings may still receive incomplete or low-quality answers. To address this, we formulate a constrained mixed-strategy GroupDRO framework for system-prompt selection. Instead of optimizing the system-prompt text, the framework assigns weights to system prompts in an existing pool to minimize the worst-case information-quality loss across evaluation metrics and groups, while constraining the mean loss to stay close to that of average-based selection. Because pool generation and selection are decoupled, the method applies to any system-prompt pool and can leverage an ensemble of complementary system prompts rather than a single one. Across five LLMs on two bilingual medical and consumer-finance benchmarks, the constrained method reduces the Overall Mean, Worst 25% Mean, and Worst by 13.1%, 13.2%, and 13.7% on average relative to no mitigation while keeping overall quality close to Average selection. Its multi-prompt weights reveal complementarity across metric-group pairs. Code and data are available at https://github.com/Rainxu09/equitable-system-prompt-selection.