WARP:用于群体意见的Wasserstein对齐检索增强生成(RAG)
WARP: Wasserstein-Aligned RAG for Population Opinions
浏览论文内容
中文总结 AI 辅助
WARP是一种用于群体意见的检索后算法,通过Wasserstein-1距离校准检索证据,在多领域实验中显著降低分布误差,提升RAG系统生成答案的偏好度。
中文摘要 AI 辅助
检索增强生成(RAG)系统越来越多地用于总结大量文档集合中的内容。用户提问“人们对X有什么看法?”,会得到一个看似代表共识的答案。但标准的Top-K检索根据查询相似度对文档进行排名,而非根据它们代表群体的忠实度,因此少数群体观点会悄然消失。现有解决方案存在不足:MMR和DPP等多样性重排器会将检索到的文档分散开,但没有目标分布作为优化方向;基于KL或JS散度的校准方法虽有目标,却将意见区间视为无序,混淆强烈正面与强烈负面的代价与相邻区间的错误相当。我们提出WARP,这是一组检索后算法,用于将检索到的证据校准至群体意见分布。WARP首先恢复余弦排名可能掩盖的代表性不足的意见,随后使用Wasserstein-1距离选择情感强度分布与群体目标匹配的文档,捕捉KL和JS散度忽略的序数结构。我们针对密集、稀疏和可变候选池开发了三种变体,在校准质量和速度间进行权衡。在涵盖35000份文档、156个查询和26个实体的三个评论领域中,WARP的领域匹配变体将分布误差降低了至少43%,延迟低于1秒。这些优势延续到生成环节:当k≤5时,由五个大模型组成的评审小组在86%的已判定比较中更偏好WARP生成的答案。
英文摘要
RAG systems are increasingly used to summarize what large collections of documents say. A user asks "What do people think about X?" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.