arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25717cs.CLcs.AI

当检索增强生成(RAG)无法实现均衡:上市公司事实问答中的地理偏差

When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies

Abhinav Havaldar, Enrico Santus

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对上市公司事实问答场景,构建含约2000家公司的基准,评估六个LLM在四种条件下的表现,发现RAG无法消除地理偏差,更大模型也无法消除结构性影响,挑战了RAG作为通用纠正手段的观点。

中文摘要 AI 辅助

检索增强生成(RAG)被广泛认为可减轻大语言模型(LLM)的事实错误,但目前尚不清楚检索是否能统一弥补知识缺失。我们在上市公司的受控事实问答场景中研究该问题,构建了涵盖全球股票指数约2000家公司的基准。我们在四种条件下针对四个基础属性评估六个LLM:无上下文、完美上下文、误导性上下文和干扰性上下文。我们发现无上下文准确率存在显著地理差异,表明参数知识分布不均。虽然完美上下文可提升性能,但无法消除这些差异:提升幅度与基线准确率相关,说明检索效果与内部表征耦合。在误导性上下文下,模型常复制错误信息。更大规模的模型整体性能更好,但无法消除这些结构性影响。这些结果对RAG作为通用纠正手段的观点提出挑战,凸显模型知识、上下文质量与实体表征之间的相互作用。

英文摘要

Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public companies, constructing a benchmark of approximately 2,000 firms across global equity indices. We evaluate six LLMs on four atomic attributes under four conditions: no-context, perfect context, misleading context, and distraction context. We find strong geographic disparities in no-context accuracy, indicating uneven parametric knowledge. While perfect context improves performance, it does not eliminate these gaps: gains are correlated with baseline accuracy, suggesting retrieval effectiveness is coupled to internal representations. Under misleading context, models frequently copy incorrect information. Larger models improve overall performance but do not remove these structural effects. These results challenge the view of RAG as a universal corrective and highlight the interaction between model knowledge, context quality, and entity representation.

发表机构

  • Bloomberg(彭博公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑