AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation
AuthorityBench: 评估LLM权威感知以实现可靠的检索增强生成
AI总结 本文提出AuthorityBench基准,通过三个数据集评估LLM权威感知能力,发现ListJudge和PairJudge方法与真实权威最相关,且权威感知对检索增强生成的准确性有显著提升。
Comments 11 pages, 4 figures. Submitted to ACL 2026