LigBench:面向基于大语言模型的研究创意生成的统一且符合人类偏好的基准
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
浏览论文内容
中文总结 AI 辅助
该研究针对现有LLM研究创意生成评估零散、缺乏统一标准的问题,提出LigBench基准及PAIR-IQ数据集,实验表明LigBench评估稳定可解释且与专家判断一致性高,PAIR-IQ训练的模型排名准确性和鲁棒性更强
中文摘要 AI 辅助
随着大语言模型(LLMs)的快速发展,研究创意生成受到越来越多的关注。现有方法可让LLMs检索相关文献并为研究领域提出新颖创意。然而,当前创意生成的评估实践仍较为零散,缺乏客观标准,通常依赖直接的LLM评分,这限制了其在生成创意的连贯分布范围内提供统一可靠评估的能力。为应对这一挑战,我们提出了LigBench,这是一个自动化评估基准,可对AI研究创意进行细粒度且可靠的评估,在不同生成分布下均能一致适用。此外,我们引入了PAIR-IQ,这是一个专门用于训练成对创意判断模型的数据集,可作为辅助参考以支持更客观的对比评估。大量实验表明,LigBench能实现稳定且可解释的评估,显著提升与专家判断的一致性。此外,在PAIR-IQ上训练的模型表现出更强的排名准确性和鲁棒性,为可扩展且客观的研究创意评估建立了有原则的标准。
英文摘要
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Innovation Institution(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。