HexEval:一种用于多维学者评估的证据驱动六边形框架
HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
AI总结:
本研究提出HexEval框架,将学者评估构建为证据驱动的推理问题,通过内在与外部两个互补证据层评估,实现可解释可审计的AI辅助学者评估,同时指出公开学术数据的局限性。
AI中文摘要:
学者评估在教师招聘、经费分配、学术晋升和人才发现中发挥着基础性作用。现有学者评估方法主要依赖文献计量指标和声誉代理,而近期基于大语言模型(LLM)的方法大多聚焦于评估单篇研究论文,而非全面评估学者。我们认为,学者评估应被构建为一个证据驱动的推理问题,需同时考虑内在研究质量和外部可验证的学术行为。为此,我们提出HexEval,一种用于多维学者评估的证据驱动六边形框架。HexEval明确将学者评估组织为两个互补的证据层:内在层对匿名化的代表性作品沿三个维度进行评估,即研究严谨性、方法创新性和科学贡献;外部层则利用从GitHub、Lens、OpenAlex及其他公开可验证来源收集的异构证据,通过知识转化、研究连贯性和学术影响力来刻画学者。HexEval不生成不透明的综合分数,而是在整个评估过程中保留中间证据、维度特定的理由和验证信号,从而生成可解释、可审计的学者档案。对所有六个维度的实验显示出与人类或外部参考标准的维度依赖一致性:结构化校准提高了内在质量的绝对一致性,而外部模块则恢复了广泛的轨迹和序数影响力信号。这些结果支持对异构学术证据进行证据驱动推理,作为可审计的AI辅助学者评估的有前景范式,同时揭示了公开学术数据的覆盖范围和归因局限性。
英文摘要:
Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent large language model (LLM)-based approaches mainly focus on evaluating individual research papers rather than comprehensively assessing scholars. We argue that scholar assessment should be formulated as an evidence-driven reasoning problem that jointly considers intrinsic research quality and externally verifiable scholarly behavior. To this end, we propose HexEval, an evidence-driven hexagonal framework for multidimensional scholar assessment. HexEval explicitly organizes scholar assessment into two complementary evidence layers. The intrinsic layer evaluates anonymized representative works along three dimensions, namely research rigor, methodological innovation, and scientific contribution, whereas the external layer characterizes scholars through knowledge translation, research coherence, and academic impact using heterogeneous evidence collected from GitHub, Lens, OpenAlex, and other publicly verifiable sources. Instead of producing opaque aggregate scores, HexEval preserves intermediate evidence, dimension-specific rationales, and verification signals throughout the evaluation process, enabling interpretable and auditable scholar profiles. Experiments across all six dimensions show dimension-dependent agreement with human or external reference criteria: structured calibration improves absolute agreement for intrinsic quality, while the external modules recover broad trajectory and ordinal impact signals. These results support evidence-driven reasoning over heterogeneous scholarly evidence as a promising paradigm for auditable AI-assisted scholar assessment, while exposing the coverage and attribution limitations of public scholarly data.