GRASP:用于无标签多主题作文评分的图检索自动评分流水线
GRASP: Graph-Retrieval Automated Scoring Pipeline for Label-Free Multi-Topic Essay Grading
- Melbourne Institute of Technology(墨尔本技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出GRASP流水线,针对无标签多主题科学考试,结合RAG与GRAG检索,用匈牙利算法分配参考节点,通过GPT-4.1-mini评分,验证图增强检索的优势。
AI中文摘要:
自动简答题评分研究以往仅聚焦于由单一主题问题构成的考试,包含多主题问题的考试自动评分仍较少被探索。本研究引入了一种用于无标签多主题科学考试评分的图检索自动评分流水线(GRASP)。无标签考试是指学生对多个不同主题的回答被合并为单个段落,无标记标签或分段指示哪一部分对应哪个问题的简答题考试。每个问题的参考答案通过 Sentence-BERT 编码为 FAISS 向量索引,并在该参考答案集合上构建语义相似度图。评分阶段,首先采用句子数量启发式方法(结合大语言模型解决模糊案例)预测学生作文中回答的不同主题数量,该过程无需训练数据或特定领域示例作文。随后通过基于余弦相似度的检索增强生成(RAG)和图检索增强生成(GRAG)检索候选参考节点,每个节点存储参考索引中的一个(问题、参考答案、两者拼接内容)。GRAG 的操作方式为:取余弦相似度最高的节点作为种子节点,再通过强边的图遍历查找 RAG 可能遗漏的其他参考节点。接着使用匈牙利算法为每个问题段最优分配一个参考节点,确保无重复参考。最后使用 GPT-4.1-mini 对每个段与其分配的参考独立评分。本实验旨在展示检索质量对评分准确率的影响,以及在不同作文复杂度下,图增强检索相较于严格余弦相似度方法的优势。
英文摘要:
Automated short-answer grading research has historically focused on exams consisting solely of questions pertaining to a single topic. Automatic grading of exams containing questions about more than one topic remains less explored. In this work, a Graph-Retrieval Automated Scoring Pipeline (GRASP) is introduced for grading label-free multi-topic science exams. Label-free exams are short-answer exams in which a student's responses to several distinct topics are merged into a single paragraph, with no markup labels or segmentation indicating which span answers which question. Reference answers for each question are encoded into a FAISS vector index via Sentence-BERT, and a semantic similarity graph is constructed over this set of reference answers. At grading time, sentence count heuristics, with a large language model used to resolve ambiguous cases, are first applied to predict how many distinct topics were answered in the student essay. This process is performed without training data or domain-specific example essays. Candidate reference nodes, each storing one (question, reference answer, concatenation of both) from the reference index, are then retrieved through cosine similarity based Retrieval-Augmented Generation (RAG) and Graph Retrieval-Augmented Generation (GRAG). GRAG operates by taking the top cosine matches as seed nodes and then performing a graph traversal over strong edges to find additional reference nodes that may have been missed by RAG. The Hungarian algorithm is then used to optimally assign one reference node per question segment such that no reference is duplicated. Each segment is then graded against its assigned reference independently using GPT-4.1-mini. This experiment is performed to show the effect of retrieval quality on grading accuracy and the benefit of graph-augmented retrieval versus strict cosine similarity methods at various levels of essay complexity.