arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STRIVE:探究分级合理性生成与评估中的推理极限

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

Bhiman Kumar Baghel, Anna Chrabaszcz, Tessa Warren, Michael Walsh Dickey, Haley C. Dresang, Xiang Lorraine Li

arXiv 2608.04567首次发表:更新:

发表机构

University of Pittsburgh; University of Wisconsin–Madison(匹兹堡大学; 威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STRIVE是基于LLM的框架,可联合生成评估跨越合理性类别与难度的受控事件集,实验中其优化后高质量事件集生成率达75.0%,为心理语言学研究减少手动工作量。

AI 中文摘要

事件知识涉及谁对谁做了什么。心理语言学家利用事件合理性判断,探究这类知识如何支撑人类语言加工。为分离合理性效应,这类研究需要受控事件集,其中一个事件槽在不同合理性水平间变化,而所有其他事件特征保持固定。手动构建此类集十分费力。因此,我们推出STRIVE,一种基于大语言模型(LLM)的框架,用于联合生成和评估跨越合理性类别(合理与不合理)及预期分类难度(易与难)的受控事件集。给定一个动词,STRIVE会构建一个共享事件框架,然后通过改变一个槽位同时固定其他所有槽位,为每个条件生成一个事件。在对60个动词、6个模型的实验中,使用基线生成提示时,GPT-5.1仅在16.7%的情况下生成高质量事件集;添加全局推理草稿板和评估器引导的优化后,该比例提升至75.0%。更多的推理努力也提高了评估器与人类判断的一致性。不过,接近合理性边界的事件仍然最具难度:它们引发最大的人类分歧,且最佳评估器在“不合理-难”条件下仅达到57%的准确率,表明仍需人类介入。总体而言,STRIVE提供了一种可扩展的方法,通过为心理语言学研究自动化初始事件集的生成与评估,减少手动工作量。

英文摘要

Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑