arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用大语言模型自动生成经复杂度验证的决策场景

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

Abdalla Doleh, Toni Somers, Ratna Babu Chinnam

arXiv 2608.08822首次发表:更新:

发表机构

Wayne State University(韦恩州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对手动生成决策场景的缺陷,开发了基于大语言模型的自动化流水线,生成并验证了4238个多领域场景,证实其具备良好心理测量学特性,可用于AI系统认知评估。

AI 中文摘要

认知决策研究依赖于复杂度经过精心控制的多样化场景,但手动生成场景的过程缓慢、不一致且存在偏差。我们开发了一条自动化流水线,利用大语言模型(LLMs)生成结构化决策场景,并通过基于成熟任务复杂度理论的复合框架验证其复杂度。我们在多个领域和复杂度层级评估了4238个场景。测量验证符合严格的心理测量学标准。五个独立模型族之间的一致性几乎达到完美,组内相关系数为0.997,kappa值为0.971。已知群组效度显示各层级间存在较大差异,eta平方为0.587,所有成对比较的p值均小于0.001。因子分析揭示了一个主导的复杂度构念,在三个框架中的载荷介于0.87至0.96之间,而交互性形成一个较弱的次级维度,载荷为0.34。复杂度与文本长度之间存在强相关,在控制层级后仍持续存在,偏相关系数为0.86,这限制了构念纯度,但并未损害该工具的层级分级功能。模型分析显示,吞吐量与方案通过率呈负相关(r=-0.967,p=0.007,n=5),表明存在速度-质量权衡,不过这一关系主要由一个高吞吐量模型驱动。Llama 4 Maverick生成场景的速度最快,为每分钟134个,而DeepSeek Chat V3.2为每分钟25个;但Llama 4 Maverick生成的高复杂度层级场景不足,而DeepSeek Chat V3.2在领域覆盖与高方案合规性之间实现了平衡。该系统展现出强大的心理测量学特性,能够可靠地将场景分类为简单、中等和复杂层级,并为AI系统的下游认知评估提供所需的测量基础设施。

英文摘要

Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios and validates their complexity through a composite framework rooted in established task-complexity theory. We evaluated 4,238 scenarios across multiple domains and complexity tiers. Measurement validation met rigorous psychometric standards. Agreement among five independent model families was nearly perfect, with an intraclass correlation coefficient of 0.997 and a kappa of 0.971. Known-groups validity demonstrated large separation between tiers, with an eta-squared of 0.587 and all pairwise comparisons significant at p less than .001. Factor analysis revealed a dominant complexity construct, with loadings between 0.87 and 0.96 across three frameworks, while interactivity formed a weaker secondary dimension at 0.34. Discriminant validity was limited by a strong relationship between complexity and text length that persisted after controlling for tier, yielding a partial correlation of 0.86. This constrains construct purity but does not undermine the instrument's tier-grading function. Model analyses showed a negative association between throughput and schema pass rate (r = -0.967, p = .007, n = 5), suggesting a speed-quality trade-off, though largely driven by one high-throughput model. Llama 4 Maverick generated scenarios fastest at 134 per minute versus 25 for DeepSeek Chat V3.2, but underproduced complex-tier scenarios, whereas DeepSeek Chat V3.2 balanced domain coverage with high schema compliance. The system demonstrated strong psychometric properties, enabling reliable classification into Simple, Moderate, and Complex tiers and providing the measurement infrastructure needed for downstream cognitive assessment of AI systems

Comments38 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑