arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09743cs.AIcs.GT

为策略制定者搭建框架:霍特林空间市场中与架构相关的推理干预

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

发表机构加州理工学院
查看机构详情
  • California Institute of Technology(加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Pratyush Singh

首次发表
浏览论文内容

中文总结 AI 辅助

研究结构化推理干预对大语言模型战略经济推理的影响及与模型架构的关系,以霍特林模型评估GPT - 4.1 - mini和GPT - 5 - mini,发现支架类型与模型架构有交叉交互作用,对抗性测试有损害,还存在陈述性 - 程序性差距。

中文摘要 AI 辅助

我们研究了结构化推理干预是否能改善大语言模型的战略经济推理,以及其效果是否依赖于模型架构。以霍特林线性城市模型作为诊断工具,我们在五个条件下评估了GPT - 4.1 - mini(标准指令跟随模型)和GPT - 5 - mini(推理优化模型),包括无支架基线和四种推理干预,涵盖八个演绎和归纳推理问题、三种提示框架,每个条件重复三次,共得到720个单独判断的回答。我们发现支架类型和模型架构之间存在统计学上显著的交叉交互作用。承诺支架提高了标准模型性能但降低了推理模型性能,原则性分离则相反。对抗性压力测试对两个模型都有损害,且对推理模型损害更大。此外,两个模型都存在陈述性 - 程序性差距,分离完全弥合了推理模型的这一差距,而没有干预措施能帮助标准模型。

英文摘要

We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture. Using Hotelling's linear city model as a diagnostic vehicle, we evaluate GPT-4.1-mini (a standard instruction-following model) and GPT-5-mini (a reasoning-optimized model) under five conditions - an unscaffolded baseline and four reasoning interventions - across eight questions spanning deductive and abductive reasoning, three prompt framings, and three repetitions per condition, yielding 720 individually judged responses. We find a statistically significant crossover interaction between scaffolding type and model architecture ($t(7) = 4.79$, $p = 0.002$, $d = 1.69$): commitment scaffolding improves the standard model ($+0.21$) while degrading the reasoning model ($-0.63$), and principled separation shows the opposite pattern ($-0.40$ vs. $+0.31$). Both crossovers are individually significant (commitment: $p = 0.040$; separation: $p = 0.002$) and hold across all eight questions with 7/8 directional consistency. Adversarial stress-testing harms both models, with $2.6\times$ greater degradation for the reasoning model ($-1.47$ vs. $-0.57$; $p = 0.038$), and the damage correlates negatively with baseline difficulty ($R^2 = 0.36$, $p = 0.014$). We further document a persistent declarative-procedural gap in which both models identify correct strategies at rates far exceeding their ability to execute them; separation fully closes this gap for the reasoning model while no intervention helps the standard model.

补充信息

↑