arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可见推理并非通用优化器:分析代码生成中依赖角色与思维的效应

Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation

Bhawani Shankar Leelar, Pawan Chorasiya, Davin Hill, Robert E. Tillman, Tamer Soliman

arXiv 2610.10639首次发表:更新:

发表机构

Optum AI(奥普腾人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对SQL-pandas基准发现可见推理无通用优势,其效应依赖角色、目标等因素,提出需结合多维度选择推理策略,还提供了评估可见推理作用的受控框架。

AI 中文摘要

可见思维链(CoT)常被视为广泛有用的推理指令,但分析代码生成结合了自然语言歧义、模式接地、目标语言约束及模型特定推理行为。由于同一分析请求可通过SQL和Python(pandas)两种不同目标语言表达,该场景为检验一项常见却未充分研究的假设提供了自然测试:当推理表示与请求目标匹配时,可见推理更有效,如“用SQL思考”或“用Python思考”。这类建议与“逐步思考”等通用指令一起,在受控的基于执行的比较下仍未得到充分评估。我们研究了一个查询匹配的SQL-pandas基准,该基准跨越角色措辞、目标语言、可见CoT格式、控制前缀、直接生成及内部推理配置。结果不支持可见CoT具有通用准确性优势,也不支持推理表示与目标语言匹配具有一致益处。相反,效应取决于角色、目标、模型配置及内部推理设置。控制消融进一步区分了推理内容与提示格式的效应。这些发现表明,应针对模型、角色、目标及内部推理配置共同选择推理策略,而非将其作为通用默认采用。更广泛而言,该研究提供了一个受控框架,用于识别可见推理何时改进可执行生成、何时主要扰动模型行为,以及何时内部推理配置是更关键的因素。

英文摘要

Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior. Because the same analytics request can be expressed in two distinct target languages-SQL and Python (pandas)-this setting provides a natural test of a common but under-examined assumption: that visible reasoning is more effective when its representation matches the requested target, as in "think in SQL" or "think in Python." Together with generic instructions such as "think step-by-step," such recommendations remain insufficiently evaluated under controlled, execution-based comparisons. We study a query matched SQL-pandas benchmark that crosses persona phrasing, target language, visible-CoT format, control prefixes, direct generation, and internal-reasoning configurations. The results do not support either a universal accuracy advantage from visible CoT or a consistent benefit from matching the reasoning representation to the target language. Instead, the effects depend on the persona, target, model configuration, and internal-reasoning setting. The control ablations further distinguish effects of reasoning content from those of prompt format. These findings indicate that reasoning strategies should be selected jointly for the model, persona, target, and internal-reasoning configuration rather than adopted as universal defaults. More broadly, the study provides a controlled framework for identifying when visible reasoning improves executable generation, when it primarily perturbs model behavior, and when the internal-reasoning configuration is the more consequential factor.

Comments26 pages, 8 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑