发表机构
Amazon Web Services(亚马逊网络服务)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出测试时增强(TTA),经匹配计算对比发现,中端大语言模型采用语义复述的TTA策略,比自洽性更高效地将计算转化为准确率,每美元准确率约为自洽性的1.8倍。
AI 中文摘要
测试时缩放可提升大语言模型(LLM)的准确率,但会增加推理成本,因此部署时每单位计算资源获得的准确率是关键指标。自洽性是已确立的方法之一,它将全部预算用于输出侧,通过采样重复推理路径实现。本研究测试时增强(TTA)扩展了自洽性,同时也对输入进行扰动,聚合输入变换版本的预测结果,并探究输入侧多样性是否比输出侧多样性更高效地将计算资源转化为准确率。我们开展了系统的匹配计算对比:在涵盖通用与多语言知识、数学推理、多模态问答及情感分类的六个数据集上,评估三种简单输入侧策略(语义复述、词汇扰动、视觉变换),并与思维链提示和自洽性对比。语义复述实现了一致且统计显著的准确率提升,同时在成本效益上帕累托优于自洽性,每美元获得的准确率约为1.8倍,且在六个任务中的五个上表现更优。我们进一步分析了增强数量、多模态策略及基础模型缩放,发现TTA在无法使用更强模型或成本过高的中端模型上成本效益最高。研究结果表明,对于当前中端LLM,改变输入比仅改变推理路径更高效地将推理计算转化为准确率。TTA实现代码可在该https URL获取。
英文摘要
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one of the established approaches, which spends this budget entirely on the output side by sampling repeated reasoning paths. We study Test-Time Augmentation (TTA), which extends self-consistency by also perturbing the input, aggregating predictions across transformed versions of the input, and ask whether input-side diversity converts compute into accuracy more efficiently than output-side diversity. We perform a systematic, matched-compute comparison: we evaluate three simple input-side strategies (semantic rephrasing, lexical perturbations, and visual transformations) across six datasets covering general and multilingual knowledge, mathematical reasoning, multi-modal question answering, and sentiment classification, against chain-of-thought prompting and self-consistency. Semantic rephrasing delivers consistent and statistically significant accuracy gains while Pareto-dominating self-consistency on cost-effectiveness, delivering roughly 1.8X more accuracy per dollar and outperforming it on five of six tasks. We further analyze the number of augmentations, multi-modal strategies, and base model scaling, finding that TTA is most cost-effective for mid-tier models where a stronger model is unavailable or too expensive. Our findings indicate that for current mid-tier LLMs, varying the input converts inference compute into accuracy more efficiently than varying the reasoning path alone. The TTA implementation is available at https://github.com/aws-samples/sample-genai-reflection-for-bedrock.
CommentsPublished at the COLM 2026 Workshop on Efficient Reasoning