发表机构
Accentrust; Georgia Institute of Technology; University of Illinois Urbana-Champaign(Accentrust; 佐治亚理工学院; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一项配对跨语言协议,分离输入语言、痕迹语言与答案实现,通过47,520次核心运行和多项假设检验,严谨评估大语言模型的标记与推理效率。
AI 中文摘要
大语言模型会产生依赖语言的表示和推理成本,但现有的比较往往混淆了输入语言、分配的可见痕迹语言和答案实现方式。我们指定了一项前瞻性配对研究,在保持语义项目、检查点和答案预言机固定的同时,分离这些接口。初始设计实例化了240个精确计分的项目,这些项目由英语和七种非英语语言的模板渲染而成,涉及三个不同谱系的开源权重检查点、三个痕迹标记预算、22种输入和痕迹语言条件、一个固定的答案储备以及一个单独计数的分隔符:在前瞻性样本量选择之前,共进行47,520次初始核心运行。RQ1-RQ3通过意向性治疗分析,采用包含失败在内的终端会计方法估计输入和痕迹效应,并通过将密封前缀和运行时原生KV状态克隆到八个交叉分支来测试答案实现。冻结前的独立语言审查、固定形式的ASCII选择器以及代码/表面/求解器一致性约束了实现测试。H1-H5共享一个Holm家族和一个全局同步分量带。一项次要的随机实验在冻结种子和预言机盲聚合下,比较了在相等痕迹配额下的一次长尝试和完整的K次短尝试策略;由于答案容量不同,这是一项完整策略对比。结果包括精确标记跨度、正确性、延迟、运行时暴露的内存,以及在预设硬件轨道上合格的同主机操作系统报告的能量。该协议分离了分词器扩展、可见痕迹成本和答案实现成本,而不将可见痕迹视为内部认知或将操作系统估计视为物理跨设备能量。未报告任何确认性模型结果;结果字段在冻结证据账本通过独立验证之前保持禁用。
英文摘要
Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these interfaces while holding the semantic item, checkpoint, and answer oracle fixed. The initial design instantiates 240 exactly scored items rendered from templates in English and seven non-English languages, three distinct-lineage open-weight checkpoints, three trace-token budgets, 22 input- and trace-language conditions, a fixed answer reserve, and a separately counted delimiter: 47,520 initial core runs before prospective sample-size selection. RQ1-RQ3 estimate input and trace effects by intention-to-treat with failure-inclusive terminal accounting and test answer realization by cloning a sealed prefix and runtime-native KV state into eight crossed branches. Pre-freeze independent language review, fixed-form ASCII selectors, and code/surface/solver agreement constrain the realization test. H1-H5 share one Holm family and a global simultaneous component band. A secondary randomized experiment compares one-long-attempt and complete K-short-attempt policies at equal trace allowance under frozen seeds and oracle-blind aggregation; it is a full-policy contrast because answer capacity differs. Outcomes include exact token spans, correctness, latency, runtime-exposed memory, and qualified same-host operating-system-reported energy over prespecified hardware rails. The protocol separates tokenizer expansion, observable-trace cost, and answer-realization cost without treating visible traces as internal cognition or operating-system estimates as physical cross-device energy. No confirmatory model outcome is reported; result fields remain disabled until the frozen evidence ledger passes independent verification.
Comments32 pages, 2 figures, 3 tables. Prospective paired study protocol; no confirmatory model outcomes are reported