arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32042cs.CL

量化阈值复现,失败模式不复现:一项从8位到2位的波兰语智能体工具使用三模型研究

Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit

Jakub Prejzner

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过PolAgentBench基准测试,探究GGUF量化对波兰语智能体工具使用的影响,发现所有模型在3位到2位精度间性能骤降,但失败模式各异,且量化阈值具有跨模型复现性。

中文摘要 AI 辅助

我们探究GGUF量化如何影响波兰语中的智能体工具使用,以及这种影响是否在不同模型间具有普遍性。我们引入了PolAgentBench,一个确定性的基准测试,包含波兰语提示和英语工具模式:一个67任务的主套件(15个对抗性探针,52个困难层级任务)和一个46任务的算术隔离阶梯。三个模型跨越两个变化轴:Bielik-11B-v3.0及其剪枝、蒸馏的子模型Bielik-Minitron-7B-v3.0隔离了模型压缩的影响,而Llama-PLLuM-8B则引入了预训练家族的改变。每个模型在六种精度下进行测量,从Q8_0到Q2_K。只有崩溃阈值是复现的。(1)所有三个模型在3位和2位之间急剧下降(11B从0.716降至0.045,7B从0.463降至0.149,PLLuM从0.224降至0.015;每对配对McNemar检验p<0.001),跨越四倍的能力差异和两个轴。(2)失败模式不复现:在2位时,7B模型失败于长序列(中位数9.1k个令牌,4步),而11B模型大多在第一步就给出虚构的最终答案(64次失败中有37次);PLLuM模型在各精度下都在内容上失败(71.8-92.0%的步骤可解析)。(3)在算术阶梯上,无脚手架的一级是底线,在一次声明无工具规则的重新运行中保持不变;四次显式调用将8位11B从1/10提升至9/10,将7B从0/10提升至7/10(在原谅了仅格式错误且具有黄金值的失败后),这是一个探索性效应,具有八个不同的基线输入,但未通过多重性校正,而顺序陷阱分支在8位时区分了模型(11B 6/6,7B 0/6)。(4)波兰语与英语之间的差距与退化或任务家族相关。我们记录了四个影响我们结论的伪影(舍入不友好的黄金值、严格的答案类型、提示中从未声明的无工具规则、优先级排序的失败标签),以严格和修正的形式报告受影响的结果,并发布基准、轨迹和提交戳记的伪影。

英文摘要

We ask how GGUF quantization affects agentic tool use in Polish and whether the effects generalize across models. We introduce PolAgentBench, a deterministic benchmark with Polish prompts and English tool schemas: a 67-task main suite (15 adversarial probes, 52 hard-tier tasks) and a 46-task arithmetic isolation ladder. Three models span two axes of variation: Bielik-11B-v3.0 and its pruned, distilled child Bielik-Minitron-7B-v3.0 isolate model compression, and Llama-PLLuM-8B adds a change of pretraining family. Each is measured at six precisions, Q8_0 to Q2_K. Only the collapse threshold replicates. (1) All three models fall off a cliff between 3-bit and 2-bit (11B 0.716 to 0.045, 7B 0.463 to 0.149, PLLuM 0.224 to 0.015; paired McNemar p < 0.001 in each), across a fourfold capability spread and both axes. (2) Failure modes do not replicate: at 2-bit the 7B fails long (median 9.1k tokens, 4 steps) while the 11B mostly answers at the first step with a confabulated final answer (37 of 64 failures); PLLuM fails on content across precisions (71.8-92.0% of steps parse). (3) On the arithmetic ladder the unscaffolded rung is a floor, left standing by a rerun that states the no-tool rule; four explicit calls lift the 8-bit 11B from 1/10 to 9/10 and the 7B from 0/10 to 7/10 after format-only failures with the gold value are forgiven, an exploratory effect with eight distinct baseline inputs that does not survive multiplicity correction, while the order-trap arm separates the models at 8-bit (11B 6/6, 7B 0/6). (4) The Polish-versus-English gap is associated with degradation or with task family. We document four artifacts that shaped our conclusions (rounding-hostile gold values, strict answer typing, a no-tool rule the prompt never stated, priority-ordered failure labels), report affected results in strict and corrected form, and release the benchmark, trajectories and commit-stamped artifacts.

补充信息

↑