发表机构
ThakiCloud(ThakiCloud)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
对韩语27B模型进行响应风格后训练,发现其脱靶效应主要影响回答倾向而非内容偏好,并通过对照实验揭示目标文本的作用。
AI 中文摘要
我们对 Qwen3.8-27B 进行了韩语响应风格的后训练,涵盖冗长程度、列表和 Markdown 的使用、话语结构及语域,并测量了目标从未涉及的两个行为:在 KoBBQ 中对模糊社会问题的弃权(不执行)(基准正确答案为 UNKNOWN),以及在证券指引中未经提示的披露。两者均发生变化,且变化主要通过模型的输出策略体现:回答的频率和内容的多少。匹配的目标形式对照表明,回答倾向取决于训练目标,而非仅由提示集或配方决定。在保持提示、配方、数据量和部署不变、仅更改目标文本的情况下,三个风格种子给出正的回答率点估计(平均 +0.82 个百分点),三个中性种子给出负估计(平均 -1.53 个百分点);观察到的种子范围不重叠,均值相差 2.34 个百分点。一个长度匹配的臂介于两者之间,而第四个在保持简短的同时保留对冲的臂在各种子间不稳定,因此形式中哪个特征起作用仍未解决。对于绝对刻板印象暴露,分解为回答倾向项和条件构成项是代数恒等式,而非发现;其经验内容在于变化发生的位置。在训练后的检查点中,变化主要由回答倾向主导,而构成项保持较小,且由于该项是在依赖处理的已回答子集上评估的,我们不将其视为潜在偏好的证据。由此得出两个测量结果。当回答状态依赖处理时,条件刻板印象份额的臂间对比无法识别条件内容偏好的变化。并且,对于同一构念的两个规则检测器之间的一致性,根据生成文本的检查点不同,范围从 0.44 到 0.99——无需任何参考标签即可观察到。
英文摘要
We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says. Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved. For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference. Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.
Comments20 pages. Korean-language evaluation (KoBBQ); uncertainty estimates over KoBBQ items are clustered on the benchmark template. v2: narrows the model-identity assertion to the KoBBQ axis, states that LoRA initialisation is unseeded, and discloses a scorer defect (adjacent-letter completions scored as the first option; incidence unrecoverable from stored outputs); numbers unchanged