arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AfriSyCo:衡量非洲语言内容周围的断言性框架、验证与措辞敏感性

AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content

David Ababio Awuni, Rose-Mary Owusuaa Mensah Gyening, Elvis Gyasi Owusu

arXiv 2609.17853首次发表:更新:

AI 中文总结

AfriSyCo通过母语与跨语言因子实验,测量断言性框架和验证对非洲语言内容答案切换的影响,发现断言显著增加错误目标选择,且效应受措辞和模型检查点强烈调节。

AI 中文摘要

AfriSyCo通过两个互补层面研究非洲语言事实内容周围的答案切换:母语后续提问和受控的跨语言因子设计,其中问题、选项和目标均保留非洲语言,而后续框架为英语。我们分析了从100个源问题中得出的1,415个首轮正确(turn-1-correct)的模型-语言-项目观测,涉及七个开放权重检查点和六种语言;首轮正确表示观测到的首次回答准确率,而非已证实的知识。在母语提示下,断言性认可(assertive endorsement)导致的任意轮次错误目标选择比提及加验证(M+V)高出29.3个百分点,即时T2对比为19.0个百分点。在预先承诺的2x2因子设计中,平均三种测试提示族,断言性框架使目标选择增加30.4个百分点(95%置信区间[28.4, 32.3]);验证使其减少17.4个百分点,而断言效应从无验证时的20.5个百分点上升到有验证时的40.2个百分点(交互作用+19.7)。在选项重新排序后正确的611个观测中,效应仍为34.8个百分点。效应大小随措辞和检查点变化显著:提示族效应范围为20.1-42.5个百分点,Twi/Qwen3释义将目标选择从70.8%变为4.2%,检查点效应范围为9.2-47.0个百分点。因此,提示实现是测量问题的一部分。

英文摘要

AfriSyCo studies answer switching around African-language factual content with two complementary layers: native-language follow-ups and a controlled cross-language factorial whose question, options, and target remain in the African language while the follow-up framing is English. We analyze 1,415 turn-1-correct model-language-item observations derived from 100 source questions across seven open-weight checkpoints and six languages; turn-1-correct denotes observed first-response accuracy, not demonstrated knowledge. Under native prompts, assertive endorsement produces 29.3 percentage points more any-turn false-target selection than mention-plus-verification (M+V), with a 19.0-point immediate T2 contrast. In the precommitted 2 x 2 factorial, averaged over three tested prompt families, assertive framing increases target selection by 30.4 points (95% CI [28.4, 32.3]); verification decreases it by 17.4 points, while the assertive effect rises from 20.5 points without verification to 40.2 with it (interaction +19.7). The effect remains 34.8 points among 611 observations correct after option reordering. Magnitude varies sharply by wording and checkpoint: prompt-family effects span 20.1-42.5 points, a Twi/Qwen3 paraphrase shifts target selection from 70.8% to 4.2%, and checkpoint effects span 9.2-47.0 points. Prompt realization is therefore part of the measurement problem.

Comments13 pages, 1 figure, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑