arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

任务能力并非遵循指令:评估小语言模型中的指令冲突行为

Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

Mahdiyeh Farajidizaji, Vatsal Raina

arXiv 2607.19608首次发表:更新:

发表机构

Khajeh Nasir Toosi University of Technology; Apta AI, Spark AI Research(哈杰·纳西尔·图西理工大学; 阿普塔人工智能公司,星火人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究小语言模型在指令冲突时的行为,通过跨任务设计,用标准准确率等指标评估指令微调的Qwen模型,发现任务能力提升不自动带来可靠行为控制,任务能力与指令遵循是不同能力,仅报标准准确率会掩盖问题。

AI 中文摘要

指令微调旨在使语言模型遵循用户请求,但尚不清楚当指令与其通常的任务行为冲突时,小模型是否会遵守。我们通过将标准指令与冲突的非标准指令(选择错误选项、输出相反情感或返回答案的两倍)配对,在多项选择题问答(MCQA)、情感分类和数学问答这三项任务中进行研究。这种跨任务设计使我们能够测试对冲突指令的抗性是否与特定任务特征相关,还是反映了更广泛的行为倾向。由于所有预测都根据原始真实情况评分,忽略非标准指令的模型看起来仍然准确。我们使用标准准确率、非标准准确率和指令遵循失败率(IFFR)来评估不同规模的指令微调Qwen模型。标准准确率和指令遵循通常随规模提高,但模式在所有任务和数据集上并不一致。小模型保持能力但经常忽略非标准指令,而大模型在两种设置之间存在明显差距。这些发现表明任务能力的提升并不能自动可靠地控制模型行为。任务能力和指令遵循是不同的能力,仅报告标准准确率会掩盖指令遵循失败情况。

英文摘要

Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts with their usual task behavior. We study this across three tasks - multiple-choice question answering (MCQA), sentiment classification, and mathematical question answering - by pairing a standard instruction with a conflicting non-standard one (select an incorrect option, output the opposite sentiment, or return twice the answer). This cross-task design allows us to test whether resistance to conflicting instructions is tied to specific task characteristics or reflects a broader behavioral tendency. As all predictions are scored against the original ground truth, a model that ignores the non-standard instruction still appears accurate. Using standard accuracy, non-standard accuracy, and an Instruction-Following Failure Rate (IFFR), we evaluate instruction-tuned Qwen models across sizes. Both standard accuracy and instruction following generally improve with scale, although the pattern is not consistent across all tasks and datasets. Small models stay competent yet routinely ignore the non-standard instruction, while larger models show a clear gap between the two settings. These findings suggest that gains in task capability do not automatically provide reliable control over model behavior. Task competence and instruction following are therefore distinct abilities, and reporting only standard accuracy hides instruction-following failures.

Comments12 pages, 4 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑