提示工程对用于安全查询路由的小型语言模型的影响
Influence of Prompt Engineering on Small Language Models for Guarded Query Routing
浏览论文内容
中文总结 AI 辅助
研究安全查询路由问题,探究紧凑型开放权重小型语言模型在延迟约束下处理任务的能力。通过实验评估发现中等规模模型在低延迟下接近前沿,提示优化技术可提升模型表现,是安全查询路由的有用第一步,部分弱模型仍需其他调整。
中文摘要 AI 辅助
我们研究安全查询路由问题,即假设用户查询首先遇到一个路由器,它要么为分布内查询确定理想端点,要么拒绝可能不安全或超出系统范围的分布外查询。我们研究紧凑型开放权重小型语言模型(SLMs)能否在延迟约束下共同处理这两项任务。我们在GQR-Bench上评估了22个模型,并以分布内和分布外准确率的调和平均值对它们进行评分。我们发现中等规模的SLMs在低得多的延迟下接近前沿模型的路由质量。然而,许多紧凑型模型失败是因为它们不能可靠地遵循所需的输出格式。不过,我们的结果表明,提示优化技术能使SLMs优雅地处理此类情况,而无需更改模型权重。此外,少样本提示优化将Mistral 7B的GQR分数从81.79提高到90.87,并将Qwen3.5 9B提高到95.74,这是我们研究中最佳的优化分数,与最强的未优化更大模型Gemma 3 27B的96.01相差不到0.3分。对于Granite 4 Tiny,无上下文示例的裸DSPy签名是最有效的策略,将其分数从54.29提高到83.05。这些结果表明,提示优化是安全查询路由的有用第一步,而较弱的模型可能仍需要权重级别的调整或模式感知训练。
英文摘要
We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in-distribution queries or rejects out-of-distribution queries that are potentially unsafe or out of the system's scope. We investigate whether compact open-weight Small Language Models (SLMs) can jointly handle both tasks under latency constraints. We evaluate 22 models on GQR-Bench and score them with the harmonic mean of in-distribution and out-of-distribution accuracy. We find that mid-scale SLMs come close to frontier model routing quality at much lower latency. Still, many compact models fail because they do not reliably follow the required output format. However, our results show that prompt optimization techniques enable SLMs to handle such cases gracefully, without changing the models' weights. Moreover, few-shot prompt optimization raises Mistral 7B from 81.79 to 90.87 GQR-Score and lifts Qwen3.5 9B to 95.74, the best optimized score in our study and within 0.3 points of the strongest unoptimized larger model: Gemma 3 27B at 96.01. The bare DSPy signature, without in-context exemplars, is the most effective strategy for Granite 4 Tiny, raising its score from 54.29 to 83.05. These results show that prompt optimization is a useful first step for guarded query routing, while weaker models may still need weight-level adaptation or schema-aware training
发表机构
- Faculty of Informatics and Information Technologies, Slovak University of Technology in Bratislava(布拉迪斯拉发斯洛伐克理工大学信息与信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。