arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28657cs.AI

TAPR:利用任务感知提示重写器提升大语言模型性能

TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter

  • University of Amsterdam(阿姆斯特丹大学)
  • Elsevier(爱思唯尔)

机构由 AI 辅助整理,请以论文原文为准。

Oliver Savolainen, Emanuele Bastianelli, Hosein Azarbonyad

AI总结:

本研究提出TAPR模型,采用带GRPO的强化学习训练,可将用户提示优化为任务适配的提示,经微调Phi-4-mini-instruct后,在多任务基准测试中提升了大语言模型的准确率。

AI中文摘要:

大语言模型(LLMs)通常需要精心设计的提示才能发挥全部潜力,这对非专业用户构成了障碍。本研究通过引入任务感知提示重写器(TAPR)来解决这一挑战,该模型可将用户提示重新表述为针对任务优化的提示,明确目标是提升下游大语言模型的性能。我们采用带组相对策略优化(GRPO)的强化学习训练TAPR,奖励值来自对重新表述的提示及对应任务输出的大语言模型评判评估。在问答、摘要、算术推理等不同任务上的实验结果表明,我们的方法在提示重写能力上相比基础模型取得了持续提升。对Phi-4-mini-instruct(作为TAPR的基础模型)进行微调后生成的提示包含更清晰、更具指导性的语言,在自然问题(Natural Questions)和GSM8K等已建立的基准测试中实现了更高的准确率。我们的代码可在此URL获取。

英文摘要:

Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformulates user prompts into task-optimized prompts with the explicit goal of improving downstream LLM performance. We train TAPR using reinforcement learning with Group Relative Policy Optimization (GRPO), where rewards are derived from LLM-as-judge evaluations of both the reformulated prompt and the corresponding task output. Experimental results on diverse tasks, such as question answering, summarization, and arithmetic reasoning, show that our method yields consistent gains over base models in prompt rewriting ability. Fine-tuning Phi-4-mini-instruct (as the base model for TAPR) produces prompts that contain clearer and more instructive language, leading to higher accuracy on established benchmarks such as Natural Questions and GSM8K. Our code is available at: https://github.com/OliverSavolainen/task-specific-prompt-rewriter

↑