查询的联谊会:学习检索动作
The Fellowship of the Query: Learning Retrieval Actions
浏览论文内容
中文总结 AI 辅助
本研究通过轨迹微调提升小型语言模型作为检索问答中的动作控制器,实验表明该方法显著提高动作预测和端到端答案生成性能。
中文摘要 AI 辅助
检索增强的问答需要控制决策,涉及何时分解问题、搜索、重新表述、提取证据、综合事实、验证进展以及停止。我们研究轨迹微调是否能提高小型语言模型(SLMs)作为下一步动作控制器的能力。我们还评估了一种低资源设置,其中单个SLM同时充当控制器和最终答案生成器。从可接受的教师搜索轨迹中,我们构建了一个七类动作预测任务,模型根据当前轨迹状态预测下一个结构化教师动作,并评估了在SLMs和xSLMs上使用LoRA监督微调作为控制器的效果。在1,646个保留的动作示例上,基于13,194个动作训练的Granite 4.1 3B达到了宏F1分数0.6536,而同一模型的零样本提示仅为0.1736,TF-IDF逻辑回归基线为0.5399。在端到端的控制器/生成器交换评估中,覆盖149个保留轨迹,使用微调模型同时担任两个角色,将精确匹配从0.7530提高到0.7946,令牌F1从0.7783提高到0.8295,相比使用基础模型同时作为控制器和生成器。跨角色条件表明,当生成器固定时,微调控制器增加了证据事实的记录,而仅控制器对最终答案的提升在统计上不显著。总体而言,在该评估流程中,轨迹监督改善了动作预测和证据记录行为。代码可在以下URL获取:https://this URL
英文摘要
Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can improve small language models (SLMs) as next-action controllers. We additionally evaluate a low-resource setting in which a single SLM serves as both the controller and the final-answer generator. From accepted teacher search traces, we build a seven-way action-prediction task, where the model predicts the next structured teacher action from the current trajectory state, and evaluate LoRA-supervised fine-tuning across SLMs and xSLMs as controllers. On 1,646 held-out action examples, Granite 4.1 3B trained on 13,194 actions reaches macro-F1 0.6536, compared with 0.1736 for zero-shot prompting of the same model and 0.5399 for a TF-IDF logistic-regression baseline. In an end-to-end controller/generator swap evaluation over 149 held-out trajectories, using the fine-tuned model for both roles improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295 compared with using the base model as both controller and generator. The cross-role conditions show that the fine-tuned controller increases evidence-fact recording when the generator is fixed, while controller-only final-answer gains are not statistically clear. Overall, trajectory supervision improves action prediction and evidence-recording behaviour in this evaluated pipeline. Code is available at https://github.com/padas-lab-de/agent-action-controller
发表机构
- University of Passau(帕绍大学)
机构由 AI 辅助整理,请以论文原文为准。