arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClusterFewshot:改进LLM工作流的少样本优化

ClusterFewshot: Improving Few-shot Optimization for LLMs workflow

Omri Bar Haim, Shahar Katz, Lior Wolf

arXiv 2609.25939首次发表:更新:

发表机构

Tel Aviv University(特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ClusterFewshot通过结合语义结构与效用感知评分来选取少样本演示,在DSPy流程中降低优化成本并提升准确率。

AI 中文摘要

大型语言模型(LLM)工作流的性能通常取决于选择少量上下文演示(in-context demonstrations)来引导模型在新任务上的行为。近期方法通过向提示中补充成功的推理路径来改进这一过程。然而,这些方法的演示选择依赖于随机采样或基于度量的排序,忽略了任务的语义结构。我们提出ClusterFewshot,一种将语义结构化与效用感知评分相结合的策略,以构建具有代表性和有效性的少样本演示集。在基于DSPy的流程中进行评估,ClusterFewshot在多个基准上大幅降低了优化成本,同时在独立提示调优和混合提示-权重优化中,相较于先前的基于引导(bootstrap)的方法,持续提升了准确率。

英文摘要

The performance of large language model (LLM) workflows often depends on selecting a small set of in-context demonstrations to guide model behavior on new tasks. Recent methods improve this process by augmenting prompts with successful reasoning paths. However, their demonstration selection relies on random sampling or metric-based rankings, overlooking the semantic structure of the task. We propose ClusterFewshot, a strategy that combines semantic structuring with utility-aware scoring to construct representative and effective few-shot demonstration sets. Evaluated within DSPy-based pipelines, ClusterFewshot substantially reduces optimization cost across multiple benchmarks, while consistently improving accuracy relative to prior bootstrap-based methods in both standalone prompt tuning and hybrid prompt-weight optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑