发表机构
Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对信息抽取问题,提出LoQ生成文档特定问题集、FeedQ反馈驱动优化方法,优化问题可训练轻量级模型,4B参数模型性能超未调大模型,发布12820个优化问题数据集。
AI 中文摘要
当信息抽取任务失败时,人们的自然反应是改进执行该任务的模型,例如扩大模型规模或优化其推理能力。本文表明,流程的另一部分同样重要:用于引出信息的查询。在四个临床基准和五个大型语言模型(LLM)上,仅改进问题设计就可使性能提升18.6个F1分数点,这比使用更大的抽取模型带来的提升更显著。为使此类问题设计可学习,我们引入了问题列表(LoQ),用于生成文档特定的问题集;还提出了FeedQ,一种反馈驱动的优化方法,可根据抽取结果迭代优化问题。由此得到的优化问题可用于训练轻量级生成器:经过微调后,40亿参数的模型可达到或超过专家设计的基准性能,且大幅超越未微调的更大模型的性能。我们发布了包含12820个优化问题的数据集,以支持信息抽取研究更广泛地将问题设计视为一级问题的转变。
英文摘要
When information extraction fails, a natural instinct is to improve the model doing it: for example, by scaling it up or refining its reasoning. In this paper, we show that another part of the pipeline matters at least as much: the queries used to elicit this information. Across four clinical benchmarks and five LLMs, improving the question design alone raises performance by 18.6 F1-score points, i.e. more than using larger extraction models. To make such question design learnable, we introduce List of Questions (LoQ), which generates document-specific question sets, and FeedQ, a feedback-driven optimization method that iteratively refines questions against extraction outcomes. The resulting optimized questions can be used to train lightweight generators: with fine-tuning, 4B-parameter models match or outperform expert-derived baselines and substantially exceed the performance of much larger untuned models. We release a dataset of 12,820 optimized questions to support a broader shift in information extraction research toward treating question design as a first-class problem.
CommentsAccepted in EMNLP-2026, 21 Pages