SCBO:面向基于LLM的社会调查的语义连贯批处理与排序
SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SCBO提出一种无需训练的框架,通过语义连贯批处理与排序,在基于LLM的社会调查中减少令牌消耗和推理时间,同时提升预测准确性。
AI中文摘要:
大型语言模型(LLMs)提供了一种可扩展的方式,利用人口统计资料和观察到的参考回答来模拟调查受访者。然而,传统方法每次提示仅预测一个问题,重复编码相同的上下文,将每个目标限制在一组狭窄的参考回答中,并阻止后续预测利用先前回答中的信息。在一次提示中预测多个问题可以减少这些成本,共享更广泛的参考池,并让后续预测基于先前的预测。这需要形成连贯的批次、选择共享参考,并有效地对问题和参考进行排序。我们提出了语义连贯批处理与排序(SCBO),一个无需训练的框架来解决这些挑战。SCBO首先使用LLM从调查项目中提取紧凑的语义表示,并过滤模板噪声。然后,它将相关问题分组为批次,并使用目标特定检索和基于质心的补全构建共享参考库。最后,它按从易到难的顺序对目标问题进行排序,并根据参考与这些问题的语义对齐程度来排列参考。在四个大规模调查数据集和四个LLM上的实验表明,与未批处理的基线相比,SCBO显著减少了令牌消耗和推理时间,同时通常提高了预测准确性。代码可在以下网址获取:此HTTPS URL。
英文摘要:
Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.