arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用语义相似度评分修正硅采样中的模式崩溃问题

Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating

Oscar Heath, Rohan Alexander

arXiv 2607.28550首次发表:更新:

AI 中文总结

该研究针对硅采样中LLMs生成回复的模式崩溃问题,提出语义相似度评分法,通过文本嵌入将LLMs生成的文本回复映射为数值尺度,提升了回复分布保真度且校准参数少。

AI 中文摘要

硅采样指利用大语言模型(LLMs)生成调查回复的方法,虽有应用前景,但生成的回复分布往往方差过低,存在模式崩溃问题。本文认为该问题源于LLMs难以生成数值数据,而文本回复更适合此任务。针对政治态度调查,本文分析语义相似度评分能否提升硅采样回复的保真度,该方法先让LLMs生成纯文本回复,再通过文本嵌入将其映射为数值尺度。研究发现,此方法既提升了硅采样回复分布的保真度,又只需校准少量参数。

英文摘要

Silicon sampling refers to the use of Large Language Models (LLMs) to generate responses to surveys. It has shown promise, but tends to generate response distributions with unrealistically low variance. We argue that this mode collapse is due to LLMs failure to generate numeric data, and that text responses may be better suited for this task. We analyze whether Semantic Similarity Rating can improve the fidelity of silicon sampling responses when asked about political attitudes. This method solicits text-only responses from LLMs, then maps this to a numeric scale using text embeddings. We find that this method both improves the fidelity of silicon sampling response distributions, and has few parameters to calibrate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑