arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提问方式决定获取内容:基于理论的建议寻求型大语言模型对话中表达清晰度的测量

How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations

Juneha Baek, Suhyeon Lee, Donghyuk Shin

arXiv 2608.29591首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究建议寻求型LLM对话中提问的表达清晰度是否为可与话题分离的潜在维度,从16447条提示词中提取出可复现的表达清晰度因子,发现其与模型回应存在关联,提出基准测试应按表达清晰度分层。

AI 中文摘要

用户会以不同方式表述相同的建议寻求请求:一些人会明确详细的约束条件,另一些人则仅模糊暗示自身需求。过往研究将这种表述差异视为需被平均消除的噪声,而我们则将其视为输入分布中一种稳定、可测量的结构。我们探究表达清晰度(即人们如何提问)是否形成与话题(即他们询问的内容)可分离的潜在维度,以及该维度是否与语言模型的回应方式存在关联。我们从公开聊天语料库(WildChat、LMSYS和ShareChat)汇总的16447条建议寻求型提示词中提取可解释特征,并恢复出少量潜在表达清晰度因子,这些因子在训练/测试拆分及不同语料库间均具有可复现性。由于该结构在很大程度上与话题分离,其定义的群体跨越不同话题,且在基于话题或任务的评估中无法被察觉。这些因子定义了少数反复出现的表达清晰度风格,其中一种风格尤为突出:篇幅较长但信息匮乏的风格,在最大型语料库中约占六分之一的提示词,在此类提示词下,模型会返回更简短、更模糊的答案,且即便存在明确需要澄清的信息不足情况,模型也不会提出澄清请求。这种对比在每个话题组和长度五分位内均成立,且并非仅因信息不足——另一种同样信息不足的风格确实会提出澄清问题。两名独立人工标注者复现了该对比结果。我们认为基准测试应按表达清晰度进行分层,且我们提供的提取结构可作为实现此目的的测量工具。

英文摘要

Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need. Prior work treats this variation as noise to be averaged away; we instead treat it as a stable, measurable structure in the input distribution. We ask whether articulation (how people ask) forms latent dimensions separable from topic (what they ask about), and whether it is associated with how language models respond. We extract interpretable features from 16,447 advice-seeking prompts pooled from public chat corpora (WildChat, LMSYS, and ShareChat) and recover a small set of latent articulation factors that replicate across train/test splits and across corpora. Because this structure is largely separable from topic, the populations it defines cut across topics and stay invisible to topic- or task-based evaluation. The factors define a handful of recurring articulation styles, one of which stands out: a long-form but information-poor style, roughly one in six prompts in the largest corpus, where models return shorter, vaguer answers and do not ask for clarification even though under-specification is exactly the condition that warrants it. The contrast holds within every topic group and length quintile, and is not under-specification alone -- a second, equally under-specified style does draw clarifying questions. Two independent human annotators reproduce this contrast. We argue that benchmarks should stratify on articulation, and we offer the extracted structure as a measurement instrument for doing so.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑