arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26167cs.AIcs.ETcs.LGeess.AS

拒绝不等同于鲁棒性:在经证实无信息的临床疼痛语音转录本上审计大型语言模型的自信编造

Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript

Sagnik De, Sreenija Pavuluri

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在TAME Pain语料库上评估7个大语言模型,发现部分模型在特定提示下会对无信息转录本产生自信编造,且拒绝不等同于鲁棒性,权威式提示会影响模型弃权行为。

中文摘要 AI 辅助

幻觉与弃权(不执行)基准很少能证明模型不可能知道正确答案,这使得区分恰当的弃权(不执行)与无依据的预测变得困难。本研究在TAME Pain语音语料库上评估了7个大型语言模型。参与者阅读语音平衡的哈佛句子,同时将一只手浸入冷水或温水中,仅在周期性疼痛陈述时报告疼痛。该协议生成了5750个无信号哈佛句子话语,其转录本不含词汇疼痛信息,以及1294个信号疼痛陈述话语,其中明确说出了疼痛评分。在无信号组中,可从声学特征中恢复疼痛(AUC为0.622,95%置信区间为0.553至0.662),而基于转录本的预测接近随机水平(AUC为0.489,95%置信区间为0.418至0.504)。由于自动语音识别会去除声学疼痛线索,因此仅从转录本推断的任何疼痛评分都缺乏可用证据支持。在协作式提示下,6个模型对几乎所有无信号转录本弃权(不执行),在阳性对照任务中正确提取口语疼痛评分,准确率范围为0.939至1.00,且保持的预期校准误差最多为0.100。在权威式提示下,弃权(不执行)变得依赖提示,同一模型在等效提示措辞下的弃权率范围为0.18至1.00。大多数模型在被迫回答时产生低置信度估计,而Gemini 2.5 Flash和Llama 3.1 8B始终产生自信的疼痛评分,自信编造率分别为0.53和0.76,而所有其他模型的该比率最多为0.15。在被迫回答中未观察到显著的人口统计学效应,所有p值均大于或等于0.20。

英文摘要

Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models were evaluated on the TAME Pain speech corpus. Participants read phonetically balanced Harvard Sentences while one hand was immersed in cold or warm water and reported pain only during periodic pain statements. This protocol generated 5,750 no signal Harvard Sentence utterances whose transcripts contained no lexical pain information and 1,294 signal pain statement utterances in which the pain rating was explicitly spoken. In the no signal arm, pain was recoverable from acoustic features (AUC 0.622, 95% CI 0.553 to 0.662), whereas transcript based prediction was near chance (AUC 0.489, 95% CI 0.418 to 0.504). Because automatic speech recognition removes the acoustic pain cues, any pain score inferred solely from the transcript is unsupported by the available evidence. Under cooperative prompting, six models abstained on nearly all no signal transcripts, correctly extracted spoken pain ratings in the positive control task with accuracies ranging from 0.939 to 1.00, and maintained an expected calibration error of at most 0.100. Under authority framed prompts, abstention became prompt dependent, with the same model ranging from 0.18 to 1.00 across equivalent prompt phrasings. Most models produced low confidence estimates when forced to answer, whereas Gemini 2.5 Flash and Llama 3.1 8B consistently generated confident pain scores with confident fabrication rates of 0.53 and 0.76, compared with at most 0.15 for all other models. No significant demographic effects were observed in forced responses, with all $p$ values greater than or equal to 0.20.

补充信息

↑