arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21721cs.AI

提问或回答:多轮健康错误信息干预的决策框架

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对对话中健康错误信息纠正的决策问题,提出RO-PnR框架,通过权衡探查收益与成本选择提问或回应,在多组数据集和模型上取得优于基线的干预效果。

中文摘要 AI 辅助

在对话中纠正健康错误信息,仅生成事实性反驳是不够的:用户的已知内容、信念及所需信息存在差异,因此有效的干预往往需要先提出恰当的澄清问题。但现有方法要么立即回应,要么无差别探查,将澄清视为不必要或始终有益。我们提出Reward-Optimized Probe-and-Respond(RO-PnR,奖励优化探查与回应)框架,用于学习何时提问的成本值得付出。每一轮中,RO-PnR会在探查更多信息与执行最终纠正之间做出选择,依据是权衡探查预期收益与交互成本的轮级奖励。为捕捉用户异质性对探查价值的影响,我们为每个模拟用户构建了沿健康素养和信念承诺维度的潜在状态模型。实验表明,RO-PnR在3个健康错误信息数据集和3个基础模型上实现了最高的成本调整效用,且比始终探查的基线方法减少了30%的轮次。

英文摘要

Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.

发表机构

  • University of North Texas(北得克萨斯大学)
  • University of Georgia(佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑