arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相关性不够:面向后果性科学问答的沟通导向检索系统

Relevance is not enough: A Communication-Oriented Retrieval System for Consequential Scientific Question Answering

Avina Nakarmi, Naga Datha Saikiran Battula, Anthony Diaz, Aritra Dasgupta

arXiv 2609.06222首次发表:更新:

发表机构

New Jersey Institute of Technology; Newark Water Coalition(新泽理工学院; 纽瓦克水联盟)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对后果性科学问答,提出沟通导向检索系统,通过分类推理类型、生成后续问题、区分已知与不确定、改写通俗语言,显著减少上下文并提升完整性,揭示完整性是当前指标未捕获的独特目标。

AI 中文摘要

人工智能系统越来越多地解答关于健康、安全和环境的科学问题。但大多数检索增强生成系统旨在提供事实正确、切题的回答,而非帮助非专家理解这些回答对其生活和决策的意义。我们聚焦于后果性科学问题,其结果直接影响人们的生活,并通过一个公共水质沟通系统进行研究,在该系统中,居民和社区领袖解读研究结果并选择行动。他们的经验表明,缺乏解释和背景的切题回答可能仍然不足,且风险信息的情感分量不可忽视。我们的系统首先按推理类型(例如,因果型与政策型)对每个问题进行分类,然后生成后续问题以识别缺失证据并检索之。一个组件明确区分已知与不确定,另一个组件将科学细节改写为通俗语言,采用基于角色的风格(如关心邻居或行政官员)来调整语气和可读性。对超过160个问题的消融研究表明,该系统使用的上下文少一个数量级,并在多种配置下提高了人工评定的完整性。与社区成员共同设计的完整性指标和微调的学习评判器显示,标准相关性得分仅能解释人类完整性评分约1%的变异,即使经过调优的评判器也只能与人类中等程度对齐,表明完整性是一个当前指标无法可靠捕获的、以人为中心的独特目标。

英文摘要

AI systems increasingly answer scientific questions about health, safety, and the environment. But most retrieval-augmented generation systems are tuned to provide factually correct, on-topic answers rather than to help non-experts understand what those answers mean for their lives and decisions. We focus on consequential scientific questions whose results directly shape people's lives and study them through a public water-quality communication system, where residents and community leaders interpret the findings and choose actions. Their experiences show that on-topic answers can still be insufficient without explanation and context and that the emotional weight of risk information cannot be ignored. Our system first classifies each question by reasoning type (for example, causal versus policy-based), then generates follow-up questions to identify missing evidence and retrieve it. One component clearly distinguishes between what is known and what is uncertain, while another rewrites scientific details into accessible language, using persona-based styles, such as a caring neighbor or an administrative official, to adapt tone and readability. Ablations on over $160$ questions show that the system uses $\textit{an order of magnitude less context}$ and, in several configurations, improves human-rated completeness. A completeness metric co-designed with community members and a fine-tuned learned judge reveal that standard relevance scores explain about $1\%$ of variation in human completeness ratings, and even the tuned judge only moderately aligns with humans, indicating that completeness is a distinct human-centered objective that current metrics do not reliably capture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑