一种用于缓解医疗查询中患者上下文歧义的知识引导智能体框架
A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries
浏览论文内容
中文总结 AI 辅助
该研究提出知识引导智能体框架,通过解析医疗查询、构建假设、识别缺失上下文并提问,缓解患者上下文歧义,在诊断检索和饮食安全分类任务中显著提升了多种语言模型的性能。
中文摘要 AI 辅助
患者经常向医疗聊天机器人提交简短、表述不明确的查询,这类查询缺乏确定合适响应所需的患者特定信息。尽管这些查询在语言层面是清晰的,但根据症状、诊断、药物、过敏或饮食限制等未公开因素,它们可能对应多个合理答案。直接用语言模型回答此类查询可能会依赖关于患者的无依据假设。我们提出一种知识引导智能体框架,用于在生成最终响应前缓解患者上下文歧义。该框架运行于患者与其他方面保持不变的下游语言模型之间,它会解析初始查询,使用特定任务的知识图谱构建一组合理假设,识别区分这些假设所需的缺失患者上下文变量,并提出针对性的跟进问题。随后将原始查询与获取的上下文结合,形成清晰的提示输入下游模型。我们使用两个受控歧义缓解基准在五种语言模型上对该框架进行评估:一是从1034个系统屏蔽了临床相关证据的症状查询中进行诊断检索,二是从487个省略了关键健康上下文的查询中进行饮食安全分类。将该框架与直接回答表述不明确的查询、不获取新患者信息仅改写同一查询的方法进行对比。在诊断检索任务中,与直接提示相比,该框架在五种被评估模型上的整体精确Top-1准确率至少提升了57.1个百分点,选择性精确Recall@5至少提升了77.7个百分点;在饮食安全分类任务中,该框架在所有五种模型上均提升了准确率,且在四种模型中达到了最高的马修斯相关系数。
英文摘要
Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on undisclosed factors such as symptoms, diagnoses, medications, allergies, or dietary restrictions. A language model answering such a query directly may therefore rely on unsupported assumptions about the patient. We introduce a knowledge-guided agentic framework for mitigating patient-context ambiguity before final response generation. The framework operates between the patient and an otherwise unchanged downstream language model. It interprets the initial query, uses a task-specific knowledge graph to construct a set of plausible hypotheses, identifies the missing patient-context variables needed to distinguish among them, and asks targeted follow-up questions. The original query and the acquired context are then combined into a clarified prompt for the downstream model. We evaluated the framework across five language models using two controlled ambiguity-mitigation benchmarks: diagnosis retrieval from 1,034 symptom queries with clinically relevant evidence systematically masked, and dietary-safety classification from 487 queries with decisive health context omitted. The framework was compared with direct answering of the underspecified query and with rephrasing the same query without acquiring new patient information. In diagnosis retrieval, it increased overall exact Top-1 accuracy by at least 57.1 percentage points and selective exact Recall@5 by at least 77.7 percentage points across the five evaluated models compared with direct prompting. In dietary-safety classification, it improved accuracy across all five models and achieved the highest Matthews correlation coefficient for four...