arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CANDI:特定领域问答的上下文对齐

CANDI: Contextual Alignment for Niche Domains Question Answering

Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth

arXiv 2607.11891首次发表:更新:

AI 中文总结

研究针对专业领域大语言模型评估问题,提出CANDI-QA数据集,含两类问答对。评估多种语言模型,给出MTSS-Net框架。揭示特定领域上下文对齐挑战及当前模型局限,为上下文感知语言模型研究提供关键基准。

AI 中文摘要

在医疗诊断和金融咨询等专业领域部署大语言模型需要评估超出常识的能力。传统问答基准往往无法捕捉这些领域所需的细微上下文基础、用户意识和领域理解。为此,我们引入了CANDI-QA数据集,用于评估大语言模型在特定设置中提供准确、上下文敏感和用户对齐答案的能力。该数据集包含专家策划的问答对,分为信息辅助问题和应用推理问题两类。我们评估了十多种不同的语言模型,并提出了MTSS-Net作为轻量级神经符号框架。研究结果凸显了在特定领域实现上下文对齐的巨大挑战,并揭示了当前大语言模型在没有增强上下文或符号集成时的局限性。最终,CANDI-QA为推进上下文感知语言模型的研究提供了关键基准,促进了高风险领域强大、可信人工智能的发展。

英文摘要

The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often fail to capture the nuanced contextual grounding, user awareness, and domain understanding these fields require. To address this, we introduce CANDI-QA (Contextual Alignment for Niche Domains Question Answering), a novel dataset evaluating LLMs on delivering accurate, context-sensitive, and user-aligned answers in specialized settings. CANDI-QA features expert-curated question-answer pairs structured into two categories: (1) Information Assistance Questions, which are direct, factual queries requiring precise extraction, and (2) Applied Inference Questions, which are multi-hop reasoning tasks needing situational inference to generate actionable insights. We evaluate over ten diverse language models, from compact open-source to state-of-the-art proprietary systems. As a robust baseline, we present MTSS-Net, a lightweight neuro-symbolic framework combining neural retrieval with rule-based reasoning. Our findings highlight the profound challenges of achieving contextual alignment in niche domains, revealing the limitations of current LLMs without enhanced contextual or symbolic integration. Ultimately, CANDI-QA serves as a critical benchmark for advancing research in context-aware language models, stimulating the development of robust, trustworthy AI for high-stakes domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑