CLARA:通过结果分析澄清自然语言癌症基因组学查询中的语言歧义
CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries
浏览论文内容
中文总结 AI 辅助
本研究提出CLARA框架,通过将自然语言癌症基因组学查询转为带类型的科学查询规范,执行多种解释并在结果分歧时澄清,在基准测试中展现出良好的歧义区分能力,揭示了安全与负担的权衡。
中文摘要 AI 辅助
自然语言界面可提升癌症基因组学数据库的易用性,但即便问题表述流畅,其科学含义仍可能存在歧义。我们提出CLARA框架,该框架将问题表示为带类型的科学查询规范,考虑多种可能的解释并执行,当估计结果出现分歧时请求澄清。在针对8个TCGA PanCancer Atlas队列和30基因panel的突变患病率对比任务中对CLARA进行评估,该基准包含330个唯一可执行的对比项,这些对比项在突变范围、检测分母和样本背景上存在差异;根据预注册定义,相对分歧大于0.10或绝对分歧大于5个百分点的对比项为结果敏感型,共115个,其余215个为结果稳定型。独立实现的pandas执行引擎完美复现了SQLite引擎的全部660个结果。在另一项由LLM生成并经人工审核的120题语言压力测试中,CLARA识别了全部60个结果敏感型对比项,对60个稳定型对比项中的13个进行了不必要的澄清,准确率为89.2%,灵敏度(召回率)为100%,特异性为78.3%。单独的机器学习整体准确率更高(97.5%),但遗漏了一个关键对比项。这表明下游执行可区分重要与不重要的歧义,并揭示了安全性与负担之间的明确权衡。
英文摘要
A natural language interface can be used to make cancer genomics databases easier to use, but even if a question is perfectly fluent, its scientific meaning can be ambiguous. We propose CLARA, a framework that represents a question as a typed scientific query specification, considers a few possible interpretations, executes them, and asks for clarification when the estimates diverge. CLARA was assessed on mutation-prevalence contrasts among eight TCGA PanCancer Atlas cohorts and a 30-gene panel. This benchmark consisted of 330 unique executable contrasts varying in mutation scope, assay denominator, and sample context; 115 contrasts were result-sensitive and 215 were result-stable, per the preregistered definition of relative divergence greater than 0.10 or absolute divergence greater than 5 percentage points. An independently implemented pandas execution engine perfectly replicated all 660 results from the SQLite engine. In a separate 120-question LLM-generated, manually vetted language stress test, CLARA recognized all 60 result-sensitive contrasts and needlessly clarified 13 of 60 stable contrasts (accuracy 89.2%, sensitivity/recall 100%, specificity 78.3%). Standalone machine learning had superior overall accuracy (97.5%) but missed one critical contrast. This demonstrates that downstream execution can distinguish consequential from inconsequential ambiguity and reveal an explicit trade-off between safety and burden.
发表机构
- Independent Researchers E-mail
机构由 AI 辅助整理,请以论文原文为准。