arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果关系如何弥合语义鸿沟

How Causality Bridges the Semantic Gap

Shuhao Zhang, Xuran Zhou, Han Guo, Pengtao Xie, Yujia Zheng

arXiv 2610.02594首次发表:更新:

发表机构

University of California, San Diego; University of Illinois Urbana-Champaign(加利福尼亚大学圣迭戈分校; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CausalBridge框架,利用因果结构而非人类知识来弥合测量与语义之间的鸿沟,通过因果图约束下的嵌入对齐为未命名变量赋予语义,实验表明其比现有方法更准确且成本更低。

AI 中文摘要

数值测量捕捉了系统的行为方式,但往往未指明其变量的含义。有些变量被测量但从未被标记,另一些则从未被测量。现有方法通过参考一般人类知识为这些变量赋予语义,但这在知识存在之处继承了其偏见,在知识缺失之处则无所作为。我们转而利用因果结构来弥合测量与其含义之间的鸿沟,通过变量对其他变量的作用方式来解读其语义。我们将此形式化为结构约束的语义对齐,即在因果图隐含的依赖关系下求解每个未命名变量的嵌入,并以少数已知名称的嵌入作为锚点。据此,我们构建了CausalBridge框架,该框架从测量中发现因果图(包括潜在变量),在这些关系下求解嵌入,并通过语言模型将其表达为名称。因果结构反映了生成测量的机制,且仅从测量中恢复,这可能使其成为唯一不受人类知识偏见影响的信息来源。我们在五份问卷和三个机器人场景上评估了CausalBridge,其中20%至90%的变量名称被遮蔽。它比依赖关联的现有方法更准确地恢复了观测变量和潜在变量的语义,且随着系统被记录的部分减少,其优势扩大。它发现的图在命名变量方面与已记录的图一样准确,新系统在几分钟内即可完成命名,成本仅为采样方法的一小部分。一旦语义鸿沟被忠实弥合,机器便能理解世界并采取因果行动。

英文摘要

Numerical measurements capture how a system behaves, but often leave the meanings of its variables unspecified. Some variables are measured but never labeled, and others are never measured at all. Existing methods assign semantics to such variables by consulting general human knowledge, but this inherits its biases where that knowledge exists and offers nothing where it does not. We bridge this gap between measurements and their meanings with causal structure instead, reading a variable's semantics from how it acts on other variables. We formalize this as structure-constrained semantic alignment, in which the embedding of each unnamed variable is solved under the dependence relations implied by the causal graph, with the embeddings of a few known names as anchors. Accordingly, we build CausalBridge, a framework that discovers the causal graph from the measurements, latent variables included, solves for the embeddings under those relations, and expresses them as names through a language model. The causal structure reflects the mechanism that generated the measurements and is recovered from the measurements alone, which may make it the one source of information free of bias from human knowledge. We evaluate CausalBridge on five questionnaires and three robotics scenarios, with 20 to 90% of the variable names masked. It recovers the semantics of observed and latent variables more accurately than existing methods that rely on association, and its lead widens as less of the system is documented. The graph it discovers names variables as accurately as the documented one, and a new system is named in minutes and at a fraction of the cost of sampling methods. Once the semantic gap is bridged faithfully, machines can understand the world and take actions causally.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑