arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRL中的块解缠:连接可辨识性与视觉状态估计

Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation

Emre Acartürk, Pranamya Kulkarni, Puranjay Datta, Karthikeyan Shanmugam, Burak Varıcı, Ali Tajer

arXiv 2610.06809首次发表:更新:

发表机构

Rensselaer Polytechnic Institute; The University of Texas at Austin; Swiss AI, EPFL; Google DeepMind; University of Delaware(伦斯勒理工学院; 德克萨斯大学奥斯汀分校; 瑞士人工智能,洛桑联邦理工学院; 谷歌DeepMind; 特拉华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文为干预性因果表示学习在弱化假设下建立块解缠的可辨识性保证,并将其应用于无标签的具身视觉状态估计,实现从图像视频恢复机器人物理变量,架起理论到实践的桥梁。

AI 中文摘要

因果表示学习(CRL)是从高维观测中恢复因果相关潜在变量的过程。作为一种无标签推断方法,CRL对于数据标签不可用或难以获取的应用尤其具有吸引力。尽管在理解CRL的可辨识性保证方面已取得显著进展,但这些保证往往在高度理想化的假设下成立,这限制了其直接应用于现实世界问题。本文对干预性CRL具有双重目标。首先,它为显著弱化的干预假设建立了可辨识性保证,从而实现因果变量的块解缠,其中块结构取决于实际可用的干预机制。其次,块解缠框架被用于具身视觉状态估计,其目标是在无标签数据的情况下,直接从视觉数据(图像和视频)中恢复机器人系统的潜在物理变量。这两个组成部分具有关键的互补性。块解缠理论在弱化假设下描绘了可辨识性保证,而应用则表明,尽管进一步违反假设,所得目标在受控的具身环境中仍然有效,从而提供了将无标签CRL的潜力转化为实际问题所需的理论到实践的桥梁。

英文摘要

Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which temper the direct application to real-world problems. This paper has a two-fold objective for interventional CRL. First, it establishes identifiability guarantees for substantially weaker interventional assumptions, resulting in block disentanglement of the causal variables, where the block structure depends on the realistically available intervention mechanisms. Secondly, the block disentanglement framework is used for embodied visual state estimation, in which the objective is to recover the latent physical variables of a robotic system directly from visual data (images and videos) without labeled data. These two components are critically complementary. The block disentanglement theory delineates identifiability guarantees under weakened assumptions, and the application demonstrates that the resulting objective remains effective in a controlled embodied setting despite further assumption violations, providing a theory-to-practice bridge needed to translate the promise of label-free CRL into practical problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑