arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31751cs.CV

EgoTSR++:面向任务进度理解的自我中心时空推理

EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding

Xiaoda Yang, Can Wang, Yuxiang Liu, Pengfei Zhou, Jianwen Lou, Shuicheng Yan

AI总结:

针对视觉语言模型在自我中心任务进度判断中的顺序偏差问题,提出EgoTSR框架,通过基准测试、双向数据构建和渐进课程训练,实现92.4%长时准确率并显著提升非单调轨迹理解能力。

AI中文摘要:

视觉语言模型(VLMs)在静态视觉理解方面取得了快速进展,但在判断自我中心任务如何进展时仍不可靠。给定一个任务指令和两个视觉观察,模型应通过分析任务相关的物体配置和空间关系来确定哪个状态更接近目标,而不是依赖时间戳或呈现顺序。这种区分在操作任务中至关重要,因为重试、纠正动作和临时倒退使得进展本质上非单调。我们引入了EgoTSR,一个用于诊断和改进顺序鲁棒的任务进度理解的统一框架。首先,SpatialLogic-Bench在短时和长时设置下,以原始和顺序交换的呈现方式评估每个物理状态对,揭示模型是遵循任务状态证据还是时间顺序捷径。其次,我们的数据构建流程将成功的、近似单调的操作和第一人称轨迹转换为双向监督;LongTag进一步保留中间子任务结构以进行长时比较,而失败感知数据则将学习扩展到倒退和恢复场景。第三,渐进式CoT-to-Tag课程首先监督对任务相关状态变化的基于证据的解释,然后通过可扩展的仅标签训练巩固比较规则。实验揭示了代表性VLM中存在显著的输入顺序偏差。EgoTSR在长时设置下达到92.4%的准确率,前向-反向差距为0.1个百分点。失败感知监督进一步将非单调轨迹上的准确率提高了11.8个百分点,恢复准确率提高了11.2个百分点,同时保持了广泛的视觉和空间能力。这些结果确立了以目标为条件的状态比较作为自我中心时空推理用于任务进度理解的显式表述。

英文摘要:

Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations and spatial relations, rather than relying on timestamps or presentation order. This distinction is critical in manipulation, where retries, corrective actions, and temporary regressions make progress inherently non-monotonic. We introduce EgoTSR, a unified framework for diagnosing and improving order-robust task-progress understanding. First, SpatialLogic-Bench evaluates each physical state pair in both original and order-swapped presentations across short- and long-horizon settings, exposing whether a model follows task-state evidence or chronological shortcuts. Second, our data construction pipeline converts successful, approximately monotonic manipulation and first-person trajectories into bidirectional supervision; LongTag further preserves intermediate subtask structure for long-horizon comparison, while failure-aware data extend learning to regressions and recoveries. Third, a progressive CoT-to-Tag curriculum first supervises evidence-grounded interpretation of task-relevant state changes and then consolidates the comparison rule through scalable label-only training. Experiments reveal substantial input-order bias in representative VLMs. EgoTSR achieves 92.4% long-horizon accuracy with a 0.1-point forward-inverse Gap. Failure-aware supervision further improves accuracy on non-monotonic trajectories by 11.8 points and Recovery Accuracy by 11.2 points, while maintaining broad visual and spatial capabilities. These results establish goal-conditioned state comparison as an explicit formulation of egocentric spatiotemporal reasoning for task-progress understanding.

↑