arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20304cs.AI

诊断、恢复、认证:隐藏动态变化下的任务就绪性

Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes

Nguyen Viet Tuan Kiet, Huynh Thi Thanh Binh

AI总结:

针对隐藏动力学变化下的任务就绪性问题,提出证据门控匹配脉冲传输的贝叶斯方法,统一诊断与恢复,提供校准下界或弃权决策,并在多模拟器基准上验证其有效性。

AI中文摘要:

已部署的控制策略可能掩盖重要的动力学变化:当策略很少激励某个执行器时,该执行器可能丧失有效性而不影响当前任务,尽管它对于尚未指定的未来任务至关重要。我们引入了休眠动力学漂移下的任务就绪性,这是一个决策问题,在有限且与任务无关的交互预算下,统一了主动变化诊断和变化后控制恢复。智能体必须识别局部动力学是否发生变化以及变化位置,在下游任务身份揭示之前,利用少量信息性交互来表征变化,随后为每个候选任务提供恢复后的策略及其可实现回报的校准下界,或提供弃权(不执行)决策以转向安全回退。我们提出了证据门控匹配脉冲传输,这是一种基于干预的贝叶斯方法,通过共享的匹配响应表示将故障定位与执行器有效性估计耦合,从而在将局部证据转化为恢复相关不确定性的同时保持诊断可靠性。该不确定性被传播到任务条件策略选择和就绪性认证中,使得部署决策能够明确权衡预期性能、置信度和回退使用。我们在涵盖多个模拟器的多样化休眠执行器基准套件上评估了所提出的框架,采用将诊断与能力恢复分离的协议,通过就绪性覆盖率、选择性风险、交互成本以及回报来评分部署,并识别了传输证据起决定性作用的故障机制。

英文摘要:

A deployed control policy can conceal consequential dynamics changes: an actuator may lose effectiveness without affecting the current task when the policy rarely excites it, despite being critical for a future task that has not yet been specified. We introduce task readiness under dormant dynamics drift, a decision problem that unifies active change diagnosis and post-change control recovery under a limited, task-agnostic interaction budget. An agent must identify whether and where local dynamics have changed, use a small number of informative interactions to characterize the change before downstream task identity is revealed, and subsequently provide each candidate task with either a recovered policy and a calibrated lower bound on its achievable return or an abstention decision to a safe fallback. We propose Evidence-Gated Matched-Pulse Transport, an intervention-based Bayesian procedure that couples fault localization with estimation of actuator effectiveness through a shared matched-response representation, thereby preserving diagnostic reliability while converting localized evidence into recovery-relevant uncertainty. This uncertainty is propagated to task-conditioned policy selection and readiness certification, enabling deployment decisions that explicitly trade off expected performance, confidence, and fallback use. We evaluate the resulting framework on a diverse suite of dormant-actuator benchmarks spanning multiple simulators, under a protocol that separates diagnosis from capability recovery, scores deployment by readiness coverage, selective risk, and interaction cost as well as return, and identifies the fault regimes in which transported evidence is decisive.

↑