AI 中文总结
该研究针对自适应模态获取中策略诱导分组破坏条件校准假设的问题,提出RouteCert方法,在临床心电图任务及掩码多模态基准上实现了高回答比例与低分歧率的良好性能。
AI 中文摘要
多模态系统在开始推理时可能仅持有部分输入,并以一定成本获取其余输入。在自适应获取场景中,策略决定最终需观测哪些输入,因此我们基于该终端输入模式给出保证。条件校准通常假设分组映射独立于校准样本固定,但策略诱导的分组不满足该条件。我们刻画了模式条件保证仍有效的情形,并提出两种有限样本构造:一是在终端模式应用校准的无阈值路由,二是同时对完整策略-模式对进行认证,该方法允许校准数据选择部署的策略。一个反例表明,针对校准独立分组映射证明的保证,在策略使终端组校准依赖后不一定能转移。我们将所提方法命名为RouteCert。在临床心电图任务中,采用分阶段、按成本排序的导联协议,该已认证策略在48.8%的预设序贯成本(获取所有阶段的成本)下,对71.2%的留存患者给出答案,且与心脏病医生诊断的观测分歧率为7.4%,三个获取阶段均有各自的认证。在掩码多模态基准上,对每个终端模式逐点认证,以全信息参考决策而非真实标签衡量的观测最坏模式选择性风险为0.034,而合并设计的对应值为0.145(上限为0.10),回答比例相当(0.350对0.342);在预算匹配的同时比较下,回答比例降至0.305。
英文摘要
A multimodal system may begin inference while holding only some of its inputs and may acquire the rest at a cost. With adaptive acquisition, the policy determines which inputs are ultimately observed, so we state risk control conditional on that terminal input pattern. Conditional calibration typically assumes that the grouping map is fixed independently of the calibration sample, a condition that policy-induced grouping does not satisfy. We characterize when pattern-conditional risk control remains valid and give two finite-sample constructions: threshold-free routing with calibration applied at the terminal pattern, and simultaneous validation of complete policy-pattern pairs, which lets calibration data select the deployed policy. A counterexample shows that validity proved for a calibration-independent grouping map need not transfer once the policy makes the terminal group calibration-dependent. We call the resulting method RouteCert. On a clinical electrocardiogram task with a staged, cost-ordered lead protocol, the deployed policy answers 71.2% of held-out patients, with an observed 7.4% disagreement with the cardiologist's diagnosis at 48.8% of the prespecified ordinal cost of acquiring each stage, and all three acquisition stages are validated separately. On masked multimodal benchmarks, validating pointwise at each terminal pattern holds observed worst-pattern selective risk, measured against the full-information reference decision rather than the true label, at 0.034, where a pooled design reaches 0.145 against a 0.10 cap, at a comparable answered fraction (0.350 vs 0.342); under the budget-matched simultaneous comparison, the answered fraction falls to 0.305.
Comments23 pages, 4 figures, 14 tables. Includes appendix with full proofs