RACER:面向人机路由的角色对齐能力估计
RACER: Role-Aligned Competence Estimation for Human-AI Routing
浏览论文内容
中文总结 AI 辅助
RACER提出角色对齐的能力估计框架,利用上下文估计未见专家能力,结合模型后验实现贝叶斯最优的人机路由,在合成和医学影像基准上表现优异。
中文摘要 AI 辅助
学习推迟(learning to defer)要求预测系统决定何时自主行动、何时将任务推迟给人类专家。种群自适应推迟(population-adaptive deferral)利用一小段专家行为上下文,将该问题扩展到未见过的专家。诸如L2D-Pop之类的神经上下文编码器可以是查询相关的,但可能学习到与绝对类别坐标绑定的路由捷径。无身份推迟(Identity-Free Deferral, IFD)通过角色索引的类别级能力画像消除了此类捷径,但其估计在各类别内是恒定的,无法捕捉实例级别的专家专业化。我们提出RACER——角色对齐能力估计用于路由(Role-Aligned Competence Estimation for Routing)——一个角色相对框架,用于从上下文中估计未见专家的能力。RACER估计在给定每个候选类别角色下专家在查询上正确的后验预测概率,然后将这些估计与模型后验结合,以获得贝叶斯相关的专家正确概率。非参数和神经核池化估计器使用候选角色关系、共享聚合和对称摘要,排除了绝对类别身份通道。我们证明了连贯的类别重标记不变性,推导出贝叶斯对齐的推迟替代损失,并给出了一个插件遗憾界,将路由遗憾与分类器和能力估计误差联系起来。在受控的合成基准上,包括一项带有模拟专家的PathMNIST组织病理学上下文缩放研究,RACER在隐藏亚型依赖下受益于额外上下文,并在CIFAR-100合成实验中单独采样的未见专家分割上给出了最强的聚合性能。在放射科医生和人机胸部X光基准(VinDr-CXR和CheXpert)上,RACER系列在预算扫描的推迟设置中具有竞争力或最佳表现,校准结果因指标和数据集而异。
英文摘要
Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be query-dependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose RACER---Role-Aligned Competence Estimation for Routing---a role-relative framework for estimating an unseen expert's competence from context. RACER estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, RACER benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human--AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the RACER family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.
发表机构
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。