发表机构
University of Southern Queensland; RMIT University(南昆士兰大学; 皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FunnelAudit提出可执行的责任审计框架,通过问责契约和分级实际责任,在多路由推荐系统中识别关键控制项,实验显示单控制消融仅能恢复少量效果,强调了显式服务语义的重要性。
AI 中文摘要
多路由推荐系统结合了检索、分配、融合和排序,使得对单个包含和排除项进行审计变得困难。路由重叠可能掩盖逐一消融实验的效果,而冻结下游阶段则会产生与线上服务行为不一致的反事实结果。我们提出了FunnelAudit,一个用于事件级责任审计的可执行框架。问责契约规定了有争议的Top-K事件、控制项和所有者、允许的参考动作以及重放语义。FunnelAudit评估每个允许的控制配置,并应用分级实际责任来找到使每个控制项成为关键的最小结果保持意外事件。其证书记录了验证判断所需的意外事件和配对的线上服务执行。我们在两阶段、九路由漏斗中实例化了该框架,使用固定并集、加权配额分配或加权倒数排名融合,随后进行SASRec排序。在来自三个真实交互数据集的258,809个用户目标事件中,4.24%至16.24%的事件承认存在负责任的控制项。在负责任的事件-控制项对中,92.55%至99.64%需要非空意外事件,因此单控制消融仅能恢复0.36%至7.45%的效果。仅在0.31%至2.39%事件上事实结果不同的策略,在匹配排除项上负责任路由集合之间的Jaccard距离为21.44%至54.05%。独立重放复现了所有9,121,792个检查的目标世界结果;穷举搜索和通用混合整数线性规划与每个抽样判断一致。这些发现证明了显式服务语义和可验证证据对于推荐系统问责的重要性。
英文摘要
Multi-route recommender systems combine retrieval, allocation, fusion, and ranking, making individual inclusions and exclusions difficult to audit. Route overlap can hide effects from one-at-a-time ablations, while freezing downstream stages produces counterfactuals inconsistent with serving behavior. We introduce FunnelAudit, an executable framework for incident-level responsibility auditing. An accountability contract specifies the disputed Top-K event, controls and owners, permitted reference actions, and replay semantics. FunnelAudit evaluates every permitted control configuration and applies graded actual responsibility to find the smallest outcome-preserving contingency that makes each control pivotal. Its certificate records the contingency and paired serving executions needed to verify the judgment. We instantiate the framework in two-stage, nine-route funnels using fixed union, weighted quota allocation, or weighted reciprocal-rank fusion, followed by SASRec ranking. Across 258,809 user-target incidents from three real interaction datasets, 4.24-16.24% admit a responsible control. Among responsible incident-control pairs, 92.55-99.64% require a nonempty contingency, so single-control ablation recovers only 0.36-7.45%. Policies differing in factual outcomes on only 0.31-2.39% of incidents yield 21.44-54.05% Jaccard distance between responsible-route sets on matched exclusions. Independent replay reproduces all 9,121,792 checked target-world outcomes; exhaustive search and a generic mixed-integer linear program agree with every sampled judgment. These findings demonstrate the importance of explicit serving semantics and checkable witnesses for recommender accountability.