冻结、验证、报告:基于公共证据审计城市站点规划
Freeze, Validate, Report: Auditing Urban Station Plans with Common Evidence
- Clausthal University of Technology(Clausthal工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出冻结-验证-报告框架,通过公共验证证据和字典序规则审计城市站点规划,在波尔图和芝加哥测试中实现可审计选择。
AI中文摘要:
交通管理部门在比较城市站点规划时,需要区分规划差异与评估输入差异。从各规划的出行结果中重建空间单元、包含的出行或候选区块,可能导致评分不可比较。冻结-验证-报告(Freeze-Validate-Report)确立了一种评估契约:目标出行、代理群体、报告单元、候选池以及字典序选择规则。一个两点示例表明,特定规划的观测选择可能逆转目标均值排序。我们比较了七个生成器,涵盖聚类、设施选址、公平性导向的改进以及网格控制。所有规划均基于公共验证证据,使用端点距离(起点和终点到最近站点的距离之和)进行评分。一个五指标规则为单一测试报告选择一个规划。在波尔图,它选择了IFkCO,一种带离群点的个体公平k-中心(individually fair $k$-center with outliers)的改进算法;测试总体和最差组的第90百分位数分别为805米和872米。在芝加哥,它选择了网格(Grid);同一天稍后测试窗口中的数值为1404米和1656米。事后检查显示输入敏感性:将候选从300扩展到600在两个城市都改变了选择,而在芝加哥三个日期保持小时固定则选择了优先级(Priority)。一个比例均值公平let诊断显示,特定规划的提案池可能改变诊断排名;可接受性和求解器界限仅适用于采样池。该软件包提供代码、哈希、决策日志和提案池。我们在声明的契约内建立可审计性,而非部署后的稳定性能;代理定义和指标优先级仍由主管部门选择。
英文摘要:
Transport authorities comparing urban station plans need to distinguish plan differences from differences in evaluation inputs. Reconstructing spatial units, included trips, or candidate blocks from each plan's access outcomes can make scores incomparable. Freeze--Validate--Report fixes an evaluation contract: target trips, proxy groups, reporting cells, candidate pools, and a lexicographic selection rule. A two-point example shows that plan-specific observation selection can reverse the target mean ordering. We compare seven generators spanning clustering, facility location, fairness-oriented adaptations, and a grid control. All plans are scored on common validation evidence using endpoint distance: the sum of origin and destination distances to their nearest stations. A five-metric rule selects one plan for a single test report. In Porto, it selects IFkCO, an adaptation of individually fair $k$-center with outliers; test overall and worst-group 90th percentiles are 805 m and 872 m. In Chicago, it selects Grid; values in a later test window on the same day are 1404 m and 1656 m. Post hoc checks show input sensitivity: expanding candidates from 300 to 600 changes selection in both cities, while holding the hour fixed across three Chicago dates selects Priority. A proportional mean fairlet diagnostic shows that plan-specific proposal pools can change diagnostic ranks; admissibility and solver bounds apply only to the sampled pool. The package supplies code, hashes, decision logs, and proposal pools. We establish auditability within a declared contract, not stable performance after deployment; proxy definitions and metric priorities remain the authority's choices.