AI 中文总结
FAIR-Compute 针对联邦数字研究基础设施的分配问题,结合多领域方法与实证研究,发现资源分配存在占用与利用率偏差、调度可优化、联邦需管控等关键问题,提出相关路线图方向。
AI 中文摘要
随着高性能计算(HPC)、高吞吐量计算及数据存储需求的增长,稀缺计算资源的分配方式(而非仅资源总量)已成为决定英国科研生产力的关键因素。FAIR-Compute 将联邦数字研究基础设施(DRI)中的分配问题,视为算法博弈论、运输经济学与 HPC 调度的交叉领域。核心观察结论为:共享系统一旦需决定谁在何时、何种条件下运行,其调度器设置便不再是纯技术问题,而成为政策问题。该项目整合三类证据:英国及国际分配实践的全景回顾(包括 EuroHPC、WLCG、ACCESS、JASMIN 与 DiRAC)及利益相关者调查;针对用户策略性与不确定性报告的分配机制设计模型;基于公开 Fresco/Anvil 工作负载轨迹及受控合成压力测试的仿真研究。三类证据均指向三项结果:其一,分配记录仅测量占用量(预留资源)而非利用率(完成的有用工作),因此系统目前无法回答英国研究与创新署(UKRI)及数字、文化、媒体与体育部(DSIT)最关注的问题——资源是否产生了生产价值;其二,简单透明的调度启发式算法经调整后,表现接近离线全信息基准,因此近期机会在于调整与监测现有调度器,而非替换;其三,联邦架构确有价值,但未受管理时会表现为不协调路由:可最小化平均延迟,却悄悄将负载集中于接收系统,可能造成损害,这与交通经济学中的经典效应(布雷斯悖论)类似,即让所有人自主选择最快路线会导致整个网络状况恶化。
英文摘要
As demand for high-performance computing (HPC), high-throughput computing and data storage grows, the way scarce compute is allocated -- not just how much exists -- has become a decisive factor in the productivity of UK research. FAIR-Compute studies allocation in a federated Digital Research Infrastructure (DRI) as a problem at the intersection of algorithmic game theory, transport economics and HPC scheduling. Our central observation is simple: once a shared system must decide who runs, when and under what evidence, its scheduler settings cease to be a purely technical matter and become policy. The project combined three strands of evidence: a landscape review of UK and international allocation practice (including EuroHPC, WLCG, ACCESS, JASMIN and DiRAC) supported by a stakeholder survey; a mechanism-design model of allocation under strategic and uncertain user reports; and a simulation study built on the public Fresco/Anvil workload trace and controlled synthetic stress tests. Three results recur across all three strands. First, allocation records measure occupancy (resources reserved) rather than utilisation (useful work done), so the system cannot currently answer the question UKRI and DSIT most want answered -- whether resources deliver productive value. Second, simple, transparent scheduling heuristics perform close to an offline full-information benchmark once tuned, so the near-term opportunity is to tune and instrument existing schedulers rather than replace them. Third, federation is genuinely valuable but behaves like uncoordinated routing when left unmanaged: it can minimise average delay while quietly concentrating load -- and potentially harm -- on the receiving system. This mirrors a classic effect from road-traffic economics (Braess's paradox), where letting everyone independently pick the fastest route can leave the whole network worse off.
CommentsFinal Report to the National Federated Compute Services (NFCS) Flexible Fund