发表机构
Purdue University; IBM Research(普渡大学; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ROAR通过关系模式和解析层统一异构AI驱动研究系统的运行数据,构建超900次运行的语料库,揭示问题依赖的搜索结构,并支持配置优化,是首个跨系统聚合分析基础设施。
AI 中文摘要
每次AI驱动研究系统(ADRS)的运行都是在广阔解空间中进行的一次昂贵搜索,而可靠的评估需要多次运行,这使得运行数据既生产成本高昂,又对大规模分析具有保留价值。然而,这些数据仍然处于碎片化状态:团队独立运作,ADRS框架以不同格式输出结果,且不存在用于跨问题和系统聚合或比较运行的共享基础设施。我们提出ROAR,一种系统性地统一和分析异构ADRS输出的解决方案。ROAR解决两个挑战:调和异构ADRS输出,以及实现跨具有不同目标和评分函数的运行的分析。我们通过关系模式和解析层实现这一点,该层在保留数据谱系和时间结构的同时规范化异构ADRS输出,并适应新系统而无需修改模式。通过构建来自多个ADRS的超过900次运行的语料库,我们展示了汇集数据如何揭示难以观察的问题景观属性。与先前工作一致,具有相同配置的运行可能收敛到不同分数。我们发现许多运行在早期实现大部分收益,且将先前解决方案纳入搜索过程的不同策略的有效性因问题而异。我们进一步表明,汇集语料库是可操作的而不仅仅是分析性的,通过使用ROAR配置ADRS运行。综合来看,这些结果说明了汇集ADRS数据如何暴露搜索行为中依赖于问题的结构,这种结构难以从任何单一系统、团队或基准中检测到。当运行仍处于孤立状态时,此类跨领域见解难以获得;ROAR是首个旨在统一它们的基础设施。
英文摘要
Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no shared infrastructure exists to aggregate or compare runs across problems and systems. We present ROAR, a solution for systematically unifying and analyzing heterogeneous ADRS outputs. ROAR addresses two challenges: reconciling heterogeneous ADRS outputs and enabling analytics across runs with different objectives and scoring functions. We achieve this through a relational schema and parsing layer that normalize heterogeneous ADRS outputs while preserving data lineage and temporal structure, and accommodating new systems without requiring schema modifications. From building a corpus of more than 900 runs from multiple ADRS, we show how pooled data can reveal properties of problem landscapes that are difficult to observe. Consistent with prior work, runs with identical configurations may converge to different scores. We find that many runs realize most gains early, and that the effectiveness of different strategies for incorporating prior solutions into the search process varies across problems. We further show that the pooled corpus is actionable and not merely analytical by using ROAR to configure ADRS runs. Together, these results illustrate how pooled ADRS data can expose problem-dependent structure in search behavior that is difficult to detect from any single system, team, or benchmark. Such cross-cutting insights are difficult to obtain while runs remain siloed; ROAR is the first infrastructure designed to unify them.