arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GraphIR:面向大语言模型引导的神经架构演化的架构级搜索状态

GraphIR: Architecture-Level Search States for LLM-Guided Neural Architecture Evolution

Zhen Liu, Wanqi Zhou, Shuanghao Bai, Yuhan Liu, Jinjun Wang, Jingwen Fu

arXiv 2608.01633首次发表:更新:

AI 中文总结

针对LLM引导NAS中代码级表示无法满足架构变异需求的问题,提出架构感知中间表示GraphIR,在六个下游基准中集成到OpenEvolve后实现最佳搜索性能。

AI 中文摘要

大语言模型(LLM)支持直接在可执行神经网络程序上进行神经架构搜索(NAS),但代码级的灵活性无法提供有效变异所需的架构状态:LLM必须从实现细节中推断张量依赖关系、可编辑组件及兼容性约束。为解决该表示不匹配问题,我们提出GraphIR,这是一种感知架构的中间表示,为可执行程序补充了与变异对齐的候选状态。GraphIR通过三个互补视图组织每个候选架构:描述张量流的计算骨架、暴露可编辑模块与操作的变异表面,以及捕获接口契约、传播形状和下游依赖的有效性包络。为评估我们的方法,我们构建了NAS-Dependency基准,包含120个问题,覆盖六个互补的依赖推理维度。诊断结果显示,GraphIR在识别确切生产者出现、追踪依赖传播及诊断接口与故障风险方面特别有效。在包括CLRS在内的六个下游基准中,GraphIR集成到OpenEvolve后,在保持可比模型规模和良好端到端NAS效率的同时,实现了最佳的整体搜索性能。这些结果表明,面向变异的架构状态为可执行神经程序与LLM引导的架构演化提供了有效接口。

英文摘要

Large language models (LLMs) enable neural architecture search (NAS) directly over executable neural network programs. However, code-level flexibility does not provide the architecture state needed for effective mutation: LLMs must infer tensor dependencies, editable components, and compatibility constraints from implementation details. To address this representation mismatch, we propose GraphIR, an architecture-aware intermediate representation that supplements executable programs with a mutation-aligned candidate state. GraphIR organizes each candidate through three complementary views: a computation skeleton describing tensor flow, a mutation surface exposing editable modules and operations, and a validity envelope capturing interface contracts, propagated shapes, and downstream dependencies. To evaluate our method, we construct NAS-Dependency, a 120-question benchmark covering six complementary dependency-reasoning dimensions. The diagnostic shows that GraphIR is particularly effective at identifying exact producer occurrences, tracing dependency propagation, and diagnosing interface and failure risks. Across six downstream benchmarks including CLRS, GraphIR achieves the best overall search performance while maintaining comparable model size and favorable end-to-end NAS efficiency when integrated into OpenEvolve. These results show that a mutation-oriented architecture state provides an effective interface between executable neural programs and LLM-guided architecture evolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑