arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00280cs.OScs.SE

在文件系统设计与实现上对大语言模型(LLMs)进行基准测试:优势、不足与缺陷

Benchmarking LLMs on File System Design and Implementation

Yuqi Xue, Daixuan Li, Jian Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出针对文件系统特定任务的LLM基准框架phi-Bench,含505项六类任务,测试开源与专有LLMs,揭示模型效率、失败原因及缓解技术,将开源phi-Bench。

中文摘要 AI 辅助

大语言模型(LLMs)正从根本上变革计算机系统的研发工作。当我们将LLMs应用于文件系统(fs)开发时,了解它们在领域特定任务中的能力、局限性及运行效率至关重要。我们提出了phi-Bench,这是一个针对fs特定任务的LLM基准测试框架。为便于开展基准测试,我们在phi-Bench中开发了六类任务:基础理解、基础实现、性能建模、调试、优化以及新功能开发。每类任务侧重LLMs的不同能力:指令遵循、知识回忆、推理或编码。为在以最少人力实现广泛覆盖的同时创建高质量任务,除了专家编写和改编自教材的任务外,我们还开发了一条新的AI辅助任务生成流水线。phi-Bench共包含505项任务,我们针对开源LLMs(DeepSeek-V4-Flash、GLM-5.1和MiniMax-M2.7)以及专有LLMs(Claude-Opus-4.7、GPT-5.2和Gemini-3.1-Pro)开展了实证研究。我们的研究揭示了模型在不同任务上的效率、fs任务失败的原因以及缓解LLM失败的技术。我们将开源phi-Bench,以推动利用LLMs进行fs开发的公共研究。

英文摘要

Large Language Models (LLMs) are fundamentally transforming computer system research and development. As we employ LLMs in file system (fs) development, it is essential to understand their capabilities, limitations, and operational efficiency for domain-specific tasks. We present ϕ-Bench, an LLM benchmarking framework for fs-specific tasks. To facilitate benchmarking, we develop six types of tasks in ϕ-Bench: basic understanding, basic implementation, performance modeling, debugging, optimization, and new feature development. Each type emphasizes different LLM capabilities: instruction following, knowledge recall, reasoning, or coding. To create high-quality tasks while achieving broad coverage with minimal human effort, we develop a new AI-assisted task generation pipeline in addition to expert-written and textbook-adapted tasks. With 505 tasks in ϕ-Bench, we conduct an empirical study with both open source (DeepSeek-V4-Flash, GLM-5.1, and MiniMax-M2.7) and proprietary (Claude-Opus-4.7, GPT-5.2, and Gemini-3.1-Pro) LLMs. Our study discloses the model efficiency for different tasks, causes of failed fs tasks, and techniques for mitigating LLM failures. We will open source ϕ-Bench to facilitate public research on using LLMs for fs development.

↑