arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Flama:用于开发和部署生产级API、机器学习及大语言模型服务的Python框架

Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

José A. Perdiguero López, Miguel A. Durán-Olivencia

arXiv 2608.18733首次发表:更新:

发表机构

Vortico Tech(沃尔蒂科科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Flama是一款开源Python框架,基于ASGI构建,统一REST API开发、ML服务与LLM推理,含七大子系统及多项内置功能,可支持多后端LLM部署等场景。

AI 中文摘要

我们提出Flama,这是一个用于开发和部署生产级Web API、机器学习服务以及大语言模型(LLM)应用的开源Python框架。Flama构建于异步服务器网关接口(ASGI)之上,提供一种以类型驱动、异步优先的编程模型,将REST API开发、预测模型服务和生成式AI推理统一在同一架构中。它围绕七个子系统构建:基于组件的依赖注入系统,在启动时从类型注解解析处理程序参数;可插拔的模式层,通过单一适配器支持Pydantic、Marshmallow和Typesystem;自动CRUD生成器,将SQLAlchemy表和模式类转换为由仓库模式和工作单元模式支持的REST端点;可移植二进制格式(.flm),打包来自scikit-learn、TensorFlow、PyTorch和Hugging Face Transformers的模型及其元数据,实现零代码部署;多后端LLM服务器,运行vLLM(Linux/CUDA)或MLX(Apple Silicon),并通过共享编解码器暴露四种有线协议(OpenAI、Anthropic、Ollama及原生流式方言);通过Maturin编译的Rust加速核心,用于路由、JSON编码、压缩和解析;以及模型上下文协议模块,将任何应用转换为基于JSON-RPC 2.0的MCP服务器。内置功能包括JWT认证、两种分页策略、线程或进程中的后台任务、WebSocket端点、服务器发送事件和NDJSON流式传输、从处理程序签名生成OpenAPI 3.2.0,以及用于运行应用、服务、打包和检查模型的命令行界面。我们描述了架构,通过示例展示编程模型,并将Flama与现有框架、模型服务平台和LLM推理引擎进行比较。

英文摘要

We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications. Built on the Asynchronous Server Gateway Interface (ASGI), Flama offers a type-driven, async-first programming model that unifies REST API development, predictive model serving, and generative AI inference in one architecture. It is organised around seven subsystems: a component-based dependency injection system resolving handler parameters from type annotations at startup; a pluggable schema layer supporting Pydantic, Marshmallow and Typesystem behind a single adapter; an automatic CRUD generator turning a SQLAlchemy table and a schema class into REST endpoints backed by the Repository and Unit of Work patterns; a portable binary format (.flm) packaging models from scikit-learn, TensorFlow, PyTorch and Hugging Face Transformers with their metadata for zero-code deployment; a multi-backend LLM server running vLLM (Linux/CUDA) or MLX (Apple Silicon) and exposing four wire protocols (OpenAI, Anthropic, Ollama, and a native streaming dialect) through a shared codec; a Rust-accelerated core compiled via Maturin for routing, JSON encoding, compression and parsing; and a Model Context Protocol module turning any application into an MCP server over JSON-RPC 2.0. Built-in capabilities include JWT authentication, two pagination strategies, background tasks in threads or processes, WebSocket endpoints, Server-Sent Event and NDJSON streaming, OpenAPI 3.2.0 generation from handler signatures, and a command-line interface for running applications and for serving, packaging and inspecting models. We describe the architecture, present the programming model through worked examples, and compare Flama with existing frameworks, model serving platforms and LLM inference engines.

Comments83 pages, 6 figures, 1 table. Software available at https://github.com/vortico/flama, up-to-date documentation at https://flama.dev

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑