arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

xMIx:用于机械可解释性应用的高性能服务时间平台

xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps

Michael Blum, Mark Silberstein, Yaniv David

arXiv 2607.22595首次发表:更新:

发表机构

Technion(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机械可解释性在生产模型服务系统中部署难的问题,提出xMIx框架,可在模型运行时预定义位置附加MI功能,支持条件调用,集成到vLLM服务系统后性能与原生相当,各指标仅有小幅下降。

AI 中文摘要

机械可解释性(MI)已成为分析和干预推理计算的有力方法,应用日益增多。但现有MI框架运行时开销过高,在生产模型服务系统中部署不实用。其根本问题是MI功能与服务模型组合不佳。我们提出xMIx,一个在生产推理服务环境中部署MI应用的原生框架。它能在模型运行时的预定义位置附加MI功能,支持条件调用,多个MI应用可在单个模型实例中部署,编译后动态激活,性能成本可忽略不计。我们将xMIx与vLLM服务系统集成并评估,结果显示其性能与原生vLLM执行相当,各指标仅有小幅下降。

英文摘要

Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbreak attempt detection, truthfulness evaluation, and hallucination detection. Unfortunately, MI deployment in production model-serving systems is currently not practical, as most existing MI frameworks introduce prohibitively high runtime overheads. The fundamental problem is that MI functions do not compose cleanly with served models: they fragment deployment, often force draining requests and rebuilding serving state, and conflict with critical performance optimizations such as continuous batching and CUDA-graph execution, essential for production deployments. We present xMIx, a serving-native framework for deploying MI applications in production inference serving environments. xMIx enables attaching MI functions to a predefined set of locations in the model runtime, interposing on activations within the layers and residual streams. xMIx supports conditional invocation of MI functions depending on the outputs in preceding model layers. Multiple MI applications can be deployed in a single model instance. xMIx compiles them all into the serving path but activates them dynamically at runtime only when necessary, with negligible performance cost, and without requiring a separate model instance or alternative execution stack. We integrate xMIx with the vLLM serving system and evaluate it across three major models and seven diverse MI applications. xMIx achieves performance comparable to native vLLM execution, incurring a slowdown of 1.3% mean inter-token latency (ITL), 1.2% for tail P99 ITL, 2.6% for mean time to first token (TTFT), and 1.6% for mean total token throughput (TTT).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑