arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31568cs.AI

DeepEdu-v1:面向越南教育的高效可扩展智能体大语言模型

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu

首次发表
浏览论文内容

中文总结 AI 辅助

DeepEdu-v1提出基于SCALE框架的越南教育AI辅导系统,通过长上下文推理引擎和自改进智能体层,实现2倍TTFT加速并将智能体准确率从70.0%提升至79.5%。

中文摘要 AI 辅助

人工智能辅导有望显著改善越南等发展中地区学生的学习成果,然而两条显而易见的路径都存在不足。云助手(如ChatGPT)将敏感的学生数据路由到外国服务器,违反了越南第53号法令等数据主权法律;并且,由于这些模型在西方中心化的语料库上预训练,并未围绕国家教科书课程组织,因此其对本地内容的知识缺乏系统性且经常产生幻觉。自托管开源模型可将数据保留在本地,但面临双重障碍:训练后量化(AWQ、GPTQ)解决了静态权重占用问题,但长辅导上下文的动态KV缓存和预填充延迟仍会导致消费级GPU上出现内存不足故障和响应缓慢,同时模型在区域特定材料上仍会产生幻觉。我们提出了DeepEdu-v1,一个基于SCALE(自我改进的上下文感知学习引擎)构建的越南教育AI辅导系统,该框架包含两项创新。首先,长上下文推理引擎将令牌选择从每个子块粒度摊销到每个簇粒度;在长上下文检索中,其发出的检索调用比最先进的选择性注意力基线少7.7倍,将预填充延迟(TTFT)降低约35%,同时匹配或提高任务准确性。其次,自我改进的智能体层持续从过去的交互中整理经过验证的操作手册,而非进行微调,这一设计旨在随着可信本地知识的积累,逐步减少对主流语言先验的依赖。在其部署配置中,DeepEdu相比标准vLLM服务实现了近2倍的TTFT加速,并将复杂任务上的智能体准确性从70.0%提升至79.5%,在金融推理和交互式智能体基准测试中取得了最强的每轨道增益。

英文摘要

AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.

发表机构

  • Posts and Telecommunications Institute of Technology(邮政与电信技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑