DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
DeepEdu-v1 tutors Vietnamese students on-premise, cutting retrieval calls 7.7x and lifting agentic accuracy from 70% to 79.5%.
Researchers present DeepEdu-v1, an on-premise Vietnamese education tutor built on the SCALE framework to avoid foreign-cloud data transfer and curriculum hallucinations. Its long-context engine selects tokens per cluster, issuing 7.7 times fewer retrieval calls than a selective-attention baseline and cutting time-to-first-token by about 35%. A self-improving agent stores a verified playbook from past interactions instead of fine-tuning. Deployed, it nearly doubles TTFT speed versus standard vLLM and raises agentic accuracy from 70.0% to 79.5%.
- On-premise design targets Vietnam Decree 53 data-sovereignty limits
- Cluster token selection issues 7.7x fewer retrieval calls
- Prefill latency falls about 35% versus selective attention
- Deployed setup nearly doubles TTFT speed versus vLLM
- Agentic accuracy rises from 70.0% to 79.5% without fine-tuning
Full article244 words · extracted from arxiv.org · click to collapse
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31568