What the paper is about
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks. The openai story also surfaces in AI Agents Going Rogue Renew Calls..., adding another angle.
What it covers
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu Affiliation: Posts and Telecommunications Institute of Technology, Ha Noi, Vietnam Email: {QuangNM.B21CN629, HieuNPT.B23CE031, HienHC.B23CE028, ToanPQ.B23CN835}@stu.ptit.edu.vn, [email protected] Abstract AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers—violating data-sovereignty laws such as Vietnam’s Decree 53—and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1 , an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues 7.7 × 7.7\times fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35 % 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly 2 × 2\times TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0 % 70.0% to 79.5 % 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks. Together, these results indicate that scalable, self-improving, curriculum-grounded AI tutoring can be democratized in resource-constrained regions. Index Terms: Large Language Models, AI Education, Long-Context Inference, Model Quantization, Data Sovereignty, Self-improving Agents I Introduction Large Language Models (LLMs) have moved beyond text generation into multimodal, tool-augmented agents capable of complex reasoning, planning, and personalized interaction [ 1 , 2 , 3 ] . In education, this puts within reach something long unavailable in resource-constrained systems: a one-on-one tutor that adapts to each student’s pace and cognitive profile [ 4 ] . Our goal is DeepEdu : an AI-tutoring platform that gives Vietnamese students a more effective, curriculum-aligned learning path, with LLM agents as its foundation. Realizing this vision, however, runs into two obstacles that a generic chatbot does not solve. A natural first question is why not simply use a cloud assistant such as ChatGPT. Two reasons rule it out. First, data sovereignty : educational records contain highly sensitive personally identifiable information, and routing them to foreign servers directly violates national data-localization laws such as Vietnam’s Decree 53 [ 5 ] , compelling institutions to self-host open models strictly on-premise. Second, curriculum grounding : pre-trained on English-dominant, Western-centric corpora [ 6 ] , such models are not organized around the Vietnamese textbook curriculum ( Sach Giao Khoa , SGK); their knowledge of local content is unsystematic and frequently hallucinated [ 7 ] . 1 1 1 For instance, early deployments of ChatGPT in Vietnam generated severe historical hallucinations, such as claiming national leaders were founders of the World Health Organization (WHO), or confidently asserting that Emperor Quang Trung and his birth name Nguyen Hue were two distinct kings from opposing dynasties [ 8 ] . Such algorithmic hallucinations highlight the severe pedagogical risks of deploying ungrounded models in regional education. A trustworthy tutor must therefore be both self-hosted and explicitly grounded in localized curricular knowledge. Both requirements are hard on the infrastructure available to public schools. Self-hosting on limited hardware creates a two-fold bottleneck. The static footprint of model weights can be tamed by post-training quantization such as AWQ [ 9 ] or GPTQ [ 10 ] ; but as a tutoring session stretches toward a 1M-token context, the dynamic KV cache surpasses the static footprint and, together with quadratic prefill latency, causes Out-Of-Memory (OOM) failures and sluggish responses on consumer-grade GPUs—a bottleneck that I/O-aware kernels such as FlashAttention [ 11 ] and PagedAttention [ 12 ] alleviate but do not remove. Grounding is equally costly the obvious way: continually fine-tuning on evolving curricula is computationally prohibitive and risks catastrophic forgetting [ 13 ] . We address both obstacles with DeepEdu-v1 , built on the SCALE framework ( S elf-improving C ontext- A ware L earning E ngine). SCALE pairs an efficient long-context inference engine—which amortizes selective sparse attention across similar query chunks to keep latency low as context grows—with a self-improving agentic layer that curates a verified, curriculum-grounded playbook from past interactions rather than updating weights. The main contributions of our work are as follows: 1. We introduce Similarity Chunk Rolling , a long-context inference engine that amortizes selective sparse attention from per-sub-chunk to per-cluster granularity. By issuing a single retrieval per cluster, it reduces retrieval invocations by 7.7 × 7.7\times and prefill latency by roughly 35 % 35% over a state-of-the-art selective-attention baseline (TokenSelect), while preserving—and on structured-retrieval tasks improving—task accuracy. 2. We design a self-improving agentic layer for DeepEdu that, through query-aware context engineering, curates a verified playbook grounding the model in local knowledge and progressively reducing reliance on dominant-language priors—without computationally expensive weight updates. 3. We show, across grounded financial-reasoning and interactive-agent benchmarks, that this self-improving loop delivers the strongest per-track gains over a strong agentic baseline, and that in its deployed configuration DeepEdu sustains a nearly 2 × 2\times TTFT speedup—a step toward democratizing localizable, self-improving AI tutoring in resource-constrained regions under data-sovereignty constraints. II Related Work To contextualize the contributions of DeepEdu, this section surveys recent advancements across the architectural evolution of LLMs and their application to education, the principal techniques for compressing models and optimizing long-context inference under hardware constraints, the challenges of deploying LLMs on low-resource languages and within localized cultural contexts, and the emerging paradigm of self-improving agentic frameworks. II-A The Evolution of Large Language Models The landscape of Natural Language Processing (NLP) has been fundamentally transformed by the scaling of neural architectures, passing through several critical milestones. This trajectory is rooted in the Transformer architecture and early scaling laws [ 14 , 15 ] . Subsequently, the community witnessed the rise of powerful open-weights foundation models that democratized high-performance language modeling [ 1 , 2 ] . To optimize the intrinsic trade-off between model capacity and inference cost, recent architectures have increasingly adopted the Mixture-of-Experts (MoE) paradigm [ 16 , 17 , 18 ] . This architectural evolution has recently culminated in highly efficient, post-training reinforced paradigms (such as DeepSeek-R1 and Qwen3) that emphasize advanced logical reasoning and test-time scaling over pure pattern matching [ 19 , 20 ] . Concurrently, a pivotal paradigm shift has occurred: moving from static sequence-to-sequence generation toward interactive, tool-augmented autonomous agents. This transition was initially catalyzed by models pre-trained on vast codebases, such as Codex [ 21 ] , which demonstrated that LLMs could translate natural language into actionable programmatic commands. Building upon this, frameworks like ReAct [ 3 ] and Toolformer [ 22 ] equipped models with the ability to synergize step-by-step reasoning with external API invocations. As reasoning engines matured, the focus shifted toward multi-agent collaboration and fully autonomous software engineering capabilities. Recent advancements, ranging from open-source frameworks like AutoGen and SWE-agent [ 23 , 24 ] to highly integrated proprietary systems like Claude Code [ 25 ] , showcase agents capable of navigating complex file systems, executing code iteratively, and autonomously self-correcting. Collectively, these advances established LLM agents as a viable substrate for complex, interactive applications — among which personalized education has emerged as a particularly compelling domain. II-B Foundation Models for Education The integration of generative AI into educational technology represents a profound shift from static, rule-based Intelligent Tutoring Systems (ITS) to dynamic, generative frameworks. Early approaches primarily relied on rigid, operation-based formalisms to parse and solve domain-specific tasks, such as mathematical word problems [ 26 ] . However, the advent of Large Language Models has fundamentally upgraded this paradigm. Comprehensive surveys now outline the transition toward utilizing foundation models as core cognitive engines for personalized education, emphasizing their capacity to dynamically tailor feedback, simulate roleplay, and adapt to individual learning styles [ 4 ] . To bridge the gap between general-purpose text generation and strict pedagogical utility, recent advancements have focused on developing specialized, agentic tutoring frameworks. For instance, models such as MathCoder [ 27 ] interleave natural language reasoning with verifiable code execution, ensuring that students receive accurate, step-by-step guidance rather than hallucinated solutions. Furthermore, the integration of LLMs with Knowledge Tracing (KT) paradigms enables these agents to actively model a student’s evolving cognitive state. By maintaining dynamic representations of a learner’s historical performance, state-of-the-art educational agents can autonomously adapt their pedagogical scaffolding—seamlessly shifting between direct hints and deeper Socratic questioning based on real-time comprehension [ 28 ] . Nevertheless, the practical deployment of such educational agents at scale remains constrained by two orthogonal challenges that have received comparatively less attention: the computational cost of running these models on accessible hardware, and their tendency to misalign with localized curricula — issues we now examine in turn. II-C Post-Training Quantization for Efficient Deployment A fundamental prerequisite for deploying state-of-the-art LLMs on consumer-grade hardware is reducing the static memory footprint of model weights. Post-training quantization (PTQ) has emerged as the dominant solution, compressing pre-trained weights into low-bit representations without expensive re-training. GPTQ [ 10 ] pioneered accurate one-shot quantization for generative transformers, employing an approximate second-order procedure to minimize layer-wise reconstruction error at 3 or 4 bits. AWQ [ 9 ] further observes that weight importance is highly non-uniform: a small fraction of salient weights, identified through activation magnitudes, disproportionately influences output quality, and selectively preserving them throughper-channel scaling achieves superior accuracy under hardware-friendly uniform bit-widths. However, quantization addresses only the static weight footprint; the dynamic memory growth of the KV cache during long-context inference remains entirely unresolved, motivating the complementary techniques surveyed next. II-D Knowledge Distillation for Compact Educational Models An alternative compression paradigm transfers the capabilities of large teacher models into smaller student architectures through knowledge distillation (KD). Traditional KD [ 29 ] matches soft probability distributions, but often induces hallucinations in generative tasks due to the mode-covering behavior of forward Kullback-Leibler divergence. MiniLLM [ 30 ] mitigates this through reverse KLD with policy gradient optimization, forcing students to concentrate probability mass on high-quality generation regions. Beyond probability matching, Hsieh et al. [ 31 ] introduced step-by-step distillation that extracts intermediate reasoning trajectories from proprietary teachers, enabling compact students to internalize multi-step problem-solving heuristics — a capability particularly valuable for pedagogical scaffolding in educational tutoring. Nevertheless, distillation pipelines remain computationally intensive at training time and produce student models with static knowledge that cannot adapt to evolving curricula without repeated re-distillation, limiting their applicability for institutions facing both hardware and content-update constraints. II-E Long-Context Inference Optimization Due to the quadratic computational complexity of standard attention, Transformer-based LLMs historically operate within limited pre-training context windows. Extending this window has been pursued through three complementary directions. First, modifications to positional encoding via interpolation and rotary embeddings enable zero-shot length generalization without re-training [ 32 , 33 , 34 , 35 , 36 ] . Second, long-context post-training explicitly adapts model weights to extended sequences, often supported by sequence parallelism techniques [ 37 , 38 , 39 , 40 ] . Third, the integration of specialized memory modules enables processing of effectively unbounded sequences [ 41 , 42 , 43 , 44 ] . However, even with these extensions, long-context inference suffers from prohibitive latency during the prefilling stage. System-level optimizations such as IO-aware FlashAttention and PagedAttention mitigate GPU memory bottlenecks but do not reduce the fundamental computational complexity of attention [ 11 , 12 ] . To process million-token contexts efficiently, a growing body of work exploits attention sparsity through input-independent patterns such as localized windows [ 45 , 46 ] and attention sinks that preserve critical initial tokens [ 47 , 48 ] . While effective, these heuristic patterns can permanently discard informative tokens, motivating dynamic KV cache selection methods that adaptively retain relevant entries [ 49 , 50 ] . State-of-the-art frameworks further refine selection granularity, progressing from block-level memory units [ 51 ] , through query-aware page-level sparsity [ 52 ] and dynamic context selection [ 53 , 54 ] , to fine-grained token-level selection that bypasses block-level inefficiencies entirely [ 55 ] . While token-level methods such as TokenSelect substantially reduce per-step attention cost, the interaction between selective sparse attention and chunked prefill at very long contexts remains an active area of investigation. II-F LLMs on Low-Resource Languages While modern LLMs exhibit remarkable zero-shot and few-shot reasoning capabilities in high-resource languages like English, their performance often degrades significantly when applied to low-resource languages (LRLs) such as Vietnamese. This degradation is primarily attributed to a severe underrepresentation of LRL tokens in the pre-training corpora, which limits the models’ native linguistic alignment. To circumvent this linguistic bottleneck without resorting to computationally prohibitive pre-training from scratch, recent literature has explored several efficient adaptation strategies. At the inference level, Nguyen et al. [ 56 ] propose leveraging the models’ dominant English capabilities through linguistically-diverse prompts. This approach acts as a cross-lingual bridge, effectively anchoring LRL inputs to the models’ highly structured English latent space to improve zero-shot generalization. Complementarily, Cahyawijaya et al. [ 6 ] demonstrate that LLMs possess latent cross-lingual capabilities that can be explicitly activated via In-Context Learning (ICL). Their empirical findings show that providing a limited number of task-specific demonstrations in the target language allows the model to dynamically adapt its reasoning pathways to the LRL without requiring any parameter updates. For applications requiring deeper structural alignment, prompt engineering alone is often insufficient. To address this, Nag et al. [ 13 ] introduce efficient continual pre-training (CPT) techniques tailored for LRLs. Their methodology natively injects new linguistic knowledge into the foundation model’s weights while actively mitigating the catastrophic forgetting of general reasoning capabilities, all under strict computational constraints. Collectively, these methodologies highlight that effective LRL adaptation can be achieved without invasive parameter updates, particularly when combined with carefully curated in-context signals. II-G Hallucination in Localized and Cultural Contexts While modern LLMs demonstrate strong generalization capabilities, their deployment in region-specific applications is frequently hindered by factual hallucinations. Ji et al. [ 57 ] identify that hallucinations fundamentally stem from data divergence, where models learn spurious correlations from imbalanced pre-training corpora. This issue is severely amplified by the Western-centric nature of foundational models. As Huang et al. [ 58 ] highlight, when LLMs process queries outside their dominant data distribution, they often generate confident confabulations rather than acknowledging knowledge boundaries. Empirical evaluations by Bang et al. [ 59 ] confirm this limitation, demonstrating that LLMs exhibit significantly higher hallucination rates and degraded reasoning when applied to low-resource languages and non-Western cultural contexts. In educational settings, relying on such ungrounded models introduces critical pedagogical risks. These findings establish that scaling pre-training alone cannot resolve the localization gap, particularly in educational deployments where factual accuracy on regional content is non-negotiable. This persistent gap motivates runtime grounding mechanisms capable of injecting localized knowledge without the prohibitive cost of continual re-training. II-H Agentic Systems and Self-Evolving Context The rapid evolution of LLMs has introduced a paradox in knowledge management: while models store vast generalized knowledge within their parametric memory, their ability to seamlessly integrate deeply specialized domain heuristics remains limited. Traditional fine-tuning is rigid, computationally expensive, and susceptible to catastrophic forgetting. Conversely, current prompt optimization methods often suffer from brevity bias — the tendency to compress prompts into short, generic instructions that strip away essential technical nuances required for high-precision tasks [ 60 ] . Furthermore, attempts at monolithic rewriting in long-context windows frequently trigger context collapse, wherein iterative LLM rewrites progressively erode accumulated domain knowledge into uninformative summaries. To overcome these barriers and build agentic systems that self-improve without weight updates, recent literature explores natural language feedback and verbal reinforcement learning [ 61 , 62 , 63 ] . Techniques leveraging reflective prompt evolution [ 64 ] and agentic context engineering [ 65 ] effectively transform the static context into a living, continuously updated "playbook." These methodologies enable LLM agents to adapt and fine-tune their behavior exclusively through dynamic context memory, preserving granular domain insights without the massive overhead of standard fine-tuning [ 66 ] . II-I Discussion and Positioning Table I summarizes where DeepEdu sits relative to the four directions above. Each addresses one facet of on-premise educational deployment but leaves the others open: quantization shrinks static weights yet does nothing for the dynamic long-context cost; selective sparse attention accelerates long contexts but neither self-improves nor localizes; agentic context engineering self-improves without weight updates but is not built for million-token efficiency or a specific curriculum; and low-resource-language adaptation localizes, but through weight updates that are costly to maintain as curricula evolve. DeepEdu is, to our knowledge, the first framework to occupy their intersection: it couples a cluster-level long-context engine with a self-improving, curriculum-grounded agentic layer, so that a single on-premise system is efficient, self-improving, and localizable at once. TABLE I : Positioning of DeepEdu against representative prior directions. ✓: directly addressed; (✓): partial; –: not addressed. Columns: on-premise/data-sovereignty fit, long-context efficiency, self-improvement without fine-tuning, and localization. Direction On-prem. Long-ctx Self-impr. Local. Quantization [ 9 , 10 ] ✓ – – – Sparse long-context [ 55 , 51 ] – ✓ – – Agentic context [ 65 , 62 ] (✓) – ✓ – LRL adaptation [ 13 , 6 ] – – – ✓ DeepEdu (ours) ✓ ✓ ✓ ✓ III Preliminaries This section establishes the notation and the two paradigms our framework builds on. We first summarize Transformer inference and Multi-Head Attention (§ III-A ) and formalize long-context inference as a Selective Sparse Attention problem (§ III-B )—the basis of our inference engine—then define Agentic Context Engineering (§ III-C ), the paradigm underlying our self-improving agentic layer. III-A LLM Inference and Multi-Head Attention Notation Throughout this paper, we adopt the following unified notation. Let d d denote the model hidden dimension, H H the number of attention heads, and d h = d / H d_{h}=d/H the per-head dimension. For a transformer layer processing n n input tokens, we denote its hidden states by 𝐗 ∈ ℝ n × d \mathbf{X}\in\mathbb{R}^{n\times d} . The KV Cache accumulated from previously processed tokens has length N N , and we use h ∈ { 1 , … , H } h\in{1,\ldots,H} to index attention heads. Multi-head attention Given hidden states 𝐗 \mathbf{X} , each head h h projects 𝐗 \mathbf{X} into per-head queries, keys, and values 𝐐 ( h ) , 𝐊 ( h ) , 𝐕 ( h ) ∈ ℝ n × d h \mathbf{Q}^{(h)},\mathbf{K}^{(h)},\mathbf{V}^{(h)}\in\mathbb{R}^{n\times d_{h}} and computes scaled dot-product attention (SDPA) [ 14 ] : 𝐎 ( h ) = softmax ( 𝐐 ( h ) ( 𝐊 ( h ) ) ⊤ d h ) 𝐕 ( h ) , \mathbf{O}^{(h)}=\text{softmax}!\left(\frac{\mathbf{Q}^{(h)}(\mathbf{K}^{(h)})^{\top}}{\sqrt{d_{h}}}\right)\mathbf{V}^{(h)}, (1) whose per-head outputs are concatenated and projected back to dimension d d . The inner softmax in Equation ( 1 ) has quadratic complexity 𝒪 ( n 2 d h ) \mathcal{O}(n^{2}d_{h}) in sequence length—the principal bottleneck of long-context inference. Prefill, decode, and the KV Cache Inference has two stages. Prefill processes the whole prompt of length n i n n_{in} in one pass and persists the per-head keys and values as a KV Cache 𝐊 cache ( h ) , 𝐕 cache ( h ) ∈ ℝ N × d h \mathbf{K}^{(h)}{\text{cache}},\mathbf{V}^{(h)}{\text{cache}}\in\mathbb{R}^{N\times d_{h}} (initially N = n i n N=n_{in} ); it is compute-bound and dominates latency for long inputs. Decode then generates one token per pass, appending its key and value to the cache so that N N grows by one each step. Each decode step is light, but the KV Cache footprint grows linearly with context, straining consumer-grade hardware. Two-level chunking for long-context prefill To bound peak memory consumption during prefill and enable processing of inputs that exceed available VRAM, modern inference frameworks such as vLLM [ 12 ] and SGLang adopt chunked prefill: the input sequence is partitioned into consecutive outer chunks, each of length at most L outer L_{\text{outer}} tokens (typically L outer = 8192 L_{\text{outer}}=8192 ). Each outer chunk is processed in a separate forward pass; the last outer chunk may be shorter when n i n n_{in} is not a multiple of L outer L_{\text{outer}} . This outer chunking caps the peak activation memory of any single forward pass but does not, by itself, reduce the cost of attending to the growing KV Cache: a naive implementation still incurs 𝒪 ( ℓ ⋅ N ⋅ d h ) \mathcal{O}(\ell\cdot N\cdot d_{h}) attention cost per outer chunk under Equation ( 1 ), where ℓ ≤ L outer \ell\leq L_{\text{outer}} denotes the actual length of the current outer chunk and N N is the cumulative KV Cache length. To further reduce per-step attention cost, recent selective sparse attention methods such as TokenSelect [ 55 ] introduce a second, finer-grained level of partitioning within each outer chunk. Specifically, the queries of an outer chunk of length ℓ \ell are subdivided into M = ⌈ ℓ / L ⌉ M=\lceil\ell/L\rceil sub-chunks, denoted { 𝐗 1 , 𝐗 2 , … , 𝐗 M } {\mathbf{X}{1},\mathbf{X}{2},\ldots,\mathbf{X}{M}} . Each sub-chunk has length L L (typically L = 512 L=512 ), except possibly the last, which has length ℓ − ( M − 1 ) L ≤ L \ell-(M-1)L\leq L . For each sub-chunk, the method invokes a query-aware selection function to identify a small subset of critical KV Cache entries, then computes sparse attention restricted to that subset. Crucially, this selection-then-attention procedure is applied sequentially to each sub-chunk within the outer chunk: M M independent retrieval calls followed by M M independent attention calls per outer chunk. While this scheme effectively bypasses the quadratic attention cost, the sequential nature of the inner loop exposes a substantial opportunity for amortization when consecutive sub-chunks share similar query representations — a redundancy our work directly targets. III-B Selective Sparse Attention The quadratic cost of SDPA in Equation ( 1 ), combined with the abundant sparsity empirically observed in LLM attention maps [ 51 , 52 , 55 ] , motivates the selective sparse attention (SSA) paradigm: rather than attending to the entire KV Cache, only a small subset of critical tokens is dynamically selected for each query. Our inference engine builds directly on this paradigm; we restate it below in our notation, following the formulation of TokenSelect [ 55 ] . Definition 1 (Selective Sparse Attention) . Consider a transformer layer at any inference step, with per-head queries 𝐐 ( h ) ∈ ℝ C × d h \mathbf{Q}^{(h)}\in\mathbb{R}^{C\times d{h}} for C C current input tokens ( C = 1 C=1 during decoding, C ≤ L C\leq L for one sub-chunk during prefill where L L is the sub-chunk size defined in § III-A ), and a KV Cache 𝐊 cache ( h ) , 𝐕 cache ( h ) ∈ ℝ N × d h \mathbf{K}^{(h)}{\text{cache}},\mathbf{V}^{(h)}{\text{cache}}\in\mathbb{R}^{N\times d_{h}} . The full attention output following Equation ( 1 ) is: 𝐎 ( h ) = softmax ( 𝐐 ( h ) ( 𝐊 cache ( h ) ) ⊤ d h ) 𝐕 cache ( h ) . \mathbf{O}^{(h)}=\text{softmax}!\left(\frac{\mathbf{Q}^{(h)}(\mathbf{K}^{(h)}{\text{cache}})^{\top}}{\sqrt{d{h}}}\right)\mathbf{V}^{(h)}{\text{cache}}. (2) Building on Equation ( 2 ), selective sparse attention restricts the softmax to a subset of k ≪ N k\ll N cache entries, yielding the sparse approximation 𝐎 ^ ( h ) \hat{\mathbf{O}}^{(h)} : 𝐎 ^ ( h ) = softmax ( 𝐐 ( h ) ( 𝐊 select ( h ) ) ⊤ d h ) 𝐕 select ( h ) , \hat{\mathbf{O}}^{(h)}=\text{softmax}!\left(\frac{\mathbf{Q}^{(h)}(\mathbf{K}^{(h)}{\text{select}})^{\top}}{\sqrt{d_{h}}}\right)\mathbf{V}^{(h)}{\text{select}}, (3) The two formulations differ only in which cache rows enter the softmax: the selected matrices 𝐊 select ( h ) , 𝐕 select ( h ) ∈ ℝ k × d h \mathbf{K}^{(h)}{\text{select}},\mathbf{V}^{(h)}{\text{select}}\in\mathbb{R}^{k\times d{h}} index 𝐊 cache ( h ) \mathbf{K}^{(h)}{\text{cache}} and 𝐕 cache ( h ) \mathbf{V}^{(h)}{\text{cache}} at the positions ℐ \mathcal{I} chosen by a selection function 𝒮 \mathcal{S} : ℐ = 𝒮 ( 𝐐 , 𝐊 cache ) , ℐ ⊆ { 1 , … , N } , | ℐ | = k . \mathcal{I}=\mathcal{S}!\left(\mathbf{Q},\mathbf{K}{\text{cache}}\right),\quad\mathcal{I}\subseteq{1,\ldots,N},\quad|\mathcal{I}|=k. (4) Here 𝐐 \mathbf{Q} and 𝐊 cache \mathbf{K}{\text{cache}} denote the multi-head queries and cached keys aggregated across all heads, allowing 𝒮 \mathcal{S} to leverage cross-head information when scoring token criticality. The design objective is to construct 𝒮 \mathcal{S} such that the approximation gap ‖ 𝐎 ( h ) − 𝐎 ^ ( h ) ‖ 2 2 \big|\mathbf{O}^{(h)}-\hat{\mathbf{O}}^{(h)}\big|_{2}^{2} is minimized while keeping k k within the hardware budget. Selection functions 𝒮 \mathcal{S} range from fixed patterns (attention sinks and recent tokens), through query-independent scoring, to query-aware token-level selection (surveyed in § II-E ). Our engine builds on the query-aware token-level scheme of TokenSelect [ 55 ] , which conditions 𝒮 \mathcal{S} on the current query 𝐐 \mathbf{Q} for high approximation quality, at the cost of a retrieval call per step. Sub-chunk redundancy in selective sparse attention Together, Equations ( 3 )–( 4 ) define a single selection-then-attention step, the unit our inference engine reorganizes. A critical observation underexplored by prior work is that, under the two-level chunking scheme of § III-A , this step—a The openai story also surfaces in Google opens early access to AI..., adding another angle. as detailed in the full paper on Arxiv The openai story also surfaces in MIT Researchers Use AI Agents to..., adding another angle.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!