Articles

2026년 9월 28일Jikai LiAGITech Talk연구

스파스 모델: 기원, 중요한 이유, 그리고 LLM의 최전선과 미래

Product

무료로 시작하기(준비 중) 엔터프라이즈 상담

Career

특화형 AI 에이전트 서비스 개발을 함께할 분을 적극 모집하고 있습니다!

Overview

As of 2026, sparsity is no longer an optional efficiency trick. It is the default architectural constraint of frontier large models, and it decides both what a large model costs to run and how accurately a long-document agent can read. The most direct evidence is DeepSeek-V4.1-Flash, released on September 10, 2026. Its backbone has 552B parameters, yet it activates only 8B per token during prefill and 16B during decode, and the global KV cache resident in GPU memory has been compressed to 890 bytes per token, 1/437 of DeepSeek-V1’s (official DeepSeek announcement, technical report arXiv:2609.19969).

Three key takeaways

  • The lineage of sparsity is continuous. From Barlow’s and Olshausen–Field’s “sparse coding in the brain,” through the “computable sparsity” of wavelets, LASSO, and compressed sensing, to the “sparsity as architecture” of MoE, sparse attention, and conditional memory, the same prior sits underneath: only a very few components truly matter. What the three successes share is that their sparsity pattern fit the hardware. Unstructured weight pruning lost not because of theory but because of hardware.

  • Why now. Scaling ran into inference cost and the memory wall, and long-context agents turned the KV cache and prefill into the dominant costs. Sparsity’s role therefore widened from “saving compute” to “saving memory, bandwidth, and storage.” Mainstream MoE in 2024 activated roughly 25% of its parameters per token (Mixtral); by 2026 the figure has fallen to 1.5%–4.5% (DeepSeek-V4.1, Kimi K3, Qwen3.5).

  • Sparsity is no panacea. The largest evaluation of sparse attention to date (Nawrot et al., Findings of ACL 2026) showed that nearly every sparse configuration suffers a large performance drop on at least one task. In interpretability, sparse autoencoders (SAEs) ran into systematic negative results in 2025. Model selection must be based not on average scores but on the tasks that are hardest for your own business.

Two threads that run through this survey

  • The five kinds of sparsity: activation sparsity, weight sparsity, structured sparsity, conditional computation (MoE), and attention sparsity. The debate over “whether sparsity helps” cannot even begin until one specifies which kind.

  • The three operators: soft thresholding (the proximal operator of L1), hard thresholding (L0), and top-k projection (projection onto the L0 ball). From wavelet denoising in 1994 to MoE routing in 2026, the step that “produces zeros” is nothing more than one of these three operators applied to a different object. Almost every modern large model uses the third, because it fixes the compute per token exactly and aligns with the parallel granularity of the hardware.

Five terms that recur throughout

TermIn a nutshell
TokenThe smallest unit of text a model processes: a word, or part of one. “How many parameters are activated per token” means how much compute is spent each time a little more text is emitted
Parameters and activated parametersParameters are all the numbers a model stores (how much it knows). Activated parameters are the part actually used in computation when processing one token (what it costs to run). Sparsification means decoupling the two
Basis and coordinatesA set of “building blocks.” Data can be written as a weighted sum of them, and the weights are the coordinates. Change the blocks and the coordinate system changes: the data stays the same, but the values you read off change
NormsL2 is length, L1 is the sum of the absolute values of the components, and L0 is the number of nonzero entries. Sparse means a small L0
FLOPs and memory bandwidthCompute throughput is the number of multiply–accumulate operations per second; bandwidth is the number of bytes that can be moved out of GPU memory per second. In the decode stage of modern large models, the bottleneck is usually the latter

A note on terminology: This survey uses “sparse” strictly to mean “most of the elements are zero.” “Low rank” means “no element is zero, but there are few degrees of freedom” (MLA, LoRA, and so on). The two are often confused; this survey keeps them consistently distinct.

How to read this survey

This survey proceeds chronologically, and in each era it answers the same four questions. What were the background and the bottleneck at the time? By what mechanism did sparsity appear? What did it solve, and at what cost? How does it relate to today’s large models? The section “Prehistory” asks why the world is sparse; “The Classical Era” shows how mathematics turned that into a tool; and “The Neural Network Era” explains why pruning lost once sparsity entered neural networks. The central section, “The Present,” discusses why large models went all in on sparsification in 2023–2026 and systematically compares each lab’s approach. “Where Sparsity Meets Deep Learning” covers the fusion of sparsity and deep learning in imaging and signal processing, and the sections “Applications,” “Benefits and Costs,” “The Future,” and “What This Means for Us” take up applications, costs, the future, and what all of this means for us. The appendices collect a timeline, data, anecdotes, and references.

Prehistory: Why the World Itself Is Sparse

Sparsity is not a property of data alone; it is a property of “data plus the way we describe it.” The world becomes sparse under the right basis, and the history of science is in part the history of searching for that basis. This section lays out three kinds of empirical evidence and one intellectual wellspring, and then marks their limits.

Figure 1. Basis functions learned by sparse coding on natural images: a reproduction of the Olshausen & Field (1996) experiment. 192 basis functions of 16 × 16 pixels were learned from the whitened natural-image set that Olshausen distributes with the original code, under only two constraints: reconstruct the patches, and keep the coefficients sparse. They emerge localized, oriented, and band-pass, like the receptive fields of V1 simple cells. Source: computed by the author for this survey following Olshausen & Field, “Emergence of simple-cell receptive field properties by learning a sparse code for natural images,” Nature 381, 607–609 (1996); a reproduction of the experiment, not the original figure.

Philosophy and information theory: intuition and framework, not technical origins

Occam’s razor is a philosophical principle that favors simplicity; it is not the technical starting point of sparse models. Shannon (1948) quantified “redundancy”: whatever part of a code length exceeds the entropy of the source is compressible redundancy, and a sparse representation is, at heart, a basis that brings that redundancy to the surface. Kolmogorov complexity (1960s) offered the ideal of “the shortest description,” but it is uncomputable; Rissanen’s MDL (1978) turned it into a practical criterion for model selection. A sparse model can be read as a special case of MDL: description length ≈ number of nonzeros × the coding cost per coefficient. These ideas should be written up as the “intellectual prehistory,” not as the “inventors of sparse models.”

Three kinds of evidence from nature

Natural images. The power spectrum of natural images approximately follows a 1/f² power law and exhibits approximate scale invariance (Field 1987; Ruderman & Bialek 1994; for a review, Simoncelli & Olshausen 2001). In a local, band-pass basis such as wavelets, the coefficients of natural images are heavy-tailed: most are close to zero, and only a few, at the edges, are large. Nonlinear approximation theory (DeVore 1998; Donoho 1993) proved that for piecewise smooth signals, the error of the best n-term approximation in a wavelet basis decays faster than in any fixed linear basis. This is the theoretical foundation of both JPEG2000 and compressed sensing, and the Haar example in Figure 3 is its smallest demonstration.

Language. Word frequencies follow Zipf’s law: a handful of words account for most occurrences, and the tail is extremely long. Bag-of-words and TF-IDF representations are therefore extremely sparse. The same statistics explain why the load on MoE experts is inherently skewed, and why DeepSeek uses an N-gram lookup table as conditional memory (Engram).

The brain. Barlow (1961) proposed the redundancy reduction hypothesis: the purpose of a sensory system is to strip out redundancy and build statistically independent representations. This is the neuroscientific source of sparse coding. Attwell & Laughlin (2001) worked out the energy budget of gray matter, and from it Lennie (2003, Current Biology 13:493–497) estimated that, because a single spike is so expensive, “the number of neurons that can be substantially active at the same time is probably severely limited to less than 1%.”

One correction is unavoidable here. The widely repeated claim that “the brain uses only 1–4% of its neurons” is a secondhand citation. When Glorot, Bordes, and Bengio (2011) cited Lennie, they wrote “1–4%,” and the figure has been reprinted ever since. Lennie’s original says “possibly fewer than 1%,” and even that is an estimated upper bound derived from energy constraints, not a measurement. Sparsity means “the fraction firing strongly at any one time is low”; it does not mean “most of the brain is idle.”

Limits and counterexamples

Sparse is not low rank. Sparse means few nonzero coefficients under some basis or dictionary (measured by L0/L1). Low rank means few singular values in a matrix (measured by rank or the nuclear norm). LoRA and MLA are low-rank; MoE, top-k attention, and SAEs are sparse. A truncated SVD yields a low-rank matrix; the only thing sparse about it is its singular-value spectrum.

The world is not always sparse. Robust PCA (Candès, Li, Ma & Wright 2011) decomposes data into “low-rank plus sparse” and showed that real data often has both: in video, the background is low-rank and the foreground is sparse. Signals dominated by texture or noise are not sparse in the commonly used bases. Feature superposition in the residual stream of an LLM means the representation is dense at the level of individual neurons and becomes sparse only under an overcomplete dictionary. That is the starting point of the SAE discussion in “The Present” (the section “Sparsity in interpretability”).

Summary of this section. Sparsity is a falsifiable prior, not a universal truth. It has been confirmed again and again in natural images, language, and neural activity, so it is worth “betting on.” But before placing the bet, one question must always be asked: sparse under which basis?

The three kinds of evidence in detail: why exactly these statistical laws

Where does the 1/f spectrum of natural images come from? There are two complementary explanations. The first is scale invariance. The sizes of objects in a natural scene span many orders of magnitude, and the statistics of an image do not change whether you zoom in or out; mathematically, that forces the power spectrum into a power law. The second is the occlusion model. A scene consists of objects overlapping one another front to back; edges produce discontinuities, and discontinuities appear in the frequency domain as a slowly decaying high-frequency tail. Both explanations arrive at the same conclusion: natural images are neither “random noise” nor “smooth everywhere,” but “wide flat regions plus a few edges.” This is precisely the structure a wavelet basis handles best. Field (1987) went further and pointed out that if the bandwidth of the receptive fields of simple cells in visual cortex matches these statistics, each cell’s response to natural images will be heavy-tailed: barely responding most of the time, firing strongly now and then. This is the original meaning of the term “sparse coding” in neuroscience.

Why is the Olshausen and Field experiment so convincing? The 1996 experiment placed only two constraints on the algorithm: reconstruct patches of natural images with a set of linear basis functions, and make the coefficients as sparse as possible. There was no prior knowledge about the brain at all. Yet the learned basis functions came out local, oriented, and band-pass, and closely matched the receptive fields of V1 simple cells in cats and monkeys. The force of the result lies in its direction: it was not that “an algorithm was built by taking inspiration from the brain,” but that “the sparsity principle independently predicted the structure of the brain.” Later, Hromádka, DeWeese, and Zador (2008) recorded highly sparse firing patterns in the auditory cortex of awake rats, and Quian Quiroga et al. (2005) discovered “concept cells” in the human hippocampus that respond selectively to particular people or landmarks. Both are taken as evidence for the sparse coding hypothesis. There is dissent. Spanne and Jörntell (2015), among others, argue that strict sparse coding is not universal in cortex and that the degree of sparsity depends on brain region and task. It should therefore be written up as a “widely supported hypothesis,” not as settled doctrine.

Why does energy force sparsity? The human brain is about 2% of body weight but consumes about 20% of resting metabolism, roughly 20 watts. Attwell and Laughlin (2001) estimated that most of the energy in gray matter goes to synaptic transmission and action potentials. Lennie (2003) worked backward from there: if a single spike costs this much, the cortex cannot afford to have a large fraction of its neurons firing at high rates simultaneously. Because the argument is an “upper bound under a budget constraint,” it yields an estimate, “probably fewer than 1%,” rather than a measurement. Its significance lies in the logic, not the specific number. In a system with limited energy or bandwidth, sparsity is not an option but a necessity. The same logic is replayed in “The Neural Network Era” as the “memory wall” on GPUs.

Zipf’s law and the long tail in language. Zipf (1935) observed that a word’s frequency is roughly inversely proportional to its rank: the second most common word appears about half as often as the first, the tenth about one-tenth as often. Mandelbrot later gave a corrected form. The direct consequence is that any text uses only a tiny fraction of the vocabulary, and most words appear only once or twice. In the era of statistical language models, this showed up as extremely sparse n-gram frequency matrices and as the smoothing problem of “zero probabilities” (Good–Turing, Kneser–Ney). In the era of large models it has returned in two forms. One is the inherently skewed routing load in MoE: popular experts take on most of the tokens, so a load-balancing mechanism has to intervene. The other is DeepSeek’s Engram, which brings the N-gram lookup table back into a frontier architecture, sending frequent patterns down a cheap lookup path and leaving only what genuinely needs reasoning to the expensive Transformer layers.

Sparsity in physics and engineering

Sparsity has a longer history in engineering than in machine learning. In scientific computing, the coefficient matrices of the finite element method and of power-flow equations are inherently sparse: each node couples only to its neighbors. Tinney and Walker (1967), solving power systems, proposed an optimal ordering that reduces “fill-in” during Gaussian elimination. That was the starting point of sparse direct methods. Rose (1972) then described the elimination process in graph-theoretic terms, Gilbert and others developed sparse LU factorization, and sparse linear algebra became a field of its own. In seismic exploration, the subsurface reflectivity series is modeled as a sparse train of spikes, and the deconvolution problem can be solved with L1 regularization, 20 years ahead of compressed sensing. In radar and array signal processing, targets are sparsely distributed in the space of angle, range, and velocity, and sparse recovery is used for super-resolution estimation.

What these fields share is that the sparsity pattern comes from physical structure (adjacency, reflections, the number of targets), which makes it predictable and exploitable. This is continuous with wavelets in “The Classical Era” (where the structure comes from piecewise smoothness) and with MoE in “The Present” (where the structure comes from expert blocks), and it also explains why unstructured weight sparsity loses in “The Neural Network Era”: its zeros have no physical structure, and the hardware cannot exploit them.

Mirrors in social phenomena, and the limits of the idea

Pareto’s 80/20 rule, Anderson’s long tail, and the factor-analysis tradition that “a few important variables explain most of the variance” are all mirrors of the sparsity prior in the social sciences. What they show is that sparsity is an assumption about the structure of the world, not a theorem. Statisticians such as Gelman have questioned the “bet on sparsity”: in many problems in social science, they argue, the true effects are dense and small, and sparse methods systematically miss them. The same criticism applies to machine learning. Features in an LLM’s residual stream are stored densely through superposition, and sparsity emerges only once we switch to an overcomplete dictionary. The position of this survey is therefore as follows. Sparsity has been confirmed repeatedly across vast amounts of natural data, so it is a prior worth betting on first. But every time we place the bet, we must answer “sparse under which basis?” and be ready to let go if the evidence does not support it.

The Classical Era (1960s–2014): How Sparsity Became a Computable Tool

The sparse methods that succeeded in the classical era share one trait: their sparsity patterns were structured and predictable, so they could be absorbed into hardware and industry standards. This foreshadows the “failure” described in “The Neural Network Era.” The table below organizes the milestones around three questions: why each appeared when it did, what the bottleneck of the day was, and what it solved.

Figure 2. Compressed sensing in one picture: a reproduction of the phantom experiment of Candès, Romberg & Tao (2006). Only 22 radial lines of the Fourier transform are measured (about 11% of the coefficients). Filling the missing coefficients with zeros gives the streaky minimum-energy image; minimizing total variation subject to the same measurements recovers the phantom exactly. Source: computed by the author for this survey following Candès, Romberg & Tao, “Robust uncertainty principles,” IEEE Trans. Inf. Theory 52(2), 489–509 (2006), Fig. 1; a reproduction, not the original figure.
MilestoneWhy it appeared thenThe bottleneck of the dayWhat it solved
Direct methods for sparse matrices (Tinney & Walker 1967)Power grids were scaling up; memory was measured in KBDense elimination is O(n³) and does not fit in memoryStore only the nonzero entries; reduce fill-in with optimal ordering
Wavelets (Daubechies 1988; Mallat 1989)The Fourier basis is not sparse for edgesEdge energy spreads across the frequency domainPiecewise smooth signals become sparse in the wavelet domain
Matching Pursuit (Mallat & Zhang 1993)The arrival of overcomplete dictionariesRepresentations are not uniqueGreedy atom selection
Wavelet thresholding denoising (Donoho & Johnstone 1994)Coefficients are sparse; noise is uniformLinear filters blur edges tooSoft thresholding = the proximal operator of L1; near minimax optimal
Breiman’s garrote 1995 → LASSO (Tibshirani 1996)High-dimensional data, p ≫ nLeast squares is ill-posed; subset selection is unstableShrinkage and variable selection at once
Basis Pursuit (Chen, Donoho & Saunders 1998)Convex optimization had maturedL0 is NP-hardConvex relaxation via L1
Nonlinear approximation (DeVore 1998)An explanation was needed for why keeping the k largest entries is optimalLinear approximation converges slowlyThe theory of best n-term approximation
JPEG (1992) / JPEG2000 (2000)Digital images went mainstreamBandwidth and storageSparse transform + quantization + entropy coding
LARS (Efron et al. 2004); Elastic Net (Zou & Hastie 2005)LASSO needed an efficient path algorithmSlow to solve; unstable with correlated variablesCompute the entire regularization path at once; L1 + L2
Compressed sensing (Candès–Romberg–Tao 2006; Donoho 2006)Sparse theory met random matrix theorySampling was bound by NyquistRecovery from m ≳ k·log(n/k) measurements under the RIP condition
K-SVD (Aharon, Elad & Bruckstein 2006)Learnable dictionaries were neededFixed wavelets do not fit specific dataAlternate sparse coding with dictionary updates
Donoho–Tanner phase transition (2009)A characterization of when L1 succeeds was neededTheoretical upper bounds were too looseA sharp phase transition in the (m/n, k/m) plane
Sparse coding (Olshausen & Field 1996); fast algorithms (Lee, Battle, Raina & Ng 2007)A response to Barlow’s hypothesisWhy are V1 receptive fields Gabor-like?Sparse dictionaries reproduce V1 receptive fields
The cat-face neuron (Le et al. 2012)The arrival of distributed trainingDo higher-order features emerge without supervision?See below
Optimal singular-value hard thresholding (Gavish & Donoho 2014)Low-rank denoising needed a principled thresholdTruncation rank was chosen by rule of thumbFor square matrices, (4/√3)√n·σ when σ is known; 2.858 × median when unknown

The “bet on sparsity” principle in statistics

Hastie, Tibshirani, and Friedman advocated the “bet on sparsity” principle in The Elements of Statistical Learning. If the true model is sparse, L1 can find it. If it is not sparse, no method works well. Therefore, bet on sparsity. High-dimensional statistics then quantified the exact price. To recover a k-sparse vector in p dimensions, n ≳ k·log p samples suffice (Wainwright 2009). The corresponding result in compressed sensing is m ≳ C·k·log(n/k) random measurements, with exact recovery by L1 whenever the RIP constant satisfies δ₂ₖ < √2−1 ≈ 0.414 (Candès 2008).

Three real-world cases (figures verified)

CS-MRI. Lustig, Donoho, and Pauly (2007) laid the foundations of compressed-sensing MRI. The clinical figures need to be quoted precisely. Sartoretti et al. (2019, PLoS ONE 14(4): e0214887) report that in routine clinical practice the mean scan time fell by 20.2%, examination duration fell by 16%, acquisition time for the applicable sequences fell by 23%–43%, and the number of examinations over the same period rose by 27% (primary source). Vranic et al. (2019, AJNR 40(1):92–98) report reductions of 25% and 35% for brain FLAIR and GRE sequences, respectively. The oft-cited “30%–60% speedup” has no backing in the primary sources.

Black hole imaging (EHT). SMILI, one of the three imaging pipelines behind the first image of M87* in 2019, descends from the sparse modeling (L1 + TV regularization) of Honma et al. (2014) and Akiyama et al. (2017). Another pipeline, CHIRP (Bouman et al. 2016), uses patch priors rather than pure compressed sensing, so the press story that “a single researcher photographed a black hole with compressed sensing” is a simplification.

The cat-face neuron. Le et al. (2012, ICML, arXiv:1112.6209) trained a nine-layer, locally connected sparse autoencoder with about one billion connections on 10 million YouTube frames, “on 1,000 machines (16,000 cores) for three days.” Without supervision, single neurons emerged that responded selectively to human faces and cat faces, and downstream on ImageNet the model reached 15.8% accuracy, a relative improvement of 70% over the previous best. This “sparse coding → unsupervised features” lineage faded after 2012 under pressure from supervised CNNs, then returned in 2023 as Anthropic’s SAEs.

A parallel line: low rank

Truncated SVD (the Eckart–Young–Mirsky theorem) is the best approximation among all matrices of rank at most k, and Gavish & Donoho (2014, IEEE T-IT 60(8):5040–5053) gave the optimal hard threshold for denoising. This is structurally isomorphic to wavelet thresholding denoising. Both follow “change basis → threshold → transform back”; the only difference is that the wavelet basis is fixed while the SVD basis is computed from the data. In the former, the signal stays sparse after thresholding; in the latter, the truncated matrix is low-rank but dense. This lineage returns to large models in “The Present,” in the form of MLA and LoRA.

Step by step: why each advance happened when it did

From Fourier to wavelets (1980s). The Fourier basis decomposes a signal into global sinusoids. That is efficient for stationary signals but extremely wasteful for edges and transients: representing a single step requires infinitely many frequency components, and the energy spreads across the frequency domain. Engineering already had workarounds, such as the Gabor transform and the short-time Fourier transform, but it lacked a unified mathematical framework. Mallat (1989) formalized multiresolution analysis, and Daubechies (1988) constructed orthogonal wavelets with compact support, making it possible to satisfy all three of locality, multiscale structure, and orthogonality at once. Each wavelet basis function is nonzero only over a finite range, and it recurs at every scale. A piecewise smooth signal therefore produces large coefficients only at edge locations and at a few scales, with everything else close to zero. DeVore (1998) and Donoho (1993) proved that this is no empirical accident: for the class of piecewise smooth functions, no fixed linear basis can match the error decay rate of the best n-term nonlinear approximation in a wavelet basis. JPEG2000 (2000) turned this theory into a standard, replacing JPEG’s DCT with wavelets and greatly reducing blocking artifacts at the same bit rate.

Figure 3. An example you can compute by hand (frame from the explainer-video series). For the signal [5,5,5,5,9,9,1,1], taking pairwise averages (blue) and half-differences (yellow) three times leaves only two nonzeros (5 and 4), and the inverse computation recovers all eight original values exactly. With orthonormal scaling (÷√2 instead of ÷2), the two values become 14.14 and 8, and energy is preserved (264 = 264). Wherever neighboring values are equal, the “difference” is zero. This is why piecewise flat signals are sparse in the wavelet domain.

From least squares to LASSO (1990s). Statistics faced a different kind of sparsity: not whether the coefficients are sparse in some basis, but “which explanatory variables actually matter.” The classical approach was subset selection (stepwise regression, AIC/BIC), but it is a combinatorial search and extremely unstable under small perturbations of the data. Breiman (1995) proposed the non-negative garrote, recasting variable selection as a continuous shrinkage problem. Inspired by it, Tibshirani (1996) proposed the LASSO: least squares with an L1 penalty. The geometry of L1 determines its behavior. The constraint region is a diamond with corners, the optimum tends to land on a corner, and each corner corresponds to a point where some coefficients are exactly zero. This is the mechanism by which “shrinkage and selection happen at once.” The LASSO’s algorithmic bottleneck was removed in 2004 by LARS (Efron, Hastie, Johnstone, and Tibshirani), which made it possible to compute the entire regularization path in one pass. In 2005, Zou and Hastie’s Elastic Net handled groups of strongly correlated variables with L1 + L2. High-dimensional statistics then supplied the theoretical guarantees. Wainwright (2009) proved that n ≳ k·log p samples suffice to recover k nonzero coefficients in p dimensions, and Bickel, Ritov, and Tsybakov (2009) established oracle inequalities for the LASSO and the Dantzig selector.

An independent discovery in signal processing (1993–1998). At almost the same time, signal processing started from “overcomplete dictionaries” and arrived at the same L1. Matching Pursuit (Mallat and Zhang 1993) greedily selects the atom most correlated with the residual, and Basis Pursuit (Chen, Donoho, and Saunders 1998) relaxed the NP-hard L0 problem of “finding the sparsest representation” into an L1 linear program. The LASSO and Basis Pursuit are mathematically two parameterizations of the same problem, yet they emerged from two communities that barely interacted. This is the first instance of the idea of sparsity “converging from multiple sources.”

Denoising: the first time sparsity directly created value (1994). Donoho and Johnstone’s wavelet thresholding denoising turned the theory above into a three-step procedure. Move to the wavelet domain with an orthogonal transform. Apply soft thresholding S_λ(w)=sign(w)·max(|w|−λ,0) to the coefficients, with threshold λ=σ√(2 log n). Return to the signal domain with the inverse transform. The key to why this works is that an orthogonal transform preserves the statistical properties of white noise. In the wavelet domain, the noise remains independent Gaussian noise of equal variance, spread evenly over all coefficients, while the signal’s energy concentrates in a few large coefficients, so it suffices to discard the small ones. The soft thresholding operator is exactly the proximal operator of the L1 penalty, which is why denoising, the LASSO, and later ISTA/FISTA (Beck and Teboulle 2009) and LISTA (Gregor and LeCun 2010, the starting point of algorithm unrolling) all share the same core operation.

Compressed sensing: turning sparsity into “measure less, get more” (2004–2009). Every method so far searched for a sparse representation on the assumption that complete data was available. Candès, Romberg, and Tao (2006) and Donoho (2006) asked the reverse question: if a signal is known to be k-sparse in some basis, how many measurements are needed for exact recovery? The answer is m ≳ C·k·log(n/k) random linear measurements, far below the Nyquist rate. The condition is that the measurement matrix satisfy the restricted isometry property (RIP), that is, approximately preserve the length of every 2k-sparse vector. Candès (2008) gave the sufficient condition δ₂ₖ < √2−1 ≈ 0.414, and Donoho and Tanner (2009) used combinatorial geometry to characterize the sharp phase transition between success and failure of L1 recovery. The historical significance of compressed sensing is that it elevated sparsity from “a property of representations” to “a principle of sampling,” directly changing the design of imaging hardware. Lustig, Donoho, and Pauly (2007) brought it to MRI, trading random undersampling trajectories for shorter scan times. Duarte et al. (2008) built the single-pixel camera, and in radio astronomy the SMILI pipeline used L1 and total variation regularization for the EHT’s black hole imaging.

Dictionary learning and sparse coding: from fixed bases to learned bases (1996–2012). Wavelets are a fixed, human-designed basis, and not necessarily optimal for specific data. Olshausen and Field (1996) learned a dictionary from data and reproduced the receptive fields of V1. K-SVD (Aharon, Elad, and Bruckstein 2006) gave a practical algorithm that alternates sparse coding with atom-by-atom SVD updates. Lee, Battle, Raina, and Ng (2007) sped up L1 sparse coding and advanced “self-taught learning.” Le et al. (2012) stacked sparse autoencoders up to nine layers and one billion connections, learning cat-face-selective neurons without supervision from 10 million YouTube frames. After AlexNet in 2012, this lineage was overshadowed for a decade by supervised convolutional networks, but it left two things behind: a representational framework, “overcomplete dictionary + sparse coefficients,” and a conviction, “dictionaries can be learned from data.” When Anthropic decoded the internal features of large models with sparse autoencoders in 2023, it cited precisely Olshausen and Field, and a co-author of the companion paper was Olshausen himself.

Low rank: sparsity’s twin (1936–2014). Eckart and Young (1936) proved that the truncated SVD is the best low-rank approximation of a matrix. Low rank and sparsity are structurally isomorphic: both “concentrate energy in a few components and cut off the tail,” but they act on different objects. Sparsity cuts coordinates; low rank cuts directions. Gavish and Donoho (2014) gave the optimal hard threshold for low-rank denoising, echoing the universal threshold of wavelet denoising. Robust PCA (Candès et al. 2011) combined the two: data = low rank + sparse. This distinction becomes critically important in “The Present.” DeepSeek’s MLA is low-rank compression; DSA and CSA2 are the sparse attention. The two are often confused.

Summary: the three legacies the classical era left to deep learning

First, finding the right basis matters more than processing the data. The same image is dense in the pixel basis and sparse in the wavelet basis; the entire difference lies in how it is described. Second, only three operations create zeros: soft thresholding (L1), hard thresholding (L0), and keeping the k largest entries (projection onto the L0 ball). Third, only structured sparsity makes it into hardware and standards. Wavelet trees made it into JPEG2000, and random undersampling made it into MRI products. All three legacies reappear, each in a new guise, in “The Neural Network Era” and “The Present.”

The Neural Network Era (1990–2022): Sparsity’s Setback and Turning Point

Sparsity did not lose. Sparsity that did not fit the hardware lost. Unstructured weight pruning succeeded in theory and was defeated on the GPU. MoE made the unit of sparsity an entire expert block, and because the inside of each expert remained a dense matrix multiplication, it won the hardware lottery. This section explains why.

Figure 4. The sparsely gated mixture-of-experts layer of 2017, embedded between LSTM layers. For each input, the gating network selects two of n experts; their outputs are weighted by the gate values and summed, and the other experts are skipped entirely. The unit of sparsity is a whole expert block, which is why this design, unlike unstructured pruning, fits the hardware. Source: Shazeer et al., “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” ICLR 2017 (arXiv:1701.06538), Figure 1. © the authors; reproduced for commentary.

The pruning lineage and the “sparsity wall”

OBD (LeCun, Denker & Solla 1990) estimated the importance of each weight from second-order information, and OBS (Hassibi & Stork 1993) refined it with the full Hessian. Deep Compression (Han, Mao & Dally 2016) combined pruning, quantization, and Huffman coding to shrink AlexNet from 240 MB to 6.9 MB. The lottery ticket hypothesis (Frankle & Carbin 2019) showed that a dense network contains sparse subnetworks that, trained in isolation, reach the same accuracy. SparseGPT (Frantar & Alistarh 2023) and Wanda (Sun et al. 2023) showed that GPT-class models can be pruned to 50% sparsity in one shot with almost no loss of accuracy.

Once the LLM era arrived, however, this lineage hit a wall. The ELSA paper of October 2025 (arXiv:2510.01650) calls it the sparsity wall: existing methods “appear unable to exceed 50–60% sparsity without a substantial loss of accuracy.” ELSA claims that with ADMM, LLaMA-2-7B can be pruned to nearly 90%, but that claim is still being debated.

Activation sparsity: the real numbers behind ReLU

Glorot, Bordes, and Bengio (2011, AISTATS) is the primary source on activation sparsity, and its numbers deserve to be quoted precisely. Immediately after uniform initialization, about 50% of hidden units output exact zeros. After training, the average strict sparsity is 83.4% on MNIST, 72.0% on CIFAR10, 68.0% on NISTP, and 73.8% on NORB. The conclusion is that “the models that generalize best have sparsity between 50% and 80%,” and that sparsity can be enforced up to about 85% without hurting performance. The claim that “ReLU is 90% sparse” has no support in the original source.

One distinction is worth making in passing. Dropout is a random regularizer applied during training; at inference the network returns to being dense. Pruning deletes weights permanently. Both can be written as “applying a mask,” but the randomness of the mask, the stage at which it acts, and its purpose are entirely different.

Why unstructured weight sparsity lost in the GPU era

The hardware lottery. By Hooker’s definition (The Hardware Lottery, 2020/2021), a research idea wins because it fits the software and hardware available, not because it is the better idea. GPUs and TPUs are designed for dense, contiguous matrix multiplication. Unstructured sparsity brings in gather/scatter, breaks the tiled dataflow of GEMM, reduces on-chip cache reuse, and usually cannot use the tensor cores. Empirically, libraries such as cuSPARSE show a clear speedup only once sparsity exceeds roughly 90%; at 50% unstructured sparsity, almost no off-the-shelf kernel beats dense GEMM on a GPU.

2:4 structured sparsity. GPUs from Ampere onward support a pattern in which exactly two of every four weights are zero, and when the pattern is met the theoretical throughput doubles. But sparsity is capped at 50%, recovering accuracy requires retraining or fine-tuning, and end-to-end speedups rarely approach 2×, so in LLM deployment it has not spread as widely as quantization.

Why quantization won. Quantization reduces the number of bits per parameter but leaves the regularity of the matrix untouched, so every GPU benefits naturally. After GPTQ (2022) and AWQ (2023), MXFP4 became the native format of 2025–2026. gpt-oss stores its MoE weights (more than 90% of all parameters) in MXFP4 (4.25 bits per parameter), which is what lets the 120b model fit on a single 80 GB GPU. Kimi K3 does quantization-aware training with MXFP4 weights and MXFP8 activations, and the expert parameters of the DeepSeek-V4 series use FP4.

The memory wall. From the V100 to the B200, growth in peak compute far outstripped growth in memory bandwidth (order-of-magnitude figures based on published vendor specifications).

GPUDense BF16/FP16 computeHBM bandwidthOperations per byte (approx.)
V100 (2017)~125 TFLOPS0.9 TB/s140
A100 (2020)~312 TFLOPS1.6–2.0 TB/s155–195
H100 SXM (2022)~989 TFLOPS3.35 TB/s295
B200 (2024)~2.2 PFLOPS (~9 PFLOPS in FP4)~8 TB/s280 (~1,100 in FP4)

During decoding, every generated token requires the weights that take part in the computation to be moved out of GPU memory. At small batch sizes the arithmetic units sit mostly idle and the bottleneck becomes bandwidth. That is why “reducing the parameters and KV that must be read” is worth more than “reducing FLOPs.” MoE reduces the bytes of weights read per token; sparse attention and KV compression reduce the bytes in the cache. They are cutting exactly the most expensive line item.

The lineage of conditional computation (the turning point)

  • 1991 Adaptive Mixtures of Local Experts by Jacobs, Jordan, Nowlan, and Hinton. It trained multiple expert networks together with a gating network, and the original motivation was to reduce interference between tasks. In 1994, Jordan & Jacobs extended it to hierarchical MoE.

  • 2013 Bengio et al. proposed conditional computation and gradient estimation through stochastic neurons (STE). Eigen, Ranzato, and Sutskever made an early attempt at MoE in deep networks.

  • 2017 Outrageously Large Neural Networks by Shazeer et al. (with Hinton and Dean as co-authors). It inserted sparsely gated MoE layers of up to 137B parameters between LSTM layers and achieved “greater than 1000x improvements in model capacity with only minor losses in computational efficiency.” Note that the original says greater than 1000x, and that 137B refers to the parameter count of the MoE layer.

  • 2020–2022 GShard brought MoE into the Transformer and scaled it past 600B, and Switch Transformer (Fedus, Zoph & Shazeer 2021) reached one trillion parameters with top-1 routing. GLaM and ST-MoE tackled stability problems (router z-loss), and Expert Choice (2022) flipped the scheme so that experts choose tokens, which balanced the load naturally.

To put this section in one sentence: from 1991 to 2022 the idea of conditional computation did not change; what changed was that the hardware and the scale to match it finally arrived.

MoE from paper to product: the engineering problems solved in 2017–2022

Shazeer et al. (2017) showed that capacity could be expanded 1000-fold, but turning MoE into a dependable product took another five years and required solving four concrete problems.

How to learn the routing. The gate outputs s = Softmax(W_g·h), the top k experts are selected, and the layer’s output is y = h + Σ_{i∈T} s_i·FFN_i(h). The problem is that top-k is a discrete choice, so its derivative with respect to s is zero almost everywhere. Gradients reach only the gate weights of the selected experts, and experts that are not selected never get to learn. Shazeer’s solution was Noisy Top-k Gating, which adds learnable Gaussian noise to the logits so that the expert that would have ranked k+1 is occasionally chosen. The more general solutions come from the straight-through estimator (STE) of Bengio et al. (2013) and, later, Gumbel-Softmax.

How to balance the load. Routers have a rich-get-richer tendency. The more popular an expert, the more often it is chosen, the better it is trained, and the more likely it is to be chosen again, until a handful of experts carry all the traffic and the model degenerates into a small one (expert collapse). GShard and Switch Transformer introduced an auxiliary load-balancing loss L_bal = α·N·Σ f_i·P_i, where f_i is the fraction of tokens assigned to expert i and P_i is its average gate probability. ST-MoE (2022) went further and used router z-loss to keep the logits from exploding. These auxiliary losses work, but they disturb the gradients of the main language-modeling objective. In 2024, DeepSeek-V3 sidestepped the trade-off with a “bias that stays out of the gradient” (see “The Present”).

How to set capacity. In distributed training, the buffer size of each expert must be fixed in advance, and that is where the capacity factor comes in: capacity per expert = capacity_factor × (number of tokens / N). Tokens beyond capacity are dropped (token dropping) and pass through the residual connection alone. A large capacity factor wastes memory; a small one hurts quality. MegaBlocks (2023) escaped this dilemma by using block-sparse matrix multiplication to build “dropless” MoE.

How to survive the communication. Experts live on separate GPUs (expert parallelism), and each MoE layer requires two data-dependent all-to-all exchanges: tokens are sent to the GPU that holds the expert, and the results are collected back. This is the main communication bottleneck in MoE training, and it is also why MoE suits large, centralized deployments. From DeepSpeed-MoE and Tutel to DeepSeek’s DeepEP, every one of these systems optimizes this path.

Milestones: GShard (2020) brought MoE into the Transformer and scaled it past 600B; Switch Transformer (2021) reached one trillion parameters with top-1 routing and pretrained 4× faster than T5-XXL. GLaM (2021) had 1.2T total parameters and about 97B activated, and consumed roughly one third of the energy of GPT-3. Expert Choice (2022) reversed the direction, letting experts choose tokens, and balanced the load naturally.

Why MoE was still not mainstream before 2022

Even with these engineering solutions in hand, MoE in 2022 was still a “big-company toy” inside Google. There were three reasons. First, the cost pressure was not yet strong enough. In the GPT-3 era, dense models were still advancing rapidly, and inference cost was not the main constraint. Second, there was no open ecosystem. With no high-quality open-weight MoE, the community could neither reproduce nor improve on it. Third, fine-tuning was hard. Routing collapsed easily on small datasets, and adaptation to downstream tasks did not go as smoothly as for dense models.

All three conditions changed at once at the end of 2023. Mixtral 8x7B (2023-12) became the first high-quality open MoE; ChatGPT-scale traffic made inference cost the most pressing problem; and DeepSeekMoE (2024-01), with fine-grained experts and shared experts, pushed the cost efficiency of MoE to a new level. This is where “The Present” begins.

Summary

The three lineages of this section, namely pruning, activation sparsity, and conditional computation, met entirely different fates before 2022. Pruning succeeded in theory and lost on hardware. Activation sparsity (ReLU) was used everywhere and exploited almost nowhere. Conditional computation had everything in place and lacked only the economic pressure. Understanding these different fates explains why the winner after 2023 was MoE rather than weight pruning. Whether sparsity can be put to practical use is decided by whether its zeros have a structure the hardware can exploit.

The Present (2023–September 2026): Why LLMs Went All In on Sparsity

By 2026, the activation ratio per token of the most advanced open models had fallen to 1.5%–4.5%, and sparsity had spread from the FFN to attention, the KV cache, and the storage of knowledge itself. This section first lays out four pressures, then takes stock of each of the five kinds of sparsity in turn, and finally turns to a unifying theory and the economics.

Figure 5. Where sparsity operates in large models (frame from the explainer-video series). (1) Computing knowledge: the FFN is replaced by MoE (conditional computation). (2) Re-reading earlier words: attention computes only a subset of query–key pairs (sparse attention). (3) The notes kept for re-reading: the KV cache is compressed (low rank) or discarded (sparse). The fourth site, the weights themselves, had largely ceded its role to quantization by the LLM era. This section proceeds in the order (1) → (2) → (3).

Why now: four pressures

  • Scaling hit the cost wall. Total parameters reached the 1–3T range (DeepSeek-V4-Pro at 1.6T, Kimi K3 at 2.8T), and training and inference with dense models became unaffordable.

  • Inference costs overtook training costs. In its inference-system overview of 2025-03-01, DeepSeek used large-scale expert parallelism spanning nodes to enlarge batches and raise GEMM efficiency. It shows that the economics of inference had become the single most important constraint on architecture design.

  • Agents made workloads “input-centric.” The V4.1 technical report opens by stating that, with the spread of long-running agents, workloads are increasingly dominated by input, and enormous KV caches keep straining the capacity and transfer bandwidth of HBM and SSDs.

  • Engineering ingenuity under export controls. DeepSeek-V3 was trained on 2,048 H800s, using roughly 2.788M GPU-hours in total. The H800 is the same chip as the H100, but its NVLink bandwidth was cut from 900 GB/s to 400 GB/s, and this directly gave rise to communication optimizations such as DeepEP.

Conditional computation: MoE becomes the default architecture

Figure 6. The intuition behind MoE (frame from the explainer-video series). For each word, a router (the dispatcher) selects a small number of experts (here, 2 of 16), and the experts not selected are skipped as entire blocks. It is the same as a hospital reception desk directing a patient only to the relevant specialists. The model retains all of its knowledge, yet each token uses only a tiny fraction of the parameters. DeepSeek-V3 selects 8 of 256, plus one shared expert.

The evolution of the DeepSeek family (all from official technical reports, model cards, and API announcements):

ModelDateTotal parameters / activated per tokenExpert configurationMain sparsification techniques
DeepSeekMoE2024-01—Fine-grained experts + shared expertsSplits experts finely, increasing the number of combinations exponentially
V22024-05236B / 21B2 shared + 160 routedMLA: low-rank joint compression of KV (low-rank, not sparse)
V32024-12671B / 37B1 shared + 256 routed, 8 selectedAuxiliary-loss-free load balancing, multi-token prediction, FP8 training
V3.22025-09/12671B / 37BSame as aboveDSA: Lightning Indexer + token-level top-k (k=2048) sparse attention
V4-Flash / V4-Pro2026-04-24284B/13B; 1.6T/49B1 shared + 256 routed, 6 selected; 384 routedAlternating Compressed Sparse Attention (CSA) and HCA; 1M context; FP4 experts
V4.1-Flash2026-09-10552B backbone + 196B Engram; prefill 8B / decode 16B1 shared + 384 routed, 6 selectedCED, CSA2, Hierarchical Sparse Indexer, FP4 KV, Engram

The five design elements of V4.1-Flash (technical report):

  • CED (Causal Encoder–Decoder). The 40 layers are split into a 20-layer encoder and a 20-layer decoder, and the decoder’s global KV is projected from the encoder’s final hidden states. Input tokens pass through the encoder only (8B activated), while output tokens pass through the decoder (16B). It is a direct answer to the fact that “agents read a lot and write a little.”

  • The three modes of CSA2. Full generates its own global KV and builds an index. Reindex reuses the previous layer’s KV but recomputes the top-k. Reuse reuses even the top-k indices as they are. The great majority of layers are Reuse, which is to say that the great majority of layers “do not re-select tokens.”

  • Hierarchical Sparse Indexer. The first Full layer of the decoder builds a candidate pool, and subsequent layers select only from within it. The cost of index computation in deep layers is decoupled from the context length.

  • The KV numbers. Global KV is 890 bytes per token, about 1/4 of V4-Flash and 1/437 of V1. Persisted KV is about 1/8 of V4-Flash. Even when the context is widened from 4K to 1M (256×), the decode FLOPs per token grow by only about 1/4. Sparse attention is trained at 64K from the start, with no dense pre-warmup.

  • Engram conditional memory. An N-gram hash table of roughly 196B parameters, accessed sparsely through per-token lookups, which can be placed in host memory and prefetched. It is not activated for every token.

Figure 7. Overall architecture of DeepSeek-V4.1-Flash (Figure 3 of the technical report). On the left is the 20-layer causal encoder: the first two layers use sliding-window attention (SWA), and the rest use CSA2(compression ratio, mode). On the right is the 20-layer decoder, whose global KV is projected from the encoder’s final hidden states. All FFNs are DeepSeekMoE. The yellow Reuse blocks clearly outnumber the green Full blocks: the great majority of layers do not re-select tokens but reuse the previous layer’s indices. The Hierarchical Sparse Indexer first builds a candidate pool, and subsequent layers select only from within it. Engram conditional memory and DSpark speculative decoding are attached at the bottom of the encoder and the top of the decoder, respectively.
Figure 8. Global KV cache per token across DeepSeek generations (log scale). From about 389 KB in V1 to 890 bytes in V4.1-Flash, a 437-fold reduction in three years. This curve is the most direct evidence that sparsity has evolved from “saving compute” to “saving memory and bandwidth.”

Performance and price: DeepSeek states that V4.1-Flash-Base reaches the level of V4-Pro-Base with 1/3 the total parameters and 1/4 the activated parameters. Peak pricing is $0.30/M for input, $0.006/M for cached input, and $1.20/M for output, with off-peak at half price. On a third-party composite index, however (40 points on Artificial Analysis, against 51 for Claude Opus 5), it still trails by about 10 points in independent evaluation. Vendor-selected benchmarks and independent composite evaluations must be written up separately. This is itself a living lesson in how “averages mislead.”

Figure 9. V4.1-Flash compared with Claude Opus 5, GPT-5.6 Sol, and DeepSeek-V4-Pro (Table 3 of the technical report, vendor-reported, maximum reasoning effort). On agentic tasks such as Terminal-Bench 2.1, DeepSWE, HLE w/ tools, and Automation-Bench it is on par or better, but on Terminal-Bench 4.0, which demands long-running execution, it trails Opus 5 by about 20 points. V4.1 activates only 8–16B per token. What this figure means is “approaching the frontier with far fewer activated parameters,” not “surpassing it across the board.”

Other major MoE models of 2025–2026 (from official model cards):

ModelDateTotal / activatedExperts (selected / total)Activation ratioOther design features
Mixtral 8x7B2023-1246.7B / 12.9B2/8About 28%The first high-quality open MoE
gpt-oss-120b2025-08116.8B / 5.1B4/1284.4%MXFP4 MoE weights; alternating sliding-window and full attention; attention sink
Qwen3-235B-A22B2025-04235B / 22B8/1289.4%—
Qwen3.5-397B-A17B2026-02397B / 17B10+1 / 5124.3%Gated DeltaNet linear attention mixed at 3:1, KV reduced to about 1/4
Kimi K2 / K2.52025-07 / 2026-011.04T / 32B8+1 / 3843.1%MuonClip, no loss spikes over 15.5T tokens
Kimi K32026-072.8T / 104B16 / 896, 2 shared3.7%KDA linear attention + Gated MLA; MXFP4 QAT
Llama 4 Maverick2025-04About 400B / 17B1 + 1 shared / 1284.3%Alternating dense and MoE layers
DeepSeek-V4.1-Flash2026-09552B / 8–16B6+1 / 3841.4%–2.9%See above

The trend: in two years the field moved from “2 of 8 experts, 25% activated” to “6–16 experts chosen from several hundred to nearly 1,000, 1.5%–4.5% activated,” with the degree of sparsity rising by roughly an order of magnitude every 1–1.5 years. Shared experts have become standard equipment, and routing now uses auxiliary-loss-free bias, hash routing (the first three layers of V4), and latent-space routing (K3). Among closed models, Google has officially acknowledged that the Gemini 1.5 and Gemini 2.5 series are sparse MoE. GPT-4’s “1.8T MoE” was no more than a 2023 rumor, and the configurations of GPT-5.x and Claude have not been disclosed. At most, write “reportedly.”

Figure 10. Activated parameters per token as a share of total parameters, for flagship models whose configurations are public. From 27.6% for Mixtral (2023-12) to 1.4% for the prefill stage of V4.1-Flash, a roughly 20-fold reduction in two years. Closed models are not included because their configurations are undisclosed.

Sparsifying attention and the KV cache

KV cache size = 2 × L × number of layers × number of KV heads × head dimension × bytes per element. Plug in 64 layers, 8 KV heads, a head dimension of 128, FP16, and a 128K context, and the result is about 31 GiB — larger than the weights of many models, and all of it must be re-read every time a single token is generated. There are three countermeasures: reduce the number of KV heads (GQA/MQA), low-rank compression (MLA), and sparse eviction or sparse computation.

Figure 11. KV cache size for the example model (frame from the explainer-video series). About 256 KB per token: 250 MB at 1,000 tokens, 7.8 GB at 32,000, and about 31 GB at 128,000. That is more than twice the weights of a 7B model, or roughly 10,000 smartphone photos.
Figure 12. Four ways to thin the notebook (frame from the explainer-video series). Several people share one notebook (GQA, 1/8); write a shorter summary (MLA, about a 93% reduction in DeepSeek-V2 — low-rank, not sparse); throw away the pages you don’t need (eviction = sparse); write in smaller letters (4-bit quantization, 1/4). These can be combined.

Sparse attention has three generations, and they can be stacked. The first generation uses fixed patterns (Longformer, BigBird, the sliding window in gpt-oss); the second is dynamic eviction at inference time with no retraining required (StreamingLLM, H2O, SnapKV, Quest); the third learns sparsity natively during training (NSA, MoBA, DSA, CSA2). NSA (Yuan et al., DeepSeek × Peking University × the University of Washington) won the Best Paper Award at ACL 2025 and, at a length of 64K, sped up decoding 11.6×, the forward pass 9.0×, and the backward pass 6.0× relative to full attention.

Figure 13. Overview of Native Sparse Attention (NSA), a third-generation design. For each query, the preceding keys and values pass through three parallel branches: compressed attention for coarse-grained context, selected attention for the most important token blocks, and sliding attention for local context. A learned gate merges the three outputs. On the right, green cells are the attention scores that must be computed and white cells are the ones that can be skipped. Source: Yuan et al., “Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention,” ACL 2025 Best Paper (arXiv:2502.11089), Figure 2. © the authors; reproduced for commentary.
Figure 14. Three ways of skipping, seen as “looking back” between words (frame from the explainer-video series, 32-word example). (1) Only the immediately preceding words (sliding window); (2) only the important words (StreamingLLM/H2O); (3) a lightweight index selects the relevant blocks (NSA/DSA). In long documents, the fraction actually computed is far smaller than in this small example.
Figure 15. Document length versus the number of look-backs (computed from the formula; frame from the explainer-video series). Full attention grows quadratically: at 128K tokens it reaches about 8.2 billion per layer per head. Looking at only the top 2,048 tokens takes about 260 million (roughly 1/32). DeepSeek’s method scores the preceding context with a lightweight “indexer” and looks back at only the top 2,048. It is like answering a question about a 500-page book by checking the table of contents first rather than rereading from the beginning every time.

A cold shower from large-scale measurement. Nawrot et al.’s The Sparse Frontier (arXiv:2504.17768, formally published in Findings of ACL 2026) evaluated training-free sparse attention on Qwen2.5 7B–72B, at 16K–128K, across nine tasks, and reached four conclusions. For very long sequences, “large and sparse” beats “small and dense,” and the crossover sits around 32–64K. The optimal sparsity for prefill is 0.80–0.93 (a budget of 1/5–1/15), and decoding is more tolerant. Almost every configuration suffers a large performance drop on at least one task. Sparse attention is no panacea. And no single strategy is best across all tasks and phases. This result is precisely the motivation that gave rise to the third generation’s “natively learnable sparsity.”

Activation sparsity

Q-Sparse (Wang, Ma, Wang & Wei, arXiv:2407.10969) applies top-k directly to the activations of every linear layer, backpropagates with STE, and demonstrated a scaling law for sparsity. The optimal sparsity ratio is about 45.58% for full-precision models and about 61.25% for 1.58-bit models, and at roughly 40% sparsity the model matches a dense model of the same size. When Mistral 7B was further trained on 40B tokens, Q-Sparse scored 63.7 with 3.8B activated parameters, against 64.6 for the dense baseline (7.0B). In ablations, removing STE or replacing top-k with ReLU each caused a clear drop in performance; moreover, with ReLU the sparsity ratio fell as training progressed, whereas with top-k it held constant. This supports the claim that “sparsity should be an architectural constraint, not a regularization term.” TEAL, ReLU Strikes Back, Deja Vu, and PowerInfer form the main lineage on the edge side.

Sparsity in interpretability: the rise of SAEs and the doubts

Anthropic’s path runs as follows. Towards Monosemanticity (2023) used sparse autoencoders to extract monosemantic features from a one-layer model. Scaling Monosemanticity (2024) trained SAEs with up to 34 million features on Claude 3 Sonnet and demonstrated feature clamping with Golden Gate Claude. Circuit Tracing and On the Biology of a Large Language Model (2025) built attribution graphs with cross-layer transcoders and traced multi-step reasoning in Claude 3.5 Haiku. At OpenAI, Gao et al. (2024) trained a 16-million-feature TopK SAE on GPT-4. This lineage is a direct descendant of Olshausen–Field dictionary learning.

The doubts arrived in a cluster in 2024–2025: feature absorption (Chanin et al. 2024); features that are not one-dimensional and linear (Engels et al. 2024); failure to beat a logistic-regression baseline on 113 probing datasets (Kantamneni et al. 2025); and different SAEs finding different features (Leask et al. 2025). In March 2025, Google DeepMind’s mechanistic interpretability team announced that it would “deprioritize fundamental SAE research for now,” saying the field had overinvested in SAEs. The relatively moderate consensus is that SAEs are well suited to discovering unknown concepts but not to serving as detectors or controllers for known ones. New work such as SharedSAE and WriteSAE continues to appear in 2026, but no new unifying consensus has emerged.

A unifying theory: three operators and learning top-k

Whether in LASSO, pruning, SAEs, or MoE routing, the step that “produces zeros” is one of three operators.

Sλ(y) = sign(y) · max(|y| − λ, 0), Hτ(y) = y · 1[|y| > τ], Pk(y) = keep only the k largest-magnitude entries

Soft thresholding is the proximal operator of L1 (LASSO, ISTA, wavelet denoising, SAEs with an L1 penalty); hard thresholding corresponds to L0 (magnitude pruning, JumpReLU SAEs); and top-k is the projection onto the set of k-sparse vectors (MoE routing, sparse attention, TopK SAEs, Q-Sparse). ReLU itself is a one-sided hard threshold with τ=0. Nearly all modern large models use the third, because it pins the compute per token to an exact constant, which makes static memory planning and dedicated kernels possible. In the era of large models, sparsity has changed from a “regularization term” into an “architectural constraint.”

Figure 16. The three operators that produce zeros, and their uses. Soft thresholding (the proximal operator of L1: LASSO, wavelet denoising, SAEs), hard thresholding (L0: magnitude pruning, JumpReLU), and top-k projection (projection onto the L0 ball: MoE routing, sparse attention, TopK SAEs). Nearly all large models use the third because it fixes the compute per token at exactly the constant k.

Top-k is not differentiable. ∂TopK/∂s is zero almost everywhere, so gradients reach only the gate weights of the selected experts, and the rest can never learn. There are four kinds of engineering fixes. Backpropagate only through the selected experts and add an auxiliary load-balancing loss (the mainstream approach). Add noise with Noisy Top-k to create exploration (Shazeer 2017). Continuous relaxations such as STE, Gumbel-Softmax, and Soft MoE. And, from DeepSeek-V3 onward, the auxiliary-loss-free bias: rank by s+b while weighting by s, and adjust b according to load, outside the gradient. On the sparse-attention side, NSA selects at the block level so that gradients can flow, and DSA first warms up with dense attention, aligning the indexer to the dense attention distribution, before switching to sparse training.

On scaling laws, Clark et al. (2022) unified the scaling of routed models and recommended 64–128 experts. Krajewski et al. (2024) showed that fine-grained experts beat dense models at every compute budget. Frantar et al. (2023) found that the optimal weight sparsity rises with the amount of data. DeepSeek’s Engram paper (arXiv:2601.07372, January 2026) posed the problem of “allocating the sparsity budget” and found a U-shaped law: pure MoE is not optimal, and the best results come when about 20–25% of the sparse parameter budget is allocated to conditional memory. Nine months later, that finding was built into the V4.1 product.

Economics and industry

  • The DeepSeek shock. On January 27, 2025, Nvidia closed down about 17%, erasing roughly $589 billion (Bloomberg) to $593 billion (Reuters) in market capitalization — the largest single-day loss of market value in the history of the US stock market.

  • A 545% theoretical profit margin. On March 1, 2025, DeepSeek disclosed that, pricing an H800 at $2 per hour, its daily cost was $87,072, and that if everything were billed at R1 prices its theoretical daily revenue would be $562,027. It stated explicitly that actual revenue is far lower and that R&D and training costs are not included. Always attach the word “theoretical” when citing this figure.

  • Inference prices. Epoch AI (March 2025) noted that the price of inference at a fixed level of capability is falling by 9× to 900× per year, with the rate varying by task. Yet enterprises’ total bills keep rising, because agents multiply the number of calls per task many times over.

  • The economics of MoE inference. Memory is determined by total parameters, and only with large batches and expert parallelism does each expert receive enough tokens. MoE is inherently suited to large centralized services, not to small-scale self-hosting. vLLM and SGLang supported MLA, DSA, and V4.1 from the day each was released.

A systematic comparison of each lab’s approach

Lining up the labs on the same criteria shows that there is no single path to sparsification; several strategies are competing in parallel.

LabDisclosure of configurationMain kind of sparsityExperts (selected / total)Attention strategyKV strategyWeight formatRepresentative modelsDistinguishing features
DeepSeekPublic, with technical reportsMoE + sparse attention + conditional memory6/384 + 1 sharedNative sparse (DSA → CSA2), trained from the startMLA low-rank + FP4 + cross-layer reuse, 890 B/tokenFP4/FP8 expertsV3, V3.2, V4, V4.1-FlashStacks three axes of sparsity; encoder–decoder division of labor for agents
Moonshot (Kimi)Weights and reports publicMoE + linear-attention hybrid16/896 + 2 sharedBlock-sparse MoBA; K3 uses KDA linear attention + Gated MLANo KV for the linear layersMXFP4 QATK2, K2.5, K3Largest expert count; stable training with MuonClip
Alibaba (Qwen)Model cards publicMoE + linear-attention hybrid10/512 + 1 sharedGated DeltaNet and Gated Attention at 3:1KV cut to about 1/4BF16/FP8Qwen3, Qwen3.5Full series from edge to flagship
MetaModel cards publicMoE1/128 + 1 shared (Maverick)Alternating dense and MoE layers; iRoPEStandard GQABF16Llama 410M context (Scout); Behemoth unreleased
OpenAIOnly gpt-oss publicMoE4/128Alternating sliding window (128) and global attention; attention sinkGQA, 8 headsMXFP4gpt-oss-120b/20bGPT-5.x undisclosed; SAE research (Gao 2024)
GoogleGemini 1.5 and 2.5 officially confirmed as MoE; details undisclosedMoE; elastic structure for the edgeUndisclosedUndisclosedUndisclosedUndisclosedGemini 1.5/2.5, Gemma 3nMatFormer nested submodels, PLE
AnthropicUndisclosedSparsity for interpretability (SAEs, transcoders)UndisclosedUndisclosedUndisclosedUndisclosedClaude seriesUses sparsity as a tool for “understanding the model”
MistralWeights publicMoE2/8StandardGQABF16Mixtral 8x7B/8x22BFirst high-quality open MoE (December 2023)
Microsoft ResearchPapers publicActivation sparsity + 1.58-bit—StandardStandard1.58-bit (BitNet)Q-Sparse, BitNet b1.58Combines sparsity with extreme low-bit quantization
Meituan / Ant / ZhipuModel cards publicMoE variantsVariousVariousVariousVariousLongCat-Flash, Ling, GLM-4.5/4.6Zero-compute experts, shortcut-connected MoE

Three observations follow. First, the main source of innovation in sparse architectures is the Chinese open-model camp. Fine-grained experts, shared experts, auxiliary-loss-free load balancing, native sparse attention, conditional memory, and linear-attention hybrids were almost all first published by DeepSeek, Qwen, and Kimi. Second, the configurations of the big closed labs are black boxes, and can only be inferred from public information such as the gpt-oss and Gemini reports. Citing “GPT-4 is a 1.8T MoE” as fact is a common error. Third, attention is splitting into two paths. DeepSeek chose “sparse attention” (computing only some pairs of tokens); Qwen and Kimi chose “linear-attention hybrids” (replacing the KV cache with a recurrent state). Both attack the same problem — the memory and bandwidth of the KV cache — but with different mathematical tools: the former is sparse, the latter closer to low-rank recurrence.

The systems stack: the engineering that makes sparsity actually run

Whether the gains of a sparse architecture are actually realized is decided in a layer that is easy to overlook: systems software.

  • Expert parallelism and communication. DeepEP (released by DeepSeek in 2025) provides high-throughput, low-latency all-to-all kernels that support FP8 and RDMA over NVSHMEM, and overlaps communication with computation. MegaBlocks (2023) rewrote MoE computation as block-sparse GEMM to achieve dropless MoE, 2.4× faster than dense Megatron-LM.

  • Attention kernels. FlashMLA is optimized for MLA. DSA and CSA2 require matching indexer and sparse-attention kernels, which DeepSeek released in FlashMLA and DeepGEMM. The three modes of CSA2 in V4.1 mean the inference engine must manage cross-layer reuse of KV and indexes.

  • Inference frameworks. vLLM and SGLang supported MLA, DSA, and V4.1-Flash from the day of release, covering H100/H200/B200/GB200 and AMD MI355X. NVIDIA NeMo AutoModel provides a fine-tuning recipe for V4.1, which requires 64 GPUs in a single NVLink domain.

  • Tiered storage. V4.1’s Engram table (about 183 GiB) is designed to sit in host memory and be prefetched over RDMA. Papers already exist that hold Engram in a CXL memory pool, moving “knowledge” from HBM to a cheaper tier. SWA Bounded Replay avoids persisting sliding-window KV to SSD.

  • Quantization formats. MXFP4 (one shared 8-bit scale per 32 elements, 4.25 bits on average) became the de facto standard for MoE expert weights in 2025–2026. Blackwell natively supports FP4 tensor cores.

The conclusion of this section echoes “The Neural Network Era.” Every victory for sparsity has come with kernels and scheduling to match. 2:4 sparsity never spread because, tensor cores notwithstanding, it lacked a mature end-to-end software stack. MoE became mainstream because it had DeepEP, MegaBlocks, and vLLM.

Where Sparsity Meets Deep Learning: Recent Advances in Imaging and Signal Processing

While large models used sparsity to save computation, imaging and signal processing took a different road: combining sparse priors with deep networks to recover signals from incomplete, noisy measurements. The results flowed back the other way, shaping how large models are interpreted and how their architectures are designed.

Figure 17. Deep Image Prior. A randomly initialized, untrained convolutional network is fitted to a single corrupted image and, stopped early, removes JPEG artifacts, fills in missing regions, super-resolves, and denoises. No training data is involved: the structure of the network alone acts as the prior. Source: Ulyanov, Vedaldi & Lempitsky, “Deep Image Prior,” CVPR 2018 (arXiv:1711.10925); teaser figure from the authors’ project repository (Apache 2.0).

Algorithm unrolling: turning iterative sparse solvers into networks

The starting point is LISTA, by Gregor and LeCun (2010). It “unrolls” the fixed iterations of ISTA into a network with a finite number of layers whose parameters can be learned. Each layer is still “a linear transform plus soft thresholding,” but the matrices and thresholds are learned from data. The result converges one to two orders of magnitude faster than hand-tuned ISTA, and it keeps an interpretable structure: each layer corresponds to one step of a proximal gradient method. Algorithm unrolling went on to become the mainstream framework in computational imaging. ADMM-Net and ISTA-Net are used to reconstruct compressed-sensing MRI, and unrolled networks have outperformed classical CS methods on clinically undersampled data. The deeper significance of unrolling is that it showed deep networks can be read as learnable versions of sparse solvers, an insight that later supplied the intuition for decoding neural networks with sparse dictionaries.

Plug-and-play priors and regularization by denoising

Classical regularization (L1, total variation) writes the prior down as an explicit penalty term. The plug-and-play (PnP) framework that emerged from 2013 onward discovered that the “apply the proximal operator” step in ADMM or proximal gradient methods can be replaced by an arbitrary denoiser, which is to say a trained deep denoising network. The prior went from “a single formula” to “a black-box denoiser,” and the sparsity prior became just one special case. Regularization by Denoising (RED) and, later, diffusion-model priors carry on this lineage. In this view, soft-thresholding denoising, BM3D, and deep denoisers are three generations of interchangeable parts that slot into the same position.

Deep Image Prior: the network’s structure is itself the prior

The Deep Image Prior of Ulyanov, Vedaldi, and Lempitsky (2018) delivered a counterintuitive result. Fit a randomly initialized, untrained convolutional network to a single degraded image, and it can denoise, super-resolve, and inpaint. The structure of a convolutional network favors the low-frequency, self-similar statistics of natural images, and it fits noise only slowly, so stopping early yields a clean image. This connects directly to the natural image statistics discussed in “Prehistory.” The network’s structure implicitly encodes the prior that natural images are sparse in some basis; the only difference is that the basis is no longer written out explicitly. Since DIP, “structure is the prior” has become a key lens for understanding the inductive biases of neural networks, and a bridge to classical sparse theory. DIP can be seen as a method that does the same thing adaptively with an implicit, learnable “frame”: it keeps only the components that genuinely represent the structure of the signal.

Learned dictionaries and structured frames

Between fixed wavelets, the unstructured dictionaries learned by K-SVD, and the implicit priors learned by deep networks, there is a middle road: preserve the mathematical structure of tight frames and unitary transforms (perfect reconstruction, energy preservation) while learning their parameters from data. Non-separable oriented symmetric lapped transforms (NSOLT) and lattice-structure unitary networks (LSUN) belong to this category. By constraining the transform to the manifold of tight frames and parameterizing it with a lattice structure, they keep the interpretability and invertibility of classical transforms while gaining adaptability to specific data. The value of this road is being rediscovered in the era of large models. SAEs need overcomplete dictionaries, and MLA needs invertible low-rank projections. Both are searching for a balance between structural constraints and adaptation to data, the very design problem that structured frames have confronted all along.

Three points of contact with large models

First, SAEs are sparse coding itself. Anthropic’s sparse autoencoders are mathematically identical to Olshausen–Field dictionary learning; only the data has changed, from image patches to residual-stream activations. Second, algorithm unrolling anticipated the reading “network = solver.” Anthropic’s transcoders and attribution graphs are, in essence, a search for a sparse, legible alternative computational graph for the Transformer. Third, the idea of structural priors has entered architecture design. MoE expert blocks, the cross-layer reuse in CSA2, and Engram’s lookup tables all write the belief “the data must have this structure” into the shape of the network itself rather than into the loss function. This is the imaging-world counterpart to the conclusion of the section “A unifying theory”: sparsity has shifted from a regularization term to an architectural constraint.

Applications: Where Sparsity Has Truly Created Value

Sparsity creates value where there is a great deal to process but only a little that truly matters. Reviewing long documents is one of the most typical examples.

Figure 18. MMLongBench-Doc, the benchmark that sets the bar for long-document understanding. (a) Long PDFs (47.5 pages on average) with single-page questions, cross-page questions, and questions designed to be unanswerable. (b) Its documents are far longer than those of earlier document-QA benchmarks. (c) Even frontier vision-language models score below 45 F1, and several do worse than feeding OCR text to a text-only model. Source: Ma et al., “MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations,” NeurIPS 2024 Datasets & Benchmarks (arXiv:2407.01523), Figure 1, from the authors’ repository. © the authors; reproduced for commentary.
  • LLM inference in the cloud. Nearly all of the low-priced frontier APIs of 2026 come from highly sparse MoE models. DeepSeek-V4.1-Flash’s off-peak output costs $0.60/M, and gpt-oss-120b runs about $0.03/M for input and about $0.17/M for output on OpenRouter.

  • Long context and agents. In the V4.1 announcement, DeepSeek stated plainly that cache-hit charges often account for the bulk of an agent’s cost, and that compressing the cache can cut this part substantially. The price of cached input ($0.006/M) is 1/50 of the cache-miss price ($0.30/M), and this is the central lever in the economics of agents.

  • Edge. Apple’s LLM in a Flash (2023) loads sparsely activated parameters from flash storage on demand. Gemma 3n uses MatFormer’s nested submodels and per-layer embeddings (PLE). MoE models with 3B activated parameters, such as Qwen3.5-35B-A3B, have become candidates for the browser and the edge.

  • Multimodal. V4.1 gives image tokens a dedicated routing bias. Visual tokens are numerous and highly redundant, which makes them an ideal target for sparse attention.

  • Science and medicine. CS-MRI has shortened average scan times in routine clinical practice by about 20%, and by 23%–43% for a single sequence. The SMILI pipeline behind the EHT’s black-hole imaging is built on sparse modeling.

Enterprise review of long documents and drawings (directly relevant to our business)

The benchmark for difficulty. MMLongBench-Doc (NeurIPS 2024 D&B) collects 135 PDFs averaging 47.5 pages and roughly 21,000 tokens of text, with 1,082 expert-written questions. Of these, 33.2% require cross-page evidence, and 22.8% are designed to be unanswerable in order to detect hallucination. Human annotators score an F1 of 66.0%, and 12 of the 14 vision-language models perform worse than when OCR text is fed to the corresponding text-only model. LongDocURL’s 396 documents average 86 pages, and 52.9% of its questions are cross-page. The length of these documents and their share of questions that span pages and elements closely resemble building-permit applications (drawings, structural calculations, and application forms running from several dozen to well over 100 pages).

The bottleneck is not context length alone. It is choosing the right top-K among scattered evidence. Retrieval (selecting the top-K pages) and sparse attention (selecting the top-K token blocks) are mathematically the same operator; the difference is whether the selection happens outside the model (auditable, able to cite provisions) or inside it (end-to-end and opaque).

Limits. Sparse Frontier shows that sparse attention can degrade performance on individual tasks even at moderate sparsity. In review work, “clause X contradicts the dimension noted on page 37” is exactly the kind of task where you “miss one and the whole answer is wrong.” A hybrid of retrieval plus re-verification in a long context is therefore safer than betting on either one alone. The 2025 research on multi-page documents, including SimpleDoc, MDocAgent, and LAD-RAG, takes this road as well.

Benefits and Costs

For weight sparsity, “sparsity is overrated” is broadly true in the LLM era. For conditional computation and attention sparsity, the opposite holds: both are already the default at the frontier. Keeping these two apart is the key to understanding the whole debate.

Figure 19. The primary source of the “545%” number: DeepSeek’s own 24-hour cost and theoretical income for serving V3 and R1 (March 1, 2025). Daily cost was about $87,072 (H800 rented at $2 per GPU-hour), against a theoretical income of $562,027 if every token had been billed at R1 API prices. DeepSeek itself notes that actual income is far lower, because web and app usage is free and off-peak API prices are discounted. Source: DeepSeek, “DeepSeek-V3/R1 Inference System Overview,” open-infra-index repository, Open Source Week day 6 (CC0 1.0).

Benefits

  • Cost: activation ratios have fallen to 1.5%–4.5%, and the price curve is dropping fast.

  • Decoupling capacity from compute: total parameters set knowledge capacity, activated parameters set the cost per token, and the two can be scaled independently. Engram goes further, moving static knowledge into cheap host memory.

  • Long context becomes practical: V4.1’s decode FLOPs rise by only about a quarter between 4K and 1M.

  • A tool for interpretability: SAEs and transcoders are used to discover unknown concepts.

Costs and risks

  • Memory is set by total parameters: V4.1’s routed experts (MXFP4) take about 259.5 GiB and the Engram table about 183 GiB; sparsity saves no storage.

  • Communication overhead: the all-to-all of expert parallelism is the main bottleneck in MoE training, and Engram needs prefetching over RDMA or CXL.

  • Training instability and load imbalance: a whole series of patches becomes necessary, among them z-loss, bias adjustment, and hash routing.

  • Fine-tuning is hard: routing collapses easily on small datasets. V4.1’s official fine-tuning procedure requires 64 GPUs in a single NVLink domain.

  • Evaluation pitfalls: Sparse Frontier’s finding that “almost every configuration loses performance on at least one task.” V4.1 matches the frontier on the vendor’s benchmarks but trails by about 10 points on an independent composite index.

  • The limits of SAEs: see the section “Sparsity in interpretability.”

  • Weight sparsity still loses on efficiency: conventional methods struggle to break through the 50%–60% sparsity wall.

Representative views (quoted from the sources)

PositionSourceView
Supportive: hardware decides who winsHooker 2021Ideas win “because they are a good fit for the software and hardware available, not because the idea is superior”
Supportive: a new axis of sparsityDeepSeek, Engram 2026“We believe conditional memory will become an indispensable modeling primitive for the next generation of sparse models”
CautiousNawrot et al. 2026“Sparse attention is not a silver bullet”
NegativeGoogle DeepMind 2025“We don’t expect SAEs to be a game-changer for interpretability, and we suspect the field has over-invested in them”
Middle groundarXiv:2506.23845SAEs are poor at acting on known concepts, but they are a powerful tool for discovering unknown ones

A verdict on each of the five kinds of sparsity (September 2026)

There is no single answer to “is sparsity overrated?”; the question has to be answered kind by kind.

Kind of sparsityStatus in 2026VerdictMain evidence
Weight sparsity (unstructured pruning)Active research, very little industrial deploymentOverrated in the LLM era; quantization wonThe 50%–60% sparsity wall; no efficient GPU kernels; MXFP4 is the de facto standard
Structured sparsity (2:4, block)Hardware support exists, but end-to-end speedups are limitedPartly realized; block sparsity won in the form of MoE2:4 rarely approaches 2× in measured results; MegaBlocks’ block-sparse GEMM became the foundation of MoE
Activation sparsity (ReLU, top-k activation)Attention at the edge and in research; absorbed into MoE in the cloudPotential underrated, but the main battleground is the edgeQ-Sparse matches dense at about 40% sparsity; TEAL and PowerInfer target memory-constrained devices
Conditional computation (MoE)The default architecture at the frontierLives up to its reputation, and still deepeningActivation ratio fell from 25% to 1.5%–4.5% in two years; adopted by every flagship-class open model of 2025–2026
Attention sparsityNatively trainable versions are in productsEstablished, but with a risk of per-task performance dropsNSA won the ACL Best Paper Award; V4.1 trains it from the start; Sparse Frontier’s “drop on at least one task”
(Parallel line) Low rankMLA and LoRA are widely usedEstablished, but should not be confused with sparsityLow-rank KV compression of up to about 93% (V2); LoRA is the standard for fine-tuning
(Tool) Sparse dictionaries / SAEsUnder debateEstablished as a discovery tool; overrated as a detector or controllerGDM deprioritized it; trails linear probes on 113 datasets; Anthropic moved to transcoders and attribution graphs

In one sentence: sparsity won on “activation” and “attention” and lost on “weights.” It won not because the theory was superior, but because the positions of the zeros had a structure the hardware could exploit.

The Future (2026–2030)

The next axis of sparsity is no longer “which parameters to compute” but “which knowledge to store, and which history to read.” Below are the research frontier, the hardware trends, and five falsifiable predictions.

Figure 20. Sparsity allocation and Engram scaling. Left: validation loss as a function of the share of the sparse parameter budget allocated to MoE experts rather than to Engram conditional memory, at two compute budgets (2e20 and 6e20 FLOPs). Both curves are U-shaped, and the hybrid beats pure MoE (the rightmost points). Right: in the infinite-memory regime, loss falls log-linearly with the number of embedding slots. Source: DeepSeek-AI (Cheng et al.), “Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models,” arXiv:2601.07372 (2026), Figure 3, from the official repository (Apache 2.0).

The research frontier

  • Multi-axis sparsity and budget allocation. The U-shaped law found in the Engram paper shows that pure MoE is not optimal and that the sparsity budget should be split between “experts” and “conditional memory.” The 27B Engram model achieved MMLU +3.4, BBH +5.0, and HumanEval +3.0 over an MoE baseline with the same parameter count and the same FLOPs. Mechanistic analysis shows that conditional memory frees the early layers from having to reassemble static knowledge.

  • Natively trainable sparsity becomes the default. V4.1 trains sparse attention from scratch at 64K, with no dense warm-up beforehand. It shows that sparse attention has gone from an inference-time approximation to a first-class citizen at training time.

  • Sparsity × inference-time compute. Reasoning models produce long outputs, and both prefill and decode are heavy. MoE provides knowledge capacity; long chains of thought provide reasoning depth. CED’s asymmetric activation is a trade-off aimed at workloads with different read/write ratios.

  • Modularity and continual learning. Experts and memory tables can be edited locally. For example, there is research that writes a user’s memory as a local parameter edit (User as Engram, arXiv:2606.19172).

  • Combining sparsity with quantization. MXFP4 experts, FP4 KV, and FP8 Engram are already used together in V4.1 and K3. Q-Sparse’s conclusion is that the more aggressive the quantization, the higher the optimal degree of sparsity.

  • Co-design with hardware. Blackwell supports FP4 natively, and tiered KV storage using CXL memory pools and SSDs is advancing. Neuromorphic computing at LLM scale remains a distant prospect.

Five falsifiable predictions (with rationale)

PredictionDeadlineRationale
Global KV falls to 500 bytes per token or less in at least one major open modelEnd of 2027A 437× reduction from V1 to V4.1, a 4× reduction in the single generation from V4 to V4.1, and FP4 already in use
At least three of the top five open models have an activation ratio per token below 3%End of 2027In 2026 there are already V4.1 (2.9%), K3 (3.7%), and Qwen3.5 (4.3%), and expert counts double every generation
At least two major labs other than DeepSeek introduce lookup-table-style conditional memory into their flagship open models2028Engram already has CXL, SSD, and training-free derivative work, and it has landed in NeMo
SAEs do not become an officially announced primary production safety-monitoring method at any frontier research lab2028GDM’s deprioritization; negative results on downstream tasks
The API unit price for a fixed level of capability keeps falling by 10× or more per year, but enterprises’ agent spend per task risesOngoingEpoch’s price trends; the doubling of agent call counts

What This Means for Us

Review is, in essence, the task of selecting the top-K out of a vast number of provisions and pages, and what the sparsification trend is driving down is precisely the unit cost of this kind of task. For a company building a review agent for long documents and drawings, there are four actionable implications.

Figure 21. MDocAgent, a multi-agent system for question answering over long multimodal documents. Text-based and image-based retrieval each select the top-k segments; a general agent and a critical agent identify the critical textual and visual evidence; specialized text and image agents answer; and a summarizing agent reconciles the answers. It is a concrete instance of the retrieve-then-verify hybrid this survey recommends for review work. Source: Han et al., “MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding,” arXiv:2503.13964 (2025), Figure 1, from the authors’ repository (MIT).
  • The cost curve. Treat “cache hit rate” and “share of off-peak batch processing” as first-class metrics. Review tasks are inherently batchable and can run offline. Place stable laws and ordinances at the front of the prompt to maximize prefix-cache hits. V4.1’s cached-input price is 1/50 of the cache-miss price, and off-peak halves it again.

  • Model selection. Choose models on our own evaluation set for “cross-page matching and clause citation,” and judge by the score on the hardest subtask, not the average. Use a highly sparse open MoE as the main workhorse model, and a closed frontier model to double-check the difficult cases.

  • Long-context strategy. Adopt a hybrid: “retrieval of the top-K pages (auditable, with citations) + re-checking in a 1M context (to avoid missing cross-page relationships).” Sparse Frontier’s conclusion applies here directly. Betting on sparse attention alone drops performance on some tasks, and review is exactly a “miss one and the whole answer is wrong” task.

  • The self-hosting decision. MoE memory is set by total parameters; V4.1’s routed experts plus the Engram table exceed 440 GiB, and the official fine-tuning procedure requires 64 GPUs in a single NVLink domain. At a scale of fewer than a few dozen GPUs it almost never pays off, so prefer the API.

A research opportunity. Existing work treats “which pages to check within a budget” as a retrieval problem. But the truly sparse object in cross-page consistency checking is not the “page” but the “page pair.” P pages have P(P−1)/2 pairs, and from these we select k pairs to cross-check. This maps directly onto the sparse coordination graphs of multi-agent systems, and it is isomorphic to the idea behind DSA, which scores every (query, key) pair with a lightweight indexer and takes the top-k. It is a direction where we can differentiate and cut in.

Appendices

A. Timeline (1948–September 2026)

YearEventReliability
1948Shannon, A Mathematical Theory of Communication: redundancy and entropy formalizedprimary paper
1961Barlow’s redundancy reduction hypothesisprimary paper
1967Tinney & Walker’s sparse-matrix ordering (power-grid computation)primary paper
1978Rissanen’s minimum description length (MDL)primary paper
1987Field: natural image statistics and cortical receptive fieldsprimary paper
1988–89Daubechies’s compactly supported orthogonal wavelets; Mallat’s multiresolution analysisprimary paper
1990–93Pruning via OBD / OBS; Matching Pursuitprimary paper
1991Jacobs–Jordan–Nowlan–Hinton’s mixtures of local expertsprimary paper
1992 / 2000JPEG / JPEG2000standard
1994–96Wavelet thresholding denoising; Breiman’s garrote; LASSO; Olshausen & Field’s sparse codingprimary paper
1998Basis Pursuit; DeVore’s nonlinear approximationprimary paper
2001–03Attwell & Laughlin’s energy budget; Lennie’s “possibly fewer than 1%”primary paper (verified)
2004–05LARS; Elastic Netprimary paper
2006–07Compressed sensing; K-SVD; Lustig’s CS-MRIprimary paper
2009Donoho–Tanner phase transitionprimary paper
2011Glorot’s ReLU (50–85% exact zeros); Robust PCAprimary paper (verified)
2012Google’s cat-face neuronprimary paper (verified)
2013–14Bengio’s conditional computation and the STE; Gavish–Donoho optimal hard thresholdprimary paper
2016Deep Compressionprimary paper
2017Shazeer’s sparsely-gated MoE (a 137B MoE layer)primary paper (verified)
2019The lottery ticket hypothesis; EHT M87*; Sartoretti’s clinical study of CS-MRIprimary paper
2020GShard; Hooker’s Hardware Lottery; Ampere’s 2:4primary paper
2021–22Switch Transformer; GLaM; ST-MoE; Expert Choice; Clark’s MoE scalingprimary paper
2023Mixtral; SparseGPT / Wanda; StreamingLLM / H2O; Towards Monosemanticityprimary paper
Jan–Feb 2024DeepSeekMoE; Gemini 1.5 officially confirmed as MoEtechnical report
May 2024DeepSeek-V2 (MLA); Scaling Monosemanticity; Golden Gate Claudetechnical report
Jun–Jul 2024Gao’s TopK SAE; Q-Sparseprimary paper (verified)
Dec 2024DeepSeek-V3 (671B/37B)technical report
Jan 27, 2025The DeepSeek shock: Nvidia’s market capitalization falls by roughly $589 billion in a single dayreliable press reports
Feb–Mar 2025NSA, MoBA; a theoretical profit margin of 545% for the inference system; GDM lowers the priority of SAEsprimary paper / official
Apr 2025Qwen3 MoE; the Sparse Frontier preprint; Llama 4official / primary paper
Jul–Aug 2025NSA wins the ACL Best Paper Award; Kimi K2; gpt-ossofficial
Sep–Dec 2025DeepSeek-V3.2, DSAtechnical report
Jan 2026Engram: conditional memory and a U-shaped sparsity budget allocationprimary paper
Feb 2026Qwen3.5-397B-A17Bofficial model card
Apr 24, 2026DeepSeek-V4 (Pro 1.6T/49B, Flash 284B/13B)technical report
Jul 2026Kimi K3 (2.8T/104B); Sparse Frontier published in Findings of ACL 2026official
Sep 10, 2026DeepSeek-V4.1-Flash (CED, CSA2, 890 B/token, Engram)official + technical report (verified)

B. Data series that can be charted

  • DeepSeek’s global KV cache (per token): V1 ≈ 389 KB (estimated) → V3.2 ≈ 48 KB → V4-Flash ≈ 3.5 KB → V4.1-Flash 890 B. Log scale.

  • Activation ratio per token: Mixtral 28% → Qwen3 9.4% → gpt-oss 4.4% → Qwen3.5 4.3% → Kimi K3 3.7% → V4.1 decode 2.9% / prefill 1.4%.

  • Number of experts (selected / total): 2/8 → 4/128 → 8/256 → 6/384 → 10/512 → 16/896.

  • DeepSeek API prices ($/M, input / output): V4-Flash 0.14 / 0.28; V4-Pro 1.74 / 3.48 (peak-hour output 3.96 from August onward); V4.1-Flash peak 0.30 / 1.20, off-peak 0.15 / 0.60, cache 0.006.

  • Sparse Frontier: optimal sparsity for prefill 0.80–0.93; decoding remains practical even at 0.95; the crossover point is 32–64K.

  • Sparsity rates in Glorot 2011: MNIST 83.4%, CIFAR10 72.0%, NISTP 68.0%, NORB 73.8%.

  • Q-Sparse: optimal sparsity rate 45.58% (full precision) / 61.25% (1.58-bit); on par with dense at around 40%.

  • Engram-27B’s gap over the MoE baseline: MMLU +3.4, CMMLU +4.0, BBH +5.0, ARC-C +3.7, HumanEval +3.0, MATH +2.4.

  • CS-MRI: scan time −20.2%, examination duration −16%, single sequences −23% to −43%, number of examinations +27%.

  • GPU compute and bandwidth: see the GPU table in “The Neural Network Era.”

C. Anecdotes and story material

  • How “Lennie’s 1%” became “1–4%”: usable as an opening that corrects a misconception.

  • The cat-face neuron (2012): 1,000 machines, over three days, taught themselves “cat” from YouTube frames.

  • Outrageously (2017): a 137B MoE layer in the LSTM era foreshadowed the landscape of 2026. The co-authorship of Hinton and Dean is emblematic of “an old idea from 1991 × Google’s compute.”

  • The DeepSeek shock (January 27, 2025) and the word “theoretical” attached to 545%.

  • Golden Gate Claude (2024) and GDM’s hard brake (March 2025).

  • CS-MRI: patients spent 20% less time lying on the scanner table, and hospitals could perform 27% more examinations.

  • 437×: the “memory” per token shrank from roughly 389 KB to 890 bytes.

  • The return of N-grams: a statistical language model once consigned to history came back to the frontier as Engram.

D. Common misconceptions and their corrections

Common phrasingAccurate statement
The brain uses only 1–4% of its neuronsLennie 2003 actually says “possibly fewer than 1%,” an upper-bound estimate based on energy. The 1–4% figure is a secondary citation via Glorot 2011
ReLU networks are 90% sparseGlorot 2011: about 50% at initialization, 50%–85% after training
GPT-4 is a 1.8T MoE with 16 expertsA 2023 rumor that OpenAI has never confirmed. The configurations of GPT-5.x and Claude are likewise undisclosed
CS-MRI makes scans 30–60% fasterSartoretti 2019: mean scan time −20.2%, single sequences −23% to −43%
Shazeer 2017: “up to 1000×”The original says “greater than 1000x.” 137B is the parameter count of the MoE layer
Sparse Frontier was presented at ICML 2025Formally published in Findings of ACL 2026
DeepSeek’s profit margin is 545%A self-reported theoretical cost-profit margin. The real figure is far lower and excludes R&D and training costs
MLA is sparse attentionMLA is low-rank compression. DSA / CSA2 are what is actually sparse
Sparse attention can be compressed without limitNearly every configuration loses significant performance on at least one task

E. References

Classical foundations: Shannon 1948; Barlow 1961; Tinney & Walker 1967; Rissanen 1978; Field 1987; Mallat 1989; LeCun, Denker & Solla 1990; Jacobs, Jordan, Nowlan & Hinton 1991; Mallat & Zhang 1993; Donoho & Johnstone 1994 (Biometrika); Breiman 1995; Tibshirani 1996 (JRSS-B); Olshausen & Field 1996 (Nature); Chen, Donoho & Saunders 1998; DeVore 1998; Attwell & Laughlin 2001; Lennie 2003 (Curr. Biol.); Efron et al. 2004; Zou & Hastie 2005; Candès, Romberg & Tao 2006; Donoho 2006; Aharon, Elad & Bruckstein 2006; Lustig, Donoho & Pauly 2007; Lee, Battle, Raina & Ng 2007; Donoho & Tanner 2009; Glorot, Bordes & Bengio 2011; Candès, Li, Ma & Wright 2011; Le et al. 2012; Bengio et al. 2013; Gavish & Donoho 2014; Han, Mao & Dally 2016; Shazeer et al. 2017; Frankle & Carbin 2019; Sartoretti et al. 2019; Vranic et al. 2019; Hooker 2021 (CACM); Fedus, Zoph & Shazeer 2021; Clark et al. 2022.

2023–2026: Frantar & Alistarh 2023 (SparseGPT); Sun et al. 2023 (Wanda); Xiao et al. 2023 (StreamingLLM); Bricken et al. 2023; Jiang et al. 2024 (Mixtral); Dai et al. 2024 (DeepSeekMoE); DeepSeek-AI 2024 (V2, V3); Templeton et al. 2024; Gao et al. 2024; Wang et al. 2024 (Q-Sparse); Krajewski et al. 2024; Yuan et al. 2025 (NSA); Nawrot et al. 2025/2026 (Sparse Frontier); Kantamneni et al. 2025; OpenAI 2025 (gpt-oss); Google DeepMind 2025 (Gemini 2.5 technical report, arXiv:2507.06261); Moonshot 2025/2026 (Kimi K2, K3); Alibaba 2026 (Qwen3.5); Cheng et al. 2026 (Engram, arXiv:2601.07372); DeepSeek-AI 2026 (V4, arXiv:2606.19348); DeepSeek-AI 2026 (V4.1-Flash, arXiv:2609.19969); Ma et al. 2024 (MMLongBench-Doc).

Textbooks and surveys: Elad, Sparse and Redundant Representations (2010); Hastie, Tibshirani & Wainwright, Statistical Learning with Sparsity (2015); Mallat, A Wavelet Tour of Signal Processing (3rd ed.); Foucart & Rauhut, A Mathematical Introduction to Compressive Sensing (2013); Hoefler et al. 2021, Sparsity in Deep Learning; Cai et al. 2024, A Survey on Mixture of Experts.

이 글을 공유하기