DeepSeek-V4-Pro-Max is the maximum reasoning effort mode of DeepSeek-V4-Pro, announced with the DeepSeek-V4 preview series. DeepSeek-V4 models expose configurable reasoning effort; at the Max setting the model spends the largest thinking budget, which significantly advances the knowledge capabilities of open-source models - top-tier performance in coding benchmarks and a significantly narrowed gap with leading closed-source models on reasoning and agentic tasks, establishing it (per DeepSeek) as the best open-source model available at release. The same mechanism gives DeepSeek-V4-Flash-Max comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller scale still trails on pure knowledge tasks and the most complex agentic workflows. Related: gpt-oss models expose a similar configurable-effort control (low/medium/high).
Related FAQs (1)
DeepSeek-V4-Pro-Max = DeepSeek-V4-Pro at maximum reasoning effort:
V4 models have a configurable reasoning-effort control; Max = largest thinking budget.
Results: top-tier open-source coding performance; significantly narrowed gap to closed-source leaders on reasoning and agentic tasks.
V4-Flash-Max at a large thinking budget matches Pro on reasoning, but trails on pure knowledge and the most complex agentic workflows.
Heavily Compressed Attention (HCA)
architecture_concepts
Heavily Compressed Attention (HCA) is one half of DeepSeek-V4's hybrid long-context attention, paired with Compressed Sparse Attention (CSA). While CSA sparsifies which positions a query attends to, HCA attacks the size of what each attended position contributes: key-value representations are compressed much more aggressively than in DeepSeek-V3.2's MLA-style latent cache, shrinking per-token KV storage for the layers where full positional fidelity matters less. Interleaving CSA and HCA layers is what lets DeepSeek-V4-Pro run 1M-token contexts with only 27% of single-token inference FLOPs and 10% of the KV cache that DeepSeek-V3.2 would need - sparsity cuts the FLOPs, heavy compression cuts the cache.
Related FAQs (1)
HCA = the KV-compression half of DeepSeek-V4's hybrid attention:
CSA (its pair) sparsifies which positions get attended.
Together at 1M-token context: DeepSeek-V4-Pro needs only 27% of inference FLOPs and 10% of KV cache vs DeepSeek-V3.2.
Gated Multi-head Latent Attention (Gated MLA)
architecture_concepts
Gated Multi-head Latent Attention (Gated MLA) is Kimi K3's precision-attention layer type: 24 of its 93 attention layers are Gated MLA, complementing the 69 Kimi Delta Attention (KDA) layers that provide cheap O(n) sequence processing. It extends Multi-head Latent Attention (MLA - key/values compressed into a small latent vector so KV-cache stays tiny) with a gating mechanism on the attention output, letting the model dynamically amplify or suppress each head's contribution. In the K3 hybrid layout, Gated MLA layers supply the precise global retrieval and exact positional binding that linear-attention state compression can lose at very long range, while KDA keeps per-token cost low across the model's 1M-token context.
Related FAQs (1)
Gated MLA = Kimi K3's gated variant of Multi-head Latent Attention:
Like MLA, it compresses KV into a latent vector (small KV cache).
Adds a gate on the attention output so each head's contribution can be amplified or suppressed.
Used in 24 of K3's 93 attention layers; the other 69 are KDA linear-attention layers - together they balance cheap long-context processing with precise global retrieval.
Attention Residuals (AttnRes)
architecture_concepts
Attention Residuals (AttnRes) is one of the two architectural components Moonshot AI built Kimi K3 on (with Kimi Delta Attention). Standard transformers add each layer's block output through a residual stream that bypasses attention; AttnRes instead feeds information from the attention pathway itself back through the residual stream - the residual connection carries attention's contribution rather than only the MLP/block output. Combined with delta-rule linear attention (KDA) and a Stable LatentMoE backbone (16 of 896 experts active), this is part of what yields Kimi K3's ~2.5x improvement in overall scaling efficiency over Kimi K2.
Related FAQs (1)
AttnRes = residual connections through the attention pathway (Kimi K3):
In standard transformers the residual stream bypasses attention; AttnRes routes attention's own output into the residual stream as well.
Lets attention contributions persist across layers instead of being overwritten by MLP paths.
Paired with KDA (delta-rule linear attention) and Stable LatentMoE in Kimi K3 (~2.5x scaling efficiency over K2).
Kimi Delta Attention (KDA)
architecture_concepts
Kimi Delta Attention (KDA) is the linear-attention mechanism Moonshot AI built Kimi K3 on (69 of its 93 attention layers). It belongs to the delta-rule family of linear attention: rather than storing every key-value pair like softmax attention, KDA maintains a compressed state that is updated with the delta rule - new information overwrites the parts of the state it can already predict, so the state carries only what is genuinely new. This gives O(1) per-token state size and O(n) sequence cost instead of quadratic attention, making million-token contexts tractable, while the remaining 24 Gated MLA layers retain precise global retrieval. Kimi K3 pairs KDA with Attention Residuals (AttnRes) and a Stable LatentMoE backbone (16 of 896 experts active), reaching ~2.5x better scaling efficiency than Kimi K2.
Related FAQs (1)
KDA = Moonshot's delta-rule linear attention (Kimi K3's main attention layer, 69/93 layers):
Maintains a compressed state updated by the delta rule: new tokens overwrite what the state can already predict - it stores only genuinely new information.
O(n) sequence cost, constant state - quadratic softmax attention can't do million-token contexts cheaply.
The other 24 layers are Gated MLA for precise global retrieval; together with AttnRes and Stable LatentMoE this gave Kimi K3 ~2.5x scaling efficiency over K2.
iHC (identity Hyper-Connections)
architecture_concepts
iHC (identity Hyper-Connections) is the residual-stream design of Tencent's Hy4 preview: a simplified, identity form of Hyper-Connections (arXiv 2409.19606). Instead of a single residual pathway, the network keeps 4 parallel residual streams and each layer's output is added into multiple streams, expanding inter-layer information flow so layers can interact through more than one path - without the learned mixing weights of full Hyper-Connections. It is the identity-matrix member of the same family as Manifold-Constrained Hyper-Connections (mHC) (DeepSeek-V4, GLM-5.3-Flash) and DeepSeek-V4's learned Hyper-Connections: where mHC constrains the mixing manifold and DeepSeek-V4 learns the mix, iHC fixes the mixing to identity, trading a little adaptivity for stability and simplicity.
The network keeps 4 parallel residual streams; layer outputs fan out into several streams.
Full Hyper-Connections learn mixing weights between streams; iHC fixes them to identity - simpler, stabler.
Same family as mHC (DeepSeek-V4, GLM-5.3-Flash) and DeepSeek-V4's learned HC, but with identity mixing.
IndexCache (cross-layer sparse index reuse)
architecture_concepts
IndexCache (arXiv 2603.12201) is a long-context efficiency technique used in Tencent's Hy4 preview: with sparse attention, each layer normally re-runs a lightweight indexer to find which key-value positions each query should attend to. IndexCache instead computes the sparse index once and reuses it across layers, eliminating redundant index computation for the same positions. Paired with Gated DSA, it makes 1M-token inference cheap: the indexer (32 heads, 128-dim, top-k 2048 in Hy4) does its work once and the cross-layer cache feeds every subsequent sparse-attention layer.
Related FAQs (1)
IndexCache = reusing a sparse attention index across layers:
A sparse indexer picks which KV positions each query attends to.
Normally every layer recomputes that index; IndexCache computes it once and shares it across layers (arXiv 2603.12201).
Used with Gated DSA in Tencent Hy4 preview to keep 1M-token contexts fast.
Gated DeepSeek Sparse Attention (Gated DSA)
architecture_concepts
Gated DeepSeek Sparse Attention (Gated DSA) is the attention module of Tencent's Hy4 preview (Hy series), adapted from DeepSeek's Sparse Attention (DSA, arXiv 2512.02556) with a gating mechanism, itself inspired by attention work in DeepSeek and GLM models. A lightweight indexer produces a sparse candidate set per query and attention runs only over those positions, cutting compute at very long context; Hy4 pairs it with IndexCache, which reuses the sparse index across layers. Hy4-preview configuration: 64 attention heads, query compression dimension 2048, key-value compression dimension 512, indexer with 32 heads / 128 head dimension and top-k 2048. Related but distinct: DeepSeek-V4's Compressed Sparse Attention (CSA) and DeepSeek-V4.1-Flash's CSA2 are the DeepSeek family's own sparsified attention schemes - DSA is the earlier scheme Hy4 gates and borrows.
Related FAQs (1)
Gated DSA = Tencent Hy4's attention module:
Based on DeepSeek Sparse Attention (arXiv 2512.02556), with an added gating mechanism.
An indexer selects a sparse candidate set (Hy4: 32 indexer heads, 128 head dim, top-k 2048) and queries attend only to those positions.
Compressed query/KV projections (2048/512 dims in Hy4) plus IndexCache cross-layer index reuse keep long-context (1M) cost low.
Not the same as DeepSeek-V4's CSA/CSA2 - those are DeepSeek's own sparse-attention generations; Gated DSA is Hy's gated adaptation of the earlier DSA.
NVFP4 quantization (NVIDIA)
architecture_concepts
NVFP4 is NVIDIA's 4-bit floating-point format (two FP4 values per byte with a shared 8-bit scale per block of 16, per the OCP MX spec) and the centerpiece of the Nemotron 3 family's quantization-aware pre-training recipe: the majority of linear layers - weights, activations, and gradients - train directly in NVFP4, while select stability-critical layers (latent projections, MTP layers, QKV/attention projections, embeddings) are kept in BF16 or MXFP8. Training in the deployed numeric format lets frontier-scale models (e.g. Nemotron-3-Ultra-550B-A55B, pre-trained on ~20T tokens) serve at a fraction of the memory and compute cost of BF16 with minimal accuracy loss - unlike post-hoc quantization, the weights never live at full precision. NVFP4 is also natively accelerated on NVIDIA Blackwell GPUs.
Format: two FP4 per byte + shared scale per 16-element block (OCP MX).
Quantization-aware pre-training: most linear layers train in NVFP4 (weights, activations, gradients); latent projections, MTP, QKV/attention and embeddings stay BF16/MXFP8 for stability.
Weights never exist at full precision - no post-hoc quantization step.
Native support on Blackwell GPUs; enables ~20T-token pre-training of a 550B model efficiently.
Latent Mixture-of-Experts (LatentMoE)
architecture_concepts
Latent Mixture-of-Experts (LatentMoE), used in NVIDIA's Nemotron 3 family, routes and computes MoE experts in a smaller latent dimension instead of the full model width: tokens are projected down into the latent space, expert routing and the experts' computation happen there, and results project back. Shrinking the routing/compute dimension improves accuracy per byte - expert capacity scales better than the parameter footprint suggests - and pairs naturally with NVIDIA's quantization-aware NVFP4 pre-training, where latent projections stay in BF16/MXFP8 while most other linear layers run NVFP4. The Nemotron 3 Ultra/Super stacks interleave Mamba-2 layers, MoE layers, and select attention layers under this latent routing scheme.
Related FAQs (1)
LatentMoE = MoE experts working in a compressed latent space:
Tokens are projected into a smaller latent dimension.
Expert routing + computation happen in that latent space.
Results project back to full width.
Why: better accuracy per byte - expert computation scales without proportional memory cost, and it composes with NVFP4 quantization (latent projections stay high-precision BF16/MXFP8 while bulk layers go 4-bit). Used in Nemotron-3-Ultra-550B-A55B and Super siblings.
MoonViT
architecture_concepts
MoonViT is Moonshot AI's vision encoder family used in Kimi multimodal models (K2.5/K2.6/K2.7). It is a ViT-style encoder adapted for native-resolution, variable-aspect image and video understanding in LLM front-ends - in Kimi K2.7 Code it contributes ~400M parameters, pairing with the 1T-parameter MoE language model (MLA attention, 384 experts) to give the coding agent the ability to read screenshots, UI captures and documents. MoonViT follows the Kimi line's philosophy of running vision natively at the model's context length rather than through a separate fixed-resolution pipeline.
Related FAQs (1)
MoonViT = Moonshot AI's vision encoder for Kimi multimodal models:
ViT-style, ~400M parameters in Kimi K2.7 Code.
Handles native-resolution, variable-aspect image + video input.
Feeds visual tokens straight into the MoE LLM's context (no fixed-resolution separate pipeline).
Used across Kimi K2.5/K2.6/K2.7 for screenshot/UI/document understanding in coding and agentic workflows.
DFlash block-diffusion drafter
architecture_concepts
DFlash (arXiv 2602.06036) is a block-diffusion model used as a lightweight speculative-decoding drafter: instead of predicting one token at a time, it proposes entire blocks of tokens (e.g. 16 tokens) in a single forward pass, which the main model then verifies in parallel - accepting correct tokens and correcting wrong ones. Because verification is parallel, output quality is identical to standard autoregressive decoding while generation is significantly faster. Muse-Glimmer-30B ships with a DFlash drafter (5 draft layers, 16-token blocks, sliding-window attention with 2048 window on all layers, 32 Q / 8 KV GQA heads, sequence length 131,072, hidden features drawn uniformly from target layers {1, 13, 25, 37, 49} of 52), delivering measured speedups of 3.1x on an NVIDIA RTX 5090, 1.8x on Apple M5 Max and 1.5x on M4 Max with quantized drafters available to cut memory overhead. Distinct from the DeepSeek-V4 "DFlash attention" sparse-attention mechanism despite the shared name.
Related FAQs (1)
DFlash (arXiv 2602.06036) drafts tokens in blocks, not one-by-one:
A small companion network predicts a 16-token block in one forward pass.
The main model verifies the whole block in parallel - accept correct tokens, fix wrong ones.
Output is bit-identical to normal decoding, just much faster.
Muse-Glimmer-30B uses a 5-layer drafter (SWA-2048, GQA 32/8) and ships quantized drafter variants: measured 3.1x on RTX 5090, 1.8x M5 Max, 1.5x M4 Max. (Not to be confused with DeepSeek-V4's DFlash attention.)
Perception Encoder (Meta)
architecture_concepts
Meta's Perception Encoder (PE, arXiv 2504.13181) is a vision-only ViT family designed as the perception backbone for multimodal LLMs - the largest variants reaching ViT-G scale. Unlike CLIP-style encoders trained with contrastive text alignment, PE is trained with core vision losses (self-supervised and weakly-supervised distillation from large vision-only teachers) followed by aligner modules that attach language - which resolves the tension between pure visual tasks (depth, segmentation) and vision-language tasks (VQA) that plagues dual-encoders. In Muse-Glimmer-30B, the perception encoder is a ~1.8B-parameter ViT-G/14 (50 layers, width 1536, patch size 14) that tokenizes interleaved text and images (up to 4,096 visual tokens per image), letting the agent interpret screenshots, charts and documents alongside conversation.
Related FAQs (1)
Perception Encoder (arXiv 2504.13181) is Meta's ViT backbone for multimodal LLMs:
Vision-only core losses (self-supervised + distillation), no contrastive text alignment - so visual understanding isn't warped by language training.
Separate aligner modules attach language afterward (contrastive aligner for retrieval, generative aligner for VQA).
In Muse-Glimmer-30B: ~1.8B-param ViT-G/14, 50 layers, width 1536, up to 4,096 visual tokens/image.
Result: strong on pure vision tasks (depth/segmentation) AND vision-language tasks simultaneously.
Configurable reasoning effort (gpt-oss)
architecture_concepts
Configurable reasoning effort in the gpt-oss models (120b and 20b) lets the caller choose how much thinking the model does per query - low, medium, or high - via a simple system prompt (Reasoning: low / medium / high) or the reasoning_effort API parameter. Lower effort answers faster with less chain-of-thought; higher effort spends more tokens reasoning before answering. The full chain-of-thought is always exposed to the developer (for debugging and trust; OpenAI notes it is not intended to be shown to end users), and the model is trained to co-exist with tool calls inside that reasoning stream. This native effort control is one of gpt-oss's signature features and the reason latency-sensitive agentic deployments pick low/medium while hard math picks high.
Related FAQs (1)
Set the effort per request - three levels:
System prompt: Reasoning: low / Reasoning: medium / Reasoning: high
API parameter: reasoning_effort in the OpenAI API
Effects: low = short/no chain-of-thought, fastest; high = long reasoning, best for hard problems. The chain-of-thought is fully exposed to the developer (not meant for end users), and reasoning interleaves with tool calls natively.
MXFP4 quantization
architecture_concepts
MXFP4 is a 4-bit floating-point microscaling format (per the OCP MX standard) that OpenAI used to post-train the gpt-oss models: the MoE (expert) weights are quantized to MXFP4 while attention and shared weights stay higher precision, packing FP4 values with per-block scaling factors. gpt-oss-120b ships this quantization natively - so the 117B-parameter model fits and runs on a single 80GB GPU (NVIDIA H100 or AMD MI300X), and gpt-oss-20b runs within 16GB of memory. All of OpenAI's published gpt-oss evals were performed with the same MXFP4 quantization, so the open weights' quality is measured exactly as distributed.
Related FAQs (1)
MXFP4 = 4-bit floating-point values in the microscaling (MX) format:
Only the MoE expert weights are quantized to 4-bit FP (with block-wise scale factors); attention and shared weights stay higher precision.
gpt-oss models are post-trained with this quantization - the shipped weights are MXFP4, not FP16 later squeezed.
Result: gpt-oss-120b runs on one 80GB GPU (H100/MI300X); gpt-oss-20b fits in 16GB.
OpenAI's published evals use the same quantized weights, so reported scores match what users run.
MiniMax Sparse Attention (MSA)
architecture_concepts
MiniMax Sparse Attention (MSA), introduced with MiniMax-M3, is a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces both the attention compute and the memory footprint while preserving model quality. On MiniMax-M3 at 1M context it delivers 9x prefill and 15x decode speedups versus the M2 generation and cuts per-token compute to 1/20, making native 1M-token multimodal contexts practical. The open-source implementation lives at github.com/MiniMax-AI/MSA.
Related FAQs (1)
MSA is MiniMax's sparse attention operator for million-token contexts:
Replaces GQA's dense attention with a sparsified pattern that cuts compute and KV memory while preserving quality.
Measured on M3 at 1M context: 9x faster prefill, 15x faster decode vs M2; per-token compute drops to 1/20.
It is the key enabler of MiniMax-M3's native 1M-token multimodal context. Code: github.com/MiniMax-AI/MSA; technical report arXiv 2606.13392.
Language World Model (LWM)
architecture_concepts
A Language World Model (LWM) is a language model trained to simulate environments - predicting how a world (terminal, web page, Android device, OS, tool API) responds to actions, not just to answer prompts. Qwen's AgentWorld line builds LWMs by injecting environment knowledge during continual pre-training, then teaching next-state-prediction reasoning (SFT) and optimizing simulation fidelity with RL (GSPO), so environment modeling is native rather than a post-hoc adaptation on a general LLM. Because the model can imagine environment transitions, it enables controllable perturbations, fictional-world construction (synthetic environments that train agents better than real ones), and zero-shot transfer to out-of-distribution environments - LWM RL warm-up on single-turn trajectories transfers to multi-turn tool-calling tasks across 7 benchmarks, including 3 entirely out-of-domain.
Related FAQs (1)
An LWM learns how environments behave, not just language:
SFT trains next-state prediction: given an action, predict the environment's response.
RL (GSPO) sharpens simulation fidelity - imagined worlds match real ones.
Why it matters: the model can imagine environments - enabling controllable perturbations, fictional worlds that train agents better than real environments, and zero-shot transfer to unseen environments (e.g. OpenClaw). Qwen-AgentWorld-35B-A3B is the open reference implementation (arXiv 2606.24597).
Manifold-Constrained Hyper-Connections (mHC)
architecture_concepts
Manifold-Constrained Hyper-Connections (mHC), used in the DeepSeek-V4 series, upgrade the plain residual stream by learning multi-channel layer-to-layer mixes under a manifold constraint that keeps the mixing weights well-conditioned. Each layer reads from and writes to several parallel connection channels with learned coefficients; the manifold constraint prevents those coefficients from drifting into degenerate configurations, so signal propagation stays stable across very deep stacks while preserving expressivity. mHC generalizes the Hyper-Connections idea (multi-channel learned residuals, e.g. hc_mult=4 channels in V4 configs, with Sinkhorn-balanced mixing) and is credited by DeepSeek for V4's stable training of its deep hybrid-attention MoE stack.
Related FAQs (1)
mHC (DeepSeek-V4) replaces x + f(x) residuals with constrained learned mixes:
Hyper-Connections: each layer reads/writes a mix of 4 parallel residual channels (config hc_mult: 4).
Manifold constraint: keeps mixing weights on a well-behaved manifold (with Sinkhorn-style balancing, hc_sinkhorn_iters: 20), avoiding degenerate drift.
Effect: stable signal propagation through the very deep V4 stacks (43+ layers) without losing expressivity - DeepSeek credits mHC (plus Muon) for V4's training stability.
Compressed Sparse Attention (CSA)
architecture_concepts
Compressed Sparse Attention (CSA), introduced with the DeepSeek-V4 series, is one half of the family's hybrid long-context attention (paired with Heavily Compressed Attention, HCA). CSA sparsifies which positions a query attends to while compressing the retained key/value representations, so both the search over the context and the cache that backs it shrink. DeepSeek reports that at 1M-token context the CSA/HCA hybrid lets DeepSeek-V4-Pro run with only 27% of single-token inference FLOPs and 10% of the KV cache versus DeepSeek-V3.2 - the enabling factor behind the series' native 1M context.
Related FAQs (1)
CSA is DeepSeek-V4's sparse-attention half (paired with HCA):
Selects a sparse subset of context positions per query (rather than attending to everything).
Compresses the kept K/V so the cache backing them is much smaller.
Combined with Heavily Compressed Attention (HCA) for the remaining global signal.
Measured effect at 1M context: 27% of per-token FLOPs and 10% of KV cache vs DeepSeek-V3.2 - making native million-token context practical for both V4-Pro (1.6T/49B active) and V4-Flash (284B/13B active).
N-gram Embedding (Qwen3.8)
architecture_concepts
N-gram Embedding, introduced in Qwen3.8-Flash-Next, scales parameters through embedding tables instead of experts. A table of 20,000,000 bigram/trigram embeddings is added at layer 2, contributing 51B parameters - four times the model's 6B activated MoE path - while adding almost no compute: lookups replace matrix-heavy FFN work, and embedding rows are trivially offloadable to CPU/cheap memory. Guided by the observation that embeddings are the most memory-efficient parameter axis, this lets a "6B-active" model carry 125B+ parameters of n-gram and MoE capacity in a way that fits memory-constrained accelerators without sacrificing quality.
Related FAQs (1)
N-gram Embedding scales parameters where it's cheap:
A 20M-entry bigram/trigram table at layer 2 adds 51B parameters (vs 6B activated).
Embedding lookup costs little compute (no big matmuls) and offloads easily to cheap memory.
The idea: embeddings are the most computation- and offload-friendly axis for parameter scaling - more efficient than scaling MoE experts on memory-constrained hardware.
In Qwen3.8-Flash-Next this gives 125B total parameters (plus 4B MTP) with only 6B activated per token.
Gated Residual (Qwen3.8)
architecture_concepts
Gated Residual, introduced in Qwen3.8-Flash-Next, replaces the plain normalized residual stream with a learned, data-dependent gating scheme: each layer's information passes through widened residual streams whose flow is modulated by an element-wise read gate (choosing what each branch takes in, per dimension) and per-branch scalar write gates (choosing how much each branch contributes back). Qwen3.8-Flash-Next uses 4 branches with a bottleneck rank of 320. Compared with standard pre-norm residuals, this gives finer-grained expressiveness across layers while preserving the training stability deep stacks need - and adds only low inference overhead.
Related FAQs (1)
Gated Residual upgrades the residual stream in Qwen3.8-Flash-Next:
Read gate (element-wise, data-dependent): each of the 4 residual branches selects, per dimension, what it reads from the stream.
Write gates (per-branch scalars): each branch controls how strongly it writes back.
Bottleneck rank 320 keeps the widened stream affordable.
Benefit: deep 48-layer hybrid stacks stay stable to train while gaining much finer control over what each layer passes forward than a fixed x + f(x).
Qwen Sparse Attention (QSA)
architecture_concepts
Qwen Sparse Attention (QSA), introduced in Qwen3.8-Flash-Next, sparsifies attention at the micro-block level rather than per token. A lightweight MQA indexer (4 query heads + 1 shared key head, head dim 128) scores relevance, but attention keeps a budget of whole 512-token blocks (up to 2048 tokens) - so every kept token's neighbors come along, preserving local structure that token-level top-k selection destroys. QSA replaces the Gated Attention layer in the Qwen3.8 hybrid stack (paired with Gated DeltaNet) and runs with 24 Q / 2 KV heads of head dim 256. The design cuts long-context latency significantly, which matters most for agentic workloads dominated by long tool-call histories.
Related FAQs (1)
QSA (Qwen3.8-Flash-Next) sparsifies attention by selecting micro-blocks, not tokens:
An MQA indexer (4 query heads + 1 shared key head, dim 128) scores candidate 512-token blocks.
Attention runs over a fixed budget of whole blocks (up to 2048 tokens).
Because blocks move together, local context around each relevant region is preserved - unlike per-token top-k, which can isolate tokens from their neighbors.
In the hybrid stack it pairs with Gated DeltaNet (12x [3 GDN + 1 QSA] layers) and cuts long-context latency for agentic workloads.
IndexShare (GLM-5.2)
architecture_concepts
IndexShare (GLM-5.2, arXiv 2603.12201) makes long-context sparse attention cheaper by sharing one indexer across groups of attention layers. A sparse-attention indexer normally selects the top-k relevant KV positions per query for each layer independently; IndexShare reuses the same indexer for every four attention layers (in GLM-5.2: 21 full + 57 shared indexers over 78 layers, index_topk_freq: 4), so 75% of layers skip their own index computation entirely. Combined with top-2048 index selection (32 index heads, dim 128) and MTP-aware index reuse (index_share_for_mtp_iteration), this reduces per-token FLOPs by 2.9x at 1M-token context - the key enabler of GLM-5.2's "solid 1M context" claim.
Related FAQs (1)
IndexShare cuts the cost of sparse attention at extreme context:
Standard DSA-style designs run a learned indexer per layer to pick top-k KV positions.
IndexShare shares one indexer across every 4 attention layers (top-k 2048; 32 index heads, dim 128).
GLM-5.2 stacks 78 layers with only 21 full indexers (57 shared), and reuses indexes across MTP iterations.
Result: 2.9x fewer per-token FLOPs at 1M context than per-layer indexing, making stable 1M-token long-horizon work practical.
DSpark (DeepSeek-V4)
architecture_concepts
DSpark is a component of the DeepSeek-V4 forward path (named in the DeepSeek-V4-Flash-Vision-Exp reference inference). Its config keys mark a parallel token stream: dspark_block_size: 5, dspark_markov_rank: 256, dspark_target_layer_ids: [40, 41, 42] (the final three of 43 layers), and a dedicated dspark_noise_token_id. Together with num_nextn_predict_layers: 3 (multi-token prediction) this forms DeepSeek-V4's speculative/parallel generation machinery: blocks of up to 5 tokens are processed through a Markov-rank-256 state on the last layers with a noise token controlling stochastic exploration. The public reference implementation documents its forward path; DeepSeek has not yet published a full paper on DSpark.
Related FAQs (1)
DSpark appears in DeepSeek-V4 configs and the Flash-Vision-Exp reference inference as a parallel-decoding component:
Processes token blocks of size 5 with a Markov state of rank 256.
Active on the final 3 layers (dspark_target_layer_ids: [40, 41, 42]).
Has a dedicated noise token id (128799) for stochastic exploration.
Works alongside num_nextn_predict_layers: 3 (MTP) for speculative generation.
Note: DeepSeek documents DSpark only via config and reference code so far; the description above reflects those artifacts, not a published paper.
Hyper-Connections (DeepSeek-V4)
architecture_concepts
Hyper-Connections (DeepSeek-V4) replace the plain residual stream with a learned mixing network: each layer's input is a weighted combination of hc_mult=4 parallel connection channels, and layer outputs are mixed back into those channels with weights produced by a tiny network (regularized by hc_eps and balanced via hc_sinkhorn_iters=20 in DeepSeek-V4's config). Instead of a fixed x + f(x) residual, the model learns how much of each channel each layer should read from and write to. This stabilizes training of very deep networks, lets different layers specialize (e.g. the 43-layer DeepSeek-V4-Flash stack), and improves gradient flow compared with standard and dense residual variants (the idea generalizes Hyper-Connections from bytedecoder/Hyper-Connections, 2024).
Related FAQs (1)
Hyper-Connections generalize the residual connection:
Standard transformer: x = x + f(x) - one fixed residual stream.
Hyper-Connections: 4 parallel channels (hc_mult: 4 in DeepSeek-V4 config); each layer reads a learned mix of the channels and writes its output back as another learned mix.
A tiny network emits the mixing weights; Sinkhorn iterations (20) keep them balanced.
Why: deeper stacks train more stably, layers can specialize on different channels, and gradients flow better - used across the 43-layer DeepSeek-V4 backbone.
DFlash attention (DeepSeek-V4)
architecture_concepts
DFlash attention is the DeepSeek-V4 family's attention mechanism (used in DeepSeek-V4-Flash and the Flash-Vision-Exp multimodal variant). It combines latent-compressed queries and outputs (q/o LoRA ranks of 1024, outputs factored into 8 groups) with a learned sparse index: a separate set of 64 index heads (dim 128) scores every position and keeps only the top-512 entries per query token for the full attention computation, while 3 hash layers and a 128-token sliding window provide additional locality. Full attention runs with 64 heads of head dim 512 but only 1 KV head. The result is near-linear cost on very long sequences while preserving retrieval quality, which together with YaRN scaling supports DeepSeek-V4's 1M-token context.
Related FAQs (1)
DFlash attention (DeepSeek-V4) makes long-context attention cheap via three tricks:
Latent compression: queries and outputs pass through LoRA-style low-rank projections (rank 1024; outputs split into 8 groups).
Learned sparse index: 64 index heads (dim 128) score all positions; only the top-512 are attended per query, plus 3 hash-attention layers and a 128-token sliding window.
Massive head dim with tiny KV: 64 heads of dim 512, but only 1 KV head.
Together with YaRN (factor 16 from 64K), this powers the DeepSeek-V4-Flash family's 1,048,576-token context at near-linear cost.
Hybrid thinking modes (Qwen3)
architecture_concepts
Hybrid thinking modes (Qwen3's signature feature) pack a thinking mode and a non-thinking mode into a single model with seamless runtime switching. Thinking mode targets complex logical reasoning, math, and coding (chain-of-thought before the answer); non-thinking mode gives efficient general-purpose dialogue. The mode is chosen per request via the chat template (enable_thinking=True/False) or the prompt (/no_think), and Qwen3's post-training includes a dedicated thinking-mode-fusion stage that aligns both modes inside one checkpoint. This removes the need to deploy separate reasoning and chat models while keeping near-parity with specialized reasoning models in thinking mode.
Related FAQs (1)
Qwen3 models run both modes in one checkpoint:
Thinking mode: chain-of-thought reasoning for math/logic/coding.
Non-thinking mode: direct answers, efficient dialogue.
Switch per call with the chat template kwarg enable_thinking (True/False), or append /no_think (or /think) in the prompt. Qwen3's four-stage post-training ends with thinking-mode fusion, so the same weights serve both modes - no separate reasoning model needed.
Agent Swarm (Kimi K2.5)
architecture_concepts
Agent Swarm is Kimi K2.5's self-directed, coordinated execution scheme: instead of scaling a single agent loop, the model decomposes a complex task into parallel sub-tasks and dynamically instantiates domain-specific agents to execute them in a swarm-like fashion. Each specialized agent handles its sub-task (e.g. visual data processing, coding, tool use), and execution is coordinated back into one coherent result. This shifts the scaling axis from single-agent chain length to parallel agent populations, improving throughput on large multi-step tasks and grounding each sub-task in an agent tuned for that domain.
Related FAQs (1)
Agent Swarm (a key feature of Kimi K2.5) means the model:
Decomposes complex tasks into parallel sub-tasks.
Dynamically instantiates domain-specific agents to execute them (coding agents, visual-data agents, tool-use agents).
Coordinates the swarm's outputs into one result.
Compared with a single long agent loop, swarm execution parallelizes work, lets each agent specialize, and mirrors how large tasks are decomposed in production pipelines. It is an inference-time orchestration capability built through the model's agentic training (15T visual-text continual pretraining + agentic paradigms).
Multi-head Latent Attention (MLA)
architecture_concepts
Multi-head Latent Attention (MLA), introduced by DeepSeek-V2, compresses the key-value cache into a low-rank latent vector. Instead of caching full K and V per head per token, each token's KV information is projected down to a shared latent (e.g. kv-lora rank 512 in Kimi K2.5 and DeepSeek-V3) and up-projected on the fly during attention. This shrinks the KV cache by an order of magnitude at long context lengths, and queries are also optionally compressed (q-lora rank 1536). MLA preserves quality close to standard multi-head attention while making million-token-scale serving and trillion-parameter MoE models practical.
Kimi K2/K2.5 uses MLA as its attention mechanism (64 heads, 128 nope + 64 rope head dims, v_head_dim 128), inheriting the DeepseekV3 backbone design.
Related FAQs (1)
MLA (DeepSeek-V2, 2024) replaces the full per-head KV cache with a compressed latent:
Each token's keys/values are jointly encoded into a small latent vector (kv-lora rank 512).
Attention re-expands the latent on demand, so the cache holds rank-512 vectors instead of full multi-head K/V.
Decoupled RoPE part (qk_rope dims) stays outside the latent to preserve positional information.
Result: 10x+ smaller KV cache at long context, near-parity quality; the enabling trick behind DeepSeek-V2/V3, Kimi K2/K2.5 and other trillion-parameter MoE models.
Shared expert (MoE)
architecture_concepts
A shared expert in a Mixture-of-Experts model is an always-on expert FFN that processes every token alongside the routed experts, while the router activates only k of the remaining N experts per token (e.g. Gemma 4 26B-A4B: 1 shared + 128 routed, top-8 active; DeepSeek-V2/V3 use the same pattern). The shared expert captures common knowledge and generic computation that every token needs, letting routed experts specialize more cleanly and reducing redundant capacity across experts.
Related FAQs (1)
In MoE models with a shared expert, each token passes through:
The shared expert - one FFN always active for every token (common/generic computation).
The top-k routed experts selected by the router for that token (specialized computation).
Benefits: routed experts specialize more (no need to relearn generic features), load balancing improves, and quality per active parameter rises. Gemma 4 26B-A4B uses 1 shared + 128 routed experts with top-8 routing (3.8B of 25.2B parameters active per token); DeepSeek-V2/V3 popularized the design.
Cross-layer KV cache sharing
architecture_concepts
Cross-layer KV cache sharing lets consecutive Transformer layers reuse the same keys/values instead of each layer computing and storing its own KV cache - the attention outputs of one layer's projection serve its neighbors. In Gemma 4 E-models, 18 of the layers share KV entries (config num_kv_shared_layers: 18), directly shrinking the KV cache that dominates long-context memory on edge devices. Combined with sliding-window attention, it keeps the E4B's 128K context practical on phones and laptops.
Related FAQs (1)
Instead of every attention layer holding a separate KV cache, adjacent layers share one set of keys/values (Gemma 4 E4B/E2B share across 18 layers, num_kv_shared_layers: 18).
Effect: the KV cache shrinks roughly by the sharing factor - critical because KV memory, not weights, dominates long-context serving on devices.
Trade-off: slightly less per-layer expressiveness; acceptable for edge models where memory is the binding constraint.
Pairs with sliding-window attention (window 512) so most layers only cache 512 positions anyway.
Encoder-free multimodal architecture
architecture_concepts
Encoder-free multimodal architecture (Gemma 4 12B "Unified") removes the dedicated vision/audio encoders that most multimodal LLMs bolt onto a text backbone. Instead, raw image patches and audio waveforms are projected directly into the LLM's embedding space through lightweight linear layers, so every modality flows into a single decoder-only transformer. This reduces multimodal latency, simplifies the pipeline (no separate ViT/AST to download or run), and allows the entire model to be fine-tuned in one pass - at the cost of putting all cross-modal work on the LLM itself.
Related FAQs (1)
Most multimodal LLMs pre-process images/audio with dedicated encoders (ViT, AST) and feed the resulting tokens to the LLM. The encoder-free (Unified) Gemma 4 12B instead:
Projects raw image patches and audio waveforms into the embedding space via lightweight linear layers.
All modalities then flow through the same decoder-only transformer - fewer components, lower multimodal latency.
The whole model can be fine-tuned in one pass (no frozen-encoder stages).
Trade-off: the LLM itself must learn cross-modal features from scratch; encoder-based models (Gemma 4 E2B/E4B/31B with ~150M-550M vision encoders) keep that workload in a specialist network.
Proportional RoPE (p-RoPE)
architecture_concepts
Proportional RoPE (p-RoPE) is the positional-encoding scheme Gemma 4 applies on its global attention layers (rope_type "proportional", base frequency 1M with partial rotary factor 0.25, versus plain RoPE at 10k on sliding-window layers). It is part of Gemma 4's long-context memory optimization: together with unified Keys and Values (K=V) on global layers, p-RoPE lets the 256K-token context window fit in much less KV-cache memory while keeping precise global retrieval.
Related FAQs (1)
p-RoPE is Gemma 4's proportional rotary variant used on the sparse global-attention layers (theta 1M, partial rotary 0.25).
Sliding-window layers use plain RoPE (theta 10k) since they only see 512-1024 tokens.
Global layers carry the long-range positional signal: p-RoPE + unified K=V heads shrink the KV cache so 256K contexts stay practical.
Introduced with the Gemma 4 family (technical report arXiv 2607.02770).
Per-Layer Embeddings (PLE)
architecture_concepts
Per-Layer Embeddings (PLE) give every decoder layer its own small per-token embedding table instead of enlarging the transformer blocks. The tables are large but are only used for cheap lookups, and each layer reads its own embedding and adds it to that layer's input. PLE decouples the "effective" compute parameters from the total stored parameters: Gemma 4 E2B has 2.3B effective parameters (5.1B with embeddings) and E4B 4.5B effective (8B with embeddings), maximizing per-layer capacity for on-device deployment without slowing inference.
Related FAQs (1)
PLE = each decoder layer gets its own small embedding table for every token.
The layer looks up its per-layer embedding and adds it to the block input, letting layers specialize without adding transformer FLOPs.
Effective vs total parameters: E2B = 2.3B effective (5.1B with embeddings), E4B = 4.5B effective (8B with embeddings) - the tables are memory but not compute.
Purpose: maximize per-layer capacity for on-device deployment (Gemma 4 E2B/E4B) while keeping inference fast.
Nemotron Elastic (model compression framework)
architecture_concepts
Nemotron Elastic is NVIDIA's compression framework for deriving smaller, deployment-efficient models from a parent model while retaining its capabilities. NVIDIA-Nemotron-3-Nano-4B was produced from the 9B NVIDIA-Nemotron-Nano-9B-v2 via this framework (arXiv 2511.16664): the elastic procedure shrinks the hybrid Mamba2-Transformer architecture and continues training so the compressed 4B model keeps the parent's reasoning behavior and accuracy profile at a fraction of the size.
Related FAQs (1)
Nemotron Elastic compresses a large parent model into a smaller elastic variant that is then continued-trained to recover capability.
Applied by NVIDIA to derive Nemotron-3-Nano-4B (3.97B) from Nemotron-Nano-9B-v2.
The compressed model keeps the parent's hybrid Mamba2-Transformer architecture family and reasoning/non-reasoning unified behavior, at a size suitable for edge platforms (Jetson Thor, GeForce RTX, DGX Spark).
Details in arXiv 2511.16664.
Mamba2
architecture_concepts
Mamba2 is a selective state-space model (SSM) layer - the successor to Mamba - that processes sequences with a linear-time recurrent scan instead of quadratic attention. Each token's hidden state is updated by input-dependent (selective) SSM dynamics, so the model compresses context into a fixed-size state: no KV cache grows with sequence length. Mamba2 simplifies Mamba's design (SSD state-space duality), enabling larger state dimensions and better hardware utilization. Hybrid stacks that interleave a few full-attention layers with many Mamba2 blocks (e.g. NVIDIA's Nemotron-H family, Zamba) get attention-like recall where needed at a fraction of the memory and compute.
Related FAQs (1)
A Mamba2 block is a selective SSM: it carries a fixed-size recurrent state through the sequence, with input-dependent transitions (the model chooses what to remember/forget per token).
Cost: linear in sequence length; memory stays constant regardless of context length (no KV cache).
Weakness: fixed state can lose precise retrieval of old tokens (needle-in-haystack recall).
Fix - hybrids: NVIDIA's Nemotron-H/Nano models use mostly Mamba2 + MLP layers with only ~4 full-attention layers, getting long-context throughput with enough precise recall for a 262K window.
Hybrid SWA-full attention (3:1)
architecture_concepts
Hybrid SWA-full attention (3:1) is Spark-X2.5's efficiency-oriented attention layout: for every full-attention layer, three sliding-window attention (SWA) layers handle the remaining positions (21 SWA + 7 full layers in the 1.7B model, SWA window 512). Local SWA layers keep the KV cache and compute small, while the sparse full-attention layers preserve global retrieval - letting the model support a native 1M-token context window with far less overhead than an all-full-attention stack. RoPE is applied with theta 5M and partial rotary factor 0.25 on full layers (10k, full rotary on SWA layers).
SWA layers (21 of 28) attend only within 512 tokens: tiny KV cache, fast TTFT.
Full-attention layers (7 of 28) give global context over the whole 1M-token window.
The mix balances performance, inference speed and cache usage for real-world deployment - the same trade-off Gemma 3 solves with 5:1 local:global interleave.
Speculative decoding
architecture_concepts
Speculative decoding accelerates autoregressive LLM inference by letting a small, fast drafter model propose several tokens which the large target model verifies in a single parallel forward pass. Accepted tokens are kept; on the first mismatch the target model's own next token is used, so outputs are identical to standard decoding. LFM2.5-8B-A1B ships a dedicated 328M-parameter drafter (LFM2.5-8B-A1B-DSpark) that pairs with it for ~2.5x faster decoding with identical outputs.
Related FAQs (1)
A small drafter model guesses the next K tokens; the large target model checks all K in one forward pass and accepts the longest correct prefix (plus its own correction token on mismatch).
Output distribution is provably identical to greedy/sampled decoding from the target model.
Speedup comes from amortizing the target's sequential decoding steps; a good drafter accepts most tokens.
Example: Liquid AI's LFM2.5-8B-A1B-DSpark (328M drafter) gives ~2.5x faster decoding with identical outputs; vLLM/SGLang/llama.cpp support spec decoding.
Grouped-Query Attention (GQA)
architecture_concepts
Grouped-Query Attention (GQA) is an attention variant where multiple query heads share a single key/value head, interpolating between multi-head attention (MHA, one KV head per query head) and multi-query attention (MQA, one KV head for all). GQA shrinks the KV cache and speeds decoding with little quality loss - e.g. LFM2.5-8B-A1B uses 32 query heads with 8 KV heads (4:1 sharing). It is standard in modern open LLMs (Llama 3, Gemma 3, Qwen3, LFM2.5).
Related FAQs (1)
GQA shares each key/value head across several query heads instead of giving every query head its own KV head.
Effect: the KV cache (and memory traffic during decoding) shrinks by the sharing ratio - e.g. LFM2.5-8B-A1B has 32 query heads / 8 KV heads, a 4x reduction vs MHA.
Quality: GQA sits between MHA (best quality, most memory) and MQA (most compression, some quality loss) and typically retains near-MHA quality.
Who uses it: Llama 3, Gemma 3, Qwen3, LFM2.5 and most modern open models.
Double-gated convolution blocks
architecture_concepts
Double-gated convolution blocks are the non-attention half of Liquid AI's LFM2/LFM2.5 hybrid architecture. Instead of attention, a token block passes through depthwise causal convolutions whose inputs and outputs are controlled by two learned gates (an input gate and an output gate), giving a fixed-size local receptive field (conv window L=3 in LFM2.5) with constant memory and compute per token. LFM2.5-8B-A1B uses 18 such conv blocks interleaved with only 6 GQA attention layers (ratio 3:1), which keeps the long-context KV-cache and attention FLOPs small - the key to its on-device throughput.
Related FAQs (1)
LFM2/LFM2.5 are hybrids: 18 double-gated convolution blocks + 6 GQA attention layers in LFM2.5-8B-A1B.
Conv blocks give local pattern mixing with a fixed window (L=3) - constant compute and no KV cache, ideal for edge devices.
"Double-gated" means learned gates modulate both the block's input and output, stabilizing the signal flow like LSTM/SSM gates.
The few attention layers provide global context (128K window), while convs handle most layers - so the model is fast on CPU and GPU (up to 18.5K output tokens/s on one H100).
Mixture of Experts (MoE)
architecture_concepts
Mixture of Experts (MoE) replaces a Transformer's dense feed-forward blocks with many parallel "expert" FFNs plus a router that activates only a few experts per token. This decouples total parameter count from per-token compute: e.g. LFM2.5-8B-A1B has 8.3B total parameters but only 1.5B active per token (top-4 routing across 32 experts). MoE yields large capacity at a fraction of the inference cost of an equally-sized dense model, at the price of higher memory footprint (all experts must be resident), routing complexity, and load-balancing challenges during training.
Related FAQs (1)
A MoE model keeps many expert FFN blocks and a router network that picks a small number of them per token.
Total vs active parameters: all experts are stored in memory, but each token computes through only a few - e.g. LFM2.5-8B-A1B: 8.3B total, 1.5B active (4 of 32 experts per token).
Why: much higher capacity per FLOP than dense models; competitive quality at far lower inference cost.
Costs: full model must fit in memory (quantization/GGUF variants help), and training needs load-balancing losses so experts specialize evenly.
SigLIP vision encoder
architecture_concepts
SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision encoder trained with a pairwise sigmoid loss instead of CLIP's softmax contrastive loss, which scales better with batch size. Gemma 3 uses a tailored, frozen 400M-parameter SigLIP ViT variant shared across its 4B, 12B and 27B models: images are resized to 896x896 and encoded into a fixed-size sequence of 256 soft tokens fed to the language model. A Pan-and-Scan (P&S) adaptive cropping scheme handles non-square and high-resolution images at inference time. (Gemma 3's encoder initializes from SigLIP v1, whereas Qwen3-VL builds on the newer SigLIP-2.)
Related FAQs (1)
Gemma 3 (4B/12B/27B) uses a tailored, frozen 400M-parameter SigLIP vision encoder - a ViT trained with a sigmoid image-text loss (instead of CLIP's softmax loss) that the Gemma team fine-tuned on visual-assistant data.
Images are resized to 896x896 and encoded into a fixed 256 soft tokens per image, which keeps image processing cost predictable. Non-square or high-resolution images are handled by Pan & Scan: the image is split into equal crops that are each encoded separately, an inference-time-only optimization.
Unlike Qwen3-VL (which continues from the newer SigLIP-2), Gemma 3 initializes its encoder from SigLIP v1 and keeps it frozen during LLM training.
5:1 Local-Global Attention Interleaving
architecture_concepts
5:1 Local-Global Attention Interleaving is Gemma 3's attention layer pattern: five local attention layers using sliding-window attention are followed by one global self-attention layer, starting with a local layer. Local layers attend within a bounded window (cheap, KV-cache-friendly), while every sixth layer attends over the full context, keeping long-range retrieval at 128K tokens while dramatically cutting compute and memory versus all-global stacks. The RoPE base frequency differs per layer type (1M on global, 10k on local).
Related FAQs (1)
Gemma 3 alternates five local attention layers with one global layer (5:1, starting local).
Local layers use sliding-window attention: each token sees only a bounded window, so their KV cache stays small and inference is cheap.
Global layers (every 6th) attend over the full context, preserving long-range retrieval at the model's 128K context.
This hybrid cuts compute and memory substantially versus making every layer global, with negligible quality loss. The RoPE base frequency is 1M on global layers and 10k on local layers, and the global span is extended via positional interpolation.
QK-Norm
architecture_concepts
QK-Norm normalizes the queries and keys in a Transformer's attention (scaling them to unit norm, typically with an RMSNorm layernorm applied to the Q and K vectors) before the attention dot product. Introduced for stabilizing training (Dehghani et al., 2023; Wortsman et al., 2023), QK-Norm prevents attention-logit blowups without the need for attention soft-capping. Google replaced Gemma 2's soft-capping with QK-Norm in Gemma 3, keeping large 128K contexts trainable and stable while preserving downstream quality.
Related FAQs (1)
QK-Norm applies normalization directly to the query and key vectors before attention, taming outlier attention logits during training.
Gemma 2 used attention soft-capping (squashing logits with tanh) for the same purpose; Gemma 3 replaced it with QK-Norm, which stabilizes training to large 128K contexts without the downsides of soft-capping (reduced expressiveness and incompatibility with some fused kernels). QK-Norm follows the findings of Dehghani et al. (ViT-22B) and Wortsman et al. (2023).
Qwen3-ViT
architecture_concepts
Qwen3-ViT is the vision encoder of the Qwen3-VL series. It initializes from SigLIP-2 (SO-400M for large LLMs; Large-300M for small LLMs) and continues training with dynamic input resolutions, 2D-RoPE and interpolated absolute position embeddings following CoMP. In Qwen3-VL-30B-A3B it is a 27-layer, 1152-hidden, 16-head encoder with 16px patches and 2x2 spatial merging; multi-level features from layers 8, 16 and 24 feed the DeepStack fusion. The Qwen3-VL Technical Report ablation shows Qwen3-ViT beats the raw SigLIP-2 encoder during CLIP-style pre-training.
Related FAQs (1)
Qwen3-ViT is the vision encoder of Qwen3-VL models. It starts from SigLIP-2 weights (SO-400M for large models like Qwen3-VL-30B-A3B) and is continued on dynamic-resolution data with 2D-RoPE and position-embedding interpolation (CoMP methodology).
In Qwen3-VL-30B-A3B the encoder has 27 layers, hidden size 1152, 16 heads, 16x16 patches and 2x2 spatial merging into the LLM's 2048 hidden dimension. Its intermediate layers (8, 16, 24) additionally feed the DeepStack feature-fusion mergers. Ablations in the Qwen3-VL Technical Report show it outperforms plain SigLIP-2.
Text-Timestamp Alignment
architecture_concepts
Text-Timestamp Alignment is Qwen3-VL's video temporal modeling technique. It moves beyond T-RoPE (timestamp rotary embedding used by earlier VLMs) to precise, timestamp-grounded event localization: textual descriptions are aligned to absolute timestamps in the video, enabling second-level indexing of long videos. Combined with Interleaved-MRoPE, it lets Qwen3-VL-30B-A3B handle hours-long video with full recall and precise temporal grounding.
Related FAQs (1)
Qwen3-VL uses Text-Timestamp Alignment: instead of encoding video time only through positional embeddings (as T-RoPE did), textual events are aligned to absolute timestamps, giving timestamp-grounded event localization.
Together with Interleaved-MRoPE (full-frequency temporal/spatial position encoding) this enables second-level indexing and full recall over hours-long videos - one of the headline Qwen3-VL capabilities.
SigLIP-2
architecture_concepts
SigLIP-2 is a family of sigmoid-loss image-text pre-trained vision encoders from Google DeepMind. It replaces the contrastive softmax loss of CLIP with a pairwise sigmoid loss that scales better and performs strongly at low resolutions. Qwen3-VL uses the SigLIP-2 architecture as its vision encoder: it continues training SigLIP2-SO-400M (large LLMs; SigLIP2-Large 300M for small LLMs) from official checkpoints with dynamic input resolutions, 2D-RoPE and interpolated absolute position embeddings following CoMP. The resulting encoder is called Qwen3-ViT.
Related FAQs (1)
SigLIP-2 is Google DeepMind's image-text pre-trained vision encoder family. Its key difference from CLIP: a pairwise sigmoid loss instead of softmax contrastive loss, which scales better with batch size and works well at low resolutions.
In Qwen3-VL, the vision encoder starts from SigLIP-2 checkpoints (SO-400M for the larger models like Qwen3-VL-30B-A3B, Large-300M for the 2B/4B variants) and is continued on dynamic-resolution data with 2D-RoPE and interpolated position embeddings (CoMP methodology). The continued-trained encoder is referred to as Qwen3-ViT and ablations in the Qwen3-VL Technical Report show it outperforms the raw SigLIP-2 initialization.
DeepStack
architecture_concepts
DeepStack is a vision-language fusion mechanism used in Qwen3-VL: instead of feeding only the final ViT layer into the language model, DeepStack injects multiple intermediate ViT feature maps into several early LLM layers through specialized mergers. Fusing multi-level visual features captures fine-grained details and sharpens image-text alignment. In Qwen3-VL-30B-A3B the vision encoder is 27 layers deep and the DeepStack visual indexes are layers 8, 16 and 24. The Qwen3-VL Technical Report ablates DeepStack with a 15B-A2B model, showing clear gains from multi-level feature fusion.
Related FAQs (1)
DeepStack is Qwen3-VL's multi-level visual feature fusion. Standard VLMs pass only the final vision-encoder layer to the LLM; DeepStack additionally injects intermediate ViT feature maps into early LLM layers via specialized mergers.
This keeps fine-grained details (text, small objects, textures) that get smoothed out in the final layer, sharpening image-text alignment. In Qwen3-VL-30B-A3B, the 27-layer ViT exposes DeepStack indexes at layers 8, 16 and 24. The Qwen3-VL Technical Report's ablation shows measurable gains from this fusion.
Interleaved-MRoPE
architecture_concepts
Interleaved-MRoPE is Qwen3-VL's redesign of multimodal rotary position embedding. Qwen2-VL's original MRoPE partitioned the embedding dimensions into contiguous temporal (t), horizontal (h) and vertical (w) subspaces, each with distinct rotary frequencies - an imbalanced frequency spectrum that degrades long-video understanding. Interleaved-MRoPE instead interleaves the t, h and w components across the embedding dimension, giving full-frequency allocation over time, width and height. This robust positional embedding enhances long-horizon video reasoning and is configured in Qwen3-VL-30B-A3B with mrope_section [24, 20, 20].
Related FAQs (1)
Interleaved-MRoPE is Qwen3-VL's multimodal rotary position embedding. Original MRoPE (Qwen2-VL) split the embedding dimensions into contiguous temporal/horizontal/vertical blocks, each with its own frequency band - an imbalanced spectrum that hurts long-video benchmarks.
Interleaved-MRoPE interleaves the t/h/w components across the embedding dimension instead, so every region of the spectrum carries temporal AND spatial frequencies. The result is robust positional encoding over time, width and height, which strengthens long-horizon video reasoning. In Qwen3-VL-30B-A3B the config sets mrope_interleaved: true with mrope_section: [24, 20, 20].
Gated Attention
architecture_concepts
Gated Attention is the full-attention branch of Qwen3-Next's hybrid attention layout. Unlike standard Transformer attention, it applies a learnable output gate to each attention head, which stabilizes training and improves long-context quality while retaining exact token-to-token attention.
In Qwen3-Next-80B-A3B-Thinking, Gated Attention layers are interleaved with Gated DeltaNet layers in a 12 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)) pattern. Each Gated Attention block uses 16 query heads and 2 KV heads with a head dimension of 256 and a rotary position embedding dimension of 64.
The gated design keeps the precise-retrieval benefits of full attention while adding the training stability that hybrid linear-attention models need for robust pre-training and post-training (together with zero-centered and weight-decayed layernorm).
Related FAQs (1)
Gated Attention is Qwen3-Next's full-attention variant: each attention head's output passes through a learnable gate before being combined. This stabilizes training and improves long-context quality compared to standard attention.
In the Qwen3-Next hybrid layout, one Gated Attention layer follows every three Gated DeltaNet (linear attention) layers, in a pattern of 12 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)). Each Gated Attention block has 16 Q heads, 2 KV heads, head dimension 256, and RoPE dimension 64.
The combination gives the model linear-attention efficiency for most of the sequence while periodically refreshing exact retrieval through full attention — enabling the native 262K context and the 1M YaRN-extended context of Qwen3-Next-80B-A3B-Thinking.
SWA Bounded Replay
architecture_concepts
SWA Bounded Replay is DeepSeek-V4.1-Flash's technique for handling sliding-window attention (SWA) KV states without persisting them.
Instead of writing every SWA KV state to SSD, the model keeps only a bounded window and reconstructs missing SWA KV states by replaying only the most recent n_win tokens when they are needed. This avoids SSD round-trips for SWA cache entirely and reduces the persistent KV cache footprint to roughly 1/8 of DeepSeek-V4-Flash.
It complements FP4 main KV caching and CSA2 in the model's overall KV-compression stack (890 bytes per token globally).
Related FAQs (1)
SWA Bounded Replay is how DeepSeek-V4.1-Flash avoids persisting sliding-window-attention KV cache. Rather than saving every SWA KV state, the model replays only the most recent n_win tokens to reconstruct missing SWA KV states on demand.
Result: no SWA KV writes to SSD, and the persistent KV footprint drops to about 1/8 of DeepSeek-V4-Flash. It is one of the four techniques (with CED, CSA2 and FP4 KV) behind the 890-bytes-per-token cache.
FP4 KV caching (E2M1)
architecture_concepts
FP4 main KV caching stores the attention KV cache in 4-bit floating point format (E2M1: 1 sign bit, 2 exponent bits, 1 mantissa bit) with one E4M3 scale factor per 16 channels, a block-scaled scheme in the spirit of NVFP4.
In DeepSeek-V4.1-Flash, FP4 main KV caching works together with CSA2 attention and SWA Bounded Replay to shrink the global KV cache to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1 - a decisive cost factor for long-context agentic workloads where KV cache memory dominates serving cost.
Related FAQs (1)
FP4 KV caching keeps the attention KV cache in 4-bit floats (E2M1 format with one E4M3 scale per 16 channels) instead of 16-bit values. In DeepSeek-V4.1-Flash this cuts the global KV cache to 890 bytes per token - about 1/4 of DeepSeek-V4-Flash and roughly 1/437 of DeepSeek-V1.
The savings matter most for agents: long prompts and tool loops make KV cache memory the dominant serving cost, and FP4 directly divides that cost.
Compressed Sparse Attention 2 (CSA2)
architecture_concepts
Compressed Sparse Attention 2 (CSA2) is the attention scheme of DeepSeek-V4.1-Flash. Each attention layer is assigned one of three static modes:
Full - standard full attention; the first Full Mode layer also builds the candidate pool for later indexing layers
Reindex - recomputes attention indices from shared state
Reuse - reuses Top-K sparse-attention indices computed by another layer
Layers share main KV and indexer K across each other, and a Hierarchical Sparse Indexer restricts later indexing layers to the candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length.
Together with FP4 main KV caching, CSA2 reduces the global KV cache footprint to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash - while supporting 1M-token contexts.
Related FAQs (1)
CSA2 is DeepSeek-V4.1-Flash's attention mechanism. Every attention layer runs in one of three static modes - Full, Reindex, or Reuse - so layers can share main KV and indexer K and reuse Top-K sparse-attention indices instead of recomputing them.
A Hierarchical Sparse Indexer further limits later indexing layers to the candidate pool built by the first Full Mode layer, so indexer cost stays bounded no matter how long the context gets. This is one of the key techniques behind the model's 890-bytes-per-token KV cache.
Causal Encoder-Decoder (CED)
architecture_concepts
Causal Encoder-Decoder (CED) is a Transformer architecture introduced with DeepSeek-V4.1-Flash that splits the network into a causal encoder followed by a decoder. Instead of deriving the decoder's global KV cache from each decoder layer's own hidden states, CED projects it from the final encoder hidden states.
This makes prefill dramatically cheaper: DeepSeek-V4.1-Flash (552B backbone parameters) activates only 8B parameters per token during prefill and 16B during decode, which is particularly cost-efficient for input-heavy agentic workloads where long prompts are processed once and short answers are generated.
The 40-layer model uses a 20-layer causal encoder plus a 20-layer decoder. Combined with CSA2 attention and FP4 KV caching, CED is one of the four techniques that cut the model's global KV cache to 890 bytes per token (~1/4 of DeepSeek-V4-Flash).
Related FAQs (1)
Causal Encoder-Decoder (CED) is the architecture behind DeepSeek-V4.1-Flash: a Transformer split into a 20-layer causal encoder and a 20-layer decoder, where the decoder's global KV cache is projected from the final encoder hidden states instead of being built from each decoder layer.
Because the expensive cache is computed once by the encoder, the model activates only 8B of its 552B backbone parameters per token during prefill (16B during decode). That makes CED especially cheap for agentic workloads with long inputs and short outputs. See the knowledge entry for details.
KV cache
architecture_concepts
During autoregressive inference, the attention keys and values of already-processed tokens are stored in a running cache so that each new token does not require recomputing attention over the whole past sequence. This key-value (KV) cache grows with sequence length and with the number of layers that cache their keys and values.
Why it is a bottleneck
For agent tasks, balancing "performance, inference speed, and cache usage" is frequently identified as a key bottleneck limiting model performance: multi-turn tool-calling and long interaction histories make the cache — not just raw compute — the limiting resource.
How architectures control it
Full attention in every layer: the cache grows with every token of context — simple, but expensive at long sequence lengths.
Sliding-window attention layers: only tokens inside a fixed window are cached, bounding the per-layer cache regardless of total sequence length.
Hybrid attention layouts: a majority of sliding-window layers plus a minority of full-attention layers keeps the cache small even at very long context lengths, while the few full-attention layers preserve global context. This balance is a key reason such models maintain strong real-world deployment behavior (e.g., favorable time-to-first-token and overall inference efficiency compared with similarly sized models).
Related FAQs (1)
By making most layers cache-cheap. In a standard transformer every layer caches full-attention keys and values, so the cache grows with every token of context — painful for agents with long multi-turn histories. A hybrid attention layout with three sliding-window attention layers per one full-attention layer solves this: the sliding layers only cache tokens inside their fixed window, while the few full-attention layers preserve global context. The KV cache stays bounded and small even at very long context lengths, which is the core tradeoff for agent tasks where performance, inference speed, and cache usage must be balanced.
Native 1M-token context window
architecture_concepts
A context capability that is built into the model itself rather than added on afterwards through external retrieval or aggressive truncation. "Native" means the model was trained to operate at that sequence length directly.
What makes long context practical
Training: a dedicated long-context training stage on large token volumes (e.g., hundreds of billions of tokens), with sequence lengths extended step by step up to the target length — this teaches the model to actually use very long inputs instead of merely accepting them.
Architecture: a hybrid attention layout (e.g., three sliding-window layers per one full-attention layer) substantially reduces the computational and cache overhead typically associated with long-context models.
Deployment
Serving frameworks such as SGLang and vLLM expose the full window via --context-length 1048576. The full setting requires sufficient device memory and can be reduced if needed.
Related FAQs (1)
Through two ingredients:
A dedicated long-context training stage — large token volumes with sequence lengths extended step by step up to the target length — teaches the model to actually use very long inputs instead of merely accepting them.
A hybrid attention architecture (three sliding-window layers per full-attention layer) keeps compute and KV-cache costs under control at those lengths.
The result is a native long context window (e.g., 1,048,576 tokens) that serving frameworks such as SGLang and vLLM expose via --context-length 1048576, requiring sufficient device memory and reducible if needed.
Sliding-window attention (SWA)
architecture_concepts
An attention variant in which each token attends only to a fixed window of neighboring tokens instead of the entire sequence. Because the window size is constant, the compute and memory of such a layer grow roughly linearly with sequence length, and its KV cache is bounded by the window size rather than the full context.
Role in hybrid layouts
Sliding-window layers are commonly combined with full-attention layers in a fixed ratio (e.g., three sliding-window layers per one full-attention layer). This pairing leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure.
The sliding-window layers process local patterns efficiently and keep per-layer cost and cache small even for very long inputs.
The full-attention layers preserve long-range accuracy and global information flow.
This improves practicality in real-world deployment scenarios where inference speed and cache usage matter — particularly for long-context and agentic workloads with long interaction histories.
Related FAQs (1)
An attention variant that restricts each token to attend only to a fixed window of nearby tokens rather than the whole sequence. This keeps the per-layer cost and KV cache small even for very long inputs.
In hybrid layouts, sliding-window layers are paired with full-attention layers (commonly three SWA layers per full-attention layer): the model keeps long-range accuracy through the full-attention layers while the sliding layers cut the overhead that normally makes long-context models slow and memory-hungry.
Hybrid attention architecture
architecture_concepts
An attention layout that interleaves full-attention layers with sliding-window attention (SWA) layers instead of using full self-attention in every layer. A common stacking pattern is one full-attention layer per three sliding-window layers: the full-attention layers give tokens access to the entire context, while the sliding-window layers handle local context with far less computation and cache.
Motivation
Systematically integrating and optimizing mature attention technologies combines the strengths of both mechanisms while avoiding the limitations of relying on a single structure.
Full attention: sees the whole context, but compute scales quadratically and the KV cache grows unboundedly as sequences grow.
Sliding-window attention: cheap and cache-bounded, but blind beyond its fixed window.
Effects
An effective balance among performance, inference efficiency, and KV-cache size — especially relevant for agent workloads with long interaction histories.
Practical support for long native context windows (e.g., up to 1M tokens) while substantially reducing the computational overhead typically associated with long-context models.
Properties
Fixed layer ratio between sliding-window and full-attention layers (commonly 3:1).
KV cache dominated by the windowed layers; grows near-linearly with sequence length.
Works with any full-attention variant (softmax attention, gated attention) as the global layers.
Related FAQs (1)
Because each mechanism fixes the other's weakness:
Full attention sees the whole context but becomes expensive and cache-hungry as sequences grow.
Sliding-window attention is cheap and cache-bounded but only sees a local neighborhood.
Stacking three sliding-window layers for every one full-attention layer processes local patterns efficiently while a minority of layers still propagates global information. The result is a balance among performance, inference efficiency, and KV-cache size — plus enough headroom for long native context windows (e.g., 1M tokens).
Multi-Token Prediction (MTP)
architecture_concepts
A training objective for Large Language Models that extends the standard next-token prediction (NTP) paradigm by requiring the model to predict multiple future tokens simultaneously from each position in the input sequence, rather than only the next token.
How it works
Introduced by Gloeckle et al. (Meta, 2024), the objective attaches n independent output heads to a shared transformer trunk. At each position, the shared trunk produces a latent representation of the context, and each of the n heads predicts one of the following n tokens using a shared unembedding matrix.
NTP: at position t, predict token t+1.
MTP: at position t, predict tokens t+1, t+2, ..., t+n simultaneously.
If n=1, the objective degenerates to standard NTP.
Key benefits
Higher sample efficiency: each training token provides n gradient signals instead of 1, extracting more learning signal per token.
Better downstream performance: improves performance on code and reasoning tasks without increasing training time or memory overhead.
Faster inference via self-speculative decoding: the extra heads can draft multiple candidate tokens per forward pass, which are then verified by the main model, enabling speculative decoding without a separate draft model.
Planning and lookahead: forces the model to plan ahead, improving coherence over longer generations.
Training and implementation
The heads are typically pretrained alongside the main model (joint training), though post-hoc approaches exist (e.g., self-distillation from the frozen target LLM).
Each head is implemented as additional transformer layers that share the trunk's representations.
The loss is the sum of cross-entropy losses across all n heads.
Post-hoc variant
Recent work (Kirchenbauer et al., 2026) shows the modules can be trained post-hoc for LLMs without official support, using self-distillation from the target LLM's native NTP capability. The target LLM remains frozen; only the extra module is trained.
References
Gloeckle et al. (2024), "Better & Faster Large Language Models via Multi-token Prediction", arXiv:2404.12040
Kirchenbauer et al. (2026), "Multi-Token Prediction via Self-Distillation", arXiv:2602.06019
Related FAQs (1)
Multi-Token Prediction (MTP) is a training objective that requires a model to predict several future tokens at every position instead of only the next one. Introduced by Gloeckle et al. (Meta, 2024), it forces the model to plan ahead and yields better pretraining performance per FLOP than standard next-token prediction.
Qwen uses MTP in two ways in the Qwen3-Next series:
Pretraining: MTP heads boost the pretrained model's performance on downstream tasks.
Inference acceleration: the MTP head doubles as a speculative-decoding drafter. With vLLM you can enable it via --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'; with SGLang via --speculative-algo NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4.
YaRN (Yet another RoPE extensioN)
architecture_concepts
A compute-efficient method to extend the context window of Large Language Models that were trained with Rotary Position Embeddings (RoPE). It requires 10x fewer tokens and 2.5x fewer training steps than previous methods, enabling models originally trained at shorter contexts to effectively utilize and extrapolate to much longer context lengths.
The extrapolation problem
When a model trained with RoPE at a maximum context length L is used with sequences longer than L, performance degrades sharply. The rotation angles for positions beyond L fall outside the range seen during training, causing the attention mechanism to produce unreliable scores.
Previous approaches
Position Interpolation (PI): compresses the entire position range by a scale factor s = L'/L, mapping positions [0, L'] to [0, L]. This uniformly scales all RoPE dimensions, which impairs the model's ability to understand small, local relationships.
NTK-aware scaling: adjusts the RoPE base frequency to stretch dimensions non-uniformly.
The method
The approach combines NTK-by-parts interpolation with an attention temperature scaling:
NTK-by-parts: instead of scaling every RoPE dimension equally, the interpolation pressure is spread across dimensions — high frequencies are scaled less (preserving local detail), while low frequencies are scaled more (enabling longer reach).
Attention temperature: a scaling factor 1/sqrt(t) is applied to the attention dot product to compensate for the changed distribution of attention scores, without modifying the attention code itself.
Static vs. dynamic scaling
Static: the scaling factor s is fixed at deployment time (e.g., s=4.0 for 1M context). It is constant regardless of input length, which may impact performance on shorter texts. All major open-source frameworks (vLLM, SGLang, TokenSpeed) implement this variant.
Dynamic: the scale factor is updated per forward pass based on actual sequence length, degrading gracefully. Not commonly used in production deployments.
References
Peng et al. (2023), "YaRN: Efficient Context Window Extension of Large Language Models", arXiv:2309.00071
YaRN (Yet another RoPE extensioN) is a compute-efficient method to extend the context window of LLMs trained with Rotary Position Embeddings (RoPE). It requires 10x fewer tokens and 2.5x fewer training steps than previous methods.
YaRN works by non-uniformly interpolating the RoPE dimensions: high-frequency dimensions (which encode local position detail) are scaled less, while low-frequency dimensions (which encode long-range position) are scaled more. This is combined with an attention temperature scaling factor.
Typical scaling factors: factor=4.0 stretches the position range 4x (e.g., 262K → 1M tokens); factor=2.0 for 524K tokens.
All open-source frameworks (vLLM, SGLang, TokenSpeed) implement static YaRN, meaning the scaling factor is fixed at deployment. This may slightly impact performance on shorter texts — the factor should only be modified when long contexts are actually needed.
See the knowledge entries on YaRN and Rotary Position Embedding (RoPE) for deeper background.
Rotary Position Embedding (RoPE)
architecture_concepts
A positional encoding scheme for Transformer models that encodes positional information by rotating query and key vectors in a 2D plane. The rotation angle depends on the token's position, so that nearby tokens have a small angular difference while distant tokens have a larger one. This naturally incorporates relative position dependency into the attention dot product without explicit position tokens.
Mathematical formulation
For a pair of features (x1, x2) at position m, a rotation by angle m*theta is applied:
The d features are organized as d/2 pairs, each rotated by a frequency that decreases exponentially. The attention dot product between positions m and n depends only on the relative position (m−n), not the absolute positions.
Key properties
Relative encoding: attention between two tokens depends on their distance, not absolute position.
No extra parameters: applied as a rotation, not learned — zero additional parameters.
Decoupled from attention: applied to Q and K before the attention computation.
Extrapolation challenge: trained at a maximum context length L, the encoding degrades when the sequence exceeds L because rotation angles fall outside the trained range.
Multimodal extension (mRoPE)
For vision-language models, the multimodal extension handles multiple dimensions (time, height, width for video). The frequency dimensions are split into sections (e.g., [11, 11, 10]) to allocate rotation capacity across spatial and temporal dimensions. mrope_interleaved mode interleaves the frequency pairs for better multimodal performance.
References
Su et al. (2021), "RoFormer: Enhanced Transformer with Rotary Position Embedding", arXiv:2104.09864
YaRN (Yet another RoPE extensioN) is a compute-efficient method to extend the context window of LLMs trained with Rotary Position Embeddings (RoPE). It requires 10x fewer tokens and 2.5x fewer training steps than previous methods.
YaRN works by non-uniformly interpolating the RoPE dimensions: high-frequency dimensions (which encode local position detail) are scaled less, while low-frequency dimensions (which encode long-range position) are scaled more. This is combined with an attention temperature scaling factor.
Typical scaling factors: factor=4.0 stretches the position range 4x (e.g., 262K → 1M tokens); factor=2.0 for 524K tokens.
All open-source frameworks (vLLM, SGLang, TokenSpeed) implement static YaRN, meaning the scaling factor is fixed at deployment. This may slightly impact performance on shorter texts — the factor should only be modified when long contexts are actually needed.
See the knowledge entries on YaRN and Rotary Position Embedding (RoPE) for deeper background.
Gated DeltaNet
architecture_concepts
A linear attention architecture that combines data-dependent gating with the delta update rule to achieve efficient, high-quality sequence modeling. Introduced by NVIDIA Research (NVlabs) as an improvement over Mamba2 and DeltaNet.
Background: linear attention
Standard Transformer attention scales quadratically with sequence length L. Linear attention replaces softmax with a kernelized dot product and reframes computation as a linear RNN with a matrix-valued state S_t, reducing inference memory to O(1) per step regardless of sequence length.
Delta rule
DeltaNet uses the delta update rule (Widrow et al., 1960): instead of accumulating all key-value pairs into the state (Hebbian learning), it selectively replaces the value associated with the current key. This is more surgical — it only modifies one key-value pair at a time while leaving others intact. The update is:
alpha_t (gate, 0 to 1): controls state decay. Setting alpha_t → 0 flushes the entire state (rapid forgetting); alpha_t → 1 reduces to the pure delta rule (selective update).
beta_t (writing strength, 0 to 1): controls how much of the new value is written.
This combination enables both bulk forgetting (via gating) and targeted updates (via delta rule), giving flexible memory control.
Hardware-efficient training
The recurrence is sequential, but a chunkwise parallel algorithm (using the WY representation and matmuls) enables GPU-efficient training with tensor cores. Within each chunk, the sequential dependencies are captured by a small C x C matrix inverse (C = chunk size).
Hybrid architectures
The layer is often combined with sliding-window attention (SWA) or Mamba2 layers in hybrid models. For example, a layout of 16 blocks of (3x gated-delta layers → 1x gated-attention layer) combines linear attention efficiency with standard attention's retrieval strength.
References
Yang et al. (2024), "Parallelizing Linear Transformers with the Delta Rule over Sequence Length", arXiv:2406.06484
This architecture combines linear attention layers with standard softmax attention layers in a fixed ratio (e.g., 3:1).
Gated DeltaNet is a linear attention mechanism that combines data-dependent gating (alpha_t) for rapid memory flushing with the delta update rule (beta_t) for targeted key-value updates. This gives O(1) inference memory per step.
Gated Attention is standard softmax attention that excels at precise retrieval.
The 3:1 ratio means 75% of layers use the efficient linear mechanism while 25% provide the retrieval strength of full attention. The Gated DeltaNet layers handle long-range modeling efficiently, while the Gated Attention layers provide precise in-context retrieval.
This hybrid design allows models to natively support long contexts (e.g., 262K) and scale to even longer (e.g., 1M tokens with YaRN) — the linear layers keep memory costs manageable while the attention layers maintain quality.
See the knowledge entries on Gated DeltaNet and Causal Language Model for deeper background.
Causal Language Model
architecture_concepts
Also called autoregressive or decoder-only language modeling: the model predicts the next token in a sequence using only the tokens that precede it. Generation runs strictly left-to-right and never conditions on future tokens.
Core principle
Given a sequence of tokens x1, x2, ..., x_{t-1}, the model estimates the conditional probability of the next token:
P(x_t | x1, x2, ..., x_{t-1})
The full sequence probability is the product of these conditional probabilities (chain rule of probability):
P(x1, ..., x_T) = product over t of P(x_t | x1, ..., x_{t-1})
Causal masking
This unidirectionality is enforced architecturally through a triangular attention mask. In standard self-attention, every token can attend to every other token. Causal masking restricts this so that position t can only attend to positions 1 through t. The mask is an upper-triangular matrix filled with -infinity values, applied to attention scores before softmax.
Key distinctions
Causal (decoder-only) LM: lower-triangular mask, sees only the past, generates left-to-right.
Masked LM: full bidirectional attention, sees the entire input at once, used for understanding tasks.
Prefix LM: hybrid — bidirectional attention on a prefix portion, autoregressive generation on the remainder.
References
Vaswani et al. (2017), "Attention Is All You Need", arXiv:1706.03762
Hugging Face Transformers Documentation: Causal language modeling
AI Wiki: https://aiwiki.ai/wiki/causal_language_model
Related FAQs (1)
This architecture combines linear attention layers with standard softmax attention layers in a fixed ratio (e.g., 3:1).
Gated DeltaNet is a linear attention mechanism that combines data-dependent gating (alpha_t) for rapid memory flushing with the delta update rule (beta_t) for targeted key-value updates. This gives O(1) inference memory per step.
Gated Attention is standard softmax attention that excels at precise retrieval.
The 3:1 ratio means 75% of layers use the efficient linear mechanism while 25% provide the retrieval strength of full attention. The Gated DeltaNet layers handle long-range modeling efficiently, while the Gated Attention layers provide precise in-context retrieval.
This hybrid design allows models to natively support long contexts (e.g., 262K) and scale to even longer (e.g., 1M tokens with YaRN) — the linear layers keep memory costs manageable while the attention layers maintain quality.
See the knowledge entries on Gated DeltaNet and Causal Language Model for deeper background.