LLM Knowledge Base — Benchmarks, Tokens & Context Explained

DeepSeek-V4-Pro-Max (max reasoning effort mode)

architecture_concepts

DeepSeek-V4-Pro-Max is the maximum reasoning effort mode of DeepSeek-V4-Pro, announced with the DeepSeek-V4 preview series. DeepSeek-V4 models expose configurable reasoning effort; at the Max setting the model spends the largest thinking budget, which significantly advances the knowledge capabilities of open-source models - top-tier performance in coding benchmarks and a significantly narrowed gap with leading closed-source models on reasoning and agentic tasks, establishing it (per DeepSeek) as the best open-source model available at release. The same mechanism gives DeepSeek-V4-Flash-Max comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller scale still trails on pure knowledge tasks and the most complex agentic workflows. Related: gpt-oss models expose a similar configurable-effort control (low/medium/high).

Related FAQs (1)

Heavily Compressed Attention (HCA)

architecture_concepts

Heavily Compressed Attention (HCA) is one half of DeepSeek-V4's hybrid long-context attention, paired with Compressed Sparse Attention (CSA). While CSA sparsifies which positions a query attends to, HCA attacks the size of what each attended position contributes: key-value representations are compressed much more aggressively than in DeepSeek-V3.2's MLA-style latent cache, shrinking per-token KV storage for the layers where full positional fidelity matters less. Interleaving CSA and HCA layers is what lets DeepSeek-V4-Pro run 1M-token contexts with only 27% of single-token inference FLOPs and 10% of the KV cache that DeepSeek-V3.2 would need - sparsity cuts the FLOPs, heavy compression cuts the cache.

Related FAQs (1)

Gated Multi-head Latent Attention (Gated MLA)

architecture_concepts

Gated Multi-head Latent Attention (Gated MLA) is Kimi K3's precision-attention layer type: 24 of its 93 attention layers are Gated MLA, complementing the 69 Kimi Delta Attention (KDA) layers that provide cheap O(n) sequence processing. It extends Multi-head Latent Attention (MLA - key/values compressed into a small latent vector so KV-cache stays tiny) with a gating mechanism on the attention output, letting the model dynamically amplify or suppress each head's contribution. In the K3 hybrid layout, Gated MLA layers supply the precise global retrieval and exact positional binding that linear-attention state compression can lose at very long range, while KDA keeps per-token cost low across the model's 1M-token context.

Related FAQs (1)

Attention Residuals (AttnRes)

architecture_concepts

Attention Residuals (AttnRes) is one of the two architectural components Moonshot AI built Kimi K3 on (with Kimi Delta Attention). Standard transformers add each layer's block output through a residual stream that bypasses attention; AttnRes instead feeds information from the attention pathway itself back through the residual stream - the residual connection carries attention's contribution rather than only the MLP/block output. Combined with delta-rule linear attention (KDA) and a Stable LatentMoE backbone (16 of 896 experts active), this is part of what yields Kimi K3's ~2.5x improvement in overall scaling efficiency over Kimi K2.

Related FAQs (1)

Kimi Delta Attention (KDA)

architecture_concepts

Kimi Delta Attention (KDA) is the linear-attention mechanism Moonshot AI built Kimi K3 on (69 of its 93 attention layers). It belongs to the delta-rule family of linear attention: rather than storing every key-value pair like softmax attention, KDA maintains a compressed state that is updated with the delta rule - new information overwrites the parts of the state it can already predict, so the state carries only what is genuinely new. This gives O(1) per-token state size and O(n) sequence cost instead of quadratic attention, making million-token contexts tractable, while the remaining 24 Gated MLA layers retain precise global retrieval. Kimi K3 pairs KDA with Attention Residuals (AttnRes) and a Stable LatentMoE backbone (16 of 896 experts active), reaching ~2.5x better scaling efficiency than Kimi K2.

Related FAQs (1)

JustRL II (critic-based RL)

training_methods

JustRL II is OpenBMB's critic-based reinforcement-learning algorithm, used in the RL stage of MiniCPM5-2B post-training and described in "JustRL II: Scaling Small LLMs to 128K Reasoning with a Critic". Instead of critic-free policy-gradient methods (GRPO-family), it trains with a learned critic that provides value estimates, which substantially improves training stability and enables long-context reasoning at 128K for small (2B-class) models. In the MiniCPM5-2B pipeline the RL teachers for math, code, agentic tasks and writing are trained with JustRL II before being merged back into the release model via On-Policy Distillation; combined RL+OPD lifted reasoning/general benchmarks by an average +10.96 points and agentic capabilities by +6.96 points.

Related FAQs (1)

UltraData Tiered Data Management

training_methods

UltraData Tiered Data Management (arXiv 2602.09003, OpenBMB) is the full-stack data practice behind the MiniCPM5 series: training data is organized into quality tiers and managed tier-by-tier across every stage of the pipeline rather than as one undifferentiated corpus.

  1. Base training uses high-quality web pre-training data (Ultra-FineWeb, Ultra-FineWeb-L3, UltraX) with stable-training and decay-training phases for core language capability.
  2. Tiered code data - UltraData-Code L0-L3 - matches code difficulty to training stage, driving large coding-capability gains.
  3. Mid-training strengthens target capabilities and adapts to the target data distribution.
  4. Post-training reuses the tiered philosophy: deep-thinking SFT (400B tokens, UltraData-SFT-2605), agent SFT (500K samples, UltraData-SFT-Agent-2609) and RL (UltraData-RL-2609, 80K+ samples).

All tiers are open-sourced in the UltraData family, making the model's full data pipeline reproducible.

Related FAQs (1)

iHC (identity Hyper-Connections)

architecture_concepts

iHC (identity Hyper-Connections) is the residual-stream design of Tencent's Hy4 preview: a simplified, identity form of Hyper-Connections (arXiv 2409.19606). Instead of a single residual pathway, the network keeps 4 parallel residual streams and each layer's output is added into multiple streams, expanding inter-layer information flow so layers can interact through more than one path - without the learned mixing weights of full Hyper-Connections. It is the identity-matrix member of the same family as Manifold-Constrained Hyper-Connections (mHC) (DeepSeek-V4, GLM-5.3-Flash) and DeepSeek-V4's learned Hyper-Connections: where mHC constrains the mixing manifold and DeepSeek-V4 learns the mix, iHC fixes the mixing to identity, trading a little adaptivity for stability and simplicity.

Related FAQs (1)

IndexCache (cross-layer sparse index reuse)

architecture_concepts

IndexCache (arXiv 2603.12201) is a long-context efficiency technique used in Tencent's Hy4 preview: with sparse attention, each layer normally re-runs a lightweight indexer to find which key-value positions each query should attend to. IndexCache instead computes the sparse index once and reuses it across layers, eliminating redundant index computation for the same positions. Paired with Gated DSA, it makes 1M-token inference cheap: the indexer (32 heads, 128-dim, top-k 2048 in Hy4) does its work once and the cross-layer cache feeds every subsequent sparse-attention layer.

Related FAQs (1)

Gated DeepSeek Sparse Attention (Gated DSA)

architecture_concepts

Gated DeepSeek Sparse Attention (Gated DSA) is the attention module of Tencent's Hy4 preview (Hy series), adapted from DeepSeek's Sparse Attention (DSA, arXiv 2512.02556) with a gating mechanism, itself inspired by attention work in DeepSeek and GLM models. A lightweight indexer produces a sparse candidate set per query and attention runs only over those positions, cutting compute at very long context; Hy4 pairs it with IndexCache, which reuses the sparse index across layers. Hy4-preview configuration: 64 attention heads, query compression dimension 2048, key-value compression dimension 512, indexer with 32 heads / 128 head dimension and top-k 2048. Related but distinct: DeepSeek-V4's Compressed Sparse Attention (CSA) and DeepSeek-V4.1-Flash's CSA2 are the DeepSeek family's own sparsified attention schemes - DSA is the earlier scheme Hy4 gates and borrows.

Related FAQs (1)

NVFP4 quantization (NVIDIA)

architecture_concepts

NVFP4 is NVIDIA's 4-bit floating-point format (two FP4 values per byte with a shared 8-bit scale per block of 16, per the OCP MX spec) and the centerpiece of the Nemotron 3 family's quantization-aware pre-training recipe: the majority of linear layers - weights, activations, and gradients - train directly in NVFP4, while select stability-critical layers (latent projections, MTP layers, QKV/attention projections, embeddings) are kept in BF16 or MXFP8. Training in the deployed numeric format lets frontier-scale models (e.g. Nemotron-3-Ultra-550B-A55B, pre-trained on ~20T tokens) serve at a fraction of the memory and compute cost of BF16 with minimal accuracy loss - unlike post-hoc quantization, the weights never live at full precision. NVFP4 is also natively accelerated on NVIDIA Blackwell GPUs.

Related FAQs (1)

Latent Mixture-of-Experts (LatentMoE)

architecture_concepts

Latent Mixture-of-Experts (LatentMoE), used in NVIDIA's Nemotron 3 family, routes and computes MoE experts in a smaller latent dimension instead of the full model width: tokens are projected down into the latent space, expert routing and the experts' computation happen there, and results project back. Shrinking the routing/compute dimension improves accuracy per byte - expert capacity scales better than the parameter footprint suggests - and pairs naturally with NVIDIA's quantization-aware NVFP4 pre-training, where latent projections stay in BF16/MXFP8 while most other linear layers run NVFP4. The Nemotron 3 Ultra/Super stacks interleave Mamba-2 layers, MoE layers, and select attention layers under this latent routing scheme.

Related FAQs (1)

MoonViT

architecture_concepts

MoonViT is Moonshot AI's vision encoder family used in Kimi multimodal models (K2.5/K2.6/K2.7). It is a ViT-style encoder adapted for native-resolution, variable-aspect image and video understanding in LLM front-ends - in Kimi K2.7 Code it contributes ~400M parameters, pairing with the 1T-parameter MoE language model (MLA attention, 384 experts) to give the coding agent the ability to read screenshots, UI captures and documents. MoonViT follows the Kimi line's philosophy of running vision natively at the model's context length rather than through a separate fixed-resolution pipeline.

Related FAQs (1)

End-to-end self-improvement (Ornith-1.5)

training_methods

End-to-end self-improvement (Ornith-1.5) extends Ornith-1.0's scaffold-rollout co-optimization by bringing task generation itself into the RL loop: the system jointly optimizes (1) generating new training tasks, (2) constructing the scaffolds/harnesses for them, and (3) the solution rollouts, instead of relying on a fixed set of human-curated tasks and manually designed harnesses. Ornith-1.5 continuously produces fresh tasks, discovers which strategies solve them, and improves the policy through reinforcement learning - so the training curriculum co-evolves with the model. The approach let a ~3B-activated MoE (35B-A3B) outperform similar-sized peers (Qwen 3.6-35B) across coding and agentic benchmarks. Reward design details for tasks, harnesses and rollouts are documented on the Ornith blog.

Related FAQs (1)

DFlash block-diffusion drafter

architecture_concepts

DFlash (arXiv 2602.06036) is a block-diffusion model used as a lightweight speculative-decoding drafter: instead of predicting one token at a time, it proposes entire blocks of tokens (e.g. 16 tokens) in a single forward pass, which the main model then verifies in parallel - accepting correct tokens and correcting wrong ones. Because verification is parallel, output quality is identical to standard autoregressive decoding while generation is significantly faster. Muse-Glimmer-30B ships with a DFlash drafter (5 draft layers, 16-token blocks, sliding-window attention with 2048 window on all layers, 32 Q / 8 KV GQA heads, sequence length 131,072, hidden features drawn uniformly from target layers {1, 13, 25, 37, 49} of 52), delivering measured speedups of 3.1x on an NVIDIA RTX 5090, 1.8x on Apple M5 Max and 1.5x on M4 Max with quantized drafters available to cut memory overhead. Distinct from the DeepSeek-V4 "DFlash attention" sparse-attention mechanism despite the shared name.

Related FAQs (1)

Perception Encoder (Meta)

architecture_concepts

Meta's Perception Encoder (PE, arXiv 2504.13181) is a vision-only ViT family designed as the perception backbone for multimodal LLMs - the largest variants reaching ViT-G scale. Unlike CLIP-style encoders trained with contrastive text alignment, PE is trained with core vision losses (self-supervised and weakly-supervised distillation from large vision-only teachers) followed by aligner modules that attach language - which resolves the tension between pure visual tasks (depth, segmentation) and vision-language tasks (VQA) that plagues dual-encoders. In Muse-Glimmer-30B, the perception encoder is a ~1.8B-parameter ViT-G/14 (50 layers, width 1536, patch size 14) that tokenizes interleaved text and images (up to 4,096 visual tokens per image), letting the agent interpret screenshots, charts and documents alongside conversation.

Related FAQs (1)

Configurable reasoning effort (gpt-oss)

architecture_concepts

Configurable reasoning effort in the gpt-oss models (120b and 20b) lets the caller choose how much thinking the model does per query - low, medium, or high - via a simple system prompt (Reasoning: low / medium / high) or the reasoning_effort API parameter. Lower effort answers faster with less chain-of-thought; higher effort spends more tokens reasoning before answering. The full chain-of-thought is always exposed to the developer (for debugging and trust; OpenAI notes it is not intended to be shown to end users), and the model is trained to co-exist with tool calls inside that reasoning stream. This native effort control is one of gpt-oss's signature features and the reason latency-sensitive agentic deployments pick low/medium while hard math picks high.

Related FAQs (1)

MXFP4 quantization

architecture_concepts

MXFP4 is a 4-bit floating-point microscaling format (per the OCP MX standard) that OpenAI used to post-train the gpt-oss models: the MoE (expert) weights are quantized to MXFP4 while attention and shared weights stay higher precision, packing FP4 values with per-block scaling factors. gpt-oss-120b ships this quantization natively - so the 117B-parameter model fits and runs on a single 80GB GPU (NVIDIA H100 or AMD MI300X), and gpt-oss-20b runs within 16GB of memory. All of OpenAI's published gpt-oss evals were performed with the same MXFP4 quantization, so the open weights' quality is measured exactly as distributed.

Related FAQs (1)

MiniMax Sparse Attention (MSA)

architecture_concepts

MiniMax Sparse Attention (MSA), introduced with MiniMax-M3, is a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces both the attention compute and the memory footprint while preserving model quality. On MiniMax-M3 at 1M context it delivers 9x prefill and 15x decode speedups versus the M2 generation and cuts per-token compute to 1/20, making native 1M-token multimodal contexts practical. The open-source implementation lives at github.com/MiniMax-AI/MSA.

Related FAQs (1)

Language World Model (LWM)

architecture_concepts

A Language World Model (LWM) is a language model trained to simulate environments - predicting how a world (terminal, web page, Android device, OS, tool API) responds to actions, not just to answer prompts. Qwen's AgentWorld line builds LWMs by injecting environment knowledge during continual pre-training, then teaching next-state-prediction reasoning (SFT) and optimizing simulation fidelity with RL (GSPO), so environment modeling is native rather than a post-hoc adaptation on a general LLM. Because the model can imagine environment transitions, it enables controllable perturbations, fictional-world construction (synthetic environments that train agents better than real ones), and zero-shot transfer to out-of-distribution environments - LWM RL warm-up on single-turn trajectories transfers to multi-turn tool-calling tasks across 7 benchmarks, including 3 entirely out-of-domain.

Related FAQs (1)

Manifold-Constrained Hyper-Connections (mHC)

architecture_concepts

Manifold-Constrained Hyper-Connections (mHC), used in the DeepSeek-V4 series, upgrade the plain residual stream by learning multi-channel layer-to-layer mixes under a manifold constraint that keeps the mixing weights well-conditioned. Each layer reads from and writes to several parallel connection channels with learned coefficients; the manifold constraint prevents those coefficients from drifting into degenerate configurations, so signal propagation stays stable across very deep stacks while preserving expressivity. mHC generalizes the Hyper-Connections idea (multi-channel learned residuals, e.g. hc_mult=4 channels in V4 configs, with Sinkhorn-balanced mixing) and is credited by DeepSeek for V4's stable training of its deep hybrid-attention MoE stack.

Related FAQs (1)

Compressed Sparse Attention (CSA)

architecture_concepts

Compressed Sparse Attention (CSA), introduced with the DeepSeek-V4 series, is one half of the family's hybrid long-context attention (paired with Heavily Compressed Attention, HCA). CSA sparsifies which positions a query attends to while compressing the retained key/value representations, so both the search over the context and the cache that backs it shrink. DeepSeek reports that at 1M-token context the CSA/HCA hybrid lets DeepSeek-V4-Pro run with only 27% of single-token inference FLOPs and 10% of the KV cache versus DeepSeek-V3.2 - the enabling factor behind the series' native 1M context.

Related FAQs (1)

Muon optimizer

training_methods

Muon (MomentUm Orthogonalized by Newton-Schulz) is a optimizer for 2D weight matrices that orthogonalizes the momentum update via a Newton-Schulz iteration before applying it, so updates spread energy evenly across weight directions instead of over-amplifying the dominant singular components that plain Adam-style updates favor. Qwen3.8-Flash-Next applies Muon to specific weight categories while using AdamW for the rest, guided by refitted scaling laws, and eliminates traditional batch-size warmups by starting directly at the target batch size - reducing total optimizer steps and safely allowing larger learning rates. Muon-type optimizers have shown stronger compute efficiency than AdamW at LLM scale (e.g. in the Moonshot/Kimi K2 training it replaced most AdamW usage).

Related FAQs (1)

N-gram Embedding (Qwen3.8)

architecture_concepts

N-gram Embedding, introduced in Qwen3.8-Flash-Next, scales parameters through embedding tables instead of experts. A table of 20,000,000 bigram/trigram embeddings is added at layer 2, contributing 51B parameters - four times the model's 6B activated MoE path - while adding almost no compute: lookups replace matrix-heavy FFN work, and embedding rows are trivially offloadable to CPU/cheap memory. Guided by the observation that embeddings are the most memory-efficient parameter axis, this lets a "6B-active" model carry 125B+ parameters of n-gram and MoE capacity in a way that fits memory-constrained accelerators without sacrificing quality.

Related FAQs (1)

Gated Residual (Qwen3.8)

architecture_concepts

Gated Residual, introduced in Qwen3.8-Flash-Next, replaces the plain normalized residual stream with a learned, data-dependent gating scheme: each layer's information passes through widened residual streams whose flow is modulated by an element-wise read gate (choosing what each branch takes in, per dimension) and per-branch scalar write gates (choosing how much each branch contributes back). Qwen3.8-Flash-Next uses 4 branches with a bottleneck rank of 320. Compared with standard pre-norm residuals, this gives finer-grained expressiveness across layers while preserving the training stability deep stacks need - and adds only low inference overhead.

Related FAQs (1)

Qwen Sparse Attention (QSA)

architecture_concepts

Qwen Sparse Attention (QSA), introduced in Qwen3.8-Flash-Next, sparsifies attention at the micro-block level rather than per token. A lightweight MQA indexer (4 query heads + 1 shared key head, head dim 128) scores relevance, but attention keeps a budget of whole 512-token blocks (up to 2048 tokens) - so every kept token's neighbors come along, preserving local structure that token-level top-k selection destroys. QSA replaces the Gated Attention layer in the Qwen3.8 hybrid stack (paired with Gated DeltaNet) and runs with 24 Q / 2 KV heads of head dim 256. The design cuts long-context latency significantly, which matters most for agentic workloads dominated by long tool-call histories.

Related FAQs (1)

IndexShare (GLM-5.2)

architecture_concepts

IndexShare (GLM-5.2, arXiv 2603.12201) makes long-context sparse attention cheaper by sharing one indexer across groups of attention layers. A sparse-attention indexer normally selects the top-k relevant KV positions per query for each layer independently; IndexShare reuses the same indexer for every four attention layers (in GLM-5.2: 21 full + 57 shared indexers over 78 layers, index_topk_freq: 4), so 75% of layers skip their own index computation entirely. Combined with top-2048 index selection (32 index heads, dim 128) and MTP-aware index reuse (index_share_for_mtp_iteration), this reduces per-token FLOPs by 2.9x at 1M-token context - the key enabler of GLM-5.2's "solid 1M context" claim.

Related FAQs (1)

DSpark (DeepSeek-V4)

architecture_concepts

DSpark is a component of the DeepSeek-V4 forward path (named in the DeepSeek-V4-Flash-Vision-Exp reference inference). Its config keys mark a parallel token stream: dspark_block_size: 5, dspark_markov_rank: 256, dspark_target_layer_ids: [40, 41, 42] (the final three of 43 layers), and a dedicated dspark_noise_token_id. Together with num_nextn_predict_layers: 3 (multi-token prediction) this forms DeepSeek-V4's speculative/parallel generation machinery: blocks of up to 5 tokens are processed through a Markov-rank-256 state on the last layers with a noise token controlling stochastic exploration. The public reference implementation documents its forward path; DeepSeek has not yet published a full paper on DSpark.

Related FAQs (1)

Hyper-Connections (DeepSeek-V4)

architecture_concepts

Hyper-Connections (DeepSeek-V4) replace the plain residual stream with a learned mixing network: each layer's input is a weighted combination of hc_mult=4 parallel connection channels, and layer outputs are mixed back into those channels with weights produced by a tiny network (regularized by hc_eps and balanced via hc_sinkhorn_iters=20 in DeepSeek-V4's config). Instead of a fixed x + f(x) residual, the model learns how much of each channel each layer should read from and write to. This stabilizes training of very deep networks, lets different layers specialize (e.g. the 43-layer DeepSeek-V4-Flash stack), and improves gradient flow compared with standard and dense residual variants (the idea generalizes Hyper-Connections from bytedecoder/Hyper-Connections, 2024).

Related FAQs (1)

DFlash attention (DeepSeek-V4)

architecture_concepts

DFlash attention is the DeepSeek-V4 family's attention mechanism (used in DeepSeek-V4-Flash and the Flash-Vision-Exp multimodal variant). It combines latent-compressed queries and outputs (q/o LoRA ranks of 1024, outputs factored into 8 groups) with a learned sparse index: a separate set of 64 index heads (dim 128) scores every position and keeps only the top-512 entries per query token for the full attention computation, while 3 hash layers and a 128-token sliding window provide additional locality. Full attention runs with 64 heads of head dim 512 but only 1 KV head. The result is near-linear cost on very long sequences while preserving retrieval quality, which together with YaRN scaling supports DeepSeek-V4's 1M-token context.

Related FAQs (1)

Scaffold-rollout co-optimization (self-improving RL)

training_methods

Scaffold-rollout co-optimization is Ornith 1.0's self-improving training framework: instead of RL over solution rollouts alone, the model learns to generate the scaffold (task setup, tools, environment, search strategy) and the rollouts (solution trajectories driven by that scaffold) jointly. Because scaffold quality determines rollout quality, optimizing both lets the model discover better search trajectories and produce higher-quality agentic coding solutions than optimizing rollouts against a fixed scaffold. Ornith-1.0 applies this RL post-training on top of Gemma 4 and Qwen 3.5 bases, reaching state-of-the-art open-source results on Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.

Related FAQs (1)

Hybrid thinking modes (Qwen3)

architecture_concepts

Hybrid thinking modes (Qwen3's signature feature) pack a thinking mode and a non-thinking mode into a single model with seamless runtime switching. Thinking mode targets complex logical reasoning, math, and coding (chain-of-thought before the answer); non-thinking mode gives efficient general-purpose dialogue. The mode is chosen per request via the chat template (enable_thinking=True/False) or the prompt (/no_think), and Qwen3's post-training includes a dedicated thinking-mode-fusion stage that aligns both modes inside one checkpoint. This removes the need to deploy separate reasoning and chat models while keeping near-parity with specialized reasoning models in thinking mode.

Related FAQs (1)

Continual pretraining

training_methods

Continual pretraining (CPT) takes an already-pretrained base model and trains it further on a new corpus - usually to add capabilities the original pretraining lacked - before fine-tuning. Kimi K2.5 is built through continual pretraining on ~15 trillion mixed visual and text tokens atop Kimi-K2-Base, adding native multimodality and agentic abilities to the text-only base. CPT differs from ordinary pretraining in that it starts from existing weights (preserving general language ability) rather than random initialization, and differs from SFT/RL in scale and objective: it is large-corpus self-supervised training, not instruction following.

Related FAQs (1)

Agent Swarm (Kimi K2.5)

architecture_concepts

Agent Swarm is Kimi K2.5's self-directed, coordinated execution scheme: instead of scaling a single agent loop, the model decomposes a complex task into parallel sub-tasks and dynamically instantiates domain-specific agents to execute them in a swarm-like fashion. Each specialized agent handles its sub-task (e.g. visual data processing, coding, tool use), and execution is coordinated back into one coherent result. This shifts the scaling axis from single-agent chain length to parallel agent populations, improving throughput on large multi-step tasks and grounding each sub-task in an agent tuned for that domain.

Related FAQs (1)

Multi-head Latent Attention (MLA)

architecture_concepts

Multi-head Latent Attention (MLA), introduced by DeepSeek-V2, compresses the key-value cache into a low-rank latent vector. Instead of caching full K and V per head per token, each token's KV information is projected down to a shared latent (e.g. kv-lora rank 512 in Kimi K2.5 and DeepSeek-V3) and up-projected on the fly during attention. This shrinks the KV cache by an order of magnitude at long context lengths, and queries are also optionally compressed (q-lora rank 1536). MLA preserves quality close to standard multi-head attention while making million-token-scale serving and trillion-parameter MoE models practical.

Kimi K2/K2.5 uses MLA as its attention mechanism (64 heads, 128 nope + 64 rope head dims, v_head_dim 128), inheriting the DeepseekV3 backbone design.

Related FAQs (1)

Shared expert (MoE)

architecture_concepts

A shared expert in a Mixture-of-Experts model is an always-on expert FFN that processes every token alongside the routed experts, while the router activates only k of the remaining N experts per token (e.g. Gemma 4 26B-A4B: 1 shared + 128 routed, top-8 active; DeepSeek-V2/V3 use the same pattern). The shared expert captures common knowledge and generic computation that every token needs, letting routed experts specialize more cleanly and reducing redundant capacity across experts.

Related FAQs (1)

Cross-layer KV cache sharing

architecture_concepts

Cross-layer KV cache sharing lets consecutive Transformer layers reuse the same keys/values instead of each layer computing and storing its own KV cache - the attention outputs of one layer's projection serve its neighbors. In Gemma 4 E-models, 18 of the layers share KV entries (config num_kv_shared_layers: 18), directly shrinking the KV cache that dominates long-context memory on edge devices. Combined with sliding-window attention, it keeps the E4B's 128K context practical on phones and laptops.

Related FAQs (1)

Encoder-free multimodal architecture

architecture_concepts

Encoder-free multimodal architecture (Gemma 4 12B "Unified") removes the dedicated vision/audio encoders that most multimodal LLMs bolt onto a text backbone. Instead, raw image patches and audio waveforms are projected directly into the LLM's embedding space through lightweight linear layers, so every modality flows into a single decoder-only transformer. This reduces multimodal latency, simplifies the pipeline (no separate ViT/AST to download or run), and allows the entire model to be fine-tuned in one pass - at the cost of putting all cross-modal work on the LLM itself.

Related FAQs (1)

Asynchronous RL

training_methods

Asynchronous RL decouples rollout generation from policy optimization: one GPU pool runs inference producing rollouts (episodes/traces) while a separate pool performs policy gradient updates, and the learner pulls fresh weights in flight instead of pausing generation for a synchronized update. This maximizes utilization of both pools and scales RL to many environments - IBM runs Granite 4.2's GRPO this way (NeMo RL, environments on NeMo Gym), and frameworks like AReaL/AsyncFlow build on the same principle.

Related FAQs (1)

Verifiable rewards (RLVR)

training_methods

Verifiable rewards (RLVR) score a model's response by checking it against ground truth: unit tests for code, exact/checked answers for math, schema validation for structured output, tool-call success flags. Because the reward is objective, RL can scale to millions of prompts without human labelers. Prompts that cannot be verified automatically are handled by a reward model instead. Granite 4.2's RL environments mostly provide verifiable rewards across math, code, science, instruction following, tool use and structured output; DeepSeek-R1 pioneered the recipe at scale.

Related FAQs (1)

Group Relative Policy Optimization (GRPO)

training_methods

Group Relative Policy Optimization (GRPO) is a reinforcement-learning algorithm (popularized by DeepSeek-R1) that replaces the value/critic network of PPO with group-relative advantage estimates: for each prompt, sample a group of responses from the current policy and compute each response's advantage by normalizing its reward against the group's mean and std. This removes the critic's memory cost while keeping stable policy updates via importance-sampling clipping and a KL penalty. Used for post-training reasoning models - e.g. IBM Granite 4.2's multi-environment RL stage and DeepSeek-R1.

Related FAQs (1)

Proportional RoPE (p-RoPE)

architecture_concepts

Proportional RoPE (p-RoPE) is the positional-encoding scheme Gemma 4 applies on its global attention layers (rope_type "proportional", base frequency 1M with partial rotary factor 0.25, versus plain RoPE at 10k on sliding-window layers). It is part of Gemma 4's long-context memory optimization: together with unified Keys and Values (K=V) on global layers, p-RoPE lets the 256K-token context window fit in much less KV-cache memory while keeping precise global retrieval.

Related FAQs (1)

Per-Layer Embeddings (PLE)

architecture_concepts

Per-Layer Embeddings (PLE) give every decoder layer its own small per-token embedding table instead of enlarging the transformer blocks. The tables are large but are only used for cheap lookups, and each layer reads its own embedding and adds it to that layer's input. PLE decouples the "effective" compute parameters from the total stored parameters: Gemma 4 E2B has 2.3B effective parameters (5.1B with embeddings) and E4B 4.5B effective (8B with embeddings), maximizing per-layer capacity for on-device deployment without slowing inference.

Related FAQs (1)

Synthetic reasoning traces (distillation)

training_methods

Synthetic reasoning traces are chain-of-thought reasoning paths generated by strong teacher models and distilled into a student's post-training data so the student learns to reason step-by-step. NVIDIA's Nemotron-3 post-training corpus includes synthetic reasoning traces from DeepSeek R1/R1-0528, Qwen3-235B-A22B, Nemotron 4 340B and Qwen2.5 models (the card notes "Improved using Qwen"); open reasoning datasets built this way include Bespoke-Stratos-17k and OpenCodeReasoning-2.

Related FAQs (1)

Nemotron Elastic (model compression framework)

architecture_concepts

Nemotron Elastic is NVIDIA's compression framework for deriving smaller, deployment-efficient models from a parent model while retaining its capabilities. NVIDIA-Nemotron-3-Nano-4B was produced from the 9B NVIDIA-Nemotron-Nano-9B-v2 via this framework (arXiv 2511.16664): the elastic procedure shrinks the hybrid Mamba2-Transformer architecture and continues training so the compressed 4B model keeps the parent's reasoning behavior and accuracy profile at a fraction of the size.

Related FAQs (1)

Mamba2

architecture_concepts

Mamba2 is a selective state-space model (SSM) layer - the successor to Mamba - that processes sequences with a linear-time recurrent scan instead of quadratic attention. Each token's hidden state is updated by input-dependent (selective) SSM dynamics, so the model compresses context into a fixed-size state: no KV cache grows with sequence length. Mamba2 simplifies Mamba's design (SSD state-space duality), enabling larger state dimensions and better hardware utilization. Hybrid stacks that interleave a few full-attention layers with many Mamba2 blocks (e.g. NVIDIA's Nemotron-H family, Zamba) get attention-like recall where needed at a fraction of the memory and compute.

Related FAQs (1)

Multi-Teacher On-Policy Distillation (MOPD)

training_methods

Multi-Teacher On-Policy Distillation (MOPD) consolidates several domain-specialized policies - each produced by large-scale RL on one capability domain (reasoning, coding, agentic tool use, ...) - into a single deployable model. The student samples on-policy trajectories and is supervised by the complementary teachers, so each teacher corrects the student only where it is strongest; the consolidated model keeps all teachers' strengths without running them all. Spark-X2.5 (XHToken/iFLYTEK) uses MOPD as its final post-training stage.

Related FAQs (1)

Hybrid SWA-full attention (3:1)

architecture_concepts

Hybrid SWA-full attention (3:1) is Spark-X2.5's efficiency-oriented attention layout: for every full-attention layer, three sliding-window attention (SWA) layers handle the remaining positions (21 SWA + 7 full layers in the 1.7B model, SWA window 512). Local SWA layers keep the KV cache and compute small, while the sparse full-attention layers preserve global retrieval - letting the model support a native 1M-token context window with far less overhead than an all-full-attention stack. RoPE is applied with theta 5M and partial rotary factor 0.25 on full layers (10k, full rotary on SWA layers).

Related FAQs (1)

Million-agent RL environments

training_methods

Million-agent RL environments refers to scaling post-training reinforcement learning across ~one million diverse agentic environments with progressively complex task distributions. Instead of RL on a handful of curated tasks, the policy is trained in parallel across a massive scaffolded environment pool (tool use, browsing, coding, embodied tasks), which forces robust generalization to unseen real-world agent settings. Qwen3.5 used this recipe, supported by asynchronous RL frameworks and massive-scale environment orchestration.

Related FAQs (1)

Early fusion (multimodal pre-training)

training_methods

Early fusion trains a single model on multimodal tokens from the start (text, image, video processed by one shared backbone), instead of bolting a vision adapter onto a text-pre-trained LLM (late fusion). Qwen3.5 is a natively multimodal family: pre-training uses early fusion on multimodal tokens, reaching near-100% multimodal training efficiency versus text-only training and outperforming the separately-trained Qwen3-VL models across reasoning, coding, agents and visual understanding.

Related FAQs (1)

On-policy distillation

training_methods

On-policy distillation distills a teacher's behavior into a student while the student generates the trajectories: the student samples actions/tokens from its own current policy, and the teacher provides per-step supervision (token-level corrections or rewards) on those on-policy samples. Unlike offline distillation on teacher-generated data, this keeps the training distribution matched to what the student will actually do at inference, correcting compounding errors early. LFM2.5-2.6B's post-training includes a multi-domain on-policy distillation stage, after per-domain teacher specialization, to transfer agent capability into the student.

Related FAQs (1)

Agentic reinforcement learning

training_methods

Agentic reinforcement learning trains a model inside the agentic harnesses and environments it will be deployed in, rather than on generic instruction data. The model is exposed to each harness's real tools, system prompts and multi-step interaction patterns, and is optimized end-to-end for task success (rewarded for correct tool calls and outcomes). Liquid AI used this in LFM2.5-2.6B's post-training so the model works reliably across popular agent environments - a general recipe for making small open models useful as on-device agents.

Related FAQs (1)

Speculative decoding

architecture_concepts

Speculative decoding accelerates autoregressive LLM inference by letting a small, fast drafter model propose several tokens which the large target model verifies in a single parallel forward pass. Accepted tokens are kept; on the first mismatch the target model's own next token is used, so outputs are identical to standard decoding. LFM2.5-8B-A1B ships a dedicated 328M-parameter drafter (LFM2.5-8B-A1B-DSpark) that pairs with it for ~2.5x faster decoding with identical outputs.

Related FAQs (1)

Grouped-Query Attention (GQA)

architecture_concepts

Grouped-Query Attention (GQA) is an attention variant where multiple query heads share a single key/value head, interpolating between multi-head attention (MHA, one KV head per query head) and multi-query attention (MQA, one KV head for all). GQA shrinks the KV cache and speeds decoding with little quality loss - e.g. LFM2.5-8B-A1B uses 32 query heads with 8 KV heads (4:1 sharing). It is standard in modern open LLMs (Llama 3, Gemma 3, Qwen3, LFM2.5).

Related FAQs (1)

Double-gated convolution blocks

architecture_concepts

Double-gated convolution blocks are the non-attention half of Liquid AI's LFM2/LFM2.5 hybrid architecture. Instead of attention, a token block passes through depthwise causal convolutions whose inputs and outputs are controlled by two learned gates (an input gate and an output gate), giving a fixed-size local receptive field (conv window L=3 in LFM2.5) with constant memory and compute per token. LFM2.5-8B-A1B uses 18 such conv blocks interleaved with only 6 GQA attention layers (ratio 3:1), which keeps the long-context KV-cache and attention FLOPs small - the key to its on-device throughput.

Related FAQs (1)

Mixture of Experts (MoE)

architecture_concepts

Mixture of Experts (MoE) replaces a Transformer's dense feed-forward blocks with many parallel "expert" FFNs plus a router that activates only a few experts per token. This decouples total parameter count from per-token compute: e.g. LFM2.5-8B-A1B has 8.3B total parameters but only 1.5B active per token (top-4 routing across 32 experts). MoE yields large capacity at a fraction of the inference cost of an equally-sized dense model, at the price of higher memory footprint (all experts must be resident), routing complexity, and load-balancing challenges during training.

Related FAQs (1)

SigLIP vision encoder

architecture_concepts

SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision encoder trained with a pairwise sigmoid loss instead of CLIP's softmax contrastive loss, which scales better with batch size. Gemma 3 uses a tailored, frozen 400M-parameter SigLIP ViT variant shared across its 4B, 12B and 27B models: images are resized to 896x896 and encoded into a fixed-size sequence of 256 soft tokens fed to the language model. A Pan-and-Scan (P&S) adaptive cropping scheme handles non-square and high-resolution images at inference time. (Gemma 3's encoder initializes from SigLIP v1, whereas Qwen3-VL builds on the newer SigLIP-2.)

Related FAQs (1)

5:1 Local-Global Attention Interleaving

architecture_concepts

5:1 Local-Global Attention Interleaving is Gemma 3's attention layer pattern: five local attention layers using sliding-window attention are followed by one global self-attention layer, starting with a local layer. Local layers attend within a bounded window (cheap, KV-cache-friendly), while every sixth layer attends over the full context, keeping long-range retrieval at 128K tokens while dramatically cutting compute and memory versus all-global stacks. The RoPE base frequency differs per layer type (1M on global, 10k on local).

Related FAQs (1)

QK-Norm

architecture_concepts

QK-Norm normalizes the queries and keys in a Transformer's attention (scaling them to unit norm, typically with an RMSNorm layernorm applied to the Q and K vectors) before the attention dot product. Introduced for stabilizing training (Dehghani et al., 2023; Wortsman et al., 2023), QK-Norm prevents attention-logit blowups without the need for attention soft-capping. Google replaced Gemma 2's soft-capping with QK-Norm in Gemma 3, keeping large 128K contexts trainable and stable while preserving downstream quality.

Related FAQs (1)

Qwen3-ViT

architecture_concepts

Qwen3-ViT is the vision encoder of the Qwen3-VL series. It initializes from SigLIP-2 (SO-400M for large LLMs; Large-300M for small LLMs) and continues training with dynamic input resolutions, 2D-RoPE and interpolated absolute position embeddings following CoMP. In Qwen3-VL-30B-A3B it is a 27-layer, 1152-hidden, 16-head encoder with 16px patches and 2x2 spatial merging; multi-level features from layers 8, 16 and 24 feed the DeepStack fusion. The Qwen3-VL Technical Report ablation shows Qwen3-ViT beats the raw SigLIP-2 encoder during CLIP-style pre-training.

Related FAQs (1)

Text-Timestamp Alignment

architecture_concepts

Text-Timestamp Alignment is Qwen3-VL's video temporal modeling technique. It moves beyond T-RoPE (timestamp rotary embedding used by earlier VLMs) to precise, timestamp-grounded event localization: textual descriptions are aligned to absolute timestamps in the video, enabling second-level indexing of long videos. Combined with Interleaved-MRoPE, it lets Qwen3-VL-30B-A3B handle hours-long video with full recall and precise temporal grounding.

Related FAQs (1)

SigLIP-2

architecture_concepts

SigLIP-2 is a family of sigmoid-loss image-text pre-trained vision encoders from Google DeepMind. It replaces the contrastive softmax loss of CLIP with a pairwise sigmoid loss that scales better and performs strongly at low resolutions. Qwen3-VL uses the SigLIP-2 architecture as its vision encoder: it continues training SigLIP2-SO-400M (large LLMs; SigLIP2-Large 300M for small LLMs) from official checkpoints with dynamic input resolutions, 2D-RoPE and interpolated absolute position embeddings following CoMP. The resulting encoder is called Qwen3-ViT.

Related FAQs (1)

DeepStack

architecture_concepts

DeepStack is a vision-language fusion mechanism used in Qwen3-VL: instead of feeding only the final ViT layer into the language model, DeepStack injects multiple intermediate ViT feature maps into several early LLM layers through specialized mergers. Fusing multi-level visual features captures fine-grained details and sharpens image-text alignment. In Qwen3-VL-30B-A3B the vision encoder is 27 layers deep and the DeepStack visual indexes are layers 8, 16 and 24. The Qwen3-VL Technical Report ablates DeepStack with a 15B-A2B model, showing clear gains from multi-level feature fusion.

Related FAQs (1)

Interleaved-MRoPE

architecture_concepts

Interleaved-MRoPE is Qwen3-VL's redesign of multimodal rotary position embedding. Qwen2-VL's original MRoPE partitioned the embedding dimensions into contiguous temporal (t), horizontal (h) and vertical (w) subspaces, each with distinct rotary frequencies - an imbalanced frequency spectrum that degrades long-video understanding. Interleaved-MRoPE instead interleaves the t, h and w components across the embedding dimension, giving full-frequency allocation over time, width and height. This robust positional embedding enhances long-horizon video reasoning and is configured in Qwen3-VL-30B-A3B with mrope_section [24, 20, 20].

Related FAQs (1)

GSPO (Group Sequence Policy Optimization)

training_methods

GSPO (Group Sequence Policy Optimization) is a reinforcement-learning algorithm introduced by the Qwen team for stable and efficient reasoning-model post-training. Instead of optimizing a per-token importance ratio as PPO-style methods do, GSPO computes a sequence-level importance ratio by averaging log-probabilities over whole sequences within a group, and then optimizes this ratio with a group-relative advantage.

Whole-sequence credit assignment makes the optimization landscape much smoother — particularly valuable for architectures like Qwen3-Next, where a hybrid attention mechanism (Gated DeltaNet + Gated Attention) combined with a high-sparsity MoE makes token-level credit assignment noisy. Qwen leveraged GSPO to post-train Qwen3-Next-80B-A3B-Thinking, which demonstrates outstanding performance on complex reasoning tasks.

Related FAQs (1)

Gated Attention

architecture_concepts

Gated Attention is the full-attention branch of Qwen3-Next's hybrid attention layout. Unlike standard Transformer attention, it applies a learnable output gate to each attention head, which stabilizes training and improves long-context quality while retaining exact token-to-token attention.

In Qwen3-Next-80B-A3B-Thinking, Gated Attention layers are interleaved with Gated DeltaNet layers in a 12 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)) pattern. Each Gated Attention block uses 16 query heads and 2 KV heads with a head dimension of 256 and a rotary position embedding dimension of 64.

The gated design keeps the precise-retrieval benefits of full attention while adding the training stability that hybrid linear-attention models need for robust pre-training and post-training (together with zero-centered and weight-decayed layernorm).

Related FAQs (1)

On-Policy Distillation (OPD)

training_methods

On-Policy Distillation (OPD) is a post-training paradigm in which a teacher model supervises a student model on the student's own rollouts (on-policy samples), rather than on a fixed corpus of teacher-generated sequences.

Because the training distribution matches the student's actual inference distribution, OPD addresses the train/inference mismatch of off-policy distillation and consolidates capabilities after SFT and RL. DeepSeek uses it as the final post-training stage: DeepSeek-V4.1-Flash follows the SFT -> RL -> OPD recipe, and DeepSeek-V3.x/V4 use the multi-teacher variant MOPD to merge domain-specialized RL teachers into one student.

OPD typically provides dense token-level supervision on student trajectories, making it more sample-efficient than reinforcement learning alone.

Related FAQs (1)

SWA Bounded Replay

architecture_concepts

SWA Bounded Replay is DeepSeek-V4.1-Flash's technique for handling sliding-window attention (SWA) KV states without persisting them.

Instead of writing every SWA KV state to SSD, the model keeps only a bounded window and reconstructs missing SWA KV states by replaying only the most recent n_win tokens when they are needed. This avoids SSD round-trips for SWA cache entirely and reduces the persistent KV cache footprint to roughly 1/8 of DeepSeek-V4-Flash.

It complements FP4 main KV caching and CSA2 in the model's overall KV-compression stack (890 bytes per token globally).

Related FAQs (1)

FP4 KV caching (E2M1)

architecture_concepts

FP4 main KV caching stores the attention KV cache in 4-bit floating point format (E2M1: 1 sign bit, 2 exponent bits, 1 mantissa bit) with one E4M3 scale factor per 16 channels, a block-scaled scheme in the spirit of NVFP4.

In DeepSeek-V4.1-Flash, FP4 main KV caching works together with CSA2 attention and SWA Bounded Replay to shrink the global KV cache to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1 - a decisive cost factor for long-context agentic workloads where KV cache memory dominates serving cost.

Related FAQs (1)

Compressed Sparse Attention 2 (CSA2)

architecture_concepts

Compressed Sparse Attention 2 (CSA2) is the attention scheme of DeepSeek-V4.1-Flash. Each attention layer is assigned one of three static modes:

  • Full - standard full attention; the first Full Mode layer also builds the candidate pool for later indexing layers
  • Reindex - recomputes attention indices from shared state
  • Reuse - reuses Top-K sparse-attention indices computed by another layer

Layers share main KV and indexer K across each other, and a Hierarchical Sparse Indexer restricts later indexing layers to the candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length.

Together with FP4 main KV caching, CSA2 reduces the global KV cache footprint to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash - while supporting 1M-token contexts.

Related FAQs (1)

Causal Encoder-Decoder (CED)

architecture_concepts

Causal Encoder-Decoder (CED) is a Transformer architecture introduced with DeepSeek-V4.1-Flash that splits the network into a causal encoder followed by a decoder. Instead of deriving the decoder's global KV cache from each decoder layer's own hidden states, CED projects it from the final encoder hidden states.

This makes prefill dramatically cheaper: DeepSeek-V4.1-Flash (552B backbone parameters) activates only 8B parameters per token during prefill and 16B during decode, which is particularly cost-efficient for input-heavy agentic workloads where long prompts are processed once and short answers are generated.

The 40-layer model uses a 20-layer causal encoder plus a 20-layer decoder. Combined with CSA2 attention and FP4 KV caching, CED is one of the four techniques that cut the model's global KV cache to 890 bytes per token (~1/4 of DeepSeek-V4-Flash).

Related FAQs (1)

KV cache

architecture_concepts

During autoregressive inference, the attention keys and values of already-processed tokens are stored in a running cache so that each new token does not require recomputing attention over the whole past sequence. This key-value (KV) cache grows with sequence length and with the number of layers that cache their keys and values.

Why it is a bottleneck

For agent tasks, balancing "performance, inference speed, and cache usage" is frequently identified as a key bottleneck limiting model performance: multi-turn tool-calling and long interaction histories make the cache — not just raw compute — the limiting resource.

How architectures control it

  • Full attention in every layer: the cache grows with every token of context — simple, but expensive at long sequence lengths.
  • Sliding-window attention layers: only tokens inside a fixed window are cached, bounding the per-layer cache regardless of total sequence length.
  • Hybrid attention layouts: a majority of sliding-window layers plus a minority of full-attention layers keeps the cache small even at very long context lengths, while the few full-attention layers preserve global context. This balance is a key reason such models maintain strong real-world deployment behavior (e.g., favorable time-to-first-token and overall inference efficiency compared with similarly sized models).

Related FAQs (1)

Native 1M-token context window

architecture_concepts

A context capability that is built into the model itself rather than added on afterwards through external retrieval or aggressive truncation. "Native" means the model was trained to operate at that sequence length directly.

What makes long context practical

  • Training: a dedicated long-context training stage on large token volumes (e.g., hundreds of billions of tokens), with sequence lengths extended step by step up to the target length — this teaches the model to actually use very long inputs instead of merely accepting them.
  • Architecture: a hybrid attention layout (e.g., three sliding-window layers per one full-attention layer) substantially reduces the computational and cache overhead typically associated with long-context models.

Deployment

Serving frameworks such as SGLang and vLLM expose the full window via --context-length 1048576. The full setting requires sufficient device memory and can be reduced if needed.

Related FAQs (1)

MOPD (Multi-Teacher On-Policy Distillation)

training_methods

A post-training paradigm that consolidates the capabilities of several domain-specialized reinforcement-learning teacher policies into one deployable student model.

Pipeline

  1. Supervised fine-tuning (SFT) on a curated corpus establishes instruction following, structured generation, and a stable policy initialization for reinforcement learning.
  2. Large-scale reinforcement learning runs across capability domains — language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following — producing a set of domain-specialized teacher policies.
  3. Multi-teacher distillation: the teachers are distilled into the student on the student's own rollouts, which keeps the training states aligned with real inference behavior and provides a dense per-token learning signal.

Effect

The combination of large-scale reinforcement learning and on-policy multi-teacher distillation enhances reasoning, coding, agentic, and instruction-following capabilities while collapsing several specialist checkpoints into a single model that behaves consistently between training and inference.

Related FAQs (1)

Sliding-window attention (SWA)

architecture_concepts

An attention variant in which each token attends only to a fixed window of neighboring tokens instead of the entire sequence. Because the window size is constant, the compute and memory of such a layer grow roughly linearly with sequence length, and its KV cache is bounded by the window size rather than the full context.

Role in hybrid layouts

Sliding-window layers are commonly combined with full-attention layers in a fixed ratio (e.g., three sliding-window layers per one full-attention layer). This pairing leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure.

  • The sliding-window layers process local patterns efficiently and keep per-layer cost and cache small even for very long inputs.
  • The full-attention layers preserve long-range accuracy and global information flow.

This improves practicality in real-world deployment scenarios where inference speed and cache usage matter — particularly for long-context and agentic workloads with long interaction histories.

Related FAQs (1)

Hybrid attention architecture

architecture_concepts

An attention layout that interleaves full-attention layers with sliding-window attention (SWA) layers instead of using full self-attention in every layer. A common stacking pattern is one full-attention layer per three sliding-window layers: the full-attention layers give tokens access to the entire context, while the sliding-window layers handle local context with far less computation and cache.

Motivation

Systematically integrating and optimizing mature attention technologies combines the strengths of both mechanisms while avoiding the limitations of relying on a single structure.

  • Full attention: sees the whole context, but compute scales quadratically and the KV cache grows unboundedly as sequences grow.
  • Sliding-window attention: cheap and cache-bounded, but blind beyond its fixed window.

Effects

  • An effective balance among performance, inference efficiency, and KV-cache size — especially relevant for agent workloads with long interaction histories.
  • Practical support for long native context windows (e.g., up to 1M tokens) while substantially reducing the computational overhead typically associated with long-context models.

Properties

  • Fixed layer ratio between sliding-window and full-attention layers (commonly 3:1).
  • KV cache dominated by the windowed layers; grows near-linearly with sequence length.
  • Works with any full-attention variant (softmax attention, gated attention) as the global layers.

Related FAQs (1)

Multi-Token Prediction (MTP)

architecture_concepts

A training objective for Large Language Models that extends the standard next-token prediction (NTP) paradigm by requiring the model to predict multiple future tokens simultaneously from each position in the input sequence, rather than only the next token.

How it works

Introduced by Gloeckle et al. (Meta, 2024), the objective attaches n independent output heads to a shared transformer trunk. At each position, the shared trunk produces a latent representation of the context, and each of the n heads predicts one of the following n tokens using a shared unembedding matrix.

  • NTP: at position t, predict token t+1.
  • MTP: at position t, predict tokens t+1, t+2, ..., t+n simultaneously.

If n=1, the objective degenerates to standard NTP.

Key benefits

  1. Higher sample efficiency: each training token provides n gradient signals instead of 1, extracting more learning signal per token.
  2. Better downstream performance: improves performance on code and reasoning tasks without increasing training time or memory overhead.
  3. Faster inference via self-speculative decoding: the extra heads can draft multiple candidate tokens per forward pass, which are then verified by the main model, enabling speculative decoding without a separate draft model.
  4. Planning and lookahead: forces the model to plan ahead, improving coherence over longer generations.

Training and implementation

  • The heads are typically pretrained alongside the main model (joint training), though post-hoc approaches exist (e.g., self-distillation from the frozen target LLM).
  • Each head is implemented as additional transformer layers that share the trunk's representations.
  • The loss is the sum of cross-entropy losses across all n heads.

Post-hoc variant

Recent work (Kirchenbauer et al., 2026) shows the modules can be trained post-hoc for LLMs without official support, using self-distillation from the target LLM's native NTP capability. The target LLM remains frozen; only the extra module is trained.

References

  • Gloeckle et al. (2024), "Better & Faster Large Language Models via Multi-token Prediction", arXiv:2404.12040
  • Kirchenbauer et al. (2026), "Multi-Token Prediction via Self-Distillation", arXiv:2602.06019

Related FAQs (1)

YaRN (Yet another RoPE extensioN)

architecture_concepts

A compute-efficient method to extend the context window of Large Language Models that were trained with Rotary Position Embeddings (RoPE). It requires 10x fewer tokens and 2.5x fewer training steps than previous methods, enabling models originally trained at shorter contexts to effectively utilize and extrapolate to much longer context lengths.

The extrapolation problem

When a model trained with RoPE at a maximum context length L is used with sequences longer than L, performance degrades sharply. The rotation angles for positions beyond L fall outside the range seen during training, causing the attention mechanism to produce unreliable scores.

Previous approaches

  1. Position Interpolation (PI): compresses the entire position range by a scale factor s = L'/L, mapping positions [0, L'] to [0, L]. This uniformly scales all RoPE dimensions, which impairs the model's ability to understand small, local relationships.
  2. NTK-aware scaling: adjusts the RoPE base frequency to stretch dimensions non-uniformly.

The method

The approach combines NTK-by-parts interpolation with an attention temperature scaling:

  • NTK-by-parts: instead of scaling every RoPE dimension equally, the interpolation pressure is spread across dimensions — high frequencies are scaled less (preserving local detail), while low frequencies are scaled more (enabling longer reach).
  • Attention temperature: a scaling factor 1/sqrt(t) is applied to the attention dot product to compensate for the changed distribution of attention scores, without modifying the attention code itself.

Static vs. dynamic scaling

  • Static: the scaling factor s is fixed at deployment time (e.g., s=4.0 for 1M context). It is constant regardless of input length, which may impact performance on shorter texts. All major open-source frameworks (vLLM, SGLang, TokenSpeed) implement this variant.
  • Dynamic: the scale factor is updated per forward pass based on actual sequence length, degrading gracefully. Not commonly used in production deployments.

References

  • Peng et al. (2023), "YaRN: Efficient Context Window Extension of Large Language Models", arXiv:2309.00071
  • Blog: https://amaarora.github.io/posts/2025-09-21-rope-context-extension.html

Related FAQs (1)

Rotary Position Embedding (RoPE)

architecture_concepts

A positional encoding scheme for Transformer models that encodes positional information by rotating query and key vectors in a 2D plane. The rotation angle depends on the token's position, so that nearby tokens have a small angular difference while distant tokens have a larger one. This naturally incorporates relative position dependency into the attention dot product without explicit position tokens.

Mathematical formulation

For a pair of features (x1, x2) at position m, a rotation by angle m*theta is applied:

(x1, x2) → (x1cos(mtheta) − x2sin(mtheta), x2cos(mtheta) + x1sin(mtheta))

The d features are organized as d/2 pairs, each rotated by a frequency that decreases exponentially. The attention dot product between positions m and n depends only on the relative position (m−n), not the absolute positions.

Key properties

  • Relative encoding: attention between two tokens depends on their distance, not absolute position.
  • No extra parameters: applied as a rotation, not learned — zero additional parameters.
  • Decoupled from attention: applied to Q and K before the attention computation.
  • Extrapolation challenge: trained at a maximum context length L, the encoding degrades when the sequence exceeds L because rotation angles fall outside the trained range.

Multimodal extension (mRoPE)

For vision-language models, the multimodal extension handles multiple dimensions (time, height, width for video). The frequency dimensions are split into sections (e.g., [11, 11, 10]) to allocate rotation capacity across spatial and temporal dimensions. mrope_interleaved mode interleaves the frequency pairs for better multimodal performance.

References

  • Su et al. (2021), "RoFormer: Enhanced Transformer with Rotary Position Embedding", arXiv:2104.09864
  • LabML implementation: https://nn.labml.ai/transformers/rope/index.html

Related FAQs (1)

Gated DeltaNet

architecture_concepts

A linear attention architecture that combines data-dependent gating with the delta update rule to achieve efficient, high-quality sequence modeling. Introduced by NVIDIA Research (NVlabs) as an improvement over Mamba2 and DeltaNet.

Background: linear attention

Standard Transformer attention scales quadratically with sequence length L. Linear attention replaces softmax with a kernelized dot product and reframes computation as a linear RNN with a matrix-valued state S_t, reducing inference memory to O(1) per step regardless of sequence length.

Delta rule

DeltaNet uses the delta update rule (Widrow et al., 1960): instead of accumulating all key-value pairs into the state (Hebbian learning), it selectively replaces the value associated with the current key. This is more surgical — it only modifies one key-value pair at a time while leaving others intact. The update is:

S_t = S_{t-1}(I − beta_t * k_t * k_t^T) + beta_t * v_t * k_t^T

This gives superior associative recall but lacks a mechanism to rapidly clear outdated information during context switches.

Gated delta rule

Unifying gating and the delta rule gives:

S_t = S_{t-1} * (alpha_t * (I − beta_t * k_t * k_t^T)) + beta_t * v_t * k_t^T

  • alpha_t (gate, 0 to 1): controls state decay. Setting alpha_t → 0 flushes the entire state (rapid forgetting); alpha_t → 1 reduces to the pure delta rule (selective update).
  • beta_t (writing strength, 0 to 1): controls how much of the new value is written.

This combination enables both bulk forgetting (via gating) and targeted updates (via delta rule), giving flexible memory control.

Hardware-efficient training

The recurrence is sequential, but a chunkwise parallel algorithm (using the WY representation and matmuls) enables GPU-efficient training with tensor cores. Within each chunk, the sequential dependencies are captured by a small C x C matrix inverse (C = chunk size).

Hybrid architectures

The layer is often combined with sliding-window attention (SWA) or Mamba2 layers in hybrid models. For example, a layout of 16 blocks of (3x gated-delta layers → 1x gated-attention layer) combines linear attention efficiency with standard attention's retrieval strength.

References

  • Yang et al. (2024), "Parallelizing Linear Transformers with the Delta Rule over Sequence Length", arXiv:2406.06484
  • NVlabs/GatedDeltaNet (ICLR 2025), arXiv:2412.06464
  • GitHub: https://github.com/NVlabs/GatedDeltaNet

Related FAQs (1)

Causal Language Model

architecture_concepts

Also called autoregressive or decoder-only language modeling: the model predicts the next token in a sequence using only the tokens that precede it. Generation runs strictly left-to-right and never conditions on future tokens.

Core principle

Given a sequence of tokens x1, x2, ..., x_{t-1}, the model estimates the conditional probability of the next token:

P(x_t | x1, x2, ..., x_{t-1})

The full sequence probability is the product of these conditional probabilities (chain rule of probability):

P(x1, ..., x_T) = product over t of P(x_t | x1, ..., x_{t-1})

Causal masking

This unidirectionality is enforced architecturally through a triangular attention mask. In standard self-attention, every token can attend to every other token. Causal masking restricts this so that position t can only attend to positions 1 through t. The mask is an upper-triangular matrix filled with -infinity values, applied to attention scores before softmax.

Key distinctions

  • Causal (decoder-only) LM: lower-triangular mask, sees only the past, generates left-to-right.
  • Masked LM: full bidirectional attention, sees the entire input at once, used for understanding tasks.
  • Prefix LM: hybrid — bidirectional attention on a prefix portion, autoregressive generation on the remainder.

References

  • Vaswani et al. (2017), "Attention Is All You Need", arXiv:1706.03762
  • Hugging Face Transformers Documentation: Causal language modeling
  • AI Wiki: https://aiwiki.ai/wiki/causal_language_model

Related FAQs (1)