LLM Knowledge Base — Benchmarks, Tokens & Context Explained

JustRL II (critic-based RL)

training_methods

JustRL II is OpenBMB's critic-based reinforcement-learning algorithm, used in the RL stage of MiniCPM5-2B post-training and described in "JustRL II: Scaling Small LLMs to 128K Reasoning with a Critic". Instead of critic-free policy-gradient methods (GRPO-family), it trains with a learned critic that provides value estimates, which substantially improves training stability and enables long-context reasoning at 128K for small (2B-class) models. In the MiniCPM5-2B pipeline the RL teachers for math, code, agentic tasks and writing are trained with JustRL II before being merged back into the release model via On-Policy Distillation; combined RL+OPD lifted reasoning/general benchmarks by an average +10.96 points and agentic capabilities by +6.96 points.

Related FAQs (1)

UltraData Tiered Data Management

training_methods

UltraData Tiered Data Management (arXiv 2602.09003, OpenBMB) is the full-stack data practice behind the MiniCPM5 series: training data is organized into quality tiers and managed tier-by-tier across every stage of the pipeline rather than as one undifferentiated corpus.

  1. Base training uses high-quality web pre-training data (Ultra-FineWeb, Ultra-FineWeb-L3, UltraX) with stable-training and decay-training phases for core language capability.
  2. Tiered code data - UltraData-Code L0-L3 - matches code difficulty to training stage, driving large coding-capability gains.
  3. Mid-training strengthens target capabilities and adapts to the target data distribution.
  4. Post-training reuses the tiered philosophy: deep-thinking SFT (400B tokens, UltraData-SFT-2605), agent SFT (500K samples, UltraData-SFT-Agent-2609) and RL (UltraData-RL-2609, 80K+ samples).

All tiers are open-sourced in the UltraData family, making the model's full data pipeline reproducible.

Related FAQs (1)

End-to-end self-improvement (Ornith-1.5)

training_methods

End-to-end self-improvement (Ornith-1.5) extends Ornith-1.0's scaffold-rollout co-optimization by bringing task generation itself into the RL loop: the system jointly optimizes (1) generating new training tasks, (2) constructing the scaffolds/harnesses for them, and (3) the solution rollouts, instead of relying on a fixed set of human-curated tasks and manually designed harnesses. Ornith-1.5 continuously produces fresh tasks, discovers which strategies solve them, and improves the policy through reinforcement learning - so the training curriculum co-evolves with the model. The approach let a ~3B-activated MoE (35B-A3B) outperform similar-sized peers (Qwen 3.6-35B) across coding and agentic benchmarks. Reward design details for tasks, harnesses and rollouts are documented on the Ornith blog.

Related FAQs (1)

Muon optimizer

training_methods

Muon (MomentUm Orthogonalized by Newton-Schulz) is a optimizer for 2D weight matrices that orthogonalizes the momentum update via a Newton-Schulz iteration before applying it, so updates spread energy evenly across weight directions instead of over-amplifying the dominant singular components that plain Adam-style updates favor. Qwen3.8-Flash-Next applies Muon to specific weight categories while using AdamW for the rest, guided by refitted scaling laws, and eliminates traditional batch-size warmups by starting directly at the target batch size - reducing total optimizer steps and safely allowing larger learning rates. Muon-type optimizers have shown stronger compute efficiency than AdamW at LLM scale (e.g. in the Moonshot/Kimi K2 training it replaced most AdamW usage).

Related FAQs (1)

Scaffold-rollout co-optimization (self-improving RL)

training_methods

Scaffold-rollout co-optimization is Ornith 1.0's self-improving training framework: instead of RL over solution rollouts alone, the model learns to generate the scaffold (task setup, tools, environment, search strategy) and the rollouts (solution trajectories driven by that scaffold) jointly. Because scaffold quality determines rollout quality, optimizing both lets the model discover better search trajectories and produce higher-quality agentic coding solutions than optimizing rollouts against a fixed scaffold. Ornith-1.0 applies this RL post-training on top of Gemma 4 and Qwen 3.5 bases, reaching state-of-the-art open-source results on Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.

Related FAQs (1)

Continual pretraining

training_methods

Continual pretraining (CPT) takes an already-pretrained base model and trains it further on a new corpus - usually to add capabilities the original pretraining lacked - before fine-tuning. Kimi K2.5 is built through continual pretraining on ~15 trillion mixed visual and text tokens atop Kimi-K2-Base, adding native multimodality and agentic abilities to the text-only base. CPT differs from ordinary pretraining in that it starts from existing weights (preserving general language ability) rather than random initialization, and differs from SFT/RL in scale and objective: it is large-corpus self-supervised training, not instruction following.

Related FAQs (1)

Asynchronous RL

training_methods

Asynchronous RL decouples rollout generation from policy optimization: one GPU pool runs inference producing rollouts (episodes/traces) while a separate pool performs policy gradient updates, and the learner pulls fresh weights in flight instead of pausing generation for a synchronized update. This maximizes utilization of both pools and scales RL to many environments - IBM runs Granite 4.2's GRPO this way (NeMo RL, environments on NeMo Gym), and frameworks like AReaL/AsyncFlow build on the same principle.

Related FAQs (1)

Verifiable rewards (RLVR)

training_methods

Verifiable rewards (RLVR) score a model's response by checking it against ground truth: unit tests for code, exact/checked answers for math, schema validation for structured output, tool-call success flags. Because the reward is objective, RL can scale to millions of prompts without human labelers. Prompts that cannot be verified automatically are handled by a reward model instead. Granite 4.2's RL environments mostly provide verifiable rewards across math, code, science, instruction following, tool use and structured output; DeepSeek-R1 pioneered the recipe at scale.

Related FAQs (1)

Group Relative Policy Optimization (GRPO)

training_methods

Group Relative Policy Optimization (GRPO) is a reinforcement-learning algorithm (popularized by DeepSeek-R1) that replaces the value/critic network of PPO with group-relative advantage estimates: for each prompt, sample a group of responses from the current policy and compute each response's advantage by normalizing its reward against the group's mean and std. This removes the critic's memory cost while keeping stable policy updates via importance-sampling clipping and a KL penalty. Used for post-training reasoning models - e.g. IBM Granite 4.2's multi-environment RL stage and DeepSeek-R1.

Related FAQs (1)

Synthetic reasoning traces (distillation)

training_methods

Synthetic reasoning traces are chain-of-thought reasoning paths generated by strong teacher models and distilled into a student's post-training data so the student learns to reason step-by-step. NVIDIA's Nemotron-3 post-training corpus includes synthetic reasoning traces from DeepSeek R1/R1-0528, Qwen3-235B-A22B, Nemotron 4 340B and Qwen2.5 models (the card notes "Improved using Qwen"); open reasoning datasets built this way include Bespoke-Stratos-17k and OpenCodeReasoning-2.

Related FAQs (1)

Multi-Teacher On-Policy Distillation (MOPD)

training_methods

Multi-Teacher On-Policy Distillation (MOPD) consolidates several domain-specialized policies - each produced by large-scale RL on one capability domain (reasoning, coding, agentic tool use, ...) - into a single deployable model. The student samples on-policy trajectories and is supervised by the complementary teachers, so each teacher corrects the student only where it is strongest; the consolidated model keeps all teachers' strengths without running them all. Spark-X2.5 (XHToken/iFLYTEK) uses MOPD as its final post-training stage.

Related FAQs (1)

Million-agent RL environments

training_methods

Million-agent RL environments refers to scaling post-training reinforcement learning across ~one million diverse agentic environments with progressively complex task distributions. Instead of RL on a handful of curated tasks, the policy is trained in parallel across a massive scaffolded environment pool (tool use, browsing, coding, embodied tasks), which forces robust generalization to unseen real-world agent settings. Qwen3.5 used this recipe, supported by asynchronous RL frameworks and massive-scale environment orchestration.

Related FAQs (1)

Early fusion (multimodal pre-training)

training_methods

Early fusion trains a single model on multimodal tokens from the start (text, image, video processed by one shared backbone), instead of bolting a vision adapter onto a text-pre-trained LLM (late fusion). Qwen3.5 is a natively multimodal family: pre-training uses early fusion on multimodal tokens, reaching near-100% multimodal training efficiency versus text-only training and outperforming the separately-trained Qwen3-VL models across reasoning, coding, agents and visual understanding.

Related FAQs (1)

On-policy distillation

training_methods

On-policy distillation distills a teacher's behavior into a student while the student generates the trajectories: the student samples actions/tokens from its own current policy, and the teacher provides per-step supervision (token-level corrections or rewards) on those on-policy samples. Unlike offline distillation on teacher-generated data, this keeps the training distribution matched to what the student will actually do at inference, correcting compounding errors early. LFM2.5-2.6B's post-training includes a multi-domain on-policy distillation stage, after per-domain teacher specialization, to transfer agent capability into the student.

Related FAQs (1)

Agentic reinforcement learning

training_methods

Agentic reinforcement learning trains a model inside the agentic harnesses and environments it will be deployed in, rather than on generic instruction data. The model is exposed to each harness's real tools, system prompts and multi-step interaction patterns, and is optimized end-to-end for task success (rewarded for correct tool calls and outcomes). Liquid AI used this in LFM2.5-2.6B's post-training so the model works reliably across popular agent environments - a general recipe for making small open models useful as on-device agents.

Related FAQs (1)

GSPO (Group Sequence Policy Optimization)

training_methods

GSPO (Group Sequence Policy Optimization) is a reinforcement-learning algorithm introduced by the Qwen team for stable and efficient reasoning-model post-training. Instead of optimizing a per-token importance ratio as PPO-style methods do, GSPO computes a sequence-level importance ratio by averaging log-probabilities over whole sequences within a group, and then optimizes this ratio with a group-relative advantage.

Whole-sequence credit assignment makes the optimization landscape much smoother — particularly valuable for architectures like Qwen3-Next, where a hybrid attention mechanism (Gated DeltaNet + Gated Attention) combined with a high-sparsity MoE makes token-level credit assignment noisy. Qwen leveraged GSPO to post-train Qwen3-Next-80B-A3B-Thinking, which demonstrates outstanding performance on complex reasoning tasks.

Related FAQs (1)

On-Policy Distillation (OPD)

training_methods

On-Policy Distillation (OPD) is a post-training paradigm in which a teacher model supervises a student model on the student's own rollouts (on-policy samples), rather than on a fixed corpus of teacher-generated sequences.

Because the training distribution matches the student's actual inference distribution, OPD addresses the train/inference mismatch of off-policy distillation and consolidates capabilities after SFT and RL. DeepSeek uses it as the final post-training stage: DeepSeek-V4.1-Flash follows the SFT -> RL -> OPD recipe, and DeepSeek-V3.x/V4 use the multi-teacher variant MOPD to merge domain-specialized RL teachers into one student.

OPD typically provides dense token-level supervision on student trajectories, making it more sample-efficient than reinforcement learning alone.

Related FAQs (1)

MOPD (Multi-Teacher On-Policy Distillation)

training_methods

A post-training paradigm that consolidates the capabilities of several domain-specialized reinforcement-learning teacher policies into one deployable student model.

Pipeline

  1. Supervised fine-tuning (SFT) on a curated corpus establishes instruction following, structured generation, and a stable policy initialization for reinforcement learning.
  2. Large-scale reinforcement learning runs across capability domains — language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following — producing a set of domain-specialized teacher policies.
  3. Multi-teacher distillation: the teachers are distilled into the student on the student's own rollouts, which keeps the training states aligned with real inference behavior and provides a dense per-token learning signal.

Effect

The combination of large-scale reinforcement learning and on-policy multi-teacher distillation enhances reasoning, coding, agentic, and instruction-following capabilities while collapsing several specialist checkpoints into a single model that behaves consistently between training and inference.

Related FAQs (1)