Parameters
78.1B total / 3.5B active
MoE: total / active
Architecture
Sparse MoE reasoning model with hybrid 4:1 sliding-window/full GQA attention
Released
03.10.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
1,048,576 tokens
About
Kolibri 1 (Aleph-Alpha/Kolibri-1) is Aleph Alpha's open-weight Mixture-of-Experts (MoE) reasoning model focused on German and English, released October 3, 2026 under Apache 2.0. It activates only 3.46B parameters per token (3,457,573,120) out of 78B total (78,103,074,560), routing each token to 6 of 384 experts plus 1 shared expert per MoE layer via token-choice top-6 sigmoid routing with Exact Quantile Balancing.
Architecture
A 50-layer transformer decoder uses 4:1 hybrid attention: 40 causal sliding-window GQA layers (window 513 tokens) interleaved with 10 full-attention GQA layers per stack. Hidden size 2,560, 48 attention heads / 4 KV heads, head dim 128, per-head RMSNorm QK normalization, SwiGLU MLPs (expert hidden 512), vocabulary 128,000. RoPE (base 10,000) is applied only in the sliding-window layers, so the context extends beyond the trained length without position scaling in principle: 262,144 tokens trained natively, validated and served up to 1,048,576 tokens (1M) — the card recommends ≤262,144 for serving efficiency and complex tasks. Weights ship as FP8 (float8_e4m3fn) in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and the MoE router stay in bfloat16. The model is served via the aleph-alpha-inference package (Kolibri vLLM plugin, --reasoning-parser kolibri1 --tool-call-parser kolibri1), needs ~78 GB FP8 weights (minimum 2× A100 80 GB, 1× H200, or 1× B200/B300).
Training
Pre-training used 20T tokens of a filtered bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code), plus 3.44T tokens of mid-training and 201B tokens of long-context extension on 768 NVIDIA B200 GPUs (392k GPU-hours pre-training, 6.4e23 FLOPS). The optimizer splits Nesterov Muon (0.001 LR, 2D backbone) from AdamW for embeddings/router/LM head, with per-head RMSNorm QK-norm. A custom UniBPE tokenizer (128k vocab) achieves higher German compression (4.7 bytes/token German, 4.2 English) with morphological awareness of German compounds. The SFT stage ran 4,000 steps on 168B tokens, selecting a 20-cluster data mix via MergeMix weighting and finishing with checkpoint model souping. The RL stage (1,000 steps, async vLLM, verifiable rewards) trained reasoning, German language consistency, hallucination abstention via the Merlin-Arthur protocol, and agentic tool-calling/software-engineering environments.
Capabilities
Kolibri supports explicit reasoning-effort control (low/medium/high/none via chat-template kwargs), Hermes-style tool calling, long-document RAG and agentic use. Recommended sampling: temperature 1.0, top_p 0.97, top_k 128. On its reported unified eval suite it reaches Overall (EN) 75.5 / Overall (DE) 70.8, GPQA Diamond (EN) 84.3, AIME 2025 (EN) 96.9, LiveCodeBench v6 85.9, SWE-Bench Verified 66.4, and strong industrial RAG (Automotive Supplier 99.0). Developed by Aleph Alpha Research GmbH; published by Aleph Alpha GmbH; tech report at aleph-alpha.com/downloads/tech-report.pdf.
Training Data Pre-training: 20T tokens of a filtered bilingual DE/EN corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings/translations and high-quality sources; plus 3.44T tokens mid-training and 201B tokens long-context extension. Post-training: bilingual SFT (168B tokens, MergeMix-weighted) and RL with verifiable rewards across reasoning, agentic and instruction-following environments.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
RGB Closed-Book
grounding_hallucination
|
100.00%
|
03.10.2026 |
|
SealQA (no distractors, >24k)
agentic
|
100.00%
|
03.10.2026 |
|
Industrial Drive Technology
agentic
|
100.00%
|
03.10.2026 |
|
TauBench V3 Banking
general_agent
|
100.00%
|
03.10.2026 |
|
LongBenchPro
long_context
|
100.00%
|
03.10.2026 |
|
Honeypot
agentic
|
100.00%
|
03.10.2026 |
|
FRAMES (<24k)
agentic
|
100.00%
|
03.10.2026 |
|
FRAMES (>24k)
agentic
|
100.00%
|
03.10.2026 |
|
OmniScience Accuracy
knowledge
|
14.80
|
03.10.2026 |
|
BFCL v4 (live)
agentic
|
100.00%
|
03.10.2026 |
|
AIME 2025 (DE)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
MMLU-ProX
multilingual
|
29.68%
|
03.10.2026 |
|
Semiconductors
agentic
|
100.00%
|
03.10.2026 |
|
HumanEval+
code
|
92.70
|
03.10.2026 |
|
Overall (EN)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
German Public Sector
agentic
|
100.00%
|
03.10.2026 |
|
Terminal Bench 2.1
coding_agent
|
30.27%
|
03.10.2026 |
|
Humanity's Last Exam (DE)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
GPQA Diamond
stem_reasoning
|
81.09%
|
03.10.2026 |
|
AIME 2026 (DE)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
OmniScience Non-Hallucination
knowledge
|
54.28%
|
03.10.2026 |
|
SWE-bench Verified
coding_agent
|
75.69%
|
03.10.2026 |
|
MMLU Pro COT
knowledge
|
100.00%
|
03.10.2026 |
|
SealQA (no distractors, <24k)
agentic
|
100.00%
|
03.10.2026 |
|
Overall (DE)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
BFCL-V4
general_agent
|
74.29%
|
03.10.2026 |
|
AA-Omniscience Index (public set)
knowledge
|
100.00%
|
03.10.2026 |
|
SealQA (12 distractors, <24k)
agentic
|
100.00%
|
03.10.2026 |
|
GPQA Diamond (DE)
stem_reasoning
|
100.00%
|
03.10.2026 |
|
SQuAD (Utility Accuracy)
grounding_hallucination
|
100.00%
|
03.10.2026 |
|
Agentic Wiki QA (DE)
agentic
|
100.00%
|
03.10.2026 |
|
IFBench (prompt loose)
instruction_following
|
80.00%
|
03.10.2026 |
|
BFCL v4 (memory)
agentic
|
100.00%
|
03.10.2026 |
|
AIME 2025
stem_reasoning
|
100.00%
|
03.10.2026 |
|
MuSiQue (EN)
agentic
|
100.00%
|
03.10.2026 |
|
BFCL v4 (multi-turn)
agentic
|
100.00%
|
03.10.2026 |
|
SQuAD (M/A Grounding Score)
grounding_hallucination
|
100.00%
|
03.10.2026 |
|
AIME 26
stem_reasoning
|
95.92%
|
03.10.2026 |
|
LiveCodeBench v6
stem_reasoning
|
89.63%
|
03.10.2026 |
|
AA-LCR
long_context
|
85.38%
|
03.10.2026 |
|
Humanity's Last Exam
stem_reasoning
|
36.74%
|
03.10.2026 |
|
BFCL v4 (non-live AST)
agentic
|
100.00%
|
03.10.2026 |
|
Aerospace
agentic
|
100.00%
|
03.10.2026 |
|
BFCL v4 (web search)
agentic
|
100.00%
|
03.10.2026 |
|
RGB Negative (Abstention)
grounding_hallucination
|
100.00%
|
03.10.2026 |
|
Automotive Supplier
agentic
|
100.00%
|
03.10.2026 |
|
SealQA (12 distractors, >24k)
agentic
|
100.00%
|
03.10.2026 |
|
BrowseComp
general_agent
|
29.85%
|
03.10.2026 |
|
RGB Fact-Check (Error Correction)
grounding_hallucination
|
100.00%
|
03.10.2026 |
Architecture
- Attention
- Hybrid Attention (48:4)
- MoE
- 384 experts · top-6 per token
- Layers
- 50
- Hidden size
- 2560
- Context
- 262K tokens
- RoPE θ
- 10K
- Parameters
- 78103.1M
- Active params
- 3457.6M
Source: Hugging Face config.json · Kolibri1ForCausalLM · exact layer pattern · model repo
Training Pipeline
-
1
pretraining
Pre-training
Trained from random initialization with a causal next-token-prediction objective on a filtered, bilingual (German/English) corpus of 20T tokens (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality curated sources. Training examples were 16,384-token sequences with multiple documents packed per sequence, run over 264,750 optimization steps with two gradient-accumulation steps across 768 NVIDIA B200 GPUs (EP8 FSDP16 DP6, TorchTitan). Data curation included URL and heuristic filtering, exact/MinHash/substring deduplication, quality classification into token-mass quantile buckets, PII redaction, and mix selection via mix-search proxy models.
-
2
cpt
Mid-training
Continued pre-training on 3.44T tokens with the sequence length increased to 65,536 tokens, using the same setup as pre-training. The data mix shifted towards instruction/reasoning, STEM QA-style, code, and agentic code/tool-use data, was globally shuffled with per-dataset and cross-dataset exact deduplication, and was decontaminated against the evaluation suite. The phase resumed the pre-training optimizer state and therefore used no additional learning-rate warmup.
-
3
cpt
Long-context extension
Final long-context adaptation stage on 201B tokens with the sequence length increased to 262,144 tokens, the model's native context length. The mix was built by length-bucketing the mid-training pool and blending it with OCR'd PDFs (fixed at 34% of tokens, with documents up to 256k tokens; ablations found the optimal blend at x=33), and was decontaminated against the evaluation suite. It again used no additional warmup and ran with FSDP128 DP6 parallelism instead of the pre-training setup.
-
4
sft
Supervised Fine-Tuning (SFT)
Fine-tuned the Kolibri Base checkpoint for reasoning and instruction following on a filtered, bilingual mix of permissively available open-source datasets, in-house generated data, and curated RL warm-start sets (e.g., 11k high-reward retrieval completions and filtered Merlin-Arthur rollouts), unified, decontaminated (0.006% of rows removed) and filtered for political bias. Mix weighting was chosen by adapting MergeMix: one specialist per 20 data cluster was trained on 4.5B packed positions, 69 Dirichlet-sampled weightings were merged and scored on 16 benchmarks across seven capabilities, and the best mixtures were validated at proxy scale before the target run, which saw the weighted mix about once over 168B tokens. The optimizer split of pre-training was kept (Muon for 2D backbone, AdamW elsewhere) at half the pre-training learning rates, and the final checkpoint was produced by model souping two checkpoints trained for the same token horizon on two different data mixtures.
-
5
rl
Reinforcement Learning (RL)
Trained in an asynchronous RL setup where vLLM inference ran in parallel to training with weight synchronization not waiting for ongoing generations, keeping the SFT optimizer split but with an adapted learning rate, and additionally using quantization-aware training to enable efficient low-precision (FP8) vLLM inference. Environments covered reasoning, agentic, software-engineering/terminal, tool-calling, retrieval-augmented QA, long-context QA, German-focused and Merlin-Arthur abstention tasks, with verifiable answer rewards plus format rewards, heavy randomization over harnesses, system prompts and languages, and on-policy self-distillation on hard tasks. Reasoning-effort control was trained across all levels with different length penalties, and a German language-consistency reward was applied to reasoning traces and answers.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Kolibri pre-training corpus (filtered bilingual DE/EN mix) | pretraining | 20T tokens (~62.5% English, ~23.9% German, ~13.6% code) | text | mixed |
| Kolibri mid-training mix | pretraining | 3.44T tokens | text | mixed |
| Kolibri long-context extension mix | pretraining | 201B tokens | text | mixed |
| Common Crawl | pretraining | — | text | automated |
| English pre-training rephrases | pretraining | — | text | synthetic |
| German pre-training rephrases | pretraining | — | text | synthetic |
| LLM-as-a-judge quality-classifier annotations | pretraining | — | text | synthetic |
| Kolibri SFT mix | finetune | 168B tokens (~1 epoch per cluster over 20 clusters) | text | mixed |
| Translated German reasoning SFT data | finetune | 10.6B tokens | text | synthetic |
| Retrieval warm-start SFT set (11k high-reward completions) | finetune | 11k completions | text | synthetic |
| Merlin-Arthur warm-start SFT rollouts | finetune | — | text | synthetic |
| Permissive open agentic SFT datasets | finetune | — | text | mixed |
| Software-engineering SFT dataset (code patch generation) | finetune | — | text | synthetic |
| Software-engineering SFT dataset (bug fixing as Terminus 2 agent) | finetune | — | text | synthetic |
| Curated long-context QA datasets (SFT) | finetune | — | text | mixed |
| Abstention data (SFT) | finetune | — | text | — |
| SFT safety data (PII extraction attack refusals) | finetune | — | text | — |
| Political-bias counter and values-alignment data (post-training) | finetune | — | text | mixed |
| RL environments (reasoning, agentic, and instruction following mix) | rl | — | text | — |
| Retrieval-augmented QA RL environments (DE/EN) | rl | — | text | synthetic |
| Merlin-Arthur hallucination RL environments | rl | — | text | synthetic |
Linked Resources
Kolibri Technical Report
https://aleph-alpha.com/downloads/tech-report.pdf
Kolibri Has Landed: A Sovereign Open-Weight Model
https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
Aleph-Alpha/aleph-alpha-inference
https://github.com/Aleph-Alpha/aleph-alpha-inference
Aleph-Alpha-Research/eval-framework
https://github.com/Aleph-Alpha-Research/eval-framework
harbor-framework/harbor
https://github.com/harbor-framework/harbor
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
https://arxiv.org/abs/2512.11614
MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
https://arxiv.org/abs/2601.17858
Training on the Party Line
https://aleph-alpha.com/en/blog/training-on-the-party-line/
Model Training as Code (Savanna)
https://aleph-alpha.com/en/blog/model-training-as-code/
Public Summary of Training Data
https://aleph-alpha.com/downloads/data-summary.pdf
Kolibri-1
https://huggingface.co/collections/Aleph-Alpha/kolibri-1