Kolibri-1

Aleph Alpha

Parameters

78.1B total / 3.5B active

MoE: total / active

Architecture

Sparse MoE reasoning model with hybrid 4:1 sliding-window/full GQA attention

Released

03.10.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16, F8_E4M3 Kolibri de en

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,048,576 tokens

Openness Index Score 100.0/100

About

Kolibri 1 (Aleph-Alpha/Kolibri-1) is Aleph Alpha's open-weight Mixture-of-Experts (MoE) reasoning model focused on German and English, released October 3, 2026 under Apache 2.0. It activates only 3.46B parameters per token (3,457,573,120) out of 78B total (78,103,074,560), routing each token to 6 of 384 experts plus 1 shared expert per MoE layer via token-choice top-6 sigmoid routing with Exact Quantile Balancing.

Architecture

A 50-layer transformer decoder uses 4:1 hybrid attention: 40 causal sliding-window GQA layers (window 513 tokens) interleaved with 10 full-attention GQA layers per stack. Hidden size 2,560, 48 attention heads / 4 KV heads, head dim 128, per-head RMSNorm QK normalization, SwiGLU MLPs (expert hidden 512), vocabulary 128,000. RoPE (base 10,000) is applied only in the sliding-window layers, so the context extends beyond the trained length without position scaling in principle: 262,144 tokens trained natively, validated and served up to 1,048,576 tokens (1M) — the card recommends ≤262,144 for serving efficiency and complex tasks. Weights ship as FP8 (float8_e4m3fn) in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and the MoE router stay in bfloat16. The model is served via the aleph-alpha-inference package (Kolibri vLLM plugin, --reasoning-parser kolibri1 --tool-call-parser kolibri1), needs ~78 GB FP8 weights (minimum 2× A100 80 GB, 1× H200, or 1× B200/B300).

Training

Pre-training used 20T tokens of a filtered bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code), plus 3.44T tokens of mid-training and 201B tokens of long-context extension on 768 NVIDIA B200 GPUs (392k GPU-hours pre-training, 6.4e23 FLOPS). The optimizer splits Nesterov Muon (0.001 LR, 2D backbone) from AdamW for embeddings/router/LM head, with per-head RMSNorm QK-norm. A custom UniBPE tokenizer (128k vocab) achieves higher German compression (4.7 bytes/token German, 4.2 English) with morphological awareness of German compounds. The SFT stage ran 4,000 steps on 168B tokens, selecting a 20-cluster data mix via MergeMix weighting and finishing with checkpoint model souping. The RL stage (1,000 steps, async vLLM, verifiable rewards) trained reasoning, German language consistency, hallucination abstention via the Merlin-Arthur protocol, and agentic tool-calling/software-engineering environments.

Capabilities

Kolibri supports explicit reasoning-effort control (low/medium/high/none via chat-template kwargs), Hermes-style tool calling, long-document RAG and agentic use. Recommended sampling: temperature 1.0, top_p 0.97, top_k 128. On its reported unified eval suite it reaches Overall (EN) 75.5 / Overall (DE) 70.8, GPQA Diamond (EN) 84.3, AIME 2025 (EN) 96.9, LiveCodeBench v6 85.9, SWE-Bench Verified 66.4, and strong industrial RAG (Automotive Supplier 99.0). Developed by Aleph Alpha Research GmbH; published by Aleph Alpha GmbH; tech report at aleph-alpha.com/downloads/tech-report.pdf.

Training Data Pre-training: 20T tokens of a filtered bilingual DE/EN corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings/translations and high-quality sources; plus 3.44T tokens mid-training and 201B tokens long-context extension. Post-training: bilingual SFT (168B tokens, MergeMix-weighted) and RL with verifiable rewards across reasoning, agentic and instruction-following environments.

Benchmark Scores

Benchmark Score Date
RGB Closed-Book
grounding_hallucination
100.00%
03.10.2026
SealQA (no distractors, >24k)
agentic
100.00%
03.10.2026
Industrial Drive Technology
agentic
100.00%
03.10.2026
TauBench V3 Banking
general_agent
100.00%
03.10.2026
LongBenchPro
long_context
100.00%
03.10.2026
Honeypot
agentic
100.00%
03.10.2026
FRAMES (<24k)
agentic
100.00%
03.10.2026
FRAMES (>24k)
agentic
100.00%
03.10.2026
OmniScience Accuracy
knowledge
14.80
03.10.2026
BFCL v4 (live)
agentic
100.00%
03.10.2026
AIME 2025 (DE)
stem_reasoning
100.00%
03.10.2026
MMLU-ProX
multilingual
29.68%
03.10.2026
Semiconductors
agentic
100.00%
03.10.2026
HumanEval+
code
92.70
03.10.2026
Overall (EN)
stem_reasoning
100.00%
03.10.2026
German Public Sector
agentic
100.00%
03.10.2026
Terminal Bench 2.1
coding_agent
30.27%
03.10.2026
Humanity's Last Exam (DE)
stem_reasoning
100.00%
03.10.2026
GPQA Diamond
stem_reasoning
81.09%
03.10.2026
AIME 2026 (DE)
stem_reasoning
100.00%
03.10.2026
OmniScience Non-Hallucination
knowledge
54.28%
03.10.2026
SWE-bench Verified
coding_agent
75.69%
03.10.2026
MMLU Pro COT
knowledge
100.00%
03.10.2026
SealQA (no distractors, <24k)
agentic
100.00%
03.10.2026
Overall (DE)
stem_reasoning
100.00%
03.10.2026
BFCL-V4
general_agent
74.29%
03.10.2026
AA-Omniscience Index (public set)
knowledge
100.00%
03.10.2026
SealQA (12 distractors, <24k)
agentic
100.00%
03.10.2026
GPQA Diamond (DE)
stem_reasoning
100.00%
03.10.2026
SQuAD (Utility Accuracy)
grounding_hallucination
100.00%
03.10.2026
Agentic Wiki QA (DE)
agentic
100.00%
03.10.2026
IFBench (prompt loose)
instruction_following
80.00%
03.10.2026
BFCL v4 (memory)
agentic
100.00%
03.10.2026
AIME 2025
stem_reasoning
100.00%
03.10.2026
MuSiQue (EN)
agentic
100.00%
03.10.2026
BFCL v4 (multi-turn)
agentic
100.00%
03.10.2026
SQuAD (M/A Grounding Score)
grounding_hallucination
100.00%
03.10.2026
AIME 26
stem_reasoning
95.92%
03.10.2026
LiveCodeBench v6
stem_reasoning
89.63%
03.10.2026
AA-LCR
long_context
85.38%
03.10.2026
Humanity's Last Exam
stem_reasoning
36.74%
03.10.2026
BFCL v4 (non-live AST)
agentic
100.00%
03.10.2026
Aerospace
agentic
100.00%
03.10.2026
BFCL v4 (web search)
agentic
100.00%
03.10.2026
RGB Negative (Abstention)
grounding_hallucination
100.00%
03.10.2026
Automotive Supplier
agentic
100.00%
03.10.2026
SealQA (12 distractors, >24k)
agentic
100.00%
03.10.2026
BrowseComp
general_agent
29.85%
03.10.2026
RGB Fact-Check (Error Correction)
grounding_hallucination
100.00%
03.10.2026

Architecture

Decoder Block ×50 input Embedding vocab 128K · d 2560 Sliding Window Attn Hybrid 48:4 · dₕ 128 · win 513 ×40 Full Attention Hybrid 48:4 · dₕ 128 · win 513 ×10 MoE FFN 384 experts · top-6 · dᴻ 512 Final Norm LM Head vocab 128K output
Attention
Hybrid Attention (48:4)
MoE
384 experts · top-6 per token
Layers
50
Hidden size
2560
Context
262K tokens
RoPE θ
10K
Parameters
78103.1M
Active params
3457.6M

Source: Hugging Face config.json · Kolibri1ForCausalLM · exact layer pattern · model repo

Type: decoder-only transformer
Attention: hybrid 4:1 sliding-window (513) GQA + full GQA; per-head RMSNorm QK-norm; RoPE only in sliding-window layers
Decoder: sparse MoE
MoE: yes (384 experts)
Routing: token-choice, top-6 sigmoid routing with Exact Quantile Balancing
Layers 50
Context length 262K
Extended context 1M
Attention heads 48
KV heads 4
Head dim 128
Hidden size 2560
Vocabulary 128K
FFN dim 512
Vision No
RoPE θ 10K
Sliding window 513
Layer Pattern swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn, swa, swa, swa, swa, full_attn
Moe {'enabled': True, 'experts_active': 6, 'experts_total': 384, 'first_k_dense': None, 'intermediate_size': 512, 'layer_freq': None, 'route_top_k': 6, 'scoring_func': None, 'shared_experts': None}
Tied embeddings No

Training Pipeline

  1. 1
    pretraining

    Pre-training

    Trained from random initialization with a causal next-token-prediction objective on a filtered, bilingual (German/English) corpus of 20T tokens (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality curated sources. Training examples were 16,384-token sequences with multiple documents packed per sequence, run over 264,750 optimization steps with two gradient-accumulation steps across 768 NVIDIA B200 GPUs (EP8 FSDP16 DP6, TorchTitan). Data curation included URL and heuristic filtering, exact/MinHash/substring deduplication, quality classification into token-mass quantile buckets, PII redaction, and mix selection via mix-search proxy models.

    511 h

  2. 2
    cpt

    Mid-training

    Continued pre-training on 3.44T tokens with the sequence length increased to 65,536 tokens, using the same setup as pre-training. The data mix shifted towards instruction/reasoning, STEM QA-style, code, and agentic code/tool-use data, was globally shuffled with per-dataset and cross-dataset exact deduplication, and was decontaminated against the evaluation suite. The phase resumed the pre-training optimizer state and therefore used no additional learning-rate warmup.

    120 h

  3. 3
    cpt

    Long-context extension

    Final long-context adaptation stage on 201B tokens with the sequence length increased to 262,144 tokens, the model's native context length. The mix was built by length-bucketing the mid-training pool and blending it with OCR'd PDFs (fixed at 34% of tokens, with documents up to 256k tokens; ablations found the optimal blend at x=33), and was decontaminated against the evaluation suite. It again used no additional warmup and ran with FSDP128 DP6 parallelism instead of the pre-training setup.

    13 h

  4. 4
    sft

    Supervised Fine-Tuning (SFT)

    Fine-tuned the Kolibri Base checkpoint for reasoning and instruction following on a filtered, bilingual mix of permissively available open-source datasets, in-house generated data, and curated RL warm-start sets (e.g., 11k high-reward retrieval completions and filtered Merlin-Arthur rollouts), unified, decontaminated (0.006% of rows removed) and filtered for political bias. Mix weighting was chosen by adapting MergeMix: one specialist per 20 data cluster was trained on 4.5B packed positions, 69 Dirichlet-sampled weightings were merged and scored on 16 benchmarks across seven capabilities, and the best mixtures were validated at proxy scale before the target run, which saw the weighted mix about once over 168B tokens. The optimizer split of pre-training was kept (Muon for 2D backbone, AdamW elsewhere) at half the pre-training learning rates, and the final checkpoint was produced by model souping two checkpoints trained for the same token horizon on two different data mixtures.

  5. 5
    rl

    Reinforcement Learning (RL)

    Trained in an asynchronous RL setup where vLLM inference ran in parallel to training with weight synchronization not waiting for ongoing generations, keeping the SFT optimizer split but with an adapted learning rate, and additionally using quantization-aware training to enable efficient low-precision (FP8) vLLM inference. Environments covered reasoning, agentic, software-engineering/terminal, tool-calling, retrieval-augmented QA, long-context QA, German-focused and Merlin-Arthur abstention tasks, with verifiable answer rewards plus format rewards, heavy randomization over harnesses, system prompts and languages, and on-policy self-distillation on hard tasks. Reasoning-effort control was trained across all levels with different length penalties, and a German language-consistency reward was applied to reasoning traces and answers.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Kolibri pre-training corpus (filtered bilingual DE/EN mix) pretraining 20T tokens (~62.5% English, ~23.9% German, ~13.6% code) text mixed
Kolibri mid-training mix pretraining 3.44T tokens text mixed
Kolibri long-context extension mix pretraining 201B tokens text mixed
Common Crawl pretraining — text automated
English pre-training rephrases pretraining — text synthetic
German pre-training rephrases pretraining — text synthetic
LLM-as-a-judge quality-classifier annotations pretraining — text synthetic
Kolibri SFT mix finetune 168B tokens (~1 epoch per cluster over 20 clusters) text mixed
Translated German reasoning SFT data finetune 10.6B tokens text synthetic
Retrieval warm-start SFT set (11k high-reward completions) finetune 11k completions text synthetic
Merlin-Arthur warm-start SFT rollouts finetune — text synthetic
Permissive open agentic SFT datasets finetune — text mixed
Software-engineering SFT dataset (code patch generation) finetune — text synthetic
Software-engineering SFT dataset (bug fixing as Terminus 2 agent) finetune — text synthetic
Curated long-context QA datasets (SFT) finetune — text mixed
Abstention data (SFT) finetune — text —
SFT safety data (PII extraction attack refusals) finetune — text —
Political-bias counter and values-alignment data (post-training) finetune — text mixed
RL environments (reasoning, agentic, and instruction following mix) rl — text —
Retrieval-augmented QA RL environments (DE/EN) rl — text synthetic
Merlin-Arthur hallucination RL environments rl — text synthetic

Related Models