Aleph Alpha Releases Open-Weight Kolibri-1 German-English MoE
Aleph Alpha released Kolibri-1, an Apache 2.0 78.1B German-English MoE with about 3.46B active parameters and context checked to 1M tokens.
Aleph Alpha released Kolibri, also listed as Kolibri-1, an Apache 2.0 open-weight mixture-of-experts model for German and English aimed at on-premise use in public administration, industry, and aerospace. Exact counts are 78,103,074,560 total parameters and 3,457,573,120 active per token—about 78.1 billion and 3.46 billion—with 384 experts per layer, one shared and six routed; some reports round the active count to 3 billion. Native context is 262,144 tokens and quality was checked to 1,048,576, though several outlets simply say a 1-million-token window; the knowledge cutoff is 18 June 2026. Pre-training is described as 20 trillion tokens on 768 NVIDIA B200 GPUs for about 21 days, plus 3.44 trillion mid-training tokens and a long-context stage totaling about 6.4e23 FLOPS, while another source says roughly 24 trillion tokens overall, with German at 21.3 percent (elsewhere “more than a fifth”), trained in Germany and Finland. FP8 weights in 128-by-128 blocks have a footprint of about 78 GB and are said to run on a single B200 or H200, with tool calling and adjustable reasoning; reported scores include AIME 2025 at 96.9, GPQA Diamond at 84.3, LiveCodeBench v6 at 85.9, and 71 percent on German benchmarks, plus faster decoding than named smaller peers. Sources differ on the public date (a Hugging Face listing on 2 October 2026 versus a 3 October release) and say the model follows Kolibri Origin, was aligned with the EU AI Act and the EU GPAI Code of Practice, and used Gemma 4, Mistral-NeMo, and Qwen models for some rephrasing, labeling, or synthetic data.
- Aleph Alpha released Kolibri (also Kolibri-1), an Apache 2.0 open-weight German-English mixture-of-experts model; a Hugging Face listing is dated 2 October 2026 and one report states a 3 October 2026 release.
- Exact size is 78,103,074,560 total parameters and 3,457,573,120 active per token (about 78.1B and 3.46B), with 384 experts per layer (one shared, six routed); some reports round active parameters to 3 billion.
- Native context is 262,144 tokens, with quality checked to 1,048,576; several outlets describe a 1-million-token window without that split. Knowledge cutoff is 18 June 2026.
- One source cites 20 trillion pre-training tokens on 768 NVIDIA B200s over about 21 days, plus 3.44 trillion mid-training tokens and a long-context stage (about 6.4e23 FLOPS); another says roughly 24 trillion tokens overall, trained in…
- German is 21.3 percent of training data in one report and “more than a fifth” in another. FP8 weights in 128-by-128 blocks have about a 78 GB footprint and are said to fit one B200 or H200.
- Cited scores: AIME 2025 96.9, GPQA Diamond 84.3, LiveCodeBench v6 85.9, and 71 percent on German benchmarks; a German tokenizer uses 11.2 percent fewer tokens than GPT-5.
- It follows Kolibri Origin (30B total, 3B active, 65k context). Aleph Alpha cites the EU AI Act and the EU GPAI Code of Practice; sources say Gemma 4, Mistral-NeMo, and Qwen models were used for some rephrasing, labeling, or synthetic data.
Coverage timelineoldest first · each row is one article
- · 6d agoAleph-Alpha/Kolibri-1 — new model trending #28 on Hugging Face
Hugging Face trending models· 72
Aleph Alpha released Kolibri-1, an Apache 2.0 78B MoE reasoning model for German and English.
- · 5d agoKolibri Has Landed: A Sovereign Open-Weight Model
Hacker News · AI· 67
Aleph Alpha released Kolibri, an Apache 2.0 English-German MoE with 78B total and 3B active parameters.
- · 5d agoShow HN: Germany's new sovereign AI model Kolibri
Hacker News · AI· 72
Aleph Alpha released Kolibri, an Apache 2.0 78B MoE model for German and English.