Aleph-Alpha/Kolibri-1 — new model trending #28 on Hugging Face
Aleph Alpha released Kolibri-1, an Apache 2.0 78B MoE reasoning model for German and English.
Aleph Alpha released Kolibri-1, an Apache 2.0 mixture-of-experts reasoning model aimed at German and English, with tool calling and an explicit reasoning mode. It has 78,103,074,560 total parameters and 3,457,573,120 active per token, using 384 experts per layer with one shared and six routed. Weights are float8 in 128 by 128 blocks; the footprint is about 78 GB. The model was pre-trained on 20 trillion tokens on 768 NVIDIA B200 GPUs for 21 days, followed by 3.44 trillion mid-training tokens and a long-context stage, for about 6.4e23 FLOPS. Native context is 262,144 tokens, with quality checked up to 1,048,576; the knowledge cutoff is 18 June 2026.
- 78.1B total parameters with about 3.46B active per token.
- Apache 2.0 FP8 weights; memory footprint is about 78 GB.
- Context validated to 1,048,576 tokens; 262,144 recommended for serving.
- Pre-training used 20T tokens on 768 NVIDIA B200s over about 21 days.
- Supports reasoning and tool calling in German and English.
Full article2,409 words · extracted from huggingface.co · click to collapse
# Kolibri
[Tech report](https://aleph-alpha.com/downloads/tech-report.pdf) | [Tech blog](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/)
Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a
focus on German and English. The model supports an explicit reasoning mode
and tool calling. It is optimized for long-context and inference efficiency.
## Model overview
| Model | Kolibri 1 |
| --- | --- |
| Model Provider | Aleph Alpha GmbH |
| Model Developer | Aleph Alpha Research GmbH |
| Architecture | Mixture-of-Experts |
| Total parameters | 78B (`78,103,074,560`) |
| Active parameters / token | 3.46B (`3,457,573,120`) |
| Languages | German, English |
| Context length | `1,048,576 tokens`; we recommend ≤262,144 tokens for serving efficiency and complex tasks |
| Precision | `float8_e4m3fn` weights in 128×128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in `bfloat16` |
| Reasoning mode | Yes |
| Tool calling | Yes |
| License | Apache 2.0 |
| Knowledge cutoff | EN: June 18, 2026, DE: June 18, 2026<br>*This only affects implicit knowledge, the model may use more recent information through tool use.* |
| Hardware requirements | Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300. |
| Best for | Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant |
| Release Date | 3rd of October 2026 |
| Code of Practice | Aleph Alpha is a signatory of the EU GPAI Code of Practice, see <https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai>. |
| Training Data | Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.<br>Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases. |
| Training method | We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed. |