
What launched
On October 3, 2026, Aleph Alpha published Kolibri under the headline "Kolibri Has Landed: A Sovereign Open-Weight Model" — timed for the Day of German Reunification. The full weights are on Hugging Face under the Apache 2.0 license: download it, run it, fine-tune it, and ship it commercially.
The spec sheet
- Architecture: English-German Mixture-of-Experts Transformer
- Total parameters: 78.1 billion
- Active per token: ~3.46 billion (a router picks 6 of 384 experts in each of 50 layers)
- Context: trained natively to 262,144 tokens, extendable to 1,048,576
- Reasoning: four effort levels (none, low, medium, high) plus tool calling
- Languages: English and German, with a German-first tokenizer
- License: Apache 2.0
The split is the point: enough capacity for serious work, but only about 3.5B parameters active per token, so it computes like a much smaller model — cheap enough to run in your own data center instead of a foreign cloud.
Trained for sovereignty
Aleph Alpha trained Kolibri on infrastructure in Germany and Finland under European and German law. The 20-trillion-token pre-training mix is 43.4% English and 23.4% German, and the tokenizer was built around German morphology rather than adapted to it. The team validated the pipeline first with a smaller sibling, Kolibri Origin (about 30B total parameters, 3B active, 65k context), then scaled the same pipeline to Kolibri in roughly three months between June and September 2026 — a stability story as much as a capability story, since pre-training reportedly stayed stable without human intervention when hardware failed.
The benchmarks — published honestly
Aleph Alpha's published numbers: 96.9 on AIME 2025, 84.3 on GPQA Diamond in English (81.3 in German), 85.9 on LiveCodeBench v6, and 66.4 on SWE-Bench Verified. Notably, the model card includes its own "best dense model" baseline — and the unnamed baseline outscores Kolibri in every category. At 1M context, Kolibri reaches 63.2% on RULER against a cited 73.1% at 512k. The framing is quality-per-serving-cost, not raw leaderboard dominance, and the company deserves credit for publishing the comparison that does not flatter it.
Who it is for
Kolibri targets mission-critical work in public administration, industry, and aerospace — customers meant to run it in their own infrastructure rather than send sensitive data to a third-party chat API. The model card explicitly lists autonomous operation without human review as outside intended use.
“Kolibri demonstrates that we have the talent and expertise in Germany to develop competitive AI models. For us, AI sovereignty means freedom of choice.”
Frequently Asked Questions
Is Kolibri really open source?
The weights and configuration files are published under Apache 2.0, which permits commercial use, modification, and redistribution — the industry's practical definition of open-weight. Training data and some pipeline details remain proprietary.
How does Kolibri compare to Chinese open models like DeepSeek and Qwen?
Strategically it is positioned as Europe's answer in the open-weight race, with a German-English bilingual focus instead of global coverage ("depth over breadth"). No independent head-to-head benchmarks are published yet, so treat the comparison as strategic, not measured.
Can I run Kolibri on consumer hardware?
Not the full model — 78B parameters need roughly 78 GB of VRAM in FP8, with listed minimums starting at 2x A100 80GB or 1x H200. The Apache 2.0 license allows community quantized or distilled variants for smaller hardware, which typically follow open releases.
What does "sovereign AI" mean here?
A model trained on EU infrastructure under EU and German law, with transparent weights, documented data curation, on-premise deployment, and EU AI Act alignment via the GPAI Code of Practice — so regulated organizations keep control of their data and their compliance story.
What is Kolibri Origin?
The smaller sibling Aleph Alpha used to validate its training pipeline: about 30B total parameters, 3B active, 65k context. Validating at 30B before scaling to 78B in three months is the project's main reliability credential.



