Aleph Alpha Launches Kolibri: A 78B Sovereign LLM Built for EU Compliance
Kolibri is a 78B parameter MoE model that dominates German reasoning and EU compliance, but you need serious GPU hardware to run it.
Models · Source: Hacker News
What happened
Aleph Alpha just dropped Kolibri. It is an open-weight large language model built specifically for German and English. The model has 78.1 billion parameters but operates as a mixture of experts. It only uses 3.46 billion parameters per token. It operates under an Apache 2.0 license for its weights and configuration files.
The model is designed from the ground up for European data sovereignty. It was trained entirely on infrastructure in Germany and Finland to comply with the strict EU AI Act. Enterprises can run it locally on their own servers so their sensitive data never leaves the building. The training process consumed 24 trillion tokens on a cluster of 768 NVIDIA B200 graphics processing units.
Kolibri brings heavy architectural changes to the table. It uses a custom tokenizer called UniBPE that handles long German compound words highly efficiently. It also features a sliding-window attention mechanism to support contexts up to one million tokens without breaking the bank. The model includes four adjustable reasoning efforts. It is also explicitly trained to admit when it does not know an answer instead of making things up.
Key facts
- 78.1 billion — Total parameters in the Kolibri model
- 3.46 billion — Active parameters used per token
- 15% — Fewer tokens needed for German legal text compared to GPT-5
- 44% — Rate at which Kolibri admits it does not know an answer
- 1,048,576 — Maximum tested context window in tokens
Why it matters
This changes the game for European builders and government contractors. You no longer have to choose between high performance and strict data privacy. Kolibri beats comparable open models in German math and reasoning while guaranteeing intellectual property safety. If you build retrieval augmented generation applications for German legal or corporate documents, this model is highly optimized for that exact pipeline. It thinks in German for German prompts instead of translating back and forth to English.
The second-order effect is a massive shift in regional artificial intelligence development. We are seeing models hyper-optimized for specific languages and regulatory environments rather than generic global models. Kolibri proves that targeted tokenizers and specialized reasoning data can drastically outperform larger models in local languages. This will push other regions to build sovereign models tailored to their own laws, hardware constraints, and linguistics. The era of one size fits all models is ending.
For builders
Cheaper German context windows
Kolibri uses a custom tokenizer that requires 15 percent fewer tokens for German legal text compared to the GPT-5 tokenizer. Builders save heavily on compute costs and fit more text into the exact same context window. European enterprises pay less for heavy document processing.
High hardware barrier to entry
Despite only using 3.5 billion parameters per token, you must load all 78 billion parameters into memory. You need at least two 80GB A100 GPUs or a single H200 to run it. Indie hackers lose out on local deployment, while well-funded enterprises and cloud providers take the advantage.
Better RAG reliability
The model was trained using the Merlin-Arthur protocol to say it does not know an answer instead of hallucinating. It admits ignorance 44 percent of the time on omniscience tests compared to 11 percent for Qwen. Founders building enterprise search tools get much higher reliability out of the box.
My take
I love seeing a model built specifically to turn a regulatory burden into a feature. Europe usually regulates instead of innovating, but Aleph Alpha actually built a technical solution to the EU AI Act. If you want to sell AI to the German government or automotive sector, this is your new baseline model. Stop fighting the laws and start building products that use them as a moat.
Original reporting: Hacker News. This is my rewrite and opinion.