Aleph Alpha Releases Kolibri, a 78B Open-Weight English-German MoE Model
Aleph Alpha released Kolibri on 3 Oct 2026: a 78.1B English-German MoE with ~3.46B active params, Apache 2.0 weights, and vendor-reported EN/DE benchmarks. On-prem sovereignty pitch; no public price.

Aleph Alpha released Kolibri on 3 October 2026, German Unity Day. The company describes it as a sovereign open-weight English-German Mixture-of-Experts Transformer, with full weights on Hugging Face under Apache 2.0.
On paper the headline specs are straightforward. Kolibri has 78.1B total parameters and about 3.46B active per token. Aleph Alpha says it was trained to 262,144 tokens of context and can serve up to 1,048,576 with explicit flags. The pitch is on-prem and regulated deployment, not another closed cloud API.
Those numbers, and the benchmark tables that accompany them, are Aleph Alpha-reported. Independent verification is not in this piece. What is confirmed is the public release page, the Apache 2.0 weight drop as Aleph-Alpha/Kolibri-1, and a serving path that needs the company’s own vLLM plugin.
What Aleph Alpha Says Kolibri Is
Kolibri is framed as a bilingual English-German MoE aimed at public administration, industrials, and aerospace. Aleph Alpha says German made up about 21.3% of pre-training tokens, roughly 4.3T of a 20T pre-train run, with translation used sparingly. The company argues that keeps cultural and administrative German registers intact instead of bolting German onto an English-first model.
Serving is not plug-and-play with stock vLLM alone. Aleph Alpha says you need the aleph-alpha-inference package or the container ghcr.io/aleph-alpha/aleph-alpha-inference, then a Kolibri-specific reasoning and tool-call parser. Recommended sampling on the release page is temperature 1.0, top_p 0.97, and top_k 128. No public list price appears on the announcement.
Reasoning effort is exposed as four levels: none, low, medium, and high. That is intended to trade latency and cost against answer quality on the same weights.
Kolibri vs Kolibri Origin
Kolibri Origin was the internal predecessor. Aleph Alpha says Origin finished pre-training on 11 June 2026 at 30.6B total and 3.27B active, with a 65,536-token trained context. Kolibri finished pre-training on 11 September 2026 and shipped publicly on 3 October.
| Spec | Kolibri Origin | Kolibri |
|---|---|---|
| Finished pre-training | 11 June 2026 | 11 September 2026 |
| Public release | None | 3 October 2026 |
| Total parameters | 30.6B | 78.1B |
| Active params / token | 3.27B | 3.46B |
| Pre-training tokens | 7.51T | 20T |
| Longest trained context | 65,536 | 262,144 |
| Reasoning effort levels | One mode | none, low, medium, high |
| Knowledge cutoff (EN/DE) | EN 1 Sept 2024; DE 1 Aug 2025 | 18 June 2026 |
Aleph Alpha also says the full stack, including mid-training and long-context adaptation, used nearly 24T tokens in total across stages, on 768 NVIDIA B200 GPUs, with infrastructure in Germany and Finland and teams under EU and German law. It says it has signed the EU GPAI Code of Practice.
Vendor Benchmarks, Labeled as Such
Aleph Alpha posts a large vendor table and claims Kolibri sits on a Pareto frontier for quality versus serving cost in English and German among the models it compared. That frontier claim is the company’s, not a third-party audit.
| Benchmark (Aleph Alpha-reported) | Kolibri | Kolibri Origin |
|---|---|---|
| AIME 2025 (EN) | 96.9 | 81.9 |
| AIME 2025 (DE) | 87.5 | 73.5 |
| GPQA diamond (EN) | 84.3 | 68.1 |
| LiveCodeBench v6 | 85.9 | 59.2 |
| HumanEval+ | 92.7 | 76.8 |
The company also highlights internal vertical proxies for German public sector, automotive, semiconductors, industrial drive technology, and aerospace, plus a Merlin-Arthur grounding and abstention protocol meant to push the model to say it does not know when evidence is missing. Those suites and the Merlin-Arthur scores are vendor-designed and vendor-scored.
Treat the leaderboard the same way you treat other self-reported open-weight drops this year, from Holo4’s OSWorld numbers to closed models that publish only their own harness results. Useful as a map of what Aleph Alpha wants you to notice. Not a substitute for your own evals.
Sovereignty Angle for Non-EU Buyers
Aleph Alpha’s sovereignty story has two parts. First, how it says the model was built: EU and German law, German and Finnish infrastructure, end-to-end pipeline control, and open weights under Apache 2.0. Second, how customers can deploy: on-prem or private infrastructure without sending internal data to a third-party inference host.
That matters outside Europe too. Teams in the USA, Canada, Australia, and India that sell into regulated EU clients, or that want an open-weight bilingual EN/DE stack they can host themselves, get a concrete option with a clear license. It does not automatically satisfy every AI Act or GDPR obligation. Compliance still depends on how you deploy, log, and govern the system.
Compared with closed frontier APIs such as Claude Sonnet 5.5 or limited-access drops like Gemini 4 Argon, Kolibri’s differentiator is less peak closed-model score and more control: weights, serving location, and bilingual German depth claimed by the vendor.
What It Means for Developers
If you run open-weight MoE stacks already, the practical checklist is short. Confirm GPU memory for 78B-total MoE with ~3.5B active, install Aleph Alpha’s inference plugin rather than assuming stock vLLM, and decide whether you need the 256k trained window or the 1M serve flags. Long-context flags can change memory and throughput sharply.
For product teams, the interesting bet is specialization. Aleph Alpha says it hill-climbed internal customer-proxy suites without training on customer data. That is a strong claim for regulated verticals. You still need your own RAG, abstention tests, and German administrative document suites before you trust it in production.
Research-minded teams may also compare Kolibri’s bilingual training story with other recent open or research releases, including Meta’s Muse Spark math workflow, which sits in a different niche but shows how vendors are packaging specialized capability claims this autumn.
What Is Confirmed vs Unconfirmed
Confirmed from Aleph Alpha’s 3 October release: the public announcement date, Apache 2.0 weights on Hugging Face as Kolibri-1, the MoE size figures Origin versus Kolibri, the context lengths and serve flags, the inference plugin requirement, the B200 count and Germany/Finland infrastructure statement, the EU GPAI Code of Practice claim, knowledge cutoff 18 June 2026, and the published sampling defaults.
Vendor-reported and not independently verified here: all benchmark scores, the Pareto frontier quality-versus-cost claim, internal vertical suite results, Merlin-Arthur grounding metrics, and any implied production readiness for public-sector or aerospace workloads.
Not disclosed on the page: commercial API pricing, hosted SaaS SKUs, or third-party safety red-team results.
Frequently Asked Questions
Is Kolibri free to download and use?
Aleph Alpha says the full weights are on Hugging Face under Apache 2.0. That covers the weights license. Serving still requires compatible hardware and the aleph-alpha-inference / vLLM plugin path described on the release page.
How big is Kolibri, and how much context does it support?
Aleph Alpha lists 78.1B total parameters with about 3.46B active per token. It says the model was trained to 262,144 tokens and can serve up to 1,048,576 with max-model-len and Hugging Face override flags.
Does Kolibri replace Claude or Gemini for general use?
Not as a like-for-like swap. Kolibri is an open-weight bilingual EN/DE MoE aimed at sovereign and on-prem deployments. Closed frontier APIs remain a different product class on pricing, tooling, and vendor support.
Are the high AIME and coding scores independently verified?
No in this article. The AIME, GPQA, LiveCodeBench, and HumanEval+ figures cited above are Aleph Alpha-reported from its own harnesses with high reasoning effort.
What is Merlin-Arthur?
Aleph Alpha describes Merlin-Arthur as an in-house grounding and abstention training protocol that tries to teach the model to withhold answers when evidence is missing. The published grounding scores are vendor metrics.
1 comment