Table of Contents
Aleph Alpha has released Kolibri, a 78B-parameter open-weight language model built for sovereign AI deployments in German and English.
The Heidelberg-based AI company formally announced Kolibri on October 5, after making the model available on October 3, Germany’s Unity Day. The model is released on Hugging Face under an Apache 2.0 license and is designed for organizations that want to run AI systems on infrastructure they control rather than relying entirely on third-party inference services.
Kolibri is a mixture-of-experts Transformer with 78.1B total parameters and 3.46B active parameters per token, according to Aleph Alpha’s technical blog and model card. That architecture gives the model the capacity of a larger system while activating only a smaller subset of parameters during inference, a design choice aimed at reducing serving costs for enterprise and government users.
Kolibri is part of Europe’s broader sovereign AI push, offering competitive models that can be inspected, deployed locally, governed under European legal requirements, and adapted for regulated workflows in sectors such as public administration, industry, and aerospace.
Why Kolibri Is Built Around Sovereign Deployment
Kolibri’s pitch goes beyond open weights. Aleph Alpha built the model for customers that need control over where AI runs, how data moves, and how compliance is documented.
Kolibri was developed and trained in Europe, using infrastructure in Germany and Finland. The model comes with documented data provenance and design decisions, with its technical report detailing measures for copyright, data protection, and EU AI Act requirements.
The release follows Aleph Alpha’s recent agreement with Cohere to build a transatlantic sovereign AI company. That transaction remains subject to regulatory approval, and Aleph Alpha continues to operate independently until closing.
Sovereignty also comes down to deployment. Kolibri’s model card lists a memory footprint of roughly 78 GB with FP8 weights, and the model can run on configurations including two A100 80 GB GPUs, two H100s, one H200, one B200, or one B300. The model is recommended for human-reviewed assistant systems, document processing, retrieval-augmented generation, structured extraction, coding, and agentic tool calling rather than unsupervised systems acting without review.
Aleph Alpha sees demand for AI systems embedded inside organizations’ own workflows, with enough control for compliance teams, IT departments, and operational leaders to evaluate how the model is used.
A German-English Model With Long-Context Support
Kolibri is bilingual. It supports German and English and is built with a 128K-vocabulary tokenizer optimized for German word structure while preserving English performance.
German-language text accounted for around 23% of Kolibri’s pre-training data. Its technical blog gives a more specific figure of 21.3% of pre-training tokens, or about 4.3T German tokens, alongside roughly 62% English and 14% code. Translation was used sparingly because translated text can carry the assumptions and phrasing of its source language.
The German-first emphasis fits public-sector and industrial deployments. Many enterprise AI systems are evaluated in English, but real operational work often involves local-language contracts, internal policies, technical manuals, and regulatory material. A model that handles German natively can reduce friction for organizations that need AI to reason over those documents without forcing workflows through English.
Kolibri also supports long-context use cases. The model was pre-trained at 16,384 tokens, mid-trained at 65,536 tokens, and extended through long-context training to 262,144 tokens. The model card says it has been validated up to 1,048,576 tokens, while recommending 262,144 tokens or less for serving efficiency and complex tasks.
The model’s architecture uses sliding-window attention across most layers and full attention every fifth layer. In practice, that means much of the model focuses on nearby context while selected layers attend across the broader input, helping keep long-context inference more manageable.
What Kolibri Can Do
Kolibri supports retrieval-augmented generation, agentic workflows, native tool calling, coding, structured extraction, and adjustable reasoning. Users can set reasoning effort levels of none, low, medium, or high, trading latency and cost against answer quality.
Kolibri is trained to abstain when supplied documents do not contain enough evidence for an answer. That behavior helps RAG systems, where enterprise users often want a model to answer from approved source material rather than improvise from its pre-training.
Kolibri is suitable for internal knowledge tools, question-answering systems over an organization’s own material, drafting systems, and advisory decision-support tools where a person reviews the output before action is taken.
Vendor-reported benchmark results place Kolibri strongly across several categories. Aleph Alpha reports scores including 96.9 on AIME 2025, 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6, 66.4 on SWE-Bench Verified, and 94.7 on Tau2-Bench Telecom. The company says Kolibri matches models with up to four times its active parameter count, including Nemotron 3 Super, across math, coding, grounding, and long-context tasks.
How Aleph Alpha Trained the Model
Kolibri was trained on 768 NVIDIA B200 GPUs across three stages. The model used 20T pre-training tokens over 21 days, followed by 3.44T mid-training tokens and 201B long-context tokens.
Kolibri is the second model from its internal “Model Factory,” following Kolibri Origin, an internal 30.6B-parameter predecessor. Compared with Kolibri Origin, the new model increases total parameters from 30.6B to 78.1B, expands the context window from 65K to as much as 1M tokens in supported serving configurations, raises the tokenizer vocabulary from 96K to 128K, and moves from 128 experts to 384 experts.
The training pipeline handled 38 unplanned interruptions during pre-training, roughly one per 10,000 GPU-hours, with automated recovery from hardware faults or connection timeouts. The company presents that as evidence that its model-building process has matured from research experimentation into repeatable infrastructure.
Compliance is also part of the pitch. All training data was screened against a blocklist of more than 4.5 million URLs, including sources from the European Commission’s Piracy Watch List. Each third-party dataset was assessed for license terms, lawful sourcing, and opt-out adherence before use.