Opens in a new tab

Aleph Alpha’s Kolibri-1 runs on one GPU and edges out same-size open models in German

05-10-2026

Aleph Alpha released Kolibri-1 on 3 October 2026, an open German-English model under Apache 2.0. Only 3.46 billion of its 78.1 billion parameters work per token, so it runs on one B200 or two H100 cards. On Aleph Alpha's own tests it scores 70.8 in German, one point above Qwen3.5 35B-A3B, but trails Qwen on German document answering and agent tasks. The model card lists no Dutch support.

Written by:

Diederik Knol

Online Marketeer at DataNorth | Passionate about AI

Sign up for our Newsletter

Published: 5 October 2026

Aleph Alpha released Kolibri-1 on 3 October 2026, an open-weight model for German and English that anyone can download and run under the Apache 2.0 licence. It fits on a single Nvidia B200 or two H100 cards and scores 70.8 on the company’s German test mix, one point ahead of Alibaba’s similar-sized Qwen3.5. For European companies, the bigger point is where it runs: on hardware they control.

Kolibri leads Qwen by one point in German and loses on company documents

Kolibri is a mixture of experts, a model where only some parts work per token. Of its 78.1 billion parameters, 3.46 billion do the work for each token, which keeps it cheap to run. The fair comparison is therefore other models with about 3 billion active parameters. Against those, the lead is real but small.

Kolibri wins most rows against Qwen3.5, but the gaps are narrow and two rows go the other way.

What is measuredKolibri-1Qwen3.5 35B-A3B
Active parameters per token3.46 billionAbout 3 billion
Overall score, German (Aleph Alpha’s test mix)70.869.8
Overall score, English (Aleph Alpha’s test mix)75.574.7
German industry RAG (answers from company documents)67.570.0
Code average, English89.385.0
Terminal-Bench 2.1 (agent tasks in a terminal)27.739.7

All figures come from Aleph Alpha’s own model card. The company says it ran every model on the same setup, using its eval-framework and the Harbor harness for Terminal-Bench, with Kolibri at its highest reasoning setting. Nobody has rerun them yet. The overall scores average Aleph Alpha’s own mix of benchmarks, so they rank models within this table but do not line up with other vendors’ charts.

The same table shows the ceiling. Qwen3.8 27B, a dense model that uses all its parameters for every token, scores 79.9 in German and 80.2 in English. It costs more to run, but if quality matters more than hardware, Kolibri is not the strongest open option. For agents that run commands on their own, the Terminal-Bench row points elsewhere too.

A single B200 or two H100 cards are enough to run it

Aleph Alpha ships the weights in FP8, a compressed 8-bit format, at about 78 GB. That is the main practical gain: a model at this level that one server can host.

  • Model: Aleph-Alpha/Kolibri-1 (FP8) and Kolibri-1-BF16 on Hugging Face
  • Licence: Apache 2.0, commercial use allowed
  • Minimum hardware: one H200, B200 or B300, or two A100 80GB or H100 SXM5 cards
  • Context window: 262,144 tokens native, up to 1,048,576 by extrapolation
  • Reasoning: four effort levels from none to high, plus tool calling
  • Hosted API and price: none listed in the model card

Read the million-token figure with care. Aleph Alpha trained the model on 262,144 tokens, stretches it beyond that, and itself recommends staying at or below the native length. At one million tokens it scores 63.2 on the RULER long-context test. Treat the larger number as a ceiling, not a working size.

The paperwork EU buyers ask for is already published

Aleph Alpha has signed the EU’s Code of Practice for general-purpose AI models and published a data summary next to the weights. The model card lists the training data mix (about 62.5% English, 23.9% German and 13.6% code), the compute used and a knowledge cutoff of 18 June 2026. That shortens a procurement review under the EU AI Act.

One fact to keep in view: Aleph Alpha is merging with Canada’s Cohere, a deal SiliconANGLE reported on 16 September. The weights are yours under Apache 2.0 whatever happens to the company, but buyers who chose Aleph Alpha for its German ownership should note the change.

Test Kolibri for German documents that must stay in-house, not for Dutch

If your organisation works with German contracts, tickets or regulations and cannot send them to a US cloud, Kolibri deserves a test this month. It runs on one server, the licence is clean and the documentation the EU AI Act expects is there. Start with retrieval: give it fifty German documents and questions you already know the answers to, and run Qwen3.5 35B-A3B on the same machine. That is the row where Aleph Alpha’s own table puts Kolibri behind.

For Dutch-language work, hold off. The model card names German and English and nothing else, and no published test measures Dutch. Dutch is close enough to German that simple tasks may work, but you would be the first to find out. For multi-step agents, pick a model with a stronger Terminal-Bench score.

For more information, visit the official announcement of Kolibri-1 in the Kolibri-1 model card.

Add DataNorth AI to your Google favorites