The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Aleph Alpha · Kolibri-1 · Hugging Face · vLLM

Aleph Alpha releases Kolibri‑1, a 78B open‑weights model with 3.5B active

The German lab's Apache 2.0 model reads 262,144 tokens, was trained on 20 trillion tokens of English, German and code, and its card says it runs on two 80GB A100 or H100 GPUs.

Aleph Alpha published Kolibri-1 on Hugging Face on 3 October, in two checkpoints: an FP8 build and a full-precision BF16 one. According to the model card, it is a 50-layer mixture-of-experts transformer with 78.1 billion parameters in total, of which 3.46 billion activate for each token. The weights carry an Apache 2.0 licence.

The card says the model has 384 routed experts per layer, picks six per token and adds one shared expert. Its native context is 262,144 tokens, which Aleph Alpha says stretches to 1,048,576 without position scaling. Post-training included four reasoning-effort settings, from none to high, plus tool calling and structured output, the card says.

Training was aimed at German and English. Aleph Alpha says it used 20 trillion tokens, about 62.5 per cent English, 23.9 per cent German and 13.6 per cent code, then 3.44 trillion tokens of mid-training and 201 billion for long context. A custom tokenizer compresses German at 4.7 bytes per token, against 4.2 for English, according to the card.

On benchmarks, the card compares Kolibri with dense models of 27 to 70 billion parameters. It reports an English overall score of 75.5 against 80.2 for the best baseline, and German 70.8 against 79.9. On maths it claims about 96.5, close to the best baseline's 97.8. These are Aleph Alpha's own figures, run at high reasoning effort.

The FP8 weights occupy about 78GB, and the card gives two 80GB A100 or H100 GPUs as the minimum. It says to serve the model with vLLM and Aleph Alpha's own inference package, using a custom reasoning and tool-call parser. The licence covers weights and configuration only, and Aleph Alpha keeps rights to its code and training methods. No independent tests have been published yet.

Sources 2 sources

  1. Source Aleph Alpha (Hugging Face model card)
  2. Source Aleph Alpha on X