AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Aleph Alpha has released Kolibri, an English-German mixture-of-experts model with 78 billion total parameters, 3 billion active parameters and a context window of up to 1 million tokens. The company says its full weights are available on Hugging Face under Apache 2.0, and presents the model for regulated and mission-critical uses. Its benchmark results and claims about sovereignty and customer-sector performance come from company-published material; independent verification is not provided in the supplied report.

Aleph Alpha has released Kolibri, an English-German mixture-of-experts model with 78 billion total parameters, of which 3 billion are active, and a context window of up to 1 million tokens. The company says the full weights are downloadable from Hugging Face under the Apache 2.0 license, a release it positions for organizations that need control over deployment and data, including public-sector, industrial and aerospace users.

Aleph Alpha describes Kolibri as a specialized language model for sovereign, mission-critical work in regulated fields. It says development focused on German language performance, reasoning, mathematics and agentic behavior, along with capabilities requested by customers. The report does not detail the model’s training data or provide independently verified assessments of its suitability for particular regulated workflows.

The company says Kolibri was trained using a pipeline developed and tested with Kolibri Origin, an earlier model with 30 billion total parameters, 3 billion active parameters and a 65,000-token context window. Aleph Alpha reports that the pipeline covered data curation, pre-training, post-training and evaluation, and supported hundreds of ablation experiments. It also says training could continue without manual intervention after hardware failures or lost data connections.

Aleph Alpha’s published benchmark table reports results across math, knowledge, coding, long-context and agentic tasks. For example, it lists Kolibri at 96.98 on AIME 2025 and 85.9 on LiveCodeBench v6, using a shared 0–100 scale. The company says Kolibri performs competitively with models that have as much as four times its active parameter count. These are the company’s reported results; the supplied material does not establish independent replication or provide enough methodology to assess every comparison.

At a glance
announcementWhen: Announced March 10, 2026, according to…
The developmentAleph Alpha announced Kolibri, an open-weight English-German model aimed at regulated and mission-critical customers.

Deployment Control for Regulated Users

The release gives organizations another option to run a capable language model using downloadable weights, rather than relying only on a hosted inference service. Aleph Alpha says Kolibri’s mixture-of-experts design, with 3 billion active parameters from 78 billion total, is intended to balance model quality and serving costs. Whether that trade-off holds for a given customer will depend on hardware, workload and deployment configuration.

For public agencies and companies handling sensitive information, the ability to deploy a model on-premises can affect data handling and compliance planning. Aleph Alpha presents deployment freedom and intellectual-property safety as parts of its sovereignty approach. Open weights can support local control, but do not by themselves establish that a deployment meets a specific legal or security requirement; organizations still need to evaluate the model, its supply chain and their own systems.

The company also argues that smaller active parameter use can make the model practical for enterprise operations while retaining broad capabilities. Its claims of economic impact and measurable return on investment are a stated goal, not a result demonstrated in the announcement. Real-world costs and benefits will need to be assessed in customer settings.

Amazon

Open-weight language model download

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Kolibri Origin to Release

Aleph Alpha frames Kolibri as the next step in an ongoing model-training effort, rather than a standalone release. Its reported predecessor, Kolibri Origin, used the same number of active parameters but had fewer total parameters and a much shorter context window. The company says work on the training pipeline helped shorten the time between the two releases.

The report also distinguishes public benchmarks from tests designed around particular industries. Aleph Alpha says it created internal evaluation suites for areas including the German public sector, aviation, manufacturing and automotive. It reports customer-proxy scores of 0.72 to 0.99 for an automotive-supplier task, 0.35 to 0.80 for semiconductors, and 0.54 to 0.70 for the German public sector. The supplied material does not define the scoring scale or provide independent validation of those internal evaluations.

Aleph Alpha says paired synthetic training environments were used to improve performance against these evaluations without training on customer data. It points readers to a technical report for further details. The announcement identifies October 3, German Reunification Day, as the intended release occasion, while the report itself is dated March 10, 2026; the supplied material does not explain that date discrepancy.

“Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active.”

— Aleph Alpha

Amazon

AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Results and Deployment Details

The announcement is based on Aleph Alpha’s own report. The supplied source does not include independent benchmark testing, full evaluation methodology, or an outside assessment of the company’s comparisons with other models. It also does not give detailed information on training data, hardware requirements, inference costs, or the practical limits of the 1-million-token context window.

It remains unclear how the model will perform across specific production tasks, languages beyond English and German, and real-world regulated workflows. The report makes broad claims about sovereignty, supply-chain integrity and intellectual-property safety, but does not provide enough detail in the supplied material to verify how those properties are implemented or to determine whether a particular deployment meets a customer’s obligations. The difference between the report’s March 10, 2026 date and its reference to a release on German Reunification Day also remains unexplained.

Amazon

enterprise AI model management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Technical Report and Customer Testing

Aleph Alpha directs readers to a technical report for fuller information on how Kolibri was built and evaluated. Further details on methodology, model access and deployment requirements would help customers assess the benchmark claims and estimate serving costs. The company’s stated next test is real-world use by customers in government and industry; the supplied announcement does not name deployments, release a customer adoption timetable or set out an independent evaluation plan.

Amazon

on-premises AI inference server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Kolibri?

Kolibri is an English-German mixture-of-experts language model released by Aleph Alpha. The company says it has 78 billion total parameters, 3 billion active parameters and supports context lengths of up to 1 million tokens.

Can the model weights be downloaded?

Aleph Alpha says the full weights are available on Hugging Face under the Apache 2.0 license. The announcement does not provide additional technical details about download size or deployment requirements.

Who is Kolibri designed for?

The company describes it as a model for regulated and mission-critical work, including public administration, industrial sectors and aerospace. Actual performance and compliance suitability will depend on the specific application and deployment.

Are the benchmark results independently verified?

The supplied material presents results reported by Aleph Alpha. It does not include independent replication or enough detail to establish how all comparisons were conducted.

What is still unknown about the release?

The announcement does not fully specify training data, hardware needs, inference costs or independent testing. It also does not explain the difference between the report’s March 10, 2026 date and its reference to a release on German Reunification Day.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is Grok 4.6 The Future Of AI? SpaceXAI’s New Model Boosts Long-Form Agent Performance

SpaceXAI’s Grok 4.6 claims a 500K context window for long-form tasks, but performance and availability details remain unverified.

Leanstral 1.5: Proof Abundance For All

Leanstral 1.5 introduces proof abundance, making proof generation more accessible. The update aims to democratize verification tools for users worldwide.

A Closer Look At How Researchers Used Claude To Hack OpenAI’s Systems

Security researchers reportedly used Anthropic’s Claude AI to breach an OpenAI product, raising concerns over AI-enabled cyberattacks and industry safety protocols.

Why OpenAI and Anthropic may struggle to float

OpenAI and Anthropic may struggle to raise funds through an initial public offering due to market and internal hurdles, experts suggest.