Urgent.News

What's breaking now, across thousands of outlets.

AI

Kolibri: A Sovereign Open-Weight Model

On the day of German reunification, the release of a new English-German mixture-of-experts transformer model, Kolibri, marked a significant milestone. Kolibri boasts an impressive 78 billion total parameters, with 3 billion active parameters, and supports context lengths of up to one million tokens. The model is open-source, downloadable with full weights on Hugging Face, and can be used under the Apache 2.0 license terms.

Kolibri is the culmination of continuous iterations in the model training process. Initially, a model training pipeline was built and validated by creating Kolibri Origin, a 30 billion total, 3 billion active parameter model with a 65k token context window. The same pipeline was then used to create Kolibri, which enabled running hundreds of ablation experiments and maintaining stable pre-training without the need for human intervention in cases of hardware failures or dropped data connections.

Regular monitoring of training metrics and custom benchmarks ensured the model's performance remained optimal.

Kolibri is a specialized language model designed for sovereign, mission-critical work in regulated areas, including public administration, industrials, and aerospace. It has been optimized for German, focusing on reasoning, math, agentic behavior, and other capabilities crucial for production. This specialization aimed to enhance performance in specific use cases for customers, enabling them to monitor the economic impact of Kolibri and ensure a growing return on investment (ROI) over time.

One of the key aspects of Kolibri's development is its sovereignty, which combines two dimensions: the model's construction and its transferability to customers. The model offers full supply-chain integrity, accounting for every decision made during data ingestion, pre- and post-training, and final evaluations. Transparency is a core principle, with customers having full freedom of deployment and intellectual-property safety, inheriting compliance as a natural property of the model.

Kolibri's small and efficient size allows customers to run it efficiently on-premise, without sending internal data to third-party inference services. The model strikes a balance between model capability and deployment costs, achieving the best quality versus serving cost trade-off on the Pareto frontier for both English and German.

Compared to other models with similar capabilities, Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super, in terms of quality across various tasks, including math, coding, grounding, and long-context tasks.

To ensure Kolibri's performance meets the specific needs of customers in various sectors, the model was developed with a focus on domain-specific language, regulatory, and procedural realities. A dedicated evaluation suite, tailored to the skills, workflows, and edge cases required in sectors such as the German public sector, aviation, manufacturing, and the automotive industry, was created.

These evaluations were performed using synthetic training environments, allowing Kolibri to improve without ever training on customer data.

Kolibri was trained with abstention data and the Merlin-Arthur protocol, enabling it to confidently say "I don't know" when the answer is not present in the context. This feature was highly valued by customers, and Kolibri's performance in tracking and validating abstention accuracy was continuously monitored. The model's bilingual German/English design, with 21.3% of pre-training tokens being German, was achieved through the use of a native German and English tokenizer and a sparing use of translation (6% overall) to include organic German data throughout the training process.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at aleph-alpha.com →

More in AI

Designing a multilingual clinic receptionist: booking, failures and handoff

Originally published as a Fonix.AI practical guide by AntEngage. This republication uses the same guide content. When the front desk cannot pick up An AI receptionist for clinics answers routine…

  • Front desk determines call routing for AI receptionist
  • AI gathers clinic info like hours, services, doctor schedules
  • AI handles bookings, offers callbacks if slots unavailable

SmartCart AI: The Automated Grocery Planner for Stress-Free Household Management

SmartCart AI: The Automated Grocery Planner for Stress-Free Household Management This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend ( https://dev.to ).

  • SmartCart AI automates grocery planning to reduce household stress
  • Predictive and emergency tracks manage monthly staples and sudden shortages
  • AI integrates with delivery platforms for price optimization and payment flexibility

I tried ChatGPT’s new try-it-on shopping function. It worked for daily outfits, but couldn't handle a Halloween costume.

I tried ChatGPT's virtual try-on feature with jewelry, daily outfits, a wedding dress, and a fairy Halloween costume. Here are the function's limits.

  • Virtual try-on feature works for accessories and simple clothing
  • Complex outfits and specific styles like wedding dresses cause inaccuracies
  • AI excels at finding matching accessories for users' selfies

QA Isn’t AI Evaluation

An AI agent prepares an internal report. The report has the right format. The figures are accurate. The conclusions sound reasonable. But the agent used a source it wasn’t allowed to access.

  • AI evaluation differs from traditional QA
  • Evaluating AI requires defining right behavior and criteria
  • Human review needed for complex decision judgments

More from Sunday 4 October →