Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenWALDO aims to blow the doors off proprietary AI training models

Contributors wanted: 167B transparent tokens have a long way to go against AI giants' trillions

OpenWALDO aims to blow the doors off proprietary AI training models

A fresh initiative aims to create an open-source AI training dataset for collaborative contributions, similar to open-source software. This project, known as OpenWALDO, is led by Gregory Kurtzer, the founder of CentOS and Rocky Linux, with support from his AI infrastructure firm, CIQ. The OpenWALDO team seeks to bring the open-source approach to AI model design, a concept that has yet to be fully implemented in the industry.

Currently, many open-weight models lack transparent training data, which hampers their true openness and collaborative potential. Kurtzer emphasizes that OpenWALDO aims to address this issue by fostering a community-driven approach to AI model development. CIQ highlights that open-weight models conceal their training data origins due to factors like copyrighted data, responses from other models, and user-generated content, making it difficult to assess the data's license and consent.

The team argues that a universally accessible training dataset would not only enhance training efficiency but also enable model improvements to benefit the entire industry. With the cost of AI services increasing and their true control remaining elusive, open AI models have gained traction, particularly those emerging from China, which are now challenging the capabilities of closed-source models like ChatGPT and Claude.

While some frontier labs warn about potential security and misuse risks associated with open-weight models, Kurtzer draws parallels to the open-source software movement, asserting that Linux overcame similar concerns through its inspectable, forkable, and community-validated nature. OpenWALDO currently comprises 167.3 billion reference tokens sourced from government records, academic papers, mailing lists, and public domain literature, a modest beginning compared to the vast amounts of data used in frontier AI models.

However, the project invites contributions and usage, with more information available on its website and GitHub page.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

More from Wednesday 12 August →