Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

OpenWALDO aims to blow the doors off proprietary AI training models

Contributors wanted: 167B transparent tokens have a long way to go against AI giants' trillions

OpenWALDO aims to blow the doors off proprietary AI training models

A new project, OpenWALDO, aims to create an open-source AI training dataset that anyone can contribute to, similar to open-source software. The project, led by Gregory Kurtzer, founder of CentOS and Rocky Linux, is backed by CIQ, Kurtzer's AI infrastructure company. OpenWALDO seeks to bring the open-source ethos to AI model design, aiming to make training data more transparent and collaborative.

Kurtzer explains that current open-weight models still have closed-source training data, which hinders true openness. CIQ argues that these open-weight models keep their foundations secret due to the nature of the data used, which can include copyrighted material, responses from other models, and potentially non-consensually gathered user-generated content.

By creating a single, shared set of public training data, OpenWALDO aims to improve efficiency, reduce duplicate work, and ensure every improvement to the dataset benefits future models. With steep prices and unclear ROI, open AI models have gained popularity, prompting businesses to question the value of paying for services they can't truly control.

While the OpenWALDO dataset currently contains 167.3 billion reference tokens, it is still a fraction of the data used in training frontier AI models. As the AI industry continues to embrace closed training data, the success of OpenWALDO remains to be seen.

Written by urgent.news from The Register Software's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

More from Wednesday 12 August →