OpenWALDO aims to blow the doors off proprietary AI training models
Contributors wanted: 167B transparent tokens have a long way to go against AI giants' trillions
A new project called OpenWALDO aims to create an open-source AI training dataset, much like open-source software. The initiative, founded by Gregory Kurtzer, aims to make training data more transparent than many proprietary models currently have. CentOS and Rocky Linux creator Kurtzer is leading the project, which received funding from his AI infrastructure company, CIQ.
Kurtzer explained that the goal is to bring the open-source model to AI model design. CIQ argues that open-weight models keep the training data hidden due to its origins, such as copyrighted data, responses from other models, and user-generated content. This lack of transparency can lead to hidden data content that may compromise the security of customer software stacks.
By creating a single, shared, public training data set, the OpenWALDO team hopes to improve efficiency, reduce duplicate work, and ensure that every improvement to the dataset benefits future models. With the rising prominence of open AI models, businesses are wondering why they should pay for AI services they don't own or control.
OpenWALDO has collected 167.3 billion reference tokens from various sources, but it is still a small fraction of the data used to train frontier AI models. The project's viability and impact are yet to be seen, but Kurtzer believes that open-source training datasets are crucial for the future of AI.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.