DiffuseDrive is filling the data gaps holding back Physical AI
Training AI to operate in the physical world requires enormous amounts of data, but many situations autonomous systems need to recognise are rare, dangerous, or hard to capture in the real world. Hung...
Training AI to function in physical environments demands vast quantities of data, yet numerous scenarios autonomous systems must recognize are infrequent, hazardous, or challenging to replicate in the real world. Hungarian startup DiffuseDrive addresses this data gap by creating synthetic training data for physical AI systems, enabling companies to expose their models to scenarios their current datasets lack.
In conversation with DiffuseDrive's team members Gabor Vecsei, Vice President of AI Engineering, and Daniel Schmid, AI Research Lead, a reporter sought to understand more about the company's innovative approach.
Vecsei leads the engineering of DiffuseDrive's AI systems and the transition of research into scalable production technology. Prior to joining the company, Vecsei worked in machine learning and AI engineering, including overseeing an ML research department that expanded from a small team to more than 30 members. He earned his degree in applied computer science from the Budapest University of Technology and Economics, specializing in computer science.
Schmid serves as DiffuseDrive's AI Research Lead, working at the intersection of the company's research and production engineering. Hailing from Germany, he has been specializing in machine learning since 2020. During his tenure at DiffuseDrive, he has contributed to various academic projects, including research experience in Korea and collaboration with the University of Oxford. Schmid completed his master’s degree at the Technical University of Munich.
The physical AI domain faces a data problem due to the scarcity and difficulty in collecting data, as well as the potential danger or complexity of reproducing certain scenarios. Schmid emphasizes that data is crucial for AI development, stating, "With language models, there is an enormous corpus of textual data available on the internet.
Physical AI is very different. Data is scarce and difficult to collect, and the scenarios you're interested in can be extremely hard—or dangerous— to reproduce, particularly in defense environments."
To tackle this issue, DiffuseDrive has developed a system capable of generating missing data points and integrating them directly into customers' datasets, thereby enhancing downstream applications such as autonomous drones, vehicles, and mining systems. For instance, a defense company with a perception system monitoring a coastline might find its data and models incomplete, as merely pointing a camera at the coastline cannot expose the system to every scenario it needs to recognize.
DiffuseDrive's platform can analyze a company's existing solution, identify data gaps, and generate data to fill those gaps. Schmid explains that one of the most successful applications of this approach lies in identifying rare edge cases, particularly those relevant to defense organizations.
Written by urgent.news from Tech.eu's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.