From lab to life: Why translating AI advances to the real world is a major challenge
All the training data in the world can’t account for every possibility when AI systems interact with people and the environment.
Bringing artificial intelligence out of the laboratory and into the real world presents significant hurdles despite impressive advancements. Systems that excel in controlled lab environments often struggle when implemented in homes, hospitals or city streets. This occurs primarily due to three key challenges: generalization, human behavior, and resource limitations.
The first challenge - the "generalization gap" - arises because most AI vision systems are trained on pristine, well-lit images usually sourced from motion capture studios or daytime recordings. However, real-world conditions often differ greatly from these idealized settings. For example, pose estimation, which involves detecting a person's body joints to understand movement, performs admirably under bright lighting but suffers greatly in low or uneven lighting.
The root cause lies in the model's reliance on clean, consistent data during training. It fails to adapt well to messy, unpredictable or dimly lit environments.
To bridge this gap, researchers have begun exploring techniques like unsupervised domain adaptation. This method involves adapting a model trained on well-lit data to handle low-light conditions without needing manual labels in low-light images. However, even with these advancements, the broader challenge remains largely unsolved.
Another major hurdle is the vast diversity of human behavior, particularly when interacting with objects. Known as human-object interaction detection in computer vision, teaching an AI system to recognize actions like cutting a tomato or passing a basketball necessitates not just object detection but understanding context and intent.
The scale of objects and interaction methods is immense, making it impractical to collect and label data for every conceivable combination. A breakthrough could come from using large vision-language models to identify object parts supporting specific actions, even without explicit training on labeled examples for that task.
The third challenge - resource limitations - stems from the difficulty in collecting and labeling data, especially under challenging conditions. For instance, accurately labeling a person's joints in low light is inherently difficult for both AI systems and human annotators. This scarcity of data makes it tough to train models that can reliably function across various environments.
Written by urgent.news from The Conversation's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.