From lab to life: Why translating AI advances to the real world is a major challenge
AI systems more reliable in everyday settings is a challenge worth solving, not just to improve gadgets and apps, but because AI has the potential to support human well-being, whether that means safer streets, smarter healthcare or more secure homes
Translating artificial intelligence (AI) research into real-world applications continues to pose significant challenges, despite the substantial advancements made in recent years. While AI systems perform well in controlled laboratory settings, their effectiveness can be severely limited when deployed in homes, hospitals, or city streets. This disparity becomes particularly evident in vision-based systems, which hold promise for enhancing safety, monitoring health, and aiding rehabilitation.
One notable application of AI technology is in evaluating a patient's gait or rehabilitation progress through camera-based systems. However, ensuring the reliable performance of assistive robots in homes, such as detecting when an elderly individual is struggling to stand or reach for an object, presents a major hurdle. These robots must rely on vision systems to perform reliably in complex and unpredictable environments, especially in low visibility conditions.
The primary challenges in translating AI advances from the lab to the real world can be summarized in three key areas: generalisation, human behavior, and resources. AI vision systems are typically trained on high-quality, well-lit images obtained from curated sources like motion capture studios or daytime recordings. However, when deployed in real-world settings like dimly lit homes, hospital rooms, or nighttime streets, these systems often struggle to adapt to the enormous variability of the real world.
For instance, human pose estimation – a task that involves detecting key points on a person's body to understand their movement – has shown impressive accuracy in well-lit environments. Yet, performance can plummet when faced with low or uneven lighting. The issue lies in the fact that these models are trained on clean, consistent data and lack the ability to generalise effectively to challenging conditions.
While collecting more training data would help, it is particularly difficult due to factors such as low light obscuring visual information and varying camera conditions.
Another major challenge lies in the vast variety of human behaviors, especially when interacting with objects. Known as human-object interaction detection, this involves recognising actions like cutting a tomato or passing a basketball, which necessitates not only object detection but also an understanding of context and intent. The scale of this problem is immense, as there are countless objects and ways humans interact with them. Collecting and labeling data for every possible combination is impractical.
To tackle the generalisation problem, researchers are exploring various solutions. One approach is unsupervised domain adaptation, where a model trained on well-lit data is adapted to handle low-light conditions without requiring manual body-joint labels for low-light images. By generating realistic low-light training images and designing the system to balance uncertain visual evidence with learned knowledge of human body structure, AI models can achieve substantial improvements in performance on low-light benchmarks.
Similarly, generalising to unseen human-object interactions presents a significant challenge. Unlike low-light pose estimation, where the system recognises a familiar movement under unfamiliar conditions, human-object interaction detection deals with new action-object combinations. Identifying the parts of an object that support particular actions, without training the system on labeled examples for that specific task, requires innovative approaches.
Large vision-language models, which combine computer vision and natural language processing, have shown promise in identifying relevant object parts for specific actions.
However, even when researchers know how to enhance models' robustness, applying these improvements can be challenging, particularly in academic or smaller research settings. Today's most powerful AI systems rely on massive computational resources, such as data centers containing numerous specialised computer chips. These resources are often prohibitively expensive for university labs, limiting the accessibility and widespread adoption of advanced AI systems.
In conclusion, while AI advances have the potential to transform various aspects of our lives, translating these breakthroughs from the lab to the real world remains a formidable challenge. Addressing issues related to generalisation in vision-based systems, understanding complex human behaviours, and overcoming resource limitations are crucial steps towards developing AI technologies that can reliably operate in diverse and unpredictable environments.
Written by urgent.news from The Hindu Health's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.