Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

Object localization in images via fusion of SFM and YOLO

Scientific Reports, Published online: 18 August 2026; doi:10.1038/s41598-026-66325-3 Object localization in images via fusion of SFM and YOLO

In recent years, the need for accurate spatial positioning of objects captured in geospatial environments has grown due to the rapid advancement of computer vision, unmanned aerial vehicle technology, and autonomous driving systems. Traditional object detection techniques, like the YOLO series algorithms, primarily identify object categories and provide two-dimensional bounding boxes, but they fall short of offering precise estimations of objects' spatial positions within three-dimensional space.

To address this limitation, a novel framework has been developed that integrates the Structure from Motion (SfM) algorithm with an enhanced YOLOv8 object detection model, aiming to achieve high-precision localization of objects of interest in image sequences.

The process begins by employing the SfM algorithm to conduct three-dimensional reconstruction on a series of multi-view images, resulting in sparse point clouds equipped with spatial coordinates. Concurrently, a YOLOv8 model enhanced with a hybrid attention mechanism (comprising both Global Average Pooling (GAP) and Channel Attention (CA)) is utilized for precise detection of target objects within the image sequences.

Ultimately, by projecting the 3D point cloud back onto the image plane and conducting spatial constraint matching between the two-dimensional detection boxes, a mapping relationship is established between the two-dimensional detection results and their corresponding three-dimensional spatial coordinates.

The efficacy of this proposed method has been validated through ablation studies conducted on the VisDrone dataset, which confirm the effectiveness of the hybrid attention mechanism incorporated within the YOLOv8 model. Furthermore, localization experiments using vehicle image sequences captured by UAVs demonstrate that the proposed approach can accurately determine the geographical coordinates of the objects captured.

This research not only broadens the functional scope of conventional object detection techniques but also offers technical support for various application scenarios, including UAV visual localization and intelligent traffic monitoring. The study was supported by the Natural Science Foundation of Guangdong Province, under Grant No. 2024A1515011569, and was conducted by researchers from the School of Geography at South China Normal University, the Beidou Research Institute at South China Normal University, and the School of Artificial Intelligence at Guangdong Mechanical and Electrical Polytechnic.

The article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, mandating the provision of credit to the original author(s) and source, a link to the Creative Commons license, and indication of any modifications made to the licensed material. Any use of adapted material derived from this article is subject to the specified conditions.

The images or other third-party materials included in this article are subject to the article's Creative Commons license, unless otherwise indicated in the credit line.

Written by urgent.news from Scientific Reports's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at nature.com →

More in AI

More from Tuesday 18 August →