{
  "id": 1597060,
  "title": "Object localization in images via fusion of SFM and YOLO",
  "url": "https://urgent.news/2026/08/18/object-localization-in-images-via-fusion-of-sfm-and-yolo",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-18T00:00:00.000Z",
  "source": {
    "name": "Scientific Reports",
    "slug": "scientific-reports",
    "url": "https://www.nature.com/articles/s41598-026-66325-3"
  },
  "original_language": "en",
  "account": "In recent years, the need for accurate spatial positioning of objects captured in geospatial environments has grown due to the rapid advancement of computer vision, unmanned aerial vehicle technology, and autonomous driving systems. Traditional object detection techniques, like the YOLO series algorithms, primarily identify object categories and provide two-dimensional bounding boxes, but they fall short of offering precise estimations of objects' spatial positions within three-dimensional space. To address this limitation, a novel framework has been developed that integrates the Structure from Motion (SfM) algorithm with an enhanced YOLOv8 object detection model, aiming to achieve high-precision localization of objects of interest in image sequences.\n\nThe process begins by employing the SfM algorithm to conduct three-dimensional reconstruction on a series of multi-view images, resulting in sparse point clouds equipped with spatial coordinates. Concurrently, a YOLOv8 model enhanced with a hybrid attention mechanism (comprising both Global Average Pooling (GAP) and Channel Attention (CA)) is utilized for precise detection of target objects within the image sequences. Ultimately, by projecting the 3D point cloud back onto the image plane and conducting spatial constraint matching between the two-dimensional detection boxes, a mapping relationship is established between the two-dimensional detection results and their corresponding three-dimensional spatial coordinates.\n\nThe efficacy of this proposed method has been validated through ablation studies conducted on the VisDrone dataset, which confirm the effectiveness of the hybrid attention mechanism incorporated within the YOLOv8 model. Furthermore, localization experiments using vehicle image sequences captured by UAVs demonstrate that the proposed approach can accurately determine the geographical coordinates of the objects captured. This research not only broadens the functional scope of conventional object detection techniques but also offers technical support for various application scenarios, including UAV visual localization and intelligent traffic monitoring. The study was supported by the Natural Science Foundation of Guangdong Province, under Grant No. 2024A1515011569, and was conducted by researchers from the School of Geography at South China Normal University, the Beidou Research Institute at South China Normal University, and the School of Artificial Intelligence at Guangdong Mechanical and Electrical Polytechnic. The article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, mandating the provision of credit to the original author(s) and source, a link to the Creative Commons license, and indication of any modifications made to the licensed material. Any use of adapted material derived from this article is subject to the specified conditions. The images or other third-party materials included in this article are subject to the article's Creative Commons license, unless otherwise indicated in the credit line.",
  "summary": "Scientific Reports, Published online: 18 August 2026; doi:10.1038/s41598-026-66325-3 Object localization in images via fusion of SFM and YOLO",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}