How Two-Stage Object Detectors Went From 47 Seconds to Real-Time-Adjacent
Object detection has two broad architectural families: one-stage detectors that predict boxes and classes directly in a single pass, and two-stage detectors that first propose candidate regions, then classify and refine them. Two-stage detectors were the dominant approach from 2014 to 2017 and are still the most accurate choice on many benchmarks when latency isn't the constraint. The interesting…
Two-stage object detectors, which emerged in 2014, transformed the field by offering significantly higher accuracy compared to their one-stage counterparts. The first of these, R-CNN, utilized selective region proposals generated by a third-party algorithm, resized them before passing through a deep convolutional neural network (CNN) for feature extraction.
A Support Vector Machine (SVM) then classified these features, with a separate regressor refining the bounding boxes. Although R-CNN improved detection accuracy by 30% over previous methods, it suffered from significant latency - a full 47 seconds per image due to the repeated CNN computations for each proposal.
Fast R-CNN addressed this issue by eliminating redundant feature computations across proposals, instead pooling those features once from the entire image. This reduction made the network 213 times quicker than R-CNN during testing, while maintaining comparable accuracy. However, Fast R-CNN still relied on the non-learned Selective Search for region proposals, which limited its speed.
Faster R-CNN revolutionized the field further by replacing the non-learned Selective Search algorithm with a Region Proposal Network (RPN), a neural network that generated proposals directly from the shared feature maps used for classification. This architecture enabled end-to-end training and a substantial reduction in the time taken for image processing.
Despite achieving a speed of approximately 5 frames per second (FPS), Faster R-CNN matched the accuracy of prior models, positioning it as a major milestone in object detection history.
While two-stage detectors like Faster R-CNN have limitations in real-time applications due to their compute-intensive proposal generation process, they remain relevant in fields where accuracy takes precedence over speed, such as medical imaging and satellite analysis. Subsequent iterations, like Cascade R-CNN, further refined this methodology by introducing multiple stages of refinement, although they still adhere to the core two-stage architecture principle of balancing speed and precision.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.