Foveation as self-supervision signal: Blur masks beat blank masks for learning robust visual representations
The primate retina is strikingly non-uniform, with receptor density falling off sharply from fovea to periphery. Every eye-movement brings a previously low-acuity peripheral area to high-acuity foveal processing. We hypothesize that this process provides a natural self-supervision signal, enabling representation learning by predicting fine details in the periphery. We implement this in a…
Recent studies have revealed that the primate retina processes visual information in a non-uniform manner, with a sharp decline in receptor density as one moves away from the fovea. This gradient in receptor density may serve as a natural self-supervision signal, allowing for representation learning through the prediction of fine details in the periphery.
To explore this hypothesis, researchers implemented a vision transformer (ViT) architecture to study its performance using different types of mask-based corruption. The study introduced three variants of the ViT: ViT-Blank, ViT-Blur, and ViT-Retina. The blank mask (ViT-Blank) replaced the standard mask used in vanilla vision transformers, while the blur mask (ViT-Blur) employed a uniform blur, and the retinal mask (ViT-Retina) featured a randomly selected foveal patch with progressively blurred patches based on eccentricity.
The pre-training process involved reconstructing the full-resolution image from the corrupted input, followed by fine-tuning on clean images for classification. The results showed that both ViT-Blur and ViT-Retina outperformed ViT-Blank on the pre-training reconstruction task while maintaining a higher level of information retention from the input, regardless of the corruption type.
More importantly, the study demonstrated that ViT-Blur and ViT-Retina outperformed ViT-Blank in downstream classification tasks, indicating that they learn representations better suited for generalization. In addition, these mask-based techniques exhibited greater robustness under various impoverishment conditions, such as higher mask ratios, fewer pre-training epochs, smaller training sets, and test-time image corruptions.
By analyzing the type of spatial frequencies to which the models were sensitive, researchers discovered that ViT-Blank relied on higher spatial frequencies compared to the blur and retinal masks. This difference in sensitivity correlated negatively with the models' classification accuracy ($r = -0.73$), suggesting that the models utilizing the blur and retinal masks might be more robust in generalizing their learned representations.
Further investigation revealed that a closed-form linear-ridge encoder derived from the same reconstruction objective could replicate the spatial frequency tuning observed in ViT-Blank ($r = 0.88$). Overall, the findings highlight the potential benefits of a biologically-inspired self-supervision objective, which could lead to the development of more robust visual representations in artificial intelligence systems.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.