Urgent.News

What's breaking now, across thousands of outlets.

Culture

Foveation as self-supervision signal: Blur masks beat blank masks for learning robust visual representations

The primate retina is strikingly non-uniform, with receptor density falling off sharply from fovea to periphery. Every eye-movement brings a previously low-acuity peripheral area to high-acuity foveal processing. We hypothesize that this process provides a natural self-supervision signal, enabling representation learning by predicting fine details in the periphery. We implement this in a…

Recent studies have revealed that the primate retina processes visual information in a non-uniform manner, with a sharp decline in receptor density as one moves away from the fovea. This gradient in receptor density may serve as a natural self-supervision signal, allowing for representation learning through the prediction of fine details in the periphery.

To explore this hypothesis, researchers implemented a vision transformer (ViT) architecture to study its performance using different types of mask-based corruption. The study introduced three variants of the ViT: ViT-Blank, ViT-Blur, and ViT-Retina. The blank mask (ViT-Blank) replaced the standard mask used in vanilla vision transformers, while the blur mask (ViT-Blur) employed a uniform blur, and the retinal mask (ViT-Retina) featured a randomly selected foveal patch with progressively blurred patches based on eccentricity.

The pre-training process involved reconstructing the full-resolution image from the corrupted input, followed by fine-tuning on clean images for classification. The results showed that both ViT-Blur and ViT-Retina outperformed ViT-Blank on the pre-training reconstruction task while maintaining a higher level of information retention from the input, regardless of the corruption type.

More importantly, the study demonstrated that ViT-Blur and ViT-Retina outperformed ViT-Blank in downstream classification tasks, indicating that they learn representations better suited for generalization. In addition, these mask-based techniques exhibited greater robustness under various impoverishment conditions, such as higher mask ratios, fewer pre-training epochs, smaller training sets, and test-time image corruptions.

By analyzing the type of spatial frequencies to which the models were sensitive, researchers discovered that ViT-Blank relied on higher spatial frequencies compared to the blur and retinal masks. This difference in sensitivity correlated negatively with the models' classification accuracy ($r = -0.73$), suggesting that the models utilizing the blur and retinal masks might be more robust in generalizing their learned representations.

Further investigation revealed that a closed-form linear-ridge encoder derived from the same reconstruction objective could replicate the spatial frequency tuning observed in ViT-Blank ($r = 0.88$). Overall, the findings highlight the potential benefits of a biologically-inspired self-supervision objective, which could lead to the development of more robust visual representations in artificial intelligence systems.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Culture

More from Thursday 1 October →