Vision-language encoding models reveal an image-computable food-quality dimension in human occipitotemporal cortex
Perceived calorie content contributes to neural representational structure in human ventral visual cortex, yet it remains unclear whether this reflects an abstract nutritional signal or whether perceived calorie is largely recoverable from the visual-semantic structure of the food image itself. In 25 female participants who passively viewed 96 food images during functional MRI, we decomposed…
Recent research has shown that the neural representation of food quality in human ventral visual cortex is primarily driven by a shared dimension that captures processedness, naturalness, and perceived healthiness. This finding comes from an analysis of perceived calorie ratings from 25 female participants who viewed 96 food images while undergoing functional MRI scans.
By breaking down the perceived calorie ratings into two components, one predicted from CLIP (Contrastive Language-Image Pretraining) image embeddings and the other as a residual component not captured by this prediction, the researchers were able to test their contributions to neural prediction using cross-validated banded ridge encoding models.
The CLIP-predictable component successfully organized foods along a processedness and naturalness dimension, distinguishing raw single-ingredient foods from prepared and energy-dense foods. Incorporating this component into a visual-semantic baseline significantly improved neural prediction, particularly in higher-level ventral temporal cortex.
These results suggest that the encoding of calorie-related information in ventral visual cortex is carried mainly by a shared food-quality axis, with a substantial portion of this information being recoverable from image-computable visual-semantic structure. However, the study acknowledges that further investigation is needed to determine the generalizability of these findings to broader populations and to larger, more diverse sets of stimuli.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.