Urgent.News

What's breaking now, across thousands of outlets.

AI

SAIL: Sparse Autoencoders for Interpretable Alignment of Human Vision and Multimodal Large Language Models

Alignment between human high-level visual representations and those of multimodal large language models (MLLMs) offers a quantitative framework for understanding information processing in human vision. However, similarities between model and brain representations do not directly reveal the semantic content underlying their correspondence, as this content is difficult to isolate in entangled MLLM…

SAIL represents a framework designed to provide an interpretable method for aligning the visual representations of human brains with those of multimodal large language models (MLLMs). This system tackles the challenge of discerning semantic content within the complex, intertwined representations found in MLLMs, which often obscure the underlying meaning.

To achieve this, SAIL employs a novel approach involving the use of unified cross-layer sparse autoencoders (SAEs). These models are trained on the representations generated by MLLMs from images in the Natural Scenes Dataset, paired with their corresponding captions. The process involves aligning each brain voxel to multiple SAE units, thus bridging the gap between neural and visual representations.

The framework further quantifies the semantic similarity between the top-response image sets of both the MLLM and brain voxels. This allows for a mapping of semantic alignment across the cortex, which is then used to project multiple semantic labels back onto the brain's cortical regions. This step effectively interprets the content of the observed alignments.

To validate these findings, category responses obtained from independent fLoc experiments are utilized. Comparisons with SAEs trained separately at each layer reveal that SAIL exhibits reduced redundancy in SAE unit response profiles and demonstrates a higher degree of semantic alignment.

The results indicate that semantic alignment is systematically organized along the ventral visual pathway in both the models and subjects studied. Projected labels from the framework reveal specific semantic preferences within key visual regions, findings that are corroborated by the results of the fLoc experiments, which show a correspondence between the predicted and measured category response patterns.

Further exploratory analyses provide insights into the cortical distribution of action-related semantics, aligning with established functional accounts of the lateral occipitotemporal and posterior parietal cortex.

In summary, SAIL offers a comprehensive, interpretable method for understanding the semantic correspondences between the representations of MLLMs and human visual responses, paving the way for deeper insights into the workings of human vision.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

More from Tuesday 6 October →