Cross-domain confidence reliability and remappability of frozen single-cell representations
Single-cell foundation models are increasingly adopted for downstream applications such as cell-type prediction. However, these predictions are often utilized without assessing their reliability, or by relying on a simple cutoff applied to the maximum softmax probability (MSP) derived from the classifier. This raises a critical question: can raw MSP be trusted as a reliable measure of confidence…
Single-cell foundation models are becoming more popular for downstream tasks like cell-type prediction. However, people often use these predictions without checking how reliable they are, or they only use a simple rule based on the highest probability score (MSP) from the classifier. This leads to an important question: can the MSP be relied upon as a trustworthy indicator of confidence when dealing with new data (the target) that has many different technical and biological features?
To answer this, we present a method to compare how confident the model was on the training set with how it performs on the test set. We find that using just the MSP does not accurately reflect the true confidence when applied to the target data. By breaking down the difference between the training and target data, we discover that the problem lies in two main areas: the way the model ranks the cells and a general shift in the probability scores.
We also show that this shift can be corrected by using a small amount of labeled data from the target set. While this correction makes the model more reliable, it also means that more cells need to be checked manually. In the end, our framework shows that making these models safe to use requires adjusting the confidence estimates specifically for the target data, using labels that accurately represent the target.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.