Sample-Conditioned Representation Selection for Audio Few-Shot Learning
Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.