The Residual Class Problem: Why a Face Shape Classifier Keeps Answering Oval
If you build a classifier with a fixed set of labels, one of those labels usually ends up doing a job nobody assigned it: catching everything that fails to match the others. Face shape detection is a clean, low-stakes example of this, and the numbers that the face shape detector measureface published about its own classifier make the effect easy to see. Six labels, one of them defined by absence…
When creating a classifier with predetermined categories, one label inevitably ends up fulfilling a purpose that was not originally intended. Face shape detection serves as a straightforward, low-stakes example of this phenomenon, and the data provided about the classifier's performance reveals this effect clearly. Out of six possible categories, one - oval - is defined by the absence of certain features. This results in a "residual class" that catches all faces that fail to meet the criteria for the other shapes.
Five of the six categories are characterized by unusual traits: square has a hard corner at the jaw, oblong is longer than it should be, diamond has a narrower forehead, heart has a pointy chin, and round is about as wide as it is long. Oval is categorized differently. Its usual description is that it is longer than wide, widest at the cheekbones, and tapers gently to a rounded chin without any sharp corners. In terms of classifier terms, oval is the residual class.
The output distribution of the classifier was measured by measureface, and the results are noteworthy. Out of 43 distinct faces, 15 were classified as oval, which is the most frequent label and nearly four times the number (4) classified as oblong. This discrepancy can be attributed to the geometry of the shapes. If the oval prototype is near the average of the other five shapes, then a face that lacks any distinct feature will be closest to oval by default.
Since most faces do not exhibit a standout characteristic, most ambiguous inputs end up in the oval category.
Another indicator of this phenomenon is the number of "tie cases," where the two closest prototypes are within a certain margin, and the detector reports a combination of labels. In this case, eight out of the 43 faces came back as paired labels, and all eight included oval. Four of these pairs were labeled as oval/round, three as oval/heart, and one as oval/diamond.
For every other shape, a single measurement plays a significant role in excluding it. For diamond, it is the forehead width, for square it is the jaw, for round it is the length compared to width, and for oblong it is the length compared to width. However, for oval, there is no dominant measurement. The forehead and jaw each account for exactly 16 out of the 43 faces, a dead tie. No single feature is deciding oval; it is essentially whatever remains after all other factors have been considered.
The implications of this residual class for those implementing classifiers are significant. The residual label carries less information than the other labels. A square result indicates that a specific measurement crossed a threshold, while an oval result mostly signifies that no threshold was crossed. When all six labels are presented with equal confidence in the user interface, users may perceive the residual label as a positive finding when it is more akin to a lack of strong signal.
This issue is not exclusive to face shape detection. Any taxonomy with an implicit "none of the above" category behaves similarly. Support ticket routing systems may have a "general" queue, document classifiers may have an "other" bucket, and sentiment models may have a "neutral" label. These residual categories absorb ambiguous inputs, their precision appears acceptable because the inputs are genuinely ambiguous, and the label quietly becomes the most common output.
To mitigate the impact of residual labels, several design choices can be made. Instead of only reporting the label with the highest score, showing the distances between the input and all prototypes allows users to see that the input was close to oval rather than distinctly oval. Additionally, explicitly emitting ties, such as a pair result within a specified margin, provides a more honest representation of the classifier's output.
Publishing the output distribution is also crucial, as it highlights when one label dominates the others.
It is essential to note that this example is based on a synthetic test set, meaning that all faces were generated by an image model and do not represent real people. Consequently, the observed rate of 15 out of 43 faces being classified as oval should not be interpreted as a population rate. Furthermore, there is no peer-reviewed data on the prevalence of these six facial styling categories, so there is no way to determine how common oval truly is.
The boundaries between facial shapes are ultimately conventions, and different prototypes could alter the counts. Nonetheless, the structural point remains valid: any tool built using the same six labels will inherit this residual behavior, regardless of whether it is explicitly reported.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.