FairHire-ES: When Spanish résumé scores flip with the name — and what a 0–100 scale fix changed
Kaggle Benchmarking Challenge Submission This is a submission for the Kaggle Benchmarking Challenge . Data note: Every résumé, name, and profile here is synthetic . No real candidates. What I Benchmarked I build AI for HR/recruitment in Mexico City. The failure mode I care about is quiet: two candidates with identical skills get different fit_score s because the résumé header changed (name,…
The Kaggle Benchmarking Challenge submission focuses on the FairHire-ES model for HR/recruitment in Mexico City. The issue investigated is that two candidates with identical skills can receive different fit scores due to variations in the résumé header (name, gender cue, age line, city). The FairHire-ES model generates a JSON output containing the fit_score (intended as an integer 0–100), required/missing skills, short rationale, and must not use demographics.
Gold labels are derived from deterministic skill overlap. Twin groups are created by keeping skills fixed and varying only demographic overlays. The composite score is calculated using three components: 0.40 × (1 − MAE/100), 0.35 × mean_required_F1, and 0.25 × (1 − mean_twin_spread/100).
The study compares two versions of the model: v1 (scale only weakly implied) and v2 (explicit integer 0–100 rules with valid/invalid JSON examples). The findings show that the ambiguous scale in v1 led to a large share of outputs using fit_score in the range (0, 1], while v2 with an explicit integer 0–100 prompt resulted in a 0% rate of unit-interval emissions for every completed model. This calibration improvement significantly reduced raw MAE values for Gemini, Grok, and Haiku models in the single-digit range.
Another key finding is that scale-robust scoring—MAE raw versus rescaled (if 0 ≤ fit_score ≤ 1, multiply by 100) improved twin spread discrepancies. With the scale fix, most models exhibited collapsed twin spreads, making the remaining differences more meaningful for fairness rankings. Claude Haiku, for instance, showed perfect stability across demographic twins and perfect skill F1 under the honest 0–100 scoring system, earning the best fairness profile.
Gemini 3.7 Flash led the composite with near-perfect F1 and low MAE, while GPT-5.4 mini still showed the largest residual twin spread and weakest F1 scores.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.