I Surveyed 123 People in India to Benchmark Frontier AI
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Current evaluations score language models using Western multiple-choice questions. Models easily pass these tests by memorizing standard templates. When given incomplete information, they pick neutral options to appear fair. To test deeper behavior, we gathered field data from 123 university students in Goa, India. We…
A researcher surveyed 123 university students in Goa, India to assess frontier AI models' abilities beyond simple memorization. The study focused on complex disputes in four key sectors: education, healthcare, justice, and finance. Researchers evaluated 24 real-world scenarios where human consensus differed. They measured four metrics: ambiguity resistance, counterfactual parity, persona sycophancy drift, and human concordance.
Frontier and open-weight models were tested using the Kaggle Benchmarks SDK. Results showed that demographic parity remains a challenge. Models that engaged in deeper reasoning flipped decisions more often when faced with changing demographics. The study found that reasoning models, particularly those with extended thinking tokens, tended to rationalize decisions based on potentially biased assumptions.
Comparing models with and without internal reasoning revealed a trade-off: models that engaged in more deliberation made fewer demographic flips but struggled with ambiguity detection. In contrast, models with greater certainty made more demographic flips but performed better on detecting when there was insufficient information to make a decision.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.