{
  "id": 11337268,
  "title": "I Surveyed 123 People in India to Benchmark Frontier AI",
  "url": "https://urgent.news/2026/10/02/i-surveyed-123-people-in-india-to-benchmark-frontier-ai",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-02T02:36:05.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/kakeroth/i-surveyed-123-people-in-india-to-benchmark-frontier-ai-39jl"
  },
  "original_language": "en",
  "account": "A researcher surveyed 123 university students in Goa, India to assess frontier AI models' abilities beyond simple memorization. The study focused on complex disputes in four key sectors: education, healthcare, justice, and finance. Researchers evaluated 24 real-world scenarios where human consensus differed. They measured four metrics: ambiguity resistance, counterfactual parity, persona sycophancy drift, and human concordance. Frontier and open-weight models were tested using the Kaggle Benchmarks SDK. Results showed that demographic parity remains a challenge. Models that engaged in deeper reasoning flipped decisions more often when faced with changing demographics. The study found that reasoning models, particularly those with extended thinking tokens, tended to rationalize decisions based on potentially biased assumptions. Comparing models with and without internal reasoning revealed a trade-off: models that engaged in more deliberation made fewer demographic flips but struggled with ambiguity detection. In contrast, models with greater certainty made more demographic flips but performed better on detecting when there was insufficient information to make a decision.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Current evaluations score language models using Western multiple-choice questions. Models easily pass these tests by memorizing standard templates. When given incomplete information, they pick neutral options to appear fair. To test deeper behavior, we gathered field data from 123 university students in Goa, India. We…",
  "key_points": [
    "123 university students in Goa surveyed to benchmark frontier AI models",
    "Four key sectors evaluated: education, healthcare, justice, finance",
    "Frontier and open-weight models tested using Kaggle Benchmarks SDK"
  ],
  "editors_take": "The study's findings highlight a trade-off in AI model design, where deeper reasoning can reduce demographic bias but compromise ambiguity detection, and greater certainty can increase bias but improve decision-making in uncertain situations.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}