Urgent.News

What's breaking now, across thousands of outlets.

AI

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent

Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while providing applicants with a flexible interview experience.

  • Amazon Connect Talent uses AI for interviews and assessments.
  • Platform streamlines hiring for retail, logistics, hospitality.
  • Recruiters focus on decisions, AI handles screening and evaluation.

More from Thursday 17 September →