Urgent.News

What's breaking now, across thousands of outlets.

AI

Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

Same Claude. Different Harness. Very Different Result.

Claude didn’t get smarter. We changed everything around it. Somehow, 6.5 points appeared between them. We beat Claude Code with Claude. Which is a slightly ridiculous sentence, but it is also a useful one. We ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock.

  • Claude's performance increased from 78.9% to 85.4% with system changes
  • System architecture, reliability, and cost-efficiency crucial for model potential
  • Cost difference of 49% between leading and submitted Terminal-Bench runs

More from Friday 18 September →