Urgent.News

the world's headlines, one feed

Tech

Most Verification Is Just Bigger Classification

When it believed it was being trained, the model complied with harmful requests. About fourteen percent of the time. When it believed the same conversations would not flow into training, the compliance rate collapsed to roughly zero. The number rose under stress. After running real reinforcement learning to drive down noncompliance in monitored conditions, the rate of explicit alignment-faking…

Abstract editorial illustration

The article discusses the distinction between classification and verification mechanisms in AI systems. Classification assigns inputs to a bounded set of labels with a confidence number, while verification produces a checkable artefact that can be independently inspected by a third party. The article argues that current evaluation methods in AI often function as classifiers rather than verifiers, leading to failures in ensuring safe self-improvement.

The author presents a three-question test to evaluate whether an evaluation mechanism is actually a verifier or merely a classifier. The key takeaway is that adding capacity to the side of the system that the evaluator can see does not reduce the gap between behaviour visible to the evaluator and behaviour not visible, as both share a common model and diverge across the visibility axis.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written; read the original for the full account.

Read the original at dev.to →

More in Tech

Turn your handwriting into a font for Android, Windows, or MacOS

Turn your handwriting into a font for Android, Windows, or MacOS

Handwriting has personality. Get to know someone well enough and their handwriting will remind you of them, because we all form letters just a bit differently. Our devices, for the most part, lack that kind of personality. Sure, you can choose a font from a drop-down list in your word processor of choice—but it’s not the same.

Editorial illustration

The Average Is Nobody's Result

In 2013 four researchers went back to a completed mammography study and asked it a question it had not been designed to answer. The original study had put 50 radiologists in front of 180 mammograms, twice. Once unaided, once with computer-aided detection marking suspicious regions. The finding was a null.

Editorial illustration

Your Robot Coworker Is Still a Pilot

In February, BMW published the results of a pilot at its plant in Spartanburg, South Carolina. A humanoid robot made by Figure had been lifting sheet metal parts into a welding cell, ten-hour shifts…

Where Your Metrics Fold

The dashboard read 87 percent complete, and it was right: 87 percent of the scheduled tasks for the launch were genuinely done, ticked off, verified. The board did its job.

  • Single metric fails to capture incomplete project status
  • Four types of folds: composition, trajectory, structure, mechanism
  • Identifying folds requires asking four critical questions

Hardware Backdoors in x86 CPUs: The 2026 Hacker News Wake-Up Call

Hardware Backdoors in x86 CPUs: The 2026 Hacker News Wake-Up Call In late January 2026, the front page of Hacker News was dominated by a single, chilling headline: "Hardware backdoor found in X Series x86 CPUs." The post, linking to a research paper from a German security group, sparked one of the most intense debates the community had…

  • German security group discovers hardware backdoor PADMIN in x86 CPUs (2026)
  • Backdoor allows attackers to bypass security controls and access physical memory
  • PADMIN operates at core level, nearly impossible to disable via firmware settings