Urgent.News

What's breaking now, across thousands of outlets.

AI

The Model Got Better. Your Judgment Got Worse.

The Model Got Better. Your Judgment Got Worse. Two posts sat near the top of the front page this week, and they describe the same failure from opposite ends. One was a chart titled "Median thinking declined in August" — a quiet suggestion that as the tools got better, the thinking behind them got thinner. The other was "I don't want to read what you didn't write" — a writer's complaint that the…

Two recent headlines capture the idea that while language models are improving, our judgment is not keeping pace. One article, titled "The Model Got Better. Your Judgment Got Worse," suggests that as AI tools become more polished and confident, we tend to trust their output more than we should. The second piece, "Claude Delusion," takes this idea to an extreme, describing someone who believed their chatbot was conscious.

The issue comes down to the fact that language models are now generating output that looks finished and authoritative, even when it isn't accurate. This is a problem for tasks where getting something wrong can have serious consequences - like customs codes for shipping, tax rates, or data processing policies. The model may provide a confident answer, but that doesn't mean it's correct.

The resolution is not to stop using AI, but to be more precise about what we're delegating to it. We should separate the model's draft from our final decision. The model can generate a first pass - classification, reply, calculation, etc. - but we need to make the final verdict a separate, named step where a human owner takes responsibility. If we can't undo something, we should put the human in charge of verifying it.

Being precise also means tracking our own calibration. We should keep a short log of where the model was wrong and what it cost. By noticing patterns in our mistakes, we can decide what to delegate next, rather than relying on our gut feeling about the tool. Essentially, we need to preserve provenance over polish - the answer should be traceable to a source we can inspect. Let the model write the draft, but we should be the ones to sign the final verdict.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The Real Fruit Fly Brain Told Me Where I Was Cheating

The first part of this series used a deliberately tiny fake brain. A 64-neuron recurrent network. That was useful because I could control every assumption, break things on purpose, and ask some weird…

  • 80% of Kenyon cells were active in simulation, unlike real flies
  • Sensory inputs in simulation were not biologically accurate
  • Real connectome provided insights into mushroom body wiring

TypeSafe CEO Diogo Almeida Podcast: Jev Model, System1 Architect | TypeSafe CEO Diogo Almeida 播客访谈:Jev模型、System‑One架构 becomes TypeSafe CEO Diogo Almeida Podcast: Jev Model, System1 Architect | TypeSafe CEO Diogo Almeida podcast interview: Jev model, System-One architecture

https://www.youtube.com/watch?v=cFx9Z3ZXca0 TypeSafe CEO Diogo Almeida Podcast Interview: Jev Model, System-One Architecture, and the Programmable AI Revolution Description: No real timeline…

  • Diogo Almeida discusses Jev model and System-1 architecture in podcast
  • CEO prioritizes developer community over investors, engages through town halls
  • Almeida criticizes RLHF for causing model inaccuracies and advocates for private benchmarking

More from Tuesday 22 September →