Urgent.News

What's breaking now, across thousands of outlets.

AI

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness

Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

I Built a World With Rules. Then I Changed Them.

The Rules Are Hidden. The Timeline Isn't. What happens when you put something inside a world whose rules it doesn't know? I wanted to build a small experiment.

  • NIXIE is a fictional reasoning benchmark for Kaggle.
  • Rules in NIXIE include ERASE or CONTINUE outcomes based on dice rolls.
  • Rules can change, making the experiment more challenging.

More from Thursday 8 October →