Urgent.News

What's breaking now, across thousands of outlets.

AI

The Load-Bearing Vocabulary of Claude

Fun (but depressing?) data analysis project from Louis Abraham, showing how repetitive Claude is in the words it chooses when submitting GitHub pull requests. Abraham, in the project’s readme: GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: eight ways of writing, and every description belongs to one of them. One of…

Data analyst Louis Abraham has compiled an intriguing, albeit somewhat disheartening, study on the vocabulary used by Claude, an AI language model. Abraham's project, featured in its README file, examines GitHub pull request descriptions generated by Claude, categorizing them based on the words employed rather than the content they address.

The findings reveal eight distinct writing styles, with one style accounting for a staggering 45% of all descriptions as of mid-2026, despite only constituting 1.0% at the beginning of 2025.

This phenomenon raises significant concerns about Claude's linguistic diversity. The situation is further exacerbated by ongoing research into AI watermarking techniques. Google's white paper on SynthID-Text, Anthropic's watermarking scheme for Claude, reveals that watermarking can lead to a "some reduction to inter-response diversity."

In simpler terms, Claude's responses to identical prompts will become increasingly similar to one another, resulting in a diminished variety within its output. This development poses a critical challenge, particularly considering Claude's existing struggle with maintaining varied responses.

Written by urgent.news from Daring Fireball's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at louisabraham.github.io →

More in AI

More from Thursday 27 August →