Urgent.News

What's breaking now, across thousands of outlets.

AI

God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

...

God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Mechanistic interpretability is a scientific approach to understanding how AI models work. Large language models are not built but rather "grown" through neural networks, with no clear understanding of their inner workings. By analyzing the neurons, connections, and activations, researchers hope to reverse engineer the AI.

In 2023, a breakthrough was made in many-to-many mappings, where neurons firing together can represent different concepts simultaneously. For example, in a 512-neuron AI, neurons 1 and 2 firing together might mean "cat," while neurons 1 and 3 firing together might mean "chair." This allowed AI to represent millions of concepts with just tens of thousands of neurons.

The excitement surrounding mechanistic interpretability grew as researchers believed it could lead to discovering the fundamentals of cognition, solving alignment issues, or even making provably non-discriminatory or honest AI systems. However, by 2024-2025, the clarity of these mappings began to dissolve, and the field faced new challenges.

One technique, linear probes, involves creating pre-defined concept-representing direction vectors in the activation space. Researchers then compare the AI's activation vector to these direction vectors to infer the AI's thoughts. While linear probes can be useful as primitive lie detectors or in specific contexts, they have limitations. The resulting concept representations may be biased or incomplete, and AI can easily evade detection by rotating or dispersing the concept.

Written by urgent.news from Astral Codex Ten's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at astralcodexten.com →

More in AI

More from Tuesday 8 September →