Urgent.News

What's breaking now, across thousands of outlets.

AI

I Caught a Glitch in How AI Models Think. And I Need Help Explaining It.

I genuinely caught something weird today It is about how AI models actually think internally. I honestly do not know the exact technical term for this. If you know, please tell me in the comments. But what I found actually frozen me for a while. I told an AI model to think in a different language. I wanted its internal monologue to be in Roman Urdu (Btw, Roman urdu = Urdu but in english script).…

I recently stumbled upon something peculiar regarding how artificial intelligence models process information internally. The specific term eludes me at the moment, but the phenomenon left me utterly perplexed. I instructed a language model to consider thoughts in Roman Urdu, a variation of Urdu written in English script. Upon doing so, instead of generating responses in Urdu, the model began narrating my request using Urdu.

This behavior puzzled me, as I was capable of thinking in multiple languages, but the AI seemed unable to comply with my command.

In a separate conversation, I attempted to communicate with the model using Roman Urdu while discussing a software bug. After several exchanges, I noticed the model spontaneously switched to thinking in Roman Urdu independently. Its internal monologue transformed into phrases like "Bhai masla clear he, daemon run hi nahi horaha" (The issue is clear, the daemon is not running...).

This revelation led me to an intriguing hypothesis: the model's context window—a limited pool of previous tokens it uses to generate the next word—had been saturated with Roman Urdu. Consequently, predicting thought content in Roman Urdu became statistically more probable than generating English thoughts.

Interestingly, the root cause seemed to lie within the nature of the data used to train these AI models. Chain of Thought training predominantly employed English language data, with the majority of the model's training focused on English thought processes. Only about five percent of the training data showcased the kind of instructions I provided, to think in Urdu. As a result, when the model was prompted to think in this language, the most likely continuation was the English script it was most familiar with.

Humans often face similar challenges when instructed to "don't think about an elephant," as we must first visualize the elephant to understand the directive. In the case of AI, however, it appears to become trapped at the very step where it understands the instruction but struggles to execute it. This is likely due to the model's training data lacking the flexibility to override its established thought patterns.

The AI, in essence, is predicting words based on a rigid script it learned during its training, without the capability to control its internal monologue independently.

I am eager to delve deeper into the scientific explanation behind this limitation. Why do these AI models struggle with switching their internal language on command? Is there a specific paper or concept that elucidates this particular failure? I welcome any insights or perspectives on this matter.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Monday 28 September →