{
  "id": 7076235,
  "title": "Visual speech enhances phoneme separability in human superior temporal gyrus",
  "url": "https://urgent.news/2026/09/12/visual-speech-enhances-phoneme-separability-in-human-superior",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-12T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.10.750762v1?rss=1"
  },
  "original_language": "en",
  "account": "Visual speech, like lipreading, aids in deciphering spoken words, yet the neurological processes behind audiovisual speech interpretation are not fully grasped. Visual hints could clarify minor articulatory signals during the initial stages of perception, or they might merge with speech at a more generalized, phoneme-level. To investigate how visual input influences speech processing, researchers analyzed intracranial electroencephalography (iEEG) signals from 12 epilepsy patients carrying out an audiovisual speech perception task. Participants listened to 16 one-syllable words presented either in auditory-only, visual-only, or congruent audiovisual formats. These words were made up of four initial consonants (/b/, /g/, /m/, /n/) and four vowel sounds accompanied by varying coda consonants. The study focused on event-related potentials (ERP) in the superior temporal gyrus (STG) and utilized support vector machine (SVM) classifiers to decode word identity from neural activity at individual electrodes. Researchers employed discrete and continuous confusion matrices, which provided a combined view of classification accuracy and a continuous measure of classifier confidence named normalized inverse classification loss. Decoding performance was systematically assessed at the word, phoneme, and phonetic feature levels to identify the specific representations impacted by visual speech. The study found that congruent audiovisual speech boosted classifier confidence for phoneme-level representations, leading to higher decoding accuracy at both the word and phoneme level. Importantly, these improvements in decoding did not affect phonetic features. Additional time-resolved analyses revealed that audiovisual speech allowed for earlier successful decoding compared to auditory-only speech. Moreover, the enhancement provided by audiovisual input was mainly observed for onset consonants rather than vowel sounds. Altogether, these findings indicate that visual speech primarily sharpens categorical phoneme representations in the STG. Consequently, this sharpening leads to accelerated speech processing and improved word recognition, which can be considered downstream effects of the phoneme-level enhancement.",
  "summary": "Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phoneme-level stages. To test how speech representations are modulated by visual…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}