My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it.
I built Podium , a tool that compares your reading of a speech with a great delivery of the same text (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range. The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged 41 "flaws" per minute on a perfectly clean…
Podium is a tool that compares a person's reading of a speech with a great delivery of the same text and highlights where the reading deviates and why. Initially, the tool performed well on the test set. However, when applied to a different speaker, it generated 41 false alarms per minute on a perfect recording. The issue was addressed with a one-line fix and improved evaluation set-up.
The process involved aligning both recordings to the same word grid, building a dataset with exact answers by injecting flaws, and separating style from flaws. By determining what shouldn't count before deciding what should, the false alarms decreased significantly. The evaluation results showed that the tool detected flaws with a mean Intersection over Union (IoU) of 0.81, an overall F1 score of 0.51, and no false alarms on clean audio of the reference speaker.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.