Urgent.News

What's breaking now, across thousands of outlets.

Tech

Is regex enough? I tested span-01 on mixed-language text

Have you ever exported a script for text-to-speech and found, later, that a single English word had slipped into it? I have — and the synthesized voice mangled that spot badly enough that I had to redo the take. All I actually want is to check "is there any English mixed into this Japanese script?" That turns out to be surprisingly annoying. Checking everything by hand is not realistic, so the…

Many users have encountered the issue of a single English word slipping into a text-to-speech script, causing the synthesized voice to produce poor results. The author faced this problem themselves and discovered that checking for mixed languages manually proved to be an impractical solution. Regular expressions (regex) were often used to filter such issues, yet the author found that regex alone was not sufficient, as it would flag brand names, URLs, personal names, and acronyms as defects. However, in actual scripts containing Latin characters, the regex performed adequately.

The author then came across respan.ai, a service that provides models built to make decisions based on plain language descriptions. Comprising span-01 and span-01-lite models, the service returns a probability of whether something matches a condition, defined in prose instead of code. This approach intrigued the author, who decided to test it, comparing it to regex and a general chat model.

They measured the performance using synthetic data of boundary cases (23 sentences) and real Japanese scripts containing Latin characters.

Upon analysis, the regex model significantly underperformed, producing 9 false positives and missing 2 actual instances of mixed languages. In contrast, the decision model (with varying instructions) achieved an impressive 0.93 F1 score, indicating that it closely approached the perfect accuracy. However, the author emphasized that this performance was unfair to the regex model, as the ground truth for the tests relied solely on the presence of Latin characters.

What surprised the author most was the impact of the instruction wording on the decision model's accuracy. A single sentence containing one English word mixed with Japanese text yielded varying results based on how the instruction was phrased. Moreover, the model exhibited a built-in prior that somewhat obstructed its ability to override the instruction, implying that crafting precise instructions is crucial for optimal performance.

Finally, the author discussed the limitations of the probability-based output from the decision model. They discovered that setting appropriate thresholds for classifying sentences as either mixed or non-mixed proved challenging. Drawing a line between 0.15 and 0.85 probability resulted in higher miss rates, while dropping the threshold to 0.15 led to increased false positives for other linguistic elements. The author concluded that a threshold of 0.50 proved to be the most effective and satisfactory choice.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Your dashboard is green and the number is wrong: the SQL checks I schedule next to every metric

Most broken dashboards don't look broken. The chart renders, the numbers are plausible, the pipeline says success. Then someone in finance asks why last month's revenue changed after the fact, and you…

  • Verify data freshness within acceptable timeframe, use matching load schedule window
  • Compare current day's volume to same weekday in past weeks, account for variations
  • Ensure data model consistency, reveal discrepancies before summing occurs

More from Monday 28 September →