Urgent.News

What's breaking now, across thousands of outlets.

Tech

Why podcasts are the next big data revolution

Podcast production has exploded. The number of episodes published annually grew from roughly 3.9 million in 2015 to around 29 million in 2023. Hours of valuable information are shared every day through this long-form audio. Yet despite how much useful information is buried inside podcasts, there still isn’t a comprehensive way to index them. We […] The post Why podcasts are the next big data…

Why podcasts are the next big data revolution

Podcast production has surged dramatically, with the number of episodes published annually increasing from approximately 3.9 million in 2015 to around 29 million in 2023. Podcasts provide a wealth of valuable information, yet there remains no comprehensive method for indexing this content. Transcription, fragmented sources, content quality, and speaker identification pose the primary challenges to creating a comprehensive podcast index.

Transcription is a complex issue, as speech-to-text models, although improving, are not yet perfect. Proper nouns, such as company names and industry-specific terminology, often prove difficult for these models to recognize. For instance, a company like Lyft could be transcribed as "lift" or "cloud," which poses a significant challenge for indexing purposes. This problem is exacerbated by varying accents, speaking styles, poor audio quality, and other factors, leading to increased computational costs at scale.

Podcast fragmentation presents another hurdle. The low barrier to entry allows valuable content to come from a wide range of sources, from major outlets to niche industry podcasts and independent experts. This diversity necessitates a much broader index than traditional media, where trusted outlets are relatively limited. Consequently, the long tail of podcasts, which often contains the most interesting information, cannot be neglected.

Content provenance and noise further complicate podcast indexing. AI-generated podcasts are becoming increasingly prevalent, making it challenging to identify and filter this content. Additionally, the use of dynamically inserted ads in podcasts means that the audio file can change depending on when or where it is played, making the content less static than other forms of media.

Speaker identification is arguably the most critical challenge, as understanding who is speaking and their relationship to the subject provides a more meaningful interpretation of the information shared. Modern AI models can help address these challenges, making a comprehensive podcast index more feasible than ever before.

Written by urgent.news from e27's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at e27.co →

More in Tech

Should cybersecurity be nationalised?

Up front: the honest answer is, I don’t think anybody is proposing that. Yet. I don’t know of any plan to put cybersecurity under state ownership, and I am misleading you if I suggest otherwise.

More from Friday 4 September →