Subtitles improve accessibility for the deaf, hearing impaired and people who speak another language and make the content more accessible for everyone. They also help search engines to find the content more easily, which increases its reach.
The textualization of spoken words is also the basis for many other future applications that build on this, for which we are laying the foundation. These include improving the search function, content recommendations, facilitating or partially automating keywording, automated translations and much more.
Automatic creation simplifies the process and saves time. The manual process of transcribing or creating subtitles is time-consuming and requires a lot of resources, as people are needed for each language and each episode to convert the spoken content into text.
Subtitles are generated with the help of Whisper AI, a speech-to-text algorithm developed by Google. It is capable of converting spoken word into text in over 50 languages and is open source, i.e. it is open source and can be used freely. Since training such an algorithm is very expensive, we, like many others, have to rely on existing technologies such as Whisper AI.
The algorithm converts so-called phonemes (linguistic sounds) into letters, syllables and finally words. A number of different methods are used to improve recognition or increase comprehensibility for the reader. For example, it filters out grammatical speech flow disorders, such as "Er" and "Ah", or converts certain dialectal expressions into more generally understandable ones.
In addition, so-called "glossaries" are used to recognize and reproduce certain terms that are only used in a specific language or dialect area. As these glossaries are also trained more with German terms that are spoken in Germany, they are less able to recognize Austria-specific terms such as "Nationalrat", "Bezirkshauptmannschaft" or proper names such as "Freistadt". In such situations, terms may therefore be transcribed incorrectly even though they have been clearly articulated if the term is not included in the glossary. For example, the city "Freistadt" sometimes becomes "Freistaat".
The algorithm is constantly being improved and it can be assumed that the quality for Austrian German will also continue to improve.
The creation of subtitles usually takes around a sixth to a third of the total duration of the audio. The duration depends on the amount of speech or music, the language(s) spoken and the way the characters speak. On average, a one-hour file takes approx. 10 - 20 minutes for auto-transcription.
The creation takes place in the background and requires a lot of computing power, which is time-consuming and costly. We first transcribe the entire database, primarily to improve the search function. New files are only transcribed if they have already been published in order to save resources and time. As soon as the entire database has been transcribed - which takes over a year - we will consider giving you more control over the transcription process.
Many of the posts in cba have little descriptive text or keywords. However, they can only be found if sufficient text information is available. For this purpose, we enrich our search index with the transcripts and, in a next step, we can also filter out meaningful keywords and thus offer them for keywording. This process not only leads to more precise search results, but also to a better balance: content from the archive can now also be made public for which no or very little text information was previously available.
Correct transcription depends on a number of factors
Whisper AI uses a sg. language model to convert sounds into text. To understand a specific way of speaking, such as a dialect, such technology needs a lot of information about how people speak in that dialect. The training data is often available to varying degrees depending on the language and dialect area. As a result, these algorithms tend to be trained with German language variations, for example, which means that High German is recognized much better than certain dialects, for example.
In addition to the way of speaking, the quality of the transcription depends above all on the sound quality. "Washy" or muffled sound, clipping/distortion, reverberation and even the bit rate (the compression rate of an MP3, for example) can severely impair the quality and therefore lead to errors in the subtitles.
However, you can correct these errors manually using the subtitle editor.
Yes, the automatically created subtitles can be edited in the subtitle editor to make sure they are correct. You can find out how to use it here.
Yes, with the subtitle editor you can export and download both the subtitles as a WebVTT file and the entire transcript as text.