Tools voor het geautomatiseerd transcriberen van audio- en videofragmenten/en

Speech recognition, speech-to-text (STT) or automated speech recognition (ASR) is a technology that makes it possible to convert spoken text in videos or audio into text, such as the automatic subtitles on YouTube or Zoom. In this overview we focus on transcribing interviews for oral history, but virtual assistants like Siri or Google Assistant are also a form of this technology.

Speech recognition is a relatively old technology. The first commercial tools appeared in the early 1990s. They use models, systems that are trained on a certain set of data to recognize patterns and make decisions without human intervention. Speech recognition models are language models trained on audio such as interviews, audiobooks, lectures and presentations. The strength of the speech recognition tool depends enormously on the model used.

Inhoud

Deel dit artikel:            

TRACKS is een samenwerking tussen deze partners: