Nvidia Accelerates Real Time Speech to Text Transcription 3500x with Kaldi
NVIDIA’s work in optimizing the Kaldi pipeline includes prior GPU optimizations to both the acoustic model and the introduction of a GPU-based Viterbi decoder in this post for the language model. You can find more detailed optimization information in the archive of our past GPU Technology Conference (GTC) presentations:
NVIDIA tested a model trained on the LibriSpeech corpus, according to the public Kaldi recipe, on both clean and noisy speech recordings. While NVIDIA focused its Kaldi acceleration work on speech inferencing alone, future work can explore additional techniques to accelerate other components, including training, of the Kaldi speech-to-text workflow.
Source: devblogs.nvidia.com