🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Faster Whisper transcription with CTranslate2
OpenVINOâ„¢ is an open source toolkit for optimizing and deploying AI inference
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Voice-to-text with push-to-talk for Wayland compositors
OpenAI Whisper ASR Webservice API
Whisper.net. Speech to text made simple using Whisper Models
A speech to text IBus engine using VOSK