yt-dlp is a feature-rich command-line audio/video downloader that grew as an active fork of youtube-dl and youtube-dlc.
Open Source: #audio
Catalog projects marked with #audio. Tags work as dedicated landing pages, so related tools are easier to find and connect.
This collection holds 10 projects with a combined 663,871 GitHub stars. Main languages: Python, C, C++.
Repositories
Whisper is OpenAI’s model and Python package for speech recognition, speech translation into English, and language identification in audio files.
FFmpeg is a set of libraries and command-line tools for processing video, audio, subtitles, and metadata.
Real-Time Voice Cloning is a research Python project for voice cloning and speech synthesis.
GPT-SoVITS is a project for few-shot speech synthesis and voice transfer from small audio samples.
whisper.cpp is a high-performance C/C++ implementation of OpenAI Whisper inference for speech recognition.
Coqui TTS is a deep-learning toolkit for text-to-speech and voice experiments.
ChatTTS is a generative speech model for dialogue and everyday audio scenarios.
Bark is Suno’s generative audio model for speech and sounds from text instructions.
OpenVoice is a model for fast voice cloning and speech synthesis experiments.