67 tools
Voice, music, and sound design tools for product audio and generative audio.
Local desktop studio for zero-shot cloning, dictation, and cinematic dubbing (646 languages)
Local AI audiobook studio with per-sentence emotion/voice control
Local-first voice cloning studio with multi-engine TTS + timeline
Local voice AI runtime: ASR, diarization, TTS, cloning + OpenAI-compatible API
Full-duplex, multi-agent, 100% local conversational voice engine
JUCE synth that uses diffusion models as oscillators
Code-first procedural SFX with MCP natural-language tools
One-command CLI for SFX, loops, songs, narration via ComfyUI
AI sample-library manager + generative “chaos mixer”
One-file Windows app for Stable Audio Open with 217 presets
Tiny ~82M Apache-2.0 TTS with strong English quality
Multilingual TTS with short-clip cloning
Open multilingual TTS + cloning from ~3s audio
Zero-shot voice cloning TTS
MIT-licensed multilingual voice cloning
Fast real-time multilingual TTS (MyShell)
Expressive TTS with emotion tags
Style-controllable TTS
Expressive TTS with non-speech sounds (laughter, etc.)
Few-shot voice cloning / conversion pipeline
Retrieval-based voice conversion
Ultra-low-latency cloud TTS (40+ languages)
TTS optimized for live voice agents + cloning
Emotionally controllable speech / empathic voice
On-prem / controlled-deployment TTS for CX
Character / interactive realtime voices
AI music with clean stem + Ableton project export
AI MIDI generator as VST3/AU/AAX for DAWs
Real-time SFX / Foley / ambience creator (VST/AU)
AI music generator; strong instrumental/mastering sound
Offline AI loop/texture VST (Stable Audio engines)
On-device Mac audio cleaner (noise, fillers, EQ)
Offline cleaner/master for Suno/Udio artifacts
Local GPU app: stems → MIDI → generation → mix
Browser/native DAW with AI gen + timeline Copilot
AI denoise/stem + DSP rack as VST3/AU/LV2
Lightweight offline TTS (rhasspy)
Alibaba open TTS with zero-shot cloning
Text-to-audio / music from Stability AI
Cloud TTS, cloning, agents, dubbing
Open-source AI music suite with layer/LEGO pipelines, cover, repaint, stems — local Suno/Udio alternative
Leading consumer AI music generator (full songs from prompts)
Expressive TTS/voice cloning with inline emotion tags; API-first
Open-source multilingual voice cloning TTS (Resemble AI)
AI voiceover studio with video timeline sync aimed at e-learning and YouTube
Enterprise-grade AI narration voices
Text-based audio/video editor with Overdub, Studio Sound, and screen recording
Licensed music + SFX + footage library with expanding AI tools
Browser podcast studio with AI enhancement and hosting
Open TTS model noted for high-fidelity (~44kHz) output
Open multi-speaker dialogue TTS with laughs/pauses
Conversational Speech Model for emotional expressiveness
OSS TTS optimized for conversational/chatbot prosody
Compact bilingual EN/ZH 0.5B TTS with prompt control
ByteDance scene audio: dialogue+SFX+ambience+music one pass
Open multimodal audio generation synced to video
CSS-like language binding sounds to UI events
AI music generation often A/B'd with Suno
AI royalty-free music tailored for videos/podcasts
AI music generator with editable structure
Stem splitter for vocals/instruments
AI stem separation + practice tools
AI filler-word/noise cleaner for voiceovers
Automated loudness leveling and podcast/audio mastering
AI noise cancellation for calls, research, and voice notes
Browser AI audio cleanup (Enhance Speech) often used alongside photo/video creator stacks
Developer-oriented AI TTS/voice API often used for product narration and apps