What each option means
Upload audio or video
Add the song, podcast, interview, or video recording you want to split. If you upload a video, the audio will be automatically extracted.
Separation model
This controls how Demucs separates music parts.
mdx_extra_q: fastest and lightest. Good when you want quicker results.
htdemucs: balanced option. Better for most people and the default choice.
htdemucs_ft: slowest and heaviest. Usually the best quality when you want the strongest separation.
htdemucs_6s: adds extra guitar and piano stems on top of the usual stems. [Default Selected]
Separate music stems
Creates files such as vocals, drums, bass, other, and also a combined background music only file with vocals removed.
Separate speakers
Tries to detect different human speakers in the recording and exports separate files for each speaker.
If there are multiple speakers, it also creates background voices only, which keeps the secondary speakers and removes the main speaker.
Estimate background noise
Creates a cleaner version of the audio and also a file containing the estimated leftover noise.
Hugging Face token for speaker diarization
This is only needed for Separate speakers.
A Hugging Face token is a private access key from Hugging Face.
Speaker diarization means detecting who spoke when, so the app can split one recording into speaker-wise tracks.
If you are not using speaker separation, you can leave this box empty.
Minimum speaker segment length
Very short speech fragments can create messy results. Increase this to ignore tiny fragments. Lower it if you want more detailed speaker clips.
Noise reduction strength
Higher values remove more noise, but can also affect the original sound more strongly.