Separate Vocals, Speakers & Music
Free, Online, in Seconds
Music playing during your shoot? Unwanted noise in the background? Drop any audio or video file below — Perso Dubbing splits it into vocals, individual speakers, and background music, and you can hear every track before signing up.
No signup needed · First 60 seconds free · Files are never stored
Click or Drag & drop your file
Separation starts instantly — no account needed (Up to 200MB)
No file handy? Try a sample:
Separating audio tracks...
Analyzing sound frequencies to separate voice from ambient background elements
Edit speaker scripts line by line in the workspace
You've reached today's free limit
You've used all 3 free separations for today. Sign up free to keep going.
The Starter plan removes the daily limit.
World-class performance — measured, not claimed
Three industry-standard public benchmarks — MUSDB18 for vocal separation, VoiceBank-DEMAND for speech denoising, and the Open ASR Leaderboard for transcription. The same datasets every research paper uses, against named engines, with per-sample data published so anyone can re-run the tests.
Vocal separation higher = better
Wins on 44 of 50 tracks — and when we lose, the gap is at most 0.66 dB.
Noise removal quality higher = better
The specialist DeepFilterNet3 leads by a hair (2.77 vs 2.64) — both far ahead of ElevenLabs.
Speech clarity higher = better
Top two are effectively tied. ElevenLabs makes speech harder to understand on half of samples — we improve it on 96%.
Voice-clone fidelity higher = better
First on both cloning systems tested — even inside ElevenLabs' own cloner. The striped bar is the clean original: the natural ceiling.
Transcription accuracy (WER) lower = better
Overall a statistical tie with Scribe v2 — but on multi-speaker content like podcasts, we come out ahead (shorter bar = fewer errors).
Bars are zoomed into the competitive range so small gaps stay visible — the exact score next to each bar is what counts.
What do these tests actually measure?
🎯 Vocal separation (SI-SDR) Higher = better
How cleanly voice and music are pulled apart — like extracting a karaoke track with zero voice left in it. Our score: 10.67 dB vs HTDemucs 8.36 dB — less leaking between tracks, and we win on 44 of 50 songs.
🔊 Noise removal (PESQ · ESTOI) Higher = better
How clear and natural speech sounds after noise is stripped out — the same scoring used for phone-call quality. We score 2.64, a hair behind the specialist DeepFilterNet3 (2.77) and well ahead of ElevenLabs (2.38). On clarity, we tie for first.
📝 Transcription accuracy (WER) Lower = better
Out of 100 spoken words, how many get written down wrong. Our 7.61% means about 92 of 100 words right — statistically the same as ElevenLabs Scribe v2 (7.52%), and ahead of it on multi-speaker recordings like podcasts.
🎤 Voice-clone fidelity (cos_sim) Higher = better
After cleanup, does a voice clone made from that audio still sound like the same person? Scored 0 to 1 against the original voice. Our 0.674 ranks first on both cloning systems tested — including inside ElevenLabs' own cloner.
Honest footnotes: vocal separation is measured on the MUSDB18 sample set (full MUSDB18-HQ rerun in progress, expected within ±0.5 dB). DeepFilterNet3 edges us on PESQ by 0.15 — we tie on clarity and lead on waveform fidelity (+18.66 vs +17.31 dB SI-SDR). MDX-Net and LALAL.AI are not yet tested, so we don't claim to beat every separator. Verified May 2026.
Three steps, under a minute
Upload your file
Drag and drop an audio or video file — MP3, WAV, M4A, MP4, MOV, or WebM, up to 200MB. No account needed for the first 60 seconds.
Preview separated tracks
The AI splits your file into individual speakers, pure background music, and background with reactions. Play each track right in the browser.
Export your mix
Pick the tracks you need and export them as one file. Sign in to download, or to process the full length of longer files.
More than a vocal remover
😂 Dual background audio modes
Pure BGM, or BGM with laughter and applause kept intact. No other separation tool offers both from one upload.
👤 Multi-speaker separation
Not just vocals vs. music — speaker separation gives every person in the recording their own track, plus a speaker-labeled transcript in 99+ languages.
🔒 Nothing is stored
Trial files are processed in temporary storage and deleted when your session ends. Never kept, never used for training.
📝 Transcription in 99+ languages
Every separation includes automatic speech-to-text with speaker labels, shown right next to your tracks. Language detection is automatic — no extra tools, no extra steps.
🎬 Works with audio & video
Upload MP3, WAV, M4A, MP4, MOV, or WebM. Export tracks with embedded subtitles or separate SRT files.
🎚 Selective mix export
Combine any tracks into one file — Background Music plus Speaker 1, for example. No other separation tool exports a custom mix in one step.
Remove background music or noise from your video — two ways
A podcast laugh track, an audience reaction, a cough during a keynote — most vocal removers can't tell these from speech. Perso Dubbing gives you both options from a single upload.
Background Music
Removes every human sound — speech, laughter, claps — leaving only the background sound. Ideal for copyright-free BGM and clean audio beds for re-dubbing.
Background with Reaction
Removes only speech, keeping laughter, applause, and crowd energy intact. Perfect for podcasts, live events, and variety shows where atmosphere matters.
One track per voice — speaker separation for interviews, podcasts & meetings
Most vocal removers stop at two stems: voice and music. Perso Dubbing's multi speaker separation goes further — the AI detects how many people are talking and splits the recording into individual speaker tracks, each with a labeled transcript in 99+ languages.
One mixed recording
An interview, podcast, or meeting recording with several people talking over music and room noise — uploaded as a single audio or video file.
A separate track for every speaker
Separate speakers from audio in one click: export a single speaker's track, or any mix you choose — no manual editing.
Who uses audio separation?
🛡 Copyright resolution
Remove copyrighted BGM while keeping dialogue intact, swap in royalty-free music, and re-upload claim-free.
🎙 Podcast editing
Cut filler words and unwanted speech while keeping audience laughter and ambient reactions untouched.
🌍 Video dubbing
Extract a clean BGM track with zero speech bleed, then lay new voice-over in any of 99+ languages on top.
💼 Meetings & conferences
Separate speakers from audio in Zoom or Meet recordings — each participant gets their own track, with speaker-labeled transcripts built in.
📱 Social media clips
Swap the original BGM in short-form videos for a trending track — without touching your voiceover.
🎤 Concerts & fancams
Strip crowd noise and venue reverb from live clips to isolate the artist's voice or the music.
📰 Journalism & interviews
Use multi-speaker separation to pull each interviewee's voice out of noisy field recordings, with clean transcripts for fact-checking.
♻️ Repurpose content
One upload becomes podcast audio, promo BGM, speaker clips for social, and a full transcript for your blog.
Do more in the Perso workspace
Frequently asked questions
Is Perso Dubbing Audio Separation free to use?
Do I need to create an account to try audio separation?
What happens if my file is longer than 60 seconds?
Are my uploaded files stored on Perso Dubbing servers?
What file formats and sizes are supported?
What is the difference between Background Music and Background with Reaction mode?
Can Perso Dubbing do multi-speaker separation, not just vocals and music?
How accurate is Perso Dubbing audio separation compared to other tools?
Can I remove copyrighted background music from my video?
How do I remove background music from a video I filmed?
How is Perso Dubbing different from LALAL.AI or Moises?
Can I combine selected audio tracks into one file?
Try it on your own file — right now
The first 60 seconds are free. No signup, no stored files, no catch.
↑ Upload a file
EN
KO
ES
PT
RU
ID
DE
TH
JP
TC