NeuralSound
15.80 dB
Highest result in this test
Company-run audio benchmark · Published July 30, 2026
Compare NeuralSound, Moises and Fadr on the same five songs. Play all 35 input, isolated-vocal and instrumental previews, inspect the objective separation scores, and download the underlying benchmark data before choosing an AI vocal remover.
Published by Neural Sound LLC · Updated July 31, 2026 · Five-song, two-stem benchmark

Disclosure: NeuralSound conducted this benchmark. The same input songs and evaluation pipeline were used for every product. These five tracks do not prove that one service will perform best on every recording.
Quick answer
NeuralSound produced the highest average objective result in this five-song test. It reached an overall SI-SDR of 15.80 dB, compared with 14.47 dB for Moises and 12.59 dB for Fadr. Higher is better for these measurements, but the result applies only to this test set and date.
That score does not establish a universal winner. The best vocal remover still depends on the recording, the required stems, the acceptable level of vocal bleed or artifacts, and whether the user needs a simple two-stem split or a broader music-practice workflow. The players below let you verify the measured result by ear.
15.80 dB
Highest result in this test
14.47 dB
Second-highest result in this test
12.59 dB
Third result in this test
These values average the complete five-track test. Review the table together with the per-track audio because one aggregate score can hide differences between songs.
| What we compared | Fadr | Moises | NeuralSound |
|---|---|---|---|
| Vocal clarity (SI-SDR dB) | 10.02 | 12.00 | 13.26 |
| Instrumental clarity (SI-SDR dB) | 15.16 | 16.94 | 18.33 |
| Cleanliness / less bleed (SI-SIR dB) | 25.55 | 30.25 | 33.30 |
| Fewer processing artifacts (SI-SAR dB) | 12.89 | 14.64 | 15.93 |
| Overall quality (SI-SDR dB) | 12.59 | 14.47 | 15.80 |
| Historical entry option (checked July 2026) | Free Basic | Free: 5 songs/month | $1.99 / 10 minutes |
Inspect the averages or all 15 product-by-track result rows.
Each score describes a different part of source-separation quality. Use the measurements alongside the listening tests.
SI-SDR
Measures how closely an estimated stem matches its reference after accounting for scale. Higher values indicate a closer overall match.
SI-SIR
Measures interference left in the stem. A higher score generally means less accompaniment in the vocal or less vocal residue in the instrumental.
SI-SAR
Measures distortion introduced by the separation process. Higher values generally indicate fewer separation artifacts.
Technical reference: SI-SDR evaluation paper.
Play all 35 benchmark previews. Start with each original mixture, then compare the isolated vocals and no-vocal instrumentals from NeuralSound, Moises and Fadr. Only the selected sample loads; starting another stops and unloads the previous request.
Track 1 of 5
Same input · Two-stem mode · Scores in dB
Original input mixture
Average SI-SDR: 14.36 dB
Isolated vocals
Instrumental
Average SI-SDR: 13.29 dB
Isolated vocals
Instrumental
Average SI-SDR: 11.78 dB
Isolated vocals
Instrumental
Track 2 of 5
Same input · Two-stem mode · Scores in dB
Original input mixture
Average SI-SDR: 18.93 dB
Isolated vocals
Instrumental
Average SI-SDR: 16.81 dB
Isolated vocals
Instrumental
Average SI-SDR: 14.27 dB
Isolated vocals
Instrumental
Track 3 of 5
Same input · Two-stem mode · Scores in dB
Original input mixture
Average SI-SDR: 13.51 dB
Isolated vocals
Instrumental
Average SI-SDR: 12.63 dB
Isolated vocals
Instrumental
Average SI-SDR: 11.78 dB
Isolated vocals
Instrumental
Track 4 of 5
Same input · Two-stem mode · Scores in dB
Original input mixture
Average SI-SDR: 12.31 dB
Isolated vocals
Instrumental
Average SI-SDR: 11.52 dB
Isolated vocals
Instrumental
Average SI-SDR: 10.13 dB
Isolated vocals
Instrumental
Track 5 of 5
Same input · Two-stem mode · Scores in dB
Original input mixture
Average SI-SDR: 19.88 dB
Isolated vocals
Instrumental
Average SI-SDR: 18.11 dB
Isolated vocals
Instrumental
Average SI-SDR: 14.97 dB
Isolated vocals
Instrumental
Technology explained
An AI vocal remover is a music source-separation system that estimates the voices and accompaniment inside a finished mix. In two-stem mode, it returns an isolated vocal and an instrumental. It does not recover the original studio multitracks; it predicts usable stems from the mixed audio, so quality varies with the song and model.
An online vocal remover processes an uploaded song or video without requiring desktop separation software. The useful outputs are normally a vocal stem for acapella work and an instrumental stem for backing tracks, rehearsal, covers or further editing.
Judge the result by more than vocal loudness. Listen for lead-vocal residue in the instrumental, drums or melody leaking into the vocal, damaged transients, metallic textures and unstable reverb. The playable comparison above exposes those tradeoffs instead of relying on a marketing claim.
A vocal remover usually answers one focused question: can the song be separated into vocals and instrumental? An AI stem splitter can divide the same recording further into vocals, drums, bass, guitar, piano and an “other” stem for more detailed control.
Use two-stem separation when the final output is an acapella or no-vocal backing track. Choose multi-stem music separation when the task requires remixing, instrument practice, arrangement analysis or independent level control. A fair comparison should evaluate equivalent modes, inputs and output conditions.
These workflows share source-separation technology but solve different tasks. The correct output depends on whether you need every major instrument, a no-vocal mix, an isolated singing voice or clearer dialogue with the music reduced.
Music separation creates multiple controllable stems from one mixed file. Two-stem output is efficient for vocals and instrumental; four- or six-stem output gives producers, DJs, musicians and educators separate rhythmic and melodic parts for remixing, rehearsal and analysis.
Vocal removal aims to reduce the lead voice while preserving the accompaniment. For karaoke, covers and rehearsal, prioritize instrumental clarity and low vocal residue. In this five-song test, NeuralSound recorded the highest average instrumental SI-SDR at 18.33 dB.
Acapella extraction keeps the singing voice and reduces the backing music. Producers and vocalists should check consonants, harmonies, reverb tails and accompaniment bleed. Dense choruses and instruments that overlap the vocal range can challenge every separation model.
A background music remover keeps the main voice while reducing the musical bed around it. It can help with interviews, narration, podcasts and video dialogue, although overlapping speech, sound effects, room echo and loud music may still produce audible residue.
Select the smallest stem layout that completes the task. A focused vocal-and-instrumental split is easier to manage for karaoke or acapella creation, while detailed production work benefits from independent instrument stems.
| Separation mode | Outputs | Best suited to |
|---|---|---|
| 2-stem separation | Vocals and instrumental | Vocal removal, acapella extraction, backing tracks and reducing music behind a voice. |
| 4-stem separation | Vocals, drums, bass and other | Remixing, DJ edits, rhythm practice, sampling and basic arrangement study. |
| 6-stem separation | Vocals, drums, bass, guitar, piano and other | Detailed production, instrument practice, stem balancing and deeper song analysis. |
The same separated stems can support rehearsal, production, performance and dialogue editing. Define the final use first, then choose the stem layout and quality criteria that matter for that workflow.
A no-vocal instrumental supports rehearsal, covers and karaoke-style practice. Singers can study timing and phrasing against the backing track, while coaches can compare the isolated vocal with the accompaniment. Preserve the key source details by starting with the highest-quality file available.
Producers may need an acapella for remixing, drums and bass for sampling, or detailed stems for arrangement study. For this audience, a multi-stem separator is usually more useful than a basic voice remover because each musical layer can be edited independently.
Separated vocals, instrumentals and rhythm stems can support transitions, vocal overlays, mashups and performance edits. Listen closely for unstable ambience and transient damage because artifacts that seem minor in isolation may become obvious through club processing or repeated looping.
When speech and music are already mixed together, background music removal can make the voice easier to edit or replace in a timeline. This is a source-separation task, not ordinary noise reduction, and results depend on how strongly the dialogue overlaps with the music and effects.
Ready to test one of these workflows? Upload a track to NeuralSound’s music separation tool.
Choose the correct process
A background music remover and a noise remover solve different audio problems. Music is structured sound with melody, harmony and rhythm, so separating it from a voice requires source separation. Hiss, hum, wind, clicks and steady room noise are usually handled by denoising or speech-enhancement models.
Use background music removal when an interview, narration or vocal recording contains a music bed. Use denoising when the problem is microphone hiss, fan noise, traffic rumble or another non-musical disturbance. A file containing both may need two stages: separate the voice from the music, then clean the resulting voice stem.
This distinction prevents the wrong model from damaging useful vocal detail. It also explains why “isolate voice from music” and “remove noise from a recording” are related searches but not interchangeable technical tasks.
NeuralSound accepts MP3, WAV, FLAC and M4A audio, plus MP4, WebM, MOV and AVI video. The source format matters because a separation model cannot restore detail that was already removed by low-bitrate or repeated lossy compression.
WAV and FLAC are preferable for controlled evaluation, remixing and detailed editing because they preserve more source information. MP3 and M4A remain convenient for everyday use. Video uploads are useful when the voice and music are combined in an existing soundtrack that will return to an editing timeline.
There is no universal best vocal remover for every recording. Use the measured scores, playable outputs, required features and real cost together rather than selecting a service from one headline number.
Prioritize instrumental SI-SDR and listen for vocal residue, damaged cymbals, missing transients and unstable stereo ambience. NeuralSound achieved the highest average instrumental SI-SDR in this five-song benchmark: 18.33 dB.
Prioritize vocal SI-SDR, interference control and natural vocal texture. NeuralSound led this test’s average vocal SI-SDR at 13.26 dB and the combined less-bleed measurement at 33.30 dB, but audition the exact song before committing to a remix.
Look beyond two-stem performance. Drummers, bassists, guitarists and pianists may need four- or six-stem separation, looping, pitch or tempo controls and synchronized practice tools. Those capabilities were not scored by this two-stem benchmark.
Compare the actual minutes, export requirements and stem types you need. The entry options in the results table were checked in July 2026 and are historical snapshots, so verify current limits and prices on each service before purchasing.
The core workflow is similar across AI vocal-removal services. Start with the highest-quality lawful source available, choose the output needed for the project, and audition both stems before downloading.
Choose the audio or video to process. For serious comparison or editing, prefer WAV or FLAC when available.
Choose vocals and instrumental for a two-stem result, or a multi-stem model when individual instruments are required.
Check for vocal bleed, accompaniment leakage, metallic artifacts, damaged transients and unstable ambience.
Use the isolated vocal for acapella work or the instrumental for backing tracks, rehearsal and further editing.
Separation becomes harder when vocals and instruments overlap in time, pitch and timbre. Dense choruses, backing vocals, distortion, long reverb, live ambience and prior compression can all make the estimated stems less clean.
Compare accompaniment inside the vocal, lead-vocal residue inside the instrumental, missing cymbals or guitar attacks, phasey stereo effects, robotic textures and pumping ambience. A quieter stem is not automatically a cleaner stem.
Cloud separation products can change without publishing a model number. A credible comparison records the test date, input, service tier, stem mode, alignment procedure, metrics and limitations so readers can understand exactly what was—and was not—measured.
We tested the first five songs in the valid folder of the MUSDB18-HQ source used for this study. The identical mixture file was processed by NeuralSound, Moises and Fadr in a two-stem mode that produced an isolated vocal and an instrumental.
The original vocal stem was the vocal reference. The instrumental reference was calculated from the complete mixture minus the vocal source. Estimated outputs were converted to mono at 44.1 kHz and aligned to the references using cross-correlation.
Timing offsets and file-length differences were corrected only to synchronize the comparison. No denoising, equalization or quality enhancement was applied before scoring SI-SDR, SI-SIR and SI-SAR.
SI-SDR summarizes scale-invariant reconstruction quality, while SI-SIR and SI-SAR separate unwanted interference from introduced artifacts. Higher values indicate a closer reference match, less source leakage or fewer artifacts, depending on the metric.
MUSDB18 contains 150 stereo tracks with isolated vocals, drums, bass and “other” stems at 44.1 kHz. This is a small company-run evaluation, not an official MUSDB18 leaderboard. Fadr Free Basic and the Moises Free entry option were the tested labels; entry options were checked in July 2026.
Sources: MUSDB18 documentation and the SI-SDR paper.