Company-run audio benchmark · Published July 30, 2026

AI Vocal Remover Benchmark 2026: NeuralSound vs Moises vs Fadr

Compare NeuralSound, Moises and Fadr on the same five songs. Play all 35 input, isolated-vocal and instrumental previews, inspect the objective separation scores, and download the underlying benchmark data before choosing an AI vocal remover.

Published by Neural Sound LLC · Updated July 31, 2026 · Five-song, two-stem benchmark

AI Vocal Remover Benchmark 2026 comparing NeuralSound, Moises and Fadr

Disclosure: NeuralSound conducted this benchmark. The same input songs and evaluation pipeline were used for every product. These five tracks do not prove that one service will perform best on every recording.

Quick answer

Quick answer: which AI vocal remover performed best?

NeuralSound produced the highest average objective result in this five-song test. It reached an overall SI-SDR of 15.80 dB, compared with 14.47 dB for Moises and 12.59 dB for Fadr. Higher is better for these measurements, but the result applies only to this test set and date.

That score does not establish a universal winner. The best vocal remover still depends on the recording, the required stems, the acceptable level of vocal bleed or artifacts, and whether the user needs a simple two-stem split or a broader music-practice workflow. The players below let you verify the measured result by ear.

NeuralSound

15.80 dB

Highest result in this test

Moises

14.47 dB

Second-highest result in this test

Fadr

12.59 dB

Third result in this test

Overall comparison across all five songs

These values average the complete five-track test. Review the table together with the per-track audio because one aggregate score can hide differences between songs.

Average benchmark results for Fadr, Moises and NeuralSound
What we comparedFadrMoisesNeuralSound
Vocal clarity (SI-SDR dB)10.0212.0013.26
Instrumental clarity (SI-SDR dB)15.1616.9418.33
Cleanliness / less bleed (SI-SIR dB)25.5530.2533.30
Fewer processing artifacts (SI-SAR dB)12.8914.6415.93
Overall quality (SI-SDR dB)12.5914.4715.80
Historical entry option (checked July 2026)Free BasicFree: 5 songs/month$1.99 / 10 minutes

Download the benchmark data

Inspect the averages or all 15 product-by-track result rows.

What SI-SDR, SI-SIR and SI-SAR measure

Each score describes a different part of source-separation quality. Use the measurements alongside the listening tests.

SI-SDR

Overall reconstruction quality

Measures how closely an estimated stem matches its reference after accounting for scale. Higher values indicate a closer overall match.

SI-SIR

Separation from unwanted sources

Measures interference left in the stem. A higher score generally means less accompaniment in the vocal or less vocal residue in the instrumental.

SI-SAR

Processing artifacts

Measures distortion introduced by the separation process. Higher values generally indicate fewer separation artifacts.

Technical reference: SI-SDR evaluation paper.

Interactive AI vocal remover audio comparison

Play all 35 benchmark previews. Start with each original mixture, then compare the isolated vocals and no-vocal instrumentals from NeuralSound, Moises and Fadr. Only the selected sample loads; starting another stops and unloads the previous request.

Track 1 of 5

The Districts - Vermont

Same input · Two-stem mode · Scores in dB

Original input mixture

NeuralSound

Average SI-SDR: 14.36 dB

Highest

Isolated vocals

SI-SDR
13.01
SI-SIR
29.76
SI-SAR
13.11

Instrumental

SI-SDR
15.70
SI-SIR
28.84
SI-SAR
15.92

Moises

Average SI-SDR: 13.29 dB

Isolated vocals

SI-SDR
11.96
SI-SIR
26.70
SI-SAR
12.11

Instrumental

SI-SDR
14.61
SI-SIR
27.50
SI-SAR
14.85

Fadr

Average SI-SDR: 11.78 dB

Isolated vocals

SI-SDR
10.40
SI-SIR
23.50
SI-SAR
10.64

Instrumental

SI-SDR
13.16
SI-SIR
23.92
SI-SAR
13.56

Track 2 of 5

The Long Wait - Back Home To Blue

Same input · Two-stem mode · Scores in dB

Original input mixture

NeuralSound

Average SI-SDR: 18.93 dB

Highest

Isolated vocals

SI-SDR
15.65
SI-SIR
40.05
SI-SAR
15.67

Instrumental

SI-SDR
22.21
SI-SIR
38.63
SI-SAR
22.31

Moises

Average SI-SDR: 16.81 dB

Isolated vocals

SI-SDR
13.53
SI-SIR
34.98
SI-SAR
13.57

Instrumental

SI-SDR
20.09
SI-SIR
34.67
SI-SAR
20.24

Fadr

Average SI-SDR: 14.27 dB

Isolated vocals

SI-SDR
10.94
SI-SIR
27.53
SI-SAR
11.04

Instrumental

SI-SDR
17.60
SI-SIR
29.57
SI-SAR
17.89

Track 3 of 5

The Scarlet Brand - Les Fleurs Du Mal

Same input · Two-stem mode · Scores in dB

Original input mixture

NeuralSound

Average SI-SDR: 13.51 dB

Highest

Isolated vocals

SI-SDR
10.59
SI-SIR
30.15
SI-SAR
10.65

Instrumental

SI-SDR
16.43
SI-SIR
28.65
SI-SAR
16.70

Moises

Average SI-SDR: 12.63 dB

Isolated vocals

SI-SDR
9.99
SI-SIR
28.51
SI-SAR
10.05

Instrumental

SI-SDR
15.28
SI-SIR
26.86
SI-SAR
15.60

Fadr

Average SI-SDR: 11.78 dB

Isolated vocals

SI-SDR
8.85
SI-SIR
24.85
SI-SAR
8.98

Instrumental

SI-SDR
14.71
SI-SIR
24.12
SI-SAR
15.26

Track 4 of 5

The So So Glos - Emergency

Same input · Two-stem mode · Scores in dB

Original input mixture

NeuralSound

Average SI-SDR: 12.31 dB

Highest

Isolated vocals

SI-SDR
9.32
SI-SIR
26.86
SI-SAR
9.41

Instrumental

SI-SDR
15.30
SI-SIR
25.88
SI-SAR
15.71

Moises

Average SI-SDR: 11.52 dB

Isolated vocals

SI-SDR
8.53
SI-SIR
24.30
SI-SAR
8.66

Instrumental

SI-SDR
14.52
SI-SIR
24.66
SI-SAR
14.97

Fadr

Average SI-SDR: 10.13 dB

Isolated vocals

SI-SDR
7.05
SI-SIR
20.79
SI-SAR
7.28

Instrumental

SI-SDR
13.22
SI-SIR
21.33
SI-SAR
13.98

Track 5 of 5

The Wrong'Uns - Rothko

Same input · Two-stem mode · Scores in dB

Original input mixture

NeuralSound

Average SI-SDR: 19.88 dB

Highest

Isolated vocals

SI-SDR
17.74
SI-SIR
41.21
SI-SAR
17.76

Instrumental

SI-SDR
22.02
SI-SIR
43.00
SI-SAR
22.05

Moises

Average SI-SDR: 18.11 dB

Isolated vocals

SI-SDR
15.99
SI-SIR
36.98
SI-SAR
16.02

Instrumental

SI-SDR
20.22
SI-SIR
37.31
SI-SAR
20.31

Fadr

Average SI-SDR: 14.97 dB

Isolated vocals

SI-SDR
12.87
SI-SIR
29.20
SI-SAR
12.98

Instrumental

SI-SDR
17.08
SI-SIR
30.71
SI-SAR
17.28

Technology explained

What is an AI vocal remover?

An AI vocal remover is a music source-separation system that estimates the voices and accompaniment inside a finished mix. In two-stem mode, it returns an isolated vocal and an instrumental. It does not recover the original studio multitracks; it predicts usable stems from the mixed audio, so quality varies with the song and model.

Remove vocals from a song online

An online vocal remover processes an uploaded song or video without requiring desktop separation software. The useful outputs are normally a vocal stem for acapella work and an instrumental stem for backing tracks, rehearsal, covers or further editing.

Judge the result by more than vocal loudness. Listen for lead-vocal residue in the instrumental, drums or melody leaking into the vocal, damaged transients, metallic textures and unstable reverb. The playable comparison above exposes those tradeoffs instead of relying on a marketing claim.

Open NeuralSound’s online AI vocal remover.

Vocal remover vs AI stem splitter

A vocal remover usually answers one focused question: can the song be separated into vocals and instrumental? An AI stem splitter can divide the same recording further into vocals, drums, bass, guitar, piano and an “other” stem for more detailed control.

Use two-stem separation when the final output is an acapella or no-vocal backing track. Choose multi-stem music separation when the task requires remixing, instrument practice, arrangement analysis or independent level control. A fair comparison should evaluate equivalent modes, inputs and output conditions.

Explore NeuralSound’s AI music separator.

Music separation, vocal removal, acapella extraction and background music removal

These workflows share source-separation technology but solve different tasks. The correct output depends on whether you need every major instrument, a no-vocal mix, an isolated singing voice or clearer dialogue with the music reduced.

Music separation for vocals and instruments

Music separation creates multiple controllable stems from one mixed file. Two-stem output is efficient for vocals and instrumental; four- or six-stem output gives producers, DJs, musicians and educators separate rhythmic and melodic parts for remixing, rehearsal and analysis.

Try NeuralSound music separation →

Vocal remover for instrumentals and backing tracks

Vocal removal aims to reduce the lead voice while preserving the accompaniment. For karaoke, covers and rehearsal, prioritize instrumental clarity and low vocal residue. In this five-song test, NeuralSound recorded the highest average instrumental SI-SDR at 18.33 dB.

Use the NeuralSound vocal remover →

Acapella extraction for isolated vocals

Acapella extraction keeps the singing voice and reduces the backing music. Producers and vocalists should check consonants, harmonies, reverb tails and accompaniment bleed. Dense choruses and instruments that overlap the vocal range can challenge every separation model.

Extract an acapella with NeuralSound →

Background music removal for speech, singing and video

A background music remover keeps the main voice while reducing the musical bed around it. It can help with interviews, narration, podcasts and video dialogue, although overlapping speech, sound effects, room echo and loud music may still produce audible residue.

Explore background music removal →

Choose the right music separation mode

Select the smallest stem layout that completes the task. A focused vocal-and-instrumental split is easier to manage for karaoke or acapella creation, while detailed production work benefits from independent instrument stems.

Separation modeOutputsBest suited to
2-stem separationVocals and instrumentalVocal removal, acapella extraction, backing tracks and reducing music behind a voice.
4-stem separationVocals, drums, bass and otherRemixing, DJ edits, rhythm practice, sampling and basic arrangement study.
6-stem separationVocals, drums, bass, guitar, piano and otherDetailed production, instrument practice, stem balancing and deeper song analysis.
NeuralSound workflow: use two-stem separation when the desired outputs are vocals and instrumental. Choose four or six stems when drums, bass, guitar, piano or other instruments must be controlled separately.

Music separation workflows for singers, producers, DJs and video creators

The same separated stems can support rehearsal, production, performance and dialogue editing. Define the final use first, then choose the stem layout and quality criteria that matter for that workflow.

Singers, vocal coaches and karaoke users

A no-vocal instrumental supports rehearsal, covers and karaoke-style practice. Singers can study timing and phrasing against the backing track, while coaches can compare the isolated vocal with the accompaniment. Preserve the key source details by starting with the highest-quality file available.

Music producers and remix creators

Producers may need an acapella for remixing, drums and bass for sampling, or detailed stems for arrangement study. For this audience, a multi-stem separator is usually more useful than a basic voice remover because each musical layer can be edited independently.

DJs, mashup artists and live performers

Separated vocals, instrumentals and rhythm stems can support transitions, vocal overlays, mashups and performance edits. Listen closely for unstable ambience and transient damage because artifacts that seem minor in isolation may become obvious through club processing or repeated looping.

Podcasts, interviews and video dialogue

When speech and music are already mixed together, background music removal can make the voice easier to edit or replace in a timeline. This is a source-separation task, not ordinary noise reduction, and results depend on how strongly the dialogue overlaps with the music and effects.

Ready to test one of these workflows? Upload a track to NeuralSound’s music separation tool.

Choose the correct process

Background music remover vs noise remover

A background music remover and a noise remover solve different audio problems. Music is structured sound with melody, harmony and rhythm, so separating it from a voice requires source separation. Hiss, hum, wind, clicks and steady room noise are usually handled by denoising or speech-enhancement models.

Use background music removal when an interview, narration or vocal recording contains a music bed. Use denoising when the problem is microphone hiss, fan noise, traffic rumble or another non-musical disturbance. A file containing both may need two stages: separate the voice from the music, then clean the resulting voice stem.

This distinction prevents the wrong model from damaging useful vocal detail. It also explains why “isolate voice from music” and “remove noise from a recording” are related searches but not interchangeable technical tasks.

MP3, WAV, FLAC and video files for vocal removal

NeuralSound accepts MP3, WAV, FLAC and M4A audio, plus MP4, WebM, MOV and AVI video. The source format matters because a separation model cannot restore detail that was already removed by low-bitrate or repeated lossy compression.

WAV and FLAC are preferable for controlled evaluation, remixing and detailed editing because they preserve more source information. MP3 and M4A remain convenient for everyday use. Video uploads are useful when the voice and music are combined in an existing soundtrack that will return to an editing timeline.

Which vocal remover is best for your use case?

There is no universal best vocal remover for every recording. Use the measured scores, playable outputs, required features and real cost together rather than selecting a service from one headline number.

Best criteria for karaoke and backing tracks

Prioritize instrumental SI-SDR and listen for vocal residue, damaged cymbals, missing transients and unstable stereo ambience. NeuralSound achieved the highest average instrumental SI-SDR in this five-song benchmark: 18.33 dB.

Best criteria for acapellas and remixing

Prioritize vocal SI-SDR, interference control and natural vocal texture. NeuralSound led this test’s average vocal SI-SDR at 13.26 dB and the combined less-bleed measurement at 33.30 dB, but audition the exact song before committing to a remix.

Best criteria for instrument practice

Look beyond two-stem performance. Drummers, bassists, guitarists and pianists may need four- or six-stem separation, looping, pitch or tempo controls and synchronized practice tools. Those capabilities were not scored by this two-stem benchmark.

Best criteria for occasional use

Compare the actual minutes, export requirements and stem types you need. The entry options in the results table were checked in July 2026 and are historical snapshots, so verify current limits and prices on each service before purchasing.

How to remove vocals from a song online

The core workflow is similar across AI vocal-removal services. Start with the highest-quality lawful source available, choose the output needed for the project, and audition both stems before downloading.

  1. 1

    Upload the original file

    Choose the audio or video to process. For serious comparison or editing, prefer WAV or FLAC when available.

  2. 2

    Select the separation mode

    Choose vocals and instrumental for a two-stem result, or a multi-stem model when individual instruments are required.

  3. 3

    Preview both outputs

    Check for vocal bleed, accompaniment leakage, metallic artifacts, damaged transients and unstable ambience.

  4. 4

    Download the required stem

    Use the isolated vocal for acapella work or the instrumental for backing tracks, rehearsal and further editing.

Rights reminder: removing or isolating vocals does not transfer copyright ownership or automatically permit commercial reuse. Users remain responsible for the licenses and permissions required for public performance, redistribution and commercial projects.

What affects AI music-separation quality?

Separation becomes harder when vocals and instruments overlap in time, pitch and timbre. Dense choruses, backing vocals, distortion, long reverb, live ambience and prior compression can all make the estimated stems less clean.

What listeners should check

Compare accompaniment inside the vocal, lead-vocal residue inside the instrumental, missing cymbals or guitar attacks, phasey stereo effects, robotic textures and pumping ambience. A quieter stem is not automatically a cleaner stem.

Why dates and methodology matter

Cloud separation products can change without publishing a model number. A credible comparison records the test date, input, service tier, stem mode, alignment procedure, metrics and limitations so readers can understand exactly what was—and was not—measured.

Methodology

We tested the first five songs in the valid folder of the MUSDB18-HQ source used for this study. The identical mixture file was processed by NeuralSound, Moises and Fadr in a two-stem mode that produced an isolated vocal and an instrumental.

The original vocal stem was the vocal reference. The instrumental reference was calculated from the complete mixture minus the vocal source. Estimated outputs were converted to mono at 44.1 kHz and aligned to the references using cross-correlation.

Timing offsets and file-length differences were corrected only to synchronize the comparison. No denoising, equalization or quality enhancement was applied before scoring SI-SDR, SI-SIR and SI-SAR.

SI-SDR summarizes scale-invariant reconstruction quality, while SI-SIR and SI-SAR separate unwanted interference from introduced artifacts. Higher values indicate a closer reference match, less source leakage or fewer artifacts, depending on the metric.

MUSDB18 contains 150 stereo tracks with isolated vocals, drums, bass and “other” stems at 44.1 kHz. This is a small company-run evaluation, not an official MUSDB18 leaderboard. Fadr Free Basic and the Moises Free entry option were the tested labels; entry options were checked in July 2026.

Sources: MUSDB18 documentation and the SI-SDR paper.

Limitations and interpretation

  • Company-run study: NeuralSound designed and conducted the benchmark, so the disclosure should be considered when interpreting the result.
  • Small sample: Five tracks cannot represent every genre, production style, recording condition or source overlap.
  • Changing services: Cloud products can update their separation systems after the July 2026 test date.
  • No universal winner: Objective scores help compare reconstruction, interference and artifacts, but listeners may prefer a different output for a particular song or workflow.

Hear how NeuralSound handles your track

Upload an audio or video file and compare the separated vocals and instruments for yourself.