Removing noise from a .wav file, and picking a single voice out of a crowd, using nothing but the Fourier transform.
Two similar problems which we will try to solve using only the Fourier Transform.
First: given a noisy recording of one person talking, strip out the noise.
Second, harder: given a recording of several people talking over each other
use a clean reference clip of the one person's voice, to pull
that person's voice back out of the mixture. Everything here is written from scratch in
Python (only scipy is used, for its fast FFT and STFT implementations). See
GitHub.
Human speech mostly lives between roughly 400 Hz and 4000 Hz, so first we apply the 'Band-Pass Filter': take the Fourier transform of the whole signal, remove all noise outside a given band (in thise case 400 Hz and 4000 Hz), and transform back. \[ \hat X(f) = \begin{cases} X(f) & \text{if } 400 < |f| \le 4000 \\ 0 & \text{otherwise}\end{cases} \] This alone removes a lot of low rrrumble and high hissssss, but it can't touch noise that happens to sit inside the speech band itself. Which is most noise.
To deal with within-band noise, the signal is first broken into overlapping short frames using the Short-Time Fourier Transform. Splitting into frames matters because noise and speech levels drift over time; a single global Fourier transform can't adapt to that. This allows us to use Band-Pass Filtering on each individual frame
For each frame, the quietest 10% of all noise is assumed to be unwanted rather than speech, and their magnitudes are averaged into a per-frame noise-floor estimate \(\hat N(\omega, k)\). That estimate is then subtracted from every bin's magnitude, floored so it can never go negative (or below some small residual \(\beta\hat N\)): \[ |\hat S(\omega,k)| = \max\bigl\{\,|Y(\omega,k)| - \alpha\hat N(\omega,k),\ \ \beta\hat N(\omega,k)\,\bigr\} \] Here \( |Y(\omega,k)|\) represents the original sound file, \(\alpha\hat N(\omega,k)\) represents the latent noise which we remove, and \(|\hat S(\omega,k)|\) represents the magnitude of the cleaned sound-file. Using the following identity \[ S(\omega,k) = |\hat S(\omega,k)|\cdot e^{i\theta(\omega,k)}, \qquad \theta = \arg Y(\omega,k), \] we convert \(|\hat S(\omega,k)|\) into \(S(\omega,k)\) and apply the inverse Fourier transform to turn it back into a sound-file.
Removing noise is a different problem from removing other people, since another voice looks statistically like speech too. This is solved with a masking approach instead of subtraction, using a short reference clip of the target speaker's voice alone.
The reference clip's STFT is averaged over time to get a spectral fingerprint, in other words which frequencies that speaker's voice tends to occupy, and normalised to a unit vector \(p(\omega)\). Then, for every frame \(t\) of the noisy mixture, a cosine similarity is computed between that frame's spectral shape and the reference fingerprint: \[ \text{gate}(t) = \frac{\sum_\omega |Y(\omega,t)|\, p(\omega)}{\lVert Y({\cdot},t)\rVert \, \lVert p \rVert} \] This gate is close to 1 when a frame sounds like the target speaker, and close to 0 when it doesn't. This ensures frames dominated by someone else's voice, or silence, get suppressed. Every frequency bin of the mixture is then scaled by both the static per-frequency fingerprint and the per-frame gate: \[ \hat Y(\omega,t) = Y(\omega,t)\cdot p(\omega)\cdot \text{gate}(t) \] before taking the inverse STFT. Because masking quietens the whole signal, the result is finally rescaled by RMS to match the loudness of the original mixture.
A single noisy recording, cleaned with only the Band-Pass Filter, and cleaned with both Band-pass and spectral subtraction.
A mixture of several people talking at once, plus a short clean reference clip of the target speaker, masked down with the cosine-similarity method above.