Voice from Crowd

Removing noise from a .wav file, and picking a single voice out of a crowd, using nothing but the Fourier transform.

Two similar problems which we will try to solve using only the Fourier Transform. First: given a noisy recording of one person talking, strip out the noise. Second, harder: given a recording of several people talking over each other use a clean reference clip of the one person's voice, to pull that person's voice back out of the mixture. Everything here is written from scratch in Python (only scipy is used, for its fast FFT and STFT implementations). See GitHub.

Band filtering

Human speech mostly lives between roughly 400 Hz and 4000 Hz, so first we apply the 'Band-Pass Filter': take the Fourier transform of the whole signal, remove all noise outside a given band (in thise case 400 Hz and 4000 Hz), and transform back. \[ \hat X(f) = \begin{cases} X(f) & \text{if } 400 < |f| \le 4000 \\ 0 & \text{otherwise}\end{cases} \] This alone removes a lot of low rrrumble and high hissssss, but it can't touch noise that happens to sit inside the speech band itself. Which is most noise.

Spectral subtraction

To deal with within-band noise, the signal is first broken into overlapping short frames using the Short-Time Fourier Transform. Splitting into frames matters because noise and speech levels drift over time; a single global Fourier transform can't adapt to that. This allows us to use Band-Pass Filtering on each individual frame

For each frame, the quietest 10% of all noise is assumed to be unwanted rather than speech, and their magnitudes are averaged into a per-frame noise-floor estimate \(\hat N(\omega, k)\). That estimate is then subtracted from every bin's magnitude, floored so it can never go negative (or below some small residual \(\beta\hat N\)): \[ |\hat S(\omega,k)| = \max\bigl\{\,|Y(\omega,k)| - \alpha\hat N(\omega,k),\ \ \beta\hat N(\omega,k)\,\bigr\} \] Here \( |Y(\omega,k)|\) represents the original sound file, \(\alpha\hat N(\omega,k)\) represents the latent noise which we remove, and \(|\hat S(\omega,k)|\) represents the magnitude of the cleaned sound-file. Using the following identity \[ S(\omega,k) = |\hat S(\omega,k)|\cdot e^{i\theta(\omega,k)}, \qquad \theta = \arg Y(\omega,k), \] we convert \(|\hat S(\omega,k)|\) into \(S(\omega,k)\) and apply the inverse Fourier transform to turn it back into a sound-file.

Picking a voice out of a crowd

Removing noise is a different problem from removing other people, since another voice looks statistically like speech too. This is solved with a masking approach instead of subtraction, using a short reference clip of the target speaker's voice alone.

The reference clip's STFT is averaged over time to get a spectral fingerprint, in other words which frequencies that speaker's voice tends to occupy, and normalised to a unit vector \(p(\omega)\). Then, for every frame \(t\) of the noisy mixture, a cosine similarity is computed between that frame's spectral shape and the reference fingerprint: \[ \text{gate}(t) = \frac{\sum_\omega |Y(\omega,t)|\, p(\omega)}{\lVert Y({\cdot},t)\rVert \, \lVert p \rVert} \] This gate is close to 1 when a frame sounds like the target speaker, and close to 0 when it doesn't. This ensures frames dominated by someone else's voice, or silence, get suppressed. Every frequency bin of the mixture is then scaled by both the static per-frequency fingerprint and the per-frame gate: \[ \hat Y(\omega,t) = Y(\omega,t)\cdot p(\omega)\cdot \text{gate}(t) \] before taking the inverse STFT. Because masking quietens the whole signal, the result is finally rescaled by RMS to match the loudness of the original mixture.

Listen: single speaker, denoised

A single noisy recording, cleaned with only the Band-Pass Filter, and cleaned with both Band-pass and spectral subtraction.

Input — noisy
Output — band filter only
Output — denoise + filter

Listen: extracting one voice from a crowd

A mixture of several people talking at once, plus a short clean reference clip of the target speaker, masked down with the cosine-similarity method above.

Input — crowd mixture
Input — reference voice
Output — extracted voice
SUPER COMPUTER 3