abracadabra: How does Shazam work? - Cameron MacLeod
Cameron MacLeod
About
CV
Projects
abracadabra: How does Shazam work?
Sat 19 February 2022<br>Tutorials
Your phone's ability to identify any song it listens to is pure technological magic. In this article, I'll show you how one of the most popular apps, Shazam, does it. The founders of Shazam released a paper in 2003 documenting how it works, and I have been working on an implementation of that paper, abracadabra.
Where the paper doesn't explain something, I will fill in the gaps with how abracadabra approaches it. I've also included links to the corresponding part of the abracadabra codebase in relevant sections so you can follow along in Python if you prefer.
The state of the art has moved on since this paper, and Shazam has probably evolved its algorithm. However, the core principles of audio identification systems haven't changed, and the accuracy you can obtain using the original Shazam method is impressive.
To get the most out of this article, you should understand:
Frequency and pitch
Waves
Graphs and axes
Quick links
What is Shazam?
Why is song recognition hard anyway?
System overview
Calculating a spectrogram<br>The Fourier transform
Spectrograms
Fingerprinting<br>Why is the fingerprint based on spectrogram peaks?
Finding peaks
Hashing
Matching
Conclusion
Enter abracadabra
Further reading
What is Shazam?#
Shazam is an app that identifies songs that are playing around you. You open the app while music is playing, and Shazam will record a few seconds of audio which it uses to search its database. Once it identifies the song that's playing, it will display the result on screen.
Shazam recognising a song
Before Shazam was an app, it was a phone number. To identify a song, you would ring up the number and hold your phone's microphone to the music. After 30 seconds, Shazam would hang up and then text you details on the song you were listening to. If you were using a mobile phone back in 2002, you'll understand that the quality of phone calls back then made this a challenging task!
Why is song recognition hard anyway?#
If you haven't done much signal processing before, it may not be obvious why this is a difficult problem to solve. To help give you an idea, take a look at the following audio:
The above graph shows what Chris Cornell's "Like a Stone" looks like when stored in a computer. Now take a look at the following section of the track:
If you wanted to tell whether this section of audio came from the track above, you could use a brute-force method. For example, you could slide the section of audio along the track and see if it matches at any point:
Matching a section of track by sliding it
This would be a bit slow, but it would work. Now imagine that you didn't know which track this audio came from, and you had a database of 10 million songs to search. This would take a lot longer!
What's worse, when you move from this toy example to samples that are recorded through a microphone you introduce background noise, frequency effects, amplitude changes and more. All of these can change the shape of the audio significantly. The sliding method just doesn't work that well for this problem.
Thankfully, Shazam's approach is a lot smarter than that. In the next section, you'll see the high-level overview of how this works.
System overview#
If Shazam doesn't take the sliding approach we described above, what does it do? Take a look at the following high-level diagram:
The first thing you will notice is that the diagram is split up into register and recognise flows. The register flow remembers a song to enable it to be recognised in the future. The recognise flow identifies a short section of audio.
Registering a song and identifying some audio share a lot of commonality. The following sections will go into more detail, but both flows have the following steps:
Calculate the spectrogram of the song/audio. This is a graph of frequency against time. We'll talk more about spectrograms later.
Find peaks in that spectrogram. These represent the loudest frequencies in the audio and will help us build a fingerprint.
Hash these peaks. In short, this means pairing peaks up to make a better fingerprint.
After calculating these hashes, the register flow will store them in the database. The recognise flow will compare them to hashes already in the database to identify which song is playing through the matching step.
In the next few sections, you'll learn more about each of these steps.
Calculating a spectrogram#
The first step for both flows is to obtain a spectrogram of the audio being registered or recognised. To understand spectrograms, you first have to understand Fourier transforms.
The Fourier transform#
A Fourier transform takes some audio and tells you which frequencies are present in that audio. For example, if you took a 20 Hertz sine wave and used the Fourier transform on it, you would see a big spike around 20 Hertz...