Decoding the DNA of Music Genres: Insights from Amplitude Envelope Analysis
As humans, we have an innate ability to recognize and categorize different styles of music. Even without formal training, most of us can tell apart a classical symphony from a rock anthem, or a jazz improvisation from an electronic dance track. But what exactly are the sonic cues that our brains latch onto to make these distinctions?
One of the most fundamental yet informative aspects of musical sound is the amplitude envelope – the shape traced out by the peaks in the audio waveform over time. Amplitude envelopes encode the dynamic contours of music – the patterns of attack, decay, sustain and release that give each note its unique expressive character. As we‘ll see in this post, these dynamic shapes can vary considerably across musical genres, providing valuable clues about the compositional and production conventions that define different style.
Using Python and the librosa library, we‘ll embark on a data-driven exploration of amplitude envelopes across a diverse spectrum of musical genres. By leveraging signal processing and machine learning techniques, we‘ll not only visualize the characteristic envelope shapes associated with each genre, but also quantify their discriminative power in automatic genre classification tasks. Along the way, we‘ll also touch on applications in music information retrieval, recommendation systems, and generative modeling.
So put on your favorite pair of headphones, fire up your Jupyter notebook, and let‘s decode the DNA of musical style, one amplitude envelope at a time!
Dissecting the Amplitude Envelope
Before we dive into the genre analysis, let‘s quickly review what amplitude envelopes are and how to extract them from audio signals. Consider the waveform plot of a simple piano note:

The amplitude envelope traces out the outer contour of the waveform, tracking the changes in peak amplitude over time. We can see the sharp attack as the hammer strikes the string, the initial decay as the string‘s vibration begins to settle, the sustained period while the note is held, and the final release as the dampener silences the string.
In general, amplitude envelopes provide a compact representation of a sound‘s dynamic profile, independent of its pitch and harmonic content. They play a key role in our perception of timbre and texture, and can vary substantially between different instruments and recording techniques.
To extract the amplitude envelope from an audio signal, we essentially need to trace out the peaks of the waveform over time. A common approach is to split the signal into short, overlapping frames (say 1024 samples with a hop size of 128 samples), and take the maximum absolute amplitude within each frame. This can be easily implemented in Python using librosa:
def amplitude_envelope(signal, frame_size=1024, hop_length=128):
return np.array([max(signal[i:i+frame_size]) for i in range(0, signal.size, hop_length)])
It‘s often helpful to normalize the envelope by the maximum amplitude of the entire signal, so that we can compare the relative shapes of envelopes across different recordings:
def normalize(envelope):
return envelope / max(envelope)
Armed with these basic tools, let‘s now load in some audio samples from different genres and take a look at their amplitude envelopes!
Exploring Genre-Specific Envelope Shapes
Using the GTZAN dataset, a widely-used collection of 30-second audio clips spanning 10 genres, I extracted amplitude envelopes for the following genres:
- Blues
- Classical
- Country
- Disco
- Hip hop
- Jazz
- Metal
- Pop
- Reggae
- Rock
Here‘s the code to load the audio files and plot their normalized envelopes:
genres = [‘blues‘, ‘classical‘, ‘country‘, ‘disco‘, ‘hiphop‘, ‘jazz‘, ‘metal‘, ‘pop‘, ‘reggae‘, ‘rock‘]
audio_files = [f‘{genre}.00000.wav‘ for genre in genres]
plt.figure(figsize=(12,10))
for i, genre in enumerate(genres):
y, sr = librosa.load(audio_files[i])
envelope = amplitude_envelope(y)
norm_env = normalize(envelope)
plt.subplot(5,2,i+1)
librosa.display.waveplot(y, alpha=0.5)
plt.plot(norm_env, color=‘r‘)
plt.title(genre)
plt.tight_layout()
plt.show()
This gives us the following plot of overlaid envelopes for each genre:

Just from a visual inspection, we can notice some striking differences between genres:
-
Classical, jazz and reggae exhibit relatively smooth, sustained envelopes with gradual attacks and releases. This points to the dominance of legato phrasing and slow-moving harmonies in these styles.
-
Metal and rock show sharper, more percussive envelopes with faster attacks and more frequent peaks. This reflects the prominent role of drums and distorted electric guitars, which often employ rapid strumming and palm muting.
-
Hip hop and electronic music (represented here by the disco sample) tend to have periodic, regular envelope patterns, mirroring the use of drum machines, sequencers and digital audio workstations in these genres.
-
Pop exhibits a balanced mix of smooth and percussive envelopes, capturing its blend of vocal-driven melodies and snappy, quantized backup instrumentation.
Of course, these are broad generalizations based on just a small sample of each genre, and there will always be plenty of individual variation and outliers. To get a more robust sense of the envelope shapes characteristic of each genre, we would need to average the envelopes over many different recordings within each genre.
Nevertheless, these examples demonstrate that amplitude envelopes do capture meaningful information about the dynamic profiles of different musical styles, which we can attempt to quantify and exploit for automatic genre classification.
Genre Classification with Envelope Features
Now that we‘ve seen some qualitative differences in amplitude envelopes across genres, let‘s try to use them as features for a simple genre classifier. The idea is to see how well we can predict the genre of an audio clip just based on a statistical summary of its amplitude envelope.
We‘ll use the following features to represent each envelope:
- Mean amplitude
- Standard deviation
- Skewness
- Kurtosis
- 95th percentile amplitude
- 90th percentile amplitude
- 80th percentile amplitude
These features capture various aspects of the envelope‘s shape and distribution, such as its average level, variability, symmetry, peakedness and dynamic range. Here‘s a function to extract these features from a normalized envelope array:
def envelope_features(norm_env):
return [
np.mean(norm_env),
np.std(norm_env),
stats.skew(norm_env),
stats.kurtosis(norm_env),
np.percentile(norm_env, 95),
np.percentile(norm_env, 90),
np.percentile(norm_env, 80)
]
To test the predictive power of these features, I first extracted them for all the audio clips in the GTZAN dataset and stored them in a Pandas DataFrame along with the corresponding genre labels. Here‘s a snippet of the resulting feature table:
| Mean | Std | Skew | Kurt | Pct95 | Pct90 | Pct80 | Genre |
|---|---|---|---|---|---|---|---|
| 0.114 | 0.119 | 2.301 | 6.830 | 0.360 | 0.309 | 0.227 | pop |
| 0.062 | 0.087 | 2.755 | 9.254 | 0.247 | 0.196 | 0.117 | classical |
| 0.096 | 0.117 | 2.427 | 7.254 | 0.334 | 0.281 | 0.195 | jazz |
| 0.147 | 0.146 | 1.896 | 4.181 | 0.449 | 0.399 | 0.312 | metal |
| 0.120 | 0.121 | 2.129 | 5.895 | 0.369 | 0.324 | 0.244 | disco |
Next, I split the data into training and testing sets, normalized the features using Sci-kit Learn‘s StandardScaler, and trained a Random Forest classifier on the training set:
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
X = df[[‘Mean‘, ‘Std‘, ‘Skew‘, ‘Kurt‘, ‘Pct95‘, ‘Pct90‘, ‘Pct80‘]]
y = df[‘Genre‘]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train_scaled, y_train)
Finally, I evaluated the classifier‘s accuracy on the held-out test set:
from sklearn.metrics import accuracy_score
y_pred = rf.predict(X_test_scaled)
accuracy = accuracy_score(y_test, y_pred)
print(f‘Test accuracy: {accuracy:.3f}‘)
# Output: Test accuracy: 0.580
The Random Forest achieved an accuracy of 58% on the test set, which is significantly better than random guessing (which would give 10% accuracy for 10 balanced classes). This suggests that amplitude envelope features do carry some discriminative information for genre classification.
However, the accuracy is still far from perfect, indicating that envelope features alone are not sufficient to fully characterize musical genres. To improve performance, we would likely need to combine envelope features with other types of features capturing different aspects of the audio signal, such as:
- Spectral features (e.g. mel-frequency cepstral coefficients, chroma features)
- Rhythmic features (e.g. tempo, beat synchronous features)
- Tonal features (e.g. key, mode, chord progressions)
We could also experiment with more sophisticated classification models, such as deep neural networks that can learn hierarchical features directly from the raw audio or spectrogram representations.
Visualizing Genre Clusters in Envelope Feature Space
Another way to explore the relationship between amplitude envelopes and genres is to visualize how the different genres cluster or separate in the space of envelope features. This can give us insight into which genres have similar envelope characteristics and which are more distinct.
A common technique for visualizing high-dimensional data in 2D is t-distributed stochastic neighbor embedding (t-SNE). t-SNE is a non-linear dimensionality reduction method that seeks to preserve the local structure of the data while spreading out the clusters for easier interpretation.
Here‘s how we can apply t-SNE to our envelope features and plot the resulting genre clusters:
from sklearn.manifold import TSNE
tsne = TSNE(n_components=2, random_state=42)
X_tsne = tsne.fit_transform(X)
plt.figure(figsize=(10,8))
for genre in genres:
mask = (df[‘Genre‘] == genre)
plt.scatter(X_tsne[mask,0], X_tsne[mask,1], label=genre, alpha=0.7)
plt.legend(title=‘Genre‘)
plt.title(‘t-SNE visualization of amplitude envelope features‘)
plt.show()

In this plot, we can see that some genres form relatively distinct clusters, such as:
- Classical (blue): Concentrated in the upper left corner, reflecting the smooth, sustained envelopes typical of orchestral recordings.
- Metal (red): Concentrated in the lower right, characterized by loud, choppy envelopes with fast attacks from distorted guitars and aggressive drumming.
- Hip hop (purple): Forming a tight cluster in the center, indicative of the steady, periodic envelopes from drum machines and samples.
Other genres show more overlap and gradation, such as:
- Jazz (green) and blues (orange): Occupying similar regions of smooth, moderate envelopes with some outliers.
- Rock (brown), pop (pink) and country (gray): Covering a wide range from smooth to percussive, reflecting the diversity of production styles within these genres.
Again, it‘s important to take these visualizations with a grain of salt, as they are based on a limited sample and a relatively coarse feature representation. But they do suggest that amplitude envelopes capture some meaningful structure in the space of musical genres, even if not perfectly separable.
Summary and Future Directions
In this post, we took a deep dive into analyzing amplitude envelopes across different musical genres. We saw how to extract envelope features from audio signals using Python and librosa, and visualized some of the characteristic envelope shapes associated with each genre.
We then used these envelope features to train a basic genre classifier, which achieved moderate accuracy and gave us a sense of their discriminative power. Finally, we visualized the genre clusters in envelope feature space using t-SNE, revealing some interesting patterns of similarity and distinction.
Of course, this is just a small slice of the much larger field of music information retrieval and genre classification. Some potential avenues for future work include:
- Combining amplitude envelope features with other feature types (e.g. spectral, rhythmic, tonal) to improve genre classification accuracy
- Experimenting with more advanced machine learning models such as convolutional and recurrent neural networks
- Analyzing envelope features for non-Western genres and lesser-known sub-genres
- Using envelope features for other MIR tasks such as music segmentation, thumbnail generation and recommendation
- Exploring generative models that can learn to synthesize genre-specific amplitude envelopes
Amplitude envelopes are a fundamental building block of musical sound, and I hope this post has given you a glimpse of their untapped potential for understanding the rich landscape of musical style. So next time you‘re listening to your favorite tunes, take a moment to appreciate the dynamic ebbs and flows that make each genre groove in its own unique way!