Tutorial: How to Visualize Feature Maps from Convolutional Neural Network Layers
Introduction to Convolutional Neural Networks
Convolutional Neural Networks, or CNNs for short, are a specialized type of deep learning model commonly used for computer vision tasks like image classification, object detection, and segmentation. Unlike regular feedforward neural networks, CNNs are designed to process data with a grid-like structure, such as 2D images.
The key building blocks of a CNN are:
-
Convolutional layers: These layers perform convolution operations to extract features from the input. Multiple filters slide across the input, computing dot products to produce feature maps.
-
Pooling layers: Pooling reduces the spatial dimensions of the feature maps, helping the network become invariant to small translations. Max pooling and average pooling are commonly used.
-
Fully connected layers: After passing through a series of convolutional and pooling layers, the final feature maps are flattened and fed into one or more fully connected layers to perform the final prediction.
As the input progresses through the layers of a CNN, the network learns to detect increasingly complex and abstract visual features. Early layers may detect simple patterns like edges and textures, while later layers can identify objects, parts, and high-level concepts.
Understanding Feature Maps
Feature maps are the outputs produced by each filter in a convolutional layer. If a conv layer has 64 filters, it will output 64 separate feature maps. Each feature map highlights the presence of a specific visual pattern learned by that filter.
Mathematically, a feature map is the result of computing the dot product between a filter‘s weights and a portion of the input. This is done repeatedly as the filter slides across the input. The feature map ends up being a 2D grid of activation values indicating how strongly the filter‘s pattern matches different regions of the input.

Visualizing these feature maps can provide valuable insights into what a CNN is learning and how it is making predictions. By examining the activation patterns, we can understand which visual features the network considers important for a given task. Feature map visualizations are a useful tool for interpreting and debugging CNNs.
Step-by-Step Tutorial
Now let‘s walk through the process of visualizing feature maps from a CNN layer by layer. We‘ll use the popular deep learning framework TensorFlow and its high-level API Keras. The code will be in Python.
1. Build a CNN Model
First, we need to define a CNN model to extract feature maps from. For this example, we‘ll create a simple model to classify images into 10 categories.
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense
model = Sequential([
Conv2D(32, (3,3), activation=‘relu‘, input_shape=(32, 32, 3)),
MaxPooling2D((2,2)),
Conv2D(64, (3,3), activation=‘relu‘),
MaxPooling2D((2,2)),
Conv2D(64, (3,3), activation=‘relu‘),
Flatten(),
Dense(64, activation=‘relu‘),
Dense(10, activation=‘softmax‘)
])
This CNN has three convolutional layers with 32 and 64 filters each, max pooling layers, and ends with fully connected layers to output class predictions. The input shape is 32×32 RGB images.
2. Access Intermediate Layer Outputs
To visualize the feature maps, we need a way to access the outputs of the intermediate layers in our model. Keras provides the Model class which allows us to specify a model‘s inputs and outputs.
We can use this to define a new model that takes the same input as our original model, but returns the outputs of all the convolutional layers:
from tensorflow.keras.models import Model
layer_outputs = [layer.output for layer in model.layers[:5]]
features_model = Model(inputs=model.input, outputs=layer_outputs)
Here we extract the outputs of the first 5 layers (the 3 conv layers and 2 max pooling layers). The features_model will return the feature maps from each of these layers.
3. Generate Feature Map Visualizations
Now we have a model that gives us access to the feature maps. To generate the visualizations, we‘ll feed an input image through the model and plot the resulting activations.
import matplotlib.pyplot as plt
import numpy as np
from tensorflow.keras.preprocessing.image import load_img, img_to_array
# Load and preprocess an input image
img = load_img(‘example_image.jpg‘, target_size=(32, 32))
img = img_to_array(img)
img = img.reshape((1, *img.shape))
img = img / 255.
# Get feature map activations
feature_maps = features_model.predict(img)
# Plot the feature maps
for layer_num, feature_map in enumerate(feature_maps, start=1):
print(f"Visualizing layer {layer_num}")
cols = 8
rows = np.ceil(feature_map.shape[-1] / cols)
fig = plt.figure(figsize=(2*cols, 2*rows))
for i in range(0, feature_map.shape[-1]):
fig.add_subplot(rows, cols, i+1)
plt.imshow(feature_map[0, :, :, i], cmap=‘viridis‘)
plt.axis(‘off‘)
plt.tight_layout()
plt.show()
We use the load_img and img_to_array functions to load and preprocess the input image. The features_model is called to get the feature map activations in feature_maps.
Then we iterate through the feature maps from each layer and plot them in a grid using Matplotlib. The viridis colormap is used to visualize the activation values.


Layer 1 feature maps tend to be simple, picking up basic visual patterns like edges and textures. Deeper layers like layer 2 show more complex features being detected.
We can repeat this for all the layers to get a full picture of how the features evolve through the network.
4. Enhancing the Visualizations (Optional)
The raw feature maps can sometimes be quite dim or low contrast, making it hard to see the activations. We can apply some transformations to make them more visually interpretable:
def to_vis(fmap):
fmap -= fmap.mean()
fmap /= fmap.std() + 1e-5
fmap *= 0.1
fmap += 0.5
fmap = np.clip(fmap, 0, 1)
fmap *= 255
fmap = np.clip(fmap, 0, 255).astype(‘uint8‘)
return fmap
for layer_num, feature_map in enumerate(feature_maps, start=1):
feature_map = to_vis(feature_map)
...
The to_vis function standardizes the feature map, increases its contrast, and clips the values to a 0-255 range. This results in clearer, brighter visualizations.
Interpreting the Visualizations
Visualizing feature maps allows us to peek inside the "black box" of a CNN and gain insights into its inner workings. Some key things to look for:
-
The complexity of the visual patterns detected by each layer. Early layers respond to simple textures and lines, while deeper layers recognize objects and parts.
-
The sparsity of the activations. Sparse feature maps with mostly dark pixels and a few strong activations may indicate the layer has learned a small set of very specific patterns. More diffuse maps suggest the layer is sensitive to a broader range of inputs.
-
The spatial localization of the activations. Feature maps preserve spatial information from the input. Activations in a certain region of a feature map mean the layer has detected a pattern in the corresponding location of the input image.
-
Similarities and differences between feature maps within a layer. Maps that look similar are detecting related patterns, while diverse maps imply the layer has learned a variety of distinct features.
By analyzing feature map visualizations in this way, we can better understand the representations learned by a CNN and diagnose issues like underutilized layers or filters.
Applications and Importance
Feature map visualization is a valuable tool with several key applications:
-
Model interpretability: Visualizations provide a way to explain and interpret the predictions of a CNN. This is crucial for building trust in models deployed in real-world settings.
-
Debugging and improving models: Examining feature maps can reveal problems like dead filters, vanishing gradients, or overfitting. Insights from visualizations can guide improvements to model architectures and hyperparameters.
-
Assessing transferability: Comparing feature maps between models trained on different datasets can indicate how well features transfer to new tasks. More general features are more likely to be useful for transfer learning.
-
Uncovering latent structure: Visualizations can reveal meaningful hierarchical or semantic structures learned by a model, such as identifying class-specific features. This can inspire scientific insights into the problem domain.
As CNNs are increasingly deployed in sensitive domains like medical diagnosis, autonomous driving, and security, the interpretability provided by feature map visualizations is becoming more critical. Regulators, users, and researchers need to be able to audit and understand the decision-making of CNN models.
Conclusion
In this tutorial, we explored feature map visualization for convolutional neural networks. We saw how to access intermediate layer activations, generate feature map visualizations, and enhance them for clarity.
Interpreting these visualizations provides valuable insights into the learned representations of a CNN. Early layers detect simple patterns, while deeper layers capture more complex, abstract features. The spatial localization and distribution of activations also reveal important characteristics of the model.
Feature map visualization is a key tool for CNN interpretability. It enables researchers and practitioners to debug models, assess transferability, and uncover meaningful latent structure. As CNNs are applied to increasingly high-stakes problems, the ability to visualize and understand their feature maps will only become more important.
With the knowledge from this tutorial, you are now equipped to visualize the feature maps of your own CNN models and gain deeper insights into how they work. By combining feature map visualization with other interpretation techniques like saliency maps and ablation studies, you can thoroughly probe and explain your models to build more robust, trustworthy CNNs.
I hope you found this guide informative and practical. Feel free to experiment further with different CNN architectures, datasets, and visualization techniques. Happy feature map visualizing!