Convolutional Neural Networks
Last updated on 2026-04-15 | Edit this page
Estimated time: 65 minutes
Overview
Questions
- Why are CNNs better suited for images than standard neural networks?
- What is a convolutional filter and how does it work?
- What is a feature map and what does it represent?
- Why is weight sharing important in CNNs?
- How do early and deeper layers differ in what they learn?
Objectives
- Understand what convolutional neural networks (CNNs) are designed for
- Explain how convolutional layers extract features from data
- Recognize why CNNs are effective for image data
- Describe how feature maps and filters work
Convolutional neural networks
A convolutional neural network, or CNN, is designed for grid-like data such as images.
Instead of treating every pixel independently, a CNN scans small filters across the input. This lets the model detect local patterns such as edges, textures, and shapes.
Why this matters - Spatial relationships: Nearby pixels in an image are often related - Translation invariance: The same pattern (e.g., an edge or object) can appear anywhere in the image - Efficiency: Reusing small filters avoids relearning the same features across the entire image
Because of these properties, CNNs became the foundation of modern computer vision systems.
That is why CNNs became such an important model family for vision.
How a CNN works
The core building block of a CNN is the convolutional layer.
- A small matrix called a filter (or kernel) slides across the input
- At each position, it performs element-wise multiplication with the input values
- The results are summed to produce a single value
- Repeating this process creates a feature map
A feature map highlights where certain patterns appear in the input.

What the network learns
- Early layers detect simple features (edges, lines, textures)
- Deeper layers combine these into complex patterns (shapes, objects)

At the start of training, filters are random. Through backpropagation and optimization, the network learns which patterns are useful and adjusts the filter values accordingly.
At the start of training, the kernel values are not meaningful. The model learns them from data by adjusting them during training.
2D convolutions in vision
For image tasks, CNNs typically use 2D convolutions.
- The kernel is a small window (e.g., 3×3 or 5×5)
- It moves across the height and width of the image
- Each pass produces a feature map
This allows the network to:
- Detect local features regardless of their position
- Build hierarchical representations (from edges → shapes → objects)


CNNs learn reusable local pattern detectors and stack them to form increasingly abstract representations.

1D convolutions for sequences and signals

Convolutions are not limited to images. A 1D convolution operates over sequences instead of 2D grids.
Instead of sliding a 2D window, it slides along a single dimension.
Common use cases - Time series data (e.g., stock prices, sensor readings) - Audio and signal processing - Spectral data - Biological sequences (e.g., DNA)
The other layers
CNNs don’t consist only of convolutional layers. They also include other types of layers that serve specific roles in helping the model learn and make predictions.
Pooling Layers (Downsampling)
Pooling layers reduce the spatial size of feature maps, which makes computation more efficient and helps the network focus on the most important features.
They also make the model more robust to small spatial changes, meaning it becomes less sensitive to the exact position of an object in an image.

Fully Connected Layers (Dense Layers)
Fully connected (dense) layers are the standard type of neural network layer where every neuron in one layer is connected to every neuron in the next.
These layers typically appear at the end of a CNN and are responsible for:
- Combining the extracted features
- Making the final prediction (e.g., classification or regression)
They act as the decision-making part of the network, taking the high-level features learned by the convolutional layers and turning them into an output.
Here is a polished and professionally formatted version of your summary. I’ve cleaned up the grammar, tightened the phrasing, and organized the comparison into a clear table to make the “division of labor” between these layers really pop. Putting It All Together
By combining the layers just mentioned with the convolutional foundations discussed earlier, you create a powerful two-phase architecture. In this setup, the Convolutional layers act as the “eyes” of the model (feature extraction), while the Dense layers act as the “brain” (final decision-making).
This standard architecture utilizes two distinct types of connectivity, which is why we conceptually separate the “convolutional” stage from the “standard neural network” stage:
| Feature | Phase 1: CNN (Feature Extraction) | Phase 2: Standard NN (Classification Head) |
|---|---|---|
| Layer Type | Convolution, Pooling | Fully Connected (Dense) |
| Connectivity | Local Connectivity: Each neuron only “sees” a small window (the filter size). | Fully Connected: Every neuron connects to every single neuron in the next layer. |
| Efficiency | High Efficiency: Uses “Weight Sharing” where one filter scans the entire image. | Low Efficiency: Parameters explode as the input grows (often millions of weights). |
| Data Structure | Operates on 2D grids / 3D volumes to preserve spatial relationships. | Operates on 1D feature vectors, effectively ignoring spatial location. |
| Function Learns what and where visual features (edges, patterns, objects) exist. | Interprets global features to make a high-level decision (e.g., “This is a cat”). |
Summary

Lastly, there are additional layers required for the construction in
a neural network 
Available demo notebooks
Two demo notebooks are available for this lesson.
CNN_from_scratch_example.ipynb: Trains you to create a CNN model from scratch.
CNN_for_time_Series.ipynb: Example of using LSTM for time series future value prediction
- CNNs are designed for grid-like data (e.g., images)
- They use filters (kernels) to scan across inputs and detect patterns
- Output of convolutions = feature maps showing where patterns occur