FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
Road map (≈90 min)
After today you can …
MLP parameter count vs a single convolutional layer for images of increasing size. Note the log scale. A 1024×1024 image with 1000 hidden neurons requires ~10⁹ weights; 64 conv filters of size 3×3 require only 576.
| Problem | MLP behaviour | CNN fix |
|---|---|---|
| Parameter explosion | \(D \times M\) weights per layer | \(C_{out}(C_{in}k^2+1)\) kernel weights |
| No spatial structure | flattened: neighbours become unrelated inputs | local receptive field: a \(k \times k\) patch |
| No translation awareness | a shifted precipitate is a new input vector | weight sharing: one kernel applied everywhere |
CNNs are MLPs with locality and weight sharing built in as hard constraints.
Convolution applied to a synthetic two-grain microstructure. Left: input image (two grains with different intensities, separated by vertical and horizontal boundaries). Centre-left: vertical Sobel kernel responds strongly at the vertical grain boundary. Centre-right: horizontal Sobel responds at the horizontal boundary. Right: Gaussian blur smooths noise. All kernels are 3×3; weights are hand-set, not trained.
For input image \(I\) and kernel \(K\):
\[ (I \star K)_{m,n} = \sum_{a=-\Delta}^{\Delta}\sum_{b=-\Delta}^{\Delta} K_{a,b}\,I_{m+a,\, n+b}, \qquad \Delta = (k-1)/2 \]
Three classic 3×3 kernels. Left: vertical Sobel — responds to left–right intensity changes. Center: horizontal Sobel — responds to top–bottom changes. Right: Laplacian — responds to any local intensity peak or boundary. Numbers are the kernel weights.
In a trained CNN, the network discovers these (and more complex) filters automatically from labelled examples — no manual kernel design needed.
Stride 1 with same-padding (left): output has the input size. Stride 2 (right): the kernel skips every other position, halving spatial resolution.
Dense layer (left): every input node connects to every output node — \(n_{in} \times n_{out}\) unique weights. Conv layer (right): the same 3-weight kernel connects each output node to only a local neighborhood of inputs — 3 shared weights total (for the 1-D case shown).
Equivariance preserves “where.” Invariance discards “where” and keeps only “what.”
Numerical check with a vertical Sobel kernel \(f\). Top: shifting the input and then convolving equals convolving and then shifting (difference exactly zero). Bottom: for a 90° rotation the two orders differ — the kernel only “knows” vertical edges. Generated by img/make_figures.py.
Input feature map (4×4):
\[ \begin{pmatrix} 1 & 3 & 2 & 4 \\ 5 & 6 & 7 & 8 \\ 3 & 2 & 1 & 0 \\ 1 & 2 & 3 & 4 \end{pmatrix} \]
After 2×2 max-pool (stride 2):
\[ \begin{pmatrix} 6 & 8 \\ 3 & 4 \end{pmatrix} \]
Top-left 2×2 block: \(\max(1,3,5,6)=6\). Top-right: \(\max(2,4,7,8)=8\).
Receptive field of a single output neuron grows with depth. One 3×3 conv layer: 3×3 input region. Two stacked layers: 5×5 region. Three layers: 7×7 region. Red star marks the output neuron; blue region is its receptive field in the input image.
Feature hierarchy on a synthetic grain microstructure. Input (left): two-grain image with boundaries. Layer 1 (centre-left): Laplacian-like edge features highlight all boundaries. Layer 2 (centre-right): neighbourhood-level grain-boundary motifs. Layer 3+ (right): coarse phase/grain-region labels.
Timeline from LeNet (1998) to AlexNet (2012) to ResNet (2015); VGG (2014) sits between AlexNet and ResNet. Each box states the year, approximate parameter count, depth, and key innovation. Read left to right as increasing depth, scale, and capability.
LeNet-5 LeCun, Yann et al., (1998)
AlexNet Krizhevsky, Alex et al., (2012)

U-Net architecture. Left column (blue): encoder — successive Conv+Pool blocks extract features while halving spatial resolution and doubling channel count. Right column (red): decoder — successive upsample+Conv blocks restore spatial resolution. Yellow dashed arrows: skip connections concatenate encoder features into corresponding decoder levels.
Encoder (context)
Decoder (resolution)
Left: IoU and Dice compare predicted and true foreground sets and ignore true-negative background. Right: on our synthetic particle images (9% foreground) predicting nothing already scores 0.91 pixel accuracy; a smoothed Otsu threshold also scores 0.91 but IoU 0.48; the tiny U-Net from the notebook reaches IoU 0.84. Generated by img/make_figures.py.
pos_weight \(=(1-f)/f\)); focal loss Lin, Tsung-Yi et al., (2017) down-weights easy pixels.Tiny U-Net (2 levels, 8–16–32 channels, BatchNorm, 29.6 k parameters) trained for 15 epochs on 256 synthetic 48×48 images of bright particles on a textured, unevenly lit support with shot noise. Rows: median, good, and worst test image. The worst case misses a low-contrast particle entirely. Generated by img/make_figures.py (same code and seeds as the notebook).
GroupKFold, Week 4).DefectSegNet predictions on STEM (DCI) images of irradiated steel: input, ground truth, prediction and comparison for precipitates and voids. From Roberts, Graham et al. (2019), doi:10.1038/s41598-019-49105-0 (Sci. Rep., CC BY 4.0).


Defect classifiers (class 1 = faint bright disk). The shortcut model was trained on images in which every defect image also had a dark 6×6 corner (a scan-start/detector-shadow artefact); the clean model on the same images without it. Test accuracy with the artefact: 0.94 vs 1.00; deployment accuracy (defects, no artefact): 0.94 vs 0.50. All three methods put the shortcut model’s evidence in the corner. Generated by img/make_figures.py.
Tiny U-Net applied to an out-of-distribution image with a bright scratch and an oversized particle. Scratch segments become false positives (red); the saliency of the summed false-positive logits sits on the scratch itself. Generated by img/make_figures.py.
Failure modes
Checklist
GroupKFold by specimen/session — not random splits.notebooks/week07_cnn_segmentation.ipynb
Material not covered in the 90-minute lecture path — for questions and self-study.
U-Net applied to TEM images of Au nanoparticles on an amorphous support. Left: input TEM image (representative Au-nanoparticle-on-amorphous-support TEM segmentation task). Centre: ground-truth binary mask (crystalline=bright, amorphous=dark). Right: U-Net prediction — pixel-wise classification matching the ground truth closely.

©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy