Data Science for Electron Microscopy
Week 7: CNNs & U-Nets for microscopy

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: Week 6 and today’s question

  • Week 6: neural networks from first principles — perceptron, MLP, activations, softmax/cross-entropy, autograd and backprop, initialisation, batch/layer norm, training diagnostics.
  • Core insight: a network is empirical-risk minimisation by gradient descent on a composition of learnable maps; backprop gives \(\nabla_\theta \hat{R}\) for free.
  • Gap: an MLP flattens every image into a vector. A 1024×1024 HAADF image into one dense layer with 1000 neurons already needs one billion weights — and the top-left and bottom-right pixels are unrelated inputs.
  • Today’s question: how do we build spatial structure into the architecture — locality, “same feature anywhere”, hierarchy — and how do we check what the trained network actually looks at?
  • Answer: convolutional networks, U-Nets for per-pixel segmentation, and saliency-type explanations.

Road map, learning outcomes and self-study

Road map (≈90 min)

  1. Why MLPs fail on images
  2. Convolution: kernels, stride/padding, channels
  3. Weight sharing, equivariance vs invariance, symmetry
  4. Pooling, receptive field, hierarchy
  5. Architectures: LeNet → AlexNet → VGG → ResNet
  6. U-Net segmentation: IoU/Dice, class imbalance, validation, EM cases
  7. Explainability: saliency, Grad-CAM, occlusion, shortcuts
  8. Failure modes, checklist, notebook, Week 8 preview

After today you can …

  • compute output shapes, parameter counts and receptive fields of a CNN;
  • distinguish translation equivariance from invariance and say which a task needs;
  • explain why skip connections make deep nets (ResNet) and sharp masks (U-Net) possible;
  • evaluate a segmentation with IoU/Dice and handle class imbalance;
  • produce and critically read saliency, Grad-CAM and occlusion maps.

The parameter explosion: a concrete count

MLP parameter count vs a single convolutional layer for images of increasing size. Note the log scale. A 1024×1024 image with 1000 hidden neurons requires ~10⁹ weights; 64 conv filters of size 3×3 require only 576.

Three failures of MLPs on images

Problem MLP behaviour CNN fix
Parameter explosion \(D \times M\) weights per layer \(C_{out}(C_{in}k^2+1)\) kernel weights
No spatial structure flattened: neighbours become unrelated inputs local receptive field: a \(k \times k\) patch
No translation awareness a shifted precipitate is a new input vector weight sharing: one kernel applied everywhere

CNNs are MLPs with locality and weight sharing built in as hard constraints.

Convolution: a sliding feature detector

Two-dimensional cross-correlation (the operation CNNs actually use): the 3×3 kernel slides over the input; at each position nine products are summed into one output value.
  • Slide a small kernel (e.g. 3×3 weights) across the image.
  • At each position: multiply with the patch, sum → one output value.
  • Output = feature map: how strongly the kernel’s pattern appears at each location.
  • \(C_{out}\) kernels in parallel → \(C_{out}\) feature maps.

Convolution on a synthetic grain image

Convolution applied to a synthetic two-grain microstructure. Left: input image (two grains with different intensities, separated by vertical and horizontal boundaries). Centre-left: vertical Sobel kernel responds strongly at the vertical grain boundary. Centre-right: horizontal Sobel responds at the horizontal boundary. Right: Gaussian blur smooths noise. All kernels are 3×3; weights are hand-set, not trained.

The discrete convolution formula

For input image \(I\) and kernel \(K\):

\[ (I \star K)_{m,n} = \sum_{a=-\Delta}^{\Delta}\sum_{b=-\Delta}^{\Delta} K_{a,b}\,I_{m+a,\, n+b}, \qquad \Delta = (k-1)/2 \]

  • Only pixels within \(\pm\Delta\) of \((m,n)\) contribute → local.
  • \(k^2\) weights, independent of image size \(H \times W\), reused at every \((m,n)\) → weight sharing.

Kernels as feature detectors

Three classic 3×3 kernels. Left: vertical Sobel — responds to left–right intensity changes. Center: horizontal Sobel — responds to top–bottom changes. Right: Laplacian — responds to any local intensity peak or boundary. Numbers are the kernel weights.

In a trained CNN, the network discovers these (and more complex) filters automatically from labelled examples — no manual kernel design needed.

Stride, padding, channels: output size and parameter count

Stride 1 with same-padding (left): output has the input size. Stride 2 (right): the kernel skips every other position, halving spatial resolution.

  • \(H_{out} = \left\lfloor\frac{H + 2p - k}{s}\right\rfloor + 1\): \(k=3, p=1, s=1\) keeps \(H\) (“same” padding); \(s=2\) halves it.
  • Kernel tensor \(C_{out}\times C_{in}\times k\times k\) → \(C_{out}(C_{in}k^2+1)\) parameters — independent of \(H, W\): the same network runs on 256² and 4096² images.
  • EM multichannel inputs: EELS/EDS spectrum images, multi-segment detector stacks.

Weight sharing: the key efficiency principle

Dense layer (left): every input node connects to every output node — \(n_{in} \times n_{out}\) unique weights. Conv layer (right): the same 3-weight kernel connects each output node to only a local neighborhood of inputs — 3 shared weights total (for the 1-D case shown).

Translation equivariance as an inductive bias

  • Translation equivariance: if the input shifts by \(\delta\) pixels, the feature map shifts by \(\delta\) pixels — the detector response follows the feature.
  • Formally: \(f(T_\delta X) = T_\delta f(X)\) for a spatial shift \(T_\delta\).
  • Translation invariance (different!): \(f(T_\delta X) = f(X)\) — the prediction ignores the shift. Built by pooling, striding, global average pooling.

Equivariance preserves “where.” Invariance discards “where” and keeps only “what.”

Beyond translation: rotations and mirrors in EM

Numerical check with a vertical Sobel kernel \(f\). Top: shifting the input and then convolving equals convolving and then shifting (difference exactly zero). Bottom: for a 90° rotation the two orders differ — the kernel only “knows” vertical edges. Generated by img/make_figures.py.

Getting symmetry: three routes

  • For a group \(G\) of transforms: equivariant \(f(g\cdot x) = g\cdot f(x)\) (masks, denoised images); invariant \(f(g\cdot x) = f(x)\) (phase label).
  • Route 1 — augmentation: train on rotated/flipped copies; approximate symmetry is learned (Week 8).
  • Route 2 — test-time augmentation: average predictions over the 8 rotations/flips of \(D_4\) (un-rotate masks first).
  • Route 3 — build it in: group-equivariant CNNs share weights across rotations too Cohen, Taco et al., (2016).
  • Rule: know the symmetry of your task first — then learn it, average over it, or hard-wire it.

Inductive bias: encoding what you know

  • An inductive bias is an assumption baked into the architecture that makes certain functions easy to learn.
  • MLP (Week 6): no structure — every function of all inputs equally easy; needs the most data.
  • CNN (today): grid + locality + translation weight sharing.
  • GNN / transformer (Week 10): graph + permutation symmetry; a transformer is a fully connected graph with learned (attention) edges.
  • Strong, correct bias helps with small datasets by ruling out physically unreasonable solutions; a wrong bias caps performance.

Pooling and the Conv–Pool block

2×2 max-pooling: keep the maximum of each non-overlapping 2×2 region; resolution halves, channel count unchanged.
  • Max pool: “did the feature appear here?” Average pool: “how strongly on average?” No learned parameters.
  • 2×2 pool halves \(H\) and \(W\) → approximate invariance to small shifts, faster receptive-field growth.
  • Standard block: Conv → ReLU → MaxPool, channels doubling per stage: 256² input → 16² × 512 after 4 blocks.
  • Global average pooling at the end → translation-invariant image-level prediction.

Max pooling: a worked example

Input feature map (4×4):

\[ \begin{pmatrix} 1 & 3 & 2 & 4 \\ 5 & 6 & 7 & 8 \\ 3 & 2 & 1 & 0 \\ 1 & 2 & 3 & 4 \end{pmatrix} \]

After 2×2 max-pool (stride 2):

\[ \begin{pmatrix} 6 & 8 \\ 3 & 4 \end{pmatrix} \]

Top-left 2×2 block: \(\max(1,3,5,6)=6\). Top-right: \(\max(2,4,7,8)=8\).

  • Output size: \(2\times2\) from \(4\times4\) — spatial dimensions halved.
  • No learned parameters — purely a max operation.
  • Approximate invariance: if the 6 moved to the adjacent cell (still in the same 2×2 window), the output is unchanged.

Receptive fields: how much context does one neuron see?

Receptive field of a single output neuron grows with depth. One 3×3 conv layer: 3×3 input region. Two stacked layers: 5×5 region. Three layers: 7×7 region. Red star marks the output neuron; blue region is its receptive field in the input image.

Feature hierarchy: edges → motifs → structures → properties

Feature hierarchy on a synthetic grain microstructure. Input (left): two-grain image with boundaries. Layer 1 (centre-left): Laplacian-like edge features highlight all boundaries. Layer 2 (centre-right): neighbourhood-level grain-boundary motifs. Layer 3+ (right): coarse phase/grain-region labels.

CNN architectures in one breath: the arc from 1998 to 2015

Timeline from LeNet (1998) to AlexNet (2012) to ResNet (2015); VGG (2014) sits between AlexNet and ResNet. Each box states the year, approximate parameter count, depth, and key innovation. Read left to right as increasing depth, scale, and capability.

LeNet and AlexNet: the template and the revolution

LeNet-5 LeCun, Yann et al., (1998)

  • Conv → Act → Pool → Conv → Act → Pool → Flatten → FC → FC → Output
  • ~60,000 parameters; learned features beat hand-crafted ones (SIFT, HOG)
  • Still the template for small micrograph classifiers

AlexNet Krizhevsky, Alex et al., (2012)

  • Won ImageNet 2012 by ~10 percentage points
  • ReLU, dropout, GPU training, ~60 M parameters
  • Opened the deep-learning era

VGG: deep networks from repeated blocks Simonyan, Karen et al., (2015)

From AlexNet (left) to VGG (right): irregular layers are replaced by repeated blocks of 3×3 convolutions followed by 2×2 max-pooling. Figure: d2l.ai (CC BY-SA 4.0).
  • VGG block: \(n\times\) (3×3 conv, pad 1, ReLU) → 2×2 max-pool; channels 64 → 128 → 256 → 512.
  • Why 3×3? Each stacked 3×3 adds 2 pixels of receptive field: three give 7×7 with \(27C^2\) instead of \(49C^2\) weights and two extra non-linearities.
  • Design lesson: think in blocks, not layers.

ResNet: skip connections make depth trainable He, Kaiming et al., (2016)

ResNet residual block: the main path learns F(x); the skip connection adds x to the output.
  • Degradation: plain stacks of 20 → 56 layers reach worse training error — not overfitting, an optimisation problem.
  • Residual block: \(\mathbf{y} = F(\mathbf{x}) + \mathbf{x}\) — learn only a correction to the identity.
  • Gradient \(\partial\mathbf{y}/\partial\mathbf{x} = \mathbf{I} + \partial F/\partial\mathbf{x}\): an unattenuated path through every block.
  • Block: conv–BN–ReLU–conv–BN, add, ReLU.
  • Skips today: additive (ResNet: optimisation) vs concatenating (U-Net: localisation).

U-Net: encoder–decoder for segmentation Ronneberger, Olaf et al., (2015)

U-Net architecture. Left column (blue): encoder — successive Conv+Pool blocks extract features while halving spatial resolution and doubling channel count. Right column (red): decoder — successive upsample+Conv blocks restore spatial resolution. Yellow dashed arrows: skip connections concatenate encoder features into corresponding decoder levels.

U-Net: encoder, decoder, per-pixel output

Encoder (context)

  • (Conv×2 → ReLU → MaxPool) × 4
  • spatial halves, channels double

Decoder (resolution)

  • (Upsample → concatenate skip → Conv×2) × 4
  • final 1×1 conv → \(K\times H\times W\) class scores
  • Mask = \(\arg\max\) over \(K\) channels; loss = per-pixel cross-entropy (often + Dice), same backprop + Adam.
  • Without skips, boundary positions are lost in the bottleneck → blurry masks.

Segmentation metrics: IoU and Dice, not accuracy

Left: IoU and Dice compare predicted and true foreground sets and ignore true-negative background. Right: on our synthetic particle images (9% foreground) predicting nothing already scores 0.91 pixel accuracy; a smoothed Otsu threshold also scores 0.91 but IoU 0.48; the tiny U-Net from the notebook reaches IoU 0.84. Generated by img/make_figures.py.

Class imbalance: losses that keep the minority visible

  • Pixel-wise BCE is an average over pixels: with 9% particles, the minority class carries 9% of the loss weight.
  • Weighted BCE (pos_weight \(=(1-f)/f\)); focal loss Lin, Tsung-Yi et al., (2017) down-weights easy pixels.
  • Soft Dice loss Milletari, Fausto et al., (2016), with \(p_i=\sigma(z_i)\): \[\mathcal{L}_\text{Dice} = 1-\frac{2\sum_i p_i y_i+\epsilon}{\sum_i p_i+\sum_i y_i+\epsilon}\] normalised by the foreground size → “predict nothing” costs ≈ 1.
  • BCE + Dice is the default. Notebook: BCE alone collapsed to “all background” (IoU 0.00); BCE + Dice reached IoU 0.84.

A tiny U-Net on synthetic particles (the notebook)

Tiny U-Net (2 levels, 8–16–32 channels, BatchNorm, 29.6 k parameters) trained for 15 epochs on 256 synthetic 48×48 images of bright particles on a textured, unevenly lit support with shot noise. Rows: median, good, and worst test image. The worst case misses a low-contrast particle entirely. Generated by img/make_figures.py (same code and seeds as the notebook).

Honest validation for segmentation

  • Tile leakage: cutting one large micrograph into overlapping 256² tiles and splitting the tiles randomly puts neighbouring pixels in train and test → inflated IoU. Split by micrograph, specimen or session (GroupKFold, Week 4).
  • Label ceiling: measure IoU between two annotators on a subset; a model cannot meaningfully beat that ceiling.
  • Report per class and per image: mean IoU hides that one rare phase (or the faint particles) is never found.
  • Downstream quantity: check the number you actually need — particle size distribution, phase fraction, grain count — not just IoU.

EM case study: defects in irradiated steel Roberts, Graham et al., (2019), doi:10.1038/s41598-019-49105-0

DefectSegNet predictions on STEM (DCI) images of irradiated steel: input, ground truth, prediction and comparison for precipitates and voids. From Roberts, Graham et al. (2019), doi:10.1038/s41598-019-49105-0 (Sci. Rep., CC BY 4.0).

  • U-Net-type encoder–decoder with dense skip connections; per-pixel labels for dislocation loops, precipitates and voids.
  • Reaches roughly inter-annotator agreement — beyond that point more model capacity buys nothing; better labels do.

EM case study: grain boundaries from synthetic Voronoi labels

Training: synthetic Voronoi grain image, free ground-truth labels, U-Net predicted boundaries.

Application: the network trained only on synthetic Voronoi images segments grain boundaries in a real SEM polycrystal image.
  • Voronoi tessellations give perfect labels for free; the boundary topology (“dark strip between two regions”) transfers to real SEM images.

Why explain a CNN in EM?

  • A segmenter or classifier with 0.99 test accuracy can still be right for the wrong reason — “shortcut learning” Geirhos, Robert et al., (2020).
  • EM shortcuts: scan-start/detector-edge artefacts, session contrast, scale bars, beam damage tied to one specimen.
  • Local explanation question: which input pixels drove this particular prediction?
  • Their job is diagnosis (find shortcuts, understand failures, decide what data to add) — not proof that a model is physically correct.

Three local explanations: saliency, Grad-CAM, occlusion

  • Gradient saliency Simonyan, Karen et al., (2014): \(S_i = |\partial z_c/\partial x_i|\) — one backward pass w.r.t. the input; pixel-level, noisy; sensitivity, not contribution.
  • Grad-CAM Selvaraju, Ramprasaath R. et al., (2017): weight the last-conv feature maps by their mean gradient, \(L^c = \mathrm{ReLU}\big(\sum_k \alpha_k^c A^k\big)\), \(\alpha_k^c = \frac{1}{Z}\sum_{i,j}\partial z_c/\partial A^k_{ij}\) — coarse, smooth regions.
  • Occlusion Zeiler, Matthew D. et al., (2014): slide a grey patch, record the drop \(p_c(\mathbf{x}) - p_c(\mathbf{x}_\text{occ})\) — model-agnostic, forward passes only.
  • For segmentation: backpropagate from a set of output pixels (e.g. all false positives).

Shortcut learning: same accuracy, different physics

Defect classifiers (class 1 = faint bright disk). The shortcut model was trained on images in which every defect image also had a dark 6×6 corner (a scan-start/detector-shadow artefact); the clean model on the same images without it. Test accuracy with the artefact: 0.94 vs 1.00; deployment accuracy (defects, no artefact): 0.94 vs 0.50. All three methods put the shortcut model’s evidence in the corner. Generated by img/make_figures.py.

Explaining a segmentation failure

Tiny U-Net applied to an out-of-distribution image with a bright scratch and an oversized particle. Scratch segments become false positives (red); the saliency of the summed false-positive logits sits on the scratch itself. Generated by img/make_figures.py.

Sanity checks and honest limits of saliency

  • Model-randomisation test Adebayo, Julius et al., (2018): randomise the weights and recompute. If the map barely changes, it shows image structure, not what the model learned (notebook: correlation ≈ 0 — passes).
  • Fragility Kindermans, Pieter-Jan et al., (2019): simple input transformations (e.g. a constant shift) can change some attribution maps without changing the prediction.
  • Sensitivity ≠ causation: a coherent map can be coherent about the wrong feature (the corner).
  • Use them as hypotheses → test them: occlude the suspected region, build a shifted test set, retrain without the cue.

Failure modes and a practical checklist

Failure modes

  • Domain shift: other microscope, detector, contrast transfer
  • Beam-damage artefacts that look like boundaries
  • Scale: trained at 0.1 nm/px, applied at 0.5 nm/px
  • Out-of-distribution microstructures

Checklist

  • Match pixel size between training and inference.
  • GroupKFold by specimen/session — not random splits.
  • Look at predicted masks; report per-class IoU/Dice.
  • Start simple; beat a thresholding baseline first.
  • Explain before deploying (+ randomisation check).

Self-study notebook

notebooks/week07_cnn_segmentation.ipynb

  1. Hand-set kernels (Sobel, Laplacian, Gaussian) as one conv layer
  2. Equivariance check: shift ✓, rotation ✗
  3. Synthetic particle images with free masks, split by “session”
  4. IoU/Dice and the accuracy trap; Otsu baseline (IoU 0.48)
  5. Tiny U-Net, BCE + soft Dice, CPU: IoU 0.84, Dice 0.91
  6. Saliency on an OOD failure (scratch)
  7. Shortcut demo: saliency, Grad-CAM, occlusion; randomisation check

Open in Colab
  • Runs end-to-end on a CPU in a few minutes.
  • Exercises: soft-Dice loss, ablations (no skips / no Dice), saliency sanity check.
  • Real data: MetalDAM (SEM, AM metal, per-pixel labels) via Ai4MatLectures.

Summary: the week in six points

  1. Convolution = locality + weight sharing: parameters independent of image size, translation-equivariant feature maps.
  2. Equivariance vs invariance: layers are shift-equivariant; invariance comes from pooling/GAP; rotations need augmentation, TTA or group-equivariant layers.
  3. Depth needs design: VGG blocks of 3×3 convs; ResNet identity skips + BN make very deep nets trainable.
  4. U-Net = encoder (context) + decoder (resolution) + concatenating skips (where) → one label per pixel.
  5. Evaluate with IoU/Dice, fight class imbalance in the loss (BCE + Dice, weights), validate by specimen/session.
  6. Explain before trusting: saliency, Grad-CAM, occlusion diagnose shortcuts and failures; sanity-check the explanation itself.

Must-know for the exam

  • Output size \(\lfloor (H+2p-k)/s\rfloor+1\) and parameter count \(C_{out}(C_{in}k^2+1)\) — independent of \(H,W\).
  • \(f(T x)=T f(x)\) (equivariance) vs \(f(Tx)=f(x)\) (invariance); CNNs are not rotation-equivariant by default.
  • Stacked 3×3 convs: receptive field grows by 2 per layer; pooling/stride multiplies the growth.
  • ResNet: \(\mathbf{y}=F(\mathbf{x})+\mathbf{x}\) eases optimisation; U-Net skips concatenate encoder features to restore localisation.
  • IoU, Dice (\(=2\mathrm{IoU}/(1+\mathrm{IoU})\)); pixel accuracy is misleading under imbalance; soft Dice / weighted BCE as remedies.
  • Saliency \(|\partial z_c/\partial \mathbf{x}|\), Grad-CAM (last conv layer), occlusion (model-agnostic); a plausible map can reveal a shortcut, and must pass a randomisation sanity check.

Next week: small data — augmentation, transfer & self-supervision

  • Next week’s question: you have 30 labelled TEM images. How do you train a reliable CNN?
  • Physics-respecting augmentation: rotations/flips from the crystal symmetry, Poisson noise at the right dose, drift, scratches — the fix for today’s scratch failure.
  • Transfer learning: start from a pretrained encoder (e.g. a ResNet) — linear probe vs fine-tune.
  • Synthetic data & sim-to-real: Voronoi grains, simulated HRTEM — free labels like today’s notebook.
  • Self-supervised learning & foundation models: SimCLR, masked autoencoders, DINO; SAM / μSAM for microscopy.

Continue

Backup slides

Material not covered in the 90-minute lecture path — for questions and self-study.

Residual blocks in practice: BN, stages, and why EM cares

  • Basic block: conv3×3 → BN → ReLU → conv3×3 → BN, add input, ReLU. If shapes change (stride 2, more channels) the skip uses a 1×1 conv.
  • Gradient view: \(\frac{\partial \mathbf{y}}{\partial \mathbf{x}} = \mathbf{I} + \frac{\partial F}{\partial \mathbf{x}}\) — the identity term carries the gradient through every block unattenuated.
  • Stages: ResNet-18/34/50 = 4 stages of residual blocks, stride-2 between stages, global average pooling + one linear layer (no big dense head: ResNet-18 ≈ 11 M params vs VGG-16 ≈ 138 M).
  • EM use: ResNet backbones are the standard encoder for U-Net-style segmenters and the usual starting point for transfer learning (Week 8).
  • Skip connections appear twice today: additive (ResNet: ease optimisation) and concatenating (U-Net: restore spatial detail).

EM case study: Au nanoparticle phase segmentation

U-Net applied to TEM images of Au nanoparticles on an amorphous support. Left: input TEM image (representative Au-nanoparticle-on-amorphous-support TEM segmentation task). Centre: ground-truth binary mask (crystalline=bright, amorphous=dark). Right: U-Net prediction — pixel-wise classification matching the ground truth closely.

References

Group equivariant convolutional networks, Proceedings of the 33rd international conference on machine learning (ICML), Taco Cohen & Max Welling.
Gradient-based learning applied to document recognition, Proceedings of the IEEE, Yann LeCun, Leon Bottou, Yoshua Bengio, & Patrick Haffner.
ImageNet classification with deep convolutional neural networks, Advances in neural information processing systems, Alex Krizhevsky, Ilya Sutskever, & Geoffrey E. Hinton.
Very deep convolutional networks for large-scale image recognition, International conference on learning representations (ICLR), Karen Simonyan & Andrew Zisserman.
Deep residual learning for image recognition, Proceedings of the IEEE conference on computer vision and pattern recognition, Kaiming He, Xiangyu Zhang, Shaoqing Ren, & Jian Sun.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, Proceedings of the 32nd international conference on machine learning (ICML), Sergey Ioffe & Christian Szegedy.
U-net: Convolutional networks for biomedical image segmentation, Medical image computing and computer-assisted intervention (MICCAI), Olaf Ronneberger, Philipp Fischer, & Thomas Brox.
Focal loss for dense object detection, Proceedings of the IEEE international conference on computer vision (ICCV), Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, & Piotr Dollár.
V-Net: Fully convolutional neural networks for volumetric medical image segmentation, Fourth international conference on 3D vision (3DV), Fausto Milletari, Nassir Navab, & Seyed-Ahmad Ahmadi.
Deep learning for semantic segmentation of defects in advanced STEM images of steels, Scientific Reports, Graham Roberts, Simon Y. Haile, Rajat Sainju, Danny J. Edwards, Brian Hutchinson, & Yuanyuan Zhu https://doi.org/10.1038/s41598-019-49105-0.
Shortcut learning in deep neural networks, Nature Machine Intelligence, Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, & Felix A. Wichmann.
Deep inside convolutional networks: Visualising image classification models and saliency maps, International conference on learning representations (ICLR), workshop track, Karen Simonyan, Andrea Vedaldi, & Andrew Zisserman.
Grad-CAM: Visual explanations from deep networks via gradient-based localization, Proceedings of the IEEE international conference on computer vision (ICCV), Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, & Dhruv Batra.
Visualizing and understanding convolutional networks, European conference on computer vision (ECCV), Matthew D. Zeiler & Rob Fergus.
Sanity checks for saliency maps, Advances in neural information processing systems (NeurIPS), Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, & Been Kim.
The (un)reliability of saliency methods, Explainable AI: Interpreting, explaining and visualizing deep learning, Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, & Been Kim.