FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
Road map (≈90 min)
By the end of today you can …
Labelled images per domain: ImageNet 14 M (crowdsourced, seconds each), medical imaging ~10⁴ (radiologists), materials EM 50–500 (hours of expert time each) Holm, Elizabeth A. et al., (2020); Sandfeld, Stefan et al., (2024).
Six views of the SAME synthetic grain layout: original, 90° rotation, horizontal flip, brightness jitter, Poisson noise (all valid for equiaxed grains) and vertical flip (invalid if a surface gradient exists).
Every augmentation is a claim that the physics has that symmetry.
| Transform | Physical claim / simulates | Illegal when … |
|---|---|---|
| Flip | mirror symmetry | feature has polarity (weld cap vs root, DS column) |
| Rotation | rotational symmetry | orientation is the signal (columnar grains, EBSD) |
| Crop / zoom | translation / scale invariance | grain size or lattice spacing is the label |
| Elastic warp | drift, sample warping | metric labels (size, aspect ratio) |
| Brightness / gamma | session / detector variation | intensity is quantitative (EELS, Z-contrast) |
| Poisson noise, blur | low dose, defocus, drift | rarely — the physically correct noise model |
| Scratches, contamination | preparation artefacts | artefact is the target |
Illegal augmentations: rotating an EBSD/IPF map (colour encodes orientation), vertically flipping directional solidification (the thermal gradient is real), intensity jitter on an EELS map (intensity encodes composition). An equiaxed polycrystal allows all of them.
Transferability vs CNN depth Yosinski, Jason et al., (2014): edges/gradients (layer 1) transfer almost fully, textures mostly, object parts and task-specific layers barely.
| Few labels (<100) | More labels (100–1 000) | |
|---|---|---|
| Small domain gap (optical ↔︎ photos) | Feature extraction (frozen backbone) | Fine-tuning, differential LRs |
| Large domain gap (photos ↔︎ HAADF) | Feature extraction + BN adaptation | Fine-tuning + gradual unfreezing |
eval() uses ImageNet statistics → mis-normalised micrographs.Fine-tuning with one large LR (red) destroys pretrained features in the first epochs; differential LRs (green) climb steadily.

U-Net trained only on simulated HRTEM, tested on real nanoparticles Rakowski, Aaron et al., (2024):
| Benchmark | Synthetic F1 | Prev. best (real) |
|---|---|---|
| Au, 5 nm | 0.92 | 0.89 |
| Au, 2.2 nm | 0.86 | 0.75 |
| CdSe, 2 nm | 0.75 | 0.59 |
Simulate the right variation, not more structures.

Contrastive: pull two views together, push other images apart. Masked modelling: hide patches, reconstruct them. Self-distillation: a student matches an EMA teacher across crops, no negatives.
Common enemy: representation collapse — every input mapped to the same vector.

For a positive pair \((i,j)\) among \(2N\) views, cosine similarity \(\mathrm{sim}(\mathbf{u},\mathbf{v})=\mathbf{u}^\top\mathbf{v}/(\|\mathbf{u}\|\|\mathbf{v}\|)\), temperature \(\tau\):
\[ \ell_{i,j} = -\log \frac{\exp\!\big(\mathrm{sim}(\mathbf{z}_i,\mathbf{z}_j)/\tau\big)}{\sum_{k\neq i}\exp\!\big(\mathrm{sim}(\mathbf{z}_i,\mathbf{z}_k)/\tau\big)} \]
Rotation + gain and flip + offset + noise give a valid positive pair (green); another micrograph is a negative. A ×2 zoom is illegal when grain size is the target.

| SimCLR | MAE | DINO / DINOv2 | |
|---|---|---|---|
| Signal | twin among negatives | reconstruct hidden patches | match EMA teacher |
| Collapse prevention | negatives | reconstruction target | EMA + centring + sharpening |
| Augmentation design | critical | minimal | important (multi-crop) |
| Frozen features | good | weaker | best |
| After fine-tuning | good | very good | very good |

μSAM: napari plugin for interactive and automatic segmentation, correction and fine-tuning; default vs microscopy-fine-tuned results for light (top) and electron (bottom) micrographs. Archit et al., Nature Methods Archit, Anwai et al., (2025), doi:10.1038/s41592-024-02580-4.
Report zero-shot and fine-tuned results side by side.
| Protocol | What is trained | Use when |
|---|---|---|
| Linear probe (LP) | logistic regression on frozen \(\mathbf{h}\) | very few labels (≲ 50); measuring embedding quality |
| Fine-tune (FT) | backbone (small \(\eta\)) + head | more labels, larger domain gap |
| LP → FT | head from LP, then FT with small \(\eta\) | default fine-tuning recipe |
| From scratch | everything, random init | many labels, no relevant backbone |
Test accuracy vs number of labelled “real-instrument” images, 3-class grain-size task: linear probes on random-init, sim-supervised and SimCLR backbones, plus LP→FT. From notebooks/week08_embeddings_transfer.ipynb.

notebooks/week08_embeddings_transfer.ipynb
Material cut from the lecture path for time — for questions and self-study.
Training and validation loss for a CNN trained from scratch on 50 EM images. Training loss falls monotonically; validation loss starts rising around epoch 40 — the model is memorising the training images, not learning to generalise. The gap is the overfitting region.
Label consistency: when a grain-boundary image is rotated 45°, the segmentation mask must be rotated by exactly the same 45°. Top row: original image, original mask, rotated image. Bottom row: correct — rotated mask (joint transform); wrong — un-rotated mask paired with the rotated image, producing misaligned ground truth.
Three-stage transfer learning recipe. Stage 1: all backbone blocks frozen (grey); only the head (red) is trained at lr=1e-3. Stage 2: last backbone block unfrozen (orange) with low lr=1e-5; head continues at 1e-3. Stage 3: gradual unfreezing, depth-graded learning rates — early layers receive the smallest lr, late layers more, head the most.
Three panels showing the sim-to-real challenge. Left: synthetic training image (clean, regular grains, no scan artefacts). Centre: real SEM image with scan distortion, vignette, and contrast drift relative to the synthetic distribution. Right: U-Net prediction — grain topology is correctly identified despite the gap, because topology is the task-relevant invariant.
Complete small-data EM workflow diagram. The labelled EM data (20–200 images) feeds augmentation and transfer learning in parallel; synthetic data feeds domain adaptation; all three converge on a fine-tuned model. The active learning loop (dotted arrow, bottom) queries the fine-tuned model for the most uncertain unlabelled images, sends them to expert annotation, and grows the labelled pool.
GroupKFold(n_splits=5).split(X, y, groups=specimen_ids) — the Week 4 lesson.
©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy