Data Science for Electron Microscopy
Week 2: Learning from EM data: signal formation, noise & loss

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: where we left off

  • Week 1 — Python crash course & the EM data deluge: NumPy arrays, PSPP/CRISP-DM frameworks, EM modality overview, why EM produces terabytes.
  • You can now load an EM image as a NumPy array, slice an ROI, and visualise it.
  • The Poisson noise in last week’s simulated STEM image was not arbitrary — it is the correct physical model.
  • Today we answer three questions: What does it mean for a machine to learn? How is an EM number formed, and why is every pixel noisy? And how does the noise model decide the loss?

Today’s question

  • What is learning? Fitting a model to data — but fitting what exactly? And how do you know it worked?
  • Where does EM noise come from? Electrons arrive one at a time; every pixel is a count; counts fluctuate — fundamentally, unavoidably.
  • Why does it matter for data science? Your noise model determines which loss function is correct. The wrong choice silently degrades every trained model on EM data.
  • Road map: learning taxonomy → model taxonomy → EM data zoo → data-formation chain → sampling & aliasing → noise models → noise model → likelihood → loss → data quality & metadata → uncertainty.

Learning outcomes

By the end of this lecture you can:

  1. Distinguish classical data analysis from machine learning and place a method in the supervised / unsupervised / self-supervised and white / grey / black-box taxonomies.
  2. Trace an EM measurement along the data-formation chain (specimen → interaction → detector → digitisation → metadata) and name where each error enters.
  3. Apply Nyquist–Shannon to choose a pixel size and recognise aliasing / Moiré artefacts.
  4. Write down the Poisson, Gaussian and mixed Poisson–Gaussian noise models and read gain and read noise off a variance–mean plot.
  5. Derive the loss from the noise model: Gaussian → MSE, Poisson → Poisson NLL, mixed → variance-stabilising transform or mixed NLL.
  6. Run a data-quality and metadata (FAIR) check before any model training.

What is learning? Data analysis vs machine learning

  • Classical data analysis: you write down the rules (equations, thresholds) and the computer applies them.
  • Machine learning: you supply examples (data) and the computer infers the rules.
  • In both cases the goal is the same: extract information from measurements.
  • The difference is who specifies the function — the human or the optimiser.

Models: all are wrong, some are useful

  • A model is a purposeful abstraction of reality built for prediction or explanation Neuer, Michael et al., (2024).
  • First-principles models: derived from physical laws; interpretable; may fail when assumptions break down.
  • Data-based models: inferred from observations; flexible; performance depends on data quality and coverage.
  • George Box’s maxim: “All models are wrong, but some are useful.” — usefulness is judged at the decision point, not by aesthetics.

When first principles are not enough

  • Complex systems can be nonlinear, high-dimensional, and partially observed.
  • Example: Electron beam scattering in a thick sample — exact multi-slice simulation exists but is too slow for real-time analysis of millions of diffraction patterns.
  • Hybrid strategy: keep trusted physics as structure; learn the residual or unknown coupling from data.
  • Result: physically consistent outputs with data-efficient learning.

Three types of learning

Supervised

  • Learn from labelled examples \((\mathbf{x}_i, y_i)\).
  • Task: predict \(y\) from new \(\mathbf{x}\).
  • Regression (continuous) or classification (discrete).
  • Example: predict phase label from diffraction pattern.

Unsupervised

  • Learn from unlabelled data \(\{\mathbf{x}_i\}\).
  • Task: find hidden structure — clusters, manifolds, compact representations.
  • Example: cluster EELS spectra into chemical phases without manual annotation.

Self-supervised

  • Generate labels from the data itself — no human annotation.
  • Task: predict one part of the input from another.
  • Example: Noise2Noise denoising — predict one noisy version from another.

The supervised learning recipe

  1. Collect labelled examples: \((\mathbf{x}_i, y_i)\) for \(i = 1, \ldots, N\) — features and targets.
  2. Choose a model class \(f_\theta\): linear, polynomial, neural network, …
  3. Define a loss function \(\ell(f_\theta(\mathbf{x}), y)\): MSE, cross-entropy, Poisson NLL.
  4. Minimise the average loss over the training set (empirical risk minimisation).
  5. Evaluate on held-out data — never use the test set during training.
  6. Deploy, monitor, retrain as data distribution shifts.

This is CRISP-DM phases 3–6, mapped onto ML terminology.

Bias–variance trade-off: the central tension

  • Bias: systematic error from overly simple assumptions — model misses real patterns.
  • Variance: sensitivity to sampling noise — model memorises training quirks.
  • Underfitting (high bias): too simple; straight-line fit to a curve.
  • Overfitting (high variance): too complex; degree-50 polynomial fits noise.
  • The “sweet spot” minimises the sum of both.
  • EM analogy: fitting an EELS edge profile — a line is high-bias; a free polynomial is high-variance; a physically-motivated Hartree–Fock profile is the sweet spot.
  • Regularisation, more training data, and physics priors all push toward the sweet spot.
  • Core message: complexity is not free — it trades bias for variance.

Model taxonomy: white / grey / black-box

  • White-box: explicit mechanism, interpretable parameters (physical laws, linear regression).
  • Black-box: high predictive flexibility, no traceable internal mechanism (deep neural networks).
  • Grey-box: partially traceable; blends physics structure with learned components (physics-informed neural networks, PINNs) Neuer, Michael et al., (2024).
  • Explainability tools can move a black-box toward the grey zone — but never all the way to white.
  • White: Bragg’s law, Beer–Lambert law, Schrödinger equation.
  • Grey: learned hardening law in crystal-plasticity simulation.
  • Black: CNN classifying phase from HAADF image.

Why black-box criticism is not just philosophy

  • Safety, traceability, and auditability requirements in engineering and science.
  • Difficulty diagnosing failure causes without interpretable intermediate representations.
  • EM example: a CNN classifies crystal phases with 97% accuracy but cannot explain why it rejected a pattern — which means you cannot tell whether it learned lattice geometry or a detector artefact.
  • Good practice: validate that the model has learned the right features, not just achieved the right score.

The EM data zoo: four modalities

  • HAADF-STEM image: intensity ∝ Z^1.6–1.9; 2-D array shape (ny, nx). Heavy atoms bright.
  • 4D-STEM: a 2-D diffraction pattern at every probe position → 4-D tensor (ny, nx, ky, kx).
  • EELS / EDS spectrum image: energy spectrum at every pixel → 3-D tensor (ny, nx, E).
  • Tilt-series tomography: N projections at different tilt angles → 3-D array (n_tilt, ny, nx).
  • Each modality has its own noise model, sampling constraints, and ML architecture family.

HAADF-STEM: what the image encodes

STEM measurement families showing HAADF, 4D-STEM, EELS, and tomography acquisition geometries.
  • Probe scans in a raster; at each position the transmitted beam is detected.
  • HAADF (High Angle Annular Dark Field): collects high-angle scattered electrons; signal ∝ Z^~1.7 (often quoted as Z² in the high-angle Rutherford limit).
  • Result: a 2-D map of projected atomic-column density — heavier atoms appear bright.
  • Each pixel = one dwell time × electron current = a count of scattered electrons.

The data-formation chain

Every EM number passes through five links. Noise and systematic errors enter at each link (red); the last link — metadata — decides whether the number can be interpreted at all.

  • Specimen → interaction → detector → digitisation → metadata: the measured array \(\mathbf{y}\) is a random draw from a physics-determined distribution, not the specimen itself.
  • A learning algorithm only ever sees the end of the chain — whatever the chain destroys or distorts, the model cannot recover without an explicit forward model or prior.

From physical signal to digital number

  • Physical interaction: electrons scatter off the sample; scattered electrons hit the detector.
  • Transduction: the detector converts particle arrivals into an electrical signal (charge, voltage).
  • Digitisation: an ADC samples the analog signal at discrete times and quantises to integer values.
  • Every step introduces noise and constraints: shot noise (quantum), read noise (electronic), quantisation (ADC resolution).
  • The digital number is never “the truth” — it is one realisation of a random variable drawn from a physics-determined distribution.

Sampling: Nyquist–Shannon theorem

  • A signal containing only frequencies up to \(f_\text{max}\) can be perfectly reconstructed from samples taken at rate \(f_s \geq 2 f_\text{max}\).
  • Nyquist frequency: \(f_\text{Nyq} = f_s / 2\) — the highest frequency the system can faithfully capture.
  • In an image: pixel pitch \(\Delta x\) sets \(f_s = 1/\Delta x\); structural features of size \(d\) need \(\Delta x \leq d/2\).
  • Key caveat: this assumes perfect band-limiting before sampling (anti-aliasing filter). Without it, high-frequency content folds into the recorded band.
  • Practical rule: sample at 3–5× the maximum frequency of interest, not the bare minimum.
  • In atomic-resolution STEM: lattice spacing \(a\) → pixel pitch ≤ \(a/4\) is standard.
  • Undersampling = aliasing (next slide).

Aliasing and Moiré fringes in TEM

Moiré pattern produced by superposing two slightly mismatched lattices (simulation; spatial frequency of the beats is the difference of the two lattice frequencies).
  • When the camera pixel pitch is too coarse to resolve the crystal lattice, lattice reflections alias into the image.
  • The alias appears as spurious long-wavelength Moiré fringes — not real structure.
  • Same artefact from two overlapping crystals with slightly different spacings: the beats have period \(1/(1/a_1 - 1/a_2)\).
  • Diagnostic: rotate the sample or change camera length — real structure rotates; aliasing pattern changes non-trivially.

Aliasing visualised: undersampled lattice

  • Imagine a crystal with lattice spacing \(a = 2\) Å sampled at pixel pitch \(\Delta x = 1.5\) Å — below Nyquist for that frequency.
  • The lattice reflections at \(k = 1/a = 5\) nm⁻¹ alias to \(k' = k - f_s = 5 - 6.67 = -1.67\) nm⁻¹ → a spurious fringe at 6 Å period appears.
  • Rule: alias frequency = \(|f - n \cdot f_s|\) for the nearest integer \(n\).
  • Prevention: always check that your camera length places the Bragg spots you care about well inside the Nyquist radius of the detector.

Sensors and detector types in EM

Scintillator + CCD/CMOS (indirect)

  • High-energy electron → scintillator → visible photons → CMOS sensor.
  • Pro: sensor survives radiation; cheap; large area.
  • Con: scintillator spreads light → broad PSF; lower DQE; mixed Gaussian+Poisson noise.

Direct electron detector (DED)

  • High-energy electron enters thinned Si layer directly → sharp PSF.
  • Pro: near-delta PSF; single-electron counting; dominant noise = pure Poisson.
  • Con: sensor damaged over time by direct irradiation; expensive.
  • Impact: the “resolution revolution” in cryo-EM (2013, Nobel Prize 2017) was enabled by DEDs.

Two fundamental noise sources

Gaussian noise (thermal / readout)

  • Many small independent fluctuations add up → Central Limit Theorem → Gaussian.
  • Read noise in CMOS/CCD electronics: standard deviation \(\sigma_r\) independent of signal.
  • Thermal noise in resistors (Johnson–Nyquist): \(\propto \sqrt{k_B T R \Delta f}\).
  • Variance is independent of the signal.

Poisson noise (shot / counting)

  • Individual electrons arrive at the detector at random times.
  • Count \(N\) in a fixed time window: \(N \sim \text{Poisson}(\lambda)\).
  • Variance equals the mean: \(\text{Var}(N) = \lambda\).
  • SNR \(= \lambda / \sqrt{\lambda} = \sqrt{\lambda}\) → doubling SNR needs 4× the dose.
  • Fundamental: cannot engineer away — it is the quantum nature of the electron.

Dose-dependent image quality: a simulation

Simulated STEM phantom at four dose levels (left to right: 5 000, 500, 50, 5 e⁻/Ų). Image-average SNR ≈ 29/9/3/1; the lowest-dose image is dominated by Poisson noise.
  • The phantom: a clean synthetic image of atomic columns on a periodic lattice.
  • Noise added via np.random.poisson(lam = dose_scale * phantom).
  • At 5 000 e⁻/Ų: columns clearly resolved; SNR ≈ 29.
  • At 500 e⁻/Ų: columns visible but noisy; SNR ≈ 9.
  • At 50 e⁻/Ų: columns barely visible; SNR ≈ 3.
  • At 5 e⁻/Ų: columns invisible to the eye; SNR ≈ 1. ML denoising required.
  • (image-average SNR = √(dose × mean_intensity); peak-column SNR is higher by ~√(1/mean_intensity) ≈ 2.5×)

Real detectors: Gaussian + Poisson mixture

  • Real EM detectors have both shot noise (Poisson) and read noise (Gaussian).
  • Mixed model: \(x_i = g \cdot \text{Poisson}(\lambda_i) + \mathcal{N}(0, \sigma_r^2)\).
    • \(g\) = detector gain (electrons per digital count), \(\sigma_r\) = read noise standard deviation.
  • Variance–mean plot: \(\text{Var}(x) = g \cdot \bar{x} + \sigma_r^2\) — a straight line with slope \(g\) and intercept \(\sigma_r^2\).
  • High-dose regime (\(g\lambda \gg \sigma_r^2\)): Poisson dominates; variance ∝ signal.
  • Low-dose regime (\(g\lambda \ll \sigma_r^2\)): Gaussian read noise dominates; variance is constant.

The complete noise budget

  • Real EM images contain several additive noise contributions:
    1. Poisson shot noise: \(\sigma^2 = \lambda\) — unavoidable quantum limit.
    2. Read noise: \(\sigma^2 = \sigma_r^2\) — independent of signal; Gaussian; reduced by cooling.
    3. Dark current: thermally generated electrons; Poisson; eliminated by cooling.
    4. Fixed-pattern noise: per-pixel gain/offset variation; removed by flat-field calibration.
    5. Quantisation noise: ADC finite resolution; negligible for 14-bit+ systems.
  • Practical characterisation: variance–mean plot — \(\mathrm{Var}(x) = g \cdot \bar{x} + \sigma_r^2\).
    • Slope = detector gain \(g\); intercept = read noise \(\sigma_r^2\).

Measuring the noise model: the variance–mean plot

Photon-transfer (variance–mean) plot of a simulated Poisson–Gaussian detector: the slope recovers the gain \(g\), the intercept the read-noise variance \(\sigma_r^2\); on log–log axes the read-noise floor and the shot-noise regime separate at \(\bar x = \sigma_r^2/g\).

  • Recipe: record flat-field frames at several illuminations → per-pixel mean and variance → linear fit. Two numbers (\(g\), \(\sigma_r\)) fix the noise model — and therefore the loss.

Learning as maximum likelihood

  • From MFML Unit 1: empirical risk minimisation chooses parameters by \[\hat{\theta} = \arg\min_\theta \hat{R}(\theta), \qquad \hat{R}(\theta) = \frac{1}{N}\sum_{i=1}^{N} \ell\big(f_\theta(\mathbf{x}_i), y_i\big).\]
  • Where does \(\ell\) come from? Write the measurement as a random draw, \(y_i \sim p(y \mid f_\theta(\mathbf{x}_i))\) — the data-formation chain as a probability model.
  • Maximum likelihood: choose \(\theta\) that makes the observed data most probable. For independent pixels: \[\hat{\theta}_\text{ML} = \arg\max_\theta \prod_i p(y_i \mid f_\theta(\mathbf{x}_i)) = \arg\min_\theta \sum_i \underbrace{-\log p(y_i \mid f_\theta(\mathbf{x}_i))}_{\ell = \text{negative log-likelihood (NLL)}}.\]
  • Recipe: noise model \(\to\) likelihood \(\to\) take \(-\log\) \(\to\) drop terms without \(\theta\) \(\to\) loss.

Gaussian noise → mean squared error

Noise model: \(y_i = \mu_i + \varepsilon_i\), \(\varepsilon_i \sim \mathcal{N}(0, \sigma^2)\), with \(\mu_i = f_\theta(\mathbf{x}_i)\)

\[p(y_i \mid \mu_i) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\!\left(-\frac{(y_i - \mu_i)^2}{2\sigma^2}\right)\]

Negative log-likelihood: \[-\log p = \frac{(y_i - \mu_i)^2}{2\sigma^2} + \tfrac{1}{2}\log(2\pi\sigma^2)\]

  • If \(\sigma^2\) is the same for every pixel, the second term and the factor \(1/2\sigma^2\) do not depend on \(\theta\) → minimising the NLL is minimising the MSE.
  • MSE silently asserts: Gaussian, independent, homoscedastic (constant-variance) residuals.

Poisson noise → Poisson negative log-likelihood

Noise model: \(y_i \sim \text{Poisson}(\lambda_i)\), \(\lambda_i = f_\theta(\mathbf{x}_i) > 0\)

\[p(y_i \mid \lambda_i) = \frac{\lambda_i^{\,y_i}\, e^{-\lambda_i}}{y_i!}\]

Negative log-likelihood: \[-\log p = \lambda_i - y_i \log\lambda_i + \log(y_i!)\]

  • \(\log(y_i!)\) does not depend on \(\theta\) → Poisson loss \(\;\ell(\lambda, y) = \lambda - y\log\lambda\).
  • Minimum at \(\lambda = y\) per pixel, but the curvature is \(y/\lambda^2\): the loss automatically weights each pixel by \(1/\text{variance}\).
  • PyTorch: nn.PoissonNLLLoss(log_input=True); predict \(\log\lambda\) or use softplus to keep \(\lambda > 0\).

Why the choice matters: dark pixels and implicit weighting

Per-pixel loss as a function of the predicted rate \(\hat\lambda\) for three observed counts. MSE is a symmetric parabola everywhere; the Poisson NLL is asymmetric, linear in \(\hat\lambda\) where nothing was observed, and diverges if the model predicts \(\hat\lambda \to 0\) where counts arrived.

  • Dark pixels (\(y=0\)): MSE punishes \(\hat\lambda^2\) (tiny near zero); Poisson punishes \(\hat\lambda\) linearly and forbids \(\hat\lambda = 0\) wherever any electron arrived.
  • Bright pixels: MSE over-weights them (largest absolute residuals); Poisson down-weights them by \(1/\lambda\) → low-count regions are fitted as carefully as bright ones.

Mixed noise: variance-stabilising transforms

Left: the Anscombe transform makes the standard deviation of Poisson data ≈ 1 for \(\lambda \gtrsim 4\); the generalised version handles added read noise. Right: the naive back-transform is biased at low counts.
  • Anscombe Anscombe, Francis J., (1948), doi:10.1093/biomet/35.3-4.246: \(\;z = 2\sqrt{y + 3/8}\) → \(z\) approximately Gaussian with \(\sigma \approx 1\).
  • Generalised Anscombe for \(y = g\,\text{Poisson}(\lambda) + \mathcal{N}(0,\sigma_r^2)\): \(z = \tfrac{2}{g}\sqrt{g\,y + \tfrac{3}{8}g^2 + \sigma_r^2}\).
  • Then: MSE on \(z\) (or any Gaussian denoiser), and invert with an unbiased inverse.
  • Breaks down below \(\sim\) 1–4 counts/pixel → use the exact Poisson or Poisson–Gaussian NLL there.

Experiment: fitting a low-count peak with four losses

Gaussian peak + background, 300 Poisson realisations per dose, fitted with MSE, the classic “weighted χ²” (weights \(1/\max(y,1)\)), Anscombe + MSE, and Poisson NLL. Middle: bias of the integrated counts (MSE and Poisson curves overlap at ≈ 0). Right: relative scatter of the fitted peak width. Generated by img/make_figures.py; reproduced in the notebook.

  • Weighted χ² (weights from the observed counts) and naive Anscombe are biased low by 10–20 % at a few counts per pixel — a systematic error in every quantified EELS/EDS map.
  • Plain MSE is unbiased here but inefficient: ≈ 1.5× larger width variance → you pay the same precision with ≈ 50 % more dose.
  • Poisson NLL is unbiased and the most precise — the maximum-likelihood estimate conserves total counts.

Cheat sheet: noise model → loss

Data / noise model Likelihood Loss PyTorch
Read-noise dominated, constant \(\sigma\) Gaussian MSE nn.MSELoss
Known per-pixel \(\sigma_i\) Gaussian, heteroscedastic weighted MSE \(\sum (y-\mu)^2/\sigma_i^2\) custom
Model predicts \(\mu\) and \(\sigma^2\) Gaussian Gaussian NLL nn.GaussianNLLLoss
Electron counts (DED, counting mode) Poisson Poisson NLL nn.PoissonNLLLoss
Counts + read noise Poisson–Gaussian VST + MSE, or mixed NLL Anscombe + MSELoss
Outliers, hot pixels, cosmic rays Laplace / heavy-tailed MAE / Huber nn.L1Loss, nn.HuberLoss
Class labels (phase, defect type) Categorical cross-entropy nn.CrossEntropyLoss
Binary masks (segmentation) Bernoulli binary cross-entropy nn.BCEWithLogitsLoss

Priors become regularisers (MAP) — a preview

  • Bayes (MFML Unit 1): \(\;p(\theta \mid \mathcal{D}) \propto p(\mathcal{D} \mid \theta)\, p(\theta)\).
  • Maximum a posteriori: \[\hat{\theta}_\text{MAP} = \arg\min_\theta \Big[\underbrace{-\log p(\mathcal{D}\mid\theta)}_{\text{data term = NLL loss}} \; \underbrace{-\,\log p(\theta)}_{\text{regulariser}}\Big]\]
  • Gaussian prior on weights → L2 / weight decay; Laplace prior → L1 / sparsity; smoothness prior → total variation; non-negativity → NMF (Week 3).
  • Physics enters twice: the detector decides the data term, the specimen knowledge decides the prior.

Data quality: failure modes in EM data

  1. Detector defects: dead / hot pixels, gain-reference errors, saturated columns, X-ray hits.
  2. Acquisition artefacts: scan distortion and drift, flyback lines, beam blanking, contamination build-up, beam damage during the series.
  3. Sampling: undersampling → aliasing / Moiré; energy-axis drift in EELS.
  1. Duplicates & near-duplicates: repeated frames, overlapping crops of the same area → leakage in Week 4.
  2. Missing / wrong metadata: unknown pixel size, dose, or camera length; units lost on export.
  3. Label problems: annotator disagreement, class definitions drifting between sessions.

Note

Garbage in, garbage out: most of these are caught only by looking — histograms, per-frame sums, FFTs — before fitting.

Metadata: the measurement you forget to save

Record with every dataset

  • Instrument, accelerating voltage, mode (STEM/TEM, 4D-STEM, EELS)
  • Probe current, dwell time, frame time → dose (e⁻/Ų)
  • Convergence & collection angles, camera length
  • Pixel size (real & reciprocal), energy dispersion, scan rotation
  • Detector name, mode (counting / integrating), gain & dark reference
  • Specimen ID, preparation, session, operator, timestamp
  • Without dose and gain you cannot convert numbers to counts → no Poisson model.
  • Without pixel size / camera length you cannot check Nyquist or tell a lattice from an alias.
  • Without specimen / session IDs you cannot split data honestly (Week 4).
  • Keep metadata with the data: vendor formats (.dm4, .emd, .mrc) + open readers (HyperSpy, RosettaSciIO); never “just the TIFF”.

FAIR data for electron microscopy

  • Findable — persistent identifier (DOI), rich metadata, indexed in a repository.
  • Accessible — retrievable by a standard protocol; metadata stays available even if data are restricted.
  • Interoperable — open formats and shared vocabularies (units, ontologies, NeXus/EMD).
  • Reusable — licence, provenance, acquisition parameters, processing history.
  • The FAIR principles Wilkinson, Mark D. et al., (2016), doi:10.1038/sdata.2016.18 are about machine-actionability — exactly what an ML pipeline needs.
  • EM practice: raw data + metadata in open formats, a processing script instead of manual clicks, deposit in Zenodo / institutional repositories / EMPIAR.
  • Payoff for ML: re-usable training data, reproducible splits, and models that can be re-validated on someone else’s microscope.

Pre-flight checklist before any ML on EM data

  1. Units & counts: offset subtracted, divided by gain → data in electrons? Dose per pixel known?
  2. Noise model: variance–mean plot made; regime (read-noise / shot-noise / mixed) identified → loss chosen accordingly.
  3. Sampling: pixel size vs smallest feature (Nyquist, ≥ 3–4× oversampling); FFT checked for aliasing.
  4. Defects: dead / hot / saturated pixels masked (not silently interpolated); drift and scan distortion checked.
  5. Duplicates & grouping: specimen / session / region IDs attached for honest splits.
  6. Metadata & provenance: native file + metadata kept; processing as a script; FAIR deposit planned.

Aleatory vs epistemic uncertainty

Aleatory (irreducible)

  • Inherent randomness of the physical process.
  • Cannot be reduced by more data or better models.
  • Examples: shot noise (quantum), thermal vibrations, radioactive decay.
  • Origin: the universe is fundamentally probabilistic.

Epistemic (reducible)

  • Uncertainty from our lack of knowledge.
  • Can be reduced by more data, better calibration, or better models.
  • Examples: unknown detector gain, uncalibrated PSF, small training dataset, wrong model class.
  • Origin: we have not yet measured or computed what we could know.

Practical implications: which uncertainty can you act on?

  • Reduce epistemic uncertainty: collect more training data, calibrate the detector, validate the PSF, improve the model architecture.
  • Accept aleatory uncertainty: report confidence intervals; use noise-matched loss functions; do not over-denoise.
  • Mixed case: a small training set means high epistemic uncertainty on top of irreducible aleatory noise. More data helps — up to the point where aleatory noise dominates.
  • Active learning principle: always prioritise experiments that reduce epistemic uncertainty. Repeating the same measurement over and over only shrinks aleatory uncertainty by \(1/\sqrt{N}\) — which is often not worth the dose budget.

Trust and limits: a teaser

  • A model can have low training loss and be physically wrong.
  • Trust requires: correct noise model, no data leakage, held-out test performance, interpretable learned features, and physical consistency checks.
  • The central tension: models that fit training data perfectly (low bias, high variance) generalise poorly. Models that are too simple (high bias, low variance) miss real patterns. Finding the balance is the bias–variance trade-off.
  • Explainability comes with each model family (Weeks 4, 5, 7, 9, 10, 11) and is synthesised in Week 13. For now: always ask what the model learned, not just how well it scored.

Summary

  • Learning = inferring the function from data; supervised / unsupervised / self-supervised by label source; white / grey / black-box by transparency.
  • Data formation: specimen → interaction → detector → digitisation → metadata. Every link adds noise or systematic error; the model only sees the end of the chain.
  • Sampling: Nyquist sets the pixel size; violating it produces aliasing / Moiré that averaging cannot remove.
  • Noise: Poisson shot noise (Var = mean, SNR = √λ) + Gaussian read noise; the variance–mean plot measures gain and read noise.
  • Noise model → likelihood → loss: Gaussian → MSE, Poisson → Poisson NLL, mixed → VST or mixed NLL; priors → regularisers.
  • Data quality & FAIR metadata are prerequisites, not afterthoughts; aleatory vs epistemic tells you which uncertainty more data can fix.

Must-know takeaways (exam)

  1. A loss function is the negative log-likelihood of an assumed noise model: MSE ⇔ Gaussian, constant variance.
  2. Poisson counts: Var = mean, SNR = √λ; doubling SNR needs 4× dose.
  3. Poisson loss: \(\ell = \lambda - y\log\lambda\); it weights each pixel by 1/variance.
  4. At low counts, observed-count-weighted χ² and naive Anscombe are biased; Poisson NLL is not.
  5. Variance–mean plot: slope = gain, intercept = read-noise variance.
  6. Sample at ≥ 2× (practically 3–4×) the highest spatial frequency; aliasing is coherent and does not average out.
  7. Without dose, gain and pixel size in the metadata, EM data cannot be modelled quantitatively (FAIR).

Self-study this week

  • Notebook: notebooks/week02_poisson_noise.ipynb — “Poisson noise, SNR & the right loss.”
    • Parts 1–6: simulate dose-dependent Poisson noise, verify the \(\sqrt{\lambda}\) law, find the dose at which a defect column disappears (CNR).
    • Part 7: variance–mean plot of a Poisson–Gaussian detector → recover gain and read noise.
    • Parts 8–9: implement the Poisson NLL yourself; fit a low-count peak with MSE, weighted χ², Anscombe + MSE and Poisson NLL → measure bias and scatter.
  • CPU only, runs in < 5 min.

Open in Colab
  • Must-know review: _shared/exam_mustknow.md, Week 2 section.
  • Miniproject: run the pre-flight checklist on your candidate dataset.

Looking ahead — Week 3

  • Topic: “Linear algebra, PCA & spectral unmixing”
  • Vectors, matrices, dot products and the SVD — the language of ML.
  • PCA on an EELS / EDS spectrum image: scree plots, which components are signal vs noise?
  • NMF and MCR-ALS: non-negative, parts-based unmixing into endmember spectra and abundance maps.
  • EELS pre-processing: background subtraction, energy alignment, Poisson weighting before PCA — today’s noise model in action.

Continue

Backup slides

Material moved out of the 90-minute path; use for questions or self-study.

Supervised learning: regression vs classification

  • Regression: predict a continuous target \(y \in \mathbb{R}\).
    • Example: predict the local strain value from a 4D-STEM diffraction pattern.
    • Loss: MSE (Gaussian noise) or Poisson NLL (Poisson noise).
  • Classification: predict a discrete class label \(y \in \{1,\ldots,K\}\).
    • Example: classify each EELS spectrum as oxide, metal, or carbide.
    • Loss: cross-entropy.
  • In EM, many tasks are neither cleanly one nor the other — strain maps are continuous, phase maps are discrete. Think carefully about the output type before choosing a loss.

Self-supervised learning: the Noise2Noise principle

  • Idea: use the data itself to generate supervision signals — no human labels.
  • Noise2Noise Lehtinen, Jaakko et al., (2018): train a denoising network using pairs of noisy images of the same scene. The network learns to predict one noisy image from another — and the expectation of Poisson noise is the clean signal.
  • Why it works for EM: at the same dose, two independently acquired STEM images have the same signal but independent Poisson noise realisations.
  • Result: denoising performance comparable to supervised training against a clean reference image.

References

Machine learning for engineers: Introduction to physics-informed, explainable learning methods for AI in engineering applications, Michael Neuer & others.
Quantitative scanning transmission electron microscopy for materials science: Imaging, diffraction, spectroscopy, and tomography, Annual Review of Materials Research, Colin Ophus.
The transformation of Poisson, binomial and negative-binomial data, Biometrika, Francis J. Anscombe https://doi.org/10.1093/biomet/35.3-4.246.
The FAIR guiding principles for scientific data management and stewardship, Scientific Data, Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, & others https://doi.org/10.1038/sdata.2016.18.
Noise2Noise: Learning image restoration without clean data, International conference on machine learning, Jaakko Lehtinen, Jacob Munkberg, Jon Hasselgren, Samuli Laine, Tero Karras, Miika Aittala, & Timo Aila.