FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
Road map (≈ 90 min)
After today you can
Self-study: notebooks/week11_gp_bo.ipynb (GP → conformal → BO, CPU, < 2 min)
Two predictive distributions with the same point prediction (450 MPa) but radically different uncertainty. Left: ±5 MPa — the part passes the safety factor. Right: ±100 MPa — the design must be rejected. The decision lives in the band, not in the mean.
Left: aleatoric — grain-to-grain scatter in grain size at a fixed annealing temperature; more measurements reveal the distribution but do not remove the scatter. Right: epistemic — a GP fitted to 3 EELS measurements is uncertain everywhere except near the measured compositions; measuring more narrows the band.
\[p(y_* \mid \mathbf{x}_*, \mathcal{D}) = \int p(y_* \mid \mathbf{x}_*, \theta)\, p(\theta \mid \mathcal{D})\, d\theta\]
\[ \underbrace{\mathrm{Var}[y_*]}_{\text{total}} = \underbrace{\mathbb{E}_{\theta}\big[\sigma_\epsilon^2(\mathbf{x}_*;\theta)\big]}_{\text{aleatoric}} + \underbrace{\mathrm{Var}_{\theta}\big[\mu(\mathbf{x}_*;\theta)\big]}_{\text{epistemic}} \]
Deep ensembles Lakshminarayanan, Balaji et al., (2017)
MC dropout Gal, Yarin et al., (2016)
model.eval() switches dropout off.For classification, per pixel \(i\): \(\bar{\mathbf{p}}_i = \frac{1}{T}\sum_t \mathbf{p}_i^{(t)}\), predictive entropy \(H_i = -\sum_c \bar p_{i,c}\log \bar p_{i,c}\).
Left: MC dropout — stochastic forward passes (light blue) and mean ±2σ. Right: a deep ensemble of 5 independently trained networks (coloured) and mean ±2σ (green). Both widen away from data; ensembles are empirically better calibrated Lakshminarayanan, Balaji et al., (2017).


Reliability diagrams. Left: a well-calibrated classifier — predicted confidence matches observed accuracy in each bin. Right: an overconfident model — bins below the diagonal; its “90%” is right far less often than 90% of the time.
Grain-size estimate vs electron dose with noise growing with dose. Left: a random forest wrapped with split conformal — target coverage 0.90 reached (0.93) with a constant band; the naive training-residual band covers only 0.78. Right: conformalized quantile regression — same guarantee (0.94) with a band that widens with dose. Notebook Part B, SEED=42.
| GP credible band | Split conformal / CQR | Needs in addition | |
|---|---|---|---|
| Guarantee | only if kernel + noise model correct | finite-sample, distribution-free, marginal | — |
| Width | varies with distance to data | constant (split) / adaptive (CQR) | — |
| Holds under shift? | no — kernel may miss the shift | no — exchangeability broken | OOD gate |
| Steers acquisition? | yes (σ is an exploration signal) | not directly | — |
| Cost | \(O(N^3)\) | sort residuals | — |
Rule of thumb: GP / ensembles to explore; conformal to certify in-distribution; an OOD score to decide whether you are in-distribution at all.
Left: GP prior — plausible curves before any EM data. Right: posterior after 6 EELS measurements (black) — curves pass near the observations; beyond x=1.0 (red shading) the band widens back toward the prior.
With noisy observations \(y_i = f(\mathbf{x}_i) + \epsilon_i\), \(\epsilon_i \sim \mathcal{N}(0,\sigma_n^2)\), kernel matrix \([\mathbf{K}]_{ij} = k(\mathbf{x}_i,\mathbf{x}_j)\) and \([\mathbf{k}_*]_i = k(\mathbf{x}_*,\mathbf{x}_i)\):
\[\boxed{\mu_*(\mathbf{x}_*) = \mathbf{k}_*^\top (\mathbf{K} + \sigma_n^2 \mathbf{I})^{-1} \mathbf{y}}\]
\[\boxed{\sigma_*^2(\mathbf{x}_*) = k(\mathbf{x}_*,\mathbf{x}_*) - \mathbf{k}_*^\top (\mathbf{K} + \sigma_n^2 \mathbf{I})^{-1} \mathbf{k}_*}\]
Illustration: left — GP fitted to 8 measurements clustered in [0, 0.7]; the 95% band widens rapidly beyond the data (red). Right — after one extra measurement at x=1.0 the band collapses locally while distant regions stay uncertain. In the notebook (EELS example, SEED=42) σ* at x=1.15 drops from 0.197 to 0.100 (−49%).
Top-left: RBF kernel for three length-scales. Bottom: GP posteriors with the same data and fixed ℓ. Short ℓ: wiggly mean, band inflates in every gap. Medium ℓ (≈ marginal-likelihood optimum): smooth mean, band opens only where data are missing. Long ℓ: over-smoothed and falsely confident extrapolation.
GaussianProcessRegressor(n_restarts_optimizer=10).| Gaussian process | Deep neural network (+ ensemble) | |
|---|---|---|
| Uncertainty | exact Bayesian (given the kernel) | add-on, approximate |
| Data regime | small, expensive (\(N \lesssim 10^3\)) | large (\(10^4\)–\(10^7\)) |
| Training cost | \(O(N^3)\) (sparse GPs: \(O(NM^2)\)) | \(O(N)\) per epoch |
| Input dimension | struggles beyond ~10–20 raw dims | images, spectra natively |
| Features | hand-chosen via the kernel | learned |
| Interpretability | \(\ell, \sigma_f, \sigma_n\) physically meaningful | opaque weights |
Experimental EM lives in the GP regime (5–500 expensive measurements) — unless inputs are images → deep kernel learning (later today).
Left: a 64-point grid over (beam voltage × convergence angle) — most measurements land in low-SNR regions. Right: BO with 11 measurements (3 initial + 8 guided) converges toward the optimum (gold star, ≈180 kV, 25 mrad).
Four snapshots of BO (UCB, κ=3, SEED=42) on a synthetic 4D-STEM SNR vs convergence-angle curve with a broad deceptive local optimum (x≈0.25, orange) and a narrow global peak (x≈0.78, green). Blue: GP ±2σ; dashed red: next query. The loop escapes the local optimum and concentrates measurements on the global peak.

DKL workflow for automated 4D-STEM. (a) Learning: sparse HAADF patches → neural network embedding → GP → scalarised property y. (b) Prediction: mean and uncertainty for all unvisited patches. (c) Measurement: the diffraction pattern at the chosen position is scalarised (centre of mass, virtual aperture). Figure: Roccapriore et al., ACS Nano 16, 7605 (2022) Roccapriore, Kevin M. et al., (2022), doi:10.1021/acsnano.1c11118.
Experiment on graphene (Nion UltraSTEM 100): HAADF with visited points (red), acquisition, predicted CoM magnitude/angle and uncertainty after 3, 10, 25 and 100 steps. The acquisition concentrates on domain boundaries. Figure: Roccapriore et al., ACS Nano 16, 7605 (2022) Roccapriore, Kevin M. et al., (2022), doi:10.1021/acsnano.1c11118.

Left: training features (penultimate-layer embedding) with the Mahalanobis accept region; new-detector and contamination inputs fall outside. Right: Mahalanobis distances for in-distribution validation data vs OOD inputs; threshold τ at the 95th percentile of in-distribution → 5% false refusals, 99% of these OOD inputs refused (synthetic illustration, img/make_figures.py).

notebooks/week11_gp_bo.ipynb — CPU, sklearn only, < 2 minReal-data extension (optional): MC-dropout U-Net on MetalDAM with reject threshold and tool-shift calibration — Ai4Mat companion notebook (ai4mat.datasets.MetalDAMDataset).
model.eval() disables dropout.Material moved out of the 90-minute lecture path; useful for questions and self-study.
Left: posterior mean and ±2σ band for 4 EELS measurements. Right: posterior standard deviation σ*(x) alone — near zero at the measurement locations, rising toward the prior σ_f in unexplored regions. This σ-map is what active learning and BO consume.
Top: GP surrogate after 3 measurements. Bottom: UCB (κ=2), EI and PI. All agree on the promising region but differ in how sharply they peak; each is maximised on a dense grid.

©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy