Data Science for Electron Microscopy
Week 4: Regression, optimisation & honest validation

Prof. Dr. Philipp Pelz

FAU Erlangen-Nürnberg

Institute of Micro- and Nanostructure Research

FAU Logo IMN Logo CENEM Logo ERC Logo Eclipse Logo

Recap: where we left off

  • Week 3: linear algebra, PCA & spectral unmixing — SVD, scree plots, NMF / MCR-ALS endmembers and abundance maps.
  • Key geometric insight: least-squares = projection of the target vector onto the column space of \(\mathbf{X}\); the condition number \(\kappa\) tells you how fragile that projection is.
  • Gap #1: the projection gives optimal weights analytically, but only when we can (stably) invert \(\mathbf{X}^T\mathbf{X}\). For large, ill-conditioned or non-linear problems we need an iterative approach — gradient descent.
  • Gap #2: fitting a model well on training data is not the same as fitting it honestly. The second half of today is about that distinction — splits, leakage and what a coefficient may be trusted to mean.

Today’s questions

  • How do we find model parameters without inverting a matrix — and when is the answer unique? Gradient descent on the loss landscape; convexity decides whether “a minimum” is “the minimum”.
  • Why does a model trained on crops from one specimen fail on a new specimen? Because crops from one specimen are correlated — random splits measure memorisation, not generalisation.
  • Road map: ERM & losses · loss landscapes & convexity · learning rate, SGD → Adam · bias–variance · Ridge/Lasso geometry · train/val/test & CV · leakage taxonomy (duplicate · temporal · group · preprocessing) · metrics · reading coefficients (E1/E2) · honest EM recipe.
  • Self-study: notebooks/week04_leakage_demo.ipynb — group leakage and preprocessing leakage, measured.

Learning outcomes

By the end of this lecture you can:

  1. Write regression as empirical risk minimisation and derive MSE / MAE / Huber from a noise model.
  2. Explain a loss landscape: convexity, local minima, saddle points, conditioning — and why optimisation ≠ generalisation.
  3. Describe what learning rate, momentum and Adam each fix.
  4. Sketch the Ridge (disc) vs Lasso (diamond) geometry and predict which gives sparse weights.
  5. Choose a split (random, grouped, chronological) and name the four leakage types with an EM example each.
  6. Report the right metric and read linear coefficients without over-interpreting them (E1/E2).

Prediction = minimising a loss

  • Dataset: \(\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^N\) — each \(\mathbf{x}_i\) is a feature vector, \(y_i\) is the target.
  • Predictor: \(\hat{y}_i = f_{\mathbf{w}}(\mathbf{x}_i)\) — parameterised by weights \(\mathbf{w}\).
  • Loss: \(L(\hat{y}_i, y_i)\) — a scalar that scores how wrong prediction \(i\) is.
  • Empirical risk (what we minimise): \(\hat{R}(\mathbf{w}) = \dfrac{1}{N}\sum_{i=1}^{N} L(f_{\mathbf{w}}(\mathbf{x}_i),\, y_i)\).
  • Goal: \(\hat{\mathbf{w}} = \arg\min_{\mathbf{w}} \hat{R}(\mathbf{w})\).

Every supervised learning algorithm is a choice of loss + a choice of optimiser.

The linear model & the normal equations

  • Linear predictor: \(\hat{y} = \mathbf{w}^T \mathbf{x} + b = \mathbf{w}^T \mathbf{x}\) (absorb \(b\) into \(\mathbf{w}\)).
  • MSE loss for linear regression: \(\hat{R}(\mathbf{w}) = \tfrac{1}{N}\|\mathbf{X}\mathbf{w} - \mathbf{y}\|^2\).
  • Analytic solution (Normal equations): \(\hat{\mathbf{w}} = (\mathbf{X}^T\mathbf{X})^{-1}\mathbf{X}^T\mathbf{y}\).
  • Geometric reading (Week 3): \(\hat{\mathbf{y}} = \mathbf{X}\hat{\mathbf{w}}\) is the projection of \(\mathbf{y}\) onto the column space of \(\mathbf{X}\).
  • Problem: inverting \(\mathbf{X}^T\mathbf{X}\) fails when \(D\) is large or features are correlated (\(\kappa(\mathbf{X}^T\mathbf{X}) \gg 1\)).

Enter gradient descent — an iterative alternative that never inverts anything.

The gradient & the gradient-descent update

  • The gradient \(\nabla_\mathbf{w} \hat{R}(\mathbf{w})\) points in the direction of steepest ascent; step against it.
  • First-order Taylor: \(\hat{R}(\mathbf{w} - \eta \nabla \hat{R}) \approx \hat{R}(\mathbf{w}) - \eta \|\nabla \hat{R}\|^2 < \hat{R}(\mathbf{w})\) for small \(\eta > 0\).
  • Update rule: \(\mathbf{w}_{t+1} = \mathbf{w}_t - \eta\,\nabla_\mathbf{w}\hat{R}(\mathbf{w}_t)\), start at \(\mathbf{w}_0\) (zeros / small random), repeat until the loss stops decreasing.
  • For MSE linear regression: \(\nabla_\mathbf{w} \hat{R} = \tfrac{2}{N}\mathbf{X}^T(\mathbf{X}\mathbf{w} - \mathbf{y})\) — a matrix-vector product, no inversion.
  • For any differentiable model: backpropagation (Week 6) computes the same gradient automatically.

Only first-order information is needed — no inverses, no second derivatives.

MSE — the default regression loss

  • \(L_{\text{MSE}}(\hat{y}, y) = (\hat{y} - y)^2\) — penalises errors quadratically.
  • Smooth, convex bowl landscape — gradient descent’s ideal setting.
  • Probabilistic identity: minimising MSE over the dataset = maximum-likelihood estimation (MLE) assuming iid Gaussian residuals \(\varepsilon \sim \mathcal{N}(0, \sigma^2)\) Bishop, Christopher M., (2006).
  • In EM: correct when your noise is additive Gaussian (readout noise, Johnson noise). For Poisson-dominated low-dose data, use Poisson NLL instead (Week 2 recap).

MSE punishes large residuals heavily — one bad crop can dominate the loss.

MAE and Huber — robust alternatives

MSE, MAE, and Huber loss as functions of the residual \(r = \hat{y}-y\). MSE grows quadratically; MAE linearly; Huber switches at \(|r|=\delta\).
  • MAE: \(L = |\hat{y} - y|\) — linear penalty, robust to outliers. Probabilistic identity: MLE under Laplacian residuals. Caveat: non-differentiable at zero → sub-gradient methods needed.
  • Huber: quadratic inside \(|r| \le \delta\), linear outside. Best of both: smooth optimisation where residuals are small, robust to spikes. Standard tool when most EM crops are clean but occasional detector artefacts occur.
  • Rule of thumb: start with MSE; switch to Huber if residual plots show heavy tails.

The loss landscape — a bowl in weight space

Gradient descent on an ill-conditioned 2D loss bowl. Left: learning rate too large — steps overshoot and bounce back and forth, then diverge. Right: good learning rate — smooth monotone descent to the minimum.
  • Loss landscape: the surface \(\hat{R}(\mathbf{w})\) over weight space; for MSE linear regression a convex bowl.
  • GD steps cross the contours at right angles (steepest descent).
  • Ill-conditioned bowl (\(\kappa(\mathbf{X}^T\mathbf{X})\gg1\): unscaled or collinear features) → elongated contours → GD zigzags. Standardise inputs.

Convexity — when “a minimum” is “the minimum”

Convex (left): every chord lies above the curve → one global minimum. Non-convex (right): the chord dips below → several local minima.
  • Definition: \(f\) is convex if for all \(\mathbf{a},\mathbf{b}\) and \(\lambda\in[0,1]\) \[f(\lambda\mathbf{a}+(1-\lambda)\mathbf{b}) \le \lambda f(\mathbf{a})+(1-\lambda)f(\mathbf{b}).\]
  • Consequence: every local minimum is global; GD with a suitable \(\eta\) reaches it from any start.
  • Twice differentiable: convex ⇔ Hessian \(\nabla^2 f \succeq 0\) everywhere. For MSE: \(\nabla^2\hat R = \tfrac{2}{N}\mathbf{X}^T\mathbf{X} \succeq 0\) ✓.
  • Convex: MSE, MAE, Huber, Poisson-NLL of a linear model, + Ridge/Lasso penalties. Non-convex: anything with a hidden layer, NMF (Week 3), ptychography.

Non-convex landscapes: saddles, and optimisation ≠ generalisation

Left: \(\hat R = w_1^2 - w_2^2 + w_2^4/4\). GD started exactly on the ridge stalls at the saddle; a tiny perturbation (the noise SGD provides) escapes to a minimum. Right: the empirical risk (training set) and the true risk (new data) have different minima — driving training loss to its minimum is not the goal.

  • Critical points (\(\nabla\hat R = 0\)): minima, maxima, saddles (mixed-sign Hessian eigenvalues) — in high dimensions saddles vastly outnumber bad minima.
  • Plateaus & vanishing gradients: gradient ≈ 0 over wide regions → GD crawls (Week 6). Remedies: SGD noise (saddles), momentum (plateaus, ravines), Adam (heterogeneous curvature).

Learning rate — too small, just right, too large

Effect of learning rate on 1D gradient descent: too small (slow), just right (smooth), too large (diverges).
  • Too small (\(\eta \ll 1/L\), \(L\) = gradient Lipschitz constant): right direction, crawls.
  • About right: loss decreases monotonically.
  • Too large (\(\eta > 2/L\)): overshoots → oscillates or diverges.
  • Practice: monitor the loss curve; schedules (step decay, cosine, reduce-on-plateau) go large early, small late.

Stochastic & minibatch gradient descent

  • Full GD averages \(\nabla L_i\) over all \(N\) samples — \(\mathcal{O}(N)\) per step, too expensive for \(N \sim 10^6\).
  • SGD / minibatch: estimate the gradient from a random batch \(\mathcal{B}_t\) of size \(b\): \[\mathbf{w}_{t+1} = \mathbf{w}_t - \frac{\eta}{b}\sum_{i \in \mathcal{B}_t}\nabla L_i(\mathbf{w}_t)\]
  • Unbiased but noisy: \(\mathbb{E}[\nabla L_i] = \nabla\hat{R}\); variance \(\propto 1/b\). The noise helps escape saddle points.
  • Practical default: \(b = 32\)–\(256\) (GPU-parallel) — the training loop of every network from Week 6 on.

Momentum — the physics intuition

  • Problem: SGD on an elongated bowl zigzags across the steep dimension while crawling along the flat one.
  • Momentum idea: accumulate a “velocity” vector that persists across steps — like a ball rolling downhill: \[\mathbf{v}_t = \beta\,\mathbf{v}_{t-1} + \nabla L(\mathbf{w}_{t-1}), \qquad \mathbf{w}_t = \mathbf{w}_{t-1} - \eta\,\mathbf{v}_t\]
  • \(\beta \approx 0.9\): 90% of previous velocity is preserved. Consistent gradient directions accumulate; oscillating directions cancel.
  • Effect: smoother, faster convergence on ill-conditioned landscapes.

1

Adam — the go-to optimiser

Optimizer trajectories on an ill-conditioned quadratic bowl (elongation ratio 20:1): GD zigzags across the steep dimension; SGD+momentum is visibly smoother; Adam reaches the minimum in far fewer steps.
  • Adam: tracks both the gradient (momentum term \(\hat{\mathbf{v}}_t\)) and the squared gradient (adaptive scaling \(\hat{\mathbf{s}}_t\)): \[\mathbf{w}_t \leftarrow \mathbf{w}_{t-1} - \frac{\eta\,\hat{\mathbf{v}}_t}{\sqrt{\hat{\mathbf{s}}_t} + \epsilon}\]
  • Adaptive scaling: each parameter gets its own effective step size.
  • Defaults (\(\eta = 10^{-3}\), \(\beta_1 = 0.9\), \(\beta_2 = 0.999\)) are the practical starting point for most EM projects.

Overfitting and underfitting

Three polynomial fits to the same noisy data: degree 1 (underfit), degree 3 (good), degree 12 (overfit). The true signal is \(\sin(x)\).
  • Underfit (high bias): too simple — train and test errors both high.
  • Good fit: train error ≈ test error; captures signal, not noise.
  • Overfit (high variance): memorises training noise — train error ≈ 0, test error ≫ 0.

Bias–variance decomposition

  • For squared-error loss, the expected test error decomposes as Bishop, Christopher M., (2006): \[\mathbb{E}\bigl[(\hat{y} - y)^2\bigr] = \underbrace{(\mathbb{E}\hat{y} - y)^2}_{\text{Bias}^2} + \underbrace{\text{Var}(\hat{y})}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise floor}}\]
  • Bias: systematic error — how far the average prediction is from the truth (underfitting).
  • Variance: sensitivity to the training set — how much the prediction changes across different training sets (overfitting).
  • Noise \(\sigma^2\): irreducible — set by the physics and detector (recall Week 2: shot noise, readout noise).
  • Diagnostic: high train + high val error → bias; low train + high val error → variance. Always plot loss vs epoch, train vs val error, and residuals.
Regime Bias Variance Cure
Underfit High Low More flexible model / features
Good fit Low Low —
Overfit Low High More data, fewer parameters, regularisation

Regularisation — controlling variance

  • Augment the loss with a penalty on weights: \[\mathcal{L}_{\text{reg}}(\mathbf{w}) = \underbrace{\frac{1}{N}\|\mathbf{Xw} - \mathbf{y}\|^2}_{\text{data term}} + \lambda\,\Omega(\mathbf{w})\]
  • Ridge (L2): \(\Omega(\mathbf{w}) = \|\mathbf{w}\|_2^2\) — shrinks all weights toward zero; keeps them non-zero. Bayesian interpretation: Gaussian prior on \(\mathbf{w}\).
  • Lasso (L1): \(\Omega(\mathbf{w}) = \|\mathbf{w}\|_1\) — drives many weights exactly to zero → automatic feature selection. Bayesian: Laplace prior.
  • \(\lambda\) is a hyperparameter — tune it by (grouped) cross-validation (§6–7).
  • In Week 3 context: Ridge adds \(\lambda\mathbf{I}\) to \(\mathbf{X}^T\mathbf{X}\), lifting all eigenvalues above \(\lambda\) — eliminates ill-conditioning.

Ridge vs Lasso — the geometry

Constrained view: minimise the MSE (grey elliptical contours around \(\hat{\mathbf{w}}_{\mathrm{OLS}}\)) subject to \(\Omega(\mathbf{w})\le t\). The first contour to touch the constraint set gives the solution. The L1 diamond has corners on the axes, so the contact point typically has \(w_2 = 0\) exactly.

  • Penalised form \(\hat R + \lambda\Omega(\mathbf{w})\) ⇔ constrained form \(\min \hat R\) s.t. \(\Omega(\mathbf{w})\le t\) (Lagrange duality; larger \(\lambda\) ⇔ smaller \(t\)).
  • Disc (L2): smooth boundary → contact anywhere → all weights shrunk, none exactly zero. Diamond (L1): corners on the axes → contact at a corner → exact zeros = feature selection Tibshirani, Robert, (1996).

Regularisation paths & choosing \(\lambda\)

Coefficient paths on a synthetic EELS-window regression (target: Fe at%). Fe-L3 and Fe-L2 are nearly collinear. Ridge shrinks them together; Lasso eventually keeps only one of the pair.

  • Always standardise first — the penalty acts on \(|w_j|\), so feature units would otherwise decide who gets penalised.
  • Pick \(\lambda\) by (grouped!) cross-validation — RidgeCV, LassoCV, or a Pipeline + GridSearchCV with the same splitter you report with.
  • Correlated groups: Lasso picks one member (arbitrarily); Ridge spreads weight; Elastic Net (\(\ell_1+\ell_2\)) keeps groups together.

Why a held-out test set is sacred

The three-way split. The test set is used exactly once, at the very end.
  • Train set: fit model parameters (\(\mathbf{w}\)).
  • Validation set: tune hyperparameters (\(\lambda\), architecture, learning rate). Can look at this repeatedly.
  • Test set: final, one-time evaluation — reports the honest generalisation score.
  • Rule: you may never use test-set information to change any modelling decision. Once you look at the test score and adjust your model, it becomes a second validation set, not a test set.

K-fold cross-validation

5-fold cross-validation: each fold serves as the test set exactly once.
  • Recipe: \(k\) folds; train on \(k-1\), test on the held-out one, rotate. Report mean \(\pm\) std across folds.
  • Every point is tested once; the std shows how stable the estimate is.
  • Why not one split? For small EM datasets (\(N<200\)) a single 80/20 split is a noisy draw.
  • Defaults: \(k=5\)–\(10\); LOOCV for \(N<30\); \(k\) trainings — use repeated holdout for slow deep models.

Data leakage — the silent score inflator

  • Definition: information that would not be available at prediction time reaches the training or model-selection process — directly or indirectly Kapoor, Sayash et al., (2023), doi:10.1016/j.patter.2023.100804. The reported performance is then optimistic by an unknown amount.
  • Symptoms:
    • CV score far above performance on a new specimen, session or lab.
    • A trivial model matches a deep network — both exploit the leak.
    • Performance drops sharply for a new sample batch or microscope.
  • Not a bug — a discipline failure. It is the norm unless actively prevented; Kapoor & Narayanan found leakage-affected papers in 17 fields.

A leakage taxonomy for EM data

Type What leaks EM example Fix
Duplicate / near-duplicate the same information appears in train and test overlapping crops, repeated acquisitions of one field of view, augmented copies made before splitting deduplicate / crop without overlap; augment after splitting
Temporal the future predicts the past random split of an in-situ video or a drifting session chronological split, gap between blocks
Group / specimen identity of the physical object crops of one specimen, spectra of one grain, frames of one tilt series on both sides GroupKFold by specimen / session / grain
Preprocessing test statistics or labels shape the features scaler, PCA, background fit or feature selection fitted on all data fit every step inside the fold (Pipeline)

Bonus — target leakage: a feature that is a by-product of the label (e.g. a hand-drawn mask outline used as input for the segmentation it defines; post-relaxation volume predicting a relaxation energy, MG U8).

The EM leakage trap: crops from one specimen

Random crop split (left, \(R^2 = 0.936\), dishonest) vs. 3-fold specimen-group split (right, mean \(R^2 = 0.169\), honest). Same synthetic dataset; the gap of 0.77 is pure leakage.
  • The scenario: 6 EM specimens, 20 crops each — 120 training examples. A per-specimen property (e.g., composition, lattice parameter, stoichiometry) is the target \(y\).
  • Random crop split (\(R^2 = 0.936\)): crops from Specimen 3 land in both train and test. The model learns “Specimen 3 looks like this” and predicts well on test crops — but it has memorised a specimen, not the property.
  • 3-fold specimen-group split (mean \(R^2 = 0.169\)): entire specimen pairs are in test only. The model must generalise across specimen identities — the honest evaluation.

The cure: specimen-level group splitting

  • Assign a group label (specimen_id) to every data point.

  • GroupKFold: the entire specimen stays in either train or test — never split across folds.

    from sklearn.model_selection import GroupKFold
    gkf = GroupKFold(n_splits=5)
    for tr, te in gkf.split(X, y, groups=specimen_id):
        model.fit(X[tr], y[tr])
        score = model.score(X[te], y[te])
  • Materials default: if there is a specimen_id column, your default CV is GroupKFold.

  • The within-specimen correlation that random CV exploits is noise from the perspective of generalisation — ignoring it inflates your score by a predictable amount.

Duplicates and “what kind of new is your test set?”

Near-duplicate leakage in practice

  • Sliding-window crops with overlap → neighbouring crops share pixels; split by image (or by region with a guard band), not by crop.
  • The same field of view recorded twice (focus series, dose series, drift-corrected frames) = one sample, not many.
  • Augment (rotate/flip/noise) after splitting and only the training set.
  • Quick check: nearest-neighbour distance from each test sample to the training set — suspiciously small ⇒ duplicate.

Axes of “new” (after MG U8)

  1. new crop of a seen image
  2. new image of a seen specimen
  3. new specimen, same session
  4. new session / operator / detector settings
  5. new microscope or lab
  6. new material system
  • A random split probes only axis 1. Choose the split that matches the deployment claim — and say which axis you tested.

Preprocessing leakage — measured

Pure-noise experiment (\(N=60\) spectra, \(D=2000\) channels, target independent of \(\mathbf{X}\)). Selecting the 20 channels most correlated with \(y\) on all data before CV fabricates \(R^2\approx 0.4\) out of nothing; the same selection inside a Pipeline gives the honest \(R^2 \le 0\).
  • Supervised preprocessing (feature selection, target scaling, choosing ROIs by looking at labels) outside the fold loop → large inflation.

  • Unsupervised preprocessing (scaler, PCA, background models) outside the loop → small, data-dependent bias; still wrong, and it grows for small \(N\) and grouped data.

  • Fix — one rule: every step that learns from data is fitted on the training fold only:

    pipe = Pipeline([("scale", StandardScaler()),
                     ("pca", PCA(10)),
                     ("model", Ridge())])
    cross_val_score(pipe, X, y, cv=GroupKFold(5), groups=specimen_id)

Spot the leak — three scenarios

For each of the three setups below, identify the leakage and the fix.

(a) You standardise all features with StandardScaler().fit_transform(X) on the full dataset, then run 5-fold cross-validation.

Pre-processing leak. The scaler saw all test-set values when computing \(\mu\) and \(\sigma\). Fix: StandardScaler().fit(X_train) inside each fold (use Pipeline).

(b) You collect 100 EBSD maps from the same 5 specimens (20 maps each). You run a random 5-fold CV and report Dice=0.91. (Dice: segmentation metric — see metrics section.) On a 6th specimen, Dice=0.51.

Group leak. Maps from the same specimen in both train and test. Fix: GroupKFold(groups=specimen_id).

(c) You record an in-situ liquid-phase TEM video (1000 frames). You randomly shuffle and split 80/20. Train \(R^2 = 0.97\), deploy \(R^2 = 0.30\).

Temporal leak. Future frames used to predict past ones. Fix: train on first 800 frames; test on last 200 (chronological split).

Pattern: every leakage scenario reduces to one sentence — test-set information influenced the training process.

Regression metrics: MAE, RMSE, \(R^2\)

Left: residual plot with conventional \(y - \hat{y}\) on the vertical axis (random scatter around zero = good model). Right: \(R^2\) — the fraction of variance in \(y\) explained by the model.
  • \(\mathrm{MAE} = \frac{1}{n}\sum|y_i - \hat{y}_i|\) — in the same units as \(y\); robust to outliers.
  • \(\mathrm{RMSE} = \sqrt{\tfrac{1}{n}\sum(y_i - \hat{y}_i)^2}\) — in the same units as \(y\); penalises large errors more.
  • \(R^2 = 1 - \mathrm{MSE}_\text{model}/\mathrm{MSE}_\text{baseline}\) — fraction of variance explained; scale-free. \(R^2 = 1\): perfect. \(R^2 = 0\): no better than predicting \(\bar{y}\). \(R^2 < 0\): worse than baseline.
  • Always report \(R^2\) on held-out data, not training data.

Classification and segmentation metrics

IoU and Dice for segmentation: good overlap, over-prediction (high recall, low precision), under-prediction (high precision, low recall).
  • Accuracy misleads on imbalanced EM data: 98% background → “always background” scores 98%.
  • Precision = TP/(TP+FP) — how clean? Recall = TP/(TP+FN) — how complete?
  • F1 = Dice = \(2PR/(P+R)\); IoU = \(|A \cap B| / |A \cup B|\) (segmentation, detection); IoU ≤ Dice.
  • Rule: defect detection → maximise recall; particle picking → balance via F1/Dice.

The explainability ladder — where today fits

  • Six levels (E1–E6) from data to decision Neuer, Michael et al., (2024): E1 data · E2 process · E3 features · E4 model · E5 prediction · E6 decision.
  • E1 — data: what data, from which specimens / sessions, how split? Today’s leakage section is E1 work.
  • E2 — process: which physical process and measurement does each feature and the target stand for (units, detector, preprocessing)?
  • A linear model is its own explanation — \(\hat y = \sum_j w_j x_j\) can be read directly (E3/E4) if E1 and E2 are in order.
  • Course thread: W5 permutation importance & SHAP · W7 saliency · W10 attention · W11 uncertainty · W13 the full ladder.

The six levels of explainability (E1–E6) Neuer, Michael et al., (2024)

Reading linear coefficients correctly

  • \(w_j\) = change in \(\hat y\) per unit change in \(x_j\) with all other features held fixed — a conditional, model-internal statement.
  • Units matter: raw \(w_j\) are in [target unit / feature unit] and cannot be compared across features. Compare standardised coefficients (features scaled to unit variance) — or report both.
  • Sign and size are only meaningful with uncertainty: refit on bootstrap resamples or CV folds and report the spread.
  • Regularisation biases coefficients toward 0 — Ridge/Lasso weights are good for prediction, shrunk as estimates.
  • Association ≠ causation: \(w_j>0\) says “the model uses \(x_j\)”, not “increasing \(x_j\) increases \(y\)” (that is an E6 claim; Week 13).

When coefficients lie: collinearity

300 bootstrap refits on \(N=40\) synthetic EELS spectra. Fe-L3 and Fe-L2 windows are nearly collinear (same element). OLS coefficients slide along a ridge and Fe-L3 flips sign in most refits; Ridge stabilises both; Lasso often zeroes one of the pair.

  • Collinear features ⇒ the data constrain only their sum, not each weight — the same flat valley we saw as an ill-conditioned bowl (\(\kappa \gg 1\)).
  • Remedies: check correlations / VIF first; merge collinear features (e.g. total Fe-L\(_{2,3}\) integral) or regress on PCA/NMF scores (Week 3); prefer Ridge for stable weights; never over-read which one Lasso kept.

The full picture — an honest EM ML recipe

  1. Collect: record specimen_id, session, acquisition time for every measurement (E1).
  2. Inspect: histogram targets; look for duplicates / overlapping crops; plot \(X\) vs \(y\) per specimen.
  3. Split to match the deployment claim: specimen → GroupKFold; time series → chronological; truly independent → KFold. Lock a test set away.
  4. Pipeline: every learned preprocessing step (scaler, PCA, feature selection) inside the Pipeline.
  5. Fit: loss from the noise model; standardise; Adam (or closed form); tune \(\lambda\) with the same grouped CV.
  6. Report: held-out metric ± fold spread, residual plot or confusion matrix, which “axis of new” was tested.
  7. Interpret with care: standardised coefficients + bootstrap spread; flag collinear groups.

Summary

  • Fit = minimise empirical risk. Loss encodes the noise model; GD needs only gradients.
  • Landscape: convex (linear + MSE/Huber/Ridge/Lasso) → one answer; non-convex → saddles, plateaus; conditioning sets the zigzag. Momentum and Adam are fixes for curvature, SGD noise for saddles.
  • Optimisation ≠ generalisation: the training-loss minimum is not the goal; bias–variance and regularisation control the gap.
  • Ridge = disc (shrink), Lasso = diamond (sparse); standardise first, tune \(\lambda\) by grouped CV.
  • Honest validation: train/val/test, K-fold, and the split must match the deployment claim.
  • Leakage taxonomy: duplicate · temporal · group · preprocessing (+ target) — one Pipeline + the right splitter prevent most of it.
  • Coefficients are conditional, unit-dependent and unstable under collinearity — the first rung (E1/E2 → E3) of the explainability ladder.

Must-know takeaways (Week 4)

  1. Minimising MSE = maximum likelihood under Gaussian noise; MAE ↔︎ Laplace, Huber = robust compromise.
  2. GD: \(\mathbf{w}_{t+1}=\mathbf{w}_t-\eta\nabla\hat R\); too large \(\eta\) diverges, too small crawls; ill-conditioning (\(\kappa\gg1\)) causes zigzag.
  3. Convex ⇒ every local minimum is global; linear models with MSE/MAE/Huber (+Ridge/Lasso) are convex, neural networks are not.
  4. Momentum damps oscillations; Adam adds per-parameter step sizes; SGD noise helps escape saddles.
  5. Bias–variance: underfit = high bias, overfit = high variance; regularisation trades a little bias for less variance.
  6. Lasso’s L1 diamond has corners ⇒ exact zeros; Ridge’s disc only shrinks.
  7. The test set is used once; hyperparameters (incl. \(\lambda\)) are tuned on validation / CV.
  8. Four leakage types: duplicate, temporal, group/specimen, preprocessing — correlated crops ⇒ GroupKFold.
  9. Every learned preprocessing step goes inside the CV fold (Pipeline); supervised selection outside CV can fake \(R^2>0\) on noise.
  10. A linear coefficient is a conditional, unit-dependent statement; compare standardised coefficients with bootstrap spread and distrust collinear pairs.

Looking ahead — Week 5

  • Topic: “From images to features: descriptors & tree ensembles”
  • From EM images to tabular descriptors: segmentation → particle / grain properties (area, eccentricity); atom-column finding → local-environment descriptors.
  • Decision trees, random forests, gradient boosting — the models that win on tabular data, and the next rung of the baseline ladder (mean → linear → trees → NN).
  • Explainability, next rung: permutation importance and SHAP — and why today’s collinearity warning applies to them too.
  • Validation carries over unchanged: GroupKFold by image, Pipeline, honest metrics.

Self-study this week

  • Notebook: notebooks/week04_leakage_demo.ipynb — Open in Colab
    • Parts 1–6: 6 synthetic specimens × 20 correlated crops; random split vs GroupKFold; the leakage gap vs within-specimen correlation.
    • Part 7 (new): preprocessing leakage — scaler + PCA and feature selection fitted before vs inside the split, on spectra with no signal.
    • Exercises: implement the grouped split; repair the leaky pipeline and assert the honest score.
  • Runs on CPU in < 1 min; first cell installs all dependencies.
  • Must-know review: _shared/exam_mustknow.md — Week 4 section.

Continue

Backup slides

Material for questions and self-study — not part of the 90-minute lecture path.

Optimiser comparison — a visual summary

Optimiser Per-step cost Adaptive \(\eta\)? Momentum? Typical use
Full GD \(\mathcal{O}(N)\) No No Tiny datasets, convex
SGD \(\mathcal{O}(1)\) No No Rarely used bare
Minibatch SGD \(\mathcal{O}(b)\) No Optional Many DL papers
SGD + Momentum \(\mathcal{O}(b)\) No Yes (\(\beta \approx 0.9\)) Fine-tuned vision models
Adam \(\mathcal{O}(b)\) Per-param Yes Default for most EM projects
  • Rule: start with Adam at its defaults. Switch to SGD+momentum only if you have a specific reason (e.g. matching a published training recipe, or if Adam converges to a sharp minimum that generalises poorly).
  • Not covered here: AdaGrad, RMSProp, AdamW, Nesterov — all first-order, all variations on the same theme.

The sklearn cross-validation pipeline

  • Gold standard: always wrap preprocessing + model in a Pipeline before passing to CV.

    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import StandardScaler
    from sklearn.linear_model import Ridge
    from sklearn.model_selection import cross_val_score, KFold, GroupKFold
    
    pipe = Pipeline([
        ("scale", StandardScaler()),   # fitted on train fold only — no leakage
        ("model", Ridge(alpha=1.0))
    ])
    
    # Random K-fold (only if data points are independent)
    scores = cross_val_score(pipe, X, y, cv=KFold(n_splits=5), scoring='r2')
    
    # Group K-fold (when specimen_id exists)
    scores = cross_val_score(pipe, X, y, cv=GroupKFold(n_splits=5),
                              groups=specimen_id, scoring='r2')
    print(f"R² = {scores.mean():.3f} ± {scores.std():.3f}")
  • Pipeline reruns StandardScaler.fit on each training fold automatically → no leakage.

Confusion matrix — the foundation of classification metrics

For a binary defect-detection task:

Predicted: no defect Predicted: defect
True: no defect TN FP (false alarm)
True: defect FN (missed!) TP

\[\text{Accuracy} = \frac{\text{TP}+\text{TN}}{\text{TP}+\text{TN}+\text{FP}+\text{FN}}\]

  • Imbalanced classes are the norm in EM: 98% background, 2% defects. A model that always predicts “no defect” scores 98% accuracy — and misses every defect.
  • Asymmetric costs: missing a defect (FN) → unsafe part ships; calling a good part bad (FP) → unnecessary scrap. Accuracy hides this asymmetry entirely.
  • Fix: use precision and recall (see “Classification and segmentation metrics”) — they separate the two error types.
  • For multi-class: the confusion matrix is \(K \times K\); off-diagonal entries are misclassifications.

References

Pattern recognition and machine learning, Christopher M. Bishop.
Regression shrinkage and selection via the lasso, Journal of the Royal Statistical Society: Series B (Methodological), Robert Tibshirani.
Ridge regression: Biased estimation for nonorthogonal problems, Technometrics, Arthur E. Hoerl & Robert W. Kennard.
Leakage and the reproducibility crisis in machine-learning-based science, Patterns, Sayash Kapoor & Arvind Narayanan https://doi.org/10.1016/j.patter.2023.100804.
Machine learning for engineers: Introduction to physics-informed, explainable learning methods for AI in engineering applications, Michael Neuer & others.