FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
notebooks/week04_leakage_demo.ipynb — group leakage and preprocessing leakage, measured.By the end of this lecture you can:
Every supervised learning algorithm is a choice of loss + a choice of optimiser.
Enter gradient descent — an iterative alternative that never inverts anything.
Only first-order information is needed — no inverses, no second derivatives.
MSE punishes large residuals heavily — one bad crop can dominate the loss.



Left: \(\hat R = w_1^2 - w_2^2 + w_2^4/4\). GD started exactly on the ridge stalls at the saddle; a tiny perturbation (the noise SGD provides) escapes to a minimum. Right: the empirical risk (training set) and the true risk (new data) have different minima — driving training loss to its minimum is not the goal.

1


| Regime | Bias | Variance | Cure |
|---|---|---|---|
| Underfit | High | Low | More flexible model / features |
| Good fit | Low | Low | — |
| Overfit | Low | High | More data, fewer parameters, regularisation |
Constrained view: minimise the MSE (grey elliptical contours around \(\hat{\mathbf{w}}_{\mathrm{OLS}}\)) subject to \(\Omega(\mathbf{w})\le t\). The first contour to touch the constraint set gives the solution. The L1 diamond has corners on the axes, so the contact point typically has \(w_2 = 0\) exactly.
Coefficient paths on a synthetic EELS-window regression (target: Fe at%). Fe-L3 and Fe-L2 are nearly collinear. Ridge shrinks them together; Lasso eventually keeps only one of the pair.
RidgeCV, LassoCV, or a Pipeline + GridSearchCV with the same splitter you report with.

| Type | What leaks | EM example | Fix |
|---|---|---|---|
| Duplicate / near-duplicate | the same information appears in train and test | overlapping crops, repeated acquisitions of one field of view, augmented copies made before splitting | deduplicate / crop without overlap; augment after splitting |
| Temporal | the future predicts the past | random split of an in-situ video or a drifting session | chronological split, gap between blocks |
| Group / specimen | identity of the physical object | crops of one specimen, spectra of one grain, frames of one tilt series on both sides | GroupKFold by specimen / session / grain |
| Preprocessing | test statistics or labels shape the features | scaler, PCA, background fit or feature selection fitted on all data | fit every step inside the fold (Pipeline) |
Bonus — target leakage: a feature that is a by-product of the label (e.g. a hand-drawn mask outline used as input for the segmentation it defines; post-relaxation volume predicting a relaxation energy, MG U8).

Assign a group label (specimen_id) to every data point.
GroupKFold: the entire specimen stays in either train or test — never split across folds.
Materials default: if there is a specimen_id column, your default CV is GroupKFold.
The within-specimen correlation that random CV exploits is noise from the perspective of generalisation — ignoring it inflates your score by a predictable amount.
Near-duplicate leakage in practice
Axes of “new” (after MG U8)

Pipeline gives the honest \(R^2 \le 0\).Supervised preprocessing (feature selection, target scaling, choosing ROIs by looking at labels) outside the fold loop → large inflation.
Unsupervised preprocessing (scaler, PCA, background models) outside the loop → small, data-dependent bias; still wrong, and it grows for small \(N\) and grouped data.
Fix — one rule: every step that learns from data is fitted on the training fold only:
For each of the three setups below, identify the leakage and the fix.
(a) You standardise all features with StandardScaler().fit_transform(X) on the full dataset, then run 5-fold cross-validation.
Pre-processing leak. The scaler saw all test-set values when computing \(\mu\) and \(\sigma\). Fix: StandardScaler().fit(X_train) inside each fold (use Pipeline).
(b) You collect 100 EBSD maps from the same 5 specimens (20 maps each). You run a random 5-fold CV and report Dice=0.91. (Dice: segmentation metric — see metrics section.) On a 6th specimen, Dice=0.51.
Group leak. Maps from the same specimen in both train and test. Fix: GroupKFold(groups=specimen_id).
(c) You record an in-situ liquid-phase TEM video (1000 frames). You randomly shuffle and split 80/20. Train \(R^2 = 0.97\), deploy \(R^2 = 0.30\).
Temporal leak. Future frames used to predict past ones. Fix: train on first 800 frames; test on last 200 (chronological split).
Pattern: every leakage scenario reduces to one sentence — test-set information influenced the training process.



300 bootstrap refits on \(N=40\) synthetic EELS spectra. Fe-L3 and Fe-L2 windows are nearly collinear (same element). OLS coefficients slide along a ridge and Fe-L3 flips sign in most refits; Ridge stabilises both; Lasso often zeroes one of the pair.
specimen_id, session, acquisition time for every measurement (E1).GroupKFold; time series → chronological; truly independent → KFold. Lock a test set away.Pipeline.Pipeline + the right splitter prevent most of it.GroupKFold.Pipeline); supervised selection outside CV can fake \(R^2>0\) on noise.GroupKFold by image, Pipeline, honest metrics.notebooks/week04_leakage_demo.ipynb — GroupKFold; the leakage gap vs within-specimen correlation._shared/exam_mustknow.md — Week 4 section.Material for questions and self-study — not part of the 90-minute lecture path.
| Optimiser | Per-step cost | Adaptive \(\eta\)? | Momentum? | Typical use |
|---|---|---|---|---|
| Full GD | \(\mathcal{O}(N)\) | No | No | Tiny datasets, convex |
| SGD | \(\mathcal{O}(1)\) | No | No | Rarely used bare |
| Minibatch SGD | \(\mathcal{O}(b)\) | No | Optional | Many DL papers |
| SGD + Momentum | \(\mathcal{O}(b)\) | No | Yes (\(\beta \approx 0.9\)) | Fine-tuned vision models |
| Adam | \(\mathcal{O}(b)\) | Per-param | Yes | Default for most EM projects |
Gold standard: always wrap preprocessing + model in a Pipeline before passing to CV.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score, KFold, GroupKFold
pipe = Pipeline([
("scale", StandardScaler()), # fitted on train fold only — no leakage
("model", Ridge(alpha=1.0))
])
# Random K-fold (only if data points are independent)
scores = cross_val_score(pipe, X, y, cv=KFold(n_splits=5), scoring='r2')
# Group K-fold (when specimen_id exists)
scores = cross_val_score(pipe, X, y, cv=GroupKFold(n_splits=5),
groups=specimen_id, scoring='r2')
print(f"R² = {scores.mean():.3f} ± {scores.std():.3f}")Pipeline reruns StandardScaler.fit on each training fold automatically → no leakage.
For a binary defect-detection task:
| Predicted: no defect | Predicted: defect | |
|---|---|---|
| True: no defect | TN | FP (false alarm) |
| True: defect | FN (missed!) | TP |
\[\text{Accuracy} = \frac{\text{TP}+\text{TN}}{\text{TP}+\text{TN}+\text{FP}+\text{FN}}\]

©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy