FAU Erlangen-Nürnberg
Institute of Micro- and Nanostructure Research
GroupKFold; the leakage taxonomy is the checklist for every evaluation.notebooks/week05_features_trees.ipynb — simulate HAADF images with vacancies and substitutions, find columns, compute descriptors, classify defects with RF/GBM under GroupKFold, and audit the model with permutation importance.By the end of this week you can:
regionprops) and name its failure modes.Synthetic particle image processed with scikit-image Walt, Stéfan van der et al., (2014), doi:10.7717/peerj.453: smoothing, Otsu threshold Otsu, Nobuyuki, (1979), doi:10.1109/TSMC.1979.4310076, binary mask, distance-transform watershed to split touching particles, and a regionprops feature table (one row per particle).
regionprops: what a particle table containsSize & shape
area, equivalent_diameter_area \(=\sqrt{4A/\pi}\)perimeter, circularity \(4\pi A/P^2\)eccentricity (of the fitted ellipse), major/minor_axis_lengthsolidity = area / convex-hull area (concavity, merged particles)orientation (angle of the major axis)Intensity & context
mean_intensity, max_intensity (needs the intensity image)A
regionpropstable is a designed representation: every column encodes a hypothesis about what matters.
clear_border) or correct statistically.Simulated HAADF-STEM image (square lattice, spacing 12 px, Poisson noise) with vacancy-containing columns (dimmer) and heavier substitutional columns (brighter). Right: detected and refined columns coloured by ground truth.
Column finding on a crop: raw counts → Gaussian smoothing (matched filter) → local maxima with a minimum separation → 2-D Gaussian refinement to sub-pixel positions (mean error ≈ 0.13–0.14 px against ground truth).
Smooth (matched filter) → local maxima with a minimum distance and an intensity threshold → 2-D Gaussian refinement. These two hyperparameters decide which objects exist.
For each peak, fit in a small window (\(9\times9\) px):
\[ I(\mathbf{r}) = A\,\exp\!\left(-\frac{|\mathbf{r}-\mathbf{r}_0|^2}{2\sigma^2}\right) + b \]
by least squares \(\min_{A,\mathbf{r}_0,\sigma,b}\sum_{\text{pixels}}\big(I_\text{obs}-I(\mathbf{r})\big)^2\).
Left: the environment of column \(i\) — neighbours \(j\) within a cutoff, bond lengths \(r_{ij}\) and angles \(\theta_{jik}\). Middle: 2-D radial symmetry function \(G_i(R_s)\) for host, vacancy-neighbourhood and substitution-neighbourhood geometries. Right: two measured descriptors already separate the three classes.
A local descriptor \(\boldsymbol\phi(\text{environment of } i)\) should be:
If a descriptor violates an invariance, the model learns coordinate-system or session accidents instead of physics.
| descriptor | definition | encodes |
|---|---|---|
| coordination \(n_i\) | \(\#\{j: r_{ij} < r_c\}\) | missing / extra neighbours |
| mean bond length \(\bar d_i\) | \(\frac{1}{n_i}\sum_j r_{ij}\) | local strain (relaxation) |
| bond-length spread \(s_i\) | std of \(r_{ij}\) | distortion, asymmetry |
| angular spread | std of angular gaps between neighbours | shear, disorder (0 for a perfect square) |
| relative intensity \(I_{\text{rel},i}\) | \(I_i / \mathrm{median}_j I_j\) | \(Z\)-contrast, occupancy |
| neighbour contrast | \(\mathrm{median}_j I_{\text{rel},j}\) | “a bright neighbour nearby” |
Radial Behler, Jörg et al., (2007), doi:10.1103/PhysRevLett.98.146401
\[ G_i^{\text{rad}}(R_s) = \sum_{j} e^{-\eta\,(r_{ij}-R_s)^2}\, f_c(r_{ij}) \]
Angular
\[ G_i^{\text{ang}} = 2^{1-\zeta}\sum_{j,k}(1+\lambda\cos\theta_{jik})^{\zeta}\, e^{-\eta'(r_{ij}^2+r_{ik}^2)}\,f_cf_c \]
SOAP Bartók, Albert P. et al., (2013), doi:10.1103/PhysRevB.87.184115
2-D version for projected columns
Dataset (notebook): 12 simulated HAADF images, 3 888 inner columns
amp, I_int, sigma, bg_local, x_posI_rel, amp_rel, n_nb, d_mean, d_std, ang_std, G_0.9a, G_1.0a, G_1.1aGroupKFold)x_pos is physically meaningless — a built-in test for importance methods.bg_local differs per image and correlates with the image’s defect fraction.A depth-2 tree on two descriptors: the root asks about the mean neighbour distance, the second level about the relative intensity. Left: the induced axis-aligned partition of feature space; right: the same tree as a flowchart.
At each node choose the (feature \(j\), threshold \(t\)) that maximises the impurity decrease
\[ \Delta = I(\text{parent}) - \frac{N_L}{N}\, I(\text{left}) - \frac{N_R}{N}\, I(\text{right}). \]
Gini, entropy and misclassification error for a two-class node as a function of \(p\). All vanish at pure nodes; only Gini and entropy are strictly concave, so they reward splits that purify a child even without changing its majority class. (Figure from the MFML course, unit 8.)
max_depth, min_samples_leaf.
Strengths
Limits
max_features="sqrt".class_weight="balanced_subsample" for imbalanced defect classes.
Fitting \(y=\sin 2x+0.3x+\varepsilon\) with gradient boosting: after 1, 3, 10 and 50 shrunken trees the ensemble approaches the target and eventually starts to follow the noise — the reason for early stopping (MFML course, unit 8).

Rules: same grouped CV folds for every rung · keep the simplest model that is not significantly worse · report all rungs in the paper / miniproject.
Balanced accuracy under GroupKFold by image (4 folds, 12 images). Left: grey = raw intensity features only; blue = with local-environment descriptors. Right: raw features only, random KFold vs GroupKFold — the random split is optimistic.
StratifiedGroupKFold), report balanced accuracy and per-class recall / confusion matrix.\[ \text{MDI}_j = \frac{1}{B}\sum_{b=1}^{B}\ \sum_{\text{nodes splitting on } j} \frac{N_\text{node}}{N}\,\Delta_\text{node}, \qquad \text{normalised to } \textstyle\sum_j \text{MDI}_j = 1 . \]
rf.feature_importances_ — free, computed during training.predict and a score → works for RF, GBM, MLP, CNN features.sklearn.inspection.permutation_importance(model, X_te, y_te, scoring="balanced_accuracy", n_repeats=10)Random forest trained on 9 images, importances for the 3 held-out images. Left: MDI (training data) gives every feature — including the meaningless column position x_pos (orange) and the width sigma — non-zero credit and ranks the geometric fingerprints G_0.9a/G_1.1a high. Right: permutation importance on held-out images puts x_pos and sigma at ≈ 0 and ranks d_mean and amp_rel first.
Left: adding a near-duplicate of I_rel (same quantity, slightly different integration window) splits its importance between the two copies; permuting both together reveals the true joint importance (≈ 2× either copy). Right: permuting d_mean alone creates columns whose bond length contradicts their own \(G(R_s)\) values — the sum of single-feature importances exceeds the importance of the geometry group permuted together.
Exact (brute-force, \(2^9\) coalitions) interventional Shapley values for \(P(\text{vacancy})\) of a gradient-boosting model with 9 descriptors, computed without the shap package. Left: mean \(|\phi_j|\) over 50 held-out columns. Right: one vacancy column — baseline 0.06, prediction 0.69; attributions sum exactly to the difference.
Material moved out of the 90-minute lecture path; covered by the notebook and self-study.
min_distance too small) → fake neighbours at \(r \ll a\); check the nearest-neighbour distance histogram.Regularisation knobs
max_leaf_nodes)Implementations
HistGradientBoostingClassifier (scikit-learn): histogram binning, native missing values, fastTuning recipe: fix \(\eta=0.05\), early-stop the number of trees on grouped validation folds, then tune depth and
min_samples_leaf.

©Philipp Pelz - FAU Erlangen-Nürnberg - Data Science for Electron Microscopy